git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH v3 0/5] unpack large objects in stream

From
HXHan Xin <chiyutianyi@gmail.com>
Date
Nov 22, 2021, 03:32 UTC
Message-ID
<20211122033220.32883-1-chiyutianyi@gmail.com>
In-Reply-To
<20211009082058.41138-1-chiyutianyi@gmail.com>
From: Han Xin <hanxin.hx@alibaba-inc.com>

Although we do not recommend users push large binary files to the git repositories, it's difficult to prevent them from doing so. Once, we found a problem with a surge in memory usage on the server. The source of the problem is that a user submitted a single object with a size of 15GB. Once someone initiates a git push, the git process will immediately allocate 15G of memory, resulting in an OOM risk.

Through further analysis, we found that when we execute git unpack-objects, in unpack_non_delta_entry(), "void *buf = get_data(size);" will directly allocate memory equal to the size of the object. This is quite a scary thing, because the pre-receive hook has not been executed at this time, and we cannot avoid this by hooks.

I got inspiration from the deflate process of zlib, maybe it would be a good idea to change unpack-objects to stream deflate.

Changes since v2:
* Rewrite commit messages and make changes suggested by Jiang Xin.
* Remove the commit "object-file.c: add dry_run mode for write_loose_object()" and
  use a new commit "unpack-objects.c: add dry_run mode for get_data()" instead.
Han Xin (5):
  object-file: refactor write_loose_object() to read buffer from stream
  object-file.c: handle undetermined oid in write_loose_object()
  object-file.c: read stream in a loop in write_loose_object()
  unpack-objects.c: add dry_run mode for get_data()
  unpack-objects: unpack_non_delta_entry() read data in a stream
 builtin/unpack-objects.c            | 92 +++++++++++++++++++++++++--
 object-file.c                       | 98 +++++++++++++++++++++++++----
 object-store.h                      |  9 +++
 t/t5590-unpack-non-delta-objects.sh | 76 ++++++++++++++++++++++
 4 files changed, 257 insertions(+), 18 deletions(-)
 create mode 100755 t/t5590-unpack-non-delta-objects.sh
Range-diff against v2:
1:  01672f50a0 ! 1:  8640b04f6d object-file: refactor write_loose_object() to support inputstream
    @@ Metadata
     Author: Han Xin <hanxin.hx@alibaba-inc.com>
     
      ## Commit message ##
    -    object-file: refactor write_loose_object() to support inputstream
    +    object-file: refactor write_loose_object() to read buffer from stream
     
    -    Refactor write_loose_object() to support inputstream, in the same way
    -    that zlib reading is chunked.
    +    We used to call "get_data()" in "unpack_non_delta_entry()" to read the
    +    entire contents of a blob object, no matter how big it is. This
    +    implementation may consume all the memory and cause OOM.
     
    -    Using "in_stream" instead of "void *buf", we needn't to allocate enough
    -    memory in advance, and only part of the contents will be read when
    -    called "in_stream.read()".
    +    This can be improved by feeding data to "write_loose_object()" in a
    +    stream. The input stream is implemented as an interface. In the first
    +    step, we make a simple implementation, feeding the entire buffer in the
    +    "stream" to "write_loose_object()" as a refactor.
     
         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>
         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>
    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam
      	return fd;
      }
      
    -+struct input_data_from_buffer {
    -+	const char *buf;
    ++struct simple_input_stream_data {
    ++	const void *buf;
     +	unsigned long len;
     +};
     +
    -+static const char *read_input_stream_from_buffer(void *data, unsigned long *len)
    ++static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)
     +{
    -+	struct input_data_from_buffer *input = (struct input_data_from_buffer *)data;
    ++	struct simple_input_stream_data *data = in_stream->data;
     +
    -+	if (input->len == 0) {
    ++	if (data->len == 0) {
     +		*len = 0;
     +		return NULL;
     +	}
    -+	*len = input->len;
    -+	input->len = 0;
    -+	return input->buf;
    ++	*len = data->len;
    ++	data->len = 0;
    ++	return data->buf;
     +}
     +
      static int write_loose_object(const struct object_id *oid, char *hdr,
    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *
      	struct object_id parano_oid;
      	static struct strbuf tmp_file = STRBUF_INIT;
      	static struct strbuf filename = STRBUF_INIT;
    -+	const char *buf;
    ++	const void *buf;
     +	unsigned long len;
      
      	loose_object_path(the_repository, &filename, oid);
    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *
      	the_hash_algo->update_fn(&c, hdr, hdrlen);
      
      	/* Then the data itself.. */
    -+	buf = in_stream->read(in_stream->data, &len);
    ++	buf = in_stream->read(in_stream, &len);
      	stream.next_in = (void *)buf;
      	stream.avail_in = len;
      	do {
    @@ object-file.c: int write_object_file_flags(const void *buf, unsigned long len,
      	char hdr[MAX_HEADER_LEN];
      	int hdrlen = sizeof(hdr);
     +	struct input_stream in_stream = {
    -+		.read = read_input_stream_from_buffer,
    -+		.data = (void *)&(struct input_data_from_buffer) {
    ++		.read = feed_simple_input_stream,
    ++		.data = (void *)&(struct simple_input_stream_data) {
     +			.buf = buf,
     +			.len = len,
     +		},
    @@ object-file.c: int hash_object_file_literally(const void *buf, unsigned long len
      	char *header;
      	int hdrlen, status = 0;
     +	struct input_stream in_stream = {
    -+		.read = read_input_stream_from_buffer,
    -+		.data = (void *)&(struct input_data_from_buffer) {
    ++		.read = feed_simple_input_stream,
    ++		.data = (void *)&(struct simple_input_stream_data) {
     +			.buf = buf,
     +			.len = len,
     +		},
    @@ object-file.c: int force_object_loose(const struct object_id *oid, time_t mtime)
      	char hdr[MAX_HEADER_LEN];
      	int hdrlen;
      	int ret;
    -+	struct input_data_from_buffer data;
    ++	struct simple_input_stream_data data;
     +	struct input_stream in_stream = {
    -+		.read = read_input_stream_from_buffer,
    ++		.read = feed_simple_input_stream,
     +		.data = &data,
     +	};
      
    @@ object-store.h: struct object_directory {
      };
      
     +struct input_stream {
    -+	const char *(*read)(void* data, unsigned long *len);
    ++	const void *(*read)(struct input_stream *, unsigned long *len);
     +	void *data;
     +};
     +
2:  a309b7e391 < -:  ---------- object-file.c: add dry_run mode for write_loose_object()
3:  b0a5b53710 ! 2:  d4a2caf2bd object-file.c: handle nil oid in write_loose_object()
    @@ Metadata
     Author: Han Xin <hanxin.hx@alibaba-inc.com>
     
      ## Commit message ##
    -    object-file.c: handle nil oid in write_loose_object()
    +    object-file.c: handle undetermined oid in write_loose_object()
     
    -    When read input stream, oid can't get before reading all, and it will be
    -    filled after reading.
    +    When streaming a large blob object to "write_loose_object()", we have no
    +    chance to run "write_object_file_prepare()" to calculate the oid in
    +    advance. So we need to handle undetermined oid in function
    +    "write_loose_object()".
    +
    +    In the original implementation, we know the oid and we can write the
    +    temporary file in the same directory as the final object, but for an
    +    object with an undetermined oid, we don't know the exact directory for
    +    the object, so we have to save the temporary file in ".git/objects/"
    +    directory instead.
     
         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>
         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>
     
      ## object-file.c ##
     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,
    - 	const char *buf;
    + 	const void *buf;
      	unsigned long len;
      
     -	loose_object_path(the_repository, &filename, oid);
    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *
     +		strbuf_reset(&filename);
     +		strbuf_addstr(&filename, the_repository->objects->odb->path);
     +		strbuf_addch(&filename, '/');
    -+	} else
    ++	} else {
     +		loose_object_path(the_repository, &filename, oid);
    ++	}
      
    - 	if (!dry_run) {
    - 		fd = create_tmpfile(&tmp_file, filename.buf);
    + 	fd = create_tmpfile(&tmp_file, filename.buf);
    + 	if (fd < 0) {
     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,
      		die(_("deflateEnd on object %s failed (%d)"), oid_to_hex(oid),
      		    ret);
    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *
      		die(_("confused by unstable object source data for %s"),
      		    oid_to_hex(oid));
      
    -@@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,
    - 
      	close_loose_object(fd);
      
     +	if (is_null_oid(oid)) {
     +		int dirlen;
     +
    -+		/* copy oid */
     +		oidcpy((struct object_id *)oid, &parano_oid);
    -+		/* We get the oid now */
     +		loose_object_path(the_repository, &filename, oid);
     +
    ++		/* We finally know the object path, and create the missing dir. */
     +		dirlen = directory_size(filename.buf);
     +		if (dirlen) {
     +			struct strbuf dir = STRBUF_INIT;
    -+			/*
    -+			 * Make sure the directory exists; note that the
    -+			 * contents of the buffer are undefined after mkstemp
    -+			 * returns an error, so we have to rewrite the whole
    -+			 * buffer from scratch.
    -+			 */
    -+			strbuf_reset(&dir);
     +			strbuf_add(&dir, filename.buf, dirlen - 1);
     +			if (mkdir(dir.buf, 0777) && errno != EEXIST)
     +				return -1;
    ++			if (adjust_shared_perm(dir.buf))
    ++				return -1;
    ++			strbuf_release(&dir);
     +		}
     +	}
     +
4:  09d438b692 ! 3:  2575900449 object-file.c: read input stream repeatedly in write_loose_object()
    @@ Metadata
     Author: Han Xin <hanxin.hx@alibaba-inc.com>
     
      ## Commit message ##
    -    object-file.c: read input stream repeatedly in write_loose_object()
    +    object-file.c: read stream in a loop in write_loose_object()
     
    -    Read input stream repeatedly in write_loose_object() unless reach the
    -    end, so that we can divide the large blob write into many small blocks.
    +    In order to prepare the stream version of "write_loose_object()", read
    +    the input stream in a loop in "write_loose_object()", so that we can
    +    feed the contents of large blob object to "write_loose_object()" using
    +    a small fixed buffer.
     
    +    Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>
         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>
     
      ## object-file.c ##
     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,
      	static struct strbuf tmp_file = STRBUF_INIT;
      	static struct strbuf filename = STRBUF_INIT;
    - 	const char *buf;
    + 	const void *buf;
     -	unsigned long len;
     +	int flush = 0;
      
    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *
      	the_hash_algo->update_fn(&c, hdr, hdrlen);
      
      	/* Then the data itself.. */
    --	buf = in_stream->read(in_stream->data, &len);
    +-	buf = in_stream->read(in_stream, &len);
     -	stream.next_in = (void *)buf;
     -	stream.avail_in = len;
      	do {
      		unsigned char *in0 = stream.next_in;
     -		ret = git_deflate(&stream, Z_FINISH);
     +		if (!stream.avail_in) {
    -+			if ((buf = in_stream->read(in_stream->data, &stream.avail_in))) {
    ++			buf = in_stream->read(in_stream, &stream.avail_in);
    ++			if (buf) {
     +				stream.next_in = (void *)buf;
     +				in0 = (unsigned char *)buf;
    -+			} else
    ++			} else {
     +				flush = Z_FINISH;
    ++			}
     +		}
     +		ret = git_deflate(&stream, flush);
      		the_hash_algo->update_fn(&c, in0, stream.next_in - in0);
    - 		if (!dry_run && write_buffer(fd, compressed, stream.next_out - compressed) < 0)
    + 		if (write_buffer(fd, compressed, stream.next_out - compressed) < 0)
      			die(_("unable to write loose object file"));
5:  9fb188d437 < -:  ---------- object-store.h: add write_loose_object()
-:  ---------- > 4:  ca93ecc780 unpack-objects.c: add dry_run mode for get_data()
6:  80468a6fbc ! 5:  39a072ee2a unpack-objects: unpack large object in stream
    @@ Metadata
     Author: Han Xin <hanxin.hx@alibaba-inc.com>
     
      ## Commit message ##
    -    unpack-objects: unpack large object in stream
    +    unpack-objects: unpack_non_delta_entry() read data in a stream
     
    -    When calling "unpack_non_delta_entry()", will allocate full memory for
    -    the whole size of the unpacked object and write the buffer to loose file
    -    on disk. This may lead to OOM for the git-unpack-objects process when
    -    unpacking a very large object.
    +    We used to call "get_data()" in "unpack_non_delta_entry()" to read the
    +    entire contents of a blob object, no matter how big it is. This
    +    implementation may consume all the memory and cause OOM.
     
    -    In function "unpack_delta_entry()", will also allocate full memory to
    -    buffer the whole delta, but since there will be no delta for an object
    -    larger than "core.bigFileThreshold", this issue is moderate.
    +    By implementing a zstream version of input_stream interface, we can use
    +    a small fixed buffer for "unpack_non_delta_entry()".
     
    -    To resolve the OOM issue in "git-unpack-objects", we can unpack large
    -    object to file in stream, and use "core.bigFileThreshold" to avoid OOM
    -    limits when called "get_data()".
    +    However, unpack non-delta objects from a stream instead of from an entrie
    +    buffer will have 10% performance penalty. Therefore, only unpack object
    +    larger than the "big_file_threshold" in zstream. See the following
    +    benchmarks:
     
    +        $ hyperfine \
    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \
    +        'git -C dest.git unpack-objects <binary_320M.pack'
    +        Benchmark 1: git -C dest.git unpack-objects <binary_320M.pack
    +          Time (mean ± σ):     10.029 s ±  0.270 s    [User: 8.265 s, System: 1.522 s]
    +          Range (min … max):    9.786 s … 10.603 s    10 runs
    +
    +        $ hyperfine \
    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \
    +        'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack'
    +        Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack
    +          Time (mean ± σ):     10.859 s ±  0.774 s    [User: 8.813 s, System: 1.898 s]
    +          Range (min … max):    9.884 s … 12.192 s    10 runs
    +
    +        $ hyperfine \
    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \
    +        'git -C dest.git unpack-objects <binary_96M.pack'
    +        Benchmark 1: git -C dest.git unpack-objects <binary_96M.pack
    +          Time (mean ± σ):      2.678 s ±  0.037 s    [User: 2.205 s, System: 0.450 s]
    +          Range (min … max):    2.639 s …  2.743 s    10 runs
    +
    +        $ hyperfine \
    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \
    +        'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack'
    +        Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack
    +          Time (mean ± σ):      2.819 s ±  0.124 s    [User: 2.216 s, System: 0.564 s]
    +          Range (min … max):    2.679 s …  3.125 s    10 runs
    +
    +    Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>
         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>
     
      ## builtin/unpack-objects.c ##
    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type
      	}
      }
      
    -+struct input_data_from_zstream {
    ++struct input_zstream_data {
     +	git_zstream *zstream;
     +	unsigned char buf[4096];
     +	int status;
     +};
     +
    -+static const char *read_inflate_in_stream(void *data, unsigned long *readlen)
    ++static const void *feed_input_zstream(struct input_stream *in_stream, unsigned long *readlen)
     +{
    -+	struct input_data_from_zstream *input = data;
    -+	git_zstream *zstream = input->zstream;
    ++	struct input_zstream_data *data = in_stream->data;
    ++	git_zstream *zstream = data->zstream;
     +	void *in = fill(1);
     +
    -+	if (!len || input->status == Z_STREAM_END) {
    ++	if (!len || data->status == Z_STREAM_END) {
     +		*readlen = 0;
     +		return NULL;
     +	}
     +
    -+	zstream->next_out = input->buf;
    -+	zstream->avail_out = sizeof(input->buf);
    ++	zstream->next_out = data->buf;
    ++	zstream->avail_out = sizeof(data->buf);
     +	zstream->next_in = in;
     +	zstream->avail_in = len;
     +
    -+	input->status = git_inflate(zstream, 0);
    ++	data->status = git_inflate(zstream, 0);
     +	use(len - zstream->avail_in);
    -+	*readlen = sizeof(input->buf) - zstream->avail_out;
    ++	*readlen = sizeof(data->buf) - zstream->avail_out;
     +
    -+	return (const char *)input->buf;
    ++	return data->buf;
     +}
     +
     +static void write_stream_blob(unsigned nr, unsigned long size)
    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type
     +	char hdr[32];
     +	int hdrlen;
     +	git_zstream zstream;
    -+	struct input_data_from_zstream data;
    ++	struct input_zstream_data data;
     +	struct input_stream in_stream = {
    -+		.read = read_inflate_in_stream,
    ++		.read = feed_input_zstream,
     +		.data = &data,
     +	};
     +	struct object_id *oid = &obj_list[nr].oid;
    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type
     +	/* Generate the header */
     +	hdrlen = xsnprintf(hdr, sizeof(hdr), "%s %"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;
     +
    -+	if ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, dry_run, 0, 0)))
    ++	if ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0)))
     +		die(_("failed to write object in stream %d"), ret);
     +
     +	if (zstream.total_out != size || data.status != Z_STREAM_END)
    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type
      static void unpack_non_delta_entry(enum object_type type, unsigned long size,
      				   unsigned nr)
      {
    --	void *buf = get_data(size);
    +-	void *buf = get_data(size, dry_run);
     +	void *buf;
     +
     +	/* Write large blob in stream without allocating full buffer. */
    -+	if (type == OBJ_BLOB && size > big_file_threshold) {
    ++	if (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {
     +		write_stream_blob(nr, size);
     +		return;
     +	}
      
    -+	buf = get_data(size);
    ++	buf = get_data(size, dry_run);
      	if (!dry_run && buf)
      		write_object(nr, type, buf, size);
      	else
     
    - ## t/t5590-receive-unpack-objects.sh (new) ##
    + ## object-file.c ##
    +@@ object-file.c: static const void *feed_simple_input_stream(struct input_stream *in_stream, unsi
    + 	return data->buf;
    + }
    + 
    +-static int write_loose_object(const struct object_id *oid, char *hdr,
    +-			      int hdrlen, struct input_stream *in_stream,
    +-			      time_t mtime, unsigned flags)
    ++int write_loose_object(const struct object_id *oid, char *hdr,
    ++		       int hdrlen, struct input_stream *in_stream,
    ++		       time_t mtime, unsigned flags)
    + {
    + 	int fd, ret;
    + 	unsigned char compressed[4096];
    +
    + ## object-store.h ##
    +@@ object-store.h: int hash_object_file(const struct git_hash_algo *algo, const void *buf,
    + 		     unsigned long len, const char *type,
    + 		     struct object_id *oid);
    + 
    ++int write_loose_object(const struct object_id *oid, char *hdr,
    ++		       int hdrlen, struct input_stream *in_stream,
    ++		       time_t mtime, unsigned flags);
    ++
    + int write_object_file_flags(const void *buf, unsigned long len,
    + 			    const char *type, struct object_id *oid,
    + 			    unsigned flags);
    +
    + ## t/t5590-unpack-non-delta-objects.sh (new) ##
     @@
     +#!/bin/sh
     +#
    @@ t/t5590-receive-unpack-objects.sh (new)
     +		cd .git &&
     +		find objects/?? -type f | sort
     +	) >expect &&
    -+	git repack -ad
    ++	PACK=$(echo main | git pack-objects --progress --revs test)
     +'
     +
     +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '
    @@ t/t5590-receive-unpack-objects.sh (new)
     +	git -C dest.git config receive.unpacklimit 100
     +'
     +
    -+test_expect_success 'fail to push: cannot allocate' '
    -+	test_must_fail git push dest.git HEAD 2>err &&
    -+	test_i18ngrep "remote: fatal: attempting to allocate" err &&
    ++test_expect_success 'fail to unpack-objects: cannot allocate' '
    ++	test_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&
    ++	test_i18ngrep "fatal: attempting to allocate" err &&
     +	(
     +		cd dest.git &&
     +		find objects/?? -type f | sort
    @@ t/t5590-receive-unpack-objects.sh (new)
     +'
     +
     +test_expect_success 'unpack big object in stream' '
    -+	git push dest.git HEAD &&
    ++	git -C dest.git unpack-objects <test-$PACK.pack &&
     +	git -C dest.git fsck &&
     +	(
     +		cd dest.git &&
    @@ t/t5590-receive-unpack-objects.sh (new)
     +'
     +
     +test_expect_success 'setup for unpack-objects dry-run test' '
    -+	PACK=$(echo main | git pack-objects --progress --revs test) &&
    -+	unset GIT_ALLOC_LIMIT &&
     +	git init --bare unpack-test.git
     +'
     +
    -+test_expect_success 'unpack-objects dry-run with large threshold' '
    -+	(
    -+		cd unpack-test.git &&
    -+		git config core.bigFileThreshold 2m &&
    -+		git unpack-objects -n <../test-$PACK.pack
    -+	) &&
    -+	(
    -+		cd unpack-test.git &&
    -+		find objects/ -type f
    -+	) >actual &&
    -+	test_must_be_empty actual
    -+'
    -+
    -+test_expect_success 'unpack-objects dry-run with small threshold' '
    ++test_expect_success 'unpack-objects dry-run' '
     +	(
     +		cd unpack-test.git &&
    -+		git config core.bigFileThreshold 1m &&
     +		git unpack-objects -n <../test-$PACK.pack
     +	) &&
     +	(
-- 
2.34.0.6.g676eedc724
Previous: Jiang XinNext: Han Xin
Message 20 of 211 in “unpack-objects: unpack large object in stream”
  1. unpack-objects: unpack large object in streamHan Xin, Oct 9, 2021
  2. Han XinOct 19, 2021
  3. Philip OakleyOct 20, 2021
  4. Han XinOct 21, 2021
  5. Philip OakleyOct 21, 2021
  6. Han XinNov 3, 2021
  7. Philip OakleyNov 3, 2021
  8. 1/6 object-file: refactor write_loose_object() to support inputstreamHan Xin, Nov 12, 2021
  9. Jiang XinNov 18, 2021
  10. Junio C HamanoNov 18, 2021
  11. 2/6 object-file.c: add dry_run mode for write_loose_object()Han Xin, Nov 12, 2021
  12. Jiang XinNov 18, 2021
  13. 3/6 object-file.c: handle nil oid in write_loose_object()Han Xin, Nov 12, 2021
  14. Jiang XinNov 18, 2021
  15. 4/6 object-file.c: read input stream repeatedly in write_loose_object()Han Xin, Nov 12, 2021
  16. Jiang XinNov 18, 2021
  17. 5/6 object-store.h: add write_loose_object()Han Xin, Nov 12, 2021
  18. 6/6 unpack-objects: unpack large object in streamHan Xin, Nov 12, 2021
  19. Jiang XinNov 18, 2021
  20. 0/5 unpack large objects in streamHan Xin, Nov 22, 2021
  21. Han XinNov 29, 2021
  22. Jeff KingNov 29, 2021
  23. Han XinNov 30, 2021
  24. 0/5 unpack large objects in streamHan Xin, Dec 3, 2021
  25. Derrick StoleeDec 7, 2021
  26. 0/6 unpack large blobs in streamHan Xin, Dec 10, 2021
  27. 0/6 unpack large blobs in streamHan Xin, Dec 17, 2021
  28. 0/5 unpack large blobs in streamHan Xin, Dec 21, 2021
  29. 1/5 unpack-objects.c: add dry_run mode for get_data()Han Xin, Dec 21, 2021
  30. Ævar Arnfjörð BjarmasonDec 21, 2021
  31. René ScharfeDec 21, 2021
  32. Ævar Arnfjörð BjarmasonDec 21, 2021
  33. Jiang XinDec 22, 2021
  34. Jiang XinDec 22, 2021
  35. Jiang XinDec 31, 2021
  36. 2/5 object-file API: add a format_object_header() functionHan Xin, Dec 21, 2021
  37. René ScharfeDec 21, 2021
  38. C99 %z (was: [PATCH v7 2/5] object-file API: add a format_object_header() function)Ævar Arnfjörð Bjarmason, Feb 1, 2022
  39. Jiang XinDec 31, 2021
  40. 3/5 object-file.c: refactor write_loose_object() to reuse in stream versionHan Xin, Dec 21, 2021
  41. Ævar Arnfjörð BjarmasonDec 21, 2021
  42. Jiang XinDec 22, 2021
  43. 4/5 object-file.c: add "write_stream_object_file()" to support read in streamHan Xin, Dec 21, 2021
  44. Ævar Arnfjörð BjarmasonDec 21, 2021
  45. Ævar Arnfjörð BjarmasonDec 21, 2021
  46. 5/5 unpack-objects: unpack_non_delta_entry() read data in a streamHan Xin, Dec 21, 2021
  47. Ævar Arnfjörð BjarmasonDec 21, 2021
  48. Jiang XinDec 31, 2021
  49. 0/6 unpack large blobs in streamHan Xin, Jan 8, 2022
  50. 1/5 unpack-objects: low memory footprint for get_data() in dry_run modeHan Xin, Jan 20, 2022
  51. 0/5 unpack large blobs in streamHan Xin, Jan 20, 2022
  52. Ævar Arnfjörð BjarmasonFeb 1, 2022
  53. Han XinFeb 2, 2022
  54. Ævar Arnfjörð BjarmasonFeb 2, 2022
  55. 0/6 unpack-objects: support streaming large objects to diskÆvar Arnfjörð Bjarmason, Feb 4, 2022
  56. 1/6 unpack-objects: low memory footprint for get_data() in dry_run modeÆvar Arnfjörð Bjarmason, Feb 4, 2022
  57. 2/6 object-file.c: do fsync() and close() before post-write die()Ævar Arnfjörð Bjarmason, Feb 4, 2022
  58. 4/6 object-file.c: add "stream_loose_object()" to handle large objectÆvar Arnfjörð Bjarmason, Feb 4, 2022
  59. 3/6 object-file.c: refactor write_loose_object() to several stepsÆvar Arnfjörð Bjarmason, Feb 4, 2022
  60. 5/6 core doc: modernize core.bigFileThreshold documentationÆvar Arnfjörð Bjarmason, Feb 4, 2022
  61. 6/6 unpack-objects: use stream_loose_object() to unpack large objectsÆvar Arnfjörð Bjarmason, Feb 4, 2022
  62. 0/8 unpack-objects: support streaming blobs to diskÆvar Arnfjörð Bjarmason, Mar 19, 2022
  63. 1/8 unpack-objects: low memory footprint for get_data() in dry_run modeÆvar Arnfjörð Bjarmason, Mar 19, 2022
  64. 2/8 object-file.c: do fsync() and close() before post-write die()Ævar Arnfjörð Bjarmason, Mar 19, 2022
  65. 3/8 object-file.c: refactor write_loose_object() to several stepsÆvar Arnfjörð Bjarmason, Mar 19, 2022
  66. René ScharfeMar 19, 2022
  67. 4/8 object-file.c: factor out deflate part of write_loose_object()Ævar Arnfjörð Bjarmason, Mar 19, 2022
  68. 5/8 object-file.c: add "stream_loose_object()" to handle large objectÆvar Arnfjörð Bjarmason, Mar 19, 2022
  69. 6/8 core doc: modernize core.bigFileThreshold documentationÆvar Arnfjörð Bjarmason, Mar 19, 2022
  70. 7/8 unpack-objects: refactor away unpack_non_delta_entry()Ævar Arnfjörð Bjarmason, Mar 19, 2022
  71. 8/8 unpack-objects: use stream_loose_object() to unpack large objectsÆvar Arnfjörð Bjarmason, Mar 19, 2022
  72. 0/8 unpack-objects: support streaming blobs to diskÆvar Arnfjörð Bjarmason, Mar 29, 2022
  73. 2/8 object-file.c: do fsync() and close() before post-write die()Ævar Arnfjörð Bjarmason, Mar 29, 2022
  74. 1/8 unpack-objects: low memory footprint for get_data() in dry_run modeÆvar Arnfjörð Bjarmason, Mar 29, 2022
  75. 3/8 object-file.c: refactor write_loose_object() to several stepsÆvar Arnfjörð Bjarmason, Mar 29, 2022
  76. Han XinMar 30, 2022
  77. Ævar Arnfjörð BjarmasonMar 30, 2022
  78. 4/8 object-file.c: factor out deflate part of write_loose_object()Ævar Arnfjörð Bjarmason, Mar 29, 2022
  79. 5/8 object-file.c: add "stream_loose_object()" to handle large objectÆvar Arnfjörð Bjarmason, Mar 29, 2022
  80. Neeraj SinghMar 31, 2022
  81. 6/8 core doc: modernize core.bigFileThreshold documentationÆvar Arnfjörð Bjarmason, Mar 29, 2022
  82. 7/8 unpack-objects: refactor away unpack_non_delta_entry()Ævar Arnfjörð Bjarmason, Mar 29, 2022
  83. René ScharfeMar 30, 2022
  84. Ævar Arnfjörð BjarmasonMar 31, 2022
  85. René ScharfeMar 31, 2022
  86. 8/8 unpack-objects: use stream_loose_object() to unpack large objectsÆvar Arnfjörð Bjarmason, Mar 29, 2022
  87. 0/7 unpack-objects: support streaming blobs to diskÆvar Arnfjörð Bjarmason, Jun 4, 2022
  88. 1/7 unpack-objects: low memory footprint for get_data() in dry_run modeÆvar Arnfjörð Bjarmason, Jun 4, 2022
  89. Junio C HamanoJun 6, 2022
  90. Han XinJun 9, 2022
  91. Junio C HamanoJun 9, 2022
  92. Han XinJun 10, 2022
  93. Ævar Arnfjörð BjarmasonJun 10, 2022
  94. Han XinJun 10, 2022
  95. 2/7 object-file.c: do fsync() and close() before post-write die()Ævar Arnfjörð Bjarmason, Jun 4, 2022
  96. Junio C HamanoJun 6, 2022
  97. 3/7 object-file.c: refactor write_loose_object() to several stepsÆvar Arnfjörð Bjarmason, Jun 4, 2022
  98. 4/7 object-file.c: factor out deflate part of write_loose_object()Ævar Arnfjörð Bjarmason, Jun 4, 2022
  99. 5/7 object-file.c: add "stream_loose_object()" to handle large objectÆvar Arnfjörð Bjarmason, Jun 4, 2022
  100. Junio C HamanoJun 6, 2022
  101. Junio C HamanoJun 6, 2022
  102. Han XinJun 9, 2022
  103. Han XinJun 9, 2022
  104. Neeraj SinghJun 7, 2022
  105. Junio C HamanoJun 8, 2022
  106. object-file.c: batched disk flushes for stream_loose_object()Han Xin, Jun 9, 2022
  107. Neeraj SinghJun 9, 2022
  108. Johannes SchindelinJun 9, 2022
  109. Han XinJun 10, 2022
  110. 7/7 unpack-objects: use stream_loose_object() to unpack large objectsÆvar Arnfjörð Bjarmason, Jun 4, 2022
  111. 6/7 core doc: modernize core.bigFileThreshold documentationÆvar Arnfjörð Bjarmason, Jun 4, 2022
  112. Junio C HamanoJun 6, 2022
  113. 0/7 unpack-objects: support streaming blobs to diskHan Xin, Jun 10, 2022
  114. 1/7 unpack-objects: low memory footprint for get_data() in dry_run modeHan Xin, Jun 10, 2022
  115. 2/7 object-file.c: do fsync() and close() before post-write die()Han Xin, Jun 10, 2022
  116. René ScharfeJun 10, 2022
  117. Junio C HamanoJun 10, 2022
  118. Han XinJun 11, 2022
  119. 3/7 object-file.c: refactor write_loose_object() to several stepsHan Xin, Jun 10, 2022
  120. 4/7 object-file.c: factor out deflate part of write_loose_object()Han Xin, Jun 10, 2022
  121. 5/7 object-file.c: add "stream_loose_object()" to handle large objectHan Xin, Jun 10, 2022
  122. 6/7 core doc: modernize core.bigFileThreshold documentationHan Xin, Jun 10, 2022
  123. Junio C HamanoJun 10, 2022
  124. 7/7 unpack-objects: use stream_loose_object() to unpack large objectsHan Xin, Jun 10, 2022
  125. 0/6 unpack-objects: support streaming blobs to diskHan Xin, Jun 11, 2022
  126. 1/6 unpack-objects: low memory footprint for get_data() in dry_run modeHan Xin, Jun 11, 2022
  127. 2/6 object-file.c: refactor write_loose_object() to several stepsHan Xin, Jun 11, 2022
  128. 3/6 object-file.c: factor out deflate part of write_loose_object()Han Xin, Jun 11, 2022
  129. 4/6 object-file.c: add "stream_loose_object()" to handle large objectHan Xin, Jun 11, 2022
  130. 5/6 core doc: modernize core.bigFileThreshold documentationHan Xin, Jun 11, 2022
  131. 6/6 unpack-objects: use stream_loose_object() to unpack large objectsHan Xin, Jun 11, 2022
  132. Junio C HamanoJul 1, 2022
  133. 0/1 unpack-objects: low memory footprint for get_data() in dry_run modeHan Xin, May 20, 2022
  134. 1/1 unpack-objects: low memory footprint for get_data() in dry_run modeHan Xin, May 20, 2022
  135. 2/5 object-file.c: refactor write_loose_object() to several stepsHan Xin, Jan 20, 2022
  136. 3/5 object-file.c: add "stream_loose_object()" to handle large objectHan Xin, Jan 20, 2022
  137. 4/5 unpack-objects: unpack_non_delta_entry() read data in a streamHan Xin, Jan 20, 2022
  138. 5/5 object-file API: add a format_object_header() functionHan Xin, Jan 20, 2022
  139. 1/6 unpack-objects: low memory footprint for get_data() in dry_run modeHan Xin, Jan 8, 2022
  140. René ScharfeJan 8, 2022
  141. Han XinJan 11, 2022
  142. 2/6 object-file.c: refactor write_loose_object() to several stepsHan Xin, Jan 8, 2022
  143. René ScharfeJan 8, 2022
  144. Han XinJan 11, 2022
  145. 3/6 object-file.c: remove the slash for directory_size()Han Xin, Jan 8, 2022
  146. René ScharfeJan 8, 2022
  147. Han XinJan 11, 2022
  148. 4/6 object-file.c: add "stream_loose_object()" to handle large objectHan Xin, Jan 8, 2022
  149. 6/6 object-file API: add a format_object_header() functionHan Xin, Jan 8, 2022
  150. 5/6 unpack-objects: unpack_non_delta_entry() read data in a streamHan Xin, Jan 8, 2022
  151. 1/6 object-file.c: release strbuf in write_loose_object()Han Xin, Dec 17, 2021
  152. René ScharfeDec 17, 2021
  153. Junio C HamanoDec 18, 2021
  154. 2/6 object-file.c: refactor object header generation into a functionHan Xin, Dec 17, 2021
  155. object-file API: add a format_loose_header() functionÆvar Arnfjörð Bjarmason, Dec 20, 2021
  156. Philip OakleyDec 20, 2021
  157. Junio C HamanoDec 20, 2021
  158. Ævar Arnfjörð BjarmasonDec 21, 2021
  159. Junio C HamanoDec 21, 2021
  160. Ævar Arnfjörð BjarmasonDec 21, 2021
  161. Han XinDec 21, 2021
  162. 3/6 object-file.c: refactor write_loose_object() to reuse in stream versionHan Xin, Dec 17, 2021
  163. 4/6 object-file.c: make "write_object_file_flags()" to support read in streamHan Xin, Dec 17, 2021
  164. René ScharfeDec 17, 2021
  165. 5/6 unpack-objects.c: add dry_run mode for get_data()Han Xin, Dec 17, 2021
  166. René ScharfeDec 17, 2021
  167. 6/6 unpack-objects: unpack_non_delta_entry() read data in a streamHan Xin, Dec 17, 2021
  168. 1/6 object-file: refactor write_loose_object() to support read from streamHan Xin, Dec 10, 2021
  169. 2/6 object-file.c: handle undetermined oid in write_loose_object()Han Xin, Dec 10, 2021
  170. Ævar Arnfjörð BjarmasonDec 13, 2021
  171. 3/6 object-file.c: read stream in a loop in write_loose_object()Han Xin, Dec 10, 2021
  172. 4/6 unpack-objects.c: add dry_run mode for get_data()Han Xin, Dec 10, 2021
  173. 5/6 object-file.c: make "write_object_file_flags()" to support "HASH_STREAM"Han Xin, Dec 10, 2021
  174. 6/6 unpack-objects: unpack_non_delta_entry() read data in a streamHan Xin, Dec 10, 2021
  175. Ævar Arnfjörð BjarmasonDec 13, 2021
  176. 1/5 object-file: refactor write_loose_object() to read buffer from streamHan Xin, Dec 3, 2021
  177. Ævar Arnfjörð BjarmasonDec 3, 2021
  178. Han XinDec 6, 2021
  179. 2/5 object-file.c: handle undetermined oid in write_loose_object()Han Xin, Dec 3, 2021
  180. Ævar Arnfjörð BjarmasonDec 3, 2021
  181. Han XinDec 6, 2021
  182. Ævar Arnfjörð BjarmasonDec 3, 2021
  183. Han XinDec 6, 2021
  184. 3/5 object-file.c: read stream in a loop in write_loose_object()Han Xin, Dec 3, 2021
  185. 4/5 unpack-objects.c: add dry_run mode for get_data()Han Xin, Dec 3, 2021
  186. Ævar Arnfjörð BjarmasonDec 3, 2021
  187. Han XinDec 6, 2021
  188. 5/5 unpack-objects: unpack_non_delta_entry() read data in a streamHan Xin, Dec 3, 2021
  189. Ævar Arnfjörð BjarmasonDec 3, 2021
  190. Han XinDec 7, 2021
  191. Ævar Arnfjörð BjarmasonDec 3, 2021
  192. Han XinDec 7, 2021
  193. Ævar Arnfjörð BjarmasonDec 3, 2021
  194. Han XinDec 7, 2021
  195. 2/5 object-file.c: handle undetermined oid in write_loose_object()Han Xin, Nov 22, 2021
  196. Derrick StoleeNov 29, 2021
  197. Junio C HamanoNov 29, 2021
  198. Derrick StoleeNov 29, 2021
  199. Han XinNov 30, 2021
  200. 1/5 object-file: refactor write_loose_object() to read buffer from streamHan Xin, Nov 22, 2021
  201. Junio C HamanoNov 23, 2021
  202. Han XinNov 24, 2021
  203. 3/5 object-file.c: read stream in a loop in write_loose_object()Han Xin, Nov 22, 2021
  204. 4/5 unpack-objects.c: add dry_run mode for get_data()Han Xin, Nov 22, 2021
  205. 5/5 unpack-objects: unpack_non_delta_entry() read data in a streamHan Xin, Nov 22, 2021
  206. Derrick StoleeNov 29, 2021
  207. Han XinNov 30, 2021
  208. Derrick StoleeNov 30, 2021
  209. "git hyperfine" (was: [PATCH v3 5/5] unpack-objects[...])Ævar Arnfjörð Bjarmason, Dec 1, 2021
  210. Han XinDec 2, 2021
  211. Derrick StoleeDec 2, 2021

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.