{"thread":{"id":"64509","subject":"[PATCH 00/18] Refactor object read streams to work via object sources","startedAt":"2025-11-19T07:47:26Z","lastAt":"2025-11-23T19:00:49Z","messageCount":85,"participants":["Patrick Steinhardt","Karthik Nayak","Justin Tobler","Junio C Hamano"],"isPatch":true,"patchVersion":1,"patchTotal":18},"messages":[{"id":"530950","messageId":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":null,"subject":"[PATCH 00/18] Refactor object read streams to work via object sources","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:00Z","receivedAt":"2025-11-19T07:47:26Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi,\n\nthe `git_istream` data structure can be used to read objects from the\nobject database in a streaming fashion. This is used for example to read\nlarge files that one doesn't want to load into memory in full.\n\nIn the current architecture, all the logic to handle these streams is\nfully self-contained in \"streaming.c\". It contains the logic to set up\nstreams for loose, packed, in-memory and filtered objects. This doesn't\nreally play all that well with pluggable object databases, as it should\nbe the responsibility of the object database source itself to handle the\nlogic.\n\nThis patch series thus revamps our object read streams: instead of being\nentirely contained in \"streaming.c\", the format-specific streams are now\ncreated by the ODB sources. This allows each source itself to decide\nwhether and, if so, how to make objects streamable.\n\nThis overall requires quite a bit of refactoring, but I think that the\nend result is an easier-to-understand infrastructure that is an\nimprovement even without pluggable object databases.\n\nThis series is built on top of v2.52.0 with ps/object-source-loose at\n3e5e360888 (object-file: refactor writing objects via a stream,\n2025-11-03) merged into it.\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (18):\n      streaming: rename `git_istream` into `odb_read_stream`\n      streaming: drop the `open()` callback function\n      streaming: propagate final object type via the stream\n      streaming: explicitly pass packfile info when streaming a packed object\n      streaming: allocate stream inside the backend-specific logic\n      streaming: create structure for in-core object streams\n      streaming: create structure for loose object streams\n      streaming: create structure for packed object streams\n      streaming: create structure for filtered object streams\n      streaming: move zlib stream into backends\n      packfile: introduce function to read object info from a store\n      streaming: rely on object sources to create object stream\n      streaming: get rid of `the_repository`\n      streaming: make the `odb_read_stream` definition public\n      streaming: move logic to read loose objects streams into backend\n      streaming: move logic to read packed objects streams into backend\n      streaming: refactor interface to be object-database-centric\n      streaming: move into object database subsystem\n\n Makefile               |   2 +-\n archive-tar.c          |  10 +-\n archive-zip.c          |  16 +-\n builtin/cat-file.c     |   4 +-\n builtin/fsck.c         |   5 +-\n builtin/index-pack.c   |  12 +-\n builtin/log.c          |   6 +-\n builtin/pack-objects.c |  20 +-\n entry.c                |   4 +-\n meson.build            |   2 +-\n object-file.c          | 179 ++++++++++++++--\n object-file.h          |  42 +---\n odb.c                  |  29 +--\n odb/streaming.c        | 299 ++++++++++++++++++++++++++\n odb/streaming.h        |  70 ++++++\n packfile.c             | 199 ++++++++++++++++--\n packfile.h             |  17 +-\n parallel-checkout.c    |   5 +-\n streaming.c            | 561 -------------------------------------------------\n streaming.h            |  21 --\n 20 files changed, 784 insertions(+), 719 deletions(-)\n\n\n---\nbase-commit: 899e578b5b7c020aec806bd694adf2563f62843c\nchange-id: 20251107-b4-pks-odb-read-stream-7ea7f0e0a8f4\n\n"},{"id":"530951","messageId":"20251119-b4-pks-odb-read-stream-v1-1-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 01/18] streaming: rename `git_istream` into `odb_read_stream`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:01Z","receivedAt":"2025-11-19T07:47:28Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In the following patches we are about to make the `git_istream` more\ngeneric so that it becomes fully controlled by the specific object\nsource that wants to create it. As part of these refactorings we'll\nfully move the structure into the object database subsystem.\n\nPrepare for this change by renaming the structure from `git_istream`\nto `odb_read_stream`. This mirrors the `odb_write_stream` structure that\nwe already have.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  2 +-\n archive-zip.c          |  2 +-\n builtin/index-pack.c   |  2 +-\n builtin/pack-objects.c |  4 ++--\n object-file.c          |  2 +-\n streaming.c            | 62 +++++++++++++++++++++++++-------------------------\n streaming.h            | 12 +++++-----\n 7 files changed, 43 insertions(+), 43 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 73b63ddc41..dc1eda09e0 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -129,7 +129,7 @@ static void write_trailer(void)\n  */\n static int stream_blocked(struct repository *r, const struct object_id *oid)\n {\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tenum object_type type;\n \tunsigned long sz;\n \tchar buf[BLOCKSIZE];\ndiff --git a/archive-zip.c b/archive-zip.c\nindex bea5bdd43d..40a9c93ff9 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -309,7 +309,7 @@ static int write_zip_entry(struct archiver_args *args,\n \tenum zip_method method;\n \tunsigned char *out;\n \tvoid *deflated = NULL;\n-\tstruct git_istream *stream = NULL;\n+\tstruct odb_read_stream *stream = NULL;\n \tunsigned long flags = 0;\n \tint is_binary = -1;\n \tconst char *path_without_prefix = path + args->baselen;\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 2b78ba7fe4..5f90f12f92 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -762,7 +762,7 @@ static void find_ref_delta_children(const struct object_id *oid,\n \n struct compare_data {\n \tstruct object_entry *entry;\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tunsigned char *buf;\n \tunsigned long buf_size;\n };\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 69e80b1443..c693d948e1 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -404,7 +404,7 @@ static unsigned long do_compress(void **pptr, unsigned long size)\n \treturn stream.total_out;\n }\n \n-static unsigned long write_large_blob_data(struct git_istream *st, struct hashfile *f,\n+static unsigned long write_large_blob_data(struct odb_read_stream *st, struct hashfile *f,\n \t\t\t\t\t   const struct object_id *oid)\n {\n \tgit_zstream stream;\n@@ -513,7 +513,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \tunsigned hdrlen;\n \tenum object_type type;\n \tvoid *buf;\n-\tstruct git_istream *st = NULL;\n+\tstruct odb_read_stream *st = NULL;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \n \tif (!usable_delta) {\ndiff --git a/object-file.c b/object-file.c\nindex 811c569ed3..b62b21a452 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -134,7 +134,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tstruct object_id real_oid;\n \tunsigned long size;\n \tenum object_type obj_type;\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tstruct git_hash_ctx c;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\ndiff --git a/streaming.c b/streaming.c\nindex 00ad649ae3..1fb4b7c1c0 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -14,17 +14,17 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*open_istream_fn)(struct git_istream *,\n+typedef int (*open_istream_fn)(struct odb_read_stream *,\n \t\t\t       struct repository *,\n \t\t\t       const struct object_id *,\n \t\t\t       enum object_type *);\n-typedef int (*close_istream_fn)(struct git_istream *);\n-typedef ssize_t (*read_istream_fn)(struct git_istream *, char *, size_t);\n+typedef int (*close_istream_fn)(struct odb_read_stream *);\n+typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n #define FILTER_BUFFER (1024*16)\n \n struct filtered_istream {\n-\tstruct git_istream *upstream;\n+\tstruct odb_read_stream *upstream;\n \tstruct stream_filter *filter;\n \tchar ibuf[FILTER_BUFFER];\n \tchar obuf[FILTER_BUFFER];\n@@ -33,7 +33,7 @@ struct filtered_istream {\n \tint input_finished;\n };\n \n-struct git_istream {\n+struct odb_read_stream {\n \topen_istream_fn open;\n \tclose_istream_fn close;\n \tread_istream_fn read;\n@@ -71,7 +71,7 @@ struct git_istream {\n  *\n  *****************************************************************/\n \n-static void close_deflated_stream(struct git_istream *st)\n+static void close_deflated_stream(struct odb_read_stream *st)\n {\n \tif (st->z_state == z_used)\n \t\tgit_inflate_end(&st->z);\n@@ -84,13 +84,13 @@ static void close_deflated_stream(struct git_istream *st)\n  *\n  *****************************************************************/\n \n-static int close_istream_filtered(struct git_istream *st)\n+static int close_istream_filtered(struct odb_read_stream *st)\n {\n \tfree_stream_filter(st->u.filtered.filter);\n \treturn close_istream(st->u.filtered.upstream);\n }\n \n-static ssize_t read_istream_filtered(struct git_istream *st, char *buf,\n+static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n \t\t\t\t     size_t sz)\n {\n \tstruct filtered_istream *fs = &(st->u.filtered);\n@@ -150,10 +150,10 @@ static ssize_t read_istream_filtered(struct git_istream *st, char *buf,\n \treturn filled;\n }\n \n-static struct git_istream *attach_stream_filter(struct git_istream *st,\n-\t\t\t\t\t\tstruct stream_filter *filter)\n+static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n+\t\t\t\t\t\t    struct stream_filter *filter)\n {\n-\tstruct git_istream *ifs = xmalloc(sizeof(*ifs));\n+\tstruct odb_read_stream *ifs = xmalloc(sizeof(*ifs));\n \tstruct filtered_istream *fs = &(ifs->u.filtered);\n \n \tifs->close = close_istream_filtered;\n@@ -173,7 +173,7 @@ static struct git_istream *attach_stream_filter(struct git_istream *st,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_loose(struct git_istream *st, char *buf, size_t sz)\n+static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t sz)\n {\n \tsize_t total_read = 0;\n \n@@ -218,14 +218,14 @@ static ssize_t read_istream_loose(struct git_istream *st, char *buf, size_t sz)\n \treturn total_read;\n }\n \n-static int close_istream_loose(struct git_istream *st)\n+static int close_istream_loose(struct odb_read_stream *st)\n {\n \tclose_deflated_stream(st);\n \tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n \treturn 0;\n }\n \n-static int open_istream_loose(struct git_istream *st, struct repository *r,\n+static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \t\t\t      const struct object_id *oid,\n \t\t\t      enum object_type *type)\n {\n@@ -277,7 +277,7 @@ static int open_istream_loose(struct git_istream *st, struct repository *r,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_pack_non_delta(struct git_istream *st, char *buf,\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf,\n \t\t\t\t\t   size_t sz)\n {\n \tsize_t total_read = 0;\n@@ -336,13 +336,13 @@ static ssize_t read_istream_pack_non_delta(struct git_istream *st, char *buf,\n \treturn total_read;\n }\n \n-static int close_istream_pack_non_delta(struct git_istream *st)\n+static int close_istream_pack_non_delta(struct odb_read_stream *st)\n {\n \tclose_deflated_stream(st);\n \treturn 0;\n }\n \n-static int open_istream_pack_non_delta(struct git_istream *st,\n+static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n \t\t\t\t       const struct object_id *oid UNUSED,\n \t\t\t\t       enum object_type *type UNUSED)\n@@ -380,13 +380,13 @@ static int open_istream_pack_non_delta(struct git_istream *st,\n  *\n  *****************************************************************/\n \n-static int close_istream_incore(struct git_istream *st)\n+static int close_istream_incore(struct odb_read_stream *st)\n {\n \tfree(st->u.incore.buf);\n \treturn 0;\n }\n \n-static ssize_t read_istream_incore(struct git_istream *st, char *buf, size_t sz)\n+static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t sz)\n {\n \tsize_t read_size = sz;\n \tsize_t remainder = st->size - st->u.incore.read_ptr;\n@@ -400,7 +400,7 @@ static ssize_t read_istream_incore(struct git_istream *st, char *buf, size_t sz)\n \treturn read_size;\n }\n \n-static int open_istream_incore(struct git_istream *st, struct repository *r,\n+static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \t\t\t       const struct object_id *oid, enum object_type *type)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n@@ -420,7 +420,7 @@ static int open_istream_incore(struct git_istream *st, struct repository *r,\n  * static helpers variables and functions for users of streaming interface\n  *****************************************************************************/\n \n-static int istream_source(struct git_istream *st,\n+static int istream_source(struct odb_read_stream *st,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid,\n \t\t\t  enum object_type *type)\n@@ -458,25 +458,25 @@ static int istream_source(struct git_istream *st,\n  * Users of streaming interface\n  ****************************************************************/\n \n-int close_istream(struct git_istream *st)\n+int close_istream(struct odb_read_stream *st)\n {\n \tint r = st->close(st);\n \tfree(st);\n \treturn r;\n }\n \n-ssize_t read_istream(struct git_istream *st, void *buf, size_t sz)\n+ssize_t read_istream(struct odb_read_stream *st, void *buf, size_t sz)\n {\n \treturn st->read(st, buf, sz);\n }\n \n-struct git_istream *open_istream(struct repository *r,\n-\t\t\t\t const struct object_id *oid,\n-\t\t\t\t enum object_type *type,\n-\t\t\t\t unsigned long *size,\n-\t\t\t\t struct stream_filter *filter)\n+struct odb_read_stream *open_istream(struct repository *r,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     enum object_type *type,\n+\t\t\t\t     unsigned long *size,\n+\t\t\t\t     struct stream_filter *filter)\n {\n-\tstruct git_istream *st = xmalloc(sizeof(*st));\n+\tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n \tint ret = istream_source(st, r, real, type);\n \n@@ -493,7 +493,7 @@ struct git_istream *open_istream(struct repository *r,\n \t}\n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n-\t\tstruct git_istream *nst = attach_stream_filter(st, filter);\n+\t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n \t\tif (!nst) {\n \t\t\tclose_istream(st);\n \t\t\treturn NULL;\n@@ -508,7 +508,7 @@ struct git_istream *open_istream(struct repository *r,\n int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter *filter,\n \t\t      int can_seek)\n {\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tenum object_type type;\n \tunsigned long sz;\n \tssize_t kept = 0;\ndiff --git a/streaming.h b/streaming.h\nindex bd27f59e57..acf4c84338 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -7,14 +7,14 @@\n #include \"object.h\"\n \n /* opaque */\n-struct git_istream;\n+struct odb_read_stream;\n struct stream_filter;\n \n-struct git_istream *open_istream(struct repository *, const struct object_id *,\n-\t\t\t\t enum object_type *, unsigned long *,\n-\t\t\t\t struct stream_filter *);\n-int close_istream(struct git_istream *);\n-ssize_t read_istream(struct git_istream *, void *, size_t);\n+struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n+\t\t\t\t       enum object_type *, unsigned long *,\n+\t\t\t\t       struct stream_filter *);\n+int close_istream(struct odb_read_stream *);\n+ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n \n int stream_blob_to_fd(int fd, const struct object_id *, struct stream_filter *, int can_seek);\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530952","messageId":"20251119-b4-pks-odb-read-stream-v1-2-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 02/18] streaming: drop the `open()` callback function","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:02Z","receivedAt":"2025-11-19T07:47:32Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating a read stream we first populate the structure with the\nopen callback function and then subsequently call the function. This\nlayout is somewhat weird though:\n\n  - The structure needs to be allocated and partially populated with the\n    open function before we can properly initialize it.\n\n  - We never use the `open()` callback after having opened it initially.\n\nEspecially the first point creates a problem for us. In subsequent\ncommits we'll want to fully move construction of the read source into\nthe respective object sources. E.g., the loose object source will be the\none that is responsible for creating the structure. But this creates a\nproblem: if we first need to create the structure so that we can call\nthe source-specific callback we cannot fully handle creation of the\nstructure in the source itself.\n\nWe could of course work around that and have the loose object source\ncreate the structure and populate it's `open()` callback, only. But\nthis doesn't really buy us anything due to the second bullet point\nabove.\n\nInstead, drop the callback entirely and refactor `istream_source()` so\nthat we open the streams immediately. This unblocks a subsequent step,\nwhere we'll also start to allocate the structure in the source-specific\nlogic.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 40 +++++++++++++++++-----------------------\n 1 file changed, 17 insertions(+), 23 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 1fb4b7c1c0..5ce6350123 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -14,10 +14,6 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*open_istream_fn)(struct odb_read_stream *,\n-\t\t\t       struct repository *,\n-\t\t\t       const struct object_id *,\n-\t\t\t       enum object_type *);\n typedef int (*close_istream_fn)(struct odb_read_stream *);\n typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n@@ -34,7 +30,6 @@ struct filtered_istream {\n };\n \n struct odb_read_stream {\n-\topen_istream_fn open;\n \tclose_istream_fn close;\n \tread_istream_fn read;\n \n@@ -437,21 +432,25 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tst->open = open_istream_loose;\n+\t\tif (open_istream_loose(st, r, oid, type) < 0)\n+\t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n-\t\tif (!oi.u.packed.is_delta &&\n-\t\t    repo_settings_get_big_file_threshold(the_repository) < size) {\n-\t\t\tst->u.in_pack.pack = oi.u.packed.pack;\n-\t\t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\t\tst->open = open_istream_pack_non_delta;\n-\t\t\treturn 0;\n-\t\t}\n-\t\t/* fallthru */\n-\tdefault:\n-\t\tst->open = open_istream_incore;\n+\t\tif (oi.u.packed.is_delta ||\n+\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\t\tbreak;\n+\n+\t\tst->u.in_pack.pack = oi.u.packed.pack;\n+\t\tst->u.in_pack.pos = oi.u.packed.offset;\n+\t\tif (open_istream_pack_non_delta(st, r, oid, type) < 0)\n+\t\t\tbreak;\n+\n \t\treturn 0;\n+\tdefault:\n+\t\tbreak;\n \t}\n+\n+\treturn open_istream_incore(st, r, oid, type);\n }\n \n /****************************************************************\n@@ -478,19 +477,14 @@ struct odb_read_stream *open_istream(struct repository *r,\n {\n \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n-\tint ret = istream_source(st, r, real, type);\n+\tint ret;\n \n+\tret = istream_source(st, r, real, type);\n \tif (ret) {\n \t\tfree(st);\n \t\treturn NULL;\n \t}\n \n-\tif (st->open(st, r, real, type)) {\n-\t\tif (open_istream_incore(st, r, real, type)) {\n-\t\t\tfree(st);\n-\t\t\treturn NULL;\n-\t\t}\n-\t}\n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n \t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530953","messageId":"20251119-b4-pks-odb-read-stream-v1-3-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 03/18] streaming: propagate final object type via the stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:03Z","receivedAt":"2025-11-19T07:47:36Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When opening the read stream for a specific object the caller is also\nexpected to pass in a pointer to the object type. This type is passed\ndown via multiple levels and will eventually be populated with the type\nof the looked-up object.\n\nThe way we propagate down the pointer though is somewhat non-obvious.\nWhile `istream_source()` still expects the pointer and looks it up via\n`odb_read_object_info_extended()`, we also pass it down even further\ninto the format-specific callbacks that perform another lookup. This is\nquite confusing overall.\n\nRefactor the code so that the responsibility to populate the object type\nrests solely with the format-specific callbacks. This will allow us to\ndrop the call to `odb_read_object_info_extended()` in `istream_source()`\nentirely in a subsequent patch.\n\nFurthermore, instead of propagating the type via an in-pointer, we now\npropagate the type via a new field in the object stream. It already has\na `size` field, so it's only natural to have a second field that\ncontains the object type.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 30 +++++++++++++++---------------\n 1 file changed, 15 insertions(+), 15 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 5ce6350123..9596a94c58 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -33,6 +33,7 @@ struct odb_read_stream {\n \tclose_istream_fn close;\n \tread_istream_fn read;\n \n+\tenum object_type type;\n \tunsigned long size; /* inflated size of full object */\n \tgit_zstream z;\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n@@ -159,6 +160,7 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \tfs->o_end = fs->o_ptr = 0;\n \tfs->input_finished = 0;\n \tifs->size = -1; /* unknown */\n+\tifs->type = st->type;\n \treturn ifs;\n }\n \n@@ -221,14 +223,13 @@ static int close_istream_loose(struct odb_read_stream *st)\n }\n \n static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n-\t\t\t      const struct object_id *oid,\n-\t\t\t      enum object_type *type)\n+\t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_source *source;\n \n \toi.sizep = &st->size;\n-\toi.typep = type;\n+\toi.typep = &st->type;\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next) {\n@@ -249,7 +250,7 @@ static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \tcase ULHR_TOO_LONG:\n \t\tgoto error;\n \t}\n-\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || *type < 0)\n+\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n \t\tgoto error;\n \n \tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n@@ -339,8 +340,7 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \n static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED,\n-\t\t\t\t       enum object_type *type UNUSED)\n+\t\t\t\t       const struct object_id *oid UNUSED)\n {\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n@@ -361,6 +361,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n+\tst->type = in_pack_type;\n \tst->z_state = z_unused;\n \tst->close = close_istream_pack_non_delta;\n \tst->read = read_istream_pack_non_delta;\n@@ -396,7 +397,7 @@ static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t\n }\n \n static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n-\t\t\t       const struct object_id *oid, enum object_type *type)\n+\t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \n@@ -404,7 +405,7 @@ static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \tst->close = close_istream_incore;\n \tst->read = read_istream_incore;\n \n-\toi.typep = type;\n+\toi.typep = &st->type;\n \toi.sizep = &st->size;\n \toi.contentp = (void **)&st->u.incore.buf;\n \treturn odb_read_object_info_extended(r->objects, oid, &oi,\n@@ -417,14 +418,12 @@ static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \n static int istream_source(struct odb_read_stream *st,\n \t\t\t  struct repository *r,\n-\t\t\t  const struct object_id *oid,\n-\t\t\t  enum object_type *type)\n+\t\t\t  const struct object_id *oid)\n {\n \tunsigned long size;\n \tint status;\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \n-\toi.typep = type;\n \toi.sizep = &size;\n \tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n \tif (status < 0)\n@@ -432,7 +431,7 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tif (open_istream_loose(st, r, oid, type) < 0)\n+\t\tif (open_istream_loose(st, r, oid) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n@@ -442,7 +441,7 @@ static int istream_source(struct odb_read_stream *st,\n \n \t\tst->u.in_pack.pack = oi.u.packed.pack;\n \t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\tif (open_istream_pack_non_delta(st, r, oid, type) < 0)\n+\t\tif (open_istream_pack_non_delta(st, r, oid) < 0)\n \t\t\tbreak;\n \n \t\treturn 0;\n@@ -450,7 +449,7 @@ static int istream_source(struct odb_read_stream *st,\n \t\tbreak;\n \t}\n \n-\treturn open_istream_incore(st, r, oid, type);\n+\treturn open_istream_incore(st, r, oid);\n }\n \n /****************************************************************\n@@ -479,7 +478,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n \tint ret;\n \n-\tret = istream_source(st, r, real, type);\n+\tret = istream_source(st, r, real);\n \tif (ret) {\n \t\tfree(st);\n \t\treturn NULL;\n@@ -496,6 +495,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t}\n \n \t*size = st->size;\n+\t*type = st->type;\n \treturn st;\n }\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530954","messageId":"20251119-b4-pks-odb-read-stream-v1-4-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 04/18] streaming: explicitly pass packfile info when streaming a packed object","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:04Z","receivedAt":"2025-11-19T07:47:40Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When streaming a packed object we first populate the stream with\ninformation about the pack that contains the object before calling\n`open_istream_pack_non_delta()`. This is done because we have already\nlooked up both the pack and the object's offset, so it would be a waste\nof time to look up this information again.\n\nBut the way this is done makes for a somewhat awkward calling interface,\nas the caller now needs to be aware of how exactly the function itself\nbehaves.\n\nRefactor the code so that we instead explicitly pass the packfile info\ninto `open_istream_pack_non_delta()`. This makes the calling convention\nexplicit, but more importantly this allows us to refactor the function\nso that it becomes its responsibility to allocate the stream itself in a\nsubsequent patch.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 20 ++++++++++----------\n 1 file changed, 10 insertions(+), 10 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 9596a94c58..d7db446d25 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -340,16 +340,18 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \n static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED)\n+\t\t\t\t       const struct object_id *oid UNUSED,\n+\t\t\t\t       struct packed_git *pack,\n+\t\t\t\t       off_t offset)\n {\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n \n \twindow = NULL;\n \n-\tin_pack_type = unpack_object_header(st->u.in_pack.pack,\n+\tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n-\t\t\t\t\t    &st->u.in_pack.pos,\n+\t\t\t\t\t    &offset,\n \t\t\t\t\t    &st->size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n@@ -365,6 +367,8 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tst->z_state = z_unused;\n \tst->close = close_istream_pack_non_delta;\n \tst->read = read_istream_pack_non_delta;\n+\tst->u.in_pack.pack = pack;\n+\tst->u.in_pack.pos = offset;\n \n \treturn 0;\n }\n@@ -436,14 +440,10 @@ static int istream_source(struct odb_read_stream *st,\n \t\treturn 0;\n \tcase OI_PACKED:\n \t\tif (oi.u.packed.is_delta ||\n-\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n+\t\t    open_istream_pack_non_delta(st, r, oid, oi.u.packed.pack,\n+\t\t\t\t\t\toi.u.packed.offset) < 0)\n \t\t\tbreak;\n-\n-\t\tst->u.in_pack.pack = oi.u.packed.pack;\n-\t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\tif (open_istream_pack_non_delta(st, r, oid) < 0)\n-\t\t\tbreak;\n-\n \t\treturn 0;\n \tdefault:\n \t\tbreak;\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530955","messageId":"20251119-b4-pks-odb-read-stream-v1-5-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 05/18] streaming: allocate stream inside the backend-specific logic","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:05Z","receivedAt":"2025-11-19T07:47:43Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating a new stream we first allocate it and then call into\nbackend-specific logic to populate the stream. This design requires that\nthe stream itself contains a `union` with backend-specific members that\nthen ultimately get populated by the backend-specific logic.\n\nThis works, but it's awkward in the context of pluggable object\ndatabases. Each backend will need its own member in that union, and as\nthe structure itself is completely opaque (it's only defined in\n\"streamgin.c\") it also has the consequence that we must have the logic\nthat is specific to backends in \"streaming.c\".\n\nIdeally though, the infrastructure would be reversed: we have a generic\n`struct odb_read_stream` and some helper functions in \"streaming.c\",\nwhereas the backend-specific logic sits in the backend's subsystem\nitself.\n\nThis can be realized by using a design that is similar to how we handle\nreference databases: instead of having a union of members, we instead\nhave backend-specific structures with a `struct odb_read_stream base`\nas its first member. The backends would thus hand out the pointer to the\nbase, but internally they know to cast back to the backend-specific\ntype.\n\nThis means though that we need to allocate different structures\ndepending on the backend. To prepare for this, move allocation of the\nstructure into the backend-specific functions that open a new stream.\nSubsequent commits will then create those new backend-specific structs.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 99 +++++++++++++++++++++++++++++++++++++++----------------------\n 1 file changed, 63 insertions(+), 36 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex d7db446d25..b8ce82483f 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -222,27 +222,34 @@ static int close_istream_loose(struct odb_read_stream *st)\n \treturn 0;\n }\n \n-static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n+static int open_istream_loose(struct odb_read_stream **out,\n+\t\t\t      struct repository *r,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_read_stream *st;\n \tstruct odb_source *source;\n-\n-\toi.sizep = &st->size;\n-\toi.typep = &st->type;\n+\tunsigned long mapsize;\n+\tvoid *mapped;\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next) {\n-\t\tst->u.loose.mapped = odb_source_loose_map_object(source, oid,\n-\t\t\t\t\t\t\t\t &st->u.loose.mapsize);\n-\t\tif (st->u.loose.mapped)\n+\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n+\t\tif (mapped)\n \t\t\tbreak;\n \t}\n-\tif (!st->u.loose.mapped)\n+\tif (!mapped)\n \t\treturn -1;\n \n-\tswitch (unpack_loose_header(&st->z, st->u.loose.mapped,\n-\t\t\t\t    st->u.loose.mapsize, st->u.loose.hdr,\n+\t/*\n+\t * Note: we must allocate this structure early even though we may still\n+\t * fail. This is because we need to initialize the zlib stream, and it\n+\t * is not possible to copy the stream around after the fact because it\n+\t * has self-referencing pointers.\n+\t */\n+\tCALLOC_ARRAY(st, 1);\n+\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->u.loose.hdr,\n \t\t\t\t    sizeof(st->u.loose.hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n@@ -250,19 +257,28 @@ static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \tcase ULHR_TOO_LONG:\n \t\tgoto error;\n \t}\n+\n+\toi.sizep = &st->size;\n+\toi.typep = &st->type;\n+\n \tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n \t\tgoto error;\n \n+\tst->u.loose.mapped = mapped;\n+\tst->u.loose.mapsize = mapsize;\n \tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n \tst->u.loose.hdr_avail = st->z.total_out;\n \tst->z_state = z_used;\n \tst->close = close_istream_loose;\n \tst->read = read_istream_loose;\n \n+\t*out = st;\n+\n \treturn 0;\n error:\n \tgit_inflate_end(&st->z);\n \tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tfree(st);\n \treturn -1;\n }\n \n@@ -338,12 +354,16 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \treturn 0;\n }\n \n-static int open_istream_pack_non_delta(struct odb_read_stream *st,\n+static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \t\t\t\t       struct repository *r UNUSED,\n \t\t\t\t       const struct object_id *oid UNUSED,\n \t\t\t\t       struct packed_git *pack,\n \t\t\t\t       off_t offset)\n {\n+\tstruct odb_read_stream stream = {\n+\t\t.close = close_istream_pack_non_delta,\n+\t\t.read = read_istream_pack_non_delta,\n+\t};\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n \n@@ -352,7 +372,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n \t\t\t\t\t    &offset,\n-\t\t\t\t\t    &st->size);\n+\t\t\t\t\t    &stream.size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n \tdefault:\n@@ -363,12 +383,13 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n-\tst->type = in_pack_type;\n-\tst->z_state = z_unused;\n-\tst->close = close_istream_pack_non_delta;\n-\tst->read = read_istream_pack_non_delta;\n-\tst->u.in_pack.pack = pack;\n-\tst->u.in_pack.pos = offset;\n+\tstream.type = in_pack_type;\n+\tstream.z_state = z_unused;\n+\tstream.u.in_pack.pack = pack;\n+\tstream.u.in_pack.pos = offset;\n+\n+\tCALLOC_ARRAY(*out, 1);\n+\t**out = stream;\n \n \treturn 0;\n }\n@@ -400,27 +421,35 @@ static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t\n \treturn read_size;\n }\n \n-static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n+static int open_istream_incore(struct odb_read_stream **out,\n+\t\t\t       struct repository *r,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_read_stream stream = {\n+\t\t.close = close_istream_incore,\n+\t\t.read = read_istream_incore,\n+\t};\n+\tint ret;\n \n-\tst->u.incore.read_ptr = 0;\n-\tst->close = close_istream_incore;\n-\tst->read = read_istream_incore;\n+\toi.typep = &stream.type;\n+\toi.sizep = &stream.size;\n+\toi.contentp = (void **)&stream.u.incore.buf;\n+\tret = odb_read_object_info_extended(r->objects, oid, &oi,\n+\t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n+\tif (ret)\n+\t\treturn ret;\n \n-\toi.typep = &st->type;\n-\toi.sizep = &st->size;\n-\toi.contentp = (void **)&st->u.incore.buf;\n-\treturn odb_read_object_info_extended(r->objects, oid, &oi,\n-\t\t\t\t\t     OBJECT_INFO_DIE_IF_CORRUPT);\n+\tCALLOC_ARRAY(*out, 1);\n+\t**out = stream;\n+\treturn 0;\n }\n \n /*****************************************************************************\n  * static helpers variables and functions for users of streaming interface\n  *****************************************************************************/\n \n-static int istream_source(struct odb_read_stream *st,\n+static int istream_source(struct odb_read_stream **out,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid)\n {\n@@ -435,13 +464,13 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tif (open_istream_loose(st, r, oid) < 0)\n+\t\tif (open_istream_loose(out, r, oid) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n \t\tif (oi.u.packed.is_delta ||\n \t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n-\t\t    open_istream_pack_non_delta(st, r, oid, oi.u.packed.pack,\n+\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n \t\t\t\t\t\toi.u.packed.offset) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n@@ -449,7 +478,7 @@ static int istream_source(struct odb_read_stream *st,\n \t\tbreak;\n \t}\n \n-\treturn open_istream_incore(st, r, oid);\n+\treturn open_istream_incore(out, r, oid);\n }\n \n /****************************************************************\n@@ -474,15 +503,13 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t\t\t\t     unsigned long *size,\n \t\t\t\t     struct stream_filter *filter)\n {\n-\tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n+\tstruct odb_read_stream *st;\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n \tint ret;\n \n-\tret = istream_source(st, r, real);\n-\tif (ret) {\n-\t\tfree(st);\n+\tret = istream_source(&st, r, real);\n+\tif (ret)\n \t\treturn NULL;\n-\t}\n \n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530956","messageId":"20251119-b4-pks-odb-read-stream-v1-6-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 06/18] streaming: create structure for in-core object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:06Z","receivedAt":"2025-11-19T07:47:47Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for in-core object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 45 +++++++++++++++++++++++++--------------------\n 1 file changed, 25 insertions(+), 20 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex b8ce82483f..9018b10b23 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,11 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tchar *buf; /* from odb_read_object_info_extended() */\n-\t\t\tunsigned long read_ptr;\n-\t\t} incore;\n-\n \t\tstruct {\n \t\t\tvoid *mapped;\n \t\t\tunsigned long mapsize;\n@@ -401,22 +396,30 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n  *\n  *****************************************************************/\n \n-static int close_istream_incore(struct odb_read_stream *st)\n+struct odb_incore_read_stream {\n+\tstruct odb_read_stream base;\n+\tchar *buf; /* from odb_read_object_info_extended() */\n+\tunsigned long read_ptr;\n+};\n+\n+static int close_istream_incore(struct odb_read_stream *_st)\n {\n-\tfree(st->u.incore.buf);\n+\tstruct odb_incore_read_stream *st = (struct odb_incore_read_stream *)_st;\n+\tfree(st->buf);\n \treturn 0;\n }\n \n-static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t sz)\n+static ssize_t read_istream_incore(struct odb_read_stream *_st, char *buf, size_t sz)\n {\n+\tstruct odb_incore_read_stream *st = (struct odb_incore_read_stream *)_st;\n \tsize_t read_size = sz;\n-\tsize_t remainder = st->size - st->u.incore.read_ptr;\n+\tsize_t remainder = st->base.size - st->read_ptr;\n \n \tif (remainder <= read_size)\n \t\tread_size = remainder;\n \tif (read_size) {\n-\t\tmemcpy(buf, st->u.incore.buf + st->u.incore.read_ptr, read_size);\n-\t\tst->u.incore.read_ptr += read_size;\n+\t\tmemcpy(buf, st->buf + st->read_ptr, read_size);\n+\t\tst->read_ptr += read_size;\n \t}\n \treturn read_size;\n }\n@@ -426,22 +429,24 @@ static int open_istream_incore(struct odb_read_stream **out,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_read_stream stream = {\n-\t\t.close = close_istream_incore,\n-\t\t.read = read_istream_incore,\n-\t};\n+\tstruct odb_incore_read_stream stream = {\n+\t\t.base.close = close_istream_incore,\n+\t\t.base.read = read_istream_incore,\n+\t}, *st;\n \tint ret;\n \n-\toi.typep = &stream.type;\n-\toi.sizep = &stream.size;\n-\toi.contentp = (void **)&stream.u.incore.buf;\n+\toi.typep = &stream.base.type;\n+\toi.sizep = &stream.base.size;\n+\toi.contentp = (void **)&stream.buf;\n \tret = odb_read_object_info_extended(r->objects, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n \n-\tCALLOC_ARRAY(*out, 1);\n-\t**out = stream;\n+\tCALLOC_ARRAY(st, 1);\n+\t*st = stream;\n+\t*out = &st->base;\n+\n \treturn 0;\n }\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530957","messageId":"20251119-b4-pks-odb-read-stream-v1-7-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 07/18] streaming: create structure for loose object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:07Z","receivedAt":"2025-11-19T07:47:51Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for loose object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 85 ++++++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 44 insertions(+), 41 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 9018b10b23..190628c767 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,14 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tvoid *mapped;\n-\t\t\tunsigned long mapsize;\n-\t\t\tchar hdr[32];\n-\t\t\tint hdr_avail;\n-\t\t\tint hdr_used;\n-\t\t} loose;\n-\n \t\tstruct {\n \t\t\tstruct packed_git *pack;\n \t\t\toff_t pos;\n@@ -165,11 +157,21 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t sz)\n+struct odb_loose_read_stream {\n+\tstruct odb_read_stream base;\n+\tvoid *mapped;\n+\tunsigned long mapsize;\n+\tchar hdr[32];\n+\tint hdr_avail;\n+\tint hdr_used;\n+};\n+\n+static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n {\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->z_state) {\n+\tswitch (st->base.z_state) {\n \tcase z_done:\n \t\treturn 0;\n \tcase z_error:\n@@ -178,42 +180,43 @@ static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t\n \t\tbreak;\n \t}\n \n-\tif (st->u.loose.hdr_used < st->u.loose.hdr_avail) {\n-\t\tsize_t to_copy = st->u.loose.hdr_avail - st->u.loose.hdr_used;\n+\tif (st->hdr_used < st->hdr_avail) {\n+\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n \t\tif (sz < to_copy)\n \t\t\tto_copy = sz;\n-\t\tmemcpy(buf, st->u.loose.hdr + st->u.loose.hdr_used, to_copy);\n-\t\tst->u.loose.hdr_used += to_copy;\n+\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n+\t\tst->hdr_used += to_copy;\n \t\ttotal_read += to_copy;\n \t}\n \n \twhile (total_read < sz) {\n \t\tint status;\n \n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->base.z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n \n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_done;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_done;\n \t\t\tbreak;\n \t\t}\n \t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_error;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_error;\n \t\t\treturn -1;\n \t\t}\n \t}\n \treturn total_read;\n }\n \n-static int close_istream_loose(struct odb_read_stream *st)\n+static int close_istream_loose(struct odb_read_stream *_st)\n {\n-\tclose_deflated_stream(st);\n-\tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tclose_deflated_stream(&st->base);\n+\tmunmap(st->mapped, st->mapsize);\n \treturn 0;\n }\n \n@@ -222,7 +225,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_read_stream *st;\n+\tstruct odb_loose_read_stream *st;\n \tstruct odb_source *source;\n \tunsigned long mapsize;\n \tvoid *mapped;\n@@ -244,8 +247,8 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t */\n \tCALLOC_ARRAY(st, 1);\n \n-\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->u.loose.hdr,\n-\t\t\t\t    sizeof(st->u.loose.hdr))) {\n+\tswitch (unpack_loose_header(&st->base.z, mapped, mapsize, st->hdr,\n+\t\t\t\t    sizeof(st->hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n \tcase ULHR_BAD:\n@@ -253,26 +256,26 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t\tgoto error;\n \t}\n \n-\toi.sizep = &st->size;\n-\toi.typep = &st->type;\n+\toi.sizep = &st->base.size;\n+\toi.typep = &st->base.type;\n \n-\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n+\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n \t\tgoto error;\n \n-\tst->u.loose.mapped = mapped;\n-\tst->u.loose.mapsize = mapsize;\n-\tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n-\tst->u.loose.hdr_avail = st->z.total_out;\n-\tst->z_state = z_used;\n-\tst->close = close_istream_loose;\n-\tst->read = read_istream_loose;\n+\tst->mapped = mapped;\n+\tst->mapsize = mapsize;\n+\tst->hdr_used = strlen(st->hdr) + 1;\n+\tst->hdr_avail = st->base.z.total_out;\n+\tst->base.z_state = z_used;\n+\tst->base.close = close_istream_loose;\n+\tst->base.read = read_istream_loose;\n \n-\t*out = st;\n+\t*out = &st->base;\n \n \treturn 0;\n error:\n-\tgit_inflate_end(&st->z);\n-\tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tgit_inflate_end(&st->base.z);\n+\tmunmap(st->mapped, st->mapsize);\n \tfree(st);\n \treturn -1;\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530958","messageId":"20251119-b4-pks-odb-read-stream-v1-8-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 08/18] streaming: create structure for packed object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:08Z","receivedAt":"2025-11-19T07:47:54Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for packed object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 75 ++++++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 40 insertions(+), 35 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 190628c767..435ead1066 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,11 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tstruct packed_git *pack;\n-\t\t\toff_t pos;\n-\t\t} in_pack;\n-\n \t\tstruct filtered_istream filtered;\n \t} u;\n };\n@@ -287,16 +282,23 @@ static int open_istream_loose(struct odb_read_stream **out,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf,\n+struct odb_packed_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct packed_git *pack;\n+\toff_t pos;\n+};\n+\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n \t\t\t\t\t   size_t sz)\n {\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->z_state) {\n+\tswitch (st->base.z_state) {\n \tcase z_unused:\n-\t\tmemset(&st->z, 0, sizeof(st->z));\n-\t\tgit_inflate_init(&st->z);\n-\t\tst->z_state = z_used;\n+\t\tmemset(&st->base.z, 0, sizeof(st->base.z));\n+\t\tgit_inflate_init(&st->base.z);\n+\t\tst->base.z_state = z_used;\n \t\tbreak;\n \tcase z_done:\n \t\treturn 0;\n@@ -311,21 +313,21 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf\n \t\tstruct pack_window *window = NULL;\n \t\tunsigned char *mapped;\n \n-\t\tmapped = use_pack(st->u.in_pack.pack, &window,\n-\t\t\t\t  st->u.in_pack.pos, &st->z.avail_in);\n+\t\tmapped = use_pack(st->pack, &window,\n+\t\t\t\t  st->pos, &st->base.z.avail_in);\n \n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tst->z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->base.z.avail_out = sz - total_read;\n+\t\tst->base.z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n \n-\t\tst->u.in_pack.pos += st->z.next_in - mapped;\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\tst->pos += st->base.z.next_in - mapped;\n+\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n \t\tunuse_pack(&window);\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_done;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_done;\n \t\t\tbreak;\n \t\t}\n \n@@ -338,17 +340,18 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf\n \t\t * or truncated), then use_pack() catches that and will die().\n \t\t */\n \t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_error;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_error;\n \t\t\treturn -1;\n \t\t}\n \t}\n \treturn total_read;\n }\n \n-static int close_istream_pack_non_delta(struct odb_read_stream *st)\n+static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n {\n-\tclose_deflated_stream(st);\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tclose_deflated_stream(&st->base);\n \treturn 0;\n }\n \n@@ -358,19 +361,17 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \t\t\t\t       struct packed_git *pack,\n \t\t\t\t       off_t offset)\n {\n-\tstruct odb_read_stream stream = {\n-\t\t.close = close_istream_pack_non_delta,\n-\t\t.read = read_istream_pack_non_delta,\n-\t};\n+\tstruct odb_packed_read_stream *stream;\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n+\tsize_t size;\n \n \twindow = NULL;\n \n \tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n \t\t\t\t\t    &offset,\n-\t\t\t\t\t    &stream.size);\n+\t\t\t\t\t    &size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n \tdefault:\n@@ -381,13 +382,17 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n-\tstream.type = in_pack_type;\n-\tstream.z_state = z_unused;\n-\tstream.u.in_pack.pack = pack;\n-\tstream.u.in_pack.pos = offset;\n \n-\tCALLOC_ARRAY(*out, 1);\n-\t**out = stream;\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.close = close_istream_pack_non_delta;\n+\tstream->base.read = read_istream_pack_non_delta;\n+\tstream->base.type = in_pack_type;\n+\tstream->base.size = size;\n+\tstream->base.z_state = z_unused;\n+\tstream->pack = pack;\n+\tstream->pos = offset;\n+\n+\t*out = &stream->base;\n \n \treturn 0;\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530959","messageId":"20251119-b4-pks-odb-read-stream-v1-9-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 09/18] streaming: create structure for filtered object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:09Z","receivedAt":"2025-11-19T07:47:58Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for filtered object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 54 +++++++++++++++++++++++++-----------------------------\n 1 file changed, 25 insertions(+), 29 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 435ead1066..8210b21b53 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -19,16 +19,6 @@ typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n #define FILTER_BUFFER (1024*16)\n \n-struct filtered_istream {\n-\tstruct odb_read_stream *upstream;\n-\tstruct stream_filter *filter;\n-\tchar ibuf[FILTER_BUFFER];\n-\tchar obuf[FILTER_BUFFER];\n-\tint i_end, i_ptr;\n-\tint o_end, o_ptr;\n-\tint input_finished;\n-};\n-\n struct odb_read_stream {\n \tclose_istream_fn close;\n \tread_istream_fn read;\n@@ -37,10 +27,6 @@ struct odb_read_stream {\n \tunsigned long size; /* inflated size of full object */\n \tgit_zstream z;\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n-\n-\tunion {\n-\t\tstruct filtered_istream filtered;\n-\t} u;\n };\n \n /*****************************************************************\n@@ -62,16 +48,28 @@ static void close_deflated_stream(struct odb_read_stream *st)\n  *\n  *****************************************************************/\n \n-static int close_istream_filtered(struct odb_read_stream *st)\n+struct odb_filtered_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct odb_read_stream *upstream;\n+\tstruct stream_filter *filter;\n+\tchar ibuf[FILTER_BUFFER];\n+\tchar obuf[FILTER_BUFFER];\n+\tint i_end, i_ptr;\n+\tint o_end, o_ptr;\n+\tint input_finished;\n+};\n+\n+static int close_istream_filtered(struct odb_read_stream *_fs)\n {\n-\tfree_stream_filter(st->u.filtered.filter);\n-\treturn close_istream(st->u.filtered.upstream);\n+\tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n+\tfree_stream_filter(fs->filter);\n+\treturn close_istream(fs->upstream);\n }\n \n-static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n+static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n \t\t\t\t     size_t sz)\n {\n-\tstruct filtered_istream *fs = &(st->u.filtered);\n+\tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n \tsize_t filled = 0;\n \n \twhile (sz) {\n@@ -131,19 +129,17 @@ static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \t\t\t\t\t\t    struct stream_filter *filter)\n {\n-\tstruct odb_read_stream *ifs = xmalloc(sizeof(*ifs));\n-\tstruct filtered_istream *fs = &(ifs->u.filtered);\n+\tstruct odb_filtered_read_stream *fs;\n \n-\tifs->close = close_istream_filtered;\n-\tifs->read = read_istream_filtered;\n+\tCALLOC_ARRAY(fs, 1);\n+\tfs->base.close = close_istream_filtered;\n+\tfs->base.read = read_istream_filtered;\n \tfs->upstream = st;\n \tfs->filter = filter;\n-\tfs->i_end = fs->i_ptr = 0;\n-\tfs->o_end = fs->o_ptr = 0;\n-\tfs->input_finished = 0;\n-\tifs->size = -1; /* unknown */\n-\tifs->type = st->type;\n-\treturn ifs;\n+\tfs->base.size = -1; /* unknown */\n+\tfs->base.type = st->type;\n+\n+\treturn &fs->base;\n }\n \n /*****************************************************************\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530960","messageId":"20251119-b4-pks-odb-read-stream-v1-10-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 10/18] streaming: move zlib stream into backends","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:10Z","receivedAt":"2025-11-19T07:48:01Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"While all backend-specific data is now contained in a backend-specific\nstructure, we still share the zlib stream across the loose and packed\nobjects.\n\nRefactor the code and move it into the specific structures so that we\nfully detangle the different backends from one another.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 104 ++++++++++++++++++++++++++++++------------------------------\n 1 file changed, 52 insertions(+), 52 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 8210b21b53..572be98248 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -25,23 +25,8 @@ struct odb_read_stream {\n \n \tenum object_type type;\n \tunsigned long size; /* inflated size of full object */\n-\tgit_zstream z;\n-\tenum { z_unused, z_used, z_done, z_error } z_state;\n };\n \n-/*****************************************************************\n- *\n- * Common helpers\n- *\n- *****************************************************************/\n-\n-static void close_deflated_stream(struct odb_read_stream *st)\n-{\n-\tif (st->z_state == z_used)\n-\t\tgit_inflate_end(&st->z);\n-}\n-\n-\n /*****************************************************************\n  *\n  * Filtered stream\n@@ -150,6 +135,12 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \n struct odb_loose_read_stream {\n \tstruct odb_read_stream base;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_LOOSE_READ_STREAM_INUSE,\n+\t\tODB_LOOSE_READ_STREAM_DONE,\n+\t\tODB_LOOSE_READ_STREAM_ERROR,\n+\t} z_state;\n \tvoid *mapped;\n \tunsigned long mapsize;\n \tchar hdr[32];\n@@ -162,10 +153,10 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n \tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->base.z_state) {\n-\tcase z_done:\n+\tswitch (st->z_state) {\n+\tcase ODB_LOOSE_READ_STREAM_DONE:\n \t\treturn 0;\n-\tcase z_error:\n+\tcase ODB_LOOSE_READ_STREAM_ERROR:\n \t\treturn -1;\n \tdefault:\n \t\tbreak;\n@@ -183,20 +174,20 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n \twhile (total_read < sz) {\n \t\tint status;\n \n-\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->base.z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n \n-\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_done;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n \t\t\tbreak;\n \t\t}\n \t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_error;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n \t\t\treturn -1;\n \t\t}\n \t}\n@@ -206,7 +197,8 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n static int close_istream_loose(struct odb_read_stream *_st)\n {\n \tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tclose_deflated_stream(&st->base);\n+\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n \tmunmap(st->mapped, st->mapsize);\n \treturn 0;\n }\n@@ -238,7 +230,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t */\n \tCALLOC_ARRAY(st, 1);\n \n-\tswitch (unpack_loose_header(&st->base.z, mapped, mapsize, st->hdr,\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n \t\t\t\t    sizeof(st->hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n@@ -256,8 +248,8 @@ static int open_istream_loose(struct odb_read_stream **out,\n \tst->mapped = mapped;\n \tst->mapsize = mapsize;\n \tst->hdr_used = strlen(st->hdr) + 1;\n-\tst->hdr_avail = st->base.z.total_out;\n-\tst->base.z_state = z_used;\n+\tst->hdr_avail = st->z.total_out;\n+\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n \tst->base.close = close_istream_loose;\n \tst->base.read = read_istream_loose;\n \n@@ -265,7 +257,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \n \treturn 0;\n error:\n-\tgit_inflate_end(&st->base.z);\n+\tgit_inflate_end(&st->z);\n \tmunmap(st->mapped, st->mapsize);\n \tfree(st);\n \treturn -1;\n@@ -281,6 +273,13 @@ static int open_istream_loose(struct odb_read_stream **out,\n struct odb_packed_read_stream {\n \tstruct odb_read_stream base;\n \tstruct packed_git *pack;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n+\t\tODB_PACKED_READ_STREAM_INUSE,\n+\t\tODB_PACKED_READ_STREAM_DONE,\n+\t\tODB_PACKED_READ_STREAM_ERROR,\n+\t} z_state;\n \toff_t pos;\n };\n \n@@ -290,17 +289,17 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->base.z_state) {\n-\tcase z_unused:\n-\t\tmemset(&st->base.z, 0, sizeof(st->base.z));\n-\t\tgit_inflate_init(&st->base.z);\n-\t\tst->base.z_state = z_used;\n+\tswitch (st->z_state) {\n+\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n+\t\tmemset(&st->z, 0, sizeof(st->z));\n+\t\tgit_inflate_init(&st->z);\n+\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n \t\tbreak;\n-\tcase z_done:\n+\tcase ODB_PACKED_READ_STREAM_DONE:\n \t\treturn 0;\n-\tcase z_error:\n+\tcase ODB_PACKED_READ_STREAM_ERROR:\n \t\treturn -1;\n-\tcase z_used:\n+\tcase ODB_PACKED_READ_STREAM_INUSE:\n \t\tbreak;\n \t}\n \n@@ -310,20 +309,20 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \t\tunsigned char *mapped;\n \n \t\tmapped = use_pack(st->pack, &window,\n-\t\t\t\t  st->pos, &st->base.z.avail_in);\n+\t\t\t\t  st->pos, &st->z.avail_in);\n \n-\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->base.z.avail_out = sz - total_read;\n-\t\tst->base.z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tst->z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n \n-\t\tst->pos += st->base.z.next_in - mapped;\n-\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n+\t\tst->pos += st->z.next_in - mapped;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n \t\tunuse_pack(&window);\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_done;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n \t\t\tbreak;\n \t\t}\n \n@@ -336,8 +335,8 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \t\t * or truncated), then use_pack() catches that and will die().\n \t\t */\n \t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_error;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n \t\t\treturn -1;\n \t\t}\n \t}\n@@ -347,7 +346,8 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n {\n \tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tclose_deflated_stream(&st->base);\n+\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n \treturn 0;\n }\n \n@@ -384,7 +384,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tstream->base.read = read_istream_pack_non_delta;\n \tstream->base.type = in_pack_type;\n \tstream->base.size = size;\n-\tstream->base.z_state = z_unused;\n+\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n \tstream->pack = pack;\n \tstream->pos = offset;\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530961","messageId":"20251119-b4-pks-odb-read-stream-v1-11-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 11/18] packfile: introduce function to read object info from a store","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:11Z","receivedAt":"2025-11-19T07:48:05Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Extract the logic to read object info for a packed object from\n`do_oid_object_into_extended()` into a standalone function that operates\non the packfile store. This function will be used in a subsequent\ncommit.\n\nNote that this change allows us to make `find_pack_entry()` an internal\nimplementation detail. As a consequence though we have to move around\n`packfile_store_freshen_object()` so that it is defined after that\nfunction.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c      | 29 ++++---------------------\n packfile.c | 71 +++++++++++++++++++++++++++++++++++++++++++++++---------------\n packfile.h | 12 ++++++++++-\n 3 files changed, 69 insertions(+), 43 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 3ec21ef24e..f4cbee4b04 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -666,8 +666,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n {\n \tstatic struct object_info blank_oi = OBJECT_INFO_INIT;\n \tconst struct cached_object *co;\n-\tstruct pack_entry e;\n-\tint rtype;\n \tconst struct object_id *real = oid;\n \tint already_retried = 0;\n \n@@ -702,8 +700,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \twhile (1) {\n \t\tstruct odb_source *source;\n \n-\t\tif (find_pack_entry(odb->repo, real, &e))\n-\t\t\tbreak;\n+\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n+\t\t\treturn 0;\n \n \t\t/* Most likely it's a loose object. */\n \t\tfor (source = odb->sources; source; source = source->next)\n@@ -713,8 +711,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t/* Not a loose object; someone else may have just packed it. */\n \t\tif (!(flags & OBJECT_INFO_QUICK)) {\n \t\t\todb_reprepare(odb->repo->objects);\n-\t\t\tif (find_pack_entry(odb->repo, real, &e))\n-\t\t\t\tbreak;\n+\t\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n+\t\t\t\treturn 0;\n \t\t}\n \n \t\t/*\n@@ -747,25 +745,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t}\n \t\treturn -1;\n \t}\n-\n-\tif (oi == &blank_oi)\n-\t\t/*\n-\t\t * We know that the caller doesn't actually need the\n-\t\t * information below, so return early.\n-\t\t */\n-\t\treturn 0;\n-\trtype = packed_object_info(odb->repo, e.p, e.offset, oi);\n-\tif (rtype < 0) {\n-\t\tmark_bad_packed_object(e.p, real);\n-\t\treturn do_oid_object_info_extended(odb, real, oi, 0);\n-\t} else if (oi->whence == OI_PACKED) {\n-\t\toi->u.packed.offset = e.offset;\n-\t\toi->u.packed.pack = e.p;\n-\t\toi->u.packed.is_delta = (rtype == OBJ_REF_DELTA ||\n-\t\t\t\t\t rtype == OBJ_OFS_DELTA);\n-\t}\n-\n-\treturn 0;\n }\n \n static int oid_object_info_convert(struct repository *r,\ndiff --git a/packfile.c b/packfile.c\nindex 40f733dd23..b4bc40d895 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -819,22 +819,6 @@ struct packed_git *packfile_store_load_pack(struct packfile_store *store,\n \treturn p;\n }\n \n-int packfile_store_freshen_object(struct packfile_store *store,\n-\t\t\t\t  const struct object_id *oid)\n-{\n-\tstruct pack_entry e;\n-\tif (!find_pack_entry(store->odb->repo, oid, &e))\n-\t\treturn 0;\n-\tif (e.p->is_cruft)\n-\t\treturn 0;\n-\tif (e.p->freshened)\n-\t\treturn 1;\n-\tif (utime(e.p->pack_name, NULL))\n-\t\treturn 0;\n-\te.p->freshened = 1;\n-\treturn 1;\n-}\n-\n void (*report_garbage)(unsigned seen_bits, const char *path);\n \n static void report_helper(const struct string_list *list,\n@@ -2064,7 +2048,9 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static int find_pack_entry(struct repository *r,\n+\t\t\t   const struct object_id *oid,\n+\t\t\t   struct pack_entry *e)\n {\n \tstruct list_head *pos;\n \n@@ -2087,6 +2073,57 @@ int find_pack_entry(struct repository *r, const struct object_id *oid, struct pa\n \treturn 0;\n }\n \n+int packfile_store_freshen_object(struct packfile_store *store,\n+\t\t\t\t  const struct object_id *oid)\n+{\n+\tstruct pack_entry e;\n+\tif (!find_pack_entry(store->odb->repo, oid, &e))\n+\t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n+\tif (e.p->freshened)\n+\t\treturn 1;\n+\tif (utime(e.p->pack_name, NULL))\n+\t\treturn 0;\n+\te.p->freshened = 1;\n+\treturn 1;\n+}\n+\n+int packfile_store_read_object_info(struct packfile_store *store,\n+\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t    struct object_info *oi,\n+\t\t\t\t    unsigned flags UNUSED)\n+{\n+\tstatic struct object_info blank_oi = OBJECT_INFO_INIT;\n+\tstruct pack_entry e;\n+\tint rtype;\n+\n+\tif (!find_pack_entry(store->odb->repo, oid, &e))\n+\t\treturn 1;\n+\n+\t/*\n+\t * We know that the caller doesn't actually need the\n+\t * information below, so return early.\n+\t */\n+\tif (oi == &blank_oi)\n+\t\treturn 0;\n+\n+\trtype = packed_object_info(store->odb->repo, e.p, e.offset, oi);\n+\tif (rtype < 0) {\n+\t\tmark_bad_packed_object(e.p, oid);\n+\t\treturn -1;\n+\t}\n+\n+\tif (oi->whence == OI_PACKED) {\n+\t\toi->u.packed.offset = e.offset;\n+\t\toi->u.packed.pack = e.p;\n+\t\toi->u.packed.is_delta = (rtype == OBJ_REF_DELTA ||\n+\t\t\t\t\t rtype == OBJ_OFS_DELTA);\n+\t}\n+\n+\treturn 0;\n+}\n+\n static void maybe_invalidate_kept_pack_cache(struct repository *r,\n \t\t\t\t\t     unsigned flags)\n {\ndiff --git a/packfile.h b/packfile.h\nindex 58fcc88e20..0a98bddd81 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -144,6 +144,17 @@ void packfile_store_add_pack(struct packfile_store *store,\n #define repo_for_each_pack(repo, p) \\\n \tfor (p = packfile_store_get_packs(repo->objects->packfiles); p; p = p->next)\n \n+/*\n+ * Try to read the object identified by its ID from the object store and\n+ * populate the object info with its data. Returns 1 in case the object was\n+ * not found, 0 if it was and read successfully, and a negative error code in\n+ * case the object was corrupted.\n+ */\n+int packfile_store_read_object_info(struct packfile_store *store,\n+\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t    struct object_info *oi,\n+\t\t\t\t    unsigned flags);\n+\n /*\n  * Get all packs managed by the given store, including packfiles that are\n  * referenced by multi-pack indices.\n@@ -357,7 +368,6 @@ const struct packed_git *has_packed_and_bad(struct repository *, const struct ob\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\n  */\n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e);\n int find_kept_pack_entry(struct repository *r, const struct object_id *oid, unsigned flags, struct pack_entry *e);\n \n int has_object_pack(struct repository *r, const struct object_id *oid);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530962","messageId":"20251119-b4-pks-odb-read-stream-v1-12-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 12/18] streaming: rely on object sources to create object stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:12Z","receivedAt":"2025-11-19T07:48:08Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating an object stream we first look up the object info and, if\nit's present, we call into the respective backend that contains the\nobject to create a new stream for it.\n\nThis has the consequence that, for loose object source, we basically\niterate through the object sources twice: we first discover that the\nfile exists as a loose object in the first place by iterating through\nall sources. And, once we have discovered it, we again walk through all\nsources to try and map the object. The same issue will eventually also\nsurface once the packfile store becomes per-object-source.\n\nFurthermore, it feels rather pointless to first look up the object only\nto then try and read it.\n\nRefactor the logic to be centered around sources instead. Instead of\nfirst reading the object, we immediately ask the source to create the\nobject stream for us. If the object exists we get stream, otherwise\nwe'll try the next source.\n\nLike this we only have to iterate through sources once. But even more\nimportantly, this change also helps us to make the whole logic\npluggable. The object read stream subsystem does not need to be aware of\nthe different source backends anymore, but eventually it'll only have to\ncall the source's callback function.\n\nNote that at the current poin in time we aren't full there yet:\n\n  - The packfile store still sits on the object database level and is\n    thus agnostic of the sources.\n\n  - We still have to call into both the packfile store and the loose\n    object source.\n\nBut both of these issues will soon be addressed.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 65 +++++++++++++++++++++++--------------------------------------\n 1 file changed, 24 insertions(+), 41 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 572be98248..bebb434cd1 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -204,21 +204,15 @@ static int close_istream_loose(struct odb_read_stream *_st)\n }\n \n static int open_istream_loose(struct odb_read_stream **out,\n-\t\t\t      struct repository *r,\n+\t\t\t      struct odb_source *source,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_loose_read_stream *st;\n-\tstruct odb_source *source;\n \tunsigned long mapsize;\n \tvoid *mapped;\n \n-\todb_prepare_alternates(r->objects);\n-\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n-\t\tif (mapped)\n-\t\t\tbreak;\n-\t}\n+\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n \tif (!mapped)\n \t\treturn -1;\n \n@@ -352,21 +346,25 @@ static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n }\n \n static int open_istream_pack_non_delta(struct odb_read_stream **out,\n-\t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED,\n-\t\t\t\t       struct packed_git *pack,\n-\t\t\t\t       off_t offset)\n+\t\t\t\t       struct object_database *odb,\n+\t\t\t\t       const struct object_id *oid)\n {\n \tstruct odb_packed_read_stream *stream;\n-\tstruct pack_window *window;\n+\tstruct pack_window *window = NULL;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n \tenum object_type in_pack_type;\n-\tsize_t size;\n+\tunsigned long size;\n \n-\twindow = NULL;\n+\toi.sizep = &size;\n+\n+\tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n+\t    oi.u.packed.is_delta ||\n+\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\treturn -1;\n \n-\tin_pack_type = unpack_object_header(pack,\n+\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n \t\t\t\t\t    &window,\n-\t\t\t\t\t    &offset,\n+\t\t\t\t\t    &oi.u.packed.offset,\n \t\t\t\t\t    &size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n@@ -385,8 +383,8 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tstream->base.type = in_pack_type;\n \tstream->base.size = size;\n \tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n-\tstream->pack = pack;\n-\tstream->pos = offset;\n+\tstream->pack = oi.u.packed.pack;\n+\tstream->pos = oi.u.packed.offset;\n \n \t*out = &stream->base;\n \n@@ -462,30 +460,15 @@ static int istream_source(struct odb_read_stream **out,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid)\n {\n-\tunsigned long size;\n-\tint status;\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\n-\toi.sizep = &size;\n-\tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n-\tif (status < 0)\n-\t\treturn status;\n+\tstruct odb_source *source;\n \n-\tswitch (oi.whence) {\n-\tcase OI_LOOSE:\n-\t\tif (open_istream_loose(out, r, oid) < 0)\n-\t\t\tbreak;\n-\t\treturn 0;\n-\tcase OI_PACKED:\n-\t\tif (oi.u.packed.is_delta ||\n-\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n-\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n-\t\t\t\t\t\toi.u.packed.offset) < 0)\n-\t\t\tbreak;\n+\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n \t\treturn 0;\n-\tdefault:\n-\t\tbreak;\n-\t}\n+\n+\todb_prepare_alternates(r->objects);\n+\tfor (source = r->objects->sources; source; source = source->next)\n+\t\tif (!open_istream_loose(out, source, oid))\n+\t\t\treturn 0;\n \n \treturn open_istream_incore(out, r, oid);\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530963","messageId":"20251119-b4-pks-odb-read-stream-v1-13-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 13/18] streaming: get rid of `the_repository`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:13Z","receivedAt":"2025-11-19T07:48:12Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Subsequent commits will move the backend-specific logic of object\nstreaming into their respective subsystems. These subsystems have gotten\nrid of `the_repository` already, but we still use it in two locations in\nthe streaming subsystem.\n\nPrepare for the move by fixing those two cases. Converting the logic in\n`open_istream_pack_non_delta()` is trivial as we already got the object\ndatabase as input.\n\nBut for `stream_blob_to_fd()` we have to add a new parameter to make it\naccessible. So, as we already have to adjust all callers anyway, rename\nthe function to `odb_stream_blob_to_fd()` to indicate it's part of the\nobject subsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/cat-file.c  |  2 +-\n builtin/fsck.c      |  3 ++-\n builtin/log.c       |  4 ++--\n entry.c             |  2 +-\n parallel-checkout.c |  3 ++-\n streaming.c         | 13 +++++++------\n streaming.h         | 18 +++++++++++++++++-\n 7 files changed, 32 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 983ecec837..120d626d66 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -95,7 +95,7 @@ static int filter_object(const char *path, unsigned mode,\n \n static int stream_blob(const struct object_id *oid)\n {\n-\tif (stream_blob_to_fd(1, oid, NULL, 0))\n+\tif (odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0))\n \t\tdie(\"unable to stream %s to stdout\", oid_to_hex(oid));\n \treturn 0;\n }\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex b1a650c673..1a348d43c2 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -340,7 +340,8 @@ static void check_unreachable_object(struct object *obj)\n \t\t\t}\n \t\t\tf = xfopen(filename, \"w\");\n \t\t\tif (obj->type == OBJ_BLOB) {\n-\t\t\t\tif (stream_blob_to_fd(fileno(f), &obj->oid, NULL, 1))\n+\t\t\t\tif (odb_stream_blob_to_fd(the_repository->objects, fileno(f),\n+\t\t\t\t\t\t\t  &obj->oid, NULL, 1))\n \t\t\t\t\tdie_errno(_(\"could not write '%s'\"), filename);\n \t\t\t} else\n \t\t\t\tfprintf(f, \"%s\\n\", describe_object(&obj->oid));\ndiff --git a/builtin/log.c b/builtin/log.c\nindex c8319b8af3..e7b83a6e00 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -584,7 +584,7 @@ static int show_blob_object(const struct object_id *oid, struct rev_info *rev, c\n \tfflush(rev->diffopt.file);\n \tif (!rev->diffopt.flags.textconv_set_via_cmdline ||\n \t    !rev->diffopt.flags.allow_textconv)\n-\t\treturn stream_blob_to_fd(1, oid, NULL, 0);\n+\t\treturn odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0);\n \n \tif (get_oid_with_context(the_repository, obj_name,\n \t\t\t\t GET_OID_RECORD_PATH,\n@@ -594,7 +594,7 @@ static int show_blob_object(const struct object_id *oid, struct rev_info *rev, c\n \t    !textconv_object(the_repository, obj_context.path,\n \t\t\t     obj_context.mode, &oidc, 1, &buf, &size)) {\n \t\tobject_context_release(&obj_context);\n-\t\treturn stream_blob_to_fd(1, oid, NULL, 0);\n+\t\treturn odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0);\n \t}\n \n \tif (!buf)\ndiff --git a/entry.c b/entry.c\nindex cae02eb503..38dfe670f7 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -139,7 +139,7 @@ static int streaming_write_entry(const struct cache_entry *ce, char *path,\n \tif (fd < 0)\n \t\treturn -1;\n \n-\tresult |= stream_blob_to_fd(fd, &ce->oid, filter, 1);\n+\tresult |= odb_stream_blob_to_fd(the_repository->objects, fd, &ce->oid, filter, 1);\n \t*fstat_done = fstat_checkout_output(fd, state, statbuf);\n \tresult |= close(fd);\n \ndiff --git a/parallel-checkout.c b/parallel-checkout.c\nindex fba6aa65a6..1cb6701b92 100644\n--- a/parallel-checkout.c\n+++ b/parallel-checkout.c\n@@ -281,7 +281,8 @@ static int write_pc_item_to_fd(struct parallel_checkout_item *pc_item, int fd,\n \n \tfilter = get_stream_filter_ca(&pc_item->ca, &pc_item->ce->oid);\n \tif (filter) {\n-\t\tif (stream_blob_to_fd(fd, &pc_item->ce->oid, filter, 1)) {\n+\t\tif (odb_stream_blob_to_fd(the_repository->objects, fd,\n+\t\t\t\t\t  &pc_item->ce->oid, filter, 1)) {\n \t\t\t/* On error, reset fd to try writing without streaming */\n \t\t\tif (reset_fd(fd, path))\n \t\t\t\treturn -1;\ndiff --git a/streaming.c b/streaming.c\nindex bebb434cd1..9e20e9a882 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -2,8 +2,6 @@\n  * Copyright (c) 2011, Google Inc.\n  */\n \n-#define USE_THE_REPOSITORY_VARIABLE\n-\n #include \"git-compat-util.h\"\n #include \"convert.h\"\n #include \"environment.h\"\n@@ -359,7 +357,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \n \tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n \t    oi.u.packed.is_delta ||\n-\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t    repo_settings_get_big_file_threshold(odb->repo) >= size)\n \t\treturn -1;\n \n \tin_pack_type = unpack_object_header(oi.u.packed.pack,\n@@ -518,8 +516,11 @@ struct odb_read_stream *open_istream(struct repository *r,\n \treturn st;\n }\n \n-int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter *filter,\n-\t\t      int can_seek)\n+int odb_stream_blob_to_fd(struct object_database *odb,\n+\t\t\t  int fd,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  struct stream_filter *filter,\n+\t\t\t  int can_seek)\n {\n \tstruct odb_read_stream *st;\n \tenum object_type type;\n@@ -527,7 +528,7 @@ int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = open_istream(the_repository, oid, &type, &sz, filter);\n+\tst = open_istream(odb->repo, oid, &type, &sz, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\ndiff --git a/streaming.h b/streaming.h\nindex acf4c84338..95c2a434fa 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -7,6 +7,7 @@\n #include \"object.h\"\n \n /* opaque */\n+struct object_database;\n struct odb_read_stream;\n struct stream_filter;\n \n@@ -16,6 +17,21 @@ struct odb_read_stream *open_istream(struct repository *, const struct object_id\n int close_istream(struct odb_read_stream *);\n ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n \n-int stream_blob_to_fd(int fd, const struct object_id *, struct stream_filter *, int can_seek);\n+/*\n+ * Look up the object by its ID and write the full contents to the file\n+ * descriptor. The object must be a blob, or the function will fail. When\n+ * provided, the filter is used to transform the blob contents.\n+ *\n+ * `can_seek` should be set to 1 in case the given file descriptor can be\n+ * seek(3p)'d on. This is used to support files with holes in case a\n+ * significant portion of the blob contains NUL bytes.\n+ *\n+ * Returns a negative error code on failure, 0 on success.\n+ */\n+int odb_stream_blob_to_fd(struct object_database *odb,\n+\t\t\t  int fd,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  struct stream_filter *filter,\n+\t\t\t  int can_seek);\n \n #endif /* STREAMING_H */\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530964","messageId":"20251119-b4-pks-odb-read-stream-v1-14-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 14/18] streaming: make the `odb_read_stream` definition public","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:14Z","receivedAt":"2025-11-19T07:48:16Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Subsequent commits will move the backend-specific logic of setting up an\nobject read stream into the specific subsystems. As the backends are now\nthe ones that are responsible for allocating the stream they'll need to\nhave the stream definition available to them.\n\nMake the stream definition public to prepare for this.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 11 -----------\n streaming.h | 15 ++++++++++++++-\n 2 files changed, 14 insertions(+), 12 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 9e20e9a882..3f94bd2a03 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -12,19 +12,8 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*close_istream_fn)(struct odb_read_stream *);\n-typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n-\n #define FILTER_BUFFER (1024*16)\n \n-struct odb_read_stream {\n-\tclose_istream_fn close;\n-\tread_istream_fn read;\n-\n-\tenum object_type type;\n-\tunsigned long size; /* inflated size of full object */\n-};\n-\n /*****************************************************************\n  *\n  * Filtered stream\ndiff --git a/streaming.h b/streaming.h\nindex 95c2a434fa..3a850e3efc 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -6,11 +6,24 @@\n \n #include \"object.h\"\n \n-/* opaque */\n struct object_database;\n struct odb_read_stream;\n struct stream_filter;\n \n+typedef int (*odb_read_stream_close_fn)(struct odb_read_stream *);\n+typedef ssize_t (*odb_read_stream_read_fn)(struct odb_read_stream *, char *, size_t);\n+\n+/*\n+ * A stream that can be used to read an object from the object database without\n+ * loading all of it into memory.\n+ */\n+struct odb_read_stream {\n+\todb_read_stream_close_fn close;\n+\todb_read_stream_read_fn read;\n+\tenum object_type type;\n+\tunsigned long size; /* inflated size of full object */\n+};\n+\n struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n \t\t\t\t       enum object_type *, unsigned long *,\n \t\t\t\t       struct stream_filter *);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530965","messageId":"20251119-b4-pks-odb-read-stream-v1-15-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 15/18] streaming: move logic to read loose objects streams into backend","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:15Z","receivedAt":"2025-11-19T07:48:19Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Move the logic to read loose object streams into the respective\nsubsystem. This allows us to make a couple of function declarations\nprivate.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-file.c | 167 ++++++++++++++++++++++++++++++++++++++++++++++++++++++----\n object-file.h |  42 ++-------------\n streaming.c   | 133 +---------------------------------------------\n 3 files changed, 164 insertions(+), 178 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b62b21a452..8c67847fea 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -234,9 +234,9 @@ static void *map_fd(int fd, const char *path, unsigned long *size)\n \treturn map;\n }\n \n-void *odb_source_loose_map_object(struct odb_source *source,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  unsigned long *size)\n+static void *odb_source_loose_map_object(struct odb_source *source,\n+\t\t\t\t\t const struct object_id *oid,\n+\t\t\t\t\t unsigned long *size)\n {\n \tconst char *p;\n \tint fd = open_loose_object(source->loose, oid, &p);\n@@ -246,11 +246,29 @@ void *odb_source_loose_map_object(struct odb_source *source,\n \treturn map_fd(fd, p, size);\n }\n \n-enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n-\t\t\t\t\t\t    unsigned char *map,\n-\t\t\t\t\t\t    unsigned long mapsize,\n-\t\t\t\t\t\t    void *buffer,\n-\t\t\t\t\t\t    unsigned long bufsiz)\n+enum unpack_loose_header_result {\n+\tULHR_OK,\n+\tULHR_BAD,\n+\tULHR_TOO_LONG,\n+};\n+\n+/**\n+ * unpack_loose_header() initializes the data stream needed to unpack\n+ * a loose object header.\n+ *\n+ * Returns:\n+ *\n+ * - ULHR_OK on success\n+ * - ULHR_BAD on error\n+ * - ULHR_TOO_LONG if the header was too long\n+ *\n+ * It will only parse up to MAX_HEADER_LEN bytes.\n+ */\n+static enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n+\t\t\t\t\t\t\t   unsigned char *map,\n+\t\t\t\t\t\t\t   unsigned long mapsize,\n+\t\t\t\t\t\t\t   void *buffer,\n+\t\t\t\t\t\t\t   unsigned long bufsiz)\n {\n \tint status;\n \n@@ -329,11 +347,18 @@ static void *unpack_loose_rest(git_zstream *stream,\n }\n \n /*\n+ * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n+ * object. If it doesn't follow that format -1 is returned. To check\n+ * the validity of the <type> populate the \"typep\" in the \"struct\n+ * object_info\". It will be OBJ_BAD if the object type is unknown. The\n+ * parsed <len> can be retrieved via \"oi->sizep\", and from there\n+ * passed to unpack_loose_rest().\n+ *\n  * We used to just use \"sscanf()\", but that's actually way\n  * too permissive for what we want to check. So do an anal\n  * object header parse by hand.\n  */\n-int parse_loose_header(const char *hdr, struct object_info *oi)\n+static int parse_loose_header(const char *hdr, struct object_info *oi)\n {\n \tconst char *type_buf = hdr;\n \tsize_t size;\n@@ -1976,3 +2001,127 @@ void odb_source_loose_free(struct odb_source_loose *loose)\n \tloose_object_map_clear(&loose->map);\n \tfree(loose);\n }\n+\n+struct odb_loose_read_stream {\n+\tstruct odb_read_stream base;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_LOOSE_READ_STREAM_INUSE,\n+\t\tODB_LOOSE_READ_STREAM_DONE,\n+\t\tODB_LOOSE_READ_STREAM_ERROR,\n+\t} z_state;\n+\tvoid *mapped;\n+\tunsigned long mapsize;\n+\tchar hdr[32];\n+\tint hdr_avail;\n+\tint hdr_used;\n+};\n+\n+static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n+{\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tsize_t total_read = 0;\n+\n+\tswitch (st->z_state) {\n+\tcase ODB_LOOSE_READ_STREAM_DONE:\n+\t\treturn 0;\n+\tcase ODB_LOOSE_READ_STREAM_ERROR:\n+\t\treturn -1;\n+\tdefault:\n+\t\tbreak;\n+\t}\n+\n+\tif (st->hdr_used < st->hdr_avail) {\n+\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n+\t\tif (sz < to_copy)\n+\t\t\tto_copy = sz;\n+\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n+\t\tst->hdr_used += to_copy;\n+\t\ttotal_read += to_copy;\n+\t}\n+\n+\twhile (total_read < sz) {\n+\t\tint status;\n+\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\n+\t\tif (status == Z_STREAM_END) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n+\t\t\tbreak;\n+\t\t}\n+\t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn total_read;\n+}\n+\n+static int close_istream_loose(struct odb_read_stream *_st)\n+{\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n+\tmunmap(st->mapped, st->mapsize);\n+\treturn 0;\n+}\n+\n+int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\tstruct odb_source *source,\n+\t\t\t\t\tconst struct object_id *oid)\n+{\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_loose_read_stream *st;\n+\tunsigned long mapsize;\n+\tvoid *mapped;\n+\n+\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n+\tif (!mapped)\n+\t\treturn -1;\n+\n+\t/*\n+\t * Note: we must allocate this structure early even though we may still\n+\t * fail. This is because we need to initialize the zlib stream, and it\n+\t * is not possible to copy the stream around after the fact because it\n+\t * has self-referencing pointers.\n+\t */\n+\tCALLOC_ARRAY(st, 1);\n+\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n+\t\t\t\t    sizeof(st->hdr))) {\n+\tcase ULHR_OK:\n+\t\tbreak;\n+\tcase ULHR_BAD:\n+\tcase ULHR_TOO_LONG:\n+\t\tgoto error;\n+\t}\n+\n+\toi.sizep = &st->base.size;\n+\toi.typep = &st->base.type;\n+\n+\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n+\t\tgoto error;\n+\n+\tst->mapped = mapped;\n+\tst->mapsize = mapsize;\n+\tst->hdr_used = strlen(st->hdr) + 1;\n+\tst->hdr_avail = st->z.total_out;\n+\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n+\tst->base.close = close_istream_loose;\n+\tst->base.read = read_istream_loose;\n+\n+\t*out = &st->base;\n+\n+\treturn 0;\n+error:\n+\tgit_inflate_end(&st->z);\n+\tmunmap(st->mapped, st->mapsize);\n+\tfree(st);\n+\treturn -1;\n+}\ndiff --git a/object-file.h b/object-file.h\nindex eeffa67bbd..1229d5f675 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -16,6 +16,8 @@ enum {\n int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n \n+struct object_info;\n+struct odb_read_stream;\n struct odb_source;\n \n struct odb_source_loose {\n@@ -47,9 +49,9 @@ int odb_source_loose_read_object_info(struct odb_source *source,\n \t\t\t\t      const struct object_id *oid,\n \t\t\t\t      struct object_info *oi, int flags);\n \n-void *odb_source_loose_map_object(struct odb_source *source,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  unsigned long *size);\n+int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\tstruct odb_source *source,\n+\t\t\t\t\tconst struct object_id *oid);\n \n /*\n  * Return true iff an object database source has a loose object\n@@ -143,40 +145,6 @@ int for_each_loose_object(struct object_database *odb,\n int format_object_header(char *str, size_t size, enum object_type type,\n \t\t\t size_t objsize);\n \n-/**\n- * unpack_loose_header() initializes the data stream needed to unpack\n- * a loose object header.\n- *\n- * Returns:\n- *\n- * - ULHR_OK on success\n- * - ULHR_BAD on error\n- * - ULHR_TOO_LONG if the header was too long\n- *\n- * It will only parse up to MAX_HEADER_LEN bytes.\n- */\n-enum unpack_loose_header_result {\n-\tULHR_OK,\n-\tULHR_BAD,\n-\tULHR_TOO_LONG,\n-};\n-enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n-\t\t\t\t\t\t    unsigned char *map,\n-\t\t\t\t\t\t    unsigned long mapsize,\n-\t\t\t\t\t\t    void *buffer,\n-\t\t\t\t\t\t    unsigned long bufsiz);\n-\n-/**\n- * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n- * object. If it doesn't follow that format -1 is returned. To check\n- * the validity of the <type> populate the \"typep\" in the \"struct\n- * object_info\". It will be OBJ_BAD if the object type is unknown. The\n- * parsed <len> can be retrieved via \"oi->sizep\", and from there\n- * passed to unpack_loose_rest().\n- */\n-struct object_info;\n-int parse_loose_header(const char *hdr, struct object_info *oi);\n-\n int force_object_loose(struct odb_source *source,\n \t\t       const struct object_id *oid, time_t mtime);\n \ndiff --git a/streaming.c b/streaming.c\nindex 3f94bd2a03..216576857f 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -114,137 +114,6 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \treturn &fs->base;\n }\n \n-/*****************************************************************\n- *\n- * Loose object stream\n- *\n- *****************************************************************/\n-\n-struct odb_loose_read_stream {\n-\tstruct odb_read_stream base;\n-\tgit_zstream z;\n-\tenum {\n-\t\tODB_LOOSE_READ_STREAM_INUSE,\n-\t\tODB_LOOSE_READ_STREAM_DONE,\n-\t\tODB_LOOSE_READ_STREAM_ERROR,\n-\t} z_state;\n-\tvoid *mapped;\n-\tunsigned long mapsize;\n-\tchar hdr[32];\n-\tint hdr_avail;\n-\tint hdr_used;\n-};\n-\n-static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n-{\n-\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tsize_t total_read = 0;\n-\n-\tswitch (st->z_state) {\n-\tcase ODB_LOOSE_READ_STREAM_DONE:\n-\t\treturn 0;\n-\tcase ODB_LOOSE_READ_STREAM_ERROR:\n-\t\treturn -1;\n-\tdefault:\n-\t\tbreak;\n-\t}\n-\n-\tif (st->hdr_used < st->hdr_avail) {\n-\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n-\t\tif (sz < to_copy)\n-\t\t\tto_copy = sz;\n-\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n-\t\tst->hdr_used += to_copy;\n-\t\ttotal_read += to_copy;\n-\t}\n-\n-\twhile (total_read < sz) {\n-\t\tint status;\n-\n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n-\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n-\n-\t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n-\t\t\tbreak;\n-\t\t}\n-\t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n-\t\t\treturn -1;\n-\t\t}\n-\t}\n-\treturn total_read;\n-}\n-\n-static int close_istream_loose(struct odb_read_stream *_st)\n-{\n-\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n-\t\tgit_inflate_end(&st->z);\n-\tmunmap(st->mapped, st->mapsize);\n-\treturn 0;\n-}\n-\n-static int open_istream_loose(struct odb_read_stream **out,\n-\t\t\t      struct odb_source *source,\n-\t\t\t      const struct object_id *oid)\n-{\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_loose_read_stream *st;\n-\tunsigned long mapsize;\n-\tvoid *mapped;\n-\n-\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n-\tif (!mapped)\n-\t\treturn -1;\n-\n-\t/*\n-\t * Note: we must allocate this structure early even though we may still\n-\t * fail. This is because we need to initialize the zlib stream, and it\n-\t * is not possible to copy the stream around after the fact because it\n-\t * has self-referencing pointers.\n-\t */\n-\tCALLOC_ARRAY(st, 1);\n-\n-\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n-\t\t\t\t    sizeof(st->hdr))) {\n-\tcase ULHR_OK:\n-\t\tbreak;\n-\tcase ULHR_BAD:\n-\tcase ULHR_TOO_LONG:\n-\t\tgoto error;\n-\t}\n-\n-\toi.sizep = &st->base.size;\n-\toi.typep = &st->base.type;\n-\n-\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n-\t\tgoto error;\n-\n-\tst->mapped = mapped;\n-\tst->mapsize = mapsize;\n-\tst->hdr_used = strlen(st->hdr) + 1;\n-\tst->hdr_avail = st->z.total_out;\n-\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n-\tst->base.close = close_istream_loose;\n-\tst->base.read = read_istream_loose;\n-\n-\t*out = &st->base;\n-\n-\treturn 0;\n-error:\n-\tgit_inflate_end(&st->z);\n-\tmunmap(st->mapped, st->mapsize);\n-\tfree(st);\n-\treturn -1;\n-}\n-\n-\n /*****************************************************************\n  *\n  * Non-delta packed object stream\n@@ -454,7 +323,7 @@ static int istream_source(struct odb_read_stream **out,\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next)\n-\t\tif (!open_istream_loose(out, source, oid))\n+\t\tif (!odb_source_loose_read_object_stream(out, source, oid))\n \t\t\treturn 0;\n \n \treturn open_istream_incore(out, r, oid);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530966","messageId":"20251119-b4-pks-odb-read-stream-v1-16-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 16/18] streaming: move logic to read packed objects streams into backend","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:16Z","receivedAt":"2025-11-19T07:48:23Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Move the logic to read packed object streams into the respective\nsubsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n packfile.c  | 128 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n packfile.h  |   5 +++\n streaming.c | 136 +-----------------------------------------------------------\n 3 files changed, 134 insertions(+), 135 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex b4bc40d895..ad56ce0b90 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -20,6 +20,7 @@\n #include \"tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"streaming.h\"\n #include \"midx.h\"\n #include \"commit-graph.h\"\n #include \"pack-revindex.h\"\n@@ -2406,3 +2407,130 @@ void packfile_store_close(struct packfile_store *store)\n \t\tclose_pack(p);\n \t}\n }\n+\n+struct odb_packed_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct packed_git *pack;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n+\t\tODB_PACKED_READ_STREAM_INUSE,\n+\t\tODB_PACKED_READ_STREAM_DONE,\n+\t\tODB_PACKED_READ_STREAM_ERROR,\n+\t} z_state;\n+\toff_t pos;\n+};\n+\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n+\t\t\t\t\t   size_t sz)\n+{\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tsize_t total_read = 0;\n+\n+\tswitch (st->z_state) {\n+\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n+\t\tmemset(&st->z, 0, sizeof(st->z));\n+\t\tgit_inflate_init(&st->z);\n+\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n+\t\tbreak;\n+\tcase ODB_PACKED_READ_STREAM_DONE:\n+\t\treturn 0;\n+\tcase ODB_PACKED_READ_STREAM_ERROR:\n+\t\treturn -1;\n+\tcase ODB_PACKED_READ_STREAM_INUSE:\n+\t\tbreak;\n+\t}\n+\n+\twhile (total_read < sz) {\n+\t\tint status;\n+\t\tstruct pack_window *window = NULL;\n+\t\tunsigned char *mapped;\n+\n+\t\tmapped = use_pack(st->pack, &window,\n+\t\t\t\t  st->pos, &st->z.avail_in);\n+\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tst->z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\n+\t\tst->pos += st->z.next_in - mapped;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\tunuse_pack(&window);\n+\n+\t\tif (status == Z_STREAM_END) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\t/*\n+\t\t * Unlike the loose object case, we do not have to worry here\n+\t\t * about running out of input bytes and spinning infinitely. If\n+\t\t * we get Z_BUF_ERROR due to too few input bytes, then we'll\n+\t\t * replenish them in the next use_pack() call when we loop. If\n+\t\t * we truly hit the end of the pack (i.e., because it's corrupt\n+\t\t * or truncated), then use_pack() catches that and will die().\n+\t\t */\n+\t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn total_read;\n+}\n+\n+static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n+{\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n+\treturn 0;\n+}\n+\n+int packfile_store_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t      struct packfile_store *store,\n+\t\t\t\t      const struct object_id *oid)\n+{\n+\tstruct odb_packed_read_stream *stream;\n+\tstruct pack_window *window = NULL;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\tenum object_type in_pack_type;\n+\tunsigned long size;\n+\n+\toi.sizep = &size;\n+\n+\tif (packfile_store_read_object_info(store, oid, &oi, 0) ||\n+\t    oi.u.packed.is_delta ||\n+\t    repo_settings_get_big_file_threshold(store->odb->repo) >= size)\n+\t\treturn -1;\n+\n+\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n+\t\t\t\t\t    &window,\n+\t\t\t\t\t    &oi.u.packed.offset,\n+\t\t\t\t\t    &size);\n+\tunuse_pack(&window);\n+\tswitch (in_pack_type) {\n+\tdefault:\n+\t\treturn -1; /* we do not do deltas for now */\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_BLOB:\n+\tcase OBJ_TAG:\n+\t\tbreak;\n+\t}\n+\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.close = close_istream_pack_non_delta;\n+\tstream->base.read = read_istream_pack_non_delta;\n+\tstream->base.type = in_pack_type;\n+\tstream->base.size = size;\n+\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n+\tstream->pack = oi.u.packed.pack;\n+\tstream->pos = oi.u.packed.offset;\n+\n+\t*out = &stream->base;\n+\n+\treturn 0;\n+}\ndiff --git a/packfile.h b/packfile.h\nindex 0a98bddd81..3fcc5ae6e0 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -8,6 +8,7 @@\n \n /* in odb.h */\n struct object_info;\n+struct odb_read_stream;\n \n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n@@ -144,6 +145,10 @@ void packfile_store_add_pack(struct packfile_store *store,\n #define repo_for_each_pack(repo, p) \\\n \tfor (p = packfile_store_get_packs(repo->objects->packfiles); p; p = p->next)\n \n+int packfile_store_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t      struct packfile_store *store,\n+\t\t\t\t      const struct object_id *oid);\n+\n /*\n  * Try to read the object identified by its ID from the object store and\n  * populate the object info with its data. Returns 1 in case the object was\ndiff --git a/streaming.c b/streaming.c\nindex 216576857f..02d790f488 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -114,140 +114,6 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \treturn &fs->base;\n }\n \n-/*****************************************************************\n- *\n- * Non-delta packed object stream\n- *\n- *****************************************************************/\n-\n-struct odb_packed_read_stream {\n-\tstruct odb_read_stream base;\n-\tstruct packed_git *pack;\n-\tgit_zstream z;\n-\tenum {\n-\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n-\t\tODB_PACKED_READ_STREAM_INUSE,\n-\t\tODB_PACKED_READ_STREAM_DONE,\n-\t\tODB_PACKED_READ_STREAM_ERROR,\n-\t} z_state;\n-\toff_t pos;\n-};\n-\n-static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n-\t\t\t\t\t   size_t sz)\n-{\n-\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tsize_t total_read = 0;\n-\n-\tswitch (st->z_state) {\n-\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n-\t\tmemset(&st->z, 0, sizeof(st->z));\n-\t\tgit_inflate_init(&st->z);\n-\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n-\t\tbreak;\n-\tcase ODB_PACKED_READ_STREAM_DONE:\n-\t\treturn 0;\n-\tcase ODB_PACKED_READ_STREAM_ERROR:\n-\t\treturn -1;\n-\tcase ODB_PACKED_READ_STREAM_INUSE:\n-\t\tbreak;\n-\t}\n-\n-\twhile (total_read < sz) {\n-\t\tint status;\n-\t\tstruct pack_window *window = NULL;\n-\t\tunsigned char *mapped;\n-\n-\t\tmapped = use_pack(st->pack, &window,\n-\t\t\t\t  st->pos, &st->z.avail_in);\n-\n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tst->z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n-\n-\t\tst->pos += st->z.next_in - mapped;\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n-\t\tunuse_pack(&window);\n-\n-\t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n-\t\t\tbreak;\n-\t\t}\n-\n-\t\t/*\n-\t\t * Unlike the loose object case, we do not have to worry here\n-\t\t * about running out of input bytes and spinning infinitely. If\n-\t\t * we get Z_BUF_ERROR due to too few input bytes, then we'll\n-\t\t * replenish them in the next use_pack() call when we loop. If\n-\t\t * we truly hit the end of the pack (i.e., because it's corrupt\n-\t\t * or truncated), then use_pack() catches that and will die().\n-\t\t */\n-\t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n-\t\t\treturn -1;\n-\t\t}\n-\t}\n-\treturn total_read;\n-}\n-\n-static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n-{\n-\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n-\t\tgit_inflate_end(&st->z);\n-\treturn 0;\n-}\n-\n-static int open_istream_pack_non_delta(struct odb_read_stream **out,\n-\t\t\t\t       struct object_database *odb,\n-\t\t\t\t       const struct object_id *oid)\n-{\n-\tstruct odb_packed_read_stream *stream;\n-\tstruct pack_window *window = NULL;\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tenum object_type in_pack_type;\n-\tunsigned long size;\n-\n-\toi.sizep = &size;\n-\n-\tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n-\t    oi.u.packed.is_delta ||\n-\t    repo_settings_get_big_file_threshold(odb->repo) >= size)\n-\t\treturn -1;\n-\n-\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n-\t\t\t\t\t    &window,\n-\t\t\t\t\t    &oi.u.packed.offset,\n-\t\t\t\t\t    &size);\n-\tunuse_pack(&window);\n-\tswitch (in_pack_type) {\n-\tdefault:\n-\t\treturn -1; /* we do not do deltas for now */\n-\tcase OBJ_COMMIT:\n-\tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n-\tcase OBJ_TAG:\n-\t\tbreak;\n-\t}\n-\n-\tCALLOC_ARRAY(stream, 1);\n-\tstream->base.close = close_istream_pack_non_delta;\n-\tstream->base.read = read_istream_pack_non_delta;\n-\tstream->base.type = in_pack_type;\n-\tstream->base.size = size;\n-\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n-\tstream->pack = oi.u.packed.pack;\n-\tstream->pos = oi.u.packed.offset;\n-\n-\t*out = &stream->base;\n-\n-\treturn 0;\n-}\n-\n-\n /*****************************************************************\n  *\n  * In-core stream\n@@ -318,7 +184,7 @@ static int istream_source(struct odb_read_stream **out,\n {\n \tstruct odb_source *source;\n \n-\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n+\tif (!packfile_store_read_object_stream(out, r->objects->packfiles, oid))\n \t\treturn 0;\n \n \todb_prepare_alternates(r->objects);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530967","messageId":"20251119-b4-pks-odb-read-stream-v1-17-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 17/18] streaming: refactor interface to be object-database-centric","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:17Z","receivedAt":"2025-11-19T07:48:27Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Refactor the streaming interface to be centered around object databases\ninstead of centered around the repository. Rename the functions\naccordingly.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  6 +++---\n archive-zip.c          | 12 ++++++------\n builtin/index-pack.c   |  8 ++++----\n builtin/pack-objects.c | 14 +++++++-------\n object-file.c          |  8 ++++----\n streaming.c            | 44 ++++++++++++++++++++++----------------------\n streaming.h            | 30 +++++++++++++++++++++++++-----\n 7 files changed, 71 insertions(+), 51 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex dc1eda09e0..4133e09ca1 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -135,16 +135,16 @@ static int stream_blocked(struct repository *r, const struct object_id *oid)\n \tchar buf[BLOCKSIZE];\n \tssize_t readlen;\n \n-\tst = open_istream(r, oid, &type, &sz, NULL);\n+\tst = odb_read_object_stream(r->objects, oid, &type, &sz, NULL);\n \tif (!st)\n \t\treturn error(_(\"cannot stream blob %s\"), oid_to_hex(oid));\n \tfor (;;) {\n-\t\treadlen = read_istream(st, buf, sizeof(buf));\n+\t\treadlen = odb_read_stream_read(st, buf, sizeof(buf));\n \t\tif (readlen <= 0)\n \t\t\tbreak;\n \t\tdo_write_blocked(buf, readlen);\n \t}\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \tif (!readlen)\n \t\tfinish_record();\n \treturn readlen;\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 40a9c93ff9..ff57f4f884 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -348,8 +348,8 @@ static int write_zip_entry(struct archiver_args *args,\n \n \t\tif (!buffer) {\n \t\t\tenum object_type type;\n-\t\t\tstream = open_istream(args->repo, oid, &type, &size,\n-\t\t\t\t\t      NULL);\n+\t\t\tstream = odb_read_object_stream(args->repo->objects, oid,\n+\t\t\t\t\t\t\t&type, &size, NULL);\n \t\t\tif (!stream)\n \t\t\t\treturn error(_(\"cannot stream blob %s\"),\n \t\t\t\t\t     oid_to_hex(oid));\n@@ -429,7 +429,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tssize_t readlen;\n \n \t\tfor (;;) {\n-\t\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\t\t\treadlen = odb_read_stream_read(stream, buf, sizeof(buf));\n \t\t\tif (readlen <= 0)\n \t\t\t\tbreak;\n \t\t\tcrc = crc32(crc, buf, readlen);\n@@ -439,7 +439,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\t\t\t\t\t    buf, readlen);\n \t\t\twrite_or_die(1, buf, readlen);\n \t\t}\n-\t\tclose_istream(stream);\n+\t\todb_read_stream_close(stream);\n \t\tif (readlen)\n \t\t\treturn readlen;\n \n@@ -462,7 +462,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tzstream.avail_out = sizeof(compressed);\n \n \t\tfor (;;) {\n-\t\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\t\t\treadlen = odb_read_stream_read(stream, buf, sizeof(buf));\n \t\t\tif (readlen <= 0)\n \t\t\t\tbreak;\n \t\t\tcrc = crc32(crc, buf, readlen);\n@@ -486,7 +486,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\t}\n \n \t\t}\n-\t\tclose_istream(stream);\n+\t\todb_read_stream_close(stream);\n \t\tif (readlen)\n \t\t\treturn readlen;\n \ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 5f90f12f92..67221dbe6a 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -779,7 +779,7 @@ static int compare_objects(const unsigned char *buf, unsigned long size,\n \t}\n \n \twhile (size) {\n-\t\tssize_t len = read_istream(data->st, data->buf, size);\n+\t\tssize_t len = odb_read_stream_read(data->st, data->buf, size);\n \t\tif (len == 0)\n \t\t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t\t    oid_to_hex(&data->entry->idx.oid));\n@@ -807,15 +807,15 @@ static int check_collison(struct object_entry *entry)\n \n \tmemset(&data, 0, sizeof(data));\n \tdata.entry = entry;\n-\tdata.st = open_istream(the_repository, &entry->idx.oid, &type, &size,\n-\t\t\t       NULL);\n+\tdata.st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n+\t\t\t\t\t &type, &size, NULL);\n \tif (!data.st)\n \t\treturn -1;\n \tif (size != entry->size || type != entry->type)\n \t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t    oid_to_hex(&entry->idx.oid));\n \tunpack_data(entry, compare_objects, &data);\n-\tclose_istream(data.st);\n+\todb_read_stream_close(data.st);\n \tfree(data.buf);\n \treturn 0;\n }\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex c693d948e1..adf267c59d 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -417,7 +417,7 @@ static unsigned long write_large_blob_data(struct odb_read_stream *st, struct ha\n \tfor (;;) {\n \t\tssize_t readlen;\n \t\tint zret = Z_OK;\n-\t\treadlen = read_istream(st, ibuf, sizeof(ibuf));\n+\t\treadlen = odb_read_stream_read(st, ibuf, sizeof(ibuf));\n \t\tif (readlen == -1)\n \t\t\tdie(_(\"unable to read %s\"), oid_to_hex(oid));\n \n@@ -520,8 +520,8 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\tif (oe_type(entry) == OBJ_BLOB &&\n \t\t    oe_size_greater_than(&to_pack, entry,\n \t\t\t\t\t repo_settings_get_big_file_threshold(the_repository)) &&\n-\t\t    (st = open_istream(the_repository, &entry->idx.oid, &type,\n-\t\t\t\t       &size, NULL)) != NULL)\n+\t\t    (st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n+\t\t\t\t\t\t &type, &size, NULL)) != NULL)\n \t\t\tbuf = NULL;\n \t\telse {\n \t\t\tbuf = odb_read_object(the_repository->objects,\n@@ -577,7 +577,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t\tdheader[--pos] = 128 | (--ofs & 127);\n \t\tif (limit && hdrlen + sizeof(dheader) - pos + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -591,7 +591,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t */\n \t\tif (limit && hdrlen + hashsz + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -601,7 +601,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t} else {\n \t\tif (limit && hdrlen + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -609,7 +609,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t}\n \tif (st) {\n \t\tdatalen = write_large_blob_data(st, f, &entry->idx.oid);\n-\t\tclose_istream(st);\n+\t\todb_read_stream_close(st);\n \t} else {\n \t\thashwrite(f, buf, datalen);\n \t\tfree(buf);\ndiff --git a/object-file.c b/object-file.c\nindex 8c67847fea..c6d2f2d953 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -139,7 +139,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \n-\tst = open_istream(r, oid, &obj_type, &size, NULL);\n+\tst = odb_read_object_stream(r->objects, oid, &obj_type, &size, NULL);\n \tif (!st)\n \t\treturn -1;\n \n@@ -151,10 +151,10 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tgit_hash_update(&c, hdr, hdrlen);\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\t\tssize_t readlen = odb_read_stream_read(st, buf, sizeof(buf));\n \n \t\tif (readlen < 0) {\n-\t\t\tclose_istream(st);\n+\t\t\todb_read_stream_close(st);\n \t\t\treturn -1;\n \t\t}\n \t\tif (!readlen)\n@@ -162,7 +162,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \t\tgit_hash_update(&c, buf, readlen);\n \t}\n \tgit_hash_final_oid(&real_oid, &c);\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \treturn !oideq(oid, &real_oid) ? -1 : 0;\n }\n \ndiff --git a/streaming.c b/streaming.c\nindex 02d790f488..5e0ff171bf 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -35,7 +35,7 @@ static int close_istream_filtered(struct odb_read_stream *_fs)\n {\n \tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n \tfree_stream_filter(fs->filter);\n-\treturn close_istream(fs->upstream);\n+\treturn odb_read_stream_close(fs->upstream);\n }\n \n static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n@@ -87,7 +87,7 @@ static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n \n \t\t/* refill the input from the upstream */\n \t\tif (!fs->input_finished) {\n-\t\t\tfs->i_end = read_istream(fs->upstream, fs->ibuf, FILTER_BUFFER);\n+\t\t\tfs->i_end = odb_read_stream_read(fs->upstream, fs->ibuf, FILTER_BUFFER);\n \t\t\tif (fs->i_end < 0)\n \t\t\t\treturn -1;\n \t\t\tif (fs->i_end)\n@@ -149,7 +149,7 @@ static ssize_t read_istream_incore(struct odb_read_stream *_st, char *buf, size_\n }\n \n static int open_istream_incore(struct odb_read_stream **out,\n-\t\t\t       struct repository *r,\n+\t\t\t       struct object_database *odb,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n@@ -162,7 +162,7 @@ static int open_istream_incore(struct odb_read_stream **out,\n \toi.typep = &stream.base.type;\n \toi.sizep = &stream.base.size;\n \toi.contentp = (void **)&stream.buf;\n-\tret = odb_read_object_info_extended(r->objects, oid, &oi,\n+\tret = odb_read_object_info_extended(odb, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n@@ -179,49 +179,49 @@ static int open_istream_incore(struct odb_read_stream **out,\n  *****************************************************************************/\n \n static int istream_source(struct odb_read_stream **out,\n-\t\t\t  struct repository *r,\n+\t\t\t  struct object_database *odb,\n \t\t\t  const struct object_id *oid)\n {\n \tstruct odb_source *source;\n \n-\tif (!packfile_store_read_object_stream(out, r->objects->packfiles, oid))\n+\tif (!packfile_store_read_object_stream(out, odb->packfiles, oid))\n \t\treturn 0;\n \n-\todb_prepare_alternates(r->objects);\n-\tfor (source = r->objects->sources; source; source = source->next)\n+\todb_prepare_alternates(odb);\n+\tfor (source = odb->sources; source; source = source->next)\n \t\tif (!odb_source_loose_read_object_stream(out, source, oid))\n \t\t\treturn 0;\n \n-\treturn open_istream_incore(out, r, oid);\n+\treturn open_istream_incore(out, odb, oid);\n }\n \n /****************************************************************\n  * Users of streaming interface\n  ****************************************************************/\n \n-int close_istream(struct odb_read_stream *st)\n+int odb_read_stream_close(struct odb_read_stream *st)\n {\n \tint r = st->close(st);\n \tfree(st);\n \treturn r;\n }\n \n-ssize_t read_istream(struct odb_read_stream *st, void *buf, size_t sz)\n+ssize_t odb_read_stream_read(struct odb_read_stream *st, void *buf, size_t sz)\n {\n \treturn st->read(st, buf, sz);\n }\n \n-struct odb_read_stream *open_istream(struct repository *r,\n-\t\t\t\t     const struct object_id *oid,\n-\t\t\t\t     enum object_type *type,\n-\t\t\t\t     unsigned long *size,\n-\t\t\t\t     struct stream_filter *filter)\n+struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n+\t\t\t\t\t       const struct object_id *oid,\n+\t\t\t\t\t       enum object_type *type,\n+\t\t\t\t\t       unsigned long *size,\n+\t\t\t\t\t       struct stream_filter *filter)\n {\n \tstruct odb_read_stream *st;\n-\tconst struct object_id *real = lookup_replace_object(r, oid);\n+\tconst struct object_id *real = lookup_replace_object(odb->repo, oid);\n \tint ret;\n \n-\tret = istream_source(&st, r, real);\n+\tret = istream_source(&st, odb, real);\n \tif (ret)\n \t\treturn NULL;\n \n@@ -229,7 +229,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n \t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n \t\tif (!nst) {\n-\t\t\tclose_istream(st);\n+\t\t\todb_read_stream_close(st);\n \t\t\treturn NULL;\n \t\t}\n \t\tst = nst;\n@@ -252,7 +252,7 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = open_istream(odb->repo, oid, &type, &sz, filter);\n+\tst = odb_read_object_stream(odb, oid, &type, &sz, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\n@@ -263,7 +263,7 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\n \t\tssize_t wrote, holeto;\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\t\tssize_t readlen = odb_read_stream_read(st, buf, sizeof(buf));\n \n \t\tif (readlen < 0)\n \t\t\tgoto close_and_exit;\n@@ -294,6 +294,6 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tresult = 0;\n \n  close_and_exit:\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \treturn result;\n }\ndiff --git a/streaming.h b/streaming.h\nindex 3a850e3efc..2dce2e359f 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -24,11 +24,31 @@ struct odb_read_stream {\n \tunsigned long size; /* inflated size of full object */\n };\n \n-struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n-\t\t\t\t       enum object_type *, unsigned long *,\n-\t\t\t\t       struct stream_filter *);\n-int close_istream(struct odb_read_stream *);\n-ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n+/*\n+ * Create a new object stream for the given object database. Populates the type\n+ * and size pointers with the object's info. An optional filter can be used to\n+ * transform the object's content.\n+ *\n+ * Returns the stream on success, a `NULL` pointer otherwise.\n+ */\n+struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n+\t\t\t\t\t       const struct object_id *oid,\n+\t\t\t\t\t       enum object_type *type,\n+\t\t\t\t\t       unsigned long *size,\n+\t\t\t\t\t       struct stream_filter *filter);\n+\n+/*\n+ * Close the given read stream and release all resources associated with it.\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+int odb_read_stream_close(struct odb_read_stream *stream);\n+\n+/*\n+ * Read data from the stream into the buffer. Returns 0 on EOF and the number\n+ * of bytes read on success. Returns a negative error code in case reading from\n+ * the stream fails.\n+ */\n+ssize_t odb_read_stream_read(struct odb_read_stream *stream, void *buf, size_t len);\n \n /*\n  * Look up the object by its ID and write the full contents to the file\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530968","messageId":"20251119-b4-pks-odb-read-stream-v1-18-adacf03c2ccf@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH 18/18] streaming: move into object database subsystem","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-19T07:47:18Z","receivedAt":"2025-11-19T07:48:30Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"The \"streaming\" terminology is somewhat generic, so it may not be\nimmediately obvious that \"streaming.{c,h}\" is specific to the object\ndatabase. Rectify this by moving it into the \"odb/\" directory so that it\ncan be immediately attributed to the object subsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Makefile                       | 2 +-\n archive-tar.c                  | 2 +-\n archive-zip.c                  | 2 +-\n builtin/cat-file.c             | 2 +-\n builtin/fsck.c                 | 2 +-\n builtin/index-pack.c           | 2 +-\n builtin/log.c                  | 2 +-\n builtin/pack-objects.c         | 2 +-\n entry.c                        | 2 +-\n meson.build                    | 2 +-\n object-file.c                  | 2 +-\n streaming.c => odb/streaming.c | 2 +-\n streaming.h => odb/streaming.h | 0\n packfile.c                     | 2 +-\n parallel-checkout.c            | 2 +-\n 15 files changed, 14 insertions(+), 14 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 7e0f77e298..6d8dcc4622 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1201,6 +1201,7 @@ LIB_OBJS += object-file.o\n LIB_OBJS += object-name.o\n LIB_OBJS += object.o\n LIB_OBJS += odb.o\n+LIB_OBJS += odb/streaming.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\n LIB_OBJS += oidset.o\n@@ -1294,7 +1295,6 @@ LIB_OBJS += split-index.o\n LIB_OBJS += stable-qsort.o\n LIB_OBJS += statinfo.o\n LIB_OBJS += strbuf.o\n-LIB_OBJS += streaming.o\n LIB_OBJS += string-list.o\n LIB_OBJS += strmap.o\n LIB_OBJS += strvec.o\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 4133e09ca1..74499c311f 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -12,8 +12,8 @@\n #include \"tar.h\"\n #include \"archive.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"strbuf.h\"\n-#include \"streaming.h\"\n #include \"run-command.h\"\n #include \"write-or-die.h\"\n \ndiff --git a/archive-zip.c b/archive-zip.c\nindex ff57f4f884..2b645f28ef 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -10,9 +10,9 @@\n #include \"gettext.h\"\n #include \"git-zlib.h\"\n #include \"hex.h\"\n-#include \"streaming.h\"\n #include \"utf8.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"strbuf.h\"\n #include \"userdiff.h\"\n #include \"write-or-die.h\"\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 120d626d66..505ddaa12f 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -18,13 +18,13 @@\n #include \"list-objects-filter-options.h\"\n #include \"parse-options.h\"\n #include \"userdiff.h\"\n-#include \"streaming.h\"\n #include \"oid-array.h\"\n #include \"packfile.h\"\n #include \"pack-bitmap.h\"\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"promisor-remote.h\"\n #include \"mailmap.h\"\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex 1a348d43c2..c7d2eea287 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -13,11 +13,11 @@\n #include \"fsck.h\"\n #include \"parse-options.h\"\n #include \"progress.h\"\n-#include \"streaming.h\"\n #include \"packfile.h\"\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"path.h\"\n #include \"read-cache-ll.h\"\n #include \"replace-object.h\"\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 67221dbe6a..6403edd3a6 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -16,12 +16,12 @@\n #include \"progress.h\"\n #include \"fsck.h\"\n #include \"strbuf.h\"\n-#include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"packfile.h\"\n #include \"pack-revindex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"oid-array.h\"\n #include \"oidset.h\"\n #include \"path.h\"\ndiff --git a/builtin/log.c b/builtin/log.c\nindex e7b83a6e00..d4cf9c59c8 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -16,6 +16,7 @@\n #include \"refs.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"pager.h\"\n #include \"color.h\"\n #include \"commit.h\"\n@@ -35,7 +36,6 @@\n #include \"parse-options.h\"\n #include \"line-log.h\"\n #include \"branch.h\"\n-#include \"streaming.h\"\n #include \"version.h\"\n #include \"mailmap.h\"\n #include \"progress.h\"\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex adf267c59d..f6c01bc4e0 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -22,7 +22,6 @@\n #include \"pack-objects.h\"\n #include \"progress.h\"\n #include \"refs.h\"\n-#include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"pack-bitmap.h\"\n #include \"delta-islands.h\"\n@@ -33,6 +32,7 @@\n #include \"packfile.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"dir.h\"\n #include \"midx.h\"\ndiff --git a/entry.c b/entry.c\nindex 38dfe670f7..7817aee362 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -2,13 +2,13 @@\n \n #include \"git-compat-util.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"dir.h\"\n #include \"environment.h\"\n #include \"gettext.h\"\n #include \"hex.h\"\n #include \"name-hash.h\"\n #include \"sparse-index.h\"\n-#include \"streaming.h\"\n #include \"submodule.h\"\n #include \"symlinks.h\"\n #include \"progress.h\"\ndiff --git a/meson.build b/meson.build\nindex 1f95a06edb..fc82929b37 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -397,6 +397,7 @@ libgit_sources = [\n   'object-name.c',\n   'object.c',\n   'odb.c',\n+  'odb/streaming.c',\n   'oid-array.c',\n   'oidmap.c',\n   'oidset.c',\n@@ -490,7 +491,6 @@ libgit_sources = [\n   'stable-qsort.c',\n   'statinfo.c',\n   'strbuf.c',\n-  'streaming.c',\n   'string-list.c',\n   'strmap.c',\n   'strvec.c',\ndiff --git a/object-file.c b/object-file.c\nindex c6d2f2d953..4b46cf5b71 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -20,13 +20,13 @@\n #include \"object-file-convert.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"oidtree.h\"\n #include \"pack.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n #include \"read-cache-ll.h\"\n #include \"setup.h\"\n-#include \"streaming.h\"\n #include \"tempfile.h\"\n #include \"tmp-objdir.h\"\n \ndiff --git a/streaming.c b/odb/streaming.c\nsimilarity index 99%\nrename from streaming.c\nrename to odb/streaming.c\nindex 5e0ff171bf..34c9582bcc 100644\n--- a/streaming.c\n+++ b/odb/streaming.c\n@@ -5,10 +5,10 @@\n #include \"git-compat-util.h\"\n #include \"convert.h\"\n #include \"environment.h\"\n-#include \"streaming.h\"\n #include \"repository.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \ndiff --git a/streaming.h b/odb/streaming.h\nsimilarity index 100%\nrename from streaming.h\nrename to odb/streaming.h\ndiff --git a/packfile.c b/packfile.c\nindex ad56ce0b90..7a16aaa90d 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -20,7 +20,7 @@\n #include \"tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n-#include \"streaming.h\"\n+#include \"odb/streaming.h\"\n #include \"midx.h\"\n #include \"commit-graph.h\"\n #include \"pack-revindex.h\"\ndiff --git a/parallel-checkout.c b/parallel-checkout.c\nindex 1cb6701b92..0bf4bd6d4a 100644\n--- a/parallel-checkout.c\n+++ b/parallel-checkout.c\n@@ -13,7 +13,7 @@\n #include \"read-cache-ll.h\"\n #include \"run-command.h\"\n #include \"sigchain.h\"\n-#include \"streaming.h\"\n+#include \"odb/streaming.h\"\n #include \"symlinks.h\"\n #include \"thread-utils.h\"\n #include \"trace2.h\"\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"530985","messageId":"CAOLa=ZRX+_NO-KqiDDtDeLWTKgwMTFDqfcgZjvOechScy+Rv3w@mail.gmail.com","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-2-adacf03c2ccf@pks.im","subject":"Re: [PATCH 02/18] streaming: drop the `open()` callback function","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2025-11-19T09:39:05Z","receivedAt":"2025-11-19T09:39:07Z","isPatch":true,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n\n> diff --git a/streaming.c b/streaming.c\n> index 1fb4b7c1c0..5ce6350123 100644\n> --- a/streaming.c\n> +++ b/streaming.c\n> @@ -14,10 +14,6 @@\n>  #include \"replace-object.h\"\n>  #include \"packfile.h\"\n>\n> -typedef int (*open_istream_fn)(struct odb_read_stream *,\n> -\t\t\t       struct repository *,\n> -\t\t\t       const struct object_id *,\n> -\t\t\t       enum object_type *);\n>  typedef int (*close_istream_fn)(struct odb_read_stream *);\n>  typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n>\n> @@ -34,7 +30,6 @@ struct filtered_istream {\n>  };\n>\n>  struct odb_read_stream {\n> -\topen_istream_fn open;\n>  \tclose_istream_fn close;\n>  \tread_istream_fn read;\n>\n> @@ -437,21 +432,25 @@ static int istream_source(struct odb_read_stream *st,\n>\n>  \tswitch (oi.whence) {\n>  \tcase OI_LOOSE:\n> -\t\tst->open = open_istream_loose;\n> +\t\tif (open_istream_loose(st, r, oid, type) < 0)\n> +\t\t\tbreak;\n\nEarlier we were checking for `if (st->open(st, r, real, type))` so there\nis a slight change in behavior here.\n\nBut both `open_istream_loose()` and `open_istream_pack_non_delta()`\nreturn either -1 or 0. So this is okay.\n\n>  \t\treturn 0;\n>  \tcase OI_PACKED:\n> -\t\tif (!oi.u.packed.is_delta &&\n> -\t\t    repo_settings_get_big_file_threshold(the_repository) < size) {\n> -\t\t\tst->u.in_pack.pack = oi.u.packed.pack;\n> -\t\t\tst->u.in_pack.pos = oi.u.packed.offset;\n> -\t\t\tst->open = open_istream_pack_non_delta;\n> -\t\t\treturn 0;\n> -\t\t}\n> -\t\t/* fallthru */\n> -\tdefault:\n> -\t\tst->open = open_istream_incore;\n> +\t\tif (oi.u.packed.is_delta ||\n> +\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n> +\t\t\tbreak;\n> +\n\nSo we switch the branch flow to break the switch early. Makes sense. The\npatch looks good.\n\n[snip]\n"},{"id":"530986","messageId":"CAOLa=ZTF+xzhZv2yXp8L_URk8cjscycheD=Xgdxd=eRGtvpt2A@mail.gmail.com","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-5-adacf03c2ccf@pks.im","subject":"Re: [PATCH 05/18] streaming: allocate stream inside the backend-specific logic","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2025-11-19T10:11:40Z","receivedAt":"2025-11-19T10:11:43Z","isPatch":true,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When creating a new stream we first allocate it and then call into\n> backend-specific logic to populate the stream. This design requires that\n> the stream itself contains a `union` with backend-specific members that\n> then ultimately get populated by the backend-specific logic.\n>\n> This works, but it's awkward in the context of pluggable object\n> databases. Each backend will need its own member in that union, and as\n> the structure itself is completely opaque (it's only defined in\n> \"streamgin.c\") it also has the consequence that we must have the logic\n\ns/streamgin/streaming\n\n> that is specific to backends in \"streaming.c\".\n>\n> Ideally though, the infrastructure would be reversed: we have a generic\n> `struct odb_read_stream` and some helper functions in \"streaming.c\",\n> whereas the backend-specific logic sits in the backend's subsystem\n> itself.\n>\n\nWill this also mean that we move the backend specific functions like\n`open_istream_loose()` away from 'streaming.c'? Let's read on.\n\n> This can be realized by using a design that is similar to how we handle\n> reference databases: instead of having a union of members, we instead\n> have backend-specific structures with a `struct odb_read_stream base`\n> as its first member. The backends would thus hand out the pointer to the\n> base, but internally they know to cast back to the backend-specific\n> type.\n>\n\nRight.\n\n> This means though that we need to allocate different structures\n> depending on the backend. To prepare for this, move allocation of the\n> structure into the backend-specific functions that open a new stream.\n> Subsequent commits will then create those new backend-specific structs.\n>\n\nWho's in charge of free'ing these structs? I see that `close_istream()`\ncalls the assigned `close()` function. So this could be handled on the\nbackend level. But it also does `free(st)`.\n\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  streaming.c | 99 +++++++++++++++++++++++++++++++++++++++----------------------\n>  1 file changed, 63 insertions(+), 36 deletions(-)\n>\n> diff --git a/streaming.c b/streaming.c\n> index d7db446d25..b8ce82483f 100644\n> --- a/streaming.c\n> +++ b/streaming.c\n> @@ -222,27 +222,34 @@ static int close_istream_loose(struct odb_read_stream *st)\n>  \treturn 0;\n>  }\n>\n> -static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n> +static int open_istream_loose(struct odb_read_stream **out,\n> +\t\t\t      struct repository *r,\n\nWe take in a double pointer now, since the allocation will be handled\ninside the function.\n\n>  \t\t\t      const struct object_id *oid)\n>  {\n>  \tstruct object_info oi = OBJECT_INFO_INIT;\n> +\tstruct odb_read_stream *st;\n>  \tstruct odb_source *source;\n> -\n> -\toi.sizep = &st->size;\n> -\toi.typep = &st->type;\n> +\tunsigned long mapsize;\n> +\tvoid *mapped;\n>\n>  \todb_prepare_alternates(r->objects);\n>  \tfor (source = r->objects->sources; source; source = source->next) {\n> -\t\tst->u.loose.mapped = odb_source_loose_map_object(source, oid,\n> -\t\t\t\t\t\t\t\t &st->u.loose.mapsize);\n> -\t\tif (st->u.loose.mapped)\n> +\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n> +\t\tif (mapped)\n>  \t\t\tbreak;\n>  \t}\n> -\tif (!st->u.loose.mapped)\n> +\tif (!mapped)\n>  \t\treturn -1;\n>\n> -\tswitch (unpack_loose_header(&st->z, st->u.loose.mapped,\n> -\t\t\t\t    st->u.loose.mapsize, st->u.loose.hdr,\n> +\t/*\n> +\t * Note: we must allocate this structure early even though we may still\n> +\t * fail. This is because we need to initialize the zlib stream, and it\n> +\t * is not possible to copy the stream around after the fact because it\n> +\t * has self-referencing pointers.\n> +\t */\n> +\tCALLOC_ARRAY(st, 1);\n> +\n> +\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->u.loose.hdr,\n>  \t\t\t\t    sizeof(st->u.loose.hdr))) {\n>  \tcase ULHR_OK:\n>  \t\tbreak;\n> @@ -250,19 +257,28 @@ static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n>  \tcase ULHR_TOO_LONG:\n>  \t\tgoto error;\n>  \t}\n> +\n> +\toi.sizep = &st->size;\n> +\toi.typep = &st->type;\n> +\n>  \tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n>  \t\tgoto error;\n>\n> +\tst->u.loose.mapped = mapped;\n> +\tst->u.loose.mapsize = mapsize;\n>  \tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n>  \tst->u.loose.hdr_avail = st->z.total_out;\n>  \tst->z_state = z_used;\n>  \tst->close = close_istream_loose;\n>  \tst->read = read_istream_loose;\n>\n> +\t*out = st;\n> +\n>  \treturn 0;\n>  error:\n>  \tgit_inflate_end(&st->z);\n>  \tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n> +\tfree(st);\n>  \treturn -1;\n>  }\n>\n> @@ -338,12 +354,16 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n>  \treturn 0;\n>  }\n>\n> -static int open_istream_pack_non_delta(struct odb_read_stream *st,\n> +static int open_istream_pack_non_delta(struct odb_read_stream **out,\n>  \t\t\t\t       struct repository *r UNUSED,\n>  \t\t\t\t       const struct object_id *oid UNUSED,\n>  \t\t\t\t       struct packed_git *pack,\n>  \t\t\t\t       off_t offset)\n>  {\n> +\tstruct odb_read_stream stream = {\n> +\t\t.close = close_istream_pack_non_delta,\n> +\t\t.read = read_istream_pack_non_delta,\n> +\t};\n\nSo this is now statically defined. Won't this cause an issue?\n\nThe rest looks good. Thanks\n"},{"id":"530987","messageId":"CAOLa=ZRwk2DPCG-kWs-g7qtjBbXc9QuZgumxA3y54JsJjGpM=g@mail.gmail.com","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-6-adacf03c2ccf@pks.im","subject":"Re: [PATCH 06/18] streaming: create structure for in-core object streams","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2025-11-19T10:14:28Z","receivedAt":"2025-11-19T10:14:31Z","isPatch":true,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n\n> @@ -426,22 +429,24 @@ static int open_istream_incore(struct odb_read_stream **out,\n>  \t\t\t       const struct object_id *oid)\n>  {\n>  \tstruct object_info oi = OBJECT_INFO_INIT;\n> -\tstruct odb_read_stream stream = {\n> -\t\t.close = close_istream_incore,\n> -\t\t.read = read_istream_incore,\n> -\t};\n> +\tstruct odb_incore_read_stream stream = {\n> +\t\t.base.close = close_istream_incore,\n> +\t\t.base.read = read_istream_incore,\n> +\t}, *st;\n\nNit: Almost missed this `*st`. I wonder if its more readable as a\nseparate line:\n\n  struct odb_incore_read_stream *st;\n\nAll good otherwise.\n\n>  \tint ret;\n>\n> -\toi.typep = &stream.type;\n> -\toi.sizep = &stream.size;\n> -\toi.contentp = (void **)&stream.u.incore.buf;\n> +\toi.typep = &stream.base.type;\n> +\toi.sizep = &stream.base.size;\n> +\toi.contentp = (void **)&stream.buf;\n>  \tret = odb_read_object_info_extended(r->objects, oid, &oi,\n>  \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n>  \tif (ret)\n>  \t\treturn ret;\n>\n> -\tCALLOC_ARRAY(*out, 1);\n> -\t**out = stream;\n> +\tCALLOC_ARRAY(st, 1);\n> +\t*st = stream;\n> +\t*out = &st->base;\n> +\n>  \treturn 0;\n>  }\n>\n>\n> --\n> 2.52.0.rc2.482.gaa765fefd0.dirty\n"},{"id":"530992","messageId":"CAOLa=ZQDqGLh3hrV6T32mdrb1Z-nrVh-zkgjgfoHJrmrTRSWFQ@mail.gmail.com","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-11-adacf03c2ccf@pks.im","subject":"Re: [PATCH 11/18] packfile: introduce function to read object info from a store","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2025-11-19T14:48:24Z","receivedAt":"2025-11-19T14:48:27Z","isPatch":true,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Extract the logic to read object info for a packed object from\n> `do_oid_object_into_extended()` into a standalone function that operates\n> on the packfile store. This function will be used in a subsequent\n> commit.\n>\n> Note that this change allows us to make `find_pack_entry()` an internal\n> implementation detail. As a consequence though we have to move around\n> `packfile_store_freshen_object()` so that it is defined after that\n> function.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb.c      | 29 ++++---------------------\n>  packfile.c | 71 +++++++++++++++++++++++++++++++++++++++++++++++---------------\n>  packfile.h | 12 ++++++++++-\n>  3 files changed, 69 insertions(+), 43 deletions(-)\n>\n> diff --git a/odb.c b/odb.c\n> index 3ec21ef24e..f4cbee4b04 100644\n> --- a/odb.c\n> +++ b/odb.c\n> @@ -666,8 +666,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n>  {\n>  \tstatic struct object_info blank_oi = OBJECT_INFO_INIT;\n>  \tconst struct cached_object *co;\n> -\tstruct pack_entry e;\n> -\tint rtype;\n>  \tconst struct object_id *real = oid;\n>  \tint already_retried = 0;\n>\n> @@ -702,8 +700,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n>  \twhile (1) {\n>  \t\tstruct odb_source *source;\n>\n> -\t\tif (find_pack_entry(odb->repo, real, &e))\n> -\t\t\tbreak;\n> +\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n> +\t\t\treturn 0;\n>\n\nEarlier we would try to find the pack entry and if we did, we would\nbreak this `while` loop and fill in the object information. Now that is\npart of the `packfile_store_read_object_info()` function. So we simply\nhave to loop until it returns a success.\n\nSpeaking of which, the loop simply exists to capture:\n1. Trying to read objects from a submodule, so we add the submodule\nsources and try everything again\n2. If its a promisor remote, we try to fetch and try everything again.\n\n[snip]\n\nThe rest looks good.\n"},{"id":"530994","messageId":"CAOLa=ZRwnsYeHDpdL+uvnw0YMTbG1Gx2SKsq+0hTWMto+QZ+Lg@mail.gmail.com","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-12-adacf03c2ccf@pks.im","subject":"Re: [PATCH 12/18] streaming: rely on object sources to create object stream","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2025-11-19T16:10:41Z","receivedAt":"2025-11-19T16:10:44Z","isPatch":true,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When creating an object stream we first look up the object info and, if\n> it's present, we call into the respective backend that contains the\n> object to create a new stream for it.\n>\n> This has the consequence that, for loose object source, we basically\n> iterate through the object sources twice: we first discover that the\n> file exists as a loose object in the first place by iterating through\n> all sources. And, once we have discovered it, we again walk through all\n> sources to try and map the object. The same issue will eventually also\n> surface once the packfile store becomes per-object-source.\n>\n> Furthermore, it feels rather pointless to first look up the object only\n> to then try and read it.\n>\n> Refactor the logic to be centered around sources instead. Instead of\n> first reading the object, we immediately ask the source to create the\n> object stream for us. If the object exists we get stream, otherwise\n> we'll try the next source.\n>\n> Like this we only have to iterate through sources once. But even more\n> importantly, this change also helps us to make the whole logic\n> pluggable. The object read stream subsystem does not need to be aware of\n> the different source backends anymore, but eventually it'll only have to\n> call the source's callback function.\n>\n> Note that at the current poin in time we aren't full there yet:\n>\n\ns/poin/point\ns/full/fully\n\n>   - The packfile store still sits on the object database level and is\n>     thus agnostic of the sources.\n>\n>   - We still have to call into both the packfile store and the loose\n>     object source.\n>\n> But both of these issues will soon be addressed.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  streaming.c | 65 +++++++++++++++++++++++--------------------------------------\n>  1 file changed, 24 insertions(+), 41 deletions(-)\n>\n> diff --git a/streaming.c b/streaming.c\n> index 572be98248..bebb434cd1 100644\n> --- a/streaming.c\n> +++ b/streaming.c\n> @@ -204,21 +204,15 @@ static int close_istream_loose(struct odb_read_stream *_st)\n>  }\n>\n>  static int open_istream_loose(struct odb_read_stream **out,\n> -\t\t\t      struct repository *r,\n> +\t\t\t      struct odb_source *source,\n>  \t\t\t      const struct object_id *oid)\n>  {\n>  \tstruct object_info oi = OBJECT_INFO_INIT;\n>  \tstruct odb_loose_read_stream *st;\n> -\tstruct odb_source *source;\n>  \tunsigned long mapsize;\n>  \tvoid *mapped;\n>\n> -\todb_prepare_alternates(r->objects);\n> -\tfor (source = r->objects->sources; source; source = source->next) {\n> -\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n> -\t\tif (mapped)\n> -\t\t\tbreak;\n> -\t}\n> +\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n>  \tif (!mapped)\n>  \t\treturn -1;\n>\n\nSo instead of going over the sources, we simply check for the given\nsource. Nice.\n\n[snip]\n\n> @@ -462,30 +460,15 @@ static int istream_source(struct odb_read_stream **out,\n>  \t\t\t  struct repository *r,\n>  \t\t\t  const struct object_id *oid)\n>  {\n> -\tunsigned long size;\n> -\tint status;\n> -\tstruct object_info oi = OBJECT_INFO_INIT;\n> -\n> -\toi.sizep = &size;\n> -\tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n> -\tif (status < 0)\n> -\t\treturn status;\n> +\tstruct odb_source *source;\n>\n> -\tswitch (oi.whence) {\n> -\tcase OI_LOOSE:\n> -\t\tif (open_istream_loose(out, r, oid) < 0)\n> -\t\t\tbreak;\n> -\t\treturn 0;\n> -\tcase OI_PACKED:\n> -\t\tif (oi.u.packed.is_delta ||\n> -\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n> -\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n> -\t\t\t\t\t\toi.u.packed.offset) < 0)\n> -\t\t\tbreak;\n> +\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n>  \t\treturn 0;\n> -\tdefault:\n> -\t\tbreak;\n> -\t}\n> +\n> +\todb_prepare_alternates(r->objects);\n> +\tfor (source = r->objects->sources; source; source = source->next)\n> +\t\tif (!open_istream_loose(out, source, oid))\n> +\t\t\treturn 0;\n>\n\nThis seem to be the crux of it, where earlier we depended on\n`odb_read_object_info_extended()` to tell us which backend to rely on\nand then we re-fetched from that backed, now we simply go over the\ndifferent sources and try to get the object stream. Makes sense.\n\n>  \treturn open_istream_incore(out, r, oid);\n>  }\n>\n> --\n> 2.52.0.rc2.482.gaa765fefd0.dirty\n"},{"id":"530996","messageId":"CAOLa=ZT_VFfbfLVdvHUqK5C6k4zROLQs0Pt5rOWL_hE_BSfGeg@mail.gmail.com","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-14-adacf03c2ccf@pks.im","subject":"Re: [PATCH 14/18] streaming: make the `odb_read_stream` definition public","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2025-11-19T16:27:29Z","receivedAt":"2025-11-19T16:27:32Z","isPatch":true,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Subsequent commits will move the backend-specific logic of setting up an\n> object read stream into the specific subsystems. As the backends are now\n> the ones that are responsible for allocating the stream they'll need to\n> have the stream definition available to them.\n>\n\nThis was a question I had in mind in one of the previous patches, looks\nlike we're going in that direction. Makes sense to me.\n\n> Make the stream definition public to prepare for this.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  streaming.c | 11 -----------\n>  streaming.h | 15 ++++++++++++++-\n>  2 files changed, 14 insertions(+), 12 deletions(-)\n>\n> diff --git a/streaming.c b/streaming.c\n> index 9e20e9a882..3f94bd2a03 100644\n> --- a/streaming.c\n> +++ b/streaming.c\n> @@ -12,19 +12,8 @@\n>  #include \"replace-object.h\"\n>  #include \"packfile.h\"\n>\n> -typedef int (*close_istream_fn)(struct odb_read_stream *);\n> -typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n> -\n>  #define FILTER_BUFFER (1024*16)\n>\n> -struct odb_read_stream {\n> -\tclose_istream_fn close;\n> -\tread_istream_fn read;\n> -\n> -\tenum object_type type;\n> -\tunsigned long size; /* inflated size of full object */\n> -};\n> -\n>  /*****************************************************************\n>   *\n>   * Filtered stream\n> diff --git a/streaming.h b/streaming.h\n> index 95c2a434fa..3a850e3efc 100644\n> --- a/streaming.h\n> +++ b/streaming.h\n> @@ -6,11 +6,24 @@\n>\n>  #include \"object.h\"\n>\n> -/* opaque */\n>  struct object_database;\n>  struct odb_read_stream;\n>  struct stream_filter;\n>\n> +typedef int (*odb_read_stream_close_fn)(struct odb_read_stream *);\n> +typedef ssize_t (*odb_read_stream_read_fn)(struct odb_read_stream *, char *, size_t);\n> +\n> +/*\n> + * A stream that can be used to read an object from the object database without\n> + * loading all of it into memory.\n> + */\n> +struct odb_read_stream {\n> +\todb_read_stream_close_fn close;\n> +\todb_read_stream_read_fn read;\n> +\tenum object_type type;\n> +\tunsigned long size; /* inflated size of full object */\n> +};\n> +\n>  struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n>  \t\t\t\t       enum object_type *, unsigned long *,\n>  \t\t\t\t       struct stream_filter *);\n>\n\nIf we're returning an `struct odb_read_stream` anyways, why take in\npointers for object size and object type? They'll be the same as\n`odb_read_stream.type` and `odb_read_stream.size` no?\n"},{"id":"531002","messageId":"2nd7qcj7jrrwc4fyhfsovs3ptrwmrdxxcap4sqadujtwwua5ha@bpbjlbnbcpiw","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-1-adacf03c2ccf@pks.im","subject":"Re: [PATCH 01/18] streaming: rename `git_istream` into `odb_read_stream`","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-11-19T18:49:22Z","receivedAt":"2025-11-19T18:49:31Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/11/19 08:47AM, Patrick Steinhardt wrote:\n> In the following patches we are about to make the `git_istream` more\n> generic so that it becomes fully controlled by the specific object\n> source that wants to create it. As part of these refactorings we'll\n> fully move the structure into the object database subsystem.\n\nOk, so looking at the current implementation of `git_istream`, it does\nappear to be already defined in a somewhat generic manner as it supports\nreading loose/packed objects. What sources are supported are all\ncentrally defined in \"streaming.c\" though. It sounds like we eventually\nwant each source to fully control this interface without having to go\nthrough \"streaming.c\" to setup each source stream type which makes\nsense.\n\n> Prepare for this change by renaming the structure from `git_istream`\n> to `odb_read_stream`. This mirrors the `odb_write_stream` structure that\n> we already have.\n> \n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n> diff --git a/streaming.h b/streaming.h\n> index bd27f59e57..acf4c84338 100644\n> --- a/streaming.h\n> +++ b/streaming.h\n> @@ -7,14 +7,14 @@\n>  #include \"object.h\"\n>  \n>  /* opaque */\n> -struct git_istream;\n> +struct odb_read_stream;\n\nThe name change here makes sense. While we are here, it might be nice to\nleave a comment annotating it's purpose in a bit more detail.\n\n-Justin\n"},{"id":"531003","messageId":"g74hupkwedtclb3gxomhxj6w4rqqzn3tsostdriauvn3gu2cw2@wxgwulitxbtq","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-2-adacf03c2ccf@pks.im","subject":"Re: [PATCH 02/18] streaming: drop the `open()` callback function","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-11-19T19:01:03Z","receivedAt":"2025-11-19T19:01:09Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/11/19 08:47AM, Patrick Steinhardt wrote:\n> When creating a read stream we first populate the structure with the\n> open callback function and then subsequently call the function. This\n> layout is somewhat weird though:\n> \n>   - The structure needs to be allocated and partially populated with the\n>     open function before we can properly initialize it.\n> \n>   - We never use the `open()` callback after having opened it initially.\n> \n> Especially the first point creates a problem for us. In subsequent\n> commits we'll want to fully move construction of the read source into\n> the respective object sources. E.g., the loose object source will be the\n> one that is responsible for creating the structure. But this creates a\n> problem: if we first need to create the structure so that we can call\n> the source-specific callback we cannot fully handle creation of the\n> structure in the source itself.\n> \n> We could of course work around that and have the loose object source\n> create the structure and populate it's `open()` callback, only. But\n\ns/it's/its/\n\n> this doesn't really buy us anything due to the second bullet point\n> above.\n> \n> Instead, drop the callback entirely and refactor `istream_source()` so\n> that we open the streams immediately. This unblocks a subsequent step,\n> where we'll also start to allocate the structure in the source-specific\n> logic.\n\nOut of curiousity, is there any reason we would ever want to delay\nopening the source read stream? If not, then I agree it makes more sense\nto just open the stream at time of its initialization.\n\n> \n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  streaming.c | 40 +++++++++++++++++-----------------------\n>  1 file changed, 17 insertions(+), 23 deletions(-)\n> \n> diff --git a/streaming.c b/streaming.c\n> index 1fb4b7c1c0..5ce6350123 100644\n> --- a/streaming.c\n> +++ b/streaming.c\n> @@ -14,10 +14,6 @@\n>  #include \"replace-object.h\"\n>  #include \"packfile.h\"\n>  \n> -typedef int (*open_istream_fn)(struct odb_read_stream *,\n> -\t\t\t       struct repository *,\n> -\t\t\t       const struct object_id *,\n> -\t\t\t       enum object_type *);\n>  typedef int (*close_istream_fn)(struct odb_read_stream *);\n>  typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n>  \n> @@ -34,7 +30,6 @@ struct filtered_istream {\n>  };\n>  \n>  struct odb_read_stream {\n> -\topen_istream_fn open;\n>  \tclose_istream_fn close;\n>  \tread_istream_fn read;\n>  \n> @@ -437,21 +432,25 @@ static int istream_source(struct odb_read_stream *st,\n>  \n>  \tswitch (oi.whence) {\n>  \tcase OI_LOOSE:\n> -\t\tst->open = open_istream_loose;\n> +\t\tif (open_istream_loose(st, r, oid, type) < 0)\n> +\t\t\tbreak;\n\nPreviously, if an error happened when executing the callback,\n`open_istream_incore()` would be invoked as a fallback. Now we handle\nthat here during initialization by breaking early. This preserves the\noriginal behavior. Makes sense. \n\n>  \t\treturn 0;\n>  \tcase OI_PACKED:\n> -\t\tif (!oi.u.packed.is_delta &&\n> -\t\t    repo_settings_get_big_file_threshold(the_repository) < size) {\n> -\t\t\tst->u.in_pack.pack = oi.u.packed.pack;\n> -\t\t\tst->u.in_pack.pos = oi.u.packed.offset;\n> -\t\t\tst->open = open_istream_pack_non_delta;\n> -\t\t\treturn 0;\n> -\t\t}\n> -\t\t/* fallthru */\n> -\tdefault:\n> -\t\tst->open = open_istream_incore;\n> +\t\tif (oi.u.packed.is_delta ||\n> +\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n> +\t\t\tbreak;\n> +\n> +\t\tst->u.in_pack.pack = oi.u.packed.pack;\n> +\t\tst->u.in_pack.pos = oi.u.packed.offset;\n> +\t\tif (open_istream_pack_non_delta(st, r, oid, type) < 0)\n> +\t\t\tbreak;\n> +\n>  \t\treturn 0;\n> +\tdefault:\n> +\t\tbreak;\n>  \t}\n> +\n> +\treturn open_istream_incore(st, r, oid, type);\n>  }\n>  \n>  /****************************************************************\n> @@ -478,19 +477,14 @@ struct odb_read_stream *open_istream(struct repository *r,\n>  {\n>  \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n>  \tconst struct object_id *real = lookup_replace_object(r, oid);\n> -\tint ret = istream_source(st, r, real, type);\n> +\tint ret;\n>  \n> +\tret = istream_source(st, r, real, type);\n>  \tif (ret) {\n>  \t\tfree(st);\n>  \t\treturn NULL;\n>  \t}\n>  \n> -\tif (st->open(st, r, real, type)) {\n> -\t\tif (open_istream_incore(st, r, real, type)) {\n> -\t\t\tfree(st);\n> -\t\t\treturn NULL;\n> -\t\t}\n> -\t}\n\nNow that opening the read stream in handled during initialization, we\ncan drop the explicit call to the open callback.\n\n-Justin\n"},{"id":"531004","messageId":"cuvoz5gl7d6xgj757jgb26kj3qeunc4w3pg72it53zi6rs5lka@2nc5x4b2e3eg","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-3-adacf03c2ccf@pks.im","subject":"Re: [PATCH 03/18] streaming: propagate final object type via the stream","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-11-19T19:25:29Z","receivedAt":"2025-11-19T19:25:39Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/11/19 08:47AM, Patrick Steinhardt wrote:\n> When opening the read stream for a specific object the caller is also\n> expected to pass in a pointer to the object type. This type is passed\n> down via multiple levels and will eventually be populated with the type\n> of the looked-up object.\n> \n> The way we propagate down the pointer though is somewhat non-obvious.\n> While `istream_source()` still expects the pointer and looks it up via\n> `odb_read_object_info_extended()`, we also pass it down even further\n> into the format-specific callbacks that perform another lookup. This is\n> quite confusing overall.\n> \n> Refactor the code so that the responsibility to populate the object type\n> rests solely with the format-specific callbacks. This will allow us to\n> drop the call to `odb_read_object_info_extended()` in `istream_source()`\n> entirely in a subsequent patch.\n> \n> Furthermore, instead of propagating the type via an in-pointer, we now\n> propagate the type via a new field in the object stream. It already has\n> a `size` field, so it's only natural to have a second field that\n> contains the object type.\n> \n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  streaming.c | 30 +++++++++++++++---------------\n>  1 file changed, 15 insertions(+), 15 deletions(-)\n> \n> diff --git a/streaming.c b/streaming.c\n> index 5ce6350123..9596a94c58 100644\n> --- a/streaming.c\n> +++ b/streaming.c\n> @@ -33,6 +33,7 @@ struct odb_read_stream {\n>  \tclose_istream_fn close;\n>  \tread_istream_fn read;\n>  \n> +\tenum object_type type;\n\nNow we are storing the object type in the stream. This avoids having to\npass the object type pointer around as much explictly. I think this is a\nnice change.\n\n>  \tunsigned long size; /* inflated size of full object */\n>  \tgit_zstream z;\n>  \tenum { z_unused, z_used, z_done, z_error } z_state;\n> @@ -159,6 +160,7 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n>  \tfs->o_end = fs->o_ptr = 0;\n>  \tfs->input_finished = 0;\n>  \tifs->size = -1; /* unknown */\n> +\tifs->type = st->type;\n>  \treturn ifs;\n>  }\n>  \n[snip]\n> @@ -496,6 +495,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n>  \t}\n>  \n>  \t*size = st->size;\n> +\t*type = st->type;\n\nSo even though `open_istream()` returns `odb_read_stream` which contains\nthe object type, this function still accepts an object type pointer. At\nfirst I thought this was a bit strange, but `odb_read_stream` is an\nopaque structure so this make sense and is also what we do for object\nsize.\n\n-Justin\n"},{"id":"531005","messageId":"xmqqikf5bxg8.fsf@gitster.g","threadId":"64509","inReplyTo":"2nd7qcj7jrrwc4fyhfsovs3ptrwmrdxxcap4sqadujtwwua5ha@bpbjlbnbcpiw","subject":"Re: [PATCH 01/18] streaming: rename `git_istream` into `odb_read_stream`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-19T20:04:55Z","receivedAt":"2025-11-19T20:04:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> On 25/11/19 08:47AM, Patrick Steinhardt wrote:\n>> In the following patches we are about to make the `git_istream` more\n>> generic so that it becomes fully controlled by the specific object\n>> source that wants to create it. As part of these refactorings we'll\n>> fully move the structure into the object database subsystem.\n>\n> Ok, so looking at the current implementation of `git_istream`, it does\n> appear to be already defined in a somewhat generic manner as it supports\n> reading loose/packed objects. What sources are supported are all\n> centrally defined in \"streaming.c\" though. It sounds like we eventually\n> want each source to fully control this interface without having to go\n> through \"streaming.c\" to setup each source stream type which makes\n> sense.\n\nAs the original inventor of git_istream abstraction, I fully agree\nwith this direction.\n\nThanks for cleaning up, and thanks for reviewing.\n\n>\n>> Prepare for this change by renaming the structure from `git_istream`\n>> to `odb_read_stream`. This mirrors the `odb_write_stream` structure that\n>> we already have.\n>> \n>> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n>> ---\n>> diff --git a/streaming.h b/streaming.h\n>> index bd27f59e57..acf4c84338 100644\n>> --- a/streaming.h\n>> +++ b/streaming.h\n>> @@ -7,14 +7,14 @@\n>>  #include \"object.h\"\n>>  \n>>  /* opaque */\n>> -struct git_istream;\n>> +struct odb_read_stream;\n>\n> The name change here makes sense. While we are here, it might be nice to\n> leave a comment annotating it's purpose in a bit more detail.\n>\n> -Justin\n"},{"id":"531100","messageId":"aSAHVZCR7U1Di-LP@pks.im","threadId":"64509","inReplyTo":"2nd7qcj7jrrwc4fyhfsovs3ptrwmrdxxcap4sqadujtwwua5ha@bpbjlbnbcpiw","subject":"Re: [PATCH 01/18] streaming: rename `git_istream` into `odb_read_stream`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T06:31:49Z","receivedAt":"2025-11-21T06:31:55Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Nov 19, 2025 at 12:49:22PM -0600, Justin Tobler wrote:\n> On 25/11/19 08:47AM, Patrick Steinhardt wrote:\n> > diff --git a/streaming.h b/streaming.h\n> > index bd27f59e57..acf4c84338 100644\n> > --- a/streaming.h\n> > +++ b/streaming.h\n> > @@ -7,14 +7,14 @@\n> >  #include \"object.h\"\n> >  \n> >  /* opaque */\n> > -struct git_istream;\n> > +struct odb_read_stream;\n> \n> The name change here makes sense. While we are here, it might be nice to\n> leave a comment annotating it's purpose in a bit more detail.\n\nI do this in a subsequent commit, so I won't add this comment here.\n\nPatrick\n"},{"id":"531101","messageId":"aSAHYNBCMwYsFMYM@pks.im","threadId":"64509","inReplyTo":"g74hupkwedtclb3gxomhxj6w4rqqzn3tsostdriauvn3gu2cw2@wxgwulitxbtq","subject":"Re: [PATCH 02/18] streaming: drop the `open()` callback function","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T06:32:00Z","receivedAt":"2025-11-21T06:32:05Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Nov 19, 2025 at 01:01:03PM -0600, Justin Tobler wrote:\n> On 25/11/19 08:47AM, Patrick Steinhardt wrote:\n> > Instead, drop the callback entirely and refactor `istream_source()` so\n> > that we open the streams immediately. This unblocks a subsequent step,\n> > where we'll also start to allocate the structure in the source-specific\n> > logic.\n> \n> Out of curiousity, is there any reason we would ever want to delay\n> opening the source read stream? If not, then I agree it makes more sense\n> to just open the stream at time of its initialization.\n\nI could not find any reason -- it's not used anywhere in our tree, and I\ncouldn't think about why one would want this, either.\n\nPatrick\n"},{"id":"531102","messageId":"aSAHhUGzG-c2o98d@pks.im","threadId":"64509","inReplyTo":"cuvoz5gl7d6xgj757jgb26kj3qeunc4w3pg72it53zi6rs5lka@2nc5x4b2e3eg","subject":"Re: [PATCH 03/18] streaming: propagate final object type via the stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T06:32:37Z","receivedAt":"2025-11-21T06:32:44Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Nov 19, 2025 at 01:25:29PM -0600, Justin Tobler wrote:\n> On 25/11/19 08:47AM, Patrick Steinhardt wrote:\n> > diff --git a/streaming.c b/streaming.c\n> > index 5ce6350123..9596a94c58 100644\n> > --- a/streaming.c\n> > +++ b/streaming.c\n> > @@ -496,6 +495,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n> >  \t}\n> >  \n> >  \t*size = st->size;\n> > +\t*type = st->type;\n> \n> So even though `open_istream()` returns `odb_read_stream` which contains\n> the object type, this function still accepts an object type pointer. At\n> first I thought this was a bit strange, but `odb_read_stream` is an\n> opaque structure so this make sense and is also what we do for object\n> size.\n\nYeah. I was a bit torn here to be honest, but ultimately decided against\ndropping the type pointer. At the end of this series we _can_ do this in\ntheory as the `struct odb_read_stream` becomes public.\n\nI'll add another patch to do this conversion at the end of this series.\n\nPatrick\n"},{"id":"531103","messageId":"aSAHjQKO7R1TvPgj@pks.im","threadId":"64509","inReplyTo":"CAOLa=ZTF+xzhZv2yXp8L_URk8cjscycheD=Xgdxd=eRGtvpt2A@mail.gmail.com","subject":"Re: [PATCH 05/18] streaming: allocate stream inside the backend-specific logic","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T06:32:45Z","receivedAt":"2025-11-21T06:32:51Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Nov 19, 2025 at 02:11:40AM -0800, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > that is specific to backends in \"streaming.c\".\n> >\n> > Ideally though, the infrastructure would be reversed: we have a generic\n> > `struct odb_read_stream` and some helper functions in \"streaming.c\",\n> > whereas the backend-specific logic sits in the backend's subsystem\n> > itself.\n> >\n> \n> Will this also mean that we move the backend specific functions like\n> `open_istream_loose()` away from 'streaming.c'? Let's read on.\n\nYup, exactly.\n\n> > This can be realized by using a design that is similar to how we handle\n> > reference databases: instead of having a union of members, we instead\n> > have backend-specific structures with a `struct odb_read_stream base`\n> > as its first member. The backends would thus hand out the pointer to the\n> > base, but internally they know to cast back to the backend-specific\n> > type.\n> >\n> \n> Right.\n> \n> > This means though that we need to allocate different structures\n> > depending on the backend. To prepare for this, move allocation of the\n> > structure into the backend-specific functions that open a new stream.\n> > Subsequent commits will then create those new backend-specific structs.\n> >\n> \n> Who's in charge of free'ing these structs? I see that `close_istream()`\n> calls the assigned `close()` function. So this could be handled on the\n> backend level. But it also does `free(st)`.\n\nYeah, this'll be changed later: the `close()` callback will then only\nclose and release the backend-specific data. `odb_read_stream_close()`\nis then responsible for freeing the stream itself.\n\n> > @@ -338,12 +354,16 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n> >  \treturn 0;\n> >  }\n> >\n> > -static int open_istream_pack_non_delta(struct odb_read_stream *st,\n> > +static int open_istream_pack_non_delta(struct odb_read_stream **out,\n> >  \t\t\t\t       struct repository *r UNUSED,\n> >  \t\t\t\t       const struct object_id *oid UNUSED,\n> >  \t\t\t\t       struct packed_git *pack,\n> >  \t\t\t\t       off_t offset)\n> >  {\n> > +\tstruct odb_read_stream stream = {\n> > +\t\t.close = close_istream_pack_non_delta,\n> > +\t\t.read = read_istream_pack_non_delta,\n> > +\t};\n> \n> So this is now statically defined. Won't this cause an issue?\n\nNo, it doesn't, as we eventually copy the local stream weh ave here into\nthe allocated `out` pointer.\n\nPatrick\n"},{"id":"531104","messageId":"aSAHlQtQuupprYw9@pks.im","threadId":"64509","inReplyTo":"CAOLa=ZRwk2DPCG-kWs-g7qtjBbXc9QuZgumxA3y54JsJjGpM=g@mail.gmail.com","subject":"Re: [PATCH 06/18] streaming: create structure for in-core object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T06:32:53Z","receivedAt":"2025-11-21T06:32:58Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Nov 19, 2025 at 10:14:28AM +0000, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > @@ -426,22 +429,24 @@ static int open_istream_incore(struct odb_read_stream **out,\n> >  \t\t\t       const struct object_id *oid)\n> >  {\n> >  \tstruct object_info oi = OBJECT_INFO_INIT;\n> > -\tstruct odb_read_stream stream = {\n> > -\t\t.close = close_istream_incore,\n> > -\t\t.read = read_istream_incore,\n> > -\t};\n> > +\tstruct odb_incore_read_stream stream = {\n> > +\t\t.base.close = close_istream_incore,\n> > +\t\t.base.read = read_istream_incore,\n> > +\t}, *st;\n> \n> Nit: Almost missed this `*st`. I wonder if its more readable as a\n> separate line:\n> \n>   struct odb_incore_read_stream *st;\n> \n> All good otherwise.\n\nFair, will adapt.\n\nPatrick\n"},{"id":"531105","messageId":"aSAHnduhZUk7gC-K@pks.im","threadId":"64509","inReplyTo":"CAOLa=ZQDqGLh3hrV6T32mdrb1Z-nrVh-zkgjgfoHJrmrTRSWFQ@mail.gmail.com","subject":"Re: [PATCH 11/18] packfile: introduce function to read object info from a store","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T06:33:01Z","receivedAt":"2025-11-21T06:33:06Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Nov 19, 2025 at 02:48:24PM +0000, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/odb.c b/odb.c\n> > index 3ec21ef24e..f4cbee4b04 100644\n> > --- a/odb.c\n> > +++ b/odb.c\n> > @@ -702,8 +700,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n> >  \twhile (1) {\n> >  \t\tstruct odb_source *source;\n> >\n> > -\t\tif (find_pack_entry(odb->repo, real, &e))\n> > -\t\t\tbreak;\n> > +\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n> > +\t\t\treturn 0;\n> >\n> \n> Earlier we would try to find the pack entry and if we did, we would\n> break this `while` loop and fill in the object information. Now that is\n> part of the `packfile_store_read_object_info()` function. So we simply\n> have to loop until it returns a success.\n> \n> Speaking of which, the loop simply exists to capture:\n> 1. Trying to read objects from a submodule, so we add the submodule\n> sources and try everything again\n> 2. If its a promisor remote, we try to fetch and try everything again.\n\nExactly. The loop will be changed somewhat to also handle the ODB\nsources. But that will be part of a later patch series that moves the\npackfile store into the ODB source.\n\nPatrick\n"},{"id":"531106","messageId":"aSAHq9_Wa2eXko1R@pks.im","threadId":"64509","inReplyTo":"CAOLa=ZT_VFfbfLVdvHUqK5C6k4zROLQs0Pt5rOWL_hE_BSfGeg@mail.gmail.com","subject":"Re: [PATCH 14/18] streaming: make the `odb_read_stream` definition public","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T06:33:15Z","receivedAt":"2025-11-21T06:33:20Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Nov 19, 2025 at 11:27:29AM -0500, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/streaming.h b/streaming.h\n> > index 95c2a434fa..3a850e3efc 100644\n> > --- a/streaming.h\n> > +++ b/streaming.h\n> > @@ -6,11 +6,24 @@\n> >\n> >  #include \"object.h\"\n> >\n> > -/* opaque */\n> >  struct object_database;\n> >  struct odb_read_stream;\n> >  struct stream_filter;\n> >\n> > +typedef int (*odb_read_stream_close_fn)(struct odb_read_stream *);\n> > +typedef ssize_t (*odb_read_stream_read_fn)(struct odb_read_stream *, char *, size_t);\n> > +\n> > +/*\n> > + * A stream that can be used to read an object from the object database without\n> > + * loading all of it into memory.\n> > + */\n> > +struct odb_read_stream {\n> > +\todb_read_stream_close_fn close;\n> > +\todb_read_stream_read_fn read;\n> > +\tenum object_type type;\n> > +\tunsigned long size; /* inflated size of full object */\n> > +};\n> > +\n> >  struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n> >  \t\t\t\t       enum object_type *, unsigned long *,\n> >  \t\t\t\t       struct stream_filter *);\n> >\n> \n> If we're returning an `struct odb_read_stream` anyways, why take in\n> pointers for object size and object type? They'll be the same as\n> `odb_read_stream.type` and `odb_read_stream.size` no?\n\nYeah, they are now, so we could change it. But I wasn't really sure\nwhether this is all that useful in the first place, and didn't quite\nfeel like doing another tree-wide change.\n\nBut I did the change now, and I think it's a net improvement. So let me\nadd it as another patch at the end of this series.\n\nThanks for your review!\n\nPatrick\n"},{"id":"531107","messageId":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH v2 00/19] Refactor object read streams to work via object sources","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:45Z","receivedAt":"2025-11-21T07:41:03Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi,\n\nthe `git_istream` data structure can be used to read objects from the\nobject database in a streaming fashion. This is used for example to read\nlarge files that one doesn't want to load into memory in full.\n\nIn the current architecture, all the logic to handle these streams is\nfully self-contained in \"streaming.c\". It contains the logic to set up\nstreams for loose, packed, in-memory and filtered objects. This doesn't\nreally play all that well with pluggable object databases, as it should\nbe the responsibility of the object database source itself to handle the\nlogic.\n\nThis patch series thus revamps our object read streams: instead of being\nentirely contained in \"streaming.c\", the format-specific streams are now\ncreated by the ODB sources. This allows each source itself to decide\nwhether and, if so, how to make objects streamable.\n\nThis overall requires quite a bit of refactoring, but I think that the\nend result is an easier-to-understand infrastructure that is an\nimprovement even without pluggable object databases.\n\nThis series is built on top of v2.52.0 with ps/object-source-loose at\n3e5e360888 (object-file: refactor writing objects via a stream,\n2025-11-03) merged into it.\n\nChanges in v2:\n  - Some commit message improvements.\n  - Drop the `type` and `size` out pointers in\n    `odb_read_object_stream()` in an additional commit.\n  - Improve a \"hidden\" variable declaration by moving it onto its own\n    line.\n  - Link to v1: https://lore.kernel.org/r/20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (19):\n      streaming: rename `git_istream` into `odb_read_stream`\n      streaming: drop the `open()` callback function\n      streaming: propagate final object type via the stream\n      streaming: explicitly pass packfile info when streaming a packed object\n      streaming: allocate stream inside the backend-specific logic\n      streaming: create structure for in-core object streams\n      streaming: create structure for loose object streams\n      streaming: create structure for packed object streams\n      streaming: create structure for filtered object streams\n      streaming: move zlib stream into backends\n      packfile: introduce function to read object info from a store\n      streaming: rely on object sources to create object stream\n      streaming: get rid of `the_repository`\n      streaming: make the `odb_read_stream` definition public\n      streaming: move logic to read loose objects streams into backend\n      streaming: move logic to read packed objects streams into backend\n      streaming: refactor interface to be object-database-centric\n      streaming: move into object database subsystem\n      streaming: drop redundant type and size pointers\n\n Makefile               |   2 +-\n archive-tar.c          |  12 +-\n archive-zip.c          |  17 +-\n builtin/cat-file.c     |   4 +-\n builtin/fsck.c         |   5 +-\n builtin/index-pack.c   |  15 +-\n builtin/log.c          |   6 +-\n builtin/pack-objects.c |  24 ++-\n entry.c                |   4 +-\n meson.build            |   2 +-\n object-file.c          | 183 ++++++++++++++--\n object-file.h          |  42 +---\n odb.c                  |  29 +--\n odb/streaming.c        | 294 ++++++++++++++++++++++++++\n odb/streaming.h        |  67 ++++++\n packfile.c             | 199 ++++++++++++++++--\n packfile.h             |  17 +-\n parallel-checkout.c    |   5 +-\n streaming.c            | 561 -------------------------------------------------\n streaming.h            |  21 --\n 20 files changed, 780 insertions(+), 729 deletions(-)\n\nRange-diff versus v1:\n\n 1:  89ec27ae18 !  1:  a6534585dd streaming: rename `git_istream` into `odb_read_stream`\n    @@ streaming.h\n     -int close_istream(struct git_istream *);\n     -ssize_t read_istream(struct git_istream *, void *, size_t);\n     +struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n    -+\t\t\t\t       enum object_type *, unsigned long *,\n    -+\t\t\t\t       struct stream_filter *);\n    ++\t\t\t\t     enum object_type *, unsigned long *,\n    ++\t\t\t\t     struct stream_filter *);\n     +int close_istream(struct odb_read_stream *);\n     +ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n      \n 2:  b4d37fd4f2 !  2:  23a8704740 streaming: drop the `open()` callback function\n    @@ Commit message\n         structure in the source itself.\n     \n         We could of course work around that and have the loose object source\n    -    create the structure and populate it's `open()` callback, only. But\n    +    create the structure and populate its `open()` callback, only. But\n         this doesn't really buy us anything due to the second bullet point\n         above.\n     \n 3:  b8bae59f58 =  3:  badcc5d72b streaming: propagate final object type via the stream\n 4:  583ed2c4f3 =  4:  09f9d2e3f2 streaming: explicitly pass packfile info when streaming a packed object\n 5:  af1a5a312a !  5:  40728b509c streaming: allocate stream inside the backend-specific logic\n    @@ Commit message\n         This works, but it's awkward in the context of pluggable object\n         databases. Each backend will need its own member in that union, and as\n         the structure itself is completely opaque (it's only defined in\n    -    \"streamgin.c\") it also has the consequence that we must have the logic\n    +    \"streaming.c\") it also has the consequence that we must have the logic\n         that is specific to backends in \"streaming.c\".\n     \n         Ideally though, the infrastructure would be reversed: we have a generic\n 6:  5c5c291bba !  6:  7d74c31e3d streaming: create structure for in-core object streams\n    @@ streaming.c: static int open_istream_incore(struct odb_read_stream **out,\n     -\tstruct odb_read_stream stream = {\n     -\t\t.close = close_istream_incore,\n     -\t\t.read = read_istream_incore,\n    --\t};\n     +\tstruct odb_incore_read_stream stream = {\n     +\t\t.base.close = close_istream_incore,\n     +\t\t.base.read = read_istream_incore,\n    -+\t}, *st;\n    + \t};\n    ++\tstruct odb_incore_read_stream *st;\n      \tint ret;\n      \n     -\toi.typep = &stream.type;\n 7:  58d214e576 =  7:  dd3440bff2 streaming: create structure for loose object streams\n 8:  7b3d095e06 =  8:  6de8cc7c9f streaming: create structure for packed object streams\n 9:  3bca3dfab5 =  9:  e00aa2b198 streaming: create structure for filtered object streams\n10:  329549b6c7 = 10:  f37441494d streaming: move zlib stream into backends\n11:  9d47d12cbf = 11:  8c62cfac57 packfile: introduce function to read object info from a store\n12:  3a5ad53484 ! 12:  82f186e8b4 streaming: rely on object sources to create object stream\n    @@ Commit message\n         the different source backends anymore, but eventually it'll only have to\n         call the source's callback function.\n     \n    -    Note that at the current poin in time we aren't full there yet:\n    +    Note that at the current point in time we aren't fully there yet:\n     \n           - The packfile store still sits on the object database level and is\n             thus agnostic of the sources.\n13:  2fa2f53ac0 = 13:  a5c1b3c717 streaming: get rid of `the_repository`\n14:  49e6fb06e8 ! 14:  5fdd600a0c streaming: make the `odb_read_stream` definition public\n    @@ streaming.h\n     +};\n     +\n      struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n    - \t\t\t\t       enum object_type *, unsigned long *,\n    - \t\t\t\t       struct stream_filter *);\n    + \t\t\t\t     enum object_type *, unsigned long *,\n    + \t\t\t\t     struct stream_filter *);\n15:  3a944f3a31 = 15:  460cab31c9 streaming: move logic to read loose objects streams into backend\n16:  60b08e3dc5 = 16:  293578ab35 streaming: move logic to read packed objects streams into backend\n17:  68ef7721b0 ! 17:  e6a242f1b8 streaming: refactor interface to be object-database-centric\n    @@ streaming.h: struct odb_read_stream {\n      };\n      \n     -struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n    --\t\t\t\t       enum object_type *, unsigned long *,\n    --\t\t\t\t       struct stream_filter *);\n    +-\t\t\t\t     enum object_type *, unsigned long *,\n    +-\t\t\t\t     struct stream_filter *);\n     -int close_istream(struct odb_read_stream *);\n     -ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n     +/*\n18:  8afda7d038 = 18:  95e7c2aa9b streaming: move into object database subsystem\n -:  ---------- > 19:  c8b2112d00 streaming: drop redundant type and size pointers\n\n---\nbase-commit: 899e578b5b7c020aec806bd694adf2563f62843c\nchange-id: 20251107-b4-pks-odb-read-stream-7ea7f0e0a8f4\n\n"},{"id":"531108","messageId":"20251121-b4-pks-odb-read-stream-v2-1-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 01/19] streaming: rename `git_istream` into `odb_read_stream`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:46Z","receivedAt":"2025-11-21T07:41:06Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In the following patches we are about to make the `git_istream` more\ngeneric so that it becomes fully controlled by the specific object\nsource that wants to create it. As part of these refactorings we'll\nfully move the structure into the object database subsystem.\n\nPrepare for this change by renaming the structure from `git_istream`\nto `odb_read_stream`. This mirrors the `odb_write_stream` structure that\nwe already have.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  2 +-\n archive-zip.c          |  2 +-\n builtin/index-pack.c   |  2 +-\n builtin/pack-objects.c |  4 ++--\n object-file.c          |  2 +-\n streaming.c            | 62 +++++++++++++++++++++++++-------------------------\n streaming.h            | 12 +++++-----\n 7 files changed, 43 insertions(+), 43 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 73b63ddc41..dc1eda09e0 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -129,7 +129,7 @@ static void write_trailer(void)\n  */\n static int stream_blocked(struct repository *r, const struct object_id *oid)\n {\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tenum object_type type;\n \tunsigned long sz;\n \tchar buf[BLOCKSIZE];\ndiff --git a/archive-zip.c b/archive-zip.c\nindex bea5bdd43d..40a9c93ff9 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -309,7 +309,7 @@ static int write_zip_entry(struct archiver_args *args,\n \tenum zip_method method;\n \tunsigned char *out;\n \tvoid *deflated = NULL;\n-\tstruct git_istream *stream = NULL;\n+\tstruct odb_read_stream *stream = NULL;\n \tunsigned long flags = 0;\n \tint is_binary = -1;\n \tconst char *path_without_prefix = path + args->baselen;\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 2b78ba7fe4..5f90f12f92 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -762,7 +762,7 @@ static void find_ref_delta_children(const struct object_id *oid,\n \n struct compare_data {\n \tstruct object_entry *entry;\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tunsigned char *buf;\n \tunsigned long buf_size;\n };\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 69e80b1443..c693d948e1 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -404,7 +404,7 @@ static unsigned long do_compress(void **pptr, unsigned long size)\n \treturn stream.total_out;\n }\n \n-static unsigned long write_large_blob_data(struct git_istream *st, struct hashfile *f,\n+static unsigned long write_large_blob_data(struct odb_read_stream *st, struct hashfile *f,\n \t\t\t\t\t   const struct object_id *oid)\n {\n \tgit_zstream stream;\n@@ -513,7 +513,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \tunsigned hdrlen;\n \tenum object_type type;\n \tvoid *buf;\n-\tstruct git_istream *st = NULL;\n+\tstruct odb_read_stream *st = NULL;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \n \tif (!usable_delta) {\ndiff --git a/object-file.c b/object-file.c\nindex 811c569ed3..b62b21a452 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -134,7 +134,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tstruct object_id real_oid;\n \tunsigned long size;\n \tenum object_type obj_type;\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tstruct git_hash_ctx c;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\ndiff --git a/streaming.c b/streaming.c\nindex 00ad649ae3..1fb4b7c1c0 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -14,17 +14,17 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*open_istream_fn)(struct git_istream *,\n+typedef int (*open_istream_fn)(struct odb_read_stream *,\n \t\t\t       struct repository *,\n \t\t\t       const struct object_id *,\n \t\t\t       enum object_type *);\n-typedef int (*close_istream_fn)(struct git_istream *);\n-typedef ssize_t (*read_istream_fn)(struct git_istream *, char *, size_t);\n+typedef int (*close_istream_fn)(struct odb_read_stream *);\n+typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n #define FILTER_BUFFER (1024*16)\n \n struct filtered_istream {\n-\tstruct git_istream *upstream;\n+\tstruct odb_read_stream *upstream;\n \tstruct stream_filter *filter;\n \tchar ibuf[FILTER_BUFFER];\n \tchar obuf[FILTER_BUFFER];\n@@ -33,7 +33,7 @@ struct filtered_istream {\n \tint input_finished;\n };\n \n-struct git_istream {\n+struct odb_read_stream {\n \topen_istream_fn open;\n \tclose_istream_fn close;\n \tread_istream_fn read;\n@@ -71,7 +71,7 @@ struct git_istream {\n  *\n  *****************************************************************/\n \n-static void close_deflated_stream(struct git_istream *st)\n+static void close_deflated_stream(struct odb_read_stream *st)\n {\n \tif (st->z_state == z_used)\n \t\tgit_inflate_end(&st->z);\n@@ -84,13 +84,13 @@ static void close_deflated_stream(struct git_istream *st)\n  *\n  *****************************************************************/\n \n-static int close_istream_filtered(struct git_istream *st)\n+static int close_istream_filtered(struct odb_read_stream *st)\n {\n \tfree_stream_filter(st->u.filtered.filter);\n \treturn close_istream(st->u.filtered.upstream);\n }\n \n-static ssize_t read_istream_filtered(struct git_istream *st, char *buf,\n+static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n \t\t\t\t     size_t sz)\n {\n \tstruct filtered_istream *fs = &(st->u.filtered);\n@@ -150,10 +150,10 @@ static ssize_t read_istream_filtered(struct git_istream *st, char *buf,\n \treturn filled;\n }\n \n-static struct git_istream *attach_stream_filter(struct git_istream *st,\n-\t\t\t\t\t\tstruct stream_filter *filter)\n+static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n+\t\t\t\t\t\t    struct stream_filter *filter)\n {\n-\tstruct git_istream *ifs = xmalloc(sizeof(*ifs));\n+\tstruct odb_read_stream *ifs = xmalloc(sizeof(*ifs));\n \tstruct filtered_istream *fs = &(ifs->u.filtered);\n \n \tifs->close = close_istream_filtered;\n@@ -173,7 +173,7 @@ static struct git_istream *attach_stream_filter(struct git_istream *st,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_loose(struct git_istream *st, char *buf, size_t sz)\n+static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t sz)\n {\n \tsize_t total_read = 0;\n \n@@ -218,14 +218,14 @@ static ssize_t read_istream_loose(struct git_istream *st, char *buf, size_t sz)\n \treturn total_read;\n }\n \n-static int close_istream_loose(struct git_istream *st)\n+static int close_istream_loose(struct odb_read_stream *st)\n {\n \tclose_deflated_stream(st);\n \tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n \treturn 0;\n }\n \n-static int open_istream_loose(struct git_istream *st, struct repository *r,\n+static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \t\t\t      const struct object_id *oid,\n \t\t\t      enum object_type *type)\n {\n@@ -277,7 +277,7 @@ static int open_istream_loose(struct git_istream *st, struct repository *r,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_pack_non_delta(struct git_istream *st, char *buf,\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf,\n \t\t\t\t\t   size_t sz)\n {\n \tsize_t total_read = 0;\n@@ -336,13 +336,13 @@ static ssize_t read_istream_pack_non_delta(struct git_istream *st, char *buf,\n \treturn total_read;\n }\n \n-static int close_istream_pack_non_delta(struct git_istream *st)\n+static int close_istream_pack_non_delta(struct odb_read_stream *st)\n {\n \tclose_deflated_stream(st);\n \treturn 0;\n }\n \n-static int open_istream_pack_non_delta(struct git_istream *st,\n+static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n \t\t\t\t       const struct object_id *oid UNUSED,\n \t\t\t\t       enum object_type *type UNUSED)\n@@ -380,13 +380,13 @@ static int open_istream_pack_non_delta(struct git_istream *st,\n  *\n  *****************************************************************/\n \n-static int close_istream_incore(struct git_istream *st)\n+static int close_istream_incore(struct odb_read_stream *st)\n {\n \tfree(st->u.incore.buf);\n \treturn 0;\n }\n \n-static ssize_t read_istream_incore(struct git_istream *st, char *buf, size_t sz)\n+static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t sz)\n {\n \tsize_t read_size = sz;\n \tsize_t remainder = st->size - st->u.incore.read_ptr;\n@@ -400,7 +400,7 @@ static ssize_t read_istream_incore(struct git_istream *st, char *buf, size_t sz)\n \treturn read_size;\n }\n \n-static int open_istream_incore(struct git_istream *st, struct repository *r,\n+static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \t\t\t       const struct object_id *oid, enum object_type *type)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n@@ -420,7 +420,7 @@ static int open_istream_incore(struct git_istream *st, struct repository *r,\n  * static helpers variables and functions for users of streaming interface\n  *****************************************************************************/\n \n-static int istream_source(struct git_istream *st,\n+static int istream_source(struct odb_read_stream *st,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid,\n \t\t\t  enum object_type *type)\n@@ -458,25 +458,25 @@ static int istream_source(struct git_istream *st,\n  * Users of streaming interface\n  ****************************************************************/\n \n-int close_istream(struct git_istream *st)\n+int close_istream(struct odb_read_stream *st)\n {\n \tint r = st->close(st);\n \tfree(st);\n \treturn r;\n }\n \n-ssize_t read_istream(struct git_istream *st, void *buf, size_t sz)\n+ssize_t read_istream(struct odb_read_stream *st, void *buf, size_t sz)\n {\n \treturn st->read(st, buf, sz);\n }\n \n-struct git_istream *open_istream(struct repository *r,\n-\t\t\t\t const struct object_id *oid,\n-\t\t\t\t enum object_type *type,\n-\t\t\t\t unsigned long *size,\n-\t\t\t\t struct stream_filter *filter)\n+struct odb_read_stream *open_istream(struct repository *r,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     enum object_type *type,\n+\t\t\t\t     unsigned long *size,\n+\t\t\t\t     struct stream_filter *filter)\n {\n-\tstruct git_istream *st = xmalloc(sizeof(*st));\n+\tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n \tint ret = istream_source(st, r, real, type);\n \n@@ -493,7 +493,7 @@ struct git_istream *open_istream(struct repository *r,\n \t}\n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n-\t\tstruct git_istream *nst = attach_stream_filter(st, filter);\n+\t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n \t\tif (!nst) {\n \t\t\tclose_istream(st);\n \t\t\treturn NULL;\n@@ -508,7 +508,7 @@ struct git_istream *open_istream(struct repository *r,\n int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter *filter,\n \t\t      int can_seek)\n {\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tenum object_type type;\n \tunsigned long sz;\n \tssize_t kept = 0;\ndiff --git a/streaming.h b/streaming.h\nindex bd27f59e57..f5ff5d7ac9 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -7,14 +7,14 @@\n #include \"object.h\"\n \n /* opaque */\n-struct git_istream;\n+struct odb_read_stream;\n struct stream_filter;\n \n-struct git_istream *open_istream(struct repository *, const struct object_id *,\n-\t\t\t\t enum object_type *, unsigned long *,\n-\t\t\t\t struct stream_filter *);\n-int close_istream(struct git_istream *);\n-ssize_t read_istream(struct git_istream *, void *, size_t);\n+struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n+\t\t\t\t     enum object_type *, unsigned long *,\n+\t\t\t\t     struct stream_filter *);\n+int close_istream(struct odb_read_stream *);\n+ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n \n int stream_blob_to_fd(int fd, const struct object_id *, struct stream_filter *, int can_seek);\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531109","messageId":"20251121-b4-pks-odb-read-stream-v2-2-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 02/19] streaming: drop the `open()` callback function","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:47Z","receivedAt":"2025-11-21T07:41:10Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating a read stream we first populate the structure with the\nopen callback function and then subsequently call the function. This\nlayout is somewhat weird though:\n\n  - The structure needs to be allocated and partially populated with the\n    open function before we can properly initialize it.\n\n  - We never use the `open()` callback after having opened it initially.\n\nEspecially the first point creates a problem for us. In subsequent\ncommits we'll want to fully move construction of the read source into\nthe respective object sources. E.g., the loose object source will be the\none that is responsible for creating the structure. But this creates a\nproblem: if we first need to create the structure so that we can call\nthe source-specific callback we cannot fully handle creation of the\nstructure in the source itself.\n\nWe could of course work around that and have the loose object source\ncreate the structure and populate its `open()` callback, only. But\nthis doesn't really buy us anything due to the second bullet point\nabove.\n\nInstead, drop the callback entirely and refactor `istream_source()` so\nthat we open the streams immediately. This unblocks a subsequent step,\nwhere we'll also start to allocate the structure in the source-specific\nlogic.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 40 +++++++++++++++++-----------------------\n 1 file changed, 17 insertions(+), 23 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 1fb4b7c1c0..5ce6350123 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -14,10 +14,6 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*open_istream_fn)(struct odb_read_stream *,\n-\t\t\t       struct repository *,\n-\t\t\t       const struct object_id *,\n-\t\t\t       enum object_type *);\n typedef int (*close_istream_fn)(struct odb_read_stream *);\n typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n@@ -34,7 +30,6 @@ struct filtered_istream {\n };\n \n struct odb_read_stream {\n-\topen_istream_fn open;\n \tclose_istream_fn close;\n \tread_istream_fn read;\n \n@@ -437,21 +432,25 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tst->open = open_istream_loose;\n+\t\tif (open_istream_loose(st, r, oid, type) < 0)\n+\t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n-\t\tif (!oi.u.packed.is_delta &&\n-\t\t    repo_settings_get_big_file_threshold(the_repository) < size) {\n-\t\t\tst->u.in_pack.pack = oi.u.packed.pack;\n-\t\t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\t\tst->open = open_istream_pack_non_delta;\n-\t\t\treturn 0;\n-\t\t}\n-\t\t/* fallthru */\n-\tdefault:\n-\t\tst->open = open_istream_incore;\n+\t\tif (oi.u.packed.is_delta ||\n+\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\t\tbreak;\n+\n+\t\tst->u.in_pack.pack = oi.u.packed.pack;\n+\t\tst->u.in_pack.pos = oi.u.packed.offset;\n+\t\tif (open_istream_pack_non_delta(st, r, oid, type) < 0)\n+\t\t\tbreak;\n+\n \t\treturn 0;\n+\tdefault:\n+\t\tbreak;\n \t}\n+\n+\treturn open_istream_incore(st, r, oid, type);\n }\n \n /****************************************************************\n@@ -478,19 +477,14 @@ struct odb_read_stream *open_istream(struct repository *r,\n {\n \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n-\tint ret = istream_source(st, r, real, type);\n+\tint ret;\n \n+\tret = istream_source(st, r, real, type);\n \tif (ret) {\n \t\tfree(st);\n \t\treturn NULL;\n \t}\n \n-\tif (st->open(st, r, real, type)) {\n-\t\tif (open_istream_incore(st, r, real, type)) {\n-\t\t\tfree(st);\n-\t\t\treturn NULL;\n-\t\t}\n-\t}\n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n \t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531110","messageId":"20251121-b4-pks-odb-read-stream-v2-3-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 03/19] streaming: propagate final object type via the stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:48Z","receivedAt":"2025-11-21T07:41:13Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When opening the read stream for a specific object the caller is also\nexpected to pass in a pointer to the object type. This type is passed\ndown via multiple levels and will eventually be populated with the type\nof the looked-up object.\n\nThe way we propagate down the pointer though is somewhat non-obvious.\nWhile `istream_source()` still expects the pointer and looks it up via\n`odb_read_object_info_extended()`, we also pass it down even further\ninto the format-specific callbacks that perform another lookup. This is\nquite confusing overall.\n\nRefactor the code so that the responsibility to populate the object type\nrests solely with the format-specific callbacks. This will allow us to\ndrop the call to `odb_read_object_info_extended()` in `istream_source()`\nentirely in a subsequent patch.\n\nFurthermore, instead of propagating the type via an in-pointer, we now\npropagate the type via a new field in the object stream. It already has\na `size` field, so it's only natural to have a second field that\ncontains the object type.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 30 +++++++++++++++---------------\n 1 file changed, 15 insertions(+), 15 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 5ce6350123..9596a94c58 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -33,6 +33,7 @@ struct odb_read_stream {\n \tclose_istream_fn close;\n \tread_istream_fn read;\n \n+\tenum object_type type;\n \tunsigned long size; /* inflated size of full object */\n \tgit_zstream z;\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n@@ -159,6 +160,7 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \tfs->o_end = fs->o_ptr = 0;\n \tfs->input_finished = 0;\n \tifs->size = -1; /* unknown */\n+\tifs->type = st->type;\n \treturn ifs;\n }\n \n@@ -221,14 +223,13 @@ static int close_istream_loose(struct odb_read_stream *st)\n }\n \n static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n-\t\t\t      const struct object_id *oid,\n-\t\t\t      enum object_type *type)\n+\t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_source *source;\n \n \toi.sizep = &st->size;\n-\toi.typep = type;\n+\toi.typep = &st->type;\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next) {\n@@ -249,7 +250,7 @@ static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \tcase ULHR_TOO_LONG:\n \t\tgoto error;\n \t}\n-\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || *type < 0)\n+\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n \t\tgoto error;\n \n \tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n@@ -339,8 +340,7 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \n static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED,\n-\t\t\t\t       enum object_type *type UNUSED)\n+\t\t\t\t       const struct object_id *oid UNUSED)\n {\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n@@ -361,6 +361,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n+\tst->type = in_pack_type;\n \tst->z_state = z_unused;\n \tst->close = close_istream_pack_non_delta;\n \tst->read = read_istream_pack_non_delta;\n@@ -396,7 +397,7 @@ static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t\n }\n \n static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n-\t\t\t       const struct object_id *oid, enum object_type *type)\n+\t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \n@@ -404,7 +405,7 @@ static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \tst->close = close_istream_incore;\n \tst->read = read_istream_incore;\n \n-\toi.typep = type;\n+\toi.typep = &st->type;\n \toi.sizep = &st->size;\n \toi.contentp = (void **)&st->u.incore.buf;\n \treturn odb_read_object_info_extended(r->objects, oid, &oi,\n@@ -417,14 +418,12 @@ static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \n static int istream_source(struct odb_read_stream *st,\n \t\t\t  struct repository *r,\n-\t\t\t  const struct object_id *oid,\n-\t\t\t  enum object_type *type)\n+\t\t\t  const struct object_id *oid)\n {\n \tunsigned long size;\n \tint status;\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \n-\toi.typep = type;\n \toi.sizep = &size;\n \tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n \tif (status < 0)\n@@ -432,7 +431,7 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tif (open_istream_loose(st, r, oid, type) < 0)\n+\t\tif (open_istream_loose(st, r, oid) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n@@ -442,7 +441,7 @@ static int istream_source(struct odb_read_stream *st,\n \n \t\tst->u.in_pack.pack = oi.u.packed.pack;\n \t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\tif (open_istream_pack_non_delta(st, r, oid, type) < 0)\n+\t\tif (open_istream_pack_non_delta(st, r, oid) < 0)\n \t\t\tbreak;\n \n \t\treturn 0;\n@@ -450,7 +449,7 @@ static int istream_source(struct odb_read_stream *st,\n \t\tbreak;\n \t}\n \n-\treturn open_istream_incore(st, r, oid, type);\n+\treturn open_istream_incore(st, r, oid);\n }\n \n /****************************************************************\n@@ -479,7 +478,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n \tint ret;\n \n-\tret = istream_source(st, r, real, type);\n+\tret = istream_source(st, r, real);\n \tif (ret) {\n \t\tfree(st);\n \t\treturn NULL;\n@@ -496,6 +495,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t}\n \n \t*size = st->size;\n+\t*type = st->type;\n \treturn st;\n }\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531111","messageId":"20251121-b4-pks-odb-read-stream-v2-4-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 04/19] streaming: explicitly pass packfile info when streaming a packed object","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:49Z","receivedAt":"2025-11-21T07:41:17Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When streaming a packed object we first populate the stream with\ninformation about the pack that contains the object before calling\n`open_istream_pack_non_delta()`. This is done because we have already\nlooked up both the pack and the object's offset, so it would be a waste\nof time to look up this information again.\n\nBut the way this is done makes for a somewhat awkward calling interface,\nas the caller now needs to be aware of how exactly the function itself\nbehaves.\n\nRefactor the code so that we instead explicitly pass the packfile info\ninto `open_istream_pack_non_delta()`. This makes the calling convention\nexplicit, but more importantly this allows us to refactor the function\nso that it becomes its responsibility to allocate the stream itself in a\nsubsequent patch.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 20 ++++++++++----------\n 1 file changed, 10 insertions(+), 10 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 9596a94c58..d7db446d25 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -340,16 +340,18 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \n static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED)\n+\t\t\t\t       const struct object_id *oid UNUSED,\n+\t\t\t\t       struct packed_git *pack,\n+\t\t\t\t       off_t offset)\n {\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n \n \twindow = NULL;\n \n-\tin_pack_type = unpack_object_header(st->u.in_pack.pack,\n+\tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n-\t\t\t\t\t    &st->u.in_pack.pos,\n+\t\t\t\t\t    &offset,\n \t\t\t\t\t    &st->size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n@@ -365,6 +367,8 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tst->z_state = z_unused;\n \tst->close = close_istream_pack_non_delta;\n \tst->read = read_istream_pack_non_delta;\n+\tst->u.in_pack.pack = pack;\n+\tst->u.in_pack.pos = offset;\n \n \treturn 0;\n }\n@@ -436,14 +440,10 @@ static int istream_source(struct odb_read_stream *st,\n \t\treturn 0;\n \tcase OI_PACKED:\n \t\tif (oi.u.packed.is_delta ||\n-\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n+\t\t    open_istream_pack_non_delta(st, r, oid, oi.u.packed.pack,\n+\t\t\t\t\t\toi.u.packed.offset) < 0)\n \t\t\tbreak;\n-\n-\t\tst->u.in_pack.pack = oi.u.packed.pack;\n-\t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\tif (open_istream_pack_non_delta(st, r, oid) < 0)\n-\t\t\tbreak;\n-\n \t\treturn 0;\n \tdefault:\n \t\tbreak;\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531112","messageId":"20251121-b4-pks-odb-read-stream-v2-5-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 05/19] streaming: allocate stream inside the backend-specific logic","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:50Z","receivedAt":"2025-11-21T07:41:21Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating a new stream we first allocate it and then call into\nbackend-specific logic to populate the stream. This design requires that\nthe stream itself contains a `union` with backend-specific members that\nthen ultimately get populated by the backend-specific logic.\n\nThis works, but it's awkward in the context of pluggable object\ndatabases. Each backend will need its own member in that union, and as\nthe structure itself is completely opaque (it's only defined in\n\"streaming.c\") it also has the consequence that we must have the logic\nthat is specific to backends in \"streaming.c\".\n\nIdeally though, the infrastructure would be reversed: we have a generic\n`struct odb_read_stream` and some helper functions in \"streaming.c\",\nwhereas the backend-specific logic sits in the backend's subsystem\nitself.\n\nThis can be realized by using a design that is similar to how we handle\nreference databases: instead of having a union of members, we instead\nhave backend-specific structures with a `struct odb_read_stream base`\nas its first member. The backends would thus hand out the pointer to the\nbase, but internally they know to cast back to the backend-specific\ntype.\n\nThis means though that we need to allocate different structures\ndepending on the backend. To prepare for this, move allocation of the\nstructure into the backend-specific functions that open a new stream.\nSubsequent commits will then create those new backend-specific structs.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 99 +++++++++++++++++++++++++++++++++++++++----------------------\n 1 file changed, 63 insertions(+), 36 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex d7db446d25..b8ce82483f 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -222,27 +222,34 @@ static int close_istream_loose(struct odb_read_stream *st)\n \treturn 0;\n }\n \n-static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n+static int open_istream_loose(struct odb_read_stream **out,\n+\t\t\t      struct repository *r,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_read_stream *st;\n \tstruct odb_source *source;\n-\n-\toi.sizep = &st->size;\n-\toi.typep = &st->type;\n+\tunsigned long mapsize;\n+\tvoid *mapped;\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next) {\n-\t\tst->u.loose.mapped = odb_source_loose_map_object(source, oid,\n-\t\t\t\t\t\t\t\t &st->u.loose.mapsize);\n-\t\tif (st->u.loose.mapped)\n+\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n+\t\tif (mapped)\n \t\t\tbreak;\n \t}\n-\tif (!st->u.loose.mapped)\n+\tif (!mapped)\n \t\treturn -1;\n \n-\tswitch (unpack_loose_header(&st->z, st->u.loose.mapped,\n-\t\t\t\t    st->u.loose.mapsize, st->u.loose.hdr,\n+\t/*\n+\t * Note: we must allocate this structure early even though we may still\n+\t * fail. This is because we need to initialize the zlib stream, and it\n+\t * is not possible to copy the stream around after the fact because it\n+\t * has self-referencing pointers.\n+\t */\n+\tCALLOC_ARRAY(st, 1);\n+\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->u.loose.hdr,\n \t\t\t\t    sizeof(st->u.loose.hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n@@ -250,19 +257,28 @@ static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \tcase ULHR_TOO_LONG:\n \t\tgoto error;\n \t}\n+\n+\toi.sizep = &st->size;\n+\toi.typep = &st->type;\n+\n \tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n \t\tgoto error;\n \n+\tst->u.loose.mapped = mapped;\n+\tst->u.loose.mapsize = mapsize;\n \tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n \tst->u.loose.hdr_avail = st->z.total_out;\n \tst->z_state = z_used;\n \tst->close = close_istream_loose;\n \tst->read = read_istream_loose;\n \n+\t*out = st;\n+\n \treturn 0;\n error:\n \tgit_inflate_end(&st->z);\n \tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tfree(st);\n \treturn -1;\n }\n \n@@ -338,12 +354,16 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \treturn 0;\n }\n \n-static int open_istream_pack_non_delta(struct odb_read_stream *st,\n+static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \t\t\t\t       struct repository *r UNUSED,\n \t\t\t\t       const struct object_id *oid UNUSED,\n \t\t\t\t       struct packed_git *pack,\n \t\t\t\t       off_t offset)\n {\n+\tstruct odb_read_stream stream = {\n+\t\t.close = close_istream_pack_non_delta,\n+\t\t.read = read_istream_pack_non_delta,\n+\t};\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n \n@@ -352,7 +372,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n \t\t\t\t\t    &offset,\n-\t\t\t\t\t    &st->size);\n+\t\t\t\t\t    &stream.size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n \tdefault:\n@@ -363,12 +383,13 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n-\tst->type = in_pack_type;\n-\tst->z_state = z_unused;\n-\tst->close = close_istream_pack_non_delta;\n-\tst->read = read_istream_pack_non_delta;\n-\tst->u.in_pack.pack = pack;\n-\tst->u.in_pack.pos = offset;\n+\tstream.type = in_pack_type;\n+\tstream.z_state = z_unused;\n+\tstream.u.in_pack.pack = pack;\n+\tstream.u.in_pack.pos = offset;\n+\n+\tCALLOC_ARRAY(*out, 1);\n+\t**out = stream;\n \n \treturn 0;\n }\n@@ -400,27 +421,35 @@ static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t\n \treturn read_size;\n }\n \n-static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n+static int open_istream_incore(struct odb_read_stream **out,\n+\t\t\t       struct repository *r,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_read_stream stream = {\n+\t\t.close = close_istream_incore,\n+\t\t.read = read_istream_incore,\n+\t};\n+\tint ret;\n \n-\tst->u.incore.read_ptr = 0;\n-\tst->close = close_istream_incore;\n-\tst->read = read_istream_incore;\n+\toi.typep = &stream.type;\n+\toi.sizep = &stream.size;\n+\toi.contentp = (void **)&stream.u.incore.buf;\n+\tret = odb_read_object_info_extended(r->objects, oid, &oi,\n+\t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n+\tif (ret)\n+\t\treturn ret;\n \n-\toi.typep = &st->type;\n-\toi.sizep = &st->size;\n-\toi.contentp = (void **)&st->u.incore.buf;\n-\treturn odb_read_object_info_extended(r->objects, oid, &oi,\n-\t\t\t\t\t     OBJECT_INFO_DIE_IF_CORRUPT);\n+\tCALLOC_ARRAY(*out, 1);\n+\t**out = stream;\n+\treturn 0;\n }\n \n /*****************************************************************************\n  * static helpers variables and functions for users of streaming interface\n  *****************************************************************************/\n \n-static int istream_source(struct odb_read_stream *st,\n+static int istream_source(struct odb_read_stream **out,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid)\n {\n@@ -435,13 +464,13 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tif (open_istream_loose(st, r, oid) < 0)\n+\t\tif (open_istream_loose(out, r, oid) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n \t\tif (oi.u.packed.is_delta ||\n \t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n-\t\t    open_istream_pack_non_delta(st, r, oid, oi.u.packed.pack,\n+\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n \t\t\t\t\t\toi.u.packed.offset) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n@@ -449,7 +478,7 @@ static int istream_source(struct odb_read_stream *st,\n \t\tbreak;\n \t}\n \n-\treturn open_istream_incore(st, r, oid);\n+\treturn open_istream_incore(out, r, oid);\n }\n \n /****************************************************************\n@@ -474,15 +503,13 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t\t\t\t     unsigned long *size,\n \t\t\t\t     struct stream_filter *filter)\n {\n-\tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n+\tstruct odb_read_stream *st;\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n \tint ret;\n \n-\tret = istream_source(st, r, real);\n-\tif (ret) {\n-\t\tfree(st);\n+\tret = istream_source(&st, r, real);\n+\tif (ret)\n \t\treturn NULL;\n-\t}\n \n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531113","messageId":"20251121-b4-pks-odb-read-stream-v2-6-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 06/19] streaming: create structure for in-core object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:51Z","receivedAt":"2025-11-21T07:41:24Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for in-core object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 44 +++++++++++++++++++++++++-------------------\n 1 file changed, 25 insertions(+), 19 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex b8ce82483f..3af2f0c776 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,11 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tchar *buf; /* from odb_read_object_info_extended() */\n-\t\t\tunsigned long read_ptr;\n-\t\t} incore;\n-\n \t\tstruct {\n \t\t\tvoid *mapped;\n \t\t\tunsigned long mapsize;\n@@ -401,22 +396,30 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n  *\n  *****************************************************************/\n \n-static int close_istream_incore(struct odb_read_stream *st)\n+struct odb_incore_read_stream {\n+\tstruct odb_read_stream base;\n+\tchar *buf; /* from odb_read_object_info_extended() */\n+\tunsigned long read_ptr;\n+};\n+\n+static int close_istream_incore(struct odb_read_stream *_st)\n {\n-\tfree(st->u.incore.buf);\n+\tstruct odb_incore_read_stream *st = (struct odb_incore_read_stream *)_st;\n+\tfree(st->buf);\n \treturn 0;\n }\n \n-static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t sz)\n+static ssize_t read_istream_incore(struct odb_read_stream *_st, char *buf, size_t sz)\n {\n+\tstruct odb_incore_read_stream *st = (struct odb_incore_read_stream *)_st;\n \tsize_t read_size = sz;\n-\tsize_t remainder = st->size - st->u.incore.read_ptr;\n+\tsize_t remainder = st->base.size - st->read_ptr;\n \n \tif (remainder <= read_size)\n \t\tread_size = remainder;\n \tif (read_size) {\n-\t\tmemcpy(buf, st->u.incore.buf + st->u.incore.read_ptr, read_size);\n-\t\tst->u.incore.read_ptr += read_size;\n+\t\tmemcpy(buf, st->buf + st->read_ptr, read_size);\n+\t\tst->read_ptr += read_size;\n \t}\n \treturn read_size;\n }\n@@ -426,22 +429,25 @@ static int open_istream_incore(struct odb_read_stream **out,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_read_stream stream = {\n-\t\t.close = close_istream_incore,\n-\t\t.read = read_istream_incore,\n+\tstruct odb_incore_read_stream stream = {\n+\t\t.base.close = close_istream_incore,\n+\t\t.base.read = read_istream_incore,\n \t};\n+\tstruct odb_incore_read_stream *st;\n \tint ret;\n \n-\toi.typep = &stream.type;\n-\toi.sizep = &stream.size;\n-\toi.contentp = (void **)&stream.u.incore.buf;\n+\toi.typep = &stream.base.type;\n+\toi.sizep = &stream.base.size;\n+\toi.contentp = (void **)&stream.buf;\n \tret = odb_read_object_info_extended(r->objects, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n \n-\tCALLOC_ARRAY(*out, 1);\n-\t**out = stream;\n+\tCALLOC_ARRAY(st, 1);\n+\t*st = stream;\n+\t*out = &st->base;\n+\n \treturn 0;\n }\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531114","messageId":"20251121-b4-pks-odb-read-stream-v2-7-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 07/19] streaming: create structure for loose object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:52Z","receivedAt":"2025-11-21T07:41:28Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for loose object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 85 ++++++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 44 insertions(+), 41 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 3af2f0c776..193405d11e 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,14 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tvoid *mapped;\n-\t\t\tunsigned long mapsize;\n-\t\t\tchar hdr[32];\n-\t\t\tint hdr_avail;\n-\t\t\tint hdr_used;\n-\t\t} loose;\n-\n \t\tstruct {\n \t\t\tstruct packed_git *pack;\n \t\t\toff_t pos;\n@@ -165,11 +157,21 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t sz)\n+struct odb_loose_read_stream {\n+\tstruct odb_read_stream base;\n+\tvoid *mapped;\n+\tunsigned long mapsize;\n+\tchar hdr[32];\n+\tint hdr_avail;\n+\tint hdr_used;\n+};\n+\n+static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n {\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->z_state) {\n+\tswitch (st->base.z_state) {\n \tcase z_done:\n \t\treturn 0;\n \tcase z_error:\n@@ -178,42 +180,43 @@ static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t\n \t\tbreak;\n \t}\n \n-\tif (st->u.loose.hdr_used < st->u.loose.hdr_avail) {\n-\t\tsize_t to_copy = st->u.loose.hdr_avail - st->u.loose.hdr_used;\n+\tif (st->hdr_used < st->hdr_avail) {\n+\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n \t\tif (sz < to_copy)\n \t\t\tto_copy = sz;\n-\t\tmemcpy(buf, st->u.loose.hdr + st->u.loose.hdr_used, to_copy);\n-\t\tst->u.loose.hdr_used += to_copy;\n+\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n+\t\tst->hdr_used += to_copy;\n \t\ttotal_read += to_copy;\n \t}\n \n \twhile (total_read < sz) {\n \t\tint status;\n \n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->base.z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n \n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_done;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_done;\n \t\t\tbreak;\n \t\t}\n \t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_error;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_error;\n \t\t\treturn -1;\n \t\t}\n \t}\n \treturn total_read;\n }\n \n-static int close_istream_loose(struct odb_read_stream *st)\n+static int close_istream_loose(struct odb_read_stream *_st)\n {\n-\tclose_deflated_stream(st);\n-\tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tclose_deflated_stream(&st->base);\n+\tmunmap(st->mapped, st->mapsize);\n \treturn 0;\n }\n \n@@ -222,7 +225,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_read_stream *st;\n+\tstruct odb_loose_read_stream *st;\n \tstruct odb_source *source;\n \tunsigned long mapsize;\n \tvoid *mapped;\n@@ -244,8 +247,8 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t */\n \tCALLOC_ARRAY(st, 1);\n \n-\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->u.loose.hdr,\n-\t\t\t\t    sizeof(st->u.loose.hdr))) {\n+\tswitch (unpack_loose_header(&st->base.z, mapped, mapsize, st->hdr,\n+\t\t\t\t    sizeof(st->hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n \tcase ULHR_BAD:\n@@ -253,26 +256,26 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t\tgoto error;\n \t}\n \n-\toi.sizep = &st->size;\n-\toi.typep = &st->type;\n+\toi.sizep = &st->base.size;\n+\toi.typep = &st->base.type;\n \n-\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n+\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n \t\tgoto error;\n \n-\tst->u.loose.mapped = mapped;\n-\tst->u.loose.mapsize = mapsize;\n-\tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n-\tst->u.loose.hdr_avail = st->z.total_out;\n-\tst->z_state = z_used;\n-\tst->close = close_istream_loose;\n-\tst->read = read_istream_loose;\n+\tst->mapped = mapped;\n+\tst->mapsize = mapsize;\n+\tst->hdr_used = strlen(st->hdr) + 1;\n+\tst->hdr_avail = st->base.z.total_out;\n+\tst->base.z_state = z_used;\n+\tst->base.close = close_istream_loose;\n+\tst->base.read = read_istream_loose;\n \n-\t*out = st;\n+\t*out = &st->base;\n \n \treturn 0;\n error:\n-\tgit_inflate_end(&st->z);\n-\tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tgit_inflate_end(&st->base.z);\n+\tmunmap(st->mapped, st->mapsize);\n \tfree(st);\n \treturn -1;\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531115","messageId":"20251121-b4-pks-odb-read-stream-v2-8-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 08/19] streaming: create structure for packed object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:53Z","receivedAt":"2025-11-21T07:41:32Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for packed object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 75 ++++++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 40 insertions(+), 35 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 193405d11e..014c9b8d90 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,11 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tstruct packed_git *pack;\n-\t\t\toff_t pos;\n-\t\t} in_pack;\n-\n \t\tstruct filtered_istream filtered;\n \t} u;\n };\n@@ -287,16 +282,23 @@ static int open_istream_loose(struct odb_read_stream **out,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf,\n+struct odb_packed_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct packed_git *pack;\n+\toff_t pos;\n+};\n+\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n \t\t\t\t\t   size_t sz)\n {\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->z_state) {\n+\tswitch (st->base.z_state) {\n \tcase z_unused:\n-\t\tmemset(&st->z, 0, sizeof(st->z));\n-\t\tgit_inflate_init(&st->z);\n-\t\tst->z_state = z_used;\n+\t\tmemset(&st->base.z, 0, sizeof(st->base.z));\n+\t\tgit_inflate_init(&st->base.z);\n+\t\tst->base.z_state = z_used;\n \t\tbreak;\n \tcase z_done:\n \t\treturn 0;\n@@ -311,21 +313,21 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf\n \t\tstruct pack_window *window = NULL;\n \t\tunsigned char *mapped;\n \n-\t\tmapped = use_pack(st->u.in_pack.pack, &window,\n-\t\t\t\t  st->u.in_pack.pos, &st->z.avail_in);\n+\t\tmapped = use_pack(st->pack, &window,\n+\t\t\t\t  st->pos, &st->base.z.avail_in);\n \n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tst->z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->base.z.avail_out = sz - total_read;\n+\t\tst->base.z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n \n-\t\tst->u.in_pack.pos += st->z.next_in - mapped;\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\tst->pos += st->base.z.next_in - mapped;\n+\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n \t\tunuse_pack(&window);\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_done;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_done;\n \t\t\tbreak;\n \t\t}\n \n@@ -338,17 +340,18 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf\n \t\t * or truncated), then use_pack() catches that and will die().\n \t\t */\n \t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_error;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_error;\n \t\t\treturn -1;\n \t\t}\n \t}\n \treturn total_read;\n }\n \n-static int close_istream_pack_non_delta(struct odb_read_stream *st)\n+static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n {\n-\tclose_deflated_stream(st);\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tclose_deflated_stream(&st->base);\n \treturn 0;\n }\n \n@@ -358,19 +361,17 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \t\t\t\t       struct packed_git *pack,\n \t\t\t\t       off_t offset)\n {\n-\tstruct odb_read_stream stream = {\n-\t\t.close = close_istream_pack_non_delta,\n-\t\t.read = read_istream_pack_non_delta,\n-\t};\n+\tstruct odb_packed_read_stream *stream;\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n+\tsize_t size;\n \n \twindow = NULL;\n \n \tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n \t\t\t\t\t    &offset,\n-\t\t\t\t\t    &stream.size);\n+\t\t\t\t\t    &size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n \tdefault:\n@@ -381,13 +382,17 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n-\tstream.type = in_pack_type;\n-\tstream.z_state = z_unused;\n-\tstream.u.in_pack.pack = pack;\n-\tstream.u.in_pack.pos = offset;\n \n-\tCALLOC_ARRAY(*out, 1);\n-\t**out = stream;\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.close = close_istream_pack_non_delta;\n+\tstream->base.read = read_istream_pack_non_delta;\n+\tstream->base.type = in_pack_type;\n+\tstream->base.size = size;\n+\tstream->base.z_state = z_unused;\n+\tstream->pack = pack;\n+\tstream->pos = offset;\n+\n+\t*out = &stream->base;\n \n \treturn 0;\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531116","messageId":"20251121-b4-pks-odb-read-stream-v2-9-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 09/19] streaming: create structure for filtered object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:54Z","receivedAt":"2025-11-21T07:41:35Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for filtered object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 54 +++++++++++++++++++++++++-----------------------------\n 1 file changed, 25 insertions(+), 29 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 014c9b8d90..45463b5c55 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -19,16 +19,6 @@ typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n #define FILTER_BUFFER (1024*16)\n \n-struct filtered_istream {\n-\tstruct odb_read_stream *upstream;\n-\tstruct stream_filter *filter;\n-\tchar ibuf[FILTER_BUFFER];\n-\tchar obuf[FILTER_BUFFER];\n-\tint i_end, i_ptr;\n-\tint o_end, o_ptr;\n-\tint input_finished;\n-};\n-\n struct odb_read_stream {\n \tclose_istream_fn close;\n \tread_istream_fn read;\n@@ -37,10 +27,6 @@ struct odb_read_stream {\n \tunsigned long size; /* inflated size of full object */\n \tgit_zstream z;\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n-\n-\tunion {\n-\t\tstruct filtered_istream filtered;\n-\t} u;\n };\n \n /*****************************************************************\n@@ -62,16 +48,28 @@ static void close_deflated_stream(struct odb_read_stream *st)\n  *\n  *****************************************************************/\n \n-static int close_istream_filtered(struct odb_read_stream *st)\n+struct odb_filtered_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct odb_read_stream *upstream;\n+\tstruct stream_filter *filter;\n+\tchar ibuf[FILTER_BUFFER];\n+\tchar obuf[FILTER_BUFFER];\n+\tint i_end, i_ptr;\n+\tint o_end, o_ptr;\n+\tint input_finished;\n+};\n+\n+static int close_istream_filtered(struct odb_read_stream *_fs)\n {\n-\tfree_stream_filter(st->u.filtered.filter);\n-\treturn close_istream(st->u.filtered.upstream);\n+\tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n+\tfree_stream_filter(fs->filter);\n+\treturn close_istream(fs->upstream);\n }\n \n-static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n+static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n \t\t\t\t     size_t sz)\n {\n-\tstruct filtered_istream *fs = &(st->u.filtered);\n+\tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n \tsize_t filled = 0;\n \n \twhile (sz) {\n@@ -131,19 +129,17 @@ static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \t\t\t\t\t\t    struct stream_filter *filter)\n {\n-\tstruct odb_read_stream *ifs = xmalloc(sizeof(*ifs));\n-\tstruct filtered_istream *fs = &(ifs->u.filtered);\n+\tstruct odb_filtered_read_stream *fs;\n \n-\tifs->close = close_istream_filtered;\n-\tifs->read = read_istream_filtered;\n+\tCALLOC_ARRAY(fs, 1);\n+\tfs->base.close = close_istream_filtered;\n+\tfs->base.read = read_istream_filtered;\n \tfs->upstream = st;\n \tfs->filter = filter;\n-\tfs->i_end = fs->i_ptr = 0;\n-\tfs->o_end = fs->o_ptr = 0;\n-\tfs->input_finished = 0;\n-\tifs->size = -1; /* unknown */\n-\tifs->type = st->type;\n-\treturn ifs;\n+\tfs->base.size = -1; /* unknown */\n+\tfs->base.type = st->type;\n+\n+\treturn &fs->base;\n }\n \n /*****************************************************************\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531117","messageId":"20251121-b4-pks-odb-read-stream-v2-10-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 10/19] streaming: move zlib stream into backends","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:55Z","receivedAt":"2025-11-21T07:41:39Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"While all backend-specific data is now contained in a backend-specific\nstructure, we still share the zlib stream across the loose and packed\nobjects.\n\nRefactor the code and move it into the specific structures so that we\nfully detangle the different backends from one another.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 104 ++++++++++++++++++++++++++++++------------------------------\n 1 file changed, 52 insertions(+), 52 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 45463b5c55..93fe72182a 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -25,23 +25,8 @@ struct odb_read_stream {\n \n \tenum object_type type;\n \tunsigned long size; /* inflated size of full object */\n-\tgit_zstream z;\n-\tenum { z_unused, z_used, z_done, z_error } z_state;\n };\n \n-/*****************************************************************\n- *\n- * Common helpers\n- *\n- *****************************************************************/\n-\n-static void close_deflated_stream(struct odb_read_stream *st)\n-{\n-\tif (st->z_state == z_used)\n-\t\tgit_inflate_end(&st->z);\n-}\n-\n-\n /*****************************************************************\n  *\n  * Filtered stream\n@@ -150,6 +135,12 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \n struct odb_loose_read_stream {\n \tstruct odb_read_stream base;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_LOOSE_READ_STREAM_INUSE,\n+\t\tODB_LOOSE_READ_STREAM_DONE,\n+\t\tODB_LOOSE_READ_STREAM_ERROR,\n+\t} z_state;\n \tvoid *mapped;\n \tunsigned long mapsize;\n \tchar hdr[32];\n@@ -162,10 +153,10 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n \tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->base.z_state) {\n-\tcase z_done:\n+\tswitch (st->z_state) {\n+\tcase ODB_LOOSE_READ_STREAM_DONE:\n \t\treturn 0;\n-\tcase z_error:\n+\tcase ODB_LOOSE_READ_STREAM_ERROR:\n \t\treturn -1;\n \tdefault:\n \t\tbreak;\n@@ -183,20 +174,20 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n \twhile (total_read < sz) {\n \t\tint status;\n \n-\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->base.z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n \n-\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_done;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n \t\t\tbreak;\n \t\t}\n \t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_error;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n \t\t\treturn -1;\n \t\t}\n \t}\n@@ -206,7 +197,8 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n static int close_istream_loose(struct odb_read_stream *_st)\n {\n \tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tclose_deflated_stream(&st->base);\n+\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n \tmunmap(st->mapped, st->mapsize);\n \treturn 0;\n }\n@@ -238,7 +230,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t */\n \tCALLOC_ARRAY(st, 1);\n \n-\tswitch (unpack_loose_header(&st->base.z, mapped, mapsize, st->hdr,\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n \t\t\t\t    sizeof(st->hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n@@ -256,8 +248,8 @@ static int open_istream_loose(struct odb_read_stream **out,\n \tst->mapped = mapped;\n \tst->mapsize = mapsize;\n \tst->hdr_used = strlen(st->hdr) + 1;\n-\tst->hdr_avail = st->base.z.total_out;\n-\tst->base.z_state = z_used;\n+\tst->hdr_avail = st->z.total_out;\n+\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n \tst->base.close = close_istream_loose;\n \tst->base.read = read_istream_loose;\n \n@@ -265,7 +257,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \n \treturn 0;\n error:\n-\tgit_inflate_end(&st->base.z);\n+\tgit_inflate_end(&st->z);\n \tmunmap(st->mapped, st->mapsize);\n \tfree(st);\n \treturn -1;\n@@ -281,6 +273,13 @@ static int open_istream_loose(struct odb_read_stream **out,\n struct odb_packed_read_stream {\n \tstruct odb_read_stream base;\n \tstruct packed_git *pack;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n+\t\tODB_PACKED_READ_STREAM_INUSE,\n+\t\tODB_PACKED_READ_STREAM_DONE,\n+\t\tODB_PACKED_READ_STREAM_ERROR,\n+\t} z_state;\n \toff_t pos;\n };\n \n@@ -290,17 +289,17 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->base.z_state) {\n-\tcase z_unused:\n-\t\tmemset(&st->base.z, 0, sizeof(st->base.z));\n-\t\tgit_inflate_init(&st->base.z);\n-\t\tst->base.z_state = z_used;\n+\tswitch (st->z_state) {\n+\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n+\t\tmemset(&st->z, 0, sizeof(st->z));\n+\t\tgit_inflate_init(&st->z);\n+\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n \t\tbreak;\n-\tcase z_done:\n+\tcase ODB_PACKED_READ_STREAM_DONE:\n \t\treturn 0;\n-\tcase z_error:\n+\tcase ODB_PACKED_READ_STREAM_ERROR:\n \t\treturn -1;\n-\tcase z_used:\n+\tcase ODB_PACKED_READ_STREAM_INUSE:\n \t\tbreak;\n \t}\n \n@@ -310,20 +309,20 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \t\tunsigned char *mapped;\n \n \t\tmapped = use_pack(st->pack, &window,\n-\t\t\t\t  st->pos, &st->base.z.avail_in);\n+\t\t\t\t  st->pos, &st->z.avail_in);\n \n-\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->base.z.avail_out = sz - total_read;\n-\t\tst->base.z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tst->z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n \n-\t\tst->pos += st->base.z.next_in - mapped;\n-\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n+\t\tst->pos += st->z.next_in - mapped;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n \t\tunuse_pack(&window);\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_done;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n \t\t\tbreak;\n \t\t}\n \n@@ -336,8 +335,8 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \t\t * or truncated), then use_pack() catches that and will die().\n \t\t */\n \t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_error;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n \t\t\treturn -1;\n \t\t}\n \t}\n@@ -347,7 +346,8 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n {\n \tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tclose_deflated_stream(&st->base);\n+\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n \treturn 0;\n }\n \n@@ -384,7 +384,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tstream->base.read = read_istream_pack_non_delta;\n \tstream->base.type = in_pack_type;\n \tstream->base.size = size;\n-\tstream->base.z_state = z_unused;\n+\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n \tstream->pack = pack;\n \tstream->pos = offset;\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531118","messageId":"20251121-b4-pks-odb-read-stream-v2-11-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 11/19] packfile: introduce function to read object info from a store","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:56Z","receivedAt":"2025-11-21T07:41:43Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Extract the logic to read object info for a packed object from\n`do_oid_object_into_extended()` into a standalone function that operates\non the packfile store. This function will be used in a subsequent\ncommit.\n\nNote that this change allows us to make `find_pack_entry()` an internal\nimplementation detail. As a consequence though we have to move around\n`packfile_store_freshen_object()` so that it is defined after that\nfunction.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c      | 29 ++++---------------------\n packfile.c | 71 +++++++++++++++++++++++++++++++++++++++++++++++---------------\n packfile.h | 12 ++++++++++-\n 3 files changed, 69 insertions(+), 43 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 3ec21ef24e..f4cbee4b04 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -666,8 +666,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n {\n \tstatic struct object_info blank_oi = OBJECT_INFO_INIT;\n \tconst struct cached_object *co;\n-\tstruct pack_entry e;\n-\tint rtype;\n \tconst struct object_id *real = oid;\n \tint already_retried = 0;\n \n@@ -702,8 +700,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \twhile (1) {\n \t\tstruct odb_source *source;\n \n-\t\tif (find_pack_entry(odb->repo, real, &e))\n-\t\t\tbreak;\n+\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n+\t\t\treturn 0;\n \n \t\t/* Most likely it's a loose object. */\n \t\tfor (source = odb->sources; source; source = source->next)\n@@ -713,8 +711,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t/* Not a loose object; someone else may have just packed it. */\n \t\tif (!(flags & OBJECT_INFO_QUICK)) {\n \t\t\todb_reprepare(odb->repo->objects);\n-\t\t\tif (find_pack_entry(odb->repo, real, &e))\n-\t\t\t\tbreak;\n+\t\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n+\t\t\t\treturn 0;\n \t\t}\n \n \t\t/*\n@@ -747,25 +745,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t}\n \t\treturn -1;\n \t}\n-\n-\tif (oi == &blank_oi)\n-\t\t/*\n-\t\t * We know that the caller doesn't actually need the\n-\t\t * information below, so return early.\n-\t\t */\n-\t\treturn 0;\n-\trtype = packed_object_info(odb->repo, e.p, e.offset, oi);\n-\tif (rtype < 0) {\n-\t\tmark_bad_packed_object(e.p, real);\n-\t\treturn do_oid_object_info_extended(odb, real, oi, 0);\n-\t} else if (oi->whence == OI_PACKED) {\n-\t\toi->u.packed.offset = e.offset;\n-\t\toi->u.packed.pack = e.p;\n-\t\toi->u.packed.is_delta = (rtype == OBJ_REF_DELTA ||\n-\t\t\t\t\t rtype == OBJ_OFS_DELTA);\n-\t}\n-\n-\treturn 0;\n }\n \n static int oid_object_info_convert(struct repository *r,\ndiff --git a/packfile.c b/packfile.c\nindex 40f733dd23..b4bc40d895 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -819,22 +819,6 @@ struct packed_git *packfile_store_load_pack(struct packfile_store *store,\n \treturn p;\n }\n \n-int packfile_store_freshen_object(struct packfile_store *store,\n-\t\t\t\t  const struct object_id *oid)\n-{\n-\tstruct pack_entry e;\n-\tif (!find_pack_entry(store->odb->repo, oid, &e))\n-\t\treturn 0;\n-\tif (e.p->is_cruft)\n-\t\treturn 0;\n-\tif (e.p->freshened)\n-\t\treturn 1;\n-\tif (utime(e.p->pack_name, NULL))\n-\t\treturn 0;\n-\te.p->freshened = 1;\n-\treturn 1;\n-}\n-\n void (*report_garbage)(unsigned seen_bits, const char *path);\n \n static void report_helper(const struct string_list *list,\n@@ -2064,7 +2048,9 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static int find_pack_entry(struct repository *r,\n+\t\t\t   const struct object_id *oid,\n+\t\t\t   struct pack_entry *e)\n {\n \tstruct list_head *pos;\n \n@@ -2087,6 +2073,57 @@ int find_pack_entry(struct repository *r, const struct object_id *oid, struct pa\n \treturn 0;\n }\n \n+int packfile_store_freshen_object(struct packfile_store *store,\n+\t\t\t\t  const struct object_id *oid)\n+{\n+\tstruct pack_entry e;\n+\tif (!find_pack_entry(store->odb->repo, oid, &e))\n+\t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n+\tif (e.p->freshened)\n+\t\treturn 1;\n+\tif (utime(e.p->pack_name, NULL))\n+\t\treturn 0;\n+\te.p->freshened = 1;\n+\treturn 1;\n+}\n+\n+int packfile_store_read_object_info(struct packfile_store *store,\n+\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t    struct object_info *oi,\n+\t\t\t\t    unsigned flags UNUSED)\n+{\n+\tstatic struct object_info blank_oi = OBJECT_INFO_INIT;\n+\tstruct pack_entry e;\n+\tint rtype;\n+\n+\tif (!find_pack_entry(store->odb->repo, oid, &e))\n+\t\treturn 1;\n+\n+\t/*\n+\t * We know that the caller doesn't actually need the\n+\t * information below, so return early.\n+\t */\n+\tif (oi == &blank_oi)\n+\t\treturn 0;\n+\n+\trtype = packed_object_info(store->odb->repo, e.p, e.offset, oi);\n+\tif (rtype < 0) {\n+\t\tmark_bad_packed_object(e.p, oid);\n+\t\treturn -1;\n+\t}\n+\n+\tif (oi->whence == OI_PACKED) {\n+\t\toi->u.packed.offset = e.offset;\n+\t\toi->u.packed.pack = e.p;\n+\t\toi->u.packed.is_delta = (rtype == OBJ_REF_DELTA ||\n+\t\t\t\t\t rtype == OBJ_OFS_DELTA);\n+\t}\n+\n+\treturn 0;\n+}\n+\n static void maybe_invalidate_kept_pack_cache(struct repository *r,\n \t\t\t\t\t     unsigned flags)\n {\ndiff --git a/packfile.h b/packfile.h\nindex 58fcc88e20..0a98bddd81 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -144,6 +144,17 @@ void packfile_store_add_pack(struct packfile_store *store,\n #define repo_for_each_pack(repo, p) \\\n \tfor (p = packfile_store_get_packs(repo->objects->packfiles); p; p = p->next)\n \n+/*\n+ * Try to read the object identified by its ID from the object store and\n+ * populate the object info with its data. Returns 1 in case the object was\n+ * not found, 0 if it was and read successfully, and a negative error code in\n+ * case the object was corrupted.\n+ */\n+int packfile_store_read_object_info(struct packfile_store *store,\n+\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t    struct object_info *oi,\n+\t\t\t\t    unsigned flags);\n+\n /*\n  * Get all packs managed by the given store, including packfiles that are\n  * referenced by multi-pack indices.\n@@ -357,7 +368,6 @@ const struct packed_git *has_packed_and_bad(struct repository *, const struct ob\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\n  */\n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e);\n int find_kept_pack_entry(struct repository *r, const struct object_id *oid, unsigned flags, struct pack_entry *e);\n \n int has_object_pack(struct repository *r, const struct object_id *oid);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531119","messageId":"20251121-b4-pks-odb-read-stream-v2-12-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 12/19] streaming: rely on object sources to create object stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:57Z","receivedAt":"2025-11-21T07:41:47Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating an object stream we first look up the object info and, if\nit's present, we call into the respective backend that contains the\nobject to create a new stream for it.\n\nThis has the consequence that, for loose object source, we basically\niterate through the object sources twice: we first discover that the\nfile exists as a loose object in the first place by iterating through\nall sources. And, once we have discovered it, we again walk through all\nsources to try and map the object. The same issue will eventually also\nsurface once the packfile store becomes per-object-source.\n\nFurthermore, it feels rather pointless to first look up the object only\nto then try and read it.\n\nRefactor the logic to be centered around sources instead. Instead of\nfirst reading the object, we immediately ask the source to create the\nobject stream for us. If the object exists we get stream, otherwise\nwe'll try the next source.\n\nLike this we only have to iterate through sources once. But even more\nimportantly, this change also helps us to make the whole logic\npluggable. The object read stream subsystem does not need to be aware of\nthe different source backends anymore, but eventually it'll only have to\ncall the source's callback function.\n\nNote that at the current point in time we aren't fully there yet:\n\n  - The packfile store still sits on the object database level and is\n    thus agnostic of the sources.\n\n  - We still have to call into both the packfile store and the loose\n    object source.\n\nBut both of these issues will soon be addressed.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 65 +++++++++++++++++++++++--------------------------------------\n 1 file changed, 24 insertions(+), 41 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 93fe72182a..fc7d88e313 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -204,21 +204,15 @@ static int close_istream_loose(struct odb_read_stream *_st)\n }\n \n static int open_istream_loose(struct odb_read_stream **out,\n-\t\t\t      struct repository *r,\n+\t\t\t      struct odb_source *source,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_loose_read_stream *st;\n-\tstruct odb_source *source;\n \tunsigned long mapsize;\n \tvoid *mapped;\n \n-\todb_prepare_alternates(r->objects);\n-\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n-\t\tif (mapped)\n-\t\t\tbreak;\n-\t}\n+\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n \tif (!mapped)\n \t\treturn -1;\n \n@@ -352,21 +346,25 @@ static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n }\n \n static int open_istream_pack_non_delta(struct odb_read_stream **out,\n-\t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED,\n-\t\t\t\t       struct packed_git *pack,\n-\t\t\t\t       off_t offset)\n+\t\t\t\t       struct object_database *odb,\n+\t\t\t\t       const struct object_id *oid)\n {\n \tstruct odb_packed_read_stream *stream;\n-\tstruct pack_window *window;\n+\tstruct pack_window *window = NULL;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n \tenum object_type in_pack_type;\n-\tsize_t size;\n+\tunsigned long size;\n \n-\twindow = NULL;\n+\toi.sizep = &size;\n+\n+\tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n+\t    oi.u.packed.is_delta ||\n+\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\treturn -1;\n \n-\tin_pack_type = unpack_object_header(pack,\n+\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n \t\t\t\t\t    &window,\n-\t\t\t\t\t    &offset,\n+\t\t\t\t\t    &oi.u.packed.offset,\n \t\t\t\t\t    &size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n@@ -385,8 +383,8 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tstream->base.type = in_pack_type;\n \tstream->base.size = size;\n \tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n-\tstream->pack = pack;\n-\tstream->pos = offset;\n+\tstream->pack = oi.u.packed.pack;\n+\tstream->pos = oi.u.packed.offset;\n \n \t*out = &stream->base;\n \n@@ -463,30 +461,15 @@ static int istream_source(struct odb_read_stream **out,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid)\n {\n-\tunsigned long size;\n-\tint status;\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\n-\toi.sizep = &size;\n-\tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n-\tif (status < 0)\n-\t\treturn status;\n+\tstruct odb_source *source;\n \n-\tswitch (oi.whence) {\n-\tcase OI_LOOSE:\n-\t\tif (open_istream_loose(out, r, oid) < 0)\n-\t\t\tbreak;\n-\t\treturn 0;\n-\tcase OI_PACKED:\n-\t\tif (oi.u.packed.is_delta ||\n-\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n-\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n-\t\t\t\t\t\toi.u.packed.offset) < 0)\n-\t\t\tbreak;\n+\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n \t\treturn 0;\n-\tdefault:\n-\t\tbreak;\n-\t}\n+\n+\todb_prepare_alternates(r->objects);\n+\tfor (source = r->objects->sources; source; source = source->next)\n+\t\tif (!open_istream_loose(out, source, oid))\n+\t\t\treturn 0;\n \n \treturn open_istream_incore(out, r, oid);\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531120","messageId":"20251121-b4-pks-odb-read-stream-v2-13-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 13/19] streaming: get rid of `the_repository`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:58Z","receivedAt":"2025-11-21T07:41:50Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Subsequent commits will move the backend-specific logic of object\nstreaming into their respective subsystems. These subsystems have gotten\nrid of `the_repository` already, but we still use it in two locations in\nthe streaming subsystem.\n\nPrepare for the move by fixing those two cases. Converting the logic in\n`open_istream_pack_non_delta()` is trivial as we already got the object\ndatabase as input.\n\nBut for `stream_blob_to_fd()` we have to add a new parameter to make it\naccessible. So, as we already have to adjust all callers anyway, rename\nthe function to `odb_stream_blob_to_fd()` to indicate it's part of the\nobject subsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/cat-file.c  |  2 +-\n builtin/fsck.c      |  3 ++-\n builtin/log.c       |  4 ++--\n entry.c             |  2 +-\n parallel-checkout.c |  3 ++-\n streaming.c         | 13 +++++++------\n streaming.h         | 18 +++++++++++++++++-\n 7 files changed, 32 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 983ecec837..120d626d66 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -95,7 +95,7 @@ static int filter_object(const char *path, unsigned mode,\n \n static int stream_blob(const struct object_id *oid)\n {\n-\tif (stream_blob_to_fd(1, oid, NULL, 0))\n+\tif (odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0))\n \t\tdie(\"unable to stream %s to stdout\", oid_to_hex(oid));\n \treturn 0;\n }\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex b1a650c673..1a348d43c2 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -340,7 +340,8 @@ static void check_unreachable_object(struct object *obj)\n \t\t\t}\n \t\t\tf = xfopen(filename, \"w\");\n \t\t\tif (obj->type == OBJ_BLOB) {\n-\t\t\t\tif (stream_blob_to_fd(fileno(f), &obj->oid, NULL, 1))\n+\t\t\t\tif (odb_stream_blob_to_fd(the_repository->objects, fileno(f),\n+\t\t\t\t\t\t\t  &obj->oid, NULL, 1))\n \t\t\t\t\tdie_errno(_(\"could not write '%s'\"), filename);\n \t\t\t} else\n \t\t\t\tfprintf(f, \"%s\\n\", describe_object(&obj->oid));\ndiff --git a/builtin/log.c b/builtin/log.c\nindex c8319b8af3..e7b83a6e00 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -584,7 +584,7 @@ static int show_blob_object(const struct object_id *oid, struct rev_info *rev, c\n \tfflush(rev->diffopt.file);\n \tif (!rev->diffopt.flags.textconv_set_via_cmdline ||\n \t    !rev->diffopt.flags.allow_textconv)\n-\t\treturn stream_blob_to_fd(1, oid, NULL, 0);\n+\t\treturn odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0);\n \n \tif (get_oid_with_context(the_repository, obj_name,\n \t\t\t\t GET_OID_RECORD_PATH,\n@@ -594,7 +594,7 @@ static int show_blob_object(const struct object_id *oid, struct rev_info *rev, c\n \t    !textconv_object(the_repository, obj_context.path,\n \t\t\t     obj_context.mode, &oidc, 1, &buf, &size)) {\n \t\tobject_context_release(&obj_context);\n-\t\treturn stream_blob_to_fd(1, oid, NULL, 0);\n+\t\treturn odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0);\n \t}\n \n \tif (!buf)\ndiff --git a/entry.c b/entry.c\nindex cae02eb503..38dfe670f7 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -139,7 +139,7 @@ static int streaming_write_entry(const struct cache_entry *ce, char *path,\n \tif (fd < 0)\n \t\treturn -1;\n \n-\tresult |= stream_blob_to_fd(fd, &ce->oid, filter, 1);\n+\tresult |= odb_stream_blob_to_fd(the_repository->objects, fd, &ce->oid, filter, 1);\n \t*fstat_done = fstat_checkout_output(fd, state, statbuf);\n \tresult |= close(fd);\n \ndiff --git a/parallel-checkout.c b/parallel-checkout.c\nindex fba6aa65a6..1cb6701b92 100644\n--- a/parallel-checkout.c\n+++ b/parallel-checkout.c\n@@ -281,7 +281,8 @@ static int write_pc_item_to_fd(struct parallel_checkout_item *pc_item, int fd,\n \n \tfilter = get_stream_filter_ca(&pc_item->ca, &pc_item->ce->oid);\n \tif (filter) {\n-\t\tif (stream_blob_to_fd(fd, &pc_item->ce->oid, filter, 1)) {\n+\t\tif (odb_stream_blob_to_fd(the_repository->objects, fd,\n+\t\t\t\t\t  &pc_item->ce->oid, filter, 1)) {\n \t\t\t/* On error, reset fd to try writing without streaming */\n \t\t\tif (reset_fd(fd, path))\n \t\t\t\treturn -1;\ndiff --git a/streaming.c b/streaming.c\nindex fc7d88e313..41c2070941 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -2,8 +2,6 @@\n  * Copyright (c) 2011, Google Inc.\n  */\n \n-#define USE_THE_REPOSITORY_VARIABLE\n-\n #include \"git-compat-util.h\"\n #include \"convert.h\"\n #include \"environment.h\"\n@@ -359,7 +357,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \n \tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n \t    oi.u.packed.is_delta ||\n-\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t    repo_settings_get_big_file_threshold(odb->repo) >= size)\n \t\treturn -1;\n \n \tin_pack_type = unpack_object_header(oi.u.packed.pack,\n@@ -519,8 +517,11 @@ struct odb_read_stream *open_istream(struct repository *r,\n \treturn st;\n }\n \n-int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter *filter,\n-\t\t      int can_seek)\n+int odb_stream_blob_to_fd(struct object_database *odb,\n+\t\t\t  int fd,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  struct stream_filter *filter,\n+\t\t\t  int can_seek)\n {\n \tstruct odb_read_stream *st;\n \tenum object_type type;\n@@ -528,7 +529,7 @@ int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = open_istream(the_repository, oid, &type, &sz, filter);\n+\tst = open_istream(odb->repo, oid, &type, &sz, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\ndiff --git a/streaming.h b/streaming.h\nindex f5ff5d7ac9..1a3de6812e 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -7,6 +7,7 @@\n #include \"object.h\"\n \n /* opaque */\n+struct object_database;\n struct odb_read_stream;\n struct stream_filter;\n \n@@ -16,6 +17,21 @@ struct odb_read_stream *open_istream(struct repository *, const struct object_id\n int close_istream(struct odb_read_stream *);\n ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n \n-int stream_blob_to_fd(int fd, const struct object_id *, struct stream_filter *, int can_seek);\n+/*\n+ * Look up the object by its ID and write the full contents to the file\n+ * descriptor. The object must be a blob, or the function will fail. When\n+ * provided, the filter is used to transform the blob contents.\n+ *\n+ * `can_seek` should be set to 1 in case the given file descriptor can be\n+ * seek(3p)'d on. This is used to support files with holes in case a\n+ * significant portion of the blob contains NUL bytes.\n+ *\n+ * Returns a negative error code on failure, 0 on success.\n+ */\n+int odb_stream_blob_to_fd(struct object_database *odb,\n+\t\t\t  int fd,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  struct stream_filter *filter,\n+\t\t\t  int can_seek);\n \n #endif /* STREAMING_H */\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531121","messageId":"20251121-b4-pks-odb-read-stream-v2-14-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 14/19] streaming: make the `odb_read_stream` definition public","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:40:59Z","receivedAt":"2025-11-21T07:41:55Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Subsequent commits will move the backend-specific logic of setting up an\nobject read stream into the specific subsystems. As the backends are now\nthe ones that are responsible for allocating the stream they'll need to\nhave the stream definition available to them.\n\nMake the stream definition public to prepare for this.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 11 -----------\n streaming.h | 15 ++++++++++++++-\n 2 files changed, 14 insertions(+), 12 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 41c2070941..586c20eac6 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -12,19 +12,8 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*close_istream_fn)(struct odb_read_stream *);\n-typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n-\n #define FILTER_BUFFER (1024*16)\n \n-struct odb_read_stream {\n-\tclose_istream_fn close;\n-\tread_istream_fn read;\n-\n-\tenum object_type type;\n-\tunsigned long size; /* inflated size of full object */\n-};\n-\n /*****************************************************************\n  *\n  * Filtered stream\ndiff --git a/streaming.h b/streaming.h\nindex 1a3de6812e..acfdef1598 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -6,11 +6,24 @@\n \n #include \"object.h\"\n \n-/* opaque */\n struct object_database;\n struct odb_read_stream;\n struct stream_filter;\n \n+typedef int (*odb_read_stream_close_fn)(struct odb_read_stream *);\n+typedef ssize_t (*odb_read_stream_read_fn)(struct odb_read_stream *, char *, size_t);\n+\n+/*\n+ * A stream that can be used to read an object from the object database without\n+ * loading all of it into memory.\n+ */\n+struct odb_read_stream {\n+\todb_read_stream_close_fn close;\n+\todb_read_stream_read_fn read;\n+\tenum object_type type;\n+\tunsigned long size; /* inflated size of full object */\n+};\n+\n struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n \t\t\t\t     enum object_type *, unsigned long *,\n \t\t\t\t     struct stream_filter *);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531122","messageId":"20251121-b4-pks-odb-read-stream-v2-15-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 15/19] streaming: move logic to read loose objects streams into backend","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:41:00Z","receivedAt":"2025-11-21T07:41:58Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Move the logic to read loose object streams into the respective\nsubsystem. This allows us to make a couple of function declarations\nprivate.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-file.c | 167 ++++++++++++++++++++++++++++++++++++++++++++++++++++++----\n object-file.h |  42 ++-------------\n streaming.c   | 133 +---------------------------------------------\n 3 files changed, 164 insertions(+), 178 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b62b21a452..8c67847fea 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -234,9 +234,9 @@ static void *map_fd(int fd, const char *path, unsigned long *size)\n \treturn map;\n }\n \n-void *odb_source_loose_map_object(struct odb_source *source,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  unsigned long *size)\n+static void *odb_source_loose_map_object(struct odb_source *source,\n+\t\t\t\t\t const struct object_id *oid,\n+\t\t\t\t\t unsigned long *size)\n {\n \tconst char *p;\n \tint fd = open_loose_object(source->loose, oid, &p);\n@@ -246,11 +246,29 @@ void *odb_source_loose_map_object(struct odb_source *source,\n \treturn map_fd(fd, p, size);\n }\n \n-enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n-\t\t\t\t\t\t    unsigned char *map,\n-\t\t\t\t\t\t    unsigned long mapsize,\n-\t\t\t\t\t\t    void *buffer,\n-\t\t\t\t\t\t    unsigned long bufsiz)\n+enum unpack_loose_header_result {\n+\tULHR_OK,\n+\tULHR_BAD,\n+\tULHR_TOO_LONG,\n+};\n+\n+/**\n+ * unpack_loose_header() initializes the data stream needed to unpack\n+ * a loose object header.\n+ *\n+ * Returns:\n+ *\n+ * - ULHR_OK on success\n+ * - ULHR_BAD on error\n+ * - ULHR_TOO_LONG if the header was too long\n+ *\n+ * It will only parse up to MAX_HEADER_LEN bytes.\n+ */\n+static enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n+\t\t\t\t\t\t\t   unsigned char *map,\n+\t\t\t\t\t\t\t   unsigned long mapsize,\n+\t\t\t\t\t\t\t   void *buffer,\n+\t\t\t\t\t\t\t   unsigned long bufsiz)\n {\n \tint status;\n \n@@ -329,11 +347,18 @@ static void *unpack_loose_rest(git_zstream *stream,\n }\n \n /*\n+ * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n+ * object. If it doesn't follow that format -1 is returned. To check\n+ * the validity of the <type> populate the \"typep\" in the \"struct\n+ * object_info\". It will be OBJ_BAD if the object type is unknown. The\n+ * parsed <len> can be retrieved via \"oi->sizep\", and from there\n+ * passed to unpack_loose_rest().\n+ *\n  * We used to just use \"sscanf()\", but that's actually way\n  * too permissive for what we want to check. So do an anal\n  * object header parse by hand.\n  */\n-int parse_loose_header(const char *hdr, struct object_info *oi)\n+static int parse_loose_header(const char *hdr, struct object_info *oi)\n {\n \tconst char *type_buf = hdr;\n \tsize_t size;\n@@ -1976,3 +2001,127 @@ void odb_source_loose_free(struct odb_source_loose *loose)\n \tloose_object_map_clear(&loose->map);\n \tfree(loose);\n }\n+\n+struct odb_loose_read_stream {\n+\tstruct odb_read_stream base;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_LOOSE_READ_STREAM_INUSE,\n+\t\tODB_LOOSE_READ_STREAM_DONE,\n+\t\tODB_LOOSE_READ_STREAM_ERROR,\n+\t} z_state;\n+\tvoid *mapped;\n+\tunsigned long mapsize;\n+\tchar hdr[32];\n+\tint hdr_avail;\n+\tint hdr_used;\n+};\n+\n+static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n+{\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tsize_t total_read = 0;\n+\n+\tswitch (st->z_state) {\n+\tcase ODB_LOOSE_READ_STREAM_DONE:\n+\t\treturn 0;\n+\tcase ODB_LOOSE_READ_STREAM_ERROR:\n+\t\treturn -1;\n+\tdefault:\n+\t\tbreak;\n+\t}\n+\n+\tif (st->hdr_used < st->hdr_avail) {\n+\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n+\t\tif (sz < to_copy)\n+\t\t\tto_copy = sz;\n+\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n+\t\tst->hdr_used += to_copy;\n+\t\ttotal_read += to_copy;\n+\t}\n+\n+\twhile (total_read < sz) {\n+\t\tint status;\n+\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\n+\t\tif (status == Z_STREAM_END) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n+\t\t\tbreak;\n+\t\t}\n+\t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn total_read;\n+}\n+\n+static int close_istream_loose(struct odb_read_stream *_st)\n+{\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n+\tmunmap(st->mapped, st->mapsize);\n+\treturn 0;\n+}\n+\n+int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\tstruct odb_source *source,\n+\t\t\t\t\tconst struct object_id *oid)\n+{\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_loose_read_stream *st;\n+\tunsigned long mapsize;\n+\tvoid *mapped;\n+\n+\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n+\tif (!mapped)\n+\t\treturn -1;\n+\n+\t/*\n+\t * Note: we must allocate this structure early even though we may still\n+\t * fail. This is because we need to initialize the zlib stream, and it\n+\t * is not possible to copy the stream around after the fact because it\n+\t * has self-referencing pointers.\n+\t */\n+\tCALLOC_ARRAY(st, 1);\n+\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n+\t\t\t\t    sizeof(st->hdr))) {\n+\tcase ULHR_OK:\n+\t\tbreak;\n+\tcase ULHR_BAD:\n+\tcase ULHR_TOO_LONG:\n+\t\tgoto error;\n+\t}\n+\n+\toi.sizep = &st->base.size;\n+\toi.typep = &st->base.type;\n+\n+\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n+\t\tgoto error;\n+\n+\tst->mapped = mapped;\n+\tst->mapsize = mapsize;\n+\tst->hdr_used = strlen(st->hdr) + 1;\n+\tst->hdr_avail = st->z.total_out;\n+\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n+\tst->base.close = close_istream_loose;\n+\tst->base.read = read_istream_loose;\n+\n+\t*out = &st->base;\n+\n+\treturn 0;\n+error:\n+\tgit_inflate_end(&st->z);\n+\tmunmap(st->mapped, st->mapsize);\n+\tfree(st);\n+\treturn -1;\n+}\ndiff --git a/object-file.h b/object-file.h\nindex eeffa67bbd..1229d5f675 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -16,6 +16,8 @@ enum {\n int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n \n+struct object_info;\n+struct odb_read_stream;\n struct odb_source;\n \n struct odb_source_loose {\n@@ -47,9 +49,9 @@ int odb_source_loose_read_object_info(struct odb_source *source,\n \t\t\t\t      const struct object_id *oid,\n \t\t\t\t      struct object_info *oi, int flags);\n \n-void *odb_source_loose_map_object(struct odb_source *source,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  unsigned long *size);\n+int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\tstruct odb_source *source,\n+\t\t\t\t\tconst struct object_id *oid);\n \n /*\n  * Return true iff an object database source has a loose object\n@@ -143,40 +145,6 @@ int for_each_loose_object(struct object_database *odb,\n int format_object_header(char *str, size_t size, enum object_type type,\n \t\t\t size_t objsize);\n \n-/**\n- * unpack_loose_header() initializes the data stream needed to unpack\n- * a loose object header.\n- *\n- * Returns:\n- *\n- * - ULHR_OK on success\n- * - ULHR_BAD on error\n- * - ULHR_TOO_LONG if the header was too long\n- *\n- * It will only parse up to MAX_HEADER_LEN bytes.\n- */\n-enum unpack_loose_header_result {\n-\tULHR_OK,\n-\tULHR_BAD,\n-\tULHR_TOO_LONG,\n-};\n-enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n-\t\t\t\t\t\t    unsigned char *map,\n-\t\t\t\t\t\t    unsigned long mapsize,\n-\t\t\t\t\t\t    void *buffer,\n-\t\t\t\t\t\t    unsigned long bufsiz);\n-\n-/**\n- * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n- * object. If it doesn't follow that format -1 is returned. To check\n- * the validity of the <type> populate the \"typep\" in the \"struct\n- * object_info\". It will be OBJ_BAD if the object type is unknown. The\n- * parsed <len> can be retrieved via \"oi->sizep\", and from there\n- * passed to unpack_loose_rest().\n- */\n-struct object_info;\n-int parse_loose_header(const char *hdr, struct object_info *oi);\n-\n int force_object_loose(struct odb_source *source,\n \t\t       const struct object_id *oid, time_t mtime);\n \ndiff --git a/streaming.c b/streaming.c\nindex 586c20eac6..cc67d56cd4 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -114,137 +114,6 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \treturn &fs->base;\n }\n \n-/*****************************************************************\n- *\n- * Loose object stream\n- *\n- *****************************************************************/\n-\n-struct odb_loose_read_stream {\n-\tstruct odb_read_stream base;\n-\tgit_zstream z;\n-\tenum {\n-\t\tODB_LOOSE_READ_STREAM_INUSE,\n-\t\tODB_LOOSE_READ_STREAM_DONE,\n-\t\tODB_LOOSE_READ_STREAM_ERROR,\n-\t} z_state;\n-\tvoid *mapped;\n-\tunsigned long mapsize;\n-\tchar hdr[32];\n-\tint hdr_avail;\n-\tint hdr_used;\n-};\n-\n-static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n-{\n-\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tsize_t total_read = 0;\n-\n-\tswitch (st->z_state) {\n-\tcase ODB_LOOSE_READ_STREAM_DONE:\n-\t\treturn 0;\n-\tcase ODB_LOOSE_READ_STREAM_ERROR:\n-\t\treturn -1;\n-\tdefault:\n-\t\tbreak;\n-\t}\n-\n-\tif (st->hdr_used < st->hdr_avail) {\n-\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n-\t\tif (sz < to_copy)\n-\t\t\tto_copy = sz;\n-\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n-\t\tst->hdr_used += to_copy;\n-\t\ttotal_read += to_copy;\n-\t}\n-\n-\twhile (total_read < sz) {\n-\t\tint status;\n-\n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n-\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n-\n-\t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n-\t\t\tbreak;\n-\t\t}\n-\t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n-\t\t\treturn -1;\n-\t\t}\n-\t}\n-\treturn total_read;\n-}\n-\n-static int close_istream_loose(struct odb_read_stream *_st)\n-{\n-\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n-\t\tgit_inflate_end(&st->z);\n-\tmunmap(st->mapped, st->mapsize);\n-\treturn 0;\n-}\n-\n-static int open_istream_loose(struct odb_read_stream **out,\n-\t\t\t      struct odb_source *source,\n-\t\t\t      const struct object_id *oid)\n-{\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_loose_read_stream *st;\n-\tunsigned long mapsize;\n-\tvoid *mapped;\n-\n-\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n-\tif (!mapped)\n-\t\treturn -1;\n-\n-\t/*\n-\t * Note: we must allocate this structure early even though we may still\n-\t * fail. This is because we need to initialize the zlib stream, and it\n-\t * is not possible to copy the stream around after the fact because it\n-\t * has self-referencing pointers.\n-\t */\n-\tCALLOC_ARRAY(st, 1);\n-\n-\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n-\t\t\t\t    sizeof(st->hdr))) {\n-\tcase ULHR_OK:\n-\t\tbreak;\n-\tcase ULHR_BAD:\n-\tcase ULHR_TOO_LONG:\n-\t\tgoto error;\n-\t}\n-\n-\toi.sizep = &st->base.size;\n-\toi.typep = &st->base.type;\n-\n-\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n-\t\tgoto error;\n-\n-\tst->mapped = mapped;\n-\tst->mapsize = mapsize;\n-\tst->hdr_used = strlen(st->hdr) + 1;\n-\tst->hdr_avail = st->z.total_out;\n-\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n-\tst->base.close = close_istream_loose;\n-\tst->base.read = read_istream_loose;\n-\n-\t*out = &st->base;\n-\n-\treturn 0;\n-error:\n-\tgit_inflate_end(&st->z);\n-\tmunmap(st->mapped, st->mapsize);\n-\tfree(st);\n-\treturn -1;\n-}\n-\n-\n /*****************************************************************\n  *\n  * Non-delta packed object stream\n@@ -455,7 +324,7 @@ static int istream_source(struct odb_read_stream **out,\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next)\n-\t\tif (!open_istream_loose(out, source, oid))\n+\t\tif (!odb_source_loose_read_object_stream(out, source, oid))\n \t\t\treturn 0;\n \n \treturn open_istream_incore(out, r, oid);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531123","messageId":"20251121-b4-pks-odb-read-stream-v2-16-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 16/19] streaming: move logic to read packed objects streams into backend","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:41:01Z","receivedAt":"2025-11-21T07:42:02Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Move the logic to read packed object streams into the respective\nsubsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n packfile.c  | 128 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n packfile.h  |   5 +++\n streaming.c | 136 +-----------------------------------------------------------\n 3 files changed, 134 insertions(+), 135 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex b4bc40d895..ad56ce0b90 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -20,6 +20,7 @@\n #include \"tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"streaming.h\"\n #include \"midx.h\"\n #include \"commit-graph.h\"\n #include \"pack-revindex.h\"\n@@ -2406,3 +2407,130 @@ void packfile_store_close(struct packfile_store *store)\n \t\tclose_pack(p);\n \t}\n }\n+\n+struct odb_packed_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct packed_git *pack;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n+\t\tODB_PACKED_READ_STREAM_INUSE,\n+\t\tODB_PACKED_READ_STREAM_DONE,\n+\t\tODB_PACKED_READ_STREAM_ERROR,\n+\t} z_state;\n+\toff_t pos;\n+};\n+\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n+\t\t\t\t\t   size_t sz)\n+{\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tsize_t total_read = 0;\n+\n+\tswitch (st->z_state) {\n+\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n+\t\tmemset(&st->z, 0, sizeof(st->z));\n+\t\tgit_inflate_init(&st->z);\n+\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n+\t\tbreak;\n+\tcase ODB_PACKED_READ_STREAM_DONE:\n+\t\treturn 0;\n+\tcase ODB_PACKED_READ_STREAM_ERROR:\n+\t\treturn -1;\n+\tcase ODB_PACKED_READ_STREAM_INUSE:\n+\t\tbreak;\n+\t}\n+\n+\twhile (total_read < sz) {\n+\t\tint status;\n+\t\tstruct pack_window *window = NULL;\n+\t\tunsigned char *mapped;\n+\n+\t\tmapped = use_pack(st->pack, &window,\n+\t\t\t\t  st->pos, &st->z.avail_in);\n+\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tst->z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\n+\t\tst->pos += st->z.next_in - mapped;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\tunuse_pack(&window);\n+\n+\t\tif (status == Z_STREAM_END) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\t/*\n+\t\t * Unlike the loose object case, we do not have to worry here\n+\t\t * about running out of input bytes and spinning infinitely. If\n+\t\t * we get Z_BUF_ERROR due to too few input bytes, then we'll\n+\t\t * replenish them in the next use_pack() call when we loop. If\n+\t\t * we truly hit the end of the pack (i.e., because it's corrupt\n+\t\t * or truncated), then use_pack() catches that and will die().\n+\t\t */\n+\t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn total_read;\n+}\n+\n+static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n+{\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n+\treturn 0;\n+}\n+\n+int packfile_store_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t      struct packfile_store *store,\n+\t\t\t\t      const struct object_id *oid)\n+{\n+\tstruct odb_packed_read_stream *stream;\n+\tstruct pack_window *window = NULL;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\tenum object_type in_pack_type;\n+\tunsigned long size;\n+\n+\toi.sizep = &size;\n+\n+\tif (packfile_store_read_object_info(store, oid, &oi, 0) ||\n+\t    oi.u.packed.is_delta ||\n+\t    repo_settings_get_big_file_threshold(store->odb->repo) >= size)\n+\t\treturn -1;\n+\n+\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n+\t\t\t\t\t    &window,\n+\t\t\t\t\t    &oi.u.packed.offset,\n+\t\t\t\t\t    &size);\n+\tunuse_pack(&window);\n+\tswitch (in_pack_type) {\n+\tdefault:\n+\t\treturn -1; /* we do not do deltas for now */\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_BLOB:\n+\tcase OBJ_TAG:\n+\t\tbreak;\n+\t}\n+\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.close = close_istream_pack_non_delta;\n+\tstream->base.read = read_istream_pack_non_delta;\n+\tstream->base.type = in_pack_type;\n+\tstream->base.size = size;\n+\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n+\tstream->pack = oi.u.packed.pack;\n+\tstream->pos = oi.u.packed.offset;\n+\n+\t*out = &stream->base;\n+\n+\treturn 0;\n+}\ndiff --git a/packfile.h b/packfile.h\nindex 0a98bddd81..3fcc5ae6e0 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -8,6 +8,7 @@\n \n /* in odb.h */\n struct object_info;\n+struct odb_read_stream;\n \n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n@@ -144,6 +145,10 @@ void packfile_store_add_pack(struct packfile_store *store,\n #define repo_for_each_pack(repo, p) \\\n \tfor (p = packfile_store_get_packs(repo->objects->packfiles); p; p = p->next)\n \n+int packfile_store_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t      struct packfile_store *store,\n+\t\t\t\t      const struct object_id *oid);\n+\n /*\n  * Try to read the object identified by its ID from the object store and\n  * populate the object info with its data. Returns 1 in case the object was\ndiff --git a/streaming.c b/streaming.c\nindex cc67d56cd4..3d80ddd757 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -114,140 +114,6 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \treturn &fs->base;\n }\n \n-/*****************************************************************\n- *\n- * Non-delta packed object stream\n- *\n- *****************************************************************/\n-\n-struct odb_packed_read_stream {\n-\tstruct odb_read_stream base;\n-\tstruct packed_git *pack;\n-\tgit_zstream z;\n-\tenum {\n-\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n-\t\tODB_PACKED_READ_STREAM_INUSE,\n-\t\tODB_PACKED_READ_STREAM_DONE,\n-\t\tODB_PACKED_READ_STREAM_ERROR,\n-\t} z_state;\n-\toff_t pos;\n-};\n-\n-static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n-\t\t\t\t\t   size_t sz)\n-{\n-\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tsize_t total_read = 0;\n-\n-\tswitch (st->z_state) {\n-\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n-\t\tmemset(&st->z, 0, sizeof(st->z));\n-\t\tgit_inflate_init(&st->z);\n-\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n-\t\tbreak;\n-\tcase ODB_PACKED_READ_STREAM_DONE:\n-\t\treturn 0;\n-\tcase ODB_PACKED_READ_STREAM_ERROR:\n-\t\treturn -1;\n-\tcase ODB_PACKED_READ_STREAM_INUSE:\n-\t\tbreak;\n-\t}\n-\n-\twhile (total_read < sz) {\n-\t\tint status;\n-\t\tstruct pack_window *window = NULL;\n-\t\tunsigned char *mapped;\n-\n-\t\tmapped = use_pack(st->pack, &window,\n-\t\t\t\t  st->pos, &st->z.avail_in);\n-\n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tst->z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n-\n-\t\tst->pos += st->z.next_in - mapped;\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n-\t\tunuse_pack(&window);\n-\n-\t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n-\t\t\tbreak;\n-\t\t}\n-\n-\t\t/*\n-\t\t * Unlike the loose object case, we do not have to worry here\n-\t\t * about running out of input bytes and spinning infinitely. If\n-\t\t * we get Z_BUF_ERROR due to too few input bytes, then we'll\n-\t\t * replenish them in the next use_pack() call when we loop. If\n-\t\t * we truly hit the end of the pack (i.e., because it's corrupt\n-\t\t * or truncated), then use_pack() catches that and will die().\n-\t\t */\n-\t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n-\t\t\treturn -1;\n-\t\t}\n-\t}\n-\treturn total_read;\n-}\n-\n-static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n-{\n-\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n-\t\tgit_inflate_end(&st->z);\n-\treturn 0;\n-}\n-\n-static int open_istream_pack_non_delta(struct odb_read_stream **out,\n-\t\t\t\t       struct object_database *odb,\n-\t\t\t\t       const struct object_id *oid)\n-{\n-\tstruct odb_packed_read_stream *stream;\n-\tstruct pack_window *window = NULL;\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tenum object_type in_pack_type;\n-\tunsigned long size;\n-\n-\toi.sizep = &size;\n-\n-\tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n-\t    oi.u.packed.is_delta ||\n-\t    repo_settings_get_big_file_threshold(odb->repo) >= size)\n-\t\treturn -1;\n-\n-\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n-\t\t\t\t\t    &window,\n-\t\t\t\t\t    &oi.u.packed.offset,\n-\t\t\t\t\t    &size);\n-\tunuse_pack(&window);\n-\tswitch (in_pack_type) {\n-\tdefault:\n-\t\treturn -1; /* we do not do deltas for now */\n-\tcase OBJ_COMMIT:\n-\tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n-\tcase OBJ_TAG:\n-\t\tbreak;\n-\t}\n-\n-\tCALLOC_ARRAY(stream, 1);\n-\tstream->base.close = close_istream_pack_non_delta;\n-\tstream->base.read = read_istream_pack_non_delta;\n-\tstream->base.type = in_pack_type;\n-\tstream->base.size = size;\n-\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n-\tstream->pack = oi.u.packed.pack;\n-\tstream->pos = oi.u.packed.offset;\n-\n-\t*out = &stream->base;\n-\n-\treturn 0;\n-}\n-\n-\n /*****************************************************************\n  *\n  * In-core stream\n@@ -319,7 +185,7 @@ static int istream_source(struct odb_read_stream **out,\n {\n \tstruct odb_source *source;\n \n-\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n+\tif (!packfile_store_read_object_stream(out, r->objects->packfiles, oid))\n \t\treturn 0;\n \n \todb_prepare_alternates(r->objects);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531124","messageId":"20251121-b4-pks-odb-read-stream-v2-17-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 17/19] streaming: refactor interface to be object-database-centric","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:41:02Z","receivedAt":"2025-11-21T07:42:06Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Refactor the streaming interface to be centered around object databases\ninstead of centered around the repository. Rename the functions\naccordingly.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  6 +++---\n archive-zip.c          | 12 ++++++------\n builtin/index-pack.c   |  8 ++++----\n builtin/pack-objects.c | 14 +++++++-------\n object-file.c          |  8 ++++----\n streaming.c            | 44 ++++++++++++++++++++++----------------------\n streaming.h            | 30 +++++++++++++++++++++++++-----\n 7 files changed, 71 insertions(+), 51 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex dc1eda09e0..4133e09ca1 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -135,16 +135,16 @@ static int stream_blocked(struct repository *r, const struct object_id *oid)\n \tchar buf[BLOCKSIZE];\n \tssize_t readlen;\n \n-\tst = open_istream(r, oid, &type, &sz, NULL);\n+\tst = odb_read_object_stream(r->objects, oid, &type, &sz, NULL);\n \tif (!st)\n \t\treturn error(_(\"cannot stream blob %s\"), oid_to_hex(oid));\n \tfor (;;) {\n-\t\treadlen = read_istream(st, buf, sizeof(buf));\n+\t\treadlen = odb_read_stream_read(st, buf, sizeof(buf));\n \t\tif (readlen <= 0)\n \t\t\tbreak;\n \t\tdo_write_blocked(buf, readlen);\n \t}\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \tif (!readlen)\n \t\tfinish_record();\n \treturn readlen;\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 40a9c93ff9..ff57f4f884 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -348,8 +348,8 @@ static int write_zip_entry(struct archiver_args *args,\n \n \t\tif (!buffer) {\n \t\t\tenum object_type type;\n-\t\t\tstream = open_istream(args->repo, oid, &type, &size,\n-\t\t\t\t\t      NULL);\n+\t\t\tstream = odb_read_object_stream(args->repo->objects, oid,\n+\t\t\t\t\t\t\t&type, &size, NULL);\n \t\t\tif (!stream)\n \t\t\t\treturn error(_(\"cannot stream blob %s\"),\n \t\t\t\t\t     oid_to_hex(oid));\n@@ -429,7 +429,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tssize_t readlen;\n \n \t\tfor (;;) {\n-\t\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\t\t\treadlen = odb_read_stream_read(stream, buf, sizeof(buf));\n \t\t\tif (readlen <= 0)\n \t\t\t\tbreak;\n \t\t\tcrc = crc32(crc, buf, readlen);\n@@ -439,7 +439,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\t\t\t\t\t    buf, readlen);\n \t\t\twrite_or_die(1, buf, readlen);\n \t\t}\n-\t\tclose_istream(stream);\n+\t\todb_read_stream_close(stream);\n \t\tif (readlen)\n \t\t\treturn readlen;\n \n@@ -462,7 +462,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tzstream.avail_out = sizeof(compressed);\n \n \t\tfor (;;) {\n-\t\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\t\t\treadlen = odb_read_stream_read(stream, buf, sizeof(buf));\n \t\t\tif (readlen <= 0)\n \t\t\t\tbreak;\n \t\t\tcrc = crc32(crc, buf, readlen);\n@@ -486,7 +486,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\t}\n \n \t\t}\n-\t\tclose_istream(stream);\n+\t\todb_read_stream_close(stream);\n \t\tif (readlen)\n \t\t\treturn readlen;\n \ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 5f90f12f92..67221dbe6a 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -779,7 +779,7 @@ static int compare_objects(const unsigned char *buf, unsigned long size,\n \t}\n \n \twhile (size) {\n-\t\tssize_t len = read_istream(data->st, data->buf, size);\n+\t\tssize_t len = odb_read_stream_read(data->st, data->buf, size);\n \t\tif (len == 0)\n \t\t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t\t    oid_to_hex(&data->entry->idx.oid));\n@@ -807,15 +807,15 @@ static int check_collison(struct object_entry *entry)\n \n \tmemset(&data, 0, sizeof(data));\n \tdata.entry = entry;\n-\tdata.st = open_istream(the_repository, &entry->idx.oid, &type, &size,\n-\t\t\t       NULL);\n+\tdata.st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n+\t\t\t\t\t &type, &size, NULL);\n \tif (!data.st)\n \t\treturn -1;\n \tif (size != entry->size || type != entry->type)\n \t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t    oid_to_hex(&entry->idx.oid));\n \tunpack_data(entry, compare_objects, &data);\n-\tclose_istream(data.st);\n+\todb_read_stream_close(data.st);\n \tfree(data.buf);\n \treturn 0;\n }\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex c693d948e1..adf267c59d 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -417,7 +417,7 @@ static unsigned long write_large_blob_data(struct odb_read_stream *st, struct ha\n \tfor (;;) {\n \t\tssize_t readlen;\n \t\tint zret = Z_OK;\n-\t\treadlen = read_istream(st, ibuf, sizeof(ibuf));\n+\t\treadlen = odb_read_stream_read(st, ibuf, sizeof(ibuf));\n \t\tif (readlen == -1)\n \t\t\tdie(_(\"unable to read %s\"), oid_to_hex(oid));\n \n@@ -520,8 +520,8 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\tif (oe_type(entry) == OBJ_BLOB &&\n \t\t    oe_size_greater_than(&to_pack, entry,\n \t\t\t\t\t repo_settings_get_big_file_threshold(the_repository)) &&\n-\t\t    (st = open_istream(the_repository, &entry->idx.oid, &type,\n-\t\t\t\t       &size, NULL)) != NULL)\n+\t\t    (st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n+\t\t\t\t\t\t &type, &size, NULL)) != NULL)\n \t\t\tbuf = NULL;\n \t\telse {\n \t\t\tbuf = odb_read_object(the_repository->objects,\n@@ -577,7 +577,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t\tdheader[--pos] = 128 | (--ofs & 127);\n \t\tif (limit && hdrlen + sizeof(dheader) - pos + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -591,7 +591,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t */\n \t\tif (limit && hdrlen + hashsz + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -601,7 +601,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t} else {\n \t\tif (limit && hdrlen + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -609,7 +609,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t}\n \tif (st) {\n \t\tdatalen = write_large_blob_data(st, f, &entry->idx.oid);\n-\t\tclose_istream(st);\n+\t\todb_read_stream_close(st);\n \t} else {\n \t\thashwrite(f, buf, datalen);\n \t\tfree(buf);\ndiff --git a/object-file.c b/object-file.c\nindex 8c67847fea..c6d2f2d953 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -139,7 +139,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \n-\tst = open_istream(r, oid, &obj_type, &size, NULL);\n+\tst = odb_read_object_stream(r->objects, oid, &obj_type, &size, NULL);\n \tif (!st)\n \t\treturn -1;\n \n@@ -151,10 +151,10 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tgit_hash_update(&c, hdr, hdrlen);\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\t\tssize_t readlen = odb_read_stream_read(st, buf, sizeof(buf));\n \n \t\tif (readlen < 0) {\n-\t\t\tclose_istream(st);\n+\t\t\todb_read_stream_close(st);\n \t\t\treturn -1;\n \t\t}\n \t\tif (!readlen)\n@@ -162,7 +162,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \t\tgit_hash_update(&c, buf, readlen);\n \t}\n \tgit_hash_final_oid(&real_oid, &c);\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \treturn !oideq(oid, &real_oid) ? -1 : 0;\n }\n \ndiff --git a/streaming.c b/streaming.c\nindex 3d80ddd757..3ac1a0c40f 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -35,7 +35,7 @@ static int close_istream_filtered(struct odb_read_stream *_fs)\n {\n \tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n \tfree_stream_filter(fs->filter);\n-\treturn close_istream(fs->upstream);\n+\treturn odb_read_stream_close(fs->upstream);\n }\n \n static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n@@ -87,7 +87,7 @@ static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n \n \t\t/* refill the input from the upstream */\n \t\tif (!fs->input_finished) {\n-\t\t\tfs->i_end = read_istream(fs->upstream, fs->ibuf, FILTER_BUFFER);\n+\t\t\tfs->i_end = odb_read_stream_read(fs->upstream, fs->ibuf, FILTER_BUFFER);\n \t\t\tif (fs->i_end < 0)\n \t\t\t\treturn -1;\n \t\t\tif (fs->i_end)\n@@ -149,7 +149,7 @@ static ssize_t read_istream_incore(struct odb_read_stream *_st, char *buf, size_\n }\n \n static int open_istream_incore(struct odb_read_stream **out,\n-\t\t\t       struct repository *r,\n+\t\t\t       struct object_database *odb,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n@@ -163,7 +163,7 @@ static int open_istream_incore(struct odb_read_stream **out,\n \toi.typep = &stream.base.type;\n \toi.sizep = &stream.base.size;\n \toi.contentp = (void **)&stream.buf;\n-\tret = odb_read_object_info_extended(r->objects, oid, &oi,\n+\tret = odb_read_object_info_extended(odb, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n@@ -180,49 +180,49 @@ static int open_istream_incore(struct odb_read_stream **out,\n  *****************************************************************************/\n \n static int istream_source(struct odb_read_stream **out,\n-\t\t\t  struct repository *r,\n+\t\t\t  struct object_database *odb,\n \t\t\t  const struct object_id *oid)\n {\n \tstruct odb_source *source;\n \n-\tif (!packfile_store_read_object_stream(out, r->objects->packfiles, oid))\n+\tif (!packfile_store_read_object_stream(out, odb->packfiles, oid))\n \t\treturn 0;\n \n-\todb_prepare_alternates(r->objects);\n-\tfor (source = r->objects->sources; source; source = source->next)\n+\todb_prepare_alternates(odb);\n+\tfor (source = odb->sources; source; source = source->next)\n \t\tif (!odb_source_loose_read_object_stream(out, source, oid))\n \t\t\treturn 0;\n \n-\treturn open_istream_incore(out, r, oid);\n+\treturn open_istream_incore(out, odb, oid);\n }\n \n /****************************************************************\n  * Users of streaming interface\n  ****************************************************************/\n \n-int close_istream(struct odb_read_stream *st)\n+int odb_read_stream_close(struct odb_read_stream *st)\n {\n \tint r = st->close(st);\n \tfree(st);\n \treturn r;\n }\n \n-ssize_t read_istream(struct odb_read_stream *st, void *buf, size_t sz)\n+ssize_t odb_read_stream_read(struct odb_read_stream *st, void *buf, size_t sz)\n {\n \treturn st->read(st, buf, sz);\n }\n \n-struct odb_read_stream *open_istream(struct repository *r,\n-\t\t\t\t     const struct object_id *oid,\n-\t\t\t\t     enum object_type *type,\n-\t\t\t\t     unsigned long *size,\n-\t\t\t\t     struct stream_filter *filter)\n+struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n+\t\t\t\t\t       const struct object_id *oid,\n+\t\t\t\t\t       enum object_type *type,\n+\t\t\t\t\t       unsigned long *size,\n+\t\t\t\t\t       struct stream_filter *filter)\n {\n \tstruct odb_read_stream *st;\n-\tconst struct object_id *real = lookup_replace_object(r, oid);\n+\tconst struct object_id *real = lookup_replace_object(odb->repo, oid);\n \tint ret;\n \n-\tret = istream_source(&st, r, real);\n+\tret = istream_source(&st, odb, real);\n \tif (ret)\n \t\treturn NULL;\n \n@@ -230,7 +230,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n \t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n \t\tif (!nst) {\n-\t\t\tclose_istream(st);\n+\t\t\todb_read_stream_close(st);\n \t\t\treturn NULL;\n \t\t}\n \t\tst = nst;\n@@ -253,7 +253,7 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = open_istream(odb->repo, oid, &type, &sz, filter);\n+\tst = odb_read_object_stream(odb, oid, &type, &sz, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\n@@ -264,7 +264,7 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\n \t\tssize_t wrote, holeto;\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\t\tssize_t readlen = odb_read_stream_read(st, buf, sizeof(buf));\n \n \t\tif (readlen < 0)\n \t\t\tgoto close_and_exit;\n@@ -295,6 +295,6 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tresult = 0;\n \n  close_and_exit:\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \treturn result;\n }\ndiff --git a/streaming.h b/streaming.h\nindex acfdef1598..2dce2e359f 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -24,11 +24,31 @@ struct odb_read_stream {\n \tunsigned long size; /* inflated size of full object */\n };\n \n-struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n-\t\t\t\t     enum object_type *, unsigned long *,\n-\t\t\t\t     struct stream_filter *);\n-int close_istream(struct odb_read_stream *);\n-ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n+/*\n+ * Create a new object stream for the given object database. Populates the type\n+ * and size pointers with the object's info. An optional filter can be used to\n+ * transform the object's content.\n+ *\n+ * Returns the stream on success, a `NULL` pointer otherwise.\n+ */\n+struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n+\t\t\t\t\t       const struct object_id *oid,\n+\t\t\t\t\t       enum object_type *type,\n+\t\t\t\t\t       unsigned long *size,\n+\t\t\t\t\t       struct stream_filter *filter);\n+\n+/*\n+ * Close the given read stream and release all resources associated with it.\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+int odb_read_stream_close(struct odb_read_stream *stream);\n+\n+/*\n+ * Read data from the stream into the buffer. Returns 0 on EOF and the number\n+ * of bytes read on success. Returns a negative error code in case reading from\n+ * the stream fails.\n+ */\n+ssize_t odb_read_stream_read(struct odb_read_stream *stream, void *buf, size_t len);\n \n /*\n  * Look up the object by its ID and write the full contents to the file\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531125","messageId":"20251121-b4-pks-odb-read-stream-v2-18-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 18/19] streaming: move into object database subsystem","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:41:03Z","receivedAt":"2025-11-21T07:42:10Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"The \"streaming\" terminology is somewhat generic, so it may not be\nimmediately obvious that \"streaming.{c,h}\" is specific to the object\ndatabase. Rectify this by moving it into the \"odb/\" directory so that it\ncan be immediately attributed to the object subsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Makefile                       | 2 +-\n archive-tar.c                  | 2 +-\n archive-zip.c                  | 2 +-\n builtin/cat-file.c             | 2 +-\n builtin/fsck.c                 | 2 +-\n builtin/index-pack.c           | 2 +-\n builtin/log.c                  | 2 +-\n builtin/pack-objects.c         | 2 +-\n entry.c                        | 2 +-\n meson.build                    | 2 +-\n object-file.c                  | 2 +-\n streaming.c => odb/streaming.c | 2 +-\n streaming.h => odb/streaming.h | 0\n packfile.c                     | 2 +-\n parallel-checkout.c            | 2 +-\n 15 files changed, 14 insertions(+), 14 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 7e0f77e298..6d8dcc4622 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1201,6 +1201,7 @@ LIB_OBJS += object-file.o\n LIB_OBJS += object-name.o\n LIB_OBJS += object.o\n LIB_OBJS += odb.o\n+LIB_OBJS += odb/streaming.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\n LIB_OBJS += oidset.o\n@@ -1294,7 +1295,6 @@ LIB_OBJS += split-index.o\n LIB_OBJS += stable-qsort.o\n LIB_OBJS += statinfo.o\n LIB_OBJS += strbuf.o\n-LIB_OBJS += streaming.o\n LIB_OBJS += string-list.o\n LIB_OBJS += strmap.o\n LIB_OBJS += strvec.o\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 4133e09ca1..74499c311f 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -12,8 +12,8 @@\n #include \"tar.h\"\n #include \"archive.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"strbuf.h\"\n-#include \"streaming.h\"\n #include \"run-command.h\"\n #include \"write-or-die.h\"\n \ndiff --git a/archive-zip.c b/archive-zip.c\nindex ff57f4f884..2b645f28ef 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -10,9 +10,9 @@\n #include \"gettext.h\"\n #include \"git-zlib.h\"\n #include \"hex.h\"\n-#include \"streaming.h\"\n #include \"utf8.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"strbuf.h\"\n #include \"userdiff.h\"\n #include \"write-or-die.h\"\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 120d626d66..505ddaa12f 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -18,13 +18,13 @@\n #include \"list-objects-filter-options.h\"\n #include \"parse-options.h\"\n #include \"userdiff.h\"\n-#include \"streaming.h\"\n #include \"oid-array.h\"\n #include \"packfile.h\"\n #include \"pack-bitmap.h\"\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"promisor-remote.h\"\n #include \"mailmap.h\"\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex 1a348d43c2..c7d2eea287 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -13,11 +13,11 @@\n #include \"fsck.h\"\n #include \"parse-options.h\"\n #include \"progress.h\"\n-#include \"streaming.h\"\n #include \"packfile.h\"\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"path.h\"\n #include \"read-cache-ll.h\"\n #include \"replace-object.h\"\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 67221dbe6a..6403edd3a6 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -16,12 +16,12 @@\n #include \"progress.h\"\n #include \"fsck.h\"\n #include \"strbuf.h\"\n-#include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"packfile.h\"\n #include \"pack-revindex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"oid-array.h\"\n #include \"oidset.h\"\n #include \"path.h\"\ndiff --git a/builtin/log.c b/builtin/log.c\nindex e7b83a6e00..d4cf9c59c8 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -16,6 +16,7 @@\n #include \"refs.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"pager.h\"\n #include \"color.h\"\n #include \"commit.h\"\n@@ -35,7 +36,6 @@\n #include \"parse-options.h\"\n #include \"line-log.h\"\n #include \"branch.h\"\n-#include \"streaming.h\"\n #include \"version.h\"\n #include \"mailmap.h\"\n #include \"progress.h\"\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex adf267c59d..f6c01bc4e0 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -22,7 +22,6 @@\n #include \"pack-objects.h\"\n #include \"progress.h\"\n #include \"refs.h\"\n-#include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"pack-bitmap.h\"\n #include \"delta-islands.h\"\n@@ -33,6 +32,7 @@\n #include \"packfile.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"dir.h\"\n #include \"midx.h\"\ndiff --git a/entry.c b/entry.c\nindex 38dfe670f7..7817aee362 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -2,13 +2,13 @@\n \n #include \"git-compat-util.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"dir.h\"\n #include \"environment.h\"\n #include \"gettext.h\"\n #include \"hex.h\"\n #include \"name-hash.h\"\n #include \"sparse-index.h\"\n-#include \"streaming.h\"\n #include \"submodule.h\"\n #include \"symlinks.h\"\n #include \"progress.h\"\ndiff --git a/meson.build b/meson.build\nindex 1f95a06edb..fc82929b37 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -397,6 +397,7 @@ libgit_sources = [\n   'object-name.c',\n   'object.c',\n   'odb.c',\n+  'odb/streaming.c',\n   'oid-array.c',\n   'oidmap.c',\n   'oidset.c',\n@@ -490,7 +491,6 @@ libgit_sources = [\n   'stable-qsort.c',\n   'statinfo.c',\n   'strbuf.c',\n-  'streaming.c',\n   'string-list.c',\n   'strmap.c',\n   'strvec.c',\ndiff --git a/object-file.c b/object-file.c\nindex c6d2f2d953..4b46cf5b71 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -20,13 +20,13 @@\n #include \"object-file-convert.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"oidtree.h\"\n #include \"pack.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n #include \"read-cache-ll.h\"\n #include \"setup.h\"\n-#include \"streaming.h\"\n #include \"tempfile.h\"\n #include \"tmp-objdir.h\"\n \ndiff --git a/streaming.c b/odb/streaming.c\nsimilarity index 99%\nrename from streaming.c\nrename to odb/streaming.c\nindex 3ac1a0c40f..a7ee50dc34 100644\n--- a/streaming.c\n+++ b/odb/streaming.c\n@@ -5,10 +5,10 @@\n #include \"git-compat-util.h\"\n #include \"convert.h\"\n #include \"environment.h\"\n-#include \"streaming.h\"\n #include \"repository.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \ndiff --git a/streaming.h b/odb/streaming.h\nsimilarity index 100%\nrename from streaming.h\nrename to odb/streaming.h\ndiff --git a/packfile.c b/packfile.c\nindex ad56ce0b90..7a16aaa90d 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -20,7 +20,7 @@\n #include \"tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n-#include \"streaming.h\"\n+#include \"odb/streaming.h\"\n #include \"midx.h\"\n #include \"commit-graph.h\"\n #include \"pack-revindex.h\"\ndiff --git a/parallel-checkout.c b/parallel-checkout.c\nindex 1cb6701b92..0bf4bd6d4a 100644\n--- a/parallel-checkout.c\n+++ b/parallel-checkout.c\n@@ -13,7 +13,7 @@\n #include \"read-cache-ll.h\"\n #include \"run-command.h\"\n #include \"sigchain.h\"\n-#include \"streaming.h\"\n+#include \"odb/streaming.h\"\n #include \"symlinks.h\"\n #include \"thread-utils.h\"\n #include \"trace2.h\"\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531126","messageId":"20251121-b4-pks-odb-read-stream-v2-19-ca8534963150@pks.im","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im","subject":"[PATCH v2 19/19] streaming: drop redundant type and size pointers","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-21T07:41:04Z","receivedAt":"2025-11-21T07:42:14Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In the preceding commits we have turned `struct odb_read_stream` into a\npublicly visible structure. Furthermore, this structure now contains the\ntype and size of the object that we are about to stream. Consequently,\nthe out-pointers that we used before to propagate the type and size of\nthe streamed object are now somewhat redundant with the data contained\nin the structure itself.\n\nDrop these out-pointers and adapt callers accordingly.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  4 +---\n archive-zip.c          |  5 ++---\n builtin/index-pack.c   |  7 ++-----\n builtin/pack-objects.c |  6 ++++--\n object-file.c          |  6 ++----\n odb/streaming.c        | 10 ++--------\n odb/streaming.h        |  7 ++-----\n 7 files changed, 15 insertions(+), 30 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 74499c311f..e34e3daec9 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -130,12 +130,10 @@ static void write_trailer(void)\n static int stream_blocked(struct repository *r, const struct object_id *oid)\n {\n \tstruct odb_read_stream *st;\n-\tenum object_type type;\n-\tunsigned long sz;\n \tchar buf[BLOCKSIZE];\n \tssize_t readlen;\n \n-\tst = odb_read_object_stream(r->objects, oid, &type, &sz, NULL);\n+\tst = odb_read_object_stream(r->objects, oid, NULL);\n \tif (!st)\n \t\treturn error(_(\"cannot stream blob %s\"), oid_to_hex(oid));\n \tfor (;;) {\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 2b645f28ef..f8d1e80671 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -347,12 +347,11 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\tmethod = ZIP_METHOD_DEFLATE;\n \n \t\tif (!buffer) {\n-\t\t\tenum object_type type;\n-\t\t\tstream = odb_read_object_stream(args->repo->objects, oid,\n-\t\t\t\t\t\t\t&type, &size, NULL);\n+\t\t\tstream = odb_read_object_stream(args->repo->objects, oid, NULL);\n \t\t\tif (!stream)\n \t\t\t\treturn error(_(\"cannot stream blob %s\"),\n \t\t\t\t\t     oid_to_hex(oid));\n+\t\t\tsize = stream->size;\n \t\t\tflags |= ZIP_STREAM;\n \t\t\tout = NULL;\n \t\t} else {\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 6403edd3a6..eb0c34b4c8 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -798,8 +798,6 @@ static int compare_objects(const unsigned char *buf, unsigned long size,\n static int check_collison(struct object_entry *entry)\n {\n \tstruct compare_data data;\n-\tenum object_type type;\n-\tunsigned long size;\n \n \tif (entry->size <= repo_settings_get_big_file_threshold(the_repository) ||\n \t    entry->type != OBJ_BLOB)\n@@ -807,11 +805,10 @@ static int check_collison(struct object_entry *entry)\n \n \tmemset(&data, 0, sizeof(data));\n \tdata.entry = entry;\n-\tdata.st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n-\t\t\t\t\t &type, &size, NULL);\n+\tdata.st = odb_read_object_stream(the_repository->objects, &entry->idx.oid, NULL);\n \tif (!data.st)\n \t\treturn -1;\n-\tif (size != entry->size || type != entry->type)\n+\tif (data.st->size != entry->size || data.st->type != entry->type)\n \t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t    oid_to_hex(&entry->idx.oid));\n \tunpack_data(entry, compare_objects, &data);\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f6c01bc4e0..2044378521 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -521,9 +521,11 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t    oe_size_greater_than(&to_pack, entry,\n \t\t\t\t\t repo_settings_get_big_file_threshold(the_repository)) &&\n \t\t    (st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n-\t\t\t\t\t\t &type, &size, NULL)) != NULL)\n+\t\t\t\t\t\t NULL)) != NULL) {\n \t\t\tbuf = NULL;\n-\t\telse {\n+\t\t\ttype = st->type;\n+\t\t\tsize = st->size;\n+\t\t} else {\n \t\t\tbuf = odb_read_object(the_repository->objects,\n \t\t\t\t\t      &entry->idx.oid, &type,\n \t\t\t\t\t      &size);\ndiff --git a/object-file.c b/object-file.c\nindex 4b46cf5b71..89ebe08b66 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -132,19 +132,17 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n int stream_object_signature(struct repository *r, const struct object_id *oid)\n {\n \tstruct object_id real_oid;\n-\tunsigned long size;\n-\tenum object_type obj_type;\n \tstruct odb_read_stream *st;\n \tstruct git_hash_ctx c;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \n-\tst = odb_read_object_stream(r->objects, oid, &obj_type, &size, NULL);\n+\tst = odb_read_object_stream(r->objects, oid, NULL);\n \tif (!st)\n \t\treturn -1;\n \n \t/* Generate the header */\n-\thdrlen = format_object_header(hdr, sizeof(hdr), obj_type, size);\n+\thdrlen = format_object_header(hdr, sizeof(hdr), st->type, st->size);\n \n \t/* Sha1.. */\n \tr->hash_algo->init_fn(&c);\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex a7ee50dc34..efd8f1f473 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -214,8 +214,6 @@ ssize_t odb_read_stream_read(struct odb_read_stream *st, void *buf, size_t sz)\n \n struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n \t\t\t\t\t       const struct object_id *oid,\n-\t\t\t\t\t       enum object_type *type,\n-\t\t\t\t\t       unsigned long *size,\n \t\t\t\t\t       struct stream_filter *filter)\n {\n \tstruct odb_read_stream *st;\n@@ -236,8 +234,6 @@ struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n \t\tst = nst;\n \t}\n \n-\t*size = st->size;\n-\t*type = st->type;\n \treturn st;\n }\n \n@@ -248,18 +244,16 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  int can_seek)\n {\n \tstruct odb_read_stream *st;\n-\tenum object_type type;\n-\tunsigned long sz;\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = odb_read_object_stream(odb, oid, &type, &sz, filter);\n+\tst = odb_read_object_stream(odb, oid, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\n \t\treturn result;\n \t}\n-\tif (type != OBJ_BLOB)\n+\tif (st->type != OBJ_BLOB)\n \t\tgoto close_and_exit;\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex 2dce2e359f..8220e8de3c 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -25,16 +25,13 @@ struct odb_read_stream {\n };\n \n /*\n- * Create a new object stream for the given object database. Populates the type\n- * and size pointers with the object's info. An optional filter can be used to\n- * transform the object's content.\n+ * Create a new object stream for the given object database. An optional filter\n+ * can be used to transform the object's content.\n  *\n  * Returns the stream on success, a `NULL` pointer otherwise.\n  */\n struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n \t\t\t\t\t       const struct object_id *oid,\n-\t\t\t\t\t       enum object_type *type,\n-\t\t\t\t\t       unsigned long *size,\n \t\t\t\t\t       struct stream_filter *filter);\n \n /*\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531150","messageId":"xmqqqztr45t5.fsf@gitster.g","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-2-ca8534963150@pks.im","subject":"Re: [PATCH v2 02/19] streaming: drop the `open()` callback function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-21T18:08:22Z","receivedAt":"2025-11-21T18:08:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When creating a read stream we first populate the structure with the\n> open callback function and then subsequently call the function. This\n> layout is somewhat weird though:\n>\n>   - The structure needs to be allocated and partially populated with the\n>     open function before we can properly initialize it.\n\nIt is unclear what are left for delayed initialization from this\ndescription.\n\n>   - We never use the `open()` callback after having opened it initially.\n\nI was not sure what this means in v1 and it still is not clear to\nme.  Naively the above reads as if it is somehow desirable if we can\ncall open() after we have already called it on an object.  The flow\nbeing a caller (e.g., stream_blob_to_fd()) first ask open_istream(),\nwhich calls the open method after figuring out which backend knows\nabout the object and how to open a stream on it, I am not sure what\nyou want your second and subsequent uses of the open() calklbacks\ndo.  Puzzled.\n\n> Instead, drop the callback entirely and refactor `istream_source()` so\n> that we open the streams immediately. This unblocks a subsequent step,\n> where we'll also start to allocate the structure in the source-specific\n> logic.\n\nBecause I do not think these open methods specific to each storage\nmechanism cascades into each other, open-coding the logic to\ndispatch into these open() methods in istream_source() itself,\ninstead of setting the method there and then have the caller call\nit, is a perfectly fine simplification, I think.\n\n> @@ -478,19 +477,14 @@ struct odb_read_stream *open_istream(struct repository *r,\n>  {\n>  \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n>  \tconst struct object_id *real = lookup_replace_object(r, oid);\n> -\tint ret = istream_source(st, r, real, type);\n> +\tint ret;\n>  \n> +\tret = istream_source(st, r, real, type);\n>  \tif (ret) {\n>  \t\tfree(st);\n>  \t\treturn NULL;\n>  \t}\n\nA patch noise?\n\n"},{"id":"531151","messageId":"xmqqldjz41wp.fsf@gitster.g","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-12-ca8534963150@pks.im","subject":"Re: [PATCH v2 12/19] streaming: rely on object sources to create object stream","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-21T19:32:38Z","receivedAt":"2025-11-21T19:32:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When creating an object stream we first look up the object info and, if\n> it's present, we call into the respective backend that contains the\n> object to create a new stream for it.\n>\n> This has the consequence that, for loose object source, we basically\n> iterate through the object sources twice: we first discover that the\n> file exists as a loose object in the first place by iterating through\n> all sources. And, once we have discovered it, we again walk through all\n> sources to try and map the object. The same issue will eventually also\n> surface once the packfile store becomes per-object-source.\n>\n> Furthermore, it feels rather pointless to first look up the object only\n> to then try and read it.\n>\n> Refactor the logic to be centered around sources instead. Instead of\n> first reading the object, we immediately ask the source to create the\n> object stream for us. If the object exists we get stream, otherwise\n> we'll try the next source.\n>\n> Like this we only have to iterate through sources once. But even more\n> importantly, this change also helps us to make the whole logic\n> pluggable. The object read stream subsystem does not need to be aware of\n> the different source backends anymore, but eventually it'll only have to\n> call the source's callback function.\n\nVery nicely done.\n\n> Note that at the current point in time we aren't fully there yet:\n>\n>   - The packfile store still sits on the object database level and is\n>     thus agnostic of the sources.\n>\n>   - We still have to call into both the packfile store and the loose\n>     object source.\n>\n> But both of these issues will soon be addressed.\n\n;-)\n\n> @@ -463,30 +461,15 @@ static int istream_source(struct odb_read_stream **out,\n>  \t\t\t  struct repository *r,\n>  \t\t\t  const struct object_id *oid)\n>  {\n> -\tunsigned long size;\n> -\tint status;\n> -\tstruct object_info oi = OBJECT_INFO_INIT;\n> -\n> -\toi.sizep = &size;\n> -\tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n> -\tif (status < 0)\n> -\t\treturn status;\n> +\tstruct odb_source *source;\n>  \n> -\tswitch (oi.whence) {\n> -\tcase OI_LOOSE:\n> -\t\tif (open_istream_loose(out, r, oid) < 0)\n> -\t\t\tbreak;\n> -\t\treturn 0;\n> -\tcase OI_PACKED:\n> -\t\tif (oi.u.packed.is_delta ||\n> -\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n> -\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n> -\t\t\t\t\t\toi.u.packed.offset) < 0)\n> -\t\t\tbreak;\n> +\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n>  \t\treturn 0;\n> -\tdefault:\n> -\t\tbreak;\n> -\t}\n> +\n> +\todb_prepare_alternates(r->objects);\n> +\tfor (source = r->objects->sources; source; source = source->next)\n> +\t\tif (!open_istream_loose(out, source, oid))\n> +\t\t\treturn 0;\n\nHmph.\n\nEarlier we let odb_read_object_info_extended() decide which one of\nthe duplicated objects (e.g., perhaps a loose object is still there\nafter packing), and then used the one it picked.  I think the\nodb_read_object_info_extended() encodes a particular order with with\nsolid reasons like \"do in-core cached one first\", \"favor objects in\npack over loose ones\".\n\nNow we instead let the first one with the object in the linked list\nof sources, which may be different, unless the linked list is\ncreated with the same \"why one source needs to be given precedence\nover the others\" reasoning.\n\nI do not know if/how it matters, this somewhat changes the\nsemantics, no?\n"},{"id":"531152","messageId":"xmqqh5un41fv.fsf@gitster.g","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-13-ca8534963150@pks.im","subject":"Re: [PATCH v2 13/19] streaming: get rid of `the_repository`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-21T19:42:44Z","receivedAt":"2025-11-21T19:42:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> diff --git a/streaming.h b/streaming.h\n> index f5ff5d7ac9..1a3de6812e 100644\n> --- a/streaming.h\n> +++ b/streaming.h\n> @@ -7,6 +7,7 @@\n>  #include \"object.h\"\n>  \n>  /* opaque */\n> +struct object_database;\n>  struct odb_read_stream;\n>  struct stream_filter;\n\nExtremely minor, but \"opaque\" used to refer to the fact that \"struct\ngit_istream\" is opaque to the users of this API.  In a later step\nyou'll remove the opaque comment (perhaps your change in that step\nwill stop it being opaque?), but until then, you'd want to keep the\ncomment next to \"struct odb_read_stream\", as it is my understanding\nthat object_database is not opaque here.\n\n"},{"id":"531156","messageId":"xmqqv7j32ahh.fsf@gitster.g","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-17-ca8534963150@pks.im","subject":"Re: [PATCH v2 17/19] streaming: refactor interface to be object-database-centric","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-22T00:10:18Z","receivedAt":"2025-11-22T00:10:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Refactor the streaming interface to be centered around object databases\n> instead of centered around the repository. Rename the functions\n> accordingly.\n\nGood idea.  \n\n> -\tst = open_istream(r, oid, &type, &sz, NULL);\n> +\tst = odb_read_object_stream(r->objects, oid, &type, &sz, NULL);\n\nCalling the thing that is returned a \"read stream\" is a lot more\ntrivially obvious than the original name \"i(nput) stream\", and I\nlike that aspect of the new name a lot better, and the structure is\nalso named appropriately (\"struct odb_read_stream\").\n\nAt least the old naming was consistent with the usual file I/O API.\nyou \"open\" istream, then \"read\" from that istream, and finally\n\"close\" that istream.  If you insist on having the noun first before\nthe verb, call them\n\n    odb_read_stream_open()\n    odb_read_stream_read()\n    odb_read_stream_close()\n\nperhaps?  I think _read and _close are already named appropriately.\n\n"},{"id":"531165","messageId":"xmqq5xb132xn.fsf@gitster.g","threadId":"64509","inReplyTo":"20251121-b4-pks-odb-read-stream-v2-18-ca8534963150@pks.im","subject":"Re: [PATCH v2 18/19] streaming: move into object database subsystem","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-23T02:20:20Z","receivedAt":"2025-11-23T02:20:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> The \"streaming\" terminology is somewhat generic, so it may not be\n> immediately obvious that \"streaming.{c,h}\" is specific to the object\n> database. Rectify this by moving it into the \"odb/\" directory so that it\n> can be immediately attributed to the object subsystem.\n\nI do not have an objection against this move.  Looking good.\n\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  Makefile                       | 2 +-\n>  archive-tar.c                  | 2 +-\n>  archive-zip.c                  | 2 +-\n>  builtin/cat-file.c             | 2 +-\n>  builtin/fsck.c                 | 2 +-\n>  builtin/index-pack.c           | 2 +-\n>  builtin/log.c                  | 2 +-\n>  builtin/pack-objects.c         | 2 +-\n>  entry.c                        | 2 +-\n>  meson.build                    | 2 +-\n>  object-file.c                  | 2 +-\n>  streaming.c => odb/streaming.c | 2 +-\n>  streaming.h => odb/streaming.h | 0\n>  packfile.c                     | 2 +-\n>  parallel-checkout.c            | 2 +-\n>  15 files changed, 14 insertions(+), 14 deletions(-)\n>\n> diff --git a/Makefile b/Makefile\n> index 7e0f77e298..6d8dcc4622 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -1201,6 +1201,7 @@ LIB_OBJS += object-file.o\n>  LIB_OBJS += object-name.o\n>  LIB_OBJS += object.o\n>  LIB_OBJS += odb.o\n> +LIB_OBJS += odb/streaming.o\n>  LIB_OBJS += oid-array.o\n>  LIB_OBJS += oidmap.o\n>  LIB_OBJS += oidset.o\n> @@ -1294,7 +1295,6 @@ LIB_OBJS += split-index.o\n>  LIB_OBJS += stable-qsort.o\n>  LIB_OBJS += statinfo.o\n>  LIB_OBJS += strbuf.o\n> -LIB_OBJS += streaming.o\n>  LIB_OBJS += string-list.o\n>  LIB_OBJS += strmap.o\n>  LIB_OBJS += strvec.o\n> diff --git a/archive-tar.c b/archive-tar.c\n> index 4133e09ca1..74499c311f 100644\n> --- a/archive-tar.c\n> +++ b/archive-tar.c\n> @@ -12,8 +12,8 @@\n>  #include \"tar.h\"\n>  #include \"archive.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"strbuf.h\"\n> -#include \"streaming.h\"\n>  #include \"run-command.h\"\n>  #include \"write-or-die.h\"\n>  \n> diff --git a/archive-zip.c b/archive-zip.c\n> index ff57f4f884..2b645f28ef 100644\n> --- a/archive-zip.c\n> +++ b/archive-zip.c\n> @@ -10,9 +10,9 @@\n>  #include \"gettext.h\"\n>  #include \"git-zlib.h\"\n>  #include \"hex.h\"\n> -#include \"streaming.h\"\n>  #include \"utf8.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"strbuf.h\"\n>  #include \"userdiff.h\"\n>  #include \"write-or-die.h\"\n> diff --git a/builtin/cat-file.c b/builtin/cat-file.c\n> index 120d626d66..505ddaa12f 100644\n> --- a/builtin/cat-file.c\n> +++ b/builtin/cat-file.c\n> @@ -18,13 +18,13 @@\n>  #include \"list-objects-filter-options.h\"\n>  #include \"parse-options.h\"\n>  #include \"userdiff.h\"\n> -#include \"streaming.h\"\n>  #include \"oid-array.h\"\n>  #include \"packfile.h\"\n>  #include \"pack-bitmap.h\"\n>  #include \"object-file.h\"\n>  #include \"object-name.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"replace-object.h\"\n>  #include \"promisor-remote.h\"\n>  #include \"mailmap.h\"\n> diff --git a/builtin/fsck.c b/builtin/fsck.c\n> index 1a348d43c2..c7d2eea287 100644\n> --- a/builtin/fsck.c\n> +++ b/builtin/fsck.c\n> @@ -13,11 +13,11 @@\n>  #include \"fsck.h\"\n>  #include \"parse-options.h\"\n>  #include \"progress.h\"\n> -#include \"streaming.h\"\n>  #include \"packfile.h\"\n>  #include \"object-file.h\"\n>  #include \"object-name.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"path.h\"\n>  #include \"read-cache-ll.h\"\n>  #include \"replace-object.h\"\n> diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> index 67221dbe6a..6403edd3a6 100644\n> --- a/builtin/index-pack.c\n> +++ b/builtin/index-pack.c\n> @@ -16,12 +16,12 @@\n>  #include \"progress.h\"\n>  #include \"fsck.h\"\n>  #include \"strbuf.h\"\n> -#include \"streaming.h\"\n>  #include \"thread-utils.h\"\n>  #include \"packfile.h\"\n>  #include \"pack-revindex.h\"\n>  #include \"object-file.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"oid-array.h\"\n>  #include \"oidset.h\"\n>  #include \"path.h\"\n> diff --git a/builtin/log.c b/builtin/log.c\n> index e7b83a6e00..d4cf9c59c8 100644\n> --- a/builtin/log.c\n> +++ b/builtin/log.c\n> @@ -16,6 +16,7 @@\n>  #include \"refs.h\"\n>  #include \"object-name.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"pager.h\"\n>  #include \"color.h\"\n>  #include \"commit.h\"\n> @@ -35,7 +36,6 @@\n>  #include \"parse-options.h\"\n>  #include \"line-log.h\"\n>  #include \"branch.h\"\n> -#include \"streaming.h\"\n>  #include \"version.h\"\n>  #include \"mailmap.h\"\n>  #include \"progress.h\"\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> index adf267c59d..f6c01bc4e0 100644\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n> @@ -22,7 +22,6 @@\n>  #include \"pack-objects.h\"\n>  #include \"progress.h\"\n>  #include \"refs.h\"\n> -#include \"streaming.h\"\n>  #include \"thread-utils.h\"\n>  #include \"pack-bitmap.h\"\n>  #include \"delta-islands.h\"\n> @@ -33,6 +32,7 @@\n>  #include \"packfile.h\"\n>  #include \"object-file.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"replace-object.h\"\n>  #include \"dir.h\"\n>  #include \"midx.h\"\n> diff --git a/entry.c b/entry.c\n> index 38dfe670f7..7817aee362 100644\n> --- a/entry.c\n> +++ b/entry.c\n> @@ -2,13 +2,13 @@\n>  \n>  #include \"git-compat-util.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"dir.h\"\n>  #include \"environment.h\"\n>  #include \"gettext.h\"\n>  #include \"hex.h\"\n>  #include \"name-hash.h\"\n>  #include \"sparse-index.h\"\n> -#include \"streaming.h\"\n>  #include \"submodule.h\"\n>  #include \"symlinks.h\"\n>  #include \"progress.h\"\n> diff --git a/meson.build b/meson.build\n> index 1f95a06edb..fc82929b37 100644\n> --- a/meson.build\n> +++ b/meson.build\n> @@ -397,6 +397,7 @@ libgit_sources = [\n>    'object-name.c',\n>    'object.c',\n>    'odb.c',\n> +  'odb/streaming.c',\n>    'oid-array.c',\n>    'oidmap.c',\n>    'oidset.c',\n> @@ -490,7 +491,6 @@ libgit_sources = [\n>    'stable-qsort.c',\n>    'statinfo.c',\n>    'strbuf.c',\n> -  'streaming.c',\n>    'string-list.c',\n>    'strmap.c',\n>    'strvec.c',\n> diff --git a/object-file.c b/object-file.c\n> index c6d2f2d953..4b46cf5b71 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -20,13 +20,13 @@\n>  #include \"object-file-convert.h\"\n>  #include \"object-file.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"oidtree.h\"\n>  #include \"pack.h\"\n>  #include \"packfile.h\"\n>  #include \"path.h\"\n>  #include \"read-cache-ll.h\"\n>  #include \"setup.h\"\n> -#include \"streaming.h\"\n>  #include \"tempfile.h\"\n>  #include \"tmp-objdir.h\"\n>  \n> diff --git a/streaming.c b/odb/streaming.c\n> similarity index 99%\n> rename from streaming.c\n> rename to odb/streaming.c\n> index 3ac1a0c40f..a7ee50dc34 100644\n> --- a/streaming.c\n> +++ b/odb/streaming.c\n> @@ -5,10 +5,10 @@\n>  #include \"git-compat-util.h\"\n>  #include \"convert.h\"\n>  #include \"environment.h\"\n> -#include \"streaming.h\"\n>  #include \"repository.h\"\n>  #include \"object-file.h\"\n>  #include \"odb.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"replace-object.h\"\n>  #include \"packfile.h\"\n>  \n> diff --git a/streaming.h b/odb/streaming.h\n> similarity index 100%\n> rename from streaming.h\n> rename to odb/streaming.h\n> diff --git a/packfile.c b/packfile.c\n> index ad56ce0b90..7a16aaa90d 100644\n> --- a/packfile.c\n> +++ b/packfile.c\n> @@ -20,7 +20,7 @@\n>  #include \"tree.h\"\n>  #include \"object-file.h\"\n>  #include \"odb.h\"\n> -#include \"streaming.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"midx.h\"\n>  #include \"commit-graph.h\"\n>  #include \"pack-revindex.h\"\n> diff --git a/parallel-checkout.c b/parallel-checkout.c\n> index 1cb6701b92..0bf4bd6d4a 100644\n> --- a/parallel-checkout.c\n> +++ b/parallel-checkout.c\n> @@ -13,7 +13,7 @@\n>  #include \"read-cache-ll.h\"\n>  #include \"run-command.h\"\n>  #include \"sigchain.h\"\n> -#include \"streaming.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"symlinks.h\"\n>  #include \"thread-utils.h\"\n>  #include \"trace2.h\"\n"},{"id":"531177","messageId":"aSNZeyoVWDVTU4X_@pks.im","threadId":"64509","inReplyTo":"xmqqh5un41fv.fsf@gitster.g","subject":"Re: [PATCH v2 13/19] streaming: get rid of `the_repository`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:07Z","receivedAt":"2025-11-23T18:59:21Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Nov 21, 2025 at 11:42:44AM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > diff --git a/streaming.h b/streaming.h\n> > index f5ff5d7ac9..1a3de6812e 100644\n> > --- a/streaming.h\n> > +++ b/streaming.h\n> > @@ -7,6 +7,7 @@\n> >  #include \"object.h\"\n> >  \n> >  /* opaque */\n> > +struct object_database;\n> >  struct odb_read_stream;\n> >  struct stream_filter;\n> \n> Extremely minor, but \"opaque\" used to refer to the fact that \"struct\n> git_istream\" is opaque to the users of this API.  In a later step\n> you'll remove the opaque comment (perhaps your change in that step\n> will stop it being opaque?), but until then, you'd want to keep the\n> comment next to \"struct odb_read_stream\", as it is my understanding\n> that object_database is not opaque here.\n\nGood point, will fix.\n\nPatrick\n"},{"id":"531178","messageId":"aSNZiaa9tRQgKbm5@pks.im","threadId":"64509","inReplyTo":"xmqqv7j32ahh.fsf@gitster.g","subject":"Re: [PATCH v2 17/19] streaming: refactor interface to be object-database-centric","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:21Z","receivedAt":"2025-11-23T18:59:26Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Nov 21, 2025 at 04:10:18PM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Refactor the streaming interface to be centered around object databases\n> > instead of centered around the repository. Rename the functions\n> > accordingly.\n> \n> Good idea.  \n> \n> > -\tst = open_istream(r, oid, &type, &sz, NULL);\n> > +\tst = odb_read_object_stream(r->objects, oid, &type, &sz, NULL);\n> \n> Calling the thing that is returned a \"read stream\" is a lot more\n> trivially obvious than the original name \"i(nput) stream\", and I\n> like that aspect of the new name a lot better, and the structure is\n> also named appropriately (\"struct odb_read_stream\").\n> \n> At least the old naming was consistent with the usual file I/O API.\n> you \"open\" istream, then \"read\" from that istream, and finally\n> \"close\" that istream.  If you insist on having the noun first before\n> the verb, call them\n> \n>     odb_read_stream_open()\n>     odb_read_stream_read()\n>     odb_read_stream_close()\n> \n> perhaps?  I think _read and _close are already named appropriately.\n\nAh, right, that makes sense. `odb_read_stream_open()` is also shorter\ncompared to `odb_read_object_stream()`. Will adapt.\n\nPatrick\n"},{"id":"531179","messageId":"aSNZklOl98TYRTUg@pks.im","threadId":"64509","inReplyTo":"xmqqqztr45t5.fsf@gitster.g","subject":"Re: [PATCH v2 02/19] streaming: drop the `open()` callback function","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:30Z","receivedAt":"2025-11-23T18:59:35Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Nov 21, 2025 at 10:08:22AM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > When creating a read stream we first populate the structure with the\n> > open callback function and then subsequently call the function. This\n> > layout is somewhat weird though:\n> >\n> >   - The structure needs to be allocated and partially populated with the\n> >     open function before we can properly initialize it.\n> \n> It is unclear what are left for delayed initialization from this\n> description.\n> \n> >   - We never use the `open()` callback after having opened it initially.\n> \n> I was not sure what this means in v1 and it still is not clear to\n> me.  Naively the above reads as if it is somehow desirable if we can\n> call open() after we have already called it on an object.  The flow\n> being a caller (e.g., stream_blob_to_fd()) first ask open_istream(),\n> which calls the open method after figuring out which backend knows\n> about the object and how to open a stream on it, I am not sure what\n> you want your second and subsequent uses of the open() calklbacks\n> do.  Puzzled.\n\nI actually mean the opposite, so exactly what you describe: why do we\nstore the `open()` callback in a member variable of the stream if it's\nonly ever called a single time, only, and is never called a second time\nthereafter?\n\nI'll rephrase this.\n\n> > Instead, drop the callback entirely and refactor `istream_source()` so\n> > that we open the streams immediately. This unblocks a subsequent step,\n> > where we'll also start to allocate the structure in the source-specific\n> > logic.\n> \n> Because I do not think these open methods specific to each storage\n> mechanism cascades into each other, open-coding the logic to\n> dispatch into these open() methods in istream_source() itself,\n> instead of setting the method there and then have the caller call\n> it, is a perfectly fine simplification, I think.\n> \n> > @@ -478,19 +477,14 @@ struct odb_read_stream *open_istream(struct repository *r,\n> >  {\n> >  \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n> >  \tconst struct object_id *real = lookup_replace_object(r, oid);\n> > -\tint ret = istream_source(st, r, real, type);\n> > +\tint ret;\n> >  \n> > +\tret = istream_source(st, r, real, type);\n> >  \tif (ret) {\n> >  \t\tfree(st);\n> >  \t\treturn NULL;\n> >  \t}\n> \n> A patch noise?\n\nWill drop this line.\n\nThanks!\n\nPatrick\n"},{"id":"531180","messageId":"aSNZmRHLIBievXkA@pks.im","threadId":"64509","inReplyTo":"xmqqldjz41wp.fsf@gitster.g","subject":"Re: [PATCH v2 12/19] streaming: rely on object sources to create object stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:37Z","receivedAt":"2025-11-23T18:59:42Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Nov 21, 2025 at 11:32:38AM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > @@ -463,30 +461,15 @@ static int istream_source(struct odb_read_stream **out,\n> >  \t\t\t  struct repository *r,\n> >  \t\t\t  const struct object_id *oid)\n> >  {\n> > -\tunsigned long size;\n> > -\tint status;\n> > -\tstruct object_info oi = OBJECT_INFO_INIT;\n> > -\n> > -\toi.sizep = &size;\n> > -\tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n> > -\tif (status < 0)\n> > -\t\treturn status;\n> > +\tstruct odb_source *source;\n> >  \n> > -\tswitch (oi.whence) {\n> > -\tcase OI_LOOSE:\n> > -\t\tif (open_istream_loose(out, r, oid) < 0)\n> > -\t\t\tbreak;\n> > -\t\treturn 0;\n> > -\tcase OI_PACKED:\n> > -\t\tif (oi.u.packed.is_delta ||\n> > -\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n> > -\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n> > -\t\t\t\t\t\toi.u.packed.offset) < 0)\n> > -\t\t\tbreak;\n> > +\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n> >  \t\treturn 0;\n> > -\tdefault:\n> > -\t\tbreak;\n> > -\t}\n> > +\n> > +\todb_prepare_alternates(r->objects);\n> > +\tfor (source = r->objects->sources; source; source = source->next)\n> > +\t\tif (!open_istream_loose(out, source, oid))\n> > +\t\t\treturn 0;\n> \n> Hmph.\n> \n> Earlier we let odb_read_object_info_extended() decide which one of\n> the duplicated objects (e.g., perhaps a loose object is still there\n> after packing), and then used the one it picked.  I think the\n> odb_read_object_info_extended() encodes a particular order with with\n> solid reasons like \"do in-core cached one first\", \"favor objects in\n> pack over loose ones\".\n> \n> Now we instead let the first one with the object in the linked list\n> of sources, which may be different, unless the linked list is\n> created with the same \"why one source needs to be given precedence\n> over the others\" reasoning.\n> \n> I do not know if/how it matters, this somewhat changes the\n> semantics, no?\n\nThe semantics are slightly different now in case multiple sources have\nthe object, true. I don't really think that this matters though: the\nstream doesn't even indicate to the caller which source the stream has\nbeen opened from, and neither does it indicate whether the object was\nloose or packed. So assuming that there is no hash collision the result\nwould be the same, as the object contents should be similar independent\nof the source.\n\nWill update the commit message and add an explanation.\n\nPatrick\n"},{"id":"531181","messageId":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im","subject":"[PATCH v3 00/19] Refactor object read streams to work via object sources","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:25Z","receivedAt":"2025-11-23T18:59:47Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi,\n\nthe `git_istream` data structure can be used to read objects from the\nobject database in a streaming fashion. This is used for example to read\nlarge files that one doesn't want to load into memory in full.\n\nIn the current architecture, all the logic to handle these streams is\nfully self-contained in \"streaming.c\". It contains the logic to set up\nstreams for loose, packed, in-memory and filtered objects. This doesn't\nreally play all that well with pluggable object databases, as it should\nbe the responsibility of the object database source itself to handle the\nlogic.\n\nThis patch series thus revamps our object read streams: instead of being\nentirely contained in \"streaming.c\", the format-specific streams are now\ncreated by the ODB sources. This allows each source itself to decide\nwhether and, if so, how to make objects streamable.\n\nThis overall requires quite a bit of refactoring, but I think that the\nend result is an easier-to-understand infrastructure that is an\nimprovement even without pluggable object databases.\n\nThis series is built on top of v2.52.0 with ps/object-source-loose at\n3e5e360888 (object-file: refactor writing objects via a stream,\n2025-11-03) merged into it.\n\nChanges in v3:\n  - Clarify why we want to get rid of the `open()` callback.\n  - Explain change in semantics now that we iterate through sources\n    first to create the read stream.\n  - Fix \"opaque\" comment applying to the correct structure.\n  - Rename `odb_read_object_stream()` to `odb_read_stream_open()`.\n  - Link to v2: https://lore.kernel.org/r/20251121-b4-pks-odb-read-stream-v2-0-ca8534963150@pks.im\n\nChanges in v2:\n  - Some commit message improvements.\n  - Drop the `type` and `size` out pointers in\n    `odb_read_object_stream()` in an additional commit.\n  - Improve a \"hidden\" variable declaration by moving it onto its own\n    line.\n  - Link to v1: https://lore.kernel.org/r/20251119-b4-pks-odb-read-stream-v1-0-adacf03c2ccf@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (19):\n      streaming: rename `git_istream` into `odb_read_stream`\n      streaming: drop the `open()` callback function\n      streaming: propagate final object type via the stream\n      streaming: explicitly pass packfile info when streaming a packed object\n      streaming: allocate stream inside the backend-specific logic\n      streaming: create structure for in-core object streams\n      streaming: create structure for loose object streams\n      streaming: create structure for packed object streams\n      streaming: create structure for filtered object streams\n      streaming: move zlib stream into backends\n      packfile: introduce function to read object info from a store\n      streaming: rely on object sources to create object stream\n      streaming: get rid of `the_repository`\n      streaming: make the `odb_read_stream` definition public\n      streaming: move logic to read loose objects streams into backend\n      streaming: move logic to read packed objects streams into backend\n      streaming: refactor interface to be object-database-centric\n      streaming: move into object database subsystem\n      streaming: drop redundant type and size pointers\n\n Makefile               |   2 +-\n archive-tar.c          |  12 +-\n archive-zip.c          |  17 +-\n builtin/cat-file.c     |   4 +-\n builtin/fsck.c         |   5 +-\n builtin/index-pack.c   |  15 +-\n builtin/log.c          |   6 +-\n builtin/pack-objects.c |  24 ++-\n entry.c                |   4 +-\n meson.build            |   2 +-\n object-file.c          | 183 ++++++++++++++--\n object-file.h          |  42 +---\n odb.c                  |  29 +--\n odb/streaming.c        | 293 ++++++++++++++++++++++++++\n odb/streaming.h        |  67 ++++++\n packfile.c             | 199 ++++++++++++++++--\n packfile.h             |  17 +-\n parallel-checkout.c    |   5 +-\n streaming.c            | 561 -------------------------------------------------\n streaming.h            |  21 --\n 20 files changed, 779 insertions(+), 729 deletions(-)\n\nRange-diff versus v2:\n\n 1:  9862db07e9 =  1:  5e1b90ccf0 streaming: rename `git_istream` into `odb_read_stream`\n 2:  42a9684d52 !  2:  c55a13abb6 streaming: drop the `open()` callback function\n    @@ Commit message\n           - The structure needs to be allocated and partially populated with the\n             open function before we can properly initialize it.\n     \n    -      - We never use the `open()` callback after having opened it initially.\n    +      - We only ever call the `open()` callback function right after having\n    +        populated the `struct odb_read_stream::open` member, and it's never\n    +        called thereafter again. So it is somewhat pointless to store the\n    +        callback in the first place.\n     \n         Especially the first point creates a problem for us. In subsequent\n         commits we'll want to fully move construction of the read source into\n    @@ streaming.c: static int istream_source(struct odb_read_stream *st,\n      \n      /****************************************************************\n     @@ streaming.c: struct odb_read_stream *open_istream(struct repository *r,\n    - {\n    - \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n    - \tconst struct object_id *real = lookup_replace_object(r, oid);\n    --\tint ret = istream_source(st, r, real, type);\n    -+\tint ret;\n    - \n    -+\tret = istream_source(st, r, real, type);\n    - \tif (ret) {\n    - \t\tfree(st);\n      \t\treturn NULL;\n      \t}\n      \n 3:  c00bef7a2d !  3:  c588ab7a66 streaming: propagate final object type via the stream\n    @@ streaming.c: static int istream_source(struct odb_read_stream *st,\n      \n      /****************************************************************\n     @@ streaming.c: struct odb_read_stream *open_istream(struct repository *r,\n    + {\n    + \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n      \tconst struct object_id *real = lookup_replace_object(r, oid);\n    - \tint ret;\n    +-\tint ret = istream_source(st, r, real, type);\n    ++\tint ret = istream_source(st, r, real);\n      \n    --\tret = istream_source(st, r, real, type);\n    -+\tret = istream_source(st, r, real);\n      \tif (ret) {\n      \t\tfree(st);\n    - \t\treturn NULL;\n     @@ streaming.c: struct odb_read_stream *open_istream(struct repository *r,\n      \t}\n      \n 4:  3d5f3ce9d2 =  4:  5b3671c699 streaming: explicitly pass packfile info when streaming a packed object\n 5:  0bd824d570 !  5:  440b858905 streaming: allocate stream inside the backend-specific logic\n    @@ streaming.c: static ssize_t read_istream_incore(struct odb_read_stream *st, char\n      \t\t\t       const struct object_id *oid)\n      {\n      \tstruct object_info oi = OBJECT_INFO_INIT;\n    +-\n    +-\tst->u.incore.read_ptr = 0;\n    +-\tst->close = close_istream_incore;\n    +-\tst->read = read_istream_incore;\n    +-\n    +-\toi.typep = &st->type;\n    +-\toi.sizep = &st->size;\n    +-\toi.contentp = (void **)&st->u.incore.buf;\n    +-\treturn odb_read_object_info_extended(r->objects, oid, &oi,\n    +-\t\t\t\t\t     OBJECT_INFO_DIE_IF_CORRUPT);\n     +\tstruct odb_read_stream stream = {\n     +\t\t.close = close_istream_incore,\n     +\t\t.read = read_istream_incore,\n     +\t};\n     +\tint ret;\n    - \n    --\tst->u.incore.read_ptr = 0;\n    --\tst->close = close_istream_incore;\n    --\tst->read = read_istream_incore;\n    ++\n     +\toi.typep = &stream.type;\n     +\toi.sizep = &stream.size;\n     +\toi.contentp = (void **)&stream.u.incore.buf;\n    @@ streaming.c: static ssize_t read_istream_incore(struct odb_read_stream *st, char\n     +\t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n     +\tif (ret)\n     +\t\treturn ret;\n    - \n    --\toi.typep = &st->type;\n    --\toi.sizep = &st->size;\n    --\toi.contentp = (void **)&st->u.incore.buf;\n    --\treturn odb_read_object_info_extended(r->objects, oid, &oi,\n    --\t\t\t\t\t     OBJECT_INFO_DIE_IF_CORRUPT);\n    ++\n     +\tCALLOC_ARRAY(*out, 1);\n     +\t**out = stream;\n     +\treturn 0;\n    @@ streaming.c: struct odb_read_stream *open_istream(struct repository *r,\n     -\tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n     +\tstruct odb_read_stream *st;\n      \tconst struct object_id *real = lookup_replace_object(r, oid);\n    - \tint ret;\n    +-\tint ret = istream_source(st, r, real);\n    ++\tint ret = istream_source(&st, r, real);\n      \n    --\tret = istream_source(st, r, real);\n     -\tif (ret) {\n     -\t\tfree(st);\n    -+\tret = istream_source(&st, r, real);\n     +\tif (ret)\n      \t\treturn NULL;\n     -\t}\n 6:  468f17442a =  6:  9107044e1a streaming: create structure for in-core object streams\n 7:  42f75b6d1f =  7:  9d2fd8212f streaming: create structure for loose object streams\n 8:  63b3dbe842 =  8:  82b994a6ca streaming: create structure for packed object streams\n 9:  e192352dc3 =  9:  96c07c0e5f streaming: create structure for filtered object streams\n10:  dd718680f6 = 10:  ccb8abf077 streaming: move zlib stream into backends\n11:  466ccbe059 = 11:  07ef79d591 packfile: introduce function to read object info from a store\n12:  ba7bddecb1 ! 12:  741414fef9 streaming: rely on object sources to create object stream\n    @@ Commit message\n     \n         But both of these issues will soon be addressed.\n     \n    +    This refactoring results in a slight change to semantics: previously, it\n    +    was `odb_read_object_info_extended()` that picked the source for us, and\n    +    it would have favored packed (non-deltified) objects over loose objects.\n    +    And while we still favor packed over loose objects for a single source\n    +    with the new logic, we'll now favor a loose object from an earlier\n    +    source over a packed object from a later source.\n    +\n    +    Ultimately this shouldn't matter though: the stream doesn't indicate to\n    +    the caller which source it is from and whether it was created from a\n    +    packed or loose object, so such details are opaque to the caller. And\n    +    other than that we should be able to assume that two objects with the\n    +    same object ID should refer to the same content, so the streamed data\n    +    would be the same, too.\n    +\n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n     \n      ## streaming.c ##\n13:  723910c871 ! 13:  39134f2260 streaming: get rid of `the_repository`\n    @@ streaming.c: int stream_blob_to_fd(int fd, const struct object_id *oid, struct s\n     \n      ## streaming.h ##\n     @@\n    + \n      #include \"object.h\"\n      \n    - /* opaque */\n     +struct object_database;\n    + /* opaque */\n      struct odb_read_stream;\n      struct stream_filter;\n    - \n     @@ streaming.h: struct odb_read_stream *open_istream(struct repository *, const struct object_id\n      int close_istream(struct odb_read_stream *);\n      ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n14:  023015855f ! 14:  12b6ff9b93 streaming: make the `odb_read_stream` definition public\n    @@ streaming.c\n     \n      ## streaming.h ##\n     @@\n    - \n      #include \"object.h\"\n      \n    --/* opaque */\n      struct object_database;\n    +-/* opaque */\n      struct odb_read_stream;\n      struct stream_filter;\n      \n15:  9439f09f8b = 15:  af51d9959f streaming: move logic to read loose objects streams into backend\n16:  e7f8c8038d = 16:  5bc76e022d streaming: move logic to read packed objects streams into backend\n17:  b8933fb980 ! 17:  1dcd53f244 streaming: refactor interface to be object-database-centric\n    @@ archive-tar.c: static int stream_blocked(struct repository *r, const struct obje\n      \tssize_t readlen;\n      \n     -\tst = open_istream(r, oid, &type, &sz, NULL);\n    -+\tst = odb_read_object_stream(r->objects, oid, &type, &sz, NULL);\n    ++\tst = odb_read_stream_open(r->objects, oid, &type, &sz, NULL);\n      \tif (!st)\n      \t\treturn error(_(\"cannot stream blob %s\"), oid_to_hex(oid));\n      \tfor (;;) {\n    @@ archive-zip.c: static int write_zip_entry(struct archiver_args *args,\n      \t\t\tenum object_type type;\n     -\t\t\tstream = open_istream(args->repo, oid, &type, &size,\n     -\t\t\t\t\t      NULL);\n    -+\t\t\tstream = odb_read_object_stream(args->repo->objects, oid,\n    -+\t\t\t\t\t\t\t&type, &size, NULL);\n    ++\t\t\tstream = odb_read_stream_open(args->repo->objects, oid,\n    ++\t\t\t\t\t\t      &type, &size, NULL);\n      \t\t\tif (!stream)\n      \t\t\t\treturn error(_(\"cannot stream blob %s\"),\n      \t\t\t\t\t     oid_to_hex(oid));\n    @@ builtin/index-pack.c: static int check_collison(struct object_entry *entry)\n      \tdata.entry = entry;\n     -\tdata.st = open_istream(the_repository, &entry->idx.oid, &type, &size,\n     -\t\t\t       NULL);\n    -+\tdata.st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n    -+\t\t\t\t\t &type, &size, NULL);\n    ++\tdata.st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n    ++\t\t\t\t       &type, &size, NULL);\n      \tif (!data.st)\n      \t\treturn -1;\n      \tif (size != entry->size || type != entry->type)\n    @@ builtin/pack-objects.c: static unsigned long write_no_reuse_object(struct hashfi\n      \t\t\t\t\t repo_settings_get_big_file_threshold(the_repository)) &&\n     -\t\t    (st = open_istream(the_repository, &entry->idx.oid, &type,\n     -\t\t\t\t       &size, NULL)) != NULL)\n    -+\t\t    (st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n    -+\t\t\t\t\t\t &type, &size, NULL)) != NULL)\n    ++\t\t    (st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n    ++\t\t\t\t\t       &type, &size, NULL)) != NULL)\n      \t\t\tbuf = NULL;\n      \t\telse {\n      \t\t\tbuf = odb_read_object(the_repository->objects,\n    @@ object-file.c: int stream_object_signature(struct repository *r, const struct ob\n      \tint hdrlen;\n      \n     -\tst = open_istream(r, oid, &obj_type, &size, NULL);\n    -+\tst = odb_read_object_stream(r->objects, oid, &obj_type, &size, NULL);\n    ++\tst = odb_read_stream_open(r->objects, oid, &obj_type, &size, NULL);\n      \tif (!st)\n      \t\treturn -1;\n      \n    @@ streaming.c: static int open_istream_incore(struct odb_read_stream **out,\n     -\t\t\t\t     enum object_type *type,\n     -\t\t\t\t     unsigned long *size,\n     -\t\t\t\t     struct stream_filter *filter)\n    -+struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n    -+\t\t\t\t\t       const struct object_id *oid,\n    -+\t\t\t\t\t       enum object_type *type,\n    -+\t\t\t\t\t       unsigned long *size,\n    -+\t\t\t\t\t       struct stream_filter *filter)\n    ++struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n    ++\t\t\t\t\t     const struct object_id *oid,\n    ++\t\t\t\t\t     enum object_type *type,\n    ++\t\t\t\t\t     unsigned long *size,\n    ++\t\t\t\t\t     struct stream_filter *filter)\n      {\n      \tstruct odb_read_stream *st;\n     -\tconst struct object_id *real = lookup_replace_object(r, oid);\n    +-\tint ret = istream_source(&st, r, real);\n     +\tconst struct object_id *real = lookup_replace_object(odb->repo, oid);\n    - \tint ret;\n    ++\tint ret = istream_source(&st, odb, real);\n      \n    --\tret = istream_source(&st, r, real);\n    -+\tret = istream_source(&st, odb, real);\n      \tif (ret)\n      \t\treturn NULL;\n    - \n     @@ streaming.c: struct odb_read_stream *open_istream(struct repository *r,\n      \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n      \t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n    @@ streaming.c: int odb_stream_blob_to_fd(struct object_database *odb,\n      \tint result = -1;\n      \n     -\tst = open_istream(odb->repo, oid, &type, &sz, filter);\n    -+\tst = odb_read_object_stream(odb, oid, &type, &sz, filter);\n    ++\tst = odb_read_stream_open(odb, oid, &type, &sz, filter);\n      \tif (!st) {\n      \t\tif (filter)\n      \t\t\tfree_stream_filter(filter);\n    @@ streaming.h: struct odb_read_stream {\n     + *\n     + * Returns the stream on success, a `NULL` pointer otherwise.\n     + */\n    -+struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n    -+\t\t\t\t\t       const struct object_id *oid,\n    -+\t\t\t\t\t       enum object_type *type,\n    -+\t\t\t\t\t       unsigned long *size,\n    -+\t\t\t\t\t       struct stream_filter *filter);\n    ++struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n    ++\t\t\t\t\t     const struct object_id *oid,\n    ++\t\t\t\t\t     enum object_type *type,\n    ++\t\t\t\t\t     unsigned long *size,\n    ++\t\t\t\t\t     struct stream_filter *filter);\n     +\n     +/*\n     + * Close the given read stream and release all resources associated with it.\n18:  9fc79d10fd = 18:  e8c4e1931c streaming: move into object database subsystem\n19:  aab61d5697 ! 19:  f8e31ef59f streaming: drop redundant type and size pointers\n    @@ archive-tar.c: static void write_trailer(void)\n      \tchar buf[BLOCKSIZE];\n      \tssize_t readlen;\n      \n    --\tst = odb_read_object_stream(r->objects, oid, &type, &sz, NULL);\n    -+\tst = odb_read_object_stream(r->objects, oid, NULL);\n    +-\tst = odb_read_stream_open(r->objects, oid, &type, &sz, NULL);\n    ++\tst = odb_read_stream_open(r->objects, oid, NULL);\n      \tif (!st)\n      \t\treturn error(_(\"cannot stream blob %s\"), oid_to_hex(oid));\n      \tfor (;;) {\n    @@ archive-zip.c: static int write_zip_entry(struct archiver_args *args,\n      \n      \t\tif (!buffer) {\n     -\t\t\tenum object_type type;\n    --\t\t\tstream = odb_read_object_stream(args->repo->objects, oid,\n    --\t\t\t\t\t\t\t&type, &size, NULL);\n    -+\t\t\tstream = odb_read_object_stream(args->repo->objects, oid, NULL);\n    +-\t\t\tstream = odb_read_stream_open(args->repo->objects, oid,\n    +-\t\t\t\t\t\t      &type, &size, NULL);\n    ++\t\t\tstream = odb_read_stream_open(args->repo->objects, oid, NULL);\n      \t\t\tif (!stream)\n      \t\t\t\treturn error(_(\"cannot stream blob %s\"),\n      \t\t\t\t\t     oid_to_hex(oid));\n    @@ builtin/index-pack.c: static int check_collison(struct object_entry *entry)\n      \n      \tmemset(&data, 0, sizeof(data));\n      \tdata.entry = entry;\n    --\tdata.st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n    --\t\t\t\t\t &type, &size, NULL);\n    -+\tdata.st = odb_read_object_stream(the_repository->objects, &entry->idx.oid, NULL);\n    +-\tdata.st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n    +-\t\t\t\t       &type, &size, NULL);\n    ++\tdata.st = odb_read_stream_open(the_repository->objects, &entry->idx.oid, NULL);\n      \tif (!data.st)\n      \t\treturn -1;\n     -\tif (size != entry->size || type != entry->type)\n    @@ builtin/pack-objects.c\n     @@ builtin/pack-objects.c: static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n      \t\t    oe_size_greater_than(&to_pack, entry,\n      \t\t\t\t\t repo_settings_get_big_file_threshold(the_repository)) &&\n    - \t\t    (st = odb_read_object_stream(the_repository->objects, &entry->idx.oid,\n    --\t\t\t\t\t\t &type, &size, NULL)) != NULL)\n    -+\t\t\t\t\t\t NULL)) != NULL) {\n    + \t\t    (st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n    +-\t\t\t\t\t       &type, &size, NULL)) != NULL)\n    ++\t\t\t\t\t       NULL)) != NULL) {\n      \t\t\tbuf = NULL;\n     -\t\telse {\n     +\t\t\ttype = st->type;\n    @@ object-file.c: int check_object_signature(struct repository *r, const struct obj\n      \tchar hdr[MAX_HEADER_LEN];\n      \tint hdrlen;\n      \n    --\tst = odb_read_object_stream(r->objects, oid, &obj_type, &size, NULL);\n    -+\tst = odb_read_object_stream(r->objects, oid, NULL);\n    +-\tst = odb_read_stream_open(r->objects, oid, &obj_type, &size, NULL);\n    ++\tst = odb_read_stream_open(r->objects, oid, NULL);\n      \tif (!st)\n      \t\treturn -1;\n      \n    @@ object-file.c: int check_object_signature(struct repository *r, const struct obj\n      ## odb/streaming.c ##\n     @@ odb/streaming.c: ssize_t odb_read_stream_read(struct odb_read_stream *st, void *buf, size_t sz)\n      \n    - struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n    - \t\t\t\t\t       const struct object_id *oid,\n    --\t\t\t\t\t       enum object_type *type,\n    --\t\t\t\t\t       unsigned long *size,\n    - \t\t\t\t\t       struct stream_filter *filter)\n    + struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n    + \t\t\t\t\t     const struct object_id *oid,\n    +-\t\t\t\t\t     enum object_type *type,\n    +-\t\t\t\t\t     unsigned long *size,\n    + \t\t\t\t\t     struct stream_filter *filter)\n      {\n      \tstruct odb_read_stream *st;\n    -@@ odb/streaming.c: struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n    +@@ odb/streaming.c: struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n      \t\tst = nst;\n      \t}\n      \n    @@ odb/streaming.c: int odb_stream_blob_to_fd(struct object_database *odb,\n      \tssize_t kept = 0;\n      \tint result = -1;\n      \n    --\tst = odb_read_object_stream(odb, oid, &type, &sz, filter);\n    -+\tst = odb_read_object_stream(odb, oid, filter);\n    +-\tst = odb_read_stream_open(odb, oid, &type, &sz, filter);\n    ++\tst = odb_read_stream_open(odb, oid, filter);\n      \tif (!st) {\n      \t\tif (filter)\n      \t\t\tfree_stream_filter(filter);\n    @@ odb/streaming.h: struct odb_read_stream {\n       *\n       * Returns the stream on success, a `NULL` pointer otherwise.\n       */\n    - struct odb_read_stream *odb_read_object_stream(struct object_database *odb,\n    - \t\t\t\t\t       const struct object_id *oid,\n    --\t\t\t\t\t       enum object_type *type,\n    --\t\t\t\t\t       unsigned long *size,\n    - \t\t\t\t\t       struct stream_filter *filter);\n    + struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n    + \t\t\t\t\t     const struct object_id *oid,\n    +-\t\t\t\t\t     enum object_type *type,\n    +-\t\t\t\t\t     unsigned long *size,\n    + \t\t\t\t\t     struct stream_filter *filter);\n      \n      /*\n\n---\nbase-commit: 899e578b5b7c020aec806bd694adf2563f62843c\nchange-id: 20251107-b4-pks-odb-read-stream-7ea7f0e0a8f4\n\n"},{"id":"531182","messageId":"20251123-b4-pks-odb-read-stream-v3-1-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 01/19] streaming: rename `git_istream` into `odb_read_stream`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:26Z","receivedAt":"2025-11-23T18:59:50Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In the following patches we are about to make the `git_istream` more\ngeneric so that it becomes fully controlled by the specific object\nsource that wants to create it. As part of these refactorings we'll\nfully move the structure into the object database subsystem.\n\nPrepare for this change by renaming the structure from `git_istream`\nto `odb_read_stream`. This mirrors the `odb_write_stream` structure that\nwe already have.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  2 +-\n archive-zip.c          |  2 +-\n builtin/index-pack.c   |  2 +-\n builtin/pack-objects.c |  4 ++--\n object-file.c          |  2 +-\n streaming.c            | 62 +++++++++++++++++++++++++-------------------------\n streaming.h            | 12 +++++-----\n 7 files changed, 43 insertions(+), 43 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 73b63ddc41..dc1eda09e0 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -129,7 +129,7 @@ static void write_trailer(void)\n  */\n static int stream_blocked(struct repository *r, const struct object_id *oid)\n {\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tenum object_type type;\n \tunsigned long sz;\n \tchar buf[BLOCKSIZE];\ndiff --git a/archive-zip.c b/archive-zip.c\nindex bea5bdd43d..40a9c93ff9 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -309,7 +309,7 @@ static int write_zip_entry(struct archiver_args *args,\n \tenum zip_method method;\n \tunsigned char *out;\n \tvoid *deflated = NULL;\n-\tstruct git_istream *stream = NULL;\n+\tstruct odb_read_stream *stream = NULL;\n \tunsigned long flags = 0;\n \tint is_binary = -1;\n \tconst char *path_without_prefix = path + args->baselen;\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 2b78ba7fe4..5f90f12f92 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -762,7 +762,7 @@ static void find_ref_delta_children(const struct object_id *oid,\n \n struct compare_data {\n \tstruct object_entry *entry;\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tunsigned char *buf;\n \tunsigned long buf_size;\n };\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 69e80b1443..c693d948e1 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -404,7 +404,7 @@ static unsigned long do_compress(void **pptr, unsigned long size)\n \treturn stream.total_out;\n }\n \n-static unsigned long write_large_blob_data(struct git_istream *st, struct hashfile *f,\n+static unsigned long write_large_blob_data(struct odb_read_stream *st, struct hashfile *f,\n \t\t\t\t\t   const struct object_id *oid)\n {\n \tgit_zstream stream;\n@@ -513,7 +513,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \tunsigned hdrlen;\n \tenum object_type type;\n \tvoid *buf;\n-\tstruct git_istream *st = NULL;\n+\tstruct odb_read_stream *st = NULL;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \n \tif (!usable_delta) {\ndiff --git a/object-file.c b/object-file.c\nindex 811c569ed3..b62b21a452 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -134,7 +134,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tstruct object_id real_oid;\n \tunsigned long size;\n \tenum object_type obj_type;\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tstruct git_hash_ctx c;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\ndiff --git a/streaming.c b/streaming.c\nindex 00ad649ae3..1fb4b7c1c0 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -14,17 +14,17 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*open_istream_fn)(struct git_istream *,\n+typedef int (*open_istream_fn)(struct odb_read_stream *,\n \t\t\t       struct repository *,\n \t\t\t       const struct object_id *,\n \t\t\t       enum object_type *);\n-typedef int (*close_istream_fn)(struct git_istream *);\n-typedef ssize_t (*read_istream_fn)(struct git_istream *, char *, size_t);\n+typedef int (*close_istream_fn)(struct odb_read_stream *);\n+typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n #define FILTER_BUFFER (1024*16)\n \n struct filtered_istream {\n-\tstruct git_istream *upstream;\n+\tstruct odb_read_stream *upstream;\n \tstruct stream_filter *filter;\n \tchar ibuf[FILTER_BUFFER];\n \tchar obuf[FILTER_BUFFER];\n@@ -33,7 +33,7 @@ struct filtered_istream {\n \tint input_finished;\n };\n \n-struct git_istream {\n+struct odb_read_stream {\n \topen_istream_fn open;\n \tclose_istream_fn close;\n \tread_istream_fn read;\n@@ -71,7 +71,7 @@ struct git_istream {\n  *\n  *****************************************************************/\n \n-static void close_deflated_stream(struct git_istream *st)\n+static void close_deflated_stream(struct odb_read_stream *st)\n {\n \tif (st->z_state == z_used)\n \t\tgit_inflate_end(&st->z);\n@@ -84,13 +84,13 @@ static void close_deflated_stream(struct git_istream *st)\n  *\n  *****************************************************************/\n \n-static int close_istream_filtered(struct git_istream *st)\n+static int close_istream_filtered(struct odb_read_stream *st)\n {\n \tfree_stream_filter(st->u.filtered.filter);\n \treturn close_istream(st->u.filtered.upstream);\n }\n \n-static ssize_t read_istream_filtered(struct git_istream *st, char *buf,\n+static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n \t\t\t\t     size_t sz)\n {\n \tstruct filtered_istream *fs = &(st->u.filtered);\n@@ -150,10 +150,10 @@ static ssize_t read_istream_filtered(struct git_istream *st, char *buf,\n \treturn filled;\n }\n \n-static struct git_istream *attach_stream_filter(struct git_istream *st,\n-\t\t\t\t\t\tstruct stream_filter *filter)\n+static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n+\t\t\t\t\t\t    struct stream_filter *filter)\n {\n-\tstruct git_istream *ifs = xmalloc(sizeof(*ifs));\n+\tstruct odb_read_stream *ifs = xmalloc(sizeof(*ifs));\n \tstruct filtered_istream *fs = &(ifs->u.filtered);\n \n \tifs->close = close_istream_filtered;\n@@ -173,7 +173,7 @@ static struct git_istream *attach_stream_filter(struct git_istream *st,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_loose(struct git_istream *st, char *buf, size_t sz)\n+static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t sz)\n {\n \tsize_t total_read = 0;\n \n@@ -218,14 +218,14 @@ static ssize_t read_istream_loose(struct git_istream *st, char *buf, size_t sz)\n \treturn total_read;\n }\n \n-static int close_istream_loose(struct git_istream *st)\n+static int close_istream_loose(struct odb_read_stream *st)\n {\n \tclose_deflated_stream(st);\n \tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n \treturn 0;\n }\n \n-static int open_istream_loose(struct git_istream *st, struct repository *r,\n+static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \t\t\t      const struct object_id *oid,\n \t\t\t      enum object_type *type)\n {\n@@ -277,7 +277,7 @@ static int open_istream_loose(struct git_istream *st, struct repository *r,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_pack_non_delta(struct git_istream *st, char *buf,\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf,\n \t\t\t\t\t   size_t sz)\n {\n \tsize_t total_read = 0;\n@@ -336,13 +336,13 @@ static ssize_t read_istream_pack_non_delta(struct git_istream *st, char *buf,\n \treturn total_read;\n }\n \n-static int close_istream_pack_non_delta(struct git_istream *st)\n+static int close_istream_pack_non_delta(struct odb_read_stream *st)\n {\n \tclose_deflated_stream(st);\n \treturn 0;\n }\n \n-static int open_istream_pack_non_delta(struct git_istream *st,\n+static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n \t\t\t\t       const struct object_id *oid UNUSED,\n \t\t\t\t       enum object_type *type UNUSED)\n@@ -380,13 +380,13 @@ static int open_istream_pack_non_delta(struct git_istream *st,\n  *\n  *****************************************************************/\n \n-static int close_istream_incore(struct git_istream *st)\n+static int close_istream_incore(struct odb_read_stream *st)\n {\n \tfree(st->u.incore.buf);\n \treturn 0;\n }\n \n-static ssize_t read_istream_incore(struct git_istream *st, char *buf, size_t sz)\n+static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t sz)\n {\n \tsize_t read_size = sz;\n \tsize_t remainder = st->size - st->u.incore.read_ptr;\n@@ -400,7 +400,7 @@ static ssize_t read_istream_incore(struct git_istream *st, char *buf, size_t sz)\n \treturn read_size;\n }\n \n-static int open_istream_incore(struct git_istream *st, struct repository *r,\n+static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \t\t\t       const struct object_id *oid, enum object_type *type)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n@@ -420,7 +420,7 @@ static int open_istream_incore(struct git_istream *st, struct repository *r,\n  * static helpers variables and functions for users of streaming interface\n  *****************************************************************************/\n \n-static int istream_source(struct git_istream *st,\n+static int istream_source(struct odb_read_stream *st,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid,\n \t\t\t  enum object_type *type)\n@@ -458,25 +458,25 @@ static int istream_source(struct git_istream *st,\n  * Users of streaming interface\n  ****************************************************************/\n \n-int close_istream(struct git_istream *st)\n+int close_istream(struct odb_read_stream *st)\n {\n \tint r = st->close(st);\n \tfree(st);\n \treturn r;\n }\n \n-ssize_t read_istream(struct git_istream *st, void *buf, size_t sz)\n+ssize_t read_istream(struct odb_read_stream *st, void *buf, size_t sz)\n {\n \treturn st->read(st, buf, sz);\n }\n \n-struct git_istream *open_istream(struct repository *r,\n-\t\t\t\t const struct object_id *oid,\n-\t\t\t\t enum object_type *type,\n-\t\t\t\t unsigned long *size,\n-\t\t\t\t struct stream_filter *filter)\n+struct odb_read_stream *open_istream(struct repository *r,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     enum object_type *type,\n+\t\t\t\t     unsigned long *size,\n+\t\t\t\t     struct stream_filter *filter)\n {\n-\tstruct git_istream *st = xmalloc(sizeof(*st));\n+\tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n \tint ret = istream_source(st, r, real, type);\n \n@@ -493,7 +493,7 @@ struct git_istream *open_istream(struct repository *r,\n \t}\n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n-\t\tstruct git_istream *nst = attach_stream_filter(st, filter);\n+\t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n \t\tif (!nst) {\n \t\t\tclose_istream(st);\n \t\t\treturn NULL;\n@@ -508,7 +508,7 @@ struct git_istream *open_istream(struct repository *r,\n int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter *filter,\n \t\t      int can_seek)\n {\n-\tstruct git_istream *st;\n+\tstruct odb_read_stream *st;\n \tenum object_type type;\n \tunsigned long sz;\n \tssize_t kept = 0;\ndiff --git a/streaming.h b/streaming.h\nindex bd27f59e57..f5ff5d7ac9 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -7,14 +7,14 @@\n #include \"object.h\"\n \n /* opaque */\n-struct git_istream;\n+struct odb_read_stream;\n struct stream_filter;\n \n-struct git_istream *open_istream(struct repository *, const struct object_id *,\n-\t\t\t\t enum object_type *, unsigned long *,\n-\t\t\t\t struct stream_filter *);\n-int close_istream(struct git_istream *);\n-ssize_t read_istream(struct git_istream *, void *, size_t);\n+struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n+\t\t\t\t     enum object_type *, unsigned long *,\n+\t\t\t\t     struct stream_filter *);\n+int close_istream(struct odb_read_stream *);\n+ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n \n int stream_blob_to_fd(int fd, const struct object_id *, struct stream_filter *, int can_seek);\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531183","messageId":"20251123-b4-pks-odb-read-stream-v3-2-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 02/19] streaming: drop the `open()` callback function","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:27Z","receivedAt":"2025-11-23T18:59:54Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating a read stream we first populate the structure with the\nopen callback function and then subsequently call the function. This\nlayout is somewhat weird though:\n\n  - The structure needs to be allocated and partially populated with the\n    open function before we can properly initialize it.\n\n  - We only ever call the `open()` callback function right after having\n    populated the `struct odb_read_stream::open` member, and it's never\n    called thereafter again. So it is somewhat pointless to store the\n    callback in the first place.\n\nEspecially the first point creates a problem for us. In subsequent\ncommits we'll want to fully move construction of the read source into\nthe respective object sources. E.g., the loose object source will be the\none that is responsible for creating the structure. But this creates a\nproblem: if we first need to create the structure so that we can call\nthe source-specific callback we cannot fully handle creation of the\nstructure in the source itself.\n\nWe could of course work around that and have the loose object source\ncreate the structure and populate its `open()` callback, only. But\nthis doesn't really buy us anything due to the second bullet point\nabove.\n\nInstead, drop the callback entirely and refactor `istream_source()` so\nthat we open the streams immediately. This unblocks a subsequent step,\nwhere we'll also start to allocate the structure in the source-specific\nlogic.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 37 +++++++++++++++----------------------\n 1 file changed, 15 insertions(+), 22 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 1fb4b7c1c0..1bb3f393b8 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -14,10 +14,6 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*open_istream_fn)(struct odb_read_stream *,\n-\t\t\t       struct repository *,\n-\t\t\t       const struct object_id *,\n-\t\t\t       enum object_type *);\n typedef int (*close_istream_fn)(struct odb_read_stream *);\n typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n@@ -34,7 +30,6 @@ struct filtered_istream {\n };\n \n struct odb_read_stream {\n-\topen_istream_fn open;\n \tclose_istream_fn close;\n \tread_istream_fn read;\n \n@@ -437,21 +432,25 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tst->open = open_istream_loose;\n+\t\tif (open_istream_loose(st, r, oid, type) < 0)\n+\t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n-\t\tif (!oi.u.packed.is_delta &&\n-\t\t    repo_settings_get_big_file_threshold(the_repository) < size) {\n-\t\t\tst->u.in_pack.pack = oi.u.packed.pack;\n-\t\t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\t\tst->open = open_istream_pack_non_delta;\n-\t\t\treturn 0;\n-\t\t}\n-\t\t/* fallthru */\n-\tdefault:\n-\t\tst->open = open_istream_incore;\n+\t\tif (oi.u.packed.is_delta ||\n+\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\t\tbreak;\n+\n+\t\tst->u.in_pack.pack = oi.u.packed.pack;\n+\t\tst->u.in_pack.pos = oi.u.packed.offset;\n+\t\tif (open_istream_pack_non_delta(st, r, oid, type) < 0)\n+\t\t\tbreak;\n+\n \t\treturn 0;\n+\tdefault:\n+\t\tbreak;\n \t}\n+\n+\treturn open_istream_incore(st, r, oid, type);\n }\n \n /****************************************************************\n@@ -485,12 +484,6 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t\treturn NULL;\n \t}\n \n-\tif (st->open(st, r, real, type)) {\n-\t\tif (open_istream_incore(st, r, real, type)) {\n-\t\t\tfree(st);\n-\t\t\treturn NULL;\n-\t\t}\n-\t}\n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n \t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531184","messageId":"20251123-b4-pks-odb-read-stream-v3-3-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 03/19] streaming: propagate final object type via the stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:28Z","receivedAt":"2025-11-23T18:59:56Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When opening the read stream for a specific object the caller is also\nexpected to pass in a pointer to the object type. This type is passed\ndown via multiple levels and will eventually be populated with the type\nof the looked-up object.\n\nThe way we propagate down the pointer though is somewhat non-obvious.\nWhile `istream_source()` still expects the pointer and looks it up via\n`odb_read_object_info_extended()`, we also pass it down even further\ninto the format-specific callbacks that perform another lookup. This is\nquite confusing overall.\n\nRefactor the code so that the responsibility to populate the object type\nrests solely with the format-specific callbacks. This will allow us to\ndrop the call to `odb_read_object_info_extended()` in `istream_source()`\nentirely in a subsequent patch.\n\nFurthermore, instead of propagating the type via an in-pointer, we now\npropagate the type via a new field in the object stream. It already has\na `size` field, so it's only natural to have a second field that\ncontains the object type.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 30 +++++++++++++++---------------\n 1 file changed, 15 insertions(+), 15 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 1bb3f393b8..665624ddc0 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -33,6 +33,7 @@ struct odb_read_stream {\n \tclose_istream_fn close;\n \tread_istream_fn read;\n \n+\tenum object_type type;\n \tunsigned long size; /* inflated size of full object */\n \tgit_zstream z;\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n@@ -159,6 +160,7 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \tfs->o_end = fs->o_ptr = 0;\n \tfs->input_finished = 0;\n \tifs->size = -1; /* unknown */\n+\tifs->type = st->type;\n \treturn ifs;\n }\n \n@@ -221,14 +223,13 @@ static int close_istream_loose(struct odb_read_stream *st)\n }\n \n static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n-\t\t\t      const struct object_id *oid,\n-\t\t\t      enum object_type *type)\n+\t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_source *source;\n \n \toi.sizep = &st->size;\n-\toi.typep = type;\n+\toi.typep = &st->type;\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next) {\n@@ -249,7 +250,7 @@ static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \tcase ULHR_TOO_LONG:\n \t\tgoto error;\n \t}\n-\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || *type < 0)\n+\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n \t\tgoto error;\n \n \tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n@@ -339,8 +340,7 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \n static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED,\n-\t\t\t\t       enum object_type *type UNUSED)\n+\t\t\t\t       const struct object_id *oid UNUSED)\n {\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n@@ -361,6 +361,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n+\tst->type = in_pack_type;\n \tst->z_state = z_unused;\n \tst->close = close_istream_pack_non_delta;\n \tst->read = read_istream_pack_non_delta;\n@@ -396,7 +397,7 @@ static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t\n }\n \n static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n-\t\t\t       const struct object_id *oid, enum object_type *type)\n+\t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \n@@ -404,7 +405,7 @@ static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \tst->close = close_istream_incore;\n \tst->read = read_istream_incore;\n \n-\toi.typep = type;\n+\toi.typep = &st->type;\n \toi.sizep = &st->size;\n \toi.contentp = (void **)&st->u.incore.buf;\n \treturn odb_read_object_info_extended(r->objects, oid, &oi,\n@@ -417,14 +418,12 @@ static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n \n static int istream_source(struct odb_read_stream *st,\n \t\t\t  struct repository *r,\n-\t\t\t  const struct object_id *oid,\n-\t\t\t  enum object_type *type)\n+\t\t\t  const struct object_id *oid)\n {\n \tunsigned long size;\n \tint status;\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \n-\toi.typep = type;\n \toi.sizep = &size;\n \tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n \tif (status < 0)\n@@ -432,7 +431,7 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tif (open_istream_loose(st, r, oid, type) < 0)\n+\t\tif (open_istream_loose(st, r, oid) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n@@ -442,7 +441,7 @@ static int istream_source(struct odb_read_stream *st,\n \n \t\tst->u.in_pack.pack = oi.u.packed.pack;\n \t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\tif (open_istream_pack_non_delta(st, r, oid, type) < 0)\n+\t\tif (open_istream_pack_non_delta(st, r, oid) < 0)\n \t\t\tbreak;\n \n \t\treturn 0;\n@@ -450,7 +449,7 @@ static int istream_source(struct odb_read_stream *st,\n \t\tbreak;\n \t}\n \n-\treturn open_istream_incore(st, r, oid, type);\n+\treturn open_istream_incore(st, r, oid);\n }\n \n /****************************************************************\n@@ -477,7 +476,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n {\n \tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n-\tint ret = istream_source(st, r, real, type);\n+\tint ret = istream_source(st, r, real);\n \n \tif (ret) {\n \t\tfree(st);\n@@ -495,6 +494,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t}\n \n \t*size = st->size;\n+\t*type = st->type;\n \treturn st;\n }\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531185","messageId":"20251123-b4-pks-odb-read-stream-v3-4-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 04/19] streaming: explicitly pass packfile info when streaming a packed object","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:29Z","receivedAt":"2025-11-23T19:00:00Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When streaming a packed object we first populate the stream with\ninformation about the pack that contains the object before calling\n`open_istream_pack_non_delta()`. This is done because we have already\nlooked up both the pack and the object's offset, so it would be a waste\nof time to look up this information again.\n\nBut the way this is done makes for a somewhat awkward calling interface,\nas the caller now needs to be aware of how exactly the function itself\nbehaves.\n\nRefactor the code so that we instead explicitly pass the packfile info\ninto `open_istream_pack_non_delta()`. This makes the calling convention\nexplicit, but more importantly this allows us to refactor the function\nso that it becomes its responsibility to allocate the stream itself in a\nsubsequent patch.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 20 ++++++++++----------\n 1 file changed, 10 insertions(+), 10 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 665624ddc0..bf277daadd 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -340,16 +340,18 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \n static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED)\n+\t\t\t\t       const struct object_id *oid UNUSED,\n+\t\t\t\t       struct packed_git *pack,\n+\t\t\t\t       off_t offset)\n {\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n \n \twindow = NULL;\n \n-\tin_pack_type = unpack_object_header(st->u.in_pack.pack,\n+\tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n-\t\t\t\t\t    &st->u.in_pack.pos,\n+\t\t\t\t\t    &offset,\n \t\t\t\t\t    &st->size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n@@ -365,6 +367,8 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tst->z_state = z_unused;\n \tst->close = close_istream_pack_non_delta;\n \tst->read = read_istream_pack_non_delta;\n+\tst->u.in_pack.pack = pack;\n+\tst->u.in_pack.pos = offset;\n \n \treturn 0;\n }\n@@ -436,14 +440,10 @@ static int istream_source(struct odb_read_stream *st,\n \t\treturn 0;\n \tcase OI_PACKED:\n \t\tif (oi.u.packed.is_delta ||\n-\t\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n+\t\t    open_istream_pack_non_delta(st, r, oid, oi.u.packed.pack,\n+\t\t\t\t\t\toi.u.packed.offset) < 0)\n \t\t\tbreak;\n-\n-\t\tst->u.in_pack.pack = oi.u.packed.pack;\n-\t\tst->u.in_pack.pos = oi.u.packed.offset;\n-\t\tif (open_istream_pack_non_delta(st, r, oid) < 0)\n-\t\t\tbreak;\n-\n \t\treturn 0;\n \tdefault:\n \t\tbreak;\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531186","messageId":"20251123-b4-pks-odb-read-stream-v3-5-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 05/19] streaming: allocate stream inside the backend-specific logic","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:30Z","receivedAt":"2025-11-23T19:00:04Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating a new stream we first allocate it and then call into\nbackend-specific logic to populate the stream. This design requires that\nthe stream itself contains a `union` with backend-specific members that\nthen ultimately get populated by the backend-specific logic.\n\nThis works, but it's awkward in the context of pluggable object\ndatabases. Each backend will need its own member in that union, and as\nthe structure itself is completely opaque (it's only defined in\n\"streaming.c\") it also has the consequence that we must have the logic\nthat is specific to backends in \"streaming.c\".\n\nIdeally though, the infrastructure would be reversed: we have a generic\n`struct odb_read_stream` and some helper functions in \"streaming.c\",\nwhereas the backend-specific logic sits in the backend's subsystem\nitself.\n\nThis can be realized by using a design that is similar to how we handle\nreference databases: instead of having a union of members, we instead\nhave backend-specific structures with a `struct odb_read_stream base`\nas its first member. The backends would thus hand out the pointer to the\nbase, but internally they know to cast back to the backend-specific\ntype.\n\nThis means though that we need to allocate different structures\ndepending on the backend. To prepare for this, move allocation of the\nstructure into the backend-specific functions that open a new stream.\nSubsequent commits will then create those new backend-specific structs.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 103 ++++++++++++++++++++++++++++++++++++++----------------------\n 1 file changed, 65 insertions(+), 38 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex bf277daadd..a2c2d88738 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -222,27 +222,34 @@ static int close_istream_loose(struct odb_read_stream *st)\n \treturn 0;\n }\n \n-static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n+static int open_istream_loose(struct odb_read_stream **out,\n+\t\t\t      struct repository *r,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_read_stream *st;\n \tstruct odb_source *source;\n-\n-\toi.sizep = &st->size;\n-\toi.typep = &st->type;\n+\tunsigned long mapsize;\n+\tvoid *mapped;\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next) {\n-\t\tst->u.loose.mapped = odb_source_loose_map_object(source, oid,\n-\t\t\t\t\t\t\t\t &st->u.loose.mapsize);\n-\t\tif (st->u.loose.mapped)\n+\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n+\t\tif (mapped)\n \t\t\tbreak;\n \t}\n-\tif (!st->u.loose.mapped)\n+\tif (!mapped)\n \t\treturn -1;\n \n-\tswitch (unpack_loose_header(&st->z, st->u.loose.mapped,\n-\t\t\t\t    st->u.loose.mapsize, st->u.loose.hdr,\n+\t/*\n+\t * Note: we must allocate this structure early even though we may still\n+\t * fail. This is because we need to initialize the zlib stream, and it\n+\t * is not possible to copy the stream around after the fact because it\n+\t * has self-referencing pointers.\n+\t */\n+\tCALLOC_ARRAY(st, 1);\n+\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->u.loose.hdr,\n \t\t\t\t    sizeof(st->u.loose.hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n@@ -250,19 +257,28 @@ static int open_istream_loose(struct odb_read_stream *st, struct repository *r,\n \tcase ULHR_TOO_LONG:\n \t\tgoto error;\n \t}\n+\n+\toi.sizep = &st->size;\n+\toi.typep = &st->type;\n+\n \tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n \t\tgoto error;\n \n+\tst->u.loose.mapped = mapped;\n+\tst->u.loose.mapsize = mapsize;\n \tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n \tst->u.loose.hdr_avail = st->z.total_out;\n \tst->z_state = z_used;\n \tst->close = close_istream_loose;\n \tst->read = read_istream_loose;\n \n+\t*out = st;\n+\n \treturn 0;\n error:\n \tgit_inflate_end(&st->z);\n \tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tfree(st);\n \treturn -1;\n }\n \n@@ -338,12 +354,16 @@ static int close_istream_pack_non_delta(struct odb_read_stream *st)\n \treturn 0;\n }\n \n-static int open_istream_pack_non_delta(struct odb_read_stream *st,\n+static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \t\t\t\t       struct repository *r UNUSED,\n \t\t\t\t       const struct object_id *oid UNUSED,\n \t\t\t\t       struct packed_git *pack,\n \t\t\t\t       off_t offset)\n {\n+\tstruct odb_read_stream stream = {\n+\t\t.close = close_istream_pack_non_delta,\n+\t\t.read = read_istream_pack_non_delta,\n+\t};\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n \n@@ -352,7 +372,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n \t\t\t\t\t    &offset,\n-\t\t\t\t\t    &st->size);\n+\t\t\t\t\t    &stream.size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n \tdefault:\n@@ -363,12 +383,13 @@ static int open_istream_pack_non_delta(struct odb_read_stream *st,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n-\tst->type = in_pack_type;\n-\tst->z_state = z_unused;\n-\tst->close = close_istream_pack_non_delta;\n-\tst->read = read_istream_pack_non_delta;\n-\tst->u.in_pack.pack = pack;\n-\tst->u.in_pack.pos = offset;\n+\tstream.type = in_pack_type;\n+\tstream.z_state = z_unused;\n+\tstream.u.in_pack.pack = pack;\n+\tstream.u.in_pack.pos = offset;\n+\n+\tCALLOC_ARRAY(*out, 1);\n+\t**out = stream;\n \n \treturn 0;\n }\n@@ -400,27 +421,35 @@ static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t\n \treturn read_size;\n }\n \n-static int open_istream_incore(struct odb_read_stream *st, struct repository *r,\n+static int open_istream_incore(struct odb_read_stream **out,\n+\t\t\t       struct repository *r,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\n-\tst->u.incore.read_ptr = 0;\n-\tst->close = close_istream_incore;\n-\tst->read = read_istream_incore;\n-\n-\toi.typep = &st->type;\n-\toi.sizep = &st->size;\n-\toi.contentp = (void **)&st->u.incore.buf;\n-\treturn odb_read_object_info_extended(r->objects, oid, &oi,\n-\t\t\t\t\t     OBJECT_INFO_DIE_IF_CORRUPT);\n+\tstruct odb_read_stream stream = {\n+\t\t.close = close_istream_incore,\n+\t\t.read = read_istream_incore,\n+\t};\n+\tint ret;\n+\n+\toi.typep = &stream.type;\n+\toi.sizep = &stream.size;\n+\toi.contentp = (void **)&stream.u.incore.buf;\n+\tret = odb_read_object_info_extended(r->objects, oid, &oi,\n+\t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\tCALLOC_ARRAY(*out, 1);\n+\t**out = stream;\n+\treturn 0;\n }\n \n /*****************************************************************************\n  * static helpers variables and functions for users of streaming interface\n  *****************************************************************************/\n \n-static int istream_source(struct odb_read_stream *st,\n+static int istream_source(struct odb_read_stream **out,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid)\n {\n@@ -435,13 +464,13 @@ static int istream_source(struct odb_read_stream *st,\n \n \tswitch (oi.whence) {\n \tcase OI_LOOSE:\n-\t\tif (open_istream_loose(st, r, oid) < 0)\n+\t\tif (open_istream_loose(out, r, oid) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n \tcase OI_PACKED:\n \t\tif (oi.u.packed.is_delta ||\n \t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n-\t\t    open_istream_pack_non_delta(st, r, oid, oi.u.packed.pack,\n+\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n \t\t\t\t\t\toi.u.packed.offset) < 0)\n \t\t\tbreak;\n \t\treturn 0;\n@@ -449,7 +478,7 @@ static int istream_source(struct odb_read_stream *st,\n \t\tbreak;\n \t}\n \n-\treturn open_istream_incore(st, r, oid);\n+\treturn open_istream_incore(out, r, oid);\n }\n \n /****************************************************************\n@@ -474,14 +503,12 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t\t\t\t     unsigned long *size,\n \t\t\t\t     struct stream_filter *filter)\n {\n-\tstruct odb_read_stream *st = xmalloc(sizeof(*st));\n+\tstruct odb_read_stream *st;\n \tconst struct object_id *real = lookup_replace_object(r, oid);\n-\tint ret = istream_source(st, r, real);\n+\tint ret = istream_source(&st, r, real);\n \n-\tif (ret) {\n-\t\tfree(st);\n+\tif (ret)\n \t\treturn NULL;\n-\t}\n \n \tif (filter) {\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531187","messageId":"20251123-b4-pks-odb-read-stream-v3-6-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 06/19] streaming: create structure for in-core object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:31Z","receivedAt":"2025-11-23T19:00:07Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for in-core object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 44 +++++++++++++++++++++++++-------------------\n 1 file changed, 25 insertions(+), 19 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex a2c2d88738..35307d7229 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,11 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tchar *buf; /* from odb_read_object_info_extended() */\n-\t\t\tunsigned long read_ptr;\n-\t\t} incore;\n-\n \t\tstruct {\n \t\t\tvoid *mapped;\n \t\t\tunsigned long mapsize;\n@@ -401,22 +396,30 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n  *\n  *****************************************************************/\n \n-static int close_istream_incore(struct odb_read_stream *st)\n+struct odb_incore_read_stream {\n+\tstruct odb_read_stream base;\n+\tchar *buf; /* from odb_read_object_info_extended() */\n+\tunsigned long read_ptr;\n+};\n+\n+static int close_istream_incore(struct odb_read_stream *_st)\n {\n-\tfree(st->u.incore.buf);\n+\tstruct odb_incore_read_stream *st = (struct odb_incore_read_stream *)_st;\n+\tfree(st->buf);\n \treturn 0;\n }\n \n-static ssize_t read_istream_incore(struct odb_read_stream *st, char *buf, size_t sz)\n+static ssize_t read_istream_incore(struct odb_read_stream *_st, char *buf, size_t sz)\n {\n+\tstruct odb_incore_read_stream *st = (struct odb_incore_read_stream *)_st;\n \tsize_t read_size = sz;\n-\tsize_t remainder = st->size - st->u.incore.read_ptr;\n+\tsize_t remainder = st->base.size - st->read_ptr;\n \n \tif (remainder <= read_size)\n \t\tread_size = remainder;\n \tif (read_size) {\n-\t\tmemcpy(buf, st->u.incore.buf + st->u.incore.read_ptr, read_size);\n-\t\tst->u.incore.read_ptr += read_size;\n+\t\tmemcpy(buf, st->buf + st->read_ptr, read_size);\n+\t\tst->read_ptr += read_size;\n \t}\n \treturn read_size;\n }\n@@ -426,22 +429,25 @@ static int open_istream_incore(struct odb_read_stream **out,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_read_stream stream = {\n-\t\t.close = close_istream_incore,\n-\t\t.read = read_istream_incore,\n+\tstruct odb_incore_read_stream stream = {\n+\t\t.base.close = close_istream_incore,\n+\t\t.base.read = read_istream_incore,\n \t};\n+\tstruct odb_incore_read_stream *st;\n \tint ret;\n \n-\toi.typep = &stream.type;\n-\toi.sizep = &stream.size;\n-\toi.contentp = (void **)&stream.u.incore.buf;\n+\toi.typep = &stream.base.type;\n+\toi.sizep = &stream.base.size;\n+\toi.contentp = (void **)&stream.buf;\n \tret = odb_read_object_info_extended(r->objects, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n \n-\tCALLOC_ARRAY(*out, 1);\n-\t**out = stream;\n+\tCALLOC_ARRAY(st, 1);\n+\t*st = stream;\n+\t*out = &st->base;\n+\n \treturn 0;\n }\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531188","messageId":"20251123-b4-pks-odb-read-stream-v3-7-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 07/19] streaming: create structure for loose object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:32Z","receivedAt":"2025-11-23T19:00:10Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for loose object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 85 ++++++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 44 insertions(+), 41 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 35307d7229..ac7b3026f5 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,14 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tvoid *mapped;\n-\t\t\tunsigned long mapsize;\n-\t\t\tchar hdr[32];\n-\t\t\tint hdr_avail;\n-\t\t\tint hdr_used;\n-\t\t} loose;\n-\n \t\tstruct {\n \t\t\tstruct packed_git *pack;\n \t\t\toff_t pos;\n@@ -165,11 +157,21 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t sz)\n+struct odb_loose_read_stream {\n+\tstruct odb_read_stream base;\n+\tvoid *mapped;\n+\tunsigned long mapsize;\n+\tchar hdr[32];\n+\tint hdr_avail;\n+\tint hdr_used;\n+};\n+\n+static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n {\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->z_state) {\n+\tswitch (st->base.z_state) {\n \tcase z_done:\n \t\treturn 0;\n \tcase z_error:\n@@ -178,42 +180,43 @@ static ssize_t read_istream_loose(struct odb_read_stream *st, char *buf, size_t\n \t\tbreak;\n \t}\n \n-\tif (st->u.loose.hdr_used < st->u.loose.hdr_avail) {\n-\t\tsize_t to_copy = st->u.loose.hdr_avail - st->u.loose.hdr_used;\n+\tif (st->hdr_used < st->hdr_avail) {\n+\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n \t\tif (sz < to_copy)\n \t\t\tto_copy = sz;\n-\t\tmemcpy(buf, st->u.loose.hdr + st->u.loose.hdr_used, to_copy);\n-\t\tst->u.loose.hdr_used += to_copy;\n+\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n+\t\tst->hdr_used += to_copy;\n \t\ttotal_read += to_copy;\n \t}\n \n \twhile (total_read < sz) {\n \t\tint status;\n \n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->base.z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n \n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_done;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_done;\n \t\t\tbreak;\n \t\t}\n \t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_error;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_error;\n \t\t\treturn -1;\n \t\t}\n \t}\n \treturn total_read;\n }\n \n-static int close_istream_loose(struct odb_read_stream *st)\n+static int close_istream_loose(struct odb_read_stream *_st)\n {\n-\tclose_deflated_stream(st);\n-\tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tclose_deflated_stream(&st->base);\n+\tmunmap(st->mapped, st->mapsize);\n \treturn 0;\n }\n \n@@ -222,7 +225,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_read_stream *st;\n+\tstruct odb_loose_read_stream *st;\n \tstruct odb_source *source;\n \tunsigned long mapsize;\n \tvoid *mapped;\n@@ -244,8 +247,8 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t */\n \tCALLOC_ARRAY(st, 1);\n \n-\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->u.loose.hdr,\n-\t\t\t\t    sizeof(st->u.loose.hdr))) {\n+\tswitch (unpack_loose_header(&st->base.z, mapped, mapsize, st->hdr,\n+\t\t\t\t    sizeof(st->hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n \tcase ULHR_BAD:\n@@ -253,26 +256,26 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t\tgoto error;\n \t}\n \n-\toi.sizep = &st->size;\n-\toi.typep = &st->type;\n+\toi.sizep = &st->base.size;\n+\toi.typep = &st->base.type;\n \n-\tif (parse_loose_header(st->u.loose.hdr, &oi) < 0 || st->type < 0)\n+\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n \t\tgoto error;\n \n-\tst->u.loose.mapped = mapped;\n-\tst->u.loose.mapsize = mapsize;\n-\tst->u.loose.hdr_used = strlen(st->u.loose.hdr) + 1;\n-\tst->u.loose.hdr_avail = st->z.total_out;\n-\tst->z_state = z_used;\n-\tst->close = close_istream_loose;\n-\tst->read = read_istream_loose;\n+\tst->mapped = mapped;\n+\tst->mapsize = mapsize;\n+\tst->hdr_used = strlen(st->hdr) + 1;\n+\tst->hdr_avail = st->base.z.total_out;\n+\tst->base.z_state = z_used;\n+\tst->base.close = close_istream_loose;\n+\tst->base.read = read_istream_loose;\n \n-\t*out = st;\n+\t*out = &st->base;\n \n \treturn 0;\n error:\n-\tgit_inflate_end(&st->z);\n-\tmunmap(st->u.loose.mapped, st->u.loose.mapsize);\n+\tgit_inflate_end(&st->base.z);\n+\tmunmap(st->mapped, st->mapsize);\n \tfree(st);\n \treturn -1;\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531189","messageId":"20251123-b4-pks-odb-read-stream-v3-8-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 08/19] streaming: create structure for packed object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:33Z","receivedAt":"2025-11-23T19:00:14Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for packed object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 75 ++++++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 40 insertions(+), 35 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex ac7b3026f5..788f04e83e 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -39,11 +39,6 @@ struct odb_read_stream {\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n \n \tunion {\n-\t\tstruct {\n-\t\t\tstruct packed_git *pack;\n-\t\t\toff_t pos;\n-\t\t} in_pack;\n-\n \t\tstruct filtered_istream filtered;\n \t} u;\n };\n@@ -287,16 +282,23 @@ static int open_istream_loose(struct odb_read_stream **out,\n  *\n  *****************************************************************/\n \n-static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf,\n+struct odb_packed_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct packed_git *pack;\n+\toff_t pos;\n+};\n+\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n \t\t\t\t\t   size_t sz)\n {\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->z_state) {\n+\tswitch (st->base.z_state) {\n \tcase z_unused:\n-\t\tmemset(&st->z, 0, sizeof(st->z));\n-\t\tgit_inflate_init(&st->z);\n-\t\tst->z_state = z_used;\n+\t\tmemset(&st->base.z, 0, sizeof(st->base.z));\n+\t\tgit_inflate_init(&st->base.z);\n+\t\tst->base.z_state = z_used;\n \t\tbreak;\n \tcase z_done:\n \t\treturn 0;\n@@ -311,21 +313,21 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf\n \t\tstruct pack_window *window = NULL;\n \t\tunsigned char *mapped;\n \n-\t\tmapped = use_pack(st->u.in_pack.pack, &window,\n-\t\t\t\t  st->u.in_pack.pos, &st->z.avail_in);\n+\t\tmapped = use_pack(st->pack, &window,\n+\t\t\t\t  st->pos, &st->base.z.avail_in);\n \n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tst->z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->base.z.avail_out = sz - total_read;\n+\t\tst->base.z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n \n-\t\tst->u.in_pack.pos += st->z.next_in - mapped;\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\tst->pos += st->base.z.next_in - mapped;\n+\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n \t\tunuse_pack(&window);\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_done;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_done;\n \t\t\tbreak;\n \t\t}\n \n@@ -338,17 +340,18 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *st, char *buf\n \t\t * or truncated), then use_pack() catches that and will die().\n \t\t */\n \t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = z_error;\n+\t\t\tgit_inflate_end(&st->base.z);\n+\t\t\tst->base.z_state = z_error;\n \t\t\treturn -1;\n \t\t}\n \t}\n \treturn total_read;\n }\n \n-static int close_istream_pack_non_delta(struct odb_read_stream *st)\n+static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n {\n-\tclose_deflated_stream(st);\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tclose_deflated_stream(&st->base);\n \treturn 0;\n }\n \n@@ -358,19 +361,17 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \t\t\t\t       struct packed_git *pack,\n \t\t\t\t       off_t offset)\n {\n-\tstruct odb_read_stream stream = {\n-\t\t.close = close_istream_pack_non_delta,\n-\t\t.read = read_istream_pack_non_delta,\n-\t};\n+\tstruct odb_packed_read_stream *stream;\n \tstruct pack_window *window;\n \tenum object_type in_pack_type;\n+\tsize_t size;\n \n \twindow = NULL;\n \n \tin_pack_type = unpack_object_header(pack,\n \t\t\t\t\t    &window,\n \t\t\t\t\t    &offset,\n-\t\t\t\t\t    &stream.size);\n+\t\t\t\t\t    &size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n \tdefault:\n@@ -381,13 +382,17 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tcase OBJ_TAG:\n \t\tbreak;\n \t}\n-\tstream.type = in_pack_type;\n-\tstream.z_state = z_unused;\n-\tstream.u.in_pack.pack = pack;\n-\tstream.u.in_pack.pos = offset;\n \n-\tCALLOC_ARRAY(*out, 1);\n-\t**out = stream;\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.close = close_istream_pack_non_delta;\n+\tstream->base.read = read_istream_pack_non_delta;\n+\tstream->base.type = in_pack_type;\n+\tstream->base.size = size;\n+\tstream->base.z_state = z_unused;\n+\tstream->pack = pack;\n+\tstream->pos = offset;\n+\n+\t*out = &stream->base;\n \n \treturn 0;\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531190","messageId":"20251123-b4-pks-odb-read-stream-v3-9-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 09/19] streaming: create structure for filtered object streams","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:34Z","receivedAt":"2025-11-23T19:00:16Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"As explained in a preceding commit, we want to get rid of the union of\nstream-type specific data in `struct odb_read_stream`. Create a new\nstructure for filtered object streams to move towards this design.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 54 +++++++++++++++++++++++++-----------------------------\n 1 file changed, 25 insertions(+), 29 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 788f04e83e..199cca5abb 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -19,16 +19,6 @@ typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n \n #define FILTER_BUFFER (1024*16)\n \n-struct filtered_istream {\n-\tstruct odb_read_stream *upstream;\n-\tstruct stream_filter *filter;\n-\tchar ibuf[FILTER_BUFFER];\n-\tchar obuf[FILTER_BUFFER];\n-\tint i_end, i_ptr;\n-\tint o_end, o_ptr;\n-\tint input_finished;\n-};\n-\n struct odb_read_stream {\n \tclose_istream_fn close;\n \tread_istream_fn read;\n@@ -37,10 +27,6 @@ struct odb_read_stream {\n \tunsigned long size; /* inflated size of full object */\n \tgit_zstream z;\n \tenum { z_unused, z_used, z_done, z_error } z_state;\n-\n-\tunion {\n-\t\tstruct filtered_istream filtered;\n-\t} u;\n };\n \n /*****************************************************************\n@@ -62,16 +48,28 @@ static void close_deflated_stream(struct odb_read_stream *st)\n  *\n  *****************************************************************/\n \n-static int close_istream_filtered(struct odb_read_stream *st)\n+struct odb_filtered_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct odb_read_stream *upstream;\n+\tstruct stream_filter *filter;\n+\tchar ibuf[FILTER_BUFFER];\n+\tchar obuf[FILTER_BUFFER];\n+\tint i_end, i_ptr;\n+\tint o_end, o_ptr;\n+\tint input_finished;\n+};\n+\n+static int close_istream_filtered(struct odb_read_stream *_fs)\n {\n-\tfree_stream_filter(st->u.filtered.filter);\n-\treturn close_istream(st->u.filtered.upstream);\n+\tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n+\tfree_stream_filter(fs->filter);\n+\treturn close_istream(fs->upstream);\n }\n \n-static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n+static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n \t\t\t\t     size_t sz)\n {\n-\tstruct filtered_istream *fs = &(st->u.filtered);\n+\tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n \tsize_t filled = 0;\n \n \twhile (sz) {\n@@ -131,19 +129,17 @@ static ssize_t read_istream_filtered(struct odb_read_stream *st, char *buf,\n static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \t\t\t\t\t\t    struct stream_filter *filter)\n {\n-\tstruct odb_read_stream *ifs = xmalloc(sizeof(*ifs));\n-\tstruct filtered_istream *fs = &(ifs->u.filtered);\n+\tstruct odb_filtered_read_stream *fs;\n \n-\tifs->close = close_istream_filtered;\n-\tifs->read = read_istream_filtered;\n+\tCALLOC_ARRAY(fs, 1);\n+\tfs->base.close = close_istream_filtered;\n+\tfs->base.read = read_istream_filtered;\n \tfs->upstream = st;\n \tfs->filter = filter;\n-\tfs->i_end = fs->i_ptr = 0;\n-\tfs->o_end = fs->o_ptr = 0;\n-\tfs->input_finished = 0;\n-\tifs->size = -1; /* unknown */\n-\tifs->type = st->type;\n-\treturn ifs;\n+\tfs->base.size = -1; /* unknown */\n+\tfs->base.type = st->type;\n+\n+\treturn &fs->base;\n }\n \n /*****************************************************************\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531191","messageId":"20251123-b4-pks-odb-read-stream-v3-10-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 10/19] streaming: move zlib stream into backends","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:35Z","receivedAt":"2025-11-23T19:00:20Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"While all backend-specific data is now contained in a backend-specific\nstructure, we still share the zlib stream across the loose and packed\nobjects.\n\nRefactor the code and move it into the specific structures so that we\nfully detangle the different backends from one another.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 104 ++++++++++++++++++++++++++++++------------------------------\n 1 file changed, 52 insertions(+), 52 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 199cca5abb..46fddaf2ca 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -25,23 +25,8 @@ struct odb_read_stream {\n \n \tenum object_type type;\n \tunsigned long size; /* inflated size of full object */\n-\tgit_zstream z;\n-\tenum { z_unused, z_used, z_done, z_error } z_state;\n };\n \n-/*****************************************************************\n- *\n- * Common helpers\n- *\n- *****************************************************************/\n-\n-static void close_deflated_stream(struct odb_read_stream *st)\n-{\n-\tif (st->z_state == z_used)\n-\t\tgit_inflate_end(&st->z);\n-}\n-\n-\n /*****************************************************************\n  *\n  * Filtered stream\n@@ -150,6 +135,12 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \n struct odb_loose_read_stream {\n \tstruct odb_read_stream base;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_LOOSE_READ_STREAM_INUSE,\n+\t\tODB_LOOSE_READ_STREAM_DONE,\n+\t\tODB_LOOSE_READ_STREAM_ERROR,\n+\t} z_state;\n \tvoid *mapped;\n \tunsigned long mapsize;\n \tchar hdr[32];\n@@ -162,10 +153,10 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n \tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->base.z_state) {\n-\tcase z_done:\n+\tswitch (st->z_state) {\n+\tcase ODB_LOOSE_READ_STREAM_DONE:\n \t\treturn 0;\n-\tcase z_error:\n+\tcase ODB_LOOSE_READ_STREAM_ERROR:\n \t\treturn -1;\n \tdefault:\n \t\tbreak;\n@@ -183,20 +174,20 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n \twhile (total_read < sz) {\n \t\tint status;\n \n-\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->base.z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n \n-\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_done;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n \t\t\tbreak;\n \t\t}\n \t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_error;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n \t\t\treturn -1;\n \t\t}\n \t}\n@@ -206,7 +197,8 @@ static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t\n static int close_istream_loose(struct odb_read_stream *_st)\n {\n \tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tclose_deflated_stream(&st->base);\n+\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n \tmunmap(st->mapped, st->mapsize);\n \treturn 0;\n }\n@@ -238,7 +230,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \t */\n \tCALLOC_ARRAY(st, 1);\n \n-\tswitch (unpack_loose_header(&st->base.z, mapped, mapsize, st->hdr,\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n \t\t\t\t    sizeof(st->hdr))) {\n \tcase ULHR_OK:\n \t\tbreak;\n@@ -256,8 +248,8 @@ static int open_istream_loose(struct odb_read_stream **out,\n \tst->mapped = mapped;\n \tst->mapsize = mapsize;\n \tst->hdr_used = strlen(st->hdr) + 1;\n-\tst->hdr_avail = st->base.z.total_out;\n-\tst->base.z_state = z_used;\n+\tst->hdr_avail = st->z.total_out;\n+\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n \tst->base.close = close_istream_loose;\n \tst->base.read = read_istream_loose;\n \n@@ -265,7 +257,7 @@ static int open_istream_loose(struct odb_read_stream **out,\n \n \treturn 0;\n error:\n-\tgit_inflate_end(&st->base.z);\n+\tgit_inflate_end(&st->z);\n \tmunmap(st->mapped, st->mapsize);\n \tfree(st);\n \treturn -1;\n@@ -281,6 +273,13 @@ static int open_istream_loose(struct odb_read_stream **out,\n struct odb_packed_read_stream {\n \tstruct odb_read_stream base;\n \tstruct packed_git *pack;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n+\t\tODB_PACKED_READ_STREAM_INUSE,\n+\t\tODB_PACKED_READ_STREAM_DONE,\n+\t\tODB_PACKED_READ_STREAM_ERROR,\n+\t} z_state;\n \toff_t pos;\n };\n \n@@ -290,17 +289,17 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n \tsize_t total_read = 0;\n \n-\tswitch (st->base.z_state) {\n-\tcase z_unused:\n-\t\tmemset(&st->base.z, 0, sizeof(st->base.z));\n-\t\tgit_inflate_init(&st->base.z);\n-\t\tst->base.z_state = z_used;\n+\tswitch (st->z_state) {\n+\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n+\t\tmemset(&st->z, 0, sizeof(st->z));\n+\t\tgit_inflate_init(&st->z);\n+\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n \t\tbreak;\n-\tcase z_done:\n+\tcase ODB_PACKED_READ_STREAM_DONE:\n \t\treturn 0;\n-\tcase z_error:\n+\tcase ODB_PACKED_READ_STREAM_ERROR:\n \t\treturn -1;\n-\tcase z_used:\n+\tcase ODB_PACKED_READ_STREAM_INUSE:\n \t\tbreak;\n \t}\n \n@@ -310,20 +309,20 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \t\tunsigned char *mapped;\n \n \t\tmapped = use_pack(st->pack, &window,\n-\t\t\t\t  st->pos, &st->base.z.avail_in);\n+\t\t\t\t  st->pos, &st->z.avail_in);\n \n-\t\tst->base.z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->base.z.avail_out = sz - total_read;\n-\t\tst->base.z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->base.z, Z_FINISH);\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tst->z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n \n-\t\tst->pos += st->base.z.next_in - mapped;\n-\t\ttotal_read = st->base.z.next_out - (unsigned char *)buf;\n+\t\tst->pos += st->z.next_in - mapped;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n \t\tunuse_pack(&window);\n \n \t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_done;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n \t\t\tbreak;\n \t\t}\n \n@@ -336,8 +335,8 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n \t\t * or truncated), then use_pack() catches that and will die().\n \t\t */\n \t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->base.z);\n-\t\t\tst->base.z_state = z_error;\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n \t\t\treturn -1;\n \t\t}\n \t}\n@@ -347,7 +346,8 @@ static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *bu\n static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n {\n \tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tclose_deflated_stream(&st->base);\n+\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n \treturn 0;\n }\n \n@@ -384,7 +384,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tstream->base.read = read_istream_pack_non_delta;\n \tstream->base.type = in_pack_type;\n \tstream->base.size = size;\n-\tstream->base.z_state = z_unused;\n+\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n \tstream->pack = pack;\n \tstream->pos = offset;\n \n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531192","messageId":"20251123-b4-pks-odb-read-stream-v3-11-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 11/19] packfile: introduce function to read object info from a store","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:36Z","receivedAt":"2025-11-23T19:00:23Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Extract the logic to read object info for a packed object from\n`do_oid_object_into_extended()` into a standalone function that operates\non the packfile store. This function will be used in a subsequent\ncommit.\n\nNote that this change allows us to make `find_pack_entry()` an internal\nimplementation detail. As a consequence though we have to move around\n`packfile_store_freshen_object()` so that it is defined after that\nfunction.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c      | 29 ++++---------------------\n packfile.c | 71 +++++++++++++++++++++++++++++++++++++++++++++++---------------\n packfile.h | 12 ++++++++++-\n 3 files changed, 69 insertions(+), 43 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 3ec21ef24e..f4cbee4b04 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -666,8 +666,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n {\n \tstatic struct object_info blank_oi = OBJECT_INFO_INIT;\n \tconst struct cached_object *co;\n-\tstruct pack_entry e;\n-\tint rtype;\n \tconst struct object_id *real = oid;\n \tint already_retried = 0;\n \n@@ -702,8 +700,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \twhile (1) {\n \t\tstruct odb_source *source;\n \n-\t\tif (find_pack_entry(odb->repo, real, &e))\n-\t\t\tbreak;\n+\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n+\t\t\treturn 0;\n \n \t\t/* Most likely it's a loose object. */\n \t\tfor (source = odb->sources; source; source = source->next)\n@@ -713,8 +711,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t/* Not a loose object; someone else may have just packed it. */\n \t\tif (!(flags & OBJECT_INFO_QUICK)) {\n \t\t\todb_reprepare(odb->repo->objects);\n-\t\t\tif (find_pack_entry(odb->repo, real, &e))\n-\t\t\t\tbreak;\n+\t\t\tif (!packfile_store_read_object_info(odb->packfiles, real, oi, flags))\n+\t\t\t\treturn 0;\n \t\t}\n \n \t\t/*\n@@ -747,25 +745,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t}\n \t\treturn -1;\n \t}\n-\n-\tif (oi == &blank_oi)\n-\t\t/*\n-\t\t * We know that the caller doesn't actually need the\n-\t\t * information below, so return early.\n-\t\t */\n-\t\treturn 0;\n-\trtype = packed_object_info(odb->repo, e.p, e.offset, oi);\n-\tif (rtype < 0) {\n-\t\tmark_bad_packed_object(e.p, real);\n-\t\treturn do_oid_object_info_extended(odb, real, oi, 0);\n-\t} else if (oi->whence == OI_PACKED) {\n-\t\toi->u.packed.offset = e.offset;\n-\t\toi->u.packed.pack = e.p;\n-\t\toi->u.packed.is_delta = (rtype == OBJ_REF_DELTA ||\n-\t\t\t\t\t rtype == OBJ_OFS_DELTA);\n-\t}\n-\n-\treturn 0;\n }\n \n static int oid_object_info_convert(struct repository *r,\ndiff --git a/packfile.c b/packfile.c\nindex 40f733dd23..b4bc40d895 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -819,22 +819,6 @@ struct packed_git *packfile_store_load_pack(struct packfile_store *store,\n \treturn p;\n }\n \n-int packfile_store_freshen_object(struct packfile_store *store,\n-\t\t\t\t  const struct object_id *oid)\n-{\n-\tstruct pack_entry e;\n-\tif (!find_pack_entry(store->odb->repo, oid, &e))\n-\t\treturn 0;\n-\tif (e.p->is_cruft)\n-\t\treturn 0;\n-\tif (e.p->freshened)\n-\t\treturn 1;\n-\tif (utime(e.p->pack_name, NULL))\n-\t\treturn 0;\n-\te.p->freshened = 1;\n-\treturn 1;\n-}\n-\n void (*report_garbage)(unsigned seen_bits, const char *path);\n \n static void report_helper(const struct string_list *list,\n@@ -2064,7 +2048,9 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static int find_pack_entry(struct repository *r,\n+\t\t\t   const struct object_id *oid,\n+\t\t\t   struct pack_entry *e)\n {\n \tstruct list_head *pos;\n \n@@ -2087,6 +2073,57 @@ int find_pack_entry(struct repository *r, const struct object_id *oid, struct pa\n \treturn 0;\n }\n \n+int packfile_store_freshen_object(struct packfile_store *store,\n+\t\t\t\t  const struct object_id *oid)\n+{\n+\tstruct pack_entry e;\n+\tif (!find_pack_entry(store->odb->repo, oid, &e))\n+\t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n+\tif (e.p->freshened)\n+\t\treturn 1;\n+\tif (utime(e.p->pack_name, NULL))\n+\t\treturn 0;\n+\te.p->freshened = 1;\n+\treturn 1;\n+}\n+\n+int packfile_store_read_object_info(struct packfile_store *store,\n+\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t    struct object_info *oi,\n+\t\t\t\t    unsigned flags UNUSED)\n+{\n+\tstatic struct object_info blank_oi = OBJECT_INFO_INIT;\n+\tstruct pack_entry e;\n+\tint rtype;\n+\n+\tif (!find_pack_entry(store->odb->repo, oid, &e))\n+\t\treturn 1;\n+\n+\t/*\n+\t * We know that the caller doesn't actually need the\n+\t * information below, so return early.\n+\t */\n+\tif (oi == &blank_oi)\n+\t\treturn 0;\n+\n+\trtype = packed_object_info(store->odb->repo, e.p, e.offset, oi);\n+\tif (rtype < 0) {\n+\t\tmark_bad_packed_object(e.p, oid);\n+\t\treturn -1;\n+\t}\n+\n+\tif (oi->whence == OI_PACKED) {\n+\t\toi->u.packed.offset = e.offset;\n+\t\toi->u.packed.pack = e.p;\n+\t\toi->u.packed.is_delta = (rtype == OBJ_REF_DELTA ||\n+\t\t\t\t\t rtype == OBJ_OFS_DELTA);\n+\t}\n+\n+\treturn 0;\n+}\n+\n static void maybe_invalidate_kept_pack_cache(struct repository *r,\n \t\t\t\t\t     unsigned flags)\n {\ndiff --git a/packfile.h b/packfile.h\nindex 58fcc88e20..0a98bddd81 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -144,6 +144,17 @@ void packfile_store_add_pack(struct packfile_store *store,\n #define repo_for_each_pack(repo, p) \\\n \tfor (p = packfile_store_get_packs(repo->objects->packfiles); p; p = p->next)\n \n+/*\n+ * Try to read the object identified by its ID from the object store and\n+ * populate the object info with its data. Returns 1 in case the object was\n+ * not found, 0 if it was and read successfully, and a negative error code in\n+ * case the object was corrupted.\n+ */\n+int packfile_store_read_object_info(struct packfile_store *store,\n+\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t    struct object_info *oi,\n+\t\t\t\t    unsigned flags);\n+\n /*\n  * Get all packs managed by the given store, including packfiles that are\n  * referenced by multi-pack indices.\n@@ -357,7 +368,6 @@ const struct packed_git *has_packed_and_bad(struct repository *, const struct ob\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\n  */\n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e);\n int find_kept_pack_entry(struct repository *r, const struct object_id *oid, unsigned flags, struct pack_entry *e);\n \n int has_object_pack(struct repository *r, const struct object_id *oid);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531193","messageId":"20251123-b4-pks-odb-read-stream-v3-12-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 12/19] streaming: rely on object sources to create object stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:37Z","receivedAt":"2025-11-23T19:00:26Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When creating an object stream we first look up the object info and, if\nit's present, we call into the respective backend that contains the\nobject to create a new stream for it.\n\nThis has the consequence that, for loose object source, we basically\niterate through the object sources twice: we first discover that the\nfile exists as a loose object in the first place by iterating through\nall sources. And, once we have discovered it, we again walk through all\nsources to try and map the object. The same issue will eventually also\nsurface once the packfile store becomes per-object-source.\n\nFurthermore, it feels rather pointless to first look up the object only\nto then try and read it.\n\nRefactor the logic to be centered around sources instead. Instead of\nfirst reading the object, we immediately ask the source to create the\nobject stream for us. If the object exists we get stream, otherwise\nwe'll try the next source.\n\nLike this we only have to iterate through sources once. But even more\nimportantly, this change also helps us to make the whole logic\npluggable. The object read stream subsystem does not need to be aware of\nthe different source backends anymore, but eventually it'll only have to\ncall the source's callback function.\n\nNote that at the current point in time we aren't fully there yet:\n\n  - The packfile store still sits on the object database level and is\n    thus agnostic of the sources.\n\n  - We still have to call into both the packfile store and the loose\n    object source.\n\nBut both of these issues will soon be addressed.\n\nThis refactoring results in a slight change to semantics: previously, it\nwas `odb_read_object_info_extended()` that picked the source for us, and\nit would have favored packed (non-deltified) objects over loose objects.\nAnd while we still favor packed over loose objects for a single source\nwith the new logic, we'll now favor a loose object from an earlier\nsource over a packed object from a later source.\n\nUltimately this shouldn't matter though: the stream doesn't indicate to\nthe caller which source it is from and whether it was created from a\npacked or loose object, so such details are opaque to the caller. And\nother than that we should be able to assume that two objects with the\nsame object ID should refer to the same content, so the streamed data\nwould be the same, too.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 65 +++++++++++++++++++++++--------------------------------------\n 1 file changed, 24 insertions(+), 41 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 46fddaf2ca..f0f7d31956 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -204,21 +204,15 @@ static int close_istream_loose(struct odb_read_stream *_st)\n }\n \n static int open_istream_loose(struct odb_read_stream **out,\n-\t\t\t      struct repository *r,\n+\t\t\t      struct odb_source *source,\n \t\t\t      const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_loose_read_stream *st;\n-\tstruct odb_source *source;\n \tunsigned long mapsize;\n \tvoid *mapped;\n \n-\todb_prepare_alternates(r->objects);\n-\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n-\t\tif (mapped)\n-\t\t\tbreak;\n-\t}\n+\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n \tif (!mapped)\n \t\treturn -1;\n \n@@ -352,21 +346,25 @@ static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n }\n \n static int open_istream_pack_non_delta(struct odb_read_stream **out,\n-\t\t\t\t       struct repository *r UNUSED,\n-\t\t\t\t       const struct object_id *oid UNUSED,\n-\t\t\t\t       struct packed_git *pack,\n-\t\t\t\t       off_t offset)\n+\t\t\t\t       struct object_database *odb,\n+\t\t\t\t       const struct object_id *oid)\n {\n \tstruct odb_packed_read_stream *stream;\n-\tstruct pack_window *window;\n+\tstruct pack_window *window = NULL;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n \tenum object_type in_pack_type;\n-\tsize_t size;\n+\tunsigned long size;\n \n-\twindow = NULL;\n+\toi.sizep = &size;\n+\n+\tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n+\t    oi.u.packed.is_delta ||\n+\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t\treturn -1;\n \n-\tin_pack_type = unpack_object_header(pack,\n+\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n \t\t\t\t\t    &window,\n-\t\t\t\t\t    &offset,\n+\t\t\t\t\t    &oi.u.packed.offset,\n \t\t\t\t\t    &size);\n \tunuse_pack(&window);\n \tswitch (in_pack_type) {\n@@ -385,8 +383,8 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \tstream->base.type = in_pack_type;\n \tstream->base.size = size;\n \tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n-\tstream->pack = pack;\n-\tstream->pos = offset;\n+\tstream->pack = oi.u.packed.pack;\n+\tstream->pos = oi.u.packed.offset;\n \n \t*out = &stream->base;\n \n@@ -463,30 +461,15 @@ static int istream_source(struct odb_read_stream **out,\n \t\t\t  struct repository *r,\n \t\t\t  const struct object_id *oid)\n {\n-\tunsigned long size;\n-\tint status;\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\n-\toi.sizep = &size;\n-\tstatus = odb_read_object_info_extended(r->objects, oid, &oi, 0);\n-\tif (status < 0)\n-\t\treturn status;\n+\tstruct odb_source *source;\n \n-\tswitch (oi.whence) {\n-\tcase OI_LOOSE:\n-\t\tif (open_istream_loose(out, r, oid) < 0)\n-\t\t\tbreak;\n-\t\treturn 0;\n-\tcase OI_PACKED:\n-\t\tif (oi.u.packed.is_delta ||\n-\t\t    repo_settings_get_big_file_threshold(the_repository) >= size ||\n-\t\t    open_istream_pack_non_delta(out, r, oid, oi.u.packed.pack,\n-\t\t\t\t\t\toi.u.packed.offset) < 0)\n-\t\t\tbreak;\n+\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n \t\treturn 0;\n-\tdefault:\n-\t\tbreak;\n-\t}\n+\n+\todb_prepare_alternates(r->objects);\n+\tfor (source = r->objects->sources; source; source = source->next)\n+\t\tif (!open_istream_loose(out, source, oid))\n+\t\t\treturn 0;\n \n \treturn open_istream_incore(out, r, oid);\n }\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531194","messageId":"20251123-b4-pks-odb-read-stream-v3-13-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 13/19] streaming: get rid of `the_repository`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:38Z","receivedAt":"2025-11-23T19:00:29Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Subsequent commits will move the backend-specific logic of object\nstreaming into their respective subsystems. These subsystems have gotten\nrid of `the_repository` already, but we still use it in two locations in\nthe streaming subsystem.\n\nPrepare for the move by fixing those two cases. Converting the logic in\n`open_istream_pack_non_delta()` is trivial as we already got the object\ndatabase as input.\n\nBut for `stream_blob_to_fd()` we have to add a new parameter to make it\naccessible. So, as we already have to adjust all callers anyway, rename\nthe function to `odb_stream_blob_to_fd()` to indicate it's part of the\nobject subsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/cat-file.c  |  2 +-\n builtin/fsck.c      |  3 ++-\n builtin/log.c       |  4 ++--\n entry.c             |  2 +-\n parallel-checkout.c |  3 ++-\n streaming.c         | 13 +++++++------\n streaming.h         | 18 +++++++++++++++++-\n 7 files changed, 32 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 983ecec837..120d626d66 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -95,7 +95,7 @@ static int filter_object(const char *path, unsigned mode,\n \n static int stream_blob(const struct object_id *oid)\n {\n-\tif (stream_blob_to_fd(1, oid, NULL, 0))\n+\tif (odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0))\n \t\tdie(\"unable to stream %s to stdout\", oid_to_hex(oid));\n \treturn 0;\n }\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex b1a650c673..1a348d43c2 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -340,7 +340,8 @@ static void check_unreachable_object(struct object *obj)\n \t\t\t}\n \t\t\tf = xfopen(filename, \"w\");\n \t\t\tif (obj->type == OBJ_BLOB) {\n-\t\t\t\tif (stream_blob_to_fd(fileno(f), &obj->oid, NULL, 1))\n+\t\t\t\tif (odb_stream_blob_to_fd(the_repository->objects, fileno(f),\n+\t\t\t\t\t\t\t  &obj->oid, NULL, 1))\n \t\t\t\t\tdie_errno(_(\"could not write '%s'\"), filename);\n \t\t\t} else\n \t\t\t\tfprintf(f, \"%s\\n\", describe_object(&obj->oid));\ndiff --git a/builtin/log.c b/builtin/log.c\nindex c8319b8af3..e7b83a6e00 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -584,7 +584,7 @@ static int show_blob_object(const struct object_id *oid, struct rev_info *rev, c\n \tfflush(rev->diffopt.file);\n \tif (!rev->diffopt.flags.textconv_set_via_cmdline ||\n \t    !rev->diffopt.flags.allow_textconv)\n-\t\treturn stream_blob_to_fd(1, oid, NULL, 0);\n+\t\treturn odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0);\n \n \tif (get_oid_with_context(the_repository, obj_name,\n \t\t\t\t GET_OID_RECORD_PATH,\n@@ -594,7 +594,7 @@ static int show_blob_object(const struct object_id *oid, struct rev_info *rev, c\n \t    !textconv_object(the_repository, obj_context.path,\n \t\t\t     obj_context.mode, &oidc, 1, &buf, &size)) {\n \t\tobject_context_release(&obj_context);\n-\t\treturn stream_blob_to_fd(1, oid, NULL, 0);\n+\t\treturn odb_stream_blob_to_fd(the_repository->objects, 1, oid, NULL, 0);\n \t}\n \n \tif (!buf)\ndiff --git a/entry.c b/entry.c\nindex cae02eb503..38dfe670f7 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -139,7 +139,7 @@ static int streaming_write_entry(const struct cache_entry *ce, char *path,\n \tif (fd < 0)\n \t\treturn -1;\n \n-\tresult |= stream_blob_to_fd(fd, &ce->oid, filter, 1);\n+\tresult |= odb_stream_blob_to_fd(the_repository->objects, fd, &ce->oid, filter, 1);\n \t*fstat_done = fstat_checkout_output(fd, state, statbuf);\n \tresult |= close(fd);\n \ndiff --git a/parallel-checkout.c b/parallel-checkout.c\nindex fba6aa65a6..1cb6701b92 100644\n--- a/parallel-checkout.c\n+++ b/parallel-checkout.c\n@@ -281,7 +281,8 @@ static int write_pc_item_to_fd(struct parallel_checkout_item *pc_item, int fd,\n \n \tfilter = get_stream_filter_ca(&pc_item->ca, &pc_item->ce->oid);\n \tif (filter) {\n-\t\tif (stream_blob_to_fd(fd, &pc_item->ce->oid, filter, 1)) {\n+\t\tif (odb_stream_blob_to_fd(the_repository->objects, fd,\n+\t\t\t\t\t  &pc_item->ce->oid, filter, 1)) {\n \t\t\t/* On error, reset fd to try writing without streaming */\n \t\t\tif (reset_fd(fd, path))\n \t\t\t\treturn -1;\ndiff --git a/streaming.c b/streaming.c\nindex f0f7d31956..807a6e03a8 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -2,8 +2,6 @@\n  * Copyright (c) 2011, Google Inc.\n  */\n \n-#define USE_THE_REPOSITORY_VARIABLE\n-\n #include \"git-compat-util.h\"\n #include \"convert.h\"\n #include \"environment.h\"\n@@ -359,7 +357,7 @@ static int open_istream_pack_non_delta(struct odb_read_stream **out,\n \n \tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n \t    oi.u.packed.is_delta ||\n-\t    repo_settings_get_big_file_threshold(the_repository) >= size)\n+\t    repo_settings_get_big_file_threshold(odb->repo) >= size)\n \t\treturn -1;\n \n \tin_pack_type = unpack_object_header(oi.u.packed.pack,\n@@ -518,8 +516,11 @@ struct odb_read_stream *open_istream(struct repository *r,\n \treturn st;\n }\n \n-int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter *filter,\n-\t\t      int can_seek)\n+int odb_stream_blob_to_fd(struct object_database *odb,\n+\t\t\t  int fd,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  struct stream_filter *filter,\n+\t\t\t  int can_seek)\n {\n \tstruct odb_read_stream *st;\n \tenum object_type type;\n@@ -527,7 +528,7 @@ int stream_blob_to_fd(int fd, const struct object_id *oid, struct stream_filter\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = open_istream(the_repository, oid, &type, &sz, filter);\n+\tst = open_istream(odb->repo, oid, &type, &sz, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\ndiff --git a/streaming.h b/streaming.h\nindex f5ff5d7ac9..148f6b3069 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -6,6 +6,7 @@\n \n #include \"object.h\"\n \n+struct object_database;\n /* opaque */\n struct odb_read_stream;\n struct stream_filter;\n@@ -16,6 +17,21 @@ struct odb_read_stream *open_istream(struct repository *, const struct object_id\n int close_istream(struct odb_read_stream *);\n ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n \n-int stream_blob_to_fd(int fd, const struct object_id *, struct stream_filter *, int can_seek);\n+/*\n+ * Look up the object by its ID and write the full contents to the file\n+ * descriptor. The object must be a blob, or the function will fail. When\n+ * provided, the filter is used to transform the blob contents.\n+ *\n+ * `can_seek` should be set to 1 in case the given file descriptor can be\n+ * seek(3p)'d on. This is used to support files with holes in case a\n+ * significant portion of the blob contains NUL bytes.\n+ *\n+ * Returns a negative error code on failure, 0 on success.\n+ */\n+int odb_stream_blob_to_fd(struct object_database *odb,\n+\t\t\t  int fd,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  struct stream_filter *filter,\n+\t\t\t  int can_seek);\n \n #endif /* STREAMING_H */\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531195","messageId":"20251123-b4-pks-odb-read-stream-v3-14-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 14/19] streaming: make the `odb_read_stream` definition public","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:39Z","receivedAt":"2025-11-23T19:00:33Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Subsequent commits will move the backend-specific logic of setting up an\nobject read stream into the specific subsystems. As the backends are now\nthe ones that are responsible for allocating the stream they'll need to\nhave the stream definition available to them.\n\nMake the stream definition public to prepare for this.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n streaming.c | 11 -----------\n streaming.h | 15 ++++++++++++++-\n 2 files changed, 14 insertions(+), 12 deletions(-)\n\ndiff --git a/streaming.c b/streaming.c\nindex 807a6e03a8..0635b7c12e 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -12,19 +12,8 @@\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \n-typedef int (*close_istream_fn)(struct odb_read_stream *);\n-typedef ssize_t (*read_istream_fn)(struct odb_read_stream *, char *, size_t);\n-\n #define FILTER_BUFFER (1024*16)\n \n-struct odb_read_stream {\n-\tclose_istream_fn close;\n-\tread_istream_fn read;\n-\n-\tenum object_type type;\n-\tunsigned long size; /* inflated size of full object */\n-};\n-\n /*****************************************************************\n  *\n  * Filtered stream\ndiff --git a/streaming.h b/streaming.h\nindex 148f6b3069..acfdef1598 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -7,10 +7,23 @@\n #include \"object.h\"\n \n struct object_database;\n-/* opaque */\n struct odb_read_stream;\n struct stream_filter;\n \n+typedef int (*odb_read_stream_close_fn)(struct odb_read_stream *);\n+typedef ssize_t (*odb_read_stream_read_fn)(struct odb_read_stream *, char *, size_t);\n+\n+/*\n+ * A stream that can be used to read an object from the object database without\n+ * loading all of it into memory.\n+ */\n+struct odb_read_stream {\n+\todb_read_stream_close_fn close;\n+\todb_read_stream_read_fn read;\n+\tenum object_type type;\n+\tunsigned long size; /* inflated size of full object */\n+};\n+\n struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n \t\t\t\t     enum object_type *, unsigned long *,\n \t\t\t\t     struct stream_filter *);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531196","messageId":"20251123-b4-pks-odb-read-stream-v3-15-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 15/19] streaming: move logic to read loose objects streams into backend","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:40Z","receivedAt":"2025-11-23T19:00:35Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Move the logic to read loose object streams into the respective\nsubsystem. This allows us to make a couple of function declarations\nprivate.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-file.c | 167 ++++++++++++++++++++++++++++++++++++++++++++++++++++++----\n object-file.h |  42 ++-------------\n streaming.c   | 133 +---------------------------------------------\n 3 files changed, 164 insertions(+), 178 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b62b21a452..8c67847fea 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -234,9 +234,9 @@ static void *map_fd(int fd, const char *path, unsigned long *size)\n \treturn map;\n }\n \n-void *odb_source_loose_map_object(struct odb_source *source,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  unsigned long *size)\n+static void *odb_source_loose_map_object(struct odb_source *source,\n+\t\t\t\t\t const struct object_id *oid,\n+\t\t\t\t\t unsigned long *size)\n {\n \tconst char *p;\n \tint fd = open_loose_object(source->loose, oid, &p);\n@@ -246,11 +246,29 @@ void *odb_source_loose_map_object(struct odb_source *source,\n \treturn map_fd(fd, p, size);\n }\n \n-enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n-\t\t\t\t\t\t    unsigned char *map,\n-\t\t\t\t\t\t    unsigned long mapsize,\n-\t\t\t\t\t\t    void *buffer,\n-\t\t\t\t\t\t    unsigned long bufsiz)\n+enum unpack_loose_header_result {\n+\tULHR_OK,\n+\tULHR_BAD,\n+\tULHR_TOO_LONG,\n+};\n+\n+/**\n+ * unpack_loose_header() initializes the data stream needed to unpack\n+ * a loose object header.\n+ *\n+ * Returns:\n+ *\n+ * - ULHR_OK on success\n+ * - ULHR_BAD on error\n+ * - ULHR_TOO_LONG if the header was too long\n+ *\n+ * It will only parse up to MAX_HEADER_LEN bytes.\n+ */\n+static enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n+\t\t\t\t\t\t\t   unsigned char *map,\n+\t\t\t\t\t\t\t   unsigned long mapsize,\n+\t\t\t\t\t\t\t   void *buffer,\n+\t\t\t\t\t\t\t   unsigned long bufsiz)\n {\n \tint status;\n \n@@ -329,11 +347,18 @@ static void *unpack_loose_rest(git_zstream *stream,\n }\n \n /*\n+ * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n+ * object. If it doesn't follow that format -1 is returned. To check\n+ * the validity of the <type> populate the \"typep\" in the \"struct\n+ * object_info\". It will be OBJ_BAD if the object type is unknown. The\n+ * parsed <len> can be retrieved via \"oi->sizep\", and from there\n+ * passed to unpack_loose_rest().\n+ *\n  * We used to just use \"sscanf()\", but that's actually way\n  * too permissive for what we want to check. So do an anal\n  * object header parse by hand.\n  */\n-int parse_loose_header(const char *hdr, struct object_info *oi)\n+static int parse_loose_header(const char *hdr, struct object_info *oi)\n {\n \tconst char *type_buf = hdr;\n \tsize_t size;\n@@ -1976,3 +2001,127 @@ void odb_source_loose_free(struct odb_source_loose *loose)\n \tloose_object_map_clear(&loose->map);\n \tfree(loose);\n }\n+\n+struct odb_loose_read_stream {\n+\tstruct odb_read_stream base;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_LOOSE_READ_STREAM_INUSE,\n+\t\tODB_LOOSE_READ_STREAM_DONE,\n+\t\tODB_LOOSE_READ_STREAM_ERROR,\n+\t} z_state;\n+\tvoid *mapped;\n+\tunsigned long mapsize;\n+\tchar hdr[32];\n+\tint hdr_avail;\n+\tint hdr_used;\n+};\n+\n+static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n+{\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tsize_t total_read = 0;\n+\n+\tswitch (st->z_state) {\n+\tcase ODB_LOOSE_READ_STREAM_DONE:\n+\t\treturn 0;\n+\tcase ODB_LOOSE_READ_STREAM_ERROR:\n+\t\treturn -1;\n+\tdefault:\n+\t\tbreak;\n+\t}\n+\n+\tif (st->hdr_used < st->hdr_avail) {\n+\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n+\t\tif (sz < to_copy)\n+\t\t\tto_copy = sz;\n+\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n+\t\tst->hdr_used += to_copy;\n+\t\ttotal_read += to_copy;\n+\t}\n+\n+\twhile (total_read < sz) {\n+\t\tint status;\n+\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\n+\t\tif (status == Z_STREAM_END) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n+\t\t\tbreak;\n+\t\t}\n+\t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn total_read;\n+}\n+\n+static int close_istream_loose(struct odb_read_stream *_st)\n+{\n+\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n+\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n+\tmunmap(st->mapped, st->mapsize);\n+\treturn 0;\n+}\n+\n+int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\tstruct odb_source *source,\n+\t\t\t\t\tconst struct object_id *oid)\n+{\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct odb_loose_read_stream *st;\n+\tunsigned long mapsize;\n+\tvoid *mapped;\n+\n+\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n+\tif (!mapped)\n+\t\treturn -1;\n+\n+\t/*\n+\t * Note: we must allocate this structure early even though we may still\n+\t * fail. This is because we need to initialize the zlib stream, and it\n+\t * is not possible to copy the stream around after the fact because it\n+\t * has self-referencing pointers.\n+\t */\n+\tCALLOC_ARRAY(st, 1);\n+\n+\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n+\t\t\t\t    sizeof(st->hdr))) {\n+\tcase ULHR_OK:\n+\t\tbreak;\n+\tcase ULHR_BAD:\n+\tcase ULHR_TOO_LONG:\n+\t\tgoto error;\n+\t}\n+\n+\toi.sizep = &st->base.size;\n+\toi.typep = &st->base.type;\n+\n+\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n+\t\tgoto error;\n+\n+\tst->mapped = mapped;\n+\tst->mapsize = mapsize;\n+\tst->hdr_used = strlen(st->hdr) + 1;\n+\tst->hdr_avail = st->z.total_out;\n+\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n+\tst->base.close = close_istream_loose;\n+\tst->base.read = read_istream_loose;\n+\n+\t*out = &st->base;\n+\n+\treturn 0;\n+error:\n+\tgit_inflate_end(&st->z);\n+\tmunmap(st->mapped, st->mapsize);\n+\tfree(st);\n+\treturn -1;\n+}\ndiff --git a/object-file.h b/object-file.h\nindex eeffa67bbd..1229d5f675 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -16,6 +16,8 @@ enum {\n int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n \n+struct object_info;\n+struct odb_read_stream;\n struct odb_source;\n \n struct odb_source_loose {\n@@ -47,9 +49,9 @@ int odb_source_loose_read_object_info(struct odb_source *source,\n \t\t\t\t      const struct object_id *oid,\n \t\t\t\t      struct object_info *oi, int flags);\n \n-void *odb_source_loose_map_object(struct odb_source *source,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  unsigned long *size);\n+int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\tstruct odb_source *source,\n+\t\t\t\t\tconst struct object_id *oid);\n \n /*\n  * Return true iff an object database source has a loose object\n@@ -143,40 +145,6 @@ int for_each_loose_object(struct object_database *odb,\n int format_object_header(char *str, size_t size, enum object_type type,\n \t\t\t size_t objsize);\n \n-/**\n- * unpack_loose_header() initializes the data stream needed to unpack\n- * a loose object header.\n- *\n- * Returns:\n- *\n- * - ULHR_OK on success\n- * - ULHR_BAD on error\n- * - ULHR_TOO_LONG if the header was too long\n- *\n- * It will only parse up to MAX_HEADER_LEN bytes.\n- */\n-enum unpack_loose_header_result {\n-\tULHR_OK,\n-\tULHR_BAD,\n-\tULHR_TOO_LONG,\n-};\n-enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n-\t\t\t\t\t\t    unsigned char *map,\n-\t\t\t\t\t\t    unsigned long mapsize,\n-\t\t\t\t\t\t    void *buffer,\n-\t\t\t\t\t\t    unsigned long bufsiz);\n-\n-/**\n- * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n- * object. If it doesn't follow that format -1 is returned. To check\n- * the validity of the <type> populate the \"typep\" in the \"struct\n- * object_info\". It will be OBJ_BAD if the object type is unknown. The\n- * parsed <len> can be retrieved via \"oi->sizep\", and from there\n- * passed to unpack_loose_rest().\n- */\n-struct object_info;\n-int parse_loose_header(const char *hdr, struct object_info *oi);\n-\n int force_object_loose(struct odb_source *source,\n \t\t       const struct object_id *oid, time_t mtime);\n \ndiff --git a/streaming.c b/streaming.c\nindex 0635b7c12e..d5acc1c396 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -114,137 +114,6 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \treturn &fs->base;\n }\n \n-/*****************************************************************\n- *\n- * Loose object stream\n- *\n- *****************************************************************/\n-\n-struct odb_loose_read_stream {\n-\tstruct odb_read_stream base;\n-\tgit_zstream z;\n-\tenum {\n-\t\tODB_LOOSE_READ_STREAM_INUSE,\n-\t\tODB_LOOSE_READ_STREAM_DONE,\n-\t\tODB_LOOSE_READ_STREAM_ERROR,\n-\t} z_state;\n-\tvoid *mapped;\n-\tunsigned long mapsize;\n-\tchar hdr[32];\n-\tint hdr_avail;\n-\tint hdr_used;\n-};\n-\n-static ssize_t read_istream_loose(struct odb_read_stream *_st, char *buf, size_t sz)\n-{\n-\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tsize_t total_read = 0;\n-\n-\tswitch (st->z_state) {\n-\tcase ODB_LOOSE_READ_STREAM_DONE:\n-\t\treturn 0;\n-\tcase ODB_LOOSE_READ_STREAM_ERROR:\n-\t\treturn -1;\n-\tdefault:\n-\t\tbreak;\n-\t}\n-\n-\tif (st->hdr_used < st->hdr_avail) {\n-\t\tsize_t to_copy = st->hdr_avail - st->hdr_used;\n-\t\tif (sz < to_copy)\n-\t\t\tto_copy = sz;\n-\t\tmemcpy(buf, st->hdr + st->hdr_used, to_copy);\n-\t\tst->hdr_used += to_copy;\n-\t\ttotal_read += to_copy;\n-\t}\n-\n-\twhile (total_read < sz) {\n-\t\tint status;\n-\n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n-\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n-\n-\t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_DONE;\n-\t\t\tbreak;\n-\t\t}\n-\t\tif (status != Z_OK && (status != Z_BUF_ERROR || total_read < sz)) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_LOOSE_READ_STREAM_ERROR;\n-\t\t\treturn -1;\n-\t\t}\n-\t}\n-\treturn total_read;\n-}\n-\n-static int close_istream_loose(struct odb_read_stream *_st)\n-{\n-\tstruct odb_loose_read_stream *st = (struct odb_loose_read_stream *)_st;\n-\tif (st->z_state == ODB_LOOSE_READ_STREAM_INUSE)\n-\t\tgit_inflate_end(&st->z);\n-\tmunmap(st->mapped, st->mapsize);\n-\treturn 0;\n-}\n-\n-static int open_istream_loose(struct odb_read_stream **out,\n-\t\t\t      struct odb_source *source,\n-\t\t\t      const struct object_id *oid)\n-{\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tstruct odb_loose_read_stream *st;\n-\tunsigned long mapsize;\n-\tvoid *mapped;\n-\n-\tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n-\tif (!mapped)\n-\t\treturn -1;\n-\n-\t/*\n-\t * Note: we must allocate this structure early even though we may still\n-\t * fail. This is because we need to initialize the zlib stream, and it\n-\t * is not possible to copy the stream around after the fact because it\n-\t * has self-referencing pointers.\n-\t */\n-\tCALLOC_ARRAY(st, 1);\n-\n-\tswitch (unpack_loose_header(&st->z, mapped, mapsize, st->hdr,\n-\t\t\t\t    sizeof(st->hdr))) {\n-\tcase ULHR_OK:\n-\t\tbreak;\n-\tcase ULHR_BAD:\n-\tcase ULHR_TOO_LONG:\n-\t\tgoto error;\n-\t}\n-\n-\toi.sizep = &st->base.size;\n-\toi.typep = &st->base.type;\n-\n-\tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n-\t\tgoto error;\n-\n-\tst->mapped = mapped;\n-\tst->mapsize = mapsize;\n-\tst->hdr_used = strlen(st->hdr) + 1;\n-\tst->hdr_avail = st->z.total_out;\n-\tst->z_state = ODB_LOOSE_READ_STREAM_INUSE;\n-\tst->base.close = close_istream_loose;\n-\tst->base.read = read_istream_loose;\n-\n-\t*out = &st->base;\n-\n-\treturn 0;\n-error:\n-\tgit_inflate_end(&st->z);\n-\tmunmap(st->mapped, st->mapsize);\n-\tfree(st);\n-\treturn -1;\n-}\n-\n-\n /*****************************************************************\n  *\n  * Non-delta packed object stream\n@@ -455,7 +324,7 @@ static int istream_source(struct odb_read_stream **out,\n \n \todb_prepare_alternates(r->objects);\n \tfor (source = r->objects->sources; source; source = source->next)\n-\t\tif (!open_istream_loose(out, source, oid))\n+\t\tif (!odb_source_loose_read_object_stream(out, source, oid))\n \t\t\treturn 0;\n \n \treturn open_istream_incore(out, r, oid);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531197","messageId":"20251123-b4-pks-odb-read-stream-v3-16-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 16/19] streaming: move logic to read packed objects streams into backend","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:41Z","receivedAt":"2025-11-23T19:00:40Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Move the logic to read packed object streams into the respective\nsubsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n packfile.c  | 128 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n packfile.h  |   5 +++\n streaming.c | 136 +-----------------------------------------------------------\n 3 files changed, 134 insertions(+), 135 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex b4bc40d895..ad56ce0b90 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -20,6 +20,7 @@\n #include \"tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"streaming.h\"\n #include \"midx.h\"\n #include \"commit-graph.h\"\n #include \"pack-revindex.h\"\n@@ -2406,3 +2407,130 @@ void packfile_store_close(struct packfile_store *store)\n \t\tclose_pack(p);\n \t}\n }\n+\n+struct odb_packed_read_stream {\n+\tstruct odb_read_stream base;\n+\tstruct packed_git *pack;\n+\tgit_zstream z;\n+\tenum {\n+\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n+\t\tODB_PACKED_READ_STREAM_INUSE,\n+\t\tODB_PACKED_READ_STREAM_DONE,\n+\t\tODB_PACKED_READ_STREAM_ERROR,\n+\t} z_state;\n+\toff_t pos;\n+};\n+\n+static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n+\t\t\t\t\t   size_t sz)\n+{\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tsize_t total_read = 0;\n+\n+\tswitch (st->z_state) {\n+\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n+\t\tmemset(&st->z, 0, sizeof(st->z));\n+\t\tgit_inflate_init(&st->z);\n+\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n+\t\tbreak;\n+\tcase ODB_PACKED_READ_STREAM_DONE:\n+\t\treturn 0;\n+\tcase ODB_PACKED_READ_STREAM_ERROR:\n+\t\treturn -1;\n+\tcase ODB_PACKED_READ_STREAM_INUSE:\n+\t\tbreak;\n+\t}\n+\n+\twhile (total_read < sz) {\n+\t\tint status;\n+\t\tstruct pack_window *window = NULL;\n+\t\tunsigned char *mapped;\n+\n+\t\tmapped = use_pack(st->pack, &window,\n+\t\t\t\t  st->pos, &st->z.avail_in);\n+\n+\t\tst->z.next_out = (unsigned char *)buf + total_read;\n+\t\tst->z.avail_out = sz - total_read;\n+\t\tst->z.next_in = mapped;\n+\t\tstatus = git_inflate(&st->z, Z_FINISH);\n+\n+\t\tst->pos += st->z.next_in - mapped;\n+\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n+\t\tunuse_pack(&window);\n+\n+\t\tif (status == Z_STREAM_END) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\t/*\n+\t\t * Unlike the loose object case, we do not have to worry here\n+\t\t * about running out of input bytes and spinning infinitely. If\n+\t\t * we get Z_BUF_ERROR due to too few input bytes, then we'll\n+\t\t * replenish them in the next use_pack() call when we loop. If\n+\t\t * we truly hit the end of the pack (i.e., because it's corrupt\n+\t\t * or truncated), then use_pack() catches that and will die().\n+\t\t */\n+\t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n+\t\t\tgit_inflate_end(&st->z);\n+\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn total_read;\n+}\n+\n+static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n+{\n+\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n+\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n+\t\tgit_inflate_end(&st->z);\n+\treturn 0;\n+}\n+\n+int packfile_store_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t      struct packfile_store *store,\n+\t\t\t\t      const struct object_id *oid)\n+{\n+\tstruct odb_packed_read_stream *stream;\n+\tstruct pack_window *window = NULL;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\tenum object_type in_pack_type;\n+\tunsigned long size;\n+\n+\toi.sizep = &size;\n+\n+\tif (packfile_store_read_object_info(store, oid, &oi, 0) ||\n+\t    oi.u.packed.is_delta ||\n+\t    repo_settings_get_big_file_threshold(store->odb->repo) >= size)\n+\t\treturn -1;\n+\n+\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n+\t\t\t\t\t    &window,\n+\t\t\t\t\t    &oi.u.packed.offset,\n+\t\t\t\t\t    &size);\n+\tunuse_pack(&window);\n+\tswitch (in_pack_type) {\n+\tdefault:\n+\t\treturn -1; /* we do not do deltas for now */\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_BLOB:\n+\tcase OBJ_TAG:\n+\t\tbreak;\n+\t}\n+\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.close = close_istream_pack_non_delta;\n+\tstream->base.read = read_istream_pack_non_delta;\n+\tstream->base.type = in_pack_type;\n+\tstream->base.size = size;\n+\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n+\tstream->pack = oi.u.packed.pack;\n+\tstream->pos = oi.u.packed.offset;\n+\n+\t*out = &stream->base;\n+\n+\treturn 0;\n+}\ndiff --git a/packfile.h b/packfile.h\nindex 0a98bddd81..3fcc5ae6e0 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -8,6 +8,7 @@\n \n /* in odb.h */\n struct object_info;\n+struct odb_read_stream;\n \n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n@@ -144,6 +145,10 @@ void packfile_store_add_pack(struct packfile_store *store,\n #define repo_for_each_pack(repo, p) \\\n \tfor (p = packfile_store_get_packs(repo->objects->packfiles); p; p = p->next)\n \n+int packfile_store_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t      struct packfile_store *store,\n+\t\t\t\t      const struct object_id *oid);\n+\n /*\n  * Try to read the object identified by its ID from the object store and\n  * populate the object info with its data. Returns 1 in case the object was\ndiff --git a/streaming.c b/streaming.c\nindex d5acc1c396..3140728a70 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -114,140 +114,6 @@ static struct odb_read_stream *attach_stream_filter(struct odb_read_stream *st,\n \treturn &fs->base;\n }\n \n-/*****************************************************************\n- *\n- * Non-delta packed object stream\n- *\n- *****************************************************************/\n-\n-struct odb_packed_read_stream {\n-\tstruct odb_read_stream base;\n-\tstruct packed_git *pack;\n-\tgit_zstream z;\n-\tenum {\n-\t\tODB_PACKED_READ_STREAM_UNINITIALIZED,\n-\t\tODB_PACKED_READ_STREAM_INUSE,\n-\t\tODB_PACKED_READ_STREAM_DONE,\n-\t\tODB_PACKED_READ_STREAM_ERROR,\n-\t} z_state;\n-\toff_t pos;\n-};\n-\n-static ssize_t read_istream_pack_non_delta(struct odb_read_stream *_st, char *buf,\n-\t\t\t\t\t   size_t sz)\n-{\n-\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tsize_t total_read = 0;\n-\n-\tswitch (st->z_state) {\n-\tcase ODB_PACKED_READ_STREAM_UNINITIALIZED:\n-\t\tmemset(&st->z, 0, sizeof(st->z));\n-\t\tgit_inflate_init(&st->z);\n-\t\tst->z_state = ODB_PACKED_READ_STREAM_INUSE;\n-\t\tbreak;\n-\tcase ODB_PACKED_READ_STREAM_DONE:\n-\t\treturn 0;\n-\tcase ODB_PACKED_READ_STREAM_ERROR:\n-\t\treturn -1;\n-\tcase ODB_PACKED_READ_STREAM_INUSE:\n-\t\tbreak;\n-\t}\n-\n-\twhile (total_read < sz) {\n-\t\tint status;\n-\t\tstruct pack_window *window = NULL;\n-\t\tunsigned char *mapped;\n-\n-\t\tmapped = use_pack(st->pack, &window,\n-\t\t\t\t  st->pos, &st->z.avail_in);\n-\n-\t\tst->z.next_out = (unsigned char *)buf + total_read;\n-\t\tst->z.avail_out = sz - total_read;\n-\t\tst->z.next_in = mapped;\n-\t\tstatus = git_inflate(&st->z, Z_FINISH);\n-\n-\t\tst->pos += st->z.next_in - mapped;\n-\t\ttotal_read = st->z.next_out - (unsigned char *)buf;\n-\t\tunuse_pack(&window);\n-\n-\t\tif (status == Z_STREAM_END) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_PACKED_READ_STREAM_DONE;\n-\t\t\tbreak;\n-\t\t}\n-\n-\t\t/*\n-\t\t * Unlike the loose object case, we do not have to worry here\n-\t\t * about running out of input bytes and spinning infinitely. If\n-\t\t * we get Z_BUF_ERROR due to too few input bytes, then we'll\n-\t\t * replenish them in the next use_pack() call when we loop. If\n-\t\t * we truly hit the end of the pack (i.e., because it's corrupt\n-\t\t * or truncated), then use_pack() catches that and will die().\n-\t\t */\n-\t\tif (status != Z_OK && status != Z_BUF_ERROR) {\n-\t\t\tgit_inflate_end(&st->z);\n-\t\t\tst->z_state = ODB_PACKED_READ_STREAM_ERROR;\n-\t\t\treturn -1;\n-\t\t}\n-\t}\n-\treturn total_read;\n-}\n-\n-static int close_istream_pack_non_delta(struct odb_read_stream *_st)\n-{\n-\tstruct odb_packed_read_stream *st = (struct odb_packed_read_stream *)_st;\n-\tif (st->z_state == ODB_PACKED_READ_STREAM_INUSE)\n-\t\tgit_inflate_end(&st->z);\n-\treturn 0;\n-}\n-\n-static int open_istream_pack_non_delta(struct odb_read_stream **out,\n-\t\t\t\t       struct object_database *odb,\n-\t\t\t\t       const struct object_id *oid)\n-{\n-\tstruct odb_packed_read_stream *stream;\n-\tstruct pack_window *window = NULL;\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tenum object_type in_pack_type;\n-\tunsigned long size;\n-\n-\toi.sizep = &size;\n-\n-\tif (packfile_store_read_object_info(odb->packfiles, oid, &oi, 0) ||\n-\t    oi.u.packed.is_delta ||\n-\t    repo_settings_get_big_file_threshold(odb->repo) >= size)\n-\t\treturn -1;\n-\n-\tin_pack_type = unpack_object_header(oi.u.packed.pack,\n-\t\t\t\t\t    &window,\n-\t\t\t\t\t    &oi.u.packed.offset,\n-\t\t\t\t\t    &size);\n-\tunuse_pack(&window);\n-\tswitch (in_pack_type) {\n-\tdefault:\n-\t\treturn -1; /* we do not do deltas for now */\n-\tcase OBJ_COMMIT:\n-\tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n-\tcase OBJ_TAG:\n-\t\tbreak;\n-\t}\n-\n-\tCALLOC_ARRAY(stream, 1);\n-\tstream->base.close = close_istream_pack_non_delta;\n-\tstream->base.read = read_istream_pack_non_delta;\n-\tstream->base.type = in_pack_type;\n-\tstream->base.size = size;\n-\tstream->z_state = ODB_PACKED_READ_STREAM_UNINITIALIZED;\n-\tstream->pack = oi.u.packed.pack;\n-\tstream->pos = oi.u.packed.offset;\n-\n-\t*out = &stream->base;\n-\n-\treturn 0;\n-}\n-\n-\n /*****************************************************************\n  *\n  * In-core stream\n@@ -319,7 +185,7 @@ static int istream_source(struct odb_read_stream **out,\n {\n \tstruct odb_source *source;\n \n-\tif (!open_istream_pack_non_delta(out, r->objects, oid))\n+\tif (!packfile_store_read_object_stream(out, r->objects->packfiles, oid))\n \t\treturn 0;\n \n \todb_prepare_alternates(r->objects);\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531198","messageId":"20251123-b4-pks-odb-read-stream-v3-17-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 17/19] streaming: refactor interface to be object-database-centric","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:42Z","receivedAt":"2025-11-23T19:00:42Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Refactor the streaming interface to be centered around object databases\ninstead of centered around the repository. Rename the functions\naccordingly.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  6 +++---\n archive-zip.c          | 12 ++++++------\n builtin/index-pack.c   |  8 ++++----\n builtin/pack-objects.c | 14 +++++++-------\n object-file.c          |  8 ++++----\n streaming.c            | 44 ++++++++++++++++++++++----------------------\n streaming.h            | 30 +++++++++++++++++++++++++-----\n 7 files changed, 71 insertions(+), 51 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex dc1eda09e0..4d87b28504 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -135,16 +135,16 @@ static int stream_blocked(struct repository *r, const struct object_id *oid)\n \tchar buf[BLOCKSIZE];\n \tssize_t readlen;\n \n-\tst = open_istream(r, oid, &type, &sz, NULL);\n+\tst = odb_read_stream_open(r->objects, oid, &type, &sz, NULL);\n \tif (!st)\n \t\treturn error(_(\"cannot stream blob %s\"), oid_to_hex(oid));\n \tfor (;;) {\n-\t\treadlen = read_istream(st, buf, sizeof(buf));\n+\t\treadlen = odb_read_stream_read(st, buf, sizeof(buf));\n \t\tif (readlen <= 0)\n \t\t\tbreak;\n \t\tdo_write_blocked(buf, readlen);\n \t}\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \tif (!readlen)\n \t\tfinish_record();\n \treturn readlen;\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 40a9c93ff9..c44684aebc 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -348,8 +348,8 @@ static int write_zip_entry(struct archiver_args *args,\n \n \t\tif (!buffer) {\n \t\t\tenum object_type type;\n-\t\t\tstream = open_istream(args->repo, oid, &type, &size,\n-\t\t\t\t\t      NULL);\n+\t\t\tstream = odb_read_stream_open(args->repo->objects, oid,\n+\t\t\t\t\t\t      &type, &size, NULL);\n \t\t\tif (!stream)\n \t\t\t\treturn error(_(\"cannot stream blob %s\"),\n \t\t\t\t\t     oid_to_hex(oid));\n@@ -429,7 +429,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tssize_t readlen;\n \n \t\tfor (;;) {\n-\t\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\t\t\treadlen = odb_read_stream_read(stream, buf, sizeof(buf));\n \t\t\tif (readlen <= 0)\n \t\t\t\tbreak;\n \t\t\tcrc = crc32(crc, buf, readlen);\n@@ -439,7 +439,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\t\t\t\t\t    buf, readlen);\n \t\t\twrite_or_die(1, buf, readlen);\n \t\t}\n-\t\tclose_istream(stream);\n+\t\todb_read_stream_close(stream);\n \t\tif (readlen)\n \t\t\treturn readlen;\n \n@@ -462,7 +462,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tzstream.avail_out = sizeof(compressed);\n \n \t\tfor (;;) {\n-\t\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\t\t\treadlen = odb_read_stream_read(stream, buf, sizeof(buf));\n \t\t\tif (readlen <= 0)\n \t\t\t\tbreak;\n \t\t\tcrc = crc32(crc, buf, readlen);\n@@ -486,7 +486,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\t}\n \n \t\t}\n-\t\tclose_istream(stream);\n+\t\todb_read_stream_close(stream);\n \t\tif (readlen)\n \t\t\treturn readlen;\n \ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 5f90f12f92..fb76ef0f4c 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -779,7 +779,7 @@ static int compare_objects(const unsigned char *buf, unsigned long size,\n \t}\n \n \twhile (size) {\n-\t\tssize_t len = read_istream(data->st, data->buf, size);\n+\t\tssize_t len = odb_read_stream_read(data->st, data->buf, size);\n \t\tif (len == 0)\n \t\t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t\t    oid_to_hex(&data->entry->idx.oid));\n@@ -807,15 +807,15 @@ static int check_collison(struct object_entry *entry)\n \n \tmemset(&data, 0, sizeof(data));\n \tdata.entry = entry;\n-\tdata.st = open_istream(the_repository, &entry->idx.oid, &type, &size,\n-\t\t\t       NULL);\n+\tdata.st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n+\t\t\t\t       &type, &size, NULL);\n \tif (!data.st)\n \t\treturn -1;\n \tif (size != entry->size || type != entry->type)\n \t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t    oid_to_hex(&entry->idx.oid));\n \tunpack_data(entry, compare_objects, &data);\n-\tclose_istream(data.st);\n+\todb_read_stream_close(data.st);\n \tfree(data.buf);\n \treturn 0;\n }\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex c693d948e1..1353c2384c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -417,7 +417,7 @@ static unsigned long write_large_blob_data(struct odb_read_stream *st, struct ha\n \tfor (;;) {\n \t\tssize_t readlen;\n \t\tint zret = Z_OK;\n-\t\treadlen = read_istream(st, ibuf, sizeof(ibuf));\n+\t\treadlen = odb_read_stream_read(st, ibuf, sizeof(ibuf));\n \t\tif (readlen == -1)\n \t\t\tdie(_(\"unable to read %s\"), oid_to_hex(oid));\n \n@@ -520,8 +520,8 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\tif (oe_type(entry) == OBJ_BLOB &&\n \t\t    oe_size_greater_than(&to_pack, entry,\n \t\t\t\t\t repo_settings_get_big_file_threshold(the_repository)) &&\n-\t\t    (st = open_istream(the_repository, &entry->idx.oid, &type,\n-\t\t\t\t       &size, NULL)) != NULL)\n+\t\t    (st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n+\t\t\t\t\t       &type, &size, NULL)) != NULL)\n \t\t\tbuf = NULL;\n \t\telse {\n \t\t\tbuf = odb_read_object(the_repository->objects,\n@@ -577,7 +577,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t\tdheader[--pos] = 128 | (--ofs & 127);\n \t\tif (limit && hdrlen + sizeof(dheader) - pos + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -591,7 +591,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t */\n \t\tif (limit && hdrlen + hashsz + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -601,7 +601,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t} else {\n \t\tif (limit && hdrlen + datalen + hashsz >= limit) {\n \t\t\tif (st)\n-\t\t\t\tclose_istream(st);\n+\t\t\t\todb_read_stream_close(st);\n \t\t\tfree(buf);\n \t\t\treturn 0;\n \t\t}\n@@ -609,7 +609,7 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t}\n \tif (st) {\n \t\tdatalen = write_large_blob_data(st, f, &entry->idx.oid);\n-\t\tclose_istream(st);\n+\t\todb_read_stream_close(st);\n \t} else {\n \t\thashwrite(f, buf, datalen);\n \t\tfree(buf);\ndiff --git a/object-file.c b/object-file.c\nindex 8c67847fea..9ba40a848c 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -139,7 +139,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \n-\tst = open_istream(r, oid, &obj_type, &size, NULL);\n+\tst = odb_read_stream_open(r->objects, oid, &obj_type, &size, NULL);\n \tif (!st)\n \t\treturn -1;\n \n@@ -151,10 +151,10 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \tgit_hash_update(&c, hdr, hdrlen);\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\t\tssize_t readlen = odb_read_stream_read(st, buf, sizeof(buf));\n \n \t\tif (readlen < 0) {\n-\t\t\tclose_istream(st);\n+\t\t\todb_read_stream_close(st);\n \t\t\treturn -1;\n \t\t}\n \t\tif (!readlen)\n@@ -162,7 +162,7 @@ int stream_object_signature(struct repository *r, const struct object_id *oid)\n \t\tgit_hash_update(&c, buf, readlen);\n \t}\n \tgit_hash_final_oid(&real_oid, &c);\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \treturn !oideq(oid, &real_oid) ? -1 : 0;\n }\n \ndiff --git a/streaming.c b/streaming.c\nindex 3140728a70..06993a751c 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -35,7 +35,7 @@ static int close_istream_filtered(struct odb_read_stream *_fs)\n {\n \tstruct odb_filtered_read_stream *fs = (struct odb_filtered_read_stream *)_fs;\n \tfree_stream_filter(fs->filter);\n-\treturn close_istream(fs->upstream);\n+\treturn odb_read_stream_close(fs->upstream);\n }\n \n static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n@@ -87,7 +87,7 @@ static ssize_t read_istream_filtered(struct odb_read_stream *_fs, char *buf,\n \n \t\t/* refill the input from the upstream */\n \t\tif (!fs->input_finished) {\n-\t\t\tfs->i_end = read_istream(fs->upstream, fs->ibuf, FILTER_BUFFER);\n+\t\t\tfs->i_end = odb_read_stream_read(fs->upstream, fs->ibuf, FILTER_BUFFER);\n \t\t\tif (fs->i_end < 0)\n \t\t\t\treturn -1;\n \t\t\tif (fs->i_end)\n@@ -149,7 +149,7 @@ static ssize_t read_istream_incore(struct odb_read_stream *_st, char *buf, size_\n }\n \n static int open_istream_incore(struct odb_read_stream **out,\n-\t\t\t       struct repository *r,\n+\t\t\t       struct object_database *odb,\n \t\t\t       const struct object_id *oid)\n {\n \tstruct object_info oi = OBJECT_INFO_INIT;\n@@ -163,7 +163,7 @@ static int open_istream_incore(struct odb_read_stream **out,\n \toi.typep = &stream.base.type;\n \toi.sizep = &stream.base.size;\n \toi.contentp = (void **)&stream.buf;\n-\tret = odb_read_object_info_extended(r->objects, oid, &oi,\n+\tret = odb_read_object_info_extended(odb, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n@@ -180,47 +180,47 @@ static int open_istream_incore(struct odb_read_stream **out,\n  *****************************************************************************/\n \n static int istream_source(struct odb_read_stream **out,\n-\t\t\t  struct repository *r,\n+\t\t\t  struct object_database *odb,\n \t\t\t  const struct object_id *oid)\n {\n \tstruct odb_source *source;\n \n-\tif (!packfile_store_read_object_stream(out, r->objects->packfiles, oid))\n+\tif (!packfile_store_read_object_stream(out, odb->packfiles, oid))\n \t\treturn 0;\n \n-\todb_prepare_alternates(r->objects);\n-\tfor (source = r->objects->sources; source; source = source->next)\n+\todb_prepare_alternates(odb);\n+\tfor (source = odb->sources; source; source = source->next)\n \t\tif (!odb_source_loose_read_object_stream(out, source, oid))\n \t\t\treturn 0;\n \n-\treturn open_istream_incore(out, r, oid);\n+\treturn open_istream_incore(out, odb, oid);\n }\n \n /****************************************************************\n  * Users of streaming interface\n  ****************************************************************/\n \n-int close_istream(struct odb_read_stream *st)\n+int odb_read_stream_close(struct odb_read_stream *st)\n {\n \tint r = st->close(st);\n \tfree(st);\n \treturn r;\n }\n \n-ssize_t read_istream(struct odb_read_stream *st, void *buf, size_t sz)\n+ssize_t odb_read_stream_read(struct odb_read_stream *st, void *buf, size_t sz)\n {\n \treturn st->read(st, buf, sz);\n }\n \n-struct odb_read_stream *open_istream(struct repository *r,\n-\t\t\t\t     const struct object_id *oid,\n-\t\t\t\t     enum object_type *type,\n-\t\t\t\t     unsigned long *size,\n-\t\t\t\t     struct stream_filter *filter)\n+struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n+\t\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t\t     enum object_type *type,\n+\t\t\t\t\t     unsigned long *size,\n+\t\t\t\t\t     struct stream_filter *filter)\n {\n \tstruct odb_read_stream *st;\n-\tconst struct object_id *real = lookup_replace_object(r, oid);\n-\tint ret = istream_source(&st, r, real);\n+\tconst struct object_id *real = lookup_replace_object(odb->repo, oid);\n+\tint ret = istream_source(&st, odb, real);\n \n \tif (ret)\n \t\treturn NULL;\n@@ -229,7 +229,7 @@ struct odb_read_stream *open_istream(struct repository *r,\n \t\t/* Add \"&& !is_null_stream_filter(filter)\" for performance */\n \t\tstruct odb_read_stream *nst = attach_stream_filter(st, filter);\n \t\tif (!nst) {\n-\t\t\tclose_istream(st);\n+\t\t\todb_read_stream_close(st);\n \t\t\treturn NULL;\n \t\t}\n \t\tst = nst;\n@@ -252,7 +252,7 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = open_istream(odb->repo, oid, &type, &sz, filter);\n+\tst = odb_read_stream_open(odb, oid, &type, &sz, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\n@@ -263,7 +263,7 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\n \t\tssize_t wrote, holeto;\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\t\tssize_t readlen = odb_read_stream_read(st, buf, sizeof(buf));\n \n \t\tif (readlen < 0)\n \t\t\tgoto close_and_exit;\n@@ -294,6 +294,6 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \tresult = 0;\n \n  close_and_exit:\n-\tclose_istream(st);\n+\todb_read_stream_close(st);\n \treturn result;\n }\ndiff --git a/streaming.h b/streaming.h\nindex acfdef1598..7cb55213b7 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -24,11 +24,31 @@ struct odb_read_stream {\n \tunsigned long size; /* inflated size of full object */\n };\n \n-struct odb_read_stream *open_istream(struct repository *, const struct object_id *,\n-\t\t\t\t     enum object_type *, unsigned long *,\n-\t\t\t\t     struct stream_filter *);\n-int close_istream(struct odb_read_stream *);\n-ssize_t read_istream(struct odb_read_stream *, void *, size_t);\n+/*\n+ * Create a new object stream for the given object database. Populates the type\n+ * and size pointers with the object's info. An optional filter can be used to\n+ * transform the object's content.\n+ *\n+ * Returns the stream on success, a `NULL` pointer otherwise.\n+ */\n+struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n+\t\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t\t     enum object_type *type,\n+\t\t\t\t\t     unsigned long *size,\n+\t\t\t\t\t     struct stream_filter *filter);\n+\n+/*\n+ * Close the given read stream and release all resources associated with it.\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+int odb_read_stream_close(struct odb_read_stream *stream);\n+\n+/*\n+ * Read data from the stream into the buffer. Returns 0 on EOF and the number\n+ * of bytes read on success. Returns a negative error code in case reading from\n+ * the stream fails.\n+ */\n+ssize_t odb_read_stream_read(struct odb_read_stream *stream, void *buf, size_t len);\n \n /*\n  * Look up the object by its ID and write the full contents to the file\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531199","messageId":"20251123-b4-pks-odb-read-stream-v3-18-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 18/19] streaming: move into object database subsystem","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:43Z","receivedAt":"2025-11-23T19:00:46Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"The \"streaming\" terminology is somewhat generic, so it may not be\nimmediately obvious that \"streaming.{c,h}\" is specific to the object\ndatabase. Rectify this by moving it into the \"odb/\" directory so that it\ncan be immediately attributed to the object subsystem.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Makefile                       | 2 +-\n archive-tar.c                  | 2 +-\n archive-zip.c                  | 2 +-\n builtin/cat-file.c             | 2 +-\n builtin/fsck.c                 | 2 +-\n builtin/index-pack.c           | 2 +-\n builtin/log.c                  | 2 +-\n builtin/pack-objects.c         | 2 +-\n entry.c                        | 2 +-\n meson.build                    | 2 +-\n object-file.c                  | 2 +-\n streaming.c => odb/streaming.c | 2 +-\n streaming.h => odb/streaming.h | 0\n packfile.c                     | 2 +-\n parallel-checkout.c            | 2 +-\n 15 files changed, 14 insertions(+), 14 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 7e0f77e298..6d8dcc4622 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1201,6 +1201,7 @@ LIB_OBJS += object-file.o\n LIB_OBJS += object-name.o\n LIB_OBJS += object.o\n LIB_OBJS += odb.o\n+LIB_OBJS += odb/streaming.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\n LIB_OBJS += oidset.o\n@@ -1294,7 +1295,6 @@ LIB_OBJS += split-index.o\n LIB_OBJS += stable-qsort.o\n LIB_OBJS += statinfo.o\n LIB_OBJS += strbuf.o\n-LIB_OBJS += streaming.o\n LIB_OBJS += string-list.o\n LIB_OBJS += strmap.o\n LIB_OBJS += strvec.o\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 4d87b28504..494b9f0667 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -12,8 +12,8 @@\n #include \"tar.h\"\n #include \"archive.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"strbuf.h\"\n-#include \"streaming.h\"\n #include \"run-command.h\"\n #include \"write-or-die.h\"\n \ndiff --git a/archive-zip.c b/archive-zip.c\nindex c44684aebc..a0bdc2fe3b 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -10,9 +10,9 @@\n #include \"gettext.h\"\n #include \"git-zlib.h\"\n #include \"hex.h\"\n-#include \"streaming.h\"\n #include \"utf8.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"strbuf.h\"\n #include \"userdiff.h\"\n #include \"write-or-die.h\"\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 120d626d66..505ddaa12f 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -18,13 +18,13 @@\n #include \"list-objects-filter-options.h\"\n #include \"parse-options.h\"\n #include \"userdiff.h\"\n-#include \"streaming.h\"\n #include \"oid-array.h\"\n #include \"packfile.h\"\n #include \"pack-bitmap.h\"\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"promisor-remote.h\"\n #include \"mailmap.h\"\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex 1a348d43c2..c7d2eea287 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -13,11 +13,11 @@\n #include \"fsck.h\"\n #include \"parse-options.h\"\n #include \"progress.h\"\n-#include \"streaming.h\"\n #include \"packfile.h\"\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"path.h\"\n #include \"read-cache-ll.h\"\n #include \"replace-object.h\"\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex fb76ef0f4c..581023495f 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -16,12 +16,12 @@\n #include \"progress.h\"\n #include \"fsck.h\"\n #include \"strbuf.h\"\n-#include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"packfile.h\"\n #include \"pack-revindex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"oid-array.h\"\n #include \"oidset.h\"\n #include \"path.h\"\ndiff --git a/builtin/log.c b/builtin/log.c\nindex e7b83a6e00..d4cf9c59c8 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -16,6 +16,7 @@\n #include \"refs.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"pager.h\"\n #include \"color.h\"\n #include \"commit.h\"\n@@ -35,7 +36,6 @@\n #include \"parse-options.h\"\n #include \"line-log.h\"\n #include \"branch.h\"\n-#include \"streaming.h\"\n #include \"version.h\"\n #include \"mailmap.h\"\n #include \"progress.h\"\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 1353c2384c..f109e26786 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -22,7 +22,6 @@\n #include \"pack-objects.h\"\n #include \"progress.h\"\n #include \"refs.h\"\n-#include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"pack-bitmap.h\"\n #include \"delta-islands.h\"\n@@ -33,6 +32,7 @@\n #include \"packfile.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"dir.h\"\n #include \"midx.h\"\ndiff --git a/entry.c b/entry.c\nindex 38dfe670f7..7817aee362 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -2,13 +2,13 @@\n \n #include \"git-compat-util.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"dir.h\"\n #include \"environment.h\"\n #include \"gettext.h\"\n #include \"hex.h\"\n #include \"name-hash.h\"\n #include \"sparse-index.h\"\n-#include \"streaming.h\"\n #include \"submodule.h\"\n #include \"symlinks.h\"\n #include \"progress.h\"\ndiff --git a/meson.build b/meson.build\nindex 1f95a06edb..fc82929b37 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -397,6 +397,7 @@ libgit_sources = [\n   'object-name.c',\n   'object.c',\n   'odb.c',\n+  'odb/streaming.c',\n   'oid-array.c',\n   'oidmap.c',\n   'oidset.c',\n@@ -490,7 +491,6 @@ libgit_sources = [\n   'stable-qsort.c',\n   'statinfo.c',\n   'strbuf.c',\n-  'streaming.c',\n   'string-list.c',\n   'strmap.c',\n   'strvec.c',\ndiff --git a/object-file.c b/object-file.c\nindex 9ba40a848c..9601fdb12d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -20,13 +20,13 @@\n #include \"object-file-convert.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"oidtree.h\"\n #include \"pack.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n #include \"read-cache-ll.h\"\n #include \"setup.h\"\n-#include \"streaming.h\"\n #include \"tempfile.h\"\n #include \"tmp-objdir.h\"\n \ndiff --git a/streaming.c b/odb/streaming.c\nsimilarity index 99%\nrename from streaming.c\nrename to odb/streaming.c\nindex 06993a751c..7ef58adaa2 100644\n--- a/streaming.c\n+++ b/odb/streaming.c\n@@ -5,10 +5,10 @@\n #include \"git-compat-util.h\"\n #include \"convert.h\"\n #include \"environment.h\"\n-#include \"streaming.h\"\n #include \"repository.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"replace-object.h\"\n #include \"packfile.h\"\n \ndiff --git a/streaming.h b/odb/streaming.h\nsimilarity index 100%\nrename from streaming.h\nrename to odb/streaming.h\ndiff --git a/packfile.c b/packfile.c\nindex ad56ce0b90..7a16aaa90d 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -20,7 +20,7 @@\n #include \"tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n-#include \"streaming.h\"\n+#include \"odb/streaming.h\"\n #include \"midx.h\"\n #include \"commit-graph.h\"\n #include \"pack-revindex.h\"\ndiff --git a/parallel-checkout.c b/parallel-checkout.c\nindex 1cb6701b92..0bf4bd6d4a 100644\n--- a/parallel-checkout.c\n+++ b/parallel-checkout.c\n@@ -13,7 +13,7 @@\n #include \"read-cache-ll.h\"\n #include \"run-command.h\"\n #include \"sigchain.h\"\n-#include \"streaming.h\"\n+#include \"odb/streaming.h\"\n #include \"symlinks.h\"\n #include \"thread-utils.h\"\n #include \"trace2.h\"\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"},{"id":"531200","messageId":"20251123-b4-pks-odb-read-stream-v3-19-1a129182822b@pks.im","threadId":"64509","inReplyTo":"20251123-b4-pks-odb-read-stream-v3-0-1a129182822b@pks.im","subject":"[PATCH v3 19/19] streaming: drop redundant type and size pointers","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-23T18:59:44Z","receivedAt":"2025-11-23T19:00:49Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In the preceding commits we have turned `struct odb_read_stream` into a\npublicly visible structure. Furthermore, this structure now contains the\ntype and size of the object that we are about to stream. Consequently,\nthe out-pointers that we used before to propagate the type and size of\nthe streamed object are now somewhat redundant with the data contained\nin the structure itself.\n\nDrop these out-pointers and adapt callers accordingly.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n archive-tar.c          |  4 +---\n archive-zip.c          |  5 ++---\n builtin/index-pack.c   |  7 ++-----\n builtin/pack-objects.c |  6 ++++--\n object-file.c          |  6 ++----\n odb/streaming.c        | 10 ++--------\n odb/streaming.h        |  7 ++-----\n 7 files changed, 15 insertions(+), 30 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 494b9f0667..0fc70d13a8 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -130,12 +130,10 @@ static void write_trailer(void)\n static int stream_blocked(struct repository *r, const struct object_id *oid)\n {\n \tstruct odb_read_stream *st;\n-\tenum object_type type;\n-\tunsigned long sz;\n \tchar buf[BLOCKSIZE];\n \tssize_t readlen;\n \n-\tst = odb_read_stream_open(r->objects, oid, &type, &sz, NULL);\n+\tst = odb_read_stream_open(r->objects, oid, NULL);\n \tif (!st)\n \t\treturn error(_(\"cannot stream blob %s\"), oid_to_hex(oid));\n \tfor (;;) {\ndiff --git a/archive-zip.c b/archive-zip.c\nindex a0bdc2fe3b..97ea8d60d6 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -347,12 +347,11 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\tmethod = ZIP_METHOD_DEFLATE;\n \n \t\tif (!buffer) {\n-\t\t\tenum object_type type;\n-\t\t\tstream = odb_read_stream_open(args->repo->objects, oid,\n-\t\t\t\t\t\t      &type, &size, NULL);\n+\t\t\tstream = odb_read_stream_open(args->repo->objects, oid, NULL);\n \t\t\tif (!stream)\n \t\t\t\treturn error(_(\"cannot stream blob %s\"),\n \t\t\t\t\t     oid_to_hex(oid));\n+\t\t\tsize = stream->size;\n \t\t\tflags |= ZIP_STREAM;\n \t\t\tout = NULL;\n \t\t} else {\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 581023495f..b01cb77f4a 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -798,8 +798,6 @@ static int compare_objects(const unsigned char *buf, unsigned long size,\n static int check_collison(struct object_entry *entry)\n {\n \tstruct compare_data data;\n-\tenum object_type type;\n-\tunsigned long size;\n \n \tif (entry->size <= repo_settings_get_big_file_threshold(the_repository) ||\n \t    entry->type != OBJ_BLOB)\n@@ -807,11 +805,10 @@ static int check_collison(struct object_entry *entry)\n \n \tmemset(&data, 0, sizeof(data));\n \tdata.entry = entry;\n-\tdata.st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n-\t\t\t\t       &type, &size, NULL);\n+\tdata.st = odb_read_stream_open(the_repository->objects, &entry->idx.oid, NULL);\n \tif (!data.st)\n \t\treturn -1;\n-\tif (size != entry->size || type != entry->type)\n+\tif (data.st->size != entry->size || data.st->type != entry->type)\n \t\tdie(_(\"SHA1 COLLISION FOUND WITH %s !\"),\n \t\t    oid_to_hex(&entry->idx.oid));\n \tunpack_data(entry, compare_objects, &data);\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f109e26786..0d1d6995bf 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -521,9 +521,11 @@ static unsigned long write_no_reuse_object(struct hashfile *f, struct object_ent\n \t\t    oe_size_greater_than(&to_pack, entry,\n \t\t\t\t\t repo_settings_get_big_file_threshold(the_repository)) &&\n \t\t    (st = odb_read_stream_open(the_repository->objects, &entry->idx.oid,\n-\t\t\t\t\t       &type, &size, NULL)) != NULL)\n+\t\t\t\t\t       NULL)) != NULL) {\n \t\t\tbuf = NULL;\n-\t\telse {\n+\t\t\ttype = st->type;\n+\t\t\tsize = st->size;\n+\t\t} else {\n \t\t\tbuf = odb_read_object(the_repository->objects,\n \t\t\t\t\t      &entry->idx.oid, &type,\n \t\t\t\t\t      &size);\ndiff --git a/object-file.c b/object-file.c\nindex 9601fdb12d..12177a7dd7 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -132,19 +132,17 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n int stream_object_signature(struct repository *r, const struct object_id *oid)\n {\n \tstruct object_id real_oid;\n-\tunsigned long size;\n-\tenum object_type obj_type;\n \tstruct odb_read_stream *st;\n \tstruct git_hash_ctx c;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \n-\tst = odb_read_stream_open(r->objects, oid, &obj_type, &size, NULL);\n+\tst = odb_read_stream_open(r->objects, oid, NULL);\n \tif (!st)\n \t\treturn -1;\n \n \t/* Generate the header */\n-\thdrlen = format_object_header(hdr, sizeof(hdr), obj_type, size);\n+\thdrlen = format_object_header(hdr, sizeof(hdr), st->type, st->size);\n \n \t/* Sha1.. */\n \tr->hash_algo->init_fn(&c);\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex 7ef58adaa2..745cd486fb 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -214,8 +214,6 @@ ssize_t odb_read_stream_read(struct odb_read_stream *st, void *buf, size_t sz)\n \n struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n \t\t\t\t\t     const struct object_id *oid,\n-\t\t\t\t\t     enum object_type *type,\n-\t\t\t\t\t     unsigned long *size,\n \t\t\t\t\t     struct stream_filter *filter)\n {\n \tstruct odb_read_stream *st;\n@@ -235,8 +233,6 @@ struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n \t\tst = nst;\n \t}\n \n-\t*size = st->size;\n-\t*type = st->type;\n \treturn st;\n }\n \n@@ -247,18 +243,16 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  int can_seek)\n {\n \tstruct odb_read_stream *st;\n-\tenum object_type type;\n-\tunsigned long sz;\n \tssize_t kept = 0;\n \tint result = -1;\n \n-\tst = odb_read_stream_open(odb, oid, &type, &sz, filter);\n+\tst = odb_read_stream_open(odb, oid, filter);\n \tif (!st) {\n \t\tif (filter)\n \t\t\tfree_stream_filter(filter);\n \t\treturn result;\n \t}\n-\tif (type != OBJ_BLOB)\n+\tif (st->type != OBJ_BLOB)\n \t\tgoto close_and_exit;\n \tfor (;;) {\n \t\tchar buf[1024 * 16];\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex 7cb55213b7..c7861f7e13 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -25,16 +25,13 @@ struct odb_read_stream {\n };\n \n /*\n- * Create a new object stream for the given object database. Populates the type\n- * and size pointers with the object's info. An optional filter can be used to\n- * transform the object's content.\n+ * Create a new object stream for the given object database. An optional filter\n+ * can be used to transform the object's content.\n  *\n  * Returns the stream on success, a `NULL` pointer otherwise.\n  */\n struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n \t\t\t\t\t     const struct object_id *oid,\n-\t\t\t\t\t     enum object_type *type,\n-\t\t\t\t\t     unsigned long *size,\n \t\t\t\t\t     struct stream_filter *filter);\n \n /*\n\n-- \n2.52.0.rc2.482.gaa765fefd0.dirty\n\n"}]}