{"thread":{"id":"65391","subject":"[PATCH 0/6] odb: add write operation to ODB transaction interface","startedAt":"2026-03-31T03:39:00Z","lastAt":"2026-05-15T03:56:17Z","messageCount":57,"participants":["Justin Tobler","Patrick Steinhardt","Junio C Hamano","Jeff King"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"540455","messageId":"20260331033835.2863514-1-jltobler@gmail.com","threadId":"65391","inReplyTo":null,"subject":"[PATCH 0/6] odb: add write operation to ODB transaction interface","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T03:38:29Z","receivedAt":"2026-03-31T03:39:00Z","isPatch":true,"body":"Greetings,\n\nThis series lays the groundwork for introducing write operations to the\nODB transaction interface. The eventual goal is for all object writes\nperformed within a transaction to go through this interface explicitly,\nrather than implicitly relying on the transaction to reconfigure ODB\nsources so that writes are redirected to a temporary location.\n\nFor now, only `odb_transaction_write_object_stream()` is implemented and\nwires up the existing logic for streaming \"large\" blobs directly into a\npackfile as part of the transaction.\n\nMost of the patches are structural refactorings to enable this, but\npatch 4 introduces a behavioral change in how packfiles that would\nexceed \"pack.packSizeLimit\" are handled.\n\nThanks,\n-Justin\n\nJustin Tobler (6):\n  odb: split `struct odb_transaction` into separate header\n  odb/transaction: use pluggable `begin_transaction()`\n  object-file: remove flags from transaction packfile writes\n  object-file: avoid fd seekback by checking object size upfront\n  object-file: generalize packfile writes to use odb_write_stream\n  odb/transaction: make `write_object_stream()` pluggable\n\n Makefile                 |   1 +\n builtin/add.c            |   1 +\n builtin/unpack-objects.c |   1 +\n builtin/update-index.c   |   1 +\n cache-tree.c             |   1 +\n meson.build              |   1 +\n object-file.c            | 249 +++++++++++++++++++++------------------\n odb.c                    |  25 ----\n odb.h                    |  31 -----\n odb/transaction.c        |  35 ++++++\n odb/transaction.h        |  57 +++++++++\n read-cache.c             |   1 +\n 12 files changed, 234 insertions(+), 170 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\n\nbase-commit: 5361983c075154725be47b65cca9a2421789e410\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540456","messageId":"20260331033835.2863514-2-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260331033835.2863514-1-jltobler@gmail.com","subject":"[PATCH 1/6] odb: split `struct odb_transaction` into separate header","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T03:38:30Z","receivedAt":"2026-03-31T03:39:01Z","isPatch":true,"body":"The current ODB transaction interface is collocated with other ODB\ninterfaces in \"odb.{c,h}\". Subsequent commits will expand `struct\nodb_transaction` to support write operations on the transaction\ndirectly. To keep things organized and prevent \"odb.{c,h}\" from becoming\nmore unwieldy, split out `struct odb_transaction` into a separate\nheader.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Makefile                 |  1 +\n builtin/add.c            |  1 +\n builtin/unpack-objects.c |  1 +\n builtin/update-index.c   |  1 +\n cache-tree.c             |  1 +\n meson.build              |  1 +\n object-file.c            |  1 +\n odb.c                    | 25 -------------------------\n odb.h                    | 31 -------------------------------\n odb/transaction.c        | 28 ++++++++++++++++++++++++++++\n odb/transaction.h        | 38 ++++++++++++++++++++++++++++++++++++++\n read-cache.c             |  1 +\n 12 files changed, 74 insertions(+), 56 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\ndiff --git a/Makefile b/Makefile\nindex dbf0022054..6342db13e5 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1219,6 +1219,7 @@ LIB_OBJS += odb.o\n LIB_OBJS += odb/source.o\n LIB_OBJS += odb/source-files.o\n LIB_OBJS += odb/streaming.o\n+LIB_OBJS += odb/transaction.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\n LIB_OBJS += oidset.o\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 7737ab878b..c859f66519 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -16,6 +16,7 @@\n #include \"run-command.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"parse-options.h\"\n #include \"path.h\"\n #include \"preload-index.h\"\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 6fc64e9e4b..bc9b1e047e 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -9,6 +9,7 @@\n #include \"hex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"object.h\"\n #include \"delta.h\"\n #include \"pack.h\"\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 8a5907767b..bcc43852ef 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -19,6 +19,7 @@\n #include \"tree-walk.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"refs.h\"\n #include \"resolve-undo.h\"\n #include \"parse-options.h\"\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 60bcc07c3b..f056869cfd 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -10,6 +10,7 @@\n #include \"cache-tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"read-cache-ll.h\"\n #include \"replace-object.h\"\n #include \"repository.h\"\ndiff --git a/meson.build b/meson.build\nindex 8309942d18..6dc23b3af2 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -405,6 +405,7 @@ libgit_sources = [\n   'odb/source.c',\n   'odb/source-files.c',\n   'odb/streaming.c',\n+  'odb/transaction.c',\n   'oid-array.c',\n   'oidmap.c',\n   'oidset.c',\ndiff --git a/object-file.c b/object-file.c\nindex f0b029ff0b..bfbb632cf8 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -21,6 +21,7 @@\n #include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/streaming.h\"\n+#include \"odb/transaction.h\"\n #include \"oidtree.h\"\n #include \"pack.h\"\n #include \"packfile.h\"\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..8c3cbc1b53 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -1069,28 +1069,3 @@ void odb_reprepare(struct object_database *o)\n \n \tobj_read_unlock();\n }\n-\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n-{\n-\tif (odb->transaction)\n-\t\treturn NULL;\n-\n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n-\n-\treturn odb->transaction;\n-}\n-\n-void odb_transaction_commit(struct odb_transaction *transaction)\n-{\n-\tif (!transaction)\n-\t\treturn;\n-\n-\t/*\n-\t * Ensure the transaction ending matches the pending transaction.\n-\t */\n-\tASSERT(transaction == transaction->source->odb->transaction);\n-\n-\ttransaction->commit(transaction);\n-\ttransaction->source->odb->transaction = NULL;\n-\tfree(transaction);\n-}\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..ec5367b13e 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -35,24 +35,6 @@ struct packed_git;\n struct packfile_store;\n struct cached_object_entry;\n \n-/*\n- * A transaction may be started for an object database prior to writing new\n- * objects via odb_transaction_begin(). These objects are not committed until\n- * odb_transaction_commit() is invoked. Only a single transaction may be pending\n- * at a time.\n- *\n- * Each ODB source is expected to implement its own transaction handling.\n- */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n-struct odb_transaction {\n-\t/* The ODB source the transaction is opened against. */\n-\tstruct odb_source *source;\n-\n-\t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n-};\n-\n /*\n  * The object database encapsulates access to objects in a repository. It\n  * manages one or more sources that store the actual objects which are\n@@ -154,19 +136,6 @@ void odb_close(struct object_database *o);\n  */\n void odb_reprepare(struct object_database *o);\n \n-/*\n- * Starts an ODB transaction. Subsequent objects are written to the transaction\n- * and not committed until odb_transaction_commit() is invoked on the\n- * transaction. If the ODB already has a pending transaction, NULL is returned.\n- */\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n-\n-/*\n- * Commits an ODB transaction making the written objects visible. If the\n- * specified transaction is NULL, the function is a no-op.\n- */\n-void odb_transaction_commit(struct odb_transaction *transaction);\n-\n /*\n  * Find source by its object directory path. Returns a `NULL` pointer in case\n  * the source could not be found.\ndiff --git a/odb/transaction.c b/odb/transaction.c\nnew file mode 100644\nindex 0000000000..9bf3f347dc\n--- /dev/null\n+++ b/odb/transaction.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include \"object-file.h\"\n+#include \"odb/transaction.h\"\n+\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n+{\n+\tif (odb->transaction)\n+\t\treturn NULL;\n+\n+\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\n+\treturn odb->transaction;\n+}\n+\n+void odb_transaction_commit(struct odb_transaction *transaction)\n+{\n+\tif (!transaction)\n+\t\treturn;\n+\n+\t/*\n+\t * Ensure the transaction ending matches the pending transaction.\n+\t */\n+\tASSERT(transaction == transaction->source->odb->transaction);\n+\n+\ttransaction->commit(transaction);\n+\ttransaction->source->odb->transaction = NULL;\n+\tfree(transaction);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nnew file mode 100644\nindex 0000000000..a56e392f21\n--- /dev/null\n+++ b/odb/transaction.h\n@@ -0,0 +1,38 @@\n+#ifndef ODB_TRANSACTION_H\n+#define ODB_TRANSACTION_H\n+\n+#include \"odb.h\"\n+#include \"odb/source.h\"\n+\n+/*\n+ * A transaction may be started for an object database prior to writing new\n+ * objects via odb_transaction_begin(). These objects are not committed until\n+ * odb_transaction_commit() is invoked. Only a single transaction may be pending\n+ * at a time.\n+ *\n+ * Each ODB source is expected to implement its own transaction handling.\n+ */\n+struct odb_transaction;\n+typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n+struct odb_transaction {\n+\t/* The ODB source the transaction is opened against. */\n+\tstruct odb_source *source;\n+\n+\t/* The ODB source specific callback invoked to commit a transaction. */\n+\todb_transaction_commit_fn commit;\n+};\n+\n+/*\n+ * Starts an ODB transaction. Subsequent objects are written to the transaction\n+ * and not committed until odb_transaction_commit() is invoked on the\n+ * transaction. If the ODB already has a pending transaction, NULL is returned.\n+ */\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n+\n+/*\n+ * Commits an ODB transaction making the written objects visible. If the\n+ * specified transaction is NULL, the function is a no-op.\n+ */\n+void odb_transaction_commit(struct odb_transaction *transaction);\n+\n+#endif\ndiff --git a/read-cache.c b/read-cache.c\nindex 5049f9baca..8147c7e94a 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -20,6 +20,7 @@\n #include \"dir.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"oid-array.h\"\n #include \"tree.h\"\n #include \"commit.h\"\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540457","messageId":"20260331033835.2863514-3-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260331033835.2863514-1-jltobler@gmail.com","subject":"[PATCH 2/6] odb/transaction: use pluggable `begin_transaction()`","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T03:38:31Z","receivedAt":"2026-03-31T03:39:02Z","isPatch":true,"body":"Each ODB source is expected to provide an ODB transaction implementation\nthat should be used when starting a transaction. With d6fc6fe6f8\n(odb/source: make `begin_transaction()` function pluggable, 2026-03-05),\nthe `struct odb_source` now provides a pluggable callback for beginning\ntransactions. Use the callback provided by the ODB source accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n odb/transaction.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 9bf3f347dc..592ac84075 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -1,5 +1,5 @@\n #include \"git-compat-util.h\"\n-#include \"object-file.h\"\n+#include \"odb/source.h\"\n #include \"odb/transaction.h\"\n \n struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n@@ -7,7 +7,7 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n \tif (odb->transaction)\n \t\treturn NULL;\n \n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\todb_source_begin_transaction(odb->sources, &odb->transaction);\n \n \treturn odb->transaction;\n }\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540458","messageId":"20260331033835.2863514-4-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260331033835.2863514-1-jltobler@gmail.com","subject":"[PATCH 3/6] object-file: remove flags from transaction packfile writes","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T03:38:32Z","receivedAt":"2026-03-31T03:39:03Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function handles streaming a\nblob from an fd to compute its object ID and conditionally writes the\nobject directly to a packfile if the INDEX_WRITE_OBJECT flag is set. A\nsubsequent commit will make these packfile object writes part of the\ntransaction interface. Consequently, having the object write be\nconditional on this flag is a bit awkward.\n\nIn preparation for this change, introduce a dedicated\n`hash_blob_stream()` helper that only computes the OID from the fd. This\nis invoked by `index_fd()` instead when the INDEX_WRITE_OBJECT is not\nset. The object write performed via `index_blob_packfile_transaction()`\nis made unconditional accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 124 +++++++++++++++++++++++++++++---------------------\n 1 file changed, 71 insertions(+), 53 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex bfbb632cf8..493173eaf4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1388,11 +1388,10 @@ static int already_written(struct odb_transaction_files *transaction,\n }\n \n /* Lazily create backing packfile for the state */\n-static void prepare_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t unsigned flags)\n+static void prepare_packfile_transaction(struct odb_transaction_files *transaction)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\tif (!(flags & INDEX_WRITE_OBJECT) || state->f)\n+\tif (state->f)\n \t\treturn;\n \n \tstate->f = create_tmp_packfile(transaction->base.source->odb->repo,\n@@ -1405,6 +1404,34 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n+static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n+\t\t\t    struct object_id *result_oid, int fd, size_t size)\n+{\n+\tunsigned char buf[16384];\n+\tstruct git_hash_ctx ctx;\n+\tunsigned header_len;\n+\n+\theader_len = format_object_header((char *)buf, sizeof(buf),\n+\t\t\t\t\t  OBJ_BLOB, size);\n+\thash_algo->init_fn(&ctx);\n+\tgit_hash_update(&ctx, buf, header_len);\n+\n+\twhile (size) {\n+\t\tsize_t rsize = size < sizeof(buf) ? size : sizeof(buf);\n+\t\tssize_t read_result = read_in_full(fd, buf, rsize);\n+\n+\t\tif ((size_t)read_result != rsize)\n+\t\t\treturn -1;\n+\n+\t\tgit_hash_update(&ctx, buf, rsize);\n+\t\tsize -= read_result;\n+\t}\n+\n+\tgit_hash_final_oid(result_oid, &ctx);\n+\n+\treturn 0;\n+}\n+\n /*\n  * Read the contents from fd for size bytes, streaming it to the\n  * packfile in state while updating the hash in ctx. Signal a failure\n@@ -1422,15 +1449,13 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n  */\n static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path,\n-\t\t\t       unsigned flags)\n+\t\t\t       int fd, size_t size, const char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\tint write_object = (flags & INDEX_WRITE_OBJECT);\n \toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n@@ -1465,20 +1490,18 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n-\t\t\tif (write_object) {\n-\t\t\t\tsize_t written = s.next_out - obuf;\n-\n-\t\t\t\t/* would we bust the size limit? */\n-\t\t\t\tif (state->nr_written &&\n-\t\t\t\t    pack_size_limit_cfg &&\n-\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\t\treturn -1;\n-\t\t\t\t}\n-\n-\t\t\t\thashwrite(state->f, obuf, written);\n-\t\t\t\tstate->offset += written;\n+\t\t\tsize_t written = s.next_out - obuf;\n+\n+\t\t\t/* would we bust the size limit? */\n+\t\t\tif (state->nr_written &&\n+\t\t\t    pack_size_limit_cfg &&\n+\t\t\t    pack_size_limit_cfg < state->offset + written) {\n+\t\t\t\tgit_deflate_abort(&s);\n+\t\t\t\treturn -1;\n \t\t\t}\n+\n+\t\t\thashwrite(state->f, obuf, written);\n+\t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n \t\t\ts.avail_out = sizeof(obuf);\n \t\t}\n@@ -1566,8 +1589,7 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  */\n static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n \t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path,\n-\t\t\t\t\t   unsigned flags)\n+\t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n \toff_t seekback, already_hashed_to;\n@@ -1575,7 +1597,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n-\tstruct pack_idx_entry *idx = NULL;\n+\tstruct pack_idx_entry *idx;\n \n \tseekback = lseek(fd, 0, SEEK_CUR);\n \tif (seekback == (off_t)-1)\n@@ -1586,33 +1608,26 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n-\t/* Note: idx is non-NULL when we are writing */\n-\tif ((flags & INDEX_WRITE_OBJECT) != 0) {\n-\t\tCALLOC_ARRAY(idx, 1);\n-\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\thashfile_checkpoint_init(state->f, &checkpoint);\n-\t}\n+\tCALLOC_ARRAY(idx, 1);\n+\tprepare_packfile_transaction(transaction);\n+\thashfile_checkpoint_init(state->f, &checkpoint);\n \n \talready_hashed_to = 0;\n \n \twhile (1) {\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\tif (idx) {\n-\t\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\t\tidx->offset = state->offset;\n-\t\t\tcrc32_begin(state->f);\n-\t\t}\n+\t\tprepare_packfile_transaction(transaction);\n+\t\thashfile_checkpoint(state->f, &checkpoint);\n+\t\tidx->offset = state->offset;\n+\t\tcrc32_begin(state->f);\n+\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path, flags))\n+\t\t\t\t\t fd, size, path))\n \t\t\tbreak;\n \t\t/*\n \t\t * Writing this object to the current pack will make\n \t\t * it too big; we need to truncate it, start a new\n \t\t * pack, and write into it.\n \t\t */\n-\t\tif (!idx)\n-\t\t\tBUG(\"should not happen\");\n \t\thashfile_truncate(state->f, &checkpoint);\n \t\tstate->offset = checkpoint.offset;\n \t\tflush_packfile_transaction(transaction);\n@@ -1620,8 +1635,6 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\treturn error(\"cannot seek back\");\n \t}\n \tgit_hash_final_oid(result_oid, &ctx);\n-\tif (!idx)\n-\t\treturn 0;\n \n \tidx->crc32 = crc32_end(state->f);\n \tif (already_written(transaction, result_oid)) {\n@@ -1642,7 +1655,7 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t     int fd, struct stat *st,\n \t     enum object_type type, const char *path, unsigned flags)\n {\n-\tint ret;\n+\tint ret = 0;\n \n \t/*\n \t * Call xsize_t() only when needed to avoid potentially unnecessary\n@@ -1659,18 +1672,23 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n \t\t\t\t type, path, flags);\n \t} else {\n-\t\tstruct object_database *odb = the_repository->objects;\n-\t\tstruct odb_transaction_files *files_transaction;\n-\t\tstruct odb_transaction *transaction;\n-\n-\t\ttransaction = odb_transaction_begin(odb);\n-\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t base);\n-\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t      path, flags);\n-\t\todb_transaction_commit(transaction);\n+\t\tif (flags & INDEX_WRITE_OBJECT) {\n+\t\t\tstruct object_database *odb = the_repository->objects;\n+\t\t\tstruct odb_transaction_files *files_transaction;\n+\t\t\tstruct odb_transaction *transaction;\n+\n+\t\t\ttransaction = odb_transaction_begin(odb);\n+\t\t\tfiles_transaction = container_of(odb->transaction,\n+\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t base);\n+\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n+\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\todb_transaction_commit(transaction);\n+\t\t} else {\n+\t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n+\t\t\t\t\t     xsize_t(st->st_size)))\n+\t\t\t\tdie(\"failed to hash blob\");\n+\t\t}\n \t}\n \n \tclose(fd);\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540459","messageId":"20260331033835.2863514-5-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260331033835.2863514-1-jltobler@gmail.com","subject":"[PATCH 4/6] object-file: avoid fd seekback by checking object size upfront","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T03:38:33Z","receivedAt":"2026-03-31T03:39:04Z","isPatch":true,"body":"In certain scenarios, Git handles writing blobs that exceed\n\"core.bigFilesThreshold\" differently by streaming the object directly\ninto a packfile. When there is an active ODB transaction, these blobs\nare streamed to the same packfile instead of using a separate packfile\nfor each. If \"pack.packSizeLimit\" is configured and streaming another\nobject causes the packfile to exceed the configured limit, the packfile\nis truncated back to the previous object and the object write is\nrestarted in a new packfile.\n\nThis works fine, but requires the fd being read from to save a\ncheckpoint so it becomes possible to rewind the input source via seeking\nback to a known offset at the beginning. In a subsequent commit, blob\nstreaming is converted to use `struct odb_write_stream` as a more\ngeneric input source instead of an fd which doesn't provide a mechanism\nfor rewinding.\n\nFor this use case though, rewinding the fd is not strictly necessary\nbecause the inflated size of the object is known and can be used to\napproximate whether writing the object would cause the packfile to\nexceed the configured limit prior to writing anything. These blobs\nwritten to the packfile are never deltafied thus the size difference\nbetween what is written versus the inflated size is due to zlib\ncompression. While this does prevent packfiles from being filled to the\npotential maximum is some cases, it should be good enough and still\nprevents the packfile from exceeding any configured limit.\n\nUse the inflated blob size to determine whether writing an object to a\npackfile will exceed the configured \"pack.packSizeLimit\".\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 82 +++++++++++++--------------------------------------\n 1 file changed, 21 insertions(+), 61 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 493173eaf4..1de2244ac5 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1434,29 +1434,17 @@ static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n \n /*\n  * Read the contents from fd for size bytes, streaming it to the\n- * packfile in state while updating the hash in ctx. Signal a failure\n- * by returning a negative value when the resulting pack would exceed\n- * the pack size limit and this is not the first object in the pack,\n- * so that the caller can discard what we wrote from the current pack\n- * by truncating it and opening a new one. The caller will then call\n- * us again after rewinding the input fd.\n- *\n- * The already_hashed_to pointer is kept untouched by the caller to\n- * make sure we do not hash the same byte when we are called\n- * again. This way, the caller does not have to checkpoint its hash\n- * status before calling us just in case we ask it to call us again\n- * with a new pack.\n+ * packfile in state while updating the hash in ctx.\n  */\n-static int stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path)\n+static void stream_blob_to_pack(struct transaction_packfile *state,\n+\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n+\t\t\t\tconst char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1473,15 +1461,10 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tif ((size_t)read_result != rsize)\n \t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n \t\t\t\t    (unsigned)rsize, path);\n-\t\t\toffset += rsize;\n-\t\t\tif (*already_hashed_to < offset) {\n-\t\t\t\tsize_t hsize = offset - *already_hashed_to;\n-\t\t\t\tif (rsize < hsize)\n-\t\t\t\t\thsize = rsize;\n-\t\t\t\tif (hsize)\n-\t\t\t\t\tgit_hash_update(ctx, ibuf, hsize);\n-\t\t\t\t*already_hashed_to = offset;\n-\t\t\t}\n+\n+\t\t\tif (rsize)\n+\t\t\t\tgit_hash_update(ctx, ibuf, rsize);\n+\n \t\t\ts.next_in = ibuf;\n \t\t\ts.avail_in = rsize;\n \t\t\tsize -= rsize;\n@@ -1492,14 +1475,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n \n-\t\t\t/* would we bust the size limit? */\n-\t\t\tif (state->nr_written &&\n-\t\t\t    pack_size_limit_cfg &&\n-\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\treturn -1;\n-\t\t\t}\n-\n \t\t\thashwrite(state->f, obuf, written);\n \t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n@@ -1516,7 +1491,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t}\n \t}\n \tgit_deflate_end(&s);\n-\treturn 0;\n }\n \n static void flush_packfile_transaction(struct odb_transaction_files *transaction)\n@@ -1592,48 +1566,34 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\toff_t seekback, already_hashed_to;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n \tstruct pack_idx_entry *idx;\n \n-\tseekback = lseek(fd, 0, SEEK_CUR);\n-\tif (seekback == (off_t)-1)\n-\t\treturn error(\"cannot find the current offset\");\n-\n \theader_len = format_object_header((char *)obuf, sizeof(obuf),\n \t\t\t\t\t  OBJ_BLOB, size);\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n+\t/*\n+\t * If writing another object to the packfile could result in it\n+\t * exceeding the configured size limit, flush the current packfile\n+\t * transaction.\n+\t */\n+\tif (state->nr_written && pack_size_limit_cfg &&\n+\t    pack_size_limit_cfg < state->offset + size)\n+\t\tflush_packfile_transaction(transaction);\n+\n \tCALLOC_ARRAY(idx, 1);\n \tprepare_packfile_transaction(transaction);\n \thashfile_checkpoint_init(state->f, &checkpoint);\n \n-\talready_hashed_to = 0;\n-\n-\twhile (1) {\n-\t\tprepare_packfile_transaction(transaction);\n-\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\tidx->offset = state->offset;\n-\t\tcrc32_begin(state->f);\n-\n-\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path))\n-\t\t\tbreak;\n-\t\t/*\n-\t\t * Writing this object to the current pack will make\n-\t\t * it too big; we need to truncate it, start a new\n-\t\t * pack, and write into it.\n-\t\t */\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tflush_packfile_transaction(transaction);\n-\t\tif (lseek(fd, seekback, SEEK_SET) == (off_t)-1)\n-\t\t\treturn error(\"cannot seek back\");\n-\t}\n+\thashfile_checkpoint(state->f, &checkpoint);\n+\tidx->offset = state->offset;\n+\tcrc32_begin(state->f);\n+\tstream_blob_to_pack(state, &ctx, fd, size, path);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540460","messageId":"20260331033835.2863514-6-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260331033835.2863514-1-jltobler@gmail.com","subject":"[PATCH 5/6] object-file: generalize packfile writes to use odb_write_stream","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T03:38:34Z","receivedAt":"2026-03-31T03:39:05Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function streams blob data\ndirectly from an fd. This makes it difficult to reuse as part of a\ngeneric transactional object writing interface.\n\nRefactor the packfile write path to operate on a `struct\nodb_write_stream`, allowing callers to supply data from arbitrary\nsources.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 99 ++++++++++++++++++++++++++++++++++++---------------\n 1 file changed, 70 insertions(+), 29 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 1de2244ac5..4c797d6498 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1433,18 +1433,18 @@ static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n }\n \n /*\n- * Read the contents from fd for size bytes, streaming it to the\n+ * Read the contents from the stream provided, streaming it to the\n  * packfile in state while updating the hash in ctx.\n  */\n static void stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n-\t\t\t\tconst char *path)\n+\t\t\t\tstruct git_hash_ctx *ctx, size_t size,\n+\t\t\t\tstruct odb_write_stream *stream)\n {\n \tgit_zstream s;\n-\tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n+\tsize_t total = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1453,24 +1453,19 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \ts.avail_out = sizeof(obuf) - hdrlen;\n \n \twhile (status != Z_STREAM_END) {\n-\t\tif (size && !s.avail_in) {\n-\t\t\tsize_t rsize = size < sizeof(ibuf) ? size : sizeof(ibuf);\n-\t\t\tssize_t read_result = read_in_full(fd, ibuf, rsize);\n-\t\t\tif (read_result < 0)\n-\t\t\t\tdie_errno(\"failed to read from '%s'\", path);\n-\t\t\tif ((size_t)read_result != rsize)\n-\t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n-\t\t\t\t    (unsigned)rsize, path);\n+\t\tif (!stream->is_finished && !s.avail_in) {\n+\t\t\tunsigned long rsize;\n+\t\t\tunsigned const char *buf = stream->read(stream, &rsize);\n \n \t\t\tif (rsize)\n-\t\t\t\tgit_hash_update(ctx, ibuf, rsize);\n+\t\t\t\tgit_hash_update(ctx, buf, rsize);\n \n-\t\t\ts.next_in = ibuf;\n+\t\t\ts.next_in = (unsigned char *)buf;\n \t\t\ts.avail_in = rsize;\n-\t\t\tsize -= rsize;\n+\t\t\ttotal += rsize;\n \t\t}\n \n-\t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n+\t\tstatus = git_deflate(&s, stream->is_finished ? Z_FINISH : 0);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n@@ -1490,6 +1485,10 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tdie(\"unexpected deflate failure: %d\", status);\n \t\t}\n \t}\n+\n+\tif (total != size)\n+\t\tdie(\"unexpected number of bytes read\");\n+\n \tgit_deflate_end(&s);\n }\n \n@@ -1543,6 +1542,40 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n \todb_reprepare(repo->objects);\n }\n \n+struct read_object_fd_data {\n+\tint fd;\n+\tsize_t size;\n+\tunsigned char buf[16384];\n+};\n+\n+static const void *read_object_fd(struct odb_write_stream *stream,\n+\t\t\t\t  unsigned long *len)\n+{\n+\tstruct read_object_fd_data *data = stream->data;\n+\tssize_t read_result;\n+\tsize_t rsize;\n+\n+\tif (stream->is_finished) {\n+\t\t*len = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\trsize = data->size < sizeof(data->buf) ? data->size : sizeof(data->buf);\n+\tread_result = read_in_full(data->fd, data->buf, rsize);\n+\tif (read_result < 0)\n+\t\tdie_errno(\"failed to read blob data\");\n+\tif ((size_t)read_result != rsize)\n+\t\tdie(\"failed to read %u bytes of blob data\", (unsigned)rsize);\n+\n+\tdata->size -= rsize;\n+\tif (!data->size)\n+\t\tstream->is_finished = 1;\n+\n+\t*len = rsize;\n+\n+\treturn data->buf;\n+}\n+\n /*\n  * This writes the specified object to a packfile. Objects written here\n  * during the same transaction are written to the same packfile. The\n@@ -1561,10 +1594,13 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path)\n+static int index_blob_packfile_transaction(struct odb_transaction *base,\n+\t\t\t\t\t   struct odb_write_stream *stream,\n+\t\t\t\t\t   size_t size, struct object_id *result_oid)\n {\n+\tstruct odb_transaction_files *transaction = container_of(base,\n+\t\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t\t base);\n \tstruct transaction_packfile *state = &transaction->packfile;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n@@ -1593,7 +1629,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \thashfile_checkpoint(state->f, &checkpoint);\n \tidx->offset = state->offset;\n \tcrc32_begin(state->f);\n-\tstream_blob_to_pack(state, &ctx, fd, size, path);\n+\tstream_blob_to_pack(state, &ctx, size, stream);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n@@ -1634,15 +1670,20 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t} else {\n \t\tif (flags & INDEX_WRITE_OBJECT) {\n \t\t\tstruct object_database *odb = the_repository->objects;\n-\t\t\tstruct odb_transaction_files *files_transaction;\n-\t\t\tstruct odb_transaction *transaction;\n-\n-\t\t\ttransaction = odb_transaction_begin(odb);\n-\t\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t\t base);\n-\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n+\t\t\tstruct read_object_fd_data data = {\n+\t\t\t\t.fd = fd,\n+\t\t\t\t.size = xsize_t(st->st_size),\n+\t\t\t};\n+\t\t\tstruct odb_write_stream in_stream = {\n+\t\t\t\t.read = read_object_fd,\n+\t\t\t\t.data = &data,\n+\t\t\t};\n+\n+\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n+\t\t\t\t\t\t\t      &in_stream,\n+\t\t\t\t\t\t\t      xsize_t(st->st_size),\n+\t\t\t\t\t\t\t      oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540461","messageId":"20260331033835.2863514-7-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260331033835.2863514-1-jltobler@gmail.com","subject":"[PATCH 6/6] odb/transaction: make `write_object_stream()` pluggable","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T03:38:35Z","receivedAt":"2026-03-31T03:39:06Z","isPatch":true,"body":"How an ODB transaction handles writing objects is expected to vary\nbetween implementations. Introduce a new `write_object_stream()`\ncallback in `struct odb_transaction` to make this function pluggable.\nWire up `index_blob_packfile_transaction()` for use with `struct\nodb_transaction_files` accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c     |  9 +++++----\n odb/transaction.c |  7 +++++++\n odb/transaction.h | 25 ++++++++++++++++++++++---\n 3 files changed, 34 insertions(+), 7 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 4c797d6498..b1c97faef3 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1680,10 +1680,10 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\t\t\t.data = &data,\n \t\t\t};\n \n-\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n-\t\t\t\t\t\t\t      &in_stream,\n-\t\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t\t      oid);\n+\t\t\tret = odb_transaction_write_object_stream(odb->transaction,\n+\t\t\t\t\t\t\t\t  &in_stream,\n+\t\t\t\t\t\t\t\t  xsize_t(st->st_size),\n+\t\t\t\t\t\t\t\t  oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n@@ -2146,6 +2146,7 @@ struct odb_transaction *odb_transaction_files_begin(struct odb_source *source)\n \ttransaction = xcalloc(1, sizeof(*transaction));\n \ttransaction->base.source = source;\n \ttransaction->base.commit = odb_transaction_files_commit;\n+\ttransaction->base.write_object_stream = index_blob_packfile_transaction;\n \n \treturn &transaction->base;\n }\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 592ac84075..b16e07aebf 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -26,3 +26,10 @@ void odb_transaction_commit(struct odb_transaction *transaction)\n \ttransaction->source->odb->transaction = NULL;\n \tfree(transaction);\n }\n+\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid)\n+{\n+\treturn transaction->write_object_stream(transaction, stream, len, oid);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nindex a56e392f21..584e8de36e 100644\n--- a/odb/transaction.h\n+++ b/odb/transaction.h\n@@ -12,14 +12,24 @@\n  *\n  * Each ODB source is expected to implement its own transaction handling.\n  */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n struct odb_transaction {\n \t/* The ODB source the transaction is opened against. */\n \tstruct odb_source *source;\n \n \t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n+\tvoid (*commit)(struct odb_transaction *transaction);\n+\n+\t/*\n+\t * This callback is expected to write the given object stream into\n+\t * the ODB transaction.\n+\t *\n+\t * The resulting object ID shall be written into the out pointer. The\n+\t * callback is expected to return 0 on success, a negative error code\n+\t * otherwise.\n+\t */\n+\tint (*write_object_stream)(struct odb_transaction *transaction,\n+\t\t\t\t   struct odb_write_stream *stream, size_t len,\n+\t\t\t\t   struct object_id *oid);\n };\n \n /*\n@@ -35,4 +45,13 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n  */\n void odb_transaction_commit(struct odb_transaction *transaction);\n \n+/*\n+ * Writes the object in the provided stream into the transaction. The resulting\n+ * object ID is written into the out pointer. Returns 0 on success, a negative\n+ * error code otherwise.\n+ */\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid);\n+\n #endif\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540484","messageId":"act8SB3hqHvleT_Z@pks.im","threadId":"65391","inReplyTo":"20260331033835.2863514-2-jltobler@gmail.com","subject":"Re: [PATCH 1/6] odb: split `struct odb_transaction` into separate header","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T07:48:24Z","receivedAt":"2026-03-31T07:48:35Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 10:38:30PM -0500, Justin Tobler wrote:\n> The current ODB transaction interface is collocated with other ODB\n\ns/collocated/colocated/\n\nOther than that this patch looks good to me. We don't yet have too much\ncode in the split-out files, but I expect that'll change over time. And\nit also aligns with the \"odb/streaming.h\" interface that we have.\n\nPatrick\n"},{"id":"540485","messageId":"act8UWTmK5iM2iT-@pks.im","threadId":"65391","inReplyTo":"20260331033835.2863514-3-jltobler@gmail.com","subject":"Re: [PATCH 2/6] odb/transaction: use pluggable `begin_transaction()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T07:48:33Z","receivedAt":"2026-03-31T07:48:37Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 10:38:31PM -0500, Justin Tobler wrote:\n> Each ODB source is expected to provide an ODB transaction implementation\n> that should be used when starting a transaction. With d6fc6fe6f8\n> (odb/source: make `begin_transaction()` function pluggable, 2026-03-05),\n> the `struct odb_source` now provides a pluggable callback for beginning\n> transactions. Use the callback provided by the ODB source accordingly.\n\nYup, this is an obvious oversight on my part.\n\n> Signed-off-by: Justin Tobler <jltobler@gmail.com>\n> ---\n>  odb/transaction.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n> \n> diff --git a/odb/transaction.c b/odb/transaction.c\n> index 9bf3f347dc..592ac84075 100644\n> --- a/odb/transaction.c\n> +++ b/odb/transaction.c\n> @@ -1,5 +1,5 @@\n>  #include \"git-compat-util.h\"\n> -#include \"object-file.h\"\n> +#include \"odb/source.h\"\n>  #include \"odb/transaction.h\"\n>  \n>  struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n\nNice to see that we don't have to care about \"object-file.h\" anymore.\n\nPatrick\n"},{"id":"540486","messageId":"act8VlnYzyTGOY7Y@pks.im","threadId":"65391","inReplyTo":"20260331033835.2863514-4-jltobler@gmail.com","subject":"Re: [PATCH 3/6] object-file: remove flags from transaction packfile writes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T07:48:38Z","receivedAt":"2026-03-31T07:48:42Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 10:38:32PM -0500, Justin Tobler wrote:\n> The `index_blob_packfile_transaction()` function handles streaming a\n> blob from an fd to compute its object ID and conditionally writes the\n> object directly to a packfile if the INDEX_WRITE_OBJECT flag is set. A\n> subsequent commit will make these packfile object writes part of the\n> transaction interface. Consequently, having the object write be\n> conditional on this flag is a bit awkward.\n\nOne could argue that we can uplift the object into the transaction\ninterface. But I have to overall agree that this would be rather awkward\nas we're now starting to mix concerns that really shouldn't be mixed.\nMost importantly, it would be weird to have a transaction where you can\nwrite some objects, but not all of them via such a flag.\n\n> In preparation for this change, introduce a dedicated\n> `hash_blob_stream()` helper that only computes the OID from the fd. This\n> is invoked by `index_fd()` instead when the INDEX_WRITE_OBJECT is not\n> set. The object write performed via `index_blob_packfile_transaction()`\n> is made unconditional accordingly.\n\nMakes sense.\n\n> diff --git a/object-file.c b/object-file.c\n> index bfbb632cf8..493173eaf4 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1405,6 +1404,34 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n>  \t\tdie_errno(\"unable to write pack header\");\n>  }\n>  \n> +static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n> +\t\t\t    struct object_id *result_oid, int fd, size_t size)\n\nWe could of course change the interface while at it to also receive the\nobject type. I'll leave it to you though whether you want to go there,\nwe don't have any use case for it right now anyway.\n\n> +{\n> +\tunsigned char buf[16384];\n> +\tstruct git_hash_ctx ctx;\n> +\tunsigned header_len;\n> +\n> +\theader_len = format_object_header((char *)buf, sizeof(buf),\n> +\t\t\t\t\t  OBJ_BLOB, size);\n> +\thash_algo->init_fn(&ctx);\n> +\tgit_hash_update(&ctx, buf, header_len);\n> +\n> +\twhile (size) {\n> +\t\tsize_t rsize = size < sizeof(buf) ? size : sizeof(buf);\n> +\t\tssize_t read_result = read_in_full(fd, buf, rsize);\n> +\n> +\t\tif ((size_t)read_result != rsize)\n> +\t\t\treturn -1;\n\nIt would be a bit cleaner to first check whether `read_result < 0`\nbefore casting.\n\n    if (read_result < 0 || (size_t) read_result != rsize)\n        return -1;\n\nDoesn't make a difference in practice though.\n\n> +\t\tgit_hash_update(&ctx, buf, rsize);\n> +\t\tsize -= read_result;\n> +\t}\n> +\n> +\tgit_hash_final_oid(result_oid, &ctx);\n> +\n> +\treturn 0;\n> +}\n\nOverall, this function really is simple enough to pull out, even if it\nduplicates a tiny amount of logic. Also, it has the benefit that we can\neasily skip deflating the data, which we used to do even if we didn't\nultimately write the data to disk, so it was just pointless busywork.\n\n> @@ -1642,7 +1655,7 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n>  \t     int fd, struct stat *st,\n>  \t     enum object_type type, const char *path, unsigned flags)\n>  {\n> -\tint ret;\n> +\tint ret = 0;\n>  \n>  \t/*\n>  \t * Call xsize_t() only when needed to avoid potentially unnecessary\n\nIn practice this doesn't have to be zero-initialized\n\n> @@ -1659,18 +1672,23 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n>  \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n>  \t\t\t\t type, path, flags);\n>  \t} else {\n> -\t\tstruct object_database *odb = the_repository->objects;\n> -\t\tstruct odb_transaction_files *files_transaction;\n> -\t\tstruct odb_transaction *transaction;\n> -\n> -\t\ttransaction = odb_transaction_begin(odb);\n> -\t\tfiles_transaction = container_of(odb->transaction,\n> -\t\t\t\t\t\t struct odb_transaction_files,\n> -\t\t\t\t\t\t base);\n> -\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> -\t\t\t\t\t\t      xsize_t(st->st_size),\n> -\t\t\t\t\t\t      path, flags);\n> -\t\todb_transaction_commit(transaction);\n> +\t\tif (flags & INDEX_WRITE_OBJECT) {\n> +\t\t\tstruct object_database *odb = the_repository->objects;\n> +\t\t\tstruct odb_transaction_files *files_transaction;\n> +\t\t\tstruct odb_transaction *transaction;\n> +\n> +\t\t\ttransaction = odb_transaction_begin(odb);\n> +\t\t\tfiles_transaction = container_of(odb->transaction,\n> +\t\t\t\t\t\t\t struct odb_transaction_files,\n> +\t\t\t\t\t\t\t base);\n> +\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> +\t\t\t\t\t\t      xsize_t(st->st_size), path);\n> +\t\t\todb_transaction_commit(transaction);\n\nOkay. It's a bit sad that we have to reach into the files backend here,\nbut we already did beforehand, and maybe a subsequent commit will fix\nthis? Reading on.\n\nOn another note, it's somewhat curious that the commit doesn't return an\nerror code. Probably something we should fix eventually.\n\nPatrick\n"},{"id":"540487","messageId":"act8W1BEg6iyUpHB@pks.im","threadId":"65391","inReplyTo":"20260331033835.2863514-5-jltobler@gmail.com","subject":"Re: [PATCH 4/6] object-file: avoid fd seekback by checking object size upfront","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T07:48:43Z","receivedAt":"2026-03-31T07:48:47Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 10:38:33PM -0500, Justin Tobler wrote:\n> In certain scenarios, Git handles writing blobs that exceed\n> \"core.bigFilesThreshold\" differently by streaming the object directly\n> into a packfile. When there is an active ODB transaction, these blobs\n> are streamed to the same packfile instead of using a separate packfile\n> for each. If \"pack.packSizeLimit\" is configured and streaming another\n> object causes the packfile to exceed the configured limit, the packfile\n> is truncated back to the previous object and the object write is\n> restarted in a new packfile.\n> \n> This works fine, but requires the fd being read from to save a\n> checkpoint so it becomes possible to rewind the input source via seeking\n> back to a known offset at the beginning. In a subsequent commit, blob\n> streaming is converted to use `struct odb_write_stream` as a more\n> generic input source instead of an fd which doesn't provide a mechanism\n> for rewinding.\n> \n> For this use case though, rewinding the fd is not strictly necessary\n> because the inflated size of the object is known and can be used to\n> approximate whether writing the object would cause the packfile to\n> exceed the configured limit prior to writing anything. These blobs\n> written to the packfile are never deltafied thus the size difference\n\ns/deltafied/deltified/\n\n> between what is written versus the inflated size is due to zlib\n> compression. While this does prevent packfiles from being filled to the\n> potential maximum is some cases, it should be good enough and still\n> prevents the packfile from exceeding any configured limit.\n> \n> Use the inflated blob size to determine whether writing an object to a\n> packfile will exceed the configured \"pack.packSizeLimit\".\n\nI agree that this is a reasonable tradeoff:\n\n  - For small objects it's probably not going to make a huge difference,\n    as we'd at most waste a couple kilobytes.\n\n  - For large objects we can expect that we wouldn't use deltification\n    anyway due to \"core.bigFileThreshold\".\n\n  - We can expect that in many cases large files will be not compress\n    well, either, as it's more likely than not that a file >512MB (our\n    default limit for \"core.bigFileThreshold\") is going to be a binary\n    file.\n\nWe also nicely document \"pack.packSizeLimit\" as \"rarely useful, and may\nresult in a larger total on-disk size\" in git-config(1), so I think it's\nfair to not bend ourselves over backwards just to make a rarely-useful\nfeature work exactly the same.\n\n> diff --git a/object-file.c b/object-file.c\n> index 493173eaf4..1de2244ac5 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1473,15 +1461,10 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n>  \t\t\tif ((size_t)read_result != rsize)\n>  \t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n>  \t\t\t\t    (unsigned)rsize, path);\n> -\t\t\toffset += rsize;\n> -\t\t\tif (*already_hashed_to < offset) {\n> -\t\t\t\tsize_t hsize = offset - *already_hashed_to;\n> -\t\t\t\tif (rsize < hsize)\n> -\t\t\t\t\thsize = rsize;\n> -\t\t\t\tif (hsize)\n> -\t\t\t\t\tgit_hash_update(ctx, ibuf, hsize);\n> -\t\t\t\t*already_hashed_to = offset;\n> -\t\t\t}\n> +\n> +\t\t\tif (rsize)\n> +\t\t\t\tgit_hash_update(ctx, ibuf, rsize);\n\nIs this guard really needed? I wouldn't expect that we ever try to read\nzero bytes into `ibuf`, and we bail in case we didn't receive the\nexpected number of bytes.\n\nAnd even if we did, `git_hash_update()` works just fine with no data.\n\n> @@ -1592,48 +1566,34 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n>  \t\t\t\t\t   size_t size, const char *path)\n>  {\n>  \tstruct transaction_packfile *state = &transaction->packfile;\n> -\toff_t seekback, already_hashed_to;\n>  \tstruct git_hash_ctx ctx;\n>  \tunsigned char obuf[16384];\n>  \tunsigned header_len;\n>  \tstruct hashfile_checkpoint checkpoint;\n>  \tstruct pack_idx_entry *idx;\n>  \n> -\tseekback = lseek(fd, 0, SEEK_CUR);\n> -\tif (seekback == (off_t)-1)\n> -\t\treturn error(\"cannot find the current offset\");\n\nOkay, no seeking necessary because we don't restart the write anymore.\n\n>  \theader_len = format_object_header((char *)obuf, sizeof(obuf),\n>  \t\t\t\t\t  OBJ_BLOB, size);\n>  \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n>  \tgit_hash_update(&ctx, obuf, header_len);\n>  \n> +\t/*\n> +\t * If writing another object to the packfile could result in it\n> +\t * exceeding the configured size limit, flush the current packfile\n> +\t * transaction.\n> +\t */\n\nDo we want to document that this intentionally works on the inflated\nsize, not the deflated one, with the arguments mentioned in the commit\nmessage?\n\n> +\tif (state->nr_written && pack_size_limit_cfg &&\n> +\t    pack_size_limit_cfg < state->offset + size)\n> +\t\tflush_packfile_transaction(transaction);\n\nAnd we now flush the packfile before writing any object that may cause\nus to bust the size limit. Makes sense.\n\n>  \tCALLOC_ARRAY(idx, 1);\n>  \tprepare_packfile_transaction(transaction);\n>  \thashfile_checkpoint_init(state->f, &checkpoint);\n>  \n> -\talready_hashed_to = 0;\n> -\n> -\twhile (1) {\n> -\t\tprepare_packfile_transaction(transaction);\n> -\t\thashfile_checkpoint(state->f, &checkpoint);\n> -\t\tidx->offset = state->offset;\n> -\t\tcrc32_begin(state->f);\n> -\n> -\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n> -\t\t\t\t\t fd, size, path))\n> -\t\t\tbreak;\n> -\t\t/*\n> -\t\t * Writing this object to the current pack will make\n> -\t\t * it too big; we need to truncate it, start a new\n> -\t\t * pack, and write into it.\n> -\t\t */\n> -\t\thashfile_truncate(state->f, &checkpoint);\n> -\t\tstate->offset = checkpoint.offset;\n> -\t\tflush_packfile_transaction(transaction);\n> -\t\tif (lseek(fd, seekback, SEEK_SET) == (off_t)-1)\n> -\t\t\treturn error(\"cannot seek back\");\n> -\t}\n\nHm. I was briefly wondering whether we'd loop indefinitely in case the\nobject alone is bigger than the packsize. But we have an escape hatch in\n`stream_blob_to_pack()` that special-cases when we haven't written any\ndata yet, so the answer is \"no\".\n\n> +\thashfile_checkpoint(state->f, &checkpoint);\n> +\tidx->offset = state->offset;\n> +\tcrc32_begin(state->f);\n> +\tstream_blob_to_pack(state, &ctx, fd, size, path);\n>  \tgit_hash_final_oid(result_oid, &ctx);\n\nThanks!\n\nPatrick\n"},{"id":"540488","messageId":"act8YM8tMeUr3cJe@pks.im","threadId":"65391","inReplyTo":"20260331033835.2863514-6-jltobler@gmail.com","subject":"Re: [PATCH 5/6] object-file: generalize packfile writes to use odb_write_stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T07:48:48Z","receivedAt":"2026-03-31T07:48:53Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 10:38:34PM -0500, Justin Tobler wrote:\n> diff --git a/object-file.c b/object-file.c\n> index 1de2244ac5..4c797d6498 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1453,24 +1453,19 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n>  \ts.avail_out = sizeof(obuf) - hdrlen;\n>  \n>  \twhile (status != Z_STREAM_END) {\n> -\t\tif (size && !s.avail_in) {\n> -\t\t\tsize_t rsize = size < sizeof(ibuf) ? size : sizeof(ibuf);\n> -\t\t\tssize_t read_result = read_in_full(fd, ibuf, rsize);\n> -\t\t\tif (read_result < 0)\n> -\t\t\t\tdie_errno(\"failed to read from '%s'\", path);\n> -\t\t\tif ((size_t)read_result != rsize)\n> -\t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n> -\t\t\t\t    (unsigned)rsize, path);\n> +\t\tif (!stream->is_finished && !s.avail_in) {\n> +\t\t\tunsigned long rsize;\n> +\t\t\tunsigned const char *buf = stream->read(stream, &rsize);\n>  \n>  \t\t\tif (rsize)\n> -\t\t\t\tgit_hash_update(ctx, ibuf, rsize);\n> +\t\t\t\tgit_hash_update(ctx, buf, rsize);\n>  \n> -\t\t\ts.next_in = ibuf;\n> +\t\t\ts.next_in = (unsigned char *)buf;\n\nA bit ugly that we have to cast away the constness, but oh, well.\n\n> @@ -1490,6 +1485,10 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n>  \t\t\tdie(\"unexpected deflate failure: %d\", status);\n>  \t\t}\n>  \t}\n> +\n> +\tif (total != size)\n> +\t\tdie(\"unexpected number of bytes read\");\n\nDo we want to mention the expected and actual number of bytes?\n\n> @@ -1543,6 +1542,40 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n>  \todb_reprepare(repo->objects);\n>  }\n>  \n> +struct read_object_fd_data {\n> +\tint fd;\n> +\tsize_t size;\n> +\tunsigned char buf[16384];\n> +};\n\nThis interface feels generally useful to me, not just in this subsystem\nhere. Would it make sense to instead expose it in \"odb/transaction.h\"\nas a new `odb_write_stream_from_fd()` function? No need to expose the\nstructure itself, I guess.\n\n> +static const void *read_object_fd(struct odb_write_stream *stream,\n> +\t\t\t\t  unsigned long *len)\n> +{\n> +\tstruct read_object_fd_data *data = stream->data;\n> +\tssize_t read_result;\n> +\tsize_t rsize;\n> +\n> +\tif (stream->is_finished) {\n> +\t\t*len = 0;\n> +\t\treturn NULL;\n> +\t}\n> +\n> +\trsize = data->size < sizeof(data->buf) ? data->size : sizeof(data->buf);\n> +\tread_result = read_in_full(data->fd, data->buf, rsize);\n> +\tif (read_result < 0)\n> +\t\tdie_errno(\"failed to read blob data\");\n\nIt's a bit unfortunate that we die here, but we don't have an easy way\nto return errors. I wonder whether we should refactor the interface a\nbit to maybe take a pointer to a buffer as well as the buffer's length\nand then return an `ssize_t`.\n\n    static ssize_t *read_object_fd(struct odb_write_stream *stream,\n                                   unsigned char *buf,\n                                   size_t buf_len);\n\nThat'd also avoid having to cast away the const-ness, and it allows the\ncaller to control how many bytes they want to read at once.\n\n> +\tif ((size_t)read_result != rsize)\n> +\t\tdie(\"failed to read %u bytes of blob data\", (unsigned)rsize);\n> +\n> +\tdata->size -= rsize;\n\nI feel like `data->size` is misleadingly named now, as it doesn't\nreflect the overall size but rather the number of remaining bytes that\nwe expect.\n\nPatrick\n"},{"id":"540489","messageId":"act8ZWi5On9uQptf@pks.im","threadId":"65391","inReplyTo":"20260331033835.2863514-7-jltobler@gmail.com","subject":"Re: [PATCH 6/6] odb/transaction: make `write_object_stream()` pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T07:48:53Z","receivedAt":"2026-03-31T07:48:58Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 10:38:35PM -0500, Justin Tobler wrote:\n> How an ODB transaction handles writing objects is expected to vary\n> between implementations. Introduce a new `write_object_stream()`\n> callback in `struct odb_transaction` to make this function pluggable.\n> Wire up `index_blob_packfile_transaction()` for use with `struct\n> odb_transaction_files` accordingly.\n> \n> Signed-off-by: Justin Tobler <jltobler@gmail.com>\n> ---\n>  object-file.c     |  9 +++++----\n>  odb/transaction.c |  7 +++++++\n>  odb/transaction.h | 25 ++++++++++++++++++++++---\n>  3 files changed, 34 insertions(+), 7 deletions(-)\n> \n> diff --git a/object-file.c b/object-file.c\n> index 4c797d6498..b1c97faef3 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1680,10 +1680,10 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n>  \t\t\t\t.data = &data,\n>  \t\t\t};\n>  \n> -\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n> -\t\t\t\t\t\t\t      &in_stream,\n> -\t\t\t\t\t\t\t      xsize_t(st->st_size),\n> -\t\t\t\t\t\t\t      oid);\n> +\t\t\tret = odb_transaction_write_object_stream(odb->transaction,\n> +\t\t\t\t\t\t\t\t  &in_stream,\n> +\t\t\t\t\t\t\t\t  xsize_t(st->st_size),\n> +\t\t\t\t\t\t\t\t  oid);\n>  \t\t\todb_transaction_commit(transaction);\n>  \t\t} else {\n>  \t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n> @@ -2146,6 +2146,7 @@ struct odb_transaction *odb_transaction_files_begin(struct odb_source *source)\n>  \ttransaction = xcalloc(1, sizeof(*transaction));\n>  \ttransaction->base.source = source;\n>  \ttransaction->base.commit = odb_transaction_files_commit;\n> +\ttransaction->base.write_object_stream = index_blob_packfile_transaction;\n>  \n>  \treturn &transaction->base;\n>  }\n\nI was originally expecting the upcast to `odb_transaction_files` in\n`index_blob_packfile_transaction()` to go away in this last step, but\nthat of course doesn't make much sense as it now _becomes_ the\nimplementation of `write_object_stream()`.\n\nBut should we rename to `odb_transaction_files_write_object_stream()`?\n\n> diff --git a/odb/transaction.h b/odb/transaction.h\n> index a56e392f21..584e8de36e 100644\n> --- a/odb/transaction.h\n> +++ b/odb/transaction.h\n> @@ -12,14 +12,24 @@\n>   *\n>   * Each ODB source is expected to implement its own transaction handling.\n>   */\n> -struct odb_transaction;\n> -typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n>  struct odb_transaction {\n>  \t/* The ODB source the transaction is opened against. */\n>  \tstruct odb_source *source;\n>  \n>  \t/* The ODB source specific callback invoked to commit a transaction. */\n> -\todb_transaction_commit_fn commit;\n> +\tvoid (*commit)(struct odb_transaction *transaction);\n> +\n> +\t/*\n> +\t * This callback is expected to write the given object stream into\n> +\t * the ODB transaction.\n\nShould we note that for now, the expectation is to always write a blob?\n\nPatrick\n"},{"id":"540517","messageId":"acvSQ_qeA79LV-8y@denethor","threadId":"65391","inReplyTo":"act8SB3hqHvleT_Z@pks.im","subject":"Re: [PATCH 1/6] odb: split `struct odb_transaction` into separate header","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T13:56:06Z","receivedAt":"2026-03-31T13:56:11Z","isPatch":true,"body":"On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> On Mon, Mar 30, 2026 at 10:38:30PM -0500, Justin Tobler wrote:\n> > The current ODB transaction interface is collocated with other ODB\n> \n> s/collocated/colocated/\n\nI wonder if this a regional spelling difference. My spell check doesn't\nseem to like this variant.\n\n-Justin\n"},{"id":"540518","messageId":"acvS74ee-N9zSYDN@denethor","threadId":"65391","inReplyTo":"act8VlnYzyTGOY7Y@pks.im","subject":"Re: [PATCH 3/6] object-file: remove flags from transaction packfile writes","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T14:10:13Z","receivedAt":"2026-03-31T14:10:17Z","isPatch":true,"body":"On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> On Mon, Mar 30, 2026 at 10:38:32PM -0500, Justin Tobler wrote:\n> > diff --git a/object-file.c b/object-file.c\n> > index bfbb632cf8..493173eaf4 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1405,6 +1404,34 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n> >  \t\tdie_errno(\"unable to write pack header\");\n> >  }\n> >  \n> > +static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n> > +\t\t\t    struct object_id *result_oid, int fd, size_t size)\n> \n> We could of course change the interface while at it to also receive the\n> object type. I'll leave it to you though whether you want to go there,\n> we don't have any use case for it right now anyway.\n\nHmmm, this particular helper is always reads the object from an fd. We\ncould of course also update this interface to use a `struct\nodb_writes_stream` to provide the object instead. For now though, since\nwe don't have a use case, I may just leave it as-is.\n\n> > +{\n> > +\tunsigned char buf[16384];\n> > +\tstruct git_hash_ctx ctx;\n> > +\tunsigned header_len;\n> > +\n> > +\theader_len = format_object_header((char *)buf, sizeof(buf),\n> > +\t\t\t\t\t  OBJ_BLOB, size);\n> > +\thash_algo->init_fn(&ctx);\n> > +\tgit_hash_update(&ctx, buf, header_len);\n> > +\n> > +\twhile (size) {\n> > +\t\tsize_t rsize = size < sizeof(buf) ? size : sizeof(buf);\n> > +\t\tssize_t read_result = read_in_full(fd, buf, rsize);\n> > +\n> > +\t\tif ((size_t)read_result != rsize)\n> > +\t\t\treturn -1;\n> \n> It would be a bit cleaner to first check whether `read_result < 0`\n> before casting.\n> \n>     if (read_result < 0 || (size_t) read_result != rsize)\n>         return -1;\n> \n> Doesn't make a difference in practice though.\n\nThat's fair, I can update in the next version to this.\n\n> > +\t\tgit_hash_update(&ctx, buf, rsize);\n> > +\t\tsize -= read_result;\n> > +\t}\n> > +\n> > +\tgit_hash_final_oid(result_oid, &ctx);\n> > +\n> > +\treturn 0;\n> > +}\n> \n> Overall, this function really is simple enough to pull out, even if it\n> duplicates a tiny amount of logic. Also, it has the benefit that we can\n> easily skip deflating the data, which we used to do even if we didn't\n> ultimately write the data to disk, so it was just pointless busywork.\n> \n> > @@ -1642,7 +1655,7 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n> >  \t     int fd, struct stat *st,\n> >  \t     enum object_type type, const char *path, unsigned flags)\n> >  {\n> > -\tint ret;\n> > +\tint ret = 0;\n> >  \n> >  \t/*\n> >  \t * Call xsize_t() only when needed to avoid potentially unnecessary\n> \n> In practice this doesn't have to be zero-initialized\n\nAh, yes. An earlier iteration of this didn't have `hash_blob_stream()`\nreturning error codes. Will update accordingly. Thanks.\n\n> > @@ -1659,18 +1672,23 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n> >  \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n> >  \t\t\t\t type, path, flags);\n> >  \t} else {\n> > -\t\tstruct object_database *odb = the_repository->objects;\n> > -\t\tstruct odb_transaction_files *files_transaction;\n> > -\t\tstruct odb_transaction *transaction;\n> > -\n> > -\t\ttransaction = odb_transaction_begin(odb);\n> > -\t\tfiles_transaction = container_of(odb->transaction,\n> > -\t\t\t\t\t\t struct odb_transaction_files,\n> > -\t\t\t\t\t\t base);\n> > -\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> > -\t\t\t\t\t\t      xsize_t(st->st_size),\n> > -\t\t\t\t\t\t      path, flags);\n> > -\t\todb_transaction_commit(transaction);\n> > +\t\tif (flags & INDEX_WRITE_OBJECT) {\n> > +\t\t\tstruct object_database *odb = the_repository->objects;\n> > +\t\t\tstruct odb_transaction_files *files_transaction;\n> > +\t\t\tstruct odb_transaction *transaction;\n> > +\n> > +\t\t\ttransaction = odb_transaction_begin(odb);\n> > +\t\t\tfiles_transaction = container_of(odb->transaction,\n> > +\t\t\t\t\t\t\t struct odb_transaction_files,\n> > +\t\t\t\t\t\t\t base);\n> > +\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> > +\t\t\t\t\t\t      xsize_t(st->st_size), path);\n> > +\t\t\todb_transaction_commit(transaction);\n> \n> Okay. It's a bit sad that we have to reach into the files backend here,\n> but we already did beforehand, and maybe a subsequent commit will fix\n> this? Reading on.\n\nYes. :)\n\n> On another note, it's somewhat curious that the commit doesn't return an\n> error code. Probably something we should fix eventually.\n\nYa, the error handling for the ODB transaction functions really should\nbubble up any errors instead of just \"die()\"ing. I plan to improve this\nin a future series.\n\nThanks,\n-Justin\n"},{"id":"540519","messageId":"acvV0_7DqGy_q9GY@denethor","threadId":"65391","inReplyTo":"act8W1BEg6iyUpHB@pks.im","subject":"Re: [PATCH 4/6] object-file: avoid fd seekback by checking object size upfront","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T14:14:02Z","receivedAt":"2026-03-31T14:14:04Z","isPatch":true,"body":"On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> On Mon, Mar 30, 2026 at 10:38:33PM -0500, Justin Tobler wrote:\n> > In certain scenarios, Git handles writing blobs that exceed\n> > \"core.bigFilesThreshold\" differently by streaming the object directly\n> > into a packfile. When there is an active ODB transaction, these blobs\n> > are streamed to the same packfile instead of using a separate packfile\n> > for each. If \"pack.packSizeLimit\" is configured and streaming another\n> > object causes the packfile to exceed the configured limit, the packfile\n> > is truncated back to the previous object and the object write is\n> > restarted in a new packfile.\n> > \n> > This works fine, but requires the fd being read from to save a\n> > checkpoint so it becomes possible to rewind the input source via seeking\n> > back to a known offset at the beginning. In a subsequent commit, blob\n> > streaming is converted to use `struct odb_write_stream` as a more\n> > generic input source instead of an fd which doesn't provide a mechanism\n> > for rewinding.\n> > \n> > For this use case though, rewinding the fd is not strictly necessary\n> > because the inflated size of the object is known and can be used to\n> > approximate whether writing the object would cause the packfile to\n> > exceed the configured limit prior to writing anything. These blobs\n> > written to the packfile are never deltafied thus the size difference\n> \n> s/deltafied/deltified/\n\nWill fix.\n\n\n> > diff --git a/object-file.c b/object-file.c\n> > index 493173eaf4..1de2244ac5 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1473,15 +1461,10 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n> >  \t\t\tif ((size_t)read_result != rsize)\n> >  \t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n> >  \t\t\t\t    (unsigned)rsize, path);\n> > -\t\t\toffset += rsize;\n> > -\t\t\tif (*already_hashed_to < offset) {\n> > -\t\t\t\tsize_t hsize = offset - *already_hashed_to;\n> > -\t\t\t\tif (rsize < hsize)\n> > -\t\t\t\t\thsize = rsize;\n> > -\t\t\t\tif (hsize)\n> > -\t\t\t\t\tgit_hash_update(ctx, ibuf, hsize);\n> > -\t\t\t\t*already_hashed_to = offset;\n> > -\t\t\t}\n> > +\n> > +\t\t\tif (rsize)\n> > +\t\t\t\tgit_hash_update(ctx, ibuf, rsize);\n> \n> Is this guard really needed? I wouldn't expect that we ever try to read\n> zero bytes into `ibuf`, and we bail in case we didn't receive the\n> expected number of bytes.\n> \n> And even if we did, `git_hash_update()` works just fine with no data.\n\nYa you are right, this guard is not needed. Will remove in the next\nversion.\n\n\n> >  \theader_len = format_object_header((char *)obuf, sizeof(obuf),\n> >  \t\t\t\t\t  OBJ_BLOB, size);\n> >  \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n> >  \tgit_hash_update(&ctx, obuf, header_len);\n> >  \n> > +\t/*\n> > +\t * If writing another object to the packfile could result in it\n> > +\t * exceeding the configured size limit, flush the current packfile\n> > +\t * transaction.\n> > +\t */\n> \n> Do we want to document that this intentionally works on the inflated\n> size, not the deflated one, with the arguments mentioned in the commit\n> message?\n\nGood suggestion. Will update.\n\nThanks,\n-Justin\n"},{"id":"540520","messageId":"acvX8wdg39xTy-Am@denethor","threadId":"65391","inReplyTo":"act8YM8tMeUr3cJe@pks.im","subject":"Re: [PATCH 5/6] object-file: generalize packfile writes to use odb_write_stream","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T14:31:25Z","receivedAt":"2026-03-31T14:31:29Z","isPatch":true,"body":"On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> On Mon, Mar 30, 2026 at 10:38:34PM -0500, Justin Tobler wrote:\n> > +\n> > +\tif (total != size)\n> > +\t\tdie(\"unexpected number of bytes read\");\n> \n> Do we want to mention the expected and actual number of bytes?\n\nYa, that sounds reasonable. Will update.\n\n> > @@ -1543,6 +1542,40 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n> >  \todb_reprepare(repo->objects);\n> >  }\n> >  \n> > +struct read_object_fd_data {\n> > +\tint fd;\n> > +\tsize_t size;\n> > +\tunsigned char buf[16384];\n> > +};\n> \n> This interface feels generally useful to me, not just in this subsystem\n> here. Would it make sense to instead expose it in \"odb/transaction.h\"\n> as a new `odb_write_stream_from_fd()` function? No need to expose the\n> structure itself, I guess.\n\nHmmm, exposing an `odb_write_stream_from_fd()` function could probably\nbe useful. Would it be better for it to be put in \"odb/streaming.h\"\nthough? Maybe the its use case would always be related to transactions?\n\n> > +static const void *read_object_fd(struct odb_write_stream *stream,\n> > +\t\t\t\t  unsigned long *len)\n> > +{\n> > +\tstruct read_object_fd_data *data = stream->data;\n> > +\tssize_t read_result;\n> > +\tsize_t rsize;\n> > +\n> > +\tif (stream->is_finished) {\n> > +\t\t*len = 0;\n> > +\t\treturn NULL;\n> > +\t}\n> > +\n> > +\trsize = data->size < sizeof(data->buf) ? data->size : sizeof(data->buf);\n> > +\tread_result = read_in_full(data->fd, data->buf, rsize);\n> > +\tif (read_result < 0)\n> > +\t\tdie_errno(\"failed to read blob data\");\n> \n> It's a bit unfortunate that we die here, but we don't have an easy way\n> to return errors. I wonder whether we should refactor the interface a\n> bit to maybe take a pointer to a buffer as well as the buffer's length\n> and then return an `ssize_t`.\n> \n>     static ssize_t *read_object_fd(struct odb_write_stream *stream,\n>                                    unsigned char *buf,\n>                                    size_t buf_len);\n> \n> That'd also avoid having to cast away the const-ness, and it allows the\n> caller to control how many bytes they want to read at once.\n\nI think the above suggestion would work. I believe there is only a\nsingle other usage of `struct odb_write_stream` so updating shouldn't be\nmuch churn.\n\n> > +\tif ((size_t)read_result != rsize)\n> > +\t\tdie(\"failed to read %u bytes of blob data\", (unsigned)rsize);\n> > +\n> > +\tdata->size -= rsize;\n> \n> I feel like `data->size` is misleadingly named now, as it doesn't\n> reflect the overall size but rather the number of remaining bytes that\n> we expect.\n\nThat's fair. The variable starts off as the initial size of the object\nbeing read, but really is just the number of unread bytes. Will update.\n\nThanks,\n-Justin\n"},{"id":"540521","messageId":"acva5CfnflhXExh4@denethor","threadId":"65391","inReplyTo":"act8ZWi5On9uQptf@pks.im","subject":"Re: [PATCH 6/6] odb/transaction: make `write_object_stream()` pluggable","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T14:40:21Z","receivedAt":"2026-03-31T14:40:23Z","isPatch":true,"body":"On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> On Mon, Mar 30, 2026 at 10:38:35PM -0500, Justin Tobler wrote:\n> > How an ODB transaction handles writing objects is expected to vary\n> > between implementations. Introduce a new `write_object_stream()`\n> > callback in `struct odb_transaction` to make this function pluggable.\n> > Wire up `index_blob_packfile_transaction()` for use with `struct\n> > odb_transaction_files` accordingly.\n> > \n> > Signed-off-by: Justin Tobler <jltobler@gmail.com>\n> > ---\n> >  object-file.c     |  9 +++++----\n> >  odb/transaction.c |  7 +++++++\n> >  odb/transaction.h | 25 ++++++++++++++++++++++---\n> >  3 files changed, 34 insertions(+), 7 deletions(-)\n> > \n> > diff --git a/object-file.c b/object-file.c\n> > index 4c797d6498..b1c97faef3 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1680,10 +1680,10 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n> >  \t\t\t\t.data = &data,\n> >  \t\t\t};\n> >  \n> > -\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n> > -\t\t\t\t\t\t\t      &in_stream,\n> > -\t\t\t\t\t\t\t      xsize_t(st->st_size),\n> > -\t\t\t\t\t\t\t      oid);\n> > +\t\t\tret = odb_transaction_write_object_stream(odb->transaction,\n> > +\t\t\t\t\t\t\t\t  &in_stream,\n> > +\t\t\t\t\t\t\t\t  xsize_t(st->st_size),\n> > +\t\t\t\t\t\t\t\t  oid);\n> >  \t\t\todb_transaction_commit(transaction);\n> >  \t\t} else {\n> >  \t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n> > @@ -2146,6 +2146,7 @@ struct odb_transaction *odb_transaction_files_begin(struct odb_source *source)\n> >  \ttransaction = xcalloc(1, sizeof(*transaction));\n> >  \ttransaction->base.source = source;\n> >  \ttransaction->base.commit = odb_transaction_files_commit;\n> > +\ttransaction->base.write_object_stream = index_blob_packfile_transaction;\n> >  \n> >  \treturn &transaction->base;\n> >  }\n> \n> I was originally expecting the upcast to `odb_transaction_files` in\n> `index_blob_packfile_transaction()` to go away in this last step, but\n> that of course doesn't make much sense as it now _becomes_ the\n> implementation of `write_object_stream()`.\n> \n> But should we rename to `odb_transaction_files_write_object_stream()`?\n\nI initially held off from changing the name because in a followup\nseries I plan to migrate `odb_source_write_object_stream()` into this\nsame function and was going to rename at that time. It doesn't really\nhurt to rename it now though. Will do. \n\n> > diff --git a/odb/transaction.h b/odb/transaction.h\n> > index a56e392f21..584e8de36e 100644\n> > --- a/odb/transaction.h\n> > +++ b/odb/transaction.h\n> > @@ -12,14 +12,24 @@\n> >   *\n> >   * Each ODB source is expected to implement its own transaction handling.\n> >   */\n> > -struct odb_transaction;\n> > -typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n> >  struct odb_transaction {\n> >  \t/* The ODB source the transaction is opened against. */\n> >  \tstruct odb_source *source;\n> >  \n> >  \t/* The ODB source specific callback invoked to commit a transaction. */\n> > -\todb_transaction_commit_fn commit;\n> > +\tvoid (*commit)(struct odb_transaction *transaction);\n> > +\n> > +\t/*\n> > +\t * This callback is expected to write the given object stream into\n> > +\t * the ODB transaction.\n> \n> Should we note that for now, the expectation is to always write a blob?\n\nI plan to drop this restriction in a future series, but it certainly\nmakes sense to document for now. Will update.\n\nThanks,\n-Justin\n"},{"id":"540530","messageId":"xmqqmrzo2dpk.fsf@gitster.g","threadId":"65391","inReplyTo":"acvSQ_qeA79LV-8y@denethor","subject":"Re: [PATCH 1/6] odb: split `struct odb_transaction` into separate header","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-31T15:58:15Z","receivedAt":"2026-03-31T15:58:18Z","isPatch":true,"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n>> On Mon, Mar 30, 2026 at 10:38:30PM -0500, Justin Tobler wrote:\n>> > The current ODB transaction interface is collocated with other ODB\n>> \n>> s/collocated/colocated/\n>\n> I wonder if this a regional spelling difference. My spell check doesn't\n> seem to like this variant.\n\nCollocate is a verb that is defined as words or items being set side\nby side. This word has been around since the early 1500s.  Colocate\nis a verb that means to place two or more items closely together,\nsometimes in order to use a shared resource.\n\nhttps://grammarist.com/spelling/collocate-vs-colocate/\n"},{"id":"540538","messageId":"acv5lsgfw2eKDCkO@denethor","threadId":"65391","inReplyTo":"xmqqmrzo2dpk.fsf@gitster.g","subject":"Re: [PATCH 1/6] odb: split `struct odb_transaction` into separate header","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T16:44:01Z","receivedAt":"2026-03-31T16:44:03Z","isPatch":true,"body":"On 26/03/31 08:58AM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> >> On Mon, Mar 30, 2026 at 10:38:30PM -0500, Justin Tobler wrote:\n> >> > The current ODB transaction interface is collocated with other ODB\n> >> \n> >> s/collocated/colocated/\n> >\n> > I wonder if this a regional spelling difference. My spell check doesn't\n> > seem to like this variant.\n> \n> Collocate is a verb that is defined as words or items being set side\n> by side. This word has been around since the early 1500s.  Colocate\n> is a verb that means to place two or more items closely together,\n> sometimes in order to use a shared resource.\n> \n> https://grammarist.com/spelling/collocate-vs-colocate/\n\nAhh, good to know. Will fix this in my next version. :)\n\nThanks,\n-Justin\n"},{"id":"540580","messageId":"acxRwaUk4XNJiDx9@pks.im","threadId":"65391","inReplyTo":"acvX8wdg39xTy-Am@denethor","subject":"Re: [PATCH 5/6] object-file: generalize packfile writes to use odb_write_stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T22:59:13Z","receivedAt":"2026-03-31T22:59:19Z","isPatch":true,"body":"On Tue, Mar 31, 2026 at 09:31:25AM -0500, Justin Tobler wrote:\n> On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> > On Mon, Mar 30, 2026 at 10:38:34PM -0500, Justin Tobler wrote:\n> > > @@ -1543,6 +1542,40 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n> > >  \todb_reprepare(repo->objects);\n> > >  }\n> > >  \n> > > +struct read_object_fd_data {\n> > > +\tint fd;\n> > > +\tsize_t size;\n> > > +\tunsigned char buf[16384];\n> > > +};\n> > \n> > This interface feels generally useful to me, not just in this subsystem\n> > here. Would it make sense to instead expose it in \"odb/transaction.h\"\n> > as a new `odb_write_stream_from_fd()` function? No need to expose the\n> > structure itself, I guess.\n> \n> Hmmm, exposing an `odb_write_stream_from_fd()` function could probably\n> be useful. Would it be better for it to be put in \"odb/streaming.h\"\n> though? Maybe the its use case would always be related to transactions?\n\nFor now it's certainly always related to writing objects, but you're\nright in that it is not necessarily related to a transaction. After all,\nwe also have `odb_write_object_stream()`.\n\nPutting it into \"odb.h\" would feel off I think, so maybe\n\"odb/streaming.h\" is a good alternative.\n\nPatrick\n"},{"id":"540583","messageId":"acxWV5U-yb2F_0lG@denethor","threadId":"65391","inReplyTo":"acxRwaUk4XNJiDx9@pks.im","subject":"Re: [PATCH 5/6] object-file: generalize packfile writes to use odb_write_stream","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-31T23:21:11Z","receivedAt":"2026-03-31T23:21:13Z","isPatch":true,"body":"On 26/04/01 12:59AM, Patrick Steinhardt wrote:\n> On Tue, Mar 31, 2026 at 09:31:25AM -0500, Justin Tobler wrote:\n> > On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> > > On Mon, Mar 30, 2026 at 10:38:34PM -0500, Justin Tobler wrote:\n> > > > @@ -1543,6 +1542,40 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n> > > >  \todb_reprepare(repo->objects);\n> > > >  }\n> > > >  \n> > > > +struct read_object_fd_data {\n> > > > +\tint fd;\n> > > > +\tsize_t size;\n> > > > +\tunsigned char buf[16384];\n> > > > +};\n> > > \n> > > This interface feels generally useful to me, not just in this subsystem\n> > > here. Would it make sense to instead expose it in \"odb/transaction.h\"\n> > > as a new `odb_write_stream_from_fd()` function? No need to expose the\n> > > structure itself, I guess.\n> > \n> > Hmmm, exposing an `odb_write_stream_from_fd()` function could probably\n> > be useful. Would it be better for it to be put in \"odb/streaming.h\"\n> > though? Maybe the its use case would always be related to transactions?\n> \n> For now it's certainly always related to writing objects, but you're\n> right in that it is not necessarily related to a transaction. After all,\n> we also have `odb_write_object_stream()`.\n> \n> Putting it into \"odb.h\" would feel off I think, so maybe\n> \"odb/streaming.h\" is a good alternative.\n\nOk, I'll put it in \"odb/streaming.h\" for now. Out of curiousity, is\nthere any reason `struct odb_write_stream` isn't currently in\n\"odb/streaming.h\" already? I was thinking it may make sense to move that\ninterface over as well.\n\nThanks,\n-Justin\n"},{"id":"540585","messageId":"acxbgmRW7LxGr5q3@pks.im","threadId":"65391","inReplyTo":"acxWV5U-yb2F_0lG@denethor","subject":"Re: [PATCH 5/6] object-file: generalize packfile writes to use odb_write_stream","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T23:40:50Z","receivedAt":"2026-03-31T23:40:56Z","isPatch":true,"body":"On Tue, Mar 31, 2026 at 06:21:11PM -0500, Justin Tobler wrote:\n> On 26/04/01 12:59AM, Patrick Steinhardt wrote:\n> > On Tue, Mar 31, 2026 at 09:31:25AM -0500, Justin Tobler wrote:\n> > > On 26/03/31 09:48AM, Patrick Steinhardt wrote:\n> > > > On Mon, Mar 30, 2026 at 10:38:34PM -0500, Justin Tobler wrote:\n> > > > > @@ -1543,6 +1542,40 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n> > > > >  \todb_reprepare(repo->objects);\n> > > > >  }\n> > > > >  \n> > > > > +struct read_object_fd_data {\n> > > > > +\tint fd;\n> > > > > +\tsize_t size;\n> > > > > +\tunsigned char buf[16384];\n> > > > > +};\n> > > > \n> > > > This interface feels generally useful to me, not just in this subsystem\n> > > > here. Would it make sense to instead expose it in \"odb/transaction.h\"\n> > > > as a new `odb_write_stream_from_fd()` function? No need to expose the\n> > > > structure itself, I guess.\n> > > \n> > > Hmmm, exposing an `odb_write_stream_from_fd()` function could probably\n> > > be useful. Would it be better for it to be put in \"odb/streaming.h\"\n> > > though? Maybe the its use case would always be related to transactions?\n> > \n> > For now it's certainly always related to writing objects, but you're\n> > right in that it is not necessarily related to a transaction. After all,\n> > we also have `odb_write_object_stream()`.\n> > \n> > Putting it into \"odb.h\" would feel off I think, so maybe\n> > \"odb/streaming.h\" is a good alternative.\n> \n> Ok, I'll put it in \"odb/streaming.h\" for now. Out of curiousity, is\n> there any reason `struct odb_write_stream` isn't currently in\n> \"odb/streaming.h\" already? I was thinking it may make sense to move that\n> interface over as well.\n\nNone that I could really think of.\n\nPatrick\n"},{"id":"540607","messageId":"20260401030316.1847362-1-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260331033835.2863514-1-jltobler@gmail.com","subject":"[PATCH v2 0/7] odb: add write operation to ODB transaction interface","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:08Z","receivedAt":"2026-04-01T03:03:22Z","isPatch":true,"body":"Greetings,\n\nThis series lays the groundwork for introducing write operations to the\nODB transaction interface. The eventual goal is for all object writes\nperformed within a transaction to go through this interface explicitly,\nrather than implicitly relying on the transaction to reconfigure ODB\nsources so that writes are redirected to a temporary location.\n\nFor now, only `odb_transaction_write_object_stream()` is implemented and\nwires up the existing logic for streaming \"large\" blobs directly into a\npackfile as part of the transaction.\n\nMost of the patches are structural refactorings to enable this, but\npatch 4 introduces a behavioral change in how packfiles that would\nexceed \"pack.packSizeLimit\" are handled.\n\nChanges since V1:\n- Fixed some typos\n- Improved error handling\n- Removed unnecessary guard statement\n- Documented in comments why inflated object size is used to approximate\n  if object write will exceed \"pack.packSizeLimit\".\n- Updated `struct odb_write_stream` read() callback to support returning\n  errors and using caller provided buffer\n- Updated the `hash_blob_stream()` function signature to operate on a\n  `struct odb_write_stream` instead of an fd directly\n- Renamed some variables/functions for better clarity\n\nThanks,\n-Justin\n\nJustin Tobler (7):\n  odb: split `struct odb_transaction` into separate header\n  odb/transaction: use pluggable `begin_transaction()`\n  odb: update `struct odb_write_stream` read() callback\n  object-file: remove flags from transaction packfile writes\n  object-file: avoid fd seekback by checking object size upfront\n  object-file: generalize packfile writes to use odb_write_stream\n  odb/transaction: make `write_object_stream()` pluggable\n\n Makefile                 |   1 +\n builtin/add.c            |   1 +\n builtin/unpack-objects.c |  20 ++--\n builtin/update-index.c   |   1 +\n cache-tree.c             |   1 +\n meson.build              |   1 +\n object-file.c            | 234 ++++++++++++++++++++-------------------\n odb.c                    |  25 -----\n odb.h                    |  33 +-----\n odb/streaming.c          |  40 +++++++\n odb/streaming.h          |   8 ++\n odb/transaction.c        |  35 ++++++\n odb/transaction.h        |  57 ++++++++++\n read-cache.c             |   1 +\n 14 files changed, 274 insertions(+), 184 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\nRange-diff against v1:\n1:  65cb557a5a ! 1:  eee372b426 odb: split `struct odb_transaction` into separate header\n    @@ Metadata\n      ## Commit message ##\n         odb: split `struct odb_transaction` into separate header\n     \n    -    The current ODB transaction interface is collocated with other ODB\n    +    The current ODB transaction interface is colocated with other ODB\n         interfaces in \"odb.{c,h}\". Subsequent commits will expand `struct\n         odb_transaction` to support write operations on the transaction\n         directly. To keep things organized and prevent \"odb.{c,h}\" from becoming\n2:  b009d08f3e = 2:  57ac075560 odb/transaction: use pluggable `begin_transaction()`\n-:  ---------- > 3:  556f003d0a odb: update `struct odb_write_stream` read() callback\n3:  ceb44a7d25 ! 4:  a9f0e5ad8a object-file: remove flags from transaction packfile writes\n    @@ Commit message\n         conditional on this flag is a bit awkward.\n     \n         In preparation for this change, introduce a dedicated\n    -    `hash_blob_stream()` helper that only computes the OID from the fd. This\n    -    is invoked by `index_fd()` instead when the INDEX_WRITE_OBJECT is not\n    -    set. The object write performed via `index_blob_packfile_transaction()`\n    -    is made unconditional accordingly.\n    +    `hash_blob_stream()` helper that only computes the OID from a `struct\n    +    odb_write_stream`. This is invoked by `index_fd()` instead when the\n    +    INDEX_WRITE_OBJECT is not set. The object write performed via\n    +    `index_blob_packfile_transaction()` is made unconditional accordingly.\n     \n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n    @@ object-file.c: static void prepare_packfile_transaction(struct odb_transaction_f\n      \t\tdie_errno(\"unable to write pack header\");\n      }\n      \n    -+static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n    -+\t\t\t    struct object_id *result_oid, int fd, size_t size)\n    ++static int hash_blob_stream(struct odb_write_stream *stream,\n    ++\t\t\t    const struct git_hash_algo *hash_algo,\n    ++\t\t\t    struct object_id *result_oid, size_t size)\n     +{\n     +\tunsigned char buf[16384];\n     +\tstruct git_hash_ctx ctx;\n     +\tunsigned header_len;\n    ++\tsize_t total = 0;\n     +\n     +\theader_len = format_object_header((char *)buf, sizeof(buf),\n     +\t\t\t\t\t  OBJ_BLOB, size);\n     +\thash_algo->init_fn(&ctx);\n     +\tgit_hash_update(&ctx, buf, header_len);\n     +\n    -+\twhile (size) {\n    -+\t\tsize_t rsize = size < sizeof(buf) ? size : sizeof(buf);\n    -+\t\tssize_t read_result = read_in_full(fd, buf, rsize);\n    ++\twhile (!stream->is_finished) {\n    ++\t\tssize_t read_result = stream->read(stream, buf, sizeof(buf));\n     +\n    -+\t\tif ((size_t)read_result != rsize)\n    ++\t\tif (read_result < 0)\n     +\t\t\treturn -1;\n     +\n    -+\t\tgit_hash_update(&ctx, buf, rsize);\n    -+\t\tsize -= read_result;\n    ++\t\tgit_hash_update(&ctx, buf, read_result);\n    ++\t\ttotal += read_result;\n     +\t}\n     +\n    ++\tif (total != size)\n    ++\t\treturn -1;\n    ++\n     +\tgit_hash_final_oid(result_oid, &ctx);\n     +\n     +\treturn 0;\n    @@ object-file.c: static int index_blob_packfile_transaction(struct odb_transaction\n      \n      \tidx->crc32 = crc32_end(state->f);\n      \tif (already_written(transaction, result_oid)) {\n    -@@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n    - \t     int fd, struct stat *st,\n    - \t     enum object_type type, const char *path, unsigned flags)\n    - {\n    --\tint ret;\n    -+\tint ret = 0;\n    - \n    - \t/*\n    - \t * Call xsize_t() only when needed to avoid potentially unnecessary\n     @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n      \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n      \t\t\t\t type, path, flags);\n    @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n     -\t\t\t\t\t\t      xsize_t(st->st_size),\n     -\t\t\t\t\t\t      path, flags);\n     -\t\todb_transaction_commit(transaction);\n    ++\t\tstruct odb_write_stream stream = { 0 };\n    ++\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n    ++\n     +\t\tif (flags & INDEX_WRITE_OBJECT) {\n     +\t\t\tstruct object_database *odb = the_repository->objects;\n     +\t\t\tstruct odb_transaction_files *files_transaction;\n    @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n     +\t\t\t\t\t\t      xsize_t(st->st_size), path);\n     +\t\t\todb_transaction_commit(transaction);\n     +\t\t} else {\n    -+\t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n    -+\t\t\t\t\t     xsize_t(st->st_size)))\n    -+\t\t\t\tdie(\"failed to hash blob\");\n    ++\t\t\tret = hash_blob_stream(&stream,\n    ++\t\t\t\t\t       the_repository->hash_algo, oid,\n    ++\t\t\t\t\t       xsize_t(st->st_size));\n     +\t\t}\n    ++\n    ++\t\tfree(stream.data);\n      \t}\n      \n      \tclose(fd);\n    +\n    + ## odb/streaming.c ##\n    +@@ odb/streaming.c: int odb_stream_blob_to_fd(struct object_database *odb,\n    + \todb_read_stream_close(st);\n    + \treturn result;\n    + }\n    ++\n    ++struct read_object_fd_data {\n    ++\tint fd;\n    ++\tsize_t remaining;\n    ++};\n    ++\n    ++static ssize_t read_object_fd(struct odb_write_stream *stream,\n    ++\t\t\t      unsigned char *buf, size_t len)\n    ++{\n    ++\tstruct read_object_fd_data *data = stream->data;\n    ++\tssize_t read_result;\n    ++\tsize_t count;\n    ++\n    ++\tif (stream->is_finished)\n    ++\t\treturn 0;\n    ++\n    ++\tcount = data->remaining < len ? data->remaining : len;\n    ++\tread_result = read_in_full(data->fd, buf, count);\n    ++\tif (read_result < 0 || (size_t)read_result != count)\n    ++\t\treturn -1;\n    ++\n    ++\tdata->remaining -= count;\n    ++\tif (!data->remaining)\n    ++\t\tstream->is_finished = 1;\n    ++\n    ++\treturn read_result;\n    ++}\n    ++\n    ++void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n    ++\t\t\t      size_t size)\n    ++{\n    ++\tstruct read_object_fd_data *data;\n    ++\n    ++\tCALLOC_ARRAY(data, 1);\n    ++\tdata->fd = fd;\n    ++\tdata->remaining = size;\n    ++\n    ++\tstream->data = data;\n    ++\tstream->read = read_object_fd;\n    ++}\n    +\n    + ## odb/streaming.h ##\n    +@@\n    + #define STREAMING_H 1\n    + \n    + #include \"object.h\"\n    ++#include \"odb.h\"\n    + \n    + struct object_database;\n    + struct odb_read_stream;\n    +@@ odb/streaming.h: int odb_stream_blob_to_fd(struct object_database *odb,\n    + \t\t\t  struct stream_filter *filter,\n    + \t\t\t  int can_seek);\n    + \n    ++/*\n    ++ * Sets up an ODB write stream that reads from an fd. The caller is expected to\n    ++ * free the underlying stream data.\n    ++ */\n    ++void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n    ++\t\t\t      size_t size);\n    ++\n    + #endif /* STREAMING_H */\n4:  8849b805e9 ! 5:  b7ac82ed7e object-file: avoid fd seekback by checking object size upfront\n    @@ Commit message\n         object-file: avoid fd seekback by checking object size upfront\n     \n         In certain scenarios, Git handles writing blobs that exceed\n    -    \"core.bigFilesThreshold\" differently by streaming the object directly\n    +    \"core.bigFileThreshold\" differently by streaming the object directly\n         into a packfile. When there is an active ODB transaction, these blobs\n         are streamed to the same packfile instead of using a separate packfile\n         for each. If \"pack.packSizeLimit\" is configured and streaming another\n    @@ Commit message\n         because the inflated size of the object is known and can be used to\n         approximate whether writing the object would cause the packfile to\n         exceed the configured limit prior to writing anything. These blobs\n    -    written to the packfile are never deltafied thus the size difference\n    +    written to the packfile are never deltified thus the size difference\n         between what is written versus the inflated size is due to zlib\n         compression. While this does prevent packfiles from being filled to the\n         potential maximum is some cases, it should be good enough and still\n    @@ Commit message\n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n      ## object-file.c ##\n    -@@ object-file.c: static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n    +@@ object-file.c: static int hash_blob_stream(struct odb_write_stream *stream,\n      \n      /*\n       * Read the contents from fd for size bytes, streaming it to the\n    @@ object-file.c: static int stream_blob_to_pack(struct transaction_packfile *state\n     -\t\t\t\t*already_hashed_to = offset;\n     -\t\t\t}\n     +\n    -+\t\t\tif (rsize)\n    -+\t\t\t\tgit_hash_update(ctx, ibuf, rsize);\n    ++\t\t\tgit_hash_update(ctx, ibuf, rsize);\n     +\n      \t\t\ts.next_in = ibuf;\n      \t\t\ts.avail_in = rsize;\n    @@ object-file.c: static int index_blob_packfile_transaction(struct odb_transaction\n     +\t * If writing another object to the packfile could result in it\n     +\t * exceeding the configured size limit, flush the current packfile\n     +\t * transaction.\n    ++\t *\n    ++\t * Note that this uses the inflated object size as an approximation.\n    ++\t * Blob objects written in this manner are not delta-compressed, so\n    ++\t * the difference between the inflated and on-disk size is limited\n    ++\t * to zlib compression and is sufficient for this check.\n     +\t */\n     +\tif (state->nr_written && pack_size_limit_cfg &&\n     +\t    pack_size_limit_cfg < state->offset + size)\n5:  a25e5a2451 ! 6:  d6c4187a0f object-file: generalize packfile writes to use odb_write_stream\n    @@ Commit message\n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n      ## object-file.c ##\n    -@@ object-file.c: static int hash_blob_stream(const struct git_hash_algo *hash_algo,\n    +@@ object-file.c: static int hash_blob_stream(struct odb_write_stream *stream,\n      }\n      \n      /*\n    @@ object-file.c: static int hash_blob_stream(const struct git_hash_algo *hash_algo\n     +\t\t\t\tstruct odb_write_stream *stream)\n      {\n      \tgit_zstream s;\n    --\tunsigned char ibuf[16384];\n    + \tunsigned char ibuf[16384];\n      \tunsigned char obuf[16384];\n      \tunsigned hdrlen;\n      \tint status = Z_OK;\n    @@ object-file.c: static void stream_blob_to_pack(struct transaction_packfile *stat\n     -\t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n     -\t\t\t\t    (unsigned)rsize, path);\n     +\t\tif (!stream->is_finished && !s.avail_in) {\n    -+\t\t\tunsigned long rsize;\n    -+\t\t\tunsigned const char *buf = stream->read(stream, &rsize);\n    ++\t\t\tssize_t rsize = stream->read(stream, ibuf, sizeof(ibuf));\n    ++\n    ++\t\t\tif (rsize < 0)\n    ++\t\t\t\tdie(\"failed to read blob data\");\n      \n    - \t\t\tif (rsize)\n    --\t\t\t\tgit_hash_update(ctx, ibuf, rsize);\n    -+\t\t\t\tgit_hash_update(ctx, buf, rsize);\n    + \t\t\tgit_hash_update(ctx, ibuf, rsize);\n      \n    --\t\t\ts.next_in = ibuf;\n    -+\t\t\ts.next_in = (unsigned char *)buf;\n    + \t\t\ts.next_in = ibuf;\n      \t\t\ts.avail_in = rsize;\n     -\t\t\tsize -= rsize;\n     +\t\t\ttotal += rsize;\n    @@ object-file.c: static void stream_blob_to_pack(struct transaction_packfile *stat\n      \t}\n     +\n     +\tif (total != size)\n    -+\t\tdie(\"unexpected number of bytes read\");\n    ++\t\tdie(\"read %\" PRIuMAX \" bytes of blob data, but expected %\" PRIuMAX \" bytes\",\n    ++\t\t    (uintmax_t)total, (uintmax_t)size);\n     +\n      \tgit_deflate_end(&s);\n      }\n      \n    -@@ object-file.c: static void flush_packfile_transaction(struct odb_transaction_files *transaction\n    - \todb_reprepare(repo->objects);\n    - }\n    - \n    -+struct read_object_fd_data {\n    -+\tint fd;\n    -+\tsize_t size;\n    -+\tunsigned char buf[16384];\n    -+};\n    -+\n    -+static const void *read_object_fd(struct odb_write_stream *stream,\n    -+\t\t\t\t  unsigned long *len)\n    -+{\n    -+\tstruct read_object_fd_data *data = stream->data;\n    -+\tssize_t read_result;\n    -+\tsize_t rsize;\n    -+\n    -+\tif (stream->is_finished) {\n    -+\t\t*len = 0;\n    -+\t\treturn NULL;\n    -+\t}\n    -+\n    -+\trsize = data->size < sizeof(data->buf) ? data->size : sizeof(data->buf);\n    -+\tread_result = read_in_full(data->fd, data->buf, rsize);\n    -+\tif (read_result < 0)\n    -+\t\tdie_errno(\"failed to read blob data\");\n    -+\tif ((size_t)read_result != rsize)\n    -+\t\tdie(\"failed to read %u bytes of blob data\", (unsigned)rsize);\n    -+\n    -+\tdata->size -= rsize;\n    -+\tif (!data->size)\n    -+\t\tstream->is_finished = 1;\n    -+\n    -+\t*len = rsize;\n    -+\n    -+\treturn data->buf;\n    -+}\n    -+\n    - /*\n    -  * This writes the specified object to a packfile. Objects written here\n    -  * during the same transaction are written to the same packfile. The\n     @@ object-file.c: static void flush_packfile_transaction(struct odb_transaction_files *transaction\n       * binary blobs, they generally do not want to get any conversion, and\n       * callers should avoid this code path when filters are requested.\n    @@ object-file.c: static int index_blob_packfile_transaction(struct odb_transaction\n      \n      \tidx->crc32 = crc32_end(state->f);\n     @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n    - \t} else {\n    + \n      \t\tif (flags & INDEX_WRITE_OBJECT) {\n      \t\t\tstruct object_database *odb = the_repository->objects;\n     -\t\t\tstruct odb_transaction_files *files_transaction;\n    @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n     -\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n     -\t\t\t\t\t\t      xsize_t(st->st_size), path);\n     +\t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n    -+\t\t\tstruct read_object_fd_data data = {\n    -+\t\t\t\t.fd = fd,\n    -+\t\t\t\t.size = xsize_t(st->st_size),\n    -+\t\t\t};\n    -+\t\t\tstruct odb_write_stream in_stream = {\n    -+\t\t\t\t.read = read_object_fd,\n    -+\t\t\t\t.data = &data,\n    -+\t\t\t};\n     +\n     +\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n    -+\t\t\t\t\t\t\t      &in_stream,\n    ++\t\t\t\t\t\t\t      &stream,\n     +\t\t\t\t\t\t\t      xsize_t(st->st_size),\n     +\t\t\t\t\t\t\t      oid);\n      \t\t\todb_transaction_commit(transaction);\n      \t\t} else {\n    - \t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n    + \t\t\tret = hash_blob_stream(&stream,\n6:  bbc62ec512 ! 7:  2b81e94677 odb/transaction: make `write_object_stream()` pluggable\n    @@ Commit message\n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n      ## object-file.c ##\n    +@@ object-file.c: static void flush_packfile_transaction(struct odb_transaction_files *transaction\n    +  * binary blobs, they generally do not want to get any conversion, and\n    +  * callers should avoid this code path when filters are requested.\n    +  */\n    +-static int index_blob_packfile_transaction(struct odb_transaction *base,\n    +-\t\t\t\t\t   struct odb_write_stream *stream,\n    +-\t\t\t\t\t   size_t size, struct object_id *result_oid)\n    ++static int odb_transaction_files_write_object_stream(struct odb_transaction *base,\n    ++\t\t\t\t\t\t     struct odb_write_stream *stream,\n    ++\t\t\t\t\t\t     size_t size,\n    ++\t\t\t\t\t\t     struct object_id *result_oid)\n    + {\n    + \tstruct odb_transaction_files *transaction = container_of(base,\n    + \t\t\t\t\t\t\t\t struct odb_transaction_files,\n     @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n    - \t\t\t\t.data = &data,\n    - \t\t\t};\n    + \t\t\tstruct object_database *odb = the_repository->objects;\n    + \t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n      \n     -\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n    --\t\t\t\t\t\t\t      &in_stream,\n    +-\t\t\t\t\t\t\t      &stream,\n     -\t\t\t\t\t\t\t      xsize_t(st->st_size),\n     -\t\t\t\t\t\t\t      oid);\n     +\t\t\tret = odb_transaction_write_object_stream(odb->transaction,\n    -+\t\t\t\t\t\t\t\t  &in_stream,\n    ++\t\t\t\t\t\t\t\t  &stream,\n     +\t\t\t\t\t\t\t\t  xsize_t(st->st_size),\n     +\t\t\t\t\t\t\t\t  oid);\n      \t\t\todb_transaction_commit(transaction);\n      \t\t} else {\n    - \t\t\tif (hash_blob_stream(the_repository->hash_algo, oid, fd,\n    + \t\t\tret = hash_blob_stream(&stream,\n     @@ object-file.c: struct odb_transaction *odb_transaction_files_begin(struct odb_source *source)\n      \ttransaction = xcalloc(1, sizeof(*transaction));\n      \ttransaction->base.source = source;\n      \ttransaction->base.commit = odb_transaction_files_commit;\n    -+\ttransaction->base.write_object_stream = index_blob_packfile_transaction;\n    ++\ttransaction->base.write_object_stream = odb_transaction_files_write_object_stream;\n      \n      \treturn &transaction->base;\n      }\n    @@ odb/transaction.h\n     +\n     +\t/*\n     +\t * This callback is expected to write the given object stream into\n    -+\t * the ODB transaction.\n    ++\t * the ODB transaction. Note that for now, only blobs support streaming.\n     +\t *\n     +\t * The resulting object ID shall be written into the out pointer. The\n     +\t * callback is expected to return 0 on success, a negative error code\n\nbase-commit: 5361983c075154725be47b65cca9a2421789e410\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540608","messageId":"20260401030316.1847362-2-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v2 1/7] odb: split `struct odb_transaction` into separate header","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:09Z","receivedAt":"2026-04-01T03:03:23Z","isPatch":true,"body":"The current ODB transaction interface is colocated with other ODB\ninterfaces in \"odb.{c,h}\". Subsequent commits will expand `struct\nodb_transaction` to support write operations on the transaction\ndirectly. To keep things organized and prevent \"odb.{c,h}\" from becoming\nmore unwieldy, split out `struct odb_transaction` into a separate\nheader.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Makefile                 |  1 +\n builtin/add.c            |  1 +\n builtin/unpack-objects.c |  1 +\n builtin/update-index.c   |  1 +\n cache-tree.c             |  1 +\n meson.build              |  1 +\n object-file.c            |  1 +\n odb.c                    | 25 -------------------------\n odb.h                    | 31 -------------------------------\n odb/transaction.c        | 28 ++++++++++++++++++++++++++++\n odb/transaction.h        | 38 ++++++++++++++++++++++++++++++++++++++\n read-cache.c             |  1 +\n 12 files changed, 74 insertions(+), 56 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\ndiff --git a/Makefile b/Makefile\nindex dbf0022054..6342db13e5 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1219,6 +1219,7 @@ LIB_OBJS += odb.o\n LIB_OBJS += odb/source.o\n LIB_OBJS += odb/source-files.o\n LIB_OBJS += odb/streaming.o\n+LIB_OBJS += odb/transaction.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\n LIB_OBJS += oidset.o\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 7737ab878b..c859f66519 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -16,6 +16,7 @@\n #include \"run-command.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"parse-options.h\"\n #include \"path.h\"\n #include \"preload-index.h\"\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 6fc64e9e4b..bc9b1e047e 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -9,6 +9,7 @@\n #include \"hex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"object.h\"\n #include \"delta.h\"\n #include \"pack.h\"\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 8a5907767b..bcc43852ef 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -19,6 +19,7 @@\n #include \"tree-walk.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"refs.h\"\n #include \"resolve-undo.h\"\n #include \"parse-options.h\"\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 60bcc07c3b..f056869cfd 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -10,6 +10,7 @@\n #include \"cache-tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"read-cache-ll.h\"\n #include \"replace-object.h\"\n #include \"repository.h\"\ndiff --git a/meson.build b/meson.build\nindex 8309942d18..6dc23b3af2 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -405,6 +405,7 @@ libgit_sources = [\n   'odb/source.c',\n   'odb/source-files.c',\n   'odb/streaming.c',\n+  'odb/transaction.c',\n   'oid-array.c',\n   'oidmap.c',\n   'oidset.c',\ndiff --git a/object-file.c b/object-file.c\nindex f0b029ff0b..bfbb632cf8 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -21,6 +21,7 @@\n #include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/streaming.h\"\n+#include \"odb/transaction.h\"\n #include \"oidtree.h\"\n #include \"pack.h\"\n #include \"packfile.h\"\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..8c3cbc1b53 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -1069,28 +1069,3 @@ void odb_reprepare(struct object_database *o)\n \n \tobj_read_unlock();\n }\n-\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n-{\n-\tif (odb->transaction)\n-\t\treturn NULL;\n-\n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n-\n-\treturn odb->transaction;\n-}\n-\n-void odb_transaction_commit(struct odb_transaction *transaction)\n-{\n-\tif (!transaction)\n-\t\treturn;\n-\n-\t/*\n-\t * Ensure the transaction ending matches the pending transaction.\n-\t */\n-\tASSERT(transaction == transaction->source->odb->transaction);\n-\n-\ttransaction->commit(transaction);\n-\ttransaction->source->odb->transaction = NULL;\n-\tfree(transaction);\n-}\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..ec5367b13e 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -35,24 +35,6 @@ struct packed_git;\n struct packfile_store;\n struct cached_object_entry;\n \n-/*\n- * A transaction may be started for an object database prior to writing new\n- * objects via odb_transaction_begin(). These objects are not committed until\n- * odb_transaction_commit() is invoked. Only a single transaction may be pending\n- * at a time.\n- *\n- * Each ODB source is expected to implement its own transaction handling.\n- */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n-struct odb_transaction {\n-\t/* The ODB source the transaction is opened against. */\n-\tstruct odb_source *source;\n-\n-\t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n-};\n-\n /*\n  * The object database encapsulates access to objects in a repository. It\n  * manages one or more sources that store the actual objects which are\n@@ -154,19 +136,6 @@ void odb_close(struct object_database *o);\n  */\n void odb_reprepare(struct object_database *o);\n \n-/*\n- * Starts an ODB transaction. Subsequent objects are written to the transaction\n- * and not committed until odb_transaction_commit() is invoked on the\n- * transaction. If the ODB already has a pending transaction, NULL is returned.\n- */\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n-\n-/*\n- * Commits an ODB transaction making the written objects visible. If the\n- * specified transaction is NULL, the function is a no-op.\n- */\n-void odb_transaction_commit(struct odb_transaction *transaction);\n-\n /*\n  * Find source by its object directory path. Returns a `NULL` pointer in case\n  * the source could not be found.\ndiff --git a/odb/transaction.c b/odb/transaction.c\nnew file mode 100644\nindex 0000000000..9bf3f347dc\n--- /dev/null\n+++ b/odb/transaction.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include \"object-file.h\"\n+#include \"odb/transaction.h\"\n+\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n+{\n+\tif (odb->transaction)\n+\t\treturn NULL;\n+\n+\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\n+\treturn odb->transaction;\n+}\n+\n+void odb_transaction_commit(struct odb_transaction *transaction)\n+{\n+\tif (!transaction)\n+\t\treturn;\n+\n+\t/*\n+\t * Ensure the transaction ending matches the pending transaction.\n+\t */\n+\tASSERT(transaction == transaction->source->odb->transaction);\n+\n+\ttransaction->commit(transaction);\n+\ttransaction->source->odb->transaction = NULL;\n+\tfree(transaction);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nnew file mode 100644\nindex 0000000000..a56e392f21\n--- /dev/null\n+++ b/odb/transaction.h\n@@ -0,0 +1,38 @@\n+#ifndef ODB_TRANSACTION_H\n+#define ODB_TRANSACTION_H\n+\n+#include \"odb.h\"\n+#include \"odb/source.h\"\n+\n+/*\n+ * A transaction may be started for an object database prior to writing new\n+ * objects via odb_transaction_begin(). These objects are not committed until\n+ * odb_transaction_commit() is invoked. Only a single transaction may be pending\n+ * at a time.\n+ *\n+ * Each ODB source is expected to implement its own transaction handling.\n+ */\n+struct odb_transaction;\n+typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n+struct odb_transaction {\n+\t/* The ODB source the transaction is opened against. */\n+\tstruct odb_source *source;\n+\n+\t/* The ODB source specific callback invoked to commit a transaction. */\n+\todb_transaction_commit_fn commit;\n+};\n+\n+/*\n+ * Starts an ODB transaction. Subsequent objects are written to the transaction\n+ * and not committed until odb_transaction_commit() is invoked on the\n+ * transaction. If the ODB already has a pending transaction, NULL is returned.\n+ */\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n+\n+/*\n+ * Commits an ODB transaction making the written objects visible. If the\n+ * specified transaction is NULL, the function is a no-op.\n+ */\n+void odb_transaction_commit(struct odb_transaction *transaction);\n+\n+#endif\ndiff --git a/read-cache.c b/read-cache.c\nindex 5049f9baca..8147c7e94a 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -20,6 +20,7 @@\n #include \"dir.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"oid-array.h\"\n #include \"tree.h\"\n #include \"commit.h\"\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540609","messageId":"20260401030316.1847362-4-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v2 3/7] odb: update `struct odb_write_stream` read() callback","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:11Z","receivedAt":"2026-04-01T03:03:24Z","isPatch":true,"body":"The `read()` callback used by `struct odb_write_stream` currently\nreturns a pointer to an internal buffer along with the number of bytes\nread. This makes buffer ownership unclear and provides no way to report\nerrors.\n\nUpdate the interface to instead require the caller to provide a buffer,\nand have the callback return the number of bytes written to it or a\nnegative value on error. Call sites are updated accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/unpack-objects.c | 19 +++++++------------\n object-file.c            | 13 ++++++++++---\n odb.h                    |  2 +-\n 3 files changed, 18 insertions(+), 16 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex bc9b1e047e..420619e2cb 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -360,24 +360,21 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \n struct input_zstream_data {\n \tgit_zstream *zstream;\n-\tunsigned char buf[8192];\n \tint status;\n };\n \n-static const void *feed_input_zstream(struct odb_write_stream *in_stream,\n-\t\t\t\t      unsigned long *readlen)\n+static ssize_t feed_input_zstream(struct odb_write_stream *in_stream,\n+\t\t\t\t  unsigned char *buf, size_t buf_len)\n {\n \tstruct input_zstream_data *data = in_stream->data;\n \tgit_zstream *zstream = data->zstream;\n \tvoid *in = fill(1);\n \n-\tif (in_stream->is_finished) {\n-\t\t*readlen = 0;\n-\t\treturn NULL;\n-\t}\n+\tif (in_stream->is_finished)\n+\t\treturn 0;\n \n-\tzstream->next_out = data->buf;\n-\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_out = buf;\n+\tzstream->avail_out = buf_len;\n \tzstream->next_in = in;\n \tzstream->avail_in = len;\n \n@@ -385,9 +382,7 @@ static const void *feed_input_zstream(struct odb_write_stream *in_stream,\n \n \tin_stream->is_finished = data->status != Z_OK;\n \tuse(len - zstream->avail_in);\n-\t*readlen = sizeof(data->buf) - zstream->avail_out;\n-\n-\treturn data->buf;\n+\treturn buf_len - zstream->avail_out;\n }\n \n static void stream_blob(unsigned long size, unsigned nr)\ndiff --git a/object-file.c b/object-file.c\nindex bfbb632cf8..f3038756fc 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1066,6 +1066,7 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \tstruct git_hash_ctx c, compat_c;\n \tstruct strbuf tmp_file = STRBUF_INIT;\n \tstruct strbuf filename = STRBUF_INIT;\n+\tunsigned char buf[8192];\n \tint dirlen;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n@@ -1098,9 +1099,15 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t\tunsigned char *in0 = stream.next_in;\n \n \t\tif (!stream.avail_in && !in_stream->is_finished) {\n-\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n-\t\t\tstream.next_in = (void *)in;\n-\t\t\tin0 = (unsigned char *)in;\n+\t\t\tssize_t read_len = in_stream->read(in_stream, buf, sizeof(buf));\n+\t\t\tif (read_len < 0) {\n+\t\t\t\terr = -1;\n+\t\t\t\tgoto cleanup;\n+\t\t\t}\n+\n+\t\t\tstream.avail_in = read_len;\n+\t\t\tstream.next_in = buf;\n+\t\t\tin0 = buf;\n \t\t\t/* All data has been read. */\n \t\t\tif (in_stream->is_finished)\n \t\t\t\tflush = 1;\ndiff --git a/odb.h b/odb.h\nindex ec5367b13e..91ec206eed 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -530,7 +530,7 @@ static inline int odb_write_object(struct object_database *odb,\n }\n \n struct odb_write_stream {\n-\tconst void *(*read)(struct odb_write_stream *, unsigned long *len);\n+\tssize_t (*read)(struct odb_write_stream *, unsigned char *, size_t len);\n \tvoid *data;\n \tint is_finished;\n };\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540610","messageId":"20260401030316.1847362-3-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v2 2/7] odb/transaction: use pluggable `begin_transaction()`","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:10Z","receivedAt":"2026-04-01T03:03:24Z","isPatch":true,"body":"Each ODB source is expected to provide an ODB transaction implementation\nthat should be used when starting a transaction. With d6fc6fe6f8\n(odb/source: make `begin_transaction()` function pluggable, 2026-03-05),\nthe `struct odb_source` now provides a pluggable callback for beginning\ntransactions. Use the callback provided by the ODB source accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n odb/transaction.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 9bf3f347dc..592ac84075 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -1,5 +1,5 @@\n #include \"git-compat-util.h\"\n-#include \"object-file.h\"\n+#include \"odb/source.h\"\n #include \"odb/transaction.h\"\n \n struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n@@ -7,7 +7,7 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n \tif (odb->transaction)\n \t\treturn NULL;\n \n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\todb_source_begin_transaction(odb->sources, &odb->transaction);\n \n \treturn odb->transaction;\n }\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540613","messageId":"20260401030316.1847362-5-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v2 4/7] object-file: remove flags from transaction packfile writes","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:12Z","receivedAt":"2026-04-01T03:03:25Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function handles streaming a\nblob from an fd to compute its object ID and conditionally writes the\nobject directly to a packfile if the INDEX_WRITE_OBJECT flag is set. A\nsubsequent commit will make these packfile object writes part of the\ntransaction interface. Consequently, having the object write be\nconditional on this flag is a bit awkward.\n\nIn preparation for this change, introduce a dedicated\n`hash_blob_stream()` helper that only computes the OID from a `struct\nodb_write_stream`. This is invoked by `index_fd()` instead when the\nINDEX_WRITE_OBJECT is not set. The object write performed via\n`index_blob_packfile_transaction()` is made unconditional accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c   | 131 +++++++++++++++++++++++++++++-------------------\n odb/streaming.c |  40 +++++++++++++++\n odb/streaming.h |   8 +++\n 3 files changed, 127 insertions(+), 52 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex f3038756fc..f317a24ccf 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1395,11 +1395,10 @@ static int already_written(struct odb_transaction_files *transaction,\n }\n \n /* Lazily create backing packfile for the state */\n-static void prepare_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t unsigned flags)\n+static void prepare_packfile_transaction(struct odb_transaction_files *transaction)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\tif (!(flags & INDEX_WRITE_OBJECT) || state->f)\n+\tif (state->f)\n \t\treturn;\n \n \tstate->f = create_tmp_packfile(transaction->base.source->odb->repo,\n@@ -1412,6 +1411,38 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n+static int hash_blob_stream(struct odb_write_stream *stream,\n+\t\t\t    const struct git_hash_algo *hash_algo,\n+\t\t\t    struct object_id *result_oid, size_t size)\n+{\n+\tunsigned char buf[16384];\n+\tstruct git_hash_ctx ctx;\n+\tunsigned header_len;\n+\tsize_t total = 0;\n+\n+\theader_len = format_object_header((char *)buf, sizeof(buf),\n+\t\t\t\t\t  OBJ_BLOB, size);\n+\thash_algo->init_fn(&ctx);\n+\tgit_hash_update(&ctx, buf, header_len);\n+\n+\twhile (!stream->is_finished) {\n+\t\tssize_t read_result = stream->read(stream, buf, sizeof(buf));\n+\n+\t\tif (read_result < 0)\n+\t\t\treturn -1;\n+\n+\t\tgit_hash_update(&ctx, buf, read_result);\n+\t\ttotal += read_result;\n+\t}\n+\n+\tif (total != size)\n+\t\treturn -1;\n+\n+\tgit_hash_final_oid(result_oid, &ctx);\n+\n+\treturn 0;\n+}\n+\n /*\n  * Read the contents from fd for size bytes, streaming it to the\n  * packfile in state while updating the hash in ctx. Signal a failure\n@@ -1429,15 +1460,13 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n  */\n static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path,\n-\t\t\t       unsigned flags)\n+\t\t\t       int fd, size_t size, const char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\tint write_object = (flags & INDEX_WRITE_OBJECT);\n \toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n@@ -1472,20 +1501,18 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n-\t\t\tif (write_object) {\n-\t\t\t\tsize_t written = s.next_out - obuf;\n-\n-\t\t\t\t/* would we bust the size limit? */\n-\t\t\t\tif (state->nr_written &&\n-\t\t\t\t    pack_size_limit_cfg &&\n-\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\t\treturn -1;\n-\t\t\t\t}\n-\n-\t\t\t\thashwrite(state->f, obuf, written);\n-\t\t\t\tstate->offset += written;\n+\t\t\tsize_t written = s.next_out - obuf;\n+\n+\t\t\t/* would we bust the size limit? */\n+\t\t\tif (state->nr_written &&\n+\t\t\t    pack_size_limit_cfg &&\n+\t\t\t    pack_size_limit_cfg < state->offset + written) {\n+\t\t\t\tgit_deflate_abort(&s);\n+\t\t\t\treturn -1;\n \t\t\t}\n+\n+\t\t\thashwrite(state->f, obuf, written);\n+\t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n \t\t\ts.avail_out = sizeof(obuf);\n \t\t}\n@@ -1573,8 +1600,7 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  */\n static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n \t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path,\n-\t\t\t\t\t   unsigned flags)\n+\t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n \toff_t seekback, already_hashed_to;\n@@ -1582,7 +1608,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n-\tstruct pack_idx_entry *idx = NULL;\n+\tstruct pack_idx_entry *idx;\n \n \tseekback = lseek(fd, 0, SEEK_CUR);\n \tif (seekback == (off_t)-1)\n@@ -1593,33 +1619,26 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n-\t/* Note: idx is non-NULL when we are writing */\n-\tif ((flags & INDEX_WRITE_OBJECT) != 0) {\n-\t\tCALLOC_ARRAY(idx, 1);\n-\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\thashfile_checkpoint_init(state->f, &checkpoint);\n-\t}\n+\tCALLOC_ARRAY(idx, 1);\n+\tprepare_packfile_transaction(transaction);\n+\thashfile_checkpoint_init(state->f, &checkpoint);\n \n \talready_hashed_to = 0;\n \n \twhile (1) {\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\tif (idx) {\n-\t\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\t\tidx->offset = state->offset;\n-\t\t\tcrc32_begin(state->f);\n-\t\t}\n+\t\tprepare_packfile_transaction(transaction);\n+\t\thashfile_checkpoint(state->f, &checkpoint);\n+\t\tidx->offset = state->offset;\n+\t\tcrc32_begin(state->f);\n+\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path, flags))\n+\t\t\t\t\t fd, size, path))\n \t\t\tbreak;\n \t\t/*\n \t\t * Writing this object to the current pack will make\n \t\t * it too big; we need to truncate it, start a new\n \t\t * pack, and write into it.\n \t\t */\n-\t\tif (!idx)\n-\t\t\tBUG(\"should not happen\");\n \t\thashfile_truncate(state->f, &checkpoint);\n \t\tstate->offset = checkpoint.offset;\n \t\tflush_packfile_transaction(transaction);\n@@ -1627,8 +1646,6 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\treturn error(\"cannot seek back\");\n \t}\n \tgit_hash_final_oid(result_oid, &ctx);\n-\tif (!idx)\n-\t\treturn 0;\n \n \tidx->crc32 = crc32_end(state->f);\n \tif (already_written(transaction, result_oid)) {\n@@ -1666,18 +1683,28 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n \t\t\t\t type, path, flags);\n \t} else {\n-\t\tstruct object_database *odb = the_repository->objects;\n-\t\tstruct odb_transaction_files *files_transaction;\n-\t\tstruct odb_transaction *transaction;\n-\n-\t\ttransaction = odb_transaction_begin(odb);\n-\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t base);\n-\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t      path, flags);\n-\t\todb_transaction_commit(transaction);\n+\t\tstruct odb_write_stream stream = { 0 };\n+\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n+\n+\t\tif (flags & INDEX_WRITE_OBJECT) {\n+\t\t\tstruct object_database *odb = the_repository->objects;\n+\t\t\tstruct odb_transaction_files *files_transaction;\n+\t\t\tstruct odb_transaction *transaction;\n+\n+\t\t\ttransaction = odb_transaction_begin(odb);\n+\t\t\tfiles_transaction = container_of(odb->transaction,\n+\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t base);\n+\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n+\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\todb_transaction_commit(transaction);\n+\t\t} else {\n+\t\t\tret = hash_blob_stream(&stream,\n+\t\t\t\t\t       the_repository->hash_algo, oid,\n+\t\t\t\t\t       xsize_t(st->st_size));\n+\t\t}\n+\n+\t\tfree(stream.data);\n \t}\n \n \tclose(fd);\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex 5927a12954..85187541c5 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -287,3 +287,43 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \todb_read_stream_close(st);\n \treturn result;\n }\n+\n+struct read_object_fd_data {\n+\tint fd;\n+\tsize_t remaining;\n+};\n+\n+static ssize_t read_object_fd(struct odb_write_stream *stream,\n+\t\t\t      unsigned char *buf, size_t len)\n+{\n+\tstruct read_object_fd_data *data = stream->data;\n+\tssize_t read_result;\n+\tsize_t count;\n+\n+\tif (stream->is_finished)\n+\t\treturn 0;\n+\n+\tcount = data->remaining < len ? data->remaining : len;\n+\tread_result = read_in_full(data->fd, buf, count);\n+\tif (read_result < 0 || (size_t)read_result != count)\n+\t\treturn -1;\n+\n+\tdata->remaining -= count;\n+\tif (!data->remaining)\n+\t\tstream->is_finished = 1;\n+\n+\treturn read_result;\n+}\n+\n+void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n+\t\t\t      size_t size)\n+{\n+\tstruct read_object_fd_data *data;\n+\n+\tCALLOC_ARRAY(data, 1);\n+\tdata->fd = fd;\n+\tdata->remaining = size;\n+\n+\tstream->data = data;\n+\tstream->read = read_object_fd;\n+}\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex c7861f7e13..e5232cd4d1 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -5,6 +5,7 @@\n #define STREAMING_H 1\n \n #include \"object.h\"\n+#include \"odb.h\"\n \n struct object_database;\n struct odb_read_stream;\n@@ -64,4 +65,11 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  struct stream_filter *filter,\n \t\t\t  int can_seek);\n \n+/*\n+ * Sets up an ODB write stream that reads from an fd. The caller is expected to\n+ * free the underlying stream data.\n+ */\n+void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n+\t\t\t      size_t size);\n+\n #endif /* STREAMING_H */\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540611","messageId":"20260401030316.1847362-6-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v2 5/7] object-file: avoid fd seekback by checking object size upfront","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:13Z","receivedAt":"2026-04-01T03:03:26Z","isPatch":true,"body":"In certain scenarios, Git handles writing blobs that exceed\n\"core.bigFileThreshold\" differently by streaming the object directly\ninto a packfile. When there is an active ODB transaction, these blobs\nare streamed to the same packfile instead of using a separate packfile\nfor each. If \"pack.packSizeLimit\" is configured and streaming another\nobject causes the packfile to exceed the configured limit, the packfile\nis truncated back to the previous object and the object write is\nrestarted in a new packfile.\n\nThis works fine, but requires the fd being read from to save a\ncheckpoint so it becomes possible to rewind the input source via seeking\nback to a known offset at the beginning. In a subsequent commit, blob\nstreaming is converted to use `struct odb_write_stream` as a more\ngeneric input source instead of an fd which doesn't provide a mechanism\nfor rewinding.\n\nFor this use case though, rewinding the fd is not strictly necessary\nbecause the inflated size of the object is known and can be used to\napproximate whether writing the object would cause the packfile to\nexceed the configured limit prior to writing anything. These blobs\nwritten to the packfile are never deltified thus the size difference\nbetween what is written versus the inflated size is due to zlib\ncompression. While this does prevent packfiles from being filled to the\npotential maximum is some cases, it should be good enough and still\nprevents the packfile from exceeding any configured limit.\n\nUse the inflated blob size to determine whether writing an object to a\npackfile will exceed the configured \"pack.packSizeLimit\".\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 86 +++++++++++++++------------------------------------\n 1 file changed, 25 insertions(+), 61 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex f317a24ccf..23229fbd95 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1445,29 +1445,17 @@ static int hash_blob_stream(struct odb_write_stream *stream,\n \n /*\n  * Read the contents from fd for size bytes, streaming it to the\n- * packfile in state while updating the hash in ctx. Signal a failure\n- * by returning a negative value when the resulting pack would exceed\n- * the pack size limit and this is not the first object in the pack,\n- * so that the caller can discard what we wrote from the current pack\n- * by truncating it and opening a new one. The caller will then call\n- * us again after rewinding the input fd.\n- *\n- * The already_hashed_to pointer is kept untouched by the caller to\n- * make sure we do not hash the same byte when we are called\n- * again. This way, the caller does not have to checkpoint its hash\n- * status before calling us just in case we ask it to call us again\n- * with a new pack.\n+ * packfile in state while updating the hash in ctx.\n  */\n-static int stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path)\n+static void stream_blob_to_pack(struct transaction_packfile *state,\n+\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n+\t\t\t\tconst char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1484,15 +1472,9 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tif ((size_t)read_result != rsize)\n \t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n \t\t\t\t    (unsigned)rsize, path);\n-\t\t\toffset += rsize;\n-\t\t\tif (*already_hashed_to < offset) {\n-\t\t\t\tsize_t hsize = offset - *already_hashed_to;\n-\t\t\t\tif (rsize < hsize)\n-\t\t\t\t\thsize = rsize;\n-\t\t\t\tif (hsize)\n-\t\t\t\t\tgit_hash_update(ctx, ibuf, hsize);\n-\t\t\t\t*already_hashed_to = offset;\n-\t\t\t}\n+\n+\t\t\tgit_hash_update(ctx, ibuf, rsize);\n+\n \t\t\ts.next_in = ibuf;\n \t\t\ts.avail_in = rsize;\n \t\t\tsize -= rsize;\n@@ -1503,14 +1485,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n \n-\t\t\t/* would we bust the size limit? */\n-\t\t\tif (state->nr_written &&\n-\t\t\t    pack_size_limit_cfg &&\n-\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\treturn -1;\n-\t\t\t}\n-\n \t\t\thashwrite(state->f, obuf, written);\n \t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n@@ -1527,7 +1501,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t}\n \t}\n \tgit_deflate_end(&s);\n-\treturn 0;\n }\n \n static void flush_packfile_transaction(struct odb_transaction_files *transaction)\n@@ -1603,48 +1576,39 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\toff_t seekback, already_hashed_to;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n \tstruct pack_idx_entry *idx;\n \n-\tseekback = lseek(fd, 0, SEEK_CUR);\n-\tif (seekback == (off_t)-1)\n-\t\treturn error(\"cannot find the current offset\");\n-\n \theader_len = format_object_header((char *)obuf, sizeof(obuf),\n \t\t\t\t\t  OBJ_BLOB, size);\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n+\t/*\n+\t * If writing another object to the packfile could result in it\n+\t * exceeding the configured size limit, flush the current packfile\n+\t * transaction.\n+\t *\n+\t * Note that this uses the inflated object size as an approximation.\n+\t * Blob objects written in this manner are not delta-compressed, so\n+\t * the difference between the inflated and on-disk size is limited\n+\t * to zlib compression and is sufficient for this check.\n+\t */\n+\tif (state->nr_written && pack_size_limit_cfg &&\n+\t    pack_size_limit_cfg < state->offset + size)\n+\t\tflush_packfile_transaction(transaction);\n+\n \tCALLOC_ARRAY(idx, 1);\n \tprepare_packfile_transaction(transaction);\n \thashfile_checkpoint_init(state->f, &checkpoint);\n \n-\talready_hashed_to = 0;\n-\n-\twhile (1) {\n-\t\tprepare_packfile_transaction(transaction);\n-\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\tidx->offset = state->offset;\n-\t\tcrc32_begin(state->f);\n-\n-\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path))\n-\t\t\tbreak;\n-\t\t/*\n-\t\t * Writing this object to the current pack will make\n-\t\t * it too big; we need to truncate it, start a new\n-\t\t * pack, and write into it.\n-\t\t */\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tflush_packfile_transaction(transaction);\n-\t\tif (lseek(fd, seekback, SEEK_SET) == (off_t)-1)\n-\t\t\treturn error(\"cannot seek back\");\n-\t}\n+\thashfile_checkpoint(state->f, &checkpoint);\n+\tidx->offset = state->offset;\n+\tcrc32_begin(state->f);\n+\tstream_blob_to_pack(state, &ctx, fd, size, path);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540612","messageId":"20260401030316.1847362-7-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v2 6/7] object-file: generalize packfile writes to use odb_write_stream","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:14Z","receivedAt":"2026-04-01T03:03:27Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function streams blob data\ndirectly from an fd. This makes it difficult to reuse as part of a\ngeneric transactional object writing interface.\n\nRefactor the packfile write path to operate on a `struct\nodb_write_stream`, allowing callers to supply data from arbitrary\nsources.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 55 +++++++++++++++++++++++++++------------------------\n 1 file changed, 29 insertions(+), 26 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 23229fbd95..f7e830c4ec 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1444,18 +1444,19 @@ static int hash_blob_stream(struct odb_write_stream *stream,\n }\n \n /*\n- * Read the contents from fd for size bytes, streaming it to the\n+ * Read the contents from the stream provided, streaming it to the\n  * packfile in state while updating the hash in ctx.\n  */\n static void stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n-\t\t\t\tconst char *path)\n+\t\t\t\tstruct git_hash_ctx *ctx, size_t size,\n+\t\t\t\tstruct odb_write_stream *stream)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n+\tsize_t total = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1464,23 +1465,20 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \ts.avail_out = sizeof(obuf) - hdrlen;\n \n \twhile (status != Z_STREAM_END) {\n-\t\tif (size && !s.avail_in) {\n-\t\t\tsize_t rsize = size < sizeof(ibuf) ? size : sizeof(ibuf);\n-\t\t\tssize_t read_result = read_in_full(fd, ibuf, rsize);\n-\t\t\tif (read_result < 0)\n-\t\t\t\tdie_errno(\"failed to read from '%s'\", path);\n-\t\t\tif ((size_t)read_result != rsize)\n-\t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n-\t\t\t\t    (unsigned)rsize, path);\n+\t\tif (!stream->is_finished && !s.avail_in) {\n+\t\t\tssize_t rsize = stream->read(stream, ibuf, sizeof(ibuf));\n+\n+\t\t\tif (rsize < 0)\n+\t\t\t\tdie(\"failed to read blob data\");\n \n \t\t\tgit_hash_update(ctx, ibuf, rsize);\n \n \t\t\ts.next_in = ibuf;\n \t\t\ts.avail_in = rsize;\n-\t\t\tsize -= rsize;\n+\t\t\ttotal += rsize;\n \t\t}\n \n-\t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n+\t\tstatus = git_deflate(&s, stream->is_finished ? Z_FINISH : 0);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n@@ -1500,6 +1498,11 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tdie(\"unexpected deflate failure: %d\", status);\n \t\t}\n \t}\n+\n+\tif (total != size)\n+\t\tdie(\"read %\" PRIuMAX \" bytes of blob data, but expected %\" PRIuMAX \" bytes\",\n+\t\t    (uintmax_t)total, (uintmax_t)size);\n+\n \tgit_deflate_end(&s);\n }\n \n@@ -1571,10 +1574,13 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path)\n+static int index_blob_packfile_transaction(struct odb_transaction *base,\n+\t\t\t\t\t   struct odb_write_stream *stream,\n+\t\t\t\t\t   size_t size, struct object_id *result_oid)\n {\n+\tstruct odb_transaction_files *transaction = container_of(base,\n+\t\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t\t base);\n \tstruct transaction_packfile *state = &transaction->packfile;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n@@ -1608,7 +1614,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \thashfile_checkpoint(state->f, &checkpoint);\n \tidx->offset = state->offset;\n \tcrc32_begin(state->f);\n-\tstream_blob_to_pack(state, &ctx, fd, size, path);\n+\tstream_blob_to_pack(state, &ctx, size, stream);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n@@ -1652,15 +1658,12 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \n \t\tif (flags & INDEX_WRITE_OBJECT) {\n \t\t\tstruct object_database *odb = the_repository->objects;\n-\t\t\tstruct odb_transaction_files *files_transaction;\n-\t\t\tstruct odb_transaction *transaction;\n-\n-\t\t\ttransaction = odb_transaction_begin(odb);\n-\t\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t\t base);\n-\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n+\n+\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n+\t\t\t\t\t\t\t      &stream,\n+\t\t\t\t\t\t\t      xsize_t(st->st_size),\n+\t\t\t\t\t\t\t      oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tret = hash_blob_stream(&stream,\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540614","messageId":"20260401030316.1847362-8-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v2 7/7] odb/transaction: make `write_object_stream()` pluggable","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T03:03:15Z","receivedAt":"2026-04-01T03:03:27Z","isPatch":true,"body":"How an ODB transaction handles writing objects is expected to vary\nbetween implementations. Introduce a new `write_object_stream()`\ncallback in `struct odb_transaction` to make this function pluggable.\nWire up `index_blob_packfile_transaction()` for use with `struct\nodb_transaction_files` accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c     | 16 +++++++++-------\n odb/transaction.c |  7 +++++++\n odb/transaction.h | 25 ++++++++++++++++++++++---\n 3 files changed, 38 insertions(+), 10 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex f7e830c4ec..45ed87c4d9 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1574,9 +1574,10 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_blob_packfile_transaction(struct odb_transaction *base,\n-\t\t\t\t\t   struct odb_write_stream *stream,\n-\t\t\t\t\t   size_t size, struct object_id *result_oid)\n+static int odb_transaction_files_write_object_stream(struct odb_transaction *base,\n+\t\t\t\t\t\t     struct odb_write_stream *stream,\n+\t\t\t\t\t\t     size_t size,\n+\t\t\t\t\t\t     struct object_id *result_oid)\n {\n \tstruct odb_transaction_files *transaction = container_of(base,\n \t\t\t\t\t\t\t\t struct odb_transaction_files,\n@@ -1660,10 +1661,10 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\t\tstruct object_database *odb = the_repository->objects;\n \t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n \n-\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n-\t\t\t\t\t\t\t      &stream,\n-\t\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t\t      oid);\n+\t\t\tret = odb_transaction_write_object_stream(odb->transaction,\n+\t\t\t\t\t\t\t\t  &stream,\n+\t\t\t\t\t\t\t\t  xsize_t(st->st_size),\n+\t\t\t\t\t\t\t\t  oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tret = hash_blob_stream(&stream,\n@@ -2128,6 +2129,7 @@ struct odb_transaction *odb_transaction_files_begin(struct odb_source *source)\n \ttransaction = xcalloc(1, sizeof(*transaction));\n \ttransaction->base.source = source;\n \ttransaction->base.commit = odb_transaction_files_commit;\n+\ttransaction->base.write_object_stream = odb_transaction_files_write_object_stream;\n \n \treturn &transaction->base;\n }\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 592ac84075..b16e07aebf 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -26,3 +26,10 @@ void odb_transaction_commit(struct odb_transaction *transaction)\n \ttransaction->source->odb->transaction = NULL;\n \tfree(transaction);\n }\n+\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid)\n+{\n+\treturn transaction->write_object_stream(transaction, stream, len, oid);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nindex a56e392f21..854fda06f5 100644\n--- a/odb/transaction.h\n+++ b/odb/transaction.h\n@@ -12,14 +12,24 @@\n  *\n  * Each ODB source is expected to implement its own transaction handling.\n  */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n struct odb_transaction {\n \t/* The ODB source the transaction is opened against. */\n \tstruct odb_source *source;\n \n \t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n+\tvoid (*commit)(struct odb_transaction *transaction);\n+\n+\t/*\n+\t * This callback is expected to write the given object stream into\n+\t * the ODB transaction. Note that for now, only blobs support streaming.\n+\t *\n+\t * The resulting object ID shall be written into the out pointer. The\n+\t * callback is expected to return 0 on success, a negative error code\n+\t * otherwise.\n+\t */\n+\tint (*write_object_stream)(struct odb_transaction *transaction,\n+\t\t\t\t   struct odb_write_stream *stream, size_t len,\n+\t\t\t\t   struct object_id *oid);\n };\n \n /*\n@@ -35,4 +45,13 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n  */\n void odb_transaction_commit(struct odb_transaction *transaction);\n \n+/*\n+ * Writes the object in the provided stream into the transaction. The resulting\n+ * object ID is written into the out pointer. Returns 0 on success, a negative\n+ * error code otherwise.\n+ */\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid);\n+\n #endif\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540632","messageId":"ac0AOEonqMnA20A1@pks.im","threadId":"65391","inReplyTo":"20260401030316.1847362-4-jltobler@gmail.com","subject":"Re: [PATCH v2 3/7] odb: update `struct odb_write_stream` read() callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-01T11:23:36Z","receivedAt":"2026-04-01T11:23:47Z","isPatch":true,"body":"On Tue, Mar 31, 2026 at 10:03:11PM -0500, Justin Tobler wrote:\n> The `read()` callback used by `struct odb_write_stream` currently\n> returns a pointer to an internal buffer along with the number of bytes\n> read. This makes buffer ownership unclear and provides no way to report\n> errors.\n\nNot only that, but it also means that it's impossible for the caller to\ncontrol the chunk size.\n\nThe changes all look straight-forward to me.\n\nPatrick\n"},{"id":"540633","messageId":"ac0AROkfM_GQ9fEW@pks.im","threadId":"65391","inReplyTo":"20260401030316.1847362-5-jltobler@gmail.com","subject":"Re: [PATCH v2 4/7] object-file: remove flags from transaction packfile writes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-01T11:23:48Z","receivedAt":"2026-04-01T11:23:53Z","isPatch":true,"body":"On Tue, Mar 31, 2026 at 10:03:12PM -0500, Justin Tobler wrote:\n> diff --git a/object-file.c b/object-file.c\n> index f3038756fc..f317a24ccf 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1412,6 +1411,38 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n>  \t\tdie_errno(\"unable to write pack header\");\n>  }\n>  \n> +static int hash_blob_stream(struct odb_write_stream *stream,\n> +\t\t\t    const struct git_hash_algo *hash_algo,\n> +\t\t\t    struct object_id *result_oid, size_t size)\n> +{\n> +\tunsigned char buf[16384];\n> +\tstruct git_hash_ctx ctx;\n> +\tunsigned header_len;\n> +\tsize_t total = 0;\n\nOne nit: I think `total` and `size` don't really give a good sense of\nwhich variable tracks what. If this was instead `bytes_hashed` and\n`size` it would become a lot more obvious.\n\n> @@ -1666,18 +1683,28 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n>  \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n>  \t\t\t\t type, path, flags);\n>  \t} else {\n> -\t\tstruct object_database *odb = the_repository->objects;\n> -\t\tstruct odb_transaction_files *files_transaction;\n> -\t\tstruct odb_transaction *transaction;\n> -\n> -\t\ttransaction = odb_transaction_begin(odb);\n> -\t\tfiles_transaction = container_of(odb->transaction,\n> -\t\t\t\t\t\t struct odb_transaction_files,\n> -\t\t\t\t\t\t base);\n> -\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> -\t\t\t\t\t\t      xsize_t(st->st_size),\n> -\t\t\t\t\t\t      path, flags);\n> -\t\todb_transaction_commit(transaction);\n> +\t\tstruct odb_write_stream stream = { 0 };\n> +\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n\nI would assume that `odb_write_stream_from_fd()` knows to fully\ninitialize the stream, so zero-initializing shouldn't be necessary,\nright?\n\n> diff --git a/odb/streaming.h b/odb/streaming.h\n> index c7861f7e13..e5232cd4d1 100644\n> --- a/odb/streaming.h\n> +++ b/odb/streaming.h\n> @@ -5,6 +5,7 @@\n>  #define STREAMING_H 1\n>  \n>  #include \"object.h\"\n> +#include \"odb.h\"\n>  \n>  struct object_database;\n>  struct odb_read_stream;\n> @@ -64,4 +65,11 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n>  \t\t\t  struct stream_filter *filter,\n>  \t\t\t  int can_seek);\n>  \n> +/*\n> + * Sets up an ODB write stream that reads from an fd. The caller is expected to\n> + * free the underlying stream data.\n> + */\n\nHm. Shouldn't we provide an interface that let's the caller do this\nwithout having to know about the stream's internals, like\n`odb_write_stream_release()`?\n\nPatrick\n"},{"id":"540634","messageId":"ac0AcARmeRamQ4Cy@pks.im","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"Re: [PATCH v2 0/7] odb: add write operation to ODB transaction interface","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-01T11:24:32Z","receivedAt":"2026-04-01T11:24:37Z","isPatch":true,"body":"On Tue, Mar 31, 2026 at 10:03:08PM -0500, Justin Tobler wrote:\n> Greetings,\n> \n> This series lays the groundwork for introducing write operations to the\n> ODB transaction interface. The eventual goal is for all object writes\n> performed within a transaction to go through this interface explicitly,\n> rather than implicitly relying on the transaction to reconfigure ODB\n> sources so that writes are redirected to a temporary location.\n> \n> For now, only `odb_transaction_write_object_stream()` is implemented and\n> wires up the existing logic for streaming \"large\" blobs directly into a\n> packfile as part of the transaction.\n> \n> Most of the patches are structural refactorings to enable this, but\n> patch 4 introduces a behavioral change in how packfiles that would\n> exceed \"pack.packSizeLimit\" are handled.\n> \n> Changes since V1:\n> - Fixed some typos\n> - Improved error handling\n> - Removed unnecessary guard statement\n> - Documented in comments why inflated object size is used to approximate\n>   if object write will exceed \"pack.packSizeLimit\".\n> - Updated `struct odb_write_stream` read() callback to support returning\n>   errors and using caller provided buffer\n> - Updated the `hash_blob_stream()` function signature to operate on a\n>   `struct odb_write_stream` instead of an fd directly\n> - Renamed some variables/functions for better clarity\n\nThanks. I've had some smaller nits, but overall I'm happy with this\nstate.\n\nPatrick\n"},{"id":"540645","messageId":"ac0irp8GJSrSD8GU@denethor","threadId":"65391","inReplyTo":"ac0AROkfM_GQ9fEW@pks.im","subject":"Re: [PATCH v2 4/7] object-file: remove flags from transaction packfile writes","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-01T14:02:18Z","receivedAt":"2026-04-01T14:02:24Z","isPatch":true,"body":"On 26/04/01 01:23PM, Patrick Steinhardt wrote:\n> On Tue, Mar 31, 2026 at 10:03:12PM -0500, Justin Tobler wrote:\n> > diff --git a/object-file.c b/object-file.c\n> > index f3038756fc..f317a24ccf 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1412,6 +1411,38 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n> >  \t\tdie_errno(\"unable to write pack header\");\n> >  }\n> >  \n> > +static int hash_blob_stream(struct odb_write_stream *stream,\n> > +\t\t\t    const struct git_hash_algo *hash_algo,\n> > +\t\t\t    struct object_id *result_oid, size_t size)\n> > +{\n> > +\tunsigned char buf[16384];\n> > +\tstruct git_hash_ctx ctx;\n> > +\tunsigned header_len;\n> > +\tsize_t total = 0;\n> \n> One nit: I think `total` and `size` don't really give a good sense of\n> which variable tracks what. If this was instead `bytes_hashed` and\n> `size` it would become a lot more obvious.\n\nThat's fair. I'll update the names in the next version.\n\n> > @@ -1666,18 +1683,28 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n> >  \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n> >  \t\t\t\t type, path, flags);\n> >  \t} else {\n> > -\t\tstruct object_database *odb = the_repository->objects;\n> > -\t\tstruct odb_transaction_files *files_transaction;\n> > -\t\tstruct odb_transaction *transaction;\n> > -\n> > -\t\ttransaction = odb_transaction_begin(odb);\n> > -\t\tfiles_transaction = container_of(odb->transaction,\n> > -\t\t\t\t\t\t struct odb_transaction_files,\n> > -\t\t\t\t\t\t base);\n> > -\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> > -\t\t\t\t\t\t      xsize_t(st->st_size),\n> > -\t\t\t\t\t\t      path, flags);\n> > -\t\todb_transaction_commit(transaction);\n> > +\t\tstruct odb_write_stream stream = { 0 };\n> > +\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n> \n> I would assume that `odb_write_stream_from_fd()` knows to fully\n> initialize the stream, so zero-initializing shouldn't be necessary,\n> right?\n\nCurrently the `is_finished` field is not being initialized, but there\nisn't really any reason we couldn't do that in\n`odb_write_stream_from_fd()` though. Will update accordingly.\n\n> > diff --git a/odb/streaming.h b/odb/streaming.h\n> > index c7861f7e13..e5232cd4d1 100644\n> > --- a/odb/streaming.h\n> > +++ b/odb/streaming.h\n> > @@ -5,6 +5,7 @@\n> >  #define STREAMING_H 1\n> >  \n> >  #include \"object.h\"\n> > +#include \"odb.h\"\n> >  \n> >  struct object_database;\n> >  struct odb_read_stream;\n> > @@ -64,4 +65,11 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n> >  \t\t\t  struct stream_filter *filter,\n> >  \t\t\t  int can_seek);\n> >  \n> > +/*\n> > + * Sets up an ODB write stream that reads from an fd. The caller is expected to\n> > + * free the underlying stream data.\n> > + */\n> \n> Hm. Shouldn't we provide an interface that let's the caller do this\n> without having to know about the stream's internals, like\n> `odb_write_stream_release()`?\n\nI thought about doint this, but hesitated because the other `struct\nodb_write_stream` usage sets up its `data` field differently and\nconsequently would have no need for a `odb_write_stream_release()`. I'm\nprobably overthinking this though and it probably makes sense just to\nadd it.\n\nThanks,\n-Justin\n"},{"id":"540800","messageId":"20260402213220.2651523-1-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260401030316.1847362-1-jltobler@gmail.com","subject":"[PATCH v3 0/7] odb: add write operation to ODB transaction interface","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:13Z","receivedAt":"2026-04-02T21:32:50Z","isPatch":true,"body":"Greetings,\n\nThis series lays the groundwork for introducing write operations to the\nODB transaction interface. The eventual goal is for all object writes\nperformed within a transaction to go through this interface explicitly,\nrather than implicitly relying on the transaction to reconfigure ODB\nsources so that writes are redirected to a temporary location.\n\nFor now, only `odb_transaction_write_object_stream()` is implemented and\nwires up the existing logic for streaming \"large\" blobs directly into a\npackfile as part of the transaction.\n\nMost of the patches are structural refactorings to enable this, but\npatch 4 introduces a behavioral change in how packfiles that would\nexceed \"pack.packSizeLimit\" are handled.\n\nChanges since V2:\n- Renamed some variables to improve clarity\n- Make `odb_write_stream_from_fd()` fully initialize the underlying\n  `struct odb_write_stream`\n- Move `struct odb_write_stream` to \"odb/streaming.h\"\n- Make the `hash_blob_stream()` helper more generic by operating on a\n  `struct odb_write_stream` instead of reading from an fd directly.\n- Introduce an `odb_write_stream_release()` helper to free the\n  underlying stream data.\n\nChanges since V1:\n- Fixed some typos\n- Improved error handling\n- Removed unnecessary guard statement\n- Documented in comments why inflated object size is used to approximate\n  if object write will exceed \"pack.packSizeLimit\".\n- Updated `struct odb_write_stream` read() callback to support returning\n  errors and using caller provided buffer\n- Updated the `hash_blob_stream()` function signature to operate on a\n  `struct odb_write_stream` instead of an fd directly\n- Renamed some variables/functions for better clarity\n\nThanks,\n-Justin\n\nJustin Tobler (7):\n  odb: split `struct odb_transaction` into separate header\n  odb/transaction: use pluggable `begin_transaction()`\n  odb: update `struct odb_write_stream` read() callback\n  object-file: remove flags from transaction packfile writes\n  object-file: avoid fd seekback by checking object size upfront\n  object-file: generalize packfile writes to use odb_write_stream\n  odb/transaction: make `write_object_stream()` pluggable\n\n Makefile                 |   1 +\n builtin/add.c            |   1 +\n builtin/unpack-objects.c |  21 ++--\n builtin/update-index.c   |   1 +\n cache-tree.c             |   1 +\n meson.build              |   1 +\n object-file.c            | 237 ++++++++++++++++++++-------------------\n odb.c                    |  25 -----\n odb.h                    |  37 +-----\n odb/streaming.c          |  51 +++++++++\n odb/streaming.h          |  30 +++++\n odb/transaction.c        |  35 ++++++\n odb/transaction.h        |  57 ++++++++++\n read-cache.c             |   1 +\n 14 files changed, 311 insertions(+), 188 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\nRange-diff against v2:\n1:  eee372b426 = 1:  eee372b426 odb: split `struct odb_transaction` into separate header\n2:  57ac075560 = 2:  57ac075560 odb/transaction: use pluggable `begin_transaction()`\n3:  556f003d0a ! 3:  11321ad607 odb: update `struct odb_write_stream` read() callback\n    @@ Commit message\n     \n         Update the interface to instead require the caller to provide a buffer,\n         and have the callback return the number of bytes written to it or a\n    -    negative value on error. Call sites are updated accordingly.\n    +    negative value on error. While at it, also move the `struct\n    +    odb_write_stream` definition to \"odb/streaming.h\". Call sites are\n    +    updated accordingly.\n     \n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n      ## builtin/unpack-objects.c ##\n    +@@\n    + #include \"hex.h\"\n    + #include \"object-file.h\"\n    + #include \"odb.h\"\n    ++#include \"odb/streaming.h\"\n    + #include \"odb/transaction.h\"\n    + #include \"object.h\"\n    + #include \"delta.h\"\n     @@ builtin/unpack-objects.c: static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n      \n      struct input_zstream_data {\n    @@ object-file.c: int odb_source_loose_write_stream(struct odb_source *source,\n     -\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n     -\t\t\tstream.next_in = (void *)in;\n     -\t\t\tin0 = (unsigned char *)in;\n    -+\t\t\tssize_t read_len = in_stream->read(in_stream, buf, sizeof(buf));\n    ++\t\t\tssize_t read_len = odb_write_stream_read(in_stream, buf,\n    ++\t\t\t\t\t\t\t\t sizeof(buf));\n     +\t\t\tif (read_len < 0) {\n     +\t\t\t\terr = -1;\n     +\t\t\t\tgoto cleanup;\n    @@ object-file.c: int odb_source_loose_write_stream(struct odb_source *source,\n     \n      ## odb.h ##\n     @@ odb.h: static inline int odb_write_object(struct object_database *odb,\n    + \treturn odb_write_object_ext(odb, buf, len, type, oid, NULL, 0);\n      }\n      \n    - struct odb_write_stream {\n    +-struct odb_write_stream {\n     -\tconst void *(*read)(struct odb_write_stream *, unsigned long *len);\n    -+\tssize_t (*read)(struct odb_write_stream *, unsigned char *, size_t len);\n    - \tvoid *data;\n    - \tint is_finished;\n    - };\n    +-\tvoid *data;\n    +-\tint is_finished;\n    +-};\n    ++struct odb_write_stream;\n    + \n    + int odb_write_object_stream(struct object_database *odb,\n    + \t\t\t    struct odb_write_stream *stream, size_t len,\n    +\n    + ## odb/streaming.c ##\n    +@@ odb/streaming.c: struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n    + \treturn st;\n    + }\n    + \n    ++ssize_t odb_write_stream_read(struct odb_write_stream *st, void *buf, size_t sz)\n    ++{\n    ++\treturn st->read(st, buf, sz);\n    ++}\n    ++\n    + int odb_stream_blob_to_fd(struct object_database *odb,\n    + \t\t\t  int fd,\n    + \t\t\t  const struct object_id *oid,\n    +\n    + ## odb/streaming.h ##\n    +@@ odb/streaming.h: int odb_read_stream_close(struct odb_read_stream *stream);\n    +  */\n    + ssize_t odb_read_stream_read(struct odb_read_stream *stream, void *buf, size_t len);\n    + \n    ++/*\n    ++ * A stream that provides an object to be written to the object database without\n    ++ * loading all of it into memory.\n    ++ */\n    ++struct odb_write_stream {\n    ++\tssize_t (*read)(struct odb_write_stream *, unsigned char *, size_t);\n    ++\tvoid *data;\n    ++\tint is_finished;\n    ++};\n    ++\n    ++/*\n    ++ * Read data from the stream into the buffer. Returns 0 when finished and the\n    ++ * number of bytes read on success. Returns a negative error code in case\n    ++ * reading from the stream fails.\n    ++ */\n    ++ssize_t odb_write_stream_read(struct odb_write_stream *stream, void *buf,\n    ++\t\t\t      size_t len);\n    ++\n    + /*\n    +  * Look up the object by its ID and write the full contents to the file\n    +  * descriptor. The object must be a blob, or the function will fail. When\n4:  a9f0e5ad8a ! 4:  72d4656eee object-file: remove flags from transaction packfile writes\n    @@ object-file.c: static void prepare_packfile_transaction(struct odb_transaction_f\n     +\tunsigned char buf[16384];\n     +\tstruct git_hash_ctx ctx;\n     +\tunsigned header_len;\n    -+\tsize_t total = 0;\n    ++\tsize_t bytes_hashed = 0;\n     +\n     +\theader_len = format_object_header((char *)buf, sizeof(buf),\n     +\t\t\t\t\t  OBJ_BLOB, size);\n    @@ object-file.c: static void prepare_packfile_transaction(struct odb_transaction_f\n     +\tgit_hash_update(&ctx, buf, header_len);\n     +\n     +\twhile (!stream->is_finished) {\n    -+\t\tssize_t read_result = stream->read(stream, buf, sizeof(buf));\n    ++\t\tssize_t read_result = odb_write_stream_read(stream, buf,\n    ++\t\t\t\t\t\t\t    sizeof(buf));\n     +\n     +\t\tif (read_result < 0)\n     +\t\t\treturn -1;\n     +\n     +\t\tgit_hash_update(&ctx, buf, read_result);\n    -+\t\ttotal += read_result;\n    ++\t\tbytes_hashed += read_result;\n     +\t}\n     +\n    -+\tif (total != size)\n    ++\tif (bytes_hashed != size)\n     +\t\treturn -1;\n     +\n     +\tgit_hash_final_oid(result_oid, &ctx);\n    @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n     -\t\t\t\t\t\t      xsize_t(st->st_size),\n     -\t\t\t\t\t\t      path, flags);\n     -\t\todb_transaction_commit(transaction);\n    -+\t\tstruct odb_write_stream stream = { 0 };\n    ++\t\tstruct odb_write_stream stream;\n     +\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n     +\n     +\t\tif (flags & INDEX_WRITE_OBJECT) {\n    @@ object-file.c: int index_fd(struct index_state *istate, struct object_id *oid,\n     +\t\t\t\t\t       xsize_t(st->st_size));\n     +\t\t}\n     +\n    -+\t\tfree(stream.data);\n    ++\t\todb_write_stream_release(&stream);\n      \t}\n      \n      \tclose(fd);\n     \n      ## odb/streaming.c ##\n    +@@ odb/streaming.c: ssize_t odb_write_stream_read(struct odb_write_stream *st, void *buf, size_t sz)\n    + \treturn st->read(st, buf, sz);\n    + }\n    + \n    ++void odb_write_stream_release(struct odb_write_stream *st)\n    ++{\n    ++\tfree(st->data);\n    ++}\n    ++\n    + int odb_stream_blob_to_fd(struct object_database *odb,\n    + \t\t\t  int fd,\n    + \t\t\t  const struct object_id *oid,\n     @@ odb/streaming.c: int odb_stream_blob_to_fd(struct object_database *odb,\n      \todb_read_stream_close(st);\n      \treturn result;\n    @@ odb/streaming.c: int odb_stream_blob_to_fd(struct object_database *odb,\n     +\n     +\tstream->data = data;\n     +\tstream->read = read_object_fd;\n    ++\tstream->is_finished = 0;\n     +}\n     \n      ## odb/streaming.h ##\n    @@ odb/streaming.h\n      \n      struct object_database;\n      struct odb_read_stream;\n    +@@ odb/streaming.h: struct odb_write_stream {\n    + ssize_t odb_write_stream_read(struct odb_write_stream *stream, void *buf,\n    + \t\t\t      size_t len);\n    + \n    ++/*\n    ++ * Releases memory allocated for underlying stream data.\n    ++ */\n    ++void odb_write_stream_release(struct odb_write_stream *stream);\n    ++\n    + /*\n    +  * Look up the object by its ID and write the full contents to the file\n    +  * descriptor. The object must be a blob, or the function will fail. When\n     @@ odb/streaming.h: int odb_stream_blob_to_fd(struct object_database *odb,\n      \t\t\t  struct stream_filter *filter,\n      \t\t\t  int can_seek);\n      \n     +/*\n    -+ * Sets up an ODB write stream that reads from an fd. The caller is expected to\n    -+ * free the underlying stream data.\n    ++ * Sets up an ODB write stream that reads from an fd.\n     + */\n     +void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n     +\t\t\t      size_t size);\n5:  b7ac82ed7e = 5:  e4896101ff object-file: avoid fd seekback by checking object size upfront\n6:  d6c4187a0f ! 6:  b3cb0a707c object-file: generalize packfile writes to use odb_write_stream\n    @@ object-file.c: static int hash_blob_stream(struct odb_write_stream *stream,\n      \tunsigned char obuf[16384];\n      \tunsigned hdrlen;\n      \tint status = Z_OK;\n    -+\tsize_t total = 0;\n    ++\tsize_t bytes_read = 0;\n      \n      \tgit_deflate_init(&s, pack_compression_level);\n      \n    @@ object-file.c: static void stream_blob_to_pack(struct transaction_packfile *stat\n     -\t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n     -\t\t\t\t    (unsigned)rsize, path);\n     +\t\tif (!stream->is_finished && !s.avail_in) {\n    -+\t\t\tssize_t rsize = stream->read(stream, ibuf, sizeof(ibuf));\n    ++\t\t\tssize_t rsize = odb_write_stream_read(stream, ibuf,\n    ++\t\t\t\t\t\t\t      sizeof(ibuf));\n     +\n     +\t\t\tif (rsize < 0)\n     +\t\t\t\tdie(\"failed to read blob data\");\n    @@ object-file.c: static void stream_blob_to_pack(struct transaction_packfile *stat\n      \t\t\ts.next_in = ibuf;\n      \t\t\ts.avail_in = rsize;\n     -\t\t\tsize -= rsize;\n    -+\t\t\ttotal += rsize;\n    ++\t\t\tbytes_read += rsize;\n      \t\t}\n      \n     -\t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n    @@ object-file.c: static void stream_blob_to_pack(struct transaction_packfile *stat\n      \t\t}\n      \t}\n     +\n    -+\tif (total != size)\n    ++\tif (bytes_read != size)\n     +\t\tdie(\"read %\" PRIuMAX \" bytes of blob data, but expected %\" PRIuMAX \" bytes\",\n    -+\t\t    (uintmax_t)total, (uintmax_t)size);\n    ++\t\t    (uintmax_t)bytes_read, (uintmax_t)size);\n     +\n      \tgit_deflate_end(&s);\n      }\n7:  2b81e94677 ! 7:  e1d292a7ed odb/transaction: make `write_object_stream()` pluggable\n    @@ Commit message\n         How an ODB transaction handles writing objects is expected to vary\n         between implementations. Introduce a new `write_object_stream()`\n         callback in `struct odb_transaction` to make this function pluggable.\n    -    Wire up `index_blob_packfile_transaction()` for use with `struct\n    -    odb_transaction_files` accordingly.\n    +    Rename `index_blob_packfile_transaction()` to\n    +    `odb_transaction_files_write_object_stream()` and wire it up for use\n    +    with `struct odb_transaction_files` accordingly.\n     \n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n\nbase-commit: 5361983c075154725be47b65cca9a2421789e410\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540799","messageId":"20260402213220.2651523-2-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v3 1/7] odb: split `struct odb_transaction` into separate header","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:14Z","receivedAt":"2026-04-02T21:32:51Z","isPatch":true,"body":"The current ODB transaction interface is colocated with other ODB\ninterfaces in \"odb.{c,h}\". Subsequent commits will expand `struct\nodb_transaction` to support write operations on the transaction\ndirectly. To keep things organized and prevent \"odb.{c,h}\" from becoming\nmore unwieldy, split out `struct odb_transaction` into a separate\nheader.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Makefile                 |  1 +\n builtin/add.c            |  1 +\n builtin/unpack-objects.c |  1 +\n builtin/update-index.c   |  1 +\n cache-tree.c             |  1 +\n meson.build              |  1 +\n object-file.c            |  1 +\n odb.c                    | 25 -------------------------\n odb.h                    | 31 -------------------------------\n odb/transaction.c        | 28 ++++++++++++++++++++++++++++\n odb/transaction.h        | 38 ++++++++++++++++++++++++++++++++++++++\n read-cache.c             |  1 +\n 12 files changed, 74 insertions(+), 56 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\ndiff --git a/Makefile b/Makefile\nindex dbf0022054..6342db13e5 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1219,6 +1219,7 @@ LIB_OBJS += odb.o\n LIB_OBJS += odb/source.o\n LIB_OBJS += odb/source-files.o\n LIB_OBJS += odb/streaming.o\n+LIB_OBJS += odb/transaction.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\n LIB_OBJS += oidset.o\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 7737ab878b..c859f66519 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -16,6 +16,7 @@\n #include \"run-command.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"parse-options.h\"\n #include \"path.h\"\n #include \"preload-index.h\"\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 6fc64e9e4b..bc9b1e047e 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -9,6 +9,7 @@\n #include \"hex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"object.h\"\n #include \"delta.h\"\n #include \"pack.h\"\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 8a5907767b..bcc43852ef 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -19,6 +19,7 @@\n #include \"tree-walk.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"refs.h\"\n #include \"resolve-undo.h\"\n #include \"parse-options.h\"\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 60bcc07c3b..f056869cfd 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -10,6 +10,7 @@\n #include \"cache-tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"read-cache-ll.h\"\n #include \"replace-object.h\"\n #include \"repository.h\"\ndiff --git a/meson.build b/meson.build\nindex 8309942d18..6dc23b3af2 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -405,6 +405,7 @@ libgit_sources = [\n   'odb/source.c',\n   'odb/source-files.c',\n   'odb/streaming.c',\n+  'odb/transaction.c',\n   'oid-array.c',\n   'oidmap.c',\n   'oidset.c',\ndiff --git a/object-file.c b/object-file.c\nindex f0b029ff0b..bfbb632cf8 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -21,6 +21,7 @@\n #include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/streaming.h\"\n+#include \"odb/transaction.h\"\n #include \"oidtree.h\"\n #include \"pack.h\"\n #include \"packfile.h\"\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..8c3cbc1b53 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -1069,28 +1069,3 @@ void odb_reprepare(struct object_database *o)\n \n \tobj_read_unlock();\n }\n-\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n-{\n-\tif (odb->transaction)\n-\t\treturn NULL;\n-\n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n-\n-\treturn odb->transaction;\n-}\n-\n-void odb_transaction_commit(struct odb_transaction *transaction)\n-{\n-\tif (!transaction)\n-\t\treturn;\n-\n-\t/*\n-\t * Ensure the transaction ending matches the pending transaction.\n-\t */\n-\tASSERT(transaction == transaction->source->odb->transaction);\n-\n-\ttransaction->commit(transaction);\n-\ttransaction->source->odb->transaction = NULL;\n-\tfree(transaction);\n-}\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..ec5367b13e 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -35,24 +35,6 @@ struct packed_git;\n struct packfile_store;\n struct cached_object_entry;\n \n-/*\n- * A transaction may be started for an object database prior to writing new\n- * objects via odb_transaction_begin(). These objects are not committed until\n- * odb_transaction_commit() is invoked. Only a single transaction may be pending\n- * at a time.\n- *\n- * Each ODB source is expected to implement its own transaction handling.\n- */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n-struct odb_transaction {\n-\t/* The ODB source the transaction is opened against. */\n-\tstruct odb_source *source;\n-\n-\t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n-};\n-\n /*\n  * The object database encapsulates access to objects in a repository. It\n  * manages one or more sources that store the actual objects which are\n@@ -154,19 +136,6 @@ void odb_close(struct object_database *o);\n  */\n void odb_reprepare(struct object_database *o);\n \n-/*\n- * Starts an ODB transaction. Subsequent objects are written to the transaction\n- * and not committed until odb_transaction_commit() is invoked on the\n- * transaction. If the ODB already has a pending transaction, NULL is returned.\n- */\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n-\n-/*\n- * Commits an ODB transaction making the written objects visible. If the\n- * specified transaction is NULL, the function is a no-op.\n- */\n-void odb_transaction_commit(struct odb_transaction *transaction);\n-\n /*\n  * Find source by its object directory path. Returns a `NULL` pointer in case\n  * the source could not be found.\ndiff --git a/odb/transaction.c b/odb/transaction.c\nnew file mode 100644\nindex 0000000000..9bf3f347dc\n--- /dev/null\n+++ b/odb/transaction.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include \"object-file.h\"\n+#include \"odb/transaction.h\"\n+\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n+{\n+\tif (odb->transaction)\n+\t\treturn NULL;\n+\n+\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\n+\treturn odb->transaction;\n+}\n+\n+void odb_transaction_commit(struct odb_transaction *transaction)\n+{\n+\tif (!transaction)\n+\t\treturn;\n+\n+\t/*\n+\t * Ensure the transaction ending matches the pending transaction.\n+\t */\n+\tASSERT(transaction == transaction->source->odb->transaction);\n+\n+\ttransaction->commit(transaction);\n+\ttransaction->source->odb->transaction = NULL;\n+\tfree(transaction);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nnew file mode 100644\nindex 0000000000..a56e392f21\n--- /dev/null\n+++ b/odb/transaction.h\n@@ -0,0 +1,38 @@\n+#ifndef ODB_TRANSACTION_H\n+#define ODB_TRANSACTION_H\n+\n+#include \"odb.h\"\n+#include \"odb/source.h\"\n+\n+/*\n+ * A transaction may be started for an object database prior to writing new\n+ * objects via odb_transaction_begin(). These objects are not committed until\n+ * odb_transaction_commit() is invoked. Only a single transaction may be pending\n+ * at a time.\n+ *\n+ * Each ODB source is expected to implement its own transaction handling.\n+ */\n+struct odb_transaction;\n+typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n+struct odb_transaction {\n+\t/* The ODB source the transaction is opened against. */\n+\tstruct odb_source *source;\n+\n+\t/* The ODB source specific callback invoked to commit a transaction. */\n+\todb_transaction_commit_fn commit;\n+};\n+\n+/*\n+ * Starts an ODB transaction. Subsequent objects are written to the transaction\n+ * and not committed until odb_transaction_commit() is invoked on the\n+ * transaction. If the ODB already has a pending transaction, NULL is returned.\n+ */\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n+\n+/*\n+ * Commits an ODB transaction making the written objects visible. If the\n+ * specified transaction is NULL, the function is a no-op.\n+ */\n+void odb_transaction_commit(struct odb_transaction *transaction);\n+\n+#endif\ndiff --git a/read-cache.c b/read-cache.c\nindex 5049f9baca..8147c7e94a 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -20,6 +20,7 @@\n #include \"dir.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"oid-array.h\"\n #include \"tree.h\"\n #include \"commit.h\"\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540798","messageId":"20260402213220.2651523-3-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v3 2/7] odb/transaction: use pluggable `begin_transaction()`","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:15Z","receivedAt":"2026-04-02T21:32:52Z","isPatch":true,"body":"Each ODB source is expected to provide an ODB transaction implementation\nthat should be used when starting a transaction. With d6fc6fe6f8\n(odb/source: make `begin_transaction()` function pluggable, 2026-03-05),\nthe `struct odb_source` now provides a pluggable callback for beginning\ntransactions. Use the callback provided by the ODB source accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n odb/transaction.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 9bf3f347dc..592ac84075 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -1,5 +1,5 @@\n #include \"git-compat-util.h\"\n-#include \"object-file.h\"\n+#include \"odb/source.h\"\n #include \"odb/transaction.h\"\n \n struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n@@ -7,7 +7,7 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n \tif (odb->transaction)\n \t\treturn NULL;\n \n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\todb_source_begin_transaction(odb->sources, &odb->transaction);\n \n \treturn odb->transaction;\n }\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540801","messageId":"20260402213220.2651523-4-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v3 3/7] odb: update `struct odb_write_stream` read() callback","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:16Z","receivedAt":"2026-04-02T21:32:53Z","isPatch":true,"body":"The `read()` callback used by `struct odb_write_stream` currently\nreturns a pointer to an internal buffer along with the number of bytes\nread. This makes buffer ownership unclear and provides no way to report\nerrors.\n\nUpdate the interface to instead require the caller to provide a buffer,\nand have the callback return the number of bytes written to it or a\nnegative value on error. While at it, also move the `struct\nodb_write_stream` definition to \"odb/streaming.h\". Call sites are\nupdated accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/unpack-objects.c | 20 ++++++++------------\n object-file.c            | 14 +++++++++++---\n odb.h                    |  6 +-----\n odb/streaming.c          |  5 +++++\n odb/streaming.h          | 18 ++++++++++++++++++\n 5 files changed, 43 insertions(+), 20 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex bc9b1e047e..64e58e79fd 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -9,6 +9,7 @@\n #include \"hex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"odb/transaction.h\"\n #include \"object.h\"\n #include \"delta.h\"\n@@ -360,24 +361,21 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \n struct input_zstream_data {\n \tgit_zstream *zstream;\n-\tunsigned char buf[8192];\n \tint status;\n };\n \n-static const void *feed_input_zstream(struct odb_write_stream *in_stream,\n-\t\t\t\t      unsigned long *readlen)\n+static ssize_t feed_input_zstream(struct odb_write_stream *in_stream,\n+\t\t\t\t  unsigned char *buf, size_t buf_len)\n {\n \tstruct input_zstream_data *data = in_stream->data;\n \tgit_zstream *zstream = data->zstream;\n \tvoid *in = fill(1);\n \n-\tif (in_stream->is_finished) {\n-\t\t*readlen = 0;\n-\t\treturn NULL;\n-\t}\n+\tif (in_stream->is_finished)\n+\t\treturn 0;\n \n-\tzstream->next_out = data->buf;\n-\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_out = buf;\n+\tzstream->avail_out = buf_len;\n \tzstream->next_in = in;\n \tzstream->avail_in = len;\n \n@@ -385,9 +383,7 @@ static const void *feed_input_zstream(struct odb_write_stream *in_stream,\n \n \tin_stream->is_finished = data->status != Z_OK;\n \tuse(len - zstream->avail_in);\n-\t*readlen = sizeof(data->buf) - zstream->avail_out;\n-\n-\treturn data->buf;\n+\treturn buf_len - zstream->avail_out;\n }\n \n static void stream_blob(unsigned long size, unsigned nr)\ndiff --git a/object-file.c b/object-file.c\nindex bfbb632cf8..0ae36314aa 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1066,6 +1066,7 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \tstruct git_hash_ctx c, compat_c;\n \tstruct strbuf tmp_file = STRBUF_INIT;\n \tstruct strbuf filename = STRBUF_INIT;\n+\tunsigned char buf[8192];\n \tint dirlen;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n@@ -1098,9 +1099,16 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t\tunsigned char *in0 = stream.next_in;\n \n \t\tif (!stream.avail_in && !in_stream->is_finished) {\n-\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n-\t\t\tstream.next_in = (void *)in;\n-\t\t\tin0 = (unsigned char *)in;\n+\t\t\tssize_t read_len = odb_write_stream_read(in_stream, buf,\n+\t\t\t\t\t\t\t\t sizeof(buf));\n+\t\t\tif (read_len < 0) {\n+\t\t\t\terr = -1;\n+\t\t\t\tgoto cleanup;\n+\t\t\t}\n+\n+\t\t\tstream.avail_in = read_len;\n+\t\t\tstream.next_in = buf;\n+\t\t\tin0 = buf;\n \t\t\t/* All data has been read. */\n \t\t\tif (in_stream->is_finished)\n \t\t\t\tflush = 1;\ndiff --git a/odb.h b/odb.h\nindex ec5367b13e..6faeaa0589 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -529,11 +529,7 @@ static inline int odb_write_object(struct object_database *odb,\n \treturn odb_write_object_ext(odb, buf, len, type, oid, NULL, 0);\n }\n \n-struct odb_write_stream {\n-\tconst void *(*read)(struct odb_write_stream *, unsigned long *len);\n-\tvoid *data;\n-\tint is_finished;\n-};\n+struct odb_write_stream;\n \n int odb_write_object_stream(struct object_database *odb,\n \t\t\t    struct odb_write_stream *stream, size_t len,\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex 5927a12954..a68dd2cbe3 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -232,6 +232,11 @@ struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n \treturn st;\n }\n \n+ssize_t odb_write_stream_read(struct odb_write_stream *st, void *buf, size_t sz)\n+{\n+\treturn st->read(st, buf, sz);\n+}\n+\n int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  int fd,\n \t\t\t  const struct object_id *oid,\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex c7861f7e13..65ced911fe 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -47,6 +47,24 @@ int odb_read_stream_close(struct odb_read_stream *stream);\n  */\n ssize_t odb_read_stream_read(struct odb_read_stream *stream, void *buf, size_t len);\n \n+/*\n+ * A stream that provides an object to be written to the object database without\n+ * loading all of it into memory.\n+ */\n+struct odb_write_stream {\n+\tssize_t (*read)(struct odb_write_stream *, unsigned char *, size_t);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n+/*\n+ * Read data from the stream into the buffer. Returns 0 when finished and the\n+ * number of bytes read on success. Returns a negative error code in case\n+ * reading from the stream fails.\n+ */\n+ssize_t odb_write_stream_read(struct odb_write_stream *stream, void *buf,\n+\t\t\t      size_t len);\n+\n /*\n  * Look up the object by its ID and write the full contents to the file\n  * descriptor. The object must be a blob, or the function will fail. When\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540802","messageId":"20260402213220.2651523-5-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v3 4/7] object-file: remove flags from transaction packfile writes","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:17Z","receivedAt":"2026-04-02T21:32:54Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function handles streaming a\nblob from an fd to compute its object ID and conditionally writes the\nobject directly to a packfile if the INDEX_WRITE_OBJECT flag is set. A\nsubsequent commit will make these packfile object writes part of the\ntransaction interface. Consequently, having the object write be\nconditional on this flag is a bit awkward.\n\nIn preparation for this change, introduce a dedicated\n`hash_blob_stream()` helper that only computes the OID from a `struct\nodb_write_stream`. This is invoked by `index_fd()` instead when the\nINDEX_WRITE_OBJECT is not set. The object write performed via\n`index_blob_packfile_transaction()` is made unconditional accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c   | 132 +++++++++++++++++++++++++++++-------------------\n odb/streaming.c |  46 +++++++++++++++++\n odb/streaming.h |  12 +++++\n 3 files changed, 138 insertions(+), 52 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 0ae36314aa..382d14c8c0 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1396,11 +1396,10 @@ static int already_written(struct odb_transaction_files *transaction,\n }\n \n /* Lazily create backing packfile for the state */\n-static void prepare_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t unsigned flags)\n+static void prepare_packfile_transaction(struct odb_transaction_files *transaction)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\tif (!(flags & INDEX_WRITE_OBJECT) || state->f)\n+\tif (state->f)\n \t\treturn;\n \n \tstate->f = create_tmp_packfile(transaction->base.source->odb->repo,\n@@ -1413,6 +1412,39 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n+static int hash_blob_stream(struct odb_write_stream *stream,\n+\t\t\t    const struct git_hash_algo *hash_algo,\n+\t\t\t    struct object_id *result_oid, size_t size)\n+{\n+\tunsigned char buf[16384];\n+\tstruct git_hash_ctx ctx;\n+\tunsigned header_len;\n+\tsize_t bytes_hashed = 0;\n+\n+\theader_len = format_object_header((char *)buf, sizeof(buf),\n+\t\t\t\t\t  OBJ_BLOB, size);\n+\thash_algo->init_fn(&ctx);\n+\tgit_hash_update(&ctx, buf, header_len);\n+\n+\twhile (!stream->is_finished) {\n+\t\tssize_t read_result = odb_write_stream_read(stream, buf,\n+\t\t\t\t\t\t\t    sizeof(buf));\n+\n+\t\tif (read_result < 0)\n+\t\t\treturn -1;\n+\n+\t\tgit_hash_update(&ctx, buf, read_result);\n+\t\tbytes_hashed += read_result;\n+\t}\n+\n+\tif (bytes_hashed != size)\n+\t\treturn -1;\n+\n+\tgit_hash_final_oid(result_oid, &ctx);\n+\n+\treturn 0;\n+}\n+\n /*\n  * Read the contents from fd for size bytes, streaming it to the\n  * packfile in state while updating the hash in ctx. Signal a failure\n@@ -1430,15 +1462,13 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n  */\n static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path,\n-\t\t\t       unsigned flags)\n+\t\t\t       int fd, size_t size, const char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\tint write_object = (flags & INDEX_WRITE_OBJECT);\n \toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n@@ -1473,20 +1503,18 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n-\t\t\tif (write_object) {\n-\t\t\t\tsize_t written = s.next_out - obuf;\n-\n-\t\t\t\t/* would we bust the size limit? */\n-\t\t\t\tif (state->nr_written &&\n-\t\t\t\t    pack_size_limit_cfg &&\n-\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\t\treturn -1;\n-\t\t\t\t}\n-\n-\t\t\t\thashwrite(state->f, obuf, written);\n-\t\t\t\tstate->offset += written;\n+\t\t\tsize_t written = s.next_out - obuf;\n+\n+\t\t\t/* would we bust the size limit? */\n+\t\t\tif (state->nr_written &&\n+\t\t\t    pack_size_limit_cfg &&\n+\t\t\t    pack_size_limit_cfg < state->offset + written) {\n+\t\t\t\tgit_deflate_abort(&s);\n+\t\t\t\treturn -1;\n \t\t\t}\n+\n+\t\t\thashwrite(state->f, obuf, written);\n+\t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n \t\t\ts.avail_out = sizeof(obuf);\n \t\t}\n@@ -1574,8 +1602,7 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  */\n static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n \t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path,\n-\t\t\t\t\t   unsigned flags)\n+\t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n \toff_t seekback, already_hashed_to;\n@@ -1583,7 +1610,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n-\tstruct pack_idx_entry *idx = NULL;\n+\tstruct pack_idx_entry *idx;\n \n \tseekback = lseek(fd, 0, SEEK_CUR);\n \tif (seekback == (off_t)-1)\n@@ -1594,33 +1621,26 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n-\t/* Note: idx is non-NULL when we are writing */\n-\tif ((flags & INDEX_WRITE_OBJECT) != 0) {\n-\t\tCALLOC_ARRAY(idx, 1);\n-\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\thashfile_checkpoint_init(state->f, &checkpoint);\n-\t}\n+\tCALLOC_ARRAY(idx, 1);\n+\tprepare_packfile_transaction(transaction);\n+\thashfile_checkpoint_init(state->f, &checkpoint);\n \n \talready_hashed_to = 0;\n \n \twhile (1) {\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\tif (idx) {\n-\t\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\t\tidx->offset = state->offset;\n-\t\t\tcrc32_begin(state->f);\n-\t\t}\n+\t\tprepare_packfile_transaction(transaction);\n+\t\thashfile_checkpoint(state->f, &checkpoint);\n+\t\tidx->offset = state->offset;\n+\t\tcrc32_begin(state->f);\n+\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path, flags))\n+\t\t\t\t\t fd, size, path))\n \t\t\tbreak;\n \t\t/*\n \t\t * Writing this object to the current pack will make\n \t\t * it too big; we need to truncate it, start a new\n \t\t * pack, and write into it.\n \t\t */\n-\t\tif (!idx)\n-\t\t\tBUG(\"should not happen\");\n \t\thashfile_truncate(state->f, &checkpoint);\n \t\tstate->offset = checkpoint.offset;\n \t\tflush_packfile_transaction(transaction);\n@@ -1628,8 +1648,6 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\treturn error(\"cannot seek back\");\n \t}\n \tgit_hash_final_oid(result_oid, &ctx);\n-\tif (!idx)\n-\t\treturn 0;\n \n \tidx->crc32 = crc32_end(state->f);\n \tif (already_written(transaction, result_oid)) {\n@@ -1667,18 +1685,28 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n \t\t\t\t type, path, flags);\n \t} else {\n-\t\tstruct object_database *odb = the_repository->objects;\n-\t\tstruct odb_transaction_files *files_transaction;\n-\t\tstruct odb_transaction *transaction;\n-\n-\t\ttransaction = odb_transaction_begin(odb);\n-\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t base);\n-\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t      path, flags);\n-\t\todb_transaction_commit(transaction);\n+\t\tstruct odb_write_stream stream;\n+\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n+\n+\t\tif (flags & INDEX_WRITE_OBJECT) {\n+\t\t\tstruct object_database *odb = the_repository->objects;\n+\t\t\tstruct odb_transaction_files *files_transaction;\n+\t\t\tstruct odb_transaction *transaction;\n+\n+\t\t\ttransaction = odb_transaction_begin(odb);\n+\t\t\tfiles_transaction = container_of(odb->transaction,\n+\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t base);\n+\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n+\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\todb_transaction_commit(transaction);\n+\t\t} else {\n+\t\t\tret = hash_blob_stream(&stream,\n+\t\t\t\t\t       the_repository->hash_algo, oid,\n+\t\t\t\t\t       xsize_t(st->st_size));\n+\t\t}\n+\n+\t\todb_write_stream_release(&stream);\n \t}\n \n \tclose(fd);\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex a68dd2cbe3..20531e864c 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -237,6 +237,11 @@ ssize_t odb_write_stream_read(struct odb_write_stream *st, void *buf, size_t sz)\n \treturn st->read(st, buf, sz);\n }\n \n+void odb_write_stream_release(struct odb_write_stream *st)\n+{\n+\tfree(st->data);\n+}\n+\n int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  int fd,\n \t\t\t  const struct object_id *oid,\n@@ -292,3 +297,44 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \todb_read_stream_close(st);\n \treturn result;\n }\n+\n+struct read_object_fd_data {\n+\tint fd;\n+\tsize_t remaining;\n+};\n+\n+static ssize_t read_object_fd(struct odb_write_stream *stream,\n+\t\t\t      unsigned char *buf, size_t len)\n+{\n+\tstruct read_object_fd_data *data = stream->data;\n+\tssize_t read_result;\n+\tsize_t count;\n+\n+\tif (stream->is_finished)\n+\t\treturn 0;\n+\n+\tcount = data->remaining < len ? data->remaining : len;\n+\tread_result = read_in_full(data->fd, buf, count);\n+\tif (read_result < 0 || (size_t)read_result != count)\n+\t\treturn -1;\n+\n+\tdata->remaining -= count;\n+\tif (!data->remaining)\n+\t\tstream->is_finished = 1;\n+\n+\treturn read_result;\n+}\n+\n+void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n+\t\t\t      size_t size)\n+{\n+\tstruct read_object_fd_data *data;\n+\n+\tCALLOC_ARRAY(data, 1);\n+\tdata->fd = fd;\n+\tdata->remaining = size;\n+\n+\tstream->data = data;\n+\tstream->read = read_object_fd;\n+\tstream->is_finished = 0;\n+}\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex 65ced911fe..2a8cac19a4 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -5,6 +5,7 @@\n #define STREAMING_H 1\n \n #include \"object.h\"\n+#include \"odb.h\"\n \n struct object_database;\n struct odb_read_stream;\n@@ -65,6 +66,11 @@ struct odb_write_stream {\n ssize_t odb_write_stream_read(struct odb_write_stream *stream, void *buf,\n \t\t\t      size_t len);\n \n+/*\n+ * Releases memory allocated for underlying stream data.\n+ */\n+void odb_write_stream_release(struct odb_write_stream *stream);\n+\n /*\n  * Look up the object by its ID and write the full contents to the file\n  * descriptor. The object must be a blob, or the function will fail. When\n@@ -82,4 +88,10 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  struct stream_filter *filter,\n \t\t\t  int can_seek);\n \n+/*\n+ * Sets up an ODB write stream that reads from an fd.\n+ */\n+void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n+\t\t\t      size_t size);\n+\n #endif /* STREAMING_H */\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540803","messageId":"20260402213220.2651523-6-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v3 5/7] object-file: avoid fd seekback by checking object size upfront","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:18Z","receivedAt":"2026-04-02T21:32:55Z","isPatch":true,"body":"In certain scenarios, Git handles writing blobs that exceed\n\"core.bigFileThreshold\" differently by streaming the object directly\ninto a packfile. When there is an active ODB transaction, these blobs\nare streamed to the same packfile instead of using a separate packfile\nfor each. If \"pack.packSizeLimit\" is configured and streaming another\nobject causes the packfile to exceed the configured limit, the packfile\nis truncated back to the previous object and the object write is\nrestarted in a new packfile.\n\nThis works fine, but requires the fd being read from to save a\ncheckpoint so it becomes possible to rewind the input source via seeking\nback to a known offset at the beginning. In a subsequent commit, blob\nstreaming is converted to use `struct odb_write_stream` as a more\ngeneric input source instead of an fd which doesn't provide a mechanism\nfor rewinding.\n\nFor this use case though, rewinding the fd is not strictly necessary\nbecause the inflated size of the object is known and can be used to\napproximate whether writing the object would cause the packfile to\nexceed the configured limit prior to writing anything. These blobs\nwritten to the packfile are never deltified thus the size difference\nbetween what is written versus the inflated size is due to zlib\ncompression. While this does prevent packfiles from being filled to the\npotential maximum is some cases, it should be good enough and still\nprevents the packfile from exceeding any configured limit.\n\nUse the inflated blob size to determine whether writing an object to a\npackfile will exceed the configured \"pack.packSizeLimit\".\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 86 +++++++++++++++------------------------------------\n 1 file changed, 25 insertions(+), 61 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 382d14c8c0..0284d5434b 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1447,29 +1447,17 @@ static int hash_blob_stream(struct odb_write_stream *stream,\n \n /*\n  * Read the contents from fd for size bytes, streaming it to the\n- * packfile in state while updating the hash in ctx. Signal a failure\n- * by returning a negative value when the resulting pack would exceed\n- * the pack size limit and this is not the first object in the pack,\n- * so that the caller can discard what we wrote from the current pack\n- * by truncating it and opening a new one. The caller will then call\n- * us again after rewinding the input fd.\n- *\n- * The already_hashed_to pointer is kept untouched by the caller to\n- * make sure we do not hash the same byte when we are called\n- * again. This way, the caller does not have to checkpoint its hash\n- * status before calling us just in case we ask it to call us again\n- * with a new pack.\n+ * packfile in state while updating the hash in ctx.\n  */\n-static int stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path)\n+static void stream_blob_to_pack(struct transaction_packfile *state,\n+\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n+\t\t\t\tconst char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1486,15 +1474,9 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tif ((size_t)read_result != rsize)\n \t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n \t\t\t\t    (unsigned)rsize, path);\n-\t\t\toffset += rsize;\n-\t\t\tif (*already_hashed_to < offset) {\n-\t\t\t\tsize_t hsize = offset - *already_hashed_to;\n-\t\t\t\tif (rsize < hsize)\n-\t\t\t\t\thsize = rsize;\n-\t\t\t\tif (hsize)\n-\t\t\t\t\tgit_hash_update(ctx, ibuf, hsize);\n-\t\t\t\t*already_hashed_to = offset;\n-\t\t\t}\n+\n+\t\t\tgit_hash_update(ctx, ibuf, rsize);\n+\n \t\t\ts.next_in = ibuf;\n \t\t\ts.avail_in = rsize;\n \t\t\tsize -= rsize;\n@@ -1505,14 +1487,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n \n-\t\t\t/* would we bust the size limit? */\n-\t\t\tif (state->nr_written &&\n-\t\t\t    pack_size_limit_cfg &&\n-\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\treturn -1;\n-\t\t\t}\n-\n \t\t\thashwrite(state->f, obuf, written);\n \t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n@@ -1529,7 +1503,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t}\n \t}\n \tgit_deflate_end(&s);\n-\treturn 0;\n }\n \n static void flush_packfile_transaction(struct odb_transaction_files *transaction)\n@@ -1605,48 +1578,39 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\toff_t seekback, already_hashed_to;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n \tstruct pack_idx_entry *idx;\n \n-\tseekback = lseek(fd, 0, SEEK_CUR);\n-\tif (seekback == (off_t)-1)\n-\t\treturn error(\"cannot find the current offset\");\n-\n \theader_len = format_object_header((char *)obuf, sizeof(obuf),\n \t\t\t\t\t  OBJ_BLOB, size);\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n+\t/*\n+\t * If writing another object to the packfile could result in it\n+\t * exceeding the configured size limit, flush the current packfile\n+\t * transaction.\n+\t *\n+\t * Note that this uses the inflated object size as an approximation.\n+\t * Blob objects written in this manner are not delta-compressed, so\n+\t * the difference between the inflated and on-disk size is limited\n+\t * to zlib compression and is sufficient for this check.\n+\t */\n+\tif (state->nr_written && pack_size_limit_cfg &&\n+\t    pack_size_limit_cfg < state->offset + size)\n+\t\tflush_packfile_transaction(transaction);\n+\n \tCALLOC_ARRAY(idx, 1);\n \tprepare_packfile_transaction(transaction);\n \thashfile_checkpoint_init(state->f, &checkpoint);\n \n-\talready_hashed_to = 0;\n-\n-\twhile (1) {\n-\t\tprepare_packfile_transaction(transaction);\n-\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\tidx->offset = state->offset;\n-\t\tcrc32_begin(state->f);\n-\n-\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path))\n-\t\t\tbreak;\n-\t\t/*\n-\t\t * Writing this object to the current pack will make\n-\t\t * it too big; we need to truncate it, start a new\n-\t\t * pack, and write into it.\n-\t\t */\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tflush_packfile_transaction(transaction);\n-\t\tif (lseek(fd, seekback, SEEK_SET) == (off_t)-1)\n-\t\t\treturn error(\"cannot seek back\");\n-\t}\n+\thashfile_checkpoint(state->f, &checkpoint);\n+\tidx->offset = state->offset;\n+\tcrc32_begin(state->f);\n+\tstream_blob_to_pack(state, &ctx, fd, size, path);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540804","messageId":"20260402213220.2651523-7-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v3 6/7] object-file: generalize packfile writes to use odb_write_stream","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:19Z","receivedAt":"2026-04-02T21:32:55Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function streams blob data\ndirectly from an fd. This makes it difficult to reuse as part of a\ngeneric transactional object writing interface.\n\nRefactor the packfile write path to operate on a `struct\nodb_write_stream`, allowing callers to supply data from arbitrary\nsources.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 56 +++++++++++++++++++++++++++------------------------\n 1 file changed, 30 insertions(+), 26 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 0284d5434b..7fa2b9239f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1446,18 +1446,19 @@ static int hash_blob_stream(struct odb_write_stream *stream,\n }\n \n /*\n- * Read the contents from fd for size bytes, streaming it to the\n+ * Read the contents from the stream provided, streaming it to the\n  * packfile in state while updating the hash in ctx.\n  */\n static void stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n-\t\t\t\tconst char *path)\n+\t\t\t\tstruct git_hash_ctx *ctx, size_t size,\n+\t\t\t\tstruct odb_write_stream *stream)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n+\tsize_t bytes_read = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1466,23 +1467,21 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \ts.avail_out = sizeof(obuf) - hdrlen;\n \n \twhile (status != Z_STREAM_END) {\n-\t\tif (size && !s.avail_in) {\n-\t\t\tsize_t rsize = size < sizeof(ibuf) ? size : sizeof(ibuf);\n-\t\t\tssize_t read_result = read_in_full(fd, ibuf, rsize);\n-\t\t\tif (read_result < 0)\n-\t\t\t\tdie_errno(\"failed to read from '%s'\", path);\n-\t\t\tif ((size_t)read_result != rsize)\n-\t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n-\t\t\t\t    (unsigned)rsize, path);\n+\t\tif (!stream->is_finished && !s.avail_in) {\n+\t\t\tssize_t rsize = odb_write_stream_read(stream, ibuf,\n+\t\t\t\t\t\t\t      sizeof(ibuf));\n+\n+\t\t\tif (rsize < 0)\n+\t\t\t\tdie(\"failed to read blob data\");\n \n \t\t\tgit_hash_update(ctx, ibuf, rsize);\n \n \t\t\ts.next_in = ibuf;\n \t\t\ts.avail_in = rsize;\n-\t\t\tsize -= rsize;\n+\t\t\tbytes_read += rsize;\n \t\t}\n \n-\t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n+\t\tstatus = git_deflate(&s, stream->is_finished ? Z_FINISH : 0);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n@@ -1502,6 +1501,11 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tdie(\"unexpected deflate failure: %d\", status);\n \t\t}\n \t}\n+\n+\tif (bytes_read != size)\n+\t\tdie(\"read %\" PRIuMAX \" bytes of blob data, but expected %\" PRIuMAX \" bytes\",\n+\t\t    (uintmax_t)bytes_read, (uintmax_t)size);\n+\n \tgit_deflate_end(&s);\n }\n \n@@ -1573,10 +1577,13 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path)\n+static int index_blob_packfile_transaction(struct odb_transaction *base,\n+\t\t\t\t\t   struct odb_write_stream *stream,\n+\t\t\t\t\t   size_t size, struct object_id *result_oid)\n {\n+\tstruct odb_transaction_files *transaction = container_of(base,\n+\t\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t\t base);\n \tstruct transaction_packfile *state = &transaction->packfile;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n@@ -1610,7 +1617,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \thashfile_checkpoint(state->f, &checkpoint);\n \tidx->offset = state->offset;\n \tcrc32_begin(state->f);\n-\tstream_blob_to_pack(state, &ctx, fd, size, path);\n+\tstream_blob_to_pack(state, &ctx, size, stream);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n@@ -1654,15 +1661,12 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \n \t\tif (flags & INDEX_WRITE_OBJECT) {\n \t\t\tstruct object_database *odb = the_repository->objects;\n-\t\t\tstruct odb_transaction_files *files_transaction;\n-\t\t\tstruct odb_transaction *transaction;\n-\n-\t\t\ttransaction = odb_transaction_begin(odb);\n-\t\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t\t base);\n-\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n+\n+\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n+\t\t\t\t\t\t\t      &stream,\n+\t\t\t\t\t\t\t      xsize_t(st->st_size),\n+\t\t\t\t\t\t\t      oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tret = hash_blob_stream(&stream,\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"540805","messageId":"20260402213220.2651523-8-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v3 7/7] odb/transaction: make `write_object_stream()` pluggable","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-02T21:32:20Z","receivedAt":"2026-04-02T21:32:56Z","isPatch":true,"body":"How an ODB transaction handles writing objects is expected to vary\nbetween implementations. Introduce a new `write_object_stream()`\ncallback in `struct odb_transaction` to make this function pluggable.\nRename `index_blob_packfile_transaction()` to\n`odb_transaction_files_write_object_stream()` and wire it up for use\nwith `struct odb_transaction_files` accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c     | 16 +++++++++-------\n odb/transaction.c |  7 +++++++\n odb/transaction.h | 25 ++++++++++++++++++++++---\n 3 files changed, 38 insertions(+), 10 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 7fa2b9239f..65356998f3 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1577,9 +1577,10 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_blob_packfile_transaction(struct odb_transaction *base,\n-\t\t\t\t\t   struct odb_write_stream *stream,\n-\t\t\t\t\t   size_t size, struct object_id *result_oid)\n+static int odb_transaction_files_write_object_stream(struct odb_transaction *base,\n+\t\t\t\t\t\t     struct odb_write_stream *stream,\n+\t\t\t\t\t\t     size_t size,\n+\t\t\t\t\t\t     struct object_id *result_oid)\n {\n \tstruct odb_transaction_files *transaction = container_of(base,\n \t\t\t\t\t\t\t\t struct odb_transaction_files,\n@@ -1663,10 +1664,10 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\t\tstruct object_database *odb = the_repository->objects;\n \t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n \n-\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n-\t\t\t\t\t\t\t      &stream,\n-\t\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t\t      oid);\n+\t\t\tret = odb_transaction_write_object_stream(odb->transaction,\n+\t\t\t\t\t\t\t\t  &stream,\n+\t\t\t\t\t\t\t\t  xsize_t(st->st_size),\n+\t\t\t\t\t\t\t\t  oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tret = hash_blob_stream(&stream,\n@@ -2131,6 +2132,7 @@ struct odb_transaction *odb_transaction_files_begin(struct odb_source *source)\n \ttransaction = xcalloc(1, sizeof(*transaction));\n \ttransaction->base.source = source;\n \ttransaction->base.commit = odb_transaction_files_commit;\n+\ttransaction->base.write_object_stream = odb_transaction_files_write_object_stream;\n \n \treturn &transaction->base;\n }\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 592ac84075..b16e07aebf 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -26,3 +26,10 @@ void odb_transaction_commit(struct odb_transaction *transaction)\n \ttransaction->source->odb->transaction = NULL;\n \tfree(transaction);\n }\n+\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid)\n+{\n+\treturn transaction->write_object_stream(transaction, stream, len, oid);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nindex a56e392f21..854fda06f5 100644\n--- a/odb/transaction.h\n+++ b/odb/transaction.h\n@@ -12,14 +12,24 @@\n  *\n  * Each ODB source is expected to implement its own transaction handling.\n  */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n struct odb_transaction {\n \t/* The ODB source the transaction is opened against. */\n \tstruct odb_source *source;\n \n \t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n+\tvoid (*commit)(struct odb_transaction *transaction);\n+\n+\t/*\n+\t * This callback is expected to write the given object stream into\n+\t * the ODB transaction. Note that for now, only blobs support streaming.\n+\t *\n+\t * The resulting object ID shall be written into the out pointer. The\n+\t * callback is expected to return 0 on success, a negative error code\n+\t * otherwise.\n+\t */\n+\tint (*write_object_stream)(struct odb_transaction *transaction,\n+\t\t\t\t   struct odb_write_stream *stream, size_t len,\n+\t\t\t\t   struct object_id *oid);\n };\n \n /*\n@@ -35,4 +45,13 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n  */\n void odb_transaction_commit(struct odb_transaction *transaction);\n \n+/*\n+ * Writes the object in the provided stream into the transaction. The resulting\n+ * object ID is written into the out pointer. Returns 0 on success, a negative\n+ * error code otherwise.\n+ */\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid);\n+\n #endif\n-- \n2.53.0.381.g628a66ccf6\n\n"},{"id":"541017","messageId":"20260406201627.GA26312@coredump.intra.peff.net","threadId":"65391","inReplyTo":"20260402213220.2651523-5-jltobler@gmail.com","subject":"Re: [PATCH v3 4/7] object-file: remove flags from transaction packfile writes","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-04-06T20:16:27Z","receivedAt":"2026-04-06T20:16:29Z","isPatch":true,"body":"On Thu, Apr 02, 2026 at 04:32:17PM -0500, Justin Tobler wrote:\n\n> @@ -1667,18 +1685,28 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n> [...]\n> +\t\tstruct odb_write_stream stream;\n> +\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n> +\n> +\t\tif (flags & INDEX_WRITE_OBJECT) {\n> +\t\t\tstruct object_database *odb = the_repository->objects;\n> +\t\t\tstruct odb_transaction_files *files_transaction;\n> +\t\t\tstruct odb_transaction *transaction;\n> +\n> +\t\t\ttransaction = odb_transaction_begin(odb);\n> +\t\t\tfiles_transaction = container_of(odb->transaction,\n> +\t\t\t\t\t\t\t struct odb_transaction_files,\n> +\t\t\t\t\t\t\t base);\n> +\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> +\t\t\t\t\t\t      xsize_t(st->st_size), path);\n> +\t\t\todb_transaction_commit(transaction);\n> +\t\t} else {\n> +\t\t\tret = hash_blob_stream(&stream,\n> +\t\t\t\t\t       the_repository->hash_algo, oid,\n> +\t\t\t\t\t       xsize_t(st->st_size));\n> +\t\t}\n> +\n> +\t\todb_write_stream_release(&stream);\n\nProbably not a big deal, but I notice that \"stream\" is not used in half\nof the conditional. Should its initialization and cleanup be pushed down\ninto the else clause?\n\n-Peff\n"},{"id":"541018","messageId":"20260406201904.GA26632@coredump.intra.peff.net","threadId":"65391","inReplyTo":"20260406201627.GA26312@coredump.intra.peff.net","subject":"Re: [PATCH v3 4/7] object-file: remove flags from transaction packfile writes","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-04-06T20:19:04Z","receivedAt":"2026-04-06T20:19:06Z","isPatch":true,"body":"On Mon, Apr 06, 2026 at 04:16:28PM -0400, Jeff King wrote:\n\n> On Thu, Apr 02, 2026 at 04:32:17PM -0500, Justin Tobler wrote:\n> \n> > @@ -1667,18 +1685,28 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n> > [...]\n> > +\t\tstruct odb_write_stream stream;\n> > +\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n> > +\n> > +\t\tif (flags & INDEX_WRITE_OBJECT) {\n> > +\t\t\tstruct object_database *odb = the_repository->objects;\n> > +\t\t\tstruct odb_transaction_files *files_transaction;\n> > +\t\t\tstruct odb_transaction *transaction;\n> > +\n> > +\t\t\ttransaction = odb_transaction_begin(odb);\n> > +\t\t\tfiles_transaction = container_of(odb->transaction,\n> > +\t\t\t\t\t\t\t struct odb_transaction_files,\n> > +\t\t\t\t\t\t\t base);\n> > +\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n> > +\t\t\t\t\t\t      xsize_t(st->st_size), path);\n> > +\t\t\todb_transaction_commit(transaction);\n> > +\t\t} else {\n> > +\t\t\tret = hash_blob_stream(&stream,\n> > +\t\t\t\t\t       the_repository->hash_algo, oid,\n> > +\t\t\t\t\t       xsize_t(st->st_size));\n> > +\t\t}\n> > +\n> > +\t\todb_write_stream_release(&stream);\n> \n> Probably not a big deal, but I notice that \"stream\" is not used in half\n> of the conditional. Should its initialization and cleanup be pushed down\n> into the else clause?\n\nAh, never mind. I was looking at this patch in isolation as a solution\nto the segfault problem. But later in the series, you end up converting\nindex_blob_packfile_transaction() to use stream, too, which seems like a\ngood direction.\n\nSo it is a little funny at this step, but it reduces the diff later.\n\n-Peff\n"},{"id":"541119","messageId":"adYC9Z1sryoepwSl@pks.im","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"Re: [PATCH v3 0/7] odb: add write operation to ODB transaction interface","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-08T07:25:41Z","receivedAt":"2026-04-08T07:25:55Z","isPatch":true,"body":"On Thu, Apr 02, 2026 at 04:32:13PM -0500, Justin Tobler wrote:\n> Greetings,\n> \n> This series lays the groundwork for introducing write operations to the\n> ODB transaction interface. The eventual goal is for all object writes\n> performed within a transaction to go through this interface explicitly,\n> rather than implicitly relying on the transaction to reconfigure ODB\n> sources so that writes are redirected to a temporary location.\n> \n> For now, only `odb_transaction_write_object_stream()` is implemented and\n> wires up the existing logic for streaming \"large\" blobs directly into a\n> packfile as part of the transaction.\n> \n> Most of the patches are structural refactorings to enable this, but\n> patch 4 introduces a behavioral change in how packfiles that would\n> exceed \"pack.packSizeLimit\" are handled.\n> \n> Changes since V2:\n> - Renamed some variables to improve clarity\n> - Make `odb_write_stream_from_fd()` fully initialize the underlying\n>   `struct odb_write_stream`\n> - Move `struct odb_write_stream` to \"odb/streaming.h\"\n> - Make the `hash_blob_stream()` helper more generic by operating on a\n>   `struct odb_write_stream` instead of reading from an fd directly.\n> - Introduce an `odb_write_stream_release()` helper to free the\n>   underlying stream data.\n\nAll these changes here look sensible to me, and the range-diff does,\ntoo. I'm happy with the state of this series, thanks!\n\nPatrick\n"},{"id":"543068","messageId":"20260511175835.GA4811@coredump.intra.peff.net","threadId":"65391","inReplyTo":"20260402213220.2651523-4-jltobler@gmail.com","subject":"Re: [PATCH v3 3/7] odb: update `struct odb_write_stream` read() callback","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-05-11T17:58:35Z","receivedAt":"2026-05-11T17:58:42Z","isPatch":true,"body":"On Thu, Apr 02, 2026 at 04:32:16PM -0500, Justin Tobler wrote:\n\n> @@ -1098,9 +1099,16 @@ int odb_source_loose_write_stream(struct odb_source *source,\n>  \t\tunsigned char *in0 = stream.next_in;\n>  \n>  \t\tif (!stream.avail_in && !in_stream->is_finished) {\n> -\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n> -\t\t\tstream.next_in = (void *)in;\n> -\t\t\tin0 = (unsigned char *)in;\n> +\t\t\tssize_t read_len = odb_write_stream_read(in_stream, buf,\n> +\t\t\t\t\t\t\t\t sizeof(buf));\n> +\t\t\tif (read_len < 0) {\n> +\t\t\t\terr = -1;\n> +\t\t\t\tgoto cleanup;\n> +\t\t\t}\n> +\n> +\t\t\tstream.avail_in = read_len;\n> +\t\t\tstream.next_in = buf;\n> +\t\t\tin0 = buf;\n\nIf we hit this \"goto cleanup\", we'll leak the \"fd\" descriptor opened\nearlier. We either need to close(fd) here, or do so in the cleanup\nhandler (but that means consistently setting fd to a sentinel value\nafter we close it, which we do not currently do).\n\nNoticed by Coverity (I guess this series just hit \"jch\", since it's\n\"new\" as of today's run).\n\n-Peff\n"},{"id":"543189","messageId":"agNEixLnzEVJRL9_@denethor","threadId":"65391","inReplyTo":"20260511175835.GA4811@coredump.intra.peff.net","subject":"Re: [PATCH v3 3/7] odb: update `struct odb_write_stream` read() callback","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-12T15:19:13Z","receivedAt":"2026-05-12T15:19:18Z","isPatch":true,"body":"On 26/05/11 01:58PM, Jeff King wrote:\n> On Thu, Apr 02, 2026 at 04:32:16PM -0500, Justin Tobler wrote:\n> \n> > @@ -1098,9 +1099,16 @@ int odb_source_loose_write_stream(struct odb_source *source,\n> >  \t\tunsigned char *in0 = stream.next_in;\n> >  \n> >  \t\tif (!stream.avail_in && !in_stream->is_finished) {\n> > -\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n> > -\t\t\tstream.next_in = (void *)in;\n> > -\t\t\tin0 = (unsigned char *)in;\n> > +\t\t\tssize_t read_len = odb_write_stream_read(in_stream, buf,\n> > +\t\t\t\t\t\t\t\t sizeof(buf));\n> > +\t\t\tif (read_len < 0) {\n> > +\t\t\t\terr = -1;\n> > +\t\t\t\tgoto cleanup;\n> > +\t\t\t}\n> > +\n> > +\t\t\tstream.avail_in = read_len;\n> > +\t\t\tstream.next_in = buf;\n> > +\t\t\tin0 = buf;\n> \n> If we hit this \"goto cleanup\", we'll leak the \"fd\" descriptor opened\n> earlier. We either need to close(fd) here, or do so in the cleanup\n> handler (but that means consistently setting fd to a sentinel value\n> after we close it, which we do not currently do).\n> \n> Noticed by Coverity (I guess this series just hit \"jch\", since it's\n> \"new\" as of today's run).\n\nThanks, I will send another version to fix this leak.\n\n-Justin\n"},{"id":"543344","messageId":"20260514183740.1505171-1-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260402213220.2651523-1-jltobler@gmail.com","subject":"[PATCH v4 0/7] odb: add write operation to ODB transaction interface","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:33Z","receivedAt":"2026-05-14T18:37:58Z","isPatch":true,"body":"Greetings,\n\nThis series lays the groundwork for introducing write operations to the\nODB transaction interface. The eventual goal is for all object writes\nperformed within a transaction to go through this interface explicitly,\nrather than implicitly relying on the transaction to reconfigure ODB\nsources so that writes are redirected to a temporary location.\n\nFor now, only `odb_transaction_write_object_stream()` is implemented and\nwires up the existing logic for streaming \"large\" blobs directly into a\npackfile as part of the transaction.\n\nMost of the patches are structural refactorings to enable this, but\npatch 4 introduces a behavioral change in how packfiles that would\nexceed \"pack.packSizeLimit\" are handled.\n\nChanges since V3:\n- Fixed leak due to an fd not being closed when exiting prior to\n  close_loose_object() being invoked.\n\nChanges since V2:\n- Renamed some variables to improve clarity\n- Make `odb_write_stream_from_fd()` fully initialize the underlying\n  `struct odb_write_stream`\n- Move `struct odb_write_stream` to \"odb/streaming.h\"\n- Make the `hash_blob_stream()` helper more generic by operating on a\n  `struct odb_write_stream` instead of reading from an fd directly.\n- Introduce an `odb_write_stream_release()` helper to free the\n  underlying stream data.\n\nChanges since V1:\n- Fixed some typos\n- Improved error handling\n- Removed unnecessary guard statement\n- Documented in comments why inflated object size is used to approximate\n  if object write will exceed \"pack.packSizeLimit\".\n- Updated `struct odb_write_stream` read() callback to support returning\n  errors and using caller provided buffer\n- Updated the `hash_blob_stream()` function signature to operate on a\n  `struct odb_write_stream` instead of an fd directly\n- Renamed some variables/functions for better clarity\n\nThanks,\n-Justin\n\nJustin Tobler (7):\n  odb: split `struct odb_transaction` into separate header\n  odb/transaction: use pluggable `begin_transaction()`\n  odb: update `struct odb_write_stream` read() callback\n  object-file: remove flags from transaction packfile writes\n  object-file: avoid fd seekback by checking object size upfront\n  object-file: generalize packfile writes to use odb_write_stream\n  odb/transaction: make `write_object_stream()` pluggable\n\n Makefile                 |   1 +\n builtin/add.c            |   1 +\n builtin/unpack-objects.c |  21 ++--\n builtin/update-index.c   |   1 +\n cache-tree.c             |   1 +\n meson.build              |   1 +\n object-file.c            | 238 ++++++++++++++++++++-------------------\n odb.c                    |  25 ----\n odb.h                    |  37 +-----\n odb/streaming.c          |  51 +++++++++\n odb/streaming.h          |  30 +++++\n odb/transaction.c        |  35 ++++++\n odb/transaction.h        |  57 ++++++++++\n read-cache.c             |   1 +\n 14 files changed, 312 insertions(+), 188 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\nRange-diff against v3:\n1:  eee372b426 = 1:  eee372b426 odb: split `struct odb_transaction` into separate header\n2:  57ac075560 = 2:  57ac075560 odb/transaction: use pluggable `begin_transaction()`\n3:  11321ad607 ! 3:  d53ad95712 odb: update `struct odb_write_stream` read() callback\n    @@ object-file.c: int odb_source_loose_write_stream(struct odb_source *source,\n     +\t\t\tssize_t read_len = odb_write_stream_read(in_stream, buf,\n     +\t\t\t\t\t\t\t\t sizeof(buf));\n     +\t\t\tif (read_len < 0) {\n    ++\t\t\t\tclose(fd);\n     +\t\t\t\terr = -1;\n     +\t\t\t\tgoto cleanup;\n     +\t\t\t}\n4:  72d4656eee = 4:  fa7a3ad5dc object-file: remove flags from transaction packfile writes\n5:  e4896101ff = 5:  1ca08e0590 object-file: avoid fd seekback by checking object size upfront\n6:  b3cb0a707c = 6:  a548401057 object-file: generalize packfile writes to use odb_write_stream\n7:  e1d292a7ed = 7:  4765b1024a odb/transaction: make `write_object_stream()` pluggable\n\nbase-commit: 5361983c075154725be47b65cca9a2421789e410\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543345","messageId":"20260514183740.1505171-2-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"[PATCH v4 1/7] odb: split `struct odb_transaction` into separate header","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:34Z","receivedAt":"2026-05-14T18:37:59Z","isPatch":true,"body":"The current ODB transaction interface is colocated with other ODB\ninterfaces in \"odb.{c,h}\". Subsequent commits will expand `struct\nodb_transaction` to support write operations on the transaction\ndirectly. To keep things organized and prevent \"odb.{c,h}\" from becoming\nmore unwieldy, split out `struct odb_transaction` into a separate\nheader.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Makefile                 |  1 +\n builtin/add.c            |  1 +\n builtin/unpack-objects.c |  1 +\n builtin/update-index.c   |  1 +\n cache-tree.c             |  1 +\n meson.build              |  1 +\n object-file.c            |  1 +\n odb.c                    | 25 -------------------------\n odb.h                    | 31 -------------------------------\n odb/transaction.c        | 28 ++++++++++++++++++++++++++++\n odb/transaction.h        | 38 ++++++++++++++++++++++++++++++++++++++\n read-cache.c             |  1 +\n 12 files changed, 74 insertions(+), 56 deletions(-)\n create mode 100644 odb/transaction.c\n create mode 100644 odb/transaction.h\n\ndiff --git a/Makefile b/Makefile\nindex dbf0022054..6342db13e5 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1219,6 +1219,7 @@ LIB_OBJS += odb.o\n LIB_OBJS += odb/source.o\n LIB_OBJS += odb/source-files.o\n LIB_OBJS += odb/streaming.o\n+LIB_OBJS += odb/transaction.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\n LIB_OBJS += oidset.o\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 7737ab878b..c859f66519 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -16,6 +16,7 @@\n #include \"run-command.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"parse-options.h\"\n #include \"path.h\"\n #include \"preload-index.h\"\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 6fc64e9e4b..bc9b1e047e 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -9,6 +9,7 @@\n #include \"hex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"object.h\"\n #include \"delta.h\"\n #include \"pack.h\"\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 8a5907767b..bcc43852ef 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -19,6 +19,7 @@\n #include \"tree-walk.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"refs.h\"\n #include \"resolve-undo.h\"\n #include \"parse-options.h\"\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 60bcc07c3b..f056869cfd 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -10,6 +10,7 @@\n #include \"cache-tree.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"read-cache-ll.h\"\n #include \"replace-object.h\"\n #include \"repository.h\"\ndiff --git a/meson.build b/meson.build\nindex 8309942d18..6dc23b3af2 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -405,6 +405,7 @@ libgit_sources = [\n   'odb/source.c',\n   'odb/source-files.c',\n   'odb/streaming.c',\n+  'odb/transaction.c',\n   'oid-array.c',\n   'oidmap.c',\n   'oidset.c',\ndiff --git a/object-file.c b/object-file.c\nindex f0b029ff0b..bfbb632cf8 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -21,6 +21,7 @@\n #include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/streaming.h\"\n+#include \"odb/transaction.h\"\n #include \"oidtree.h\"\n #include \"pack.h\"\n #include \"packfile.h\"\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..8c3cbc1b53 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -1069,28 +1069,3 @@ void odb_reprepare(struct object_database *o)\n \n \tobj_read_unlock();\n }\n-\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n-{\n-\tif (odb->transaction)\n-\t\treturn NULL;\n-\n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n-\n-\treturn odb->transaction;\n-}\n-\n-void odb_transaction_commit(struct odb_transaction *transaction)\n-{\n-\tif (!transaction)\n-\t\treturn;\n-\n-\t/*\n-\t * Ensure the transaction ending matches the pending transaction.\n-\t */\n-\tASSERT(transaction == transaction->source->odb->transaction);\n-\n-\ttransaction->commit(transaction);\n-\ttransaction->source->odb->transaction = NULL;\n-\tfree(transaction);\n-}\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..ec5367b13e 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -35,24 +35,6 @@ struct packed_git;\n struct packfile_store;\n struct cached_object_entry;\n \n-/*\n- * A transaction may be started for an object database prior to writing new\n- * objects via odb_transaction_begin(). These objects are not committed until\n- * odb_transaction_commit() is invoked. Only a single transaction may be pending\n- * at a time.\n- *\n- * Each ODB source is expected to implement its own transaction handling.\n- */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n-struct odb_transaction {\n-\t/* The ODB source the transaction is opened against. */\n-\tstruct odb_source *source;\n-\n-\t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n-};\n-\n /*\n  * The object database encapsulates access to objects in a repository. It\n  * manages one or more sources that store the actual objects which are\n@@ -154,19 +136,6 @@ void odb_close(struct object_database *o);\n  */\n void odb_reprepare(struct object_database *o);\n \n-/*\n- * Starts an ODB transaction. Subsequent objects are written to the transaction\n- * and not committed until odb_transaction_commit() is invoked on the\n- * transaction. If the ODB already has a pending transaction, NULL is returned.\n- */\n-struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n-\n-/*\n- * Commits an ODB transaction making the written objects visible. If the\n- * specified transaction is NULL, the function is a no-op.\n- */\n-void odb_transaction_commit(struct odb_transaction *transaction);\n-\n /*\n  * Find source by its object directory path. Returns a `NULL` pointer in case\n  * the source could not be found.\ndiff --git a/odb/transaction.c b/odb/transaction.c\nnew file mode 100644\nindex 0000000000..9bf3f347dc\n--- /dev/null\n+++ b/odb/transaction.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include \"object-file.h\"\n+#include \"odb/transaction.h\"\n+\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n+{\n+\tif (odb->transaction)\n+\t\treturn NULL;\n+\n+\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\n+\treturn odb->transaction;\n+}\n+\n+void odb_transaction_commit(struct odb_transaction *transaction)\n+{\n+\tif (!transaction)\n+\t\treturn;\n+\n+\t/*\n+\t * Ensure the transaction ending matches the pending transaction.\n+\t */\n+\tASSERT(transaction == transaction->source->odb->transaction);\n+\n+\ttransaction->commit(transaction);\n+\ttransaction->source->odb->transaction = NULL;\n+\tfree(transaction);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nnew file mode 100644\nindex 0000000000..a56e392f21\n--- /dev/null\n+++ b/odb/transaction.h\n@@ -0,0 +1,38 @@\n+#ifndef ODB_TRANSACTION_H\n+#define ODB_TRANSACTION_H\n+\n+#include \"odb.h\"\n+#include \"odb/source.h\"\n+\n+/*\n+ * A transaction may be started for an object database prior to writing new\n+ * objects via odb_transaction_begin(). These objects are not committed until\n+ * odb_transaction_commit() is invoked. Only a single transaction may be pending\n+ * at a time.\n+ *\n+ * Each ODB source is expected to implement its own transaction handling.\n+ */\n+struct odb_transaction;\n+typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n+struct odb_transaction {\n+\t/* The ODB source the transaction is opened against. */\n+\tstruct odb_source *source;\n+\n+\t/* The ODB source specific callback invoked to commit a transaction. */\n+\todb_transaction_commit_fn commit;\n+};\n+\n+/*\n+ * Starts an ODB transaction. Subsequent objects are written to the transaction\n+ * and not committed until odb_transaction_commit() is invoked on the\n+ * transaction. If the ODB already has a pending transaction, NULL is returned.\n+ */\n+struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n+\n+/*\n+ * Commits an ODB transaction making the written objects visible. If the\n+ * specified transaction is NULL, the function is a no-op.\n+ */\n+void odb_transaction_commit(struct odb_transaction *transaction);\n+\n+#endif\ndiff --git a/read-cache.c b/read-cache.c\nindex 5049f9baca..8147c7e94a 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -20,6 +20,7 @@\n #include \"dir.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/transaction.h\"\n #include \"oid-array.h\"\n #include \"tree.h\"\n #include \"commit.h\"\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543346","messageId":"20260514183740.1505171-3-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"[PATCH v4 2/7] odb/transaction: use pluggable `begin_transaction()`","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:35Z","receivedAt":"2026-05-14T18:37:59Z","isPatch":true,"body":"Each ODB source is expected to provide an ODB transaction implementation\nthat should be used when starting a transaction. With d6fc6fe6f8\n(odb/source: make `begin_transaction()` function pluggable, 2026-03-05),\nthe `struct odb_source` now provides a pluggable callback for beginning\ntransactions. Use the callback provided by the ODB source accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n odb/transaction.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 9bf3f347dc..592ac84075 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -1,5 +1,5 @@\n #include \"git-compat-util.h\"\n-#include \"object-file.h\"\n+#include \"odb/source.h\"\n #include \"odb/transaction.h\"\n \n struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n@@ -7,7 +7,7 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb)\n \tif (odb->transaction)\n \t\treturn NULL;\n \n-\todb->transaction = odb_transaction_files_begin(odb->sources);\n+\todb_source_begin_transaction(odb->sources, &odb->transaction);\n \n \treturn odb->transaction;\n }\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543347","messageId":"20260514183740.1505171-4-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"[PATCH v4 3/7] odb: update `struct odb_write_stream` read() callback","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:36Z","receivedAt":"2026-05-14T18:38:00Z","isPatch":true,"body":"The `read()` callback used by `struct odb_write_stream` currently\nreturns a pointer to an internal buffer along with the number of bytes\nread. This makes buffer ownership unclear and provides no way to report\nerrors.\n\nUpdate the interface to instead require the caller to provide a buffer,\nand have the callback return the number of bytes written to it or a\nnegative value on error. While at it, also move the `struct\nodb_write_stream` definition to \"odb/streaming.h\". Call sites are\nupdated accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/unpack-objects.c | 20 ++++++++------------\n object-file.c            | 15 ++++++++++++---\n odb.h                    |  6 +-----\n odb/streaming.c          |  5 +++++\n odb/streaming.h          | 18 ++++++++++++++++++\n 5 files changed, 44 insertions(+), 20 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex bc9b1e047e..64e58e79fd 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -9,6 +9,7 @@\n #include \"hex.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n+#include \"odb/streaming.h\"\n #include \"odb/transaction.h\"\n #include \"object.h\"\n #include \"delta.h\"\n@@ -360,24 +361,21 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \n struct input_zstream_data {\n \tgit_zstream *zstream;\n-\tunsigned char buf[8192];\n \tint status;\n };\n \n-static const void *feed_input_zstream(struct odb_write_stream *in_stream,\n-\t\t\t\t      unsigned long *readlen)\n+static ssize_t feed_input_zstream(struct odb_write_stream *in_stream,\n+\t\t\t\t  unsigned char *buf, size_t buf_len)\n {\n \tstruct input_zstream_data *data = in_stream->data;\n \tgit_zstream *zstream = data->zstream;\n \tvoid *in = fill(1);\n \n-\tif (in_stream->is_finished) {\n-\t\t*readlen = 0;\n-\t\treturn NULL;\n-\t}\n+\tif (in_stream->is_finished)\n+\t\treturn 0;\n \n-\tzstream->next_out = data->buf;\n-\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_out = buf;\n+\tzstream->avail_out = buf_len;\n \tzstream->next_in = in;\n \tzstream->avail_in = len;\n \n@@ -385,9 +383,7 @@ static const void *feed_input_zstream(struct odb_write_stream *in_stream,\n \n \tin_stream->is_finished = data->status != Z_OK;\n \tuse(len - zstream->avail_in);\n-\t*readlen = sizeof(data->buf) - zstream->avail_out;\n-\n-\treturn data->buf;\n+\treturn buf_len - zstream->avail_out;\n }\n \n static void stream_blob(unsigned long size, unsigned nr)\ndiff --git a/object-file.c b/object-file.c\nindex bfbb632cf8..a1afca23c5 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1066,6 +1066,7 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \tstruct git_hash_ctx c, compat_c;\n \tstruct strbuf tmp_file = STRBUF_INIT;\n \tstruct strbuf filename = STRBUF_INIT;\n+\tunsigned char buf[8192];\n \tint dirlen;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n@@ -1098,9 +1099,17 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t\tunsigned char *in0 = stream.next_in;\n \n \t\tif (!stream.avail_in && !in_stream->is_finished) {\n-\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n-\t\t\tstream.next_in = (void *)in;\n-\t\t\tin0 = (unsigned char *)in;\n+\t\t\tssize_t read_len = odb_write_stream_read(in_stream, buf,\n+\t\t\t\t\t\t\t\t sizeof(buf));\n+\t\t\tif (read_len < 0) {\n+\t\t\t\tclose(fd);\n+\t\t\t\terr = -1;\n+\t\t\t\tgoto cleanup;\n+\t\t\t}\n+\n+\t\t\tstream.avail_in = read_len;\n+\t\t\tstream.next_in = buf;\n+\t\t\tin0 = buf;\n \t\t\t/* All data has been read. */\n \t\t\tif (in_stream->is_finished)\n \t\t\t\tflush = 1;\ndiff --git a/odb.h b/odb.h\nindex ec5367b13e..6faeaa0589 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -529,11 +529,7 @@ static inline int odb_write_object(struct object_database *odb,\n \treturn odb_write_object_ext(odb, buf, len, type, oid, NULL, 0);\n }\n \n-struct odb_write_stream {\n-\tconst void *(*read)(struct odb_write_stream *, unsigned long *len);\n-\tvoid *data;\n-\tint is_finished;\n-};\n+struct odb_write_stream;\n \n int odb_write_object_stream(struct object_database *odb,\n \t\t\t    struct odb_write_stream *stream, size_t len,\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex 5927a12954..a68dd2cbe3 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -232,6 +232,11 @@ struct odb_read_stream *odb_read_stream_open(struct object_database *odb,\n \treturn st;\n }\n \n+ssize_t odb_write_stream_read(struct odb_write_stream *st, void *buf, size_t sz)\n+{\n+\treturn st->read(st, buf, sz);\n+}\n+\n int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  int fd,\n \t\t\t  const struct object_id *oid,\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex c7861f7e13..65ced911fe 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -47,6 +47,24 @@ int odb_read_stream_close(struct odb_read_stream *stream);\n  */\n ssize_t odb_read_stream_read(struct odb_read_stream *stream, void *buf, size_t len);\n \n+/*\n+ * A stream that provides an object to be written to the object database without\n+ * loading all of it into memory.\n+ */\n+struct odb_write_stream {\n+\tssize_t (*read)(struct odb_write_stream *, unsigned char *, size_t);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n+/*\n+ * Read data from the stream into the buffer. Returns 0 when finished and the\n+ * number of bytes read on success. Returns a negative error code in case\n+ * reading from the stream fails.\n+ */\n+ssize_t odb_write_stream_read(struct odb_write_stream *stream, void *buf,\n+\t\t\t      size_t len);\n+\n /*\n  * Look up the object by its ID and write the full contents to the file\n  * descriptor. The object must be a blob, or the function will fail. When\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543348","messageId":"20260514183740.1505171-5-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"[PATCH v4 4/7] object-file: remove flags from transaction packfile writes","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:37Z","receivedAt":"2026-05-14T18:38:01Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function handles streaming a\nblob from an fd to compute its object ID and conditionally writes the\nobject directly to a packfile if the INDEX_WRITE_OBJECT flag is set. A\nsubsequent commit will make these packfile object writes part of the\ntransaction interface. Consequently, having the object write be\nconditional on this flag is a bit awkward.\n\nIn preparation for this change, introduce a dedicated\n`hash_blob_stream()` helper that only computes the OID from a `struct\nodb_write_stream`. This is invoked by `index_fd()` instead when the\nINDEX_WRITE_OBJECT is not set. The object write performed via\n`index_blob_packfile_transaction()` is made unconditional accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c   | 132 +++++++++++++++++++++++++++++-------------------\n odb/streaming.c |  46 +++++++++++++++++\n odb/streaming.h |  12 +++++\n 3 files changed, 138 insertions(+), 52 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex a1afca23c5..a59030911f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1397,11 +1397,10 @@ static int already_written(struct odb_transaction_files *transaction,\n }\n \n /* Lazily create backing packfile for the state */\n-static void prepare_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t unsigned flags)\n+static void prepare_packfile_transaction(struct odb_transaction_files *transaction)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\tif (!(flags & INDEX_WRITE_OBJECT) || state->f)\n+\tif (state->f)\n \t\treturn;\n \n \tstate->f = create_tmp_packfile(transaction->base.source->odb->repo,\n@@ -1414,6 +1413,39 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n+static int hash_blob_stream(struct odb_write_stream *stream,\n+\t\t\t    const struct git_hash_algo *hash_algo,\n+\t\t\t    struct object_id *result_oid, size_t size)\n+{\n+\tunsigned char buf[16384];\n+\tstruct git_hash_ctx ctx;\n+\tunsigned header_len;\n+\tsize_t bytes_hashed = 0;\n+\n+\theader_len = format_object_header((char *)buf, sizeof(buf),\n+\t\t\t\t\t  OBJ_BLOB, size);\n+\thash_algo->init_fn(&ctx);\n+\tgit_hash_update(&ctx, buf, header_len);\n+\n+\twhile (!stream->is_finished) {\n+\t\tssize_t read_result = odb_write_stream_read(stream, buf,\n+\t\t\t\t\t\t\t    sizeof(buf));\n+\n+\t\tif (read_result < 0)\n+\t\t\treturn -1;\n+\n+\t\tgit_hash_update(&ctx, buf, read_result);\n+\t\tbytes_hashed += read_result;\n+\t}\n+\n+\tif (bytes_hashed != size)\n+\t\treturn -1;\n+\n+\tgit_hash_final_oid(result_oid, &ctx);\n+\n+\treturn 0;\n+}\n+\n /*\n  * Read the contents from fd for size bytes, streaming it to the\n  * packfile in state while updating the hash in ctx. Signal a failure\n@@ -1431,15 +1463,13 @@ static void prepare_packfile_transaction(struct odb_transaction_files *transacti\n  */\n static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path,\n-\t\t\t       unsigned flags)\n+\t\t\t       int fd, size_t size, const char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\tint write_object = (flags & INDEX_WRITE_OBJECT);\n \toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n@@ -1474,20 +1504,18 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n-\t\t\tif (write_object) {\n-\t\t\t\tsize_t written = s.next_out - obuf;\n-\n-\t\t\t\t/* would we bust the size limit? */\n-\t\t\t\tif (state->nr_written &&\n-\t\t\t\t    pack_size_limit_cfg &&\n-\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\t\treturn -1;\n-\t\t\t\t}\n-\n-\t\t\t\thashwrite(state->f, obuf, written);\n-\t\t\t\tstate->offset += written;\n+\t\t\tsize_t written = s.next_out - obuf;\n+\n+\t\t\t/* would we bust the size limit? */\n+\t\t\tif (state->nr_written &&\n+\t\t\t    pack_size_limit_cfg &&\n+\t\t\t    pack_size_limit_cfg < state->offset + written) {\n+\t\t\t\tgit_deflate_abort(&s);\n+\t\t\t\treturn -1;\n \t\t\t}\n+\n+\t\t\thashwrite(state->f, obuf, written);\n+\t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n \t\t\ts.avail_out = sizeof(obuf);\n \t\t}\n@@ -1575,8 +1603,7 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  */\n static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n \t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path,\n-\t\t\t\t\t   unsigned flags)\n+\t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n \toff_t seekback, already_hashed_to;\n@@ -1584,7 +1611,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n-\tstruct pack_idx_entry *idx = NULL;\n+\tstruct pack_idx_entry *idx;\n \n \tseekback = lseek(fd, 0, SEEK_CUR);\n \tif (seekback == (off_t)-1)\n@@ -1595,33 +1622,26 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n-\t/* Note: idx is non-NULL when we are writing */\n-\tif ((flags & INDEX_WRITE_OBJECT) != 0) {\n-\t\tCALLOC_ARRAY(idx, 1);\n-\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\thashfile_checkpoint_init(state->f, &checkpoint);\n-\t}\n+\tCALLOC_ARRAY(idx, 1);\n+\tprepare_packfile_transaction(transaction);\n+\thashfile_checkpoint_init(state->f, &checkpoint);\n \n \talready_hashed_to = 0;\n \n \twhile (1) {\n-\t\tprepare_packfile_transaction(transaction, flags);\n-\t\tif (idx) {\n-\t\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\t\tidx->offset = state->offset;\n-\t\t\tcrc32_begin(state->f);\n-\t\t}\n+\t\tprepare_packfile_transaction(transaction);\n+\t\thashfile_checkpoint(state->f, &checkpoint);\n+\t\tidx->offset = state->offset;\n+\t\tcrc32_begin(state->f);\n+\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path, flags))\n+\t\t\t\t\t fd, size, path))\n \t\t\tbreak;\n \t\t/*\n \t\t * Writing this object to the current pack will make\n \t\t * it too big; we need to truncate it, start a new\n \t\t * pack, and write into it.\n \t\t */\n-\t\tif (!idx)\n-\t\t\tBUG(\"should not happen\");\n \t\thashfile_truncate(state->f, &checkpoint);\n \t\tstate->offset = checkpoint.offset;\n \t\tflush_packfile_transaction(transaction);\n@@ -1629,8 +1649,6 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\treturn error(\"cannot seek back\");\n \t}\n \tgit_hash_final_oid(result_oid, &ctx);\n-\tif (!idx)\n-\t\treturn 0;\n \n \tidx->crc32 = crc32_end(state->f);\n \tif (already_written(transaction, result_oid)) {\n@@ -1668,18 +1686,28 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n \t\t\t\t type, path, flags);\n \t} else {\n-\t\tstruct object_database *odb = the_repository->objects;\n-\t\tstruct odb_transaction_files *files_transaction;\n-\t\tstruct odb_transaction *transaction;\n-\n-\t\ttransaction = odb_transaction_begin(odb);\n-\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t base);\n-\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t      path, flags);\n-\t\todb_transaction_commit(transaction);\n+\t\tstruct odb_write_stream stream;\n+\t\todb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));\n+\n+\t\tif (flags & INDEX_WRITE_OBJECT) {\n+\t\t\tstruct object_database *odb = the_repository->objects;\n+\t\t\tstruct odb_transaction_files *files_transaction;\n+\t\t\tstruct odb_transaction *transaction;\n+\n+\t\t\ttransaction = odb_transaction_begin(odb);\n+\t\t\tfiles_transaction = container_of(odb->transaction,\n+\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t base);\n+\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n+\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\todb_transaction_commit(transaction);\n+\t\t} else {\n+\t\t\tret = hash_blob_stream(&stream,\n+\t\t\t\t\t       the_repository->hash_algo, oid,\n+\t\t\t\t\t       xsize_t(st->st_size));\n+\t\t}\n+\n+\t\todb_write_stream_release(&stream);\n \t}\n \n \tclose(fd);\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex a68dd2cbe3..20531e864c 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -237,6 +237,11 @@ ssize_t odb_write_stream_read(struct odb_write_stream *st, void *buf, size_t sz)\n \treturn st->read(st, buf, sz);\n }\n \n+void odb_write_stream_release(struct odb_write_stream *st)\n+{\n+\tfree(st->data);\n+}\n+\n int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  int fd,\n \t\t\t  const struct object_id *oid,\n@@ -292,3 +297,44 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \todb_read_stream_close(st);\n \treturn result;\n }\n+\n+struct read_object_fd_data {\n+\tint fd;\n+\tsize_t remaining;\n+};\n+\n+static ssize_t read_object_fd(struct odb_write_stream *stream,\n+\t\t\t      unsigned char *buf, size_t len)\n+{\n+\tstruct read_object_fd_data *data = stream->data;\n+\tssize_t read_result;\n+\tsize_t count;\n+\n+\tif (stream->is_finished)\n+\t\treturn 0;\n+\n+\tcount = data->remaining < len ? data->remaining : len;\n+\tread_result = read_in_full(data->fd, buf, count);\n+\tif (read_result < 0 || (size_t)read_result != count)\n+\t\treturn -1;\n+\n+\tdata->remaining -= count;\n+\tif (!data->remaining)\n+\t\tstream->is_finished = 1;\n+\n+\treturn read_result;\n+}\n+\n+void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n+\t\t\t      size_t size)\n+{\n+\tstruct read_object_fd_data *data;\n+\n+\tCALLOC_ARRAY(data, 1);\n+\tdata->fd = fd;\n+\tdata->remaining = size;\n+\n+\tstream->data = data;\n+\tstream->read = read_object_fd;\n+\tstream->is_finished = 0;\n+}\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex 65ced911fe..2a8cac19a4 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -5,6 +5,7 @@\n #define STREAMING_H 1\n \n #include \"object.h\"\n+#include \"odb.h\"\n \n struct object_database;\n struct odb_read_stream;\n@@ -65,6 +66,11 @@ struct odb_write_stream {\n ssize_t odb_write_stream_read(struct odb_write_stream *stream, void *buf,\n \t\t\t      size_t len);\n \n+/*\n+ * Releases memory allocated for underlying stream data.\n+ */\n+void odb_write_stream_release(struct odb_write_stream *stream);\n+\n /*\n  * Look up the object by its ID and write the full contents to the file\n  * descriptor. The object must be a blob, or the function will fail. When\n@@ -82,4 +88,10 @@ int odb_stream_blob_to_fd(struct object_database *odb,\n \t\t\t  struct stream_filter *filter,\n \t\t\t  int can_seek);\n \n+/*\n+ * Sets up an ODB write stream that reads from an fd.\n+ */\n+void odb_write_stream_from_fd(struct odb_write_stream *stream, int fd,\n+\t\t\t      size_t size);\n+\n #endif /* STREAMING_H */\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543349","messageId":"20260514183740.1505171-6-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"[PATCH v4 5/7] object-file: avoid fd seekback by checking object size upfront","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:38Z","receivedAt":"2026-05-14T18:38:02Z","isPatch":true,"body":"In certain scenarios, Git handles writing blobs that exceed\n\"core.bigFileThreshold\" differently by streaming the object directly\ninto a packfile. When there is an active ODB transaction, these blobs\nare streamed to the same packfile instead of using a separate packfile\nfor each. If \"pack.packSizeLimit\" is configured and streaming another\nobject causes the packfile to exceed the configured limit, the packfile\nis truncated back to the previous object and the object write is\nrestarted in a new packfile.\n\nThis works fine, but requires the fd being read from to save a\ncheckpoint so it becomes possible to rewind the input source via seeking\nback to a known offset at the beginning. In a subsequent commit, blob\nstreaming is converted to use `struct odb_write_stream` as a more\ngeneric input source instead of an fd which doesn't provide a mechanism\nfor rewinding.\n\nFor this use case though, rewinding the fd is not strictly necessary\nbecause the inflated size of the object is known and can be used to\napproximate whether writing the object would cause the packfile to\nexceed the configured limit prior to writing anything. These blobs\nwritten to the packfile are never deltified thus the size difference\nbetween what is written versus the inflated size is due to zlib\ncompression. While this does prevent packfiles from being filled to the\npotential maximum is some cases, it should be good enough and still\nprevents the packfile from exceeding any configured limit.\n\nUse the inflated blob size to determine whether writing an object to a\npackfile will exceed the configured \"pack.packSizeLimit\".\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 86 +++++++++++++++------------------------------------\n 1 file changed, 25 insertions(+), 61 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex a59030911f..6d7afdb723 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1448,29 +1448,17 @@ static int hash_blob_stream(struct odb_write_stream *stream,\n \n /*\n  * Read the contents from fd for size bytes, streaming it to the\n- * packfile in state while updating the hash in ctx. Signal a failure\n- * by returning a negative value when the resulting pack would exceed\n- * the pack size limit and this is not the first object in the pack,\n- * so that the caller can discard what we wrote from the current pack\n- * by truncating it and opening a new one. The caller will then call\n- * us again after rewinding the input fd.\n- *\n- * The already_hashed_to pointer is kept untouched by the caller to\n- * make sure we do not hash the same byte when we are called\n- * again. This way, the caller does not have to checkpoint its hash\n- * status before calling us just in case we ask it to call us again\n- * with a new pack.\n+ * packfile in state while updating the hash in ctx.\n  */\n-static int stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t       struct git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path)\n+static void stream_blob_to_pack(struct transaction_packfile *state,\n+\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n+\t\t\t\tconst char *path)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n-\toff_t offset = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1487,15 +1475,9 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tif ((size_t)read_result != rsize)\n \t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n \t\t\t\t    (unsigned)rsize, path);\n-\t\t\toffset += rsize;\n-\t\t\tif (*already_hashed_to < offset) {\n-\t\t\t\tsize_t hsize = offset - *already_hashed_to;\n-\t\t\t\tif (rsize < hsize)\n-\t\t\t\t\thsize = rsize;\n-\t\t\t\tif (hsize)\n-\t\t\t\t\tgit_hash_update(ctx, ibuf, hsize);\n-\t\t\t\t*already_hashed_to = offset;\n-\t\t\t}\n+\n+\t\t\tgit_hash_update(ctx, ibuf, rsize);\n+\n \t\t\ts.next_in = ibuf;\n \t\t\ts.avail_in = rsize;\n \t\t\tsize -= rsize;\n@@ -1506,14 +1488,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n \n-\t\t\t/* would we bust the size limit? */\n-\t\t\tif (state->nr_written &&\n-\t\t\t    pack_size_limit_cfg &&\n-\t\t\t    pack_size_limit_cfg < state->offset + written) {\n-\t\t\t\tgit_deflate_abort(&s);\n-\t\t\t\treturn -1;\n-\t\t\t}\n-\n \t\t\thashwrite(state->f, obuf, written);\n \t\t\tstate->offset += written;\n \t\t\ts.next_out = obuf;\n@@ -1530,7 +1504,6 @@ static int stream_blob_to_pack(struct transaction_packfile *state,\n \t\t}\n \t}\n \tgit_deflate_end(&s);\n-\treturn 0;\n }\n \n static void flush_packfile_transaction(struct odb_transaction_files *transaction)\n@@ -1606,48 +1579,39 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \t\t\t\t\t   size_t size, const char *path)\n {\n \tstruct transaction_packfile *state = &transaction->packfile;\n-\toff_t seekback, already_hashed_to;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint;\n \tstruct pack_idx_entry *idx;\n \n-\tseekback = lseek(fd, 0, SEEK_CUR);\n-\tif (seekback == (off_t)-1)\n-\t\treturn error(\"cannot find the current offset\");\n-\n \theader_len = format_object_header((char *)obuf, sizeof(obuf),\n \t\t\t\t\t  OBJ_BLOB, size);\n \ttransaction->base.source->odb->repo->hash_algo->init_fn(&ctx);\n \tgit_hash_update(&ctx, obuf, header_len);\n \n+\t/*\n+\t * If writing another object to the packfile could result in it\n+\t * exceeding the configured size limit, flush the current packfile\n+\t * transaction.\n+\t *\n+\t * Note that this uses the inflated object size as an approximation.\n+\t * Blob objects written in this manner are not delta-compressed, so\n+\t * the difference between the inflated and on-disk size is limited\n+\t * to zlib compression and is sufficient for this check.\n+\t */\n+\tif (state->nr_written && pack_size_limit_cfg &&\n+\t    pack_size_limit_cfg < state->offset + size)\n+\t\tflush_packfile_transaction(transaction);\n+\n \tCALLOC_ARRAY(idx, 1);\n \tprepare_packfile_transaction(transaction);\n \thashfile_checkpoint_init(state->f, &checkpoint);\n \n-\talready_hashed_to = 0;\n-\n-\twhile (1) {\n-\t\tprepare_packfile_transaction(transaction);\n-\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\tidx->offset = state->offset;\n-\t\tcrc32_begin(state->f);\n-\n-\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path))\n-\t\t\tbreak;\n-\t\t/*\n-\t\t * Writing this object to the current pack will make\n-\t\t * it too big; we need to truncate it, start a new\n-\t\t * pack, and write into it.\n-\t\t */\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tflush_packfile_transaction(transaction);\n-\t\tif (lseek(fd, seekback, SEEK_SET) == (off_t)-1)\n-\t\t\treturn error(\"cannot seek back\");\n-\t}\n+\thashfile_checkpoint(state->f, &checkpoint);\n+\tidx->offset = state->offset;\n+\tcrc32_begin(state->f);\n+\tstream_blob_to_pack(state, &ctx, fd, size, path);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543350","messageId":"20260514183740.1505171-7-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"[PATCH v4 6/7] object-file: generalize packfile writes to use odb_write_stream","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:39Z","receivedAt":"2026-05-14T18:38:03Z","isPatch":true,"body":"The `index_blob_packfile_transaction()` function streams blob data\ndirectly from an fd. This makes it difficult to reuse as part of a\ngeneric transactional object writing interface.\n\nRefactor the packfile write path to operate on a `struct\nodb_write_stream`, allowing callers to supply data from arbitrary\nsources.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c | 56 +++++++++++++++++++++++++++------------------------\n 1 file changed, 30 insertions(+), 26 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 6d7afdb723..0d492e6962 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1447,18 +1447,19 @@ static int hash_blob_stream(struct odb_write_stream *stream,\n }\n \n /*\n- * Read the contents from fd for size bytes, streaming it to the\n+ * Read the contents from the stream provided, streaming it to the\n  * packfile in state while updating the hash in ctx.\n  */\n static void stream_blob_to_pack(struct transaction_packfile *state,\n-\t\t\t\tstruct git_hash_ctx *ctx, int fd, size_t size,\n-\t\t\t\tconst char *path)\n+\t\t\t\tstruct git_hash_ctx *ctx, size_t size,\n+\t\t\t\tstruct odb_write_stream *stream)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n \tunsigned char obuf[16384];\n \tunsigned hdrlen;\n \tint status = Z_OK;\n+\tsize_t bytes_read = 0;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -1467,23 +1468,21 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \ts.avail_out = sizeof(obuf) - hdrlen;\n \n \twhile (status != Z_STREAM_END) {\n-\t\tif (size && !s.avail_in) {\n-\t\t\tsize_t rsize = size < sizeof(ibuf) ? size : sizeof(ibuf);\n-\t\t\tssize_t read_result = read_in_full(fd, ibuf, rsize);\n-\t\t\tif (read_result < 0)\n-\t\t\t\tdie_errno(\"failed to read from '%s'\", path);\n-\t\t\tif ((size_t)read_result != rsize)\n-\t\t\t\tdie(\"failed to read %u bytes from '%s'\",\n-\t\t\t\t    (unsigned)rsize, path);\n+\t\tif (!stream->is_finished && !s.avail_in) {\n+\t\t\tssize_t rsize = odb_write_stream_read(stream, ibuf,\n+\t\t\t\t\t\t\t      sizeof(ibuf));\n+\n+\t\t\tif (rsize < 0)\n+\t\t\t\tdie(\"failed to read blob data\");\n \n \t\t\tgit_hash_update(ctx, ibuf, rsize);\n \n \t\t\ts.next_in = ibuf;\n \t\t\ts.avail_in = rsize;\n-\t\t\tsize -= rsize;\n+\t\t\tbytes_read += rsize;\n \t\t}\n \n-\t\tstatus = git_deflate(&s, size ? 0 : Z_FINISH);\n+\t\tstatus = git_deflate(&s, stream->is_finished ? Z_FINISH : 0);\n \n \t\tif (!s.avail_out || status == Z_STREAM_END) {\n \t\t\tsize_t written = s.next_out - obuf;\n@@ -1503,6 +1502,11 @@ static void stream_blob_to_pack(struct transaction_packfile *state,\n \t\t\tdie(\"unexpected deflate failure: %d\", status);\n \t\t}\n \t}\n+\n+\tif (bytes_read != size)\n+\t\tdie(\"read %\" PRIuMAX \" bytes of blob data, but expected %\" PRIuMAX \" bytes\",\n+\t\t    (uintmax_t)bytes_read, (uintmax_t)size);\n+\n \tgit_deflate_end(&s);\n }\n \n@@ -1574,10 +1578,13 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_blob_packfile_transaction(struct odb_transaction_files *transaction,\n-\t\t\t\t\t   struct object_id *result_oid, int fd,\n-\t\t\t\t\t   size_t size, const char *path)\n+static int index_blob_packfile_transaction(struct odb_transaction *base,\n+\t\t\t\t\t   struct odb_write_stream *stream,\n+\t\t\t\t\t   size_t size, struct object_id *result_oid)\n {\n+\tstruct odb_transaction_files *transaction = container_of(base,\n+\t\t\t\t\t\t\t\t struct odb_transaction_files,\n+\t\t\t\t\t\t\t\t base);\n \tstruct transaction_packfile *state = &transaction->packfile;\n \tstruct git_hash_ctx ctx;\n \tunsigned char obuf[16384];\n@@ -1611,7 +1618,7 @@ static int index_blob_packfile_transaction(struct odb_transaction_files *transac\n \thashfile_checkpoint(state->f, &checkpoint);\n \tidx->offset = state->offset;\n \tcrc32_begin(state->f);\n-\tstream_blob_to_pack(state, &ctx, fd, size, path);\n+\tstream_blob_to_pack(state, &ctx, size, stream);\n \tgit_hash_final_oid(result_oid, &ctx);\n \n \tidx->crc32 = crc32_end(state->f);\n@@ -1655,15 +1662,12 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \n \t\tif (flags & INDEX_WRITE_OBJECT) {\n \t\t\tstruct object_database *odb = the_repository->objects;\n-\t\t\tstruct odb_transaction_files *files_transaction;\n-\t\t\tstruct odb_transaction *transaction;\n-\n-\t\t\ttransaction = odb_transaction_begin(odb);\n-\t\t\tfiles_transaction = container_of(odb->transaction,\n-\t\t\t\t\t\t\t struct odb_transaction_files,\n-\t\t\t\t\t\t\t base);\n-\t\t\tret = index_blob_packfile_transaction(files_transaction, oid, fd,\n-\t\t\t\t\t\t      xsize_t(st->st_size), path);\n+\t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n+\n+\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n+\t\t\t\t\t\t\t      &stream,\n+\t\t\t\t\t\t\t      xsize_t(st->st_size),\n+\t\t\t\t\t\t\t      oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tret = hash_blob_stream(&stream,\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543351","messageId":"20260514183740.1505171-8-jltobler@gmail.com","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"[PATCH v4 7/7] odb/transaction: make `write_object_stream()` pluggable","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-05-14T18:37:40Z","receivedAt":"2026-05-14T18:38:04Z","isPatch":true,"body":"How an ODB transaction handles writing objects is expected to vary\nbetween implementations. Introduce a new `write_object_stream()`\ncallback in `struct odb_transaction` to make this function pluggable.\nRename `index_blob_packfile_transaction()` to\n`odb_transaction_files_write_object_stream()` and wire it up for use\nwith `struct odb_transaction_files` accordingly.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n object-file.c     | 16 +++++++++-------\n odb/transaction.c |  7 +++++++\n odb/transaction.h | 25 ++++++++++++++++++++++---\n 3 files changed, 38 insertions(+), 10 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 0d492e6962..23f665df90 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1578,9 +1578,10 @@ static void flush_packfile_transaction(struct odb_transaction_files *transaction\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_blob_packfile_transaction(struct odb_transaction *base,\n-\t\t\t\t\t   struct odb_write_stream *stream,\n-\t\t\t\t\t   size_t size, struct object_id *result_oid)\n+static int odb_transaction_files_write_object_stream(struct odb_transaction *base,\n+\t\t\t\t\t\t     struct odb_write_stream *stream,\n+\t\t\t\t\t\t     size_t size,\n+\t\t\t\t\t\t     struct object_id *result_oid)\n {\n \tstruct odb_transaction_files *transaction = container_of(base,\n \t\t\t\t\t\t\t\t struct odb_transaction_files,\n@@ -1664,10 +1665,10 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\t\tstruct object_database *odb = the_repository->objects;\n \t\t\tstruct odb_transaction *transaction = odb_transaction_begin(odb);\n \n-\t\t\tret = index_blob_packfile_transaction(odb->transaction,\n-\t\t\t\t\t\t\t      &stream,\n-\t\t\t\t\t\t\t      xsize_t(st->st_size),\n-\t\t\t\t\t\t\t      oid);\n+\t\t\tret = odb_transaction_write_object_stream(odb->transaction,\n+\t\t\t\t\t\t\t\t  &stream,\n+\t\t\t\t\t\t\t\t  xsize_t(st->st_size),\n+\t\t\t\t\t\t\t\t  oid);\n \t\t\todb_transaction_commit(transaction);\n \t\t} else {\n \t\t\tret = hash_blob_stream(&stream,\n@@ -2132,6 +2133,7 @@ struct odb_transaction *odb_transaction_files_begin(struct odb_source *source)\n \ttransaction = xcalloc(1, sizeof(*transaction));\n \ttransaction->base.source = source;\n \ttransaction->base.commit = odb_transaction_files_commit;\n+\ttransaction->base.write_object_stream = odb_transaction_files_write_object_stream;\n \n \treturn &transaction->base;\n }\ndiff --git a/odb/transaction.c b/odb/transaction.c\nindex 592ac84075..b16e07aebf 100644\n--- a/odb/transaction.c\n+++ b/odb/transaction.c\n@@ -26,3 +26,10 @@ void odb_transaction_commit(struct odb_transaction *transaction)\n \ttransaction->source->odb->transaction = NULL;\n \tfree(transaction);\n }\n+\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid)\n+{\n+\treturn transaction->write_object_stream(transaction, stream, len, oid);\n+}\ndiff --git a/odb/transaction.h b/odb/transaction.h\nindex a56e392f21..854fda06f5 100644\n--- a/odb/transaction.h\n+++ b/odb/transaction.h\n@@ -12,14 +12,24 @@\n  *\n  * Each ODB source is expected to implement its own transaction handling.\n  */\n-struct odb_transaction;\n-typedef void (*odb_transaction_commit_fn)(struct odb_transaction *transaction);\n struct odb_transaction {\n \t/* The ODB source the transaction is opened against. */\n \tstruct odb_source *source;\n \n \t/* The ODB source specific callback invoked to commit a transaction. */\n-\todb_transaction_commit_fn commit;\n+\tvoid (*commit)(struct odb_transaction *transaction);\n+\n+\t/*\n+\t * This callback is expected to write the given object stream into\n+\t * the ODB transaction. Note that for now, only blobs support streaming.\n+\t *\n+\t * The resulting object ID shall be written into the out pointer. The\n+\t * callback is expected to return 0 on success, a negative error code\n+\t * otherwise.\n+\t */\n+\tint (*write_object_stream)(struct odb_transaction *transaction,\n+\t\t\t\t   struct odb_write_stream *stream, size_t len,\n+\t\t\t\t   struct object_id *oid);\n };\n \n /*\n@@ -35,4 +45,13 @@ struct odb_transaction *odb_transaction_begin(struct object_database *odb);\n  */\n void odb_transaction_commit(struct odb_transaction *transaction);\n \n+/*\n+ * Writes the object in the provided stream into the transaction. The resulting\n+ * object ID is written into the out pointer. Returns 0 on success, a negative\n+ * error code otherwise.\n+ */\n+int odb_transaction_write_object_stream(struct odb_transaction *transaction,\n+\t\t\t\t\tstruct odb_write_stream *stream,\n+\t\t\t\t\tsize_t len, struct object_id *oid);\n+\n #endif\n-- \n2.54.0.105.g59ff4886a5\n\n"},{"id":"543371","messageId":"20260515035616.GB75627@coredump.intra.peff.net","threadId":"65391","inReplyTo":"20260514183740.1505171-1-jltobler@gmail.com","subject":"Re: [PATCH v4 0/7] odb: add write operation to ODB transaction interface","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-05-15T03:56:16Z","receivedAt":"2026-05-15T03:56:17Z","isPatch":true,"body":"On Thu, May 14, 2026 at 01:37:33PM -0500, Justin Tobler wrote:\n\n> Changes since V3:\n> - Fixed leak due to an fd not being closed when exiting prior to\n>   close_loose_object() being invoked.\n> [...]\n> 3:  11321ad607 ! 3:  d53ad95712 odb: update `struct odb_write_stream` read() callback\n>     @@ object-file.c: int odb_source_loose_write_stream(struct odb_source *source,\n>      +\t\t\tssize_t read_len = odb_write_stream_read(in_stream, buf,\n>      +\t\t\t\t\t\t\t\t sizeof(buf));\n>      +\t\t\tif (read_len < 0) {\n>     ++\t\t\t\tclose(fd);\n>      +\t\t\t\terr = -1;\n>      +\t\t\t\tgoto cleanup;\n>      +\t\t\t}\n\nThis fix looks good to me (and I think is the best way to write it,\ngiven the rest of the function).\n\nI briefly wondered whether callers might care about errno being\npreserved, but I couldn't find any indication that they do.\n\n-Peff\n"}]}