{"thread":{"id":"60318","subject":"[PATCH 0/7] merge-ort: implement support for packing objects together","startedAt":"2023-10-06T22:01:51Z","lastAt":"2024-01-29T23:58:44Z","messageCount":89,"participants":["Taylor Blau","Junio C Hamano","Eric Biederman","Elijah Newren","Jeff King","Patrick Steinhardt","SZEDER Gábor"],"isPatch":true,"patchVersion":1,"patchTotal":7},"messages":[{"id":"482763","messageId":"cover.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":null,"subject":"[PATCH 0/7] merge-ort: implement support for packing objects together","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:01:39Z","receivedAt":"2023-10-06T22:01:51Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"(Based on the tip of 'eb/limit-bulk-checkin-to-blobs'.)\n\nThis series implements support for a new merge-tree option,\n`--write-pack`, which causes any newly-written objects to be packed\ntogether instead of being stored individually as loose.\n\nThe motivating use-case behind these changes is to better support\nrepositories who invoke merge-tree frequently, generating a potentially\nlarge number of loose objects, resulting in a possible adverse effect on\nperformance.\n\nThe majority of the changes here are preparatory to refactor common\nroutines out of the bulk-checkin machinery to prepare for indexing\ndifferent types of objects whose contents can be held in-core.\n\nAlso worth noting is the relative ease this series can be adapted to\nsupport the $NEW_HASH interop work in 'eb/hash-transition-rfc'. For more\ndetails on the relatively small number of changes necessary to make that\nwork, see the log message of the second-to-last patch.\n\nThanks in advance for your review!\n\nTaylor Blau (7):\n  bulk-checkin: factor out `format_object_header_hash()`\n  bulk-checkin: factor out `prepare_checkpoint()`\n  bulk-checkin: factor out `truncate_checkpoint()`\n  bulk-checkin: factor our `finalize_checkpoint()`\n  bulk-checkin: introduce `index_blob_bulk_checkin_incore()`\n  bulk-checkin: introduce `index_tree_bulk_checkin_incore()`\n  builtin/merge-tree.c: implement support for `--write-pack`\n\n Documentation/git-merge-tree.txt |   4 +\n builtin/merge-tree.c             |   5 +\n bulk-checkin.c                   | 249 ++++++++++++++++++++++++++-----\n bulk-checkin.h                   |   8 +\n merge-ort.c                      |  43 ++++--\n merge-recursive.h                |   1 +\n t/t4301-merge-tree-write-tree.sh |  93 ++++++++++++\n 7 files changed, 355 insertions(+), 48 deletions(-)\n\n-- \n2.42.0.8.g7a7e1e881e.dirty\n"},{"id":"482764","messageId":"37f407281596dd596e49c847c35fdf163977b479.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH 1/7] bulk-checkin: factor out `format_object_header_hash()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:01:50Z","receivedAt":"2023-10-06T22:01:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Before deflating a blob into a pack, the bulk-checkin mechanism prepares\nthe pack object header by calling `format_object_header()`, and writing\ninto a scratch buffer, the contents of which eventually makes its way\ninto the pack.\n\nFuture commits will add support for deflating multiple kinds of objects\ninto a pack, and will likewise need to perform a similar operation as\nbelow.\n\nThis is a mostly straightforward extraction, with one notable exception.\nInstead of hard-coding `the_hash_algo`, pass it in to the new function\nas an argument. This isn't strictly necessary for our immediate purposes\nhere, but will prove useful in the future if/when the bulk-checkin\nmechanism grows support for the hash transition plan.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 20 ++++++++++++++------\n 1 file changed, 14 insertions(+), 6 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 223562b4e7..0aac3dfe31 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -247,6 +247,19 @@ static void prepare_to_stream(struct bulk_checkin_packfile *state,\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n+static void format_object_header_hash(const struct git_hash_algo *algop,\n+\t\t\t\t      git_hash_ctx *ctx, enum object_type type,\n+\t\t\t\t      size_t size)\n+{\n+\tunsigned char header[16384];\n+\tunsigned header_len = format_object_header((char *)header,\n+\t\t\t\t\t\t   sizeof(header),\n+\t\t\t\t\t\t   type, size);\n+\n+\talgop->init_fn(ctx);\n+\talgop->update_fn(ctx, header, header_len);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -254,8 +267,6 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n {\n \toff_t seekback, already_hashed_to;\n \tgit_hash_ctx ctx;\n-\tunsigned char obuf[16384];\n-\tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint = {0};\n \tstruct pack_idx_entry *idx = NULL;\n \n@@ -263,10 +274,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \tif (seekback == (off_t) -1)\n \t\treturn error(\"cannot find the current offset\");\n \n-\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n-\t\t\t\t\t  OBJ_BLOB, size);\n-\tthe_hash_algo->init_fn(&ctx);\n-\tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n+\tformat_object_header_hash(the_hash_algo, &ctx, OBJ_BLOB, size);\n \n \t/* Note: idx is non-NULL when we are writing */\n \tif ((flags & HASH_WRITE_OBJECT) != 0)\n-- \n2.42.0.8.g7a7e1e881e.dirty\n\n"},{"id":"482765","messageId":"9cc1f3014abe7fec997a99b6ac93d8ebb5455fa6.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH 2/7] bulk-checkin: factor out `prepare_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:01:53Z","receivedAt":"2023-10-06T22:02:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as the previous commit, factor out the routine to\nprepare streaming into a bulk-checkin pack into its own function. Unlike\nthe previous patch, this is a verbatim copy and paste.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 20 ++++++++++++++------\n 1 file changed, 14 insertions(+), 6 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 0aac3dfe31..377c41f3ad 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -260,6 +260,19 @@ static void format_object_header_hash(const struct git_hash_algo *algop,\n \talgop->update_fn(ctx, header, header_len);\n }\n \n+static void prepare_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t       struct hashfile_checkpoint *checkpoint,\n+\t\t\t       struct pack_idx_entry *idx,\n+\t\t\t       unsigned flags)\n+{\n+\tprepare_to_stream(state, flags);\n+\tif (idx) {\n+\t\thashfile_checkpoint(state->f, checkpoint);\n+\t\tidx->offset = state->offset;\n+\t\tcrc32_begin(state->f);\n+\t}\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -283,12 +296,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \talready_hashed_to = 0;\n \n \twhile (1) {\n-\t\tprepare_to_stream(state, flags);\n-\t\tif (idx) {\n-\t\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\t\tidx->offset = state->offset;\n-\t\t\tcrc32_begin(state->f);\n-\t\t}\n+\t\tprepare_checkpoint(state, &checkpoint, idx, flags);\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n \t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n-- \n2.42.0.8.g7a7e1e881e.dirty\n\n"},{"id":"482766","messageId":"f392ed2211d6d0d2dc99941f7fb44dcd978affb5.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH 3/7] bulk-checkin: factor out `truncate_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:01:55Z","receivedAt":"2023-10-06T22:02:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as previous commits, factor our the routine to\ntruncate a bulk-checkin packfile when writing past the pack size limit.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 27 +++++++++++++++++----------\n 1 file changed, 17 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 377c41f3ad..2dae8be461 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -273,6 +273,22 @@ static void prepare_checkpoint(struct bulk_checkin_packfile *state,\n \t}\n }\n \n+static void truncate_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t\tstruct hashfile_checkpoint *checkpoint,\n+\t\t\t\tstruct pack_idx_entry *idx)\n+{\n+\t/*\n+\t * Writing this object to the current pack will make\n+\t * it too big; we need to truncate it, start a new\n+\t * pack, and write into it.\n+\t */\n+\tif (!idx)\n+\t\tBUG(\"should not happen\");\n+\thashfile_truncate(state->f, checkpoint);\n+\tstate->offset = checkpoint->offset;\n+\tflush_bulk_checkin_packfile(state);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -300,16 +316,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n \t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n-\t\t/*\n-\t\t * Writing this object to the current pack will make\n-\t\t * it too big; we need to truncate it, start a new\n-\t\t * pack, and write into it.\n-\t\t */\n-\t\tif (!idx)\n-\t\t\tBUG(\"should not happen\");\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tflush_bulk_checkin_packfile(state);\n+\t\ttruncate_checkpoint(state, &checkpoint, idx);\n \t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n-- \n2.42.0.8.g7a7e1e881e.dirty\n\n"},{"id":"482767","messageId":"9c6ca564adf297e77e3304ef06692b8c82cddfd6.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH 4/7] bulk-checkin: factor our `finalize_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:01:58Z","receivedAt":"2023-10-06T22:02:06Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as previous commits, factor out the routine to\nfinalize the just-written object from the bulk-checkin mechanism.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 41 +++++++++++++++++++++++++----------------\n 1 file changed, 25 insertions(+), 16 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 2dae8be461..a9497fcb28 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -289,6 +289,30 @@ static void truncate_checkpoint(struct bulk_checkin_packfile *state,\n \tflush_bulk_checkin_packfile(state);\n }\n \n+static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t\tgit_hash_ctx *ctx,\n+\t\t\t\tstruct hashfile_checkpoint *checkpoint,\n+\t\t\t\tstruct pack_idx_entry *idx,\n+\t\t\t\tstruct object_id *result_oid)\n+{\n+\tthe_hash_algo->final_oid_fn(result_oid, ctx);\n+\tif (!idx)\n+\t\treturn;\n+\n+\tidx->crc32 = crc32_end(state->f);\n+\tif (already_written(state, result_oid)) {\n+\t\thashfile_truncate(state->f, checkpoint);\n+\t\tstate->offset = checkpoint->offset;\n+\t\tfree(idx);\n+\t} else {\n+\t\toidcpy(&idx->oid, result_oid);\n+\t\tALLOC_GROW(state->written,\n+\t\t\t   state->nr_written + 1,\n+\t\t\t   state->alloc_written);\n+\t\tstate->written[state->nr_written++] = idx;\n+\t}\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -320,22 +344,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n-\tthe_hash_algo->final_oid_fn(result_oid, &ctx);\n-\tif (!idx)\n-\t\treturn 0;\n-\n-\tidx->crc32 = crc32_end(state->f);\n-\tif (already_written(state, result_oid)) {\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tfree(idx);\n-\t} else {\n-\t\toidcpy(&idx->oid, result_oid);\n-\t\tALLOC_GROW(state->written,\n-\t\t\t   state->nr_written + 1,\n-\t\t\t   state->alloc_written);\n-\t\tstate->written[state->nr_written++] = idx;\n-\t}\n+\tfinalize_checkpoint(state, &ctx, &checkpoint, idx, result_oid);\n \treturn 0;\n }\n \n-- \n2.42.0.8.g7a7e1e881e.dirty\n\n"},{"id":"482768","messageId":"30ca7334c7605a81b9a6bbb386627e436bf8ab33.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH 5/7] bulk-checkin: introduce `index_blob_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:02:01Z","receivedAt":"2023-10-06T22:02:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that we have factored out many of the common routines necessary to\nindex a new object into a pack created by the bulk-checkin machinery, we\ncan introduce a variant of `index_blob_bulk_checkin()` that acts on\nblobs whose contents we can fit in memory.\n\nThis will be useful in a couple of more commits in order to provide the\n`merge-tree` builtin with a mechanism to create a new pack containing\nany objects it created during the merge, instead of storing those\nobjects individually as loose.\n\nSimilar to the existing `index_blob_bulk_checkin()` function, the\nentrypoint delegates to `deflate_blob_to_pack_incore()`, which is\nresponsible for formatting the pack header and then deflating the\ncontents into the pack. The latter is accomplished by calling\ndeflate_blob_contents_to_pack_incore(), which takes advantage of the\nearlier refactoring and is responsible for writing the object to the\npack and handling any overage from pack.packSizeLimit.\n\nThe bulk of the new functionality is implemented in the function\n`stream_obj_to_pack_incore()`, which is a generic implementation for\nwriting objects of arbitrary type (whose contents we can fit in-core)\ninto a bulk-checkin pack.\n\nThe new function shares an unfortunate degree of similarity to the\nexisting `stream_blob_to_pack()` function. But DRY-ing up these two\nwould likely be more trouble than it's worth, since the latter has to\ndeal with reading and writing the contents of the object.\n\nConsistent with the rest of the bulk-checkin mechanism, there are no\ndirect tests here. In future commits when we expose this new\nfunctionality via the `merge-tree` builtin, we will test it indirectly\nthere.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 116 +++++++++++++++++++++++++++++++++++++++++++++++++\n bulk-checkin.h |   4 ++\n 2 files changed, 120 insertions(+)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex a9497fcb28..319921efe7 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -140,6 +140,69 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n \treturn 0;\n }\n \n+static int stream_obj_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t     git_hash_ctx *ctx,\n+\t\t\t\t     off_t *already_hashed_to,\n+\t\t\t\t     const void *buf, size_t size,\n+\t\t\t\t     enum object_type type,\n+\t\t\t\t     const char *path, unsigned flags)\n+{\n+\tgit_zstream s;\n+\tunsigned char obuf[16384];\n+\tunsigned hdrlen;\n+\tint status = Z_OK;\n+\tint write_object = (flags & HASH_WRITE_OBJECT);\n+\n+\tgit_deflate_init(&s, pack_compression_level);\n+\n+\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), type, size);\n+\ts.next_out = obuf + hdrlen;\n+\ts.avail_out = sizeof(obuf) - hdrlen;\n+\n+\tif (*already_hashed_to < size) {\n+\t\tsize_t hsize = size - *already_hashed_to;\n+\t\tif (hsize) {\n+\t\t\tthe_hash_algo->update_fn(ctx, buf, hsize);\n+\t\t}\n+\t\t*already_hashed_to = size;\n+\t}\n+\ts.next_in = (void *)buf;\n+\ts.avail_in = size;\n+\n+\twhile (status != Z_STREAM_END) {\n+\t\tstatus = git_deflate(&s, Z_FINISH);\n+\t\tif (!s.avail_out || status == Z_STREAM_END) {\n+\t\t\tif (write_object) {\n+\t\t\t\tsize_t written = s.next_out - obuf;\n+\n+\t\t\t\t/* would we bust the size limit? */\n+\t\t\t\tif (state->nr_written &&\n+\t\t\t\t    pack_size_limit_cfg &&\n+\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n+\t\t\t\t\tgit_deflate_abort(&s);\n+\t\t\t\t\treturn -1;\n+\t\t\t\t}\n+\n+\t\t\t\thashwrite(state->f, obuf, written);\n+\t\t\t\tstate->offset += written;\n+\t\t\t}\n+\t\t\ts.next_out = obuf;\n+\t\t\ts.avail_out = sizeof(obuf);\n+\t\t}\n+\n+\t\tswitch (status) {\n+\t\tcase Z_OK:\n+\t\tcase Z_BUF_ERROR:\n+\t\tcase Z_STREAM_END:\n+\t\t\tcontinue;\n+\t\tdefault:\n+\t\t\tdie(\"unexpected deflate failure: %d\", status);\n+\t\t}\n+\t}\n+\tgit_deflate_end(&s);\n+\treturn 0;\n+}\n+\n /*\n  * Read the contents from fd for size bytes, streaming it to the\n  * packfile in state while updating the hash in ctx. Signal a failure\n@@ -313,6 +376,48 @@ static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n \t}\n }\n \n+static int deflate_obj_contents_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t\t       git_hash_ctx *ctx,\n+\t\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t\t       enum object_type type,\n+\t\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tstruct hashfile_checkpoint checkpoint = {0};\n+\tstruct pack_idx_entry *idx = NULL;\n+\toff_t already_hashed_to = 0;\n+\n+\t/* Note: idx is non-NULL when we are writing */\n+\tif (flags & HASH_WRITE_OBJECT)\n+\t\tCALLOC_ARRAY(idx, 1);\n+\n+\twhile (1) {\n+\t\tprepare_checkpoint(state, &checkpoint, idx, flags);\n+\t\tif (!stream_obj_to_pack_incore(state, ctx, &already_hashed_to,\n+\t\t\t\t\t       buf, size, type, path, flags))\n+\t\t\tbreak;\n+\t\ttruncate_checkpoint(state, &checkpoint, idx);\n+\t}\n+\n+\tfinalize_checkpoint(state, ctx, &checkpoint, idx, result_oid);\n+\n+\treturn 0;\n+}\n+\n+static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tgit_hash_ctx ctx;\n+\n+\tformat_object_header_hash(the_hash_algo, &ctx, OBJ_BLOB, size);\n+\n+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, result_oid,\n+\t\t\t\t\t\t   buf, size, OBJ_BLOB, path,\n+\t\t\t\t\t\t   flags);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -392,6 +497,17 @@ int index_blob_bulk_checkin(struct object_id *oid,\n \treturn status;\n }\n \n+int index_blob_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags)\n+{\n+\tint status = deflate_blob_to_pack_incore(&bulk_checkin_packfile, oid,\n+\t\t\t\t\t\t buf, size, path, flags);\n+\tif (!odb_transaction_nesting)\n+\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n+\treturn status;\n+}\n+\n void begin_odb_transaction(void)\n {\n \todb_transaction_nesting += 1;\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex aa7286a7b3..1b91daeaee 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -13,6 +13,10 @@ int index_blob_bulk_checkin(struct object_id *oid,\n \t\t\t    int fd, size_t size,\n \t\t\t    const char *path, unsigned flags);\n \n+int index_blob_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags);\n+\n /*\n  * Tell the object database to optimize for adding\n  * multiple objects. end_odb_transaction must be called\n-- \n2.42.0.8.g7a7e1e881e.dirty\n\n"},{"id":"482769","messageId":"cb0f79cabb7921ab7e334ad8a467ae84853bbd39.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH 6/7] bulk-checkin: introduce `index_tree_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:02:04Z","receivedAt":"2023-10-06T22:02:21Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The remaining missing piece in order to teach the `merge-tree` builtin\nhow to write the contents of a merge into a pack is a function to index\ntree objects into a bulk-checkin pack.\n\nThis patch implements that missing piece, which is a thin wrapper around\nall of the functionality introduced in previous commits.\n\nIf and when Git gains support for a \"compatibility\" hash algorithm, the\nchanges to support that here will be minimal. The bulk-checkin machinery\nwill need to convert the incoming tree to compute its length under the\ncompatibility hash, necessary to reconstruct its header. With that\ninformation (and the converted contents of the tree), the bulk-checkin\nmachinery will have enough to keep track of the converted object's hash\nin order to update the compatibility mapping.\n\nWithin `deflate_tree_to_pack_incore()`, the changes should be limited\nto something like:\n\n    if (the_repository->compat_hash_algo) {\n      struct strbuf converted = STRBUF_INIT;\n      if (convert_object_file(&compat_obj,\n                              the_repository->hash_algo,\n                              the_repository->compat_hash_algo, ...) < 0)\n        die(...);\n\n      format_object_header_hash(the_repository->compat_hash_algo,\n                                OBJ_TREE, size);\n\n      strbuf_release(&converted);\n    }\n\n, assuming related changes throughout the rest of the bulk-checkin\nmachinery necessary to update the hash of the converted object, which\nare likewise minimal in size.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 25 +++++++++++++++++++++++++\n bulk-checkin.h |  4 ++++\n 2 files changed, 29 insertions(+)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 319921efe7..d7d46f1dac 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -418,6 +418,20 @@ static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n \t\t\t\t\t\t   flags);\n }\n \n+static int deflate_tree_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tgit_hash_ctx ctx;\n+\n+\tformat_object_header_hash(the_hash_algo, &ctx, OBJ_TREE, size);\n+\n+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, result_oid,\n+\t\t\t\t\t\t   buf, size, OBJ_TREE, path,\n+\t\t\t\t\t\t   flags);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -508,6 +522,17 @@ int index_blob_bulk_checkin_incore(struct object_id *oid,\n \treturn status;\n }\n \n+int index_tree_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags)\n+{\n+\tint status = deflate_tree_to_pack_incore(&bulk_checkin_packfile, oid,\n+\t\t\t\t\t\t buf, size, path, flags);\n+\tif (!odb_transaction_nesting)\n+\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n+\treturn status;\n+}\n+\n void begin_odb_transaction(void)\n {\n \todb_transaction_nesting += 1;\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 1b91daeaee..89786b3954 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -17,6 +17,10 @@ int index_blob_bulk_checkin_incore(struct object_id *oid,\n \t\t\t\t   const void *buf, size_t size,\n \t\t\t\t   const char *path, unsigned flags);\n \n+int index_tree_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags);\n+\n /*\n  * Tell the object database to optimize for adding\n  * multiple objects. end_odb_transaction must be called\n-- \n2.42.0.8.g7a7e1e881e.dirty\n\n"},{"id":"482770","messageId":"e96921014557edb41dd73d93a8c3cf6cfaf0c719.1696629697.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T22:02:07Z","receivedAt":"2023-10-06T22:02:24Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using merge-tree often within a repository[^1], it is possible to\ngenerate a relatively large number of loose objects, which can result in\ndegraded performance, and inode exhaustion in extreme cases.\n\nBuilding on the functionality introduced in previous commits, the\nbulk-checkin machinery now has support to write arbitrary blob and tree\nobjects which are small enough to be held in-core. We can use this to\nwrite any blob/tree objects generated by ORT into a separate pack\ninstead of writing them out individually as loose.\n\nThis functionality is gated behind a new `--write-pack` option to\n`merge-tree` that works with the (non-deprecated) `--write-tree` mode.\n\nThe implementation is relatively straightforward. There are two spots\nwithin the ORT mechanism where we call `write_object_file()`, one for\ncontent differences within blobs, and another to assemble any new trees\nnecessary to construct the merge. In each of those locations,\nconditionally replace calls to `write_object_file()` with\n`index_blob_bulk_checkin_incore()` or `index_tree_bulk_checkin_incore()`\ndepending on which kind of object we are writing.\n\nThe only remaining task is to begin and end the transaction necessary to\ninitialize the bulk-checkin machinery, and move any new pack(s) it\ncreated into the main object store.\n\n[^1]: Such is the case at GitHub, where we run presumptive \"test merges\"\n  on open pull requests to see whether or not we can light up the merge\n  button green depending on whether or not the presumptive merge was\n  conflicted.\n\n  This is done in response to a number of user-initiated events,\n  including viewing an open pull request whose last test merge is stale\n  with respect to the current base and tip of the pull request. As a\n  result, merge-tree can be run very frequently on large, active\n  repositories.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-merge-tree.txt |  4 ++\n builtin/merge-tree.c             |  5 ++\n merge-ort.c                      | 43 +++++++++++----\n merge-recursive.h                |  1 +\n t/t4301-merge-tree-write-tree.sh | 93 ++++++++++++++++++++++++++++++++\n 5 files changed, 136 insertions(+), 10 deletions(-)\n\ndiff --git a/Documentation/git-merge-tree.txt b/Documentation/git-merge-tree.txt\nindex ffc4fbf7e8..9d37609ef1 100644\n--- a/Documentation/git-merge-tree.txt\n+++ b/Documentation/git-merge-tree.txt\n@@ -69,6 +69,10 @@ OPTIONS\n \tspecify a merge-base for the merge, and specifying multiple bases is\n \tcurrently not supported. This option is incompatible with `--stdin`.\n \n+--write-pack::\n+\tWrite any new objects into a separate packfile instead of as\n+\tindividual loose objects.\n+\n [[OUTPUT]]\n OUTPUT\n ------\ndiff --git a/builtin/merge-tree.c b/builtin/merge-tree.c\nindex 0de42aecf4..672ebd4c54 100644\n--- a/builtin/merge-tree.c\n+++ b/builtin/merge-tree.c\n@@ -18,6 +18,7 @@\n #include \"quote.h\"\n #include \"tree.h\"\n #include \"config.h\"\n+#include \"bulk-checkin.h\"\n \n static int line_termination = '\\n';\n \n@@ -414,6 +415,7 @@ struct merge_tree_options {\n \tint show_messages;\n \tint name_only;\n \tint use_stdin;\n+\tint write_pack;\n };\n \n static int real_merge(struct merge_tree_options *o,\n@@ -440,6 +442,7 @@ static int real_merge(struct merge_tree_options *o,\n \tinit_merge_options(&opt, the_repository);\n \n \topt.show_rename_progress = 0;\n+\topt.write_pack = o->write_pack;\n \n \topt.branch1 = branch1;\n \topt.branch2 = branch2;\n@@ -548,6 +551,8 @@ int cmd_merge_tree(int argc, const char **argv, const char *prefix)\n \t\t\t   &merge_base,\n \t\t\t   N_(\"commit\"),\n \t\t\t   N_(\"specify a merge-base for the merge\")),\n+\t\tOPT_BOOL(0, \"write-pack\", &o.write_pack,\n+\t\t\t N_(\"write new objects to a pack instead of as loose\")),\n \t\tOPT_END()\n \t};\n \ndiff --git a/merge-ort.c b/merge-ort.c\nindex 8631c99700..85d8c5c6b3 100644\n--- a/merge-ort.c\n+++ b/merge-ort.c\n@@ -48,6 +48,7 @@\n #include \"tree.h\"\n #include \"unpack-trees.h\"\n #include \"xdiff-interface.h\"\n+#include \"bulk-checkin.h\"\n \n /*\n  * We have many arrays of size 3.  Whenever we have such an array, the\n@@ -2124,11 +2125,19 @@ static int handle_content_merge(struct merge_options *opt,\n \t\tif ((merge_status < 0) || !result_buf.ptr)\n \t\t\tret = err(opt, _(\"Failed to execute internal merge\"));\n \n-\t\tif (!ret &&\n-\t\t    write_object_file(result_buf.ptr, result_buf.size,\n-\t\t\t\t      OBJ_BLOB, &result->oid))\n-\t\t\tret = err(opt, _(\"Unable to add %s to database\"),\n-\t\t\t\t  path);\n+\t\tif (!ret) {\n+\t\t\tret = opt->write_pack\n+\t\t\t\t? index_blob_bulk_checkin_incore(&result->oid,\n+\t\t\t\t\t\t\t\t result_buf.ptr,\n+\t\t\t\t\t\t\t\t result_buf.size,\n+\t\t\t\t\t\t\t\t path, 1)\n+\t\t\t\t: write_object_file(result_buf.ptr,\n+\t\t\t\t\t\t    result_buf.size,\n+\t\t\t\t\t\t    OBJ_BLOB, &result->oid);\n+\t\t\tif (ret)\n+\t\t\t\tret = err(opt, _(\"Unable to add %s to database\"),\n+\t\t\t\t\t  path);\n+\t\t}\n \n \t\tfree(result_buf.ptr);\n \t\tif (ret)\n@@ -3618,7 +3627,8 @@ static int tree_entry_order(const void *a_, const void *b_)\n \t\t\t\t b->string, strlen(b->string), bmi->result.mode);\n }\n \n-static int write_tree(struct object_id *result_oid,\n+static int write_tree(struct merge_options *opt,\n+\t\t      struct object_id *result_oid,\n \t\t      struct string_list *versions,\n \t\t      unsigned int offset,\n \t\t      size_t hash_size)\n@@ -3652,8 +3662,14 @@ static int write_tree(struct object_id *result_oid,\n \t}\n \n \t/* Write this object file out, and record in result_oid */\n-\tif (write_object_file(buf.buf, buf.len, OBJ_TREE, result_oid))\n+\tret = opt->write_pack\n+\t\t? index_tree_bulk_checkin_incore(result_oid,\n+\t\t\t\t\t\t buf.buf, buf.len, \"\", 1)\n+\t\t: write_object_file(buf.buf, buf.len, OBJ_TREE, result_oid);\n+\n+\tif (ret)\n \t\tret = -1;\n+\n \tstrbuf_release(&buf);\n \treturn ret;\n }\n@@ -3818,8 +3834,8 @@ static int write_completed_directory(struct merge_options *opt,\n \t\t */\n \t\tdir_info->is_null = 0;\n \t\tdir_info->result.mode = S_IFDIR;\n-\t\tif (write_tree(&dir_info->result.oid, &info->versions, offset,\n-\t\t\t       opt->repo->hash_algo->rawsz) < 0)\n+\t\tif (write_tree(opt, &dir_info->result.oid, &info->versions,\n+\t\t\t       offset, opt->repo->hash_algo->rawsz) < 0)\n \t\t\tret = -1;\n \t}\n \n@@ -4353,9 +4369,13 @@ static int process_entries(struct merge_options *opt,\n \t\tfflush(stdout);\n \t\tBUG(\"dir_metadata accounting completely off; shouldn't happen\");\n \t}\n-\tif (write_tree(result_oid, &dir_metadata.versions, 0,\n+\tif (write_tree(opt, result_oid, &dir_metadata.versions, 0,\n \t\t       opt->repo->hash_algo->rawsz) < 0)\n \t\tret = -1;\n+\n+\tif (opt->write_pack)\n+\t\tend_odb_transaction();\n+\n cleanup:\n \tstring_list_clear(&plist, 0);\n \tstring_list_clear(&dir_metadata.versions, 0);\n@@ -4899,6 +4919,9 @@ static void merge_start(struct merge_options *opt, struct merge_result *result)\n \t */\n \tstrmap_init(&opt->priv->conflicts);\n \n+\tif (opt->write_pack)\n+\t\tbegin_odb_transaction();\n+\n \ttrace2_region_leave(\"merge\", \"allocate/init\", opt->repo);\n }\n \ndiff --git a/merge-recursive.h b/merge-recursive.h\nindex b88000e3c2..156e160876 100644\n--- a/merge-recursive.h\n+++ b/merge-recursive.h\n@@ -48,6 +48,7 @@ struct merge_options {\n \tunsigned renormalize : 1;\n \tunsigned record_conflict_msgs_as_headers : 1;\n \tconst char *msg_header_prefix;\n+\tunsigned write_pack : 1;\n \n \t/* internal fields used by the implementation */\n \tstruct merge_options_internal *priv;\ndiff --git a/t/t4301-merge-tree-write-tree.sh b/t/t4301-merge-tree-write-tree.sh\nindex 250f721795..2d81ff4de5 100755\n--- a/t/t4301-merge-tree-write-tree.sh\n+++ b/t/t4301-merge-tree-write-tree.sh\n@@ -922,4 +922,97 @@ test_expect_success 'check the input format when --stdin is passed' '\n \ttest_cmp expect actual\n '\n \n+packdir=\".git/objects/pack\"\n+\n+test_expect_success 'merge-tree can pack its result with --write-pack' '\n+\ttest_when_finished \"rm -rf repo\" &&\n+\tgit init repo &&\n+\n+\t# base has lines [3, 4, 5]\n+\t#   - side adds to the beginning, resulting in [1, 2, 3, 4, 5]\n+\t#   - other adds to the end, resulting in [3, 4, 5, 6, 7]\n+\t#\n+\t# merging the two should result in a new blob object containing\n+\t# [1, 2, 3, 4, 5, 6, 7], along with a new tree.\n+\ttest_commit -C repo base file \"$(test_seq 3 5)\" &&\n+\tgit -C repo branch -M main &&\n+\tgit -C repo checkout -b side main &&\n+\ttest_commit -C repo side file \"$(test_seq 1 5)\" &&\n+\tgit -C repo checkout -b other main &&\n+\ttest_commit -C repo other file \"$(test_seq 3 7)\" &&\n+\n+\tfind repo/$packdir -type f -name \"pack-*.idx\" >packs.before &&\n+\ttree=\"$(git -C repo merge-tree --write-pack \\\n+\t\trefs/tags/side refs/tags/other)\" &&\n+\tblob=\"$(git -C repo rev-parse $tree:file)\" &&\n+\tfind repo/$packdir -type f -name \"pack-*.idx\" >packs.after &&\n+\n+\ttest_must_be_empty packs.before &&\n+\ttest_line_count = 1 packs.after &&\n+\n+\tgit show-index <$(cat packs.after) >objects &&\n+\ttest_line_count = 2 objects &&\n+\tgrep \"^[1-9][0-9]* $tree\" objects &&\n+\tgrep \"^[1-9][0-9]* $blob\" objects\n+'\n+\n+test_expect_success 'merge-tree can write multiple packs with --write-pack' '\n+\ttest_when_finished \"rm -rf repo\" &&\n+\tgit init repo &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\tgit config pack.packSizeLimit 512 &&\n+\n+\t\ttest_seq 512 >f &&\n+\n+\t\t# \"f\" contains roughly ~2,000 bytes.\n+\t\t#\n+\t\t# Each side (\"foo\" and \"bar\") adds a small amount of data at the\n+\t\t# beginning and end of \"base\", respectively.\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m base &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout -b foo main &&\n+\t\t{\n+\t\t\techo foo && cat f\n+\t\t} >f.tmp &&\n+\t\tmv f.tmp f &&\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m foo &&\n+\n+\t\tgit checkout -b bar main &&\n+\t\techo bar >>f &&\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m bar &&\n+\n+\t\tfind $packdir -type f -name \"pack-*.idx\" >packs.before &&\n+\t\t# Merging either side should result in a new object which is\n+\t\t# larger than 1M, thus the result should be split into two\n+\t\t# separate packs.\n+\t\ttree=\"$(git merge-tree --write-pack \\\n+\t\t\trefs/heads/foo refs/heads/bar)\" &&\n+\t\tblob=\"$(git rev-parse $tree:f)\" &&\n+\t\tfind $packdir -type f -name \"pack-*.idx\" >packs.after &&\n+\n+\t\ttest_must_be_empty packs.before &&\n+\t\ttest_line_count = 2 packs.after &&\n+\t\tfor idx in $(cat packs.after)\n+\t\tdo\n+\t\t\tgit show-index <$idx || return 1\n+\t\tdone >objects &&\n+\n+\t\t# The resulting set of packs should contain one copy of both\n+\t\t# objects, each in a separate pack.\n+\t\ttest_line_count = 2 objects &&\n+\t\tgrep \"^[1-9][0-9]* $tree\" objects &&\n+\t\tgrep \"^[1-9][0-9]* $blob\" objects\n+\n+\t)\n+'\n+\n test_done\n-- \n2.42.0.8.g7a7e1e881e.dirty\n"},{"id":"482772","messageId":"xmqqil7j751u.fsf@gitster.g","threadId":"60318","inReplyTo":"e96921014557edb41dd73d93a8c3cf6cfaf0c719.1696629697.git.me@ttaylorr.com","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-06T22:35:25Z","receivedAt":"2023-10-06T22:35:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> When using merge-tree often within a repository[^1], it is possible to\n> generate a relatively large number of loose objects, which can result in\n> degraded performance, and inode exhaustion in extreme cases.\n\nWell, be it \"git merge-tree\" or \"git merge\", new loose objects tend\nto accumulate until \"gc\" kicks in, so it is not a new problem for\nmere mortals, is it?\n\nAs one \"interesting\" use case of \"merge-tree\" is for a Git hosting\nsite with bare repositories to offer trial merges, without which\nmajority of the object their repositories acquire would have been in\npacks pushed by their users, \"Gee, loose objects consume many inodes\nin exchange for easier selective pruning\" becomes an issue, right?\n\nJust like it hurts performance to have too many loose object files,\npresumably it would also hurt performance to keep too many packs,\neach came from such a trial merge.  Do we have a \"gc\" story offered\nfor these packs created by the new feature?  E.g., \"once merge-tree\nis done creating a trial merge, we can discard the objects created\nin the pack, because we never expose new objects in the pack to the\noutside, processes running simultaneously, so instead closing the\nnew packfile by calling flush_bulk_checkin_packfile(), we can safely\nunlink the temporary pack.  We do not even need to spend cycles to\nrun a gc that requires cycles to enumerate what is still reachable\",\nor something like that?\n\nThanks.\n"},{"id":"482776","messageId":"ZSCR7e6KKqFv8mZk@nand.local","threadId":"60318","inReplyTo":"xmqqil7j751u.fsf@gitster.g","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-06T23:02:05Z","receivedAt":"2023-10-06T23:02:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Oct 06, 2023 at 03:35:25PM -0700, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > When using merge-tree often within a repository[^1], it is possible to\n> > generate a relatively large number of loose objects, which can result in\n> > degraded performance, and inode exhaustion in extreme cases.\n>\n> Well, be it \"git merge-tree\" or \"git merge\", new loose objects tend\n> to accumulate until \"gc\" kicks in, so it is not a new problem for\n> mere mortals, is it?\n\nYeah, I would definitely suspect that this is more of an issue for\nforges than individual Git users.\n\n> As one \"interesting\" use case of \"merge-tree\" is for a Git hosting\n> site with bare repositories to offer trial merges, without which\n> majority of the object their repositories acquire would have been in\n> packs pushed by their users, \"Gee, loose objects consume many inodes\n> in exchange for easier selective pruning\" becomes an issue, right?\n\nRight.\n\n> Just like it hurts performance to have too many loose object files,\n> presumably it would also hurt performance to keep too many packs,\n> each came from such a trial merge.  Do we have a \"gc\" story offered\n> for these packs created by the new feature?  E.g., \"once merge-tree\n> is done creating a trial merge, we can discard the objects created\n> in the pack, because we never expose new objects in the pack to the\n> outside, processes running simultaneously, so instead closing the\n> new packfile by calling flush_bulk_checkin_packfile(), we can safely\n> unlink the temporary pack.  We do not even need to spend cycles to\n> run a gc that requires cycles to enumerate what is still reachable\",\n> or something like that?\n\nI know Johannes worked on something like this recently. IIRC, it\neffectively does something like:\n\n    struct tmp_objdir *tmp_objdir = tmp_objdir_create(...);\n    tmp_objdir_replace_primary_odb(tmp_objdir, 1);\n\nat the beginning of a merge operation, and:\n\n    tmp_objdir_discard_objects(tmp_objdir);\n\nat the end. I haven't followed that work off-list very closely, but it\nis only possible for GitHub to discard certain niche kinds of\nmerges/rebases, since in general we make the objects created during test\nmerges available via refs/pull/N/{merge,rebase}.\n\nI think that like anything, this is a trade-off. Having lots of packs\ncan be a performance hindrance just like having lots of loose objects.\nBut since we can represent more objects with fewer inodes when packed,\nstoring those objects together in a pack is preferable when (a) you're\ndoing lots of test-merges, and (b) you want to keep those objects\naround, e.g., because they are reachable.\n\nThanks,\nTaylor\n"},{"id":"482784","messageId":"E81727B0-A523-4A45-A606-61442357291D@gmail.com","threadId":"60318","inReplyTo":"cb0f79cabb7921ab7e334ad8a467ae84853bbd39.1696629697.git.me@ttaylorr.com","subject":"Re: [PATCH 6/7] bulk-checkin: introduce `index_tree_bulk_checkin_incore()`","fromName":"Eric Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-07T03:07:20Z","receivedAt":"2023-10-07T03:07:34Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"\n\nOn October 6, 2023 5:02:04 PM CDT, Taylor Blau <me@ttaylorr.com> wrote:\n>\n>Within `deflate_tree_to_pack_incore()`, the changes should be limited\n>to something like:\n>\n>    if (the_repository->compat_hash_algo) {\n>      struct strbuf converted = STRBUF_INIT;\n>      if (convert_object_file(&compat_obj,\n>                              the_repository->hash_algo,\n>                              the_repository->compat_hash_algo, ...) < 0)\n>        die(...);\n>\n>      format_object_header_hash(the_repository->compat_hash_algo,\n>                                OBJ_TREE, size);\n>\n>      strbuf_release(&converted);\n>    }\n>\n>, assuming related changes throughout the rest of the bulk-checkin\n>machinery necessary to update the hash of the converted object, which\n>are likewise minimal in size.\n\nSo this is close.   Just in case someone wants to\ngo down this path I want to point out that\nthe converted object need to have the compat hash computed over it.\n\nWhich means that the strbuf_release in your example comes a bit early.\n\n\nEric\n"},{"id":"482818","messageId":"CABPp-BE+mJ4e==fWNqUNi5RVkoui_xeZN+axnM6vBykDqAzHiA@mail.gmail.com","threadId":"60318","inReplyTo":"ZSCR7e6KKqFv8mZk@nand.local","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2023-10-08T07:02:27Z","receivedAt":"2023-10-08T07:02:48Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Oct 6, 2023 at 4:02 PM Taylor Blau <me@ttaylorr.com> wrote:\n>\n> On Fri, Oct 06, 2023 at 03:35:25PM -0700, Junio C Hamano wrote:\n> > Taylor Blau <me@ttaylorr.com> writes:\n> >\n> > > When using merge-tree often within a repository[^1], it is possible to\n> > > generate a relatively large number of loose objects, which can result in\n> > > degraded performance, and inode exhaustion in extreme cases.\n> >\n> > Well, be it \"git merge-tree\" or \"git merge\", new loose objects tend\n> > to accumulate until \"gc\" kicks in, so it is not a new problem for\n> > mere mortals, is it?\n>\n> Yeah, I would definitely suspect that this is more of an issue for\n> forges than individual Git users.\n\nIt may still be nice to also do this optimization for plain \"git\nmerge\" as well.  I had it in my list of ideas somewhere to do a\n\"fast-import-like\" thing to avoid writing loose objects, as I\nsuspected that'd actually be a performance impediment.\n\n> > As one \"interesting\" use case of \"merge-tree\" is for a Git hosting\n> > site with bare repositories to offer trial merges, without which\n> > majority of the object their repositories acquire would have been in\n> > packs pushed by their users, \"Gee, loose objects consume many inodes\n> > in exchange for easier selective pruning\" becomes an issue, right?\n>\n> Right.\n>\n> > Just like it hurts performance to have too many loose object files,\n> > presumably it would also hurt performance to keep too many packs,\n> > each came from such a trial merge.  Do we have a \"gc\" story offered\n> > for these packs created by the new feature?  E.g., \"once merge-tree\n> > is done creating a trial merge, we can discard the objects created\n> > in the pack, because we never expose new objects in the pack to the\n> > outside, processes running simultaneously, so instead closing the\n> > new packfile by calling flush_bulk_checkin_packfile(), we can safely\n> > unlink the temporary pack.  We do not even need to spend cycles to\n> > run a gc that requires cycles to enumerate what is still reachable\",\n> > or something like that?\n>\n> I know Johannes worked on something like this recently. IIRC, it\n> effectively does something like:\n>\n>     struct tmp_objdir *tmp_objdir = tmp_objdir_create(...);\n>     tmp_objdir_replace_primary_odb(tmp_objdir, 1);\n>\n> at the beginning of a merge operation, and:\n>\n>     tmp_objdir_discard_objects(tmp_objdir);\n>\n> at the end. I haven't followed that work off-list very closely, but it\n> is only possible for GitHub to discard certain niche kinds of\n> merges/rebases, since in general we make the objects created during test\n> merges available via refs/pull/N/{merge,rebase}.\n\nOh, at the contributor summit, Johannes said he only needed pass/fail,\nnot the actual commits, which is why I suggested this route.  If you\nneed to keep the actual commits, then this won't help.\n\nI was interested in the same question as Junio, but from a different\nangle.  fast-import documentation points out that the packs it creates\nare suboptimal with poorer delta choices.  Are the packs created by\nbulk-checkin prone to the same issues?  When I was thinking in terms\nof having \"git merge\" use fast-import for pack creation instead of\nwriting loose objects (an idea I never investigated very far), I was\nwondering if I'd need to mark those packs as \"less optimal\" and do\nsomething to make sure they were more likely to be repacked.\n\nI believe geometric repacking didn't exist back when I was thinking\nabout this, and perhaps geometric repacking automatically handles\nthings nicely for us.  Does it, or are we risking retaining\nsub-optimal deltas from the bulk-checkin code?\n\n(I've never really cracked open the pack code, so I have absolutely no\nidea; I'm just curious.)\n\n> I think that like anything, this is a trade-off. Having lots of packs\n> can be a performance hindrance just like having lots of loose objects.\n> But since we can represent more objects with fewer inodes when packed,\n> storing those objects together in a pack is preferable when (a) you're\n> doing lots of test-merges, and (b) you want to keep those objects\n> around, e.g., because they are reachable.\n\nA couple of the comments earlier in the series suggested this was\nabout streaming blobs to a pack in the bulk checkin code.  Are tree\nand commit objects also put in the pack, or will those continue to be\nwritten loosely?\n"},{"id":"482824","messageId":"ZSLS9G1lHruig48a@nand.local","threadId":"60318","inReplyTo":"CABPp-BE+mJ4e==fWNqUNi5RVkoui_xeZN+axnM6vBykDqAzHiA@mail.gmail.com","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-08T16:04:04Z","receivedAt":"2023-10-08T16:04:10Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sun, Oct 08, 2023 at 12:02:27AM -0700, Elijah Newren wrote:\n> On Fri, Oct 6, 2023 at 4:02 PM Taylor Blau <me@ttaylorr.com> wrote:\n> >\n> > On Fri, Oct 06, 2023 at 03:35:25PM -0700, Junio C Hamano wrote:\n> > > Taylor Blau <me@ttaylorr.com> writes:\n> > >\n> > > > When using merge-tree often within a repository[^1], it is possible to\n> > > > generate a relatively large number of loose objects, which can result in\n> > > > degraded performance, and inode exhaustion in extreme cases.\n> > >\n> > > Well, be it \"git merge-tree\" or \"git merge\", new loose objects tend\n> > > to accumulate until \"gc\" kicks in, so it is not a new problem for\n> > > mere mortals, is it?\n> >\n> > Yeah, I would definitely suspect that this is more of an issue for\n> > forges than individual Git users.\n>\n> It may still be nice to also do this optimization for plain \"git\n> merge\" as well.  I had it in my list of ideas somewhere to do a\n> \"fast-import-like\" thing to avoid writing loose objects, as I\n> suspected that'd actually be a performance impediment.\n\nI think that would be worth doing, definitely. I do worry a little bit\nabout locking in low-quality deltas (or lack thereof), but more on that\nbelow...\n\n> Oh, at the contributor summit, Johannes said he only needed pass/fail,\n> not the actual commits, which is why I suggested this route.  If you\n> need to keep the actual commits, then this won't help.\n\nYep, agreed. Like I said earlier, I think there are some niche scenarios\nwhere we just care about \"would this merge cleanly?\", but in most other\ncases we want to keep around the actual tree.\n\n> I was interested in the same question as Junio, but from a different\n> angle.  fast-import documentation points out that the packs it creates\n> are suboptimal with poorer delta choices.  Are the packs created by\n> bulk-checkin prone to the same issues?  When I was thinking in terms\n> of having \"git merge\" use fast-import for pack creation instead of\n> writing loose objects (an idea I never investigated very far), I was\n> wondering if I'd need to mark those packs as \"less optimal\" and do\n> something to make sure they were more likely to be repacked.\n>\n> I believe geometric repacking didn't exist back when I was thinking\n> about this, and perhaps geometric repacking automatically handles\n> things nicely for us.  Does it, or are we risking retaining\n> sub-optimal deltas from the bulk-checkin code?\n>\n> (I've never really cracked open the pack code, so I have absolutely no\n> idea; I'm just curious.)\n\nYes, the bulk-checkin mechanism suffers from an even worse problem which\nis the pack it creates will contain no deltas whatsoever. The contents\nof the pack are just getting written as-is, so there's no fancy\ndelta-ficiation going on.\n\nI think Michael Haggerty (?) suggested to me off-list that it might be\ninteresting to have a flag that we could mark packs with bad/no deltas\nas such so that we don't implicitly trust their contents as having high\nquality deltas.\n\n> > I think that like anything, this is a trade-off. Having lots of packs\n> > can be a performance hindrance just like having lots of loose objects.\n> > But since we can represent more objects with fewer inodes when packed,\n> > storing those objects together in a pack is preferable when (a) you're\n> > doing lots of test-merges, and (b) you want to keep those objects\n> > around, e.g., because they are reachable.\n>\n> A couple of the comments earlier in the series suggested this was\n> about streaming blobs to a pack in the bulk checkin code.  Are tree\n> and commit objects also put in the pack, or will those continue to be\n> written loosely?\n\nThis covers both blobs and trees, since IIUC that's all we'd need to\nimplement support for merge-tree to be able to write any objects it\ncreates into a pack. AFAIK merge-tree never generates any commit\nobjects. But teaching 'merge' to perform the same bulk-checkin trick\nwould just require us implementing index_bulk_commit_checkin_in_core()\nor similar, which is straightforward to do on top of the existing code.\n\nThanks,\nTaylor\n"},{"id":"482828","messageId":"20231008173329.GA1557002@coredump.intra.peff.net","threadId":"60318","inReplyTo":"ZSLS9G1lHruig48a@nand.local","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-10-08T17:33:29Z","receivedAt":"2023-10-08T17:33:44Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Oct 08, 2023 at 12:04:04PM -0400, Taylor Blau wrote:\n\n> > I was interested in the same question as Junio, but from a different\n> > angle.  fast-import documentation points out that the packs it creates\n> > are suboptimal with poorer delta choices.  Are the packs created by\n> > bulk-checkin prone to the same issues?  When I was thinking in terms\n> > of having \"git merge\" use fast-import for pack creation instead of\n> > writing loose objects (an idea I never investigated very far), I was\n> > wondering if I'd need to mark those packs as \"less optimal\" and do\n> > something to make sure they were more likely to be repacked.\n> >\n> > I believe geometric repacking didn't exist back when I was thinking\n> > about this, and perhaps geometric repacking automatically handles\n> > things nicely for us.  Does it, or are we risking retaining\n> > sub-optimal deltas from the bulk-checkin code?\n> >\n> > (I've never really cracked open the pack code, so I have absolutely no\n> > idea; I'm just curious.)\n> \n> Yes, the bulk-checkin mechanism suffers from an even worse problem which\n> is the pack it creates will contain no deltas whatsoever. The contents\n> of the pack are just getting written as-is, so there's no fancy\n> delta-ficiation going on.\n\nI wonder how big a deal this would be in practice for merges.\npack-objects will look for deltas between any two candidates objects,\nbut in practice I think most deltas are between objects from multiple\ncommits (across the \"time\" dimension, if you will) rather than within a\nsingle tree (the \"space\" dimension). And a merge operation is generally\ncreating a single new tree (recursive merging may create intermediate\nstates which would delta, but we don't actually need to keep those\nintermediate ones. I won't be surprised if we do, though).\n\nWe should be able to test that theory by looking at existing deltas.\nHere's a script which builds an index of blobs and trees to the commits\nthat introduce them:\n\n  git rev-list HEAD |\n  git diff-tree --stdin -r -m -t --raw |\n  perl -lne '\n    if (/^[0-9a-f]/) {\n      $commit = $_;\n    } elsif (/^:\\S+ \\S+ \\S+ (\\S+)/) {\n      $h{$1} = $commit;\n    }\n    END { print \"$_ $h{$_}\" for keys(%h) }\n  ' >commits.db\n\nAnd then we can see which deltas come from the same commit:\n\n  git cat-file --batch-all-objects --batch-check='%(objectname) %(deltabase)' |\n  perl -alne '\n    BEGIN {\n      open(my $fh, \"<\", \"commits.db\");\n      %commit = map { chomp; split } <$fh>;\n    }\n    next if $F[1] =~ /0{40}/; # not a delta\n    next unless defined $commit{$F[0]}; # not in index\n    print $commit{$F[0]} eq $commit{$F[1]} ? \"inner\" : \"outer\", \" \", $_;\n  '\n\nIn git.git, I see 460 \"inner\" deltas, and 193,081 \"outer\" ones. The\ninner ones are mostly small single-file test vectors, which makes sense.\nIt's possible to have a merge result that does conflict resolution in\ntwo such files (that would then delta), but it seems like a fairly\nunlikely case. Numbers for linux.git are similar.\n\nSo it might just not be a big issue at all for this use case.\n\n> I think Michael Haggerty (?) suggested to me off-list that it might be\n> interesting to have a flag that we could mark packs with bad/no deltas\n> as such so that we don't implicitly trust their contents as having high\n> quality deltas.\n\nI was going to suggest the same thing. ;) Unfortunately it's a bit\ntricky to do as we have no room in the file format for an optional flag.\nYou'd have to add a \".mediocre-delta\" file or something.\n\nBut here's another approach. I recall discussing a while back the idea\nthat we should not necessarily trust the quality of deltas in packs that\nare pushed (and I think Thomas Gummerer even did some experiments inside\nGitHub with those, though I don't remember the results). And one way\naround that is during geometric repacking to consider the biggest/oldest\npack as \"preferred\", reuse its deltas, but always compute from scratch\nwith the others (neither reusing on-disk deltas, nor skipping\ntry_delta() when two objects come from the same pack).\n\nThat same strategy would work here (and for incremental fast-import\npacks, though of course not if your fast-import pack is the \"big\" one\nafter you do a from-scratch import).\n\n> > A couple of the comments earlier in the series suggested this was\n> > about streaming blobs to a pack in the bulk checkin code.  Are tree\n> > and commit objects also put in the pack, or will those continue to be\n> > written loosely?\n> \n> This covers both blobs and trees, since IIUC that's all we'd need to\n> implement support for merge-tree to be able to write any objects it\n> creates into a pack. AFAIK merge-tree never generates any commit\n> objects. But teaching 'merge' to perform the same bulk-checkin trick\n> would just require us implementing index_bulk_commit_checkin_in_core()\n> or similar, which is straightforward to do on top of the existing code.\n\nThis is a bit of a devil's advocate question, but: would it make sense\nto implement this as a general of Git's object-writing code, and not tie\nit to a specific command? That is, what if a user could do:\n\n  git --write-to-pack merge ...\n\nbut also:\n\n  git --write-to-pack add ...\n\nand the object-writing code would just write to a pack instead of\nwriting loose objects. That lets the caller decide when it is or is not\na good idea to use this mode. And if making loose objects gives bad\nperformance for merges, wouldn't the same be true of other operations\nwhich generate many objects?\n\nPossibly it exacerbates the \"no deltas\" issue from above (though it\nwould depend on the command).  The bigger question to me is one of\ncheckpointing. When do we finish off the pack with a .idx and make it\navailable to other readers? We could do it at program exit, but I\nsuspect there are some commands that really want to make objects\navailable sooner (e.g., as soon as \"git add\" writes an index, we'd want\nthose objects to already be available). Probably every program that\nwrites objects would need to be annotated with a checkpoint call (which\nwould be a noop in loose object mode).\n\nSo maybe it's a dumb direction. I dunno.\n\n-Peff\n"},{"id":"482840","messageId":"ZSNX5CQPWOy5B+cg@nand.local","threadId":"60318","inReplyTo":"E81727B0-A523-4A45-A606-61442357291D@gmail.com","subject":"Re: [PATCH 6/7] bulk-checkin: introduce `index_tree_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-09T01:31:16Z","receivedAt":"2023-10-09T01:31:22Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Oct 06, 2023 at 10:07:20PM -0500, Eric Biederman wrote:\n> On October 6, 2023 5:02:04 PM CDT, Taylor Blau <me@ttaylorr.com> wrote:\n> >Within `deflate_tree_to_pack_incore()`, the changes should be limited\n> >to something like:\n> >\n> >    if (the_repository->compat_hash_algo) {\n> >      struct strbuf converted = STRBUF_INIT;\n> >      if (convert_object_file(&compat_obj,\n> >                              the_repository->hash_algo,\n> >                              the_repository->compat_hash_algo, ...) < 0)\n> >        die(...);\n> >\n> >      format_object_header_hash(the_repository->compat_hash_algo,\n> >                                OBJ_TREE, size);\n> >\n> >      strbuf_release(&converted);\n> >    }\n> >\n> >, assuming related changes throughout the rest of the bulk-checkin\n> >machinery necessary to update the hash of the converted object, which\n> >are likewise minimal in size.\n>\n> So this is close.   Just in case someone wants to\n> go down this path I want to point out that\n> the converted object need to have the compat hash computed over it.\n>\n> Which means that the strbuf_release in your example comes a bit early.\n\nDoh. You're absolutely right. Let's fix this if/when we cross that\nbridge ;-).\n\nThanks,\nTaylor\n"},{"id":"482841","messageId":"ZSNZZrWyCqRH+0Bd@nand.local","threadId":"60318","inReplyTo":"20231008173329.GA1557002@coredump.intra.peff.net","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-09T01:37:42Z","receivedAt":"2023-10-09T01:37:49Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sun, Oct 08, 2023 at 01:33:29PM -0400, Jeff King wrote:\n> On Sun, Oct 08, 2023 at 12:04:04PM -0400, Taylor Blau wrote:\n>\n> > > I was interested in the same question as Junio, but from a different\n> > > angle.  fast-import documentation points out that the packs it creates\n> > > are suboptimal with poorer delta choices.  Are the packs created by\n> > > bulk-checkin prone to the same issues?  When I was thinking in terms\n> > > of having \"git merge\" use fast-import for pack creation instead of\n> > > writing loose objects (an idea I never investigated very far), I was\n> > > wondering if I'd need to mark those packs as \"less optimal\" and do\n> > > something to make sure they were more likely to be repacked.\n> > >\n> > > I believe geometric repacking didn't exist back when I was thinking\n> > > about this, and perhaps geometric repacking automatically handles\n> > > things nicely for us.  Does it, or are we risking retaining\n> > > sub-optimal deltas from the bulk-checkin code?\n> > >\n> > > (I've never really cracked open the pack code, so I have absolutely no\n> > > idea; I'm just curious.)\n> >\n> > Yes, the bulk-checkin mechanism suffers from an even worse problem which\n> > is the pack it creates will contain no deltas whatsoever. The contents\n> > of the pack are just getting written as-is, so there's no fancy\n> > delta-ficiation going on.\n>\n> I wonder how big a deal this would be in practice for merges.\n> pack-objects will look for deltas between any two candidates objects,\n> but in practice I think most deltas are between objects from multiple\n> commits (across the \"time\" dimension, if you will) rather than within a\n> single tree (the \"space\" dimension). And a merge operation is generally\n> creating a single new tree (recursive merging may create intermediate\n> states which would delta, but we don't actually need to keep those\n> intermediate ones. I won't be surprised if we do, though).\n>\n> We should be able to test that theory by looking at existing deltas.\n> Here's a script which builds an index of blobs and trees to the commits\n> that introduce them:\n>\n>   git rev-list HEAD |\n>   git diff-tree --stdin -r -m -t --raw |\n>   perl -lne '\n>     if (/^[0-9a-f]/) {\n>       $commit = $_;\n>     } elsif (/^:\\S+ \\S+ \\S+ (\\S+)/) {\n>       $h{$1} = $commit;\n>     }\n>     END { print \"$_ $h{$_}\" for keys(%h) }\n>   ' >commits.db\n>\n> And then we can see which deltas come from the same commit:\n>\n>   git cat-file --batch-all-objects --batch-check='%(objectname) %(deltabase)' |\n>   perl -alne '\n>     BEGIN {\n>       open(my $fh, \"<\", \"commits.db\");\n>       %commit = map { chomp; split } <$fh>;\n>     }\n>     next if $F[1] =~ /0{40}/; # not a delta\n>     next unless defined $commit{$F[0]}; # not in index\n>     print $commit{$F[0]} eq $commit{$F[1]} ? \"inner\" : \"outer\", \" \", $_;\n>   '\n>\n> In git.git, I see 460 \"inner\" deltas, and 193,081 \"outer\" ones. The\n> inner ones are mostly small single-file test vectors, which makes sense.\n> It's possible to have a merge result that does conflict resolution in\n> two such files (that would then delta), but it seems like a fairly\n> unlikely case. Numbers for linux.git are similar.\n>\n> So it might just not be a big issue at all for this use case.\n\nVery interesting, thanks for running (and documenting!) this experiment.\nI'm mostly with you that it probably doesn't make a huge difference in\npractice here.\n\nOne thing that I'm not entirely clear on is how we'd treat objects that\ncould be good delta candidates for each other between two packs. For\ninstance, if I write a tree corresponding to the merge between two\nbranches, it's likely that the resulting tree would be a good delta\ncandidate against either of the trees at the tips of those two refs.\n\nBut we won't pack those trees (the ones at the tips of the refs) in the\nsame pack as the tree containing their merge. If we later on tried to\nrepack, would we evaluate the tip trees as possible delta candidates\nagainst the merged tree? Or would we look at the merged tree, realize it\nisn't delta'd with anything, and then not attempt to find any\ncandidates?\n\n> > I think Michael Haggerty (?) suggested to me off-list that it might be\n> > interesting to have a flag that we could mark packs with bad/no deltas\n> > as such so that we don't implicitly trust their contents as having high\n> > quality deltas.\n>\n> I was going to suggest the same thing. ;) Unfortunately it's a bit\n> tricky to do as we have no room in the file format for an optional flag.\n> You'd have to add a \".mediocre-delta\" file or something.\n\nYeah, I figured that we'd add a new \".baddeltas\" file or something. (As\nan aside, we probably should have an optional flags section in the .pack\nformat, since we seem to have a lot of optional pack extensions: .rev,\n.bitmap, .keep, .promisor, etc.)\n\n> But here's another approach. I recall discussing a while back the idea\n> that we should not necessarily trust the quality of deltas in packs that\n> are pushed (and I think Thomas Gummerer even did some experiments inside\n> GitHub with those, though I don't remember the results). And one way\n> around that is during geometric repacking to consider the biggest/oldest\n> pack as \"preferred\", reuse its deltas, but always compute from scratch\n> with the others (neither reusing on-disk deltas, nor skipping\n> try_delta() when two objects come from the same pack).\n>\n> That same strategy would work here (and for incremental fast-import\n> packs, though of course not if your fast-import pack is the \"big\" one\n> after you do a from-scratch import).\n\nYeah, definitely. I think that that code is still living in GitHub's\nfork, but inactive since we haven't set any of the relevant\nconfiguration in GitHub's production environment.\n\n> Possibly it exacerbates the \"no deltas\" issue from above (though it\n> would depend on the command).  The bigger question to me is one of\n> checkpointing. When do we finish off the pack with a .idx and make it\n> available to other readers? We could do it at program exit, but I\n> suspect there are some commands that really want to make objects\n> available sooner (e.g., as soon as \"git add\" writes an index, we'd want\n> those objects to already be available). Probably every program that\n> writes objects would need to be annotated with a checkpoint call (which\n> would be a noop in loose object mode).\n>\n> So maybe it's a dumb direction. I dunno.\n\nI wouldn't say it's a dumb direction ;-). But I'd be hesitant pursuing\nit without solving the \"no deltas\" question from earlier.\n\nThanks,\nTaylor\n"},{"id":"482860","messageId":"ZSPb1OYRrQSUugtg@tanuki","threadId":"60318","inReplyTo":"ZSCR7e6KKqFv8mZk@nand.local","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2023-10-09T10:54:12Z","receivedAt":"2023-10-09T11:00:05Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Oct 06, 2023 at 07:02:05PM -0400, Taylor Blau wrote:\n> On Fri, Oct 06, 2023 at 03:35:25PM -0700, Junio C Hamano wrote:\n> > Taylor Blau <me@ttaylorr.com> writes:\n> >\n> > > When using merge-tree often within a repository[^1], it is possible to\n> > > generate a relatively large number of loose objects, which can result in\n> > > degraded performance, and inode exhaustion in extreme cases.\n> >\n> > Well, be it \"git merge-tree\" or \"git merge\", new loose objects tend\n> > to accumulate until \"gc\" kicks in, so it is not a new problem for\n> > mere mortals, is it?\n> \n> Yeah, I would definitely suspect that this is more of an issue for\n> forges than individual Git users.\n> \n> > As one \"interesting\" use case of \"merge-tree\" is for a Git hosting\n> > site with bare repositories to offer trial merges, without which\n> > majority of the object their repositories acquire would have been in\n> > packs pushed by their users, \"Gee, loose objects consume many inodes\n> > in exchange for easier selective pruning\" becomes an issue, right?\n> \n> Right.\n> \n> > Just like it hurts performance to have too many loose object files,\n> > presumably it would also hurt performance to keep too many packs,\n> > each came from such a trial merge.  Do we have a \"gc\" story offered\n> > for these packs created by the new feature?  E.g., \"once merge-tree\n> > is done creating a trial merge, we can discard the objects created\n> > in the pack, because we never expose new objects in the pack to the\n> > outside, processes running simultaneously, so instead closing the\n> > new packfile by calling flush_bulk_checkin_packfile(), we can safely\n> > unlink the temporary pack.  We do not even need to spend cycles to\n> > run a gc that requires cycles to enumerate what is still reachable\",\n> > or something like that?\n> \n> I know Johannes worked on something like this recently. IIRC, it\n> effectively does something like:\n> \n>     struct tmp_objdir *tmp_objdir = tmp_objdir_create(...);\n>     tmp_objdir_replace_primary_odb(tmp_objdir, 1);\n> \n> at the beginning of a merge operation, and:\n> \n>     tmp_objdir_discard_objects(tmp_objdir);\n> \n> at the end. I haven't followed that work off-list very closely, but it\n> is only possible for GitHub to discard certain niche kinds of\n> merges/rebases, since in general we make the objects created during test\n> merges available via refs/pull/N/{merge,rebase}.\n> \n> I think that like anything, this is a trade-off. Having lots of packs\n> can be a performance hindrance just like having lots of loose objects.\n> But since we can represent more objects with fewer inodes when packed,\n> storing those objects together in a pack is preferable when (a) you're\n> doing lots of test-merges, and (b) you want to keep those objects\n> around, e.g., because they are reachable.\n\nIn Gitaly, we usually set up quarantine directories for all operations\nthat create objects. This allows us to discard any newly written objects\nin case either the RPC call gets cancelled or in case our access checks\ndetermine that the change should not be allowed. The logic is rather\nsimple:\n\n    1. Create a new temporary directory.\n\n    2. Set up the new temporary directory as main object database via\n       the `GIT_OBJECT_DIRECTORY` environment variable.\n\n    3. Set up the main repository's object database via the\n       `GIT_ALTERNATE_OBJECT_DIRECTORIES` environment variable.\n\n    4. Execute Git commands that write objects with these environment\n       variables set up. The new objects will end up neatly contained in\n       the temporary directory.\n\n    5. Once done, either discard the temporary object database or\n       migrate objects into the main object daatabase.\n\nI wonder whether this would be a viable approach for you, as well.\n\nPatrick\n"},{"id":"482871","messageId":"ZSQlgLDv+MrYmSp8@nand.local","threadId":"60318","inReplyTo":"ZSPb1OYRrQSUugtg@tanuki","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-09T16:08:32Z","receivedAt":"2023-10-09T16:08:37Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Oct 09, 2023 at 12:54:12PM +0200, Patrick Steinhardt wrote:\n> In Gitaly, we usually set up quarantine directories for all operations\n> that create objects. This allows us to discard any newly written objects\n> in case either the RPC call gets cancelled or in case our access checks\n> determine that the change should not be allowed. The logic is rather\n> simple:\n>\n>     1. Create a new temporary directory.\n>\n>     2. Set up the new temporary directory as main object database via\n>        the `GIT_OBJECT_DIRECTORY` environment variable.\n>\n>     3. Set up the main repository's object database via the\n>        `GIT_ALTERNATE_OBJECT_DIRECTORIES` environment variable.\n\nIs there a reason not to run Git in the quarantine environment and list\nthe main repository as an alternate via $GIT_DIR/objects/info/alternates\ninstead of the GIT_ALTERNATE_OBJECT_DIRECTORIES environment variable?\n\n>     4. Execute Git commands that write objects with these environment\n>        variables set up. The new objects will end up neatly contained in\n>        the temporary directory.\n>\n>     5. Once done, either discard the temporary object database or\n>        migrate objects into the main object daatabase.\n\nInteresting. I'm curious why you don't use the builtin tmp_objdir\nmechanism in Git itself. Do you need to run more than one command in the\nquarantine environment? If so, that makes sense that you'd want to have\na scratch repository that lasts beyond the lifetime of a single process.\n\n> I wonder whether this would be a viable approach for you, as well.\n\nI think that the main problem that we are trying to solve with this\nseries is creating a potentially large number of loose objects. I think\nthat you could do something like what you propose above, with a 'git\nrepacks -adk' before moving its objects over back to the main repository.\n\nBut since we're working in a single process only when doing a merge-tree\noperation, I think it is probably more expedient to write the pack's\nbytes directly.\n\n> Patrick\n\n\nThanks,\nTaylor\n"},{"id":"482880","messageId":"xmqqy1gb1zfd.fsf@gitster.g","threadId":"60318","inReplyTo":"20231008173329.GA1557002@coredump.intra.peff.net","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-09T17:24:54Z","receivedAt":"2023-10-09T17:25:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> Yes, the bulk-checkin mechanism suffers from an even worse problem which\n>> is the pack it creates will contain no deltas whatsoever. The contents\n>> of the pack are just getting written as-is, so there's no fancy\n>> delta-ficiation going on.\n>\n> I wonder how big a deal this would be in practice for merges.\n> ...\n\nThanks for your experiments ;-)\n\nThe reason why bulk-checkin mechanism does not attempt deltifying\n(as opposed to fast-import that attempts to deltify with the\nimmediately previous object and only with that single object) is\nexactly the same.  It was done to support the initial check-in,\nwhich by definition lacks the delta opportunity along the time axis.\n\nAs you describe, such a delta-less pack would risk missed\ndeltification opportunity when running a repack (without \"-f\"), as\nthe opposite of the well known \"reuse delta\" heuristics, aka \"this\nobject was stored in the base form, it is likely that the previous\npack-object tried but did not find a good delta base for it, let's\nnot waste time retrying that\" heuristics would get in the way.\n\n"},{"id":"482901","messageId":"20231009202149.GA3281325@coredump.intra.peff.net","threadId":"60318","inReplyTo":"ZSNZZrWyCqRH+0Bd@nand.local","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-10-09T20:21:49Z","receivedAt":"2023-10-09T20:21:55Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Oct 08, 2023 at 09:37:42PM -0400, Taylor Blau wrote:\n\n> Very interesting, thanks for running (and documenting!) this experiment.\n> I'm mostly with you that it probably doesn't make a huge difference in\n> practice here.\n> \n> One thing that I'm not entirely clear on is how we'd treat objects that\n> could be good delta candidates for each other between two packs. For\n> instance, if I write a tree corresponding to the merge between two\n> branches, it's likely that the resulting tree would be a good delta\n> candidate against either of the trees at the tips of those two refs.\n> \n> But we won't pack those trees (the ones at the tips of the refs) in the\n> same pack as the tree containing their merge. If we later on tried to\n> repack, would we evaluate the tip trees as possible delta candidates\n> against the merged tree? Or would we look at the merged tree, realize it\n> isn't delta'd with anything, and then not attempt to find any\n> candidates?\n\nWhen we repack (either all-into-one, or in a geometric roll-up), we\nshould consider those trees as candidates. The only deltas we don't\nconsider are:\n\n  - if something is already a delta in a pack, then we will usually\n    reuse that delta verbatim (so you might get fooled by a mediocre\n    delta and not go to the trouble to search again. But I don't think\n    that applies here; there is no other tree in your new pack to make\n    such a mediocre delta from, and anyway you are skipping deltas\n    entirely)\n\n  - if two objects are in the same pack but there's no delta\n    relationship, the try_delta() heuristics will skip them immediately\n    (under the assumption that we tried during the last repack and\n    didn't find anything good).\n\nSo yes, if you had a big old pack with the original trees, and then a\nnew pack with the merge result, we should try to delta the new merge\nresult tree against the others, just as we would if it were loose.\n\n> > I was going to suggest the same thing. ;) Unfortunately it's a bit\n> > tricky to do as we have no room in the file format for an optional flag.\n> > You'd have to add a \".mediocre-delta\" file or something.\n> \n> Yeah, I figured that we'd add a new \".baddeltas\" file or something. (As\n> an aside, we probably should have an optional flags section in the .pack\n> format, since we seem to have a lot of optional pack extensions: .rev,\n> .bitmap, .keep, .promisor, etc.)\n\nYes, though since packv2 is the on-the-wire format, it's very hard to\nchange now. It might be easier to put an annotation into the .idx file.\n\n-Peff\n"},{"id":"482950","messageId":"ZSTw04yxPg3NiBOs@tanuki","threadId":"60318","inReplyTo":"ZSQlgLDv+MrYmSp8@nand.local","subject":"Re: [PATCH 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2023-10-10T06:36:03Z","receivedAt":"2023-10-10T06:36:11Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Oct 09, 2023 at 12:08:32PM -0400, Taylor Blau wrote:\n> On Mon, Oct 09, 2023 at 12:54:12PM +0200, Patrick Steinhardt wrote:\n> > In Gitaly, we usually set up quarantine directories for all operations\n> > that create objects. This allows us to discard any newly written objects\n> > in case either the RPC call gets cancelled or in case our access checks\n> > determine that the change should not be allowed. The logic is rather\n> > simple:\n> >\n> >     1. Create a new temporary directory.\n> >\n> >     2. Set up the new temporary directory as main object database via\n> >        the `GIT_OBJECT_DIRECTORY` environment variable.\n> >\n> >     3. Set up the main repository's object database via the\n> >        `GIT_ALTERNATE_OBJECT_DIRECTORIES` environment variable.\n> \n> Is there a reason not to run Git in the quarantine environment and list\n> the main repository as an alternate via $GIT_DIR/objects/info/alternates\n> instead of the GIT_ALTERNATE_OBJECT_DIRECTORIES environment variable?\n\nThe quarantine environment as we use it is really only a single object\ndatabase, so you cannot run Git inside of it directly.\n\n> >     4. Execute Git commands that write objects with these environment\n> >        variables set up. The new objects will end up neatly contained in\n> >        the temporary directory.\n> >\n> >     5. Once done, either discard the temporary object database or\n> >        migrate objects into the main object daatabase.\n> \n> Interesting. I'm curious why you don't use the builtin tmp_objdir\n> mechanism in Git itself. Do you need to run more than one command in the\n> quarantine environment? If so, that makes sense that you'd want to have\n> a scratch repository that lasts beyond the lifetime of a single process.\n\nIt's a mixture of things:\n\n    - Many commands simply don't set up a temporary object directory.\n\n    - We want to check the result after the objects have been generated.\n      Many of the commands don't provide hooks to do so in a reasonable\n      way. So we want to check the result _after_ the command has exited\n      already, and objects should not yet have been migrated into the\n      target object database at that point.\n\n    - Sometimes we indeed want to run multiple Git commands. We use this\n      e.g. for worktreeless rebases, where we run a succession of\n      commands to rebase every single commit.\n\nSo ultimately, our own quarantine directory sits at a conceptually\nhigher level than what Git commands would be able to provide.\n\n> > I wonder whether this would be a viable approach for you, as well.\n> \n> I think that the main problem that we are trying to solve with this\n> series is creating a potentially large number of loose objects. I think\n> that you could do something like what you propose above, with a 'git\n> repacks -adk' before moving its objects over back to the main repository.\n\nYeah, that's certainly possible. I had the feeling that there are two\ndifferent problems that we're trying to solve in this thread:\n\n    - The problem that the objects are of temporary nature. Regardless\n      of whether they are in a packfile or loose, it doesn't make much\n      sense to put them into the main object database as they would be\n      unreachable anyway and thus get pruned after some time.\n\n    - The problem that writing many loose objects can be slow.\n\nMy proposal only addresses the first problem.\n\n> But since we're working in a single process only when doing a merge-tree\n> operation, I think it is probably more expedient to write the pack's\n> bytes directly.\n\nIf it's about efficiency and not about how we discard those objects\nafter the command has ran to completion then yes, I agree.\n\nPatrick\n"},{"id":"483344","messageId":"cover.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH v2 0/7] merge-ort: implement support for packing objects together","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:08Z","receivedAt":"2023-10-17T16:31:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"(Previously based on 'eb/limit-bulk-checkin-to-blobs', which has since\nbeen merged. This series is now based on the tip of 'master', which is\na9ecda2788 (The eighteenth batch, 2023-10-13) at the time of writing).\n\nThis series implements support for a new merge-tree option,\n`--write-pack`, which causes any newly-written objects to be packed\ntogether instead of being stored individually as loose.\n\nMuch is unchanged since last time, except for a small tweak to one of\nthe commit messages in response to feedback from Eric W. Biederman. The\nseries has also been rebased onto 'master', which had a couple of\nconflicts that I resolved pertaining to:\n\n  - 9eb5419799 (bulk-checkin: only support blobs in index_bulk_checkin,\n    2023-09-26)\n  - e0b8c84240 (treewide: fix various bugs w/ OpenSSL 3+ EVP API,\n    2023-09-01)\n\nThey were mostly trivial resolutions, and the results can be viewed in\nthe range-diff included below.\n\n(From last time: the motivating use-case behind these changes is to\nbetter support repositories who invoke merge-tree frequently, generating\na potentially large number of loose objects, resulting in a possible\nadverse effect on performance.)\n\nThanks in advance for any review!\n\nTaylor Blau (7):\n  bulk-checkin: factor out `format_object_header_hash()`\n  bulk-checkin: factor out `prepare_checkpoint()`\n  bulk-checkin: factor out `truncate_checkpoint()`\n  bulk-checkin: factor our `finalize_checkpoint()`\n  bulk-checkin: introduce `index_blob_bulk_checkin_incore()`\n  bulk-checkin: introduce `index_tree_bulk_checkin_incore()`\n  builtin/merge-tree.c: implement support for `--write-pack`\n\n Documentation/git-merge-tree.txt |   4 +\n builtin/merge-tree.c             |   5 +\n bulk-checkin.c                   | 258 ++++++++++++++++++++++++++-----\n bulk-checkin.h                   |   8 +\n merge-ort.c                      |  42 +++--\n merge-recursive.h                |   1 +\n t/t4301-merge-tree-write-tree.sh |  93 +++++++++++\n 7 files changed, 363 insertions(+), 48 deletions(-)\n\nRange-diff against v1:\n1:  37f4072815 ! 1:  edf1cbafc1 bulk-checkin: factor out `format_object_header_hash()`\n    @@ bulk-checkin.c: static void prepare_to_stream(struct bulk_checkin_packfile *stat\n      }\n      \n     +static void format_object_header_hash(const struct git_hash_algo *algop,\n    -+\t\t\t\t      git_hash_ctx *ctx, enum object_type type,\n    ++\t\t\t\t      git_hash_ctx *ctx,\n    ++\t\t\t\t      struct hashfile_checkpoint *checkpoint,\n    ++\t\t\t\t      enum object_type type,\n     +\t\t\t\t      size_t size)\n     +{\n     +\tunsigned char header[16384];\n    @@ bulk-checkin.c: static void prepare_to_stream(struct bulk_checkin_packfile *stat\n     +\n     +\talgop->init_fn(ctx);\n     +\talgop->update_fn(ctx, header, header_len);\n    ++\talgop->init_fn(&checkpoint->ctx);\n     +}\n     +\n      static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n    @@ bulk-checkin.c: static int deflate_blob_to_pack(struct bulk_checkin_packfile *st\n     -\t\t\t\t\t  OBJ_BLOB, size);\n     -\tthe_hash_algo->init_fn(&ctx);\n     -\tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n    -+\tformat_object_header_hash(the_hash_algo, &ctx, OBJ_BLOB, size);\n    +-\tthe_hash_algo->init_fn(&checkpoint.ctx);\n    ++\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n    ++\t\t\t\t  size);\n      \n      \t/* Note: idx is non-NULL when we are writing */\n      \tif ((flags & HASH_WRITE_OBJECT) != 0)\n2:  9cc1f3014a ! 2:  b3f89d5853 bulk-checkin: factor out `prepare_checkpoint()`\n    @@ Commit message\n     \n      ## bulk-checkin.c ##\n     @@ bulk-checkin.c: static void format_object_header_hash(const struct git_hash_algo *algop,\n    - \talgop->update_fn(ctx, header, header_len);\n    + \talgop->init_fn(&checkpoint->ctx);\n      }\n      \n     +static void prepare_checkpoint(struct bulk_checkin_packfile *state,\n3:  f392ed2211 = 3:  abe4fb0a59 bulk-checkin: factor out `truncate_checkpoint()`\n4:  9c6ca564ad = 4:  0b855a6eb7 bulk-checkin: factor our `finalize_checkpoint()`\n5:  30ca7334c7 ! 5:  239bf39bfb bulk-checkin: introduce `index_blob_bulk_checkin_incore()`\n    @@ bulk-checkin.c: static void finalize_checkpoint(struct bulk_checkin_packfile *st\n      \n     +static int deflate_obj_contents_to_pack_incore(struct bulk_checkin_packfile *state,\n     +\t\t\t\t\t       git_hash_ctx *ctx,\n    ++\t\t\t\t\t       struct hashfile_checkpoint *checkpoint,\n     +\t\t\t\t\t       struct object_id *result_oid,\n     +\t\t\t\t\t       const void *buf, size_t size,\n     +\t\t\t\t\t       enum object_type type,\n     +\t\t\t\t\t       const char *path, unsigned flags)\n     +{\n    -+\tstruct hashfile_checkpoint checkpoint = {0};\n     +\tstruct pack_idx_entry *idx = NULL;\n     +\toff_t already_hashed_to = 0;\n     +\n    @@ bulk-checkin.c: static void finalize_checkpoint(struct bulk_checkin_packfile *st\n     +\t\tCALLOC_ARRAY(idx, 1);\n     +\n     +\twhile (1) {\n    -+\t\tprepare_checkpoint(state, &checkpoint, idx, flags);\n    ++\t\tprepare_checkpoint(state, checkpoint, idx, flags);\n     +\t\tif (!stream_obj_to_pack_incore(state, ctx, &already_hashed_to,\n     +\t\t\t\t\t       buf, size, type, path, flags))\n     +\t\t\tbreak;\n    -+\t\ttruncate_checkpoint(state, &checkpoint, idx);\n    ++\t\ttruncate_checkpoint(state, checkpoint, idx);\n     +\t}\n     +\n    -+\tfinalize_checkpoint(state, ctx, &checkpoint, idx, result_oid);\n    ++\tfinalize_checkpoint(state, ctx, checkpoint, idx, result_oid);\n     +\n     +\treturn 0;\n     +}\n    @@ bulk-checkin.c: static void finalize_checkpoint(struct bulk_checkin_packfile *st\n     +\t\t\t\t       const char *path, unsigned flags)\n     +{\n     +\tgit_hash_ctx ctx;\n    ++\tstruct hashfile_checkpoint checkpoint = {0};\n     +\n    -+\tformat_object_header_hash(the_hash_algo, &ctx, OBJ_BLOB, size);\n    ++\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n    ++\t\t\t\t  size);\n     +\n    -+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, result_oid,\n    -+\t\t\t\t\t\t   buf, size, OBJ_BLOB, path,\n    -+\t\t\t\t\t\t   flags);\n    ++\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n    ++\t\t\t\t\t\t   result_oid, buf, size,\n    ++\t\t\t\t\t\t   OBJ_BLOB, path, flags);\n     +}\n     +\n      static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n6:  cb0f79cabb ! 6:  57613807d8 bulk-checkin: introduce `index_tree_bulk_checkin_incore()`\n    @@ Commit message\n         Within `deflate_tree_to_pack_incore()`, the changes should be limited\n         to something like:\n     \n    +        struct strbuf converted = STRBUF_INIT;\n             if (the_repository->compat_hash_algo) {\n    -          struct strbuf converted = STRBUF_INIT;\n               if (convert_object_file(&compat_obj,\n                                       the_repository->hash_algo,\n                                       the_repository->compat_hash_algo, ...) < 0)\n    @@ Commit message\n     \n               format_object_header_hash(the_repository->compat_hash_algo,\n                                         OBJ_TREE, size);\n    -\n    -          strbuf_release(&converted);\n             }\n    +        /* compute the converted tree's hash using the compat algorithm */\n    +        strbuf_release(&converted);\n     \n         , assuming related changes throughout the rest of the bulk-checkin\n         machinery necessary to update the hash of the converted object, which\n    @@ Commit message\n     \n      ## bulk-checkin.c ##\n     @@ bulk-checkin.c: static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n    - \t\t\t\t\t\t   flags);\n    + \t\t\t\t\t\t   OBJ_BLOB, path, flags);\n      }\n      \n     +static int deflate_tree_to_pack_incore(struct bulk_checkin_packfile *state,\n    @@ bulk-checkin.c: static int deflate_blob_to_pack_incore(struct bulk_checkin_packf\n     +\t\t\t\t       const char *path, unsigned flags)\n     +{\n     +\tgit_hash_ctx ctx;\n    ++\tstruct hashfile_checkpoint checkpoint = {0};\n     +\n    -+\tformat_object_header_hash(the_hash_algo, &ctx, OBJ_TREE, size);\n    ++\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_TREE,\n    ++\t\t\t\t  size);\n     +\n    -+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, result_oid,\n    -+\t\t\t\t\t\t   buf, size, OBJ_TREE, path,\n    -+\t\t\t\t\t\t   flags);\n    ++\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n    ++\t\t\t\t\t\t   result_oid, buf, size,\n    ++\t\t\t\t\t\t   OBJ_TREE, path, flags);\n     +}\n     +\n      static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n7:  e969210145 ! 7:  f21400f56c builtin/merge-tree.c: implement support for `--write-pack`\n    @@ merge-ort.c\n       * We have many arrays of size 3.  Whenever we have such an array, the\n     @@ merge-ort.c: static int handle_content_merge(struct merge_options *opt,\n      \t\tif ((merge_status < 0) || !result_buf.ptr)\n    - \t\t\tret = err(opt, _(\"Failed to execute internal merge\"));\n    + \t\t\tret = error(_(\"failed to execute internal merge\"));\n      \n     -\t\tif (!ret &&\n     -\t\t    write_object_file(result_buf.ptr, result_buf.size,\n     -\t\t\t\t      OBJ_BLOB, &result->oid))\n    --\t\t\tret = err(opt, _(\"Unable to add %s to database\"),\n    --\t\t\t\t  path);\n    +-\t\t\tret = error(_(\"unable to add %s to database\"), path);\n     +\t\tif (!ret) {\n     +\t\t\tret = opt->write_pack\n     +\t\t\t\t? index_blob_bulk_checkin_incore(&result->oid,\n    @@ merge-ort.c: static int handle_content_merge(struct merge_options *opt,\n     +\t\t\t\t\t\t    result_buf.size,\n     +\t\t\t\t\t\t    OBJ_BLOB, &result->oid);\n     +\t\t\tif (ret)\n    -+\t\t\t\tret = err(opt, _(\"Unable to add %s to database\"),\n    -+\t\t\t\t\t  path);\n    ++\t\t\t\tret = error(_(\"unable to add %s to database\"),\n    ++\t\t\t\t\t    path);\n     +\t\t}\n      \n      \t\tfree(result_buf.ptr);\n-- \n2.42.0.405.gdb2a2f287e\n"},{"id":"483345","messageId":"edf1cbafc166bce619541cadc1d6fc2c1d9b84d6.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697560266.git.me@ttaylorr.com","subject":"[PATCH v2 1/7] bulk-checkin: factor out `format_object_header_hash()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:12Z","receivedAt":"2023-10-17T16:31:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Before deflating a blob into a pack, the bulk-checkin mechanism prepares\nthe pack object header by calling `format_object_header()`, and writing\ninto a scratch buffer, the contents of which eventually makes its way\ninto the pack.\n\nFuture commits will add support for deflating multiple kinds of objects\ninto a pack, and will likewise need to perform a similar operation as\nbelow.\n\nThis is a mostly straightforward extraction, with one notable exception.\nInstead of hard-coding `the_hash_algo`, pass it in to the new function\nas an argument. This isn't strictly necessary for our immediate purposes\nhere, but will prove useful in the future if/when the bulk-checkin\nmechanism grows support for the hash transition plan.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 25 ++++++++++++++++++-------\n 1 file changed, 18 insertions(+), 7 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6ce62999e5..fd3c110d1c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -247,6 +247,22 @@ static void prepare_to_stream(struct bulk_checkin_packfile *state,\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n+static void format_object_header_hash(const struct git_hash_algo *algop,\n+\t\t\t\t      git_hash_ctx *ctx,\n+\t\t\t\t      struct hashfile_checkpoint *checkpoint,\n+\t\t\t\t      enum object_type type,\n+\t\t\t\t      size_t size)\n+{\n+\tunsigned char header[16384];\n+\tunsigned header_len = format_object_header((char *)header,\n+\t\t\t\t\t\t   sizeof(header),\n+\t\t\t\t\t\t   type, size);\n+\n+\talgop->init_fn(ctx);\n+\talgop->update_fn(ctx, header, header_len);\n+\talgop->init_fn(&checkpoint->ctx);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -254,8 +270,6 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n {\n \toff_t seekback, already_hashed_to;\n \tgit_hash_ctx ctx;\n-\tunsigned char obuf[16384];\n-\tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint = {0};\n \tstruct pack_idx_entry *idx = NULL;\n \n@@ -263,11 +277,8 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \tif (seekback == (off_t) -1)\n \t\treturn error(\"cannot find the current offset\");\n \n-\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n-\t\t\t\t\t  OBJ_BLOB, size);\n-\tthe_hash_algo->init_fn(&ctx);\n-\tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n-\tthe_hash_algo->init_fn(&checkpoint.ctx);\n+\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n+\t\t\t\t  size);\n \n \t/* Note: idx is non-NULL when we are writing */\n \tif ((flags & HASH_WRITE_OBJECT) != 0)\n-- \n2.42.0.405.gdb2a2f287e\n\n"},{"id":"483346","messageId":"b3f89d5853174d2bbf694f6a3b89dc700575cf24.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697560266.git.me@ttaylorr.com","subject":"[PATCH v2 2/7] bulk-checkin: factor out `prepare_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:15Z","receivedAt":"2023-10-17T16:31:19Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as the previous commit, factor out the routine to\nprepare streaming into a bulk-checkin pack into its own function. Unlike\nthe previous patch, this is a verbatim copy and paste.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 20 ++++++++++++++------\n 1 file changed, 14 insertions(+), 6 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex fd3c110d1c..c1f5450583 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -263,6 +263,19 @@ static void format_object_header_hash(const struct git_hash_algo *algop,\n \talgop->init_fn(&checkpoint->ctx);\n }\n \n+static void prepare_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t       struct hashfile_checkpoint *checkpoint,\n+\t\t\t       struct pack_idx_entry *idx,\n+\t\t\t       unsigned flags)\n+{\n+\tprepare_to_stream(state, flags);\n+\tif (idx) {\n+\t\thashfile_checkpoint(state->f, checkpoint);\n+\t\tidx->offset = state->offset;\n+\t\tcrc32_begin(state->f);\n+\t}\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -287,12 +300,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \talready_hashed_to = 0;\n \n \twhile (1) {\n-\t\tprepare_to_stream(state, flags);\n-\t\tif (idx) {\n-\t\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\t\tidx->offset = state->offset;\n-\t\t\tcrc32_begin(state->f);\n-\t\t}\n+\t\tprepare_checkpoint(state, &checkpoint, idx, flags);\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n \t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n-- \n2.42.0.405.gdb2a2f287e\n\n"},{"id":"483347","messageId":"abe4fb0a59468c95db1003d87f8f26ba937c2e4e.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697560266.git.me@ttaylorr.com","subject":"[PATCH v2 3/7] bulk-checkin: factor out `truncate_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:19Z","receivedAt":"2023-10-17T16:31:24Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as previous commits, factor our the routine to\ntruncate a bulk-checkin packfile when writing past the pack size limit.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 27 +++++++++++++++++----------\n 1 file changed, 17 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex c1f5450583..b92d7a6f5a 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -276,6 +276,22 @@ static void prepare_checkpoint(struct bulk_checkin_packfile *state,\n \t}\n }\n \n+static void truncate_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t\tstruct hashfile_checkpoint *checkpoint,\n+\t\t\t\tstruct pack_idx_entry *idx)\n+{\n+\t/*\n+\t * Writing this object to the current pack will make\n+\t * it too big; we need to truncate it, start a new\n+\t * pack, and write into it.\n+\t */\n+\tif (!idx)\n+\t\tBUG(\"should not happen\");\n+\thashfile_truncate(state->f, checkpoint);\n+\tstate->offset = checkpoint->offset;\n+\tflush_bulk_checkin_packfile(state);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -304,16 +320,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n \t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n-\t\t/*\n-\t\t * Writing this object to the current pack will make\n-\t\t * it too big; we need to truncate it, start a new\n-\t\t * pack, and write into it.\n-\t\t */\n-\t\tif (!idx)\n-\t\t\tBUG(\"should not happen\");\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tflush_bulk_checkin_packfile(state);\n+\t\ttruncate_checkpoint(state, &checkpoint, idx);\n \t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n-- \n2.42.0.405.gdb2a2f287e\n\n"},{"id":"483348","messageId":"0b855a6eb7f147a9fc4c41dd183b768162345220.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697560266.git.me@ttaylorr.com","subject":"[PATCH v2 4/7] bulk-checkin: factor our `finalize_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:22Z","receivedAt":"2023-10-17T16:31:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as previous commits, factor out the routine to\nfinalize the just-written object from the bulk-checkin mechanism.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 41 +++++++++++++++++++++++++----------------\n 1 file changed, 25 insertions(+), 16 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b92d7a6f5a..f4914fb6d1 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -292,6 +292,30 @@ static void truncate_checkpoint(struct bulk_checkin_packfile *state,\n \tflush_bulk_checkin_packfile(state);\n }\n \n+static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t\tgit_hash_ctx *ctx,\n+\t\t\t\tstruct hashfile_checkpoint *checkpoint,\n+\t\t\t\tstruct pack_idx_entry *idx,\n+\t\t\t\tstruct object_id *result_oid)\n+{\n+\tthe_hash_algo->final_oid_fn(result_oid, ctx);\n+\tif (!idx)\n+\t\treturn;\n+\n+\tidx->crc32 = crc32_end(state->f);\n+\tif (already_written(state, result_oid)) {\n+\t\thashfile_truncate(state->f, checkpoint);\n+\t\tstate->offset = checkpoint->offset;\n+\t\tfree(idx);\n+\t} else {\n+\t\toidcpy(&idx->oid, result_oid);\n+\t\tALLOC_GROW(state->written,\n+\t\t\t   state->nr_written + 1,\n+\t\t\t   state->alloc_written);\n+\t\tstate->written[state->nr_written++] = idx;\n+\t}\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -324,22 +348,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n-\tthe_hash_algo->final_oid_fn(result_oid, &ctx);\n-\tif (!idx)\n-\t\treturn 0;\n-\n-\tidx->crc32 = crc32_end(state->f);\n-\tif (already_written(state, result_oid)) {\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tfree(idx);\n-\t} else {\n-\t\toidcpy(&idx->oid, result_oid);\n-\t\tALLOC_GROW(state->written,\n-\t\t\t   state->nr_written + 1,\n-\t\t\t   state->alloc_written);\n-\t\tstate->written[state->nr_written++] = idx;\n-\t}\n+\tfinalize_checkpoint(state, &ctx, &checkpoint, idx, result_oid);\n \treturn 0;\n }\n \n-- \n2.42.0.405.gdb2a2f287e\n\n"},{"id":"483349","messageId":"239bf39bfb21ef621a15839bade34446dcbc3103.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697560266.git.me@ttaylorr.com","subject":"[PATCH v2 5/7] bulk-checkin: introduce `index_blob_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:26Z","receivedAt":"2023-10-17T16:31:30Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that we have factored out many of the common routines necessary to\nindex a new object into a pack created by the bulk-checkin machinery, we\ncan introduce a variant of `index_blob_bulk_checkin()` that acts on\nblobs whose contents we can fit in memory.\n\nThis will be useful in a couple of more commits in order to provide the\n`merge-tree` builtin with a mechanism to create a new pack containing\nany objects it created during the merge, instead of storing those\nobjects individually as loose.\n\nSimilar to the existing `index_blob_bulk_checkin()` function, the\nentrypoint delegates to `deflate_blob_to_pack_incore()`, which is\nresponsible for formatting the pack header and then deflating the\ncontents into the pack. The latter is accomplished by calling\ndeflate_blob_contents_to_pack_incore(), which takes advantage of the\nearlier refactoring and is responsible for writing the object to the\npack and handling any overage from pack.packSizeLimit.\n\nThe bulk of the new functionality is implemented in the function\n`stream_obj_to_pack_incore()`, which is a generic implementation for\nwriting objects of arbitrary type (whose contents we can fit in-core)\ninto a bulk-checkin pack.\n\nThe new function shares an unfortunate degree of similarity to the\nexisting `stream_blob_to_pack()` function. But DRY-ing up these two\nwould likely be more trouble than it's worth, since the latter has to\ndeal with reading and writing the contents of the object.\n\nConsistent with the rest of the bulk-checkin mechanism, there are no\ndirect tests here. In future commits when we expose this new\nfunctionality via the `merge-tree` builtin, we will test it indirectly\nthere.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 118 +++++++++++++++++++++++++++++++++++++++++++++++++\n bulk-checkin.h |   4 ++\n 2 files changed, 122 insertions(+)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f4914fb6d1..25cd1ffa25 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -140,6 +140,69 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n \treturn 0;\n }\n \n+static int stream_obj_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t     git_hash_ctx *ctx,\n+\t\t\t\t     off_t *already_hashed_to,\n+\t\t\t\t     const void *buf, size_t size,\n+\t\t\t\t     enum object_type type,\n+\t\t\t\t     const char *path, unsigned flags)\n+{\n+\tgit_zstream s;\n+\tunsigned char obuf[16384];\n+\tunsigned hdrlen;\n+\tint status = Z_OK;\n+\tint write_object = (flags & HASH_WRITE_OBJECT);\n+\n+\tgit_deflate_init(&s, pack_compression_level);\n+\n+\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), type, size);\n+\ts.next_out = obuf + hdrlen;\n+\ts.avail_out = sizeof(obuf) - hdrlen;\n+\n+\tif (*already_hashed_to < size) {\n+\t\tsize_t hsize = size - *already_hashed_to;\n+\t\tif (hsize) {\n+\t\t\tthe_hash_algo->update_fn(ctx, buf, hsize);\n+\t\t}\n+\t\t*already_hashed_to = size;\n+\t}\n+\ts.next_in = (void *)buf;\n+\ts.avail_in = size;\n+\n+\twhile (status != Z_STREAM_END) {\n+\t\tstatus = git_deflate(&s, Z_FINISH);\n+\t\tif (!s.avail_out || status == Z_STREAM_END) {\n+\t\t\tif (write_object) {\n+\t\t\t\tsize_t written = s.next_out - obuf;\n+\n+\t\t\t\t/* would we bust the size limit? */\n+\t\t\t\tif (state->nr_written &&\n+\t\t\t\t    pack_size_limit_cfg &&\n+\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n+\t\t\t\t\tgit_deflate_abort(&s);\n+\t\t\t\t\treturn -1;\n+\t\t\t\t}\n+\n+\t\t\t\thashwrite(state->f, obuf, written);\n+\t\t\t\tstate->offset += written;\n+\t\t\t}\n+\t\t\ts.next_out = obuf;\n+\t\t\ts.avail_out = sizeof(obuf);\n+\t\t}\n+\n+\t\tswitch (status) {\n+\t\tcase Z_OK:\n+\t\tcase Z_BUF_ERROR:\n+\t\tcase Z_STREAM_END:\n+\t\t\tcontinue;\n+\t\tdefault:\n+\t\t\tdie(\"unexpected deflate failure: %d\", status);\n+\t\t}\n+\t}\n+\tgit_deflate_end(&s);\n+\treturn 0;\n+}\n+\n /*\n  * Read the contents from fd for size bytes, streaming it to the\n  * packfile in state while updating the hash in ctx. Signal a failure\n@@ -316,6 +379,50 @@ static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n \t}\n }\n \n+static int deflate_obj_contents_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t\t       git_hash_ctx *ctx,\n+\t\t\t\t\t       struct hashfile_checkpoint *checkpoint,\n+\t\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t\t       enum object_type type,\n+\t\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tstruct pack_idx_entry *idx = NULL;\n+\toff_t already_hashed_to = 0;\n+\n+\t/* Note: idx is non-NULL when we are writing */\n+\tif (flags & HASH_WRITE_OBJECT)\n+\t\tCALLOC_ARRAY(idx, 1);\n+\n+\twhile (1) {\n+\t\tprepare_checkpoint(state, checkpoint, idx, flags);\n+\t\tif (!stream_obj_to_pack_incore(state, ctx, &already_hashed_to,\n+\t\t\t\t\t       buf, size, type, path, flags))\n+\t\t\tbreak;\n+\t\ttruncate_checkpoint(state, checkpoint, idx);\n+\t}\n+\n+\tfinalize_checkpoint(state, ctx, checkpoint, idx, result_oid);\n+\n+\treturn 0;\n+}\n+\n+static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tgit_hash_ctx ctx;\n+\tstruct hashfile_checkpoint checkpoint = {0};\n+\n+\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n+\t\t\t\t  size);\n+\n+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n+\t\t\t\t\t\t   result_oid, buf, size,\n+\t\t\t\t\t\t   OBJ_BLOB, path, flags);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -396,6 +503,17 @@ int index_blob_bulk_checkin(struct object_id *oid,\n \treturn status;\n }\n \n+int index_blob_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags)\n+{\n+\tint status = deflate_blob_to_pack_incore(&bulk_checkin_packfile, oid,\n+\t\t\t\t\t\t buf, size, path, flags);\n+\tif (!odb_transaction_nesting)\n+\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n+\treturn status;\n+}\n+\n void begin_odb_transaction(void)\n {\n \todb_transaction_nesting += 1;\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex aa7286a7b3..1b91daeaee 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -13,6 +13,10 @@ int index_blob_bulk_checkin(struct object_id *oid,\n \t\t\t    int fd, size_t size,\n \t\t\t    const char *path, unsigned flags);\n \n+int index_blob_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags);\n+\n /*\n  * Tell the object database to optimize for adding\n  * multiple objects. end_odb_transaction must be called\n-- \n2.42.0.405.gdb2a2f287e\n\n"},{"id":"483350","messageId":"57613807d84d19fa2691fcf7fe81c4aa9a575d4b.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697560266.git.me@ttaylorr.com","subject":"[PATCH v2 6/7] bulk-checkin: introduce `index_tree_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:29Z","receivedAt":"2023-10-17T16:31:34Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The remaining missing piece in order to teach the `merge-tree` builtin\nhow to write the contents of a merge into a pack is a function to index\ntree objects into a bulk-checkin pack.\n\nThis patch implements that missing piece, which is a thin wrapper around\nall of the functionality introduced in previous commits.\n\nIf and when Git gains support for a \"compatibility\" hash algorithm, the\nchanges to support that here will be minimal. The bulk-checkin machinery\nwill need to convert the incoming tree to compute its length under the\ncompatibility hash, necessary to reconstruct its header. With that\ninformation (and the converted contents of the tree), the bulk-checkin\nmachinery will have enough to keep track of the converted object's hash\nin order to update the compatibility mapping.\n\nWithin `deflate_tree_to_pack_incore()`, the changes should be limited\nto something like:\n\n    struct strbuf converted = STRBUF_INIT;\n    if (the_repository->compat_hash_algo) {\n      if (convert_object_file(&compat_obj,\n                              the_repository->hash_algo,\n                              the_repository->compat_hash_algo, ...) < 0)\n        die(...);\n\n      format_object_header_hash(the_repository->compat_hash_algo,\n                                OBJ_TREE, size);\n    }\n    /* compute the converted tree's hash using the compat algorithm */\n    strbuf_release(&converted);\n\n, assuming related changes throughout the rest of the bulk-checkin\nmachinery necessary to update the hash of the converted object, which\nare likewise minimal in size.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 27 +++++++++++++++++++++++++++\n bulk-checkin.h |  4 ++++\n 2 files changed, 31 insertions(+)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 25cd1ffa25..fe13100e04 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -423,6 +423,22 @@ static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n \t\t\t\t\t\t   OBJ_BLOB, path, flags);\n }\n \n+static int deflate_tree_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tgit_hash_ctx ctx;\n+\tstruct hashfile_checkpoint checkpoint = {0};\n+\n+\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_TREE,\n+\t\t\t\t  size);\n+\n+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n+\t\t\t\t\t\t   result_oid, buf, size,\n+\t\t\t\t\t\t   OBJ_TREE, path, flags);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -514,6 +530,17 @@ int index_blob_bulk_checkin_incore(struct object_id *oid,\n \treturn status;\n }\n \n+int index_tree_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags)\n+{\n+\tint status = deflate_tree_to_pack_incore(&bulk_checkin_packfile, oid,\n+\t\t\t\t\t\t buf, size, path, flags);\n+\tif (!odb_transaction_nesting)\n+\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n+\treturn status;\n+}\n+\n void begin_odb_transaction(void)\n {\n \todb_transaction_nesting += 1;\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 1b91daeaee..89786b3954 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -17,6 +17,10 @@ int index_blob_bulk_checkin_incore(struct object_id *oid,\n \t\t\t\t   const void *buf, size_t size,\n \t\t\t\t   const char *path, unsigned flags);\n \n+int index_tree_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags);\n+\n /*\n  * Tell the object database to optimize for adding\n  * multiple objects. end_odb_transaction must be called\n-- \n2.42.0.405.gdb2a2f287e\n\n"},{"id":"483351","messageId":"f21400f56c6e32cd4a0d760ba6524c2e56d0d5ab.1697560266.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697560266.git.me@ttaylorr.com","subject":"[PATCH v2 7/7] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-17T16:31:32Z","receivedAt":"2023-10-17T16:31:37Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using merge-tree often within a repository[^1], it is possible to\ngenerate a relatively large number of loose objects, which can result in\ndegraded performance, and inode exhaustion in extreme cases.\n\nBuilding on the functionality introduced in previous commits, the\nbulk-checkin machinery now has support to write arbitrary blob and tree\nobjects which are small enough to be held in-core. We can use this to\nwrite any blob/tree objects generated by ORT into a separate pack\ninstead of writing them out individually as loose.\n\nThis functionality is gated behind a new `--write-pack` option to\n`merge-tree` that works with the (non-deprecated) `--write-tree` mode.\n\nThe implementation is relatively straightforward. There are two spots\nwithin the ORT mechanism where we call `write_object_file()`, one for\ncontent differences within blobs, and another to assemble any new trees\nnecessary to construct the merge. In each of those locations,\nconditionally replace calls to `write_object_file()` with\n`index_blob_bulk_checkin_incore()` or `index_tree_bulk_checkin_incore()`\ndepending on which kind of object we are writing.\n\nThe only remaining task is to begin and end the transaction necessary to\ninitialize the bulk-checkin machinery, and move any new pack(s) it\ncreated into the main object store.\n\n[^1]: Such is the case at GitHub, where we run presumptive \"test merges\"\n  on open pull requests to see whether or not we can light up the merge\n  button green depending on whether or not the presumptive merge was\n  conflicted.\n\n  This is done in response to a number of user-initiated events,\n  including viewing an open pull request whose last test merge is stale\n  with respect to the current base and tip of the pull request. As a\n  result, merge-tree can be run very frequently on large, active\n  repositories.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-merge-tree.txt |  4 ++\n builtin/merge-tree.c             |  5 ++\n merge-ort.c                      | 42 +++++++++++----\n merge-recursive.h                |  1 +\n t/t4301-merge-tree-write-tree.sh | 93 ++++++++++++++++++++++++++++++++\n 5 files changed, 136 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/git-merge-tree.txt b/Documentation/git-merge-tree.txt\nindex ffc4fbf7e8..9d37609ef1 100644\n--- a/Documentation/git-merge-tree.txt\n+++ b/Documentation/git-merge-tree.txt\n@@ -69,6 +69,10 @@ OPTIONS\n \tspecify a merge-base for the merge, and specifying multiple bases is\n \tcurrently not supported. This option is incompatible with `--stdin`.\n \n+--write-pack::\n+\tWrite any new objects into a separate packfile instead of as\n+\tindividual loose objects.\n+\n [[OUTPUT]]\n OUTPUT\n ------\ndiff --git a/builtin/merge-tree.c b/builtin/merge-tree.c\nindex 0de42aecf4..672ebd4c54 100644\n--- a/builtin/merge-tree.c\n+++ b/builtin/merge-tree.c\n@@ -18,6 +18,7 @@\n #include \"quote.h\"\n #include \"tree.h\"\n #include \"config.h\"\n+#include \"bulk-checkin.h\"\n \n static int line_termination = '\\n';\n \n@@ -414,6 +415,7 @@ struct merge_tree_options {\n \tint show_messages;\n \tint name_only;\n \tint use_stdin;\n+\tint write_pack;\n };\n \n static int real_merge(struct merge_tree_options *o,\n@@ -440,6 +442,7 @@ static int real_merge(struct merge_tree_options *o,\n \tinit_merge_options(&opt, the_repository);\n \n \topt.show_rename_progress = 0;\n+\topt.write_pack = o->write_pack;\n \n \topt.branch1 = branch1;\n \topt.branch2 = branch2;\n@@ -548,6 +551,8 @@ int cmd_merge_tree(int argc, const char **argv, const char *prefix)\n \t\t\t   &merge_base,\n \t\t\t   N_(\"commit\"),\n \t\t\t   N_(\"specify a merge-base for the merge\")),\n+\t\tOPT_BOOL(0, \"write-pack\", &o.write_pack,\n+\t\t\t N_(\"write new objects to a pack instead of as loose\")),\n \t\tOPT_END()\n \t};\n \ndiff --git a/merge-ort.c b/merge-ort.c\nindex 7857ce9fbd..e198d2bc2b 100644\n--- a/merge-ort.c\n+++ b/merge-ort.c\n@@ -48,6 +48,7 @@\n #include \"tree.h\"\n #include \"unpack-trees.h\"\n #include \"xdiff-interface.h\"\n+#include \"bulk-checkin.h\"\n \n /*\n  * We have many arrays of size 3.  Whenever we have such an array, the\n@@ -2107,10 +2108,19 @@ static int handle_content_merge(struct merge_options *opt,\n \t\tif ((merge_status < 0) || !result_buf.ptr)\n \t\t\tret = error(_(\"failed to execute internal merge\"));\n \n-\t\tif (!ret &&\n-\t\t    write_object_file(result_buf.ptr, result_buf.size,\n-\t\t\t\t      OBJ_BLOB, &result->oid))\n-\t\t\tret = error(_(\"unable to add %s to database\"), path);\n+\t\tif (!ret) {\n+\t\t\tret = opt->write_pack\n+\t\t\t\t? index_blob_bulk_checkin_incore(&result->oid,\n+\t\t\t\t\t\t\t\t result_buf.ptr,\n+\t\t\t\t\t\t\t\t result_buf.size,\n+\t\t\t\t\t\t\t\t path, 1)\n+\t\t\t\t: write_object_file(result_buf.ptr,\n+\t\t\t\t\t\t    result_buf.size,\n+\t\t\t\t\t\t    OBJ_BLOB, &result->oid);\n+\t\t\tif (ret)\n+\t\t\t\tret = error(_(\"unable to add %s to database\"),\n+\t\t\t\t\t    path);\n+\t\t}\n \n \t\tfree(result_buf.ptr);\n \t\tif (ret)\n@@ -3596,7 +3606,8 @@ static int tree_entry_order(const void *a_, const void *b_)\n \t\t\t\t b->string, strlen(b->string), bmi->result.mode);\n }\n \n-static int write_tree(struct object_id *result_oid,\n+static int write_tree(struct merge_options *opt,\n+\t\t      struct object_id *result_oid,\n \t\t      struct string_list *versions,\n \t\t      unsigned int offset,\n \t\t      size_t hash_size)\n@@ -3630,8 +3641,14 @@ static int write_tree(struct object_id *result_oid,\n \t}\n \n \t/* Write this object file out, and record in result_oid */\n-\tif (write_object_file(buf.buf, buf.len, OBJ_TREE, result_oid))\n+\tret = opt->write_pack\n+\t\t? index_tree_bulk_checkin_incore(result_oid,\n+\t\t\t\t\t\t buf.buf, buf.len, \"\", 1)\n+\t\t: write_object_file(buf.buf, buf.len, OBJ_TREE, result_oid);\n+\n+\tif (ret)\n \t\tret = -1;\n+\n \tstrbuf_release(&buf);\n \treturn ret;\n }\n@@ -3796,8 +3813,8 @@ static int write_completed_directory(struct merge_options *opt,\n \t\t */\n \t\tdir_info->is_null = 0;\n \t\tdir_info->result.mode = S_IFDIR;\n-\t\tif (write_tree(&dir_info->result.oid, &info->versions, offset,\n-\t\t\t       opt->repo->hash_algo->rawsz) < 0)\n+\t\tif (write_tree(opt, &dir_info->result.oid, &info->versions,\n+\t\t\t       offset, opt->repo->hash_algo->rawsz) < 0)\n \t\t\tret = -1;\n \t}\n \n@@ -4331,9 +4348,13 @@ static int process_entries(struct merge_options *opt,\n \t\tfflush(stdout);\n \t\tBUG(\"dir_metadata accounting completely off; shouldn't happen\");\n \t}\n-\tif (write_tree(result_oid, &dir_metadata.versions, 0,\n+\tif (write_tree(opt, result_oid, &dir_metadata.versions, 0,\n \t\t       opt->repo->hash_algo->rawsz) < 0)\n \t\tret = -1;\n+\n+\tif (opt->write_pack)\n+\t\tend_odb_transaction();\n+\n cleanup:\n \tstring_list_clear(&plist, 0);\n \tstring_list_clear(&dir_metadata.versions, 0);\n@@ -4877,6 +4898,9 @@ static void merge_start(struct merge_options *opt, struct merge_result *result)\n \t */\n \tstrmap_init(&opt->priv->conflicts);\n \n+\tif (opt->write_pack)\n+\t\tbegin_odb_transaction();\n+\n \ttrace2_region_leave(\"merge\", \"allocate/init\", opt->repo);\n }\n \ndiff --git a/merge-recursive.h b/merge-recursive.h\nindex b88000e3c2..156e160876 100644\n--- a/merge-recursive.h\n+++ b/merge-recursive.h\n@@ -48,6 +48,7 @@ struct merge_options {\n \tunsigned renormalize : 1;\n \tunsigned record_conflict_msgs_as_headers : 1;\n \tconst char *msg_header_prefix;\n+\tunsigned write_pack : 1;\n \n \t/* internal fields used by the implementation */\n \tstruct merge_options_internal *priv;\ndiff --git a/t/t4301-merge-tree-write-tree.sh b/t/t4301-merge-tree-write-tree.sh\nindex 250f721795..2d81ff4de5 100755\n--- a/t/t4301-merge-tree-write-tree.sh\n+++ b/t/t4301-merge-tree-write-tree.sh\n@@ -922,4 +922,97 @@ test_expect_success 'check the input format when --stdin is passed' '\n \ttest_cmp expect actual\n '\n \n+packdir=\".git/objects/pack\"\n+\n+test_expect_success 'merge-tree can pack its result with --write-pack' '\n+\ttest_when_finished \"rm -rf repo\" &&\n+\tgit init repo &&\n+\n+\t# base has lines [3, 4, 5]\n+\t#   - side adds to the beginning, resulting in [1, 2, 3, 4, 5]\n+\t#   - other adds to the end, resulting in [3, 4, 5, 6, 7]\n+\t#\n+\t# merging the two should result in a new blob object containing\n+\t# [1, 2, 3, 4, 5, 6, 7], along with a new tree.\n+\ttest_commit -C repo base file \"$(test_seq 3 5)\" &&\n+\tgit -C repo branch -M main &&\n+\tgit -C repo checkout -b side main &&\n+\ttest_commit -C repo side file \"$(test_seq 1 5)\" &&\n+\tgit -C repo checkout -b other main &&\n+\ttest_commit -C repo other file \"$(test_seq 3 7)\" &&\n+\n+\tfind repo/$packdir -type f -name \"pack-*.idx\" >packs.before &&\n+\ttree=\"$(git -C repo merge-tree --write-pack \\\n+\t\trefs/tags/side refs/tags/other)\" &&\n+\tblob=\"$(git -C repo rev-parse $tree:file)\" &&\n+\tfind repo/$packdir -type f -name \"pack-*.idx\" >packs.after &&\n+\n+\ttest_must_be_empty packs.before &&\n+\ttest_line_count = 1 packs.after &&\n+\n+\tgit show-index <$(cat packs.after) >objects &&\n+\ttest_line_count = 2 objects &&\n+\tgrep \"^[1-9][0-9]* $tree\" objects &&\n+\tgrep \"^[1-9][0-9]* $blob\" objects\n+'\n+\n+test_expect_success 'merge-tree can write multiple packs with --write-pack' '\n+\ttest_when_finished \"rm -rf repo\" &&\n+\tgit init repo &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\tgit config pack.packSizeLimit 512 &&\n+\n+\t\ttest_seq 512 >f &&\n+\n+\t\t# \"f\" contains roughly ~2,000 bytes.\n+\t\t#\n+\t\t# Each side (\"foo\" and \"bar\") adds a small amount of data at the\n+\t\t# beginning and end of \"base\", respectively.\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m base &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout -b foo main &&\n+\t\t{\n+\t\t\techo foo && cat f\n+\t\t} >f.tmp &&\n+\t\tmv f.tmp f &&\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m foo &&\n+\n+\t\tgit checkout -b bar main &&\n+\t\techo bar >>f &&\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m bar &&\n+\n+\t\tfind $packdir -type f -name \"pack-*.idx\" >packs.before &&\n+\t\t# Merging either side should result in a new object which is\n+\t\t# larger than 1M, thus the result should be split into two\n+\t\t# separate packs.\n+\t\ttree=\"$(git merge-tree --write-pack \\\n+\t\t\trefs/heads/foo refs/heads/bar)\" &&\n+\t\tblob=\"$(git rev-parse $tree:f)\" &&\n+\t\tfind $packdir -type f -name \"pack-*.idx\" >packs.after &&\n+\n+\t\ttest_must_be_empty packs.before &&\n+\t\ttest_line_count = 2 packs.after &&\n+\t\tfor idx in $(cat packs.after)\n+\t\tdo\n+\t\t\tgit show-index <$idx || return 1\n+\t\tdone >objects &&\n+\n+\t\t# The resulting set of packs should contain one copy of both\n+\t\t# objects, each in a separate pack.\n+\t\ttest_line_count = 2 objects &&\n+\t\tgrep \"^[1-9][0-9]* $tree\" objects &&\n+\t\tgrep \"^[1-9][0-9]* $blob\" objects\n+\n+\t)\n+'\n+\n test_done\n-- \n2.42.0.405.gdb2a2f287e\n"},{"id":"483379","messageId":"xmqq5y34wu5f.fsf@gitster.g","threadId":"60318","inReplyTo":"239bf39bfb21ef621a15839bade34446dcbc3103.1697560266.git.me@ttaylorr.com","subject":"Re: [PATCH v2 5/7] bulk-checkin: introduce `index_blob_bulk_checkin_incore()`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-18T02:18:04Z","receivedAt":"2023-10-18T02:18:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n>  bulk-checkin.c | 118 +++++++++++++++++++++++++++++++++++++++++++++++++\n>  bulk-checkin.h |   4 ++\n>  2 files changed, 122 insertions(+)\n\nUnlike the previous four, which were very clear refactoring to\ncreate reusable helper functions, this step leaves a bad aftertaste\nafter reading twice, and I think what is disturbing is that many new\nlines are pretty much literally copied from stream_blob_to_pack().\n\nI wonder if we can introduce an \"input\" source abstraction, that\nreplaces \"fd\" and \"size\" (and \"path\" for error reporting) parameters\nto the stream_blob_to_pack(), so that the bulk of the implementation\nof stream_blob_to_pack() can call its .read() method to read bytes\nup to \"size\" from such an abstracted interface?  That would be a\ngood sized first half of this change.  Then in the second half, you\ncan add another \"input\" source that works with in-core \"buf\" and\n\"size\", whose .read() method will merely be a memcpy().\n\nThat way, we will have two functions, one for stream_obj_to_pack()\nthat reads from an open file descriptor, and the other for\nstream_obj_to_pack_incore() that reads from an in-core buffer,\nsharing the bulk of the implementation that is extracted from the\ncurrent code, which hopefully be easier to audit.\n\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index f4914fb6d1..25cd1ffa25 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -140,6 +140,69 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n>  \treturn 0;\n>  }\n>  \n> +static int stream_obj_to_pack_incore(struct bulk_checkin_packfile *state,\n> +\t\t\t\t     git_hash_ctx *ctx,\n> +\t\t\t\t     off_t *already_hashed_to,\n> +\t\t\t\t     const void *buf, size_t size,\n> +\t\t\t\t     enum object_type type,\n> +\t\t\t\t     const char *path, unsigned flags)\n> +{\n> +\tgit_zstream s;\n> +\tunsigned char obuf[16384];\n> +\tunsigned hdrlen;\n> +\tint status = Z_OK;\n> +\tint write_object = (flags & HASH_WRITE_OBJECT);\n> +\n> +\tgit_deflate_init(&s, pack_compression_level);\n> +\n> +\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), type, size);\n> +\ts.next_out = obuf + hdrlen;\n> +\ts.avail_out = sizeof(obuf) - hdrlen;\n> +\n> +\tif (*already_hashed_to < size) {\n> +\t\tsize_t hsize = size - *already_hashed_to;\n> +\t\tif (hsize) {\n> +\t\t\tthe_hash_algo->update_fn(ctx, buf, hsize);\n> +\t\t}\n> +\t\t*already_hashed_to = size;\n> +\t}\n> +\ts.next_in = (void *)buf;\n> +\ts.avail_in = size;\n> +\n> +\twhile (status != Z_STREAM_END) {\n> +\t\tstatus = git_deflate(&s, Z_FINISH);\n> +\t\tif (!s.avail_out || status == Z_STREAM_END) {\n> +\t\t\tif (write_object) {\n> +\t\t\t\tsize_t written = s.next_out - obuf;\n> +\n> +\t\t\t\t/* would we bust the size limit? */\n> +\t\t\t\tif (state->nr_written &&\n> +\t\t\t\t    pack_size_limit_cfg &&\n> +\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n> +\t\t\t\t\tgit_deflate_abort(&s);\n> +\t\t\t\t\treturn -1;\n> +\t\t\t\t}\n> +\n> +\t\t\t\thashwrite(state->f, obuf, written);\n> +\t\t\t\tstate->offset += written;\n> +\t\t\t}\n> +\t\t\ts.next_out = obuf;\n> +\t\t\ts.avail_out = sizeof(obuf);\n> +\t\t}\n> +\n> +\t\tswitch (status) {\n> +\t\tcase Z_OK:\n> +\t\tcase Z_BUF_ERROR:\n> +\t\tcase Z_STREAM_END:\n> +\t\t\tcontinue;\n> +\t\tdefault:\n> +\t\t\tdie(\"unexpected deflate failure: %d\", status);\n> +\t\t}\n> +\t}\n> +\tgit_deflate_end(&s);\n> +\treturn 0;\n> +}\n> +\n>  /*\n>   * Read the contents from fd for size bytes, streaming it to the\n>   * packfile in state while updating the hash in ctx. Signal a failure\n> @@ -316,6 +379,50 @@ static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n>  \t}\n>  }\n>  \n> +static int deflate_obj_contents_to_pack_incore(struct bulk_checkin_packfile *state,\n> +\t\t\t\t\t       git_hash_ctx *ctx,\n> +\t\t\t\t\t       struct hashfile_checkpoint *checkpoint,\n> +\t\t\t\t\t       struct object_id *result_oid,\n> +\t\t\t\t\t       const void *buf, size_t size,\n> +\t\t\t\t\t       enum object_type type,\n> +\t\t\t\t\t       const char *path, unsigned flags)\n> +{\n> +\tstruct pack_idx_entry *idx = NULL;\n> +\toff_t already_hashed_to = 0;\n> +\n> +\t/* Note: idx is non-NULL when we are writing */\n> +\tif (flags & HASH_WRITE_OBJECT)\n> +\t\tCALLOC_ARRAY(idx, 1);\n> +\n> +\twhile (1) {\n> +\t\tprepare_checkpoint(state, checkpoint, idx, flags);\n> +\t\tif (!stream_obj_to_pack_incore(state, ctx, &already_hashed_to,\n> +\t\t\t\t\t       buf, size, type, path, flags))\n> +\t\t\tbreak;\n> +\t\ttruncate_checkpoint(state, checkpoint, idx);\n> +\t}\n> +\n> +\tfinalize_checkpoint(state, ctx, checkpoint, idx, result_oid);\n> +\n> +\treturn 0;\n> +}\n> +\n> +static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n> +\t\t\t\t       struct object_id *result_oid,\n> +\t\t\t\t       const void *buf, size_t size,\n> +\t\t\t\t       const char *path, unsigned flags)\n> +{\n> +\tgit_hash_ctx ctx;\n> +\tstruct hashfile_checkpoint checkpoint = {0};\n> +\n> +\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n> +\t\t\t\t  size);\n> +\n> +\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n> +\t\t\t\t\t\t   result_oid, buf, size,\n> +\t\t\t\t\t\t   OBJ_BLOB, path, flags);\n> +}\n> +\n>  static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n>  \t\t\t\tstruct object_id *result_oid,\n>  \t\t\t\tint fd, size_t size,\n> @@ -396,6 +503,17 @@ int index_blob_bulk_checkin(struct object_id *oid,\n>  \treturn status;\n>  }\n>  \n> +int index_blob_bulk_checkin_incore(struct object_id *oid,\n> +\t\t\t\t   const void *buf, size_t size,\n> +\t\t\t\t   const char *path, unsigned flags)\n> +{\n> +\tint status = deflate_blob_to_pack_incore(&bulk_checkin_packfile, oid,\n> +\t\t\t\t\t\t buf, size, path, flags);\n> +\tif (!odb_transaction_nesting)\n> +\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n> +\treturn status;\n> +}\n> +\n>  void begin_odb_transaction(void)\n>  {\n>  \todb_transaction_nesting += 1;\n> diff --git a/bulk-checkin.h b/bulk-checkin.h\n> index aa7286a7b3..1b91daeaee 100644\n> --- a/bulk-checkin.h\n> +++ b/bulk-checkin.h\n> @@ -13,6 +13,10 @@ int index_blob_bulk_checkin(struct object_id *oid,\n>  \t\t\t    int fd, size_t size,\n>  \t\t\t    const char *path, unsigned flags);\n>  \n> +int index_blob_bulk_checkin_incore(struct object_id *oid,\n> +\t\t\t\t   const void *buf, size_t size,\n> +\t\t\t\t   const char *path, unsigned flags);\n> +\n>  /*\n>   * Tell the object database to optimize for adding\n>   * multiple objects. end_odb_transaction must be called\n"},{"id":"483408","messageId":"ZTAJFK8BLFEb9FFq@nand.local","threadId":"60318","inReplyTo":"xmqq5y34wu5f.fsf@gitster.g","subject":"Re: [PATCH v2 5/7] bulk-checkin: introduce `index_blob_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T16:34:28Z","receivedAt":"2023-10-18T16:34:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Oct 17, 2023 at 07:18:04PM -0700, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> >  bulk-checkin.c | 118 +++++++++++++++++++++++++++++++++++++++++++++++++\n> >  bulk-checkin.h |   4 ++\n> >  2 files changed, 122 insertions(+)\n>\n> Unlike the previous four, which were very clear refactoring to\n> create reusable helper functions, this step leaves a bad aftertaste\n> after reading twice, and I think what is disturbing is that many new\n> lines are pretty much literally copied from stream_blob_to_pack().\n>\n> I wonder if we can introduce an \"input\" source abstraction, that\n> replaces \"fd\" and \"size\" (and \"path\" for error reporting) parameters\n> to the stream_blob_to_pack(), so that the bulk of the implementation\n> of stream_blob_to_pack() can call its .read() method to read bytes\n> up to \"size\" from such an abstracted interface?  That would be a\n> good sized first half of this change.  Then in the second half, you\n> can add another \"input\" source that works with in-core \"buf\" and\n> \"size\", whose .read() method will merely be a memcpy().\n\nThanks, I like this idea. I had initially avoided it in the first couple\nof rounds, because the abstraction felt clunky and involved an\nunnecessary extra memcpy().\n\nBut having applied your suggestion here, I think that the price is well\nworth the result, which is that `stream_blob_to_pack()` does not have to\nbe implemented twice with very subtle differences.\n\nThanks again for the suggestion, I'm really pleased with how it came\nout. Reroll coming shortly...\n\nThanks,\nTaylor\n"},{"id":"483409","messageId":"20c32d2178560180692327d8b93fe2a7adcf6ffd.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 03/10] bulk-checkin: factor out `truncate_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:07:56Z","receivedAt":"2023-10-18T17:08:39Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as previous commits, factor our the routine to\ntruncate a bulk-checkin packfile when writing past the pack size limit.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 27 +++++++++++++++++----------\n 1 file changed, 17 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex c1f5450583..b92d7a6f5a 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -276,6 +276,22 @@ static void prepare_checkpoint(struct bulk_checkin_packfile *state,\n \t}\n }\n \n+static void truncate_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t\tstruct hashfile_checkpoint *checkpoint,\n+\t\t\t\tstruct pack_idx_entry *idx)\n+{\n+\t/*\n+\t * Writing this object to the current pack will make\n+\t * it too big; we need to truncate it, start a new\n+\t * pack, and write into it.\n+\t */\n+\tif (!idx)\n+\t\tBUG(\"should not happen\");\n+\thashfile_truncate(state->f, checkpoint);\n+\tstate->offset = checkpoint->offset;\n+\tflush_bulk_checkin_packfile(state);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -304,16 +320,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n \t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n-\t\t/*\n-\t\t * Writing this object to the current pack will make\n-\t\t * it too big; we need to truncate it, start a new\n-\t\t * pack, and write into it.\n-\t\t */\n-\t\tif (!idx)\n-\t\t\tBUG(\"should not happen\");\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tflush_bulk_checkin_packfile(state);\n+\t\ttruncate_checkpoint(state, &checkpoint, idx);\n \t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483410","messageId":"ae70508037265ed220d1d33543d61c5d9f0721e0.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 10/10] builtin/merge-tree.c: implement support for `--write-pack`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:08:17Z","receivedAt":"2023-10-18T17:08:49Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using merge-tree often within a repository[^1], it is possible to\ngenerate a relatively large number of loose objects, which can result in\ndegraded performance, and inode exhaustion in extreme cases.\n\nBuilding on the functionality introduced in previous commits, the\nbulk-checkin machinery now has support to write arbitrary blob and tree\nobjects which are small enough to be held in-core. We can use this to\nwrite any blob/tree objects generated by ORT into a separate pack\ninstead of writing them out individually as loose.\n\nThis functionality is gated behind a new `--write-pack` option to\n`merge-tree` that works with the (non-deprecated) `--write-tree` mode.\n\nThe implementation is relatively straightforward. There are two spots\nwithin the ORT mechanism where we call `write_object_file()`, one for\ncontent differences within blobs, and another to assemble any new trees\nnecessary to construct the merge. In each of those locations,\nconditionally replace calls to `write_object_file()` with\n`index_blob_bulk_checkin_incore()` or `index_tree_bulk_checkin_incore()`\ndepending on which kind of object we are writing.\n\nThe only remaining task is to begin and end the transaction necessary to\ninitialize the bulk-checkin machinery, and move any new pack(s) it\ncreated into the main object store.\n\n[^1]: Such is the case at GitHub, where we run presumptive \"test merges\"\n  on open pull requests to see whether or not we can light up the merge\n  button green depending on whether or not the presumptive merge was\n  conflicted.\n\n  This is done in response to a number of user-initiated events,\n  including viewing an open pull request whose last test merge is stale\n  with respect to the current base and tip of the pull request. As a\n  result, merge-tree can be run very frequently on large, active\n  repositories.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-merge-tree.txt |  4 ++\n builtin/merge-tree.c             |  5 ++\n merge-ort.c                      | 42 +++++++++++----\n merge-recursive.h                |  1 +\n t/t4301-merge-tree-write-tree.sh | 93 ++++++++++++++++++++++++++++++++\n 5 files changed, 136 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/git-merge-tree.txt b/Documentation/git-merge-tree.txt\nindex ffc4fbf7e8..9d37609ef1 100644\n--- a/Documentation/git-merge-tree.txt\n+++ b/Documentation/git-merge-tree.txt\n@@ -69,6 +69,10 @@ OPTIONS\n \tspecify a merge-base for the merge, and specifying multiple bases is\n \tcurrently not supported. This option is incompatible with `--stdin`.\n \n+--write-pack::\n+\tWrite any new objects into a separate packfile instead of as\n+\tindividual loose objects.\n+\n [[OUTPUT]]\n OUTPUT\n ------\ndiff --git a/builtin/merge-tree.c b/builtin/merge-tree.c\nindex 0de42aecf4..672ebd4c54 100644\n--- a/builtin/merge-tree.c\n+++ b/builtin/merge-tree.c\n@@ -18,6 +18,7 @@\n #include \"quote.h\"\n #include \"tree.h\"\n #include \"config.h\"\n+#include \"bulk-checkin.h\"\n \n static int line_termination = '\\n';\n \n@@ -414,6 +415,7 @@ struct merge_tree_options {\n \tint show_messages;\n \tint name_only;\n \tint use_stdin;\n+\tint write_pack;\n };\n \n static int real_merge(struct merge_tree_options *o,\n@@ -440,6 +442,7 @@ static int real_merge(struct merge_tree_options *o,\n \tinit_merge_options(&opt, the_repository);\n \n \topt.show_rename_progress = 0;\n+\topt.write_pack = o->write_pack;\n \n \topt.branch1 = branch1;\n \topt.branch2 = branch2;\n@@ -548,6 +551,8 @@ int cmd_merge_tree(int argc, const char **argv, const char *prefix)\n \t\t\t   &merge_base,\n \t\t\t   N_(\"commit\"),\n \t\t\t   N_(\"specify a merge-base for the merge\")),\n+\t\tOPT_BOOL(0, \"write-pack\", &o.write_pack,\n+\t\t\t N_(\"write new objects to a pack instead of as loose\")),\n \t\tOPT_END()\n \t};\n \ndiff --git a/merge-ort.c b/merge-ort.c\nindex 7857ce9fbd..e198d2bc2b 100644\n--- a/merge-ort.c\n+++ b/merge-ort.c\n@@ -48,6 +48,7 @@\n #include \"tree.h\"\n #include \"unpack-trees.h\"\n #include \"xdiff-interface.h\"\n+#include \"bulk-checkin.h\"\n \n /*\n  * We have many arrays of size 3.  Whenever we have such an array, the\n@@ -2107,10 +2108,19 @@ static int handle_content_merge(struct merge_options *opt,\n \t\tif ((merge_status < 0) || !result_buf.ptr)\n \t\t\tret = error(_(\"failed to execute internal merge\"));\n \n-\t\tif (!ret &&\n-\t\t    write_object_file(result_buf.ptr, result_buf.size,\n-\t\t\t\t      OBJ_BLOB, &result->oid))\n-\t\t\tret = error(_(\"unable to add %s to database\"), path);\n+\t\tif (!ret) {\n+\t\t\tret = opt->write_pack\n+\t\t\t\t? index_blob_bulk_checkin_incore(&result->oid,\n+\t\t\t\t\t\t\t\t result_buf.ptr,\n+\t\t\t\t\t\t\t\t result_buf.size,\n+\t\t\t\t\t\t\t\t path, 1)\n+\t\t\t\t: write_object_file(result_buf.ptr,\n+\t\t\t\t\t\t    result_buf.size,\n+\t\t\t\t\t\t    OBJ_BLOB, &result->oid);\n+\t\t\tif (ret)\n+\t\t\t\tret = error(_(\"unable to add %s to database\"),\n+\t\t\t\t\t    path);\n+\t\t}\n \n \t\tfree(result_buf.ptr);\n \t\tif (ret)\n@@ -3596,7 +3606,8 @@ static int tree_entry_order(const void *a_, const void *b_)\n \t\t\t\t b->string, strlen(b->string), bmi->result.mode);\n }\n \n-static int write_tree(struct object_id *result_oid,\n+static int write_tree(struct merge_options *opt,\n+\t\t      struct object_id *result_oid,\n \t\t      struct string_list *versions,\n \t\t      unsigned int offset,\n \t\t      size_t hash_size)\n@@ -3630,8 +3641,14 @@ static int write_tree(struct object_id *result_oid,\n \t}\n \n \t/* Write this object file out, and record in result_oid */\n-\tif (write_object_file(buf.buf, buf.len, OBJ_TREE, result_oid))\n+\tret = opt->write_pack\n+\t\t? index_tree_bulk_checkin_incore(result_oid,\n+\t\t\t\t\t\t buf.buf, buf.len, \"\", 1)\n+\t\t: write_object_file(buf.buf, buf.len, OBJ_TREE, result_oid);\n+\n+\tif (ret)\n \t\tret = -1;\n+\n \tstrbuf_release(&buf);\n \treturn ret;\n }\n@@ -3796,8 +3813,8 @@ static int write_completed_directory(struct merge_options *opt,\n \t\t */\n \t\tdir_info->is_null = 0;\n \t\tdir_info->result.mode = S_IFDIR;\n-\t\tif (write_tree(&dir_info->result.oid, &info->versions, offset,\n-\t\t\t       opt->repo->hash_algo->rawsz) < 0)\n+\t\tif (write_tree(opt, &dir_info->result.oid, &info->versions,\n+\t\t\t       offset, opt->repo->hash_algo->rawsz) < 0)\n \t\t\tret = -1;\n \t}\n \n@@ -4331,9 +4348,13 @@ static int process_entries(struct merge_options *opt,\n \t\tfflush(stdout);\n \t\tBUG(\"dir_metadata accounting completely off; shouldn't happen\");\n \t}\n-\tif (write_tree(result_oid, &dir_metadata.versions, 0,\n+\tif (write_tree(opt, result_oid, &dir_metadata.versions, 0,\n \t\t       opt->repo->hash_algo->rawsz) < 0)\n \t\tret = -1;\n+\n+\tif (opt->write_pack)\n+\t\tend_odb_transaction();\n+\n cleanup:\n \tstring_list_clear(&plist, 0);\n \tstring_list_clear(&dir_metadata.versions, 0);\n@@ -4877,6 +4898,9 @@ static void merge_start(struct merge_options *opt, struct merge_result *result)\n \t */\n \tstrmap_init(&opt->priv->conflicts);\n \n+\tif (opt->write_pack)\n+\t\tbegin_odb_transaction();\n+\n \ttrace2_region_leave(\"merge\", \"allocate/init\", opt->repo);\n }\n \ndiff --git a/merge-recursive.h b/merge-recursive.h\nindex b88000e3c2..156e160876 100644\n--- a/merge-recursive.h\n+++ b/merge-recursive.h\n@@ -48,6 +48,7 @@ struct merge_options {\n \tunsigned renormalize : 1;\n \tunsigned record_conflict_msgs_as_headers : 1;\n \tconst char *msg_header_prefix;\n+\tunsigned write_pack : 1;\n \n \t/* internal fields used by the implementation */\n \tstruct merge_options_internal *priv;\ndiff --git a/t/t4301-merge-tree-write-tree.sh b/t/t4301-merge-tree-write-tree.sh\nindex 250f721795..2d81ff4de5 100755\n--- a/t/t4301-merge-tree-write-tree.sh\n+++ b/t/t4301-merge-tree-write-tree.sh\n@@ -922,4 +922,97 @@ test_expect_success 'check the input format when --stdin is passed' '\n \ttest_cmp expect actual\n '\n \n+packdir=\".git/objects/pack\"\n+\n+test_expect_success 'merge-tree can pack its result with --write-pack' '\n+\ttest_when_finished \"rm -rf repo\" &&\n+\tgit init repo &&\n+\n+\t# base has lines [3, 4, 5]\n+\t#   - side adds to the beginning, resulting in [1, 2, 3, 4, 5]\n+\t#   - other adds to the end, resulting in [3, 4, 5, 6, 7]\n+\t#\n+\t# merging the two should result in a new blob object containing\n+\t# [1, 2, 3, 4, 5, 6, 7], along with a new tree.\n+\ttest_commit -C repo base file \"$(test_seq 3 5)\" &&\n+\tgit -C repo branch -M main &&\n+\tgit -C repo checkout -b side main &&\n+\ttest_commit -C repo side file \"$(test_seq 1 5)\" &&\n+\tgit -C repo checkout -b other main &&\n+\ttest_commit -C repo other file \"$(test_seq 3 7)\" &&\n+\n+\tfind repo/$packdir -type f -name \"pack-*.idx\" >packs.before &&\n+\ttree=\"$(git -C repo merge-tree --write-pack \\\n+\t\trefs/tags/side refs/tags/other)\" &&\n+\tblob=\"$(git -C repo rev-parse $tree:file)\" &&\n+\tfind repo/$packdir -type f -name \"pack-*.idx\" >packs.after &&\n+\n+\ttest_must_be_empty packs.before &&\n+\ttest_line_count = 1 packs.after &&\n+\n+\tgit show-index <$(cat packs.after) >objects &&\n+\ttest_line_count = 2 objects &&\n+\tgrep \"^[1-9][0-9]* $tree\" objects &&\n+\tgrep \"^[1-9][0-9]* $blob\" objects\n+'\n+\n+test_expect_success 'merge-tree can write multiple packs with --write-pack' '\n+\ttest_when_finished \"rm -rf repo\" &&\n+\tgit init repo &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\tgit config pack.packSizeLimit 512 &&\n+\n+\t\ttest_seq 512 >f &&\n+\n+\t\t# \"f\" contains roughly ~2,000 bytes.\n+\t\t#\n+\t\t# Each side (\"foo\" and \"bar\") adds a small amount of data at the\n+\t\t# beginning and end of \"base\", respectively.\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m base &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout -b foo main &&\n+\t\t{\n+\t\t\techo foo && cat f\n+\t\t} >f.tmp &&\n+\t\tmv f.tmp f &&\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m foo &&\n+\n+\t\tgit checkout -b bar main &&\n+\t\techo bar >>f &&\n+\t\tgit add f &&\n+\t\ttest_tick &&\n+\t\tgit commit -m bar &&\n+\n+\t\tfind $packdir -type f -name \"pack-*.idx\" >packs.before &&\n+\t\t# Merging either side should result in a new object which is\n+\t\t# larger than 1M, thus the result should be split into two\n+\t\t# separate packs.\n+\t\ttree=\"$(git merge-tree --write-pack \\\n+\t\t\trefs/heads/foo refs/heads/bar)\" &&\n+\t\tblob=\"$(git rev-parse $tree:f)\" &&\n+\t\tfind $packdir -type f -name \"pack-*.idx\" >packs.after &&\n+\n+\t\ttest_must_be_empty packs.before &&\n+\t\ttest_line_count = 2 packs.after &&\n+\t\tfor idx in $(cat packs.after)\n+\t\tdo\n+\t\t\tgit show-index <$idx || return 1\n+\t\tdone >objects &&\n+\n+\t\t# The resulting set of packs should contain one copy of both\n+\t\t# objects, each in a separate pack.\n+\t\ttest_line_count = 2 objects &&\n+\t\tgrep \"^[1-9][0-9]* $tree\" objects &&\n+\t\tgrep \"^[1-9][0-9]* $blob\" objects\n+\n+\t)\n+'\n+\n test_done\n-- \n2.42.0.408.g97fac66ae4\n"},{"id":"483411","messageId":"04ec74e3574b8e0cfc503c46fa3481ef196348ac.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 07/10] bulk-checkin: generify `stream_blob_to_pack()` for arbitrary types","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:08:08Z","receivedAt":"2023-10-18T17:09:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The existing `stream_blob_to_pack()` function is named based on the fact\nthat it knows only how to stream blobs into a bulk-checkin pack.\n\nBut there is no longer anything in this function which prevents us from\nwriting objects of arbitrary types to the bulk-checkin pack. Prepare to\nwrite OBJ_TREEs by removing this assumption, adding an `enum\nobject_type` parameter to this function's argument list, and renaming it\nto `stream_obj_to_pack()` accordingly.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 15 +++++++--------\n 1 file changed, 7 insertions(+), 8 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 133e02ce36..f0115efb2e 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -204,10 +204,10 @@ static ssize_t bulk_checkin_source_read(struct bulk_checkin_source *source,\n  * status before calling us just in case we ask it to call us again\n  * with a new pack.\n  */\n-static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n-\t\t\t       git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       struct bulk_checkin_source *source,\n-\t\t\t       unsigned flags)\n+static int stream_obj_to_pack(struct bulk_checkin_packfile *state,\n+\t\t\t      git_hash_ctx *ctx, off_t *already_hashed_to,\n+\t\t\t      struct bulk_checkin_source *source,\n+\t\t\t      enum object_type type, unsigned flags)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n@@ -220,8 +220,7 @@ static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n-\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), OBJ_BLOB,\n-\t\t\t\t\t      size);\n+\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), type, size);\n \ts.next_out = obuf + hdrlen;\n \ts.avail_out = sizeof(obuf) - hdrlen;\n \n@@ -402,8 +401,8 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \n \twhile (1) {\n \t\tprepare_checkpoint(state, &checkpoint, idx, flags);\n-\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t &source, flags))\n+\t\tif (!stream_obj_to_pack(state, &ctx, &already_hashed_to,\n+\t\t\t\t\t&source, OBJ_BLOB, flags))\n \t\t\tbreak;\n \t\ttruncate_checkpoint(state, &checkpoint, idx);\n \t\tif (bulk_checkin_source_seek_to(&source, seekback) == (off_t)-1)\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483412","messageId":"8667b763652ffa71b52b7bd78821e46a6e5fe5a9.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 08/10] bulk-checkin: introduce `index_blob_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:08:11Z","receivedAt":"2023-10-18T17:09:48Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that we have factored out many of the common routines necessary to\nindex a new object into a pack created by the bulk-checkin machinery, we\ncan introduce a variant of `index_blob_bulk_checkin()` that acts on\nblobs whose contents we can fit in memory.\n\nThis will be useful in a couple of more commits in order to provide the\n`merge-tree` builtin with a mechanism to create a new pack containing\nany objects it created during the merge, instead of storing those\nobjects individually as loose.\n\nSimilar to the existing `index_blob_bulk_checkin()` function, the\nentrypoint delegates to `deflate_blob_to_pack_incore()`, which is\nresponsible for formatting the pack header and then deflating the\ncontents into the pack. The latter is accomplished by calling\ndeflate_obj_contents_to_pack_incore(), which takes advantage of the\nearlier refactorings and is responsible for writing the object to the\npack and handling any overage from pack.packSizeLimit.\n\nThe bulk of the new functionality is implemented in the function\n`stream_obj_to_pack()`, which can handle streaming objects from memory\nto the bulk-checkin pack as a result of the earlier refactoring.\n\nConsistent with the rest of the bulk-checkin mechanism, there are no\ndirect tests here. In future commits when we expose this new\nfunctionality via the `merge-tree` builtin, we will test it indirectly\nthere.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 64 ++++++++++++++++++++++++++++++++++++++++++++++++++\n bulk-checkin.h |  4 ++++\n 2 files changed, 68 insertions(+)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f0115efb2e..9ae43648ba 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -370,6 +370,59 @@ static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n \t}\n }\n \n+static int deflate_obj_contents_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t\t       git_hash_ctx *ctx,\n+\t\t\t\t\t       struct hashfile_checkpoint *checkpoint,\n+\t\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t\t       enum object_type type,\n+\t\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tstruct pack_idx_entry *idx = NULL;\n+\toff_t already_hashed_to = 0;\n+\tstruct bulk_checkin_source source = {\n+\t\t.type = SOURCE_INCORE,\n+\t\t.buf = buf,\n+\t\t.size = size,\n+\t\t.read = 0,\n+\t\t.path = path,\n+\t};\n+\n+\t/* Note: idx is non-NULL when we are writing */\n+\tif (flags & HASH_WRITE_OBJECT)\n+\t\tCALLOC_ARRAY(idx, 1);\n+\n+\twhile (1) {\n+\t\tprepare_checkpoint(state, checkpoint, idx, flags);\n+\n+\t\tif (!stream_obj_to_pack(state, ctx, &already_hashed_to, &source,\n+\t\t\t\t\ttype, flags))\n+\t\t\tbreak;\n+\t\ttruncate_checkpoint(state, checkpoint, idx);\n+\t\tbulk_checkin_source_seek_to(&source, 0);\n+\t}\n+\n+\tfinalize_checkpoint(state, ctx, checkpoint, idx, result_oid);\n+\n+\treturn 0;\n+}\n+\n+static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tgit_hash_ctx ctx;\n+\tstruct hashfile_checkpoint checkpoint = {0};\n+\n+\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n+\t\t\t\t  size);\n+\n+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n+\t\t\t\t\t\t   result_oid, buf, size,\n+\t\t\t\t\t\t   OBJ_BLOB, path, flags);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -456,6 +509,17 @@ int index_blob_bulk_checkin(struct object_id *oid,\n \treturn status;\n }\n \n+int index_blob_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags)\n+{\n+\tint status = deflate_blob_to_pack_incore(&bulk_checkin_packfile, oid,\n+\t\t\t\t\t\t buf, size, path, flags);\n+\tif (!odb_transaction_nesting)\n+\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n+\treturn status;\n+}\n+\n void begin_odb_transaction(void)\n {\n \todb_transaction_nesting += 1;\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex aa7286a7b3..1b91daeaee 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -13,6 +13,10 @@ int index_blob_bulk_checkin(struct object_id *oid,\n \t\t\t    int fd, size_t size,\n \t\t\t    const char *path, unsigned flags);\n \n+int index_blob_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags);\n+\n /*\n  * Tell the object database to optimize for adding\n  * multiple objects. end_odb_transaction must be called\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483414","messageId":"cover.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH v3 00/10] merge-ort: implement support for packing objects together","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:07:45Z","receivedAt":"2023-10-18T17:10:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This series implements support for a new merge-tree option,\n`--write-pack`, which causes any newly-written objects to be packed\ntogether instead of being stored individually as loose.\n\nThe notable change from last time is in response to a suggestion[1] from\nJunio to factor out an abstract bulk-checkin \"source\", which ended up\nreducing the duplication between a couple of functions in the earlier\nround by a significant degree.\n\nBeyond that, the changes since last time can be viewed in the range-diff\nbelow. Thanks in advance for any review!\n\n[1]: https://lore.kernel.org/git/xmqq5y34wu5f.fsf@gitster.g/\n\nTaylor Blau (10):\n  bulk-checkin: factor out `format_object_header_hash()`\n  bulk-checkin: factor out `prepare_checkpoint()`\n  bulk-checkin: factor out `truncate_checkpoint()`\n  bulk-checkin: factor out `finalize_checkpoint()`\n  bulk-checkin: extract abstract `bulk_checkin_source`\n  bulk-checkin: implement `SOURCE_INCORE` mode for `bulk_checkin_source`\n  bulk-checkin: generify `stream_blob_to_pack()` for arbitrary types\n  bulk-checkin: introduce `index_blob_bulk_checkin_incore()`\n  bulk-checkin: introduce `index_tree_bulk_checkin_incore()`\n  builtin/merge-tree.c: implement support for `--write-pack`\n\n Documentation/git-merge-tree.txt |   4 +\n builtin/merge-tree.c             |   5 +\n bulk-checkin.c                   | 288 +++++++++++++++++++++++++------\n bulk-checkin.h                   |   8 +\n merge-ort.c                      |  42 ++++-\n merge-recursive.h                |   1 +\n t/t4301-merge-tree-write-tree.sh |  93 ++++++++++\n 7 files changed, 381 insertions(+), 60 deletions(-)\n\nRange-diff against v2:\n 1:  edf1cbafc1 =  1:  2dffa45183 bulk-checkin: factor out `format_object_header_hash()`\n 2:  b3f89d5853 =  2:  7a10dc794a bulk-checkin: factor out `prepare_checkpoint()`\n 3:  abe4fb0a59 =  3:  20c32d2178 bulk-checkin: factor out `truncate_checkpoint()`\n 4:  0b855a6eb7 !  4:  893051d0b7 bulk-checkin: factor our `finalize_checkpoint()`\n    @@ Metadata\n     Author: Taylor Blau <me@ttaylorr.com>\n     \n      ## Commit message ##\n    -    bulk-checkin: factor our `finalize_checkpoint()`\n    +    bulk-checkin: factor out `finalize_checkpoint()`\n     \n         In a similar spirit as previous commits, factor out the routine to\n         finalize the just-written object from the bulk-checkin mechanism.\n -:  ---------- >  5:  da52ec8380 bulk-checkin: extract abstract `bulk_checkin_source`\n -:  ---------- >  6:  4e9bac5bc1 bulk-checkin: implement `SOURCE_INCORE` mode for `bulk_checkin_source`\n -:  ---------- >  7:  04ec74e357 bulk-checkin: generify `stream_blob_to_pack()` for arbitrary types\n 5:  239bf39bfb !  8:  8667b76365 bulk-checkin: introduce `index_blob_bulk_checkin_incore()`\n    @@ Commit message\n         entrypoint delegates to `deflate_blob_to_pack_incore()`, which is\n         responsible for formatting the pack header and then deflating the\n         contents into the pack. The latter is accomplished by calling\n    -    deflate_blob_contents_to_pack_incore(), which takes advantage of the\n    -    earlier refactoring and is responsible for writing the object to the\n    +    deflate_obj_contents_to_pack_incore(), which takes advantage of the\n    +    earlier refactorings and is responsible for writing the object to the\n         pack and handling any overage from pack.packSizeLimit.\n     \n         The bulk of the new functionality is implemented in the function\n    -    `stream_obj_to_pack_incore()`, which is a generic implementation for\n    -    writing objects of arbitrary type (whose contents we can fit in-core)\n    -    into a bulk-checkin pack.\n    -\n    -    The new function shares an unfortunate degree of similarity to the\n    -    existing `stream_blob_to_pack()` function. But DRY-ing up these two\n    -    would likely be more trouble than it's worth, since the latter has to\n    -    deal with reading and writing the contents of the object.\n    +    `stream_obj_to_pack()`, which can handle streaming objects from memory\n    +    to the bulk-checkin pack as a result of the earlier refactoring.\n     \n         Consistent with the rest of the bulk-checkin mechanism, there are no\n         direct tests here. In future commits when we expose this new\n    @@ Commit message\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n      ## bulk-checkin.c ##\n    -@@ bulk-checkin.c: static int already_written(struct bulk_checkin_packfile *state, struct object_id\n    - \treturn 0;\n    - }\n    - \n    -+static int stream_obj_to_pack_incore(struct bulk_checkin_packfile *state,\n    -+\t\t\t\t     git_hash_ctx *ctx,\n    -+\t\t\t\t     off_t *already_hashed_to,\n    -+\t\t\t\t     const void *buf, size_t size,\n    -+\t\t\t\t     enum object_type type,\n    -+\t\t\t\t     const char *path, unsigned flags)\n    -+{\n    -+\tgit_zstream s;\n    -+\tunsigned char obuf[16384];\n    -+\tunsigned hdrlen;\n    -+\tint status = Z_OK;\n    -+\tint write_object = (flags & HASH_WRITE_OBJECT);\n    -+\n    -+\tgit_deflate_init(&s, pack_compression_level);\n    -+\n    -+\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), type, size);\n    -+\ts.next_out = obuf + hdrlen;\n    -+\ts.avail_out = sizeof(obuf) - hdrlen;\n    -+\n    -+\tif (*already_hashed_to < size) {\n    -+\t\tsize_t hsize = size - *already_hashed_to;\n    -+\t\tif (hsize) {\n    -+\t\t\tthe_hash_algo->update_fn(ctx, buf, hsize);\n    -+\t\t}\n    -+\t\t*already_hashed_to = size;\n    -+\t}\n    -+\ts.next_in = (void *)buf;\n    -+\ts.avail_in = size;\n    -+\n    -+\twhile (status != Z_STREAM_END) {\n    -+\t\tstatus = git_deflate(&s, Z_FINISH);\n    -+\t\tif (!s.avail_out || status == Z_STREAM_END) {\n    -+\t\t\tif (write_object) {\n    -+\t\t\t\tsize_t written = s.next_out - obuf;\n    -+\n    -+\t\t\t\t/* would we bust the size limit? */\n    -+\t\t\t\tif (state->nr_written &&\n    -+\t\t\t\t    pack_size_limit_cfg &&\n    -+\t\t\t\t    pack_size_limit_cfg < state->offset + written) {\n    -+\t\t\t\t\tgit_deflate_abort(&s);\n    -+\t\t\t\t\treturn -1;\n    -+\t\t\t\t}\n    -+\n    -+\t\t\t\thashwrite(state->f, obuf, written);\n    -+\t\t\t\tstate->offset += written;\n    -+\t\t\t}\n    -+\t\t\ts.next_out = obuf;\n    -+\t\t\ts.avail_out = sizeof(obuf);\n    -+\t\t}\n    -+\n    -+\t\tswitch (status) {\n    -+\t\tcase Z_OK:\n    -+\t\tcase Z_BUF_ERROR:\n    -+\t\tcase Z_STREAM_END:\n    -+\t\t\tcontinue;\n    -+\t\tdefault:\n    -+\t\t\tdie(\"unexpected deflate failure: %d\", status);\n    -+\t\t}\n    -+\t}\n    -+\tgit_deflate_end(&s);\n    -+\treturn 0;\n    -+}\n    -+\n    - /*\n    -  * Read the contents from fd for size bytes, streaming it to the\n    -  * packfile in state while updating the hash in ctx. Signal a failure\n     @@ bulk-checkin.c: static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n      \t}\n      }\n    @@ bulk-checkin.c: static void finalize_checkpoint(struct bulk_checkin_packfile *st\n     +{\n     +\tstruct pack_idx_entry *idx = NULL;\n     +\toff_t already_hashed_to = 0;\n    ++\tstruct bulk_checkin_source source = {\n    ++\t\t.type = SOURCE_INCORE,\n    ++\t\t.buf = buf,\n    ++\t\t.size = size,\n    ++\t\t.read = 0,\n    ++\t\t.path = path,\n    ++\t};\n     +\n     +\t/* Note: idx is non-NULL when we are writing */\n     +\tif (flags & HASH_WRITE_OBJECT)\n    @@ bulk-checkin.c: static void finalize_checkpoint(struct bulk_checkin_packfile *st\n     +\n     +\twhile (1) {\n     +\t\tprepare_checkpoint(state, checkpoint, idx, flags);\n    -+\t\tif (!stream_obj_to_pack_incore(state, ctx, &already_hashed_to,\n    -+\t\t\t\t\t       buf, size, type, path, flags))\n    ++\n    ++\t\tif (!stream_obj_to_pack(state, ctx, &already_hashed_to, &source,\n    ++\t\t\t\t\ttype, flags))\n     +\t\t\tbreak;\n     +\t\ttruncate_checkpoint(state, checkpoint, idx);\n    ++\t\tbulk_checkin_source_seek_to(&source, 0);\n     +\t}\n     +\n     +\tfinalize_checkpoint(state, ctx, checkpoint, idx, result_oid);\n 6:  57613807d8 =  9:  cba043ef14 bulk-checkin: introduce `index_tree_bulk_checkin_incore()`\n 7:  f21400f56c = 10:  ae70508037 builtin/merge-tree.c: implement support for `--write-pack`\n-- \n2.42.0.408.g97fac66ae4\n"},{"id":"483415","messageId":"cba043ef14fbf6fdeacc213669bb95f1e6f81f8a.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 09/10] bulk-checkin: introduce `index_tree_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:08:14Z","receivedAt":"2023-10-18T17:10:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The remaining missing piece in order to teach the `merge-tree` builtin\nhow to write the contents of a merge into a pack is a function to index\ntree objects into a bulk-checkin pack.\n\nThis patch implements that missing piece, which is a thin wrapper around\nall of the functionality introduced in previous commits.\n\nIf and when Git gains support for a \"compatibility\" hash algorithm, the\nchanges to support that here will be minimal. The bulk-checkin machinery\nwill need to convert the incoming tree to compute its length under the\ncompatibility hash, necessary to reconstruct its header. With that\ninformation (and the converted contents of the tree), the bulk-checkin\nmachinery will have enough to keep track of the converted object's hash\nin order to update the compatibility mapping.\n\nWithin `deflate_tree_to_pack_incore()`, the changes should be limited\nto something like:\n\n    struct strbuf converted = STRBUF_INIT;\n    if (the_repository->compat_hash_algo) {\n      if (convert_object_file(&compat_obj,\n                              the_repository->hash_algo,\n                              the_repository->compat_hash_algo, ...) < 0)\n        die(...);\n\n      format_object_header_hash(the_repository->compat_hash_algo,\n                                OBJ_TREE, size);\n    }\n    /* compute the converted tree's hash using the compat algorithm */\n    strbuf_release(&converted);\n\n, assuming related changes throughout the rest of the bulk-checkin\nmachinery necessary to update the hash of the converted object, which\nare likewise minimal in size.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 27 +++++++++++++++++++++++++++\n bulk-checkin.h |  4 ++++\n 2 files changed, 31 insertions(+)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 9ae43648ba..d088a9c10b 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -423,6 +423,22 @@ static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n \t\t\t\t\t\t   OBJ_BLOB, path, flags);\n }\n \n+static int deflate_tree_to_pack_incore(struct bulk_checkin_packfile *state,\n+\t\t\t\t       struct object_id *result_oid,\n+\t\t\t\t       const void *buf, size_t size,\n+\t\t\t\t       const char *path, unsigned flags)\n+{\n+\tgit_hash_ctx ctx;\n+\tstruct hashfile_checkpoint checkpoint = {0};\n+\n+\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_TREE,\n+\t\t\t\t  size);\n+\n+\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n+\t\t\t\t\t\t   result_oid, buf, size,\n+\t\t\t\t\t\t   OBJ_TREE, path, flags);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -520,6 +536,17 @@ int index_blob_bulk_checkin_incore(struct object_id *oid,\n \treturn status;\n }\n \n+int index_tree_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags)\n+{\n+\tint status = deflate_tree_to_pack_incore(&bulk_checkin_packfile, oid,\n+\t\t\t\t\t\t buf, size, path, flags);\n+\tif (!odb_transaction_nesting)\n+\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n+\treturn status;\n+}\n+\n void begin_odb_transaction(void)\n {\n \todb_transaction_nesting += 1;\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 1b91daeaee..89786b3954 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -17,6 +17,10 @@ int index_blob_bulk_checkin_incore(struct object_id *oid,\n \t\t\t\t   const void *buf, size_t size,\n \t\t\t\t   const char *path, unsigned flags);\n \n+int index_tree_bulk_checkin_incore(struct object_id *oid,\n+\t\t\t\t   const void *buf, size_t size,\n+\t\t\t\t   const char *path, unsigned flags);\n+\n /*\n  * Tell the object database to optimize for adding\n  * multiple objects. end_odb_transaction must be called\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483416","messageId":"2dffa4518339a7b96a885db4c64431276bfeb4d6.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 01/10] bulk-checkin: factor out `format_object_header_hash()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:07:48Z","receivedAt":"2023-10-18T17:10:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Before deflating a blob into a pack, the bulk-checkin mechanism prepares\nthe pack object header by calling `format_object_header()`, and writing\ninto a scratch buffer, the contents of which eventually makes its way\ninto the pack.\n\nFuture commits will add support for deflating multiple kinds of objects\ninto a pack, and will likewise need to perform a similar operation as\nbelow.\n\nThis is a mostly straightforward extraction, with one notable exception.\nInstead of hard-coding `the_hash_algo`, pass it in to the new function\nas an argument. This isn't strictly necessary for our immediate purposes\nhere, but will prove useful in the future if/when the bulk-checkin\nmechanism grows support for the hash transition plan.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 25 ++++++++++++++++++-------\n 1 file changed, 18 insertions(+), 7 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6ce62999e5..fd3c110d1c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -247,6 +247,22 @@ static void prepare_to_stream(struct bulk_checkin_packfile *state,\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n+static void format_object_header_hash(const struct git_hash_algo *algop,\n+\t\t\t\t      git_hash_ctx *ctx,\n+\t\t\t\t      struct hashfile_checkpoint *checkpoint,\n+\t\t\t\t      enum object_type type,\n+\t\t\t\t      size_t size)\n+{\n+\tunsigned char header[16384];\n+\tunsigned header_len = format_object_header((char *)header,\n+\t\t\t\t\t\t   sizeof(header),\n+\t\t\t\t\t\t   type, size);\n+\n+\talgop->init_fn(ctx);\n+\talgop->update_fn(ctx, header, header_len);\n+\talgop->init_fn(&checkpoint->ctx);\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -254,8 +270,6 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n {\n \toff_t seekback, already_hashed_to;\n \tgit_hash_ctx ctx;\n-\tunsigned char obuf[16384];\n-\tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint = {0};\n \tstruct pack_idx_entry *idx = NULL;\n \n@@ -263,11 +277,8 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \tif (seekback == (off_t) -1)\n \t\treturn error(\"cannot find the current offset\");\n \n-\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n-\t\t\t\t\t  OBJ_BLOB, size);\n-\tthe_hash_algo->init_fn(&ctx);\n-\tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n-\tthe_hash_algo->init_fn(&checkpoint.ctx);\n+\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n+\t\t\t\t  size);\n \n \t/* Note: idx is non-NULL when we are writing */\n \tif ((flags & HASH_WRITE_OBJECT) != 0)\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483417","messageId":"7a10dc794aad20cfc226184acda1d40b191164d5.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 02/10] bulk-checkin: factor out `prepare_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:07:52Z","receivedAt":"2023-10-18T17:10:21Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as the previous commit, factor out the routine to\nprepare streaming into a bulk-checkin pack into its own function. Unlike\nthe previous patch, this is a verbatim copy and paste.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 20 ++++++++++++++------\n 1 file changed, 14 insertions(+), 6 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex fd3c110d1c..c1f5450583 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -263,6 +263,19 @@ static void format_object_header_hash(const struct git_hash_algo *algop,\n \talgop->init_fn(&checkpoint->ctx);\n }\n \n+static void prepare_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t       struct hashfile_checkpoint *checkpoint,\n+\t\t\t       struct pack_idx_entry *idx,\n+\t\t\t       unsigned flags)\n+{\n+\tprepare_to_stream(state, flags);\n+\tif (idx) {\n+\t\thashfile_checkpoint(state->f, checkpoint);\n+\t\tidx->offset = state->offset;\n+\t\tcrc32_begin(state->f);\n+\t}\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -287,12 +300,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \talready_hashed_to = 0;\n \n \twhile (1) {\n-\t\tprepare_to_stream(state, flags);\n-\t\tif (idx) {\n-\t\t\thashfile_checkpoint(state->f, &checkpoint);\n-\t\t\tidx->offset = state->offset;\n-\t\t\tcrc32_begin(state->f);\n-\t\t}\n+\t\tprepare_checkpoint(state, &checkpoint, idx, flags);\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n \t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483418","messageId":"893051d0b7aa162396778cd696e98ae507d7f3d6.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 04/10] bulk-checkin: factor out `finalize_checkpoint()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:07:59Z","receivedAt":"2023-10-18T17:10:24Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a similar spirit as previous commits, factor out the routine to\nfinalize the just-written object from the bulk-checkin mechanism.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 41 +++++++++++++++++++++++++----------------\n 1 file changed, 25 insertions(+), 16 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b92d7a6f5a..f4914fb6d1 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -292,6 +292,30 @@ static void truncate_checkpoint(struct bulk_checkin_packfile *state,\n \tflush_bulk_checkin_packfile(state);\n }\n \n+static void finalize_checkpoint(struct bulk_checkin_packfile *state,\n+\t\t\t\tgit_hash_ctx *ctx,\n+\t\t\t\tstruct hashfile_checkpoint *checkpoint,\n+\t\t\t\tstruct pack_idx_entry *idx,\n+\t\t\t\tstruct object_id *result_oid)\n+{\n+\tthe_hash_algo->final_oid_fn(result_oid, ctx);\n+\tif (!idx)\n+\t\treturn;\n+\n+\tidx->crc32 = crc32_end(state->f);\n+\tif (already_written(state, result_oid)) {\n+\t\thashfile_truncate(state->f, checkpoint);\n+\t\tstate->offset = checkpoint->offset;\n+\t\tfree(idx);\n+\t} else {\n+\t\toidcpy(&idx->oid, result_oid);\n+\t\tALLOC_GROW(state->written,\n+\t\t\t   state->nr_written + 1,\n+\t\t\t   state->alloc_written);\n+\t\tstate->written[state->nr_written++] = idx;\n+\t}\n+}\n+\n static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tstruct object_id *result_oid,\n \t\t\t\tint fd, size_t size,\n@@ -324,22 +348,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n-\tthe_hash_algo->final_oid_fn(result_oid, &ctx);\n-\tif (!idx)\n-\t\treturn 0;\n-\n-\tidx->crc32 = crc32_end(state->f);\n-\tif (already_written(state, result_oid)) {\n-\t\thashfile_truncate(state->f, &checkpoint);\n-\t\tstate->offset = checkpoint.offset;\n-\t\tfree(idx);\n-\t} else {\n-\t\toidcpy(&idx->oid, result_oid);\n-\t\tALLOC_GROW(state->written,\n-\t\t\t   state->nr_written + 1,\n-\t\t\t   state->alloc_written);\n-\t\tstate->written[state->nr_written++] = idx;\n-\t}\n+\tfinalize_checkpoint(state, &ctx, &checkpoint, idx, result_oid);\n \treturn 0;\n }\n \n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483419","messageId":"da52ec838025a59a3f4f4ffaf2e6f9098a37547e.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 05/10] bulk-checkin: extract abstract `bulk_checkin_source`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:08:02Z","receivedAt":"2023-10-18T17:10:26Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A future commit will want to implement a very similar routine as in\n`stream_blob_to_pack()` with two notable changes:\n\n  - Instead of streaming just OBJ_BLOBs, this new function may want to\n    stream objects of arbitrary type.\n\n  - Instead of streaming the object's contents from an open\n    file-descriptor, this new function may want to \"stream\" its contents\n    from memory.\n\nTo avoid duplicating a significant chunk of code between the existing\n`stream_blob_to_pack()`, extract an abstract `bulk_checkin_source`. This\nconcept currently is a thin layer of `lseek()` and `read_in_full()`, but\nwill grow to understand how to perform analogous operations when writing\nout an object's contents from memory.\n\nSuggested-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 61 +++++++++++++++++++++++++++++++++++++++++++-------\n 1 file changed, 53 insertions(+), 8 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f4914fb6d1..fc1d902018 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -140,8 +140,41 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n \treturn 0;\n }\n \n+struct bulk_checkin_source {\n+\tenum { SOURCE_FILE } type;\n+\n+\t/* SOURCE_FILE fields */\n+\tint fd;\n+\n+\t/* common fields */\n+\tsize_t size;\n+\tconst char *path;\n+};\n+\n+static off_t bulk_checkin_source_seek_to(struct bulk_checkin_source *source,\n+\t\t\t\t\t off_t offset)\n+{\n+\tswitch (source->type) {\n+\tcase SOURCE_FILE:\n+\t\treturn lseek(source->fd, offset, SEEK_SET);\n+\tdefault:\n+\t\tBUG(\"unknown bulk-checkin source: %d\", source->type);\n+\t}\n+}\n+\n+static ssize_t bulk_checkin_source_read(struct bulk_checkin_source *source,\n+\t\t\t\t\tvoid *buf, size_t nr)\n+{\n+\tswitch (source->type) {\n+\tcase SOURCE_FILE:\n+\t\treturn read_in_full(source->fd, buf, nr);\n+\tdefault:\n+\t\tBUG(\"unknown bulk-checkin source: %d\", source->type);\n+\t}\n+}\n+\n /*\n- * Read the contents from fd for size bytes, streaming it to the\n+ * Read the contents from 'source' for 'size' bytes, streaming it to the\n  * packfile in state while updating the hash in ctx. Signal a failure\n  * by returning a negative value when the resulting pack would exceed\n  * the pack size limit and this is not the first object in the pack,\n@@ -157,7 +190,7 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n  */\n static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t       git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t       int fd, size_t size, const char *path,\n+\t\t\t       struct bulk_checkin_source *source,\n \t\t\t       unsigned flags)\n {\n \tgit_zstream s;\n@@ -167,22 +200,28 @@ static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n \tint status = Z_OK;\n \tint write_object = (flags & HASH_WRITE_OBJECT);\n \toff_t offset = 0;\n+\tsize_t size = source->size;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n-\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), OBJ_BLOB, size);\n+\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), OBJ_BLOB,\n+\t\t\t\t\t      size);\n \ts.next_out = obuf + hdrlen;\n \ts.avail_out = sizeof(obuf) - hdrlen;\n \n \twhile (status != Z_STREAM_END) {\n \t\tif (size && !s.avail_in) {\n \t\t\tssize_t rsize = size < sizeof(ibuf) ? size : sizeof(ibuf);\n-\t\t\tssize_t read_result = read_in_full(fd, ibuf, rsize);\n+\t\t\tssize_t read_result;\n+\n+\t\t\tread_result = bulk_checkin_source_read(source, ibuf,\n+\t\t\t\t\t\t\t       rsize);\n \t\t\tif (read_result < 0)\n-\t\t\t\tdie_errno(\"failed to read from '%s'\", path);\n+\t\t\t\tdie_errno(\"failed to read from '%s'\",\n+\t\t\t\t\t  source->path);\n \t\t\tif (read_result != rsize)\n \t\t\t\tdie(\"failed to read %d bytes from '%s'\",\n-\t\t\t\t    (int)rsize, path);\n+\t\t\t\t    (int)rsize, source->path);\n \t\t\toffset += rsize;\n \t\t\tif (*already_hashed_to < offset) {\n \t\t\t\tsize_t hsize = offset - *already_hashed_to;\n@@ -325,6 +364,12 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \tgit_hash_ctx ctx;\n \tstruct hashfile_checkpoint checkpoint = {0};\n \tstruct pack_idx_entry *idx = NULL;\n+\tstruct bulk_checkin_source source = {\n+\t\t.type = SOURCE_FILE,\n+\t\t.fd = fd,\n+\t\t.size = size,\n+\t\t.path = path,\n+\t};\n \n \tseekback = lseek(fd, 0, SEEK_CUR);\n \tif (seekback == (off_t) -1)\n@@ -342,10 +387,10 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \twhile (1) {\n \t\tprepare_checkpoint(state, &checkpoint, idx, flags);\n \t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t\t fd, size, path, flags))\n+\t\t\t\t\t &source, flags))\n \t\t\tbreak;\n \t\ttruncate_checkpoint(state, &checkpoint, idx);\n-\t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n+\t\tif (bulk_checkin_source_seek_to(&source, seekback) == (off_t)-1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n \tfinalize_checkpoint(state, &ctx, &checkpoint, idx, result_oid);\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483420","messageId":"4e9bac5bc1a49ca7a96aaee84a46b389c6bfe99b.1697648864.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697648864.git.me@ttaylorr.com","subject":"[PATCH v3 06/10] bulk-checkin: implement `SOURCE_INCORE` mode for `bulk_checkin_source`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T17:08:05Z","receivedAt":"2023-10-18T17:10:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Continue to prepare for streaming an object's contents directly from\nmemory by teaching `bulk_checkin_source` how to perform reads and seeks\nbased on an address in memory.\n\nUnlike file descriptors, which manage their own offset internally, we\nhave to keep track of how many bytes we've read out of the buffer, and\nmake sure we don't read past the end of the buffer.\n\nSuggested-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bulk-checkin.c | 18 +++++++++++++++++-\n 1 file changed, 17 insertions(+), 1 deletion(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex fc1d902018..133e02ce36 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -141,11 +141,15 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n }\n \n struct bulk_checkin_source {\n-\tenum { SOURCE_FILE } type;\n+\tenum { SOURCE_FILE, SOURCE_INCORE } type;\n \n \t/* SOURCE_FILE fields */\n \tint fd;\n \n+\t/* SOURCE_INCORE fields */\n+\tconst void *buf;\n+\tsize_t read;\n+\n \t/* common fields */\n \tsize_t size;\n \tconst char *path;\n@@ -157,6 +161,11 @@ static off_t bulk_checkin_source_seek_to(struct bulk_checkin_source *source,\n \tswitch (source->type) {\n \tcase SOURCE_FILE:\n \t\treturn lseek(source->fd, offset, SEEK_SET);\n+\tcase SOURCE_INCORE:\n+\t\tif (!(0 <= offset && offset < source->size))\n+\t\t\treturn (off_t)-1;\n+\t\tsource->read = offset;\n+\t\treturn source->read;\n \tdefault:\n \t\tBUG(\"unknown bulk-checkin source: %d\", source->type);\n \t}\n@@ -168,6 +177,13 @@ static ssize_t bulk_checkin_source_read(struct bulk_checkin_source *source,\n \tswitch (source->type) {\n \tcase SOURCE_FILE:\n \t\treturn read_in_full(source->fd, buf, nr);\n+\tcase SOURCE_INCORE:\n+\t\tassert(source->read <= source->size);\n+\t\tif (nr > source->size - source->read)\n+\t\t\tnr = source->size - source->read;\n+\t\tmemcpy(buf, (unsigned char *)source->buf + source->read, nr);\n+\t\tsource->read += nr;\n+\t\treturn nr;\n \tdefault:\n \t\tBUG(\"unknown bulk-checkin source: %d\", source->type);\n \t}\n-- \n2.42.0.408.g97fac66ae4\n\n"},{"id":"483426","messageId":"cover.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1696629697.git.me@ttaylorr.com","subject":"[PATCH v4 00/17] bloom: changed-path Bloom filters v2 (& sundries)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:22Z","receivedAt":"2023-10-18T18:32:31Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"(Rebased onto the tip of 'master', which is 3a06386e31 (The fifteenth\nbatch, 2023-10-04), at the time of writing).\n\nThis series is a reroll of the combined efforts of [1] and [2] to\nintroduce the v2 changed-path Bloom filters, which fixes a bug in our\nexisting implementation of murmur3 paths with non-ASCII characters (when\nthe \"char\" type is signed).\n\nIn large part, this is the same as the previous round. But this round\nincludes some extra bits that address issues pointed out by SZEDER\nGábor, which are:\n\n  - not reading Bloom filters for root commits\n  - corrupting Bloom filter reads by tweaking the filter settings\n    between layers.\n\nThese issues were discussed in (among other places) [3], and [4],\nrespectively.\n\nThanks to Jonathan, Peff, and SZEDER who have helped a great deal in\nassembling these patches. As usual, a range-diff is included below.\nThanks in advance for your\nreview!\n\n[1]: https://lore.kernel.org/git/cover.1684790529.git.jonathantanmy@google.com/\n[2]: https://lore.kernel.org/git/cover.1691426160.git.me@ttaylorr.com/\n[3]: https://public-inbox.org/git/20201015132147.GB24954@szeder.dev/\n[4]: https://lore.kernel.org/git/20230830200218.GA5147@szeder.dev/\n\nJonathan Tan (4):\n  gitformat-commit-graph: describe version 2 of BDAT\n  t4216: test changed path filters with high bit paths\n  repo-settings: introduce commitgraph.changedPathsVersion\n  commit-graph: new filter ver. that fixes murmur3\n\nTaylor Blau (13):\n  t/t4216-log-bloom.sh: harden `test_bloom_filters_not_used()`\n  revision.c: consult Bloom filters for root commits\n  commit-graph: ensure Bloom filters are read with consistent settings\n  t/helper/test-read-graph.c: extract `dump_graph_info()`\n  bloom.h: make `load_bloom_filter_from_graph()` public\n  t/helper/test-read-graph: implement `bloom-filters` mode\n  bloom: annotate filters with hash version\n  bloom: prepare to discard incompatible Bloom filters\n  commit-graph.c: unconditionally load Bloom filters\n  commit-graph: drop unnecessary `graph_read_bloom_data_context`\n  object.h: fix mis-aligned flag bits table\n  commit-graph: reuse existing Bloom filters where possible\n  bloom: introduce `deinit_bloom_filters()`\n\n Documentation/config/commitgraph.txt     |  26 ++-\n Documentation/gitformat-commit-graph.txt |   9 +-\n bloom.c                                  | 208 ++++++++++++++++-\n bloom.h                                  |  38 ++-\n commit-graph.c                           |  61 ++++-\n object.h                                 |   3 +-\n oss-fuzz/fuzz-commit-graph.c             |   2 +-\n repo-settings.c                          |   6 +-\n repository.h                             |   2 +-\n revision.c                               |  26 ++-\n t/helper/test-bloom.c                    |   9 +-\n t/helper/test-read-graph.c               |  67 ++++--\n t/t0095-bloom.sh                         |   8 +\n t/t4216-log-bloom.sh                     | 282 ++++++++++++++++++++++-\n 14 files changed, 692 insertions(+), 55 deletions(-)\n\nRange-diff against v3:\n 1:  fe671d616c =  1:  e0fc51c3fb t/t4216-log-bloom.sh: harden `test_bloom_filters_not_used()`\n 2:  7d0fa93543 =  2:  87b09e6266 revision.c: consult Bloom filters for root commits\n 3:  2ecc0a2d58 !  3:  46d8a41005 commit-graph: ensure Bloom filters are read with consistent settings\n    @@ t/t4216-log-bloom.sh: test_expect_success 'Bloom generation backfills empty comm\n     +\tdone\n     +'\n     +\n    -+test_expect_success 'split' '\n    ++test_expect_success 'ensure incompatible Bloom filters are ignored' '\n     +\t# Compute Bloom filters with \"unusual\" settings.\n     +\tgit -C $repo rev-parse one >in &&\n     +\tGIT_TEST_BLOOM_SETTINGS_NUM_HASHES=3 git -C $repo commit-graph write \\\n    @@ t/t4216-log-bloom.sh: test_expect_success 'Bloom generation backfills empty comm\n     +\n     +test_expect_success 'merge graph layers with incompatible Bloom settings' '\n     +\t# Ensure that incompatible Bloom filters are ignored when\n    -+\t# generating new layers.\n    ++\t# merging existing layers.\n     +\tgit -C $repo commit-graph write --reachable --changed-paths 2>err &&\n     +\tgrep \"disabling Bloom filters for commit-graph layer .$layer.\" err &&\n     +\n     +\ttest_path_is_file $repo/$graph &&\n     +\ttest_dir_is_empty $repo/$graphdir &&\n     +\n    -+\t# ...and merging existing ones.\n    -+\tgit -C $repo -c core.commitGraph=false log --oneline --no-decorate -- file \\\n    -+\t\t>expect 2>err &&\n    -+\tGIT_TRACE2_PERF=\"$(pwd)/trace.perf\" \\\n    ++\tgit -C $repo -c core.commitGraph=false log --oneline --no-decorate -- \\\n    ++\t\tfile >expect &&\n    ++\ttrace_out=\"$(pwd)/trace.perf\" &&\n    ++\tGIT_TRACE2_PERF=\"$trace_out\" \\\n     +\t\tgit -C $repo log --oneline --no-decorate -- file >actual 2>err &&\n     +\n    -+\ttest_cmp expect actual && cat err &&\n    -+\tgrep \"statistics:{\\\"filter_not_present\\\":0\" trace.perf &&\n    -+\t! grep \"disabling Bloom filters\" err\n    ++\ttest_cmp expect actual &&\n    ++\tgrep \"statistics:{\\\"filter_not_present\\\":0,\" trace.perf &&\n    ++\ttest_must_be_empty err\n     +'\n     +\n      test_done\n 4:  17703ed89a =  4:  4d0190a992 gitformat-commit-graph: describe version 2 of BDAT\n 5:  94552abf45 =  5:  3c2057c11c t/helper/test-read-graph.c: extract `dump_graph_info()`\n 6:  3d81efa27b =  6:  e002e35004 bloom.h: make `load_bloom_filter_from_graph()` public\n 7:  d23cd89037 =  7:  c7016f51cd t/helper/test-read-graph: implement `bloom-filters` mode\n 8:  cba766f224 !  8:  cef2aac8ba t4216: test changed path filters with high bit paths\n    @@ Commit message\n     \n      ## t/t4216-log-bloom.sh ##\n     @@ t/t4216-log-bloom.sh: test_expect_success 'merge graph layers with incompatible Bloom settings' '\n    - \t! grep \"disabling Bloom filters\" err\n    + \ttest_must_be_empty err\n      '\n      \n     +get_first_changed_path_filter () {\n    @@ t/t4216-log-bloom.sh: test_expect_success 'merge graph layers with incompatible\n     +\t(\n     +\t\tcd highbit1 &&\n     +\t\techo \"52a9\" >expect &&\n    -+\t\tget_first_changed_path_filter >actual &&\n    -+\t\ttest_cmp expect actual\n    ++\t\tget_first_changed_path_filter >actual\n     +\t)\n     +'\n     +\n 9:  a08a961f41 =  9:  36d4e2202e repo-settings: introduce commitgraph.changedPathsVersion\n10:  61d44519a5 ! 10:  f6ab427ead commit-graph: new filter ver. that fixes murmur3\n    @@ t/t4216-log-bloom.sh: test_expect_success 'version 1 changed-path used when vers\n     +\ttest_commit -C doublewrite c \"$CENT\" &&\n     +\tgit -C doublewrite config --add commitgraph.changedPathsVersion 1 &&\n     +\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n    ++\tfor v in -2 3\n    ++\tdo\n    ++\t\tgit -C doublewrite config --add commitgraph.changedPathsVersion $v &&\n    ++\t\tgit -C doublewrite commit-graph write --reachable --changed-paths 2>err &&\n    ++\t\tcat >expect <<-EOF &&\n    ++\t\twarning: attempting to write a commit-graph, but ${SQ}commitgraph.changedPathsVersion${SQ} ($v) is not supported\n    ++\t\tEOF\n    ++\t\ttest_cmp expect err || return 1\n    ++\tdone &&\n     +\tgit -C doublewrite config --add commitgraph.changedPathsVersion 2 &&\n     +\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n     +\t(\n11:  a8c10f8de8 = 11:  dc69b28329 bloom: annotate filters with hash version\n12:  2ba10a4b4b = 12:  85dbdc4ed2 bloom: prepare to discard incompatible Bloom filters\n13:  09d8669c3a = 13:  3ff669a622 commit-graph.c: unconditionally load Bloom filters\n14:  0d4f9dc4ee = 14:  1c78e3d178 commit-graph: drop unnecessary `graph_read_bloom_data_context`\n15:  1f7f27bc47 = 15:  a289514faa object.h: fix mis-aligned flag bits table\n16:  abbef95ae8 ! 16:  6a12e39e7f commit-graph: reuse existing Bloom filters where possible\n    @@ t/t4216-log-bloom.sh: test_expect_success 'when writing another commit graph, pr\n      \ttest_commit -C doublewrite c \"$CENT\" &&\n     +\n      \tgit -C doublewrite config --add commitgraph.changedPathsVersion 1 &&\n    --\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n     +\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n     +\t\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n     +\ttest_filter_computed 1 trace2.txt &&\n     +\ttest_filter_upgraded 0 trace2.txt &&\n    ++\n    + \tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n    + \tfor v in -2 3\n    + \tdo\n    +@@ t/t4216-log-bloom.sh: test_expect_success 'when writing commit graph, do not reuse changed-path of ano\n    + \t\tEOF\n    + \t\ttest_cmp expect err || return 1\n    + \tdone &&\n     +\n      \tgit -C doublewrite config --add commitgraph.changedPathsVersion 2 &&\n     -\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n17:  ca362408d5 ! 17:  8942f205c8 bloom: introduce `deinit_bloom_filters()`\n    @@ bloom.h: void add_key_to_filter(const struct bloom_key *key,\n      \tBLOOM_NOT_COMPUTED = (1 << 0),\n     \n      ## commit-graph.c ##\n    -@@ commit-graph.c: static void close_commit_graph_one(struct commit_graph *g)\n    +@@ commit-graph.c: struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n      void close_commit_graph(struct raw_object_store *o)\n      {\n    - \tclose_commit_graph_one(o->commit_graph);\n    + \tclear_commit_graph_data_slab(&commit_graph_data_slab);\n     +\tdeinit_bloom_filters();\n    + \tfree_commit_graph(o->commit_graph);\n      \to->commit_graph = NULL;\n      }\n    - \n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n      \n      \tres = write_commit_graph_file(ctx);\n-- \n2.42.0.415.g8942f205c8\n"},{"id":"483427","messageId":"e0fc51c3fb345c7e9ee3a64dca94e87ba2378382.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 01/17] t/t4216-log-bloom.sh: harden `test_bloom_filters_not_used()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:29Z","receivedAt":"2023-10-18T18:32:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The existing implementation of test_bloom_filters_not_used() asserts\nthat the Bloom filter sub-system has not been initialized at all, by\nchecking for the absence of any data from it from trace2.\n\nIn the following commit, it will become possible to load Bloom filters\nwithout using them (e.g., because `commitGraph.changedPathVersion` is\nincompatible with the hash version with which the commit-graph's Bloom\nfilters were written).\n\nWhen this is the case, it's possible to initialize the Bloom filter\nsub-system, while still not using any Bloom filters. When this is the\ncase, check that the data dump from the Bloom sub-system is all zeros,\nindicating that no filters were used.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/t4216-log-bloom.sh | 14 +++++++++++++-\n 1 file changed, 13 insertions(+), 1 deletion(-)\n\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex fa9d32facf..487fc3d6b9 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -81,7 +81,19 @@ test_bloom_filters_used () {\n test_bloom_filters_not_used () {\n \tlog_args=$1\n \tsetup \"$log_args\" &&\n-\t! grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\" &&\n+\n+\tif grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\"\n+\tthen\n+\t\t# if the Bloom filter system is initialized, ensure that no\n+\t\t# filters were used\n+\t\tdata=\"statistics:{\"\n+\t\tdata=\"$data\\\"filter_not_present\\\":0,\"\n+\t\tdata=\"$data\\\"maybe\\\":0,\"\n+\t\tdata=\"$data\\\"definitely_not\\\":0,\"\n+\t\tdata=\"$data\\\"false_positive\\\":0}\"\n+\n+\t\tgrep -q \"$data\" \"$TRASH_DIRECTORY/trace.perf\"\n+\tfi &&\n \ttest_cmp log_wo_bloom log_w_bloom\n }\n \n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483428","messageId":"87b09e6266a01e7fa4480d37f22e1ac3f4be6bc3.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 02/17] revision.c: consult Bloom filters for root commits","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:33Z","receivedAt":"2023-10-18T18:32:37Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The commit-graph stores changed-path Bloom filters which represent the\nset of paths included in a tree-level diff between a commit's root tree\nand that of its parent.\n\nWhen a commit has no parents, the tree-diff is computed against that\ncommit's root tree and the empty tree. In other words, every path in\nthat commit's tree is stored in the Bloom filter (since they all appear\nin the diff).\n\nConsult these filters during pathspec-limited traversals in the function\n`rev_same_tree_as_empty()`. Doing so yields a performance improvement\nwhere we can avoid enumerating the full set of paths in a parentless\ncommit's root tree when we know that the path(s) of interest were not\nlisted in that commit's changed-path Bloom filter.\n\nSuggested-by: SZEDER Gábor <szeder.dev@gmail.com>\nOriginal-patch-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n revision.c           | 26 ++++++++++++++++++++++----\n t/t4216-log-bloom.sh |  8 ++++++--\n 2 files changed, 28 insertions(+), 6 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 219dc76716..6e9da518d9 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -834,17 +834,28 @@ static int rev_compare_tree(struct rev_info *revs,\n \treturn tree_difference;\n }\n \n-static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit)\n+static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n+\t\t\t\t  int nth_parent)\n {\n \tstruct tree *t1 = repo_get_commit_tree(the_repository, commit);\n+\tint bloom_ret = 1;\n \n \tif (!t1)\n \t\treturn 0;\n \n+\tif (nth_parent == 1 && revs->bloom_keys_nr) {\n+\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n+\t\tif (!bloom_ret)\n+\t\t\treturn 1;\n+\t}\n+\n \ttree_difference = REV_TREE_SAME;\n \trevs->pruning.flags.has_changes = 0;\n \tdiff_tree_oid(NULL, &t1->object.oid, \"\", &revs->pruning);\n \n+\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n+\t\tcount_bloom_filter_false_positive++;\n+\n \treturn tree_difference == REV_TREE_SAME;\n }\n \n@@ -882,7 +893,7 @@ static int compact_treesame(struct rev_info *revs, struct commit *commit, unsign\n \t\tif (nth_parent != 0)\n \t\t\tdie(\"compact_treesame %u\", nth_parent);\n \t\told_same = !!(commit->object.flags & TREESAME);\n-\t\tif (rev_same_tree_as_empty(revs, commit))\n+\t\tif (rev_same_tree_as_empty(revs, commit, nth_parent))\n \t\t\tcommit->object.flags |= TREESAME;\n \t\telse\n \t\t\tcommit->object.flags &= ~TREESAME;\n@@ -978,7 +989,14 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \t\treturn;\n \n \tif (!commit->parents) {\n-\t\tif (rev_same_tree_as_empty(revs, commit))\n+\t\t/*\n+\t\t * Pretend as if we are comparing ourselves to the\n+\t\t * (non-existent) first parent of this commit object. Even\n+\t\t * though no such parent exists, its changed-path Bloom filter\n+\t\t * (if one exists) is relative to the empty tree, using Bloom\n+\t\t * filters is allowed here.\n+\t\t */\n+\t\tif (rev_same_tree_as_empty(revs, commit, 1))\n \t\t\tcommit->object.flags |= TREESAME;\n \t\treturn;\n \t}\n@@ -1059,7 +1077,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \n \t\tcase REV_TREE_NEW:\n \t\t\tif (revs->remove_empty_trees &&\n-\t\t\t    rev_same_tree_as_empty(revs, p)) {\n+\t\t\t    rev_same_tree_as_empty(revs, p, nth_parent)) {\n \t\t\t\t/* We are adding all the specified\n \t\t\t\t * paths from this parent, so the\n \t\t\t\t * history beyond this parent is not\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 487fc3d6b9..322640feeb 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -87,7 +87,11 @@ test_bloom_filters_not_used () {\n \t\t# if the Bloom filter system is initialized, ensure that no\n \t\t# filters were used\n \t\tdata=\"statistics:{\"\n-\t\tdata=\"$data\\\"filter_not_present\\\":0,\"\n+\t\t# unusable filters (e.g., those computed with a\n+\t\t# different value of commitGraph.changedPathsVersion)\n+\t\t# are counted in the filter_not_present bucket, so any\n+\t\t# value is OK there.\n+\t\tdata=\"$data\\\"filter_not_present\\\":[0-9][0-9]*,\"\n \t\tdata=\"$data\\\"maybe\\\":0,\"\n \t\tdata=\"$data\\\"definitely_not\\\":0,\"\n \t\tdata=\"$data\\\"false_positive\\\":0}\"\n@@ -174,7 +178,7 @@ test_expect_success 'setup - add commit-graph to the chain with Bloom filters' '\n \n test_bloom_filters_used_when_some_filters_are_missing () {\n \tlog_args=$1\n-\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"maybe\\\":6,\\\"definitely_not\\\":9\"\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"maybe\\\":6,\\\"definitely_not\\\":10\"\n \tsetup \"$log_args\" &&\n \tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" &&\n \ttest_cmp log_wo_bloom log_w_bloom\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483429","messageId":"46d8a41005a0431a4f03b5fced7e1bb3705d5ed9.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 03/17] commit-graph: ensure Bloom filters are read with consistent settings","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:36Z","receivedAt":"2023-10-18T18:32:41Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The changed-path Bloom filter mechanism is parameterized by a couple of\nvariables, notably the number of bits per hash (typically \"m\" in Bloom\nfilter literature) and the number of hashes themselves (typically \"k\").\n\nIt is critically important that filters are read with the Bloom filter\nsettings that they were written with. Failing to do so would mean that\neach query is liable to compute different fingerprints, meaning that the\nfilter itself could return a false negative. This goes against a basic\nassumption of using Bloom filters (that they may return false positives,\nbut never false negatives) and can lead to incorrect results.\n\nWe have some existing logic to carry forward existing Bloom filter\nsettings from one layer to the next. In `write_commit_graph()`, we have\nsomething like:\n\n    if (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\n        struct commit_graph *g = ctx->r->objects->commit_graph;\n\n        /* We have changed-paths already. Keep them in the next graph */\n        if (g && g->chunk_bloom_data) {\n            ctx->changed_paths = 1;\n            ctx->bloom_settings = g->bloom_filter_settings;\n        }\n    }\n\n, which drags forward Bloom filter settings across adjacent layers.\n\nThis doesn't quite address all cases, however, since it is possible for\nintermediate layers to contain no Bloom filters at all. For example,\nsuppose we have two layers in a commit-graph chain, say, {G1, G2}. If G1\ncontains Bloom filters, but G2 doesn't, a new G3 (whose base graph is\nG2) may be written with arbitrary Bloom filter settings, because we only\ncheck the immediately adjacent layer's settings for compatibility.\n\nThis behavior has existed since the introduction of changed-path Bloom\nfilters. But in practice, this is not such a big deal, since the only\nway up until this point to modify the Bloom filter settings at write\ntime is with the undocumented environment variables:\n\n  - GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\n  - GIT_TEST_BLOOM_SETTINGS_NUM_HASHES\n  - GIT_TEST_BLOOM_SETTINGS_MAX_CHANGED_PATHS\n\n(it is still possible to tweak MAX_CHANGED_PATHS between layers, but\nthis does not affect reads, so is allowed to differ across multiple\ngraph layers).\n\nBut in future commits, we will introduce another parameter to change the\nhash algorithm used to compute Bloom fingerprints itself. This will be\nexposed via a configuration setting, making this foot-gun easier to use.\n\nTo prevent this potential issue, validate that all layers of a split\ncommit-graph have compatible settings with the newest layer which\ncontains Bloom filters.\n\nReported-by: SZEDER Gábor <szeder.dev@gmail.com>\nOriginal-test-by: SZEDER Gábor <szeder.dev@gmail.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n commit-graph.c       | 25 +++++++++++++++++\n t/t4216-log-bloom.sh | 64 ++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 89 insertions(+)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex fd2f700b2e..0ac79aff5a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -498,6 +498,30 @@ static int validate_mixed_generation_chain(struct commit_graph *g)\n \treturn 0;\n }\n \n+static void validate_mixed_bloom_settings(struct commit_graph *g)\n+{\n+\tstruct bloom_filter_settings *settings = NULL;\n+\tfor (; g; g = g->base_graph) {\n+\t\tif (!g->bloom_filter_settings)\n+\t\t\tcontinue;\n+\t\tif (!settings) {\n+\t\t\tsettings = g->bloom_filter_settings;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (g->bloom_filter_settings->bits_per_entry != settings->bits_per_entry ||\n+\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes) {\n+\t\t\tg->chunk_bloom_indexes = NULL;\n+\t\t\tg->chunk_bloom_data = NULL;\n+\t\t\tFREE_AND_NULL(g->bloom_filter_settings);\n+\n+\t\t\twarning(_(\"disabling Bloom filters for commit-graph \"\n+\t\t\t\t  \"layer '%s' due to incompatible settings\"),\n+\t\t\t\toid_to_hex(&g->oid));\n+\t\t}\n+\t}\n+}\n+\n static int add_graph_to_chain(struct commit_graph *g,\n \t\t\t      struct commit_graph *chain,\n \t\t\t      struct object_id *oids,\n@@ -616,6 +640,7 @@ struct commit_graph *load_commit_graph_chain_fd_st(struct repository *r,\n \t}\n \n \tvalidate_mixed_generation_chain(graph_chain);\n+\tvalidate_mixed_bloom_settings(graph_chain);\n \n \tfree(oids);\n \tfclose(fp);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 322640feeb..2ea5e90955 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -420,4 +420,68 @@ test_expect_success 'Bloom generation backfills empty commits' '\n \t)\n '\n \n+graph=.git/objects/info/commit-graph\n+graphdir=.git/objects/info/commit-graphs\n+chain=$graphdir/commit-graph-chain\n+\n+test_expect_success 'setup for mixed Bloom setting tests' '\n+\trepo=mixed-bloom-settings &&\n+\n+\tgit init $repo &&\n+\tfor i in one two three\n+\tdo\n+\t\ttest_commit -C $repo $i file || return 1\n+\tdone\n+'\n+\n+test_expect_success 'ensure incompatible Bloom filters are ignored' '\n+\t# Compute Bloom filters with \"unusual\" settings.\n+\tgit -C $repo rev-parse one >in &&\n+\tGIT_TEST_BLOOM_SETTINGS_NUM_HASHES=3 git -C $repo commit-graph write \\\n+\t\t--stdin-commits --changed-paths --split <in &&\n+\tlayer=$(head -n 1 $repo/$chain) &&\n+\n+\t# A commit-graph layer without Bloom filters \"hides\" the layers\n+\t# below ...\n+\tgit -C $repo rev-parse two >in &&\n+\tgit -C $repo commit-graph write --stdin-commits --no-changed-paths \\\n+\t\t--split=no-merge <in &&\n+\n+\t# Another commit-graph layer that has Bloom filters, but with\n+\t# standard settings, and is thus incompatible with the base\n+\t# layer written above.\n+\tgit -C $repo rev-parse HEAD >in &&\n+\tgit -C $repo commit-graph write --stdin-commits --changed-paths \\\n+\t\t--split=no-merge <in &&\n+\n+\ttest_line_count = 3 $repo/$chain &&\n+\n+\t# Ensure that incompatible Bloom filters are ignored.\n+\tgit -C $repo -c core.commitGraph=false log --oneline --no-decorate -- file \\\n+\t\t>expect 2>err &&\n+\tgit -C $repo log --oneline --no-decorate -- file >actual 2>err &&\n+\ttest_cmp expect actual &&\n+\tgrep \"disabling Bloom filters for commit-graph layer .$layer.\" err\n+'\n+\n+test_expect_success 'merge graph layers with incompatible Bloom settings' '\n+\t# Ensure that incompatible Bloom filters are ignored when\n+\t# merging existing layers.\n+\tgit -C $repo commit-graph write --reachable --changed-paths 2>err &&\n+\tgrep \"disabling Bloom filters for commit-graph layer .$layer.\" err &&\n+\n+\ttest_path_is_file $repo/$graph &&\n+\ttest_dir_is_empty $repo/$graphdir &&\n+\n+\tgit -C $repo -c core.commitGraph=false log --oneline --no-decorate -- \\\n+\t\tfile >expect &&\n+\ttrace_out=\"$(pwd)/trace.perf\" &&\n+\tGIT_TRACE2_PERF=\"$trace_out\" \\\n+\t\tgit -C $repo log --oneline --no-decorate -- file >actual 2>err &&\n+\n+\ttest_cmp expect actual &&\n+\tgrep \"statistics:{\\\"filter_not_present\\\":0,\" trace.perf &&\n+\ttest_must_be_empty err\n+'\n+\n test_done\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483430","messageId":"4d0190a9926bddc691cfa5b856d02b7bcc3a1d81.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 04/17] gitformat-commit-graph: describe version 2 of BDAT","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:40Z","receivedAt":"2023-10-18T18:32:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jonathan Tan <jonathantanmy@google.com>\n\nThe code change to Git to support version 2 will be done in subsequent\ncommits.\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/gitformat-commit-graph.txt | 9 ++++++---\n 1 file changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/gitformat-commit-graph.txt b/Documentation/gitformat-commit-graph.txt\nindex 31cad585e2..3e906e8030 100644\n--- a/Documentation/gitformat-commit-graph.txt\n+++ b/Documentation/gitformat-commit-graph.txt\n@@ -142,13 +142,16 @@ All multi-byte numbers are in network byte order.\n \n ==== Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n     * It starts with header consisting of three unsigned 32-bit integers:\n-      - Version of the hash algorithm being used. We currently only support\n-\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n+      - Version of the hash algorithm being used. We currently support\n+\tvalue 2 which corresponds to the 32-bit version of the murmur3 hash\n \timplemented exactly as described in\n \thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n \thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n \tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n-\tin Probabilistic Verification\"\n+\tin Probabilistic Verification\". Version 1 Bloom filters have a bug that appears\n+\twhen char is signed and the repository has path names that have characters >=\n+\t0x80; Git supports reading and writing them, but this ability will be removed\n+\tin a future version of Git.\n       - The number of times a path is hashed and hence the number of bit positions\n \t      that cumulatively determine whether a file is present in the commit.\n       - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483431","messageId":"3c2057c11c9229794e9d410e34260b2e92b2907e.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 05/17] t/helper/test-read-graph.c: extract `dump_graph_info()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:43Z","receivedAt":"2023-10-18T18:32:47Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Prepare for the 'read-graph' test helper to perform other tasks besides\ndumping high-level information about the commit-graph by extracting its\nmain routine into a separate function.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/helper/test-read-graph.c | 31 ++++++++++++++++++-------------\n 1 file changed, 18 insertions(+), 13 deletions(-)\n\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 8c7a83f578..3375392f6c 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -5,20 +5,8 @@\n #include \"bloom.h\"\n #include \"setup.h\"\n \n-int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n+static void dump_graph_info(struct commit_graph *graph)\n {\n-\tstruct commit_graph *graph = NULL;\n-\tstruct object_directory *odb;\n-\n-\tsetup_git_directory();\n-\todb = the_repository->objects->odb;\n-\n-\tprepare_repo_settings(the_repository);\n-\n-\tgraph = read_commit_graph_one(the_repository, odb);\n-\tif (!graph)\n-\t\treturn 1;\n-\n \tprintf(\"header: %08x %d %d %d %d\\n\",\n \t\tntohl(*(uint32_t*)graph->data),\n \t\t*(unsigned char*)(graph->data + 4),\n@@ -57,6 +45,23 @@ int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n \tif (graph->topo_levels)\n \t\tprintf(\" topo_levels\");\n \tprintf(\"\\n\");\n+}\n+\n+int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n+{\n+\tstruct commit_graph *graph = NULL;\n+\tstruct object_directory *odb;\n+\n+\tsetup_git_directory();\n+\todb = the_repository->objects->odb;\n+\n+\tprepare_repo_settings(the_repository);\n+\n+\tgraph = read_commit_graph_one(the_repository, odb);\n+\tif (!graph)\n+\t\treturn 1;\n+\n+\tdump_graph_info(graph);\n \n \tUNLEAK(graph);\n \n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483432","messageId":"e002e350044ecf2b141ba2c71b6ce81fadeeefc4.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 06/17] bloom.h: make `load_bloom_filter_from_graph()` public","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:46Z","receivedAt":"2023-10-18T18:32:50Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Prepare for a future commit to use the load_bloom_filter_from_graph()\nfunction directly to load specific Bloom filters out of the commit-graph\nfor manual inspection (to be used during tests).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c | 6 +++---\n bloom.h | 5 +++++\n 2 files changed, 8 insertions(+), 3 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex aef6b5fea2..3e78cfe79d 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -29,9 +29,9 @@ static inline unsigned char get_bitmask(uint32_t pos)\n \treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n }\n \n-static int load_bloom_filter_from_graph(struct commit_graph *g,\n-\t\t\t\t\tstruct bloom_filter *filter,\n-\t\t\t\t\tuint32_t graph_pos)\n+int load_bloom_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t struct bloom_filter *filter,\n+\t\t\t\t uint32_t graph_pos)\n {\n \tuint32_t lex_pos, start_index, end_index;\n \ndiff --git a/bloom.h b/bloom.h\nindex adde6dfe21..1e4f612d2c 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -3,6 +3,7 @@\n \n struct commit;\n struct repository;\n+struct commit_graph;\n \n struct bloom_filter_settings {\n \t/*\n@@ -68,6 +69,10 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+int load_bloom_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t struct bloom_filter *filter,\n+\t\t\t\t uint32_t graph_pos);\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483433","messageId":"c7016f51cddf892fa96e40db896e8fe96281ffd9.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 07/17] t/helper/test-read-graph: implement `bloom-filters` mode","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:49Z","receivedAt":"2023-10-18T18:32:53Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Implement a mode of the \"read-graph\" test helper to dump out the\nhexadecimal contents of the Bloom filter(s) contained in a commit-graph.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/helper/test-read-graph.c | 44 +++++++++++++++++++++++++++++++++-----\n 1 file changed, 39 insertions(+), 5 deletions(-)\n\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 3375392f6c..da9ac8584d 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -47,10 +47,32 @@ static void dump_graph_info(struct commit_graph *graph)\n \tprintf(\"\\n\");\n }\n \n-int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n+static void dump_graph_bloom_filters(struct commit_graph *graph)\n+{\n+\tuint32_t i;\n+\n+\tfor (i = 0; i < graph->num_commits + graph->num_commits_in_base; i++) {\n+\t\tstruct bloom_filter filter = { 0 };\n+\t\tsize_t j;\n+\n+\t\tif (load_bloom_filter_from_graph(graph, &filter, i) < 0) {\n+\t\t\tfprintf(stderr, \"missing Bloom filter for graph \"\n+\t\t\t\t\"position %\"PRIu32\"\\n\", i);\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tfor (j = 0; j < filter.len; j++)\n+\t\t\tprintf(\"%02x\", filter.data[j]);\n+\t\tif (filter.len)\n+\t\t\tprintf(\"\\n\");\n+\t}\n+}\n+\n+int cmd__read_graph(int argc, const char **argv)\n {\n \tstruct commit_graph *graph = NULL;\n \tstruct object_directory *odb;\n+\tint ret = 0;\n \n \tsetup_git_directory();\n \todb = the_repository->objects->odb;\n@@ -58,12 +80,24 @@ int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n \tprepare_repo_settings(the_repository);\n \n \tgraph = read_commit_graph_one(the_repository, odb);\n-\tif (!graph)\n-\t\treturn 1;\n+\tif (!graph) {\n+\t\tret = 1;\n+\t\tgoto done;\n+\t}\n \n-\tdump_graph_info(graph);\n+\tif (argc <= 1)\n+\t\tdump_graph_info(graph);\n+\telse if (!strcmp(argv[1], \"bloom-filters\"))\n+\t\tdump_graph_bloom_filters(graph);\n+\telse {\n+\t\tfprintf(stderr, \"unknown sub-command: '%s'\\n\", argv[1]);\n+\t\tret = 1;\n+\t}\n \n+done:\n \tUNLEAK(graph);\n \n-\treturn 0;\n+\treturn ret;\n }\n+\n+\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483434","messageId":"cef2aac8ba01051ebc6194a5dea28964f76a5243.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 08/17] t4216: test changed path filters with high bit paths","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:52Z","receivedAt":"2023-10-18T18:32:56Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jonathan Tan <jonathantanmy@google.com>\n\nSubsequent commits will teach Git another version of changed path\nfilter that has different behavior with paths that contain at least\none character with its high bit set, so test the existing behavior as\na baseline.\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/t4216-log-bloom.sh | 51 ++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 51 insertions(+)\n\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 2ea5e90955..400dce2193 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -484,4 +484,55 @@ test_expect_success 'merge graph layers with incompatible Bloom settings' '\n \ttest_must_be_empty err\n '\n \n+get_first_changed_path_filter () {\n+\ttest-tool read-graph bloom-filters >filters.dat &&\n+\thead -n 1 filters.dat\n+}\n+\n+# chosen to be the same under all Unicode normalization forms\n+CENT=$(printf \"\\302\\242\")\n+\n+test_expect_success 'set up repo with high bit path, version 1 changed-path' '\n+\tgit init highbit1 &&\n+\ttest_commit -C highbit1 c1 \"$CENT\" &&\n+\tgit -C highbit1 commit-graph write --reachable --changed-paths\n+'\n+\n+test_expect_success 'setup check value of version 1 changed-path' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\techo \"52a9\" >expect &&\n+\t\tget_first_changed_path_filter >actual\n+\t)\n+'\n+\n+# expect will not match actual if char is unsigned by default. Write the test\n+# in this way, so that a user running this test script can still see if the two\n+# files match. (It will appear as an ordinary success if they match, and a skip\n+# if not.)\n+if test_cmp highbit1/expect highbit1/actual\n+then\n+\ttest_set_prereq SIGNED_CHAR_BY_DEFAULT\n+fi\n+test_expect_success SIGNED_CHAR_BY_DEFAULT 'check value of version 1 changed-path' '\n+\t# Only the prereq matters for this test.\n+\ttrue\n+'\n+\n+test_expect_success 'setup make another commit' '\n+\t# \"git log\" does not use Bloom filters for root commits - see how, in\n+\t# revision.c, rev_compare_tree() (the only code path that eventually calls\n+\t# get_bloom_filter()) is only called by try_to_simplify_commit() when the commit\n+\t# has one parent. Therefore, make another commit so that we perform the tests on\n+\t# a non-root commit.\n+\ttest_commit -C highbit1 anotherc1 \"another$CENT\"\n+'\n+\n+test_expect_success 'version 1 changed-path used when version 1 requested' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n test_done\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483435","messageId":"36d4e2202e88aa61a2d7a76df33395186e6b71be.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 09/17] repo-settings: introduce commitgraph.changedPathsVersion","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:56Z","receivedAt":"2023-10-18T18:33:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jonathan Tan <jonathantanmy@google.com>\n\nA subsequent commit will introduce another version of the changed-path\nfilter in the commit graph file. In order to control which version to\nwrite (and read), a config variable is needed.\n\nTherefore, introduce this config variable. For forwards compatibility,\nteach Git to not read commit graphs when the config variable\nis set to an unsupported version. Because we teach Git this,\ncommitgraph.readChangedPaths is now redundant, so deprecate it and\ndefine its behavior in terms of the config variable we introduce.\n\nThis commit does not change the behavior of writing (Git writes changed\npath filters when explicitly instructed regardless of any config\nvariable), but a subsequent commit will restrict Git such that it will\nonly write when commitgraph.changedPathsVersion is a recognized value.\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/commitgraph.txt | 23 ++++++++++++++++++++---\n commit-graph.c                       |  2 +-\n oss-fuzz/fuzz-commit-graph.c         |  2 +-\n repo-settings.c                      |  6 +++++-\n repository.h                         |  2 +-\n 5 files changed, 28 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config/commitgraph.txt b/Documentation/config/commitgraph.txt\nindex 30604e4a4c..2dc9170622 100644\n--- a/Documentation/config/commitgraph.txt\n+++ b/Documentation/config/commitgraph.txt\n@@ -9,6 +9,23 @@ commitGraph.maxNewFilters::\n \tcommit-graph write` (c.f., linkgit:git-commit-graph[1]).\n \n commitGraph.readChangedPaths::\n-\tIf true, then git will use the changed-path Bloom filters in the\n-\tcommit-graph file (if it exists, and they are present). Defaults to\n-\ttrue. See linkgit:git-commit-graph[1] for more information.\n+\tDeprecated. Equivalent to commitGraph.changedPathsVersion=-1 if true, and\n+\tcommitGraph.changedPathsVersion=0 if false. (If commitGraph.changedPathVersion\n+\tis also set, commitGraph.changedPathsVersion takes precedence.)\n+\n+commitGraph.changedPathsVersion::\n+\tSpecifies the version of the changed-path Bloom filters that Git will read and\n+\twrite. May be -1, 0 or 1.\n++\n+Defaults to -1.\n++\n+If -1, Git will use the version of the changed-path Bloom filters in the\n+repository, defaulting to 1 if there are none.\n++\n+If 0, Git will not read any Bloom filters, and will write version 1 Bloom\n+filters when instructed to write.\n++\n+If 1, Git will only read version 1 Bloom filters, and will write version 1\n+Bloom filters.\n++\n+See linkgit:git-commit-graph[1] for more information.\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 0ac79aff5a..bcc9a15cfa 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -411,7 +411,7 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s,\n \t\t\tgraph->read_generation_data = 1;\n \t}\n \n-\tif (s->commit_graph_read_changed_paths) {\n+\tif (s->commit_graph_changed_paths_version) {\n \t\tpair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n \t\t\t   &graph->chunk_bloom_indexes);\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\ndiff --git a/oss-fuzz/fuzz-commit-graph.c b/oss-fuzz/fuzz-commit-graph.c\nindex 2992079dd9..325c0b991a 100644\n--- a/oss-fuzz/fuzz-commit-graph.c\n+++ b/oss-fuzz/fuzz-commit-graph.c\n@@ -19,7 +19,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n \t * possible.\n \t */\n \tthe_repository->settings.commit_graph_generation_version = 2;\n-\tthe_repository->settings.commit_graph_read_changed_paths = 1;\n+\tthe_repository->settings.commit_graph_changed_paths_version = 1;\n \tg = parse_commit_graph(&the_repository->settings, (void *)data, size);\n \trepo_clear(the_repository);\n \tfree_commit_graph(g);\ndiff --git a/repo-settings.c b/repo-settings.c\nindex 525f69c0c7..db8fe817f3 100644\n--- a/repo-settings.c\n+++ b/repo-settings.c\n@@ -24,6 +24,7 @@ void prepare_repo_settings(struct repository *r)\n \tint value;\n \tconst char *strval;\n \tint manyfiles;\n+\tint read_changed_paths;\n \n \tif (!r->gitdir)\n \t\tBUG(\"Cannot add settings for uninitialized repository\");\n@@ -54,7 +55,10 @@ void prepare_repo_settings(struct repository *r)\n \t/* Commit graph config or default, does not cascade (simple) */\n \trepo_cfg_bool(r, \"core.commitgraph\", &r->settings.core_commit_graph, 1);\n \trepo_cfg_int(r, \"commitgraph.generationversion\", &r->settings.commit_graph_generation_version, 2);\n-\trepo_cfg_bool(r, \"commitgraph.readchangedpaths\", &r->settings.commit_graph_read_changed_paths, 1);\n+\trepo_cfg_bool(r, \"commitgraph.readchangedpaths\", &read_changed_paths, 1);\n+\trepo_cfg_int(r, \"commitgraph.changedpathsversion\",\n+\t\t     &r->settings.commit_graph_changed_paths_version,\n+\t\t     read_changed_paths ? -1 : 0);\n \trepo_cfg_bool(r, \"gc.writecommitgraph\", &r->settings.gc_write_commit_graph, 1);\n \trepo_cfg_bool(r, \"fetch.writecommitgraph\", &r->settings.fetch_write_commit_graph, 0);\n \ndiff --git a/repository.h b/repository.h\nindex 5f18486f64..f71154e12c 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -29,7 +29,7 @@ struct repo_settings {\n \n \tint core_commit_graph;\n \tint commit_graph_generation_version;\n-\tint commit_graph_read_changed_paths;\n+\tint commit_graph_changed_paths_version;\n \tint gc_write_commit_graph;\n \tint fetch_write_commit_graph;\n \tint command_requires_full_index;\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483437","messageId":"f6ab427ead86bc82284b2c721f3c177947ece3c9.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 10/17] commit-graph: new filter ver. that fixes murmur3","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:32:59Z","receivedAt":"2023-10-18T18:33:06Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jonathan Tan <jonathantanmy@google.com>\n\nThe murmur3 implementation in bloom.c has a bug when converting series\nof 4 bytes into network-order integers when char is signed (which is\ncontrollable by a compiler option, and the default signedness of char is\nplatform-specific). When a string contains characters with the high bit\nset, this bug causes results that, although internally consistent within\nGit, does not accord with other implementations of murmur3 (thus,\nthe changed path filters wouldn't be readable by other off-the-shelf\nimplementatios of murmur3) and even with Git binaries that were compiled\nwith different signedness of char. This bug affects both how Git writes\nchanged path filters to disk and how Git interprets changed path filters\non disk.\n\nTherefore, introduce a new version (2) of changed path filters that\ncorrects this problem. The existing version (1) is still supported and\nis still the default, but users should migrate away from it as soon\nas possible.\n\nBecause this bug only manifests with characters that have the high bit\nset, it may be possible that some (or all) commits in a given repo would\nhave the same changed path filter both before and after this fix is\napplied. However, in order to determine whether this is the case, the\nchanged paths would first have to be computed, at which point it is not\nmuch more expensive to just compute a new changed path filter.\n\nSo this patch does not include any mechanism to \"salvage\" changed path\nfilters from repositories. There is also no \"mixed\" mode - for each\ninvocation of Git, reading and writing changed path filters are done\nwith the same version number; this version number may be explicitly\nstated (typically if the user knows which version they need) or\nautomatically determined from the version of the existing changed path\nfilters in the repository.\n\nThere is a change in write_commit_graph(). graph_read_bloom_data()\nmakes it possible for chunk_bloom_data to be non-NULL but\nbloom_filter_settings to be NULL, which causes a segfault later on. I\nproduced such a segfault while developing this patch, but couldn't find\na way to reproduce it neither after this complete patch (or before),\nbut in any case it seemed like a good thing to include that might help\nfuture patch authors.\n\nThe value in t0095 was obtained from another murmur3 implementation\nusing the following Go source code:\n\n  package main\n\n  import \"fmt\"\n  import \"github.com/spaolacci/murmur3\"\n\n  func main() {\n          fmt.Printf(\"%x\\n\", murmur3.Sum32([]byte(\"Hello world!\")))\n          fmt.Printf(\"%x\\n\", murmur3.Sum32([]byte{0x99, 0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff}))\n  }\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/commitgraph.txt |   5 +-\n bloom.c                              |  69 +++++++++++++++-\n bloom.h                              |   8 +-\n commit-graph.c                       |  32 ++++++--\n t/helper/test-bloom.c                |   9 ++-\n t/t0095-bloom.sh                     |   8 ++\n t/t4216-log-bloom.sh                 | 114 +++++++++++++++++++++++++++\n 7 files changed, 232 insertions(+), 13 deletions(-)\n\ndiff --git a/Documentation/config/commitgraph.txt b/Documentation/config/commitgraph.txt\nindex 2dc9170622..acc74a2f27 100644\n--- a/Documentation/config/commitgraph.txt\n+++ b/Documentation/config/commitgraph.txt\n@@ -15,7 +15,7 @@ commitGraph.readChangedPaths::\n \n commitGraph.changedPathsVersion::\n \tSpecifies the version of the changed-path Bloom filters that Git will read and\n-\twrite. May be -1, 0 or 1.\n+\twrite. May be -1, 0, 1, or 2.\n +\n Defaults to -1.\n +\n@@ -28,4 +28,7 @@ filters when instructed to write.\n If 1, Git will only read version 1 Bloom filters, and will write version 1\n Bloom filters.\n +\n+If 2, Git will only read version 2 Bloom filters, and will write version 2\n+Bloom filters.\n++\n See linkgit:git-commit-graph[1] for more information.\ndiff --git a/bloom.c b/bloom.c\nindex 3e78cfe79d..ebef5cfd2f 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -66,7 +66,64 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n  * Not considered to be cryptographically secure.\n  * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  */\n-uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len)\n+uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n+{\n+\tconst uint32_t c1 = 0xcc9e2d51;\n+\tconst uint32_t c2 = 0x1b873593;\n+\tconst uint32_t r1 = 15;\n+\tconst uint32_t r2 = 13;\n+\tconst uint32_t m = 5;\n+\tconst uint32_t n = 0xe6546b64;\n+\tint i;\n+\tuint32_t k1 = 0;\n+\tconst char *tail;\n+\n+\tint len4 = len / sizeof(uint32_t);\n+\n+\tuint32_t k;\n+\tfor (i = 0; i < len4; i++) {\n+\t\tuint32_t byte1 = (uint32_t)(unsigned char)data[4*i];\n+\t\tuint32_t byte2 = ((uint32_t)(unsigned char)data[4*i + 1]) << 8;\n+\t\tuint32_t byte3 = ((uint32_t)(unsigned char)data[4*i + 2]) << 16;\n+\t\tuint32_t byte4 = ((uint32_t)(unsigned char)data[4*i + 3]) << 24;\n+\t\tk = byte1 | byte2 | byte3 | byte4;\n+\t\tk *= c1;\n+\t\tk = rotate_left(k, r1);\n+\t\tk *= c2;\n+\n+\t\tseed ^= k;\n+\t\tseed = rotate_left(seed, r2) * m + n;\n+\t}\n+\n+\ttail = (data + len4 * sizeof(uint32_t));\n+\n+\tswitch (len & (sizeof(uint32_t) - 1)) {\n+\tcase 3:\n+\t\tk1 ^= ((uint32_t)(unsigned char)tail[2]) << 16;\n+\t\t/*-fallthrough*/\n+\tcase 2:\n+\t\tk1 ^= ((uint32_t)(unsigned char)tail[1]) << 8;\n+\t\t/*-fallthrough*/\n+\tcase 1:\n+\t\tk1 ^= ((uint32_t)(unsigned char)tail[0]) << 0;\n+\t\tk1 *= c1;\n+\t\tk1 = rotate_left(k1, r1);\n+\t\tk1 *= c2;\n+\t\tseed ^= k1;\n+\t\tbreak;\n+\t}\n+\n+\tseed ^= (uint32_t)len;\n+\tseed ^= (seed >> 16);\n+\tseed *= 0x85ebca6b;\n+\tseed ^= (seed >> 13);\n+\tseed *= 0xc2b2ae35;\n+\tseed ^= (seed >> 16);\n+\n+\treturn seed;\n+}\n+\n+static uint32_t murmur3_seeded_v1(uint32_t seed, const char *data, size_t len)\n {\n \tconst uint32_t c1 = 0xcc9e2d51;\n \tconst uint32_t c2 = 0x1b873593;\n@@ -131,8 +188,14 @@ void fill_bloom_key(const char *data,\n \tint i;\n \tconst uint32_t seed0 = 0x293ae76f;\n \tconst uint32_t seed1 = 0x7e646e2c;\n-\tconst uint32_t hash0 = murmur3_seeded(seed0, data, len);\n-\tconst uint32_t hash1 = murmur3_seeded(seed1, data, len);\n+\tuint32_t hash0, hash1;\n+\tif (settings->hash_version == 2) {\n+\t\thash0 = murmur3_seeded_v2(seed0, data, len);\n+\t\thash1 = murmur3_seeded_v2(seed1, data, len);\n+\t} else {\n+\t\thash0 = murmur3_seeded_v1(seed0, data, len);\n+\t\thash1 = murmur3_seeded_v1(seed1, data, len);\n+\t}\n \n \tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n \tfor (i = 0; i < settings->num_hashes; i++)\ndiff --git a/bloom.h b/bloom.h\nindex 1e4f612d2c..138d57a86b 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -8,9 +8,11 @@ struct commit_graph;\n struct bloom_filter_settings {\n \t/*\n \t * The version of the hashing technique being used.\n-\t * We currently only support version = 1 which is\n+\t * The newest version is 2, which is\n \t * the seeded murmur3 hashing technique implemented\n-\t * in bloom.c.\n+\t * in bloom.c. Bloom filters of version 1 were created\n+\t * with prior versions of Git, which had a bug in the\n+\t * implementation of the hash function.\n \t */\n \tuint32_t hash_version;\n \n@@ -80,7 +82,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n  * Not considered to be cryptographically secure.\n  * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  */\n-uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len);\n+uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len);\n \n void fill_bloom_key(const char *data,\n \t\t    size_t len,\ndiff --git a/commit-graph.c b/commit-graph.c\nindex bcc9a15cfa..6b21b17b20 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -314,17 +314,26 @@ static int graph_read_oid_lookup(const unsigned char *chunk_start,\n \treturn 0;\n }\n \n+struct graph_read_bloom_data_context {\n+\tstruct commit_graph *g;\n+\tint *commit_graph_changed_paths_version;\n+};\n+\n static int graph_read_bloom_data(const unsigned char *chunk_start,\n \t\t\t\t  size_t chunk_size, void *data)\n {\n-\tstruct commit_graph *g = data;\n+\tstruct graph_read_bloom_data_context *c = data;\n+\tstruct commit_graph *g = c->g;\n \tuint32_t hash_version;\n-\tg->chunk_bloom_data = chunk_start;\n \thash_version = get_be32(chunk_start);\n \n-\tif (hash_version != 1)\n+\tif (*c->commit_graph_changed_paths_version == -1) {\n+\t\t*c->commit_graph_changed_paths_version = hash_version;\n+\t} else if (hash_version != *c->commit_graph_changed_paths_version) {\n \t\treturn 0;\n+\t}\n \n+\tg->chunk_bloom_data = chunk_start;\n \tg->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n \tg->bloom_filter_settings->hash_version = hash_version;\n \tg->bloom_filter_settings->num_hashes = get_be32(chunk_start + 4);\n@@ -412,10 +421,14 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s,\n \t}\n \n \tif (s->commit_graph_changed_paths_version) {\n+\t\tstruct graph_read_bloom_data_context context = {\n+\t\t\t.g = graph,\n+\t\t\t.commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version\n+\t\t};\n \t\tpair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n \t\t\t   &graph->chunk_bloom_indexes);\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n-\t\t\t   graph_read_bloom_data, graph);\n+\t\t\t   graph_read_bloom_data, &context);\n \t}\n \n \tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n@@ -2436,6 +2449,13 @@ int write_commit_graph(struct object_directory *odb,\n \t}\n \tif (!commit_graph_compatible(r))\n \t\treturn 0;\n+\tif (r->settings.commit_graph_changed_paths_version < -1\n+\t    || r->settings.commit_graph_changed_paths_version > 2) {\n+\t\twarning(_(\"attempting to write a commit-graph, but \"\n+\t\t\t  \"'commitgraph.changedPathsVersion' (%d) is not supported\"),\n+\t\t\tr->settings.commit_graph_changed_paths_version);\n+\t\treturn 0;\n+\t}\n \n \tCALLOC_ARRAY(ctx, 1);\n \tctx->r = r;\n@@ -2448,6 +2468,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->write_generation_data = (get_configured_generation_version(r) == 2);\n \tctx->num_generation_data_overflows = 0;\n \n+\tbloom_settings.hash_version = r->settings.commit_graph_changed_paths_version == 2\n+\t\t? 2 : 1;\n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n \tbloom_settings.num_hashes = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_NUM_HASHES\",\n@@ -2477,7 +2499,7 @@ int write_commit_graph(struct object_directory *odb,\n \t\tg = ctx->r->objects->commit_graph;\n \n \t\t/* We have changed-paths already. Keep them in the next graph */\n-\t\tif (g && g->chunk_bloom_data) {\n+\t\tif (g && g->bloom_filter_settings) {\n \t\t\tctx->changed_paths = 1;\n \t\t\tctx->bloom_settings = g->bloom_filter_settings;\n \t\t}\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex aabe31d724..3cbc0a5b50 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -50,6 +50,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n \n static const char *bloom_usage = \"\\n\"\n \"  test-tool bloom get_murmur3 <string>\\n\"\n+\"  test-tool bloom get_murmur3_seven_highbit\\n\"\n \"  test-tool bloom generate_filter <string> [<string>...]\\n\"\n \"  test-tool bloom get_filter_for_commit <commit-hex>\\n\";\n \n@@ -64,7 +65,13 @@ int cmd__bloom(int argc, const char **argv)\n \t\tuint32_t hashed;\n \t\tif (argc < 3)\n \t\t\tusage(bloom_usage);\n-\t\thashed = murmur3_seeded(0, argv[2], strlen(argv[2]));\n+\t\thashed = murmur3_seeded_v2(0, argv[2], strlen(argv[2]));\n+\t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n+\t}\n+\n+\tif (!strcmp(argv[1], \"get_murmur3_seven_highbit\")) {\n+\t\tuint32_t hashed;\n+\t\thashed = murmur3_seeded_v2(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nindex b567383eb8..c8d84ab606 100755\n--- a/t/t0095-bloom.sh\n+++ b/t/t0095-bloom.sh\n@@ -29,6 +29,14 @@ test_expect_success 'compute unseeded murmur3 hash for test string 2' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'compute unseeded murmur3 hash for test string 3' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0xa183ccfd\n+\tEOF\n+\ttest-tool bloom get_murmur3_seven_highbit >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success 'compute bloom key for empty string' '\n \tcat >expect <<-\\EOF &&\n \tHashes:0x5615800c|0x5b966560|0x61174ab4|0x66983008|0x6c19155c|0x7199fab0|0x771ae004|\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 400dce2193..68066b7928 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -535,4 +535,118 @@ test_expect_success 'version 1 changed-path used when version 1 requested' '\n \t)\n '\n \n+test_expect_success 'version 1 changed-path not used when version 2 requested' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\tgit config --add commitgraph.changedPathsVersion 2 &&\n+\t\ttest_bloom_filters_not_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'version 1 changed-path used when autodetect requested' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'when writing another commit graph, preserve existing version 1 of changed-path' '\n+\ttest_commit -C highbit1 c1double \"$CENT$CENT\" &&\n+\tgit -C highbit1 commit-graph write --reachable --changed-paths &&\n+\t(\n+\t\tcd highbit1 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\techo \"options: bloom(1,10,7) read_generation_data\" >expect &&\n+\t\ttest-tool read-graph >full &&\n+\t\tgrep options full >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'set up repo with high bit path, version 2 changed-path' '\n+\tgit init highbit2 &&\n+\tgit -C highbit2 config --add commitgraph.changedPathsVersion 2 &&\n+\ttest_commit -C highbit2 c2 \"$CENT\" &&\n+\tgit -C highbit2 commit-graph write --reachable --changed-paths\n+'\n+\n+test_expect_success 'check value of version 2 changed-path' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\techo \"c01f\" >expect &&\n+\t\tget_first_changed_path_filter >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'setup make another commit' '\n+\t# \"git log\" does not use Bloom filters for root commits - see how, in\n+\t# revision.c, rev_compare_tree() (the only code path that eventually calls\n+\t# get_bloom_filter()) is only called by try_to_simplify_commit() when the commit\n+\t# has one parent. Therefore, make another commit so that we perform the tests on\n+\t# a non-root commit.\n+\ttest_commit -C highbit2 anotherc2 \"another$CENT\"\n+'\n+\n+test_expect_success 'version 2 changed-path used when version 2 requested' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'version 2 changed-path not used when version 1 requested' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\tgit config --add commitgraph.changedPathsVersion 1 &&\n+\t\ttest_bloom_filters_not_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'version 2 changed-path used when autodetect requested' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'when writing another commit graph, preserve existing version 2 of changed-path' '\n+\ttest_commit -C highbit2 c2double \"$CENT$CENT\" &&\n+\tgit -C highbit2 commit-graph write --reachable --changed-paths &&\n+\t(\n+\t\tcd highbit2 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\techo \"options: bloom(2,10,7) read_generation_data\" >expect &&\n+\t\ttest-tool read-graph >full &&\n+\t\tgrep options full >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'when writing commit graph, do not reuse changed-path of another version' '\n+\tgit init doublewrite &&\n+\ttest_commit -C doublewrite c \"$CENT\" &&\n+\tgit -C doublewrite config --add commitgraph.changedPathsVersion 1 &&\n+\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\tfor v in -2 3\n+\tdo\n+\t\tgit -C doublewrite config --add commitgraph.changedPathsVersion $v &&\n+\t\tgit -C doublewrite commit-graph write --reachable --changed-paths 2>err &&\n+\t\tcat >expect <<-EOF &&\n+\t\twarning: attempting to write a commit-graph, but ${SQ}commitgraph.changedPathsVersion${SQ} ($v) is not supported\n+\t\tEOF\n+\t\ttest_cmp expect err || return 1\n+\tdone &&\n+\tgit -C doublewrite config --add commitgraph.changedPathsVersion 2 &&\n+\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\t(\n+\t\tcd doublewrite &&\n+\t\techo \"c01f\" >expect &&\n+\t\tget_first_changed_path_filter >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_done\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483436","messageId":"dc69b28329557a6a090cd54dd34d7032462a7155.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 11/17] bloom: annotate filters with hash version","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:33:03Z","receivedAt":"2023-10-18T18:33:07Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In subsequent commits, we will want to load existing Bloom filters out\nof a commit-graph, even when the hash version they were computed with\ndoes not match the value of `commitGraph.changedPathVersion`.\n\nIn order to differentiate between the two, add a \"version\" field to each\nBloom filter.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c | 11 ++++++++---\n bloom.h |  1 +\n 2 files changed, 9 insertions(+), 3 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex ebef5cfd2f..9b6a30f6f6 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -55,6 +55,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n \t\t\t\t\tsizeof(unsigned char) * start_index +\n \t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n+\tfilter->version = g->bloom_filter_settings->hash_version;\n \n \treturn 1;\n }\n@@ -240,11 +241,13 @@ static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \treturn strcmp(e1->path, e2->path);\n }\n \n-static void init_truncated_large_filter(struct bloom_filter *filter)\n+static void init_truncated_large_filter(struct bloom_filter *filter,\n+\t\t\t\t\tint version)\n {\n \tfilter->data = xmalloc(1);\n \tfilter->data[0] = 0xFF;\n \tfilter->len = 1;\n+\tfilter->version = version;\n }\n \n struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n@@ -329,13 +332,15 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t}\n \n \t\tif (hashmap_get_size(&pathmap) > settings->max_changed_paths) {\n-\t\t\tinit_truncated_large_filter(filter);\n+\t\t\tinit_truncated_large_filter(filter,\n+\t\t\t\t\t\t    settings->hash_version);\n \t\t\tif (computed)\n \t\t\t\t*computed |= BLOOM_TRUNC_LARGE;\n \t\t\tgoto cleanup;\n \t\t}\n \n \t\tfilter->len = (hashmap_get_size(&pathmap) * settings->bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter->version = settings->hash_version;\n \t\tif (!filter->len) {\n \t\t\tif (computed)\n \t\t\t\t*computed |= BLOOM_TRUNC_EMPTY;\n@@ -355,7 +360,7 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t} else {\n \t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n \t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n-\t\tinit_truncated_large_filter(filter);\n+\t\tinit_truncated_large_filter(filter, settings->hash_version);\n \n \t\tif (computed)\n \t\t\t*computed |= BLOOM_TRUNC_LARGE;\ndiff --git a/bloom.h b/bloom.h\nindex 138d57a86b..330a140520 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -55,6 +55,7 @@ struct bloom_filter_settings {\n struct bloom_filter {\n \tunsigned char *data;\n \tsize_t len;\n+\tint version;\n };\n \n /*\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483438","messageId":"85dbdc4ed28058bcf4d1eff8f134fd2299bd18fb.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 12/17] bloom: prepare to discard incompatible Bloom filters","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:33:07Z","receivedAt":"2023-10-18T18:33:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Callers use the inline `get_bloom_filter()` implementation as a thin\nwrapper around `get_or_compute_bloom_filter()`. The former calls the\nlatter with a value of \"0\" for `compute_if_not_present`, making\n`get_bloom_filter()` the default read-only path for fetching an existing\nBloom filter.\n\nCallers expect the value returned from `get_bloom_filter()` is usable,\nthat is that it's compatible with the configured value corresponding to\n`commitGraph.changedPathsVersion`.\n\nThis is OK, since the commit-graph machinery only initializes its BDAT\nchunk (thereby enabling it to service Bloom filter queries) when the\nBloom filter hash_version is compatible with our settings. So any value\nreturned by `get_bloom_filter()` is trivially useable.\n\nHowever, subsequent commits will load the BDAT chunk even when the Bloom\nfilters are built with incompatible hash versions. Prepare to handle\nthis by teaching `get_bloom_filter()` to discard filters that are\nincompatible with the configured hash version.\n\nCallers who wish to read incompatible filters (e.g., for upgrading\nfilters from v1 to v2) may use the lower level routine,\n`get_or_compute_bloom_filter()`.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c | 20 +++++++++++++++++++-\n bloom.h | 20 ++++++++++++++++++--\n 2 files changed, 37 insertions(+), 3 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 9b6a30f6f6..739fa093ba 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -250,6 +250,23 @@ static void init_truncated_large_filter(struct bloom_filter *filter,\n \tfilter->version = version;\n }\n \n+struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c)\n+{\n+\tstruct bloom_filter *filter;\n+\tint hash_version;\n+\n+\tfilter = get_or_compute_bloom_filter(r, c, 0, NULL, NULL);\n+\tif (!filter)\n+\t\treturn NULL;\n+\n+\tprepare_repo_settings(r);\n+\thash_version = r->settings.commit_graph_changed_paths_version;\n+\n+\tif (!(hash_version == -1 || hash_version == filter->version))\n+\t\treturn NULL; /* unusable filter */\n+\treturn filter;\n+}\n+\n struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t struct commit *c,\n \t\t\t\t\t\t int compute_if_not_present,\n@@ -275,7 +292,8 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t     filter, graph_pos);\n \t}\n \n-\tif (filter->data && filter->len)\n+\tif ((filter->data && filter->len) &&\n+\t    (!settings || settings->hash_version == filter->version))\n \t\treturn filter;\n \tif (!compute_if_not_present)\n \t\treturn NULL;\ndiff --git a/bloom.h b/bloom.h\nindex 330a140520..bfe389e29c 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -110,8 +110,24 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t const struct bloom_filter_settings *settings,\n \t\t\t\t\t\t enum bloom_filter_computed *computed);\n \n-#define get_bloom_filter(r, c) get_or_compute_bloom_filter( \\\n-\t(r), (c), 0, NULL, NULL)\n+/*\n+ * Find the Bloom filter associated with the given commit \"c\".\n+ *\n+ * If any of the following are true\n+ *\n+ *   - the repository does not have a commit-graph, or\n+ *   - the repository disables reading from the commit-graph, or\n+ *   - the given commit does not have a Bloom filter computed, or\n+ *   - there is a Bloom filter for commit \"c\", but it cannot be read\n+ *     because the filter uses an incompatible version of murmur3\n+ *\n+ * , then `get_bloom_filter()` will return NULL. Otherwise, the corresponding\n+ * Bloom filter will be returned.\n+ *\n+ * For callers who wish to inspect Bloom filters with incompatible hash\n+ * versions, use get_or_compute_bloom_filter().\n+ */\n+struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c);\n \n int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483439","messageId":"3ff669a622355bd8eb63494cd90d241d47dd2d83.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 13/17] commit-graph.c: unconditionally load Bloom filters","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:33:10Z","receivedAt":"2023-10-18T18:33:14Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In 9e4df4da07 (commit-graph: new filter ver. that fixes murmur3,\n2023-08-01), we began ignoring the Bloom data (\"BDAT\") chunk for\ncommit-graphs whose Bloom filters were computed using a hash version\nincompatible with the value of `commitGraph.changedPathVersion`.\n\nNow that the Bloom API has been hardened to discard these incompatible\nfilters (with the exception of low-level APIs), we can safely load these\nBloom filters unconditionally.\n\nWe no longer want to return early from `graph_read_bloom_data()`, and\nsimilarly do not want to set the bloom_settings' `hash_version` field as\na side-effect. The latter is because we want to wait until we know which\nBloom settings we're using (either the defaults, from the GIT_TEST\nvariables, or from the previous commit-graph layer) before deciding what\nhash_version to use.\n\nIf we detect an existing BDAT chunk, we'll infer the rest of the\nsettings (e.g., number of hashes, bits per entry, and maximum number of\nchanged paths) from the earlier graph layer. The hash_version will be\ninferred from the previous layer as well, unless one has already been\nspecified via configuration.\n\nOnce all of that is done, we normalize the value of the hash_version to\neither \"1\" or \"2\".\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n commit-graph.c | 19 ++++++++++---------\n 1 file changed, 10 insertions(+), 9 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 6b21b17b20..7d0fb32107 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -327,12 +327,6 @@ static int graph_read_bloom_data(const unsigned char *chunk_start,\n \tuint32_t hash_version;\n \thash_version = get_be32(chunk_start);\n \n-\tif (*c->commit_graph_changed_paths_version == -1) {\n-\t\t*c->commit_graph_changed_paths_version = hash_version;\n-\t} else if (hash_version != *c->commit_graph_changed_paths_version) {\n-\t\treturn 0;\n-\t}\n-\n \tg->chunk_bloom_data = chunk_start;\n \tg->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n \tg->bloom_filter_settings->hash_version = hash_version;\n@@ -2468,8 +2462,7 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->write_generation_data = (get_configured_generation_version(r) == 2);\n \tctx->num_generation_data_overflows = 0;\n \n-\tbloom_settings.hash_version = r->settings.commit_graph_changed_paths_version == 2\n-\t\t? 2 : 1;\n+\tbloom_settings.hash_version = r->settings.commit_graph_changed_paths_version;\n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n \tbloom_settings.num_hashes = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_NUM_HASHES\",\n@@ -2501,10 +2494,18 @@ int write_commit_graph(struct object_directory *odb,\n \t\t/* We have changed-paths already. Keep them in the next graph */\n \t\tif (g && g->bloom_filter_settings) {\n \t\t\tctx->changed_paths = 1;\n-\t\t\tctx->bloom_settings = g->bloom_filter_settings;\n+\n+\t\t\t/* don't propagate the hash_version unless unspecified */\n+\t\t\tif (bloom_settings.hash_version == -1)\n+\t\t\t\tbloom_settings.hash_version = g->bloom_filter_settings->hash_version;\n+\t\t\tbloom_settings.bits_per_entry = g->bloom_filter_settings->bits_per_entry;\n+\t\t\tbloom_settings.num_hashes = g->bloom_filter_settings->num_hashes;\n+\t\t\tbloom_settings.max_changed_paths = g->bloom_filter_settings->max_changed_paths;\n \t\t}\n \t}\n \n+\tbloom_settings.hash_version = bloom_settings.hash_version == 2 ? 2 : 1;\n+\n \tif (ctx->split) {\n \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n \n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483440","messageId":"1c78e3d17804f64797fc13cd31dadfd51e550bf4.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 14/17] commit-graph: drop unnecessary `graph_read_bloom_data_context`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:33:14Z","receivedAt":"2023-10-18T18:33:18Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The `graph_read_bloom_data_context` struct was introduced in an earlier\ncommit in order to pass pointers to the commit-graph and changed-path\nBloom filter version when reading the BDAT chunk.\n\nThe previous commit no longer writes through the changed_paths_version\npointer, making the surrounding context structure unnecessary. Drop it\nand pass a pointer to the commit-graph directly when reading the BDAT\nchunk.\n\nNoticed-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n commit-graph.c | 14 ++------------\n 1 file changed, 2 insertions(+), 12 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 7d0fb32107..b70d57b085 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -314,16 +314,10 @@ static int graph_read_oid_lookup(const unsigned char *chunk_start,\n \treturn 0;\n }\n \n-struct graph_read_bloom_data_context {\n-\tstruct commit_graph *g;\n-\tint *commit_graph_changed_paths_version;\n-};\n-\n static int graph_read_bloom_data(const unsigned char *chunk_start,\n \t\t\t\t  size_t chunk_size, void *data)\n {\n-\tstruct graph_read_bloom_data_context *c = data;\n-\tstruct commit_graph *g = c->g;\n+\tstruct commit_graph *g = data;\n \tuint32_t hash_version;\n \thash_version = get_be32(chunk_start);\n \n@@ -415,14 +409,10 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s,\n \t}\n \n \tif (s->commit_graph_changed_paths_version) {\n-\t\tstruct graph_read_bloom_data_context context = {\n-\t\t\t.g = graph,\n-\t\t\t.commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version\n-\t\t};\n \t\tpair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n \t\t\t   &graph->chunk_bloom_indexes);\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n-\t\t\t   graph_read_bloom_data, &context);\n+\t\t\t   graph_read_bloom_data, graph);\n \t}\n \n \tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483441","messageId":"a289514faabb7749e2f422ab779a8cbd62521824.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 15/17] object.h: fix mis-aligned flag bits table","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:33:17Z","receivedAt":"2023-10-18T18:33:21Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Bit position 23 is one column too far to the left.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/object.h b/object.h\nindex 114d45954d..db25714b4e 100644\n--- a/object.h\n+++ b/object.h\n@@ -62,7 +62,7 @@ void object_array_init(struct object_array *array);\n \n /*\n  * object flag allocation:\n- * revision.h:               0---------10         15             23------27\n+ * revision.h:               0---------10         15               23------27\n  * fetch-pack.c:             01    67\n  * negotiator/default.c:       2--5\n  * walker.c:                 0-2\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483442","messageId":"6a12e39e7f5e2d68a23811a98276740646874662.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 16/17] commit-graph: reuse existing Bloom filters where possible","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:33:21Z","receivedAt":"2023-10-18T18:33:26Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In 9e4df4da07 (commit-graph: new filter ver. that fixes murmur3,\n2023-08-01), a bug was described where it's possible for Git to produce\nnon-murmur3 hashes when the platform's \"char\" type is signed, and there\nare paths with characters whose highest bit is set (i.e. all characters\n>= 0x80).\n\nThat patch allows the caller to control which version of Bloom filters\nare read and written. However, even on platforms with a signed \"char\"\ntype, it is possible to reuse existing Bloom filters if and only if\nthere are no changed paths in any commit's first parent tree-diff whose\ncharacters have their highest bit set.\n\nWhen this is the case, we can reuse the existing filter without having\nto compute a new one. This is done by marking trees which are known to\nhave (or not have) any such paths. When a commit's root tree is verified\nto not have any such paths, we mark it as such and declare that the\ncommit's Bloom filter is reusable.\n\nNote that this heuristic only goes in one direction. If neither a commit\nnor its first parent have any paths in their trees with non-ASCII\ncharacters, then we know for certain that a path with non-ASCII\ncharacters will not appear in a tree-diff against that commit's first\nparent. The reverse isn't necessarily true: just because the tree-diff\ndoesn't contain any such paths does not imply that no such paths exist\nin either tree.\n\nSo we end up recomputing some Bloom filters that we don't strictly have\nto (i.e. their bits are the same no matter which version of murmur3 we\nuse). But culling these out is impossible, since we'd have to perform\nthe full tree-diff, which is the same effort as computing the Bloom\nfilter from scratch.\n\nBut because we can cache our results in each tree's flag bits, we can\noften avoid recomputing many filters, thereby reducing the time it takes\nto run\n\n    $ git commit-graph write --changed-paths --reachable\n\nwhen upgrading from v1 to v2 Bloom filters.\n\nTo benchmark this, let's generate a commit-graph in linux.git with v1\nchanged-paths in generation order[^1]:\n\n    $ git clone git@github.com:torvalds/linux.git\n    $ cd linux\n    $ git commit-graph write --reachable --changed-paths\n    $ graph=\".git/objects/info/commit-graph\"\n    $ mv $graph{,.bak}\n\nThen let's time how long it takes to go from v1 to v2 filters (with and\nwithout the upgrade path enabled), resetting the state of the\ncommit-graph each time:\n\n    $ git config commitGraph.changedPathsVersion 2\n    $ hyperfine -p 'cp -f $graph.bak $graph' -L v 0,1 \\\n        'GIT_TEST_UPGRADE_BLOOM_FILTERS={v} git.compile commit-graph write --reachable --changed-paths'\n\nOn linux.git (where there aren't any non-ASCII paths), the timings\nindicate that this patch represents a speed-up over recomputing all\nBloom filters from scratch:\n\n    Benchmark 1: GIT_TEST_UPGRADE_BLOOM_FILTERS=0 git.compile commit-graph write --reachable --changed-paths\n      Time (mean ± σ):     124.873 s ±  0.316 s    [User: 124.081 s, System: 0.643 s]\n      Range (min … max):   124.621 s … 125.227 s    3 runs\n\n    Benchmark 2: GIT_TEST_UPGRADE_BLOOM_FILTERS=1 git.compile commit-graph write --reachable --changed-paths\n      Time (mean ± σ):     79.271 s ±  0.163 s    [User: 74.611 s, System: 4.521 s]\n      Range (min … max):   79.112 s … 79.437 s    3 runs\n\n    Summary\n      'GIT_TEST_UPGRADE_BLOOM_FILTERS=1 git.compile commit-graph write --reachable --changed-paths' ran\n        1.58 ± 0.01 times faster than 'GIT_TEST_UPGRADE_BLOOM_FILTERS=0 git.compile commit-graph write --reachable --changed-paths'\n\nOn git.git, we do have some non-ASCII paths, giving us a more modest\nimprovement from 4.163 seconds to 3.348 seconds, for a 1.24x speed-up.\nOn my machine, the stats for git.git are:\n\n  - 8,285 Bloom filters computed from scratch\n  - 10 Bloom filters generated as empty\n  - 4 Bloom filters generated as truncated due to too many changed paths\n  - 65,114 Bloom filters were reused when transitioning from v1 to v2.\n\n[^1]: Note that this is is important, since `--stdin-packs` or\n  `--stdin-commits` orders commits in the commit-graph by their pack\n  position (with `--stdin-packs`) or in the raw input (with\n  `--stdin-commits`).\n\n  Since we compute Bloom filters in the same order that commits appear\n  in the graph, we must see a commit's (first) parent before we process\n  the commit itself. This is only guaranteed to happen when sorting\n  commits by their generation number.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c              | 90 ++++++++++++++++++++++++++++++++++++++++++--\n bloom.h              |  1 +\n commit-graph.c       |  5 +++\n object.h             |  1 +\n t/t4216-log-bloom.sh | 35 ++++++++++++++++-\n 5 files changed, 128 insertions(+), 4 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 739fa093ba..24dd874e46 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -7,6 +7,9 @@\n #include \"commit-graph.h\"\n #include \"commit.h\"\n #include \"commit-slab.h\"\n+#include \"tree.h\"\n+#include \"tree-walk.h\"\n+#include \"config.h\"\n \n define_commit_slab(bloom_filter_slab, struct bloom_filter);\n \n@@ -250,6 +253,73 @@ static void init_truncated_large_filter(struct bloom_filter *filter,\n \tfilter->version = version;\n }\n \n+#define VISITED   (1u<<21)\n+#define HIGH_BITS (1u<<22)\n+\n+static int has_entries_with_high_bit(struct repository *r, struct tree *t)\n+{\n+\tif (parse_tree(t))\n+\t\treturn 1;\n+\n+\tif (!(t->object.flags & VISITED)) {\n+\t\tstruct tree_desc desc;\n+\t\tstruct name_entry entry;\n+\n+\t\tinit_tree_desc(&desc, t->buffer, t->size);\n+\t\twhile (tree_entry(&desc, &entry)) {\n+\t\t\tsize_t i;\n+\t\t\tfor (i = 0; i < entry.pathlen; i++) {\n+\t\t\t\tif (entry.path[i] & 0x80) {\n+\t\t\t\t\tt->object.flags |= HIGH_BITS;\n+\t\t\t\t\tgoto done;\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t\tif (S_ISDIR(entry.mode)) {\n+\t\t\t\tstruct tree *sub = lookup_tree(r, &entry.oid);\n+\t\t\t\tif (sub && has_entries_with_high_bit(r, sub)) {\n+\t\t\t\t\tt->object.flags |= HIGH_BITS;\n+\t\t\t\t\tgoto done;\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t}\n+\n+done:\n+\t\tt->object.flags |= VISITED;\n+\t}\n+\n+\treturn !!(t->object.flags & HIGH_BITS);\n+}\n+\n+static int commit_tree_has_high_bit_paths(struct repository *r,\n+\t\t\t\t\t  struct commit *c)\n+{\n+\tstruct tree *t;\n+\tif (repo_parse_commit(r, c))\n+\t\treturn 1;\n+\tt = repo_get_commit_tree(r, c);\n+\tif (!t)\n+\t\treturn 1;\n+\treturn has_entries_with_high_bit(r, t);\n+}\n+\n+static struct bloom_filter *upgrade_filter(struct repository *r, struct commit *c,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   int hash_version)\n+{\n+\tstruct commit_list *p = c->parents;\n+\tif (commit_tree_has_high_bit_paths(r, c))\n+\t\treturn NULL;\n+\n+\tif (p && commit_tree_has_high_bit_paths(r, p->item))\n+\t\treturn NULL;\n+\n+\tfilter->version = hash_version;\n+\n+\treturn filter;\n+}\n+\n struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c)\n {\n \tstruct bloom_filter *filter;\n@@ -292,9 +362,23 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t     filter, graph_pos);\n \t}\n \n-\tif ((filter->data && filter->len) &&\n-\t    (!settings || settings->hash_version == filter->version))\n-\t\treturn filter;\n+\tif (filter->data && filter->len) {\n+\t\tstruct bloom_filter *upgrade;\n+\t\tif (!settings || settings->hash_version == filter->version)\n+\t\t\treturn filter;\n+\n+\t\t/* version mismatch, see if we can upgrade */\n+\t\tif (compute_if_not_present &&\n+\t\t    git_env_bool(\"GIT_TEST_UPGRADE_BLOOM_FILTERS\", 1)) {\n+\t\t\tupgrade = upgrade_filter(r, c, filter,\n+\t\t\t\t\t\t settings->hash_version);\n+\t\t\tif (upgrade) {\n+\t\t\t\tif (computed)\n+\t\t\t\t\t*computed |= BLOOM_UPGRADED;\n+\t\t\t\treturn upgrade;\n+\t\t\t}\n+\t\t}\n+\t}\n \tif (!compute_if_not_present)\n \t\treturn NULL;\n \ndiff --git a/bloom.h b/bloom.h\nindex bfe389e29c..e3a9b68905 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -102,6 +102,7 @@ enum bloom_filter_computed {\n \tBLOOM_COMPUTED     = (1 << 1),\n \tBLOOM_TRUNC_LARGE  = (1 << 2),\n \tBLOOM_TRUNC_EMPTY  = (1 << 3),\n+\tBLOOM_UPGRADED     = (1 << 4),\n };\n \n struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\ndiff --git a/commit-graph.c b/commit-graph.c\nindex b70d57b085..50dcbb4d9b 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1102,6 +1102,7 @@ struct write_commit_graph_context {\n \tint count_bloom_filter_not_computed;\n \tint count_bloom_filter_trunc_empty;\n \tint count_bloom_filter_trunc_large;\n+\tint count_bloom_filter_upgraded;\n };\n \n static int write_graph_chunk_fanout(struct hashfile *f,\n@@ -1709,6 +1710,8 @@ static void trace2_bloom_filter_write_statistics(struct write_commit_graph_conte\n \t\t\t   ctx->count_bloom_filter_trunc_empty);\n \ttrace2_data_intmax(\"commit-graph\", ctx->r, \"filter-trunc-large\",\n \t\t\t   ctx->count_bloom_filter_trunc_large);\n+\ttrace2_data_intmax(\"commit-graph\", ctx->r, \"filter-upgraded\",\n+\t\t\t   ctx->count_bloom_filter_upgraded);\n }\n \n static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n@@ -1750,6 +1753,8 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \t\t\t\tctx->count_bloom_filter_trunc_empty++;\n \t\t\tif (computed & BLOOM_TRUNC_LARGE)\n \t\t\t\tctx->count_bloom_filter_trunc_large++;\n+\t\t} else if (computed & BLOOM_UPGRADED) {\n+\t\t\tctx->count_bloom_filter_upgraded++;\n \t\t} else if (computed & BLOOM_NOT_COMPUTED)\n \t\t\tctx->count_bloom_filter_not_computed++;\n \t\tctx->total_bloom_filter_data_size += filter\ndiff --git a/object.h b/object.h\nindex db25714b4e..2e5e08725f 100644\n--- a/object.h\n+++ b/object.h\n@@ -75,6 +75,7 @@ void object_array_init(struct object_array *array);\n  * commit-reach.c:                                  16-----19\n  * sha1-name.c:                                              20\n  * list-objects-filter.c:                                      21\n+ * bloom.c:                                                    2122\n  * builtin/fsck.c:           0--3\n  * builtin/gc.c:             0\n  * builtin/index-pack.c:                                     2021\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 68066b7928..569f2b6f8b 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -221,6 +221,10 @@ test_filter_trunc_large () {\n \tgrep \"\\\"key\\\":\\\"filter-trunc-large\\\",\\\"value\\\":\\\"$1\\\"\" $2\n }\n \n+test_filter_upgraded () {\n+\tgrep \"\\\"key\\\":\\\"filter-upgraded\\\",\\\"value\\\":\\\"$1\\\"\" $2\n+}\n+\n test_expect_success 'correctly report changes over limit' '\n \tgit init limits &&\n \t(\n@@ -628,7 +632,13 @@ test_expect_success 'when writing another commit graph, preserve existing versio\n test_expect_success 'when writing commit graph, do not reuse changed-path of another version' '\n \tgit init doublewrite &&\n \ttest_commit -C doublewrite c \"$CENT\" &&\n+\n \tgit -C doublewrite config --add commitgraph.changedPathsVersion 1 &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 1 trace2.txt &&\n+\ttest_filter_upgraded 0 trace2.txt &&\n+\n \tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n \tfor v in -2 3\n \tdo\n@@ -639,8 +649,13 @@ test_expect_success 'when writing commit graph, do not reuse changed-path of ano\n \t\tEOF\n \t\ttest_cmp expect err || return 1\n \tdone &&\n+\n \tgit -C doublewrite config --add commitgraph.changedPathsVersion 2 &&\n-\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 1 trace2.txt &&\n+\ttest_filter_upgraded 0 trace2.txt &&\n+\n \t(\n \t\tcd doublewrite &&\n \t\techo \"c01f\" >expect &&\n@@ -649,4 +664,22 @@ test_expect_success 'when writing commit graph, do not reuse changed-path of ano\n \t)\n '\n \n+test_expect_success 'when writing commit graph, reuse changed-path of another version where possible' '\n+\tgit init upgrade &&\n+\n+\ttest_commit -C upgrade base no-high-bits &&\n+\n+\tgit -C upgrade config --add commitgraph.changedPathsVersion 1 &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C upgrade commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 1 trace2.txt &&\n+\ttest_filter_upgraded 0 trace2.txt &&\n+\n+\tgit -C upgrade config --add commitgraph.changedPathsVersion 2 &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C upgrade commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 0 trace2.txt &&\n+\ttest_filter_upgraded 1 trace2.txt\n+'\n+\n test_done\n-- \n2.42.0.415.g8942f205c8\n\n"},{"id":"483443","messageId":"8942f205c84d204376033a931b1ba8a575181f1f.1697653929.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v4 17/17] bloom: introduce `deinit_bloom_filters()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-18T18:33:24Z","receivedAt":"2023-10-18T18:33:28Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"After we are done using Bloom filters, we do not currently clean up any\nmemory allocated by the commit slab used to store those filters in the\nfirst place.\n\nBesides the bloom_filter structures themselves, there is mostly nothing\nto free() in the first place, since in the read-only path all Bloom\nfilter's `data` members point to a memory mapped region in the\ncommit-graph file itself.\n\nBut when generating Bloom filters from scratch (or initializing\ntruncated filters) we allocate additional memory to store the filter's\ndata.\n\nKeep track of when we need to free() this additional chunk of memory by\nusing an extra pointer `to_free`. Most of the time this will be NULL\n(indicating that we are representing an existing Bloom filter stored in\na memory mapped region). When it is non-NULL, free it before discarding\nthe Bloom filters slab.\n\nSuggested-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c        | 16 +++++++++++++++-\n bloom.h        |  3 +++\n commit-graph.c |  4 ++++\n 3 files changed, 22 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 24dd874e46..ff131893cd 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -59,6 +59,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t\tsizeof(unsigned char) * start_index +\n \t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n \tfilter->version = g->bloom_filter_settings->hash_version;\n+\tfilter->to_free = NULL;\n \n \treturn 1;\n }\n@@ -231,6 +232,18 @@ void init_bloom_filters(void)\n \tinit_bloom_filter_slab(&bloom_filters);\n }\n \n+static void free_one_bloom_filter(struct bloom_filter *filter)\n+{\n+\tif (!filter)\n+\t\treturn;\n+\tfree(filter->to_free);\n+}\n+\n+void deinit_bloom_filters(void)\n+{\n+\tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -247,7 +260,7 @@ static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n static void init_truncated_large_filter(struct bloom_filter *filter,\n \t\t\t\t\tint version)\n {\n-\tfilter->data = xmalloc(1);\n+\tfilter->data = filter->to_free = xmalloc(1);\n \tfilter->data[0] = 0xFF;\n \tfilter->len = 1;\n \tfilter->version = version;\n@@ -449,6 +462,7 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\tfilter->len = 1;\n \t\t}\n \t\tCALLOC_ARRAY(filter->data, filter->len);\n+\t\tfilter->to_free = filter->data;\n \n \t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n \t\t\tstruct bloom_key key;\ndiff --git a/bloom.h b/bloom.h\nindex e3a9b68905..d20e64bfbb 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -56,6 +56,8 @@ struct bloom_filter {\n \tunsigned char *data;\n \tsize_t len;\n \tint version;\n+\n+\tvoid *to_free;\n };\n \n /*\n@@ -96,6 +98,7 @@ void add_key_to_filter(const struct bloom_key *key,\n \t\t       const struct bloom_filter_settings *settings);\n \n void init_bloom_filters(void);\n+void deinit_bloom_filters(void);\n \n enum bloom_filter_computed {\n \tBLOOM_NOT_COMPUTED = (1 << 0),\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 50dcbb4d9b..60fa64d956 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -779,6 +779,7 @@ struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n void close_commit_graph(struct raw_object_store *o)\n {\n \tclear_commit_graph_data_slab(&commit_graph_data_slab);\n+\tdeinit_bloom_filters();\n \tfree_commit_graph(o->commit_graph);\n \to->commit_graph = NULL;\n }\n@@ -2583,6 +2584,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \tres = write_commit_graph_file(ctx);\n \n+\tif (ctx->changed_paths)\n+\t\tdeinit_bloom_filters();\n+\n \tif (ctx->split)\n \t\tmark_commit_graphs(ctx);\n \n-- \n2.42.0.415.g8942f205c8\n"},{"id":"483459","messageId":"xmqqa5sfplvw.fsf@gitster.g","threadId":"60318","inReplyTo":"da52ec838025a59a3f4f4ffaf2e6f9098a37547e.1697648864.git.me@ttaylorr.com","subject":"Re: [PATCH v3 05/10] bulk-checkin: extract abstract `bulk_checkin_source`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-18T23:10:43Z","receivedAt":"2023-10-18T23:10:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> A future commit will want to implement a very similar routine as in\n> `stream_blob_to_pack()` with two notable changes:\n>\n>   - Instead of streaming just OBJ_BLOBs, this new function may want to\n>     stream objects of arbitrary type.\n>\n>   - Instead of streaming the object's contents from an open\n>     file-descriptor, this new function may want to \"stream\" its contents\n>     from memory.\n>\n> To avoid duplicating a significant chunk of code between the existing\n> `stream_blob_to_pack()`, extract an abstract `bulk_checkin_source`. This\n> concept currently is a thin layer of `lseek()` and `read_in_full()`, but\n> will grow to understand how to perform analogous operations when writing\n> out an object's contents from memory.\n>\n> Suggested-by: Junio C Hamano <gitster@pobox.com>\n> Signed-off-by: Taylor Blau <me@ttaylorr.com>\n> ---\n>  bulk-checkin.c | 61 +++++++++++++++++++++++++++++++++++++++++++-------\n>  1 file changed, 53 insertions(+), 8 deletions(-)\n\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index f4914fb6d1..fc1d902018 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -140,8 +140,41 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n>  \treturn 0;\n>  }\n>  \n> +struct bulk_checkin_source {\n> +\tenum { SOURCE_FILE } type;\n> +\n> +\t/* SOURCE_FILE fields */\n> +\tint fd;\n> +\n> +\t/* common fields */\n> +\tsize_t size;\n> +\tconst char *path;\n> +};\n\nLooks OK, even though I expected to see a bit more involved object\norientation with something like\n\n\tstruct bulk_checkin_source {\n\t\toff_t (*read)(struct bulk_checkin_source *, void *, size_t);\n\t\toff_t (*seek)(struct bulk_checkin_source *, off_t);\n\t\tunion {\n\t\t\tstruct {\n\t\t\t\tint fd;\n\t\t\t\tsize_t size;\n\t\t\t\tconst char *path;\n\t\t\t} from_fd;\n\t\t\tstruct {\n\t\t\t\t...\n\t\t\t} incore;\n\t\t} data;\n\t};\n\nAs there will only be two subclasses of this thing, it may not\nmatter all that much right now, but it would be much nicer as your\nmethods do not have to care about \"switch (enum) { case FILE: ... }\".\n\n"},{"id":"483460","messageId":"xmqq34y7plj4.fsf@gitster.g","threadId":"60318","inReplyTo":"8667b763652ffa71b52b7bd78821e46a6e5fe5a9.1697648864.git.me@ttaylorr.com","subject":"Re: [PATCH v3 08/10] bulk-checkin: introduce `index_blob_bulk_checkin_incore()`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-18T23:18:23Z","receivedAt":"2023-10-18T23:19:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> Now that we have factored out many of the common routines necessary to\n> index a new object into a pack created by the bulk-checkin machinery, we\n> can introduce a variant of `index_blob_bulk_checkin()` that acts on\n> blobs whose contents we can fit in memory.\n\nHmph.\n\nDoesn't the duplication between the main loop of the new\ndeflate_obj_contents_to_pack() with existing deflate_blob_to_pack()\nbother you?\n\nA similar duplication in the previous round resulted in a nice\nrefactoring of patches 5 and 6 in this round.  Compared to that,\nthe differences in the set-up code between the two functions may be\nmuch larger, but subtle and unnecessary differences between the code\nthat was copied and pasted (e.g., we do not check the errors from\nthe seek_to method here, but the original does) is a sign that over\ntime a fix to one will need to be carried over to the other, adding\nunnecessary maintenance burden, isn't it?\n\n> +static int deflate_obj_contents_to_pack_incore(struct bulk_checkin_packfile *state,\n> +\t\t\t\t\t       git_hash_ctx *ctx,\n> +\t\t\t\t\t       struct hashfile_checkpoint *checkpoint,\n> +\t\t\t\t\t       struct object_id *result_oid,\n> +\t\t\t\t\t       const void *buf, size_t size,\n> +\t\t\t\t\t       enum object_type type,\n> +\t\t\t\t\t       const char *path, unsigned flags)\n> +{\n> +\tstruct pack_idx_entry *idx = NULL;\n> +\toff_t already_hashed_to = 0;\n> +\tstruct bulk_checkin_source source = {\n> +\t\t.type = SOURCE_INCORE,\n> +\t\t.buf = buf,\n> +\t\t.size = size,\n> +\t\t.read = 0,\n> +\t\t.path = path,\n> +\t};\n> +\n> +\t/* Note: idx is non-NULL when we are writing */\n> +\tif (flags & HASH_WRITE_OBJECT)\n> +\t\tCALLOC_ARRAY(idx, 1);\n> +\n> +\twhile (1) {\n> +\t\tprepare_checkpoint(state, checkpoint, idx, flags);\n> +\n> +\t\tif (!stream_obj_to_pack(state, ctx, &already_hashed_to, &source,\n> +\t\t\t\t\ttype, flags))\n> +\t\t\tbreak;\n> +\t\ttruncate_checkpoint(state, checkpoint, idx);\n> +\t\tbulk_checkin_source_seek_to(&source, 0);\n> +\t}\n> +\n> +\tfinalize_checkpoint(state, ctx, checkpoint, idx, result_oid);\n> +\n> +\treturn 0;\n> +}\n> +\n> +static int deflate_blob_to_pack_incore(struct bulk_checkin_packfile *state,\n> +\t\t\t\t       struct object_id *result_oid,\n> +\t\t\t\t       const void *buf, size_t size,\n> +\t\t\t\t       const char *path, unsigned flags)\n> +{\n> +\tgit_hash_ctx ctx;\n> +\tstruct hashfile_checkpoint checkpoint = {0};\n> +\n> +\tformat_object_header_hash(the_hash_algo, &ctx, &checkpoint, OBJ_BLOB,\n> +\t\t\t\t  size);\n> +\n> +\treturn deflate_obj_contents_to_pack_incore(state, &ctx, &checkpoint,\n> +\t\t\t\t\t\t   result_oid, buf, size,\n> +\t\t\t\t\t\t   OBJ_BLOB, path, flags);\n> +}\n> +\n>  static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n>  \t\t\t\tstruct object_id *result_oid,\n>  \t\t\t\tint fd, size_t size,\n> @@ -456,6 +509,17 @@ int index_blob_bulk_checkin(struct object_id *oid,\n>  \treturn status;\n>  }\n>  \n> +int index_blob_bulk_checkin_incore(struct object_id *oid,\n> +\t\t\t\t   const void *buf, size_t size,\n> +\t\t\t\t   const char *path, unsigned flags)\n> +{\n> +\tint status = deflate_blob_to_pack_incore(&bulk_checkin_packfile, oid,\n> +\t\t\t\t\t\t buf, size, path, flags);\n> +\tif (!odb_transaction_nesting)\n> +\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n> +\treturn status;\n> +}\n> +\n>  void begin_odb_transaction(void)\n>  {\n>  \todb_transaction_nesting += 1;\n> diff --git a/bulk-checkin.h b/bulk-checkin.h\n> index aa7286a7b3..1b91daeaee 100644\n> --- a/bulk-checkin.h\n> +++ b/bulk-checkin.h\n> @@ -13,6 +13,10 @@ int index_blob_bulk_checkin(struct object_id *oid,\n>  \t\t\t    int fd, size_t size,\n>  \t\t\t    const char *path, unsigned flags);\n>  \n> +int index_blob_bulk_checkin_incore(struct object_id *oid,\n> +\t\t\t\t   const void *buf, size_t size,\n> +\t\t\t\t   const char *path, unsigned flags);\n> +\n>  /*\n>   * Tell the object database to optimize for adding\n>   * multiple objects. end_odb_transaction must be called\n"},{"id":"483461","messageId":"xmqq34y71phj.fsf@gitster.g","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"Re: [PATCH v4 00/17] bloom: changed-path Bloom filters v2 (& sundries)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-18T23:26:48Z","receivedAt":"2023-10-18T23:26:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> (Rebased onto the tip of 'master', which is 3a06386e31 (The fifteenth\n> batch, 2023-10-04), at the time of writing).\n\nJudging from 17/17 that has a free_commit_graph() call in\nclose_commit_graph(), that was merged in the eighteenth batch,\nthe above is probably untrue.  I'll apply to the current master and\nsee how it goes instead.\n\n> Thanks to Jonathan, Peff, and SZEDER who have helped a great deal in\n> assembling these patches. As usual, a range-diff is included below.\n> Thanks in advance for your\n> review!\n\nThanks.\n"},{"id":"483489","messageId":"ZTFI++b51Cj+Sto9@nand.local","threadId":"60318","inReplyTo":"xmqqa5sfplvw.fsf@gitster.g","subject":"Re: [PATCH v3 05/10] bulk-checkin: extract abstract `bulk_checkin_source`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-19T15:19:23Z","receivedAt":"2023-10-19T15:19:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Oct 18, 2023 at 04:10:43PM -0700, Junio C Hamano wrote:\n> Looks OK, even though I expected to see a bit more involved object\n> orientation with something like\n>\n> \tstruct bulk_checkin_source {\n> \t\toff_t (*read)(struct bulk_checkin_source *, void *, size_t);\n> \t\toff_t (*seek)(struct bulk_checkin_source *, off_t);\n> \t\tunion {\n> \t\t\tstruct {\n> \t\t\t\tint fd;\n> \t\t\t\tsize_t size;\n> \t\t\t\tconst char *path;\n> \t\t\t} from_fd;\n> \t\t\tstruct {\n> \t\t\t\t...\n> \t\t\t} incore;\n> \t\t} data;\n> \t};\n>\n> As there will only be two subclasses of this thing, it may not\n> matter all that much right now, but it would be much nicer as your\n> methods do not have to care about \"switch (enum) { case FILE: ... }\".\n\nI want to be cautious of going too far in this direction. I anticipate\nthat \"two\" is probably the maximum number of kinds of sources we can\nreasonably expect for the foreseeable future. If that changes, it's easy\nenough to convert from the existing implementation to something closer\nto the above.\n\nThanks,\nTaylor\n"},{"id":"483494","messageId":"ZTFLrOpGaTQHEGcT@nand.local","threadId":"60318","inReplyTo":"xmqq34y7plj4.fsf@gitster.g","subject":"Re: [PATCH v3 08/10] bulk-checkin: introduce `index_blob_bulk_checkin_incore()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-19T15:30:52Z","receivedAt":"2023-10-19T15:30:56Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Oct 18, 2023 at 04:18:23PM -0700, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > Now that we have factored out many of the common routines necessary to\n> > index a new object into a pack created by the bulk-checkin machinery, we\n> > can introduce a variant of `index_blob_bulk_checkin()` that acts on\n> > blobs whose contents we can fit in memory.\n>\n> Hmph.\n>\n> Doesn't the duplication between the main loop of the new\n> deflate_obj_contents_to_pack() with existing deflate_blob_to_pack()\n> bother you?\n\nYeah, I am not sure how I missed seeing the opportunity to clean that\nup. Another (hopefully final...) reroll incoming.\n\nThanks,\nTaylor\n"},{"id":"483508","messageId":"xmqq4jimxzrp.fsf@gitster.g","threadId":"60318","inReplyTo":"ZTFI++b51Cj+Sto9@nand.local","subject":"Re: [PATCH v3 05/10] bulk-checkin: extract abstract `bulk_checkin_source`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-19T17:55:54Z","receivedAt":"2023-10-19T17:55:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> I want to be cautious of going too far in this direction.\n\nThat's fine.  Thanks.\n"},{"id":"483576","messageId":"ZTK4ZKESDVghzSH8@nand.local","threadId":"60318","inReplyTo":"xmqq34y71phj.fsf@gitster.g","subject":"Re: [PATCH v4 00/17] bloom: changed-path Bloom filters v2 (& sundries)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-20T17:27:00Z","receivedAt":"2023-10-20T17:27:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Oct 18, 2023 at 04:26:48PM -0700, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > (Rebased onto the tip of 'master', which is 3a06386e31 (The fifteenth\n> > batch, 2023-10-04), at the time of writing).\n>\n> Judging from 17/17 that has a free_commit_graph() call in\n> close_commit_graph(), that was merged in the eighteenth batch,\n> the above is probably untrue.  I'll apply to the current master and\n> see how it goes instead.\n\nWorse than that, I sent this `--in-reply-to` the wrong thread :-<.\n\nSorry about that, and indeed you are right that the correct base for\nthis round should be a9ecda2788 (The eighteenth batch, 2023-10-13).\n\nI'm optimistic that with the amount of careful review that this topic\nhas already received, that this round should do the trick. But if there\nare more comments and we end up re-rolling it, I'll break this thread\nand split out the v5 into it's thread to avoid further confusion.\n\n> > Thanks to Jonathan, Peff, and SZEDER who have helped a great deal in\n> > assembling these patches. As usual, a range-diff is included below.\n> > Thanks in advance for your\n> > review!\n>\n> Thanks.\n\nThank you, and sorry for the mistake on my end.\n\nThanks,\nTaylor\n"},{"id":"483722","messageId":"20231023202212.GA5470@szeder.dev","threadId":"60318","inReplyTo":"ZTK4ZKESDVghzSH8@nand.local","subject":"Re: [PATCH v4 00/17] bloom: changed-path Bloom filters v2 (& sundries)","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2023-10-23T20:22:12Z","receivedAt":"2023-10-23T20:22:42Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Fri, Oct 20, 2023 at 01:27:00PM -0400, Taylor Blau wrote:\n> On Wed, Oct 18, 2023 at 04:26:48PM -0700, Junio C Hamano wrote:\n> > Taylor Blau <me@ttaylorr.com> writes:\n> >\n> > > (Rebased onto the tip of 'master', which is 3a06386e31 (The fifteenth\n> > > batch, 2023-10-04), at the time of writing).\n> >\n> > Judging from 17/17 that has a free_commit_graph() call in\n> > close_commit_graph(), that was merged in the eighteenth batch,\n> > the above is probably untrue.  I'll apply to the current master and\n> > see how it goes instead.\n> \n> Worse than that, I sent this `--in-reply-to` the wrong thread :-<.\n> \n> Sorry about that, and indeed you are right that the correct base for\n> this round should be a9ecda2788 (The eighteenth batch, 2023-10-13).\n> \n> I'm optimistic that with the amount of careful review that this topic\n> has already received, that this round should do the trick.\n\nUnfortunately, I can't share this optimism.  This series still lacks\ntests exercising the interaction of different versions of Bloom\nfilters and split commit graphs, and the one such test that I sent a\nwhile ago demonstrates that it's still broken.  And it's getting\nworse: back then I didn't send the related test that merged\ncommit-graph layers containing different Bloom filter versions,\nbecause happened to succeed even back then; but, alas, with this\nseries even that test fails.\n\n"},{"id":"484149","messageId":"ZUARCJ1MmqgXfS4i@nand.local","threadId":"60318","inReplyTo":"20231023202212.GA5470@szeder.dev","subject":"Re: [PATCH v4 00/17] bloom: changed-path Bloom filters v2 (& sundries)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-30T20:24:40Z","receivedAt":"2023-10-30T20:24:44Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Oct 23, 2023 at 10:22:12PM +0200, SZEDER Gábor wrote:\n> On Fri, Oct 20, 2023 at 01:27:00PM -0400, Taylor Blau wrote:\n> > I'm optimistic that with the amount of careful review that this topic\n> > has already received, that this round should do the trick.\n>\n> Unfortunately, I can't share this optimism.  This series still lacks\n> tests exercising the interaction of different versions of Bloom\n> filters and split commit graphs, and the one such test that I sent a\n> while ago demonstrates that it's still broken.  And it's getting\n> worse: back then I didn't send the related test that merged\n> commit-graph layers containing different Bloom filter versions,\n> because happened to succeed even back then; but, alas, with this\n> series even that test fails.\n\nI am very confused here, the tests that you're referring to have been\nadded to (and pass in) this series. What am I missing here?\n\nThanks,\nTaylor\n"},{"id":"486872","messageId":"cover.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1697653929.git.me@ttaylorr.com","subject":"[PATCH v5 00/17] bloom: changed-path Bloom filters v2 (& sundries)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:08:56Z","receivedAt":"2024-01-16T22:09:05Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"(Rebased onto the tip of 'master', which is d4dbce1db5 (The seventh\nbatch, 2024-01-12), at the time of writing).\n\nThis series is a reroll of the combined efforts of [1] and [2] to\nintroduce the v2 changed-path Bloom filters, which fixes a bug in our\nexisting implementation of murmur3 paths with non-ASCII characters (when\nthe \"char\" type is signed).\n\nIn large part, this is the same as the previous round. Like last time,\nthis round addresses the remaining additional issues pointed out by\nSZEDER Gábor. The remaining issues which have been addressed by this\nseries are:\n\n  - Incorrectly reading Bloom filters computed with differing hash\n    versions. This has been corrected by discarding them when a version\n    mismatch is detected.\n\n  - Added a note about the new `commitGraph.changedPathVersion`\n    configuration variable which can cause (un-fixable, see [3])\n    issues in earlier versions of Git which do not yet understand them.\n\nThanks to Jonathan, Peff, and SZEDER who have helped a great deal in\nassembling these patches. As usual, a range-diff is included below.\n\nThanks in advance for your review!\n\n[1]: https://lore.kernel.org/git/cover.1684790529.git.jonathantanmy@google.com/\n[2]: https://lore.kernel.org/git/cover.1691426160.git.me@ttaylorr.com/\n[3]: https://lore.kernel.org/git/Zabr1Glljjgl%2FUUB@nand.local/\n\nJonathan Tan (1):\n  gitformat-commit-graph: describe version 2 of BDAT\n\nTaylor Blau (16):\n  t/t4216-log-bloom.sh: harden `test_bloom_filters_not_used()`\n  revision.c: consult Bloom filters for root commits\n  commit-graph: ensure Bloom filters are read with consistent settings\n  t/helper/test-read-graph.c: extract `dump_graph_info()`\n  bloom.h: make `load_bloom_filter_from_graph()` public\n  t/helper/test-read-graph: implement `bloom-filters` mode\n  t4216: test changed path filters with high bit paths\n  repo-settings: introduce commitgraph.changedPathsVersion\n  commit-graph: new Bloom filter version that fixes murmur3\n  bloom: annotate filters with hash version\n  bloom: prepare to discard incompatible Bloom filters\n  commit-graph.c: unconditionally load Bloom filters\n  commit-graph: drop unnecessary `graph_read_bloom_data_context`\n  object.h: fix mis-aligned flag bits table\n  commit-graph: reuse existing Bloom filters where possible\n  bloom: introduce `deinit_bloom_filters()`\n\n Documentation/config/commitgraph.txt     |  29 ++-\n Documentation/gitformat-commit-graph.txt |   9 +-\n bloom.c                                  | 208 ++++++++++++++-\n bloom.h                                  |  38 ++-\n commit-graph.c                           |  64 ++++-\n object.h                                 |   3 +-\n oss-fuzz/fuzz-commit-graph.c             |   2 +-\n repo-settings.c                          |   6 +-\n repository.h                             |   2 +-\n revision.c                               |  26 +-\n t/helper/test-bloom.c                    |   9 +-\n t/helper/test-read-graph.c               |  67 ++++-\n t/t0095-bloom.sh                         |   8 +\n t/t4216-log-bloom.sh                     | 310 ++++++++++++++++++++++-\n 14 files changed, 724 insertions(+), 57 deletions(-)\n\nRange-diff against v4:\n 1:  e0fc51c3fb !  1:  c5e1b3e507 t/t4216-log-bloom.sh: harden `test_bloom_filters_not_used()`\n    @@ Commit message\n         checking for the absence of any data from it from trace2.\n     \n         In the following commit, it will become possible to load Bloom filters\n    -    without using them (e.g., because `commitGraph.changedPathVersion` is\n    -    incompatible with the hash version with which the commit-graph's Bloom\n    -    filters were written).\n    +    without using them (e.g., because the `commitGraph.changedPathVersion`\n    +    introduced later in this series is incompatible with the hash version\n    +    with which the commit-graph's Bloom filters were written).\n     \n         When this is the case, it's possible to initialize the Bloom filter\n         sub-system, while still not using any Bloom filters. When this is the\n 2:  87b09e6266 !  2:  8f32fd5f46 revision.c: consult Bloom filters for root commits\n    @@ revision.c: static int rev_compare_tree(struct rev_info *revs,\n     +\t\t\t\t  int nth_parent)\n      {\n      \tstruct tree *t1 = repo_get_commit_tree(the_repository, commit);\n    -+\tint bloom_ret = 1;\n    ++\tint bloom_ret = -1;\n      \n      \tif (!t1)\n      \t\treturn 0;\n      \n    -+\tif (nth_parent == 1 && revs->bloom_keys_nr) {\n    ++\tif (!nth_parent && revs->bloom_keys_nr) {\n     +\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n     +\t\tif (!bloom_ret)\n     +\t\t\treturn 1;\n    @@ revision.c: static void try_to_simplify_commit(struct rev_info *revs, struct com\n     +\t\t * (if one exists) is relative to the empty tree, using Bloom\n     +\t\t * filters is allowed here.\n     +\t\t */\n    -+\t\tif (rev_same_tree_as_empty(revs, commit, 1))\n    ++\t\tif (rev_same_tree_as_empty(revs, commit, 0))\n      \t\t\tcommit->object.flags |= TREESAME;\n      \t\treturn;\n      \t}\n 3:  46d8a41005 !  3:  285b25f1b7 commit-graph: ensure Bloom filters are read with consistent settings\n    @@ t/t4216-log-bloom.sh: test_expect_success 'Bloom generation backfills empty comm\n     +\ttest_must_be_empty err\n     +'\n     +\n    - test_done\n    + corrupt_graph () {\n    +-\tgraph=.git/objects/info/commit-graph &&\n    + \ttest_when_finished \"rm -rf $graph\" &&\n    + \tgit commit-graph write --reachable --changed-paths &&\n    + \tcorrupt_chunk_file $graph \"$@\"\n 4:  4d0190a992 =  4:  0cee8078d4 gitformat-commit-graph: describe version 2 of BDAT\n 5:  3c2057c11c =  5:  1fc8d2828d t/helper/test-read-graph.c: extract `dump_graph_info()`\n 6:  e002e35004 !  6:  03dd7cf30a bloom.h: make `load_bloom_filter_from_graph()` public\n    @@ Commit message\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n      ## bloom.c ##\n    -@@ bloom.c: static inline unsigned char get_bitmask(uint32_t pos)\n    - \treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n    +@@ bloom.c: static int check_bloom_offset(struct commit_graph *g, uint32_t pos,\n    + \treturn -1;\n      }\n      \n     -static int load_bloom_filter_from_graph(struct commit_graph *g,\n 7:  c7016f51cd =  7:  dd9193e404 t/helper/test-read-graph: implement `bloom-filters` mode\n 8:  cef2aac8ba !  8:  aa2416795d t4216: test changed path filters with high bit paths\n    @@\n      ## Metadata ##\n    -Author: Jonathan Tan <jonathantanmy@google.com>\n    +Author: Taylor Blau <me@ttaylorr.com>\n     \n      ## Commit message ##\n         t4216: test changed path filters with high bit paths\n    @@ t/t4216-log-bloom.sh: test_expect_success 'merge graph layers with incompatible\n     +\t)\n     +'\n     +\n    - test_done\n    + corrupt_graph () {\n    + \ttest_when_finished \"rm -rf $graph\" &&\n    + \tgit commit-graph write --reachable --changed-paths &&\n 9:  36d4e2202e !  9:  a77dcc99b4 repo-settings: introduce commitgraph.changedPathsVersion\n    @@\n      ## Metadata ##\n    -Author: Jonathan Tan <jonathantanmy@google.com>\n    +Author: Taylor Blau <me@ttaylorr.com>\n     \n      ## Commit message ##\n         repo-settings: introduce commitgraph.changedPathsVersion\n    @@ Documentation/config/commitgraph.txt: commitGraph.maxNewFilters::\n     +\n     +commitGraph.changedPathsVersion::\n     +\tSpecifies the version of the changed-path Bloom filters that Git will read and\n    -+\twrite. May be -1, 0 or 1.\n    ++\twrite. May be -1, 0 or 1. Note that values greater than 1 may be\n    ++\tincompatible with older versions of Git which do not yet understand\n    ++\tthose versions. Use caution when operating in a mixed-version\n    ++\tenvironment.\n     ++\n     +Defaults to -1.\n     ++\n    @@ commit-graph.c: struct commit_graph *parse_commit_graph(struct repo_settings *s,\n      \n     -\tif (s->commit_graph_read_changed_paths) {\n     +\tif (s->commit_graph_changed_paths_version) {\n    - \t\tpair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n    - \t\t\t   &graph->chunk_bloom_indexes);\n    + \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n    + \t\t\t   graph_read_bloom_index, graph);\n      \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n    +@@ commit-graph.c: static void validate_mixed_bloom_settings(struct commit_graph *g)\n    + \t\t}\n    + \n    + \t\tif (g->bloom_filter_settings->bits_per_entry != settings->bits_per_entry ||\n    +-\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes) {\n    ++\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes ||\n    ++\t\t    g->bloom_filter_settings->hash_version != settings->hash_version) {\n    + \t\t\tg->chunk_bloom_indexes = NULL;\n    + \t\t\tg->chunk_bloom_data = NULL;\n    + \t\t\tFREE_AND_NULL(g->bloom_filter_settings);\n     \n      ## oss-fuzz/fuzz-commit-graph.c ##\n     @@ oss-fuzz/fuzz-commit-graph.c: int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n    @@ repository.h: struct repo_settings {\n      \tint gc_write_commit_graph;\n      \tint fetch_write_commit_graph;\n      \tint command_requires_full_index;\n    +\n    + ## t/t4216-log-bloom.sh ##\n    +@@ t/t4216-log-bloom.sh: test_expect_success 'setup for mixed Bloom setting tests' '\n    + \tdone\n    + '\n    + \n    +-test_expect_success 'ensure incompatible Bloom filters are ignored' '\n    ++test_expect_success 'ensure Bloom filters with incompatible settings are ignored' '\n    + \t# Compute Bloom filters with \"unusual\" settings.\n    + \tgit -C $repo rev-parse one >in &&\n    + \tGIT_TEST_BLOOM_SETTINGS_NUM_HASHES=3 git -C $repo commit-graph write \\\n    +@@ t/t4216-log-bloom.sh: test_expect_success 'merge graph layers with incompatible Bloom settings' '\n    + \ttest_must_be_empty err\n    + '\n    + \n    ++test_expect_success 'ensure Bloom filter with incompatible versions are ignored' '\n    ++\trm \"$repo/$graph\" &&\n    ++\n    ++\tgit -C $repo log --oneline --no-decorate -- $CENT >expect &&\n    ++\n    ++\t# Compute v1 Bloom filters for commits at the bottom.\n    ++\tgit -C $repo rev-parse HEAD^ >in &&\n    ++\tgit -C $repo commit-graph write --stdin-commits --changed-paths \\\n    ++\t\t--split <in &&\n    ++\n    ++\t# Compute v2 Bloomfilters for the rest of the commits at the top.\n    ++\tgit -C $repo rev-parse HEAD >in &&\n    ++\tgit -C $repo -c commitGraph.changedPathsVersion=2 commit-graph write \\\n    ++\t\t--stdin-commits --changed-paths --split=no-merge <in &&\n    ++\n    ++\ttest_line_count = 2 $repo/$chain &&\n    ++\n    ++\tgit -C $repo log --oneline --no-decorate -- $CENT >actual 2>err &&\n    ++\ttest_cmp expect actual &&\n    ++\n    ++\tlayer=\"$(head -n 1 $repo/$chain)\" &&\n    ++\tcat >expect.err <<-EOF &&\n    ++\twarning: disabling Bloom filters for commit-graph layer $SQ$layer$SQ due to incompatible settings\n    ++\tEOF\n    ++\ttest_cmp expect.err err\n    ++'\n    ++\n    + get_first_changed_path_filter () {\n    + \ttest-tool read-graph bloom-filters >filters.dat &&\n    + \thead -n 1 filters.dat\n10:  f6ab427ead ! 10:  f0f22e852c commit-graph: new filter ver. that fixes murmur3\n    @@\n      ## Metadata ##\n    -Author: Jonathan Tan <jonathantanmy@google.com>\n    +Author: Taylor Blau <me@ttaylorr.com>\n     \n      ## Commit message ##\n    -    commit-graph: new filter ver. that fixes murmur3\n    +    commit-graph: new Bloom filter version that fixes murmur3\n     \n         The murmur3 implementation in bloom.c has a bug when converting series\n         of 4 bytes into network-order integers when char is signed (which is\n    @@ Documentation/config/commitgraph.txt: commitGraph.readChangedPaths::\n      \n      commitGraph.changedPathsVersion::\n      \tSpecifies the version of the changed-path Bloom filters that Git will read and\n    --\twrite. May be -1, 0 or 1.\n    -+\twrite. May be -1, 0, 1, or 2.\n    - +\n    - Defaults to -1.\n    - +\n    +-\twrite. May be -1, 0 or 1. Note that values greater than 1 may be\n    ++\twrite. May be -1, 0, 1, or 2. Note that values greater than 1 may be\n    + \tincompatible with older versions of Git which do not yet understand\n    + \tthose versions. Use caution when operating in a mixed-version\n    + \tenvironment.\n     @@ Documentation/config/commitgraph.txt: filters when instructed to write.\n      If 1, Git will only read version 1 Bloom filters, and will write version 1\n      Bloom filters.\n    @@ bloom.h: int load_bloom_filter_from_graph(struct commit_graph *g,\n      \t\t    size_t len,\n     \n      ## commit-graph.c ##\n    -@@ commit-graph.c: static int graph_read_oid_lookup(const unsigned char *chunk_start,\n    +@@ commit-graph.c: static int graph_read_bloom_index(const unsigned char *chunk_start,\n      \treturn 0;\n      }\n      \n    @@ commit-graph.c: static int graph_read_oid_lookup(const unsigned char *chunk_star\n     +\tstruct graph_read_bloom_data_context *c = data;\n     +\tstruct commit_graph *g = c->g;\n      \tuint32_t hash_version;\n    --\tg->chunk_bloom_data = chunk_start;\n    - \thash_version = get_be32(chunk_start);\n      \n    --\tif (hash_version != 1)\n    -+\tif (*c->commit_graph_changed_paths_version == -1) {\n    + \tif (chunk_size < BLOOMDATA_CHUNK_HEADER_SIZE) {\n    +@@ commit-graph.c: static int graph_read_bloom_data(const unsigned char *chunk_start,\n    + \t\treturn -1;\n    + \t}\n    + \n    ++\thash_version = get_be32(chunk_start);\n    ++\n    ++\tif (*c->commit_graph_changed_paths_version == -1)\n     +\t\t*c->commit_graph_changed_paths_version = hash_version;\n    -+\t} else if (hash_version != *c->commit_graph_changed_paths_version) {\n    - \t\treturn 0;\n    -+\t}\n    - \n    -+\tg->chunk_bloom_data = chunk_start;\n    ++\telse if (hash_version != *c->commit_graph_changed_paths_version)\n    ++\t\treturn 0;\n    ++\n    + \tg->chunk_bloom_data = chunk_start;\n    + \tg->chunk_bloom_data_size = chunk_size;\n    +-\thash_version = get_be32(chunk_start);\n    +-\n    +-\tif (hash_version != 1)\n    +-\t\treturn 0;\n    +-\n      \tg->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n      \tg->bloom_filter_settings->hash_version = hash_version;\n      \tg->bloom_filter_settings->num_hashes = get_be32(chunk_start + 4);\n    @@ commit-graph.c: struct commit_graph *parse_commit_graph(struct repo_settings *s,\n     +\t\t\t.g = graph,\n     +\t\t\t.commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version\n     +\t\t};\n    - \t\tpair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n    - \t\t\t   &graph->chunk_bloom_indexes);\n    + \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n    + \t\t\t   graph_read_bloom_index, graph);\n      \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n     -\t\t\t   graph_read_bloom_data, graph);\n     +\t\t\t   graph_read_bloom_data, &context);\n    @@ t/t4216-log-bloom.sh: test_expect_success 'version 1 changed-path used when vers\n     +\t)\n     +'\n     +\n    - test_done\n    + corrupt_graph () {\n    + \ttest_when_finished \"rm -rf $graph\" &&\n    + \tgit commit-graph write --reachable --changed-paths &&\n11:  dc69b28329 = 11:  b56e94cad7 bloom: annotate filters with hash version\n12:  85dbdc4ed2 = 12:  ddfd1ba32a bloom: prepare to discard incompatible Bloom filters\n13:  3ff669a622 ! 13:  72aabd289b commit-graph.c: unconditionally load Bloom filters\n    @@ Metadata\n      ## Commit message ##\n         commit-graph.c: unconditionally load Bloom filters\n     \n    -    In 9e4df4da07 (commit-graph: new filter ver. that fixes murmur3,\n    -    2023-08-01), we began ignoring the Bloom data (\"BDAT\") chunk for\n    -    commit-graphs whose Bloom filters were computed using a hash version\n    -    incompatible with the value of `commitGraph.changedPathVersion`.\n    +    In an earlier commit, we began ignoring the Bloom data (\"BDAT\") chunk\n    +    for commit-graphs whose Bloom filters were computed using a hash version\n    +      incompatible with the value of `commitGraph.changedPathVersion`.\n     \n         Now that the Bloom API has been hardened to discard these incompatible\n         filters (with the exception of low-level APIs), we can safely load these\n    @@ Commit message\n     \n      ## commit-graph.c ##\n     @@ commit-graph.c: static int graph_read_bloom_data(const unsigned char *chunk_start,\n    - \tuint32_t hash_version;\n    + \n      \thash_version = get_be32(chunk_start);\n      \n    --\tif (*c->commit_graph_changed_paths_version == -1) {\n    +-\tif (*c->commit_graph_changed_paths_version == -1)\n     -\t\t*c->commit_graph_changed_paths_version = hash_version;\n    --\t} else if (hash_version != *c->commit_graph_changed_paths_version) {\n    +-\telse if (hash_version != *c->commit_graph_changed_paths_version)\n     -\t\treturn 0;\n    --\t}\n     -\n      \tg->chunk_bloom_data = chunk_start;\n    + \tg->chunk_bloom_data_size = chunk_size;\n      \tg->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n    - \tg->bloom_filter_settings->hash_version = hash_version;\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n      \tctx->write_generation_data = (get_configured_generation_version(r) == 2);\n      \tctx->num_generation_data_overflows = 0;\n14:  1c78e3d178 ! 14:  526beb9766 commit-graph: drop unnecessary `graph_read_bloom_data_context`\n    @@ Commit message\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n      ## commit-graph.c ##\n    -@@ commit-graph.c: static int graph_read_oid_lookup(const unsigned char *chunk_start,\n    +@@ commit-graph.c: static int graph_read_bloom_index(const unsigned char *chunk_start,\n      \treturn 0;\n      }\n      \n    @@ commit-graph.c: static int graph_read_oid_lookup(const unsigned char *chunk_star\n     -\tstruct commit_graph *g = c->g;\n     +\tstruct commit_graph *g = data;\n      \tuint32_t hash_version;\n    - \thash_version = get_be32(chunk_start);\n      \n    + \tif (chunk_size < BLOOMDATA_CHUNK_HEADER_SIZE) {\n     @@ commit-graph.c: struct commit_graph *parse_commit_graph(struct repo_settings *s,\n      \t}\n      \n    @@ commit-graph.c: struct commit_graph *parse_commit_graph(struct repo_settings *s,\n     -\t\t\t.g = graph,\n     -\t\t\t.commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version\n     -\t\t};\n    - \t\tpair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n    - \t\t\t   &graph->chunk_bloom_indexes);\n    + \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n    + \t\t\t   graph_read_bloom_index, graph);\n      \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n     -\t\t\t   graph_read_bloom_data, &context);\n     +\t\t\t   graph_read_bloom_data, graph);\n15:  a289514faa = 15:  c683697efa object.h: fix mis-aligned flag bits table\n16:  6a12e39e7f ! 16:  4bf043be9a commit-graph: reuse existing Bloom filters where possible\n    @@ Metadata\n      ## Commit message ##\n         commit-graph: reuse existing Bloom filters where possible\n     \n    -    In 9e4df4da07 (commit-graph: new filter ver. that fixes murmur3,\n    -    2023-08-01), a bug was described where it's possible for Git to produce\n    -    non-murmur3 hashes when the platform's \"char\" type is signed, and there\n    -    are paths with characters whose highest bit is set (i.e. all characters\n    -    >= 0x80).\n    +    In an earlier commit, a bug was described where it's possible for Git to\n    +    produce non-murmur3 hashes when the platform's \"char\" type is signed,\n    +    and there are paths with characters whose highest bit is set (i.e. all\n    +    characters >= 0x80).\n     \n         That patch allows the caller to control which version of Bloom filters\n         are read and written. However, even on platforms with a signed \"char\"\n    @@ t/t4216-log-bloom.sh: test_expect_success 'when writing commit graph, do not reu\n     +\ttest_filter_upgraded 1 trace2.txt\n     +'\n     +\n    - test_done\n    + corrupt_graph () {\n    + \ttest_when_finished \"rm -rf $graph\" &&\n    + \tgit commit-graph write --reachable --changed-paths &&\n17:  8942f205c8 = 17:  7daa0d8833 bloom: introduce `deinit_bloom_filters()`\n\nbase-commit: d4dbce1db5cd227a57074bcfc7ec9f0655961bba\n-- \n2.43.0.334.gd4dbce1db5.dirty\n"},{"id":"486873","messageId":"c5e1b3e507b3cb8fd3faac7056ada82d45cb7a03.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 01/17] t/t4216-log-bloom.sh: harden `test_bloom_filters_not_used()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:06Z","receivedAt":"2024-01-16T22:09:08Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The existing implementation of test_bloom_filters_not_used() asserts\nthat the Bloom filter sub-system has not been initialized at all, by\nchecking for the absence of any data from it from trace2.\n\nIn the following commit, it will become possible to load Bloom filters\nwithout using them (e.g., because the `commitGraph.changedPathVersion`\nintroduced later in this series is incompatible with the hash version\nwith which the commit-graph's Bloom filters were written).\n\nWhen this is the case, it's possible to initialize the Bloom filter\nsub-system, while still not using any Bloom filters. When this is the\ncase, check that the data dump from the Bloom sub-system is all zeros,\nindicating that no filters were used.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/t4216-log-bloom.sh | 14 +++++++++++++-\n 1 file changed, 13 insertions(+), 1 deletion(-)\n\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 2ba0324a69..b7baf49d62 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -82,7 +82,19 @@ test_bloom_filters_used () {\n test_bloom_filters_not_used () {\n \tlog_args=$1\n \tsetup \"$log_args\" &&\n-\t! grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\" &&\n+\n+\tif grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\"\n+\tthen\n+\t\t# if the Bloom filter system is initialized, ensure that no\n+\t\t# filters were used\n+\t\tdata=\"statistics:{\"\n+\t\tdata=\"$data\\\"filter_not_present\\\":0,\"\n+\t\tdata=\"$data\\\"maybe\\\":0,\"\n+\t\tdata=\"$data\\\"definitely_not\\\":0,\"\n+\t\tdata=\"$data\\\"false_positive\\\":0}\"\n+\n+\t\tgrep -q \"$data\" \"$TRASH_DIRECTORY/trace.perf\"\n+\tfi &&\n \ttest_cmp log_wo_bloom log_w_bloom\n }\n \n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486874","messageId":"8f32fd5f460accd07f0d8c0a63f1a4580817a341.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 02/17] revision.c: consult Bloom filters for root commits","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:08Z","receivedAt":"2024-01-16T22:09:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The commit-graph stores changed-path Bloom filters which represent the\nset of paths included in a tree-level diff between a commit's root tree\nand that of its parent.\n\nWhen a commit has no parents, the tree-diff is computed against that\ncommit's root tree and the empty tree. In other words, every path in\nthat commit's tree is stored in the Bloom filter (since they all appear\nin the diff).\n\nConsult these filters during pathspec-limited traversals in the function\n`rev_same_tree_as_empty()`. Doing so yields a performance improvement\nwhere we can avoid enumerating the full set of paths in a parentless\ncommit's root tree when we know that the path(s) of interest were not\nlisted in that commit's changed-path Bloom filter.\n\nSuggested-by: SZEDER Gábor <szeder.dev@gmail.com>\nOriginal-patch-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n revision.c           | 26 ++++++++++++++++++++++----\n t/t4216-log-bloom.sh |  8 ++++++--\n 2 files changed, 28 insertions(+), 6 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 2424c9bd67..0e6f7d02b6 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -833,17 +833,28 @@ static int rev_compare_tree(struct rev_info *revs,\n \treturn tree_difference;\n }\n \n-static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit)\n+static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n+\t\t\t\t  int nth_parent)\n {\n \tstruct tree *t1 = repo_get_commit_tree(the_repository, commit);\n+\tint bloom_ret = -1;\n \n \tif (!t1)\n \t\treturn 0;\n \n+\tif (!nth_parent && revs->bloom_keys_nr) {\n+\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n+\t\tif (!bloom_ret)\n+\t\t\treturn 1;\n+\t}\n+\n \ttree_difference = REV_TREE_SAME;\n \trevs->pruning.flags.has_changes = 0;\n \tdiff_tree_oid(NULL, &t1->object.oid, \"\", &revs->pruning);\n \n+\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n+\t\tcount_bloom_filter_false_positive++;\n+\n \treturn tree_difference == REV_TREE_SAME;\n }\n \n@@ -881,7 +892,7 @@ static int compact_treesame(struct rev_info *revs, struct commit *commit, unsign\n \t\tif (nth_parent != 0)\n \t\t\tdie(\"compact_treesame %u\", nth_parent);\n \t\told_same = !!(commit->object.flags & TREESAME);\n-\t\tif (rev_same_tree_as_empty(revs, commit))\n+\t\tif (rev_same_tree_as_empty(revs, commit, nth_parent))\n \t\t\tcommit->object.flags |= TREESAME;\n \t\telse\n \t\t\tcommit->object.flags &= ~TREESAME;\n@@ -977,7 +988,14 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \t\treturn;\n \n \tif (!commit->parents) {\n-\t\tif (rev_same_tree_as_empty(revs, commit))\n+\t\t/*\n+\t\t * Pretend as if we are comparing ourselves to the\n+\t\t * (non-existent) first parent of this commit object. Even\n+\t\t * though no such parent exists, its changed-path Bloom filter\n+\t\t * (if one exists) is relative to the empty tree, using Bloom\n+\t\t * filters is allowed here.\n+\t\t */\n+\t\tif (rev_same_tree_as_empty(revs, commit, 0))\n \t\t\tcommit->object.flags |= TREESAME;\n \t\treturn;\n \t}\n@@ -1058,7 +1076,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \n \t\tcase REV_TREE_NEW:\n \t\t\tif (revs->remove_empty_trees &&\n-\t\t\t    rev_same_tree_as_empty(revs, p)) {\n+\t\t\t    rev_same_tree_as_empty(revs, p, nth_parent)) {\n \t\t\t\t/* We are adding all the specified\n \t\t\t\t * paths from this parent, so the\n \t\t\t\t * history beyond this parent is not\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex b7baf49d62..cc6ebc8140 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -88,7 +88,11 @@ test_bloom_filters_not_used () {\n \t\t# if the Bloom filter system is initialized, ensure that no\n \t\t# filters were used\n \t\tdata=\"statistics:{\"\n-\t\tdata=\"$data\\\"filter_not_present\\\":0,\"\n+\t\t# unusable filters (e.g., those computed with a\n+\t\t# different value of commitGraph.changedPathsVersion)\n+\t\t# are counted in the filter_not_present bucket, so any\n+\t\t# value is OK there.\n+\t\tdata=\"$data\\\"filter_not_present\\\":[0-9][0-9]*,\"\n \t\tdata=\"$data\\\"maybe\\\":0,\"\n \t\tdata=\"$data\\\"definitely_not\\\":0,\"\n \t\tdata=\"$data\\\"false_positive\\\":0}\"\n@@ -175,7 +179,7 @@ test_expect_success 'setup - add commit-graph to the chain with Bloom filters' '\n \n test_bloom_filters_used_when_some_filters_are_missing () {\n \tlog_args=$1\n-\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"maybe\\\":6,\\\"definitely_not\\\":9\"\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"maybe\\\":6,\\\"definitely_not\\\":10\"\n \tsetup \"$log_args\" &&\n \tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" &&\n \ttest_cmp log_wo_bloom log_w_bloom\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486875","messageId":"285b25f1b7cca0a64ec410613135d58f73133722.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 03/17] commit-graph: ensure Bloom filters are read with consistent settings","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:11Z","receivedAt":"2024-01-16T22:09:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The changed-path Bloom filter mechanism is parameterized by a couple of\nvariables, notably the number of bits per hash (typically \"m\" in Bloom\nfilter literature) and the number of hashes themselves (typically \"k\").\n\nIt is critically important that filters are read with the Bloom filter\nsettings that they were written with. Failing to do so would mean that\neach query is liable to compute different fingerprints, meaning that the\nfilter itself could return a false negative. This goes against a basic\nassumption of using Bloom filters (that they may return false positives,\nbut never false negatives) and can lead to incorrect results.\n\nWe have some existing logic to carry forward existing Bloom filter\nsettings from one layer to the next. In `write_commit_graph()`, we have\nsomething like:\n\n    if (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\n        struct commit_graph *g = ctx->r->objects->commit_graph;\n\n        /* We have changed-paths already. Keep them in the next graph */\n        if (g && g->chunk_bloom_data) {\n            ctx->changed_paths = 1;\n            ctx->bloom_settings = g->bloom_filter_settings;\n        }\n    }\n\n, which drags forward Bloom filter settings across adjacent layers.\n\nThis doesn't quite address all cases, however, since it is possible for\nintermediate layers to contain no Bloom filters at all. For example,\nsuppose we have two layers in a commit-graph chain, say, {G1, G2}. If G1\ncontains Bloom filters, but G2 doesn't, a new G3 (whose base graph is\nG2) may be written with arbitrary Bloom filter settings, because we only\ncheck the immediately adjacent layer's settings for compatibility.\n\nThis behavior has existed since the introduction of changed-path Bloom\nfilters. But in practice, this is not such a big deal, since the only\nway up until this point to modify the Bloom filter settings at write\ntime is with the undocumented environment variables:\n\n  - GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\n  - GIT_TEST_BLOOM_SETTINGS_NUM_HASHES\n  - GIT_TEST_BLOOM_SETTINGS_MAX_CHANGED_PATHS\n\n(it is still possible to tweak MAX_CHANGED_PATHS between layers, but\nthis does not affect reads, so is allowed to differ across multiple\ngraph layers).\n\nBut in future commits, we will introduce another parameter to change the\nhash algorithm used to compute Bloom fingerprints itself. This will be\nexposed via a configuration setting, making this foot-gun easier to use.\n\nTo prevent this potential issue, validate that all layers of a split\ncommit-graph have compatible settings with the newest layer which\ncontains Bloom filters.\n\nReported-by: SZEDER Gábor <szeder.dev@gmail.com>\nOriginal-test-by: SZEDER Gábor <szeder.dev@gmail.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n commit-graph.c       | 25 +++++++++++++++++\n t/t4216-log-bloom.sh | 65 +++++++++++++++++++++++++++++++++++++++++++-\n 2 files changed, 89 insertions(+), 1 deletion(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex bba316913c..00113b0f62 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -543,6 +543,30 @@ static int validate_mixed_generation_chain(struct commit_graph *g)\n \treturn 0;\n }\n \n+static void validate_mixed_bloom_settings(struct commit_graph *g)\n+{\n+\tstruct bloom_filter_settings *settings = NULL;\n+\tfor (; g; g = g->base_graph) {\n+\t\tif (!g->bloom_filter_settings)\n+\t\t\tcontinue;\n+\t\tif (!settings) {\n+\t\t\tsettings = g->bloom_filter_settings;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (g->bloom_filter_settings->bits_per_entry != settings->bits_per_entry ||\n+\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes) {\n+\t\t\tg->chunk_bloom_indexes = NULL;\n+\t\t\tg->chunk_bloom_data = NULL;\n+\t\t\tFREE_AND_NULL(g->bloom_filter_settings);\n+\n+\t\t\twarning(_(\"disabling Bloom filters for commit-graph \"\n+\t\t\t\t  \"layer '%s' due to incompatible settings\"),\n+\t\t\t\toid_to_hex(&g->oid));\n+\t\t}\n+\t}\n+}\n+\n static int add_graph_to_chain(struct commit_graph *g,\n \t\t\t      struct commit_graph *chain,\n \t\t\t      struct object_id *oids,\n@@ -666,6 +690,7 @@ struct commit_graph *load_commit_graph_chain_fd_st(struct repository *r,\n \t}\n \n \tvalidate_mixed_generation_chain(graph_chain);\n+\tvalidate_mixed_bloom_settings(graph_chain);\n \n \tfree(oids);\n \tfclose(fp);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex cc6ebc8140..20b0cf0c0e 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -421,8 +421,71 @@ test_expect_success 'Bloom generation backfills empty commits' '\n \t)\n '\n \n+graph=.git/objects/info/commit-graph\n+graphdir=.git/objects/info/commit-graphs\n+chain=$graphdir/commit-graph-chain\n+\n+test_expect_success 'setup for mixed Bloom setting tests' '\n+\trepo=mixed-bloom-settings &&\n+\n+\tgit init $repo &&\n+\tfor i in one two three\n+\tdo\n+\t\ttest_commit -C $repo $i file || return 1\n+\tdone\n+'\n+\n+test_expect_success 'ensure incompatible Bloom filters are ignored' '\n+\t# Compute Bloom filters with \"unusual\" settings.\n+\tgit -C $repo rev-parse one >in &&\n+\tGIT_TEST_BLOOM_SETTINGS_NUM_HASHES=3 git -C $repo commit-graph write \\\n+\t\t--stdin-commits --changed-paths --split <in &&\n+\tlayer=$(head -n 1 $repo/$chain) &&\n+\n+\t# A commit-graph layer without Bloom filters \"hides\" the layers\n+\t# below ...\n+\tgit -C $repo rev-parse two >in &&\n+\tgit -C $repo commit-graph write --stdin-commits --no-changed-paths \\\n+\t\t--split=no-merge <in &&\n+\n+\t# Another commit-graph layer that has Bloom filters, but with\n+\t# standard settings, and is thus incompatible with the base\n+\t# layer written above.\n+\tgit -C $repo rev-parse HEAD >in &&\n+\tgit -C $repo commit-graph write --stdin-commits --changed-paths \\\n+\t\t--split=no-merge <in &&\n+\n+\ttest_line_count = 3 $repo/$chain &&\n+\n+\t# Ensure that incompatible Bloom filters are ignored.\n+\tgit -C $repo -c core.commitGraph=false log --oneline --no-decorate -- file \\\n+\t\t>expect 2>err &&\n+\tgit -C $repo log --oneline --no-decorate -- file >actual 2>err &&\n+\ttest_cmp expect actual &&\n+\tgrep \"disabling Bloom filters for commit-graph layer .$layer.\" err\n+'\n+\n+test_expect_success 'merge graph layers with incompatible Bloom settings' '\n+\t# Ensure that incompatible Bloom filters are ignored when\n+\t# merging existing layers.\n+\tgit -C $repo commit-graph write --reachable --changed-paths 2>err &&\n+\tgrep \"disabling Bloom filters for commit-graph layer .$layer.\" err &&\n+\n+\ttest_path_is_file $repo/$graph &&\n+\ttest_dir_is_empty $repo/$graphdir &&\n+\n+\tgit -C $repo -c core.commitGraph=false log --oneline --no-decorate -- \\\n+\t\tfile >expect &&\n+\ttrace_out=\"$(pwd)/trace.perf\" &&\n+\tGIT_TRACE2_PERF=\"$trace_out\" \\\n+\t\tgit -C $repo log --oneline --no-decorate -- file >actual 2>err &&\n+\n+\ttest_cmp expect actual &&\n+\tgrep \"statistics:{\\\"filter_not_present\\\":0,\" trace.perf &&\n+\ttest_must_be_empty err\n+'\n+\n corrupt_graph () {\n-\tgraph=.git/objects/info/commit-graph &&\n \ttest_when_finished \"rm -rf $graph\" &&\n \tgit commit-graph write --reachable --changed-paths &&\n \tcorrupt_chunk_file $graph \"$@\"\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486876","messageId":"0cee8078d42ffccc410588a14ff184edbe07b7d9.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 04/17] gitformat-commit-graph: describe version 2 of BDAT","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:14Z","receivedAt":"2024-01-16T22:09:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jonathan Tan <jonathantanmy@google.com>\n\nThe code change to Git to support version 2 will be done in subsequent\ncommits.\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/gitformat-commit-graph.txt | 9 ++++++---\n 1 file changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/gitformat-commit-graph.txt b/Documentation/gitformat-commit-graph.txt\nindex 31cad585e2..3e906e8030 100644\n--- a/Documentation/gitformat-commit-graph.txt\n+++ b/Documentation/gitformat-commit-graph.txt\n@@ -142,13 +142,16 @@ All multi-byte numbers are in network byte order.\n \n ==== Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n     * It starts with header consisting of three unsigned 32-bit integers:\n-      - Version of the hash algorithm being used. We currently only support\n-\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n+      - Version of the hash algorithm being used. We currently support\n+\tvalue 2 which corresponds to the 32-bit version of the murmur3 hash\n \timplemented exactly as described in\n \thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n \thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n \tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n-\tin Probabilistic Verification\"\n+\tin Probabilistic Verification\". Version 1 Bloom filters have a bug that appears\n+\twhen char is signed and the repository has path names that have characters >=\n+\t0x80; Git supports reading and writing them, but this ability will be removed\n+\tin a future version of Git.\n       - The number of times a path is hashed and hence the number of bit positions\n \t      that cumulatively determine whether a file is present in the commit.\n       - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486877","messageId":"1fc8d2828d8a40ce04cea646b43d03871b6a224b.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 05/17] t/helper/test-read-graph.c: extract `dump_graph_info()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:17Z","receivedAt":"2024-01-16T22:09:19Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Prepare for the 'read-graph' test helper to perform other tasks besides\ndumping high-level information about the commit-graph by extracting its\nmain routine into a separate function.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/helper/test-read-graph.c | 31 ++++++++++++++++++-------------\n 1 file changed, 18 insertions(+), 13 deletions(-)\n\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 8c7a83f578..3375392f6c 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -5,20 +5,8 @@\n #include \"bloom.h\"\n #include \"setup.h\"\n \n-int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n+static void dump_graph_info(struct commit_graph *graph)\n {\n-\tstruct commit_graph *graph = NULL;\n-\tstruct object_directory *odb;\n-\n-\tsetup_git_directory();\n-\todb = the_repository->objects->odb;\n-\n-\tprepare_repo_settings(the_repository);\n-\n-\tgraph = read_commit_graph_one(the_repository, odb);\n-\tif (!graph)\n-\t\treturn 1;\n-\n \tprintf(\"header: %08x %d %d %d %d\\n\",\n \t\tntohl(*(uint32_t*)graph->data),\n \t\t*(unsigned char*)(graph->data + 4),\n@@ -57,6 +45,23 @@ int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n \tif (graph->topo_levels)\n \t\tprintf(\" topo_levels\");\n \tprintf(\"\\n\");\n+}\n+\n+int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n+{\n+\tstruct commit_graph *graph = NULL;\n+\tstruct object_directory *odb;\n+\n+\tsetup_git_directory();\n+\todb = the_repository->objects->odb;\n+\n+\tprepare_repo_settings(the_repository);\n+\n+\tgraph = read_commit_graph_one(the_repository, odb);\n+\tif (!graph)\n+\t\treturn 1;\n+\n+\tdump_graph_info(graph);\n \n \tUNLEAK(graph);\n \n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486878","messageId":"03dd7cf30a5c35ad0bffaf5c8141fbf59ae5c84b.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 06/17] bloom.h: make `load_bloom_filter_from_graph()` public","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:20Z","receivedAt":"2024-01-16T22:09:22Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Prepare for a future commit to use the load_bloom_filter_from_graph()\nfunction directly to load specific Bloom filters out of the commit-graph\nfor manual inspection (to be used during tests).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c | 6 +++---\n bloom.h | 5 +++++\n 2 files changed, 8 insertions(+), 3 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex e529f7605c..401999ed3c 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -48,9 +48,9 @@ static int check_bloom_offset(struct commit_graph *g, uint32_t pos,\n \treturn -1;\n }\n \n-static int load_bloom_filter_from_graph(struct commit_graph *g,\n-\t\t\t\t\tstruct bloom_filter *filter,\n-\t\t\t\t\tuint32_t graph_pos)\n+int load_bloom_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t struct bloom_filter *filter,\n+\t\t\t\t uint32_t graph_pos)\n {\n \tuint32_t lex_pos, start_index, end_index;\n \ndiff --git a/bloom.h b/bloom.h\nindex adde6dfe21..1e4f612d2c 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -3,6 +3,7 @@\n \n struct commit;\n struct repository;\n+struct commit_graph;\n \n struct bloom_filter_settings {\n \t/*\n@@ -68,6 +69,10 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+int load_bloom_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t struct bloom_filter *filter,\n+\t\t\t\t uint32_t graph_pos);\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486879","messageId":"dd9193e404ef7e896d5a7e40788a47d76872b3b9.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 07/17] t/helper/test-read-graph: implement `bloom-filters` mode","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:22Z","receivedAt":"2024-01-16T22:09:24Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Implement a mode of the \"read-graph\" test helper to dump out the\nhexadecimal contents of the Bloom filter(s) contained in a commit-graph.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/helper/test-read-graph.c | 44 +++++++++++++++++++++++++++++++++-----\n 1 file changed, 39 insertions(+), 5 deletions(-)\n\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 3375392f6c..da9ac8584d 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -47,10 +47,32 @@ static void dump_graph_info(struct commit_graph *graph)\n \tprintf(\"\\n\");\n }\n \n-int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n+static void dump_graph_bloom_filters(struct commit_graph *graph)\n+{\n+\tuint32_t i;\n+\n+\tfor (i = 0; i < graph->num_commits + graph->num_commits_in_base; i++) {\n+\t\tstruct bloom_filter filter = { 0 };\n+\t\tsize_t j;\n+\n+\t\tif (load_bloom_filter_from_graph(graph, &filter, i) < 0) {\n+\t\t\tfprintf(stderr, \"missing Bloom filter for graph \"\n+\t\t\t\t\"position %\"PRIu32\"\\n\", i);\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tfor (j = 0; j < filter.len; j++)\n+\t\t\tprintf(\"%02x\", filter.data[j]);\n+\t\tif (filter.len)\n+\t\t\tprintf(\"\\n\");\n+\t}\n+}\n+\n+int cmd__read_graph(int argc, const char **argv)\n {\n \tstruct commit_graph *graph = NULL;\n \tstruct object_directory *odb;\n+\tint ret = 0;\n \n \tsetup_git_directory();\n \todb = the_repository->objects->odb;\n@@ -58,12 +80,24 @@ int cmd__read_graph(int argc UNUSED, const char **argv UNUSED)\n \tprepare_repo_settings(the_repository);\n \n \tgraph = read_commit_graph_one(the_repository, odb);\n-\tif (!graph)\n-\t\treturn 1;\n+\tif (!graph) {\n+\t\tret = 1;\n+\t\tgoto done;\n+\t}\n \n-\tdump_graph_info(graph);\n+\tif (argc <= 1)\n+\t\tdump_graph_info(graph);\n+\telse if (!strcmp(argv[1], \"bloom-filters\"))\n+\t\tdump_graph_bloom_filters(graph);\n+\telse {\n+\t\tfprintf(stderr, \"unknown sub-command: '%s'\\n\", argv[1]);\n+\t\tret = 1;\n+\t}\n \n+done:\n \tUNLEAK(graph);\n \n-\treturn 0;\n+\treturn ret;\n }\n+\n+\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486880","messageId":"aa2416795da980cf5ef1000a3dfc1bc04fa7710f.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 08/17] t4216: test changed path filters with high bit paths","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:25Z","receivedAt":"2024-01-16T22:09:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Subsequent commits will teach Git another version of changed path\nfilter that has different behavior with paths that contain at least\none character with its high bit set, so test the existing behavior as\na baseline.\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/t4216-log-bloom.sh | 51 ++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 51 insertions(+)\n\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 20b0cf0c0e..484dd093cd 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -485,6 +485,57 @@ test_expect_success 'merge graph layers with incompatible Bloom settings' '\n \ttest_must_be_empty err\n '\n \n+get_first_changed_path_filter () {\n+\ttest-tool read-graph bloom-filters >filters.dat &&\n+\thead -n 1 filters.dat\n+}\n+\n+# chosen to be the same under all Unicode normalization forms\n+CENT=$(printf \"\\302\\242\")\n+\n+test_expect_success 'set up repo with high bit path, version 1 changed-path' '\n+\tgit init highbit1 &&\n+\ttest_commit -C highbit1 c1 \"$CENT\" &&\n+\tgit -C highbit1 commit-graph write --reachable --changed-paths\n+'\n+\n+test_expect_success 'setup check value of version 1 changed-path' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\techo \"52a9\" >expect &&\n+\t\tget_first_changed_path_filter >actual\n+\t)\n+'\n+\n+# expect will not match actual if char is unsigned by default. Write the test\n+# in this way, so that a user running this test script can still see if the two\n+# files match. (It will appear as an ordinary success if they match, and a skip\n+# if not.)\n+if test_cmp highbit1/expect highbit1/actual\n+then\n+\ttest_set_prereq SIGNED_CHAR_BY_DEFAULT\n+fi\n+test_expect_success SIGNED_CHAR_BY_DEFAULT 'check value of version 1 changed-path' '\n+\t# Only the prereq matters for this test.\n+\ttrue\n+'\n+\n+test_expect_success 'setup make another commit' '\n+\t# \"git log\" does not use Bloom filters for root commits - see how, in\n+\t# revision.c, rev_compare_tree() (the only code path that eventually calls\n+\t# get_bloom_filter()) is only called by try_to_simplify_commit() when the commit\n+\t# has one parent. Therefore, make another commit so that we perform the tests on\n+\t# a non-root commit.\n+\ttest_commit -C highbit1 anotherc1 \"another$CENT\"\n+'\n+\n+test_expect_success 'version 1 changed-path used when version 1 requested' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n corrupt_graph () {\n \ttest_when_finished \"rm -rf $graph\" &&\n \tgit commit-graph write --reachable --changed-paths &&\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486881","messageId":"a77dcc99b4eb0a19dc6c09a40a84785413502126.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 09/17] repo-settings: introduce commitgraph.changedPathsVersion","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:28Z","receivedAt":"2024-01-16T22:09:30Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A subsequent commit will introduce another version of the changed-path\nfilter in the commit graph file. In order to control which version to\nwrite (and read), a config variable is needed.\n\nTherefore, introduce this config variable. For forwards compatibility,\nteach Git to not read commit graphs when the config variable\nis set to an unsupported version. Because we teach Git this,\ncommitgraph.readChangedPaths is now redundant, so deprecate it and\ndefine its behavior in terms of the config variable we introduce.\n\nThis commit does not change the behavior of writing (Git writes changed\npath filters when explicitly instructed regardless of any config\nvariable), but a subsequent commit will restrict Git such that it will\nonly write when commitgraph.changedPathsVersion is a recognized value.\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/commitgraph.txt | 26 ++++++++++++++++++++++---\n commit-graph.c                       |  5 +++--\n oss-fuzz/fuzz-commit-graph.c         |  2 +-\n repo-settings.c                      |  6 +++++-\n repository.h                         |  2 +-\n t/t4216-log-bloom.sh                 | 29 +++++++++++++++++++++++++++-\n 6 files changed, 61 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config/commitgraph.txt b/Documentation/config/commitgraph.txt\nindex 30604e4a4c..e68cdededa 100644\n--- a/Documentation/config/commitgraph.txt\n+++ b/Documentation/config/commitgraph.txt\n@@ -9,6 +9,26 @@ commitGraph.maxNewFilters::\n \tcommit-graph write` (c.f., linkgit:git-commit-graph[1]).\n \n commitGraph.readChangedPaths::\n-\tIf true, then git will use the changed-path Bloom filters in the\n-\tcommit-graph file (if it exists, and they are present). Defaults to\n-\ttrue. See linkgit:git-commit-graph[1] for more information.\n+\tDeprecated. Equivalent to commitGraph.changedPathsVersion=-1 if true, and\n+\tcommitGraph.changedPathsVersion=0 if false. (If commitGraph.changedPathVersion\n+\tis also set, commitGraph.changedPathsVersion takes precedence.)\n+\n+commitGraph.changedPathsVersion::\n+\tSpecifies the version of the changed-path Bloom filters that Git will read and\n+\twrite. May be -1, 0 or 1. Note that values greater than 1 may be\n+\tincompatible with older versions of Git which do not yet understand\n+\tthose versions. Use caution when operating in a mixed-version\n+\tenvironment.\n++\n+Defaults to -1.\n++\n+If -1, Git will use the version of the changed-path Bloom filters in the\n+repository, defaulting to 1 if there are none.\n++\n+If 0, Git will not read any Bloom filters, and will write version 1 Bloom\n+filters when instructed to write.\n++\n+If 1, Git will only read version 1 Bloom filters, and will write version 1\n+Bloom filters.\n++\n+See linkgit:git-commit-graph[1] for more information.\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 00113b0f62..91c98ebc6c 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -459,7 +459,7 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s,\n \t\t\tgraph->read_generation_data = 1;\n \t}\n \n-\tif (s->commit_graph_read_changed_paths) {\n+\tif (s->commit_graph_changed_paths_version) {\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n \t\t\t   graph_read_bloom_index, graph);\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n@@ -555,7 +555,8 @@ static void validate_mixed_bloom_settings(struct commit_graph *g)\n \t\t}\n \n \t\tif (g->bloom_filter_settings->bits_per_entry != settings->bits_per_entry ||\n-\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes) {\n+\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes ||\n+\t\t    g->bloom_filter_settings->hash_version != settings->hash_version) {\n \t\t\tg->chunk_bloom_indexes = NULL;\n \t\t\tg->chunk_bloom_data = NULL;\n \t\t\tFREE_AND_NULL(g->bloom_filter_settings);\ndiff --git a/oss-fuzz/fuzz-commit-graph.c b/oss-fuzz/fuzz-commit-graph.c\nindex 2992079dd9..325c0b991a 100644\n--- a/oss-fuzz/fuzz-commit-graph.c\n+++ b/oss-fuzz/fuzz-commit-graph.c\n@@ -19,7 +19,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n \t * possible.\n \t */\n \tthe_repository->settings.commit_graph_generation_version = 2;\n-\tthe_repository->settings.commit_graph_read_changed_paths = 1;\n+\tthe_repository->settings.commit_graph_changed_paths_version = 1;\n \tg = parse_commit_graph(&the_repository->settings, (void *)data, size);\n \trepo_clear(the_repository);\n \tfree_commit_graph(g);\ndiff --git a/repo-settings.c b/repo-settings.c\nindex 30cd478762..c821583fe5 100644\n--- a/repo-settings.c\n+++ b/repo-settings.c\n@@ -23,6 +23,7 @@ void prepare_repo_settings(struct repository *r)\n \tint value;\n \tconst char *strval;\n \tint manyfiles;\n+\tint read_changed_paths;\n \n \tif (!r->gitdir)\n \t\tBUG(\"Cannot add settings for uninitialized repository\");\n@@ -53,7 +54,10 @@ void prepare_repo_settings(struct repository *r)\n \t/* Commit graph config or default, does not cascade (simple) */\n \trepo_cfg_bool(r, \"core.commitgraph\", &r->settings.core_commit_graph, 1);\n \trepo_cfg_int(r, \"commitgraph.generationversion\", &r->settings.commit_graph_generation_version, 2);\n-\trepo_cfg_bool(r, \"commitgraph.readchangedpaths\", &r->settings.commit_graph_read_changed_paths, 1);\n+\trepo_cfg_bool(r, \"commitgraph.readchangedpaths\", &read_changed_paths, 1);\n+\trepo_cfg_int(r, \"commitgraph.changedpathsversion\",\n+\t\t     &r->settings.commit_graph_changed_paths_version,\n+\t\t     read_changed_paths ? -1 : 0);\n \trepo_cfg_bool(r, \"gc.writecommitgraph\", &r->settings.gc_write_commit_graph, 1);\n \trepo_cfg_bool(r, \"fetch.writecommitgraph\", &r->settings.fetch_write_commit_graph, 0);\n \ndiff --git a/repository.h b/repository.h\nindex 5f18486f64..f71154e12c 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -29,7 +29,7 @@ struct repo_settings {\n \n \tint core_commit_graph;\n \tint commit_graph_generation_version;\n-\tint commit_graph_read_changed_paths;\n+\tint commit_graph_changed_paths_version;\n \tint gc_write_commit_graph;\n \tint fetch_write_commit_graph;\n \tint command_requires_full_index;\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 484dd093cd..642b960893 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -435,7 +435,7 @@ test_expect_success 'setup for mixed Bloom setting tests' '\n \tdone\n '\n \n-test_expect_success 'ensure incompatible Bloom filters are ignored' '\n+test_expect_success 'ensure Bloom filters with incompatible settings are ignored' '\n \t# Compute Bloom filters with \"unusual\" settings.\n \tgit -C $repo rev-parse one >in &&\n \tGIT_TEST_BLOOM_SETTINGS_NUM_HASHES=3 git -C $repo commit-graph write \\\n@@ -485,6 +485,33 @@ test_expect_success 'merge graph layers with incompatible Bloom settings' '\n \ttest_must_be_empty err\n '\n \n+test_expect_success 'ensure Bloom filter with incompatible versions are ignored' '\n+\trm \"$repo/$graph\" &&\n+\n+\tgit -C $repo log --oneline --no-decorate -- $CENT >expect &&\n+\n+\t# Compute v1 Bloom filters for commits at the bottom.\n+\tgit -C $repo rev-parse HEAD^ >in &&\n+\tgit -C $repo commit-graph write --stdin-commits --changed-paths \\\n+\t\t--split <in &&\n+\n+\t# Compute v2 Bloomfilters for the rest of the commits at the top.\n+\tgit -C $repo rev-parse HEAD >in &&\n+\tgit -C $repo -c commitGraph.changedPathsVersion=2 commit-graph write \\\n+\t\t--stdin-commits --changed-paths --split=no-merge <in &&\n+\n+\ttest_line_count = 2 $repo/$chain &&\n+\n+\tgit -C $repo log --oneline --no-decorate -- $CENT >actual 2>err &&\n+\ttest_cmp expect actual &&\n+\n+\tlayer=\"$(head -n 1 $repo/$chain)\" &&\n+\tcat >expect.err <<-EOF &&\n+\twarning: disabling Bloom filters for commit-graph layer $SQ$layer$SQ due to incompatible settings\n+\tEOF\n+\ttest_cmp expect.err err\n+'\n+\n get_first_changed_path_filter () {\n \ttest-tool read-graph bloom-filters >filters.dat &&\n \thead -n 1 filters.dat\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486882","messageId":"f0f22e852cd616591fd9717c041d5fa5d6bf7381.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 10/17] commit-graph: new Bloom filter version that fixes murmur3","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:31Z","receivedAt":"2024-01-16T22:09:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The murmur3 implementation in bloom.c has a bug when converting series\nof 4 bytes into network-order integers when char is signed (which is\ncontrollable by a compiler option, and the default signedness of char is\nplatform-specific). When a string contains characters with the high bit\nset, this bug causes results that, although internally consistent within\nGit, does not accord with other implementations of murmur3 (thus,\nthe changed path filters wouldn't be readable by other off-the-shelf\nimplementatios of murmur3) and even with Git binaries that were compiled\nwith different signedness of char. This bug affects both how Git writes\nchanged path filters to disk and how Git interprets changed path filters\non disk.\n\nTherefore, introduce a new version (2) of changed path filters that\ncorrects this problem. The existing version (1) is still supported and\nis still the default, but users should migrate away from it as soon\nas possible.\n\nBecause this bug only manifests with characters that have the high bit\nset, it may be possible that some (or all) commits in a given repo would\nhave the same changed path filter both before and after this fix is\napplied. However, in order to determine whether this is the case, the\nchanged paths would first have to be computed, at which point it is not\nmuch more expensive to just compute a new changed path filter.\n\nSo this patch does not include any mechanism to \"salvage\" changed path\nfilters from repositories. There is also no \"mixed\" mode - for each\ninvocation of Git, reading and writing changed path filters are done\nwith the same version number; this version number may be explicitly\nstated (typically if the user knows which version they need) or\nautomatically determined from the version of the existing changed path\nfilters in the repository.\n\nThere is a change in write_commit_graph(). graph_read_bloom_data()\nmakes it possible for chunk_bloom_data to be non-NULL but\nbloom_filter_settings to be NULL, which causes a segfault later on. I\nproduced such a segfault while developing this patch, but couldn't find\na way to reproduce it neither after this complete patch (or before),\nbut in any case it seemed like a good thing to include that might help\nfuture patch authors.\n\nThe value in t0095 was obtained from another murmur3 implementation\nusing the following Go source code:\n\n  package main\n\n  import \"fmt\"\n  import \"github.com/spaolacci/murmur3\"\n\n  func main() {\n          fmt.Printf(\"%x\\n\", murmur3.Sum32([]byte(\"Hello world!\")))\n          fmt.Printf(\"%x\\n\", murmur3.Sum32([]byte{0x99, 0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff}))\n  }\n\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/commitgraph.txt |   5 +-\n bloom.c                              |  69 +++++++++++++++-\n bloom.h                              |   8 +-\n commit-graph.c                       |  37 +++++++--\n t/helper/test-bloom.c                |   9 ++-\n t/t0095-bloom.sh                     |   8 ++\n t/t4216-log-bloom.sh                 | 114 +++++++++++++++++++++++++++\n 7 files changed, 234 insertions(+), 16 deletions(-)\n\ndiff --git a/Documentation/config/commitgraph.txt b/Documentation/config/commitgraph.txt\nindex e68cdededa..7f8c9d6638 100644\n--- a/Documentation/config/commitgraph.txt\n+++ b/Documentation/config/commitgraph.txt\n@@ -15,7 +15,7 @@ commitGraph.readChangedPaths::\n \n commitGraph.changedPathsVersion::\n \tSpecifies the version of the changed-path Bloom filters that Git will read and\n-\twrite. May be -1, 0 or 1. Note that values greater than 1 may be\n+\twrite. May be -1, 0, 1, or 2. Note that values greater than 1 may be\n \tincompatible with older versions of Git which do not yet understand\n \tthose versions. Use caution when operating in a mixed-version\n \tenvironment.\n@@ -31,4 +31,7 @@ filters when instructed to write.\n If 1, Git will only read version 1 Bloom filters, and will write version 1\n Bloom filters.\n +\n+If 2, Git will only read version 2 Bloom filters, and will write version 2\n+Bloom filters.\n++\n See linkgit:git-commit-graph[1] for more information.\ndiff --git a/bloom.c b/bloom.c\nindex 401999ed3c..e9361b1c1f 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -99,7 +99,64 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n  * Not considered to be cryptographically secure.\n  * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  */\n-uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len)\n+uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n+{\n+\tconst uint32_t c1 = 0xcc9e2d51;\n+\tconst uint32_t c2 = 0x1b873593;\n+\tconst uint32_t r1 = 15;\n+\tconst uint32_t r2 = 13;\n+\tconst uint32_t m = 5;\n+\tconst uint32_t n = 0xe6546b64;\n+\tint i;\n+\tuint32_t k1 = 0;\n+\tconst char *tail;\n+\n+\tint len4 = len / sizeof(uint32_t);\n+\n+\tuint32_t k;\n+\tfor (i = 0; i < len4; i++) {\n+\t\tuint32_t byte1 = (uint32_t)(unsigned char)data[4*i];\n+\t\tuint32_t byte2 = ((uint32_t)(unsigned char)data[4*i + 1]) << 8;\n+\t\tuint32_t byte3 = ((uint32_t)(unsigned char)data[4*i + 2]) << 16;\n+\t\tuint32_t byte4 = ((uint32_t)(unsigned char)data[4*i + 3]) << 24;\n+\t\tk = byte1 | byte2 | byte3 | byte4;\n+\t\tk *= c1;\n+\t\tk = rotate_left(k, r1);\n+\t\tk *= c2;\n+\n+\t\tseed ^= k;\n+\t\tseed = rotate_left(seed, r2) * m + n;\n+\t}\n+\n+\ttail = (data + len4 * sizeof(uint32_t));\n+\n+\tswitch (len & (sizeof(uint32_t) - 1)) {\n+\tcase 3:\n+\t\tk1 ^= ((uint32_t)(unsigned char)tail[2]) << 16;\n+\t\t/*-fallthrough*/\n+\tcase 2:\n+\t\tk1 ^= ((uint32_t)(unsigned char)tail[1]) << 8;\n+\t\t/*-fallthrough*/\n+\tcase 1:\n+\t\tk1 ^= ((uint32_t)(unsigned char)tail[0]) << 0;\n+\t\tk1 *= c1;\n+\t\tk1 = rotate_left(k1, r1);\n+\t\tk1 *= c2;\n+\t\tseed ^= k1;\n+\t\tbreak;\n+\t}\n+\n+\tseed ^= (uint32_t)len;\n+\tseed ^= (seed >> 16);\n+\tseed *= 0x85ebca6b;\n+\tseed ^= (seed >> 13);\n+\tseed *= 0xc2b2ae35;\n+\tseed ^= (seed >> 16);\n+\n+\treturn seed;\n+}\n+\n+static uint32_t murmur3_seeded_v1(uint32_t seed, const char *data, size_t len)\n {\n \tconst uint32_t c1 = 0xcc9e2d51;\n \tconst uint32_t c2 = 0x1b873593;\n@@ -164,8 +221,14 @@ void fill_bloom_key(const char *data,\n \tint i;\n \tconst uint32_t seed0 = 0x293ae76f;\n \tconst uint32_t seed1 = 0x7e646e2c;\n-\tconst uint32_t hash0 = murmur3_seeded(seed0, data, len);\n-\tconst uint32_t hash1 = murmur3_seeded(seed1, data, len);\n+\tuint32_t hash0, hash1;\n+\tif (settings->hash_version == 2) {\n+\t\thash0 = murmur3_seeded_v2(seed0, data, len);\n+\t\thash1 = murmur3_seeded_v2(seed1, data, len);\n+\t} else {\n+\t\thash0 = murmur3_seeded_v1(seed0, data, len);\n+\t\thash1 = murmur3_seeded_v1(seed1, data, len);\n+\t}\n \n \tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n \tfor (i = 0; i < settings->num_hashes; i++)\ndiff --git a/bloom.h b/bloom.h\nindex 1e4f612d2c..138d57a86b 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -8,9 +8,11 @@ struct commit_graph;\n struct bloom_filter_settings {\n \t/*\n \t * The version of the hashing technique being used.\n-\t * We currently only support version = 1 which is\n+\t * The newest version is 2, which is\n \t * the seeded murmur3 hashing technique implemented\n-\t * in bloom.c.\n+\t * in bloom.c. Bloom filters of version 1 were created\n+\t * with prior versions of Git, which had a bug in the\n+\t * implementation of the hash function.\n \t */\n \tuint32_t hash_version;\n \n@@ -80,7 +82,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n  * Not considered to be cryptographically secure.\n  * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  */\n-uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len);\n+uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len);\n \n void fill_bloom_key(const char *data,\n \t\t    size_t len,\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 91c98ebc6c..22237e7dfc 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -340,10 +340,16 @@ static int graph_read_bloom_index(const unsigned char *chunk_start,\n \treturn 0;\n }\n \n+struct graph_read_bloom_data_context {\n+\tstruct commit_graph *g;\n+\tint *commit_graph_changed_paths_version;\n+};\n+\n static int graph_read_bloom_data(const unsigned char *chunk_start,\n \t\t\t\t  size_t chunk_size, void *data)\n {\n-\tstruct commit_graph *g = data;\n+\tstruct graph_read_bloom_data_context *c = data;\n+\tstruct commit_graph *g = c->g;\n \tuint32_t hash_version;\n \n \tif (chunk_size < BLOOMDATA_CHUNK_HEADER_SIZE) {\n@@ -354,13 +360,15 @@ static int graph_read_bloom_data(const unsigned char *chunk_start,\n \t\treturn -1;\n \t}\n \n+\thash_version = get_be32(chunk_start);\n+\n+\tif (*c->commit_graph_changed_paths_version == -1)\n+\t\t*c->commit_graph_changed_paths_version = hash_version;\n+\telse if (hash_version != *c->commit_graph_changed_paths_version)\n+\t\treturn 0;\n+\n \tg->chunk_bloom_data = chunk_start;\n \tg->chunk_bloom_data_size = chunk_size;\n-\thash_version = get_be32(chunk_start);\n-\n-\tif (hash_version != 1)\n-\t\treturn 0;\n-\n \tg->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n \tg->bloom_filter_settings->hash_version = hash_version;\n \tg->bloom_filter_settings->num_hashes = get_be32(chunk_start + 4);\n@@ -460,10 +468,14 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s,\n \t}\n \n \tif (s->commit_graph_changed_paths_version) {\n+\t\tstruct graph_read_bloom_data_context context = {\n+\t\t\t.g = graph,\n+\t\t\t.commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version\n+\t\t};\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n \t\t\t   graph_read_bloom_index, graph);\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n-\t\t\t   graph_read_bloom_data, graph);\n+\t\t\t   graph_read_bloom_data, &context);\n \t}\n \n \tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n@@ -2501,6 +2513,13 @@ int write_commit_graph(struct object_directory *odb,\n \t}\n \tif (!commit_graph_compatible(r))\n \t\treturn 0;\n+\tif (r->settings.commit_graph_changed_paths_version < -1\n+\t    || r->settings.commit_graph_changed_paths_version > 2) {\n+\t\twarning(_(\"attempting to write a commit-graph, but \"\n+\t\t\t  \"'commitgraph.changedPathsVersion' (%d) is not supported\"),\n+\t\t\tr->settings.commit_graph_changed_paths_version);\n+\t\treturn 0;\n+\t}\n \n \tCALLOC_ARRAY(ctx, 1);\n \tctx->r = r;\n@@ -2513,6 +2532,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->write_generation_data = (get_configured_generation_version(r) == 2);\n \tctx->num_generation_data_overflows = 0;\n \n+\tbloom_settings.hash_version = r->settings.commit_graph_changed_paths_version == 2\n+\t\t? 2 : 1;\n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n \tbloom_settings.num_hashes = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_NUM_HASHES\",\n@@ -2542,7 +2563,7 @@ int write_commit_graph(struct object_directory *odb,\n \t\tg = ctx->r->objects->commit_graph;\n \n \t\t/* We have changed-paths already. Keep them in the next graph */\n-\t\tif (g && g->chunk_bloom_data) {\n+\t\tif (g && g->bloom_filter_settings) {\n \t\t\tctx->changed_paths = 1;\n \t\t\tctx->bloom_settings = g->bloom_filter_settings;\n \t\t}\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 1281e66876..eefc1668c7 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -49,6 +49,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n \n static const char *bloom_usage = \"\\n\"\n \"  test-tool bloom get_murmur3 <string>\\n\"\n+\"  test-tool bloom get_murmur3_seven_highbit\\n\"\n \"  test-tool bloom generate_filter <string> [<string>...]\\n\"\n \"  test-tool bloom get_filter_for_commit <commit-hex>\\n\";\n \n@@ -63,7 +64,13 @@ int cmd__bloom(int argc, const char **argv)\n \t\tuint32_t hashed;\n \t\tif (argc < 3)\n \t\t\tusage(bloom_usage);\n-\t\thashed = murmur3_seeded(0, argv[2], strlen(argv[2]));\n+\t\thashed = murmur3_seeded_v2(0, argv[2], strlen(argv[2]));\n+\t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n+\t}\n+\n+\tif (!strcmp(argv[1], \"get_murmur3_seven_highbit\")) {\n+\t\tuint32_t hashed;\n+\t\thashed = murmur3_seeded_v2(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nindex b567383eb8..c8d84ab606 100755\n--- a/t/t0095-bloom.sh\n+++ b/t/t0095-bloom.sh\n@@ -29,6 +29,14 @@ test_expect_success 'compute unseeded murmur3 hash for test string 2' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'compute unseeded murmur3 hash for test string 3' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0xa183ccfd\n+\tEOF\n+\ttest-tool bloom get_murmur3_seven_highbit >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success 'compute bloom key for empty string' '\n \tcat >expect <<-\\EOF &&\n \tHashes:0x5615800c|0x5b966560|0x61174ab4|0x66983008|0x6c19155c|0x7199fab0|0x771ae004|\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 642b960893..a7bf3a7dca 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -563,6 +563,120 @@ test_expect_success 'version 1 changed-path used when version 1 requested' '\n \t)\n '\n \n+test_expect_success 'version 1 changed-path not used when version 2 requested' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\tgit config --add commitgraph.changedPathsVersion 2 &&\n+\t\ttest_bloom_filters_not_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'version 1 changed-path used when autodetect requested' '\n+\t(\n+\t\tcd highbit1 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'when writing another commit graph, preserve existing version 1 of changed-path' '\n+\ttest_commit -C highbit1 c1double \"$CENT$CENT\" &&\n+\tgit -C highbit1 commit-graph write --reachable --changed-paths &&\n+\t(\n+\t\tcd highbit1 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\techo \"options: bloom(1,10,7) read_generation_data\" >expect &&\n+\t\ttest-tool read-graph >full &&\n+\t\tgrep options full >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'set up repo with high bit path, version 2 changed-path' '\n+\tgit init highbit2 &&\n+\tgit -C highbit2 config --add commitgraph.changedPathsVersion 2 &&\n+\ttest_commit -C highbit2 c2 \"$CENT\" &&\n+\tgit -C highbit2 commit-graph write --reachable --changed-paths\n+'\n+\n+test_expect_success 'check value of version 2 changed-path' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\techo \"c01f\" >expect &&\n+\t\tget_first_changed_path_filter >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'setup make another commit' '\n+\t# \"git log\" does not use Bloom filters for root commits - see how, in\n+\t# revision.c, rev_compare_tree() (the only code path that eventually calls\n+\t# get_bloom_filter()) is only called by try_to_simplify_commit() when the commit\n+\t# has one parent. Therefore, make another commit so that we perform the tests on\n+\t# a non-root commit.\n+\ttest_commit -C highbit2 anotherc2 \"another$CENT\"\n+'\n+\n+test_expect_success 'version 2 changed-path used when version 2 requested' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'version 2 changed-path not used when version 1 requested' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\tgit config --add commitgraph.changedPathsVersion 1 &&\n+\t\ttest_bloom_filters_not_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'version 2 changed-path used when autodetect requested' '\n+\t(\n+\t\tcd highbit2 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\ttest_bloom_filters_used \"-- another$CENT\"\n+\t)\n+'\n+\n+test_expect_success 'when writing another commit graph, preserve existing version 2 of changed-path' '\n+\ttest_commit -C highbit2 c2double \"$CENT$CENT\" &&\n+\tgit -C highbit2 commit-graph write --reachable --changed-paths &&\n+\t(\n+\t\tcd highbit2 &&\n+\t\tgit config --add commitgraph.changedPathsVersion -1 &&\n+\t\techo \"options: bloom(2,10,7) read_generation_data\" >expect &&\n+\t\ttest-tool read-graph >full &&\n+\t\tgrep options full >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'when writing commit graph, do not reuse changed-path of another version' '\n+\tgit init doublewrite &&\n+\ttest_commit -C doublewrite c \"$CENT\" &&\n+\tgit -C doublewrite config --add commitgraph.changedPathsVersion 1 &&\n+\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\tfor v in -2 3\n+\tdo\n+\t\tgit -C doublewrite config --add commitgraph.changedPathsVersion $v &&\n+\t\tgit -C doublewrite commit-graph write --reachable --changed-paths 2>err &&\n+\t\tcat >expect <<-EOF &&\n+\t\twarning: attempting to write a commit-graph, but ${SQ}commitgraph.changedPathsVersion${SQ} ($v) is not supported\n+\t\tEOF\n+\t\ttest_cmp expect err || return 1\n+\tdone &&\n+\tgit -C doublewrite config --add commitgraph.changedPathsVersion 2 &&\n+\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\t(\n+\t\tcd doublewrite &&\n+\t\techo \"c01f\" >expect &&\n+\t\tget_first_changed_path_filter >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n corrupt_graph () {\n \ttest_when_finished \"rm -rf $graph\" &&\n \tgit commit-graph write --reachable --changed-paths &&\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486883","messageId":"b56e94cad7379e229b4a915d378ca2a036864c73.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 11/17] bloom: annotate filters with hash version","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:34Z","receivedAt":"2024-01-16T22:09:36Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In subsequent commits, we will want to load existing Bloom filters out\nof a commit-graph, even when the hash version they were computed with\ndoes not match the value of `commitGraph.changedPathVersion`.\n\nIn order to differentiate between the two, add a \"version\" field to each\nBloom filter.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c | 11 ++++++++---\n bloom.h |  1 +\n 2 files changed, 9 insertions(+), 3 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex e9361b1c1f..9284b88e95 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -88,6 +88,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n \t\t\t\t\tsizeof(unsigned char) * start_index +\n \t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n+\tfilter->version = g->bloom_filter_settings->hash_version;\n \n \treturn 1;\n }\n@@ -273,11 +274,13 @@ static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \treturn strcmp(e1->path, e2->path);\n }\n \n-static void init_truncated_large_filter(struct bloom_filter *filter)\n+static void init_truncated_large_filter(struct bloom_filter *filter,\n+\t\t\t\t\tint version)\n {\n \tfilter->data = xmalloc(1);\n \tfilter->data[0] = 0xFF;\n \tfilter->len = 1;\n+\tfilter->version = version;\n }\n \n struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n@@ -362,13 +365,15 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t}\n \n \t\tif (hashmap_get_size(&pathmap) > settings->max_changed_paths) {\n-\t\t\tinit_truncated_large_filter(filter);\n+\t\t\tinit_truncated_large_filter(filter,\n+\t\t\t\t\t\t    settings->hash_version);\n \t\t\tif (computed)\n \t\t\t\t*computed |= BLOOM_TRUNC_LARGE;\n \t\t\tgoto cleanup;\n \t\t}\n \n \t\tfilter->len = (hashmap_get_size(&pathmap) * settings->bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter->version = settings->hash_version;\n \t\tif (!filter->len) {\n \t\t\tif (computed)\n \t\t\t\t*computed |= BLOOM_TRUNC_EMPTY;\n@@ -388,7 +393,7 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t} else {\n \t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n \t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n-\t\tinit_truncated_large_filter(filter);\n+\t\tinit_truncated_large_filter(filter, settings->hash_version);\n \n \t\tif (computed)\n \t\t\t*computed |= BLOOM_TRUNC_LARGE;\ndiff --git a/bloom.h b/bloom.h\nindex 138d57a86b..330a140520 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -55,6 +55,7 @@ struct bloom_filter_settings {\n struct bloom_filter {\n \tunsigned char *data;\n \tsize_t len;\n+\tint version;\n };\n \n /*\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486884","messageId":"ddfd1ba32a0ea03eb8297f6700b24690eb63e684.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 12/17] bloom: prepare to discard incompatible Bloom filters","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:37Z","receivedAt":"2024-01-16T22:09:39Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Callers use the inline `get_bloom_filter()` implementation as a thin\nwrapper around `get_or_compute_bloom_filter()`. The former calls the\nlatter with a value of \"0\" for `compute_if_not_present`, making\n`get_bloom_filter()` the default read-only path for fetching an existing\nBloom filter.\n\nCallers expect the value returned from `get_bloom_filter()` is usable,\nthat is that it's compatible with the configured value corresponding to\n`commitGraph.changedPathsVersion`.\n\nThis is OK, since the commit-graph machinery only initializes its BDAT\nchunk (thereby enabling it to service Bloom filter queries) when the\nBloom filter hash_version is compatible with our settings. So any value\nreturned by `get_bloom_filter()` is trivially useable.\n\nHowever, subsequent commits will load the BDAT chunk even when the Bloom\nfilters are built with incompatible hash versions. Prepare to handle\nthis by teaching `get_bloom_filter()` to discard filters that are\nincompatible with the configured hash version.\n\nCallers who wish to read incompatible filters (e.g., for upgrading\nfilters from v1 to v2) may use the lower level routine,\n`get_or_compute_bloom_filter()`.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c | 20 +++++++++++++++++++-\n bloom.h | 20 ++++++++++++++++++--\n 2 files changed, 37 insertions(+), 3 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 9284b88e95..323d8012b8 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -283,6 +283,23 @@ static void init_truncated_large_filter(struct bloom_filter *filter,\n \tfilter->version = version;\n }\n \n+struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c)\n+{\n+\tstruct bloom_filter *filter;\n+\tint hash_version;\n+\n+\tfilter = get_or_compute_bloom_filter(r, c, 0, NULL, NULL);\n+\tif (!filter)\n+\t\treturn NULL;\n+\n+\tprepare_repo_settings(r);\n+\thash_version = r->settings.commit_graph_changed_paths_version;\n+\n+\tif (!(hash_version == -1 || hash_version == filter->version))\n+\t\treturn NULL; /* unusable filter */\n+\treturn filter;\n+}\n+\n struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t struct commit *c,\n \t\t\t\t\t\t int compute_if_not_present,\n@@ -308,7 +325,8 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t     filter, graph_pos);\n \t}\n \n-\tif (filter->data && filter->len)\n+\tif ((filter->data && filter->len) &&\n+\t    (!settings || settings->hash_version == filter->version))\n \t\treturn filter;\n \tif (!compute_if_not_present)\n \t\treturn NULL;\ndiff --git a/bloom.h b/bloom.h\nindex 330a140520..bfe389e29c 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -110,8 +110,24 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t const struct bloom_filter_settings *settings,\n \t\t\t\t\t\t enum bloom_filter_computed *computed);\n \n-#define get_bloom_filter(r, c) get_or_compute_bloom_filter( \\\n-\t(r), (c), 0, NULL, NULL)\n+/*\n+ * Find the Bloom filter associated with the given commit \"c\".\n+ *\n+ * If any of the following are true\n+ *\n+ *   - the repository does not have a commit-graph, or\n+ *   - the repository disables reading from the commit-graph, or\n+ *   - the given commit does not have a Bloom filter computed, or\n+ *   - there is a Bloom filter for commit \"c\", but it cannot be read\n+ *     because the filter uses an incompatible version of murmur3\n+ *\n+ * , then `get_bloom_filter()` will return NULL. Otherwise, the corresponding\n+ * Bloom filter will be returned.\n+ *\n+ * For callers who wish to inspect Bloom filters with incompatible hash\n+ * versions, use get_or_compute_bloom_filter().\n+ */\n+struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c);\n \n int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486885","messageId":"72aabd289b9e455b5fa0331fe27f73d4e6792794.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 13/17] commit-graph.c: unconditionally load Bloom filters","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:40Z","receivedAt":"2024-01-16T22:09:42Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In an earlier commit, we began ignoring the Bloom data (\"BDAT\") chunk\nfor commit-graphs whose Bloom filters were computed using a hash version\n  incompatible with the value of `commitGraph.changedPathVersion`.\n\nNow that the Bloom API has been hardened to discard these incompatible\nfilters (with the exception of low-level APIs), we can safely load these\nBloom filters unconditionally.\n\nWe no longer want to return early from `graph_read_bloom_data()`, and\nsimilarly do not want to set the bloom_settings' `hash_version` field as\na side-effect. The latter is because we want to wait until we know which\nBloom settings we're using (either the defaults, from the GIT_TEST\nvariables, or from the previous commit-graph layer) before deciding what\nhash_version to use.\n\nIf we detect an existing BDAT chunk, we'll infer the rest of the\nsettings (e.g., number of hashes, bits per entry, and maximum number of\nchanged paths) from the earlier graph layer. The hash_version will be\ninferred from the previous layer as well, unless one has already been\nspecified via configuration.\n\nOnce all of that is done, we normalize the value of the hash_version to\neither \"1\" or \"2\".\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n commit-graph.c | 18 ++++++++++--------\n 1 file changed, 10 insertions(+), 8 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 22237e7dfc..a2063d5f91 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -362,11 +362,6 @@ static int graph_read_bloom_data(const unsigned char *chunk_start,\n \n \thash_version = get_be32(chunk_start);\n \n-\tif (*c->commit_graph_changed_paths_version == -1)\n-\t\t*c->commit_graph_changed_paths_version = hash_version;\n-\telse if (hash_version != *c->commit_graph_changed_paths_version)\n-\t\treturn 0;\n-\n \tg->chunk_bloom_data = chunk_start;\n \tg->chunk_bloom_data_size = chunk_size;\n \tg->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n@@ -2532,8 +2527,7 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->write_generation_data = (get_configured_generation_version(r) == 2);\n \tctx->num_generation_data_overflows = 0;\n \n-\tbloom_settings.hash_version = r->settings.commit_graph_changed_paths_version == 2\n-\t\t? 2 : 1;\n+\tbloom_settings.hash_version = r->settings.commit_graph_changed_paths_version;\n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n \tbloom_settings.num_hashes = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_NUM_HASHES\",\n@@ -2565,10 +2559,18 @@ int write_commit_graph(struct object_directory *odb,\n \t\t/* We have changed-paths already. Keep them in the next graph */\n \t\tif (g && g->bloom_filter_settings) {\n \t\t\tctx->changed_paths = 1;\n-\t\t\tctx->bloom_settings = g->bloom_filter_settings;\n+\n+\t\t\t/* don't propagate the hash_version unless unspecified */\n+\t\t\tif (bloom_settings.hash_version == -1)\n+\t\t\t\tbloom_settings.hash_version = g->bloom_filter_settings->hash_version;\n+\t\t\tbloom_settings.bits_per_entry = g->bloom_filter_settings->bits_per_entry;\n+\t\t\tbloom_settings.num_hashes = g->bloom_filter_settings->num_hashes;\n+\t\t\tbloom_settings.max_changed_paths = g->bloom_filter_settings->max_changed_paths;\n \t\t}\n \t}\n \n+\tbloom_settings.hash_version = bloom_settings.hash_version == 2 ? 2 : 1;\n+\n \tif (ctx->split) {\n \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n \n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486886","messageId":"526beb9766a25dc97b9a4913fc02701908c0612e.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 14/17] commit-graph: drop unnecessary `graph_read_bloom_data_context`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:44Z","receivedAt":"2024-01-16T22:09:46Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"The `graph_read_bloom_data_context` struct was introduced in an earlier\ncommit in order to pass pointers to the commit-graph and changed-path\nBloom filter version when reading the BDAT chunk.\n\nThe previous commit no longer writes through the changed_paths_version\npointer, making the surrounding context structure unnecessary. Drop it\nand pass a pointer to the commit-graph directly when reading the BDAT\nchunk.\n\nNoticed-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n commit-graph.c | 14 ++------------\n 1 file changed, 2 insertions(+), 12 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a2063d5f91..a02556716d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -340,16 +340,10 @@ static int graph_read_bloom_index(const unsigned char *chunk_start,\n \treturn 0;\n }\n \n-struct graph_read_bloom_data_context {\n-\tstruct commit_graph *g;\n-\tint *commit_graph_changed_paths_version;\n-};\n-\n static int graph_read_bloom_data(const unsigned char *chunk_start,\n \t\t\t\t  size_t chunk_size, void *data)\n {\n-\tstruct graph_read_bloom_data_context *c = data;\n-\tstruct commit_graph *g = c->g;\n+\tstruct commit_graph *g = data;\n \tuint32_t hash_version;\n \n \tif (chunk_size < BLOOMDATA_CHUNK_HEADER_SIZE) {\n@@ -463,14 +457,10 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s,\n \t}\n \n \tif (s->commit_graph_changed_paths_version) {\n-\t\tstruct graph_read_bloom_data_context context = {\n-\t\t\t.g = graph,\n-\t\t\t.commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version\n-\t\t};\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n \t\t\t   graph_read_bloom_index, graph);\n \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n-\t\t\t   graph_read_bloom_data, &context);\n+\t\t\t   graph_read_bloom_data, graph);\n \t}\n \n \tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486887","messageId":"c683697efacc1c8f53951bf28895d83ce436b5d8.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 15/17] object.h: fix mis-aligned flag bits table","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:47Z","receivedAt":"2024-01-16T22:09:49Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Bit position 23 is one column too far to the left.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/object.h b/object.h\nindex 114d45954d..db25714b4e 100644\n--- a/object.h\n+++ b/object.h\n@@ -62,7 +62,7 @@ void object_array_init(struct object_array *array);\n \n /*\n  * object flag allocation:\n- * revision.h:               0---------10         15             23------27\n+ * revision.h:               0---------10         15               23------27\n  * fetch-pack.c:             01    67\n  * negotiator/default.c:       2--5\n  * walker.c:                 0-2\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486888","messageId":"4bf043be9aff322279854ce96afe697f362ac60b.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 16/17] commit-graph: reuse existing Bloom filters where possible","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:50Z","receivedAt":"2024-01-16T22:09:52Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In an earlier commit, a bug was described where it's possible for Git to\nproduce non-murmur3 hashes when the platform's \"char\" type is signed,\nand there are paths with characters whose highest bit is set (i.e. all\ncharacters >= 0x80).\n\nThat patch allows the caller to control which version of Bloom filters\nare read and written. However, even on platforms with a signed \"char\"\ntype, it is possible to reuse existing Bloom filters if and only if\nthere are no changed paths in any commit's first parent tree-diff whose\ncharacters have their highest bit set.\n\nWhen this is the case, we can reuse the existing filter without having\nto compute a new one. This is done by marking trees which are known to\nhave (or not have) any such paths. When a commit's root tree is verified\nto not have any such paths, we mark it as such and declare that the\ncommit's Bloom filter is reusable.\n\nNote that this heuristic only goes in one direction. If neither a commit\nnor its first parent have any paths in their trees with non-ASCII\ncharacters, then we know for certain that a path with non-ASCII\ncharacters will not appear in a tree-diff against that commit's first\nparent. The reverse isn't necessarily true: just because the tree-diff\ndoesn't contain any such paths does not imply that no such paths exist\nin either tree.\n\nSo we end up recomputing some Bloom filters that we don't strictly have\nto (i.e. their bits are the same no matter which version of murmur3 we\nuse). But culling these out is impossible, since we'd have to perform\nthe full tree-diff, which is the same effort as computing the Bloom\nfilter from scratch.\n\nBut because we can cache our results in each tree's flag bits, we can\noften avoid recomputing many filters, thereby reducing the time it takes\nto run\n\n    $ git commit-graph write --changed-paths --reachable\n\nwhen upgrading from v1 to v2 Bloom filters.\n\nTo benchmark this, let's generate a commit-graph in linux.git with v1\nchanged-paths in generation order[^1]:\n\n    $ git clone git@github.com:torvalds/linux.git\n    $ cd linux\n    $ git commit-graph write --reachable --changed-paths\n    $ graph=\".git/objects/info/commit-graph\"\n    $ mv $graph{,.bak}\n\nThen let's time how long it takes to go from v1 to v2 filters (with and\nwithout the upgrade path enabled), resetting the state of the\ncommit-graph each time:\n\n    $ git config commitGraph.changedPathsVersion 2\n    $ hyperfine -p 'cp -f $graph.bak $graph' -L v 0,1 \\\n        'GIT_TEST_UPGRADE_BLOOM_FILTERS={v} git.compile commit-graph write --reachable --changed-paths'\n\nOn linux.git (where there aren't any non-ASCII paths), the timings\nindicate that this patch represents a speed-up over recomputing all\nBloom filters from scratch:\n\n    Benchmark 1: GIT_TEST_UPGRADE_BLOOM_FILTERS=0 git.compile commit-graph write --reachable --changed-paths\n      Time (mean ± σ):     124.873 s ±  0.316 s    [User: 124.081 s, System: 0.643 s]\n      Range (min … max):   124.621 s … 125.227 s    3 runs\n\n    Benchmark 2: GIT_TEST_UPGRADE_BLOOM_FILTERS=1 git.compile commit-graph write --reachable --changed-paths\n      Time (mean ± σ):     79.271 s ±  0.163 s    [User: 74.611 s, System: 4.521 s]\n      Range (min … max):   79.112 s … 79.437 s    3 runs\n\n    Summary\n      'GIT_TEST_UPGRADE_BLOOM_FILTERS=1 git.compile commit-graph write --reachable --changed-paths' ran\n        1.58 ± 0.01 times faster than 'GIT_TEST_UPGRADE_BLOOM_FILTERS=0 git.compile commit-graph write --reachable --changed-paths'\n\nOn git.git, we do have some non-ASCII paths, giving us a more modest\nimprovement from 4.163 seconds to 3.348 seconds, for a 1.24x speed-up.\nOn my machine, the stats for git.git are:\n\n  - 8,285 Bloom filters computed from scratch\n  - 10 Bloom filters generated as empty\n  - 4 Bloom filters generated as truncated due to too many changed paths\n  - 65,114 Bloom filters were reused when transitioning from v1 to v2.\n\n[^1]: Note that this is is important, since `--stdin-packs` or\n  `--stdin-commits` orders commits in the commit-graph by their pack\n  position (with `--stdin-packs`) or in the raw input (with\n  `--stdin-commits`).\n\n  Since we compute Bloom filters in the same order that commits appear\n  in the graph, we must see a commit's (first) parent before we process\n  the commit itself. This is only guaranteed to happen when sorting\n  commits by their generation number.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c              | 90 ++++++++++++++++++++++++++++++++++++++++++--\n bloom.h              |  1 +\n commit-graph.c       |  5 +++\n object.h             |  1 +\n t/t4216-log-bloom.sh | 35 ++++++++++++++++-\n 5 files changed, 128 insertions(+), 4 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 323d8012b8..a1c616bc71 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -6,6 +6,9 @@\n #include \"commit-graph.h\"\n #include \"commit.h\"\n #include \"commit-slab.h\"\n+#include \"tree.h\"\n+#include \"tree-walk.h\"\n+#include \"config.h\"\n \n define_commit_slab(bloom_filter_slab, struct bloom_filter);\n \n@@ -283,6 +286,73 @@ static void init_truncated_large_filter(struct bloom_filter *filter,\n \tfilter->version = version;\n }\n \n+#define VISITED   (1u<<21)\n+#define HIGH_BITS (1u<<22)\n+\n+static int has_entries_with_high_bit(struct repository *r, struct tree *t)\n+{\n+\tif (parse_tree(t))\n+\t\treturn 1;\n+\n+\tif (!(t->object.flags & VISITED)) {\n+\t\tstruct tree_desc desc;\n+\t\tstruct name_entry entry;\n+\n+\t\tinit_tree_desc(&desc, t->buffer, t->size);\n+\t\twhile (tree_entry(&desc, &entry)) {\n+\t\t\tsize_t i;\n+\t\t\tfor (i = 0; i < entry.pathlen; i++) {\n+\t\t\t\tif (entry.path[i] & 0x80) {\n+\t\t\t\t\tt->object.flags |= HIGH_BITS;\n+\t\t\t\t\tgoto done;\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t\tif (S_ISDIR(entry.mode)) {\n+\t\t\t\tstruct tree *sub = lookup_tree(r, &entry.oid);\n+\t\t\t\tif (sub && has_entries_with_high_bit(r, sub)) {\n+\t\t\t\t\tt->object.flags |= HIGH_BITS;\n+\t\t\t\t\tgoto done;\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t}\n+\n+done:\n+\t\tt->object.flags |= VISITED;\n+\t}\n+\n+\treturn !!(t->object.flags & HIGH_BITS);\n+}\n+\n+static int commit_tree_has_high_bit_paths(struct repository *r,\n+\t\t\t\t\t  struct commit *c)\n+{\n+\tstruct tree *t;\n+\tif (repo_parse_commit(r, c))\n+\t\treturn 1;\n+\tt = repo_get_commit_tree(r, c);\n+\tif (!t)\n+\t\treturn 1;\n+\treturn has_entries_with_high_bit(r, t);\n+}\n+\n+static struct bloom_filter *upgrade_filter(struct repository *r, struct commit *c,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   int hash_version)\n+{\n+\tstruct commit_list *p = c->parents;\n+\tif (commit_tree_has_high_bit_paths(r, c))\n+\t\treturn NULL;\n+\n+\tif (p && commit_tree_has_high_bit_paths(r, p->item))\n+\t\treturn NULL;\n+\n+\tfilter->version = hash_version;\n+\n+\treturn filter;\n+}\n+\n struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c)\n {\n \tstruct bloom_filter *filter;\n@@ -325,9 +395,23 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\t\t\t\t     filter, graph_pos);\n \t}\n \n-\tif ((filter->data && filter->len) &&\n-\t    (!settings || settings->hash_version == filter->version))\n-\t\treturn filter;\n+\tif (filter->data && filter->len) {\n+\t\tstruct bloom_filter *upgrade;\n+\t\tif (!settings || settings->hash_version == filter->version)\n+\t\t\treturn filter;\n+\n+\t\t/* version mismatch, see if we can upgrade */\n+\t\tif (compute_if_not_present &&\n+\t\t    git_env_bool(\"GIT_TEST_UPGRADE_BLOOM_FILTERS\", 1)) {\n+\t\t\tupgrade = upgrade_filter(r, c, filter,\n+\t\t\t\t\t\t settings->hash_version);\n+\t\t\tif (upgrade) {\n+\t\t\t\tif (computed)\n+\t\t\t\t\t*computed |= BLOOM_UPGRADED;\n+\t\t\t\treturn upgrade;\n+\t\t\t}\n+\t\t}\n+\t}\n \tif (!compute_if_not_present)\n \t\treturn NULL;\n \ndiff --git a/bloom.h b/bloom.h\nindex bfe389e29c..e3a9b68905 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -102,6 +102,7 @@ enum bloom_filter_computed {\n \tBLOOM_COMPUTED     = (1 << 1),\n \tBLOOM_TRUNC_LARGE  = (1 << 2),\n \tBLOOM_TRUNC_EMPTY  = (1 << 3),\n+\tBLOOM_UPGRADED     = (1 << 4),\n };\n \n struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a02556716d..b285e32043 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1167,6 +1167,7 @@ struct write_commit_graph_context {\n \tint count_bloom_filter_not_computed;\n \tint count_bloom_filter_trunc_empty;\n \tint count_bloom_filter_trunc_large;\n+\tint count_bloom_filter_upgraded;\n };\n \n static int write_graph_chunk_fanout(struct hashfile *f,\n@@ -1774,6 +1775,8 @@ static void trace2_bloom_filter_write_statistics(struct write_commit_graph_conte\n \t\t\t   ctx->count_bloom_filter_trunc_empty);\n \ttrace2_data_intmax(\"commit-graph\", ctx->r, \"filter-trunc-large\",\n \t\t\t   ctx->count_bloom_filter_trunc_large);\n+\ttrace2_data_intmax(\"commit-graph\", ctx->r, \"filter-upgraded\",\n+\t\t\t   ctx->count_bloom_filter_upgraded);\n }\n \n static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n@@ -1815,6 +1818,8 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \t\t\t\tctx->count_bloom_filter_trunc_empty++;\n \t\t\tif (computed & BLOOM_TRUNC_LARGE)\n \t\t\t\tctx->count_bloom_filter_trunc_large++;\n+\t\t} else if (computed & BLOOM_UPGRADED) {\n+\t\t\tctx->count_bloom_filter_upgraded++;\n \t\t} else if (computed & BLOOM_NOT_COMPUTED)\n \t\t\tctx->count_bloom_filter_not_computed++;\n \t\tctx->total_bloom_filter_data_size += filter\ndiff --git a/object.h b/object.h\nindex db25714b4e..2e5e08725f 100644\n--- a/object.h\n+++ b/object.h\n@@ -75,6 +75,7 @@ void object_array_init(struct object_array *array);\n  * commit-reach.c:                                  16-----19\n  * sha1-name.c:                                              20\n  * list-objects-filter.c:                                      21\n+ * bloom.c:                                                    2122\n  * builtin/fsck.c:           0--3\n  * builtin/gc.c:             0\n  * builtin/index-pack.c:                                     2021\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex a7bf3a7dca..823d1cf773 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -222,6 +222,10 @@ test_filter_trunc_large () {\n \tgrep \"\\\"key\\\":\\\"filter-trunc-large\\\",\\\"value\\\":\\\"$1\\\"\" $2\n }\n \n+test_filter_upgraded () {\n+\tgrep \"\\\"key\\\":\\\"filter-upgraded\\\",\\\"value\\\":\\\"$1\\\"\" $2\n+}\n+\n test_expect_success 'correctly report changes over limit' '\n \tgit init limits &&\n \t(\n@@ -656,7 +660,13 @@ test_expect_success 'when writing another commit graph, preserve existing versio\n test_expect_success 'when writing commit graph, do not reuse changed-path of another version' '\n \tgit init doublewrite &&\n \ttest_commit -C doublewrite c \"$CENT\" &&\n+\n \tgit -C doublewrite config --add commitgraph.changedPathsVersion 1 &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 1 trace2.txt &&\n+\ttest_filter_upgraded 0 trace2.txt &&\n+\n \tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n \tfor v in -2 3\n \tdo\n@@ -667,8 +677,13 @@ test_expect_success 'when writing commit graph, do not reuse changed-path of ano\n \t\tEOF\n \t\ttest_cmp expect err || return 1\n \tdone &&\n+\n \tgit -C doublewrite config --add commitgraph.changedPathsVersion 2 &&\n-\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C doublewrite commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 1 trace2.txt &&\n+\ttest_filter_upgraded 0 trace2.txt &&\n+\n \t(\n \t\tcd doublewrite &&\n \t\techo \"c01f\" >expect &&\n@@ -677,6 +692,24 @@ test_expect_success 'when writing commit graph, do not reuse changed-path of ano\n \t)\n '\n \n+test_expect_success 'when writing commit graph, reuse changed-path of another version where possible' '\n+\tgit init upgrade &&\n+\n+\ttest_commit -C upgrade base no-high-bits &&\n+\n+\tgit -C upgrade config --add commitgraph.changedPathsVersion 1 &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C upgrade commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 1 trace2.txt &&\n+\ttest_filter_upgraded 0 trace2.txt &&\n+\n+\tgit -C upgrade config --add commitgraph.changedPathsVersion 2 &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C upgrade commit-graph write --reachable --changed-paths &&\n+\ttest_filter_computed 0 trace2.txt &&\n+\ttest_filter_upgraded 1 trace2.txt\n+'\n+\n corrupt_graph () {\n \ttest_when_finished \"rm -rf $graph\" &&\n \tgit commit-graph write --reachable --changed-paths &&\n-- \n2.43.0.334.gd4dbce1db5.dirty\n\n"},{"id":"486889","messageId":"7daa0d8833ebe9aba0de90f43f279dd6d87d634f.1705442923.git.me@ttaylorr.com","threadId":"60318","inReplyTo":"cover.1705442923.git.me@ttaylorr.com","subject":"[PATCH v5 17/17] bloom: introduce `deinit_bloom_filters()`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-16T22:09:53Z","receivedAt":"2024-01-16T22:09:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"After we are done using Bloom filters, we do not currently clean up any\nmemory allocated by the commit slab used to store those filters in the\nfirst place.\n\nBesides the bloom_filter structures themselves, there is mostly nothing\nto free() in the first place, since in the read-only path all Bloom\nfilter's `data` members point to a memory mapped region in the\ncommit-graph file itself.\n\nBut when generating Bloom filters from scratch (or initializing\ntruncated filters) we allocate additional memory to store the filter's\ndata.\n\nKeep track of when we need to free() this additional chunk of memory by\nusing an extra pointer `to_free`. Most of the time this will be NULL\n(indicating that we are representing an existing Bloom filter stored in\na memory mapped region). When it is non-NULL, free it before discarding\nthe Bloom filters slab.\n\nSuggested-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n bloom.c        | 16 +++++++++++++++-\n bloom.h        |  3 +++\n commit-graph.c |  4 ++++\n 3 files changed, 22 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex a1c616bc71..dbcc0f4f50 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -92,6 +92,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t\tsizeof(unsigned char) * start_index +\n \t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n \tfilter->version = g->bloom_filter_settings->hash_version;\n+\tfilter->to_free = NULL;\n \n \treturn 1;\n }\n@@ -264,6 +265,18 @@ void init_bloom_filters(void)\n \tinit_bloom_filter_slab(&bloom_filters);\n }\n \n+static void free_one_bloom_filter(struct bloom_filter *filter)\n+{\n+\tif (!filter)\n+\t\treturn;\n+\tfree(filter->to_free);\n+}\n+\n+void deinit_bloom_filters(void)\n+{\n+\tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -280,7 +293,7 @@ static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n static void init_truncated_large_filter(struct bloom_filter *filter,\n \t\t\t\t\tint version)\n {\n-\tfilter->data = xmalloc(1);\n+\tfilter->data = filter->to_free = xmalloc(1);\n \tfilter->data[0] = 0xFF;\n \tfilter->len = 1;\n \tfilter->version = version;\n@@ -482,6 +495,7 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \t\t\tfilter->len = 1;\n \t\t}\n \t\tCALLOC_ARRAY(filter->data, filter->len);\n+\t\tfilter->to_free = filter->data;\n \n \t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n \t\t\tstruct bloom_key key;\ndiff --git a/bloom.h b/bloom.h\nindex e3a9b68905..d20e64bfbb 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -56,6 +56,8 @@ struct bloom_filter {\n \tunsigned char *data;\n \tsize_t len;\n \tint version;\n+\n+\tvoid *to_free;\n };\n \n /*\n@@ -96,6 +98,7 @@ void add_key_to_filter(const struct bloom_key *key,\n \t\t       const struct bloom_filter_settings *settings);\n \n void init_bloom_filters(void);\n+void deinit_bloom_filters(void);\n \n enum bloom_filter_computed {\n \tBLOOM_NOT_COMPUTED = (1 << 0),\ndiff --git a/commit-graph.c b/commit-graph.c\nindex b285e32043..1841638801 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -830,6 +830,7 @@ struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n void close_commit_graph(struct raw_object_store *o)\n {\n \tclear_commit_graph_data_slab(&commit_graph_data_slab);\n+\tdeinit_bloom_filters();\n \tfree_commit_graph(o->commit_graph);\n \to->commit_graph = NULL;\n }\n@@ -2648,6 +2649,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \tres = write_commit_graph_file(ctx);\n \n+\tif (ctx->changed_paths)\n+\t\tdeinit_bloom_filters();\n+\n \tif (ctx->split)\n \t\tmark_commit_graphs(ctx);\n \n-- \n2.43.0.334.gd4dbce1db5.dirty\n"},{"id":"487555","messageId":"20240129212614.GB9612@szeder.dev","threadId":"60318","inReplyTo":"a77dcc99b4eb0a19dc6c09a40a84785413502126.1705442923.git.me@ttaylorr.com","subject":"Re: [PATCH v5 09/17] repo-settings: introduce commitgraph.changedPathsVersion","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2024-01-29T21:26:14Z","receivedAt":"2024-01-29T21:26:18Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Tue, Jan 16, 2024 at 05:09:28PM -0500, Taylor Blau wrote:\n> A subsequent commit will introduce another version of the changed-path\n> filter in the commit graph file. In order to control which version to\n> write (and read), a config variable is needed.\n> \n> Therefore, introduce this config variable. For forwards compatibility,\n> teach Git to not read commit graphs when the config variable\n> is set to an unsupported version. Because we teach Git this,\n> commitgraph.readChangedPaths is now redundant, so deprecate it and\n> define its behavior in terms of the config variable we introduce.\n> \n> This commit does not change the behavior of writing (Git writes changed\n> path filters when explicitly instructed regardless of any config\n> variable), but a subsequent commit will restrict Git such that it will\n> only write when commitgraph.changedPathsVersion is a recognized value.\n> \n> Signed-off-by: Jonathan Tan <jonathantanmy@google.com>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> Signed-off-by: Taylor Blau <me@ttaylorr.com>\n> ---\n>  Documentation/config/commitgraph.txt | 26 ++++++++++++++++++++++---\n>  commit-graph.c                       |  5 +++--\n>  oss-fuzz/fuzz-commit-graph.c         |  2 +-\n>  repo-settings.c                      |  6 +++++-\n>  repository.h                         |  2 +-\n>  t/t4216-log-bloom.sh                 | 29 +++++++++++++++++++++++++++-\n>  6 files changed, 61 insertions(+), 9 deletions(-)\n> \n> diff --git a/Documentation/config/commitgraph.txt b/Documentation/config/commitgraph.txt\n> index 30604e4a4c..e68cdededa 100644\n> --- a/Documentation/config/commitgraph.txt\n> +++ b/Documentation/config/commitgraph.txt\n> @@ -9,6 +9,26 @@ commitGraph.maxNewFilters::\n>  \tcommit-graph write` (c.f., linkgit:git-commit-graph[1]).\n>  \n>  commitGraph.readChangedPaths::\n> -\tIf true, then git will use the changed-path Bloom filters in the\n> -\tcommit-graph file (if it exists, and they are present). Defaults to\n> -\ttrue. See linkgit:git-commit-graph[1] for more information.\n> +\tDeprecated. Equivalent to commitGraph.changedPathsVersion=-1 if true, and\n> +\tcommitGraph.changedPathsVersion=0 if false. (If commitGraph.changedPathVersion\n> +\tis also set, commitGraph.changedPathsVersion takes precedence.)\n> +\n> +commitGraph.changedPathsVersion::\n> +\tSpecifies the version of the changed-path Bloom filters that Git will read and\n> +\twrite. May be -1, 0 or 1. Note that values greater than 1 may be\n> +\tincompatible with older versions of Git which do not yet understand\n> +\tthose versions. Use caution when operating in a mixed-version\n> +\tenvironment.\n> ++\n> +Defaults to -1.\n> ++\n> +If -1, Git will use the version of the changed-path Bloom filters in the\n> +repository, defaulting to 1 if there are none.\n> ++\n> +If 0, Git will not read any Bloom filters, and will write version 1 Bloom\n> +filters when instructed to write.\n> ++\n> +If 1, Git will only read version 1 Bloom filters, and will write version 1\n> +Bloom filters.\n> ++\n> +See linkgit:git-commit-graph[1] for more information.\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 00113b0f62..91c98ebc6c 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -459,7 +459,7 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s,\n>  \t\t\tgraph->read_generation_data = 1;\n>  \t}\n>  \n> -\tif (s->commit_graph_read_changed_paths) {\n> +\tif (s->commit_graph_changed_paths_version) {\n>  \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES,\n>  \t\t\t   graph_read_bloom_index, graph);\n>  \t\tread_chunk(cf, GRAPH_CHUNKID_BLOOMDATA,\n> @@ -555,7 +555,8 @@ static void validate_mixed_bloom_settings(struct commit_graph *g)\n>  \t\t}\n>  \n>  \t\tif (g->bloom_filter_settings->bits_per_entry != settings->bits_per_entry ||\n> -\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes) {\n> +\t\t    g->bloom_filter_settings->num_hashes != settings->num_hashes ||\n> +\t\t    g->bloom_filter_settings->hash_version != settings->hash_version) {\n>  \t\t\tg->chunk_bloom_indexes = NULL;\n>  \t\t\tg->chunk_bloom_data = NULL;\n>  \t\t\tFREE_AND_NULL(g->bloom_filter_settings);\n> diff --git a/oss-fuzz/fuzz-commit-graph.c b/oss-fuzz/fuzz-commit-graph.c\n> index 2992079dd9..325c0b991a 100644\n> --- a/oss-fuzz/fuzz-commit-graph.c\n> +++ b/oss-fuzz/fuzz-commit-graph.c\n> @@ -19,7 +19,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n>  \t * possible.\n>  \t */\n>  \tthe_repository->settings.commit_graph_generation_version = 2;\n> -\tthe_repository->settings.commit_graph_read_changed_paths = 1;\n> +\tthe_repository->settings.commit_graph_changed_paths_version = 1;\n>  \tg = parse_commit_graph(&the_repository->settings, (void *)data, size);\n>  \trepo_clear(the_repository);\n>  \tfree_commit_graph(g);\n> diff --git a/repo-settings.c b/repo-settings.c\n> index 30cd478762..c821583fe5 100644\n> --- a/repo-settings.c\n> +++ b/repo-settings.c\n> @@ -23,6 +23,7 @@ void prepare_repo_settings(struct repository *r)\n>  \tint value;\n>  \tconst char *strval;\n>  \tint manyfiles;\n> +\tint read_changed_paths;\n>  \n>  \tif (!r->gitdir)\n>  \t\tBUG(\"Cannot add settings for uninitialized repository\");\n> @@ -53,7 +54,10 @@ void prepare_repo_settings(struct repository *r)\n>  \t/* Commit graph config or default, does not cascade (simple) */\n>  \trepo_cfg_bool(r, \"core.commitgraph\", &r->settings.core_commit_graph, 1);\n>  \trepo_cfg_int(r, \"commitgraph.generationversion\", &r->settings.commit_graph_generation_version, 2);\n> -\trepo_cfg_bool(r, \"commitgraph.readchangedpaths\", &r->settings.commit_graph_read_changed_paths, 1);\n> +\trepo_cfg_bool(r, \"commitgraph.readchangedpaths\", &read_changed_paths, 1);\n> +\trepo_cfg_int(r, \"commitgraph.changedpathsversion\",\n> +\t\t     &r->settings.commit_graph_changed_paths_version,\n> +\t\t     read_changed_paths ? -1 : 0);\n>  \trepo_cfg_bool(r, \"gc.writecommitgraph\", &r->settings.gc_write_commit_graph, 1);\n>  \trepo_cfg_bool(r, \"fetch.writecommitgraph\", &r->settings.fetch_write_commit_graph, 0);\n>  \n> diff --git a/repository.h b/repository.h\n> index 5f18486f64..f71154e12c 100644\n> --- a/repository.h\n> +++ b/repository.h\n> @@ -29,7 +29,7 @@ struct repo_settings {\n>  \n>  \tint core_commit_graph;\n>  \tint commit_graph_generation_version;\n> -\tint commit_graph_read_changed_paths;\n> +\tint commit_graph_changed_paths_version;\n>  \tint gc_write_commit_graph;\n>  \tint fetch_write_commit_graph;\n>  \tint command_requires_full_index;\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> index 484dd093cd..642b960893 100755\n> --- a/t/t4216-log-bloom.sh\n> +++ b/t/t4216-log-bloom.sh\n> @@ -435,7 +435,7 @@ test_expect_success 'setup for mixed Bloom setting tests' '\n>  \tdone\n>  '\n>  \n> -test_expect_success 'ensure incompatible Bloom filters are ignored' '\n> +test_expect_success 'ensure Bloom filters with incompatible settings are ignored' '\n>  \t# Compute Bloom filters with \"unusual\" settings.\n>  \tgit -C $repo rev-parse one >in &&\n>  \tGIT_TEST_BLOOM_SETTINGS_NUM_HASHES=3 git -C $repo commit-graph write \\\n> @@ -485,6 +485,33 @@ test_expect_success 'merge graph layers with incompatible Bloom settings' '\n>  \ttest_must_be_empty err\n>  '\n>  \n> +test_expect_success 'ensure Bloom filter with incompatible versions are ignored' '\n> +\trm \"$repo/$graph\" &&\n> +\n> +\tgit -C $repo log --oneline --no-decorate -- $CENT >expect &&\n> +\n> +\t# Compute v1 Bloom filters for commits at the bottom.\n> +\tgit -C $repo rev-parse HEAD^ >in &&\n> +\tgit -C $repo commit-graph write --stdin-commits --changed-paths \\\n> +\t\t--split <in &&\n> +\n> +\t# Compute v2 Bloomfilters for the rest of the commits at the top.\n> +\tgit -C $repo rev-parse HEAD >in &&\n> +\tgit -C $repo -c commitGraph.changedPathsVersion=2 commit-graph write \\\n> +\t\t--stdin-commits --changed-paths --split=no-merge <in &&\n> +\n> +\ttest_line_count = 2 $repo/$chain &&\n> +\n> +\tgit -C $repo log --oneline --no-decorate -- $CENT >actual 2>err &&\n> +\ttest_cmp expect actual &&\n> +\n> +\tlayer=\"$(head -n 1 $repo/$chain)\" &&\n> +\tcat >expect.err <<-EOF &&\n> +\twarning: disabling Bloom filters for commit-graph layer $SQ$layer$SQ due to incompatible settings\n> +\tEOF\n> +\ttest_cmp expect.err err\n> +'\n\nAt this point in the series this test fails with:\n\n  + test_cmp expect.err err\n  + test 2 -ne 2\n  + eval diff -u \"$@\"\n  + diff -u expect.err err\n  --- expect.err  2024-01-29 21:02:57.927462620 +0000\n  +++ err 2024-01-29 21:02:57.923462642 +0000\n  @@ -1 +0,0 @@\n  -warning: disabling Bloom filters for commit-graph layer 'e338a7a1b4cfa5f6bcd31aea3e027df67d06442a' due to incompatible settings\n  error: last command exited with $?=1\n\n\n> +\n>  get_first_changed_path_filter () {\n>  \ttest-tool read-graph bloom-filters >filters.dat &&\n>  \thead -n 1 filters.dat\n> -- \n> 2.43.0.334.gd4dbce1db5.dirty\n> \n"},{"id":"487564","messageId":"Zbg7so2b4puSEWNK@nand.local","threadId":"60318","inReplyTo":"20240129212614.GB9612@szeder.dev","subject":"Re: [PATCH v5 09/17] repo-settings: introduce commitgraph.changedPathsVersion","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2024-01-29T23:58:42Z","receivedAt":"2024-01-29T23:58:44Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Jan 29, 2024 at 10:26:14PM +0100, SZEDER Gábor wrote:\n> At this point in the series this test fails with:\n>\n>   + test_cmp expect.err err\n>   + test 2 -ne 2\n>   + eval diff -u \"$@\"\n>   + diff -u expect.err err\n>   --- expect.err  2024-01-29 21:02:57.927462620 +0000\n>   +++ err 2024-01-29 21:02:57.923462642 +0000\n>   @@ -1 +0,0 @@\n>   -warning: disabling Bloom filters for commit-graph layer 'e338a7a1b4cfa5f6bcd31aea3e027df67d06442a' due to incompatible settings\n>   error: last command exited with $?=1\n\nVery good catch, thanks, I'm not sure how this one slipped through.\n\nThe fix should be mostly trivial, but I'll have to reroll the series\nsince it has some minor fallout outside of just this patch.\n\nJunio, please hold off on merging this to 'next' until I've had a chance\nto send out a new round.\n\nThanks,\nTaylor\n"}]}