{"thread":{"id":"55012","subject":"[PATCH 00/10] repack: support repacking into a geometric sequence","startedAt":"2021-01-19T23:24:52Z","lastAt":"2021-03-04T22:01:48Z","messageCount":120,"participants":["Taylor Blau","Derrick Stolee","Junio C Hamano","Jeff King","Eric Sunshine","Martin Fick"],"isPatch":true,"patchVersion":1,"patchTotal":10},"messages":[{"id":"414751","messageId":"cover.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":null,"subject":"[PATCH 00/10] repack: support repacking into a geometric sequence","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:23:55Z","receivedAt":"2021-01-19T23:24:52Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This series introduces a new mode of 'git repack' where (instead of packing just\nloose objects or packing everything together into one pack), the set of packs\nleft forms a geometric progression by object count.\n\nIt does not depend on either series of the revindex patches I sent recently.\n\nRoughly speaking, for a given factor, say \"d\", each pack has at least \"d\" times\nthe number of objects as the next largest pack. So, if there are \"N\" packs,\n\"P1\", \"P2\", ..., \"PN\" ordered by object count (where \"PN\" has the most objects,\nand \"P1\" the fewest), then:\n\n  objects(Pi) > d * objects(P(i-1))\n\nfor all 1 < i <= N.\n\nThis is done by first ordering packs by object count, and then determining the\nlongest sequence of large packs which already form a geometric progression. All\npacks on the small side of that cut must be repacked together, and so we check\nthat the existing progression can be maintained with the new pack, and adjust as\nnecessary.\n\nIn actuality, this is approximated in order for 'git repack' to have to create\nat most one new pack. The details of this approximation are discussed at length\nin the final patch.\n\n'git repack' implements this new option by marking the packs that don't need to\nbe touched as \"frozen\" and it does this by marking them as pack_keep_in_core,\nand then using a new option pack-objects option '--assume-kept-packs-closed' to\nstop the reachability traversal once it encounters any objects in the kept\npacks.\n\nWhen repacking in this mode, the caller implicitly trusts that the unchanged\npacks are closed under reachability, and thus they can halt the traversal as\nsoon as an object in any one of those packs is found.\n\nThe first three patches introduce the new revision and pack-objects options\nnecessary for this to work. The next four patches introduce an MRU cache for\nkept packs only. Then a new pack-objects mode is introduced to allow callers to\nspecify the list of kept packs over stdin in case they are too long to be listed\nas arguments. Finally, geometric repacking is introduced\n\nThanks in advance for your review.\n\nJeff King (4):\n  p5303: add missing &&-chains\n  p5303: measure time to repack with keep\n  pack-objects: rewrite honor-pack-keep logic\n  packfile: add kept-pack cache for find_kept_pack_entry()\n\nTaylor Blau (6):\n  packfile: introduce 'find_kept_pack_entry()'\n  revision: learn '--no-kept-objects'\n  builtin/pack-objects.c: learn '--assume-kept-packs-closed'\n  builtin/pack-objects.c: teach '--keep-pack-stdin'\n  builtin/repack.c: extract loose object handling\n  builtin/repack.c: add '--geometric' option\n\n Documentation/git-pack-objects.txt |  19 +++\n Documentation/git-repack.txt       |  11 ++\n Documentation/rev-list-options.txt |   7 +\n builtin/pack-objects.c             | 161 ++++++++++++++--------\n builtin/repack.c                   | 206 ++++++++++++++++++++++++++---\n list-objects.c                     |   7 +\n object-store.h                     |  10 ++\n packfile.c                         |  69 ++++++++++\n packfile.h                         |   2 +\n revision.c                         |  15 +++\n revision.h                         |   4 +\n t/perf/p5303-many-packs.sh         |  18 ++-\n t/t6114-keep-packs.sh              | 128 ++++++++++++++++++\n t/t7703-repack-geometric.sh        |  81 ++++++++++++\n 14 files changed, 663 insertions(+), 75 deletions(-)\n create mode 100755 t/t6114-keep-packs.sh\n create mode 100755 t/t7703-repack-geometric.sh\n\n-- \n2.30.0.138.g6d7191ea01\n"},{"id":"414752","messageId":"4184529648abe7451b5c7b772493df8c067cec82.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 02/10] revision: learn '--no-kept-objects'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:05Z","receivedAt":"2021-01-19T23:25:30Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Some callers want to perform a reachability traversal that terminates\nwhen an object is found in a kept pack. The closest existing option is\n'--honor-pack-keep', but this isn't quite what we want. Instead of\nhalting the traversal midway through, a full traversal is always\nperformed, and the results are only trimmed afterwords.\n\nBesides needing to introduce a new flag (since culling results\npost-facto can be different than halting the traversal as it's\nhappening), there is an additional wrinkle handling the distinction\nin-core and on-disk kept packs. That is: what kinds of kept pack should\nstop the traversal?\n\nIntroduce '--no-kept-objects[=<on-disk|in-core>]' to specify which kinds\nof kept packs, if any, should stop a traversal. This can be useful for\ncallers that want to perform a reachability analysis, but want to leave\ncertain packs alone (for e.g., when doing a geometric repack that has\nsome \"large\" packs it wants to leave alone).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/rev-list-options.txt |  7 +++\n list-objects.c                     |  7 +++\n revision.c                         | 15 +++++++\n revision.h                         |  4 ++\n t/t6114-keep-packs.sh              | 69 ++++++++++++++++++++++++++++++\n 5 files changed, 102 insertions(+)\n create mode 100755 t/t6114-keep-packs.sh\n\ndiff --git a/Documentation/rev-list-options.txt b/Documentation/rev-list-options.txt\nindex 002379056a..817419d552 100644\n--- a/Documentation/rev-list-options.txt\n+++ b/Documentation/rev-list-options.txt\n@@ -856,6 +856,13 @@ ifdef::git-rev-list[]\n \tOnly useful with `--objects`; print the object IDs that are not\n \tin packs.\n \n+--no-kept-objects[=<kind>]::\n+\tHalts the traversal as soon as an object in a kept pack is\n+\tfound. If `<kind>` is `on-disk`, only packs with a corresponding\n+\t`*.keep` file are ignored. If `<kind>` is `in-core`, only packs\n+\twith their in-core kept state set are ignored. Otherwise, both\n+\tkinds of kept packs are ignored.\n+\n --object-names::\n \tOnly useful with `--objects`; print the names of the object IDs\n \tthat are found. This is the default behavior.\ndiff --git a/list-objects.c b/list-objects.c\nindex e19589baa0..b06c3bfeba 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -338,6 +338,13 @@ static void traverse_trees_and_blobs(struct traversal_context *ctx,\n \t\t\tctx->show_object(obj, name, ctx->show_data);\n \t\t\tcontinue;\n \t\t}\n+\t\tif (ctx->revs->no_kept_objects) {\n+\t\t\tstruct pack_entry e;\n+\t\t\tif (find_kept_pack_entry(ctx->revs->repo, &obj->oid,\n+\t\t\t\t\t\t ctx->revs->keep_pack_cache_flags,\n+\t\t\t\t\t\t &e))\n+\t\t\t\tcontinue;\n+\t\t}\n \t\tif (!path)\n \t\t\tpath = \"\";\n \t\tif (obj->type == OBJ_TREE) {\ndiff --git a/revision.c b/revision.c\nindex 1bb590ece7..ff1ea77224 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -2334,6 +2334,16 @@ static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n \t\trevs->unpacked = 1;\n \t} else if (starts_with(arg, \"--unpacked=\")) {\n \t\tdie(_(\"--unpacked=<packfile> no longer supported\"));\n+\t} else if (!strcmp(arg, \"--no-kept-objects\")) {\n+\t\trevs->no_kept_objects = 1;\n+\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t} else if (skip_prefix(arg, \"--no-kept-objects=\", &optarg)) {\n+\t\trevs->no_kept_objects = 1;\n+\t\tif (!strcmp(optarg, \"in-core\"))\n+\t\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\tif (!strcmp(optarg, \"on-disk\"))\n+\t\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n \t} else if (!strcmp(arg, \"-r\")) {\n \t\trevs->diff = 1;\n \t\trevs->diffopt.flags.recursive = 1;\n@@ -3822,6 +3832,11 @@ enum commit_action get_commit_action(struct rev_info *revs, struct commit *commi\n \t\treturn commit_ignore;\n \tif (revs->unpacked && has_object_pack(&commit->object.oid))\n \t\treturn commit_ignore;\n+\tif (revs->no_kept_objects) {\n+\t\tif (has_object_kept_pack(&commit->object.oid,\n+\t\t\t\t\t revs->keep_pack_cache_flags))\n+\t\t\treturn commit_ignore;\n+\t}\n \tif (commit->object.flags & UNINTERESTING)\n \t\treturn commit_ignore;\n \tif (revs->line_level_traverse && !want_ancestry(revs)) {\ndiff --git a/revision.h b/revision.h\nindex 086ff10280..15d0e6aee5 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -148,6 +148,7 @@ struct rev_info {\n \t\t\tedge_hint_aggressive:1,\n \t\t\tlimited:1,\n \t\t\tunpacked:1,\n+\t\t\tno_kept_objects:1,\n \t\t\tboundary:2,\n \t\t\tcount:1,\n \t\t\tleft_right:1,\n@@ -312,6 +313,9 @@ struct rev_info {\n \t * This is loaded from the commit-graph being used.\n \t */\n \tstruct bloom_filter_settings *bloom_filter_settings;\n+\n+\t/* misc. flags related to '--no-kept-objects' */\n+\tunsigned keep_pack_cache_flags;\n };\n \n int ref_excluded(struct string_list *, const char *path);\ndiff --git a/t/t6114-keep-packs.sh b/t/t6114-keep-packs.sh\nnew file mode 100755\nindex 0000000000..9239d8aa46\n--- /dev/null\n+++ b/t/t6114-keep-packs.sh\n@@ -0,0 +1,69 @@\n+#!/bin/sh\n+\n+test_description='rev-list with .keep packs'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup' '\n+\ttest_commit loose &&\n+\ttest_commit packed &&\n+\ttest_commit kept &&\n+\n+\tKEPT_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/kept\n+\t^refs/tags/packed\n+\tEOF\n+\t) &&\n+\tMISC_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/packed\n+\t^refs/tags/loose\n+\tEOF\n+\t) &&\n+\n+\ttouch .git/objects/pack/pack-$KEPT_PACK.keep\n+'\n+\n+rev_list_objects () {\n+\tgit rev-list \"$@\" >out &&\n+\tsort out\n+}\n+\n+idx_objects () {\n+\tgit show-index <$1 >expect-idx &&\n+\tcut -d\" \" -f2 <expect-idx | sort\n+}\n+\n+test_expect_success '--no-kept-objects excludes trees and blobs in .keep packs' '\n+\trev_list_objects --objects --all --no-object-names >kept &&\n+\trev_list_objects --objects --all --no-object-names --no-kept-objects >no-kept &&\n+\n+\tidx_objects .git/objects/pack/pack-$KEPT_PACK.idx >expect &&\n+\tcomm -3 kept no-kept >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success '--no-kept-objects excludes kept non-MIDX object' '\n+\ttest_config core.multiPackIndex true &&\n+\n+\t# Create a pack with just the commit object in pack, and do not mark it\n+\t# as kept (even though it appears in $KEPT_PACK, which does have a .keep\n+\t# file).\n+\tMIDX_PACK=$(git pack-objects .git/objects/pack/pack <<-EOF\n+\t$(git rev-parse kept)\n+\tEOF\n+\t) &&\n+\n+\t# Write a MIDX containing all packs, but use the version of the commit\n+\t# at \"kept\" in a non-kept pack by touching $MIDX_PACK.\n+\ttouch .git/objects/pack/pack-$MIDX_PACK.pack &&\n+\tgit multi-pack-index write &&\n+\n+\trev_list_objects --objects --no-object-names --no-kept-objects HEAD >actual &&\n+\t(\n+\t\tidx_objects .git/objects/pack/pack-$MISC_PACK.idx &&\n+\t\tgit rev-list --objects --no-object-names refs/tags/loose\n+\t) | sort >expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414753","messageId":"dc7fa4c7a61f657e779e10385d3e8076d6dac36c.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 01/10] packfile: introduce 'find_kept_pack_entry()'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:00Z","receivedAt":"2021-01-19T23:26:52Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Future callers will want a function to fill a 'struct pack_entry' for a\ngiven object id but _only_ from its position in any kept pack(s). They\ncould accomplish this by calling 'find_pack_entry()' and checking\nwhether the found pack is kept or not, but this is insufficient, since\nthere may be duplicate objects (and the mru cache makes it unpredictable\nwhich variant we'll get).\n\nTeach this new function to treat the two different kinds of kept packs\n(on disk ones with .keep files, as well as in-core ones which are set by\nmanually poking the 'pack_keep_in_core' bit) separately. This will\nbecome important for callers that only want to respect a certain kind of\nkept pack.\n\nIntroduce 'find_kept_pack_entry()' which behaves like\n'find_pack_entry()', except that it skips over packs which are not\nmarked kept. Callers will be added in subsequent patches.\n\nCo-authored-by: Jeff King <peff@peff.net>\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n packfile.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++-----\n packfile.h |  6 +++++\n 2 files changed, 65 insertions(+), 5 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex 62d92e0c7c..30f43a1a35 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2015,7 +2015,10 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static int find_one_pack_entry(struct repository *r,\n+\t\t\t       const struct object_id *oid,\n+\t\t\t       struct pack_entry *e,\n+\t\t\t       int kept_only)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2025,26 +2028,77 @@ int find_pack_entry(struct repository *r, const struct object_id *oid, struct pa\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (fill_midx_entry(r, oid, e, m))\n+\t\tif (!(fill_midx_entry(r, oid, e, m)))\n+\t\t\tcontinue;\n+\n+\t\tif (!kept_only)\n+\t\t\treturn 1;\n+\n+\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n+\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n-\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\treturn 1;\n+\t\tif (p->multi_pack_index && !kept_only) {\n+\t\t\t/*\n+\t\t\t * If this pack is covered by the MIDX, we'd have found\n+\t\t\t * the object already in the loop above if it was here,\n+\t\t\t * so don't bother looking.\n+\t\t\t *\n+\t\t\t * The exception is if we are looking only at kept\n+\t\t\t * packs. An object can be present in two packs covered\n+\t\t\t * by the MIDX, one kept and one not-kept. And as the\n+\t\t\t * MIDX points to only one copy of each object, it might\n+\t\t\t * have returned only the non-kept version above. We\n+\t\t\t * have to check again to be thorough.\n+\t\t\t */\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!kept_only ||\n+\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n+\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n+\t\t\tif (fill_pack_entry(oid, e, p)) {\n+\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\t\treturn 1;\n+\t\t\t}\n \t\t}\n \t}\n \treturn 0;\n }\n \n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+{\n+\treturn find_one_pack_entry(r, oid, e, 0);\n+}\n+\n+int find_kept_pack_entry(struct repository *r,\n+\t\t\t const struct object_id *oid,\n+\t\t\t unsigned flags,\n+\t\t\t struct pack_entry *e)\n+{\n+\t/*\n+\t * Load all packs, including midx packs, since our \"kept\" strategy\n+\t * relies on that. We're relying on the side effect of it setting up\n+\t * r->objects->packed_git, which is a little ugly.\n+\t */\n+\tget_all_packs(r);\n+\treturn find_one_pack_entry(r, oid, e, flags);\n+}\n+\n int has_object_pack(const struct object_id *oid)\n {\n \tstruct pack_entry e;\n \treturn find_pack_entry(the_repository, oid, &e);\n }\n \n+int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n+{\n+\tstruct pack_entry e;\n+\treturn find_kept_pack_entry(the_repository, oid, flags, &e);\n+}\n+\n int has_pack_index(const unsigned char *sha1)\n {\n \tstruct stat st;\ndiff --git a/packfile.h b/packfile.h\nindex a58fc738e0..624327f64d 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -161,13 +161,19 @@ int packed_object_info(struct repository *r,\n void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n \n+#define ON_DISK_KEEP_PACKS 1\n+#define IN_CORE_KEEP_PACKS 2\n+#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n+\n /*\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\n  */\n int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e);\n+int find_kept_pack_entry(struct repository *r, const struct object_id *oid, unsigned flags, struct pack_entry *e);\n \n int has_object_pack(const struct object_id *oid);\n+int has_object_kept_pack(const struct object_id *oid, unsigned flags);\n \n int has_pack_index(const unsigned char *sha1);\n \n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414754","messageId":"182664e1a9c107fd830f9eb02bfa25fe9679e9a7.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 07/10] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:25Z","receivedAt":"2021-01-19T23:26:52Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nIn a recent patch we added a function 'find_kept_pack_entry()' to look\nfor an object only among kept packs.\n\nWhile this function avoids doing any lookup work in non-kept packs, it\nis still linear in the number of packs, since we have to traverse the\nlinked list of packs once per object. Let's cache a reduced version of\nthat list to save us time.\n\nNote that this cache will last the lifetime of the program. We could\ninvalidate it on reprepare_packed_git(), but there's not much point in\nbeing rigorous here:\n\n  - we might already fail to notice new .keep packs showing up after the\n    program starts. We only reprepare_packed_git() when we fail to find\n    an object. But adding a new pack won't cause that to happen.\n    Somebody repacking could add a new pack and delete an old one, but\n    most of the time we'd have a descriptor or mmap open to the old\n    pack anyway, so we might not even notice.\n\n  - in pack-objects we already cache the .keep state at startup, since\n    56dfeb6263 (pack-objects: compute local/ignore_pack_keep early,\n    2016-07-29). So this is just extending that concept further.\n\n  - we don't have to worry about any packed_git being removed; we always\n    keep the old structs around, even after reprepare_packed_git()\n\nHere are p5303 results (as always, measured against the kernel):\n\n  Test                               HEAD^                  HEAD\n  ------------------------------------------------------------------------------------\n  5303.5: repack (1)                 56.87(54.63+10.48)     56.63(54.41+10.36) -0.4%\n  5303.6: repack with keep (1)       1.26(1.19+0.06)        1.25(1.19+0.05) -0.8%\n  5303.10: repack (50)               89.35(132.42+6.25)     89.49(132.31+6.31) +0.2%\n  5303.11: repack with keep (50)     6.73(26.61+0.59)       6.72(26.70+0.53) -0.1%\n  5303.15: repack (1000)             217.25(494.38+15.24)   218.69(495.62+14.99) +0.7%\n  5303.16: repack with keep (1000)   133.12(311.80+8.44)    128.79(306.96+8.55) -3.3%\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c |   4 +-\n object-store.h         |  10 ++++\n packfile.c             | 103 +++++++++++++++++++++++------------------\n packfile.h             |   4 --\n revision.c             |   8 ++--\n 5 files changed, 75 insertions(+), 54 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex c84642df98..f2c7a1e35b 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1215,9 +1215,9 @@ static int want_found_object(const struct object_id *oid, int exclude,\n \t\t */\n \t\tunsigned flags = 0;\n \t\tif (ignore_packed_keep_on_disk)\n-\t\t\tflags |= ON_DISK_KEEP_PACKS;\n+\t\t\tflags |= CACHE_ON_DISK_KEEP_PACKS;\n \t\tif (ignore_packed_keep_in_core)\n-\t\t\tflags |= IN_CORE_KEEP_PACKS;\n+\t\t\tflags |= CACHE_IN_CORE_KEEP_PACKS;\n \n \t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n \t\t\treturn 0;\ndiff --git a/object-store.h b/object-store.h\nindex c4fc9dd74e..4cbe8eae3c 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -105,6 +105,14 @@ static inline int pack_map_entry_cmp(const void *unused_cmp_data,\n \treturn strcmp(pg1->pack_name, key ? key : pg2->pack_name);\n }\n \n+#define CACHE_ON_DISK_KEEP_PACKS 1\n+#define CACHE_IN_CORE_KEEP_PACKS 2\n+\n+struct kept_pack_cache {\n+\tstruct packed_git **packs;\n+\tunsigned flags;\n+};\n+\n struct raw_object_store {\n \t/*\n \t * Set of all object directories; the main directory is first (and\n@@ -150,6 +158,8 @@ struct raw_object_store {\n \t/* A most-recently-used ordered version of the packed_git list. */\n \tstruct list_head packed_git_mru;\n \n+\tstruct kept_pack_cache *kept_pack_cache;\n+\n \t/*\n \t * A map of packfiles to packed_git structs for tracking which\n \t * packs have been loaded already.\ndiff --git a/packfile.c b/packfile.c\nindex 30f43a1a35..25f5407ed0 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2015,10 +2015,7 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int find_one_pack_entry(struct repository *r,\n-\t\t\t       const struct object_id *oid,\n-\t\t\t       struct pack_entry *e,\n-\t\t\t       int kept_only)\n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2028,49 +2025,64 @@ static int find_one_pack_entry(struct repository *r,\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (!(fill_midx_entry(r, oid, e, m)))\n-\t\t\tcontinue;\n-\n-\t\tif (!kept_only)\n-\t\t\treturn 1;\n-\n-\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n-\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n+\t\tif (fill_midx_entry(r, oid, e, m))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (p->multi_pack_index && !kept_only) {\n-\t\t\t/*\n-\t\t\t * If this pack is covered by the MIDX, we'd have found\n-\t\t\t * the object already in the loop above if it was here,\n-\t\t\t * so don't bother looking.\n-\t\t\t *\n-\t\t\t * The exception is if we are looking only at kept\n-\t\t\t * packs. An object can be present in two packs covered\n-\t\t\t * by the MIDX, one kept and one not-kept. And as the\n-\t\t\t * MIDX points to only one copy of each object, it might\n-\t\t\t * have returned only the non-kept version above. We\n-\t\t\t * have to check again to be thorough.\n-\t\t\t */\n-\t\t\tcontinue;\n-\t\t}\n-\t\tif (!kept_only ||\n-\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n-\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n-\t\t\tif (fill_pack_entry(oid, e, p)) {\n-\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\t\treturn 1;\n-\t\t\t}\n+\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n+\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\treturn 1;\n \t\t}\n \t}\n \treturn 0;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static void maybe_invalidate_kept_pack_cache(struct repository *r,\n+\t\t\t\t\t     unsigned flags)\n {\n-\treturn find_one_pack_entry(r, oid, e, 0);\n+\tif (!r->objects->kept_pack_cache)\n+\t\treturn;\n+\tif (r->objects->kept_pack_cache->flags == flags)\n+\t\treturn;\n+\tfree(r->objects->kept_pack_cache->packs);\n+\tFREE_AND_NULL(r->objects->kept_pack_cache);\n+}\n+\n+static struct packed_git **kept_pack_cache(struct repository *r, unsigned flags)\n+{\n+\tmaybe_invalidate_kept_pack_cache(r, flags);\n+\n+\tif (!r->objects->kept_pack_cache) {\n+\t\tstruct packed_git **packs = NULL;\n+\t\tsize_t nr = 0, alloc = 0;\n+\t\tstruct packed_git *p;\n+\n+\t\t/*\n+\t\t * We want \"all\" packs here, because we need to cover ones that\n+\t\t * are used by a midx, as well. We need to look in every one of\n+\t\t * them (instead of the midx itself) to cover duplicates. It's\n+\t\t * possible that an object is found in two packs that the midx\n+\t\t * covers, one kept and one not kept, but the midx returns only\n+\t\t * the non-kept version.\n+\t\t */\n+\t\tfor (p = get_all_packs(r); p; p = p->next) {\n+\t\t\tif ((p->pack_keep && (flags & CACHE_ON_DISK_KEEP_PACKS)) ||\n+\t\t\t    (p->pack_keep_in_core && (flags & CACHE_IN_CORE_KEEP_PACKS))) {\n+\t\t\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\t\t\tpacks[nr++] = p;\n+\t\t\t}\n+\t\t}\n+\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\tpacks[nr] = NULL;\n+\n+\t\tr->objects->kept_pack_cache = xmalloc(sizeof(*r->objects->kept_pack_cache));\n+\t\tr->objects->kept_pack_cache->packs = packs;\n+\t\tr->objects->kept_pack_cache->flags = flags;\n+\t}\n+\n+\treturn r->objects->kept_pack_cache->packs;\n }\n \n int find_kept_pack_entry(struct repository *r,\n@@ -2078,13 +2090,15 @@ int find_kept_pack_entry(struct repository *r,\n \t\t\t unsigned flags,\n \t\t\t struct pack_entry *e)\n {\n-\t/*\n-\t * Load all packs, including midx packs, since our \"kept\" strategy\n-\t * relies on that. We're relying on the side effect of it setting up\n-\t * r->objects->packed_git, which is a little ugly.\n-\t */\n-\tget_all_packs(r);\n-\treturn find_one_pack_entry(r, oid, e, flags);\n+\tstruct packed_git **cache;\n+\n+\tfor (cache = kept_pack_cache(r, flags); *cache; cache++) {\n+\t\tstruct packed_git *p = *cache;\n+\t\tif (fill_pack_entry(oid, e, p))\n+\t\t\treturn 1;\n+\t}\n+\n+\treturn 0;\n }\n \n int has_object_pack(const struct object_id *oid)\n@@ -2093,7 +2107,8 @@ int has_object_pack(const struct object_id *oid)\n \treturn find_pack_entry(the_repository, oid, &e);\n }\n \n-int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n+int has_object_kept_pack(const struct object_id *oid,\n+\t\t\t unsigned flags)\n {\n \tstruct pack_entry e;\n \treturn find_kept_pack_entry(the_repository, oid, flags, &e);\ndiff --git a/packfile.h b/packfile.h\nindex 624327f64d..eb56db2a7b 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -161,10 +161,6 @@ int packed_object_info(struct repository *r,\n void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n \n-#define ON_DISK_KEEP_PACKS 1\n-#define IN_CORE_KEEP_PACKS 2\n-#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n-\n /*\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\ndiff --git a/revision.c b/revision.c\nindex ff1ea77224..ce87081b8e 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -2336,14 +2336,14 @@ static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n \t\tdie(_(\"--unpacked=<packfile> no longer supported\"));\n \t} else if (!strcmp(arg, \"--no-kept-objects\")) {\n \t\trevs->no_kept_objects = 1;\n-\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n-\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= CACHE_ON_DISK_KEEP_PACKS;\n \t} else if (skip_prefix(arg, \"--no-kept-objects=\", &optarg)) {\n \t\trevs->no_kept_objects = 1;\n \t\tif (!strcmp(optarg, \"in-core\"))\n-\t\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\t\trevs->keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n \t\tif (!strcmp(optarg, \"on-disk\"))\n-\t\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t\t\trevs->keep_pack_cache_flags |= CACHE_ON_DISK_KEEP_PACKS;\n \t} else if (!strcmp(arg, \"-r\")) {\n \t\trevs->diff = 1;\n \t\trevs->diffopt.flags.recursive = 1;\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414755","messageId":"4dd5076fcc94fca906b1e5471f08288b4d225cab.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 06/10] pack-objects: rewrite honor-pack-keep logic","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:21Z","receivedAt":"2021-01-19T23:26:52Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nNow that we have find_kept_pack_entry(), we don't have to manually keep\nhunting through every pack to find a possible \"kept\" duplicate of the\nobject. This should be faster, assuming only a portion of your total\npacks are actually kept.\n\nNote that we have to re-order the logic a bit here; we can deal with the\n\"kept\" situation completely, and then just fall back to the \"--local\"\nquestion. It might be worth having a similar optimized function to look\nat only local packs.\n\nHere are the results from p5303 (measurements taken on git.git):\n\n  Test                               HEAD^                  HEAD\n  ------------------------------------------------------------------------------------\n  5303.5: repack (1)                 57.29(54.88+10.39)     56.87(54.63+10.48) -0.7%\n  5303.6: repack with keep (1)       1.25(1.19+0.05)        1.26(1.19+0.06) +0.8%\n  5303.10: repack (50)               89.71(132.78+6.14)     89.35(132.42+6.25) -0.4%\n  5303.11: repack with keep (50)     6.92(26.93+0.58)       6.73(26.61+0.59) -2.7%\n  5303.15: repack (1000)             217.14(493.76+15.29)   217.25(494.38+15.24) +0.1%\n  5303.16: repack with keep (1000)   209.46(387.83+8.42)    133.12(311.80+8.44) -36.4%\n\nSo our case with many packs and a .keep is finally now faster than the\nnon-keep case (because it gets the speed benefit of looking at fewer\nobjects, but not as big a penalty for looking at many packs).\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 125 ++++++++++++++++++++++++-----------------\n 1 file changed, 73 insertions(+), 52 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex a5dcd66f52..c84642df98 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1178,7 +1178,8 @@ static int have_duplicate_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int want_found_object(int exclude, struct packed_git *p)\n+static int want_found_object(const struct object_id *oid, int exclude,\n+\t\t\t     struct packed_git *p)\n {\n \tif (exclude)\n \t\treturn 1;\n@@ -1199,22 +1200,73 @@ static int want_found_object(int exclude, struct packed_git *p)\n \t * Otherwise, we signal \"-1\" at the end to tell the caller that we do\n \t * not know either way, and it needs to check more packs.\n \t */\n-\tif (!ignore_packed_keep_on_disk &&\n-\t    !ignore_packed_keep_in_core &&\n-\t    (!local || !have_non_local_packs))\n+\n+\t/*\n+\t * Handle .keep first, as we have a fast(er) path there.\n+\t */\n+\tif (ignore_packed_keep_on_disk || ignore_packed_keep_in_core) {\n+\t\t/*\n+\t\t * Set the flags for the kept-pack cache to be the ones we want\n+\t\t * to ignore.\n+\t\t *\n+\t\t * That is, if we are ignoring objects in on-disk keep packs,\n+\t\t * then we want to search through the on-disk keep and ignore\n+\t\t * the in-core ones.\n+\t\t */\n+\t\tunsigned flags = 0;\n+\t\tif (ignore_packed_keep_on_disk)\n+\t\t\tflags |= ON_DISK_KEEP_PACKS;\n+\t\tif (ignore_packed_keep_in_core)\n+\t\t\tflags |= IN_CORE_KEEP_PACKS;\n+\n+\t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n+\t\t\treturn 0;\n+\t\tif (ignore_packed_keep_in_core && p->pack_keep_in_core)\n+\t\t\treturn 0;\n+\t\tif (has_object_kept_pack(oid, flags))\n+\t\t\treturn 0;\n+\t}\n+\n+\t/*\n+\t * At this point we know definitively that either we don't care about\n+\t * keep-packs, or the object is not in one. Keep checking other\n+\t * conditions...\n+\t */\n+\n+\tif (!local || !have_non_local_packs)\n \t\treturn 1;\n-\n \tif (local && !p->pack_local)\n \t\treturn 0;\n-\tif (p->pack_local &&\n-\t    ((ignore_packed_keep_on_disk && p->pack_keep) ||\n-\t     (ignore_packed_keep_in_core && p->pack_keep_in_core)))\n-\t\treturn 0;\n \n \t/* we don't know yet; keep looking for more packs */\n \treturn -1;\n }\n \n+static int want_object_in_pack_one(struct packed_git *p,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   int exclude,\n+\t\t\t\t   struct packed_git **found_pack,\n+\t\t\t\t   off_t *found_offset)\n+{\n+\toff_t offset;\n+\n+\tif (p == *found_pack)\n+\t\toffset = *found_offset;\n+\telse\n+\t\toffset = find_pack_entry_one(oid->hash, p);\n+\n+\tif (offset) {\n+\t\tif (!*found_pack) {\n+\t\t\tif (!is_pack_valid(p))\n+\t\t\t\treturn -1;\n+\t\t\t*found_offset = offset;\n+\t\t\t*found_pack = p;\n+\t\t}\n+\t\treturn want_found_object(oid, exclude, p);\n+\t}\n+\treturn -1;\n+}\n+\n /*\n  * Check whether we want the object in the pack (e.g., we do not want\n  * objects found in non-local stores if the \"--local\" option was used).\n@@ -1242,7 +1294,7 @@ static int want_object_in_pack(const struct object_id *oid,\n \t * are present we will determine the answer right now.\n \t */\n \tif (*found_pack) {\n-\t\twant = want_found_object(exclude, *found_pack);\n+\t\twant = want_found_object(oid, exclude, *found_pack);\n \t\tif (want != -1)\n \t\t\treturn want;\n \t}\n@@ -1250,53 +1302,22 @@ static int want_object_in_pack(const struct object_id *oid,\n \tfor (m = get_multi_pack_index(the_repository); m; m = m->next) {\n \t\tstruct pack_entry e;\n \t\tif (fill_midx_entry(the_repository, oid, &e, m)) {\n-\t\t\tstruct packed_git *p = e.p;\n-\t\t\toff_t offset;\n-\n-\t\t\tif (p == *found_pack)\n-\t\t\t\toffset = *found_offset;\n-\t\t\telse\n-\t\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\t\tif (offset) {\n-\t\t\t\tif (!*found_pack) {\n-\t\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\t\tcontinue;\n-\t\t\t\t\t*found_offset = offset;\n-\t\t\t\t\t*found_pack = p;\n-\t\t\t\t}\n-\t\t\t\twant = want_found_object(exclude, p);\n-\t\t\t\tif (want != -1)\n-\t\t\t\t\treturn want;\n-\t\t\t}\n-\t\t}\n-\t}\n-\n-\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n-\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\toff_t offset;\n-\n-\t\tif (p == *found_pack)\n-\t\t\toffset = *found_offset;\n-\t\telse\n-\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\tif (offset) {\n-\t\t\tif (!*found_pack) {\n-\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\tcontinue;\n-\t\t\t\t*found_offset = offset;\n-\t\t\t\t*found_pack = p;\n-\t\t\t}\n-\t\t\twant = want_found_object(exclude, p);\n-\t\t\tif (!exclude && want > 0)\n-\t\t\t\tlist_move(&p->mru,\n-\t\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\t\twant = want_object_in_pack_one(e.p, oid, exclude, found_pack, found_offset);\n \t\t\tif (want != -1)\n \t\t\t\treturn want;\n \t\t}\n \t}\n \n+\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n+\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n+\t\twant = want_object_in_pack_one(p, oid, exclude, found_pack, found_offset);\n+\t\tif (!exclude && want > 0)\n+\t\t\tlist_move(&p->mru,\n+\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\tif (want != -1)\n+\t\t\treturn want;\n+\t}\n+\n \tif (uri_protocols.nr) {\n \t\tstruct configured_exclusion *ex =\n \t\t\toidmap_get(&configured_exclusions, oid);\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414756","messageId":"b3b2574d4d9d10f226b52d81fe0e6bf1f761504e.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 05/10] p5303: measure time to repack with keep","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:16Z","receivedAt":"2021-01-19T23:27:52Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nThis is the same as the regular repack test, except that we mark the\nsingle base pack as \"kept\" and use --assume-kept-packs-closed. The\ntheory is that this should be faster than the normal repack, because\nwe'll have fewer objects to traverse and process.\n\nAnd indeed, it is much faster in the single-pack case (all timings\nmeasured on the kernel):\n\n  5303.5: repack (1)                 57.29(54.88+10.39)\n  5303.6: repack with keep (1)       1.25(1.19+0.05)\n\nand in the 50-pack case:\n\n  5303.10: repack (50)               89.71(132.78+6.14)\n  5303.11: repack with keep (50)     6.92(26.93+0.58)\n\nbut our improvements vanish as we approach 1000 packs.\n\n  5303.15: repack (1000)             217.14(493.76+15.29)\n  5303.16: repack with keep (1000)   209.46(387.83+8.42)\n\nThat's because the code paths around handling .keep files are known to\nscale badly; they look in every single pack file to find each object.\nOur solution to that was to notice that most repos don't have keep\nfiles, and to make that case a fast path. But as soon as you add a\nsingle .keep, that part of pack-objects slows down again (even if we\nhave fewer objects total to look at).\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 16 ++++++++++++++--\n 1 file changed, 14 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex 277d22ec4b..85b077b72b 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -27,8 +27,11 @@ repack_into_n () {\n \t>pushes &&\n \n \t# create base packfile\n-\thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack &&\n+\tbase_pack=$(\n+\t\thead -n 1 pushes |\n+\t\tgit pack-objects --delta-base-offset --revs staging/pack\n+\t) &&\n+\ttest_export base_pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n@@ -87,6 +90,15 @@ do\n \t\t  --reflog --indexed-objects --delta-base-offset \\\n \t\t  --stdout </dev/null >/dev/null\n \t'\n+\n+\ttest_perf \"repack with keep ($nr_packs)\" '\n+\t\tgit pack-objects --keep-true-parents \\\n+\t\t  --honor-pack-keep --assume-kept-packs-closed \\\n+\t\t  --keep-pack=pack-$base_pack.pack \\\n+\t\t  --non-empty --all \\\n+\t\t  --reflog --indexed-objects --delta-base-offset \\\n+\t\t  --stdout </dev/null >/dev/null\n+\t'\n done\n \n # Measure pack loading with 10,000 packs.\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414757","messageId":"6547c082f8f2b696f8711295bdeb4a24a09dffe5.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 08/10] builtin/pack-objects.c: teach '--keep-pack-stdin'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:29Z","receivedAt":"2021-01-19T23:28:06Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Add a shortcut to specify '--keep-pack=<pack-name>' arguments over\nstdin, in case a caller wishes to indicate more kept packs than the\nargument limit will allow.\n\nPassing this option overrides any other option to 'git pack-objects'\nthat takes input over stdin. For example, '--revs' still forces a\nreachability traversal, but will not accept any revision arguments over\nstdin. Use of '--keep-pack-stdin' within Git is limited to one caller\n(added in a subsequent patch) which does not pass any other input over\nstdin.\n\nNo new tests are added here, since a caller from 'git repack' will\nexercise these options in a subsequent patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  8 ++++++++\n builtin/pack-objects.c             | 23 ++++++++++++++++++++---\n 2 files changed, 28 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex cbe08e7415..45ecc4e9e5 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -135,6 +135,14 @@ depth is 4095.\n \tleading directory (e.g. `pack-123.pack`). The option could be\n \tspecified multiple times to keep multiple packs.\n \n+--keep-pack-stdin::\n+\tTake a list of line-delimited `<pack-name>` arguments, treating\n+\tthem as if they were each passed as `--keep-pack=<pack-name>`.\n+\tUseful for when many packs are being kept to avoid argument\n+\tlength limitations. Requires that `--revs` be passed or implied,\n+\tbut does not allow the caller to pass additional traversal\n+\targuments over standard input.\n+\n --assume-kept-packs-closed::\n \tThis flag causes `git rev-list` to halt the object traversal\n \twhen it encounters an object found in a kept pack. This is\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f2c7a1e35b..f528a07d78 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3343,7 +3343,7 @@ static void record_recent_commit(struct commit *commit, void *data)\n \toid_array_append(&recent_objects, &commit->object.oid);\n }\n \n-static void get_object_list(int ac, const char **av)\n+static void get_object_list(int ac, const char **av, int read_from_stdin)\n {\n \tstruct rev_info revs;\n \tstruct setup_revision_opt s_r_opt = {\n@@ -3363,7 +3363,7 @@ static void get_object_list(int ac, const char **av)\n \tsave_warning = warn_on_object_refname_ambiguity;\n \twarn_on_object_refname_ambiguity = 0;\n \n-\twhile (fgets(line, sizeof(line), stdin) != NULL) {\n+\twhile (read_from_stdin && fgets(line, sizeof(line), stdin) != NULL) {\n \t\tint len = strlen(line);\n \t\tif (len && line[len - 1] == '\\n')\n \t\t\tline[--len] = 0;\n@@ -3487,6 +3487,15 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n \treturn 0;\n }\n \n+static void collect_kept_packs(struct string_list *keep_pack_list)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\twhile (strbuf_getline(&buf, stdin) != EOF)\n+\t\tstring_list_append(keep_pack_list,\n+\t\t\t\t   strbuf_detach(&buf, NULL));\n+\tstrbuf_release(&buf);\n+}\n+\n int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n {\n \tint use_internal_rev_list = 0;\n@@ -3496,6 +3505,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tint rev_list_unpacked = 0, rev_list_all = 0, rev_list_reflog = 0;\n \tint rev_list_index = 0;\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n+\tint keep_pack_stdin = 0;\n \tstruct option pack_objects_options[] = {\n \t\tOPT_SET_INT('q', \"quiet\", &progress,\n \t\t\t    N_(\"do not show progress meter\"), 0),\n@@ -3568,6 +3578,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t N_(\"assume the union of kept packs is closed under reachability\")),\n \t\tOPT_STRING_LIST(0, \"keep-pack\", &keep_pack_list, N_(\"name\"),\n \t\t\t\tN_(\"ignore this pack\")),\n+\t\tOPT_BOOL(0, \"keep-pack-stdin\", &keep_pack_stdin,\n+\t\t\t N_(\"read the list of kept packs from stdin\")),\n \t\tOPT_INTEGER(0, \"compression\", &pack_compression_level,\n \t\t\t    N_(\"pack compression level\")),\n \t\tOPT_SET_INT(0, \"keep-true-parents\", &grafts_replace_parents,\n@@ -3728,6 +3740,11 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (progress && all_progress_implied)\n \t\tprogress = 2;\n \n+\tif (keep_pack_stdin) {\n+\t\tif (!use_internal_rev_list)\n+\t\t\tdie(_(\"--keep-pack-stdin requires --revs\"));\n+\t\tcollect_kept_packs(&keep_pack_list);\n+\t}\n \tadd_extra_kept_packs(&keep_pack_list);\n \tif (ignore_packed_keep_on_disk) {\n \t\tstruct packed_git *p;\n@@ -3769,7 +3786,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (!use_internal_rev_list)\n \t\tread_object_list_from_stdin();\n \telse {\n-\t\tget_object_list(rp.nr, rp.v);\n+\t\tget_object_list(rp.nr, rp.v, !keep_pack_stdin);\n \t\tstrvec_clear(&rp);\n \t}\n \tcleanup_preferred_base();\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414758","messageId":"f853087216cc7f1af8d376b5a8a9c86086ef8df0.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 10/10] builtin/repack.c: add '--geometric' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:37Z","receivedAt":"2021-01-19T23:28:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Often it is useful to both:\n\n  - have relatively few packfiles in a repository, and\n\n  - avoid having so few packfiles in a repository that we repack its\n    entire contents regularly\n\nThis patch implements a '--geometric=<n>' option in 'git repack'. This\nallows the caller to specify that they would like each pack to be at\nleast a factor times as large as the previous largest pack (by object\ncount).\n\nConcretely, say that a repository has 'n' packfiles, labeled P1, P2,\n..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\nWith a geometric factor of 'r', it should be that:\n\n  objects(Pi) > r*objects(P(i-1))\n\nfor all i in [1, n], where the packs are sorted by\n\n  objects(P1) <= objects(P2) <= ... <= objects(Pn).\n\nSince finding a true optimal repacking is NP-hard, we approximate it\nalong two directions:\n\n  1. We assume that there is a cutoff of packs _before starting the\n     repack_ where everything to the right of that cut-off already forms\n     a geometric progression (or no cutoff exists and everything must be\n     repacked).\n\n  2. We assume that everything smaller than the cutoff count must be\n     repacked. This forms our base assumption, but it can also cause\n     even the \"heavy\" packs to get repacked, for e.g., if we have 6\n     packs containing the following number of objects:\n\n       1, 1, 1, 2, 4, 32\n\n     then we would place the cutoff between '1, 1' and '1, 2, 4, 32',\n     rolling up the first two packs into a pack with 2 objects. That\n     breaks our progression and leaves us:\n\n       2, 1, 2, 4, 32\n         ^\n\n     (where the '^' indicates the position of our split). To restore a\n     progression, we move the split forward (towards larger packs)\n     joining each pack into our new pack until a geometric progression\n     is restored. Here, that looks like:\n\n       2, 1, 2, 4, 32  ~>  3, 2, 4, 32  ~>  5, 4, 32  ~> ... ~> 9, 32\n         ^                   ^                ^                   ^\n\nThis has the advantage of not repacking the heavy-side of packs too\noften while also only creating one new pack at a time. Another wrinkle\nis that we assume that loose, indexed, and reflog'd objects are\ninsignificant, and lump them into any new pack that we create. This can\nlead to non-idempotent results.\n\nSuggested-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt |  11 +++\n builtin/repack.c             | 165 ++++++++++++++++++++++++++++++++++-\n t/t7703-repack-geometric.sh  |  81 +++++++++++++++++\n 3 files changed, 256 insertions(+), 1 deletion(-)\n create mode 100755 t/t7703-repack-geometric.sh\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex 92f146d27d..b1ffcfd974 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -165,6 +165,17 @@ depth is 4095.\n \tPass the `--delta-islands` option to `git-pack-objects`, see\n \tlinkgit:git-pack-objects[1].\n \n+-g=<factor>::\n+--geometric=<factor>::\n+\tArrange resulting pack structure so that each successive pack\n+\tcontains at least `<factor>` times the number of objects as the\n+\tnext-largest pack.\n++\n+`git repack` ensures this by determining a \"cut\" of packfiles that need to be\n+repacked into one in order to ensure a geometric progression. It picks the\n+smallest set of packfiles such that as many of the larger packfiles (by count of\n+objects contained in that pack) may be left intact.\n+\n Configuration\n -------------\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 664863111b..083088ae1f 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -298,6 +298,116 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n #define ALL_INTO_ONE 1\n #define LOOSEN_UNREACHABLE 2\n \n+struct pack_geometry {\n+\tstruct packed_git **pack;\n+\tuint32_t pack_nr, pack_alloc;\n+\tuint32_t split;\n+};\n+\n+static uint32_t geometry_pack_weight(struct packed_git *p)\n+{\n+\tif (open_pack_index(p))\n+\t\tdie(_(\"cannot open index for %s\"), p->pack_name);\n+\treturn p->num_objects;\n+}\n+\n+static int geometry_cmp(const void *va, const void *vb)\n+{\n+\tuint32_t aw = geometry_pack_weight(*(struct packed_git **)va),\n+\t\t bw = geometry_pack_weight(*(struct packed_git **)vb);\n+\n+\tif (aw < bw)\n+\t\treturn -1;\n+\tif (aw > bw)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static void init_pack_geometry(struct pack_geometry **geometry_p)\n+{\n+\tstruct packed_git *p;\n+\tstruct pack_geometry *geometry;\n+\n+\t*geometry_p = xcalloc(1, sizeof(struct pack_geometry));\n+\tgeometry = *geometry_p;\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tALLOC_GROW(geometry->pack,\n+\t\t\t   geometry->pack_nr + 1,\n+\t\t\t   geometry->pack_alloc);\n+\n+\t\tgeometry->pack[geometry->pack_nr] = p;\n+\t\tgeometry->pack_nr++;\n+\t}\n+\n+\tQSORT(geometry->pack, geometry->pack_nr, geometry_cmp);\n+}\n+\n+static void split_pack_geometry(struct pack_geometry *geometry, int factor)\n+{\n+\tuint32_t i;\n+\tuint32_t split;\n+\toff_t total_size = 0;\n+\n+\tsplit = geometry->pack_nr - 1;\n+\n+\t/*\n+\t * First, count the number of packs (in descending order of size) which\n+\t * already form a geometric progression.\n+\t */\n+\tfor (i = geometry->pack_nr - 1; i > 0; i--) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tstruct packed_git *prev = geometry->pack[i - 1];\n+\t\tif (geometry_pack_weight(ours) >= factor * geometry_pack_weight(prev))\n+\t\t\tsplit--;\n+\t\telse\n+\t\t\tbreak;\n+\t}\n+\n+\tif (split) {\n+\t\t/*\n+\t\t * Move the split one to the right, since the top element in the\n+\t\t * last-compared pair can't be in the progression. Only do this\n+\t\t * when we split in the middle of the array (otherwise if we got\n+\t\t * to the end, then the split is in the right place).\n+\t\t */\n+\t\tsplit++;\n+\t}\n+\n+\t/*\n+\t * Then, anything to the left of 'split' must be in a new pack. But,\n+\t * creating that new pack may cause packs in the heavy half to no longer\n+\t * form a geometric progression.\n+\t *\n+\t * Compute an expected size of the new pack, and then determine how many\n+\t * packs in the heavy half need to be joined into it (if any) to restore\n+\t * the geometric progression.\n+\t */\n+\tfor (i = 0; i < split; i++)\n+\t\ttotal_size += geometry_pack_weight(geometry->pack[i]);\n+\tfor (i = split; i < geometry->pack_nr; i++) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tif (geometry_pack_weight(ours) < factor * total_size) {\n+\t\t\tsplit++;\n+\t\t\ttotal_size += geometry_pack_weight(ours);\n+\t\t} else\n+\t\t\tbreak;\n+\t}\n+\n+\tgeometry->split = split;\n+}\n+\n+static void clear_pack_geometry(struct pack_geometry *geometry)\n+{\n+\tif (!geometry)\n+\t\treturn;\n+\n+\tfree(geometry->pack);\n+\tgeometry->pack_nr = 0;\n+\tgeometry->pack_alloc = 0;\n+\tgeometry->split = 0;\n+}\n+\n static void handle_loose_and_reachable(struct child_process *cmd,\n \t\t\t\t       const char *unpack_unreachable,\n \t\t\t\t       int pack_everything,\n@@ -326,6 +436,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list names = STRING_LIST_INIT_DUP;\n \tstruct string_list rollback = STRING_LIST_INIT_NODUP;\n \tstruct string_list existing_packs = STRING_LIST_INIT_DUP;\n+\tstruct pack_geometry *geometry = NULL;\n \tstruct strbuf line = STRBUF_INIT;\n \tint i, ext, ret;\n \tFILE *out;\n@@ -338,6 +449,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tint no_update_server_info = 0;\n \tstruct pack_objects_args po_args = {NULL};\n+\tint geometric_factor = 0;\n \n \tstruct option builtin_repack_options[] = {\n \t\tOPT_BIT('a', NULL, &pack_everything,\n@@ -378,6 +490,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t\tN_(\"repack objects in packs marked with .keep\")),\n \t\tOPT_STRING_LIST(0, \"keep-pack\", &keep_pack_list, N_(\"name\"),\n \t\t\t\tN_(\"do not repack this pack\")),\n+\t\tOPT_INTEGER('g', \"geometric\", &geometric_factor,\n+\t\t\t    N_(\"find a geometric progression with factor <N>\")),\n \t\tOPT_END()\n \t};\n \n@@ -404,6 +518,11 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (write_bitmaps && !(pack_everything & ALL_INTO_ONE))\n \t\tdie(_(incremental_bitmap_conflict_error));\n \n+\tif (geometric_factor) {\n+\t\tinit_pack_geometry(&geometry);\n+\t\tsplit_pack_geometry(geometry, geometric_factor);\n+\t}\n+\n \tpackdir = mkpathdup(\"%s/pack\", get_object_directory());\n \tpacktmp = mkpathdup(\"%s/.tmp-%d-pack\", packdir, (int)getpid());\n \n@@ -439,17 +558,41 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\thandle_loose_and_reachable(&cmd, unpack_unreachable,\n \t\t\t\t\t\t   pack_everything,\n \t\t\t\t\t\t   keep_unreachable);\n+\t} else if (geometry) {\n+\t\tstrvec_push(&cmd.args, \"--keep-pack-stdin\");\n+\t\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n+\t\tstrvec_push(&cmd.args, \"--assume-kept-packs-closed\");\n+\t\tif (delete_redundant)\n+\t\t\thandle_loose_and_reachable(&cmd, unpack_unreachable,\n+\t\t\t\t\t\t   pack_everything,\n+\t\t\t\t\t\t   keep_unreachable);\n \t} else {\n \t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t\tstrvec_push(&cmd.args, \"--incremental\");\n \t}\n \n-\tcmd.no_stdin = 1;\n+\tif (geometry)\n+\t\tcmd.in = -1;\n+\telse\n+\t\tcmd.no_stdin = 1;\n \n \tret = start_command(&cmd);\n \tif (ret)\n \t\treturn ret;\n \n+\tif (geometry) {\n+\t\tFILE *in = xfdopen(cmd.in, \"w\");\n+\t\t/*\n+\t\t * Tell 'git pack-objects' to avoid tampering with the structure\n+\t\t * with the packs that already form a geometric progression.\n+\t\t *\n+\t\t * Everything else will get picked up by the reachability walk.\n+\t\t */\n+\t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n+\t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n+\t\tfclose(in);\n+\t}\n+\n \tout = xfdopen(cmd.out, \"r\");\n \twhile (strbuf_getline_lf(&line, out) != EOF) {\n \t\tif (line.len != the_hash_algo->hexsz)\n@@ -517,6 +660,25 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tif (!string_list_has_string(&names, sha1))\n \t\t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n+\n+\t\tif (geometry) {\n+\t\t\tstruct strbuf buf = STRBUF_INIT;\n+\n+\t\t\tuint32_t i;\n+\t\t\tfor (i = 0; i < geometry->split; i++) {\n+\t\t\t\tstruct packed_git *p = geometry->pack[i];\n+\t\t\t\tif (string_list_has_string(&names,\n+\t\t\t\t\t\t\t   hash_to_hex(p->hash)))\n+\t\t\t\t\tcontinue;\n+\n+\t\t\t\tstrbuf_reset(&buf);\n+\t\t\t\tstrbuf_addstr(&buf, pack_basename(p));\n+\t\t\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\n+\t\t\t\tremove_redundant_pack(packdir, buf.buf);\n+\t\t\t}\n+\t\t\tstrbuf_release(&buf);\n+\t\t}\n \t\tif (!po_args.quiet && isatty(2))\n \t\t\topts |= PRUNE_PACKED_VERBOSE;\n \t\tprune_packed_objects(opts);\n@@ -538,6 +700,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstring_list_clear(&names, 0);\n \tstring_list_clear(&rollback, 0);\n \tstring_list_clear(&existing_packs, 0);\n+\tclear_pack_geometry(geometry);\n \tstrbuf_release(&line);\n \n \treturn 0;\ndiff --git a/t/t7703-repack-geometric.sh b/t/t7703-repack-geometric.sh\nnew file mode 100755\nindex 0000000000..39cef892f8\n--- /dev/null\n+++ b/t/t7703-repack-geometric.sh\n@@ -0,0 +1,81 @@\n+#!/bin/sh\n+\n+test_description='git repack --geometric works correctly'\n+\n+. ./test-lib.sh\n+\n+GIT_TEST_MULTI_PACK_INDEX=0\n+\n+objdir=.git/objects\n+midx=$objdir/pack/multi-pack-index\n+\n+test_expect_success '--geometric with an intact progression' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# These packs already form a geometric progression.\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 2 && # 6 objects\n+\t\ttest_commit_bulk --start=4 4 && # 12 objects\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >expect &&\n+\t\tGIT_TEST_MULTI_PACK_BITMAP=0 git repack --geometric 2 -d &&\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >actual &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success '--geometric with small-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >small &&\n+\t\ttest_commit_bulk --start=3 4 && # 12 objects\n+\t\ttest_commit_bulk --start=7 8 && # 24 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tGIT_TEST_MULTI_PACK_BITMAP=0 git repack --geometric 2 -d &&\n+\n+\t\t# Three packs in total; two of the existing large ones, and one\n+\t\t# new one.\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\ttest_line_count = 3 after &&\n+\t\tcomm -3 small before | tr -d \"\\t\" >large &&\n+\t\tgrep -qFf large after\n+\t)\n+'\n+\n+test_expect_success '--geometric with small- and large-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# size(small1) + size(small2) > size(medium) / 2\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 3 && # 7 objects\n+\t\ttest_commit_bulk --start=6 9 && # 27 objects &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tGIT_TEST_MULTI_PACK_BITMAP=0 git repack --geometric 2 -d &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\tcomm -12 before after >untouched &&\n+\n+\t\t# Two packs in total; the largest pack from before running \"git\n+\t\t# repack\", and one new one.\n+\t\ttest_line_count = 1 untouched &&\n+\t\ttest_line_count = 2 after\n+\t)\n+'\n+\n+test_done\n-- \n2.30.0.138.g6d7191ea01\n"},{"id":"414759","messageId":"2da42e9ca26c9ef914b8b044047d505f00a27e20.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:08Z","receivedAt":"2021-01-19T23:29:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Teach pack-objects an option to imply the revision machinery's new\n'--no-kept-objects' option when doing a reachability traversal.\n\nWhen '--assume-kept-packs-closed' is given as an argument to\npack-objects, it behaves differently (i.e., passes different options to\nthe ensuing revision walk) depending on whether or not other arguments\nare passed:\n\n  - If the caller also specifies a '--keep-pack' argument (to mark a\n    pack as kept in-core), then assume that this combination means to\n    stop traversal only at in-core packs.\n\n  - If instead the caller passes '--honor-pack-keep', then assume that\n    the caller wants to stop traversal only at packs with a\n    corresponding .keep file (consistent with the original meaning which\n    only refers to packs with a .keep file).\n\n  - If both '--keep-pack' and '--honor-pack-keep' are passed, then\n    assume the caller wants to stop traversal at either kind of kept\n    pack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt | 11 ++++++\n builtin/pack-objects.c             | 13 +++++++\n t/t6114-keep-packs.sh              | 59 ++++++++++++++++++++++++++++++\n 3 files changed, 83 insertions(+)\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex 54d715ead1..cbe08e7415 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -135,6 +135,17 @@ depth is 4095.\n \tleading directory (e.g. `pack-123.pack`). The option could be\n \tspecified multiple times to keep multiple packs.\n \n+--assume-kept-packs-closed::\n+\tThis flag causes `git rev-list` to halt the object traversal\n+\twhen it encounters an object found in a kept pack. This is\n+\tdissimilar to `--honor-pack-keep`, which only prunes unwanted\n+\tresults after the full traversal is completed.\n++\n+Without any `--keep-pack=<pack-name>` arguments, only packs with an\n+on-disk `*.keep` files are used when considering when to halt the\n+traversal. If other packs are artificially marked as \"kept\" with\n+`--keep-pack`, then those are considered as well.\n+\n --incremental::\n \tThis flag causes an object already in a pack to be ignored\n \teven if it would have otherwise been packed.\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 2a00358f34..a5dcd66f52 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -78,6 +78,7 @@ static int have_non_local_packs;\n static int incremental;\n static int ignore_packed_keep_on_disk;\n static int ignore_packed_keep_in_core;\n+static int assume_kept_packs_closed;\n static int allow_ofs_delta;\n static struct pack_idx_option pack_idx_opts;\n static const char *base_name;\n@@ -3542,6 +3543,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t N_(\"create packs suitable for shallow fetches\")),\n \t\tOPT_BOOL(0, \"honor-pack-keep\", &ignore_packed_keep_on_disk,\n \t\t\t N_(\"ignore packs that have companion .keep file\")),\n+\t\tOPT_BOOL(0, \"assume-kept-packs-closed\", &assume_kept_packs_closed,\n+\t\t\t N_(\"assume the union of kept packs is closed under reachability\")),\n \t\tOPT_STRING_LIST(0, \"keep-pack\", &keep_pack_list, N_(\"name\"),\n \t\t\t\tN_(\"ignore this pack\")),\n \t\tOPT_INTEGER(0, \"compression\", &pack_compression_level,\n@@ -3631,6 +3634,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tuse_internal_rev_list = 1;\n \t\tstrvec_push(&rp, \"--unpacked\");\n \t}\n+\tif (assume_kept_packs_closed)\n+\t\tuse_internal_rev_list = 1;\n \n \tif (exclude_promisor_objects) {\n \t\tuse_internal_rev_list = 1;\n@@ -3711,6 +3716,14 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tif (!p) /* no keep-able packs found */\n \t\t\tignore_packed_keep_on_disk = 0;\n \t}\n+\tif (assume_kept_packs_closed) {\n+\t\tif (ignore_packed_keep_on_disk && ignore_packed_keep_in_core)\n+\t\t\tstrvec_push(&rp, \"--no-kept-objects\");\n+\t\telse if (ignore_packed_keep_on_disk)\n+\t\t\tstrvec_push(&rp, \"--no-kept-objects=on-disk\");\n+\t\telse if (ignore_packed_keep_in_core)\n+\t\t\tstrvec_push(&rp, \"--no-kept-objects=in-core\");\n+\t}\n \tif (local) {\n \t\t/*\n \t\t * unlike ignore_packed_keep_on_disk above, we do not\ndiff --git a/t/t6114-keep-packs.sh b/t/t6114-keep-packs.sh\nindex 9239d8aa46..0861305a04 100755\n--- a/t/t6114-keep-packs.sh\n+++ b/t/t6114-keep-packs.sh\n@@ -66,4 +66,63 @@ test_expect_success '--no-kept-objects excludes kept non-MIDX object' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success '--no-kept-objects can respect only in-core keep packs' '\n+\ttest_when_finished \"rm -fr actual-*.idx actual-*.pack\" &&\n+\t(\n+\t\tgit rev-list --objects --no-object-names packed..kept &&\n+\t\tgit rev-list --objects --no-object-names loose\n+\t) | sort >expect &&\n+\n+\tgit pack-objects \\\n+\t  --assume-kept-packs-closed \\\n+\t  --keep-pack=pack-$MISC_PACK.pack \\\n+\t  --all actual </dev/null &&\n+\tidx_objects actual-*.idx >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'setup additional --no-kept-objects tests' '\n+\ttest_commit additional &&\n+\n+\tADDITIONAL_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/additional\n+\t^refs/tags/kept\n+\tEOF\n+\t)\n+'\n+\n+test_expect_success '--no-kept-objects can respect only on-disk keep packs' '\n+\ttest_when_finished \"rm -fr actual-*.idx actual-*.pack\" &&\n+\t(\n+\t\tgit rev-list --objects --no-object-names kept..additional &&\n+\t\tgit rev-list --objects --no-object-names packed\n+\t) | sort >expect &&\n+\n+\tgit pack-objects \\\n+\t  --assume-kept-packs-closed \\\n+\t  --honor-pack-keep \\\n+\t  --all actual </dev/null &&\n+\tidx_objects actual-*.idx >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success '--no-kept-objects can respect mixed kept packs' '\n+\ttest_when_finished \"rm -fr actual-*.idx actual-*.pack\" &&\n+\t(\n+\t\tgit rev-list --objects --no-object-names kept..additional &&\n+\t\tgit rev-list --objects --no-object-names loose\n+\t) | sort >expect &&\n+\n+\tgit pack-objects \\\n+\t  --assume-kept-packs-closed \\\n+\t  --honor-pack-keep \\\n+\t  --keep-pack=pack-$MISC_PACK.pack \\\n+\t  --all actual </dev/null &&\n+\tidx_objects actual-*.idx >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414760","messageId":"a808fbdf31afc9ad9ba0ab27ce889e5a2d1a01ae.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 09/10] builtin/repack.c: extract loose object handling","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:33Z","receivedAt":"2021-01-19T23:29:39Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"'git repack -g' will have to learn about unreachable loose objects that\nneed to be removed in a separate path from the existing checks.\n\nExtract that check into a function so it can be called from multiple\nplaces.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c | 41 +++++++++++++++++++++++++----------------\n 1 file changed, 25 insertions(+), 16 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 279be11a16..664863111b 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -298,6 +298,27 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n #define ALL_INTO_ONE 1\n #define LOOSEN_UNREACHABLE 2\n \n+static void handle_loose_and_reachable(struct child_process *cmd,\n+\t\t\t\t       const char *unpack_unreachable,\n+\t\t\t\t       int pack_everything,\n+\t\t\t\t       int keep_unreachable)\n+{\n+\tif (unpack_unreachable) {\n+\t\tstrvec_pushf(&cmd->args,\n+\t\t\t     \"--unpack-unreachable=%s\",\n+\t\t\t     unpack_unreachable);\n+\t\tstrvec_push(&cmd->env_array, \"GIT_REF_PARANOIA=1\");\n+\t} else if (pack_everything & LOOSEN_UNREACHABLE) {\n+\t\tstrvec_push(&cmd->args,\n+\t\t\t    \"--unpack-unreachable\");\n+\t} else if (keep_unreachable) {\n+\t\tstrvec_push(&cmd->args, \"--keep-unreachable\");\n+\t\tstrvec_push(&cmd->args, \"--pack-loose-unreachable\");\n+\t} else {\n+\t\tstrvec_push(&cmd->env_array, \"GIT_REF_PARANOIA=1\");\n+\t}\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -414,22 +435,10 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \n \t\trepack_promisor_objects(&po_args, &names);\n \n-\t\tif (existing_packs.nr && delete_redundant) {\n-\t\t\tif (unpack_unreachable) {\n-\t\t\t\tstrvec_pushf(&cmd.args,\n-\t\t\t\t\t     \"--unpack-unreachable=%s\",\n-\t\t\t\t\t     unpack_unreachable);\n-\t\t\t\tstrvec_push(&cmd.env_array, \"GIT_REF_PARANOIA=1\");\n-\t\t\t} else if (pack_everything & LOOSEN_UNREACHABLE) {\n-\t\t\t\tstrvec_push(&cmd.args,\n-\t\t\t\t\t    \"--unpack-unreachable\");\n-\t\t\t} else if (keep_unreachable) {\n-\t\t\t\tstrvec_push(&cmd.args, \"--keep-unreachable\");\n-\t\t\t\tstrvec_push(&cmd.args, \"--pack-loose-unreachable\");\n-\t\t\t} else {\n-\t\t\t\tstrvec_push(&cmd.env_array, \"GIT_REF_PARANOIA=1\");\n-\t\t\t}\n-\t\t}\n+\t\tif (existing_packs.nr && delete_redundant)\n+\t\t\thandle_loose_and_reachable(&cmd, unpack_unreachable,\n+\t\t\t\t\t\t   pack_everything,\n+\t\t\t\t\t\t   keep_unreachable);\n \t} else {\n \t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t\tstrvec_push(&cmd.args, \"--incremental\");\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414761","messageId":"26b46dff15ce89f8ccab3866a0e230d99c538697.1611098616.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH 04/10] p5303: add missing &&-chains","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-19T23:24:13Z","receivedAt":"2021-01-19T23:29:44Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nThese are in a helper function, so the usual chain-lint doesn't notice\nthem. This function is still not perfect, as it has some git invocations\non the left-hand-side of the pipe, but it's primary purpose is timing,\nnot finding bugs or correctness issues.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex f4c2ab0584..277d22ec4b 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -24,11 +24,11 @@ repack_into_n () {\n \tsed -n '1~5p' |\n \thead -n \"$1\" |\n \tperl -e 'print reverse <>' \\\n-\t>pushes\n+\t>pushes &&\n \n \t# create base packfile\n \thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack\n+\tgit pack-objects --delta-base-offset --revs staging/pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n-- \n2.30.0.138.g6d7191ea01\n\n"},{"id":"414794","messageId":"YAg/WU01bvfsMxgX@nand.local","threadId":"55012","inReplyTo":"98c65017-8c22-a21f-0e86-a15d91bd7f70@gmail.com","subject":"Re: [PATCH 09/10] builtin/repack.c: extract loose object handling","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-20T14:34:01Z","receivedAt":"2021-01-20T14:35:54Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 20, 2021 at 08:59:48AM -0500, Derrick Stolee wrote:\n> On 1/19/2021 6:24 PM, Taylor Blau wrote:\n> > 'git repack -g' will have to learn about unreachable loose objects that\n>\n> This reference to the '-g' option is one patch too early. Perhaps\n> say\n>\n>   An upcoming patch will introduce geometric repacking. This will\n>   require removing unreachable loose objects in a separate path\n>   from the existing checks.\n>\n> or similar?\n\nMmm. I had imagined that this would be read either in the context of\nthis series, or by someone in the future long after 'git repack -g' had\nbeen introduced.\n\nI could see that it's confusing, though, and I do agree your wording\nmakes clearer that the option doesn't exist yet.\n\nI'm happy to send a replacement or reroll if you feel strongly, but in\neither case I'll wait for a little more review first.\n\nThanks,\nTaylor\n"},{"id":"414796","messageId":"YAhAWUw6Hzs9nG8Z@nand.local","threadId":"55012","inReplyTo":"607e7ebd-240d-f2dc-42ef-1d5a5a0b7f51@gmail.com","subject":"Re: [PATCH 01/10] packfile: introduce 'find_kept_pack_entry()'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-20T14:38:17Z","receivedAt":"2021-01-20T14:51:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 20, 2021 at 08:40:22AM -0500, Derrick Stolee wrote:\n> On 1/19/2021 6:24 PM, Taylor Blau wrote:\n> >  \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n> > -\t\tif (fill_midx_entry(r, oid, e, m))\n> > +\t\tif (!(fill_midx_entry(r, oid, e, m)))\n>\n> nit: we don't need extra parens around fill_midx_entry().\n\nYep. I checked whether we should have written this as \"if\n(fill_midx_entry(...) < 0)\", but fill_midx_entry returns a positive\nnumber on error, so checking \"!fill_midx_entry\" is certainly what we\nshould be doing.\n\n> > -\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n> > -\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n> > -\t\t\treturn 1;\n> > +\t\tif (p->multi_pack_index && !kept_only) {\n> > +\t\t\t/*\n> > +\t\t\t * If this pack is covered by the MIDX, we'd have found\n> > +\t\t\t * the object already in the loop above if it was here,\n> > +\t\t\t * so don't bother looking.\n> > +\t\t\t *\n> > +\t\t\t * The exception is if we are looking only at kept\n> > +\t\t\t * packs. An object can be present in two packs covered\n> > +\t\t\t * by the MIDX, one kept and one not-kept. And as the\n> > +\t\t\t * MIDX points to only one copy of each object, it might\n> > +\t\t\t * have returned only the non-kept version above. We\n> > +\t\t\t * have to check again to be thorough.\n> > +\t\t\t */\n> > +\t\t\tcontinue;\n> > +\t\t}\n> > +\t\tif (!kept_only ||\n> > +\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n> > +\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n> > +\t\t\tif (fill_pack_entry(oid, e, p)) {\n> > +\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n> > +\t\t\t\treturn 1;\n> > +\t\t\t}\n>\n> Here is the meat of your patch. The comment helps a lot.\n>\n> This might have been easier if the MIDX had preferred kept packs\n> over non-kept packs (before sorting by modified time). Perhaps\n> the MIDX could get an extra field to say \"I preferred kept packs\"\n> which would let us trust the MIDX return here without the pack\n> loop.\n>\n> (Note: we can't just change the MIDX selection and then start\n> trusting all MIDXs to have the right tie-breakers because of\n> existing files in the wild.)\n\nYeah, that is what makes it tricky. Changing the code isn't so hard: a\nnew field that we check and do one of two things when we're breaking\nties.\n\nBut I think the cognitive load is high, and I'm not sure that the\nbenefit (skipping another linear pass through non-MIDX'd packs when\nlooking up an object in kept packs only _and_ that object is duplicated)\nis worth the extra hassle with the MIDX code.\n\nAll of that said, I do think that it's worth revisiting this and giving\nit some more thought after multi-pack bitmaps to see whether we feel the\nsame or not.\n\nThanks,\nTaylor\n"},{"id":"414803","messageId":"ae13318c-f894-6c9a-414c-e7911abdef76@gmail.com","threadId":"55012","inReplyTo":"YAg/WU01bvfsMxgX@nand.local","subject":"Re: [PATCH 09/10] builtin/repack.c: extract loose object handling","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-01-20T15:51:02Z","receivedAt":"2021-01-20T15:53:55Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/20/2021 9:34 AM, Taylor Blau wrote:\n> On Wed, Jan 20, 2021 at 08:59:48AM -0500, Derrick Stolee wrote:\n>> On 1/19/2021 6:24 PM, Taylor Blau wrote:\n>>> 'git repack -g' will have to learn about unreachable loose objects that\n>>\n>> This reference to the '-g' option is one patch too early. Perhaps\n>> say\n>>\n>>   An upcoming patch will introduce geometric repacking. This will\n>>   require removing unreachable loose objects in a separate path\n>>   from the existing checks.\n>>\n>> or similar?\n> \n> Mmm. I had imagined that this would be read either in the context of\n> this series, or by someone in the future long after 'git repack -g' had\n> been introduced.\n> \n> I could see that it's confusing, though, and I do agree your wording\n> makes clearer that the option doesn't exist yet.\n> \n> I'm happy to send a replacement or reroll if you feel strongly, but in\n> either case I'll wait for a little more review first.\n\nDefinitely don't rush a re-roll for my nit-picks.\n\nThanks,\n-Stolee\n"},{"id":"414854","messageId":"f477fd3d-1725-9792-7907-780afe91e37b@gmail.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"Re: [PATCH 00/10] repack: support repacking into a geometric sequence","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-01-20T14:05:59Z","receivedAt":"2021-01-20T20:08:09Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/19/2021 6:23 PM, Taylor Blau wrote:\n> This series introduces a new mode of 'git repack' where (instead of packing just\n> loose objects or packing everything together into one pack), the set of packs\n> left forms a geometric progression by object count.\n...\n> Thanks in advance for your review.\n\nI had the pleasure of reading an early version of this series, but it's\nbeen a while. Upon a fresh reading, I only had nitpicks. Otherwise, this\nLGTM.\n\nI encourage other reviewers to read patch 10 carefully, as that is the\nmost math-heavy of all of them.\n\nThanks,\n-Stolee\n"},{"id":"414858","messageId":"98c65017-8c22-a21f-0e86-a15d91bd7f70@gmail.com","threadId":"55012","inReplyTo":"a808fbdf31afc9ad9ba0ab27ce889e5a2d1a01ae.1611098616.git.me@ttaylorr.com","subject":"Re: [PATCH 09/10] builtin/repack.c: extract loose object handling","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-01-20T13:59:48Z","receivedAt":"2021-01-20T20:17:31Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/19/2021 6:24 PM, Taylor Blau wrote:\n> 'git repack -g' will have to learn about unreachable loose objects that\n\nThis reference to the '-g' option is one patch too early. Perhaps\nsay\n\n  An upcoming patch will introduce geometric repacking. This will\n  require removing unreachable loose objects in a separate path\n  from the existing checks.\n\nor similar?\n\nThanks,\n-Stolee\n\n"},{"id":"414861","messageId":"607e7ebd-240d-f2dc-42ef-1d5a5a0b7f51@gmail.com","threadId":"55012","inReplyTo":"dc7fa4c7a61f657e779e10385d3e8076d6dac36c.1611098616.git.me@ttaylorr.com","subject":"Re: [PATCH 01/10] packfile: introduce 'find_kept_pack_entry()'","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-01-20T13:40:22Z","receivedAt":"2021-01-20T21:36:34Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/19/2021 6:24 PM, Taylor Blau wrote:\n>  \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n> -\t\tif (fill_midx_entry(r, oid, e, m))\n> +\t\tif (!(fill_midx_entry(r, oid, e, m)))\n\nnit: we don't need extra parens around fill_midx_entry().\n\n> -\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n> -\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n> -\t\t\treturn 1;\n> +\t\tif (p->multi_pack_index && !kept_only) {\n> +\t\t\t/*\n> +\t\t\t * If this pack is covered by the MIDX, we'd have found\n> +\t\t\t * the object already in the loop above if it was here,\n> +\t\t\t * so don't bother looking.\n> +\t\t\t *\n> +\t\t\t * The exception is if we are looking only at kept\n> +\t\t\t * packs. An object can be present in two packs covered\n> +\t\t\t * by the MIDX, one kept and one not-kept. And as the\n> +\t\t\t * MIDX points to only one copy of each object, it might\n> +\t\t\t * have returned only the non-kept version above. We\n> +\t\t\t * have to check again to be thorough.\n> +\t\t\t */\n> +\t\t\tcontinue;\n> +\t\t}\n> +\t\tif (!kept_only ||\n> +\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n> +\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n> +\t\t\tif (fill_pack_entry(oid, e, p)) {\n> +\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n> +\t\t\t\treturn 1;\n> +\t\t\t}\n\nHere is the meat of your patch. The comment helps a lot.\n\nThis might have been easier if the MIDX had preferred kept packs\nover non-kept packs (before sorting by modified time). Perhaps\nthe MIDX could get an extra field to say \"I preferred kept packs\"\nwhich would let us trust the MIDX return here without the pack\nloop.\n\n(Note: we can't just change the MIDX selection and then start\ntrusting all MIDXs to have the right tie-breakers because of\nexisting files in the wild.)\n\nThanks,\n-Stolee\n"},{"id":"414882","messageId":"xmqqk0s7c5rm.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"98c65017-8c22-a21f-0e86-a15d91bd7f70@gmail.com","subject":"Re: [PATCH 09/10] builtin/repack.c: extract loose object handling","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-21T03:45:17Z","receivedAt":"2021-01-21T03:51:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> On 1/19/2021 6:24 PM, Taylor Blau wrote:\n>> 'git repack -g' will have to learn about unreachable loose objects that\n>\n> This reference to the '-g' option is one patch too early. Perhaps\n> say\n>\n>   An upcoming patch will introduce geometric repacking. This will\n>   require removing unreachable loose objects in a separate path\n>   from the existing checks.\n>\n> or similar?\n\nYeah, sounds like a trivially obvious improvement to me.  \n\nIt does not matter to reviewers who are very well aware that the\nseries is about adding \"repack -g\", but it may end up being\nconfusing when somebody tries to see what commit the feature was\nadded later when the help from the cover letter is not available.\n"},{"id":"415557","messageId":"xmqqwnvwtqu1.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"dc7fa4c7a61f657e779e10385d3e8076d6dac36c.1611098616.git.me@ttaylorr.com","subject":"Re: [PATCH 01/10] packfile: introduce 'find_kept_pack_entry()'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T02:33:10Z","receivedAt":"2021-01-29T02:34:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> Future callers will want a function to fill a 'struct pack_entry' for a\n> given object id but _only_ from its position in any kept pack(s). They\n> could accomplish this by calling 'find_pack_entry()' and checking\n> whether the found pack is kept or not, but this is insufficient, since\n> there may be duplicate objects (and the mru cache makes it unpredictable\n> which variant we'll get).\n\nI wonder if we eventually need a callback interface to walk _all_\npack entries for a given object, so that \"I am only interested in\ninstances in kept packs\" will be under total control of the callers.\nAs it stands, it is \"just grab any one that is in a kept pack, any\none of them is fine\", which is almost just of as narrow utility as\nthe original's \"just grab the first one---any one of them is fine\",\nthe latter of which is \"insufficient\" as the log message says.\n\nBut this (in the context of the remainder of the series) might be\nsufficient, at least for now.\n\n> Teach this new function to treat the two different kinds of kept packs\n> (on disk ones with .keep files, as well as in-core ones which are set by\n> manually poking the 'pack_keep_in_core' bit) separately. This will\n> become important for callers that only want to respect a certain kind of\n> kept pack.\n\nOr maybe not ;-)\n\nIf there are notable relationship between on-disk and in-core kept\npacks (e.g. \"the set of on-disk kept packs is a subset of in-core\nkept packs\", \"usually on-disk kept packs get in-core kept bit upon\ntheir packed_git instances are populated, but we can drop the bit at\nruntime, so on-disk and in-core are pretty much independent and\nthere is no notable relationship\"), it must be explained upfront to\nhelp the reader form a sensible world view.\n\n> Introduce 'find_kept_pack_entry()' which behaves like\n> 'find_pack_entry()', except that it skips over packs which are not\n> marked kept. Callers will be added in subsequent patches.\n>\n> Co-authored-by: Jeff King <peff@peff.net>\n> Signed-off-by: Jeff King <peff@peff.net>\n> Signed-off-by: Taylor Blau <me@ttaylorr.com>\n\nFun.\n\nThanks.\n"},{"id":"415559","messageId":"xmqqo8h8tp4j.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"4184529648abe7451b5c7b772493df8c067cec82.1611098616.git.me@ttaylorr.com","subject":"Re: [PATCH 02/10] revision: learn '--no-kept-objects'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T03:10:04Z","receivedAt":"2021-01-29T03:11:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> Some callers want to perform a reachability traversal that terminates\n> when an object is found in a kept pack. The closest existing option is\n> '--honor-pack-keep', but this isn't quite what we want. Instead of\n> halting the traversal midway through, a full traversal is always\n> performed, and the results are only trimmed afterwords.\n\nTrue.  \n\nIs there a reason to keep both kinds?  It is obvious that stopping\ntraversal once we hit a kept pack would be more time and space\nefficient (I presume that the reason why .kept pack matters is\nbecause we are repacking everything else) to enumerate the objects\nthat need to be repacked than traversing all the way and filtering\nout objects that appear in .kept packs, but would there be some\ncorrectness implications to replace the existing use of\n\"--honor-pack-keep\" with \"--no-kept-objects=on-disk\"?  \n\nWhat it means to be excluded by the former is quite clear: any\nobject that appears in a kept pack, whether another copy of it\nappears elsewhere, is excluded from getting enumerated for\nrepacking.  It is quite unclear what it means to enumerate objects\nwith \"--no-kept-objects\".  It is clear from the implementation side\nof the thing (stop traversal at objects that appear in any kept\npack), but it is totally unclear what such a meaning defined\noperationally affects the resulting enumeration.  We know that the\nenumerated objects do not appear in any of the kept pack, but it\ndoes not mean all objects that are reachable/in-use that are not in\nany kept packs are enumerated.\n\n> diff --git a/Documentation/rev-list-options.txt b/Documentation/rev-list-options.txt\n> index 002379056a..817419d552 100644\n> --- a/Documentation/rev-list-options.txt\n> +++ b/Documentation/rev-list-options.txt\n> @@ -856,6 +856,13 @@ ifdef::git-rev-list[]\n>  \tOnly useful with `--objects`; print the object IDs that are not\n>  \tin packs.\n>  \n> +--no-kept-objects[=<kind>]::\n> +\tHalts the traversal as soon as an object in a kept pack is\n> +\tfound. If `<kind>` is `on-disk`, only packs with a corresponding\n> +\t`*.keep` file are ignored. If `<kind>` is `in-core`, only packs\n> +\twith their in-core kept state set are ignored. Otherwise, both\n> +\tkinds of kept packs are ignored.\n\nIs it explained anywhere how \"in-core kept state\" is bootstrapped,\nmodified and maintained?\n\nThe patch to C-part itself is a trivially correct implementation of\n\"stop at an object that can be found in a kept pack\", and there is\nno comment, but it is not clear to me what we want to achieve by\nthis.  Is the underlying assumption that no objects in .kept pack\nwould refer to outside world, either loose or packs that are not\nkept?  How are we guaranteeing it?\n\n"},{"id":"415560","messageId":"xmqqk0rwtom2.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"2da42e9ca26c9ef914b8b044047d505f00a27e20.1611098616.git.me@ttaylorr.com","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T03:21:09Z","receivedAt":"2021-01-29T03:22:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> Teach pack-objects an option to imply the revision machinery's new\n> '--no-kept-objects' option when doing a reachability traversal.\n>\n> When '--assume-kept-packs-closed' is given as an argument to\n> pack-objects, it behaves differently (i.e., passes different options to\n> the ensuing revision walk) depending on whether or not other arguments\n> are passed:\n>\n>   - If the caller also specifies a '--keep-pack' argument (to mark a\n>     pack as kept in-core), then assume that this combination means to\n>     stop traversal only at in-core packs.\n>\n>   - If instead the caller passes '--honor-pack-keep', then assume that\n>     the caller wants to stop traversal only at packs with a\n>     corresponding .keep file (consistent with the original meaning which\n>     only refers to packs with a .keep file).\n>\n>   - If both '--keep-pack' and '--honor-pack-keep' are passed, then\n>     assume the caller wants to stop traversal at either kind of kept\n>     pack.\n\nIf there is an out-of-band guarantee that .kept packs won't refer to\noutside world, then we can obtain identical results to what existing\n--honor-pack-keep (which traverses everything and then filteres out\nwhat is in .keep pack) does by just stopping traversal when we see\nan object that is found in a .keep pack.  OK, I guess that it\nanswers the correctness question I asked about [02/10].\n\nIt still is curious how we can safely \"assume\", but presumably we\nwill see how in a patch that appears later in the series.\n\nHow \"closed\" are these kept packs supposed to be?  When there are\ntwo .keep packs, should objects in each of the packs never refer to\noutside their own pack, or is it OK for objects in one kept pack to\nrefer to another object in the other kept pack?  Readers and those\nwho want to understand and extend this code in the future would need\nto know what definition of \"closed\" you are using here.\n\nThanks.\n"},{"id":"415561","messageId":"xmqqft2ktnpj.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"b3b2574d4d9d10f226b52d81fe0e6bf1f761504e.1611098616.git.me@ttaylorr.com","subject":"Re: [PATCH 05/10] p5303: measure time to repack with keep","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T03:40:40Z","receivedAt":"2021-01-29T03:41:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> From: Jeff King <peff@peff.net>\n\nNot a fault of this series at all, but before the precontext of the\nfirst hunk, there is  \n\n\n> diff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\n> index 277d22ec4b..85b077b72b 100755\n> --- a/t/perf/p5303-many-packs.sh\n> +++ b/t/perf/p5303-many-packs.sh\n> @@ -27,8 +27,11 @@ repack_into_n () {\n\nthis construct:\n\n\t... |\n\tsed -n '1~5p' |\n\thead -n \"$1\" |\n        ...\n\nwhich is a GNUism.  Peff often says that very small population\nactually run our perf suite, and this seems to corroborate the\nconjecture.\n\n>  \t>pushes &&\n>  \n>  \t# create base packfile\n> -\thead -n 1 pushes |\n> -\tgit pack-objects --delta-base-offset --revs staging/pack &&\n> +\tbase_pack=$(\n> +\t\thead -n 1 pushes |\n> +\t\tgit pack-objects --delta-base-offset --revs staging/pack\n> +\t) &&\n> +\ttest_export base_pack &&\n>  \n>  \t# and then incrementals between each pair of commits\n>  \tlast= &&\n> @@ -87,6 +90,15 @@ do\n>  \t\t  --reflog --indexed-objects --delta-base-offset \\\n>  \t\t  --stdout </dev/null >/dev/null\n>  \t'\n> +\n> +\ttest_perf \"repack with keep ($nr_packs)\" '\n> +\t\tgit pack-objects --keep-true-parents \\\n> +\t\t  --honor-pack-keep --assume-kept-packs-closed \\\n> +\t\t  --keep-pack=pack-$base_pack.pack \\\n> +\t\t  --non-empty --all \\\n> +\t\t  --reflog --indexed-objects --delta-base-offset \\\n> +\t\t  --stdout </dev/null >/dev/null\n> +\t'\n>  done\n>  \n>  # Measure pack loading with 10,000 packs.\n"},{"id":"415581","messageId":"YBRWPnTxeFPBft7y@nand.local","threadId":"55012","inReplyTo":"xmqqwnvwtqu1.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 01/10] packfile: introduce 'find_kept_pack_entry()'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-29T18:38:54Z","receivedAt":"2021-01-29T18:39:41Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Jan 28, 2021 at 06:33:10PM -0800, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > Future callers will want a function to fill a 'struct pack_entry' for a\n> > given object id but _only_ from its position in any kept pack(s). They\n> > could accomplish this by calling 'find_pack_entry()' and checking\n> > whether the found pack is kept or not, but this is insufficient, since\n> > there may be duplicate objects (and the mru cache makes it unpredictable\n> > which variant we'll get).\n>\n> I wonder if we eventually need a callback interface to walk _all_\n> pack entries for a given object, so that \"I am only interested in\n> instances in kept packs\" will be under total control of the callers.\n> As it stands, it is \"just grab any one that is in a kept pack, any\n> one of them is fine\", which is almost just of as narrow utility as\n> the original's \"just grab the first one---any one of them is fine\",\n> the latter of which is \"insufficient\" as the log message says.\n>\n> But this (in the context of the remainder of the series) might be\n> sufficient, at least for now.\n\nAs you note, it's more about \"can I find this object in any kept pack\n(of a certain kind)\" versus, \"show me this object in a pack\" (and hope\nthat if it appears in a kept pack, that that's the copy that is picked).\n\n> > Teach this new function to treat the two different kinds of kept packs\n> > (on disk ones with .keep files, as well as in-core ones which are set by\n> > manually poking the 'pack_keep_in_core' bit) separately. This will\n> > become important for callers that only want to respect a certain kind of\n> > kept pack.\n>\n> Or maybe not ;-)\n\n:-). The difference here is that we will only want to stop the traversal\nat packs which are considered to be stable from the perspective of a\ngeometric repack.\n\nWe mark those packs as \"stable\" by setting their in-core kept bit, but\nwe don't write \".keep\" files (which would make them on-disk kept). The\nlatter is up to the user, not us.\n\n> If there are notable relationship between on-disk and in-core kept\n> packs (e.g. \"the set of on-disk kept packs is a subset of in-core\n> kept packs\", \"usually on-disk kept packs get in-core kept bit upon\n> their packed_git instances are populated, but we can drop the bit at\n> runtime, so on-disk and in-core are pretty much independent and\n> there is no notable relationship\"), it must be explained upfront to\n> help the reader form a sensible world view.\n\nUnfortunately, I don't think that there is a sensible world-view here\nto be formed. Honestly, the distinction between .keep packs and in-core\nkept packs is incredibly narrow, and I find our separate handling of\nthem awkward and error-prone.\n\nBut, it is sort of what you'd want here (i.e., a way to mark all objects\nin a pack as ignored without actually writing the physical file that\nsays \"ignore all objects in this pack\").\n\nThanks,\nTaylor\n"},{"id":"415582","messageId":"YBReZxK3WbEAn5PV@nand.local","threadId":"55012","inReplyTo":"xmqqo8h8tp4j.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 02/10] revision: learn '--no-kept-objects'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-29T19:13:43Z","receivedAt":"2021-01-29T19:16:50Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Jan 28, 2021 at 07:10:04PM -0800, Junio C Hamano wrote:\n> We know that the enumerated objects do not appear in any of the kept\n> pack, but it does not mean all objects that are reachable/in-use that\n> are not in any kept packs are enumerated.\n\nYou raise a very valid point. FWIW, I originally wrote these patches as\njust \"enumerate the objects in these small packs, make a new pack out of\nthose, and then (optionally) delete the small ones. I abandoned that\nidea because it needs special handling for loose objects, and it has no\nidea which objects are unreachable, etc.\n\nBut maybe it is time to go back to the drawing board there. Perhaps a\n`--geometric` repack implies that we keep unreachable objects in effect,\nand that a full repack (i.e., one that does reachability analysis) is\nrequired to drop them.\n\nOther ideas are welcome.\n\nThanks,\nTaylor\n"},{"id":"415583","messageId":"YBRfvZh86Z8wAnxZ@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqqk0rwtom2.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T19:19:25Z","receivedAt":"2021-01-29T19:22:42Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jan 28, 2021 at 07:21:09PM -0800, Junio C Hamano wrote:\n\n> Taylor Blau <me@ttaylorr.com> writes:\n> \n> > Teach pack-objects an option to imply the revision machinery's new\n> > '--no-kept-objects' option when doing a reachability traversal.\n> >\n> > When '--assume-kept-packs-closed' is given as an argument to\n> > pack-objects, it behaves differently (i.e., passes different options to\n> > the ensuing revision walk) depending on whether or not other arguments\n> > are passed:\n> >\n> >   - If the caller also specifies a '--keep-pack' argument (to mark a\n> >     pack as kept in-core), then assume that this combination means to\n> >     stop traversal only at in-core packs.\n> >\n> >   - If instead the caller passes '--honor-pack-keep', then assume that\n> >     the caller wants to stop traversal only at packs with a\n> >     corresponding .keep file (consistent with the original meaning which\n> >     only refers to packs with a .keep file).\n> >\n> >   - If both '--keep-pack' and '--honor-pack-keep' are passed, then\n> >     assume the caller wants to stop traversal at either kind of kept\n> >     pack.\n> \n> If there is an out-of-band guarantee that .kept packs won't refer to\n> outside world, then we can obtain identical results to what existing\n> --honor-pack-keep (which traverses everything and then filteres out\n> what is in .keep pack) does by just stopping traversal when we see\n> an object that is found in a .keep pack.  OK, I guess that it\n> answers the correctness question I asked about [02/10].\n> \n> It still is curious how we can safely \"assume\", but presumably we\n> will see how in a patch that appears later in the series.\n\nI think this would generally happen if the .keep packs are generated\nusing something like \"git repack -a\", which packs everything reachable\ntogether. So if you do:\n\n  git repack -ad\n  touch .git/objects/pack/pack-whatever.keep\n  ... some more packs come in, perhaps via pushes ...\n  # imagine repack knew how to pass this along...\n  git repack -a --assume-kept-packs-closed\n\nthen you'd repack just the objects that aren't in the big pack.\n\n(And other kept packs; this probably would not want to be used with\n.keep packs, since those can be racily created by receive-pack. Rather\nyou'd want to say \"consider this specific list of packs to be closed and\nkept\").\n\nIt is dangerous, though. If your assumption is somehow wrong, then you'd\npotentially corrupt the repository (because you'd stop traversing, but\nperhaps delete a pack that contained a useful object you would have\nreached that you don't have elsewhere).\n\nThe overall goal here is being able to roll up loose objects and smaller\npacks without having to pay the cost of a full reachability traversal\n(which can take several minutes on large repositories). Another\nvery-different direction there is to just enumerate those objects\nwithout respect to reachability, stick them in a pack, and then delete\nthe originals. That does imply something like \"repack -k\", though, and\ninteracts weirdly with letting unreachable objects age out via their\nmtimes (we'd constantly suck them back into fresh packs).\n\nThat would work better if we our unreachable \"aging out\" storage was\nmarked as such (say, in a pack marked with a \".cruft\" file, rather than\njust a regular loose object that might be new or might be cruft). Then a\nroll-up repack would leave cruft packs alone (neither rolling them up,\nnor deleting them). A \"real\" repack would eventually delete them, but\nonly after having done an actual reachability traversal, which make sure\nthere are no objects within them that need rescued.\n\n> How \"closed\" are these kept packs supposed to be?  When there are\n> two .keep packs, should objects in each of the packs never refer to\n> outside their own pack, or is it OK for objects in one kept pack to\n> refer to another object in the other kept pack?  Readers and those\n> who want to understand and extend this code in the future would need\n> to know what definition of \"closed\" you are using here.\n\nI think it would want to be \"the set of all .keep packs is closed\". In a\n\"roll all into one\" scenario like above, there is only one .keep pack.\nBut in a geometric progression, that single pack which constitutes your\nbase set could be multiple packs (the last whole \"git repack -ad\", but\nthen a sequence of roll-ups that came on top of it).\n\n-Peff\n"},{"id":"415584","messageId":"YBRirz8xAB4Swf8X@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqqwnvwtqu1.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 01/10] packfile: introduce 'find_kept_pack_entry()'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T19:31:59Z","receivedAt":"2021-01-29T19:33:10Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jan 28, 2021 at 06:33:10PM -0800, Junio C Hamano wrote:\n\n> Taylor Blau <me@ttaylorr.com> writes:\n> \n> > Future callers will want a function to fill a 'struct pack_entry' for a\n> > given object id but _only_ from its position in any kept pack(s). They\n> > could accomplish this by calling 'find_pack_entry()' and checking\n> > whether the found pack is kept or not, but this is insufficient, since\n> > there may be duplicate objects (and the mru cache makes it unpredictable\n> > which variant we'll get).\n> \n> I wonder if we eventually need a callback interface to walk _all_\n> pack entries for a given object, so that \"I am only interested in\n> instances in kept packs\" will be under total control of the callers.\n> As it stands, it is \"just grab any one that is in a kept pack, any\n> one of them is fine\", which is almost just of as narrow utility as\n> the original's \"just grab the first one---any one of them is fine\",\n> the latter of which is \"insufficient\" as the log message says.\n\nWe do that already in pack-objects, and that's the problem: it's really\nslow. So if you have few kept packs, but a lot of other ones, you'd like\nto pre-split the packs into two lists, and not bother walking the one\nyou know won't turn up interesting results.\n\nI think the commit message here doesn't emphasize that reasoning enough.\nIt talks about using \"find_pack_entry()\", and that is definitely not\nsufficient for our purposes. But the interesting part is replacing the\nexisting \"walk all packs and see if any were kept\" logic, which happens\nin patch 6.\n\nSo the more compelling argument, I think, is something like:\n\n  - you sometimes want to know if object X is any kept packs\n\n  - you can't use find_pack_entry(), because it only gives you the first\n    pack it finds\n\n  - you can walk over all packs and look for the object in each.\n    pack-objects does this. But it's slow, because you are looking in\n    packs you don't care about.\n\n  - so it's helpful for the lookup to know up front which packs are\n    interesting to find objects in and which are not, to avoid looking\n    in the uninteresting ones\n\n-Peff\n"},{"id":"415585","messageId":"YBRi4v/AeDD/Zc9X@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqqft2ktnpj.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 05/10] p5303: measure time to repack with keep","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T19:32:50Z","receivedAt":"2021-01-29T19:33:34Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jan 28, 2021 at 07:40:40PM -0800, Junio C Hamano wrote:\n\n> > diff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\n> > index 277d22ec4b..85b077b72b 100755\n> > --- a/t/perf/p5303-many-packs.sh\n> > +++ b/t/perf/p5303-many-packs.sh\n> > @@ -27,8 +27,11 @@ repack_into_n () {\n> \n> this construct:\n> \n> \t... |\n> \tsed -n '1~5p' |\n> \thead -n \"$1\" |\n>         ...\n> \n> which is a GNUism.  Peff often says that very small population\n> actually run our perf suite, and this seems to corroborate the\n> conjecture.\n\nOops. Looks like I was the one who introduced that. Nobody seems to have\ncomplained, so I'm somewhat tempted to leave it. But it would not be too\nhard to replace with perl, I think.\n\n-Peff\n"},{"id":"415586","messageId":"YBRprCmIX4IrHAi0@nand.local","threadId":"55012","inReplyTo":"YBRfvZh86Z8wAnxZ@coredump.intra.peff.net","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-29T20:01:48Z","receivedAt":"2021-01-29T20:03:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Jan 29, 2021 at 02:19:25PM -0500, Jeff King wrote:\n> The overall goal here is being able to roll up loose objects and smaller\n> packs without having to pay the cost of a full reachability traversal\n> (which can take several minutes on large repositories). Another\n> very-different direction there is to just enumerate those objects\n> without respect to reachability, stick them in a pack, and then delete\n> the originals. That does imply something like \"repack -k\", though, and\n> interacts weirdly with letting unreachable objects age out via their\n> mtimes (we'd constantly suck them back into fresh packs).\n\nAs I mentioned in an earlier response to Junio, this was the original\napproach that I took when implementing this, but ultimately decided\nagainst it because it means that we'll never let unreachable objects age\nout (as you note).\n\nI wonder if we need our assumption that the union of kept packs is\nclosed under reachability to be specified as an option. If the option is\npassed, then we stop the traversal as soon as we hit an object in the\nfrozen packs. If not passed, then we do a full traversal but pass\n--honor-pack-keep to drop out objects in the frozen packs after the\nfact.\n\nThoughts?\n\n> I think it would want to be \"the set of all .keep packs is closed\". In a\n> \"roll all into one\" scenario like above, there is only one .keep pack.\n> But in a geometric progression, that single pack which constitutes your\n> base set could be multiple packs (the last whole \"git repack -ad\", but\n> then a sequence of roll-ups that came on top of it).\n\nI don't think having a roll-up strategy of \"all-except-one\" simplifies\nthings. Or, if it does, then I don't understand it. Isn't this the exact\nsame thing as a geometric repack which decides to keep only one pack?\n\nISTM that you would be susceptible to the same problems in this case,\ntoo.\n\n\nThanks,\nTaylor\n"},{"id":"415587","messageId":"YBRqOI84MkU+HNzt@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YBRi4v/AeDD/Zc9X@coredump.intra.peff.net","subject":"[PATCH] p5303: avoid sed GNU-ism","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T20:04:08Z","receivedAt":"2021-01-29T20:05:31Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 02:32:50PM -0500, Jeff King wrote:\n\n> > this construct:\n> > \n> > \t... |\n> > \tsed -n '1~5p' |\n> > \thead -n \"$1\" |\n> >         ...\n> > \n> > which is a GNUism.  Peff often says that very small population\n> > actually run our perf suite, and this seems to corroborate the\n> > conjecture.\n> \n> Oops. Looks like I was the one who introduced that. Nobody seems to have\n> complained, so I'm somewhat tempted to leave it. But it would not be too\n> hard to replace with perl, I think.\n\nMaybe worth doing this?\n\n-- >8 --\nSubject: [PATCH] p5303: avoid sed GNU-ism\n\nUsing \"1~5\" isn't portable. Nobody seems to have noticed, since perhaps\npeople don't tend to run the perf suite on more exotic platforms. Still,\nit's better to set a good example.\n\nWe can use:\n\n  perl -ne 'print if $. % 5 == 1'\n\ninstead. But we can further observe that perl does a good job of the\nother parts of this pipeline, and fold the whole thing together.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n t/perf/p5303-many-packs.sh | 12 ++++++++----\n 1 file changed, 8 insertions(+), 4 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex f4c2ab0584..ce0c42cc9f 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -21,10 +21,14 @@ repack_into_n () {\n \tmkdir staging &&\n \n \tgit rev-list --first-parent HEAD |\n-\tsed -n '1~5p' |\n-\thead -n \"$1\" |\n-\tperl -e 'print reverse <>' \\\n-\t>pushes\n+\tperl -e '\n+\t\tmy $n = shift;\n+\t\twhile (<>) {\n+\t\t\tlast unless @commits < $n;\n+\t\t\tpush @commits, $_ if $. % 5 == 1;\n+\t\t}\n+\t\tprint reverse @commits;\n+\t' \"$1\" >pushes\n \n \t# create base packfile\n \thead -n 1 pushes |\n-- \n2.30.0.759.g69d54d14a7\n\n"},{"id":"415588","messageId":"CAPig+cTc3RGEu4anOmnYBBQxXgB0UL8Z_vJfnV=FoNTOFYEpTQ@mail.gmail.com","threadId":"55012","inReplyTo":"YBRqOI84MkU+HNzt@coredump.intra.peff.net","subject":"Re: [PATCH] p5303: avoid sed GNU-ism","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2021-01-29T20:19:31Z","receivedAt":"2021-01-29T20:20:39Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Fri, Jan 29, 2021 at 3:07 PM Jeff King <peff@peff.net> wrote:\n> Subject: [PATCH] p5303: avoid sed GNU-ism\n>\n> Using \"1~5\" isn't portable. Nobody seems to have noticed, since perhaps\n> people don't tend to run the perf suite on more exotic platforms. Still,\n> it's better to set a good example.\n\nIt's not just exotic platforms on which this can be a problem. BSD\nlineage `sed`, such as stock `sed` on macOS, doesn't understand this\nnotation.\n\nThanks for eliminating this particular GNU-ism.\n"},{"id":"415589","messageId":"xmqq7dnvtrzh.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"YBRirz8xAB4Swf8X@coredump.intra.peff.net","subject":"Re: [PATCH 01/10] packfile: introduce 'find_kept_pack_entry()'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T20:20:34Z","receivedAt":"2021-01-29T20:22:16Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> So the more compelling argument, I think, is something like:\n>\n>   - you sometimes want to know if object X is any kept packs\n>\n>   - you can't use find_pack_entry(), because it only gives you the first\n>     pack it finds\n>\n>   - you can walk over all packs and look for the object in each.\n>     pack-objects does this. But it's slow, because you are looking in\n>     packs you don't care about.\n>\n>   - so it's helpful for the lookup to know up front which packs are\n>     interesting to find objects in and which are not, to avoid looking\n>     in the uninteresting ones\n\nThat does make sense.\n"},{"id":"415591","messageId":"YBRvQdHoslnF0OXr@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YBRprCmIX4IrHAi0@nand.local","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T20:25:37Z","receivedAt":"2021-01-29T20:29:43Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 03:01:48PM -0500, Taylor Blau wrote:\n\n> On Fri, Jan 29, 2021 at 02:19:25PM -0500, Jeff King wrote:\n> > The overall goal here is being able to roll up loose objects and smaller\n> > packs without having to pay the cost of a full reachability traversal\n> > (which can take several minutes on large repositories). Another\n> > very-different direction there is to just enumerate those objects\n> > without respect to reachability, stick them in a pack, and then delete\n> > the originals. That does imply something like \"repack -k\", though, and\n> > interacts weirdly with letting unreachable objects age out via their\n> > mtimes (we'd constantly suck them back into fresh packs).\n> \n> As I mentioned in an earlier response to Junio, this was the original\n> approach that I took when implementing this, but ultimately decided\n> against it because it means that we'll never let unreachable objects age\n> out (as you note).\n\nRight. But that's no different than using \"-k\" most of the time, and\nthen occasionally doing a more careful repack with short expiration\ntimes and a full reachability check. As you know, this is basically what\nwe do at GitHub.\n\nSo it may be reasonable to go that direction, which is really defining a\ntotally separate strategy from git-gc's \"repack, and occasionally\nobjects age out\". Especially if we find that the\nassume-kept-packs-closed route is too risky (i.e., has too many cases\nwhere it's possible to cause corruption if our assumptions isn't met).\n\nI'm not convinced either way at this point, but just thinking out loud\non the options (and trying to give some context to the list).\n\n> I wonder if we need our assumption that the union of kept packs is\n> closed under reachability to be specified as an option. If the option is\n> passed, then we stop the traversal as soon as we hit an object in the\n> frozen packs. If not passed, then we do a full traversal but pass\n> --honor-pack-keep to drop out objects in the frozen packs after the\n> fact.\n> \n> Thoughts?\n\nI'm confused. I thought the whole idea was to pass it as an option (the\nuser telling Git \"I know these packs are supposed to be closed; trust\nme\")?\n\n> > I think it would want to be \"the set of all .keep packs is closed\". In a\n> > \"roll all into one\" scenario like above, there is only one .keep pack.\n> > But in a geometric progression, that single pack which constitutes your\n> > base set could be multiple packs (the last whole \"git repack -ad\", but\n> > then a sequence of roll-ups that came on top of it).\n> \n> I don't think having a roll-up strategy of \"all-except-one\" simplifies\n> things. Or, if it does, then I don't understand it. Isn't this the exact\n> same thing as a geometric repack which decides to keep only one pack?\n> \n> ISTM that you would be susceptible to the same problems in this case,\n> too.\n\nI wasn't trying to argue that all-except-one avoids any problems. I was\nsaying that the example I gave above was an all-into-one, but if you\nwant to extend the concept to multiple packs, it has to cover the whole\nset. I.e., answering Junio's:\n\n  > is it OK for objects in one kept pack to refer to another object in\n  > the other kept pack?\n\nwith \"yes\".\n\n-Peff\n"},{"id":"415592","messageId":"YBRvy3s5gnsrIBEB@coredump.intra.peff.net","threadId":"55012","inReplyTo":"CAPig+cTc3RGEu4anOmnYBBQxXgB0UL8Z_vJfnV=FoNTOFYEpTQ@mail.gmail.com","subject":"Re: [PATCH] p5303: avoid sed GNU-ism","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T20:27:55Z","receivedAt":"2021-01-29T20:31:38Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 03:19:31PM -0500, Eric Sunshine wrote:\n\n> On Fri, Jan 29, 2021 at 3:07 PM Jeff King <peff@peff.net> wrote:\n> > Subject: [PATCH] p5303: avoid sed GNU-ism\n> >\n> > Using \"1~5\" isn't portable. Nobody seems to have noticed, since perhaps\n> > people don't tend to run the perf suite on more exotic platforms. Still,\n> > it's better to set a good example.\n> \n> It's not just exotic platforms on which this can be a problem. BSD\n> lineage `sed`, such as stock `sed` on macOS, doesn't understand this\n> notation.\n> \n> Thanks for eliminating this particular GNU-ism.\n\nOK, then I'm doubly surprised nobody has noticed and complained about\nthis. :)\n\n-Peff\n"},{"id":"415593","messageId":"xmqq35yjtrip.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"YBRfvZh86Z8wAnxZ@coredump.intra.peff.net","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T20:30:38Z","receivedAt":"2021-01-29T20:34:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> If there is an out-of-band guarantee that .kept packs won't refer to\n>> outside world, then we can obtain identical results to what existing\n>> --honor-pack-keep (which traverses everything and then filteres out\n>> what is in .keep pack) does by just stopping traversal when we see\n>> an object that is found in a .keep pack.  OK, I guess that it\n>> answers the correctness question I asked about [02/10].\n>> \n>> It still is curious how we can safely \"assume\", but presumably we\n>> will see how in a patch that appears later in the series.\n>\n> I think this would generally happen if the .keep packs are generated\n> using something like \"git repack -a\", which packs everything reachable\n> together. So if you do:\n>\n>   git repack -ad\n>   touch .git/objects/pack/pack-whatever.keep\n>   ... some more packs come in, perhaps via pushes ...\n>   # imagine repack knew how to pass this along...\n>   git repack -a --assume-kept-packs-closed\n>\n> then you'd repack just the objects that aren't in the big pack.\n\nYeah.  As a tool to help the above workflow, where you are only\ncreating another .keep out of youngest objects (i.e. those that are\neither loose or in non-kept packs), because by definition anything\nin .keep cannot be pointing back at these younger objects, it does\nmake sense to take advantage of \"the set of packs with .keep as a\nwhole is closed\".\n\nIt may become tricky once we start talking about creating a new\n.keep out of youngest objects PLUS a few young keep packs, though.\n\nStarting from all on-disk .keep packs, you'd mark them as in-core\nkeep bit, then drop in-core keep bit from the few young keep packs\nthat you intend to coalesce with the youngest objects---that is how\nI would imagine your repacking strategy would go.  The set of all\nthe on-disk .keep packs may give us \"closed\" guarantee, but if we \nexclude a few latest packs from that set, would the remainder still\ngive us the \"closed\" guarantee we can take advantage of, in order to\npack these youngest objects (including the ones in the kept packs\nthat we are coalescing)?\n"},{"id":"415594","messageId":"CAPig+cS9Fp3N9XNux8=JZo-T3QWdr7O3+NbChhhs62hh6D_5tQ@mail.gmail.com","threadId":"55012","inReplyTo":"YBRvy3s5gnsrIBEB@coredump.intra.peff.net","subject":"Re: [PATCH] p5303: avoid sed GNU-ism","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2021-01-29T20:36:01Z","receivedAt":"2021-01-29T20:37:19Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Fri, Jan 29, 2021 at 3:28 PM Jeff King <peff@peff.net> wrote:\n> On Fri, Jan 29, 2021 at 03:19:31PM -0500, Eric Sunshine wrote:\n> > It's not just exotic platforms on which this can be a problem. BSD\n> > lineage `sed`, such as stock `sed` on macOS, doesn't understand this\n> > notation.\n>\n> OK, then I'm doubly surprised nobody has noticed and complained about\n> this. :)\n\nAside from there possibly being relatively few regular Git developers\nusing macOS, it could also be because it's difficult to run the perf\ntests on macOS in the first place due to the GNU prerequisites. For\ninstance, the perf tests have an unconditional dependency on GNU\n`time` which is not installed on macOS by default, and it's not always\neasy to figure out how to obtain it.\n"},{"id":"415595","messageId":"xmqqy2gbsclr.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"YBRi4v/AeDD/Zc9X@coredump.intra.peff.net","subject":"Re: [PATCH 05/10] p5303: measure time to repack with keep","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T20:38:08Z","receivedAt":"2021-01-29T20:39:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Thu, Jan 28, 2021 at 07:40:40PM -0800, Junio C Hamano wrote:\n>\n>> > diff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\n>> > index 277d22ec4b..85b077b72b 100755\n>> > --- a/t/perf/p5303-many-packs.sh\n>> > +++ b/t/perf/p5303-many-packs.sh\n>> > @@ -27,8 +27,11 @@ repack_into_n () {\n>> \n>> this construct:\n>> \n>> \t... |\n>> \tsed -n '1~5p' |\n>> \thead -n \"$1\" |\n>>         ...\n>> \n>> which is a GNUism.  Peff often says that very small population\n>> actually run our perf suite, and this seems to corroborate the\n>> conjecture.\n>\n> Oops. Looks like I was the one who introduced that. Nobody seems to have\n> complained, so I'm somewhat tempted to leave it. But it would not be too\n> hard to replace with perl, I think.\n\nYeah, but would it be worth it?  I am actually OK to say that you\nneed GNU sed if you want to run perf.  We already rely on GNU time\nto run perf tests, no?\n\n"},{"id":"415602","messageId":"YBSHv3TWleRxM1+/@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqqy2gbsclr.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 05/10] p5303: measure time to repack with keep","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T22:10:07Z","receivedAt":"2021-01-29T22:11:07Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 12:38:08PM -0800, Junio C Hamano wrote:\n\n> > Oops. Looks like I was the one who introduced that. Nobody seems to have\n> > complained, so I'm somewhat tempted to leave it. But it would not be too\n> > hard to replace with perl, I think.\n> \n> Yeah, but would it be worth it?  I am actually OK to say that you\n> need GNU sed if you want to run perf.  We already rely on GNU time\n> to run perf tests, no?\n\nTrue. This one is a little worse because it's subtle, and somebody might\ncopy it unknowingly into the regular test suite.\n\nI am happy to leave it, or for you to pick up the patch I sent earlier\n(which I did verify produces identical output).\n\n-Peff\n"},{"id":"415603","messageId":"YBSHzG9T72nYYVt4@nand.local","threadId":"55012","inReplyTo":"YBRvQdHoslnF0OXr@coredump.intra.peff.net","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-29T22:10:20Z","receivedAt":"2021-01-29T22:11:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Jan 29, 2021 at 03:25:37PM -0500, Jeff King wrote:\n> So it may be reasonable to go that direction, which is really defining a\n> totally separate strategy from git-gc's \"repack, and occasionally\n> objects age out\". Especially if we find that the\n> assume-kept-packs-closed route is too risky (i.e., has too many cases\n> where it's possible to cause corruption if our assumptions isn't met).\n\nYeah, this whole conversation has made me very nervous about using\nreachability. Fundamentally, this isn't about reachability at all. The\noperation is as simple as telling pack-objects a list of packs that you\ndo and don't want objects from, making a new pack out of that, and then\noptionally dropping the packs that you rolled up.\n\nSo, I think that teaching pack-objects a way to understand a caller that\nsays \"include objects from packs X, Y, and Z, but not if they appear in\npacks A, B, or C, and also pull in any loose objects\" is the best way\nforward here.\n\nOf course, you're going to be dragging along unreachable objects until\nyou decide to do a full repack, but I'm OK with that since we wouldn't\nexpect anybody to be solely relying on geometric repacks without\noccasionally running 'git repack -ad'.\n\nJunio: I don't think that you have picked this up yet, but please avoid\ndoing so for now, and I'll send a new series that goes in the direction\nI outlined above.\n\nThanks,\nTaylor\n"},{"id":"415604","messageId":"YBSIE7Om1SGR83LJ@nand.local","threadId":"55012","inReplyTo":"CAPig+cS9Fp3N9XNux8=JZo-T3QWdr7O3+NbChhhs62hh6D_5tQ@mail.gmail.com","subject":"Re: [PATCH] p5303: avoid sed GNU-ism","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-29T22:11:31Z","receivedAt":"2021-01-29T22:12:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Jan 29, 2021 at 03:36:01PM -0500, Eric Sunshine wrote:\n> On Fri, Jan 29, 2021 at 3:28 PM Jeff King <peff@peff.net> wrote:\n> > On Fri, Jan 29, 2021 at 03:19:31PM -0500, Eric Sunshine wrote:\n> > > It's not just exotic platforms on which this can be a problem. BSD\n> > > lineage `sed`, such as stock `sed` on macOS, doesn't understand this\n> > > notation.\n> >\n> > OK, then I'm doubly surprised nobody has noticed and complained about\n> > this. :)\n>\n> Aside from there possibly being relatively few regular Git developers\n> using macOS, it could also be because it's difficult to run the perf\n> tests on macOS in the first place due to the GNU prerequisites. For\n> instance, the perf tests have an unconditional dependency on GNU\n> `time` which is not installed on macOS by default, and it's not always\n> easy to figure out how to obtain it.\n\nYep, I agree completely. I was going to say that this would produce a\nconflict (albeit, a trivial one) with the series that this came out of.\n\nBut I think that we're better off abandoning that series for now until I\nsend a different version, so I think we should just go ahead an apply\nthis.\n\n\nThanks,\nTaylor\n"},{"id":"415605","messageId":"xmqqpn1ns86k.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"YBRvQdHoslnF0OXr@coredump.intra.peff.net","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T22:13:39Z","receivedAt":"2021-01-29T22:15:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> I wonder if we need our assumption that the union of kept packs is\n>> closed under reachability to be specified as an option. If the option is\n>> passed, then we stop the traversal as soon as we hit an object in the\n>> frozen packs. If not passed, then we do a full traversal but pass\n>> --honor-pack-keep to drop out objects in the frozen packs after the\n>> fact.\n>> \n>> Thoughts?\n>\n> I'm confused. I thought the whole idea was to pass it as an option (the\n> user telling Git \"I know these packs are supposed to be closed; trust\n> me\")?\n\nYes, that is how I read these patches, and it sounds like an assumption\nthat we can make under many scenarios/repacking strategies.\n"},{"id":"415607","messageId":"YBSPlO/ki5vRNX0T@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqq35yjtrip.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T22:43:32Z","receivedAt":"2021-01-29T22:46:02Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 12:30:38PM -0800, Junio C Hamano wrote:\n\n> > I think this would generally happen if the .keep packs are generated\n> > using something like \"git repack -a\", which packs everything reachable\n> > together. So if you do:\n> >\n> >   git repack -ad\n> >   touch .git/objects/pack/pack-whatever.keep\n> >   ... some more packs come in, perhaps via pushes ...\n> >   # imagine repack knew how to pass this along...\n> >   git repack -a --assume-kept-packs-closed\n> >\n> > then you'd repack just the objects that aren't in the big pack.\n> \n> Yeah.  As a tool to help the above workflow, where you are only\n> creating another .keep out of youngest objects (i.e. those that are\n> either loose or in non-kept packs), because by definition anything\n> in .keep cannot be pointing back at these younger objects, it does\n> make sense to take advantage of \"the set of packs with .keep as a\n> whole is closed\".\n> \n> It may become tricky once we start talking about creating a new\n> .keep out of youngest objects PLUS a few young keep packs, though.\n\nRight. You'd have to make sure the younger packs were also created with\nthat reachability in mind. I.e., if you are in a situation where you've\ngot:\n\n  - a big \"old\" pack\n  - N new packs from pushes\n\nthen you can't assume anything about the reachability for those\nindividual N packs. It would be wrong to split any of them into the\n\"keep\" side. They need to have all their objects traversed until we hit\nsomething in the old pack.\n\nBut if you have a situation with:\n\n  - a big \"old\" pack\n  - M packs made from previous rollups on top of the old pack\n  - N new packs from pushes\n\nThen I think you can still take the old pack + the M packs as a\ncohesive unit closed under reachability.\n\nThe tricky part is knowing which packs are which (size is a heuristic,\nbut it can be wrong; people may make a big push to a previously-small\nrepository).\n\n> Starting from all on-disk .keep packs, you'd mark them as in-core\n> keep bit, then drop in-core keep bit from the few young keep packs\n> that you intend to coalesce with the youngest objects---that is how\n> I would imagine your repacking strategy would go.  The set of all\n> the on-disk .keep packs may give us \"closed\" guarantee, but if we \n> exclude a few latest packs from that set, would the remainder still\n> give us the \"closed\" guarantee we can take advantage of, in order to\n> pack these youngest objects (including the ones in the kept packs\n> that we are coalescing)?\n\nYeah, I think we are going along the same lines. Except I think it is\ndangerous to use on-disk \".keep\" as your marker, because we will racily\nsee incoming push packs with a \".keep\" (which receive-pack/index-pack\nuse as a lockfile until the refs are updated).\n\nSo repack has to \"somehow\" get the list of which is which.\n\nNone of which is disputing your \"it may become tricky\", of course. ;) It\nis exactly this trickiness that I am worried about. And I am not being\ncoy with \"somehow\", as if we have some custom not-yet-shared layer on\ntop of repack that tracks this. We are still figuring out whether this\nis a good direction in the first place. :)\n\nOne of the things that led us to this reachability traversal, away from\n\"just suck up all of the objects from these packs\", is that this is how\n--unpacked works. I had always assumed it was implemented as \"if it's\nloose, then put it in the pack\". But it's not. It's attached to the\nrevision traversal. And it actually gets some cases wrong!\n\nIt will walk every commit, so you don't have to worry about a packed\ncommit referring to an unpacked one. But it doesn't look at the trees of\nthe packed commits (for the quite obvious reason that doing so is orders\nof magnitude more expensive). That means that if there is a packed\ncommit that refers to an unpacked blob (which is not referenced by an\nunpacked commit), then \"rev-list --unpacked\" will not report it (and\nlikewise \"git repack -d\" would not pack it).\n\nIt's easy to create such a situation manually, but I've included a\nmore plausible sequence involving \"--amend\" and push/fetch unpackLimit\nat the end of this email.\n\nAt the core, --unpacked is assuming certain things about reachability of\nloose/packed objects that aren't necessarily true. And this\n--assume-kept-pack-closed stuff is basically doing the same thing for a\nparticular set of packs (albeit more so; I believe the patches here cut\noff traversal of parent pointers, not just commit-to-tree pointers).\n\nOne of the reasons I think nobody noticed with --unpacked is that the\nstakes are pretty low. If our assumption is wrong, the worst case is\nthat a loose object remains unpacked in \"repack -d\". But we'd never\ndelete it based on that information (instead, git prune would do its own\ntraversal to find the correct reachability). And it would eventually get\npicked up by \"git repack -ad\".\n\nBut for repacking, the general strategy is to put things you want to\nkeep into the new pack, and then delete the old ones (not marked as\nkeep, of course). So if our assumption is ever wrong, it means we'd\npotentially drop packs that have reachable objects not found elsewhere,\nand we'd end up corrupting the repository.\n\nSo I think the paths forward are either:\n\n  - come up with an air-tight system of making sure that we know packs\n    we claim are closed under reachability really are (perhaps some\n    marker that says \"I was generated by repack -a\")\n\n  - have a \"roll-up\" mode that does not care about reachability at all,\n    and just takes any objects from a particular set of packs (plus\n    probably loose objects)\n\nI'm still thinking aloud here, and not really sure which is a better\npath. I do feel like the failure modes for the second one are less\nrisky.\n\nAnyway, here's the --unpacked example, if you're curious. It's based on\nfetching, but you could invert it to do pushes (in which case it is\nrepacking in \"parent\" that gets the wrong result).\n\n-- >8 --\n# two repos, one a clone of the other\ngit init parent\ngit -C parent commit --allow-empty -m base\ngit clone parent child\n\n# now there's a small fetch, which will get\n# exploded into loose objects.\n(\n\tcd parent\n\techo small >small\n\tgit add small\n\tgit commit -m small\n)\ngit -C child fetch\n\n# We can verify that \"rev-list --unpacked\" reports these\n# objects.\ngit -C child rev-list --objects --unpacked origin\n\n# and now a bigger one that will remain a pack (we'll\n# tweak unpackLimit instead of making a really big commit,\n# but the concept is the same)\n#\n# There are two key things here:\n#\n#   - the \"small\" commit is no longer reachable, but the big one\n#     still contains the \"small\" blob object. Using --amend is\n#     a plausible mechanism for this happening.\n#\n#   - we are using bitmaps, which give us an exact answer for the set of\n#     objects to send. Otherwise, pack-objects on the server actually\n#     fails to notice the other side has told us it has the small blob, and\n#     sends another copy of it.\n(\n\tcd parent\n\tgit repack -adb\n\techo big >big\n\tgit add big\n\tgit commit --amend -m big\n)\ngit -C child -c fetch.unpackLimit=1 fetch\n\n# So now in the child we have a packed object\n# whose ancestor is an unpacked one. rev-list\n# now won't report the \"small\" blob (ac790413).\ngit.compile -C child rev-list --objects --unpacked origin\n\n# Even though we can see that it's present only as an\n# unpacked object.\nshow_objects() {\n\tfor i in child/.git/objects/pack/*.idx; do\n\t\tgit show-index <$i\n\tdone | cut -d' ' -f2 | sed 's/^/packed: /'\n\tfind child/.git/objects/??/* |\n\t\tperl -F/ -alne 'print \" loose: $F[-2]$F[-1]\"'\n}\n\n# If we were to do an incremental repack now, it wouldn't be packed.\n# (Note we have to kill off the reflog, which still references the\n# rewound commit).\nrm -rf child/.git/logs\ngit -C child repack -d\nshow_objects\n"},{"id":"415608","messageId":"YBSSBviXOL8rM3Ao@nand.local","threadId":"55012","inReplyTo":"YBSPlO/ki5vRNX0T@coredump.intra.peff.net","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-29T22:53:58Z","receivedAt":"2021-01-29T22:55:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Jan 29, 2021 at 05:43:32PM -0500, Jeff King wrote:\n> So I think the paths forward are either:\n>\n>   - come up with an air-tight system of making sure that we know packs\n>     we claim are closed under reachability really are (perhaps some\n>     marker that says \"I was generated by repack -a\")\n>\n>   - have a \"roll-up\" mode that does not care about reachability at all,\n>     and just takes any objects from a particular set of packs (plus\n>     probably loose objects)\n>\n> I'm still thinking aloud here, and not really sure which is a better\n> path. I do feel like the failure modes for the second one are less\n> risky.\n\nThe more I think about it, the more I feel that the second option is the\nright approach. It seems like if you were naïvely implementing this from\nscratch, that you'd pick the second one (i.e., have pack-objects\nunderstand a new input mode, and then make a pack based on that).\n\nI am leery that we'd be able to get the first option \"right\" without\nattaching some sort of marker to each pack, especially given how\ndifficult I think that this is to reason about precisely. I suppose you\ncould have a .closed file corresponding to each pack, or alternatively a\n$objdir/pack/pack-geometry file which specifies the same thing, but both\nof these feel overly restrictive.\n\nBesides having to special case the loose objects, is there any downside\nto doing the simpler thing here?\n\n> Anyway, here's the --unpacked example, if you're curious. It's based on\n> fetching, but you could invert it to do pushes (in which case it is\n> repacking in \"parent\" that gets the wrong result).\n\nFascinating indeed :-).\n\nThanks,\nTaylor\n"},{"id":"415610","messageId":"YBSSvXHIwUe/8rVj@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YBSHzG9T72nYYVt4@nand.local","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T22:57:01Z","receivedAt":"2021-01-29T22:57:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 05:10:20PM -0500, Taylor Blau wrote:\n\n> On Fri, Jan 29, 2021 at 03:25:37PM -0500, Jeff King wrote:\n> > So it may be reasonable to go that direction, which is really defining a\n> > totally separate strategy from git-gc's \"repack, and occasionally\n> > objects age out\". Especially if we find that the\n> > assume-kept-packs-closed route is too risky (i.e., has too many cases\n> > where it's possible to cause corruption if our assumptions isn't met).\n> \n> Yeah, this whole conversation has made me very nervous about using\n> reachability. Fundamentally, this isn't about reachability at all. The\n> operation is as simple as telling pack-objects a list of packs that you\n> do and don't want objects from, making a new pack out of that, and then\n> optionally dropping the packs that you rolled up.\n> \n> So, I think that teaching pack-objects a way to understand a caller that\n> says \"include objects from packs X, Y, and Z, but not if they appear in\n> packs A, B, or C, and also pull in any loose objects\" is the best way\n> forward here.\n> \n> Of course, you're going to be dragging along unreachable objects until\n> you decide to do a full repack, but I'm OK with that since we wouldn't\n> expect anybody to be solely relying on geometric repacks without\n> occasionally running 'git repack -ad'.\n\nWhile writing my other response, I had some thoughts that this \"dragging\nalong\" might not be so bad.\n\nJust to lay out the problem as I see it, if you do:\n\n  - frequently roll up all small packs and loose objects into a new\n    pack, without regard to reachability\n\n  - occasionally run \"git repack -ad\" to do a real traversal\n\nthen the problem is that unreachable objects never age out:\n\n  - a loose unreachable object starts with a recent-ish mtime\n\n  - the frequent roll-up rolls it into a pack, freshening its mtime\n\n  - the full \"repack -ad\" doesn't delete it, because its pack mtime is\n    too recent. It explodes it loose again.\n\n  - repeat forever\n\nWe know that \"repack -d\" is not 100% accurate because of similar \"closed\nunder reachability\" assumptions (see my other email). But it's OK,\nbecause the worst case is an object that doesn't quite get packed yet,\nnot that it gets deleted.\n\nSo you could do something like:\n\n  - roll up loose objects into a pack with \"repack -d\"; mostly accurate,\n    but doesn't suck up unreachable objects\n\n  - roll up small packs into a bigger pack without regard for\n    reachability. This includes the pack created in the first step, but\n    we know everything in it is actually reachable.\n\n  - eventually run \"repack -ad\" to do a real traversal\n\nThat would extend the lifetime of unreachable objects which were found\nin a pack (they get dragged forward during the rollups). But they'd\neventually get exploded loose during a \"repack -ad\", and then _not_\nsucked back into a roll-up pack. And then eventually \"repack -ad\"\nremoves them.\n\nThe downsides are:\n\n  - doing a separate \"repack -d\" plus a roll-up repack is wasted work.\n    But I think they could be combined into a single step (at the cost\n    of some extra complexity in the implementation).\n\n  - using \"--unpacked\" still means traversing every commit. That's much\n    faster than traversing the whole object graph, but still scales with\n    the size of the repo, not the size of the new objects. That might be\n    acceptable, though.\n\nI do think the original problem goes away entirely if we can keep better\ntrack of the mtimes. I.e., if we had packs marked with \".cruft\" instead\nof exploding loose, then the logic is:\n\n  - roll up all loose objects and any objects in a pack that isn't\n    marked as cruft (or keep); never delete a cruft pack at this stage\n\n  - occasionally \"repack -ad\"; this does delete old cruft packs (because\n    we'd have rescued any reachable objects they might have contained)\n\nI'm not sure I want to block this topic on having cruft packs, though.\nOf course there are tons of _other_ reasons to want them (like not\ncausing operational headaches when a repo's disk and inode usage grows\nby 10x due to exploding loose objects). So maybe it's not a bad idea to\nwork on them together. I dunno.\n\n-Peff\n"},{"id":"415611","messageId":"YBSTjkuW8Mib4o5A@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YBSSBviXOL8rM3Ao@nand.local","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T23:00:30Z","receivedAt":"2021-01-29T23:01:21Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 05:53:58PM -0500, Taylor Blau wrote:\n\n> > I'm still thinking aloud here, and not really sure which is a better\n> > path. I do feel like the failure modes for the second one are less\n> > risky.\n> \n> The more I think about it, the more I feel that the second option is the\n> right approach. It seems like if you were naïvely implementing this from\n> scratch, that you'd pick the second one (i.e., have pack-objects\n> understand a new input mode, and then make a pack based on that).\n> \n> I am leery that we'd be able to get the first option \"right\" without\n> attaching some sort of marker to each pack, especially given how\n> difficult I think that this is to reason about precisely. I suppose you\n> could have a .closed file corresponding to each pack, or alternatively a\n> $objdir/pack/pack-geometry file which specifies the same thing, but both\n> of these feel overly restrictive.\n\nYeah, I think my gut feeling matches yours.\n\n> Besides having to special case the loose objects, is there any downside\n> to doing the simpler thing here?\n\nThe other downside I can think of is that you can't just run \"git repack\n--geometric\" every time, and eventually get a good result (or one that\nasymptotically approaches good ;) ). I.e., you now have two types of\nrepacks: quick and dirty rollups, and \"real\" ones that do reachability.\nSo you need some heuristics about how often you do one versus the other.\n\nI'm definitely OK with that outcome. And I think we could even bake\nthose heuristics into a script or mode of repack (e.g., maybe \"gc\n--auto\" would trigger a bigger repack every N times or something). But\nthat's what I came up with by brainstorming. :)\n\n-Peff\n"},{"id":"415612","messageId":"xmqqh7mzs5w3.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"YBSHzG9T72nYYVt4@nand.local","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T23:03:08Z","receivedAt":"2021-01-29T23:04:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> So, I think that teaching pack-objects a way to understand a caller that\n> says \"include objects from packs X, Y, and Z, but not if they appear in\n> packs A, B, or C, and also pull in any loose objects\" is the best way\n> forward here.\n\nAre our goals still include that the resulting packfile has good\ndelta compression and object locality?  Reachability traversal\ndiscovers which commit comes close to which other commits to help\npack-objects to arrange the resulting pack so that objects that\nappear close together in history appears close together.  It also\ngives each object a pathname hint to help group objects of the same\ntype (either blobs or trees) with like-paths together for better\ndeltification.\n\nWithout reachability traversal, I would imagine that it would become\nquite important to keep the order in which objects appear in the\noriginal pack, and existing delta chain, as much as possible, or\nwe'd be seeing a horribly inefficient pack like fast-import would\nproduce.\n\nThanks.\n"},{"id":"415613","messageId":"xmqqczxns5j8.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"YBSSBviXOL8rM3Ao@nand.local","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T23:10:51Z","receivedAt":"2021-01-29T23:11:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> On Fri, Jan 29, 2021 at 05:43:32PM -0500, Jeff King wrote:\n>> So I think the paths forward are either:\n>>\n>>   - come up with an air-tight system of making sure that we know packs\n>>     we claim are closed under reachability really are (perhaps some\n>>     marker that says \"I was generated by repack -a\")\n>>\n>>   - have a \"roll-up\" mode that does not care about reachability at all,\n>>     and just takes any objects from a particular set of packs (plus\n>>     probably loose objects)\n>>\n>> I'm still thinking aloud here, and not really sure which is a better\n>> path. I do feel like the failure modes for the second one are less\n>> risky.\n>\n> The more I think about it, the more I feel that the second option is the\n> right approach. It seems like if you were naïvely implementing this from\n> scratch, that you'd pick the second one (i.e., have pack-objects\n> understand a new input mode, and then make a pack based on that).\n\nYes, \"roll-up\" mode would be a sensible thing to have, as long as we\ncan keep pruning out of the picture for now.  But in the end, I do\nthink \"stop at any object in this frozen pack---these objects go all\nthe way down to root and we know they are reachable\" optimization\nthat would give 'prune' a performance boost with small margin of\nfalse positive about reachability (i.e. we may never be able to\nprune away an object in such a pack, even when it becomes\nunreachable) would be a valuable thing to have in a practical\nsystem, so from that point of view, the work done in these patches\nare not lost ;-)\n\nThe efficiency issue of the resulting pack I mentioned earlier in a\nseparate message is there in the \"roll-up\" mode, though.\n"},{"id":"415614","messageId":"xmqq8s8bs5ft.fsf@gitster.c.googlers.com","threadId":"55012","inReplyTo":"YBSHv3TWleRxM1+/@coredump.intra.peff.net","subject":"Re: [PATCH 05/10] p5303: measure time to repack with keep","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-29T23:12:54Z","receivedAt":"2021-01-29T23:13:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Fri, Jan 29, 2021 at 12:38:08PM -0800, Junio C Hamano wrote:\n>\n>> > Oops. Looks like I was the one who introduced that. Nobody seems to have\n>> > complained, so I'm somewhat tempted to leave it. But it would not be too\n>> > hard to replace with perl, I think.\n>> \n>> Yeah, but would it be worth it?  I am actually OK to say that you\n>> need GNU sed if you want to run perf.  We already rely on GNU time\n>> to run perf tests, no?\n>\n> True. This one is a little worse because it's subtle, and somebody might\n> copy it unknowingly into the regular test suite.\n>\n> I am happy to leave it, or for you to pick up the patch I sent earlier\n> (which I did verify produces identical output).\n\nYeah, I would be very unhappy if somebody copied-and-pasted it, but\nsomehow I didn't think too many people moved code in that direction\n;-)\n\nWill apply the portability fix, then.\n\nThanks.\n"},{"id":"415616","messageId":"YBSaHHKV5ncjjJum@nand.local","threadId":"55012","inReplyTo":"xmqqh7mzs5w3.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-29T23:28:28Z","receivedAt":"2021-01-29T23:29:35Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Jan 29, 2021 at 03:03:08PM -0800, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > So, I think that teaching pack-objects a way to understand a caller that\n> > says \"include objects from packs X, Y, and Z, but not if they appear in\n> > packs A, B, or C, and also pull in any loose objects\" is the best way\n> > forward here.\n>\n> Are our goals still include that the resulting packfile has good\n> delta compression and object locality?  Reachability traversal\n> discovers which commit comes close to which other commits to help\n> pack-objects to arrange the resulting pack so that objects that\n> appear close together in history appears close together.  It also\n> gives each object a pathname hint to help group objects of the same\n> type (either blobs or trees) with like-paths together for better\n> deltification.\n\nI think our goals here are somewhere between having fewer packfiles\nwhile also ensuring that the packfiles we had to create don't have\nhorrible delta compression and locality.\n\nBut now that you do mention it, I remember the reachability traversal's\nbringing in object names was a reason that we decided to implement this\nseries using a reachability traversal in the first place.\n\n> Without reachability traversal, I would imagine that it would become\n> quite important to keep the order in which objects appear in the\n> original pack, and existing delta chain, as much as possible, or\n> we'd be seeing a horribly inefficient pack like fast-import would\n> produce.\n\nYeah; we'd definitely want to feed the objects to pack-objects in the\norder that they appear in the original pack. Maybe that's not that bad a\ntradeoff to make, though...\n\n> Thanks.\n\nThanks,\nTaylor\n"},{"id":"415617","messageId":"YBSazIgZds5ZWGlW@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqqh7mzs5w3.fsf@gitster.c.googlers.com","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-29T23:31:24Z","receivedAt":"2021-01-29T23:32:39Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 29, 2021 at 03:03:08PM -0800, Junio C Hamano wrote:\n\n> Taylor Blau <me@ttaylorr.com> writes:\n> \n> > So, I think that teaching pack-objects a way to understand a caller that\n> > says \"include objects from packs X, Y, and Z, but not if they appear in\n> > packs A, B, or C, and also pull in any loose objects\" is the best way\n> > forward here.\n> \n> Are our goals still include that the resulting packfile has good\n> delta compression and object locality?  Reachability traversal\n> discovers which commit comes close to which other commits to help\n> pack-objects to arrange the resulting pack so that objects that\n> appear close together in history appears close together.  It also\n> gives each object a pathname hint to help group objects of the same\n> type (either blobs or trees) with like-paths together for better\n> deltification.\n> \n> Without reachability traversal, I would imagine that it would become\n> quite important to keep the order in which objects appear in the\n> original pack, and existing delta chain, as much as possible, or\n> we'd be seeing a horribly inefficient pack like fast-import would\n> produce.\n\nThanks, that's another good point we discussed a while ago (off-list),\nbut hasn't come up in this discussion yet.\n\nAnother option here is not to roll up packs at all, but instead to use a\nmidx to cover them all[1]. That solves the issue where object lookup is\nO(nr_packs), and you retain the same locality and delta characteristics.\n\nBut I think part of the goal is to actually improve the deltas, in two\nways:\n\n  - we'd hopefully find new delta opportunities between objects in the\n    various packs\n\n  - we'll drop some objects that are duplicated in other packs.\n    Definitely we have to to avoid duplicates in the roll-up pack, but I\n    think we'd want to even for objects that are in the \"big\" kept pack.\n    These are likely bases of deltas in our roll-up pack, since the\n    common cause there is --fix-thin adding them to complete the pack.\n    But we really prefer to serve fetches using the ones out of the main\n    pack, since they may already themselves be deltas (which makes them\n    way cheaper; we can send the delta straight off the disk, rather\n    than looking for a new possible base).\n\nSo I would anticipate the delta-compression phase actually trying to do\nsome new work. I do worry that the lack of pathname hints may make the\ndeltas we find much more worse (or cause us to spend excessive CPU\nsearching for them). It's possible we could do a \"best effort\" traversal\nwhere we walk new commits to find newly added pathnames, but don't\nbother crossing into trees/commits that aren't in the set of objects to\nbe packed. It's OK to optimize for speed there, because it's just\nfeeding the delta heuristic, not the set of objects we'd plan to pack.\n\n-Peff\n\n[1] Our end-game plan is actually to _also_ use a midx to cover the\n    roll-ups and the \"big\" pack, since we'd want to generate bitmaps for\n    the new objects, too.'\n"},{"id":"415872","messageId":"YBjBWD8Lz81P9ElM@nand.local","threadId":"55012","inReplyTo":"YBSaHHKV5ncjjJum@nand.local","subject":"Re: [PATCH 03/10] builtin/pack-objects.c: learn '--assume-kept-packs-closed'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-02T03:04:56Z","receivedAt":"2021-02-02T03:06:09Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Jan 29, 2021 at 06:28:28PM -0500, Taylor Blau wrote:\n> On Fri, Jan 29, 2021 at 03:03:08PM -0800, Junio C Hamano wrote:\n> > Are our goals still include that the resulting packfile has good\n> > delta compression and object locality?  Reachability traversal\n> > discovers which commit comes close to which other commits to help\n> > pack-objects to arrange the resulting pack so that objects that\n> > appear close together in history appears close together.  It also\n> > gives each object a pathname hint to help group objects of the same\n> > type (either blobs or trees) with like-paths together for better\n> > deltification.\n>\n> I think our goals here are somewhere between having fewer packfiles\n> while also ensuring that the packfiles we had to create don't have\n> horrible delta compression and locality.\n>\n> But now that you do mention it, I remember the reachability traversal's\n> bringing in object names was a reason that we decided to implement this\n> series using a reachability traversal in the first place.\n\nPeff shared a very clever idea with me today. Like in the naive\napproach, we fill the list of \"objects to pack\" with everything in the\npacks that are about to get rolled up, excluding anything that appears\nin the large packs.\n\nBut we do a reachability traversal whose starting points are all of the\ncommits in the packs that are about to be rolled up, filling in the\nnamehash of the objects we encounter along the way.\n\nLike in the original version of this series, we'll stop early once we\nencounter an object in any of the frozen packs (which are marked as kept\nin core), and so we might not traverse through everything. But that's\ncompletely OK, since we know we have the right list of objects to pack\n(at worst, we would having some zero'd namehashes and come up with\nslightly worse deltas).\n\nBut, I think that this is a nice middle-ground (and it allows us to\nreuse lots of work from the original version), so I'm quite happy.\n\nIt's in my fork [1] in the tb/geometric-repack.wip branch, but I'll try\nand clean those patches up tomorrow and send a v2 to the list.\n\nThanks,\nTaylor\n\n[1]: https://github.com/ttaylorr/git\n"},{"id":"416127","messageId":"cover.1612411123.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH v2 0/8] repack: support repacking into a geometric sequence","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:58:45Z","receivedAt":"2021-02-04T04:00:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Here is an updated version of mine and Peff's series to add a new 'git repack\n--geometric' mode which supports repacking a repository into a geometric\nprogression of packs by object count.\n\nThis version depends on jk/p5303-sed-portability-fix, but it could be applied\nonto 'master' after resolving a trivial conflict.\n\nAs a reminder, here is a description from the original cover letter [1] which\noutlines what the geometric mode entails:\n\n  Roughly speaking, for a given factor, say \"d\", each pack has at least \"d\"\n  times the number of objects as the next largest pack. So, if there are \"N\"\n  packs, \"P1\", \"P2\", ..., \"PN\" ordered by object count (where \"PN\" has the most\n  objects, and \"P1\" the fewest), then:\n\n    objects(Pi) > d * objects(P(i-1))\n\n  for all 1 < i <= N.\n\n  This is done by first ordering packs by object count, and then determining the\n  longest sequence of large packs which already form a geometric progression.\n  All packs on the small side of that cut must be repacked together, and so we\n  check that the existing progression can be maintained with the new pack, and\n  adjust as necessary.\n\nSince last time, the series has been reworked substantially. In the previous\nversion, a single reachability traversal was performed to determine the set of\nobjects to pack. That traversal halted upon encountering any objects found in a\nkept pack, but this led to serious correctness problems (if, for e.g., an object\nwe would like to pack is an ancestor of some other object in a kept pack, and\nthus isn't picked up).\n\nThe details of the new approach can be found in the third patch, but the gist is\nas follows:\n\n  - 'git repack --geometric' calls 'git pack-objects --stdin-packs', which\n    expects input like:\n\n        pack-xyz.pack\n        pack-abc.pack\n        ^pack-exclude.pack\n\n    'git pack-objects' determines the set of objects to pack by iterating all of\n    the objects in the listed packs, and then removing any objects found in the\n    packs which are prefixed with '^'.\n\n  - To improve the delta selection process, the same reachability traversal from\n    the original version of this series is performed. But, the set of objects to\n    pack is already known, so we don't run the risk of the correctness bugs from\n    before.\n\n  - In this reachability traversal, visited objects get their namehash field\n    set, which helps drive the heuristics that power delta selection. It's\n    possible that we may not visit all of the objects to pack, but that's OK\n    since this process is only additive (again, the set of objects to pack is\n    known up-front independent of the reachability traversal).\n\nSo, this strikes a happy medium between not relying on reachability so much that\nwe run the risk of corrupting the repository, but relying on it enough that we\ncan aid in the delta selection process.\n\nBecause we reuse the same \"halt the traversal when encountering objects in kept\npacks\" mechanism, a lot of the patches are able to be reused. The structure of\nthe series is as follows:\n\n  - The first three patches introduce new infrastructure, and implement 'git\n    pack-objects --stdin-packs'.\n\n  - The next four patches introduce and use a kept-pack cache, which improves\n    the performance of 'git pack-objects --stdin-packs' substantially.\n\n  - The final patch implements 'git repack --geometric'.\n\nLet me know what you think of this new approach, and thanks in advance for your\nreview.\n\n[1]: https://lore.kernel.org/git/cover.1611098616.git.me@ttaylorr.com/\n\nThanks in advance for your review.\n\nJeff King (4):\n  p5303: add missing &&-chains\n  p5303: measure time to repack with keep\n  builtin/pack-objects.c: rewrite honor-pack-keep logic\n  packfile: add kept-pack cache for find_kept_pack_entry()\n\nTaylor Blau (4):\n  packfile: introduce 'find_kept_pack_entry()'\n  revision: learn '--no-kept-objects'\n  builtin/pack-objects.c: add '--stdin-packs' option\n  builtin/repack.c: add '--geometric' option\n\n Documentation/git-pack-objects.txt |  10 +\n Documentation/git-repack.txt       |  11 ++\n Documentation/rev-list-options.txt |   7 +\n builtin/pack-objects.c             | 301 +++++++++++++++++++++++------\n builtin/repack.c                   | 187 +++++++++++++++++-\n list-objects.c                     |   7 +\n object-store.h                     |  10 +\n packfile.c                         |  69 +++++++\n packfile.h                         |   2 +\n revision.c                         |  15 ++\n revision.h                         |   4 +\n t/perf/p5303-many-packs.sh         |  24 ++-\n t/t5300-pack-object.sh             |  97 ++++++++++\n t/t6114-keep-packs.sh              |  69 +++++++\n t/t7703-repack-geometric.sh        | 137 +++++++++++++\n 15 files changed, 889 insertions(+), 61 deletions(-)\n create mode 100755 t/t6114-keep-packs.sh\n create mode 100755 t/t7703-repack-geometric.sh\n\nRange-diff against v1:\n 1:  dc7fa4c7a6 !  1:  f7186147eb packfile: introduce 'find_kept_pack_entry()'\n    @@ Commit message\n         packfile: introduce 'find_kept_pack_entry()'\n     \n         Future callers will want a function to fill a 'struct pack_entry' for a\n    -    given object id but _only_ from its position in any kept pack(s). They\n    -    could accomplish this by calling 'find_pack_entry()' and checking\n    -    whether the found pack is kept or not, but this is insufficient, since\n    -    there may be duplicate objects (and the mru cache makes it unpredictable\n    -    which variant we'll get).\n    -\n    -    Teach this new function to treat the two different kinds of kept packs\n    -    (on disk ones with .keep files, as well as in-core ones which are set by\n    -    manually poking the 'pack_keep_in_core' bit) separately. This will\n    -    become important for callers that only want to respect a certain kind of\n    -    kept pack.\n    -\n    -    Introduce 'find_kept_pack_entry()' which behaves like\n    -    'find_pack_entry()', except that it skips over packs which are not\n    -    marked kept. Callers will be added in subsequent patches.\n    +    given object id but _only_ from its position in any kept pack(s).\n    +\n    +    In particular, an new 'git repack' mode which ensures the resulting\n    +    packs form a geometric progress by object count will mark packs that it\n    +    does not want to repack as \"kept in-core\", and it will want to halt a\n    +    reachability traversal as soon as it visits an object in any of the kept\n    +    packs. But, it does not want to halt the traversal at non-kept, or\n    +    .keep packs.\n    +\n    +    The obvious alternative is 'find_pack_entry()', but this doesn't quite\n    +    suffice since it only returns the first pack it finds, which may or may\n    +    not be kept (and the mru cache makes it unpredictable which one you'll\n    +    get if there are options).\n    +\n    +    Short of that, you could walk over all packs looking for the object in\n    +    each one, but it scales with the number of packs, which may be\n    +    prohibitive.\n    +\n    +    Introduce 'find_kept_pack_entry()', a function which is like\n    +    'find_pack_entry()', but only fills in objects in the kept packs.\n    +\n    +    Handle packs which have .keep files, as well as in-core kept packs\n    +    separately, since certain callers will want to distinguish one from the\n    +    other. (Though on-disk and in-core kept packs share the adjective\n    +    \"kept\", it is best to think of the two sets as independent.)\n    +\n    +    There is a gotcha when looking up objects that are duplicated in kept\n    +    and non-kept packs, particularly when the MIDX stores the non-kept\n    +    version and the caller asked for kept objects only. This could be\n    +    resolved by teaching the MIDX to resolve duplicates by always favoring\n    +    the kept pack (if one exists), but this breaks an assumption in existing\n    +    MIDXs, and so it would require a format change.\n    +\n    +    The benefit to changing the MIDX in this way is marginal, so we instead\n    +    have a more thorough check here which is explained with a comment.\n    +\n    +    Callers will be added in subsequent patches.\n     \n         Co-authored-by: Jeff King <peff@peff.net>\n         Signed-off-by: Jeff King <peff@peff.net>\n    @@ packfile.c: int find_pack_entry(struct repository *r, const struct object_id *oi\n      \n      \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n     -\t\tif (fill_midx_entry(r, oid, e, m))\n    -+\t\tif (!(fill_midx_entry(r, oid, e, m)))\n    ++\t\tif (!fill_midx_entry(r, oid, e, m))\n     +\t\t\tcontinue;\n     +\n     +\t\tif (!kept_only)\n 2:  4184529648 !  2:  ddc2896caa revision: learn '--no-kept-objects'\n    @@ Metadata\n      ## Commit message ##\n         revision: learn '--no-kept-objects'\n     \n    -    Some callers want to perform a reachability traversal that terminates\n    -    when an object is found in a kept pack. The closest existing option is\n    -    '--honor-pack-keep', but this isn't quite what we want. Instead of\n    -    halting the traversal midway through, a full traversal is always\n    -    performed, and the results are only trimmed afterwords.\n    +    A future caller will want to be able to perform a reachability traversal\n    +    which terminates when visiting an object found in a kept pack. The\n    +    closest existing option is '--honor-pack-keep', but this isn't quite\n    +    what we want. Instead of halting the traversal midway through, a full\n    +    traversal is always performed, and the results are only trimmed\n    +    afterwords.\n     \n         Besides needing to introduce a new flag (since culling results\n         post-facto can be different than halting the traversal as it's\n    @@ Commit message\n         of kept packs, if any, should stop a traversal. This can be useful for\n         callers that want to perform a reachability analysis, but want to leave\n         certain packs alone (for e.g., when doing a geometric repack that has\n    -    some \"large\" packs it wants to leave alone).\n    +    some \"large\" packs which are kept in-core that it wants to leave alone).\n     \n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n 3:  2da42e9ca2 <  -:  ---------- builtin/pack-objects.c: learn '--assume-kept-packs-closed'\n -:  ---------- >  3:  c96b1bf995 builtin/pack-objects.c: add '--stdin-packs' option\n 4:  26b46dff15 !  4:  a46b7002b4 p5303: add missing &&-chains\n    @@ Commit message\n     \n      ## t/perf/p5303-many-packs.sh ##\n     @@ t/perf/p5303-many-packs.sh: repack_into_n () {\n    - \tsed -n '1~5p' |\n    - \thead -n \"$1\" |\n    - \tperl -e 'print reverse <>' \\\n    --\t>pushes\n    -+\t>pushes &&\n    + \t\t\tpush @commits, $_ if $. % 5 == 1;\n    + \t\t}\n    + \t\tprint reverse @commits;\n    +-\t' \"$1\" >pushes\n    ++\t' \"$1\" >pushes &&\n      \n      \t# create base packfile\n      \thead -n 1 pushes |\n 5:  b3b2574d4d <  -:  ---------- p5303: measure time to repack with keep\n -:  ---------- >  5:  b5081c01b5 p5303: measure time to repack with keep\n 6:  4dd5076fcc !  6:  c3868c7df9 pack-objects: rewrite honor-pack-keep logic\n    @@ Metadata\n     Author: Jeff King <peff@peff.net>\n     \n      ## Commit message ##\n    -    pack-objects: rewrite honor-pack-keep logic\n    +    builtin/pack-objects.c: rewrite honor-pack-keep logic\n     \n         Now that we have find_kept_pack_entry(), we don't have to manually keep\n         hunting through every pack to find a possible \"kept\" duplicate of the\n    @@ Commit message\n         question. It might be worth having a similar optimized function to look\n         at only local packs.\n     \n    -    Here are the results from p5303 (measurements taken on git.git):\n    +    Here are the results from p5303 (measurements again taken on the\n    +    kernel):\n     \n    -      Test                               HEAD^                  HEAD\n    -      ------------------------------------------------------------------------------------\n    -      5303.5: repack (1)                 57.29(54.88+10.39)     56.87(54.63+10.48) -0.7%\n    -      5303.6: repack with keep (1)       1.25(1.19+0.05)        1.26(1.19+0.06) +0.8%\n    -      5303.10: repack (50)               89.71(132.78+6.14)     89.35(132.42+6.25) -0.4%\n    -      5303.11: repack with keep (50)     6.92(26.93+0.58)       6.73(26.61+0.59) -2.7%\n    -      5303.15: repack (1000)             217.14(493.76+15.29)   217.25(494.38+15.24) +0.1%\n    -      5303.16: repack with keep (1000)   209.46(387.83+8.42)    133.12(311.80+8.44) -36.4%\n    +      Test                                        HEAD^                    HEAD\n    +      -----------------------------------------------------------------------------------------------\n    +      5303.5: repack (1)                          57.42(54.88+10.64)       57.44(54.71+10.78) +0.0%\n    +      5303.6: repack with --stdin-packs (1)       0.01(0.01+0.00)          0.01(0.00+0.01) +0.0%\n    +      5303.10: repack (50)                        71.26(88.24+4.96)        71.32(88.38+4.90) +0.1%\n    +      5303.11: repack with --stdin-packs (50)     3.49(11.82+0.28)         3.43(11.81+0.22) -1.7%\n    +      5303.15: repack (1000)                      215.64(491.33+14.80)     215.59(493.75+14.62) -0.0%\n    +      5303.16: repack with --stdin-packs (1000)   198.79(380.51+7.97)      131.44(314.24+8.11) -33.9%\n     \n    -    So our case with many packs and a .keep is finally now faster than the\n    +    So our --stdin-packs case with many packs is now finally faster than the\n         non-keep case (because it gets the speed benefit of looking at fewer\n         objects, but not as big a penalty for looking at many packs).\n     \n 7:  182664e1a9 !  7:  f1c07324f6 packfile: add kept-pack cache for find_kept_pack_entry()\n    @@ Commit message\n     \n         Here are p5303 results (as always, measured against the kernel):\n     \n    -      Test                               HEAD^                  HEAD\n    -      ------------------------------------------------------------------------------------\n    -      5303.5: repack (1)                 56.87(54.63+10.48)     56.63(54.41+10.36) -0.4%\n    -      5303.6: repack with keep (1)       1.26(1.19+0.06)        1.25(1.19+0.05) -0.8%\n    -      5303.10: repack (50)               89.35(132.42+6.25)     89.49(132.31+6.31) +0.2%\n    -      5303.11: repack with keep (50)     6.73(26.61+0.59)       6.72(26.70+0.53) -0.1%\n    -      5303.15: repack (1000)             217.25(494.38+15.24)   218.69(495.62+14.99) +0.7%\n    -      5303.16: repack with keep (1000)   133.12(311.80+8.44)    128.79(306.96+8.55) -3.3%\n    +      Test                                        HEAD^                   HEAD\n    +      ----------------------------------------------------------------------------------------------\n    +      5303.5: repack (1)                          57.44(54.71+10.78)      57.06(54.29+10.96) -0.7%\n    +      5303.6: repack with --stdin-packs (1)       0.01(0.00+0.01)         0.01(0.01+0.00) +0.0%\n    +      5303.10: repack (50)                        71.32(88.38+4.90)       71.47(88.60+5.04) +0.2%\n    +      5303.11: repack with --stdin-packs (50)     3.43(11.81+0.22)        3.49(12.21+0.26) +1.7%\n    +      5303.15: repack (1000)                      215.59(493.75+14.62)    217.41(495.36+14.85) +0.8%\n    +      5303.16: repack with --stdin-packs (1000)   131.44(314.24+8.11)     126.75(309.88+8.09) -3.6%\n     \n         Signed-off-by: Jeff King <peff@peff.net>\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n    @@ builtin/pack-objects.c: static int want_found_object(const struct object_id *oid\n      \n      \t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n      \t\t\treturn 0;\n    +@@ builtin/pack-objects.c: static void read_packs_list_from_stdin(void)\n    + \t * an optimization during delta selection.\n    + \t */\n    + \trevs.no_kept_objects = 1;\n    +-\trevs.keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n    ++\trevs.keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n    + \trevs.blob_objects = 1;\n    + \trevs.tree_objects = 1;\n    + \trevs.tag_objects = 1;\n     \n      ## object-store.h ##\n     @@ object-store.h: static inline int pack_map_entry_cmp(const void *unused_cmp_data,\n    @@ packfile.c: static int find_one_pack_entry(struct repository *r,\n      \t\treturn 0;\n      \n      \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n    --\t\tif (!(fill_midx_entry(r, oid, e, m)))\n    +-\t\tif (!fill_midx_entry(r, oid, e, m))\n     -\t\t\tcontinue;\n     -\n     -\t\tif (!kept_only)\n 8:  6547c082f8 <  -:  ---------- builtin/pack-objects.c: teach '--keep-pack-stdin'\n 9:  a808fbdf31 <  -:  ---------- builtin/repack.c: extract loose object handling\n10:  f853087216 !  8:  d5561585c2 builtin/repack.c: add '--geometric' option\n    @@ builtin/repack.c: static void repack_promisor_objects(const struct pack_objects_\n     +\tgeometry = *geometry_p;\n     +\n     +\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n    ++\t\tif (!pack_kept_objects && p->pack_keep)\n    ++\t\t\tcontinue;\n    ++\n     +\t\tALLOC_GROW(geometry->pack,\n     +\t\t\t   geometry->pack_nr + 1,\n     +\t\t\t   geometry->pack_alloc);\n    @@ builtin/repack.c: static void repack_promisor_objects(const struct pack_objects_\n     +\tuint32_t split;\n     +\toff_t total_size = 0;\n     +\n    ++\tif (geometry->pack_nr <= 1) {\n    ++\t\tgeometry->split = geometry->pack_nr;\n    ++\t\treturn;\n    ++\t}\n    ++\n     +\tsplit = geometry->pack_nr - 1;\n     +\n     +\t/*\n    @@ builtin/repack.c: static void repack_promisor_objects(const struct pack_objects_\n     +\tgeometry->split = 0;\n     +}\n     +\n    - static void handle_loose_and_reachable(struct child_process *cmd,\n    - \t\t\t\t       const char *unpack_unreachable,\n    - \t\t\t\t       int pack_everything,\n    + int cmd_repack(int argc, const char **argv, const char *prefix)\n    + {\n    + \tstruct child_process cmd = CHILD_PROCESS_INIT;\n     @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n      \tstruct string_list names = STRING_LIST_INIT_DUP;\n      \tstruct string_list rollback = STRING_LIST_INIT_NODUP;\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \t\tdie(_(incremental_bitmap_conflict_error));\n      \n     +\tif (geometric_factor) {\n    ++\t\tif (pack_everything)\n    ++\t\t\tdie(_(\"--geometric is incompatible with -A, -a\"));\n     +\t\tinit_pack_geometry(&geometry);\n     +\t\tsplit_pack_geometry(geometry, geometric_factor);\n     +\t}\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \tpacktmp = mkpathdup(\"%s/.tmp-%d-pack\", packdir, (int)getpid());\n      \n     @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n    - \t\t\thandle_loose_and_reachable(&cmd, unpack_unreachable,\n    - \t\t\t\t\t\t   pack_everything,\n    - \t\t\t\t\t\t   keep_unreachable);\n    + \t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s\",\n    + \t\t\t     keep_pack_list.items[i].string);\n    + \tstrvec_push(&cmd.args, \"--non-empty\");\n    +-\tstrvec_push(&cmd.args, \"--all\");\n    +-\tstrvec_push(&cmd.args, \"--reflog\");\n    +-\tstrvec_push(&cmd.args, \"--indexed-objects\");\n    ++\tif (!geometry) {\n    ++\t\t/*\n    ++\t\t * 'git pack-objects' will up all objects loose or packed\n    ++\t\t * (either rolling them up or leaving them alone), so don't pass\n    ++\t\t * these options.\n    ++\t\t *\n    ++\t\t * The implementation of 'git pack-objects --stdin-packs'\n    ++\t\t * makes them redundant (and the two are incompatible).\n    ++\t\t */\n    ++\t\tstrvec_push(&cmd.args, \"--all\");\n    ++\t\tstrvec_push(&cmd.args, \"--reflog\");\n    ++\t\tstrvec_push(&cmd.args, \"--indexed-objects\");\n    ++\t}\n    + \tif (has_promisor_remote())\n    + \t\tstrvec_push(&cmd.args, \"--exclude-promisor-objects\");\n    + \tif (write_bitmaps > 0)\n    +@@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n    + \t\t\t\tstrvec_push(&cmd.env_array, \"GIT_REF_PARANOIA=1\");\n    + \t\t\t}\n    + \t\t}\n     +\t} else if (geometry) {\n    -+\t\tstrvec_push(&cmd.args, \"--keep-pack-stdin\");\n    -+\t\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n    -+\t\tstrvec_push(&cmd.args, \"--assume-kept-packs-closed\");\n    -+\t\tif (delete_redundant)\n    -+\t\t\thandle_loose_and_reachable(&cmd, unpack_unreachable,\n    -+\t\t\t\t\t\t   pack_everything,\n    -+\t\t\t\t\t\t   keep_unreachable);\n    ++\t\tstrvec_push(&cmd.args, \"--stdin-packs\");\n    ++\t\tstrvec_push(&cmd.args, \"--unpacked\");\n      \t} else {\n      \t\tstrvec_push(&cmd.args, \"--unpacked\");\n      \t\tstrvec_push(&cmd.args, \"--incremental\");\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n     +\tif (geometry) {\n     +\t\tFILE *in = xfdopen(cmd.in, \"w\");\n     +\t\t/*\n    -+\t\t * Tell 'git pack-objects' to avoid tampering with the structure\n    -+\t\t * with the packs that already form a geometric progression.\n    -+\t\t *\n    -+\t\t * Everything else will get picked up by the reachability walk.\n    ++\t\t * The resulting pack should contain all objects in packs that\n    ++\t\t * are going to be rolled up, but exclude objects in packs which\n    ++\t\t * are being left alone.\n     +\t\t */\n    -+\t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n    ++\t\tfor (i = 0; i < geometry->split; i++)\n     +\t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n    ++\t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n    ++\t\t\tfprintf(in, \"^%s\\n\", pack_basename(geometry->pack[i]));\n     +\t\tfclose(in);\n     +\t}\n     +\n    @@ t/t7703-repack-geometric.sh (new)\n     +objdir=.git/objects\n     +midx=$objdir/pack/multi-pack-index\n     +\n    ++test_expect_success '--geometric with no packs' '\n    ++\tgit init geometric &&\n    ++\ttest_when_finished \"rm -fr geometric\" &&\n    ++\t(\n    ++\t\tcd geometric &&\n    ++\n    ++\t\tgit repack --geometric 2 >out &&\n    ++\t\ttest_i18ngrep \"Nothing new to pack\" out\n    ++\t)\n    ++'\n    ++\n     +test_expect_success '--geometric with an intact progression' '\n     +\tgit init geometric &&\n     +\ttest_when_finished \"rm -fr geometric\" &&\n    @@ t/t7703-repack-geometric.sh (new)\n     +\t\ttest_commit_bulk --start=4 4 && # 12 objects\n     +\n     +\t\tfind $objdir/pack -name \"*.pack\" | sort >expect &&\n    -+\t\tGIT_TEST_MULTI_PACK_BITMAP=0 git repack --geometric 2 -d &&\n    ++\t\tgit repack --geometric 2 -d &&\n     +\t\tfind $objdir/pack -name \"*.pack\" | sort >actual &&\n     +\n     +\t\ttest_cmp expect actual\n    @@ t/t7703-repack-geometric.sh (new)\n     +\t\ttest_commit_bulk --start=7 8 && # 24 objects\n     +\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n     +\n    -+\t\tGIT_TEST_MULTI_PACK_BITMAP=0 git repack --geometric 2 -d &&\n    ++\t\tgit repack --geometric 2 -d &&\n     +\n     +\t\t# Three packs in total; two of the existing large ones, and one\n     +\t\t# new one.\n    @@ t/t7703-repack-geometric.sh (new)\n     +\n     +\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n     +\n    -+\t\tGIT_TEST_MULTI_PACK_BITMAP=0 git repack --geometric 2 -d &&\n    ++\t\tgit repack --geometric 2 -d &&\n     +\n     +\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n     +\t\tcomm -12 before after >untouched &&\n    @@ t/t7703-repack-geometric.sh (new)\n     +\t)\n     +'\n     +\n    ++test_expect_success '--geometric ignores kept packs' '\n    ++\tgit init geometric &&\n    ++\ttest_when_finished \"rm -fr geometric\" &&\n    ++\t(\n    ++\t\tcd geometric &&\n    ++\n    ++\t\ttest_commit kept && # 3 objects\n    ++\t\ttest_commit pack && # 3 objects\n    ++\n    ++\t\tKEPT=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n    ++\t\trefs/tags/kept\n    ++\t\tEOF\n    ++\t\t) &&\n    ++\t\tPACK=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n    ++\t\trefs/tags/pack\n    ++\t\t^refs/tags/kept\n    ++\t\tEOF\n    ++\t\t) &&\n    ++\n    ++\t\t# neither pack contains more than twice the number of objects in\n    ++\t\t# the other, so they should be combined. but, marking one as\n    ++\t\t# .kept on disk will \"freeze\" it, so the pack structure should\n    ++\t\t# remain unchanged.\n    ++\t\ttouch $objdir/pack/pack-$KEPT.keep &&\n    ++\n    ++\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n    ++\t\tgit repack --geometric 2 -d &&\n    ++\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n    ++\n    ++\t\t# both packs should still exist\n    ++\t\ttest_path_is_file $objdir/pack/pack-$KEPT.pack &&\n    ++\t\ttest_path_is_file $objdir/pack/pack-$PACK.pack &&\n    ++\n    ++\t\t# and no new packs should be created\n    ++\t\ttest_cmp before after &&\n    ++\n    ++\t\t# Passing --pack-kept-objects causes packs with a .keep file to\n    ++\t\t# be repacked, too.\n    ++\t\tgit repack --geometric 2 -d --pack-kept-objects &&\n    ++\n    ++\t\tfind $objdir/pack -name \"*.pack\" >after &&\n    ++\t\ttest_line_count = 1 after\n    ++\t)\n    ++'\n    ++\n     +test_done\n-- \n2.30.0.533.g2f8b6b552f.dirty\n"},{"id":"416128","messageId":"a46b7002b42a9f948154c4d276bf31a0a3e7b552.1612411124.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 4/8] p5303: add missing &&-chains","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:59:09Z","receivedAt":"2021-02-04T04:00:21Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nThese are in a helper function, so the usual chain-lint doesn't notice\nthem. This function is still not perfect, as it has some git invocations\non the left-hand-side of the pipe, but it's primary purpose is timing,\nnot finding bugs or correctness issues.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex ce0c42cc9f..d90d714923 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -28,11 +28,11 @@ repack_into_n () {\n \t\t\tpush @commits, $_ if $. % 5 == 1;\n \t\t}\n \t\tprint reverse @commits;\n-\t' \"$1\" >pushes\n+\t' \"$1\" >pushes &&\n \n \t# create base packfile\n \thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack\n+\tgit pack-objects --delta-base-offset --revs staging/pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n-- \n2.30.0.533.g2f8b6b552f.dirty\n\n"},{"id":"416129","messageId":"c96b1bf99582beefb96c3774b13a4f5a12fc61cc.1612411124.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:59:03Z","receivedAt":"2021-02-04T04:00:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In an upcoming commit, 'git repack' will want to create a pack comprised\nof all of the objects in some packs (the included packs) excluding any\nobjects in some other packs (the excluded packs).\n\nThis caller could iterate those packs themselves and feed the objects it\nfinds to 'git pack-objects' directly over stdin, but this approach has a\nfew downsides:\n\n  - It requires every caller that wants to drive 'git pack-objects' in\n    this way to implement pack iteration themselves. This forces the\n    caller to think about details like what order objects are fed to\n    pack-objects, which callers would likely rather not do.\n\n  - If the set of objects in included packs is large, it requires\n    sending a lot of data over a pipe, which is inefficient.\n\n  - The caller is forced to keep track of the excluded objects, too, and\n    make sure that it doesn't send any objects that appear in both\n    included and excluded packs.\n\nBut the biggest downside is the lack of a reachability traversal.\nBecause the caller passes in a list of objects directly, those objects\ndon't get a namehash assigned to them, which can have a negative impact\non the delta selection process, causing 'git pack-objects' to fail to\nfind good deltas even when they exist.\n\nThe caller could formulate a reachability traversal themselves, but the\nonly way to drive 'git pack-objects' in this way is to do a full\ntraversal, and then remove objects in the excluded packs after the\ntraversal is complete. This can be detrimental to callers who care\nabout performance, especially in repositories with many objects.\n\nIntroduce 'git pack-objects --stdin-packs' which remedies these four\nconcerns.\n\n'git pack-objects --stdin-packs' expects a list of pack names on stdin,\nwhere 'pack-xyz.pack' denotes that pack as included, and\n'^pack-xyz.pack' denotes it as excluded. The resulting pack includes all\nobjects that are present in at least one included pack, and aren't\npresent in any excluded pack.\n\nTo address the delta selection problem, 'git pack-objects --stdin-packs'\nworks as follows. First, it assembles a list of objects that it is going\nto pack, as above. Then, a reachability traversal is started, whose tips\nare any commits mentioned in included packs. Upon visiting an object, we\nfind its corresponding object_entry in the to_pack list, and set its\nnamehash parameter appropriately.\n\nTo avoid the traversal visiting more objects than it needs to, the\ntraversal is halted upon encountering an object which can be found in an\nexcluded pack (by marking the excluded packs as kept in-core, and\npassing --no-kept-objects=in-core to the revision machinery).\n\nThis can cause the traversal to halt early, for example if an object in\nan included pack is an ancestor of ones in excluded packs. But stopping\nearly is OK, since filling in the namehash fields of objects in the\nto_pack list is only additive (i.e., having it helps the delta selection\nprocess, but leaving it blank doesn't impact the correctness of the\nresulting pack).\n\nEven still, it is unlikely that this hurts us much in practice, since\nthe 'git repack --geometric' caller (which is introduced in a later\ncommit) marks small packs as included, and large ones as excluded.\nDuring ordinary use, the small packs usually represent pushes after a\nlarge repack, and so are unlikely to be ancestors of objects that\nalready exist in the repository.\n\n(I found it convenient while developing this patch to have 'git\npack-objects' report the number of objects which were visited and got\ntheir namehash fields filled in during traversal. This is also included\nin the below patch via trace2 data lines).\n\nSuggested-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  10 ++\n builtin/pack-objects.c             | 176 ++++++++++++++++++++++++++++-\n t/t5300-pack-object.sh             |  97 ++++++++++++++++\n 3 files changed, 281 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex 54d715ead1..92733f6bf5 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -85,6 +85,16 @@ base-name::\n \treference was included in the resulting packfile.  This\n \tcan be useful to send new tags to native Git clients.\n \n+--stdin-packs::\n+\tRead the basenames of packfiles from the standard input, instead\n+\tof object names or revision arguments. The resulting pack\n+\tcontains all objects listed in the included packs (those not\n+\tbeginning with `^`), excluding any objects listed in the\n+\texcluded packs (beginning with `^`).\n++\n+Incompatible with `--revs`, or options that imply `--revs` (such as\n+`--all`), with the exception of `--unpacked`, which is compatible.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 13cde5896a..6d19eb000a 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -2979,6 +2979,164 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \treturn git_default_config(k, v, cb);\n }\n \n+static int stdin_packs_found_nr;\n+static int stdin_packs_hints_nr;\n+\n+static int add_object_entry_from_pack(const struct object_id *oid,\n+\t\t\t\t      struct packed_git *p,\n+\t\t\t\t      uint32_t pos,\n+\t\t\t\t      void *_data)\n+{\n+\tstruct rev_info *revs = _data;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\toff_t ofs;\n+\tenum object_type type;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tofs = nth_packed_object_offset(p, pos);\n+\n+\toi.typep = &type;\n+\tif (packed_object_info(the_repository, p, ofs, &oi) < 0)\n+\t\tdie(_(\"could not get type of object %s in pack %s\"),\n+\t\t    oid_to_hex(oid), p->pack_name);\n+\telse if (type == OBJ_COMMIT) {\n+\t\t/*\n+\t\t * commits in included packs are used as starting points for the\n+\t\t * subsequent revision walk\n+\t\t */\n+\t\tadd_pending_oid(revs, NULL, oid, 0);\n+\t}\n+\n+\tif (have_duplicate_entry(oid, 0))\n+\t\treturn 0;\n+\n+\tif (!want_object_in_pack(oid, 0, &p, &ofs))\n+\t\treturn 0;\n+\n+\tstdin_packs_found_nr++;\n+\n+\tcreate_object_entry(oid, type, 0, 0, 0, p, ofs);\n+\n+\treturn 0;\n+}\n+\n+static void show_commit_pack_hint(struct commit *commit, void *_data)\n+{\n+}\n+\n+static void show_object_pack_hint(struct object *object, const char *name,\n+\t\t\t\t  void *_data)\n+{\n+\tstruct object_entry *oe = packlist_find(&to_pack, &object->oid);\n+\tif (!oe)\n+\t\treturn;\n+\n+\t/*\n+\t * Our 'to_pack' list was constructed by iterating all objects packed in\n+\t * included packs, and so doesn't have a non-zero hash field that you\n+\t * would typically pick up during a reachability traversal.\n+\t *\n+\t * Make a best-effort attempt to fill in the ->hash and ->no_try_delta\n+\t * here using a now in order to perhaps improve the delta selection\n+\t * process.\n+\t */\n+\toe->hash = pack_name_hash(name);\n+\toe->no_try_delta = name && no_try_delta(name);\n+\n+\tstdin_packs_hints_nr++;\n+}\n+\n+static void read_packs_list_from_stdin(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list include_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list exclude_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list_item *item = NULL;\n+\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\t/*\n+\t * Use a revision walk to fill in the namehash of objects in the include\n+\t * packs. To save time, we'll avoid traversing through objects that are\n+\t * in excluded packs.\n+\t *\n+\t * That may cause us to avoid populating all of the namehash fields of\n+\t * all included objects, but our goal is best-effort, since this is only\n+\t * an optimization during delta selection.\n+\t */\n+\trevs.no_kept_objects = 1;\n+\trevs.keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\trevs.blob_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.tag_objects = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '^')\n+\t\t\tstring_list_append(&exclude_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&include_packs, buf.buf);\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&include_packs);\n+\tstring_list_sort(&exclude_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\n+\t\titem = string_list_lookup(&include_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&exclude_packs, pack_name);\n+\n+\t\tif (item)\n+\t\t\titem->util = p;\n+\t}\n+\n+\t/*\n+\t * First handle all of the excluded packs, marking them as kept in-core\n+\t * so that later calls to add_object_entry() discards any objects that\n+\t * are also found in excluded packs.\n+\t */\n+\tfor_each_string_list_item(item, &exclude_packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = 1;\n+\t}\n+\tfor_each_string_list_item(item, &include_packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tfor_each_object_in_pack(p,\n+\t\t\t\t\tadd_object_entry_from_pack,\n+\t\t\t\t\t&revs,\n+\t\t\t\t\tFOR_EACH_OBJECT_PACK_ORDER);\n+\t}\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\ttraverse_commit_list(&revs,\n+\t\t\t     show_commit_pack_hint,\n+\t\t\t     show_object_pack_hint,\n+\t\t\t     NULL);\n+\n+\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_found\",\n+\t\t\t   stdin_packs_found_nr);\n+\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_hints\",\n+\t\t\t   stdin_packs_hints_nr);\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&include_packs, 0);\n+\tstring_list_clear(&exclude_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3482,6 +3640,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tstruct strvec rp = STRVEC_INIT;\n \tint rev_list_unpacked = 0, rev_list_all = 0, rev_list_reflog = 0;\n \tint rev_list_index = 0;\n+\tint stdin_packs = 0;\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tstruct option pack_objects_options[] = {\n \t\tOPT_SET_INT('q', \"quiet\", &progress,\n@@ -3532,6 +3691,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_SET_INT_F(0, \"indexed-objects\", &rev_list_index,\n \t\t\t      N_(\"include objects referred to by the index\"),\n \t\t\t      1, PARSE_OPT_NONEG),\n+\t\tOPT_BOOL(0, \"stdin-packs\", &stdin_packs,\n+\t\t\t N_(\"read packs from stdin\")),\n \t\tOPT_BOOL(0, \"stdout\", &pack_to_stdout,\n \t\t\t N_(\"output pack to stdout\")),\n \t\tOPT_BOOL(0, \"include-tag\", &include_tag,\n@@ -3636,7 +3797,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tuse_internal_rev_list = 1;\n \t\tstrvec_push(&rp, \"--indexed-objects\");\n \t}\n-\tif (rev_list_unpacked) {\n+\tif (rev_list_unpacked && !stdin_packs) {\n \t\tuse_internal_rev_list = 1;\n \t\tstrvec_push(&rp, \"--unpacked\");\n \t}\n@@ -3681,8 +3842,13 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (filter_options.choice) {\n \t\tif (!pack_to_stdout)\n \t\t\tdie(_(\"cannot use --filter without --stdout\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --filter with --stdin-packs\"));\n \t}\n \n+\tif (stdin_packs && use_internal_rev_list)\n+\t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -3741,7 +3907,13 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (progress)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n-\tif (!use_internal_rev_list)\n+\tif (stdin_packs) {\n+\t\t/* avoids adding objects in excluded packs */\n+\t\tignore_packed_keep_in_core = 1;\n+\t\tread_packs_list_from_stdin();\n+\t\tif (rev_list_unpacked)\n+\t\t\tadd_unreachable_loose_objects();\n+\t} else if (!use_internal_rev_list)\n \t\tread_object_list_from_stdin();\n \telse {\n \t\tget_object_list(rp.nr, rp.v);\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex 392201cabd..7138a54595 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -532,4 +532,101 @@ test_expect_success 'prefetch objects' '\n \ttest_line_count = 1 donelines\n '\n \n+test_expect_success 'setup for --stdin-packs tests' '\n+\tgit init stdin-packs &&\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\ttest_commit A &&\n+\t\ttest_commit B &&\n+\t\ttest_commit C &&\n+\n+\t\tfor id in A B C\n+\t\tdo\n+\t\t\tgit pack-objects .git/objects/pack/pack-$id \\\n+\t\t\t\t--incremental --revs <<-EOF\n+\t\t\trefs/tags/$id\n+\t\t\tEOF\n+\t\tdone &&\n+\n+\t\tls -la .git/objects/pack\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs with excluded packs' '\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\tPACK_A=\"$(basename .git/objects/pack/pack-A-*.pack)\" &&\n+\t\tPACK_B=\"$(basename .git/objects/pack/pack-B-*.pack)\" &&\n+\t\tPACK_C=\"$(basename .git/objects/pack/pack-C-*.pack)\" &&\n+\n+\t\tgit pack-objects test --stdin-packs <<-EOF &&\n+\t\t$PACK_A\n+\t\t^$PACK_B\n+\t\t$PACK_C\n+\t\tEOF\n+\n+\t\t(\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-A-*.idx) &&\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-C-*.idx)\n+\t\t) >expect.raw &&\n+\t\tgit show-index <$(ls test-*.idx) >actual.raw &&\n+\n+\t\tcut -d\" \" -f2 <expect.raw | sort >expect &&\n+\t\tcut -d\" \" -f2 <actual.raw | sort >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs is incompatible with --filter' '\n+\t(\n+\t\tcd stdin-packs &&\n+\t\ttest_must_fail git pack-objects --stdin-packs --stdout \\\n+\t\t\t--filter=blob:none </dev/null 2>err &&\n+\t\ttest_i18ngrep \"cannot use --filter with --stdin-packs\" err\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs is incompatible with --revs' '\n+\t(\n+\t\tcd stdin-packs &&\n+\t\ttest_must_fail git pack-objects --stdin-packs --revs out \\\n+\t\t\t</dev/null 2>err &&\n+\t\ttest_i18ngrep \"cannot use internal rev list with --stdin-packs\" err\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs with loose objects' '\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\tPACK_A=\"$(basename .git/objects/pack/pack-A-*.pack)\" &&\n+\t\tPACK_B=\"$(basename .git/objects/pack/pack-B-*.pack)\" &&\n+\t\tPACK_C=\"$(basename .git/objects/pack/pack-C-*.pack)\" &&\n+\n+\t\ttest_commit D && # loose\n+\n+\t\tgit pack-objects test2 --stdin-packs --unpacked <<-EOF &&\n+\t\t$PACK_A\n+\t\t^$PACK_B\n+\t\t$PACK_C\n+\t\tEOF\n+\n+\t\t(\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-A-*.idx) &&\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-C-*.idx) &&\n+\t\t\tgit rev-list --objects --no-object-names \\\n+\t\t\t\trefs/tags/C..refs/tags/D\n+\n+\t\t) >expect.raw &&\n+\t\tls -la . &&\n+\t\tgit show-index <$(ls test2-*.idx) >actual.raw &&\n+\n+\t\tcut -d\" \" -f2 <expect.raw | sort >expect &&\n+\t\tcut -d\" \" -f2 <actual.raw | sort >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_done\n-- \n2.30.0.533.g2f8b6b552f.dirty\n\n"},{"id":"416130","messageId":"ddc2896caa13b9f1cdccb2f0a5892143fa98237c.1612411123.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 2/8] revision: learn '--no-kept-objects'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:58:57Z","receivedAt":"2021-02-04T04:00:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A future caller will want to be able to perform a reachability traversal\nwhich terminates when visiting an object found in a kept pack. The\nclosest existing option is '--honor-pack-keep', but this isn't quite\nwhat we want. Instead of halting the traversal midway through, a full\ntraversal is always performed, and the results are only trimmed\nafterwords.\n\nBesides needing to introduce a new flag (since culling results\npost-facto can be different than halting the traversal as it's\nhappening), there is an additional wrinkle handling the distinction\nin-core and on-disk kept packs. That is: what kinds of kept pack should\nstop the traversal?\n\nIntroduce '--no-kept-objects[=<on-disk|in-core>]' to specify which kinds\nof kept packs, if any, should stop a traversal. This can be useful for\ncallers that want to perform a reachability analysis, but want to leave\ncertain packs alone (for e.g., when doing a geometric repack that has\nsome \"large\" packs which are kept in-core that it wants to leave alone).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/rev-list-options.txt |  7 +++\n list-objects.c                     |  7 +++\n revision.c                         | 15 +++++++\n revision.h                         |  4 ++\n t/t6114-keep-packs.sh              | 69 ++++++++++++++++++++++++++++++\n 5 files changed, 102 insertions(+)\n create mode 100755 t/t6114-keep-packs.sh\n\ndiff --git a/Documentation/rev-list-options.txt b/Documentation/rev-list-options.txt\nindex 96cc89d157..f611832277 100644\n--- a/Documentation/rev-list-options.txt\n+++ b/Documentation/rev-list-options.txt\n@@ -861,6 +861,13 @@ ifdef::git-rev-list[]\n \tOnly useful with `--objects`; print the object IDs that are not\n \tin packs.\n \n+--no-kept-objects[=<kind>]::\n+\tHalts the traversal as soon as an object in a kept pack is\n+\tfound. If `<kind>` is `on-disk`, only packs with a corresponding\n+\t`*.keep` file are ignored. If `<kind>` is `in-core`, only packs\n+\twith their in-core kept state set are ignored. Otherwise, both\n+\tkinds of kept packs are ignored.\n+\n --object-names::\n \tOnly useful with `--objects`; print the names of the object IDs\n \tthat are found. This is the default behavior.\ndiff --git a/list-objects.c b/list-objects.c\nindex e19589baa0..b06c3bfeba 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -338,6 +338,13 @@ static void traverse_trees_and_blobs(struct traversal_context *ctx,\n \t\t\tctx->show_object(obj, name, ctx->show_data);\n \t\t\tcontinue;\n \t\t}\n+\t\tif (ctx->revs->no_kept_objects) {\n+\t\t\tstruct pack_entry e;\n+\t\t\tif (find_kept_pack_entry(ctx->revs->repo, &obj->oid,\n+\t\t\t\t\t\t ctx->revs->keep_pack_cache_flags,\n+\t\t\t\t\t\t &e))\n+\t\t\t\tcontinue;\n+\t\t}\n \t\tif (!path)\n \t\t\tpath = \"\";\n \t\tif (obj->type == OBJ_TREE) {\ndiff --git a/revision.c b/revision.c\nindex fbc3e607fd..4c5adb90b1 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -2336,6 +2336,16 @@ static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n \t\trevs->unpacked = 1;\n \t} else if (starts_with(arg, \"--unpacked=\")) {\n \t\tdie(_(\"--unpacked=<packfile> no longer supported\"));\n+\t} else if (!strcmp(arg, \"--no-kept-objects\")) {\n+\t\trevs->no_kept_objects = 1;\n+\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t} else if (skip_prefix(arg, \"--no-kept-objects=\", &optarg)) {\n+\t\trevs->no_kept_objects = 1;\n+\t\tif (!strcmp(optarg, \"in-core\"))\n+\t\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\tif (!strcmp(optarg, \"on-disk\"))\n+\t\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n \t} else if (!strcmp(arg, \"-r\")) {\n \t\trevs->diff = 1;\n \t\trevs->diffopt.flags.recursive = 1;\n@@ -3797,6 +3807,11 @@ enum commit_action get_commit_action(struct rev_info *revs, struct commit *commi\n \t\treturn commit_ignore;\n \tif (revs->unpacked && has_object_pack(&commit->object.oid))\n \t\treturn commit_ignore;\n+\tif (revs->no_kept_objects) {\n+\t\tif (has_object_kept_pack(&commit->object.oid,\n+\t\t\t\t\t revs->keep_pack_cache_flags))\n+\t\t\treturn commit_ignore;\n+\t}\n \tif (commit->object.flags & UNINTERESTING)\n \t\treturn commit_ignore;\n \tif (revs->line_level_traverse && !want_ancestry(revs)) {\ndiff --git a/revision.h b/revision.h\nindex e6be3c845e..a20a530d52 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -148,6 +148,7 @@ struct rev_info {\n \t\t\tedge_hint_aggressive:1,\n \t\t\tlimited:1,\n \t\t\tunpacked:1,\n+\t\t\tno_kept_objects:1,\n \t\t\tboundary:2,\n \t\t\tcount:1,\n \t\t\tleft_right:1,\n@@ -317,6 +318,9 @@ struct rev_info {\n \t * This is loaded from the commit-graph being used.\n \t */\n \tstruct bloom_filter_settings *bloom_filter_settings;\n+\n+\t/* misc. flags related to '--no-kept-objects' */\n+\tunsigned keep_pack_cache_flags;\n };\n \n int ref_excluded(struct string_list *, const char *path);\ndiff --git a/t/t6114-keep-packs.sh b/t/t6114-keep-packs.sh\nnew file mode 100755\nindex 0000000000..9239d8aa46\n--- /dev/null\n+++ b/t/t6114-keep-packs.sh\n@@ -0,0 +1,69 @@\n+#!/bin/sh\n+\n+test_description='rev-list with .keep packs'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup' '\n+\ttest_commit loose &&\n+\ttest_commit packed &&\n+\ttest_commit kept &&\n+\n+\tKEPT_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/kept\n+\t^refs/tags/packed\n+\tEOF\n+\t) &&\n+\tMISC_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/packed\n+\t^refs/tags/loose\n+\tEOF\n+\t) &&\n+\n+\ttouch .git/objects/pack/pack-$KEPT_PACK.keep\n+'\n+\n+rev_list_objects () {\n+\tgit rev-list \"$@\" >out &&\n+\tsort out\n+}\n+\n+idx_objects () {\n+\tgit show-index <$1 >expect-idx &&\n+\tcut -d\" \" -f2 <expect-idx | sort\n+}\n+\n+test_expect_success '--no-kept-objects excludes trees and blobs in .keep packs' '\n+\trev_list_objects --objects --all --no-object-names >kept &&\n+\trev_list_objects --objects --all --no-object-names --no-kept-objects >no-kept &&\n+\n+\tidx_objects .git/objects/pack/pack-$KEPT_PACK.idx >expect &&\n+\tcomm -3 kept no-kept >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success '--no-kept-objects excludes kept non-MIDX object' '\n+\ttest_config core.multiPackIndex true &&\n+\n+\t# Create a pack with just the commit object in pack, and do not mark it\n+\t# as kept (even though it appears in $KEPT_PACK, which does have a .keep\n+\t# file).\n+\tMIDX_PACK=$(git pack-objects .git/objects/pack/pack <<-EOF\n+\t$(git rev-parse kept)\n+\tEOF\n+\t) &&\n+\n+\t# Write a MIDX containing all packs, but use the version of the commit\n+\t# at \"kept\" in a non-kept pack by touching $MIDX_PACK.\n+\ttouch .git/objects/pack/pack-$MIDX_PACK.pack &&\n+\tgit multi-pack-index write &&\n+\n+\trev_list_objects --objects --no-object-names --no-kept-objects HEAD >actual &&\n+\t(\n+\t\tidx_objects .git/objects/pack/pack-$MISC_PACK.idx &&\n+\t\tgit rev-list --objects --no-object-names refs/tags/loose\n+\t) | sort >expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n-- \n2.30.0.533.g2f8b6b552f.dirty\n\n"},{"id":"416131","messageId":"f7186147ebb0b2d01d8f1e0f742f367204d7d9c9.1612411123.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 1/8] packfile: introduce 'find_kept_pack_entry()'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:58:50Z","receivedAt":"2021-02-04T04:00:47Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Future callers will want a function to fill a 'struct pack_entry' for a\ngiven object id but _only_ from its position in any kept pack(s).\n\nIn particular, an new 'git repack' mode which ensures the resulting\npacks form a geometric progress by object count will mark packs that it\ndoes not want to repack as \"kept in-core\", and it will want to halt a\nreachability traversal as soon as it visits an object in any of the kept\npacks. But, it does not want to halt the traversal at non-kept, or\n.keep packs.\n\nThe obvious alternative is 'find_pack_entry()', but this doesn't quite\nsuffice since it only returns the first pack it finds, which may or may\nnot be kept (and the mru cache makes it unpredictable which one you'll\nget if there are options).\n\nShort of that, you could walk over all packs looking for the object in\neach one, but it scales with the number of packs, which may be\nprohibitive.\n\nIntroduce 'find_kept_pack_entry()', a function which is like\n'find_pack_entry()', but only fills in objects in the kept packs.\n\nHandle packs which have .keep files, as well as in-core kept packs\nseparately, since certain callers will want to distinguish one from the\nother. (Though on-disk and in-core kept packs share the adjective\n\"kept\", it is best to think of the two sets as independent.)\n\nThere is a gotcha when looking up objects that are duplicated in kept\nand non-kept packs, particularly when the MIDX stores the non-kept\nversion and the caller asked for kept objects only. This could be\nresolved by teaching the MIDX to resolve duplicates by always favoring\nthe kept pack (if one exists), but this breaks an assumption in existing\nMIDXs, and so it would require a format change.\n\nThe benefit to changing the MIDX in this way is marginal, so we instead\nhave a more thorough check here which is explained with a comment.\n\nCallers will be added in subsequent patches.\n\nCo-authored-by: Jeff King <peff@peff.net>\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n packfile.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++-----\n packfile.h |  6 +++++\n 2 files changed, 65 insertions(+), 5 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex 4b938b4372..5f35cfe788 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2031,7 +2031,10 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static int find_one_pack_entry(struct repository *r,\n+\t\t\t       const struct object_id *oid,\n+\t\t\t       struct pack_entry *e,\n+\t\t\t       int kept_only)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2041,26 +2044,77 @@ int find_pack_entry(struct repository *r, const struct object_id *oid, struct pa\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (fill_midx_entry(r, oid, e, m))\n+\t\tif (!fill_midx_entry(r, oid, e, m))\n+\t\t\tcontinue;\n+\n+\t\tif (!kept_only)\n+\t\t\treturn 1;\n+\n+\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n+\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n-\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\treturn 1;\n+\t\tif (p->multi_pack_index && !kept_only) {\n+\t\t\t/*\n+\t\t\t * If this pack is covered by the MIDX, we'd have found\n+\t\t\t * the object already in the loop above if it was here,\n+\t\t\t * so don't bother looking.\n+\t\t\t *\n+\t\t\t * The exception is if we are looking only at kept\n+\t\t\t * packs. An object can be present in two packs covered\n+\t\t\t * by the MIDX, one kept and one not-kept. And as the\n+\t\t\t * MIDX points to only one copy of each object, it might\n+\t\t\t * have returned only the non-kept version above. We\n+\t\t\t * have to check again to be thorough.\n+\t\t\t */\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!kept_only ||\n+\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n+\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n+\t\t\tif (fill_pack_entry(oid, e, p)) {\n+\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\t\treturn 1;\n+\t\t\t}\n \t\t}\n \t}\n \treturn 0;\n }\n \n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+{\n+\treturn find_one_pack_entry(r, oid, e, 0);\n+}\n+\n+int find_kept_pack_entry(struct repository *r,\n+\t\t\t const struct object_id *oid,\n+\t\t\t unsigned flags,\n+\t\t\t struct pack_entry *e)\n+{\n+\t/*\n+\t * Load all packs, including midx packs, since our \"kept\" strategy\n+\t * relies on that. We're relying on the side effect of it setting up\n+\t * r->objects->packed_git, which is a little ugly.\n+\t */\n+\tget_all_packs(r);\n+\treturn find_one_pack_entry(r, oid, e, flags);\n+}\n+\n int has_object_pack(const struct object_id *oid)\n {\n \tstruct pack_entry e;\n \treturn find_pack_entry(the_repository, oid, &e);\n }\n \n+int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n+{\n+\tstruct pack_entry e;\n+\treturn find_kept_pack_entry(the_repository, oid, flags, &e);\n+}\n+\n int has_pack_index(const unsigned char *sha1)\n {\n \tstruct stat st;\ndiff --git a/packfile.h b/packfile.h\nindex a58fc738e0..624327f64d 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -161,13 +161,19 @@ int packed_object_info(struct repository *r,\n void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n \n+#define ON_DISK_KEEP_PACKS 1\n+#define IN_CORE_KEEP_PACKS 2\n+#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n+\n /*\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\n  */\n int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e);\n+int find_kept_pack_entry(struct repository *r, const struct object_id *oid, unsigned flags, struct pack_entry *e);\n \n int has_object_pack(const struct object_id *oid);\n+int has_object_kept_pack(const struct object_id *oid, unsigned flags);\n \n int has_pack_index(const unsigned char *sha1);\n \n-- \n2.30.0.533.g2f8b6b552f.dirty\n\n"},{"id":"416132","messageId":"b5081c01b53beb568ef2e59956d25b36be9f24d0.1612411124.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 5/8] p5303: measure time to repack with keep","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:59:13Z","receivedAt":"2021-02-04T04:01:03Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nThis is the same as the regular repack test, except that we mark the\nsingle base pack as \"kept\" and use --assume-kept-packs-closed. The\ntheory is that this should be faster than the normal repack, because\nwe'll have fewer objects to traverse and process.\n\nHere are some timings on a recent clone of the kernel. In the\nsingle-pack case, there is nothing do since there are no non-excluded\npacks:\n\n  5303.5: repack (1)                          57.42(54.88+10.64)\n  5303.6: repack with --stdin-packs (1)       0.01(0.01+0.00)\n\nand in the 50-pack case, it is much faster to use `--stdin-packs`, since\nwe avoid having to consider any objects in the excluded pack:\n\n  5303.10: repack (50)                        71.26(88.24+4.96)\n  5303.11: repack with --stdin-packs (50)     3.49(11.82+0.28)\n\nbut our improvements vanish as we approach 1000 packs.\n\n  5303.15: repack (1000)                      215.64(491.33+14.80)\n  5303.16: repack with --stdin-packs (1000)   198.79(380.51+7.97)\n\nThat's because the code paths around handling .keep files are known to\nscale badly; they look in every single pack file to find each object.\nOur solution to that was to notice that most repos don't have keep\nfiles, and to make that case a fast path. But as soon as you add a\nsingle .keep, that part of pack-objects slows down again (even if we\nhave fewer objects total to look at).\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 22 ++++++++++++++++++++--\n 1 file changed, 20 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex d90d714923..b76a6efe00 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -31,8 +31,11 @@ repack_into_n () {\n \t' \"$1\" >pushes &&\n \n \t# create base packfile\n-\thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack &&\n+\tbase_pack=$(\n+\t\thead -n 1 pushes |\n+\t\tgit pack-objects --delta-base-offset --revs staging/pack\n+\t) &&\n+\ttest_export base_pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n@@ -49,6 +52,12 @@ repack_into_n () {\n \t\tlast=$rev\n \tdone <pushes &&\n \n+\t(\n+\t\tfind staging -type f -name 'pack-*.pack' |\n+\t\t\txargs -n 1 basename | grep -v \"$base_pack\" &&\n+\t\tprintf \"^pack-%s.pack\\n\" $base_pack\n+\t) >stdin.packs\n+\n \t# and install the whole thing\n \trm -f .git/objects/pack/* &&\n \tmv staging/* .git/objects/pack/\n@@ -91,6 +100,15 @@ do\n \t\t  --reflog --indexed-objects --delta-base-offset \\\n \t\t  --stdout </dev/null >/dev/null\n \t'\n+\n+\ttest_perf \"repack with --stdin-packs ($nr_packs)\" '\n+\t\tgit pack-objects \\\n+\t\t  --keep-true-parents \\\n+\t\t  --stdin-packs \\\n+\t\t  --non-empty \\\n+\t\t  --delta-base-offset \\\n+\t\t  --stdout <stdin.packs >/dev/null\n+\t'\n done\n \n # Measure pack loading with 10,000 packs.\n-- \n2.30.0.533.g2f8b6b552f.dirty\n\n"},{"id":"416133","messageId":"f1c07324f62cf4d087c41165cefed98f554cfd78.1612411124.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:59:21Z","receivedAt":"2021-02-04T04:01:23Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nIn a recent patch we added a function 'find_kept_pack_entry()' to look\nfor an object only among kept packs.\n\nWhile this function avoids doing any lookup work in non-kept packs, it\nis still linear in the number of packs, since we have to traverse the\nlinked list of packs once per object. Let's cache a reduced version of\nthat list to save us time.\n\nNote that this cache will last the lifetime of the program. We could\ninvalidate it on reprepare_packed_git(), but there's not much point in\nbeing rigorous here:\n\n  - we might already fail to notice new .keep packs showing up after the\n    program starts. We only reprepare_packed_git() when we fail to find\n    an object. But adding a new pack won't cause that to happen.\n    Somebody repacking could add a new pack and delete an old one, but\n    most of the time we'd have a descriptor or mmap open to the old\n    pack anyway, so we might not even notice.\n\n  - in pack-objects we already cache the .keep state at startup, since\n    56dfeb6263 (pack-objects: compute local/ignore_pack_keep early,\n    2016-07-29). So this is just extending that concept further.\n\n  - we don't have to worry about any packed_git being removed; we always\n    keep the old structs around, even after reprepare_packed_git()\n\nHere are p5303 results (as always, measured against the kernel):\n\n  Test                                        HEAD^                   HEAD\n  ----------------------------------------------------------------------------------------------\n  5303.5: repack (1)                          57.44(54.71+10.78)      57.06(54.29+10.96) -0.7%\n  5303.6: repack with --stdin-packs (1)       0.01(0.00+0.01)         0.01(0.01+0.00) +0.0%\n  5303.10: repack (50)                        71.32(88.38+4.90)       71.47(88.60+5.04) +0.2%\n  5303.11: repack with --stdin-packs (50)     3.43(11.81+0.22)        3.49(12.21+0.26) +1.7%\n  5303.15: repack (1000)                      215.59(493.75+14.62)    217.41(495.36+14.85) +0.8%\n  5303.16: repack with --stdin-packs (1000)   131.44(314.24+8.11)     126.75(309.88+8.09) -3.6%\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c |   6 +--\n object-store.h         |  10 ++++\n packfile.c             | 103 +++++++++++++++++++++++------------------\n packfile.h             |   4 --\n revision.c             |   8 ++--\n 5 files changed, 76 insertions(+), 55 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex fbd7b54d70..b2ba5aa14f 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1225,9 +1225,9 @@ static int want_found_object(const struct object_id *oid, int exclude,\n \t\t */\n \t\tunsigned flags = 0;\n \t\tif (ignore_packed_keep_on_disk)\n-\t\t\tflags |= ON_DISK_KEEP_PACKS;\n+\t\t\tflags |= CACHE_ON_DISK_KEEP_PACKS;\n \t\tif (ignore_packed_keep_in_core)\n-\t\t\tflags |= IN_CORE_KEEP_PACKS;\n+\t\t\tflags |= CACHE_IN_CORE_KEEP_PACKS;\n \n \t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n \t\t\treturn 0;\n@@ -3089,7 +3089,7 @@ static void read_packs_list_from_stdin(void)\n \t * an optimization during delta selection.\n \t */\n \trevs.no_kept_objects = 1;\n-\trevs.keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\trevs.keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n \trevs.blob_objects = 1;\n \trevs.tree_objects = 1;\n \trevs.tag_objects = 1;\ndiff --git a/object-store.h b/object-store.h\nindex c4fc9dd74e..4cbe8eae3c 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -105,6 +105,14 @@ static inline int pack_map_entry_cmp(const void *unused_cmp_data,\n \treturn strcmp(pg1->pack_name, key ? key : pg2->pack_name);\n }\n \n+#define CACHE_ON_DISK_KEEP_PACKS 1\n+#define CACHE_IN_CORE_KEEP_PACKS 2\n+\n+struct kept_pack_cache {\n+\tstruct packed_git **packs;\n+\tunsigned flags;\n+};\n+\n struct raw_object_store {\n \t/*\n \t * Set of all object directories; the main directory is first (and\n@@ -150,6 +158,8 @@ struct raw_object_store {\n \t/* A most-recently-used ordered version of the packed_git list. */\n \tstruct list_head packed_git_mru;\n \n+\tstruct kept_pack_cache *kept_pack_cache;\n+\n \t/*\n \t * A map of packfiles to packed_git structs for tracking which\n \t * packs have been loaded already.\ndiff --git a/packfile.c b/packfile.c\nindex 5f35cfe788..2a139c907b 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2031,10 +2031,7 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int find_one_pack_entry(struct repository *r,\n-\t\t\t       const struct object_id *oid,\n-\t\t\t       struct pack_entry *e,\n-\t\t\t       int kept_only)\n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2044,49 +2041,64 @@ static int find_one_pack_entry(struct repository *r,\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (!fill_midx_entry(r, oid, e, m))\n-\t\t\tcontinue;\n-\n-\t\tif (!kept_only)\n-\t\t\treturn 1;\n-\n-\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n-\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n+\t\tif (fill_midx_entry(r, oid, e, m))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (p->multi_pack_index && !kept_only) {\n-\t\t\t/*\n-\t\t\t * If this pack is covered by the MIDX, we'd have found\n-\t\t\t * the object already in the loop above if it was here,\n-\t\t\t * so don't bother looking.\n-\t\t\t *\n-\t\t\t * The exception is if we are looking only at kept\n-\t\t\t * packs. An object can be present in two packs covered\n-\t\t\t * by the MIDX, one kept and one not-kept. And as the\n-\t\t\t * MIDX points to only one copy of each object, it might\n-\t\t\t * have returned only the non-kept version above. We\n-\t\t\t * have to check again to be thorough.\n-\t\t\t */\n-\t\t\tcontinue;\n-\t\t}\n-\t\tif (!kept_only ||\n-\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n-\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n-\t\t\tif (fill_pack_entry(oid, e, p)) {\n-\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\t\treturn 1;\n-\t\t\t}\n+\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n+\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\treturn 1;\n \t\t}\n \t}\n \treturn 0;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static void maybe_invalidate_kept_pack_cache(struct repository *r,\n+\t\t\t\t\t     unsigned flags)\n {\n-\treturn find_one_pack_entry(r, oid, e, 0);\n+\tif (!r->objects->kept_pack_cache)\n+\t\treturn;\n+\tif (r->objects->kept_pack_cache->flags == flags)\n+\t\treturn;\n+\tfree(r->objects->kept_pack_cache->packs);\n+\tFREE_AND_NULL(r->objects->kept_pack_cache);\n+}\n+\n+static struct packed_git **kept_pack_cache(struct repository *r, unsigned flags)\n+{\n+\tmaybe_invalidate_kept_pack_cache(r, flags);\n+\n+\tif (!r->objects->kept_pack_cache) {\n+\t\tstruct packed_git **packs = NULL;\n+\t\tsize_t nr = 0, alloc = 0;\n+\t\tstruct packed_git *p;\n+\n+\t\t/*\n+\t\t * We want \"all\" packs here, because we need to cover ones that\n+\t\t * are used by a midx, as well. We need to look in every one of\n+\t\t * them (instead of the midx itself) to cover duplicates. It's\n+\t\t * possible that an object is found in two packs that the midx\n+\t\t * covers, one kept and one not kept, but the midx returns only\n+\t\t * the non-kept version.\n+\t\t */\n+\t\tfor (p = get_all_packs(r); p; p = p->next) {\n+\t\t\tif ((p->pack_keep && (flags & CACHE_ON_DISK_KEEP_PACKS)) ||\n+\t\t\t    (p->pack_keep_in_core && (flags & CACHE_IN_CORE_KEEP_PACKS))) {\n+\t\t\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\t\t\tpacks[nr++] = p;\n+\t\t\t}\n+\t\t}\n+\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\tpacks[nr] = NULL;\n+\n+\t\tr->objects->kept_pack_cache = xmalloc(sizeof(*r->objects->kept_pack_cache));\n+\t\tr->objects->kept_pack_cache->packs = packs;\n+\t\tr->objects->kept_pack_cache->flags = flags;\n+\t}\n+\n+\treturn r->objects->kept_pack_cache->packs;\n }\n \n int find_kept_pack_entry(struct repository *r,\n@@ -2094,13 +2106,15 @@ int find_kept_pack_entry(struct repository *r,\n \t\t\t unsigned flags,\n \t\t\t struct pack_entry *e)\n {\n-\t/*\n-\t * Load all packs, including midx packs, since our \"kept\" strategy\n-\t * relies on that. We're relying on the side effect of it setting up\n-\t * r->objects->packed_git, which is a little ugly.\n-\t */\n-\tget_all_packs(r);\n-\treturn find_one_pack_entry(r, oid, e, flags);\n+\tstruct packed_git **cache;\n+\n+\tfor (cache = kept_pack_cache(r, flags); *cache; cache++) {\n+\t\tstruct packed_git *p = *cache;\n+\t\tif (fill_pack_entry(oid, e, p))\n+\t\t\treturn 1;\n+\t}\n+\n+\treturn 0;\n }\n \n int has_object_pack(const struct object_id *oid)\n@@ -2109,7 +2123,8 @@ int has_object_pack(const struct object_id *oid)\n \treturn find_pack_entry(the_repository, oid, &e);\n }\n \n-int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n+int has_object_kept_pack(const struct object_id *oid,\n+\t\t\t unsigned flags)\n {\n \tstruct pack_entry e;\n \treturn find_kept_pack_entry(the_repository, oid, flags, &e);\ndiff --git a/packfile.h b/packfile.h\nindex 624327f64d..eb56db2a7b 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -161,10 +161,6 @@ int packed_object_info(struct repository *r,\n void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n \n-#define ON_DISK_KEEP_PACKS 1\n-#define IN_CORE_KEEP_PACKS 2\n-#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n-\n /*\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\ndiff --git a/revision.c b/revision.c\nindex 4c5adb90b1..41c0478705 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -2338,14 +2338,14 @@ static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n \t\tdie(_(\"--unpacked=<packfile> no longer supported\"));\n \t} else if (!strcmp(arg, \"--no-kept-objects\")) {\n \t\trevs->no_kept_objects = 1;\n-\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n-\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= CACHE_ON_DISK_KEEP_PACKS;\n \t} else if (skip_prefix(arg, \"--no-kept-objects=\", &optarg)) {\n \t\trevs->no_kept_objects = 1;\n \t\tif (!strcmp(optarg, \"in-core\"))\n-\t\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\t\trevs->keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n \t\tif (!strcmp(optarg, \"on-disk\"))\n-\t\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t\t\trevs->keep_pack_cache_flags |= CACHE_ON_DISK_KEEP_PACKS;\n \t} else if (!strcmp(arg, \"-r\")) {\n \t\trevs->diff = 1;\n \t\trevs->diffopt.flags.recursive = 1;\n-- \n2.30.0.533.g2f8b6b552f.dirty\n\n"},{"id":"416134","messageId":"d5561585c2221a9635eb0fc7a65298ee8a2b6348.1612411124.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 8/8] builtin/repack.c: add '--geometric' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:59:25Z","receivedAt":"2021-02-04T04:01:57Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Often it is useful to both:\n\n  - have relatively few packfiles in a repository, and\n\n  - avoid having so few packfiles in a repository that we repack its\n    entire contents regularly\n\nThis patch implements a '--geometric=<n>' option in 'git repack'. This\nallows the caller to specify that they would like each pack to be at\nleast a factor times as large as the previous largest pack (by object\ncount).\n\nConcretely, say that a repository has 'n' packfiles, labeled P1, P2,\n..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\nWith a geometric factor of 'r', it should be that:\n\n  objects(Pi) > r*objects(P(i-1))\n\nfor all i in [1, n], where the packs are sorted by\n\n  objects(P1) <= objects(P2) <= ... <= objects(Pn).\n\nSince finding a true optimal repacking is NP-hard, we approximate it\nalong two directions:\n\n  1. We assume that there is a cutoff of packs _before starting the\n     repack_ where everything to the right of that cut-off already forms\n     a geometric progression (or no cutoff exists and everything must be\n     repacked).\n\n  2. We assume that everything smaller than the cutoff count must be\n     repacked. This forms our base assumption, but it can also cause\n     even the \"heavy\" packs to get repacked, for e.g., if we have 6\n     packs containing the following number of objects:\n\n       1, 1, 1, 2, 4, 32\n\n     then we would place the cutoff between '1, 1' and '1, 2, 4, 32',\n     rolling up the first two packs into a pack with 2 objects. That\n     breaks our progression and leaves us:\n\n       2, 1, 2, 4, 32\n         ^\n\n     (where the '^' indicates the position of our split). To restore a\n     progression, we move the split forward (towards larger packs)\n     joining each pack into our new pack until a geometric progression\n     is restored. Here, that looks like:\n\n       2, 1, 2, 4, 32  ~>  3, 2, 4, 32  ~>  5, 4, 32  ~> ... ~> 9, 32\n         ^                   ^                ^                   ^\n\nThis has the advantage of not repacking the heavy-side of packs too\noften while also only creating one new pack at a time. Another wrinkle\nis that we assume that loose, indexed, and reflog'd objects are\ninsignificant, and lump them into any new pack that we create. This can\nlead to non-idempotent results.\n\nSuggested-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt |  11 +++\n builtin/repack.c             | 187 ++++++++++++++++++++++++++++++++++-\n t/t7703-repack-geometric.sh  | 137 +++++++++++++++++++++++++\n 3 files changed, 331 insertions(+), 4 deletions(-)\n create mode 100755 t/t7703-repack-geometric.sh\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex 92f146d27d..b1ffcfd974 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -165,6 +165,17 @@ depth is 4095.\n \tPass the `--delta-islands` option to `git-pack-objects`, see\n \tlinkgit:git-pack-objects[1].\n \n+-g=<factor>::\n+--geometric=<factor>::\n+\tArrange resulting pack structure so that each successive pack\n+\tcontains at least `<factor>` times the number of objects as the\n+\tnext-largest pack.\n++\n+`git repack` ensures this by determining a \"cut\" of packfiles that need to be\n+repacked into one in order to ensure a geometric progression. It picks the\n+smallest set of packfiles such that as many of the larger packfiles (by count of\n+objects contained in that pack) may be left intact.\n+\n Configuration\n -------------\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 2158b48f4c..b4e0e69661 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -296,6 +296,124 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n #define ALL_INTO_ONE 1\n #define LOOSEN_UNREACHABLE 2\n \n+struct pack_geometry {\n+\tstruct packed_git **pack;\n+\tuint32_t pack_nr, pack_alloc;\n+\tuint32_t split;\n+};\n+\n+static uint32_t geometry_pack_weight(struct packed_git *p)\n+{\n+\tif (open_pack_index(p))\n+\t\tdie(_(\"cannot open index for %s\"), p->pack_name);\n+\treturn p->num_objects;\n+}\n+\n+static int geometry_cmp(const void *va, const void *vb)\n+{\n+\tuint32_t aw = geometry_pack_weight(*(struct packed_git **)va),\n+\t\t bw = geometry_pack_weight(*(struct packed_git **)vb);\n+\n+\tif (aw < bw)\n+\t\treturn -1;\n+\tif (aw > bw)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static void init_pack_geometry(struct pack_geometry **geometry_p)\n+{\n+\tstruct packed_git *p;\n+\tstruct pack_geometry *geometry;\n+\n+\t*geometry_p = xcalloc(1, sizeof(struct pack_geometry));\n+\tgeometry = *geometry_p;\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tif (!pack_kept_objects && p->pack_keep)\n+\t\t\tcontinue;\n+\n+\t\tALLOC_GROW(geometry->pack,\n+\t\t\t   geometry->pack_nr + 1,\n+\t\t\t   geometry->pack_alloc);\n+\n+\t\tgeometry->pack[geometry->pack_nr] = p;\n+\t\tgeometry->pack_nr++;\n+\t}\n+\n+\tQSORT(geometry->pack, geometry->pack_nr, geometry_cmp);\n+}\n+\n+static void split_pack_geometry(struct pack_geometry *geometry, int factor)\n+{\n+\tuint32_t i;\n+\tuint32_t split;\n+\toff_t total_size = 0;\n+\n+\tif (geometry->pack_nr <= 1) {\n+\t\tgeometry->split = geometry->pack_nr;\n+\t\treturn;\n+\t}\n+\n+\tsplit = geometry->pack_nr - 1;\n+\n+\t/*\n+\t * First, count the number of packs (in descending order of size) which\n+\t * already form a geometric progression.\n+\t */\n+\tfor (i = geometry->pack_nr - 1; i > 0; i--) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tstruct packed_git *prev = geometry->pack[i - 1];\n+\t\tif (geometry_pack_weight(ours) >= factor * geometry_pack_weight(prev))\n+\t\t\tsplit--;\n+\t\telse\n+\t\t\tbreak;\n+\t}\n+\n+\tif (split) {\n+\t\t/*\n+\t\t * Move the split one to the right, since the top element in the\n+\t\t * last-compared pair can't be in the progression. Only do this\n+\t\t * when we split in the middle of the array (otherwise if we got\n+\t\t * to the end, then the split is in the right place).\n+\t\t */\n+\t\tsplit++;\n+\t}\n+\n+\t/*\n+\t * Then, anything to the left of 'split' must be in a new pack. But,\n+\t * creating that new pack may cause packs in the heavy half to no longer\n+\t * form a geometric progression.\n+\t *\n+\t * Compute an expected size of the new pack, and then determine how many\n+\t * packs in the heavy half need to be joined into it (if any) to restore\n+\t * the geometric progression.\n+\t */\n+\tfor (i = 0; i < split; i++)\n+\t\ttotal_size += geometry_pack_weight(geometry->pack[i]);\n+\tfor (i = split; i < geometry->pack_nr; i++) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tif (geometry_pack_weight(ours) < factor * total_size) {\n+\t\t\tsplit++;\n+\t\t\ttotal_size += geometry_pack_weight(ours);\n+\t\t} else\n+\t\t\tbreak;\n+\t}\n+\n+\tgeometry->split = split;\n+}\n+\n+static void clear_pack_geometry(struct pack_geometry *geometry)\n+{\n+\tif (!geometry)\n+\t\treturn;\n+\n+\tfree(geometry->pack);\n+\tgeometry->pack_nr = 0;\n+\tgeometry->pack_alloc = 0;\n+\tgeometry->split = 0;\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -303,6 +421,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list names = STRING_LIST_INIT_DUP;\n \tstruct string_list rollback = STRING_LIST_INIT_NODUP;\n \tstruct string_list existing_packs = STRING_LIST_INIT_DUP;\n+\tstruct pack_geometry *geometry = NULL;\n \tstruct strbuf line = STRBUF_INIT;\n \tint i, ext, ret;\n \tFILE *out;\n@@ -315,6 +434,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tint no_update_server_info = 0;\n \tstruct pack_objects_args po_args = {NULL};\n+\tint geometric_factor = 0;\n \n \tstruct option builtin_repack_options[] = {\n \t\tOPT_BIT('a', NULL, &pack_everything,\n@@ -355,6 +475,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t\tN_(\"repack objects in packs marked with .keep\")),\n \t\tOPT_STRING_LIST(0, \"keep-pack\", &keep_pack_list, N_(\"name\"),\n \t\t\t\tN_(\"do not repack this pack\")),\n+\t\tOPT_INTEGER('g', \"geometric\", &geometric_factor,\n+\t\t\t    N_(\"find a geometric progression with factor <N>\")),\n \t\tOPT_END()\n \t};\n \n@@ -381,6 +503,13 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (write_bitmaps && !(pack_everything & ALL_INTO_ONE))\n \t\tdie(_(incremental_bitmap_conflict_error));\n \n+\tif (geometric_factor) {\n+\t\tif (pack_everything)\n+\t\t\tdie(_(\"--geometric is incompatible with -A, -a\"));\n+\t\tinit_pack_geometry(&geometry);\n+\t\tsplit_pack_geometry(geometry, geometric_factor);\n+\t}\n+\n \tpackdir = mkpathdup(\"%s/pack\", get_object_directory());\n \tpacktmp = mkpathdup(\"%s/.tmp-%d-pack\", packdir, (int)getpid());\n \n@@ -395,9 +524,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s\",\n \t\t\t     keep_pack_list.items[i].string);\n \tstrvec_push(&cmd.args, \"--non-empty\");\n-\tstrvec_push(&cmd.args, \"--all\");\n-\tstrvec_push(&cmd.args, \"--reflog\");\n-\tstrvec_push(&cmd.args, \"--indexed-objects\");\n+\tif (!geometry) {\n+\t\t/*\n+\t\t * 'git pack-objects' will up all objects loose or packed\n+\t\t * (either rolling them up or leaving them alone), so don't pass\n+\t\t * these options.\n+\t\t *\n+\t\t * The implementation of 'git pack-objects --stdin-packs'\n+\t\t * makes them redundant (and the two are incompatible).\n+\t\t */\n+\t\tstrvec_push(&cmd.args, \"--all\");\n+\t\tstrvec_push(&cmd.args, \"--reflog\");\n+\t\tstrvec_push(&cmd.args, \"--indexed-objects\");\n+\t}\n \tif (has_promisor_remote())\n \t\tstrvec_push(&cmd.args, \"--exclude-promisor-objects\");\n \tif (write_bitmaps > 0)\n@@ -428,17 +567,37 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t\tstrvec_push(&cmd.env_array, \"GIT_REF_PARANOIA=1\");\n \t\t\t}\n \t\t}\n+\t} else if (geometry) {\n+\t\tstrvec_push(&cmd.args, \"--stdin-packs\");\n+\t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t} else {\n \t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t\tstrvec_push(&cmd.args, \"--incremental\");\n \t}\n \n-\tcmd.no_stdin = 1;\n+\tif (geometry)\n+\t\tcmd.in = -1;\n+\telse\n+\t\tcmd.no_stdin = 1;\n \n \tret = start_command(&cmd);\n \tif (ret)\n \t\treturn ret;\n \n+\tif (geometry) {\n+\t\tFILE *in = xfdopen(cmd.in, \"w\");\n+\t\t/*\n+\t\t * The resulting pack should contain all objects in packs that\n+\t\t * are going to be rolled up, but exclude objects in packs which\n+\t\t * are being left alone.\n+\t\t */\n+\t\tfor (i = 0; i < geometry->split; i++)\n+\t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n+\t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n+\t\t\tfprintf(in, \"^%s\\n\", pack_basename(geometry->pack[i]));\n+\t\tfclose(in);\n+\t}\n+\n \tout = xfdopen(cmd.out, \"r\");\n \twhile (strbuf_getline_lf(&line, out) != EOF) {\n \t\tif (line.len != the_hash_algo->hexsz)\n@@ -506,6 +665,25 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tif (!string_list_has_string(&names, sha1))\n \t\t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n+\n+\t\tif (geometry) {\n+\t\t\tstruct strbuf buf = STRBUF_INIT;\n+\n+\t\t\tuint32_t i;\n+\t\t\tfor (i = 0; i < geometry->split; i++) {\n+\t\t\t\tstruct packed_git *p = geometry->pack[i];\n+\t\t\t\tif (string_list_has_string(&names,\n+\t\t\t\t\t\t\t   hash_to_hex(p->hash)))\n+\t\t\t\t\tcontinue;\n+\n+\t\t\t\tstrbuf_reset(&buf);\n+\t\t\t\tstrbuf_addstr(&buf, pack_basename(p));\n+\t\t\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\n+\t\t\t\tremove_redundant_pack(packdir, buf.buf);\n+\t\t\t}\n+\t\t\tstrbuf_release(&buf);\n+\t\t}\n \t\tif (!po_args.quiet && isatty(2))\n \t\t\topts |= PRUNE_PACKED_VERBOSE;\n \t\tprune_packed_objects(opts);\n@@ -527,6 +705,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstring_list_clear(&names, 0);\n \tstring_list_clear(&rollback, 0);\n \tstring_list_clear(&existing_packs, 0);\n+\tclear_pack_geometry(geometry);\n \tstrbuf_release(&line);\n \n \treturn 0;\ndiff --git a/t/t7703-repack-geometric.sh b/t/t7703-repack-geometric.sh\nnew file mode 100755\nindex 0000000000..96917fc163\n--- /dev/null\n+++ b/t/t7703-repack-geometric.sh\n@@ -0,0 +1,137 @@\n+#!/bin/sh\n+\n+test_description='git repack --geometric works correctly'\n+\n+. ./test-lib.sh\n+\n+GIT_TEST_MULTI_PACK_INDEX=0\n+\n+objdir=.git/objects\n+midx=$objdir/pack/multi-pack-index\n+\n+test_expect_success '--geometric with no packs' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\tgit repack --geometric 2 >out &&\n+\t\ttest_i18ngrep \"Nothing new to pack\" out\n+\t)\n+'\n+\n+test_expect_success '--geometric with an intact progression' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# These packs already form a geometric progression.\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 2 && # 6 objects\n+\t\ttest_commit_bulk --start=4 4 && # 12 objects\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >expect &&\n+\t\tgit repack --geometric 2 -d &&\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >actual &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success '--geometric with small-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >small &&\n+\t\ttest_commit_bulk --start=3 4 && # 12 objects\n+\t\ttest_commit_bulk --start=7 8 && # 24 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tgit repack --geometric 2 -d &&\n+\n+\t\t# Three packs in total; two of the existing large ones, and one\n+\t\t# new one.\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\ttest_line_count = 3 after &&\n+\t\tcomm -3 small before | tr -d \"\\t\" >large &&\n+\t\tgrep -qFf large after\n+\t)\n+'\n+\n+test_expect_success '--geometric with small- and large-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# size(small1) + size(small2) > size(medium) / 2\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 3 && # 7 objects\n+\t\ttest_commit_bulk --start=6 9 && # 27 objects &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tgit repack --geometric 2 -d &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\tcomm -12 before after >untouched &&\n+\n+\t\t# Two packs in total; the largest pack from before running \"git\n+\t\t# repack\", and one new one.\n+\t\ttest_line_count = 1 untouched &&\n+\t\ttest_line_count = 2 after\n+\t)\n+'\n+\n+test_expect_success '--geometric ignores kept packs' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\ttest_commit kept && # 3 objects\n+\t\ttest_commit pack && # 3 objects\n+\n+\t\tKEPT=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n+\t\trefs/tags/kept\n+\t\tEOF\n+\t\t) &&\n+\t\tPACK=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n+\t\trefs/tags/pack\n+\t\t^refs/tags/kept\n+\t\tEOF\n+\t\t) &&\n+\n+\t\t# neither pack contains more than twice the number of objects in\n+\t\t# the other, so they should be combined. but, marking one as\n+\t\t# .kept on disk will \"freeze\" it, so the pack structure should\n+\t\t# remain unchanged.\n+\t\ttouch $objdir/pack/pack-$KEPT.keep &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\t\tgit repack --geometric 2 -d &&\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\n+\t\t# both packs should still exist\n+\t\ttest_path_is_file $objdir/pack/pack-$KEPT.pack &&\n+\t\ttest_path_is_file $objdir/pack/pack-$PACK.pack &&\n+\n+\t\t# and no new packs should be created\n+\t\ttest_cmp before after &&\n+\n+\t\t# Passing --pack-kept-objects causes packs with a .keep file to\n+\t\t# be repacked, too.\n+\t\tgit repack --geometric 2 -d --pack-kept-objects &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" >after &&\n+\t\ttest_line_count = 1 after\n+\t)\n+'\n+\n+test_done\n-- \n2.30.0.533.g2f8b6b552f.dirty\n"},{"id":"416135","messageId":"c3868c7df92484f0527ce500ad1156275be334e8.1612411124.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"[PATCH v2 6/8] builtin/pack-objects.c: rewrite honor-pack-keep logic","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-04T03:59:17Z","receivedAt":"2021-02-04T04:02:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nNow that we have find_kept_pack_entry(), we don't have to manually keep\nhunting through every pack to find a possible \"kept\" duplicate of the\nobject. This should be faster, assuming only a portion of your total\npacks are actually kept.\n\nNote that we have to re-order the logic a bit here; we can deal with the\n\"kept\" situation completely, and then just fall back to the \"--local\"\nquestion. It might be worth having a similar optimized function to look\nat only local packs.\n\nHere are the results from p5303 (measurements again taken on the\nkernel):\n\n  Test                                        HEAD^                    HEAD\n  -----------------------------------------------------------------------------------------------\n  5303.5: repack (1)                          57.42(54.88+10.64)       57.44(54.71+10.78) +0.0%\n  5303.6: repack with --stdin-packs (1)       0.01(0.01+0.00)          0.01(0.00+0.01) +0.0%\n  5303.10: repack (50)                        71.26(88.24+4.96)        71.32(88.38+4.90) +0.1%\n  5303.11: repack with --stdin-packs (50)     3.49(11.82+0.28)         3.43(11.81+0.22) -1.7%\n  5303.15: repack (1000)                      215.64(491.33+14.80)     215.59(493.75+14.62) -0.0%\n  5303.16: repack with --stdin-packs (1000)   198.79(380.51+7.97)      131.44(314.24+8.11) -33.9%\n\nSo our --stdin-packs case with many packs is now finally faster than the\nnon-keep case (because it gets the speed benefit of looking at fewer\nobjects, but not as big a penalty for looking at many packs).\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 125 ++++++++++++++++++++++++-----------------\n 1 file changed, 73 insertions(+), 52 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 6d19eb000a..fbd7b54d70 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1188,7 +1188,8 @@ static int have_duplicate_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int want_found_object(int exclude, struct packed_git *p)\n+static int want_found_object(const struct object_id *oid, int exclude,\n+\t\t\t     struct packed_git *p)\n {\n \tif (exclude)\n \t\treturn 1;\n@@ -1209,22 +1210,73 @@ static int want_found_object(int exclude, struct packed_git *p)\n \t * Otherwise, we signal \"-1\" at the end to tell the caller that we do\n \t * not know either way, and it needs to check more packs.\n \t */\n-\tif (!ignore_packed_keep_on_disk &&\n-\t    !ignore_packed_keep_in_core &&\n-\t    (!local || !have_non_local_packs))\n+\n+\t/*\n+\t * Handle .keep first, as we have a fast(er) path there.\n+\t */\n+\tif (ignore_packed_keep_on_disk || ignore_packed_keep_in_core) {\n+\t\t/*\n+\t\t * Set the flags for the kept-pack cache to be the ones we want\n+\t\t * to ignore.\n+\t\t *\n+\t\t * That is, if we are ignoring objects in on-disk keep packs,\n+\t\t * then we want to search through the on-disk keep and ignore\n+\t\t * the in-core ones.\n+\t\t */\n+\t\tunsigned flags = 0;\n+\t\tif (ignore_packed_keep_on_disk)\n+\t\t\tflags |= ON_DISK_KEEP_PACKS;\n+\t\tif (ignore_packed_keep_in_core)\n+\t\t\tflags |= IN_CORE_KEEP_PACKS;\n+\n+\t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n+\t\t\treturn 0;\n+\t\tif (ignore_packed_keep_in_core && p->pack_keep_in_core)\n+\t\t\treturn 0;\n+\t\tif (has_object_kept_pack(oid, flags))\n+\t\t\treturn 0;\n+\t}\n+\n+\t/*\n+\t * At this point we know definitively that either we don't care about\n+\t * keep-packs, or the object is not in one. Keep checking other\n+\t * conditions...\n+\t */\n+\n+\tif (!local || !have_non_local_packs)\n \t\treturn 1;\n-\n \tif (local && !p->pack_local)\n \t\treturn 0;\n-\tif (p->pack_local &&\n-\t    ((ignore_packed_keep_on_disk && p->pack_keep) ||\n-\t     (ignore_packed_keep_in_core && p->pack_keep_in_core)))\n-\t\treturn 0;\n \n \t/* we don't know yet; keep looking for more packs */\n \treturn -1;\n }\n \n+static int want_object_in_pack_one(struct packed_git *p,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   int exclude,\n+\t\t\t\t   struct packed_git **found_pack,\n+\t\t\t\t   off_t *found_offset)\n+{\n+\toff_t offset;\n+\n+\tif (p == *found_pack)\n+\t\toffset = *found_offset;\n+\telse\n+\t\toffset = find_pack_entry_one(oid->hash, p);\n+\n+\tif (offset) {\n+\t\tif (!*found_pack) {\n+\t\t\tif (!is_pack_valid(p))\n+\t\t\t\treturn -1;\n+\t\t\t*found_offset = offset;\n+\t\t\t*found_pack = p;\n+\t\t}\n+\t\treturn want_found_object(oid, exclude, p);\n+\t}\n+\treturn -1;\n+}\n+\n /*\n  * Check whether we want the object in the pack (e.g., we do not want\n  * objects found in non-local stores if the \"--local\" option was used).\n@@ -1252,7 +1304,7 @@ static int want_object_in_pack(const struct object_id *oid,\n \t * are present we will determine the answer right now.\n \t */\n \tif (*found_pack) {\n-\t\twant = want_found_object(exclude, *found_pack);\n+\t\twant = want_found_object(oid, exclude, *found_pack);\n \t\tif (want != -1)\n \t\t\treturn want;\n \t}\n@@ -1260,53 +1312,22 @@ static int want_object_in_pack(const struct object_id *oid,\n \tfor (m = get_multi_pack_index(the_repository); m; m = m->next) {\n \t\tstruct pack_entry e;\n \t\tif (fill_midx_entry(the_repository, oid, &e, m)) {\n-\t\t\tstruct packed_git *p = e.p;\n-\t\t\toff_t offset;\n-\n-\t\t\tif (p == *found_pack)\n-\t\t\t\toffset = *found_offset;\n-\t\t\telse\n-\t\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\t\tif (offset) {\n-\t\t\t\tif (!*found_pack) {\n-\t\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\t\tcontinue;\n-\t\t\t\t\t*found_offset = offset;\n-\t\t\t\t\t*found_pack = p;\n-\t\t\t\t}\n-\t\t\t\twant = want_found_object(exclude, p);\n-\t\t\t\tif (want != -1)\n-\t\t\t\t\treturn want;\n-\t\t\t}\n-\t\t}\n-\t}\n-\n-\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n-\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\toff_t offset;\n-\n-\t\tif (p == *found_pack)\n-\t\t\toffset = *found_offset;\n-\t\telse\n-\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\tif (offset) {\n-\t\t\tif (!*found_pack) {\n-\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\tcontinue;\n-\t\t\t\t*found_offset = offset;\n-\t\t\t\t*found_pack = p;\n-\t\t\t}\n-\t\t\twant = want_found_object(exclude, p);\n-\t\t\tif (!exclude && want > 0)\n-\t\t\t\tlist_move(&p->mru,\n-\t\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\t\twant = want_object_in_pack_one(e.p, oid, exclude, found_pack, found_offset);\n \t\t\tif (want != -1)\n \t\t\t\treturn want;\n \t\t}\n \t}\n \n+\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n+\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n+\t\twant = want_object_in_pack_one(p, oid, exclude, found_pack, found_offset);\n+\t\tif (!exclude && want > 0)\n+\t\t\tlist_move(&p->mru,\n+\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\tif (want != -1)\n+\t\t\treturn want;\n+\t}\n+\n \tif (uri_protocols.nr) {\n \t\tstruct configured_exclusion *ex =\n \t\t\toidmap_get(&configured_exclusions, oid);\n-- \n2.30.0.533.g2f8b6b552f.dirty\n\n"},{"id":"417110","messageId":"YCw8TmDdEfcnZSOo@coredump.intra.peff.net","threadId":"55012","inReplyTo":"f7186147ebb0b2d01d8f1e0f742f367204d7d9c9.1612411123.git.me@ttaylorr.com","subject":"Re: [PATCH v2 1/8] packfile: introduce 'find_kept_pack_entry()'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-16T21:42:38Z","receivedAt":"2021-02-16T21:43:21Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:58:50PM -0500, Taylor Blau wrote:\n\n> Future callers will want a function to fill a 'struct pack_entry' for a\n> given object id but _only_ from its position in any kept pack(s).\n> \n> In particular, an new 'git repack' mode which ensures the resulting\n\nNit (not worth re-rolling): s/an new/a new/\n\n> There is a gotcha when looking up objects that are duplicated in kept\n> and non-kept packs, particularly when the MIDX stores the non-kept\n> version and the caller asked for kept objects only. This could be\n> resolved by teaching the MIDX to resolve duplicates by always favoring\n> the kept pack (if one exists), but this breaks an assumption in existing\n> MIDXs, and so it would require a format change.\n\nI don't think this would be possible without a major rethink of how\nmidxs work. The \"keep\" property of a pack is not set in stone when the\nmidx is created. You could add a \".keep\" file to one of its packs later,\nor even mark one as an in-core keep on the fly. But the duplicate\nresolution happens at creation.\n\nSo maybe your \"breaks an assumption\" is the notion that we do not store\nduplicate information at all in the midx. If so, then I agree. :) But\nI'd also call fixing that more than just a format change.\n\n(None of which changes your point, which isn't that it isn't worth\npursuing that direction).\n\n-Peff\n"},{"id":"417111","messageId":"YCw9oX9EEEzo5Kaj@nand.local","threadId":"55012","inReplyTo":"YCw8TmDdEfcnZSOo@coredump.intra.peff.net","subject":"Re: [PATCH v2 1/8] packfile: introduce 'find_kept_pack_entry()'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-16T21:48:17Z","receivedAt":"2021-02-16T21:49:06Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Feb 16, 2021 at 04:42:38PM -0500, Jeff King wrote:\n> On Wed, Feb 03, 2021 at 10:58:50PM -0500, Taylor Blau wrote:\n>\n> > Future callers will want a function to fill a 'struct pack_entry' for a\n> > given object id but _only_ from its position in any kept pack(s).\n> >\n> > In particular, an new 'git repack' mode which ensures the resulting\n>\n> Nit (not worth re-rolling): s/an new/a new/\n\nOops. Good eyes.\n\n> > There is a gotcha when looking up objects that are duplicated in kept\n> > and non-kept packs, particularly when the MIDX stores the non-kept\n> > version and the caller asked for kept objects only. This could be\n> > resolved by teaching the MIDX to resolve duplicates by always favoring\n> > the kept pack (if one exists), but this breaks an assumption in existing\n> > MIDXs, and so it would require a format change.\n>\n> I don't think this would be possible without a major rethink of how\n> midxs work. The \"keep\" property of a pack is not set in stone when the\n> midx is created. You could add a \".keep\" file to one of its packs later,\n> or even mark one as an in-core keep on the fly. But the duplicate\n> resolution happens at creation.\n>\n> So maybe your \"breaks an assumption\" is the notion that we do not store\n> duplicate information at all in the midx. If so, then I agree. :) But\n> I'd also call fixing that more than just a format change.\n\nThat's part of it, indeed. The part that I was referring to is that\nexisting MIDX readers expect duplicates to be resolved in a certain way\n(effectively in favor of the pack with the lowest mtime). So the easy\npart is indicating a format change which tells new readers how to expect\nties to be broken.\n\nBut (as you note) that's only part of the problem: even if we say \"ties\nare resolved in favor of the lowest mtime pack, or a .keep one, if it\nexists\", then which ones are kept and which aren't? Even *if* we wrote\nthat down (which I'm not suggesting we do), kept-ness isn't an immutable\nproperty of the pack, and so I think relying on it is a tricky direction\nto take.\n\n> (None of which changes your point, which isn't that it isn't worth\n> pursuing that direction).\n\nYeah; my hope in writing some of this down in the above paragraph is\nthat it would make clear to future readers that such a MIDX change would\nresolve some complexity here, but the complexity it adds in the MIDX\ncode isn't worth the tradeoff.\n\n> -Peff\n\nThanks,\nTaylor\n"},{"id":"417121","messageId":"YCxSlNK4Ug2exDbv@coredump.intra.peff.net","threadId":"55012","inReplyTo":"ddc2896caa13b9f1cdccb2f0a5892143fa98237c.1612411123.git.me@ttaylorr.com","subject":"Re: [PATCH v2 2/8] revision: learn '--no-kept-objects'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-16T23:17:40Z","receivedAt":"2021-02-16T23:18:38Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:58:57PM -0500, Taylor Blau wrote:\n\n> @@ -3797,6 +3807,11 @@ enum commit_action get_commit_action(struct rev_info *revs, struct commit *commi\n>  \t\treturn commit_ignore;\n>  \tif (revs->unpacked && has_object_pack(&commit->object.oid))\n>  \t\treturn commit_ignore;\n> +\tif (revs->no_kept_objects) {\n> +\t\tif (has_object_kept_pack(&commit->object.oid,\n> +\t\t\t\t\t revs->keep_pack_cache_flags))\n> +\t\t\treturn commit_ignore;\n> +\t}\n\nOK, so this has the same \"problems\" as --unpacked, which is that we can\nmiss some objects (i.e., things that are reachable but not-kept may not\nbe reported). But it should be OK in this version of the series, because\nwe will not be relying on it for selection of objects, but only to fill\nin ordering / namehash fields.\n\nShould we warn people about that, either as a comment or in the commit\nmessage?\n\n> +--no-kept-objects[=<kind>]::\n> +\tHalts the traversal as soon as an object in a kept pack is\n> +\tfound. If `<kind>` is `on-disk`, only packs with a corresponding\n> +\t`*.keep` file are ignored. If `<kind>` is `in-core`, only packs\n> +\twith their in-core kept state set are ignored. Otherwise, both\n> +\tkinds of kept packs are ignored.\n\nLikewise, I wonder whether we need to expose this mode to users.\nNormally I'm a fan of doing so, because it allows scripted callers\naccess to more of the internals, but:\n\n  - the semantics are kind of weird about where we draw the line between\n    performance and absolute correctness\n\n  - the \"in-core\" thing is a bit weird for callers of rev-list; how do I\n    as a caller mark a pack as kept-in-core? I think it's only an\n    internal pack-objects thing.\n\nOnce we support this in rev-list, we'll have to do it forever (or deal\nwith deprecation, etc). If we just need it internally, maybe it's wise\nto leave it as a something you ask for by manipulating rev_info\ndirectly. Or perhaps leave it as an undocumented interface we use for\ntesting, and not something we promise to keep working.\n\n> --- a/list-objects.c\n> +++ b/list-objects.c\n> @@ -338,6 +338,13 @@ static void traverse_trees_and_blobs(struct traversal_context *ctx,\n>  \t\t\tctx->show_object(obj, name, ctx->show_data);\n>  \t\t\tcontinue;\n>  \t\t}\n> +\t\tif (ctx->revs->no_kept_objects) {\n> +\t\t\tstruct pack_entry e;\n> +\t\t\tif (find_kept_pack_entry(ctx->revs->repo, &obj->oid,\n> +\t\t\t\t\t\t ctx->revs->keep_pack_cache_flags,\n> +\t\t\t\t\t\t &e))\n> +\t\t\t\tcontinue;\n> +\t\t}\n\nThis hunk is interesting.\n\nThere is no similar check for revs->unpacked in list-objects.c to cut\noff the traversal. And indeed, running \"rev-list --unpacked\" will\ngenerally look at the _whole_ tree for a commit that is unpacked, even\nif all of the tree entries are packed. That's something we might\nconsider changing in the name of performance (though it does increase\nthe number of cases where --unpacked will fail to find an unpacked but\nreachable object).\n\nBut this is a funny place to put it. If I understand it correctly, it is\ncutting off the traversal at the very top of the tree. I.e., if we had a\ncommit that is not-kept, we'd queue it's root tree. And then we might\nfind that the root tree is kept, and avoid traversing it. But if we _do_\ntraverse it, we would look at every subtree it contains, even if they\nare kept! That's because we recurse the tree via the recursive\nprocess_tree(), not by queueing more objects in the pending array here.\n\nSo this check seems to exist in a funny middle ground. I think it's\nunlikely to catch anything useful (usually commits have a unique root\ntree; it's all of the untouched parts of the subtrees that will be in\nthe kept packs). IMHO we should either drop it (and act like\n\"--unpacked\", accepting that we may traverse some extra tree objects),\nor we should go all-in on performance and cut it off in the top of\nprocess_tree().\n\n-Peff\n"},{"id":"417122","messageId":"YCxZcz5JGtxObOF3@coredump.intra.peff.net","threadId":"55012","inReplyTo":"c96b1bf99582beefb96c3774b13a4f5a12fc61cc.1612411124.git.me@ttaylorr.com","subject":"Re: [PATCH v2 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-16T23:46:59Z","receivedAt":"2021-02-16T23:48:00Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:59:03PM -0500, Taylor Blau wrote:\n\n> In an upcoming commit, 'git repack' will want to create a pack comprised\n> of all of the objects in some packs (the included packs) excluding any\n> objects in some other packs (the excluded packs).\n> \n> This caller could iterate those packs themselves and feed the objects it\n> finds to 'git pack-objects' directly over stdin, but this approach has a\n> few downsides:\n> \n>   - It requires every caller that wants to drive 'git pack-objects' in\n>     this way to implement pack iteration themselves. This forces the\n>     caller to think about details like what order objects are fed to\n>     pack-objects, which callers would likely rather not do.\n> \n>   - If the set of objects in included packs is large, it requires\n>     sending a lot of data over a pipe, which is inefficient.\n> \n>   - The caller is forced to keep track of the excluded objects, too, and\n>     make sure that it doesn't send any objects that appear in both\n>     included and excluded packs.\n> \n> But the biggest downside is the lack of a reachability traversal.\n> Because the caller passes in a list of objects directly, those objects\n> don't get a namehash assigned to them, which can have a negative impact\n> on the delta selection process, causing 'git pack-objects' to fail to\n> find good deltas even when they exist.\n>\n> The caller could formulate a reachability traversal themselves, but the\n> only way to drive 'git pack-objects' in this way is to do a full\n> traversal, and then remove objects in the excluded packs after the\n> traversal is complete. This can be detrimental to callers who care\n> about performance, especially in repositories with many objects.\n\nYep, I think this is a good summary of the problem space, and why this\ncomplexity should be pushed into pack-objects and not the caller.\n\n> To address the delta selection problem, 'git pack-objects --stdin-packs'\n> works as follows. First, it assembles a list of objects that it is going\n> to pack, as above. Then, a reachability traversal is started, whose tips\n> are any commits mentioned in included packs. Upon visiting an object, we\n> find its corresponding object_entry in the to_pack list, and set its\n> namehash parameter appropriately.\n> \n> To avoid the traversal visiting more objects than it needs to, the\n> traversal is halted upon encountering an object which can be found in an\n> excluded pack (by marking the excluded packs as kept in-core, and\n> passing --no-kept-objects=in-core to the revision machinery).\n> \n> This can cause the traversal to halt early, for example if an object in\n> an included pack is an ancestor of ones in excluded packs. But stopping\n> early is OK, since filling in the namehash fields of objects in the\n> to_pack list is only additive (i.e., having it helps the delta selection\n> process, but leaving it blank doesn't impact the correctness of the\n> resulting pack).\n\nOK, good. Definitely worth calling out this subtle distinction of\ncorrectness versus the heuristic.\n\nDo we use this partial traversal to impact the write order at all? That\nwould be a nice-to-have, but I suspect that just concatenating the packs\n(presumably by descending mtime) ends up with a similar result.\n\n> --- a/Documentation/git-pack-objects.txt\n> +++ b/Documentation/git-pack-objects.txt\n> @@ -85,6 +85,16 @@ base-name::\n>  \treference was included in the resulting packfile.  This\n>  \tcan be useful to send new tags to native Git clients.\n>  \n> +--stdin-packs::\n> +\tRead the basenames of packfiles from the standard input, instead\n> +\tof object names or revision arguments. The resulting pack\n> +\tcontains all objects listed in the included packs (those not\n> +\tbeginning with `^`), excluding any objects listed in the\n> +\texcluded packs (beginning with `^`).\n> ++\n> +Incompatible with `--revs`, or options that imply `--revs` (such as\n> +`--all`), with the exception of `--unpacked`, which is compatible.\n\nI know you say \"basename\" here, but I wonder if it is worth giving an\nexample (`pack-1234abcd.pack`) to make it clear in what form we expect\nit. Or possibly something in the `EXAMPLES` section.\n\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n> @@ -2979,6 +2979,164 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n>  \treturn git_default_config(k, v, cb);\n>  }\n>  \n> +static int stdin_packs_found_nr;\n> +static int stdin_packs_hints_nr;\n\nI scratched my head at these until I looked further in the code. They're\nthe counters for the trace output. Might be worth a brief comment above\nthem. (I do approve of adding this kind of trace debugging info; I'm\npretty accustomed to using gdb or adding one-off debug statements, but\nwe really could do a better job in general of making these kinds of\ninternals visible to mere mortal admins).\n\n> +static int add_object_entry_from_pack(const struct object_id *oid,\n> +\t\t\t\t      struct packed_git *p,\n> +\t\t\t\t      uint32_t pos,\n> +\t\t\t\t      void *_data)\n> +{\n> +\tstruct rev_info *revs = _data;\n> +\tstruct object_info oi = OBJECT_INFO_INIT;\n> +\toff_t ofs;\n> +\tenum object_type type;\n> +\n> +\tdisplay_progress(progress_state, ++nr_seen);\n> +\n> +\tofs = nth_packed_object_offset(p, pos);\n> +\n> +\toi.typep = &type;\n> +\tif (packed_object_info(the_repository, p, ofs, &oi) < 0)\n> +\t\tdie(_(\"could not get type of object %s in pack %s\"),\n> +\t\t    oid_to_hex(oid), p->pack_name);\n\nCalling out for other reviewers: the oi.typep field will be filled in\nthe with _real_ type of the object, even if it's a delta. This is as\nopposed to the return value of packed_object_info(), which may be\nOFS_DELTA or REF_DELTA.\n\nAnd that real type is what we want here:\n\n> +\telse if (type == OBJ_COMMIT) {\n> +\t\t/*\n> +\t\t * commits in included packs are used as starting points for the\n> +\t\t * subsequent revision walk\n> +\t\t */\n> +\t\tadd_pending_oid(revs, NULL, oid, 0);\n> +\t}\n\nAnd later when we call create_object_entry().\n\nI wondered whether it would be worth adding other objects we might find,\nlike trees, in order to increase our traversal. But that doesn't make\nany sense. The whole point is to find the paths, which come from\ntraversing from the root trees. And we can only find the root trees by\nstarting at commits. Adding any random tree we found would defeat the\npurpose (most of them are sub-trees and would give us a useless partial\npath).\n\nShould we avoid adding the commit as a tip for walking if it won't end\nup in the resulting pack? I.e., should we check these:\n\n> +\tif (have_duplicate_entry(oid, 0))\n> +\t\treturn 0;\n> +\n> +\tif (!want_object_in_pack(oid, 0, &p, &ofs))\n> +\t\treturn 0;\n\n...first? I guess it probably doesn't matter too much since we'd\ntruncate the traversal as soon as we saw it was in a kept pack anyway.\n\n> +static void show_commit_pack_hint(struct commit *commit, void *_data)\n> +{\n> +}\n\nNothing to do here, since commits don't have a name field. Makes sense.\n\n> +static void show_object_pack_hint(struct object *object, const char *name,\n> +\t\t\t\t  void *_data)\n> +{\n> +\tstruct object_entry *oe = packlist_find(&to_pack, &object->oid);\n> +\tif (!oe)\n> +\t\treturn;\n> +\n> +\t/*\n> +\t * Our 'to_pack' list was constructed by iterating all objects packed in\n> +\t * included packs, and so doesn't have a non-zero hash field that you\n> +\t * would typically pick up during a reachability traversal.\n> +\t *\n> +\t * Make a best-effort attempt to fill in the ->hash and ->no_try_delta\n> +\t * here using a now in order to perhaps improve the delta selection\n> +\t * process.\n> +\t */\n> +\toe->hash = pack_name_hash(name);\n> +\toe->no_try_delta = name && no_try_delta(name);\n> +\n> +\tstdin_packs_hints_nr++;\n> +}\n\nBut for actual objects, we do fill in the hash. I wonder if it's\npossible for oe->hash to have been already filled. I don't think it\nreally matters, though. Any value we get is equally valid, so\noverwriting is OK in that case.\n\n> +\tstring_list_sort(&include_packs);\n> +\tstring_list_sort(&exclude_packs);\n> +\n> +\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n> +\t\tconst char *pack_name = pack_basename(p);\n> +\n> +\t\titem = string_list_lookup(&include_packs, pack_name);\n> +\t\tif (!item)\n> +\t\t\titem = string_list_lookup(&exclude_packs, pack_name);\n> +\n> +\t\tif (item)\n> +\t\t\titem->util = p;\n> +\t}\n\nOK, here we're just filling in the util field with each found pack. So\nwe wouldn't notice a pack that we didn't find, but we will in the\nsubsequent loops. Makes sense.\n\nI think you could do without string lists at all by using the recent-ish\npack-hash to efficiently look up the names, but I'm perfectly content to\nsee it all handled within this function.\n\n> +\t/*\n> +\t * First handle all of the excluded packs, marking them as kept in-core\n> +\t * so that later calls to add_object_entry() discards any objects that\n> +\t * are also found in excluded packs.\n> +\t */\n> +\tfor_each_string_list_item(item, &exclude_packs) {\n> +\t\tstruct packed_git *p = item->util;\n> +\t\tif (!p)\n> +\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n> +\t\tp->pack_keep_in_core = 1;\n> +\t}\n> +\tfor_each_string_list_item(item, &include_packs) {\n> +\t\tstruct packed_git *p = item->util;\n> +\t\tif (!p)\n> +\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n> +\t\tfor_each_object_in_pack(p,\n> +\t\t\t\t\tadd_object_entry_from_pack,\n> +\t\t\t\t\t&revs,\n> +\t\t\t\t\tFOR_EACH_OBJECT_PACK_ORDER);\n> +\t}\n\nYeah, this ordering makes sense.\n\n> +\tif (prepare_revision_walk(&revs))\n> +\t\tdie(_(\"revision walk setup failed\"));\n> +\ttraverse_commit_list(&revs,\n> +\t\t\t     show_commit_pack_hint,\n> +\t\t\t     show_object_pack_hint,\n> +\t\t\t     NULL);\n\nAnd this traversal is pretty straight-forward. Looks good.\n\n> +\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_found\",\n> +\t\t\t   stdin_packs_found_nr);\n\nI wonder if it makes sense to report the actual set of packs via trace\n(obviously not as an int, but as a list). That's less helpful for\ndebugging pack-objects, if you just fed it the input anyway, but if you\nwere debugging \"git repack --geometric\" it might be useful to see which\npacks it thought were which (though arguably that would be a useful\ntrace in builtin/repack.c instead).\n\n> @@ -3636,7 +3797,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n>  \t\tuse_internal_rev_list = 1;\n>  \t\tstrvec_push(&rp, \"--indexed-objects\");\n>  \t}\n> -\tif (rev_list_unpacked) {\n> +\tif (rev_list_unpacked && !stdin_packs) {\n>  \t\tuse_internal_rev_list = 1;\n>  \t\tstrvec_push(&rp, \"--unpacked\");\n>  \t}\n\nOK, this is necessary to avoid triggering the internal rev-list, because\nwe handle --unpacked ourselves specially later here...\n\n> @@ -3741,7 +3907,13 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n>  \n>  \tif (progress)\n>  \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n> -\tif (!use_internal_rev_list)\n> +\tif (stdin_packs) {\n> +\t\t/* avoids adding objects in excluded packs */\n> +\t\tignore_packed_keep_in_core = 1;\n> +\t\tread_packs_list_from_stdin();\n> +\t\tif (rev_list_unpacked)\n> +\t\t\tadd_unreachable_loose_objects();\n\nWhich isn't quite behaving like normal --unpacked (in that we are adding\nall loose objects, not just reachable ones). I think we actually could\njust add --unpacked as part of our heuristic traversal. It's not\nperfect, but unlike the packed objects, it's OK for us to miss some\ncorner cases (they just end up not getting packed; they don't get\ndeleted).\n\nI'm OK to consider that an implementation detail for now, though. We can\nchange it later without impacting the interface.\n\n> +\t\tif (rev_list_unpacked)\n> +\t\t\tadd_unreachable_loose_objects();\n\nDespite the name, that function is adding both reachable and unreachable\nones. So it is doing what you want. It might be worth renaming, but it's\nnot too big a deal since it's local to this file.\n\n-Peff\n"},{"id":"417125","messageId":"YCxcGKo7kyLwVvw+@coredump.intra.peff.net","threadId":"55012","inReplyTo":"b5081c01b53beb568ef2e59956d25b36be9f24d0.1612411124.git.me@ttaylorr.com","subject":"Re: [PATCH v2 5/8] p5303: measure time to repack with keep","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-16T23:58:16Z","receivedAt":"2021-02-16T23:59:01Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:59:13PM -0500, Taylor Blau wrote:\n\n> From: Jeff King <peff@peff.net>\n> \n> This is the same as the regular repack test, except that we mark the\n> single base pack as \"kept\" and use --assume-kept-packs-closed. The\n\nI don't think that option exists anymore. I guess we are just using\n--stdin-packs, which causes us to mark a pack as kept.\n\nI think we could just mark it in the filesystem and use\n--honor-pack-keep, which would make it independent of your new feature.\nAt first I was going to say \"but it doesn't matter either way\", but...\n\n> theory is that this should be faster than the normal repack, because\n> we'll have fewer objects to traverse and process.\n> \n> Here are some timings on a recent clone of the kernel. In the\n> single-pack case, there is nothing do since there are no non-excluded\n> packs:\n> \n>   5303.5: repack (1)                          57.42(54.88+10.64)\n>   5303.6: repack with --stdin-packs (1)       0.01(0.01+0.00)\n> \n> and in the 50-pack case, it is much faster to use `--stdin-packs`, since\n> we avoid having to consider any objects in the excluded pack:\n> \n>   5303.10: repack (50)                        71.26(88.24+4.96)\n>   5303.11: repack with --stdin-packs (50)     3.49(11.82+0.28)\n> \n> but our improvements vanish as we approach 1000 packs.\n> \n>   5303.15: repack (1000)                      215.64(491.33+14.80)\n>   5303.16: repack with --stdin-packs (1000)   198.79(380.51+7.97)\n> \n> That's because the code paths around handling .keep files are known to\n> scale badly; they look in every single pack file to find each object.\n\nWell, part of it is just that with 1000 packs we have 20 times as many\nobjects that are actually getting packed with --stdin-packs, compared to\nthe 50-pack case. IIRC, each pack is a fixed-size slice and then the\nresidual is put into the .keep pack. So the fact that the time gets\ncloser to a full repack as we add more packs is expected: we are asking\npack-objects to do more work!\n\nFor showing the impact of the optimizations in patches 7 and 8, I think\ndoing a full repack with --honor-pack-keep is a better test. Because\nthen we're always doing a full traversal, and most of the work continues\nto scale with the repo size (though obviously not the actual shuffling\nof packed bytes around). That would get rid of the weird \"no work to do\"\ncase in the single-pack tests, too.\n\n-Peff\n"},{"id":"417126","messageId":"YCxcyQcNO9LO2n9m@coredump.intra.peff.net","threadId":"55012","inReplyTo":"cover.1612411123.git.me@ttaylorr.com","subject":"Re: [PATCH v2 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T00:01:13Z","receivedAt":"2021-02-17T00:01:56Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:58:45PM -0500, Taylor Blau wrote:\n\n> The details of the new approach can be found in the third patch, but the gist is\n> as follows:\n> [...]\n\nI think this turned out very nice (and less complicated than I feared it\nmight). I've read up through patch 5. I think the overall approach is\ngood, but I had various small-to-medium comments.\n\nI'll try to pick up reviewing the rest tomorrow, though it may make\nsense to resolve the earlier comments first.\n\n-Peff\n"},{"id":"417127","messageId":"YCxdL9Xi4nAcnqIg@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YCxcGKo7kyLwVvw+@coredump.intra.peff.net","subject":"Re: [PATCH v2 5/8] p5303: measure time to repack with keep","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T00:02:55Z","receivedAt":"2021-02-17T00:03:38Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 16, 2021 at 06:58:16PM -0500, Jeff King wrote:\n\n> For showing the impact of the optimizations in patches 7 and 8, I think\n> doing a full repack with --honor-pack-keep is a better test. Because\n> then we're always doing a full traversal, and most of the work continues\n> to scale with the repo size (though obviously not the actual shuffling\n> of packed bytes around). That would get rid of the weird \"no work to do\"\n> case in the single-pack tests, too.\n\nI meant to add: but I do like that we are timing --stdin-packs, too. We\nmay actually want to time both.\n\nAnother thing we _could_ do, if we have --honor-pack-keep perf tests, is\nto shuffle patches 5, 6, and 7 towards the front of the series. They\nshould be able to show off the improvement even without the\n--stdin-packs feature.\n\n-Peff\n"},{"id":"417159","messageId":"YC0+wlRksoqm0xLO@coredump.intra.peff.net","threadId":"55012","inReplyTo":"c3868c7df92484f0527ce500ad1156275be334e8.1612411124.git.me@ttaylorr.com","subject":"Re: [PATCH v2 6/8] builtin/pack-objects.c: rewrite honor-pack-keep logic","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T16:05:22Z","receivedAt":"2021-02-17T16:06:22Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:59:17PM -0500, Taylor Blau wrote:\n\n> @@ -1209,22 +1210,73 @@ static int want_found_object(int exclude, struct packed_git *p)\n>  \t * Otherwise, we signal \"-1\" at the end to tell the caller that we do\n>  \t * not know either way, and it needs to check more packs.\n>  \t */\n> -\tif (!ignore_packed_keep_on_disk &&\n> -\t    !ignore_packed_keep_in_core &&\n> -\t    (!local || !have_non_local_packs))\n> +\n> +\t/*\n> +\t * Handle .keep first, as we have a fast(er) path there.\n> +\t */\n> +\tif (ignore_packed_keep_on_disk || ignore_packed_keep_in_core) {\n> +\t\t/*\n> +\t\t * Set the flags for the kept-pack cache to be the ones we want\n> +\t\t * to ignore.\n> +\t\t *\n> +\t\t * That is, if we are ignoring objects in on-disk keep packs,\n> +\t\t * then we want to search through the on-disk keep and ignore\n> +\t\t * the in-core ones.\n> +\t\t */\n> +\t\tunsigned flags = 0;\n> +\t\tif (ignore_packed_keep_on_disk)\n> +\t\t\tflags |= ON_DISK_KEEP_PACKS;\n> +\t\tif (ignore_packed_keep_in_core)\n> +\t\t\tflags |= IN_CORE_KEEP_PACKS;\n> +\n> +\t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n> +\t\t\treturn 0;\n> +\t\tif (ignore_packed_keep_in_core && p->pack_keep_in_core)\n> +\t\t\treturn 0;\n> +\t\tif (has_object_kept_pack(oid, flags))\n> +\t\t\treturn 0;\n> +\t}\n> +\n> +\t/*\n> +\t * At this point we know definitively that either we don't care about\n> +\t * keep-packs, or the object is not in one. Keep checking other\n> +\t * conditions...\n> +\t */\n> +\n> +\tif (!local || !have_non_local_packs)\n>  \t\treturn 1;\n> -\n>  \tif (local && !p->pack_local)\n>  \t\treturn 0;\n> -\tif (p->pack_local &&\n> -\t    ((ignore_packed_keep_on_disk && p->pack_keep) ||\n> -\t     (ignore_packed_keep_in_core && p->pack_keep_in_core)))\n> -\t\treturn 0;\n>  \n>  \t/* we don't know yet; keep looking for more packs */\n>  \treturn -1;\n\nI know I wrote this patch, but just looking it over again with a\ncritical eye: it looks like more re-ordering could avoid work in some\ncases.\n\nIn particular, has_object_kept_pack() is a potentially expensive call.\nBut if \"(local && !p->pack_local)\" is true, then we could cheaply exit\nthe function with \"0\", regardless of what the keep requirement says.\n\nThat's not a case that I think anybody cares that deeply about (and it\ncertainly is not covered by t/perf). But I think it does regress in this\npatch. Prior to the patch, we'd check that condition before returning\n-1, and it was the caller who would then continue to search through all\nthe kept packs. Now we do it preemptively.\n\nI think just bumping that:\n\n  if (local && !p->pack_local)\n\treturn 0;\n\nabove the new code would fix it. Or to lay out the logic more fully, the\norder of checks should be:\n\n  - does _this_ pack we found the object in disqualify it. If so, we can\n    cheaply return 0. And that applies to both keep and local rules.\n\n  - otherwise, check all packs via has_object_kept_pack(), which is\n    cheaper than continuing to iterate through all packs by returning\n    -1.\n\n  - once we know definitively about keep-packs, then check any shortcuts\n    related to local packs (like !have_non_local_packs)\n\n  - and then if no shortcuts, we return -1\n\nI think that might be easier to express by rewriting the patch. :)\n\n-Peff\n"},{"id":"417162","messageId":"YC1OJDFXPnxGMHPK@coredump.intra.peff.net","threadId":"55012","inReplyTo":"f1c07324f62cf4d087c41165cefed98f554cfd78.1612411124.git.me@ttaylorr.com","subject":"Re: [PATCH v2 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T17:11:00Z","receivedAt":"2021-02-17T17:11:45Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:59:21PM -0500, Taylor Blau wrote:\n\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> index fbd7b54d70..b2ba5aa14f 100644\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n> @@ -1225,9 +1225,9 @@ static int want_found_object(const struct object_id *oid, int exclude,\n>  \t\t */\n>  \t\tunsigned flags = 0;\n>  \t\tif (ignore_packed_keep_on_disk)\n> -\t\t\tflags |= ON_DISK_KEEP_PACKS;\n> +\t\t\tflags |= CACHE_ON_DISK_KEEP_PACKS;\n>  \t\tif (ignore_packed_keep_in_core)\n> -\t\t\tflags |= IN_CORE_KEEP_PACKS;\n> +\t\t\tflags |= CACHE_IN_CORE_KEEP_PACKS;\n\nWhy are we renaming the constants in this patch?\n\nI know I'm listed as the author, but I think this came out of some\noff-list back and forth between us. It seems like the existing constants\nwould have been fine.\n\n> +static void maybe_invalidate_kept_pack_cache(struct repository *r,\n> +\t\t\t\t\t     unsigned flags)\n>  {\n> -\treturn find_one_pack_entry(r, oid, e, 0);\n> +\tif (!r->objects->kept_pack_cache)\n> +\t\treturn;\n> +\tif (r->objects->kept_pack_cache->flags == flags)\n> +\t\treturn;\n> +\tfree(r->objects->kept_pack_cache->packs);\n> +\tFREE_AND_NULL(r->objects->kept_pack_cache);\n> +}\n\nOK, so we keep a single cache based on the flags, and then if somebody\never asks for different flags, we throw it away. That's probably OK for\nour purposes, since we wouldn't expect multiple callers within a single\nprocess.\n\nI wondered if it would be simpler to just keep two lists, one for\nin-core keeps and one for on-disk keeps. And then just walk over each\nlist separately based on the query flags. That makes things more robust\n_and_ I think would be less code. It does mean that a pack could appear\nin both lists, though, which means we might do a lookup in it twice.\nThat doesn't seem all that likely, but it is working against our goal\nhere.\n\nAnother option is to keep 3 caches (two separate and one combined),\nrather than flipping between them. I'm not sure if that would be less\ncode or not (it gets rid of the \"invalidate\" function, but you do have\nto pick the right cache depending on the query flags).\n\nYet another option is to keep a cache of any that are marked as _either_\nin core or on-disk keeps, and then decide to look up the object based on\nthe query flags. Then you just pay the cost to iterate over the list and\ncheck the flags (which really is all this cache is helping with in the\nfirst place).\n\nI dunno. TBH, I kind of wonder if this whole patch is worth doing at\nall, giving the underwhelming performance benefit (3% on the\npathological 1000-pack case). When I had timed this strategy initially,\nit was more like 15%. I'm not sure where the savings went in the\ninterim, or if it was a timing fluke.\n\n> +static struct packed_git **kept_pack_cache(struct repository *r, unsigned flags)\n> +{\n> +\tmaybe_invalidate_kept_pack_cache(r, flags);\n> +\n> +\tif (!r->objects->kept_pack_cache) {\n> +\t\tstruct packed_git **packs = NULL;\n> +\t\tsize_t nr = 0, alloc = 0;\n> +\t\tstruct packed_git *p;\n> +\n> +\t\t/*\n> +\t\t * We want \"all\" packs here, because we need to cover ones that\n> +\t\t * are used by a midx, as well. We need to look in every one of\n> +\t\t * them (instead of the midx itself) to cover duplicates. It's\n> +\t\t * possible that an object is found in two packs that the midx\n> +\t\t * covers, one kept and one not kept, but the midx returns only\n> +\t\t * the non-kept version.\n> +\t\t */\n> +\t\tfor (p = get_all_packs(r); p; p = p->next) {\n> +\t\t\tif ((p->pack_keep && (flags & CACHE_ON_DISK_KEEP_PACKS)) ||\n> +\t\t\t    (p->pack_keep_in_core && (flags & CACHE_IN_CORE_KEEP_PACKS))) {\n> +\t\t\t\tALLOC_GROW(packs, nr + 1, alloc);\n> +\t\t\t\tpacks[nr++] = p;\n> +\t\t\t}\n> +\t\t}\n> +\t\tALLOC_GROW(packs, nr + 1, alloc);\n> +\t\tpacks[nr] = NULL;\n> +\n> +\t\tr->objects->kept_pack_cache = xmalloc(sizeof(*r->objects->kept_pack_cache));\n> +\t\tr->objects->kept_pack_cache->packs = packs;\n> +\t\tr->objects->kept_pack_cache->flags = flags;\n> +\t}\n\nIs there any reason not to just embed the kept_pack_cache struct inside\nthe object_store? It's one less pointer to deal with. I wonder if this\nis a holdover from an attempt to have multiple caches.\n\n(I also think it would be reasonable if we wanted to hide the definition\nof the cache struct from callers, but we don't seem do to that).\n\n> @@ -2109,7 +2123,8 @@ int has_object_pack(const struct object_id *oid)\n>  \treturn find_pack_entry(the_repository, oid, &e);\n>  }\n>  \n> -int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n> +int has_object_kept_pack(const struct object_id *oid,\n> +\t\t\t unsigned flags)\n>  {\n>  \tstruct pack_entry e;\n>  \treturn find_kept_pack_entry(the_repository, oid, flags, &e);\n\nThis seems like a stray change.\n\n> diff --git a/packfile.h b/packfile.h\n> index 624327f64d..eb56db2a7b 100644\n> --- a/packfile.h\n> +++ b/packfile.h\n> @@ -161,10 +161,6 @@ int packed_object_info(struct repository *r,\n>  void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n>  const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n>  \n> -#define ON_DISK_KEEP_PACKS 1\n> -#define IN_CORE_KEEP_PACKS 2\n> -#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n\nI notice that when the constants moved, we didn't keep an equivalent of\nALL_KEEP_PACKS. Maybe we didn't need it in the first place in patch 1?\n\n  BTW, I absolutely hate the complication that all of this on-disk\n  versus in-core keep distinction brings to this code. And I wondered\n  what it was really doing for us and whether we could get rid of it.\n  But I think we do need it: a common case may be to avoid using\n  --honor-pack-keep (because you don't want to deal with racy .keep\n  writes from incoming receive-pack processes), but use in-core ones for\n  something like --stdin-packs. So we do need to respect one and not the\n  other.\n\n  I do wonder if things would be simpler if pack-objects simply kept its\n  own list of \"in core\" packs in a separate array. But that is really\n  just another form of the same problem, I guess.\n\n-Peff\n"},{"id":"417167","messageId":"YC1drGrIEg0C7Zo5@coredump.intra.peff.net","threadId":"55012","inReplyTo":"d5561585c2221a9635eb0fc7a65298ee8a2b6348.1612411124.git.me@ttaylorr.com","subject":"Re: [PATCH v2 8/8] builtin/repack.c: add '--geometric' option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T18:17:16Z","receivedAt":"2021-02-17T18:18:21Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 03, 2021 at 10:59:25PM -0500, Taylor Blau wrote:\n\n> Often it is useful to both:\n> \n>   - have relatively few packfiles in a repository, and\n> \n>   - avoid having so few packfiles in a repository that we repack its\n>     entire contents regularly\n> \n> This patch implements a '--geometric=<n>' option in 'git repack'. This\n> allows the caller to specify that they would like each pack to be at\n> least a factor times as large as the previous largest pack (by object\n> count).\n> \n> Concretely, say that a repository has 'n' packfiles, labeled P1, P2,\n> ..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\n> With a geometric factor of 'r', it should be that:\n> \n>   objects(Pi) > r*objects(P(i-1))\n> \n> for all i in [1, n], where the packs are sorted by\n> \n>   objects(P1) <= objects(P2) <= ... <= objects(Pn).\n\nJust devil's advocating for a moment.\n\nI think in this kind of geometric roll-up strategy, you want to imagine\nthat you are rolling up recent pushes but leaving untouched a good\n\"base\" pack that you previously created.\n\nAnd that will usually be true if you are doing the rollup based on\nnumber of objects (or size, etc). But it won't always be (e.g., for some\nreason somebody makes a very large push relative to the current\nrepository size). What happens when this assumption is violated?\n\nIn some ways, it is a good thing to drift away from this \"base pack\"\nview of the world. If we're trying to amortize the per-object work done,\nthen we are better off rolling up the small things into the large,\nregardless of where they came from.\n\nBut the base pack may also have other properties we want to retain. Two\nI can think of:\n\n  - it may have a .bitmap that we'll be throwing away, without\n    generating a new one. I know that your end-game involves writing a\n    midx with bitmaps that covers all of the packs, so this would become\n    a non-issue in that strategy.\n\n  - it may have been more carefully packed (e.g., with a larger window\n    size, using \"-f\", etc) than the packs we got from pushes. We do\n    _mostly_ retain the deltas when we roll up the packs, so it probably\n    only has a small impact in practice (I'd expect in a few cases we'd\n    throw away deltas because a pushed pack contains a duplicate of its\n    base object that we added via --fix-thin).\n\nSo I suspect it's probably OK in practice. These cases would happen\nrarely, and the impact would not be all that big. The bitmap thing I'd\nworry the most about. As part of a larger strategy involving a midx it\nis taken care of, but people using just this new feature may not realize\nthat. The bitmaps of course are \"just\" an optimization, but it's hard to\nsay how dire things are when they don't exist. For many situations,\nprobably not very dire. But I know that on our servers, when repos lack\nbitmaps, people notice the performance degradation.\n\nOn the other hand, by definition this happens in a case where there are\nmore objects that have just been pushed (and are therefore not\nbitmapped) than existed already. So you _already_ have a performance\nproblem either way until you get bitmap coverage of those new objects.\n\n> --- a/Documentation/git-repack.txt\n> +++ b/Documentation/git-repack.txt\n> @@ -165,6 +165,17 @@ depth is 4095.\n>  \tPass the `--delta-islands` option to `git-pack-objects`, see\n>  \tlinkgit:git-pack-objects[1].\n>  \n> +-g=<factor>::\n> +--geometric=<factor>::\n> +\tArrange resulting pack structure so that each successive pack\n> +\tcontains at least `<factor>` times the number of objects as the\n> +\tnext-largest pack.\n> ++\n> +`git repack` ensures this by determining a \"cut\" of packfiles that need to be\n> +repacked into one in order to ensure a geometric progression. It picks the\n> +smallest set of packfiles such that as many of the larger packfiles (by count of\n> +objects contained in that pack) may be left intact.\n\nI think we might need to make clear in the documentation how this\ndiffers from other repacks, in that it is not considering reachability\nat all. I like the term \"roll up\" to describe what is happening, but we\nprobably need to define that term clearly, as well.\n\nEspecially important, I think, is that we talk about what's happening\nwith loose objects, which are part of the rollup here. And IMHO we\nshould make clear that for now we include them all, without\nconsideration of their reachability, but that this may change in the\nfuture.\n\nLikewise, are there any options that are incompatible with \"-g\"? I have\nto imagine that \"--write-bitmap-index\" would not work very well. I don't\nknow that we need to enumerate them all, but I'm wondering if a blanket\n\"this may not play well with other options\" warning may be advisable.\n\n> +static void split_pack_geometry(struct pack_geometry *geometry, int factor)\n> [...]\n\nI'll admit I didn't carefully think about the math of the progression\nhere. IMHO the exact split is the least interesting part of this whole\nseries (compared to the general idea of \"rolling up some packs\" versus a\nwhole repack). Between the comments and the tests, I'll assume it's\ngenerally behaving as advertised. (I of course did look for any obvious\ncoding errors, but didn't see any).\n\n-Peff\n"},{"id":"417168","messageId":"YC1eD+m5acCpyEf2@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YCxcyQcNO9LO2n9m@coredump.intra.peff.net","subject":"Re: [PATCH v2 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T18:18:55Z","receivedAt":"2021-02-17T18:19:39Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 16, 2021 at 07:01:13PM -0500, Jeff King wrote:\n\n> On Wed, Feb 03, 2021 at 10:58:45PM -0500, Taylor Blau wrote:\n> \n> > The details of the new approach can be found in the third patch, but the gist is\n> > as follows:\n> > [...]\n> \n> I think this turned out very nice (and less complicated than I feared it\n> might). I've read up through patch 5. I think the overall approach is\n> good, but I had various small-to-medium comments.\n> \n> I'll try to pick up reviewing the rest tomorrow, though it may make\n> sense to resolve the earlier comments first.\n\nOK, I finished reading the rest and left a few more comments. The short\nof it is that I really like the new direction, but I think there are\nenough small comments to merit a re-roll, which I hope would probably be\nthe final.\n\n-Peff\n"},{"id":"417170","messageId":"YC1iBEFyQnv2URSV@nand.local","threadId":"55012","inReplyTo":"YCxSlNK4Ug2exDbv@coredump.intra.peff.net","subject":"Re: [PATCH v2 2/8] revision: learn '--no-kept-objects'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T18:35:48Z","receivedAt":"2021-02-17T18:36:51Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Feb 16, 2021 at 06:17:40PM -0500, Jeff King wrote:\n> On Wed, Feb 03, 2021 at 10:58:57PM -0500, Taylor Blau wrote:\n>\n> > @@ -3797,6 +3807,11 @@ enum commit_action get_commit_action(struct rev_info *revs, struct commit *commi\n> >  \t\treturn commit_ignore;\n> >  \tif (revs->unpacked && has_object_pack(&commit->object.oid))\n> >  \t\treturn commit_ignore;\n> > +\tif (revs->no_kept_objects) {\n> > +\t\tif (has_object_kept_pack(&commit->object.oid,\n> > +\t\t\t\t\t revs->keep_pack_cache_flags))\n> > +\t\t\treturn commit_ignore;\n> > +\t}\n>\n> OK, so this has the same \"problems\" as --unpacked, which is that we can\n> miss some objects (i.e., things that are reachable but not-kept may not\n> be reported). But it should be OK in this version of the series, because\n> we will not be relying on it for selection of objects, but only to fill\n> in ordering / namehash fields.\n>\n> Should we warn people about that, either as a comment or in the commit\n> message?\n\nYeah, let's warn about it in the commit message. We could put it in the\ndocumentation, but...\n\n> > +--no-kept-objects[=<kind>]::\n> > +\tHalts the traversal as soon as an object in a kept pack is\n> > +\tfound. If `<kind>` is `on-disk`, only packs with a corresponding\n> > +\t`*.keep` file are ignored. If `<kind>` is `in-core`, only packs\n> > +\twith their in-core kept state set are ignored. Otherwise, both\n> > +\tkinds of kept packs are ignored.\n>\n> Likewise, I wonder whether we need to expose this mode to users.\n> Normally I'm a fan of doing so, because it allows scripted callers\n> access to more of the internals, but:\n>\n>   - the semantics are kind of weird about where we draw the line between\n>     performance and absolute correctness\n>\n>   - the \"in-core\" thing is a bit weird for callers of rev-list; how do I\n>     as a caller mark a pack as kept-in-core? I think it's only an\n>     internal pack-objects thing.\n>\n> Once we support this in rev-list, we'll have to do it forever (or deal\n> with deprecation, etc). If we just need it internally, maybe it's wise\n> to leave it as a something you ask for by manipulating rev_info\n> directly. Or perhaps leave it as an undocumented interface we use for\n> testing, and not something we promise to keep working.\n\nI think that you raise a good point about not advertising this option,\nsince doing so paints us into a corner that we have to keep it working\nand behaving consistently forever.\n\nI'm not opposed to the idea that we may eventually want to do so, but I\nthink that this is too early for that. As you note, we *could* just\nexpose it in rev_info flags, but that makes it much more difficult to\ntest some of the tricky cases that are added in t6114, so I think a\nmiddle ground of having an undocumented option satisfies both of our\nwants.\n\n> > --- a/list-objects.c\n> > +++ b/list-objects.c\n> > @@ -338,6 +338,13 @@ static void traverse_trees_and_blobs(struct traversal_context *ctx,\n> >  \t\t\tctx->show_object(obj, name, ctx->show_data);\n> >  \t\t\tcontinue;\n> >  \t\t}\n> > +\t\tif (ctx->revs->no_kept_objects) {\n> > +\t\t\tstruct pack_entry e;\n> > +\t\t\tif (find_kept_pack_entry(ctx->revs->repo, &obj->oid,\n> > +\t\t\t\t\t\t ctx->revs->keep_pack_cache_flags,\n> > +\t\t\t\t\t\t &e))\n> > +\t\t\t\tcontinue;\n> > +\t\t}\n>\n> This hunk is interesting.\n>\n> There is no similar check for revs->unpacked in list-objects.c to cut\n> off the traversal. And indeed, running \"rev-list --unpacked\" will\n> generally look at the _whole_ tree for a commit that is unpacked, even\n> if all of the tree entries are packed. That's something we might\n> consider changing in the name of performance (though it does increase\n> the number of cases where --unpacked will fail to find an unpacked but\n> reachable object).\n>\n> But this is a funny place to put it. If I understand it correctly, it is\n> cutting off the traversal at the very top of the tree. I.e., if we had a\n> commit that is not-kept, we'd queue it's root tree. And then we might\n> find that the root tree is kept, and avoid traversing it. But if we _do_\n> traverse it, we would look at every subtree it contains, even if they\n> are kept! That's because we recurse the tree via the recursive\n> process_tree(), not by queueing more objects in the pending array here.\n>\n> So this check seems to exist in a funny middle ground. I think it's\n> unlikely to catch anything useful (usually commits have a unique root\n> tree; it's all of the untouched parts of the subtrees that will be in\n> the kept packs). IMHO we should either drop it (and act like\n> \"--unpacked\", accepting that we may traverse some extra tree objects),\n> or we should go all-in on performance and cut it off in the top of\n> process_tree().\n\nAgreed. Let's drop it.\n\n> -Peff\n\nThanks,\nTaylor\n"},{"id":"417173","messageId":"YC1nfH356wfmAKE2@nand.local","threadId":"55012","inReplyTo":"YCxZcz5JGtxObOF3@coredump.intra.peff.net","subject":"Re: [PATCH v2 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T18:59:08Z","receivedAt":"2021-02-17T19:00:10Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Feb 16, 2021 at 06:46:59PM -0500, Jeff King wrote:\n> Do we use this partial traversal to impact the write order at all? That\n> would be a nice-to-have, but I suspect that just concatenating the packs\n> (presumably by descending mtime) ends up with a similar result.\n\nWe don't; the objects are written in pack order. In the version of the\npatch you reviewed, the order of packs was determined by their hash (due\nto the string_list_sort()), but the version I just prepared re-sorts by\nmtime.\n\nIt's kind of gross, since we need to use QSORT directly on the\nstring_list internals in order to have access to the ->util field of the\nstring_list_items (string_list_sort() only lets you compare strings\ndirectly for obvious reasons).\n\nI added a comment describing this hack.\n\n> > +--stdin-packs::\n> > +\tRead the basenames of packfiles from the standard input, instead\n> > +\tof object names or revision arguments. The resulting pack\n> > +\tcontains all objects listed in the included packs (those not\n> > +\tbeginning with `^`), excluding any objects listed in the\n> > +\texcluded packs (beginning with `^`).\n> > ++\n> > +Incompatible with `--revs`, or options that imply `--revs` (such as\n> > +`--all`), with the exception of `--unpacked`, which is compatible.\n>\n> I know you say \"basename\" here, but I wonder if it is worth giving an\n> example (`pack-1234abcd.pack`) to make it clear in what form we expect\n> it. Or possibly something in the `EXAMPLES` section.\n\nGood idea, thanks.\n\n> > --- a/builtin/pack-objects.c\n> > +++ b/builtin/pack-objects.c\n> > @@ -2979,6 +2979,164 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n> >  \treturn git_default_config(k, v, cb);\n> >  }\n> >\n> > +static int stdin_packs_found_nr;\n> > +static int stdin_packs_hints_nr;\n>\n> I scratched my head at these until I looked further in the code. They're\n> the counters for the trace output. Might be worth a brief comment above\n> them. (I do approve of adding this kind of trace debugging info; I'm\n> pretty accustomed to using gdb or adding one-off debug statements, but\n> we really could do a better job in general of making these kinds of\n> internals visible to mere mortal admins).\n\nGood call.\n\n> > +static int add_object_entry_from_pack(const struct object_id *oid,\n> > +\t\t\t\t      struct packed_git *p,\n> > +\t\t\t\t      uint32_t pos,\n> > +\t\t\t\t      void *_data)\n> > +{\n> > +\tstruct rev_info *revs = _data;\n> > +\tstruct object_info oi = OBJECT_INFO_INIT;\n> > +\toff_t ofs;\n> > +\tenum object_type type;\n> > +\n> > +\tdisplay_progress(progress_state, ++nr_seen);\n> > +\n> > +\tofs = nth_packed_object_offset(p, pos);\n> > +\n> > +\toi.typep = &type;\n> > +\tif (packed_object_info(the_repository, p, ofs, &oi) < 0)\n> > +\t\tdie(_(\"could not get type of object %s in pack %s\"),\n> > +\t\t    oid_to_hex(oid), p->pack_name);\n>\n> Calling out for other reviewers: the oi.typep field will be filled in\n> the with _real_ type of the object, even if it's a delta. This is as\n> opposed to the return value of packed_object_info(), which may be\n> OFS_DELTA or REF_DELTA.\n>\n> And that real type is what we want here:\n>\n> > +\telse if (type == OBJ_COMMIT) {\n> > +\t\t/*\n> > +\t\t * commits in included packs are used as starting points for the\n> > +\t\t * subsequent revision walk\n> > +\t\t */\n> > +\t\tadd_pending_oid(revs, NULL, oid, 0);\n> > +\t}\n>\n> And later when we call create_object_entry().\n\n:-). Yes indeed. As I'm sure that you will recall, the pack-objects\ncode _does not_ behave well when you give it the packed type of an\nobject (which is not entirely unexpected, since the pack-objects code\nonly operates on the true type, so passing the packed type--as I did\nwhen originally writing this patch--is a bug).\n\n> I wondered whether it would be worth adding other objects we might find,\n> like trees, in order to increase our traversal. But that doesn't make\n> any sense. The whole point is to find the paths, which come from\n> traversing from the root trees. And we can only find the root trees by\n> starting at commits. Adding any random tree we found would defeat the\n> purpose (most of them are sub-trees and would give us a useless partial\n> path).\n\nRight.\n\n> Should we avoid adding the commit as a tip for walking if it won't end\n> up in the resulting pack? I.e., should we check these:\n>\n> > +\tif (have_duplicate_entry(oid, 0))\n> > +\t\treturn 0;\n> > +\n> > +\tif (!want_object_in_pack(oid, 0, &p, &ofs))\n> > +\t\treturn 0;\n>\n> ...first? I guess it probably doesn't matter too much since we'd\n> truncate the traversal as soon as we saw it was in a kept pack anyway.\n\nI agree it doesn't make a difference, but I think placing the extra\nguards first makes it easier to read (since the reader doesn't have to\nconsider how the subsequent traversal would treat it).\n\n> > +static void show_commit_pack_hint(struct commit *commit, void *_data)\n> > +{\n> > +}\n>\n> Nothing to do here, since commits don't have a name field. Makes sense.\n\nYeah. I added a comment to say the same thing, just for extra clarity.\n\n>\n> > +static void show_object_pack_hint(struct object *object, const char *name,\n> > +\t\t\t\t  void *_data)\n> > +{\n> > +\tstruct object_entry *oe = packlist_find(&to_pack, &object->oid);\n> > +\tif (!oe)\n> > +\t\treturn;\n> > +\n> > +\t/*\n> > +\t * Our 'to_pack' list was constructed by iterating all objects packed in\n> > +\t * included packs, and so doesn't have a non-zero hash field that you\n> > +\t * would typically pick up during a reachability traversal.\n> > +\t *\n> > +\t * Make a best-effort attempt to fill in the ->hash and ->no_try_delta\n> > +\t * here using a now in order to perhaps improve the delta selection\n> > +\t * process.\n> > +\t */\n> > +\toe->hash = pack_name_hash(name);\n> > +\toe->no_try_delta = name && no_try_delta(name);\n> > +\n> > +\tstdin_packs_hints_nr++;\n> > +}\n>\n> But for actual objects, we do fill in the hash. I wonder if it's\n> possible for oe->hash to have been already filled. I don't think it\n> really matters, though. Any value we get is equally valid, so\n> overwriting is OK in that case.\n\nRight.\n\n> > +\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_found\",\n> > +\t\t\t   stdin_packs_found_nr);\n>\n> I wonder if it makes sense to report the actual set of packs via trace\n> (obviously not as an int, but as a list). That's less helpful for\n> debugging pack-objects, if you just fed it the input anyway, but if you\n> were debugging \"git repack --geometric\" it might be useful to see which\n> packs it thought were which (though arguably that would be a useful\n> trace in builtin/repack.c instead).\n\nI could see an argument in both ways. I'd rather pass for now until we\nhave a clearer need for it.\n\n> [passing --unpacked to the namehash traversal]\n>\n> I'm OK to consider that an implementation detail for now, though. We can\n> change it later without impacting the interface.\n\nAgreed.\n\n> > +\t\tif (rev_list_unpacked)\n> > +\t\t\tadd_unreachable_loose_objects();\n>\n> Despite the name, that function is adding both reachable and unreachable\n> ones. So it is doing what you want. It might be worth renaming, but it's\n> not too big a deal since it's local to this file.\n\nYeah, I tend to err on the side of \"it's fine as-is\" since this isn't\nexposed outside of pack-objects internals. If you feel strongly I'm\nhappy to change it, but I suspect you don't.\n\nThanks,\nTaylor\n"},{"id":"417175","messageId":"YC1q64fQxHBMxmyw@nand.local","threadId":"55012","inReplyTo":"YCxcGKo7kyLwVvw+@coredump.intra.peff.net","subject":"Re: [PATCH v2 5/8] p5303: measure time to repack with keep","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T19:13:47Z","receivedAt":"2021-02-17T19:14:50Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Feb 16, 2021 at 06:58:16PM -0500, Jeff King wrote:\n> On Wed, Feb 03, 2021 at 10:59:13PM -0500, Taylor Blau wrote:\n>\n> > From: Jeff King <peff@peff.net>\n> >\n> > This is the same as the regular repack test, except that we mark the\n> > single base pack as \"kept\" and use --assume-kept-packs-closed. The\n>\n> I don't think that option exists anymore. I guess we are just using\n> --stdin-packs, which causes us to mark a pack as kept.\n>\n> I think we could just mark it in the filesystem and use\n> --honor-pack-keep, which would make it independent of your new feature.\n> At first I was going to say \"but it doesn't matter either way\", but...\n>\n> > theory is that this should be faster than the normal repack, because\n> > we'll have fewer objects to traverse and process.\n> >\n> > Here are some timings on a recent clone of the kernel. In the\n> > single-pack case, there is nothing do since there are no non-excluded\n> > packs:\n> >\n> >   5303.5: repack (1)                          57.42(54.88+10.64)\n> >   5303.6: repack with --stdin-packs (1)       0.01(0.01+0.00)\n> >\n> > and in the 50-pack case, it is much faster to use `--stdin-packs`, since\n> > we avoid having to consider any objects in the excluded pack:\n> >\n> >   5303.10: repack (50)                        71.26(88.24+4.96)\n> >   5303.11: repack with --stdin-packs (50)     3.49(11.82+0.28)\n> >\n> > but our improvements vanish as we approach 1000 packs.\n> >\n> >   5303.15: repack (1000)                      215.64(491.33+14.80)\n> >   5303.16: repack with --stdin-packs (1000)   198.79(380.51+7.97)\n> >\n> > That's because the code paths around handling .keep files are known to\n> > scale badly; they look in every single pack file to find each object.\n>\n> Well, part of it is just that with 1000 packs we have 20 times as many\n> objects that are actually getting packed with --stdin-packs, compared to\n> the 50-pack case. IIRC, each pack is a fixed-size slice and then the\n> residual is put into the .keep pack. So the fact that the time gets\n> closer to a full repack as we add more packs is expected: we are asking\n> pack-objects to do more work!\n\nNo, the residual base pack isn't marked as kept on-disk. But the\n--stdin-packs test treats it as such, by passing '^pack-$base_pack.pack'\nas input to '--stdin-packs' (thus marking it as kept in-core).\n\n> For showing the impact of the optimizations in patches 7 and 8, I think\n> doing a full repack with --honor-pack-keep is a better test. Because\n> then we're always doing a full traversal, and most of the work continues\n> to scale with the repo size (though obviously not the actual shuffling\n> of packed bytes around). That would get rid of the weird \"no work to do\"\n> case in the single-pack tests, too.\n\nI think you're suggesting that we change the \"repack ($nr_packs)\" test\nto have the residual pack marked as kept (so we're measuring time it\ntakes to repack everything that _isn't_ in the base pack)?\n\nThat would allow a more direct comparison, but I think it's loosing out\non an important aspect which is how long it takes to pack the entire\nrepository. Maybe we want three.\n\nWhat do you think?\n\n> -Peff\n\nThanks,\nTaylor\n"},{"id":"417176","messageId":"YC1sthj/3xFyoInq@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YC1nfH356wfmAKE2@nand.local","subject":"Re: [PATCH v2 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T19:21:26Z","receivedAt":"2021-02-17T19:22:16Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 17, 2021 at 01:59:08PM -0500, Taylor Blau wrote:\n\n> > > +\t\tif (rev_list_unpacked)\n> > > +\t\t\tadd_unreachable_loose_objects();\n> >\n> > Despite the name, that function is adding both reachable and unreachable\n> > ones. So it is doing what you want. It might be worth renaming, but it's\n> > not too big a deal since it's local to this file.\n> \n> Yeah, I tend to err on the side of \"it's fine as-is\" since this isn't\n> exposed outside of pack-objects internals. If you feel strongly I'm\n> happy to change it, but I suspect you don't.\n\nYeah, I don't feel strongly (and if we did change it, it should be in a\nseparate patch anyway).\n\n-Peff\n"},{"id":"417177","messageId":"YC1tLzivHqnoV6U7@nand.local","threadId":"55012","inReplyTo":"YC0+wlRksoqm0xLO@coredump.intra.peff.net","subject":"Re: [PATCH v2 6/8] builtin/pack-objects.c: rewrite honor-pack-keep logic","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T19:23:27Z","receivedAt":"2021-02-17T19:24:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 17, 2021 at 11:05:22AM -0500, Jeff King wrote:\n> I think just bumping that:\n>\n>   if (local && !p->pack_local)\n> \treturn 0;\n\n> above the new code would fix it. Or to lay out the logic more fully, the\n> order of checks should be:\n\n>   - does _this_ pack we found the object in disqualify it. If so, we can\n>     cheaply return 0. And that applies to both keep and local rules.\n>\n>   - otherwise, check all packs via has_object_kept_pack(), which is\n>     cheaper than continuing to iterate through all packs by returning\n>     -1.\n>\n>   - once we know definitively about keep-packs, then check any shortcuts\n>     related to local packs (like !have_non_local_packs)\n>\n>   - and then if no shortcuts, we return -1\n\nI don't understand what you're suggesting. Is the (local &&\n!p->pack_local) a disqualifying condition? Reading the comment, I think\nit is, and so we could do something like:\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 36c2fa3aff..be3ba60bc2 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1205,14 +1205,21 @@ static int want_found_object(const struct object_id *oid, int exclude,\n         * make sure no copy of this object appears in _any_ pack that makes us\n         * to omit the object, so we need to check all the packs.\n         *\n-        * We can however first check whether these options can possible matter;\n+        * We can however first check whether these options can possibly matter;\n         * if they do not matter we know we want the object in generated pack.\n         * Otherwise, we signal \"-1\" at the end to tell the caller that we do\n         * not know either way, and it needs to check more packs.\n         */\n\n        /*\n-        * Handle .keep first, as we have a fast(er) path there.\n+        * Objects in packs borrowed from elsewhere are discarded regardless of\n+        * if they appear in other packs that weren't borrowed.\n+        */\n+       if (local && !p->pack_local)\n+               return 0;\n+\n+       /*\n+        * Then handle .keep first, as we have a fast(er) path there.\n         */\n        if (ignore_packed_keep_on_disk || ignore_packed_keep_in_core) {\n                /*\n@@ -1242,11 +1249,8 @@ static int want_found_object(const struct object_id *oid, int exclude,\n         * keep-packs, or the object is not in one. Keep checking other\n         * conditions...\n         */\n-\n        if (!local || !have_non_local_packs)\n                return 1;\n-       if (local && !p->pack_local)\n-               return 0;\n\n        /* we don't know yet; keep looking for more packs */\n        return -1;\n\nBut your \"check any shortcuts related to local packs\" makes me think\nthat we should leave the code as-is.\n\nWhich are you suggesting?\n\nThanks,\nTaylor\n"},{"id":"417178","messageId":"YC1ttSf1xZLbCSH8@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YC1q64fQxHBMxmyw@nand.local","subject":"Re: [PATCH v2 5/8] p5303: measure time to repack with keep","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T19:25:41Z","receivedAt":"2021-02-17T19:26:35Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 17, 2021 at 02:13:47PM -0500, Taylor Blau wrote:\n\n> > > That's because the code paths around handling .keep files are known to\n> > > scale badly; they look in every single pack file to find each object.\n> >\n> > Well, part of it is just that with 1000 packs we have 20 times as many\n> > objects that are actually getting packed with --stdin-packs, compared to\n> > the 50-pack case. IIRC, each pack is a fixed-size slice and then the\n> > residual is put into the .keep pack. So the fact that the time gets\n> > closer to a full repack as we add more packs is expected: we are asking\n> > pack-objects to do more work!\n> \n> No, the residual base pack isn't marked as kept on-disk. But the\n> --stdin-packs test treats it as such, by passing '^pack-$base_pack.pack'\n> as input to '--stdin-packs' (thus marking it as kept in-core).\n\nSorry, I perhaps shouldn't have said \".keep\" here. But it's the same\nthing, isn't it? The 50 pack case is packing 50*pack_size objects\n(because it's excluding everything else that is in the base pack we mark\nas keep-in-core), and the 1000-pack case is packing 1000*pack_size\nobjects (for the same reason).\n\nSo any patterns we see between them have more to do with that, than how\nthe keep-handling code scales with the number of non-kept packs.\n\n> > For showing the impact of the optimizations in patches 7 and 8, I think\n> > doing a full repack with --honor-pack-keep is a better test. Because\n> > then we're always doing a full traversal, and most of the work continues\n> > to scale with the repo size (though obviously not the actual shuffling\n> > of packed bytes around). That would get rid of the weird \"no work to do\"\n> > case in the single-pack tests, too.\n> \n> I think you're suggesting that we change the \"repack ($nr_packs)\" test\n> to have the residual pack marked as kept (so we're measuring time it\n> takes to repack everything that _isn't_ in the base pack)?\n> \n> That would allow a more direct comparison, but I think it's loosing out\n> on an important aspect which is how long it takes to pack the entire\n> repository. Maybe we want three.\n\nThat was what I was suggesting, but I think it's equivalent to what your\n--stdin-packs is testing. I guess the most interesting thing would\nactually be an _additional_ pack mark as .keep (and that pack does not\neven have to contain anything interesting -- the point is how much\neffort it costs to find that out. Of course the bigger it is the more\npronounced the effect of avoiding lookups in it).\n\n-Peff\n"},{"id":"417180","messageId":"YC1ugXFk32xHY4k0@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YC1tLzivHqnoV6U7@nand.local","subject":"Re: [PATCH v2 6/8] builtin/pack-objects.c: rewrite honor-pack-keep logic","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T19:29:05Z","receivedAt":"2021-02-17T19:30:06Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 17, 2021 at 02:23:27PM -0500, Taylor Blau wrote:\n\n> On Wed, Feb 17, 2021 at 11:05:22AM -0500, Jeff King wrote:\n> > I think just bumping that:\n> >\n> >   if (local && !p->pack_local)\n> > \treturn 0;\n> \n> > above the new code would fix it. Or to lay out the logic more fully, the\n> > order of checks should be:\n> \n> >   - does _this_ pack we found the object in disqualify it. If so, we can\n> >     cheaply return 0. And that applies to both keep and local rules.\n> >\n> >   - otherwise, check all packs via has_object_kept_pack(), which is\n> >     cheaper than continuing to iterate through all packs by returning\n> >     -1.\n> >\n> >   - once we know definitively about keep-packs, then check any shortcuts\n> >     related to local packs (like !have_non_local_packs)\n> >\n> >   - and then if no shortcuts, we return -1\n> \n> I don't understand what you're suggesting. Is the (local &&\n> !p->pack_local) a disqualifying condition? Reading the comment, I think\n> it is, and so we could do something like:\n\nThat's exactly what I'm suggesting. If we have a non-local pack and were\ngiven --local, then we can shortcut immediately without caring about\nkept packs: we know that we do not want the object.\n\n> [...]\n> But your \"check any shortcuts related to local packs\" makes me think\n> that we should leave the code as-is.\n\nNo, the \"shortcuts\" there is the opposite:\n\n  if (!local || !have_non_local_packs)\n\treturn 1;\n\nIf either of those is true, we can say \"definitely include\" but only\nwith respect to the --local requirement. So we _can't_ bump that up, but\nmust check it only after we've definitively resolved the keep-pack\nrequirement.\n\n-Peff\n"},{"id":"417199","messageId":"YC10eZkpqtzLlJUP@nand.local","threadId":"55012","inReplyTo":"YC1OJDFXPnxGMHPK@coredump.intra.peff.net","subject":"Re: [PATCH v2 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T19:54:33Z","receivedAt":"2021-02-17T19:55:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 17, 2021 at 12:11:00PM -0500, Jeff King wrote:\n> On Wed, Feb 03, 2021 at 10:59:21PM -0500, Taylor Blau wrote:\n>\n> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> > index fbd7b54d70..b2ba5aa14f 100644\n> > --- a/builtin/pack-objects.c\n> > +++ b/builtin/pack-objects.c\n> > @@ -1225,9 +1225,9 @@ static int want_found_object(const struct object_id *oid, int exclude,\n> >  \t\t */\n> >  \t\tunsigned flags = 0;\n> >  \t\tif (ignore_packed_keep_on_disk)\n> > -\t\t\tflags |= ON_DISK_KEEP_PACKS;\n> > +\t\t\tflags |= CACHE_ON_DISK_KEEP_PACKS;\n> >  \t\tif (ignore_packed_keep_in_core)\n> > -\t\t\tflags |= IN_CORE_KEEP_PACKS;\n> > +\t\t\tflags |= CACHE_IN_CORE_KEEP_PACKS;\n>\n> Why are we renaming the constants in this patch?\n>\n> I know I'm listed as the author, but I think this came out of some\n> off-list back and forth between us. It seems like the existing constants\n> would have been fine.\n\nYeah, they would have been fine. They were renamed because this patch\nmakes them only used for the kept pack cache, but I agree the existing\nnames are fine, too.\n\nIn any case, they make an easier-to-read diff, so I'm perfectly happy to\nun-rename them ;).\n\n> > +static void maybe_invalidate_kept_pack_cache(struct repository *r,\n> > +\t\t\t\t\t     unsigned flags)\n> >  {\n> > -\treturn find_one_pack_entry(r, oid, e, 0);\n> > +\tif (!r->objects->kept_pack_cache)\n> > +\t\treturn;\n> > +\tif (r->objects->kept_pack_cache->flags == flags)\n> > +\t\treturn;\n> > +\tfree(r->objects->kept_pack_cache->packs);\n> > +\tFREE_AND_NULL(r->objects->kept_pack_cache);\n> > +}\n>\n> OK, so we keep a single cache based on the flags, and then if somebody\n> ever asks for different flags, we throw it away. That's probably OK for\n> our purposes, since we wouldn't expect multiple callers within a single\n> process.\n>\n> I wondered if it would be simpler to just keep two lists, one for\n> in-core keeps and one for on-disk keeps. And then just walk over each\n> list separately based on the query flags. That makes things more robust\n> _and_ I think would be less code. It does mean that a pack could appear\n> in both lists, though, which means we might do a lookup in it twice.\n> That doesn't seem all that likely, but it is working against our goal\n> here.\n>\n> Another option is to keep 3 caches (two separate and one combined),\n> rather than flipping between them. I'm not sure if that would be less\n> code or not (it gets rid of the \"invalidate\" function, but you do have\n> to pick the right cache depending on the query flags).\n>\n> Yet another option is to keep a cache of any that are marked as _either_\n> in core or on-disk keeps, and then decide to look up the object based on\n> the query flags. Then you just pay the cost to iterate over the list and\n> check the flags (which really is all this cache is helping with in the\n> first place).\n\nAll interesting ideas. In this patch (and by the end of the series)\ncallers that use the kept pack cache never ask for the cache with a\ndifferent set of flags. IOW, there isn't a situation where a caller\nwould populate the in-core kept pack cache, and then suddenly ask for\nboth in-core and on-disk packs to be kept.\n\nSo all of this code is defensive in case that were to change, and\nsuddenly we'd be returning subtly wrong results. I could imagine that\nbeing kind of a nasty bug to track down, so detecting and invalidating\nthe cache would make it a non-issue.\n\nI'll note it in the commit message, though, since it's good for future\nreaders to be aware, too.\n\n> I dunno. TBH, I kind of wonder if this whole patch is worth doing at\n> all, giving the underwhelming performance benefit (3% on the\n> pathological 1000-pack case). When I had timed this strategy initially,\n> it was more like 15%. I'm not sure where the savings went in the\n> interim, or if it was a timing fluke.\n\nYeah, I dunno. It's certainly not hurting (I don't think the extra code\nis all that complex, and the savings is at least non-zero), so I'm\ninclined to keep it.\n\n> > +static struct packed_git **kept_pack_cache(struct repository *r, unsigned flags)\n> > +{\n> > +\tmaybe_invalidate_kept_pack_cache(r, flags);\n> > +\n> > +\tif (!r->objects->kept_pack_cache) {\n> > +\t\tstruct packed_git **packs = NULL;\n> > +\t\tsize_t nr = 0, alloc = 0;\n> > +\t\tstruct packed_git *p;\n> > +\n> > +\t\t/*\n> > +\t\t * We want \"all\" packs here, because we need to cover ones that\n> > +\t\t * are used by a midx, as well. We need to look in every one of\n> > +\t\t * them (instead of the midx itself) to cover duplicates. It's\n> > +\t\t * possible that an object is found in two packs that the midx\n> > +\t\t * covers, one kept and one not kept, but the midx returns only\n> > +\t\t * the non-kept version.\n> > +\t\t */\n> > +\t\tfor (p = get_all_packs(r); p; p = p->next) {\n> > +\t\t\tif ((p->pack_keep && (flags & CACHE_ON_DISK_KEEP_PACKS)) ||\n> > +\t\t\t    (p->pack_keep_in_core && (flags & CACHE_IN_CORE_KEEP_PACKS))) {\n> > +\t\t\t\tALLOC_GROW(packs, nr + 1, alloc);\n> > +\t\t\t\tpacks[nr++] = p;\n> > +\t\t\t}\n> > +\t\t}\n> > +\t\tALLOC_GROW(packs, nr + 1, alloc);\n> > +\t\tpacks[nr] = NULL;\n> > +\n> > +\t\tr->objects->kept_pack_cache = xmalloc(sizeof(*r->objects->kept_pack_cache));\n> > +\t\tr->objects->kept_pack_cache->packs = packs;\n> > +\t\tr->objects->kept_pack_cache->flags = flags;\n> > +\t}\n>\n> Is there any reason not to just embed the kept_pack_cache struct inside\n> the object_store? It's one less pointer to deal with. I wonder if this\n> is a holdover from an attempt to have multiple caches.\n>\n> (I also think it would be reasonable if we wanted to hide the definition\n> of the cache struct from callers, but we don't seem do to that).\n\nNot a holdover, just designed to avoid adding too many extra fields to\nthe object-store. I don't feel strongly, but I do think hiding the\ndefinition is a good idea, so I'll inline it.\n\n> > @@ -2109,7 +2123,8 @@ int has_object_pack(const struct object_id *oid)\n> >  \treturn find_pack_entry(the_repository, oid, &e);\n> >  }\n> >\n> > -int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n> > +int has_object_kept_pack(const struct object_id *oid,\n> > +\t\t\t unsigned flags)\n> >  {\n> >  \tstruct pack_entry e;\n> >  \treturn find_kept_pack_entry(the_repository, oid, flags, &e);\n>\n> This seems like a stray change.\n\nGood eyes, thanks.\n\n>\n> > diff --git a/packfile.h b/packfile.h\n> > index 624327f64d..eb56db2a7b 100644\n> > --- a/packfile.h\n> > +++ b/packfile.h\n> > @@ -161,10 +161,6 @@ int packed_object_info(struct repository *r,\n> >  void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n> >  const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n> >\n> > -#define ON_DISK_KEEP_PACKS 1\n> > -#define IN_CORE_KEEP_PACKS 2\n> > -#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n>\n> I notice that when the constants moved, we didn't keep an equivalent of\n> ALL_KEEP_PACKS. Maybe we didn't need it in the first place in patch 1?\n\nYeah, we didn't need it to begin with. I'll drop it accordingly.\n\n>   BTW, I absolutely hate the complication that all of this on-disk\n>   versus in-core keep distinction brings to this code. And I wondered\n>   what it was really doing for us and whether we could get rid of it.\n>   But I think we do need it: a common case may be to avoid using\n>   --honor-pack-keep (because you don't want to deal with racy .keep\n>   writes from incoming receive-pack processes), but use in-core ones for\n>   something like --stdin-packs. So we do need to respect one and not the\n>   other.\n>\n>   I do wonder if things would be simpler if pack-objects simply kept its\n>   own list of \"in core\" packs in a separate array. But that is really\n>   just another form of the same problem, I guess.\n\nYeah, the complexity is awfully hard to reason about, but you're right\nthat here it is necessary.\n\n> -Peff\n\nThanks,\nTaylor\n"},{"id":"417206","messageId":"YC12EnHZCsCPwiay@nand.local","threadId":"55012","inReplyTo":"YC1drGrIEg0C7Zo5@coredump.intra.peff.net","subject":"Re: [PATCH v2 8/8] builtin/repack.c: add '--geometric' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T20:01:22Z","receivedAt":"2021-02-17T20:02:09Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 17, 2021 at 01:17:16PM -0500, Jeff King wrote:\n> On Wed, Feb 03, 2021 at 10:59:25PM -0500, Taylor Blau wrote:\n>\n> > Often it is useful to both:\n> >\n> >   - have relatively few packfiles in a repository, and\n> >\n> >   - avoid having so few packfiles in a repository that we repack its\n> >     entire contents regularly\n> >\n> > This patch implements a '--geometric=<n>' option in 'git repack'. This\n> > allows the caller to specify that they would like each pack to be at\n> > least a factor times as large as the previous largest pack (by object\n> > count).\n> >\n> > Concretely, say that a repository has 'n' packfiles, labeled P1, P2,\n> > ..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\n> > With a geometric factor of 'r', it should be that:\n> >\n> >   objects(Pi) > r*objects(P(i-1))\n> >\n> > for all i in [1, n], where the packs are sorted by\n> >\n> >   objects(P1) <= objects(P2) <= ... <= objects(Pn).\n>\n> Just devil's advocating for a moment.\n>\n> [large push becoming the biggest pack in a repository]\n>\n>   - it may have been more carefully packed (e.g., with a larger window\n>     size, using \"-f\", etc) than the packs we got from pushes. We do\n>     _mostly_ retain the deltas when we roll up the packs, so it probably\n>     only has a small impact in practice (I'd expect in a few cases we'd\n>     throw away deltas because a pushed pack contains a duplicate of its\n>     base object that we added via --fix-thin).\n\nYeah, agreed.\n\n> So I suspect it's probably OK in practice. These cases would happen\n> rarely, and the impact would not be all that big. The bitmap thing I'd\n> worry the most about. As part of a larger strategy involving a midx it\n> is taken care of, but people using just this new feature may not realize\n> that. The bitmaps of course are \"just\" an optimization, but it's hard to\n> say how dire things are when they don't exist. For many situations,\n> probably not very dire. But I know that on our servers, when repos lack\n> bitmaps, people notice the performance degradation.\n>\n> On the other hand, by definition this happens in a case where there are\n> more objects that have just been pushed (and are therefore not\n> bitmapped) than existed already. So you _already_ have a performance\n> problem either way until you get bitmap coverage of those new objects.\n\nI almost split my reply between this and the above paragraph to say\nexactly this. I think in this case you'd want to rewrite your bitmap\nfrom scratch either way (whether you were using multi-pack or\ntraditional reachability bitmaps).\n\n> > --- a/Documentation/git-repack.txt\n> > +++ b/Documentation/git-repack.txt\n> > @@ -165,6 +165,17 @@ depth is 4095.\n> >  \tPass the `--delta-islands` option to `git-pack-objects`, see\n> >  \tlinkgit:git-pack-objects[1].\n> >\n> > +-g=<factor>::\n> > +--geometric=<factor>::\n> > +\tArrange resulting pack structure so that each successive pack\n> > +\tcontains at least `<factor>` times the number of objects as the\n> > +\tnext-largest pack.\n> > ++\n> > +`git repack` ensures this by determining a \"cut\" of packfiles that need to be\n> > +repacked into one in order to ensure a geometric progression. It picks the\n> > +smallest set of packfiles such that as many of the larger packfiles (by count of\n> > +objects contained in that pack) may be left intact.\n>\n> I think we might need to make clear in the documentation how this\n> differs from other repacks, in that it is not considering reachability\n> at all. I like the term \"roll up\" to describe what is happening, but we\n> probably need to define that term clearly, as well.\n\nAll fair suggestions, thanks.\n\nThanks,\nTaylor\n"},{"id":"417210","messageId":"YC17rflmxAAdBBCd@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YC10eZkpqtzLlJUP@nand.local","subject":"Re: [PATCH v2 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T20:25:17Z","receivedAt":"2021-02-17T20:26:29Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 17, 2021 at 02:54:33PM -0500, Taylor Blau wrote:\n\n> > OK, so we keep a single cache based on the flags, and then if somebody\n> > ever asks for different flags, we throw it away. That's probably OK for\n> > our purposes, since we wouldn't expect multiple callers within a single\n> > process.\n> > [...some alternatives]\n> \n> All interesting ideas. In this patch (and by the end of the series)\n> callers that use the kept pack cache never ask for the cache with a\n> different set of flags. IOW, there isn't a situation where a caller\n> would populate the in-core kept pack cache, and then suddenly ask for\n> both in-core and on-disk packs to be kept.\n> \n> So all of this code is defensive in case that were to change, and\n> suddenly we'd be returning subtly wrong results. I could imagine that\n> being kind of a nasty bug to track down, so detecting and invalidating\n> the cache would make it a non-issue.\n\nYeah, I agree that the current crop of callers does not care. And I am\nglad we are not leaving a booby-trap for later programmers with respect\nto correctness (by virtue of the invalidation function). But it does\nfeel like we are leaving one for performance, which they very well might\nnot realize the cache is doing worse-than-nothing.\n\nWould just doing:\n\n  if (cache.packs && cache.flags != flags)\n\tBUG(\"kept-pack-cache cannot handle multiple queries in a single process\");\n\nbe a better solution? That is not helping anyone towards a world where\nwe gracefully handle back-and-forth queries. But it makes it abundantly\nclear when such a thing would become necessary.\n\n> > Is there any reason not to just embed the kept_pack_cache struct inside\n> > the object_store? It's one less pointer to deal with. I wonder if this\n> > is a holdover from an attempt to have multiple caches.\n> >\n> > (I also think it would be reasonable if we wanted to hide the definition\n> > of the cache struct from callers, but we don't seem do to that).\n> \n> Not a holdover, just designed to avoid adding too many extra fields to\n> the object-store. I don't feel strongly, but I do think hiding the\n> definition is a good idea, so I'll inline it.\n\nThis response confuses me a bit. Hiding the definition from callers\nwould mean _keeping_ it as a pointer, but putting the definition into\npackfile.c, where nobody outside that file could see it (at least that\nis what I meant by hiding).\n\nBut inlining it to me implies embedding the struct (not a pointer to it)\nin \"struct object_store\", defining the struct at the point we define the\nstruct field which uses it.\n\nI am fine with either, to be clear. I'm just confused which you are\nproposing to do. :)\n\n-Peff\n"},{"id":"417211","messageId":"YC18vmTYKo/lEaB7@nand.local","threadId":"55012","inReplyTo":"YC17rflmxAAdBBCd@coredump.intra.peff.net","subject":"Re: [PATCH v2 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T20:29:50Z","receivedAt":"2021-02-17T20:33:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 17, 2021 at 03:25:17PM -0500, Jeff King wrote:\n> Would just doing:\n>\n>   if (cache.packs && cache.flags != flags)\n> \tBUG(\"kept-pack-cache cannot handle multiple queries in a single process\");\n>\n> be a better solution? That is not helping anyone towards a world where\n> we gracefully handle back-and-forth queries. But it makes it abundantly\n> clear when such a thing would become necessary.\n\nI dunno. I can certainly see its merits, but I have to imagine that\nanybody who cares enough about the performance will be able to find our\nconversation here. Assuming that's the case, I would rather have the\nkept-pack cache handle multiple queries before BUG()-ing.\n\n> > > Is there any reason not to just embed the kept_pack_cache struct inside\n> > > the object_store? It's one less pointer to deal with. I wonder if this\n> > > is a holdover from an attempt to have multiple caches.\n> > >\n> > > (I also think it would be reasonable if we wanted to hide the definition\n> > > of the cache struct from callers, but we don't seem do to that).\n> >\n> > Not a holdover, just designed to avoid adding too many extra fields to\n> > the object-store. I don't feel strongly, but I do think hiding the\n> > definition is a good idea, so I'll inline it.\n>\n> This response confuses me a bit. Hiding the definition from callers\n> would mean _keeping_ it as a pointer, but putting the definition into\n> packfile.c, where nobody outside that file could see it (at least that\n> is what I meant by hiding).\n>\n> But inlining it to me implies embedding the struct (not a pointer to it)\n> in \"struct object_store\", defining the struct at the point we define the\n> struct field which uses it.\n>\n> I am fine with either, to be clear. I'm just confused which you are\n> proposing to do. :)\n\nProbably because I changed my mind in the middle of writing it ;). I'm\nproposing embedding the definition of the struct into the definition of\nobject_store, and then operating on its fields (from within packfile.c).\n\nThanks,\nTaylor\n"},{"id":"417224","messageId":"YC2N///TCMK65XNr@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YC18vmTYKo/lEaB7@nand.local","subject":"Re: [PATCH v2 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T21:43:27Z","receivedAt":"2021-02-17T21:44:30Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 17, 2021 at 03:29:50PM -0500, Taylor Blau wrote:\n\n> On Wed, Feb 17, 2021 at 03:25:17PM -0500, Jeff King wrote:\n> > Would just doing:\n> >\n> >   if (cache.packs && cache.flags != flags)\n> > \tBUG(\"kept-pack-cache cannot handle multiple queries in a single process\");\n> >\n> > be a better solution? That is not helping anyone towards a world where\n> > we gracefully handle back-and-forth queries. But it makes it abundantly\n> > clear when such a thing would become necessary.\n> \n> I dunno. I can certainly see its merits, but I have to imagine that\n> anybody who cares enough about the performance will be able to find our\n> conversation here. Assuming that's the case, I would rather have the\n> kept-pack cache handle multiple queries before BUG()-ing.\n\nOK. I am on the fence, and you are the author, so I'm happy to go with\nyour preference.\n\nI'm not quite as optimistic that somebody would find this conversation,\nif only because they have to know to look for it. I could easily see\nsomebody adding a find_kept_in_pack() without thinking too hard about\nit. OTOH, I find it quite unlikely that anybody would use a different\nset of flags within the same process, so it would probably Just Work for\nthem regardless. :)\n\n> > This response confuses me a bit. Hiding the definition from callers\n> > would mean _keeping_ it as a pointer, but putting the definition into\n> > packfile.c, where nobody outside that file could see it (at least that\n> > is what I meant by hiding).\n> >\n> > But inlining it to me implies embedding the struct (not a pointer to it)\n> > in \"struct object_store\", defining the struct at the point we define the\n> > struct field which uses it.\n> >\n> > I am fine with either, to be clear. I'm just confused which you are\n> > proposing to do. :)\n> \n> Probably because I changed my mind in the middle of writing it ;). I'm\n> proposing embedding the definition of the struct into the definition of\n> object_store, and then operating on its fields (from within packfile.c).\n\nOK, that sounds great to me (and arguably produces more efficient code,\nsince we avoid a pointer dereference, though I doubt it matters in\npractice). Thanks for clarifying.\n\n-Peff\n"},{"id":"417272","messageId":"aa94edf39b957f3547c625fe11938082cdc30421.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 1/8] packfile: introduce 'find_kept_pack_entry()'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:16Z","receivedAt":"2021-02-18T03:15:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Future callers will want a function to fill a 'struct pack_entry' for a\ngiven object id but _only_ from its position in any kept pack(s).\n\nIn particular, an new 'git repack' mode which ensures the resulting\npacks form a geometric progress by object count will mark packs that it\ndoes not want to repack as \"kept in-core\", and it will want to halt a\nreachability traversal as soon as it visits an object in any of the kept\npacks. But, it does not want to halt the traversal at non-kept, or\n.keep packs.\n\nThe obvious alternative is 'find_pack_entry()', but this doesn't quite\nsuffice since it only returns the first pack it finds, which may or may\nnot be kept (and the mru cache makes it unpredictable which one you'll\nget if there are options).\n\nShort of that, you could walk over all packs looking for the object in\neach one, but it scales with the number of packs, which may be\nprohibitive.\n\nIntroduce 'find_kept_pack_entry()', a function which is like\n'find_pack_entry()', but only fills in objects in the kept packs.\n\nHandle packs which have .keep files, as well as in-core kept packs\nseparately, since certain callers will want to distinguish one from the\nother. (Though on-disk and in-core kept packs share the adjective\n\"kept\", it is best to think of the two sets as independent.)\n\nThere is a gotcha when looking up objects that are duplicated in kept\nand non-kept packs, particularly when the MIDX stores the non-kept\nversion and the caller asked for kept objects only. This could be\nresolved by teaching the MIDX to resolve duplicates by always favoring\nthe kept pack (if one exists), but this breaks an assumption in existing\nMIDXs, and so it would require a format change.\n\nThe benefit to changing the MIDX in this way is marginal, so we instead\nhave a more thorough check here which is explained with a comment.\n\nCallers will be added in subsequent patches.\n\nCo-authored-by: Jeff King <peff@peff.net>\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n packfile.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++-----\n packfile.h |  5 +++++\n 2 files changed, 64 insertions(+), 5 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex 1fec12ac5f..7f84f221ce 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2042,7 +2042,10 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static int find_one_pack_entry(struct repository *r,\n+\t\t\t       const struct object_id *oid,\n+\t\t\t       struct pack_entry *e,\n+\t\t\t       int kept_only)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2052,26 +2055,77 @@ int find_pack_entry(struct repository *r, const struct object_id *oid, struct pa\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (fill_midx_entry(r, oid, e, m))\n+\t\tif (!fill_midx_entry(r, oid, e, m))\n+\t\t\tcontinue;\n+\n+\t\tif (!kept_only)\n+\t\t\treturn 1;\n+\n+\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n+\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n-\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\treturn 1;\n+\t\tif (p->multi_pack_index && !kept_only) {\n+\t\t\t/*\n+\t\t\t * If this pack is covered by the MIDX, we'd have found\n+\t\t\t * the object already in the loop above if it was here,\n+\t\t\t * so don't bother looking.\n+\t\t\t *\n+\t\t\t * The exception is if we are looking only at kept\n+\t\t\t * packs. An object can be present in two packs covered\n+\t\t\t * by the MIDX, one kept and one not-kept. And as the\n+\t\t\t * MIDX points to only one copy of each object, it might\n+\t\t\t * have returned only the non-kept version above. We\n+\t\t\t * have to check again to be thorough.\n+\t\t\t */\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!kept_only ||\n+\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n+\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n+\t\t\tif (fill_pack_entry(oid, e, p)) {\n+\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\t\treturn 1;\n+\t\t\t}\n \t\t}\n \t}\n \treturn 0;\n }\n \n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+{\n+\treturn find_one_pack_entry(r, oid, e, 0);\n+}\n+\n+int find_kept_pack_entry(struct repository *r,\n+\t\t\t const struct object_id *oid,\n+\t\t\t unsigned flags,\n+\t\t\t struct pack_entry *e)\n+{\n+\t/*\n+\t * Load all packs, including midx packs, since our \"kept\" strategy\n+\t * relies on that. We're relying on the side effect of it setting up\n+\t * r->objects->packed_git, which is a little ugly.\n+\t */\n+\tget_all_packs(r);\n+\treturn find_one_pack_entry(r, oid, e, flags);\n+}\n+\n int has_object_pack(const struct object_id *oid)\n {\n \tstruct pack_entry e;\n \treturn find_pack_entry(the_repository, oid, &e);\n }\n \n+int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n+{\n+\tstruct pack_entry e;\n+\treturn find_kept_pack_entry(the_repository, oid, flags, &e);\n+}\n+\n int has_pack_index(const unsigned char *sha1)\n {\n \tstruct stat st;\ndiff --git a/packfile.h b/packfile.h\nindex 4cfec9e8d3..3ae117a8ae 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -162,13 +162,18 @@ int packed_object_info(struct repository *r,\n void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n \n+#define ON_DISK_KEEP_PACKS 1\n+#define IN_CORE_KEEP_PACKS 2\n+\n /*\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\n  */\n int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e);\n+int find_kept_pack_entry(struct repository *r, const struct object_id *oid, unsigned flags, struct pack_entry *e);\n \n int has_object_pack(const struct object_id *oid);\n+int has_object_kept_pack(const struct object_id *oid, unsigned flags);\n \n int has_pack_index(const unsigned char *sha1);\n \n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417273","messageId":"cover.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH v3 0/8] repack: support repacking into a geometric sequence","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:11Z","receivedAt":"2021-02-18T03:15:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Here is another updated version of mine and Peff's series to add a new 'git\nrepack --geometric' mode which supports repacking a repository into a geometric\nprogression of packs by object count.\n\n(A previous version of this series depended on 'jk/p5303-sed-portability-fix',\nbut that topic has since been merged to 'master'. This series has been updated\nto apply based on 'master' accordingly).\n\nThe series has not changed substantially since v2, but a range-diff is included\nbelow for convenience. The most notable change is the new tests in p5303 were\nreworked to provide a more equivalent comparison.\n\nBeyond that, some minor code clean-up (embedding the kept-pack cache, making the\n'--no-kept-packs' option of 'rev-list' undocumented, etc) has been applied to\naddress Peff's review.\n\nThanks in advance for another look at this series. I'm hopeful that this version\nis in a good state to be queued so that it can make the 2.31 release, and users\ncan start playing with it.\n\nJeff King (4):\n  p5303: add missing &&-chains\n  p5303: measure time to repack with keep\n  builtin/pack-objects.c: rewrite honor-pack-keep logic\n  packfile: add kept-pack cache for find_kept_pack_entry()\n\nTaylor Blau (4):\n  packfile: introduce 'find_kept_pack_entry()'\n  revision: learn '--no-kept-objects'\n  builtin/pack-objects.c: add '--stdin-packs' option\n  builtin/repack.c: add '--geometric' option\n\n Documentation/git-pack-objects.txt |  10 +\n Documentation/git-repack.txt       |  22 ++\n builtin/pack-objects.c             | 329 ++++++++++++++++++++++++-----\n builtin/repack.c                   | 187 +++++++++++++++-\n object-store.h                     |   5 +\n packfile.c                         |  67 ++++++\n packfile.h                         |   5 +\n revision.c                         |  15 ++\n revision.h                         |   4 +\n t/perf/p5303-many-packs.sh         |  36 +++-\n t/t5300-pack-object.sh             |  97 +++++++++\n t/t6114-keep-packs.sh              |  69 ++++++\n t/t7703-repack-geometric.sh        | 137 ++++++++++++\n 13 files changed, 921 insertions(+), 62 deletions(-)\n create mode 100755 t/t6114-keep-packs.sh\n create mode 100755 t/t7703-repack-geometric.sh\n\nRange-diff against v2:\n[rebased onto 'master']\n13:  f7186147eb !  1:  aa94edf39b packfile: introduce 'find_kept_pack_entry()'\n    @@ packfile.h: int packed_object_info(struct repository *r,\n\n     +#define ON_DISK_KEEP_PACKS 1\n     +#define IN_CORE_KEEP_PACKS 2\n    -+#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n     +\n      /*\n       * Iff a pack file in the given repository contains the object named by sha1,\n14:  ddc2896caa !  2:  82f6b45463 revision: learn '--no-kept-objects'\n    @@ Commit message\n         certain packs alone (for e.g., when doing a geometric repack that has\n         some \"large\" packs which are kept in-core that it wants to leave alone).\n\n    +    Note that this option is not guaranteed to produce exactly the set of\n    +    objects that aren't in kept packs, since it's possible the traversal\n    +    order may end up in a situation where a non-kept ancestor was \"cut off\"\n    +    by a kept object (at which point we would stop traversing). But, we\n    +    don't care about absolute correctness here, since this will eventually\n    +    be used as a purely additive guide in an upcoming new repack mode.\n    +\n    +    Explicitly avoid documenting this new flag, since it is only used\n    +    internally. In theory we could avoid even adding it rev-list, but being\n    +    able to spell this option out on the command-line makes some special\n    +    cases easier to test without promising to keep it behaving consistently\n    +    forever. Those tricky cases are exercised in t6114.\n    +\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n\n    - ## Documentation/rev-list-options.txt ##\n    -@@ Documentation/rev-list-options.txt: ifdef::git-rev-list[]\n    - \tOnly useful with `--objects`; print the object IDs that are not\n    - \tin packs.\n    -\n    -+--no-kept-objects[=<kind>]::\n    -+\tHalts the traversal as soon as an object in a kept pack is\n    -+\tfound. If `<kind>` is `on-disk`, only packs with a corresponding\n    -+\t`*.keep` file are ignored. If `<kind>` is `in-core`, only packs\n    -+\twith their in-core kept state set are ignored. Otherwise, both\n    -+\tkinds of kept packs are ignored.\n    -+\n    - --object-names::\n    - \tOnly useful with `--objects`; print the names of the object IDs\n    - \tthat are found. This is the default behavior.\n    -\n    - ## list-objects.c ##\n    -@@ list-objects.c: static void traverse_trees_and_blobs(struct traversal_context *ctx,\n    - \t\t\tctx->show_object(obj, name, ctx->show_data);\n    - \t\t\tcontinue;\n    - \t\t}\n    -+\t\tif (ctx->revs->no_kept_objects) {\n    -+\t\t\tstruct pack_entry e;\n    -+\t\t\tif (find_kept_pack_entry(ctx->revs->repo, &obj->oid,\n    -+\t\t\t\t\t\t ctx->revs->keep_pack_cache_flags,\n    -+\t\t\t\t\t\t &e))\n    -+\t\t\t\tcontinue;\n    -+\t\t}\n    - \t\tif (!path)\n    - \t\t\tpath = \"\";\n    - \t\tif (obj->type == OBJ_TREE) {\n    -\n      ## revision.c ##\n     @@ revision.c: static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n      \t\trevs->unpacked = 1;\n15:  c96b1bf995 !  3:  033e4e3f67 builtin/pack-objects.c: add '--stdin-packs' option\n    @@ Documentation/git-pack-objects.txt: base-name::\n      \tcan be useful to send new tags to native Git clients.\n\n     +--stdin-packs::\n    -+\tRead the basenames of packfiles from the standard input, instead\n    -+\tof object names or revision arguments. The resulting pack\n    -+\tcontains all objects listed in the included packs (those not\n    -+\tbeginning with `^`), excluding any objects listed in the\n    -+\texcluded packs (beginning with `^`).\n    ++\tRead the basenames of packfiles (e.g., `pack-1234abcd.pack`)\n    ++\tfrom the standard input, instead of object names or revision\n    ++\targuments. The resulting pack contains all objects listed in the\n    ++\tincluded packs (those not beginning with `^`), excluding any\n    ++\tobjects listed in the excluded packs (beginning with `^`).\n     ++\n     +Incompatible with `--revs`, or options that imply `--revs` (such as\n     +`--all`), with the exception of `--unpacked`, which is compatible.\n    @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n      \treturn git_default_config(k, v, cb);\n      }\n\n    ++/* Counters for trace2 output when in --stdin-packs mode. */\n     +static int stdin_packs_found_nr;\n     +static int stdin_packs_hints_nr;\n     +\n    @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n     +\n     +\tdisplay_progress(progress_state, ++nr_seen);\n     +\n    ++\tif (have_duplicate_entry(oid, 0))\n    ++\t\treturn 0;\n    ++\n     +\tofs = nth_packed_object_offset(p, pos);\n    ++\tif (!want_object_in_pack(oid, 0, &p, &ofs))\n    ++\t\treturn 0;\n     +\n     +\toi.typep = &type;\n     +\tif (packed_object_info(the_repository, p, ofs, &oi) < 0)\n    @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n     +\t\tadd_pending_oid(revs, NULL, oid, 0);\n     +\t}\n     +\n    -+\tif (have_duplicate_entry(oid, 0))\n    -+\t\treturn 0;\n    -+\n    -+\tif (!want_object_in_pack(oid, 0, &p, &ofs))\n    -+\t\treturn 0;\n    -+\n     +\tstdin_packs_found_nr++;\n     +\n     +\tcreate_object_entry(oid, type, 0, 0, 0, p, ofs);\n    @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n     +\n     +static void show_commit_pack_hint(struct commit *commit, void *_data)\n     +{\n    ++\t/* nothing to do; commits don't have a namehash */\n     +}\n     +\n     +static void show_object_pack_hint(struct object *object, const char *name,\n    @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n     +\tstdin_packs_hints_nr++;\n     +}\n     +\n    ++static int pack_mtime_cmp(const void *_a, const void *_b)\n    ++{\n    ++\tstruct packed_git *a = ((const struct string_list_item*)_a)->util;\n    ++\tstruct packed_git *b = ((const struct string_list_item*)_b)->util;\n    ++\n    ++\tif (a->mtime < b->mtime)\n    ++\t\treturn -1;\n    ++\telse if (b->mtime < a->mtime)\n    ++\t\treturn 1;\n    ++\telse\n    ++\t\treturn 0;\n    ++}\n    ++\n     +static void read_packs_list_from_stdin(void)\n     +{\n     +\tstruct strbuf buf = STRBUF_INIT;\n    @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n     +\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n     +\t\tp->pack_keep_in_core = 1;\n     +\t}\n    ++\n    ++\t/*\n    ++\t * Order packs by ascending mtime; use QSORT directly to access the\n    ++\t * string_list_item's ->util pointer, which string_list_sort() does not\n    ++\t * provide.\n    ++\t */\n    ++\tQSORT(include_packs.items, include_packs.nr, pack_mtime_cmp);\n    ++\n     +\tfor_each_string_list_item(item, &include_packs) {\n     +\t\tstruct packed_git *p = item->util;\n     +\t\tif (!p)\n16:  a46b7002b4 =  4:  f9a5faf773 p5303: add missing &&-chains\n17:  b5081c01b5 !  5:  181c104a03 p5303: measure time to repack with keep\n    @@ Metadata\n      ## Commit message ##\n         p5303: measure time to repack with keep\n\n    -    This is the same as the regular repack test, except that we mark the\n    -    single base pack as \"kept\" and use --assume-kept-packs-closed. The\n    -    theory is that this should be faster than the normal repack, because\n    -    we'll have fewer objects to traverse and process.\n    +    Add two new tests to measure repack performance. Both test split the\n    +    repository into synthetic \"pushes\", and then leave the remaining objects\n    +    in a big base pack.\n\n    -    Here are some timings on a recent clone of the kernel. In the\n    -    single-pack case, there is nothing do since there are no non-excluded\n    -    packs:\n    +    The first new test marks an empty pack as \"kept\" and then passes\n    +    --honor-pack-keep to avoid including objects in it. That doesn't change\n    +    the resulting pack, but it does let us compare to the normal repack case\n    +    to see how much overhead we add to check whether objects are kept or\n    +    not.\n\n    -      5303.5: repack (1)                          57.42(54.88+10.64)\n    -      5303.6: repack with --stdin-packs (1)       0.01(0.01+0.00)\n    +    The other test is of --stdin-packs, which gives us a sense of how that\n    +    number scales based on the number of packs we provide as input. In each\n    +    of those tests, the empty pack isn't considered, but the residual pack\n    +    (objects that were left over and not included in one of the synthetic\n    +    push packs) is marked as kept.\n\n    -    and in the 50-pack case, it is much faster to use `--stdin-packs`, since\n    -    we avoid having to consider any objects in the excluded pack:\n    +    (Note that in the single-pack case of the --stdin-packs test, there is\n    +    nothing do since there are no non-excluded packs).\n\n    -      5303.10: repack (50)                        71.26(88.24+4.96)\n    -      5303.11: repack with --stdin-packs (50)     3.49(11.82+0.28)\n    +    Here are some timings on a recent clone of the kernel:\n\n    -    but our improvements vanish as we approach 1000 packs.\n    +      5303.5: repack (1)                          57.26(54.59+10.84)\n    +      5303.6: repack with kept (1)                57.33(54.80+10.51)\n\n    -      5303.15: repack (1000)                      215.64(491.33+14.80)\n    -      5303.16: repack with --stdin-packs (1000)   198.79(380.51+7.97)\n    +    in the 50-pack case, things start to slow down:\n    +\n    +      5303.11: repack (50)                        71.54(88.57+4.84)\n    +      5303.12: repack with kept (50)              85.12(102.05+4.94)\n    +\n    +    and by the time we hit 1,000 packs, things are substantially worse, even\n    +    though the resulting pack produced is the same:\n    +\n    +      5303.17: repack (1000)                      216.87(490.79+14.57)\n    +      5303.18: repack with kept (1000)            665.63(938.87+15.76)\n    +\n    +    Likewise, the scaling is pretty extreme on --stdin-packs:\n    +\n    +      5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)\n    +      5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)\n    +      5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)\n\n         That's because the code paths around handling .keep files are known to\n         scale badly; they look in every single pack file to find each object.\n    @@ t/perf/p5303-many-packs.sh: repack_into_n () {\n     +\t\tgit pack-objects --delta-base-offset --revs staging/pack\n     +\t) &&\n     +\ttest_export base_pack &&\n    ++\n    ++\t# create an empty packfile\n    ++\tempty_pack=$(git pack-objects staging/pack </dev/null) &&\n    ++\ttest_export empty_pack &&\n\n      \t# and then incrementals between each pair of commits\n      \tlast= &&\n    @@ t/perf/p5303-many-packs.sh: do\n      \t\t  --stdout </dev/null >/dev/null\n      \t'\n     +\n    ++\ttest_perf \"repack with kept ($nr_packs)\" '\n    ++\t\tgit pack-objects --keep-true-parents \\\n    ++\t\t  --keep-pack=pack-$empty_pack.pack \\\n    ++\t\t  --honor-pack-keep --non-empty --all \\\n    ++\t\t  --reflog --indexed-objects --delta-base-offset \\\n    ++\t\t  --stdout </dev/null >/dev/null\n    ++\t'\n    ++\n     +\ttest_perf \"repack with --stdin-packs ($nr_packs)\" '\n     +\t\tgit pack-objects \\\n     +\t\t  --keep-true-parents \\\n18:  c3868c7df9 !  6:  67af143fd1 builtin/pack-objects.c: rewrite honor-pack-keep logic\n    @@ Commit message\n         packs are actually kept.\n\n         Note that we have to re-order the logic a bit here; we can deal with the\n    -    \"kept\" situation completely, and then just fall back to the \"--local\"\n    -    question. It might be worth having a similar optimized function to look\n    -    at only local packs.\n    +    disqualifying situations first (e.g., finding the object in a non-local\n    +    pack with --local), then \"kept\" situation(s), and then just fall back to\n    +    other \"--local\" conditions.\n\n         Here are the results from p5303 (measurements again taken on the\n         kernel):\n\n    -      Test                                        HEAD^                    HEAD\n    +      Test                                        HEAD^                   HEAD\n           -----------------------------------------------------------------------------------------------\n    -      5303.5: repack (1)                          57.42(54.88+10.64)       57.44(54.71+10.78) +0.0%\n    -      5303.6: repack with --stdin-packs (1)       0.01(0.01+0.00)          0.01(0.00+0.01) +0.0%\n    -      5303.10: repack (50)                        71.26(88.24+4.96)        71.32(88.38+4.90) +0.1%\n    -      5303.11: repack with --stdin-packs (50)     3.49(11.82+0.28)         3.43(11.81+0.22) -1.7%\n    -      5303.15: repack (1000)                      215.64(491.33+14.80)     215.59(493.75+14.62) -0.0%\n    -      5303.16: repack with --stdin-packs (1000)   198.79(380.51+7.97)      131.44(314.24+8.11) -33.9%\n    -\n    -    So our --stdin-packs case with many packs is now finally faster than the\n    -    non-keep case (because it gets the speed benefit of looking at fewer\n    -    objects, but not as big a penalty for looking at many packs).\n    +      5303.5: repack (1)                          57.26(54.59+10.84)      57.34(54.66+10.88) +0.1%\n    +      5303.6: repack with kept (1)                57.33(54.80+10.51)      57.38(54.83+10.49) +0.1%\n    +      5303.11: repack (50)                        71.54(88.57+4.84)       71.70(88.99+4.74) +0.2%\n    +      5303.12: repack with kept (50)              85.12(102.05+4.94)      72.58(89.61+4.78) -14.7%\n    +      5303.17: repack (1000)                      216.87(490.79+14.57)    217.19(491.72+14.25) +0.1%\n    +      5303.18: repack with kept (1000)            665.63(938.87+15.76)    246.12(520.07+14.93) -63.0%\n    +\n    +    and the --stdin-packs timings:\n    +\n    +      5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)         0.00(0.00+0.00) -100.0%\n    +      5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)        3.43(11.75+0.24) -2.8%\n    +      5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)     130.50(307.15+7.66) -33.4%\n    +\n    +    So our repack with an empty .keep pack is roughly as fast as one without\n    +    a .keep pack up to 50 packs. But the --stdin-packs case scales a little\n    +    better, too.\n    +\n    +    Notably, it is faster than a repack of the same size and a kept pack. It\n    +    looks at fewer objects, of course, but the penalty for looking at many\n    +    packs isn't as costly.\n\n         Signed-off-by: Jeff King <peff@peff.net>\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n    @@ builtin/pack-objects.c: static int have_duplicate_entry(const struct object_id *\n      \tif (exclude)\n      \t\treturn 1;\n     @@ builtin/pack-objects.c: static int want_found_object(int exclude, struct packed_git *p)\n    + \t * make sure no copy of this object appears in _any_ pack that makes us\n    + \t * to omit the object, so we need to check all the packs.\n    + \t *\n    +-\t * We can however first check whether these options can possible matter;\n    ++\t * We can however first check whether these options can possibly matter;\n    + \t * if they do not matter we know we want the object in generated pack.\n      \t * Otherwise, we signal \"-1\" at the end to tell the caller that we do\n      \t * not know either way, and it needs to check more packs.\n      \t */\n     -\tif (!ignore_packed_keep_on_disk &&\n     -\t    !ignore_packed_keep_in_core &&\n     -\t    (!local || !have_non_local_packs))\n    +-\t\treturn 1;\n    +\n    ++\t/*\n    ++\t * Objects in packs borrowed from elsewhere are discarded regardless of\n    ++\t * if they appear in other packs that weren't borrowed.\n    ++\t */\n    + \tif (local && !p->pack_local)\n    + \t\treturn 0;\n    +-\tif (p->pack_local &&\n    +-\t    ((ignore_packed_keep_on_disk && p->pack_keep) ||\n    +-\t     (ignore_packed_keep_in_core && p->pack_keep_in_core)))\n    +-\t\treturn 0;\n     +\n     +\t/*\n    -+\t * Handle .keep first, as we have a fast(er) path there.\n    ++\t * Then handle .keep first, as we have a fast(er) path there.\n     +\t */\n     +\tif (ignore_packed_keep_on_disk || ignore_packed_keep_in_core) {\n     +\t\t/*\n    @@ builtin/pack-objects.c: static int want_found_object(int exclude, struct packed_\n     +\t * keep-packs, or the object is not in one. Keep checking other\n     +\t * conditions...\n     +\t */\n    -+\n     +\tif (!local || !have_non_local_packs)\n    - \t\treturn 1;\n    --\n    - \tif (local && !p->pack_local)\n    - \t\treturn 0;\n    --\tif (p->pack_local &&\n    --\t    ((ignore_packed_keep_on_disk && p->pack_keep) ||\n    --\t     (ignore_packed_keep_in_core && p->pack_keep_in_core)))\n    --\t\treturn 0;\n    ++\t\treturn 1;\n\n      \t/* we don't know yet; keep looking for more packs */\n      \treturn -1;\n19:  f1c07324f6 !  7:  e9e04b95e7 packfile: add kept-pack cache for find_kept_pack_entry()\n    @@ Commit message\n           - we don't have to worry about any packed_git being removed; we always\n             keep the old structs around, even after reprepare_packed_git()\n\n    +    We do defensively invalidate the cache in case the set of kept packs\n    +    being asked for changes (e.g., only in-core kept packs were cached, but\n    +    suddenly the caller also wants on-disk kept packs, too). In theory we\n    +    could build all three caches and switch between them, but it's not\n    +    necessary, since this patch (and series) never changes the set of kept\n    +    packs that it wants to inspect from the cache.\n    +\n    +    So that \"optimization\" is more about being defensive in the face of\n    +    future changes than it is about asking for multiple kinds of kept packs\n    +    in this patch.\n    +\n         Here are p5303 results (as always, measured against the kernel):\n\n           Test                                        HEAD^                   HEAD\n    -      ----------------------------------------------------------------------------------------------\n    -      5303.5: repack (1)                          57.44(54.71+10.78)      57.06(54.29+10.96) -0.7%\n    -      5303.6: repack with --stdin-packs (1)       0.01(0.00+0.01)         0.01(0.01+0.00) +0.0%\n    -      5303.10: repack (50)                        71.32(88.38+4.90)       71.47(88.60+5.04) +0.2%\n    -      5303.11: repack with --stdin-packs (50)     3.43(11.81+0.22)        3.49(12.21+0.26) +1.7%\n    -      5303.15: repack (1000)                      215.59(493.75+14.62)    217.41(495.36+14.85) +0.8%\n    -      5303.16: repack with --stdin-packs (1000)   131.44(314.24+8.11)     126.75(309.88+8.09) -3.6%\n    +      -----------------------------------------------------------------------------------------------\n    +      5303.5: repack (1)                          57.34(54.66+10.88)      56.98(54.36+10.98) -0.6%\n    +      5303.6: repack with kept (1)                57.38(54.83+10.49)      57.17(54.97+10.26) -0.4%\n    +      5303.11: repack (50)                        71.70(88.99+4.74)       71.62(88.48+5.08) -0.1%\n    +      5303.12: repack with kept (50)              72.58(89.61+4.78)       71.56(88.80+4.59) -1.4%\n    +      5303.17: repack (1000)                      217.19(491.72+14.25)    217.31(490.82+14.53) +0.1%\n    +      5303.18: repack with kept (1000)            246.12(520.07+14.93)    217.08(490.37+15.10) -11.8%\n    +\n    +    and the --stdin-packs case, which scales a little bit better (although\n    +    not by that much even at 1,000 packs):\n    +\n    +      5303.7: repack with --stdin-packs (1)       0.00(0.00+0.00)         0.00(0.00+0.00) =\n    +      5303.13: repack with --stdin-packs (50)     3.43(11.75+0.24)        3.43(11.69+0.30) +0.0%\n    +      5303.19: repack with --stdin-packs (1000)   130.50(307.15+7.66)     125.13(301.36+8.04) -4.1%\n\n         Signed-off-by: Jeff King <peff@peff.net>\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n\n    - ## builtin/pack-objects.c ##\n    -@@ builtin/pack-objects.c: static int want_found_object(const struct object_id *oid, int exclude,\n    - \t\t */\n    - \t\tunsigned flags = 0;\n    - \t\tif (ignore_packed_keep_on_disk)\n    --\t\t\tflags |= ON_DISK_KEEP_PACKS;\n    -+\t\t\tflags |= CACHE_ON_DISK_KEEP_PACKS;\n    - \t\tif (ignore_packed_keep_in_core)\n    --\t\t\tflags |= IN_CORE_KEEP_PACKS;\n    -+\t\t\tflags |= CACHE_IN_CORE_KEEP_PACKS;\n    -\n    - \t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n    - \t\t\treturn 0;\n    -@@ builtin/pack-objects.c: static void read_packs_list_from_stdin(void)\n    - \t * an optimization during delta selection.\n    - \t */\n    - \trevs.no_kept_objects = 1;\n    --\trevs.keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n    -+\trevs.keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n    - \trevs.blob_objects = 1;\n    - \trevs.tree_objects = 1;\n    - \trevs.tag_objects = 1;\n    -\n      ## object-store.h ##\n    -@@ object-store.h: static inline int pack_map_entry_cmp(const void *unused_cmp_data,\n    - \treturn strcmp(pg1->pack_name, key ? key : pg2->pack_name);\n    - }\n    -\n    -+#define CACHE_ON_DISK_KEEP_PACKS 1\n    -+#define CACHE_IN_CORE_KEEP_PACKS 2\n    -+\n    -+struct kept_pack_cache {\n    -+\tstruct packed_git **packs;\n    -+\tunsigned flags;\n    -+};\n    -+\n    - struct raw_object_store {\n    - \t/*\n    - \t * Set of all object directories; the main directory is first (and\n     @@ object-store.h: struct raw_object_store {\n      \t/* A most-recently-used ordered version of the packed_git list. */\n      \tstruct list_head packed_git_mru;\n\n    -+\tstruct kept_pack_cache *kept_pack_cache;\n    ++\tstruct {\n    ++\t\tstruct packed_git **packs;\n    ++\t\tunsigned flags;\n    ++\t} kept_pack_cache;\n     +\n      \t/*\n      \t * A map of packfiles to packed_git structs for tracking which\n    @@ packfile.c: static int find_one_pack_entry(struct repository *r,\n     +\t\t\t\t\t     unsigned flags)\n      {\n     -\treturn find_one_pack_entry(r, oid, e, 0);\n    -+\tif (!r->objects->kept_pack_cache)\n    ++\tif (!r->objects->kept_pack_cache.packs)\n     +\t\treturn;\n    -+\tif (r->objects->kept_pack_cache->flags == flags)\n    ++\tif (r->objects->kept_pack_cache.flags == flags)\n     +\t\treturn;\n    -+\tfree(r->objects->kept_pack_cache->packs);\n    -+\tFREE_AND_NULL(r->objects->kept_pack_cache);\n    ++\tFREE_AND_NULL(r->objects->kept_pack_cache.packs);\n    ++\tr->objects->kept_pack_cache.flags = 0;\n     +}\n     +\n     +static struct packed_git **kept_pack_cache(struct repository *r, unsigned flags)\n     +{\n     +\tmaybe_invalidate_kept_pack_cache(r, flags);\n     +\n    -+\tif (!r->objects->kept_pack_cache) {\n    ++\tif (!r->objects->kept_pack_cache.packs) {\n     +\t\tstruct packed_git **packs = NULL;\n     +\t\tsize_t nr = 0, alloc = 0;\n     +\t\tstruct packed_git *p;\n    @@ packfile.c: static int find_one_pack_entry(struct repository *r,\n     +\t\t * the non-kept version.\n     +\t\t */\n     +\t\tfor (p = get_all_packs(r); p; p = p->next) {\n    -+\t\t\tif ((p->pack_keep && (flags & CACHE_ON_DISK_KEEP_PACKS)) ||\n    -+\t\t\t    (p->pack_keep_in_core && (flags & CACHE_IN_CORE_KEEP_PACKS))) {\n    ++\t\t\tif ((p->pack_keep && (flags & ON_DISK_KEEP_PACKS)) ||\n    ++\t\t\t    (p->pack_keep_in_core && (flags & IN_CORE_KEEP_PACKS))) {\n     +\t\t\t\tALLOC_GROW(packs, nr + 1, alloc);\n     +\t\t\t\tpacks[nr++] = p;\n     +\t\t\t}\n    @@ packfile.c: static int find_one_pack_entry(struct repository *r,\n     +\t\tALLOC_GROW(packs, nr + 1, alloc);\n     +\t\tpacks[nr] = NULL;\n     +\n    -+\t\tr->objects->kept_pack_cache = xmalloc(sizeof(*r->objects->kept_pack_cache));\n    -+\t\tr->objects->kept_pack_cache->packs = packs;\n    -+\t\tr->objects->kept_pack_cache->flags = flags;\n    ++\t\tr->objects->kept_pack_cache.packs = packs;\n    ++\t\tr->objects->kept_pack_cache.flags = flags;\n     +\t}\n     +\n    -+\treturn r->objects->kept_pack_cache->packs;\n    ++\treturn r->objects->kept_pack_cache.packs;\n      }\n\n      int find_kept_pack_entry(struct repository *r,\n    @@ packfile.c: int find_kept_pack_entry(struct repository *r,\n      }\n\n      int has_object_pack(const struct object_id *oid)\n    -@@ packfile.c: int has_object_pack(const struct object_id *oid)\n    - \treturn find_pack_entry(the_repository, oid, &e);\n    - }\n    -\n    --int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n    -+int has_object_kept_pack(const struct object_id *oid,\n    -+\t\t\t unsigned flags)\n    - {\n    - \tstruct pack_entry e;\n    - \treturn find_kept_pack_entry(the_repository, oid, flags, &e);\n    -\n    - ## packfile.h ##\n    -@@ packfile.h: int packed_object_info(struct repository *r,\n    - void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n    - const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n    -\n    --#define ON_DISK_KEEP_PACKS 1\n    --#define IN_CORE_KEEP_PACKS 2\n    --#define ALL_KEEP_PACKS (ON_DISK_KEEP_PACKS | IN_CORE_KEEP_PACKS)\n    --\n    - /*\n    -  * Iff a pack file in the given repository contains the object named by sha1,\n    -  * return true and store its location to e.\n    -\n    - ## revision.c ##\n    -@@ revision.c: static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n    - \t\tdie(_(\"--unpacked=<packfile> no longer supported\"));\n    - \t} else if (!strcmp(arg, \"--no-kept-objects\")) {\n    - \t\trevs->no_kept_objects = 1;\n    --\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n    --\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n    -+\t\trevs->keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n    -+\t\trevs->keep_pack_cache_flags |= CACHE_ON_DISK_KEEP_PACKS;\n    - \t} else if (skip_prefix(arg, \"--no-kept-objects=\", &optarg)) {\n    - \t\trevs->no_kept_objects = 1;\n    - \t\tif (!strcmp(optarg, \"in-core\"))\n    --\t\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n    -+\t\t\trevs->keep_pack_cache_flags |= CACHE_IN_CORE_KEEP_PACKS;\n    - \t\tif (!strcmp(optarg, \"on-disk\"))\n    --\t\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n    -+\t\t\trevs->keep_pack_cache_flags |= CACHE_ON_DISK_KEEP_PACKS;\n    - \t} else if (!strcmp(arg, \"-r\")) {\n    - \t\trevs->diff = 1;\n    - \t\trevs->diffopt.flags.recursive = 1;\n20:  d5561585c2 !  8:  bd492ec142 builtin/repack.c: add '--geometric' option\n    @@ Documentation/git-repack.txt: depth is 4095.\n     +\tcontains at least `<factor>` times the number of objects as the\n     +\tnext-largest pack.\n     ++\n    -+`git repack` ensures this by determining a \"cut\" of packfiles that need to be\n    -+repacked into one in order to ensure a geometric progression. It picks the\n    -+smallest set of packfiles such that as many of the larger packfiles (by count of\n    -+objects contained in that pack) may be left intact.\n    ++`git repack` ensures this by determining a \"cut\" of packfiles that need\n    ++to be repacked into one in order to ensure a geometric progression. It\n    ++picks the smallest set of packfiles such that as many of the larger\n    ++packfiles (by count of objects contained in that pack) may be left\n    ++intact.\n    +++\n    ++Unlike other repack modes, the set of objects to pack is determined\n    ++uniquely by the set of packs being \"rolled-up\"; in other words, the\n    ++packs determined to need to be combined in order to restore a geometric\n    ++progression.\n    +++\n    ++Loose objects are implicitly included in this \"roll-up\", without respect\n    ++to their reachability. This is subject to change in the future. This\n    ++option (implying a drastically different repack mode) is not guarenteed\n    ++to work with all other combinations of option to `git repack`).\n     +\n      Configuration\n      -------------\n--\n2.30.0.667.g81c0cbc6fd\n"},{"id":"417274","messageId":"82f6b45463586a90b95d7457a463eb7c50838648.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 2/8] revision: learn '--no-kept-objects'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:20Z","receivedAt":"2021-02-18T03:15:21Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A future caller will want to be able to perform a reachability traversal\nwhich terminates when visiting an object found in a kept pack. The\nclosest existing option is '--honor-pack-keep', but this isn't quite\nwhat we want. Instead of halting the traversal midway through, a full\ntraversal is always performed, and the results are only trimmed\nafterwords.\n\nBesides needing to introduce a new flag (since culling results\npost-facto can be different than halting the traversal as it's\nhappening), there is an additional wrinkle handling the distinction\nin-core and on-disk kept packs. That is: what kinds of kept pack should\nstop the traversal?\n\nIntroduce '--no-kept-objects[=<on-disk|in-core>]' to specify which kinds\nof kept packs, if any, should stop a traversal. This can be useful for\ncallers that want to perform a reachability analysis, but want to leave\ncertain packs alone (for e.g., when doing a geometric repack that has\nsome \"large\" packs which are kept in-core that it wants to leave alone).\n\nNote that this option is not guaranteed to produce exactly the set of\nobjects that aren't in kept packs, since it's possible the traversal\norder may end up in a situation where a non-kept ancestor was \"cut off\"\nby a kept object (at which point we would stop traversing). But, we\ndon't care about absolute correctness here, since this will eventually\nbe used as a purely additive guide in an upcoming new repack mode.\n\nExplicitly avoid documenting this new flag, since it is only used\ninternally. In theory we could avoid even adding it rev-list, but being\nable to spell this option out on the command-line makes some special\ncases easier to test without promising to keep it behaving consistently\nforever. Those tricky cases are exercised in t6114.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n revision.c            | 15 ++++++++++\n revision.h            |  4 +++\n t/t6114-keep-packs.sh | 69 +++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 88 insertions(+)\n create mode 100755 t/t6114-keep-packs.sh\n\ndiff --git a/revision.c b/revision.c\nindex 3efd994160..f9311f3a73 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -2336,6 +2336,16 @@ static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n \t\trevs->unpacked = 1;\n \t} else if (starts_with(arg, \"--unpacked=\")) {\n \t\tdie(_(\"--unpacked=<packfile> no longer supported\"));\n+\t} else if (!strcmp(arg, \"--no-kept-objects\")) {\n+\t\trevs->no_kept_objects = 1;\n+\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t} else if (skip_prefix(arg, \"--no-kept-objects=\", &optarg)) {\n+\t\trevs->no_kept_objects = 1;\n+\t\tif (!strcmp(optarg, \"in-core\"))\n+\t\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\tif (!strcmp(optarg, \"on-disk\"))\n+\t\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n \t} else if (!strcmp(arg, \"-r\")) {\n \t\trevs->diff = 1;\n \t\trevs->diffopt.flags.recursive = 1;\n@@ -3792,6 +3802,11 @@ enum commit_action get_commit_action(struct rev_info *revs, struct commit *commi\n \t\treturn commit_ignore;\n \tif (revs->unpacked && has_object_pack(&commit->object.oid))\n \t\treturn commit_ignore;\n+\tif (revs->no_kept_objects) {\n+\t\tif (has_object_kept_pack(&commit->object.oid,\n+\t\t\t\t\t revs->keep_pack_cache_flags))\n+\t\t\treturn commit_ignore;\n+\t}\n \tif (commit->object.flags & UNINTERESTING)\n \t\treturn commit_ignore;\n \tif (revs->line_level_traverse && !want_ancestry(revs)) {\ndiff --git a/revision.h b/revision.h\nindex e6be3c845e..a20a530d52 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -148,6 +148,7 @@ struct rev_info {\n \t\t\tedge_hint_aggressive:1,\n \t\t\tlimited:1,\n \t\t\tunpacked:1,\n+\t\t\tno_kept_objects:1,\n \t\t\tboundary:2,\n \t\t\tcount:1,\n \t\t\tleft_right:1,\n@@ -317,6 +318,9 @@ struct rev_info {\n \t * This is loaded from the commit-graph being used.\n \t */\n \tstruct bloom_filter_settings *bloom_filter_settings;\n+\n+\t/* misc. flags related to '--no-kept-objects' */\n+\tunsigned keep_pack_cache_flags;\n };\n \n int ref_excluded(struct string_list *, const char *path);\ndiff --git a/t/t6114-keep-packs.sh b/t/t6114-keep-packs.sh\nnew file mode 100755\nindex 0000000000..9239d8aa46\n--- /dev/null\n+++ b/t/t6114-keep-packs.sh\n@@ -0,0 +1,69 @@\n+#!/bin/sh\n+\n+test_description='rev-list with .keep packs'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup' '\n+\ttest_commit loose &&\n+\ttest_commit packed &&\n+\ttest_commit kept &&\n+\n+\tKEPT_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/kept\n+\t^refs/tags/packed\n+\tEOF\n+\t) &&\n+\tMISC_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/packed\n+\t^refs/tags/loose\n+\tEOF\n+\t) &&\n+\n+\ttouch .git/objects/pack/pack-$KEPT_PACK.keep\n+'\n+\n+rev_list_objects () {\n+\tgit rev-list \"$@\" >out &&\n+\tsort out\n+}\n+\n+idx_objects () {\n+\tgit show-index <$1 >expect-idx &&\n+\tcut -d\" \" -f2 <expect-idx | sort\n+}\n+\n+test_expect_success '--no-kept-objects excludes trees and blobs in .keep packs' '\n+\trev_list_objects --objects --all --no-object-names >kept &&\n+\trev_list_objects --objects --all --no-object-names --no-kept-objects >no-kept &&\n+\n+\tidx_objects .git/objects/pack/pack-$KEPT_PACK.idx >expect &&\n+\tcomm -3 kept no-kept >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success '--no-kept-objects excludes kept non-MIDX object' '\n+\ttest_config core.multiPackIndex true &&\n+\n+\t# Create a pack with just the commit object in pack, and do not mark it\n+\t# as kept (even though it appears in $KEPT_PACK, which does have a .keep\n+\t# file).\n+\tMIDX_PACK=$(git pack-objects .git/objects/pack/pack <<-EOF\n+\t$(git rev-parse kept)\n+\tEOF\n+\t) &&\n+\n+\t# Write a MIDX containing all packs, but use the version of the commit\n+\t# at \"kept\" in a non-kept pack by touching $MIDX_PACK.\n+\ttouch .git/objects/pack/pack-$MIDX_PACK.pack &&\n+\tgit multi-pack-index write &&\n+\n+\trev_list_objects --objects --no-object-names --no-kept-objects HEAD >actual &&\n+\t(\n+\t\tidx_objects .git/objects/pack/pack-$MISC_PACK.idx &&\n+\t\tgit rev-list --objects --no-object-names refs/tags/loose\n+\t) | sort >expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417275","messageId":"f9a5faf77322f6735a0283e793711629ebf3b65a.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 4/8] p5303: add missing &&-chains","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:26Z","receivedAt":"2021-02-18T03:15:24Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nThese are in a helper function, so the usual chain-lint doesn't notice\nthem. This function is still not perfect, as it has some git invocations\non the left-hand-side of the pipe, but it's primary purpose is timing,\nnot finding bugs or correctness issues.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex ce0c42cc9f..d90d714923 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -28,11 +28,11 @@ repack_into_n () {\n \t\t\tpush @commits, $_ if $. % 5 == 1;\n \t\t}\n \t\tprint reverse @commits;\n-\t' \"$1\" >pushes\n+\t' \"$1\" >pushes &&\n \n \t# create base packfile\n \thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack\n+\tgit pack-objects --delta-base-offset --revs staging/pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417276","messageId":"033e4e3f67b96489c3ba1b2ab7977e23fac34189.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:23Z","receivedAt":"2021-02-18T03:15:25Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In an upcoming commit, 'git repack' will want to create a pack comprised\nof all of the objects in some packs (the included packs) excluding any\nobjects in some other packs (the excluded packs).\n\nThis caller could iterate those packs themselves and feed the objects it\nfinds to 'git pack-objects' directly over stdin, but this approach has a\nfew downsides:\n\n  - It requires every caller that wants to drive 'git pack-objects' in\n    this way to implement pack iteration themselves. This forces the\n    caller to think about details like what order objects are fed to\n    pack-objects, which callers would likely rather not do.\n\n  - If the set of objects in included packs is large, it requires\n    sending a lot of data over a pipe, which is inefficient.\n\n  - The caller is forced to keep track of the excluded objects, too, and\n    make sure that it doesn't send any objects that appear in both\n    included and excluded packs.\n\nBut the biggest downside is the lack of a reachability traversal.\nBecause the caller passes in a list of objects directly, those objects\ndon't get a namehash assigned to them, which can have a negative impact\non the delta selection process, causing 'git pack-objects' to fail to\nfind good deltas even when they exist.\n\nThe caller could formulate a reachability traversal themselves, but the\nonly way to drive 'git pack-objects' in this way is to do a full\ntraversal, and then remove objects in the excluded packs after the\ntraversal is complete. This can be detrimental to callers who care\nabout performance, especially in repositories with many objects.\n\nIntroduce 'git pack-objects --stdin-packs' which remedies these four\nconcerns.\n\n'git pack-objects --stdin-packs' expects a list of pack names on stdin,\nwhere 'pack-xyz.pack' denotes that pack as included, and\n'^pack-xyz.pack' denotes it as excluded. The resulting pack includes all\nobjects that are present in at least one included pack, and aren't\npresent in any excluded pack.\n\nTo address the delta selection problem, 'git pack-objects --stdin-packs'\nworks as follows. First, it assembles a list of objects that it is going\nto pack, as above. Then, a reachability traversal is started, whose tips\nare any commits mentioned in included packs. Upon visiting an object, we\nfind its corresponding object_entry in the to_pack list, and set its\nnamehash parameter appropriately.\n\nTo avoid the traversal visiting more objects than it needs to, the\ntraversal is halted upon encountering an object which can be found in an\nexcluded pack (by marking the excluded packs as kept in-core, and\npassing --no-kept-objects=in-core to the revision machinery).\n\nThis can cause the traversal to halt early, for example if an object in\nan included pack is an ancestor of ones in excluded packs. But stopping\nearly is OK, since filling in the namehash fields of objects in the\nto_pack list is only additive (i.e., having it helps the delta selection\nprocess, but leaving it blank doesn't impact the correctness of the\nresulting pack).\n\nEven still, it is unlikely that this hurts us much in practice, since\nthe 'git repack --geometric' caller (which is introduced in a later\ncommit) marks small packs as included, and large ones as excluded.\nDuring ordinary use, the small packs usually represent pushes after a\nlarge repack, and so are unlikely to be ancestors of objects that\nalready exist in the repository.\n\n(I found it convenient while developing this patch to have 'git\npack-objects' report the number of objects which were visited and got\ntheir namehash fields filled in during traversal. This is also included\nin the below patch via trace2 data lines).\n\nSuggested-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  10 ++\n builtin/pack-objects.c             | 198 ++++++++++++++++++++++++++++-\n t/t5300-pack-object.sh             |  97 ++++++++++++++\n 3 files changed, 303 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex 54d715ead1..df533c3b19 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -85,6 +85,16 @@ base-name::\n \treference was included in the resulting packfile.  This\n \tcan be useful to send new tags to native Git clients.\n \n+--stdin-packs::\n+\tRead the basenames of packfiles (e.g., `pack-1234abcd.pack`)\n+\tfrom the standard input, instead of object names or revision\n+\targuments. The resulting pack contains all objects listed in the\n+\tincluded packs (those not beginning with `^`), excluding any\n+\tobjects listed in the excluded packs (beginning with `^`).\n++\n+Incompatible with `--revs`, or options that imply `--revs` (such as\n+`--all`), with the exception of `--unpacked`, which is compatible.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 6d62aaf59a..e766a4a43b 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -2986,6 +2986,186 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \treturn git_default_config(k, v, cb);\n }\n \n+/* Counters for trace2 output when in --stdin-packs mode. */\n+static int stdin_packs_found_nr;\n+static int stdin_packs_hints_nr;\n+\n+static int add_object_entry_from_pack(const struct object_id *oid,\n+\t\t\t\t      struct packed_git *p,\n+\t\t\t\t      uint32_t pos,\n+\t\t\t\t      void *_data)\n+{\n+\tstruct rev_info *revs = _data;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\toff_t ofs;\n+\tenum object_type type;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tif (have_duplicate_entry(oid, 0))\n+\t\treturn 0;\n+\n+\tofs = nth_packed_object_offset(p, pos);\n+\tif (!want_object_in_pack(oid, 0, &p, &ofs))\n+\t\treturn 0;\n+\n+\toi.typep = &type;\n+\tif (packed_object_info(the_repository, p, ofs, &oi) < 0)\n+\t\tdie(_(\"could not get type of object %s in pack %s\"),\n+\t\t    oid_to_hex(oid), p->pack_name);\n+\telse if (type == OBJ_COMMIT) {\n+\t\t/*\n+\t\t * commits in included packs are used as starting points for the\n+\t\t * subsequent revision walk\n+\t\t */\n+\t\tadd_pending_oid(revs, NULL, oid, 0);\n+\t}\n+\n+\tstdin_packs_found_nr++;\n+\n+\tcreate_object_entry(oid, type, 0, 0, 0, p, ofs);\n+\n+\treturn 0;\n+}\n+\n+static void show_commit_pack_hint(struct commit *commit, void *_data)\n+{\n+\t/* nothing to do; commits don't have a namehash */\n+}\n+\n+static void show_object_pack_hint(struct object *object, const char *name,\n+\t\t\t\t  void *_data)\n+{\n+\tstruct object_entry *oe = packlist_find(&to_pack, &object->oid);\n+\tif (!oe)\n+\t\treturn;\n+\n+\t/*\n+\t * Our 'to_pack' list was constructed by iterating all objects packed in\n+\t * included packs, and so doesn't have a non-zero hash field that you\n+\t * would typically pick up during a reachability traversal.\n+\t *\n+\t * Make a best-effort attempt to fill in the ->hash and ->no_try_delta\n+\t * here using a now in order to perhaps improve the delta selection\n+\t * process.\n+\t */\n+\toe->hash = pack_name_hash(name);\n+\toe->no_try_delta = name && no_try_delta(name);\n+\n+\tstdin_packs_hints_nr++;\n+}\n+\n+static int pack_mtime_cmp(const void *_a, const void *_b)\n+{\n+\tstruct packed_git *a = ((const struct string_list_item*)_a)->util;\n+\tstruct packed_git *b = ((const struct string_list_item*)_b)->util;\n+\n+\tif (a->mtime < b->mtime)\n+\t\treturn -1;\n+\telse if (b->mtime < a->mtime)\n+\t\treturn 1;\n+\telse\n+\t\treturn 0;\n+}\n+\n+static void read_packs_list_from_stdin(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list include_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list exclude_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list_item *item = NULL;\n+\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\t/*\n+\t * Use a revision walk to fill in the namehash of objects in the include\n+\t * packs. To save time, we'll avoid traversing through objects that are\n+\t * in excluded packs.\n+\t *\n+\t * That may cause us to avoid populating all of the namehash fields of\n+\t * all included objects, but our goal is best-effort, since this is only\n+\t * an optimization during delta selection.\n+\t */\n+\trevs.no_kept_objects = 1;\n+\trevs.keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\trevs.blob_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.tag_objects = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '^')\n+\t\t\tstring_list_append(&exclude_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&include_packs, buf.buf);\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&include_packs);\n+\tstring_list_sort(&exclude_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\n+\t\titem = string_list_lookup(&include_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&exclude_packs, pack_name);\n+\n+\t\tif (item)\n+\t\t\titem->util = p;\n+\t}\n+\n+\t/*\n+\t * First handle all of the excluded packs, marking them as kept in-core\n+\t * so that later calls to add_object_entry() discards any objects that\n+\t * are also found in excluded packs.\n+\t */\n+\tfor_each_string_list_item(item, &exclude_packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = 1;\n+\t}\n+\n+\t/*\n+\t * Order packs by ascending mtime; use QSORT directly to access the\n+\t * string_list_item's ->util pointer, which string_list_sort() does not\n+\t * provide.\n+\t */\n+\tQSORT(include_packs.items, include_packs.nr, pack_mtime_cmp);\n+\n+\tfor_each_string_list_item(item, &include_packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tfor_each_object_in_pack(p,\n+\t\t\t\t\tadd_object_entry_from_pack,\n+\t\t\t\t\t&revs,\n+\t\t\t\t\tFOR_EACH_OBJECT_PACK_ORDER);\n+\t}\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\ttraverse_commit_list(&revs,\n+\t\t\t     show_commit_pack_hint,\n+\t\t\t     show_object_pack_hint,\n+\t\t\t     NULL);\n+\n+\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_found\",\n+\t\t\t   stdin_packs_found_nr);\n+\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_hints\",\n+\t\t\t   stdin_packs_hints_nr);\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&include_packs, 0);\n+\tstring_list_clear(&exclude_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3489,6 +3669,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tstruct strvec rp = STRVEC_INIT;\n \tint rev_list_unpacked = 0, rev_list_all = 0, rev_list_reflog = 0;\n \tint rev_list_index = 0;\n+\tint stdin_packs = 0;\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tstruct option pack_objects_options[] = {\n \t\tOPT_SET_INT('q', \"quiet\", &progress,\n@@ -3539,6 +3720,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_SET_INT_F(0, \"indexed-objects\", &rev_list_index,\n \t\t\t      N_(\"include objects referred to by the index\"),\n \t\t\t      1, PARSE_OPT_NONEG),\n+\t\tOPT_BOOL(0, \"stdin-packs\", &stdin_packs,\n+\t\t\t N_(\"read packs from stdin\")),\n \t\tOPT_BOOL(0, \"stdout\", &pack_to_stdout,\n \t\t\t N_(\"output pack to stdout\")),\n \t\tOPT_BOOL(0, \"include-tag\", &include_tag,\n@@ -3645,7 +3828,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tuse_internal_rev_list = 1;\n \t\tstrvec_push(&rp, \"--indexed-objects\");\n \t}\n-\tif (rev_list_unpacked) {\n+\tif (rev_list_unpacked && !stdin_packs) {\n \t\tuse_internal_rev_list = 1;\n \t\tstrvec_push(&rp, \"--unpacked\");\n \t}\n@@ -3690,8 +3873,13 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (filter_options.choice) {\n \t\tif (!pack_to_stdout)\n \t\t\tdie(_(\"cannot use --filter without --stdout\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --filter with --stdin-packs\"));\n \t}\n \n+\tif (stdin_packs && use_internal_rev_list)\n+\t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -3750,7 +3938,13 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (progress)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n-\tif (!use_internal_rev_list)\n+\tif (stdin_packs) {\n+\t\t/* avoids adding objects in excluded packs */\n+\t\tignore_packed_keep_in_core = 1;\n+\t\tread_packs_list_from_stdin();\n+\t\tif (rev_list_unpacked)\n+\t\t\tadd_unreachable_loose_objects();\n+\t} else if (!use_internal_rev_list)\n \t\tread_object_list_from_stdin();\n \telse {\n \t\tget_object_list(rp.nr, rp.v);\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex 392201cabd..7138a54595 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -532,4 +532,101 @@ test_expect_success 'prefetch objects' '\n \ttest_line_count = 1 donelines\n '\n \n+test_expect_success 'setup for --stdin-packs tests' '\n+\tgit init stdin-packs &&\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\ttest_commit A &&\n+\t\ttest_commit B &&\n+\t\ttest_commit C &&\n+\n+\t\tfor id in A B C\n+\t\tdo\n+\t\t\tgit pack-objects .git/objects/pack/pack-$id \\\n+\t\t\t\t--incremental --revs <<-EOF\n+\t\t\trefs/tags/$id\n+\t\t\tEOF\n+\t\tdone &&\n+\n+\t\tls -la .git/objects/pack\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs with excluded packs' '\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\tPACK_A=\"$(basename .git/objects/pack/pack-A-*.pack)\" &&\n+\t\tPACK_B=\"$(basename .git/objects/pack/pack-B-*.pack)\" &&\n+\t\tPACK_C=\"$(basename .git/objects/pack/pack-C-*.pack)\" &&\n+\n+\t\tgit pack-objects test --stdin-packs <<-EOF &&\n+\t\t$PACK_A\n+\t\t^$PACK_B\n+\t\t$PACK_C\n+\t\tEOF\n+\n+\t\t(\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-A-*.idx) &&\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-C-*.idx)\n+\t\t) >expect.raw &&\n+\t\tgit show-index <$(ls test-*.idx) >actual.raw &&\n+\n+\t\tcut -d\" \" -f2 <expect.raw | sort >expect &&\n+\t\tcut -d\" \" -f2 <actual.raw | sort >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs is incompatible with --filter' '\n+\t(\n+\t\tcd stdin-packs &&\n+\t\ttest_must_fail git pack-objects --stdin-packs --stdout \\\n+\t\t\t--filter=blob:none </dev/null 2>err &&\n+\t\ttest_i18ngrep \"cannot use --filter with --stdin-packs\" err\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs is incompatible with --revs' '\n+\t(\n+\t\tcd stdin-packs &&\n+\t\ttest_must_fail git pack-objects --stdin-packs --revs out \\\n+\t\t\t</dev/null 2>err &&\n+\t\ttest_i18ngrep \"cannot use internal rev list with --stdin-packs\" err\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs with loose objects' '\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\tPACK_A=\"$(basename .git/objects/pack/pack-A-*.pack)\" &&\n+\t\tPACK_B=\"$(basename .git/objects/pack/pack-B-*.pack)\" &&\n+\t\tPACK_C=\"$(basename .git/objects/pack/pack-C-*.pack)\" &&\n+\n+\t\ttest_commit D && # loose\n+\n+\t\tgit pack-objects test2 --stdin-packs --unpacked <<-EOF &&\n+\t\t$PACK_A\n+\t\t^$PACK_B\n+\t\t$PACK_C\n+\t\tEOF\n+\n+\t\t(\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-A-*.idx) &&\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-C-*.idx) &&\n+\t\t\tgit rev-list --objects --no-object-names \\\n+\t\t\t\trefs/tags/C..refs/tags/D\n+\n+\t\t) >expect.raw &&\n+\t\tls -la . &&\n+\t\tgit show-index <$(ls test2-*.idx) >actual.raw &&\n+\n+\t\tcut -d\" \" -f2 <expect.raw | sort >expect &&\n+\t\tcut -d\" \" -f2 <actual.raw | sort >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_done\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417277","messageId":"67af143fd1f3bcece0a8b27894cbdcdc5ae60ae8.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 6/8] builtin/pack-objects.c: rewrite honor-pack-keep logic","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:33Z","receivedAt":"2021-02-18T03:15:41Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nNow that we have find_kept_pack_entry(), we don't have to manually keep\nhunting through every pack to find a possible \"kept\" duplicate of the\nobject. This should be faster, assuming only a portion of your total\npacks are actually kept.\n\nNote that we have to re-order the logic a bit here; we can deal with the\ndisqualifying situations first (e.g., finding the object in a non-local\npack with --local), then \"kept\" situation(s), and then just fall back to\nother \"--local\" conditions.\n\nHere are the results from p5303 (measurements again taken on the\nkernel):\n\n  Test                                        HEAD^                   HEAD\n  -----------------------------------------------------------------------------------------------\n  5303.5: repack (1)                          57.26(54.59+10.84)      57.34(54.66+10.88) +0.1%\n  5303.6: repack with kept (1)                57.33(54.80+10.51)      57.38(54.83+10.49) +0.1%\n  5303.11: repack (50)                        71.54(88.57+4.84)       71.70(88.99+4.74) +0.2%\n  5303.12: repack with kept (50)              85.12(102.05+4.94)      72.58(89.61+4.78) -14.7%\n  5303.17: repack (1000)                      216.87(490.79+14.57)    217.19(491.72+14.25) +0.1%\n  5303.18: repack with kept (1000)            665.63(938.87+15.76)    246.12(520.07+14.93) -63.0%\n\nand the --stdin-packs timings:\n\n  5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)         0.00(0.00+0.00) -100.0%\n  5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)        3.43(11.75+0.24) -2.8%\n  5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)     130.50(307.15+7.66) -33.4%\n\nSo our repack with an empty .keep pack is roughly as fast as one without\na .keep pack up to 50 packs. But the --stdin-packs case scales a little\nbetter, too.\n\nNotably, it is faster than a repack of the same size and a kept pack. It\nlooks at fewer objects, of course, but the penalty for looking at many\npacks isn't as costly.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 131 ++++++++++++++++++++++++-----------------\n 1 file changed, 78 insertions(+), 53 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex e766a4a43b..be3ba60bc2 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1188,7 +1188,8 @@ static int have_duplicate_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int want_found_object(int exclude, struct packed_git *p)\n+static int want_found_object(const struct object_id *oid, int exclude,\n+\t\t\t     struct packed_git *p)\n {\n \tif (exclude)\n \t\treturn 1;\n@@ -1204,27 +1205,82 @@ static int want_found_object(int exclude, struct packed_git *p)\n \t * make sure no copy of this object appears in _any_ pack that makes us\n \t * to omit the object, so we need to check all the packs.\n \t *\n-\t * We can however first check whether these options can possible matter;\n+\t * We can however first check whether these options can possibly matter;\n \t * if they do not matter we know we want the object in generated pack.\n \t * Otherwise, we signal \"-1\" at the end to tell the caller that we do\n \t * not know either way, and it needs to check more packs.\n \t */\n-\tif (!ignore_packed_keep_on_disk &&\n-\t    !ignore_packed_keep_in_core &&\n-\t    (!local || !have_non_local_packs))\n-\t\treturn 1;\n \n+\t/*\n+\t * Objects in packs borrowed from elsewhere are discarded regardless of\n+\t * if they appear in other packs that weren't borrowed.\n+\t */\n \tif (local && !p->pack_local)\n \t\treturn 0;\n-\tif (p->pack_local &&\n-\t    ((ignore_packed_keep_on_disk && p->pack_keep) ||\n-\t     (ignore_packed_keep_in_core && p->pack_keep_in_core)))\n-\t\treturn 0;\n+\n+\t/*\n+\t * Then handle .keep first, as we have a fast(er) path there.\n+\t */\n+\tif (ignore_packed_keep_on_disk || ignore_packed_keep_in_core) {\n+\t\t/*\n+\t\t * Set the flags for the kept-pack cache to be the ones we want\n+\t\t * to ignore.\n+\t\t *\n+\t\t * That is, if we are ignoring objects in on-disk keep packs,\n+\t\t * then we want to search through the on-disk keep and ignore\n+\t\t * the in-core ones.\n+\t\t */\n+\t\tunsigned flags = 0;\n+\t\tif (ignore_packed_keep_on_disk)\n+\t\t\tflags |= ON_DISK_KEEP_PACKS;\n+\t\tif (ignore_packed_keep_in_core)\n+\t\t\tflags |= IN_CORE_KEEP_PACKS;\n+\n+\t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n+\t\t\treturn 0;\n+\t\tif (ignore_packed_keep_in_core && p->pack_keep_in_core)\n+\t\t\treturn 0;\n+\t\tif (has_object_kept_pack(oid, flags))\n+\t\t\treturn 0;\n+\t}\n+\n+\t/*\n+\t * At this point we know definitively that either we don't care about\n+\t * keep-packs, or the object is not in one. Keep checking other\n+\t * conditions...\n+\t */\n+\tif (!local || !have_non_local_packs)\n+\t\treturn 1;\n \n \t/* we don't know yet; keep looking for more packs */\n \treturn -1;\n }\n \n+static int want_object_in_pack_one(struct packed_git *p,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   int exclude,\n+\t\t\t\t   struct packed_git **found_pack,\n+\t\t\t\t   off_t *found_offset)\n+{\n+\toff_t offset;\n+\n+\tif (p == *found_pack)\n+\t\toffset = *found_offset;\n+\telse\n+\t\toffset = find_pack_entry_one(oid->hash, p);\n+\n+\tif (offset) {\n+\t\tif (!*found_pack) {\n+\t\t\tif (!is_pack_valid(p))\n+\t\t\t\treturn -1;\n+\t\t\t*found_offset = offset;\n+\t\t\t*found_pack = p;\n+\t\t}\n+\t\treturn want_found_object(oid, exclude, p);\n+\t}\n+\treturn -1;\n+}\n+\n /*\n  * Check whether we want the object in the pack (e.g., we do not want\n  * objects found in non-local stores if the \"--local\" option was used).\n@@ -1252,7 +1308,7 @@ static int want_object_in_pack(const struct object_id *oid,\n \t * are present we will determine the answer right now.\n \t */\n \tif (*found_pack) {\n-\t\twant = want_found_object(exclude, *found_pack);\n+\t\twant = want_found_object(oid, exclude, *found_pack);\n \t\tif (want != -1)\n \t\t\treturn want;\n \t}\n@@ -1260,53 +1316,22 @@ static int want_object_in_pack(const struct object_id *oid,\n \tfor (m = get_multi_pack_index(the_repository); m; m = m->next) {\n \t\tstruct pack_entry e;\n \t\tif (fill_midx_entry(the_repository, oid, &e, m)) {\n-\t\t\tstruct packed_git *p = e.p;\n-\t\t\toff_t offset;\n-\n-\t\t\tif (p == *found_pack)\n-\t\t\t\toffset = *found_offset;\n-\t\t\telse\n-\t\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\t\tif (offset) {\n-\t\t\t\tif (!*found_pack) {\n-\t\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\t\tcontinue;\n-\t\t\t\t\t*found_offset = offset;\n-\t\t\t\t\t*found_pack = p;\n-\t\t\t\t}\n-\t\t\t\twant = want_found_object(exclude, p);\n-\t\t\t\tif (want != -1)\n-\t\t\t\t\treturn want;\n-\t\t\t}\n-\t\t}\n-\t}\n-\n-\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n-\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\toff_t offset;\n-\n-\t\tif (p == *found_pack)\n-\t\t\toffset = *found_offset;\n-\t\telse\n-\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\tif (offset) {\n-\t\t\tif (!*found_pack) {\n-\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\tcontinue;\n-\t\t\t\t*found_offset = offset;\n-\t\t\t\t*found_pack = p;\n-\t\t\t}\n-\t\t\twant = want_found_object(exclude, p);\n-\t\t\tif (!exclude && want > 0)\n-\t\t\t\tlist_move(&p->mru,\n-\t\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\t\twant = want_object_in_pack_one(e.p, oid, exclude, found_pack, found_offset);\n \t\t\tif (want != -1)\n \t\t\t\treturn want;\n \t\t}\n \t}\n \n+\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n+\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n+\t\twant = want_object_in_pack_one(p, oid, exclude, found_pack, found_offset);\n+\t\tif (!exclude && want > 0)\n+\t\t\tlist_move(&p->mru,\n+\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\tif (want != -1)\n+\t\t\treturn want;\n+\t}\n+\n \tif (uri_protocols.nr) {\n \t\tstruct configured_exclusion *ex =\n \t\t\toidmap_get(&configured_exclusions, oid);\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417278","messageId":"181c104a038f1ad8c30e68013bcdbc79cf394ea4.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 5/8] p5303: measure time to repack with keep","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:30Z","receivedAt":"2021-02-18T03:15:42Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nAdd two new tests to measure repack performance. Both test split the\nrepository into synthetic \"pushes\", and then leave the remaining objects\nin a big base pack.\n\nThe first new test marks an empty pack as \"kept\" and then passes\n--honor-pack-keep to avoid including objects in it. That doesn't change\nthe resulting pack, but it does let us compare to the normal repack case\nto see how much overhead we add to check whether objects are kept or\nnot.\n\nThe other test is of --stdin-packs, which gives us a sense of how that\nnumber scales based on the number of packs we provide as input. In each\nof those tests, the empty pack isn't considered, but the residual pack\n(objects that were left over and not included in one of the synthetic\npush packs) is marked as kept.\n\n(Note that in the single-pack case of the --stdin-packs test, there is\nnothing do since there are no non-excluded packs).\n\nHere are some timings on a recent clone of the kernel:\n\n  5303.5: repack (1)                          57.26(54.59+10.84)\n  5303.6: repack with kept (1)                57.33(54.80+10.51)\n\nin the 50-pack case, things start to slow down:\n\n  5303.11: repack (50)                        71.54(88.57+4.84)\n  5303.12: repack with kept (50)              85.12(102.05+4.94)\n\nand by the time we hit 1,000 packs, things are substantially worse, even\nthough the resulting pack produced is the same:\n\n  5303.17: repack (1000)                      216.87(490.79+14.57)\n  5303.18: repack with kept (1000)            665.63(938.87+15.76)\n\nLikewise, the scaling is pretty extreme on --stdin-packs:\n\n  5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)\n  5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)\n  5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)\n\nThat's because the code paths around handling .keep files are known to\nscale badly; they look in every single pack file to find each object.\nOur solution to that was to notice that most repos don't have keep\nfiles, and to make that case a fast path. But as soon as you add a\nsingle .keep, that part of pack-objects slows down again (even if we\nhave fewer objects total to look at).\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 34 ++++++++++++++++++++++++++++++++--\n 1 file changed, 32 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex d90d714923..35c0cbdf49 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -31,8 +31,15 @@ repack_into_n () {\n \t' \"$1\" >pushes &&\n \n \t# create base packfile\n-\thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack &&\n+\tbase_pack=$(\n+\t\thead -n 1 pushes |\n+\t\tgit pack-objects --delta-base-offset --revs staging/pack\n+\t) &&\n+\ttest_export base_pack &&\n+\n+\t# create an empty packfile\n+\tempty_pack=$(git pack-objects staging/pack </dev/null) &&\n+\ttest_export empty_pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n@@ -49,6 +56,12 @@ repack_into_n () {\n \t\tlast=$rev\n \tdone <pushes &&\n \n+\t(\n+\t\tfind staging -type f -name 'pack-*.pack' |\n+\t\t\txargs -n 1 basename | grep -v \"$base_pack\" &&\n+\t\tprintf \"^pack-%s.pack\\n\" $base_pack\n+\t) >stdin.packs\n+\n \t# and install the whole thing\n \trm -f .git/objects/pack/* &&\n \tmv staging/* .git/objects/pack/\n@@ -91,6 +104,23 @@ do\n \t\t  --reflog --indexed-objects --delta-base-offset \\\n \t\t  --stdout </dev/null >/dev/null\n \t'\n+\n+\ttest_perf \"repack with kept ($nr_packs)\" '\n+\t\tgit pack-objects --keep-true-parents \\\n+\t\t  --keep-pack=pack-$empty_pack.pack \\\n+\t\t  --honor-pack-keep --non-empty --all \\\n+\t\t  --reflog --indexed-objects --delta-base-offset \\\n+\t\t  --stdout </dev/null >/dev/null\n+\t'\n+\n+\ttest_perf \"repack with --stdin-packs ($nr_packs)\" '\n+\t\tgit pack-objects \\\n+\t\t  --keep-true-parents \\\n+\t\t  --stdin-packs \\\n+\t\t  --non-empty \\\n+\t\t  --delta-base-offset \\\n+\t\t  --stdout <stdin.packs >/dev/null\n+\t'\n done\n \n # Measure pack loading with 10,000 packs.\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417279","messageId":"e9e04b95e7b44a7a08920ee8a68fa837318f0e65.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:37Z","receivedAt":"2021-02-18T03:15:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nIn a recent patch we added a function 'find_kept_pack_entry()' to look\nfor an object only among kept packs.\n\nWhile this function avoids doing any lookup work in non-kept packs, it\nis still linear in the number of packs, since we have to traverse the\nlinked list of packs once per object. Let's cache a reduced version of\nthat list to save us time.\n\nNote that this cache will last the lifetime of the program. We could\ninvalidate it on reprepare_packed_git(), but there's not much point in\nbeing rigorous here:\n\n  - we might already fail to notice new .keep packs showing up after the\n    program starts. We only reprepare_packed_git() when we fail to find\n    an object. But adding a new pack won't cause that to happen.\n    Somebody repacking could add a new pack and delete an old one, but\n    most of the time we'd have a descriptor or mmap open to the old\n    pack anyway, so we might not even notice.\n\n  - in pack-objects we already cache the .keep state at startup, since\n    56dfeb6263 (pack-objects: compute local/ignore_pack_keep early,\n    2016-07-29). So this is just extending that concept further.\n\n  - we don't have to worry about any packed_git being removed; we always\n    keep the old structs around, even after reprepare_packed_git()\n\nWe do defensively invalidate the cache in case the set of kept packs\nbeing asked for changes (e.g., only in-core kept packs were cached, but\nsuddenly the caller also wants on-disk kept packs, too). In theory we\ncould build all three caches and switch between them, but it's not\nnecessary, since this patch (and series) never changes the set of kept\npacks that it wants to inspect from the cache.\n\nSo that \"optimization\" is more about being defensive in the face of\nfuture changes than it is about asking for multiple kinds of kept packs\nin this patch.\n\nHere are p5303 results (as always, measured against the kernel):\n\n  Test                                        HEAD^                   HEAD\n  -----------------------------------------------------------------------------------------------\n  5303.5: repack (1)                          57.34(54.66+10.88)      56.98(54.36+10.98) -0.6%\n  5303.6: repack with kept (1)                57.38(54.83+10.49)      57.17(54.97+10.26) -0.4%\n  5303.11: repack (50)                        71.70(88.99+4.74)       71.62(88.48+5.08) -0.1%\n  5303.12: repack with kept (50)              72.58(89.61+4.78)       71.56(88.80+4.59) -1.4%\n  5303.17: repack (1000)                      217.19(491.72+14.25)    217.31(490.82+14.53) +0.1%\n  5303.18: repack with kept (1000)            246.12(520.07+14.93)    217.08(490.37+15.10) -11.8%\n\nand the --stdin-packs case, which scales a little bit better (although\nnot by that much even at 1,000 packs):\n\n  5303.7: repack with --stdin-packs (1)       0.00(0.00+0.00)         0.00(0.00+0.00) =\n  5303.13: repack with --stdin-packs (50)     3.43(11.75+0.24)        3.43(11.69+0.30) +0.0%\n  5303.19: repack with --stdin-packs (1000)   130.50(307.15+7.66)     125.13(301.36+8.04) -4.1%\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-store.h |  5 +++\n packfile.c     | 99 ++++++++++++++++++++++++++++----------------------\n 2 files changed, 61 insertions(+), 43 deletions(-)\n\ndiff --git a/object-store.h b/object-store.h\nindex 541dab0858..ec32c23dcb 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -153,6 +153,11 @@ struct raw_object_store {\n \t/* A most-recently-used ordered version of the packed_git list. */\n \tstruct list_head packed_git_mru;\n \n+\tstruct {\n+\t\tstruct packed_git **packs;\n+\t\tunsigned flags;\n+\t} kept_pack_cache;\n+\n \t/*\n \t * A map of packfiles to packed_git structs for tracking which\n \t * packs have been loaded already.\ndiff --git a/packfile.c b/packfile.c\nindex 7f84f221ce..57d5b436fb 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2042,10 +2042,7 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int find_one_pack_entry(struct repository *r,\n-\t\t\t       const struct object_id *oid,\n-\t\t\t       struct pack_entry *e,\n-\t\t\t       int kept_only)\n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2055,49 +2052,63 @@ static int find_one_pack_entry(struct repository *r,\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (!fill_midx_entry(r, oid, e, m))\n-\t\t\tcontinue;\n-\n-\t\tif (!kept_only)\n-\t\t\treturn 1;\n-\n-\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n-\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n+\t\tif (fill_midx_entry(r, oid, e, m))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (p->multi_pack_index && !kept_only) {\n-\t\t\t/*\n-\t\t\t * If this pack is covered by the MIDX, we'd have found\n-\t\t\t * the object already in the loop above if it was here,\n-\t\t\t * so don't bother looking.\n-\t\t\t *\n-\t\t\t * The exception is if we are looking only at kept\n-\t\t\t * packs. An object can be present in two packs covered\n-\t\t\t * by the MIDX, one kept and one not-kept. And as the\n-\t\t\t * MIDX points to only one copy of each object, it might\n-\t\t\t * have returned only the non-kept version above. We\n-\t\t\t * have to check again to be thorough.\n-\t\t\t */\n-\t\t\tcontinue;\n-\t\t}\n-\t\tif (!kept_only ||\n-\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n-\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n-\t\t\tif (fill_pack_entry(oid, e, p)) {\n-\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\t\treturn 1;\n-\t\t\t}\n+\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n+\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\treturn 1;\n \t\t}\n \t}\n \treturn 0;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static void maybe_invalidate_kept_pack_cache(struct repository *r,\n+\t\t\t\t\t     unsigned flags)\n {\n-\treturn find_one_pack_entry(r, oid, e, 0);\n+\tif (!r->objects->kept_pack_cache.packs)\n+\t\treturn;\n+\tif (r->objects->kept_pack_cache.flags == flags)\n+\t\treturn;\n+\tFREE_AND_NULL(r->objects->kept_pack_cache.packs);\n+\tr->objects->kept_pack_cache.flags = 0;\n+}\n+\n+static struct packed_git **kept_pack_cache(struct repository *r, unsigned flags)\n+{\n+\tmaybe_invalidate_kept_pack_cache(r, flags);\n+\n+\tif (!r->objects->kept_pack_cache.packs) {\n+\t\tstruct packed_git **packs = NULL;\n+\t\tsize_t nr = 0, alloc = 0;\n+\t\tstruct packed_git *p;\n+\n+\t\t/*\n+\t\t * We want \"all\" packs here, because we need to cover ones that\n+\t\t * are used by a midx, as well. We need to look in every one of\n+\t\t * them (instead of the midx itself) to cover duplicates. It's\n+\t\t * possible that an object is found in two packs that the midx\n+\t\t * covers, one kept and one not kept, but the midx returns only\n+\t\t * the non-kept version.\n+\t\t */\n+\t\tfor (p = get_all_packs(r); p; p = p->next) {\n+\t\t\tif ((p->pack_keep && (flags & ON_DISK_KEEP_PACKS)) ||\n+\t\t\t    (p->pack_keep_in_core && (flags & IN_CORE_KEEP_PACKS))) {\n+\t\t\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\t\t\tpacks[nr++] = p;\n+\t\t\t}\n+\t\t}\n+\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\tpacks[nr] = NULL;\n+\n+\t\tr->objects->kept_pack_cache.packs = packs;\n+\t\tr->objects->kept_pack_cache.flags = flags;\n+\t}\n+\n+\treturn r->objects->kept_pack_cache.packs;\n }\n \n int find_kept_pack_entry(struct repository *r,\n@@ -2105,13 +2116,15 @@ int find_kept_pack_entry(struct repository *r,\n \t\t\t unsigned flags,\n \t\t\t struct pack_entry *e)\n {\n-\t/*\n-\t * Load all packs, including midx packs, since our \"kept\" strategy\n-\t * relies on that. We're relying on the side effect of it setting up\n-\t * r->objects->packed_git, which is a little ugly.\n-\t */\n-\tget_all_packs(r);\n-\treturn find_one_pack_entry(r, oid, e, flags);\n+\tstruct packed_git **cache;\n+\n+\tfor (cache = kept_pack_cache(r, flags); *cache; cache++) {\n+\t\tstruct packed_git *p = *cache;\n+\t\tif (fill_pack_entry(oid, e, p))\n+\t\t\treturn 1;\n+\t}\n+\n+\treturn 0;\n }\n \n int has_object_pack(const struct object_id *oid)\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417280","messageId":"bd492ec1429e86a5fc836a474f3dbf87b8723be5.1613618042.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"[PATCH v3 8/8] builtin/repack.c: add '--geometric' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-18T03:14:41Z","receivedAt":"2021-02-18T03:15:46Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Often it is useful to both:\n\n  - have relatively few packfiles in a repository, and\n\n  - avoid having so few packfiles in a repository that we repack its\n    entire contents regularly\n\nThis patch implements a '--geometric=<n>' option in 'git repack'. This\nallows the caller to specify that they would like each pack to be at\nleast a factor times as large as the previous largest pack (by object\ncount).\n\nConcretely, say that a repository has 'n' packfiles, labeled P1, P2,\n..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\nWith a geometric factor of 'r', it should be that:\n\n  objects(Pi) > r*objects(P(i-1))\n\nfor all i in [1, n], where the packs are sorted by\n\n  objects(P1) <= objects(P2) <= ... <= objects(Pn).\n\nSince finding a true optimal repacking is NP-hard, we approximate it\nalong two directions:\n\n  1. We assume that there is a cutoff of packs _before starting the\n     repack_ where everything to the right of that cut-off already forms\n     a geometric progression (or no cutoff exists and everything must be\n     repacked).\n\n  2. We assume that everything smaller than the cutoff count must be\n     repacked. This forms our base assumption, but it can also cause\n     even the \"heavy\" packs to get repacked, for e.g., if we have 6\n     packs containing the following number of objects:\n\n       1, 1, 1, 2, 4, 32\n\n     then we would place the cutoff between '1, 1' and '1, 2, 4, 32',\n     rolling up the first two packs into a pack with 2 objects. That\n     breaks our progression and leaves us:\n\n       2, 1, 2, 4, 32\n         ^\n\n     (where the '^' indicates the position of our split). To restore a\n     progression, we move the split forward (towards larger packs)\n     joining each pack into our new pack until a geometric progression\n     is restored. Here, that looks like:\n\n       2, 1, 2, 4, 32  ~>  3, 2, 4, 32  ~>  5, 4, 32  ~> ... ~> 9, 32\n         ^                   ^                ^                   ^\n\nThis has the advantage of not repacking the heavy-side of packs too\noften while also only creating one new pack at a time. Another wrinkle\nis that we assume that loose, indexed, and reflog'd objects are\ninsignificant, and lump them into any new pack that we create. This can\nlead to non-idempotent results.\n\nSuggested-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt |  22 +++++\n builtin/repack.c             | 187 ++++++++++++++++++++++++++++++++++-\n t/t7703-repack-geometric.sh  | 137 +++++++++++++++++++++++++\n 3 files changed, 342 insertions(+), 4 deletions(-)\n create mode 100755 t/t7703-repack-geometric.sh\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex 92f146d27d..21c7068925 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -165,6 +165,28 @@ depth is 4095.\n \tPass the `--delta-islands` option to `git-pack-objects`, see\n \tlinkgit:git-pack-objects[1].\n \n+-g=<factor>::\n+--geometric=<factor>::\n+\tArrange resulting pack structure so that each successive pack\n+\tcontains at least `<factor>` times the number of objects as the\n+\tnext-largest pack.\n++\n+`git repack` ensures this by determining a \"cut\" of packfiles that need\n+to be repacked into one in order to ensure a geometric progression. It\n+picks the smallest set of packfiles such that as many of the larger\n+packfiles (by count of objects contained in that pack) may be left\n+intact.\n++\n+Unlike other repack modes, the set of objects to pack is determined\n+uniquely by the set of packs being \"rolled-up\"; in other words, the\n+packs determined to need to be combined in order to restore a geometric\n+progression.\n++\n+Loose objects are implicitly included in this \"roll-up\", without respect\n+to their reachability. This is subject to change in the future. This\n+option (implying a drastically different repack mode) is not guarenteed\n+to work with all other combinations of option to `git repack`).\n+\n Configuration\n -------------\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 01440de2d5..bcf280b10d 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -297,6 +297,124 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n #define ALL_INTO_ONE 1\n #define LOOSEN_UNREACHABLE 2\n \n+struct pack_geometry {\n+\tstruct packed_git **pack;\n+\tuint32_t pack_nr, pack_alloc;\n+\tuint32_t split;\n+};\n+\n+static uint32_t geometry_pack_weight(struct packed_git *p)\n+{\n+\tif (open_pack_index(p))\n+\t\tdie(_(\"cannot open index for %s\"), p->pack_name);\n+\treturn p->num_objects;\n+}\n+\n+static int geometry_cmp(const void *va, const void *vb)\n+{\n+\tuint32_t aw = geometry_pack_weight(*(struct packed_git **)va),\n+\t\t bw = geometry_pack_weight(*(struct packed_git **)vb);\n+\n+\tif (aw < bw)\n+\t\treturn -1;\n+\tif (aw > bw)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static void init_pack_geometry(struct pack_geometry **geometry_p)\n+{\n+\tstruct packed_git *p;\n+\tstruct pack_geometry *geometry;\n+\n+\t*geometry_p = xcalloc(1, sizeof(struct pack_geometry));\n+\tgeometry = *geometry_p;\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tif (!pack_kept_objects && p->pack_keep)\n+\t\t\tcontinue;\n+\n+\t\tALLOC_GROW(geometry->pack,\n+\t\t\t   geometry->pack_nr + 1,\n+\t\t\t   geometry->pack_alloc);\n+\n+\t\tgeometry->pack[geometry->pack_nr] = p;\n+\t\tgeometry->pack_nr++;\n+\t}\n+\n+\tQSORT(geometry->pack, geometry->pack_nr, geometry_cmp);\n+}\n+\n+static void split_pack_geometry(struct pack_geometry *geometry, int factor)\n+{\n+\tuint32_t i;\n+\tuint32_t split;\n+\toff_t total_size = 0;\n+\n+\tif (geometry->pack_nr <= 1) {\n+\t\tgeometry->split = geometry->pack_nr;\n+\t\treturn;\n+\t}\n+\n+\tsplit = geometry->pack_nr - 1;\n+\n+\t/*\n+\t * First, count the number of packs (in descending order of size) which\n+\t * already form a geometric progression.\n+\t */\n+\tfor (i = geometry->pack_nr - 1; i > 0; i--) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tstruct packed_git *prev = geometry->pack[i - 1];\n+\t\tif (geometry_pack_weight(ours) >= factor * geometry_pack_weight(prev))\n+\t\t\tsplit--;\n+\t\telse\n+\t\t\tbreak;\n+\t}\n+\n+\tif (split) {\n+\t\t/*\n+\t\t * Move the split one to the right, since the top element in the\n+\t\t * last-compared pair can't be in the progression. Only do this\n+\t\t * when we split in the middle of the array (otherwise if we got\n+\t\t * to the end, then the split is in the right place).\n+\t\t */\n+\t\tsplit++;\n+\t}\n+\n+\t/*\n+\t * Then, anything to the left of 'split' must be in a new pack. But,\n+\t * creating that new pack may cause packs in the heavy half to no longer\n+\t * form a geometric progression.\n+\t *\n+\t * Compute an expected size of the new pack, and then determine how many\n+\t * packs in the heavy half need to be joined into it (if any) to restore\n+\t * the geometric progression.\n+\t */\n+\tfor (i = 0; i < split; i++)\n+\t\ttotal_size += geometry_pack_weight(geometry->pack[i]);\n+\tfor (i = split; i < geometry->pack_nr; i++) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tif (geometry_pack_weight(ours) < factor * total_size) {\n+\t\t\tsplit++;\n+\t\t\ttotal_size += geometry_pack_weight(ours);\n+\t\t} else\n+\t\t\tbreak;\n+\t}\n+\n+\tgeometry->split = split;\n+}\n+\n+static void clear_pack_geometry(struct pack_geometry *geometry)\n+{\n+\tif (!geometry)\n+\t\treturn;\n+\n+\tfree(geometry->pack);\n+\tgeometry->pack_nr = 0;\n+\tgeometry->pack_alloc = 0;\n+\tgeometry->split = 0;\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -304,6 +422,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list names = STRING_LIST_INIT_DUP;\n \tstruct string_list rollback = STRING_LIST_INIT_NODUP;\n \tstruct string_list existing_packs = STRING_LIST_INIT_DUP;\n+\tstruct pack_geometry *geometry = NULL;\n \tstruct strbuf line = STRBUF_INIT;\n \tint i, ext, ret;\n \tFILE *out;\n@@ -316,6 +435,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tint no_update_server_info = 0;\n \tstruct pack_objects_args po_args = {NULL};\n+\tint geometric_factor = 0;\n \n \tstruct option builtin_repack_options[] = {\n \t\tOPT_BIT('a', NULL, &pack_everything,\n@@ -356,6 +476,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t\tN_(\"repack objects in packs marked with .keep\")),\n \t\tOPT_STRING_LIST(0, \"keep-pack\", &keep_pack_list, N_(\"name\"),\n \t\t\t\tN_(\"do not repack this pack\")),\n+\t\tOPT_INTEGER('g', \"geometric\", &geometric_factor,\n+\t\t\t    N_(\"find a geometric progression with factor <N>\")),\n \t\tOPT_END()\n \t};\n \n@@ -382,6 +504,13 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (write_bitmaps && !(pack_everything & ALL_INTO_ONE))\n \t\tdie(_(incremental_bitmap_conflict_error));\n \n+\tif (geometric_factor) {\n+\t\tif (pack_everything)\n+\t\t\tdie(_(\"--geometric is incompatible with -A, -a\"));\n+\t\tinit_pack_geometry(&geometry);\n+\t\tsplit_pack_geometry(geometry, geometric_factor);\n+\t}\n+\n \tpackdir = mkpathdup(\"%s/pack\", get_object_directory());\n \tpacktmp = mkpathdup(\"%s/.tmp-%d-pack\", packdir, (int)getpid());\n \n@@ -396,9 +525,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s\",\n \t\t\t     keep_pack_list.items[i].string);\n \tstrvec_push(&cmd.args, \"--non-empty\");\n-\tstrvec_push(&cmd.args, \"--all\");\n-\tstrvec_push(&cmd.args, \"--reflog\");\n-\tstrvec_push(&cmd.args, \"--indexed-objects\");\n+\tif (!geometry) {\n+\t\t/*\n+\t\t * 'git pack-objects' will up all objects loose or packed\n+\t\t * (either rolling them up or leaving them alone), so don't pass\n+\t\t * these options.\n+\t\t *\n+\t\t * The implementation of 'git pack-objects --stdin-packs'\n+\t\t * makes them redundant (and the two are incompatible).\n+\t\t */\n+\t\tstrvec_push(&cmd.args, \"--all\");\n+\t\tstrvec_push(&cmd.args, \"--reflog\");\n+\t\tstrvec_push(&cmd.args, \"--indexed-objects\");\n+\t}\n \tif (has_promisor_remote())\n \t\tstrvec_push(&cmd.args, \"--exclude-promisor-objects\");\n \tif (write_bitmaps > 0)\n@@ -429,17 +568,37 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t\tstrvec_push(&cmd.env_array, \"GIT_REF_PARANOIA=1\");\n \t\t\t}\n \t\t}\n+\t} else if (geometry) {\n+\t\tstrvec_push(&cmd.args, \"--stdin-packs\");\n+\t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t} else {\n \t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t\tstrvec_push(&cmd.args, \"--incremental\");\n \t}\n \n-\tcmd.no_stdin = 1;\n+\tif (geometry)\n+\t\tcmd.in = -1;\n+\telse\n+\t\tcmd.no_stdin = 1;\n \n \tret = start_command(&cmd);\n \tif (ret)\n \t\treturn ret;\n \n+\tif (geometry) {\n+\t\tFILE *in = xfdopen(cmd.in, \"w\");\n+\t\t/*\n+\t\t * The resulting pack should contain all objects in packs that\n+\t\t * are going to be rolled up, but exclude objects in packs which\n+\t\t * are being left alone.\n+\t\t */\n+\t\tfor (i = 0; i < geometry->split; i++)\n+\t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n+\t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n+\t\t\tfprintf(in, \"^%s\\n\", pack_basename(geometry->pack[i]));\n+\t\tfclose(in);\n+\t}\n+\n \tout = xfdopen(cmd.out, \"r\");\n \twhile (strbuf_getline_lf(&line, out) != EOF) {\n \t\tif (line.len != the_hash_algo->hexsz)\n@@ -507,6 +666,25 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tif (!string_list_has_string(&names, sha1))\n \t\t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n+\n+\t\tif (geometry) {\n+\t\t\tstruct strbuf buf = STRBUF_INIT;\n+\n+\t\t\tuint32_t i;\n+\t\t\tfor (i = 0; i < geometry->split; i++) {\n+\t\t\t\tstruct packed_git *p = geometry->pack[i];\n+\t\t\t\tif (string_list_has_string(&names,\n+\t\t\t\t\t\t\t   hash_to_hex(p->hash)))\n+\t\t\t\t\tcontinue;\n+\n+\t\t\t\tstrbuf_reset(&buf);\n+\t\t\t\tstrbuf_addstr(&buf, pack_basename(p));\n+\t\t\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\n+\t\t\t\tremove_redundant_pack(packdir, buf.buf);\n+\t\t\t}\n+\t\t\tstrbuf_release(&buf);\n+\t\t}\n \t\tif (!po_args.quiet && isatty(2))\n \t\t\topts |= PRUNE_PACKED_VERBOSE;\n \t\tprune_packed_objects(opts);\n@@ -528,6 +706,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstring_list_clear(&names, 0);\n \tstring_list_clear(&rollback, 0);\n \tstring_list_clear(&existing_packs, 0);\n+\tclear_pack_geometry(geometry);\n \tstrbuf_release(&line);\n \n \treturn 0;\ndiff --git a/t/t7703-repack-geometric.sh b/t/t7703-repack-geometric.sh\nnew file mode 100755\nindex 0000000000..96917fc163\n--- /dev/null\n+++ b/t/t7703-repack-geometric.sh\n@@ -0,0 +1,137 @@\n+#!/bin/sh\n+\n+test_description='git repack --geometric works correctly'\n+\n+. ./test-lib.sh\n+\n+GIT_TEST_MULTI_PACK_INDEX=0\n+\n+objdir=.git/objects\n+midx=$objdir/pack/multi-pack-index\n+\n+test_expect_success '--geometric with no packs' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\tgit repack --geometric 2 >out &&\n+\t\ttest_i18ngrep \"Nothing new to pack\" out\n+\t)\n+'\n+\n+test_expect_success '--geometric with an intact progression' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# These packs already form a geometric progression.\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 2 && # 6 objects\n+\t\ttest_commit_bulk --start=4 4 && # 12 objects\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >expect &&\n+\t\tgit repack --geometric 2 -d &&\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >actual &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success '--geometric with small-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >small &&\n+\t\ttest_commit_bulk --start=3 4 && # 12 objects\n+\t\ttest_commit_bulk --start=7 8 && # 24 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tgit repack --geometric 2 -d &&\n+\n+\t\t# Three packs in total; two of the existing large ones, and one\n+\t\t# new one.\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\ttest_line_count = 3 after &&\n+\t\tcomm -3 small before | tr -d \"\\t\" >large &&\n+\t\tgrep -qFf large after\n+\t)\n+'\n+\n+test_expect_success '--geometric with small- and large-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# size(small1) + size(small2) > size(medium) / 2\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 3 && # 7 objects\n+\t\ttest_commit_bulk --start=6 9 && # 27 objects &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tgit repack --geometric 2 -d &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\tcomm -12 before after >untouched &&\n+\n+\t\t# Two packs in total; the largest pack from before running \"git\n+\t\t# repack\", and one new one.\n+\t\ttest_line_count = 1 untouched &&\n+\t\ttest_line_count = 2 after\n+\t)\n+'\n+\n+test_expect_success '--geometric ignores kept packs' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\ttest_commit kept && # 3 objects\n+\t\ttest_commit pack && # 3 objects\n+\n+\t\tKEPT=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n+\t\trefs/tags/kept\n+\t\tEOF\n+\t\t) &&\n+\t\tPACK=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n+\t\trefs/tags/pack\n+\t\t^refs/tags/kept\n+\t\tEOF\n+\t\t) &&\n+\n+\t\t# neither pack contains more than twice the number of objects in\n+\t\t# the other, so they should be combined. but, marking one as\n+\t\t# .kept on disk will \"freeze\" it, so the pack structure should\n+\t\t# remain unchanged.\n+\t\ttouch $objdir/pack/pack-$KEPT.keep &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\t\tgit repack --geometric 2 -d &&\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\n+\t\t# both packs should still exist\n+\t\ttest_path_is_file $objdir/pack/pack-$KEPT.pack &&\n+\t\ttest_path_is_file $objdir/pack/pack-$PACK.pack &&\n+\n+\t\t# and no new packs should be created\n+\t\ttest_cmp before after &&\n+\n+\t\t# Passing --pack-kept-objects causes packs with a .keep file to\n+\t\t# be repacked, too.\n+\t\tgit repack --geometric 2 -d --pack-kept-objects &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" >after &&\n+\t\ttest_line_count = 1 after\n+\t)\n+'\n+\n+test_done\n-- \n2.30.0.667.g81c0cbc6fd\n"},{"id":"417509","messageId":"YDRM0E+GjLlXoSwC@coredump.intra.peff.net","threadId":"55012","inReplyTo":"cover.1613618042.git.me@ttaylorr.com","subject":"Re: [PATCH v3 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-23T00:31:12Z","receivedAt":"2021-02-23T00:31:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 17, 2021 at 10:14:11PM -0500, Taylor Blau wrote:\n\n> Here is another updated version of mine and Peff's series to add a new 'git\n> repack --geometric' mode which supports repacking a repository into a geometric\n> progression of packs by object count.\n\nThanks. This version looks pretty good to me. I have a few inline\ncomments below. Mostly just observations, but there a couple tiny nits\nthat I think may justify one more re-roll.\n\n> 14:  ddc2896caa !  2:  82f6b45463 revision: learn '--no-kept-objects'\n>     @@ Commit message\n>          certain packs alone (for e.g., when doing a geometric repack that has\n>          some \"large\" packs which are kept in-core that it wants to leave alone).\n> \n>     +    Note that this option is not guaranteed to produce exactly the set of\n>     +    objects that aren't in kept packs, since it's possible the traversal\n>     +    order may end up in a situation where a non-kept ancestor was \"cut off\"\n>     +    by a kept object (at which point we would stop traversing). But, we\n>     +    don't care about absolute correctness here, since this will eventually\n>     +    be used as a purely additive guide in an upcoming new repack mode.\n>     +\n>     +    Explicitly avoid documenting this new flag, since it is only used\n>     +    internally. In theory we could avoid even adding it rev-list, but being\n>     +    able to spell this option out on the command-line makes some special\n>     +    cases easier to test without promising to keep it behaving consistently\n>     +    forever. Those tricky cases are exercised in t6114.\n\nWe don't have a real procedure for marking something as \"off limits\" for\nusers. IMHO omitting it from the documentation and putting an explicit\nnote in the commit message is probably enough. It would be perhaps\nstronger to mark it explicitly as \"do not touch\" in the documentation,\nbut then we are polluting the documentation. :)\n\n>     @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n>      +\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n>      +\t\tp->pack_keep_in_core = 1;\n>      +\t}\n>     ++\n>     ++\t/*\n>     ++\t * Order packs by ascending mtime; use QSORT directly to access the\n>     ++\t * string_list_item's ->util pointer, which string_list_sort() does not\n>     ++\t * provide.\n>     ++\t */\n>     ++\tQSORT(include_packs.items, include_packs.nr, pack_mtime_cmp);\n>     ++\n\nI wondered briefly if we should accept the order from the caller, and\nmake it responsible for any sorting. But in other instances, we are\nhappy to reorder objects internally for the sake of optimization, so it\nprobably makes sense here.\n\nI also wondered if we could piggy-back on the sorting of packed_git,\nwhich is already in reverse chronological order. But here our primary\nstructure is the string-list, so we lose that order.\n\nI'm not sure if your sort function is going the right way, though. It\ndoes:\n\n>     ++static int pack_mtime_cmp(const void *_a, const void *_b)\n>     ++{\n>     ++        struct packed_git *a = ((const struct string_list_item*)_a)->util;\n>     ++        struct packed_git *b = ((const struct string_list_item*)_b)->util;\n>     ++\n>     ++        if (a->mtime < b->mtime)\n>     ++                return -1;\n>     ++        else if (b->mtime < a->mtime)\n>     ++                return 1;\n>     ++        else\n>     ++                return 0;\n>     ++}\n>     ++\n\nDoes that give us the packs in increasing chronological order, but then\ndecreasing chronological order within the packs themselves?\n\n> 17:  b5081c01b5 !  5:  181c104a03 p5303: measure time to repack with keep\n>     @@ Metadata\n>       ## Commit message ##\n>          p5303: measure time to repack with keep\n> \n>     -    This is the same as the regular repack test, except that we mark the\n>     -    single base pack as \"kept\" and use --assume-kept-packs-closed. The\n>     -    theory is that this should be faster than the normal repack, because\n>     -    we'll have fewer objects to traverse and process.\n>     +    Add two new tests to measure repack performance. Both test split the\n\ns/test split/tests split/, I think.\n\n>     +    in the 50-pack case, things start to slow down:\n>     +\n>     +      5303.11: repack (50)                        71.54(88.57+4.84)\n>     +      5303.12: repack with kept (50)              85.12(102.05+4.94)\n>     +\n>     +    and by the time we hit 1,000 packs, things are substantially worse, even\n>     +    though the resulting pack produced is the same:\n>     +\n>     +      5303.17: repack (1000)                      216.87(490.79+14.57)\n>     +      5303.18: repack with kept (1000)            665.63(938.87+15.76)\n\nOK, that's the kind of horrendous slowdown I knew we could demonstrate. :)\nI'm excited to see the numbers improve in the next patch.\n\n>     +    Likewise, the scaling is pretty extreme on --stdin-packs:\n>     +\n>     +      5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)\n>     +      5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)\n>     +      5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)\n> \n>          That's because the code paths around handling .keep files are known to\n>          scale badly; they look in every single pack file to find each object.\n\nYour \"that's because\" is a little confusing to me. It certainly applies\nto the repack vs repack-with-kept comparisons for a given number of\npacks. But the scaling on the three --stdin-packs tests is high because\neach subsequent test is being asked to do a lot more work. But they're\nstill cheaper than the matching \"repack\" case with a given number of\npacks. Just not _as_ cheap as they would be if the kept code weren't so\nslow.\n\nWould it make sense to reorder those two paragraphs?\n\n>     ++\ttest_perf \"repack with kept ($nr_packs)\" '\n>     ++\t\tgit pack-objects --keep-true-parents \\\n>     ++\t\t  --keep-pack=pack-$empty_pack.pack \\\n>     ++\t\t  --honor-pack-keep --non-empty --all \\\n>     ++\t\t  --reflog --indexed-objects --delta-base-offset \\\n>     ++\t\t  --stdout </dev/null >/dev/null\n>     ++\t'\n\nThe new test itself looks sensible. I like using --keep-pack here to\navoid needing to do any other setup/cleanup. (It does assume that\non-disk and in-core keeps behave the same, but I'm fine with that\nwhite-box assumption, especially for a perf test).\n\n>     +      5303.5: repack (1)                          57.26(54.59+10.84)      57.34(54.66+10.88) +0.1%\n>     +      5303.6: repack with kept (1)                57.33(54.80+10.51)      57.38(54.83+10.49) +0.1%\n>     +      5303.11: repack (50)                        71.54(88.57+4.84)       71.70(88.99+4.74) +0.2%\n>     +      5303.12: repack with kept (50)              85.12(102.05+4.94)      72.58(89.61+4.78) -14.7%\n>     +      5303.17: repack (1000)                      216.87(490.79+14.57)    217.19(491.72+14.25) +0.1%\n>     +      5303.18: repack with kept (1000)            665.63(938.87+15.76)    246.12(520.07+14.93) -63.0%\n\nNice. In each amount we are recovering almost all of the kept slowdown\nseen between the repack and repack-with-kept cases. The remaining\nslowdown is just from iterating that N-pack linked list, even though we\ndon't look in any of its .idx files.\n\n>     +    and the --stdin-packs timings:\n>     +\n>     +      5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)         0.00(0.00+0.00) -100.0%\n>     +      5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)        3.43(11.75+0.24) -2.8%\n>     +      5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)     130.50(307.15+7.66) -33.4%\n\nAnd of course we see an improvement here, too (as expected, but not as\ndramatic because we are doing less work overall).\n\n> 19:  f1c07324f6 !  7:  e9e04b95e7 packfile: add kept-pack cache for find_kept_pack_entry()\n> [...]\n>     +      5303.5: repack (1)                          57.34(54.66+10.88)      56.98(54.36+10.98) -0.6%\n>     +      5303.6: repack with kept (1)                57.38(54.83+10.49)      57.17(54.97+10.26) -0.4%\n>     +      5303.11: repack (50)                        71.70(88.99+4.74)       71.62(88.48+5.08) -0.1%\n>     +      5303.12: repack with kept (50)              72.58(89.61+4.78)       71.56(88.80+4.59) -1.4%\n>     +      5303.17: repack (1000)                      217.19(491.72+14.25)    217.31(490.82+14.53) +0.1%\n>     +      5303.18: repack with kept (1000)            246.12(520.07+14.93)    217.08(490.37+15.10) -11.8%\n\nAnd now we can see this patch carrying its weight much more than in the\nprevious iteration of the series. Good. Our N-pack linked list is now a\nsingle element (just the kept pack), so we expect our repack-with-kept\ntimes to match their non-kept partners. And they do.\n\n>     +    and the --stdin-packs case, which scales a little bit better (although\n>     +    not by that much even at 1,000 packs):\n>     +\n>     +      5303.7: repack with --stdin-packs (1)       0.00(0.00+0.00)         0.00(0.00+0.00) =\n>     +      5303.13: repack with --stdin-packs (50)     3.43(11.75+0.24)        3.43(11.69+0.30) +0.0%\n>     +      5303.19: repack with --stdin-packs (1000)   130.50(307.15+7.66)     125.13(301.36+8.04) -4.1%\n\nAnd likewise this is less dramatic, but still nice to see.\n\n> 20:  d5561585c2 !  8:  bd492ec142 builtin/repack.c: add '--geometric' option\n>     @@ Documentation/git-repack.txt: depth is 4095.\n> [...]\n>     ++Unlike other repack modes, the set of objects to pack is determined\n>     ++uniquely by the set of packs being \"rolled-up\"; in other words, the\n>     ++packs determined to need to be combined in order to restore a geometric\n>     ++progression.\n\nAnd this is the \"clarify roll-up\" bit I asked for. Looks good.\n\n>     ++Loose objects are implicitly included in this \"roll-up\", without respect\n>     ++to their reachability. This is subject to change in the future. This\n>     ++option (implying a drastically different repack mode) is not guarenteed\n>     ++to work with all other combinations of option to `git repack`).\n\nLikewise, this is a big improvement. But should it make it clear that\ntouching loose objects requires --unpacked? I.e., something like:\n\n  When `--unpacked` is specified, loose objects are included in this\n  \"roll-up\" without respect to their reachability...\n\nAlso, s/guarenteed/guaranteed/.\n\n-Peff\n"},{"id":"417513","messageId":"YDRVCIdwRTw4PoMR@nand.local","threadId":"55012","inReplyTo":"YDRM0E+GjLlXoSwC@coredump.intra.peff.net","subject":"Re: [PATCH v3 0/8] repack: support repacking into a geometric sequence","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T01:06:16Z","receivedAt":"2021-02-23T01:07:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Feb 22, 2021 at 07:31:12PM -0500, Jeff King wrote:\n> On Wed, Feb 17, 2021 at 10:14:11PM -0500, Taylor Blau wrote:\n>\n> > Here is another updated version of mine and Peff's series to add a new 'git\n> > repack --geometric' mode which supports repacking a repository into a geometric\n> > progression of packs by object count.\n>\n> Thanks. This version looks pretty good to me. I have a few inline\n> comments below. Mostly just observations, but there a couple tiny nits\n> that I think may justify one more re-roll.\n\nThanks for taking a look; I agree that your comments do justify a\nre-roll. But I think that one can be done without touching any of the\ncode (or maybe one line of code), depending on my question below.\n\nLet's see...\n\n> > [snip documentation]\n>\n> We don't have a real procedure for marking something as \"off limits\" for\n> users. IMHO omitting it from the documentation and putting an explicit\n> note in the commit message is probably enough. It would be perhaps\n> stronger to mark it explicitly as \"do not touch\" in the documentation,\n> but then we are polluting the documentation. :)\n\nI agree; and the second paragraph in the quoted snippet is the \"do not\ntouch\" one. So I think this one is good as-is.\n\n> I also wondered if we could piggy-back on the sorting of packed_git,\n> which is already in reverse chronological order. But here our primary\n> structure is the string-list, so we lose that order.\n>\n> I'm not sure if your sort function is going the right way, though. It\n> does:\n>\n> >     ++static int pack_mtime_cmp(const void *_a, const void *_b)\n> >     ++{\n> >     ++        struct packed_git *a = ((const struct string_list_item*)_a)->util;\n> >     ++        struct packed_git *b = ((const struct string_list_item*)_b)->util;\n> >     ++\n> >     ++        if (a->mtime < b->mtime)\n> >     ++                return -1;\n> >     ++        else if (b->mtime < a->mtime)\n> >     ++                return 1;\n> >     ++        else\n> >     ++                return 0;\n> >     ++}\n> >     ++\n>\n> Does that give us the packs in increasing chronological order, but then\n> decreasing chronological order within the packs themselves?\n\nI agree we should be sorting and not blindly accepting the order that\nthe caller gave us, but...\n\n\"chronological order within the packs themselves\" confuses me. I think\nthat you mean ordering objects within a pack by their offsets. If so,\nthen yes: this gives you the oldest pack first (and all of its objects\nin their original order), then the second oldest (and all of its\nobjects) and so on.\n\nCould you clarify a bit how you'd expect to sort the objects in two\npacks?\n\n> > 17:  b5081c01b5 !  5:  181c104a03 p5303: measure time to repack with keep\n> >     @@ Metadata\n> >       ## Commit message ##\n> >          p5303: measure time to repack with keep\n> >\n> >     -    This is the same as the regular repack test, except that we mark the\n> >     -    single base pack as \"kept\" and use --assume-kept-packs-closed. The\n> >     -    theory is that this should be faster than the normal repack, because\n> >     -    we'll have fewer objects to traverse and process.\n> >     +    Add two new tests to measure repack performance. Both test split the\n>\n> s/test split/tests split/, I think.\n\nGood eyes, thanks.\n\n> >     +    Likewise, the scaling is pretty extreme on --stdin-packs:\n> >     +\n> >     +      5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)\n> >     +      5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)\n> >     +      5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)\n> >\n> >          That's because the code paths around handling .keep files are known to\n> >          scale badly; they look in every single pack file to find each object.\n>\n> Your \"that's because\" is a little confusing to me. It certainly applies\n> to the repack vs repack-with-kept comparisons for a given number of\n> packs. But the scaling on the three --stdin-packs tests is high because\n> each subsequent test is being asked to do a lot more work. But they're\n> still cheaper than the matching \"repack\" case with a given number of\n> packs. Just not _as_ cheap as they would be if the kept code weren't so\n> slow.\n>\n> Would it make sense to reorder those two paragraphs?\n\nI think so. I did add a tiny parenthetical after my \"Likewise, the\nscaling is pretty extreme [...]\" to say \"(but each subsequent test is\nalso being asked to do more work)\".\n\n> >     ++Loose objects are implicitly included in this \"roll-up\", without respect\n> >     ++to their reachability. This is subject to change in the future. This\n> >     ++option (implying a drastically different repack mode) is not guarenteed\n> >     ++to work with all other combinations of option to `git repack`).\n>\n> Likewise, this is a big improvement. But should it make it clear that\n> touching loose objects requires --unpacked? I.e., something like:\n>\n>   When `--unpacked` is specified, loose objects are included in this\n>   \"roll-up\" without respect to their reachability...\n>\n> Also, s/guarenteed/guaranteed/.\n\nAgreed on both, thanks.\n\nThanks,\nTaylor\n"},{"id":"417519","messageId":"YDRdmh8oS5/xq4rB@coredump.intra.peff.net","threadId":"55012","inReplyTo":"YDRVCIdwRTw4PoMR@nand.local","subject":"Re: [PATCH v3 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-23T01:42:50Z","receivedAt":"2021-02-23T01:43:37Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Feb 22, 2021 at 08:06:16PM -0500, Taylor Blau wrote:\n\n> > >     ++static int pack_mtime_cmp(const void *_a, const void *_b)\n> > >     ++{\n> > >     ++        struct packed_git *a = ((const struct string_list_item*)_a)->util;\n> > >     ++        struct packed_git *b = ((const struct string_list_item*)_b)->util;\n> > >     ++\n> > >     ++        if (a->mtime < b->mtime)\n> > >     ++                return -1;\n> > >     ++        else if (b->mtime < a->mtime)\n> > >     ++                return 1;\n> > >     ++        else\n> > >     ++                return 0;\n> > >     ++}\n> > >     ++\n> >\n> > Does that give us the packs in increasing chronological order, but then\n> > decreasing chronological order within the packs themselves?\n> \n> I agree we should be sorting and not blindly accepting the order that\n> the caller gave us, but...\n> \n> \"chronological order within the packs themselves\" confuses me. I think\n> that you mean ordering objects within a pack by their offsets. If so,\n> then yes: this gives you the oldest pack first (and all of its objects\n> in their original order), then the second oldest (and all of its\n> objects) and so on.\n> \n> Could you clarify a bit how you'd expect to sort the objects in two\n> packs?\n\nYes, by \"within the packs themselves\" I meant the physical order of\nobjects within an individual pack (sorted by their offsets, as we'd get\nfrom for_each_object_in_pack). We would generally expect that to be\n\"newest first\" within a given pack (modulo some other heuristics, but we\ngenerally follow traversal order from rev-list).\n\nSo if the packs themselves are in oldest-first order, won't that create\na weird discontinuity at the pack boundaries?\n\nE.g., imagine we have a linear sequence of commits A..Z in chronological\norder, stored in two packs of equal size. Something like:\n\n  tick=1234567890\n  commit() {\n    tick=$((tick+10))\n    export GIT_COMMITTER_DATE=\"@$tick +0000\"\n    git commit --allow-empty -m $1\n  }\n\n  for i in $(perl -le 'print for A..M'); do commit $i; done\n  git repack -d\n  sleep 5\n  for i in $(perl -le 'print for N..Z'); do commit $i; done\n  git repack -d\n\nSince \"repack -d\" will use a traversal to decide which objects to pack,\nthe two packs will have their commits in reverse chronological order:\nM..A and Z..N. You can verify that with:\n\n  for idx in $(ls -rt .git/objects/pack/*.idx); do\n    stat --format='==> %y %n' $idx\n    git show-index <$idx |\n    sort -n |\n    awk '{print $2}' |\n    git --no-pager log --no-walk=unsorted --stdin --format=%s\n  done\n\nAnd if we then ran \"git repack -ad\" to make a new pack, it would be in\nnewest-to-oldest Z..A order.\n\nBut if instead we concatenate the packs after sorting them in\noldest-first order, we'll end up with a pack that contains M..A, then\nZ..N. We instead want newest packs first (and then newest objects within\nthat pack, which is the pack order), then oldest.\n\nIn other words, I think your comparison function should be reversed\n(return \"1\" when a->mtime < b->mtime).\n\n(Of course these orders aren't perfect; in a real pack you'd have\nnon-commit objects, and we'd tweak the write order to keep delta\nfamilies together, etc. But our \"best guess\" should keep packs and\nobjects-within-packs consistent in newest-first order).\n\n-Peff\n"},{"id":"417522","messageId":"bb674e511914b7070a8f985a30dd09f9a3ed58db.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 1/8] packfile: introduce 'find_kept_pack_entry()'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:03Z","receivedAt":"2021-02-23T02:26:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Future callers will want a function to fill a 'struct pack_entry' for a\ngiven object id but _only_ from its position in any kept pack(s).\n\nIn particular, an new 'git repack' mode which ensures the resulting\npacks form a geometric progress by object count will mark packs that it\ndoes not want to repack as \"kept in-core\", and it will want to halt a\nreachability traversal as soon as it visits an object in any of the kept\npacks. But, it does not want to halt the traversal at non-kept, or\n.keep packs.\n\nThe obvious alternative is 'find_pack_entry()', but this doesn't quite\nsuffice since it only returns the first pack it finds, which may or may\nnot be kept (and the mru cache makes it unpredictable which one you'll\nget if there are options).\n\nShort of that, you could walk over all packs looking for the object in\neach one, but it scales with the number of packs, which may be\nprohibitive.\n\nIntroduce 'find_kept_pack_entry()', a function which is like\n'find_pack_entry()', but only fills in objects in the kept packs.\n\nHandle packs which have .keep files, as well as in-core kept packs\nseparately, since certain callers will want to distinguish one from the\nother. (Though on-disk and in-core kept packs share the adjective\n\"kept\", it is best to think of the two sets as independent.)\n\nThere is a gotcha when looking up objects that are duplicated in kept\nand non-kept packs, particularly when the MIDX stores the non-kept\nversion and the caller asked for kept objects only. This could be\nresolved by teaching the MIDX to resolve duplicates by always favoring\nthe kept pack (if one exists), but this breaks an assumption in existing\nMIDXs, and so it would require a format change.\n\nThe benefit to changing the MIDX in this way is marginal, so we instead\nhave a more thorough check here which is explained with a comment.\n\nCallers will be added in subsequent patches.\n\nCo-authored-by: Jeff King <peff@peff.net>\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n packfile.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++-----\n packfile.h |  5 +++++\n 2 files changed, 64 insertions(+), 5 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex 1fec12ac5f..7f84f221ce 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2042,7 +2042,10 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static int find_one_pack_entry(struct repository *r,\n+\t\t\t       const struct object_id *oid,\n+\t\t\t       struct pack_entry *e,\n+\t\t\t       int kept_only)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2052,26 +2055,77 @@ int find_pack_entry(struct repository *r, const struct object_id *oid, struct pa\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (fill_midx_entry(r, oid, e, m))\n+\t\tif (!fill_midx_entry(r, oid, e, m))\n+\t\t\tcontinue;\n+\n+\t\tif (!kept_only)\n+\t\t\treturn 1;\n+\n+\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n+\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n-\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\treturn 1;\n+\t\tif (p->multi_pack_index && !kept_only) {\n+\t\t\t/*\n+\t\t\t * If this pack is covered by the MIDX, we'd have found\n+\t\t\t * the object already in the loop above if it was here,\n+\t\t\t * so don't bother looking.\n+\t\t\t *\n+\t\t\t * The exception is if we are looking only at kept\n+\t\t\t * packs. An object can be present in two packs covered\n+\t\t\t * by the MIDX, one kept and one not-kept. And as the\n+\t\t\t * MIDX points to only one copy of each object, it might\n+\t\t\t * have returned only the non-kept version above. We\n+\t\t\t * have to check again to be thorough.\n+\t\t\t */\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!kept_only ||\n+\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n+\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n+\t\t\tif (fill_pack_entry(oid, e, p)) {\n+\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\t\treturn 1;\n+\t\t\t}\n \t\t}\n \t}\n \treturn 0;\n }\n \n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+{\n+\treturn find_one_pack_entry(r, oid, e, 0);\n+}\n+\n+int find_kept_pack_entry(struct repository *r,\n+\t\t\t const struct object_id *oid,\n+\t\t\t unsigned flags,\n+\t\t\t struct pack_entry *e)\n+{\n+\t/*\n+\t * Load all packs, including midx packs, since our \"kept\" strategy\n+\t * relies on that. We're relying on the side effect of it setting up\n+\t * r->objects->packed_git, which is a little ugly.\n+\t */\n+\tget_all_packs(r);\n+\treturn find_one_pack_entry(r, oid, e, flags);\n+}\n+\n int has_object_pack(const struct object_id *oid)\n {\n \tstruct pack_entry e;\n \treturn find_pack_entry(the_repository, oid, &e);\n }\n \n+int has_object_kept_pack(const struct object_id *oid, unsigned flags)\n+{\n+\tstruct pack_entry e;\n+\treturn find_kept_pack_entry(the_repository, oid, flags, &e);\n+}\n+\n int has_pack_index(const unsigned char *sha1)\n {\n \tstruct stat st;\ndiff --git a/packfile.h b/packfile.h\nindex 4cfec9e8d3..3ae117a8ae 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -162,13 +162,18 @@ int packed_object_info(struct repository *r,\n void mark_bad_packed_object(struct packed_git *p, const unsigned char *sha1);\n const struct packed_git *has_packed_and_bad(struct repository *r, const unsigned char *sha1);\n \n+#define ON_DISK_KEEP_PACKS 1\n+#define IN_CORE_KEEP_PACKS 2\n+\n /*\n  * Iff a pack file in the given repository contains the object named by sha1,\n  * return true and store its location to e.\n  */\n int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e);\n+int find_kept_pack_entry(struct repository *r, const struct object_id *oid, unsigned flags, struct pack_entry *e);\n \n int has_object_pack(const struct object_id *oid);\n+int has_object_kept_pack(const struct object_id *oid, unsigned flags);\n \n int has_pack_index(const unsigned char *sha1);\n \n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417523","messageId":"cover.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1611098616.git.me@ttaylorr.com","subject":"[PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:24:59Z","receivedAt":"2021-02-23T02:26:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Here's a very lightly modified version on v3 of mine and Peff's series\nto add a new 'git repack --geometric' mode. Almost nothing has changed\nsince last time, with the exception of:\n\n  - Packs listed over standard input to 'git pack-objects --stdin-packs'\n    are sorted in descending mtime order (and objects are strung\n    together in pack order as before) so that objects are laid out\n    roughly newest-to-oldest in the resulting pack.\n\n  - Swapped the order of two paragraphs in patch 5 to make the perf\n    results clearer.\n\n  - Mention '--unpacked' specifically in the documentation for 'git\n    repack --geometric'.\n\n  - Typo fixes.\n\nRange-diff is below. It would be good to start merging this down since\nwe have a release candidate coming up soon, and I'd rather focus future\nreviewer efforts on the multi-pack reverse index and bitmaps series\ninstead of this one.\n\nJeff King (4):\n  p5303: add missing &&-chains\n  p5303: measure time to repack with keep\n  builtin/pack-objects.c: rewrite honor-pack-keep logic\n  packfile: add kept-pack cache for find_kept_pack_entry()\n\nTaylor Blau (4):\n  packfile: introduce 'find_kept_pack_entry()'\n  revision: learn '--no-kept-objects'\n  builtin/pack-objects.c: add '--stdin-packs' option\n  builtin/repack.c: add '--geometric' option\n\n Documentation/git-pack-objects.txt |  10 +\n Documentation/git-repack.txt       |  23 ++\n builtin/pack-objects.c             | 333 ++++++++++++++++++++++++-----\n builtin/repack.c                   | 187 +++++++++++++++-\n object-store.h                     |   5 +\n packfile.c                         |  67 ++++++\n packfile.h                         |   5 +\n revision.c                         |  15 ++\n revision.h                         |   4 +\n t/perf/p5303-many-packs.sh         |  36 +++-\n t/t5300-pack-object.sh             |  97 +++++++++\n t/t6114-keep-packs.sh              |  69 ++++++\n t/t7703-repack-geometric.sh        | 137 ++++++++++++\n 13 files changed, 926 insertions(+), 62 deletions(-)\n create mode 100755 t/t6114-keep-packs.sh\n create mode 100755 t/t7703-repack-geometric.sh\n\nRange-diff against v3:\n1:  aa94edf39b = 1:  bb674e5119 packfile: introduce 'find_kept_pack_entry()'\n2:  82f6b45463 = 2:  c85a915597 revision: learn '--no-kept-objects'\n3:  033e4e3f67 ! 3:  649cf9020b builtin/pack-objects.c: add '--stdin-packs' option\n    @@ builtin/pack-objects.c: static int git_pack_config(const char *k, const char *v,\n     +\tstruct packed_git *a = ((const struct string_list_item*)_a)->util;\n     +\tstruct packed_git *b = ((const struct string_list_item*)_b)->util;\n     +\n    ++\t/*\n    ++\t * order packs by descending mtime so that objects are laid out\n    ++\t * roughly as newest-to-oldest\n    ++\t */\n     +\tif (a->mtime < b->mtime)\n    -+\t\treturn -1;\n    -+\telse if (b->mtime < a->mtime)\n     +\t\treturn 1;\n    ++\telse if (b->mtime < a->mtime)\n    ++\t\treturn -1;\n     +\telse\n     +\t\treturn 0;\n     +}\n4:  f9a5faf773 = 4:  6de9f0c52b p5303: add missing &&-chains\n5:  181c104a03 ! 5:  94e4f3ee3a p5303: measure time to repack with keep\n    @@ Metadata\n      ## Commit message ##\n         p5303: measure time to repack with keep\n     \n    -    Add two new tests to measure repack performance. Both test split the\n    +    Add two new tests to measure repack performance. Both tests split the\n         repository into synthetic \"pushes\", and then leave the remaining objects\n         in a big base pack.\n     \n    @@ Commit message\n           5303.17: repack (1000)                      216.87(490.79+14.57)\n           5303.18: repack with kept (1000)            665.63(938.87+15.76)\n     \n    -    Likewise, the scaling is pretty extreme on --stdin-packs:\n    -\n    -      5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)\n    -      5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)\n    -      5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)\n    -\n         That's because the code paths around handling .keep files are known to\n         scale badly; they look in every single pack file to find each object.\n         Our solution to that was to notice that most repos don't have keep\n    @@ Commit message\n         single .keep, that part of pack-objects slows down again (even if we\n         have fewer objects total to look at).\n     \n    +    Likewise, the scaling is pretty extreme on --stdin-packs (but each\n    +    subsequent test is also being asked to do more work):\n    +\n    +      5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)\n    +      5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)\n    +      5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)\n    +\n         Signed-off-by: Jeff King <peff@peff.net>\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n6:  67af143fd1 = 6:  a116587fb2 builtin/pack-objects.c: rewrite honor-pack-keep logic\n7:  e9e04b95e7 = 7:  db9f07ec1a packfile: add kept-pack cache for find_kept_pack_entry()\n8:  bd492ec142 ! 8:  51f57d5da2 builtin/repack.c: add '--geometric' option\n    @@ Documentation/git-repack.txt: depth is 4095.\n     +packs determined to need to be combined in order to restore a geometric\n     +progression.\n     ++\n    -+Loose objects are implicitly included in this \"roll-up\", without respect\n    -+to their reachability. This is subject to change in the future. This\n    -+option (implying a drastically different repack mode) is not guarenteed\n    -+to work with all other combinations of option to `git repack`).\n    ++When `--unpacked` is specified, loose objects are implicitly included in\n    ++this \"roll-up\", without respect to their reachability. This is subject\n    ++to change in the future. This option (implying a drastically different\n    ++repack mode) is not guaranteed to work with all other combinations of\n    ++option to `git repack`).\n     +\n      Configuration\n      -------------\n-- \n2.30.0.667.g81c0cbc6fd\n"},{"id":"417524","messageId":"c85a9155970aa3f59990d3678da21bb88a9cac08.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 2/8] revision: learn '--no-kept-objects'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:07Z","receivedAt":"2021-02-23T02:26:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A future caller will want to be able to perform a reachability traversal\nwhich terminates when visiting an object found in a kept pack. The\nclosest existing option is '--honor-pack-keep', but this isn't quite\nwhat we want. Instead of halting the traversal midway through, a full\ntraversal is always performed, and the results are only trimmed\nafterwords.\n\nBesides needing to introduce a new flag (since culling results\npost-facto can be different than halting the traversal as it's\nhappening), there is an additional wrinkle handling the distinction\nin-core and on-disk kept packs. That is: what kinds of kept pack should\nstop the traversal?\n\nIntroduce '--no-kept-objects[=<on-disk|in-core>]' to specify which kinds\nof kept packs, if any, should stop a traversal. This can be useful for\ncallers that want to perform a reachability analysis, but want to leave\ncertain packs alone (for e.g., when doing a geometric repack that has\nsome \"large\" packs which are kept in-core that it wants to leave alone).\n\nNote that this option is not guaranteed to produce exactly the set of\nobjects that aren't in kept packs, since it's possible the traversal\norder may end up in a situation where a non-kept ancestor was \"cut off\"\nby a kept object (at which point we would stop traversing). But, we\ndon't care about absolute correctness here, since this will eventually\nbe used as a purely additive guide in an upcoming new repack mode.\n\nExplicitly avoid documenting this new flag, since it is only used\ninternally. In theory we could avoid even adding it rev-list, but being\nable to spell this option out on the command-line makes some special\ncases easier to test without promising to keep it behaving consistently\nforever. Those tricky cases are exercised in t6114.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n revision.c            | 15 ++++++++++\n revision.h            |  4 +++\n t/t6114-keep-packs.sh | 69 +++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 88 insertions(+)\n create mode 100755 t/t6114-keep-packs.sh\n\ndiff --git a/revision.c b/revision.c\nindex b78733f508..dca2b8c801 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -2336,6 +2336,16 @@ static int handle_revision_opt(struct rev_info *revs, int argc, const char **arg\n \t\trevs->unpacked = 1;\n \t} else if (starts_with(arg, \"--unpacked=\")) {\n \t\tdie(_(\"--unpacked=<packfile> no longer supported\"));\n+\t} else if (!strcmp(arg, \"--no-kept-objects\")) {\n+\t\trevs->no_kept_objects = 1;\n+\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n+\t} else if (skip_prefix(arg, \"--no-kept-objects=\", &optarg)) {\n+\t\trevs->no_kept_objects = 1;\n+\t\tif (!strcmp(optarg, \"in-core\"))\n+\t\t\trevs->keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\t\tif (!strcmp(optarg, \"on-disk\"))\n+\t\t\trevs->keep_pack_cache_flags |= ON_DISK_KEEP_PACKS;\n \t} else if (!strcmp(arg, \"-r\")) {\n \t\trevs->diff = 1;\n \t\trevs->diffopt.flags.recursive = 1;\n@@ -3795,6 +3805,11 @@ enum commit_action get_commit_action(struct rev_info *revs, struct commit *commi\n \t\treturn commit_ignore;\n \tif (revs->unpacked && has_object_pack(&commit->object.oid))\n \t\treturn commit_ignore;\n+\tif (revs->no_kept_objects) {\n+\t\tif (has_object_kept_pack(&commit->object.oid,\n+\t\t\t\t\t revs->keep_pack_cache_flags))\n+\t\t\treturn commit_ignore;\n+\t}\n \tif (commit->object.flags & UNINTERESTING)\n \t\treturn commit_ignore;\n \tif (revs->line_level_traverse && !want_ancestry(revs)) {\ndiff --git a/revision.h b/revision.h\nindex e6be3c845e..a20a530d52 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -148,6 +148,7 @@ struct rev_info {\n \t\t\tedge_hint_aggressive:1,\n \t\t\tlimited:1,\n \t\t\tunpacked:1,\n+\t\t\tno_kept_objects:1,\n \t\t\tboundary:2,\n \t\t\tcount:1,\n \t\t\tleft_right:1,\n@@ -317,6 +318,9 @@ struct rev_info {\n \t * This is loaded from the commit-graph being used.\n \t */\n \tstruct bloom_filter_settings *bloom_filter_settings;\n+\n+\t/* misc. flags related to '--no-kept-objects' */\n+\tunsigned keep_pack_cache_flags;\n };\n \n int ref_excluded(struct string_list *, const char *path);\ndiff --git a/t/t6114-keep-packs.sh b/t/t6114-keep-packs.sh\nnew file mode 100755\nindex 0000000000..9239d8aa46\n--- /dev/null\n+++ b/t/t6114-keep-packs.sh\n@@ -0,0 +1,69 @@\n+#!/bin/sh\n+\n+test_description='rev-list with .keep packs'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup' '\n+\ttest_commit loose &&\n+\ttest_commit packed &&\n+\ttest_commit kept &&\n+\n+\tKEPT_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/kept\n+\t^refs/tags/packed\n+\tEOF\n+\t) &&\n+\tMISC_PACK=$(git pack-objects --revs .git/objects/pack/pack <<-EOF\n+\trefs/tags/packed\n+\t^refs/tags/loose\n+\tEOF\n+\t) &&\n+\n+\ttouch .git/objects/pack/pack-$KEPT_PACK.keep\n+'\n+\n+rev_list_objects () {\n+\tgit rev-list \"$@\" >out &&\n+\tsort out\n+}\n+\n+idx_objects () {\n+\tgit show-index <$1 >expect-idx &&\n+\tcut -d\" \" -f2 <expect-idx | sort\n+}\n+\n+test_expect_success '--no-kept-objects excludes trees and blobs in .keep packs' '\n+\trev_list_objects --objects --all --no-object-names >kept &&\n+\trev_list_objects --objects --all --no-object-names --no-kept-objects >no-kept &&\n+\n+\tidx_objects .git/objects/pack/pack-$KEPT_PACK.idx >expect &&\n+\tcomm -3 kept no-kept >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success '--no-kept-objects excludes kept non-MIDX object' '\n+\ttest_config core.multiPackIndex true &&\n+\n+\t# Create a pack with just the commit object in pack, and do not mark it\n+\t# as kept (even though it appears in $KEPT_PACK, which does have a .keep\n+\t# file).\n+\tMIDX_PACK=$(git pack-objects .git/objects/pack/pack <<-EOF\n+\t$(git rev-parse kept)\n+\tEOF\n+\t) &&\n+\n+\t# Write a MIDX containing all packs, but use the version of the commit\n+\t# at \"kept\" in a non-kept pack by touching $MIDX_PACK.\n+\ttouch .git/objects/pack/pack-$MIDX_PACK.pack &&\n+\tgit multi-pack-index write &&\n+\n+\trev_list_objects --objects --no-object-names --no-kept-objects HEAD >actual &&\n+\t(\n+\t\tidx_objects .git/objects/pack/pack-$MISC_PACK.idx &&\n+\t\tgit rev-list --objects --no-object-names refs/tags/loose\n+\t) | sort >expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417525","messageId":"649cf9020bfdae9f48e3efbfbc52429cefd31432.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:10Z","receivedAt":"2021-02-23T02:26:21Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In an upcoming commit, 'git repack' will want to create a pack comprised\nof all of the objects in some packs (the included packs) excluding any\nobjects in some other packs (the excluded packs).\n\nThis caller could iterate those packs themselves and feed the objects it\nfinds to 'git pack-objects' directly over stdin, but this approach has a\nfew downsides:\n\n  - It requires every caller that wants to drive 'git pack-objects' in\n    this way to implement pack iteration themselves. This forces the\n    caller to think about details like what order objects are fed to\n    pack-objects, which callers would likely rather not do.\n\n  - If the set of objects in included packs is large, it requires\n    sending a lot of data over a pipe, which is inefficient.\n\n  - The caller is forced to keep track of the excluded objects, too, and\n    make sure that it doesn't send any objects that appear in both\n    included and excluded packs.\n\nBut the biggest downside is the lack of a reachability traversal.\nBecause the caller passes in a list of objects directly, those objects\ndon't get a namehash assigned to them, which can have a negative impact\non the delta selection process, causing 'git pack-objects' to fail to\nfind good deltas even when they exist.\n\nThe caller could formulate a reachability traversal themselves, but the\nonly way to drive 'git pack-objects' in this way is to do a full\ntraversal, and then remove objects in the excluded packs after the\ntraversal is complete. This can be detrimental to callers who care\nabout performance, especially in repositories with many objects.\n\nIntroduce 'git pack-objects --stdin-packs' which remedies these four\nconcerns.\n\n'git pack-objects --stdin-packs' expects a list of pack names on stdin,\nwhere 'pack-xyz.pack' denotes that pack as included, and\n'^pack-xyz.pack' denotes it as excluded. The resulting pack includes all\nobjects that are present in at least one included pack, and aren't\npresent in any excluded pack.\n\nTo address the delta selection problem, 'git pack-objects --stdin-packs'\nworks as follows. First, it assembles a list of objects that it is going\nto pack, as above. Then, a reachability traversal is started, whose tips\nare any commits mentioned in included packs. Upon visiting an object, we\nfind its corresponding object_entry in the to_pack list, and set its\nnamehash parameter appropriately.\n\nTo avoid the traversal visiting more objects than it needs to, the\ntraversal is halted upon encountering an object which can be found in an\nexcluded pack (by marking the excluded packs as kept in-core, and\npassing --no-kept-objects=in-core to the revision machinery).\n\nThis can cause the traversal to halt early, for example if an object in\nan included pack is an ancestor of ones in excluded packs. But stopping\nearly is OK, since filling in the namehash fields of objects in the\nto_pack list is only additive (i.e., having it helps the delta selection\nprocess, but leaving it blank doesn't impact the correctness of the\nresulting pack).\n\nEven still, it is unlikely that this hurts us much in practice, since\nthe 'git repack --geometric' caller (which is introduced in a later\ncommit) marks small packs as included, and large ones as excluded.\nDuring ordinary use, the small packs usually represent pushes after a\nlarge repack, and so are unlikely to be ancestors of objects that\nalready exist in the repository.\n\n(I found it convenient while developing this patch to have 'git\npack-objects' report the number of objects which were visited and got\ntheir namehash fields filled in during traversal. This is also included\nin the below patch via trace2 data lines).\n\nSuggested-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  10 ++\n builtin/pack-objects.c             | 202 ++++++++++++++++++++++++++++-\n t/t5300-pack-object.sh             |  97 ++++++++++++++\n 3 files changed, 307 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex 54d715ead1..df533c3b19 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -85,6 +85,16 @@ base-name::\n \treference was included in the resulting packfile.  This\n \tcan be useful to send new tags to native Git clients.\n \n+--stdin-packs::\n+\tRead the basenames of packfiles (e.g., `pack-1234abcd.pack`)\n+\tfrom the standard input, instead of object names or revision\n+\targuments. The resulting pack contains all objects listed in the\n+\tincluded packs (those not beginning with `^`), excluding any\n+\tobjects listed in the excluded packs (beginning with `^`).\n++\n+Incompatible with `--revs`, or options that imply `--revs` (such as\n+`--all`), with the exception of `--unpacked`, which is compatible.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 6d62aaf59a..6ee8e40665 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -2986,6 +2986,190 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \treturn git_default_config(k, v, cb);\n }\n \n+/* Counters for trace2 output when in --stdin-packs mode. */\n+static int stdin_packs_found_nr;\n+static int stdin_packs_hints_nr;\n+\n+static int add_object_entry_from_pack(const struct object_id *oid,\n+\t\t\t\t      struct packed_git *p,\n+\t\t\t\t      uint32_t pos,\n+\t\t\t\t      void *_data)\n+{\n+\tstruct rev_info *revs = _data;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\toff_t ofs;\n+\tenum object_type type;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tif (have_duplicate_entry(oid, 0))\n+\t\treturn 0;\n+\n+\tofs = nth_packed_object_offset(p, pos);\n+\tif (!want_object_in_pack(oid, 0, &p, &ofs))\n+\t\treturn 0;\n+\n+\toi.typep = &type;\n+\tif (packed_object_info(the_repository, p, ofs, &oi) < 0)\n+\t\tdie(_(\"could not get type of object %s in pack %s\"),\n+\t\t    oid_to_hex(oid), p->pack_name);\n+\telse if (type == OBJ_COMMIT) {\n+\t\t/*\n+\t\t * commits in included packs are used as starting points for the\n+\t\t * subsequent revision walk\n+\t\t */\n+\t\tadd_pending_oid(revs, NULL, oid, 0);\n+\t}\n+\n+\tstdin_packs_found_nr++;\n+\n+\tcreate_object_entry(oid, type, 0, 0, 0, p, ofs);\n+\n+\treturn 0;\n+}\n+\n+static void show_commit_pack_hint(struct commit *commit, void *_data)\n+{\n+\t/* nothing to do; commits don't have a namehash */\n+}\n+\n+static void show_object_pack_hint(struct object *object, const char *name,\n+\t\t\t\t  void *_data)\n+{\n+\tstruct object_entry *oe = packlist_find(&to_pack, &object->oid);\n+\tif (!oe)\n+\t\treturn;\n+\n+\t/*\n+\t * Our 'to_pack' list was constructed by iterating all objects packed in\n+\t * included packs, and so doesn't have a non-zero hash field that you\n+\t * would typically pick up during a reachability traversal.\n+\t *\n+\t * Make a best-effort attempt to fill in the ->hash and ->no_try_delta\n+\t * here using a now in order to perhaps improve the delta selection\n+\t * process.\n+\t */\n+\toe->hash = pack_name_hash(name);\n+\toe->no_try_delta = name && no_try_delta(name);\n+\n+\tstdin_packs_hints_nr++;\n+}\n+\n+static int pack_mtime_cmp(const void *_a, const void *_b)\n+{\n+\tstruct packed_git *a = ((const struct string_list_item*)_a)->util;\n+\tstruct packed_git *b = ((const struct string_list_item*)_b)->util;\n+\n+\t/*\n+\t * order packs by descending mtime so that objects are laid out\n+\t * roughly as newest-to-oldest\n+\t */\n+\tif (a->mtime < b->mtime)\n+\t\treturn 1;\n+\telse if (b->mtime < a->mtime)\n+\t\treturn -1;\n+\telse\n+\t\treturn 0;\n+}\n+\n+static void read_packs_list_from_stdin(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list include_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list exclude_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list_item *item = NULL;\n+\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\t/*\n+\t * Use a revision walk to fill in the namehash of objects in the include\n+\t * packs. To save time, we'll avoid traversing through objects that are\n+\t * in excluded packs.\n+\t *\n+\t * That may cause us to avoid populating all of the namehash fields of\n+\t * all included objects, but our goal is best-effort, since this is only\n+\t * an optimization during delta selection.\n+\t */\n+\trevs.no_kept_objects = 1;\n+\trevs.keep_pack_cache_flags |= IN_CORE_KEEP_PACKS;\n+\trevs.blob_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.tag_objects = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '^')\n+\t\t\tstring_list_append(&exclude_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&include_packs, buf.buf);\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&include_packs);\n+\tstring_list_sort(&exclude_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\n+\t\titem = string_list_lookup(&include_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&exclude_packs, pack_name);\n+\n+\t\tif (item)\n+\t\t\titem->util = p;\n+\t}\n+\n+\t/*\n+\t * First handle all of the excluded packs, marking them as kept in-core\n+\t * so that later calls to add_object_entry() discards any objects that\n+\t * are also found in excluded packs.\n+\t */\n+\tfor_each_string_list_item(item, &exclude_packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = 1;\n+\t}\n+\n+\t/*\n+\t * Order packs by ascending mtime; use QSORT directly to access the\n+\t * string_list_item's ->util pointer, which string_list_sort() does not\n+\t * provide.\n+\t */\n+\tQSORT(include_packs.items, include_packs.nr, pack_mtime_cmp);\n+\n+\tfor_each_string_list_item(item, &include_packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tfor_each_object_in_pack(p,\n+\t\t\t\t\tadd_object_entry_from_pack,\n+\t\t\t\t\t&revs,\n+\t\t\t\t\tFOR_EACH_OBJECT_PACK_ORDER);\n+\t}\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\ttraverse_commit_list(&revs,\n+\t\t\t     show_commit_pack_hint,\n+\t\t\t     show_object_pack_hint,\n+\t\t\t     NULL);\n+\n+\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_found\",\n+\t\t\t   stdin_packs_found_nr);\n+\ttrace2_data_intmax(\"pack-objects\", the_repository, \"stdin_packs_hints\",\n+\t\t\t   stdin_packs_hints_nr);\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&include_packs, 0);\n+\tstring_list_clear(&exclude_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3489,6 +3673,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tstruct strvec rp = STRVEC_INIT;\n \tint rev_list_unpacked = 0, rev_list_all = 0, rev_list_reflog = 0;\n \tint rev_list_index = 0;\n+\tint stdin_packs = 0;\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tstruct option pack_objects_options[] = {\n \t\tOPT_SET_INT('q', \"quiet\", &progress,\n@@ -3539,6 +3724,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_SET_INT_F(0, \"indexed-objects\", &rev_list_index,\n \t\t\t      N_(\"include objects referred to by the index\"),\n \t\t\t      1, PARSE_OPT_NONEG),\n+\t\tOPT_BOOL(0, \"stdin-packs\", &stdin_packs,\n+\t\t\t N_(\"read packs from stdin\")),\n \t\tOPT_BOOL(0, \"stdout\", &pack_to_stdout,\n \t\t\t N_(\"output pack to stdout\")),\n \t\tOPT_BOOL(0, \"include-tag\", &include_tag,\n@@ -3645,7 +3832,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tuse_internal_rev_list = 1;\n \t\tstrvec_push(&rp, \"--indexed-objects\");\n \t}\n-\tif (rev_list_unpacked) {\n+\tif (rev_list_unpacked && !stdin_packs) {\n \t\tuse_internal_rev_list = 1;\n \t\tstrvec_push(&rp, \"--unpacked\");\n \t}\n@@ -3690,8 +3877,13 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (filter_options.choice) {\n \t\tif (!pack_to_stdout)\n \t\t\tdie(_(\"cannot use --filter without --stdout\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --filter with --stdin-packs\"));\n \t}\n \n+\tif (stdin_packs && use_internal_rev_list)\n+\t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -3750,7 +3942,13 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (progress)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n-\tif (!use_internal_rev_list)\n+\tif (stdin_packs) {\n+\t\t/* avoids adding objects in excluded packs */\n+\t\tignore_packed_keep_in_core = 1;\n+\t\tread_packs_list_from_stdin();\n+\t\tif (rev_list_unpacked)\n+\t\t\tadd_unreachable_loose_objects();\n+\t} else if (!use_internal_rev_list)\n \t\tread_object_list_from_stdin();\n \telse {\n \t\tget_object_list(rp.nr, rp.v);\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex 392201cabd..7138a54595 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -532,4 +532,101 @@ test_expect_success 'prefetch objects' '\n \ttest_line_count = 1 donelines\n '\n \n+test_expect_success 'setup for --stdin-packs tests' '\n+\tgit init stdin-packs &&\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\ttest_commit A &&\n+\t\ttest_commit B &&\n+\t\ttest_commit C &&\n+\n+\t\tfor id in A B C\n+\t\tdo\n+\t\t\tgit pack-objects .git/objects/pack/pack-$id \\\n+\t\t\t\t--incremental --revs <<-EOF\n+\t\t\trefs/tags/$id\n+\t\t\tEOF\n+\t\tdone &&\n+\n+\t\tls -la .git/objects/pack\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs with excluded packs' '\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\tPACK_A=\"$(basename .git/objects/pack/pack-A-*.pack)\" &&\n+\t\tPACK_B=\"$(basename .git/objects/pack/pack-B-*.pack)\" &&\n+\t\tPACK_C=\"$(basename .git/objects/pack/pack-C-*.pack)\" &&\n+\n+\t\tgit pack-objects test --stdin-packs <<-EOF &&\n+\t\t$PACK_A\n+\t\t^$PACK_B\n+\t\t$PACK_C\n+\t\tEOF\n+\n+\t\t(\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-A-*.idx) &&\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-C-*.idx)\n+\t\t) >expect.raw &&\n+\t\tgit show-index <$(ls test-*.idx) >actual.raw &&\n+\n+\t\tcut -d\" \" -f2 <expect.raw | sort >expect &&\n+\t\tcut -d\" \" -f2 <actual.raw | sort >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs is incompatible with --filter' '\n+\t(\n+\t\tcd stdin-packs &&\n+\t\ttest_must_fail git pack-objects --stdin-packs --stdout \\\n+\t\t\t--filter=blob:none </dev/null 2>err &&\n+\t\ttest_i18ngrep \"cannot use --filter with --stdin-packs\" err\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs is incompatible with --revs' '\n+\t(\n+\t\tcd stdin-packs &&\n+\t\ttest_must_fail git pack-objects --stdin-packs --revs out \\\n+\t\t\t</dev/null 2>err &&\n+\t\ttest_i18ngrep \"cannot use internal rev list with --stdin-packs\" err\n+\t)\n+'\n+\n+test_expect_success '--stdin-packs with loose objects' '\n+\t(\n+\t\tcd stdin-packs &&\n+\n+\t\tPACK_A=\"$(basename .git/objects/pack/pack-A-*.pack)\" &&\n+\t\tPACK_B=\"$(basename .git/objects/pack/pack-B-*.pack)\" &&\n+\t\tPACK_C=\"$(basename .git/objects/pack/pack-C-*.pack)\" &&\n+\n+\t\ttest_commit D && # loose\n+\n+\t\tgit pack-objects test2 --stdin-packs --unpacked <<-EOF &&\n+\t\t$PACK_A\n+\t\t^$PACK_B\n+\t\t$PACK_C\n+\t\tEOF\n+\n+\t\t(\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-A-*.idx) &&\n+\t\t\tgit show-index <$(ls .git/objects/pack/pack-C-*.idx) &&\n+\t\t\tgit rev-list --objects --no-object-names \\\n+\t\t\t\trefs/tags/C..refs/tags/D\n+\n+\t\t) >expect.raw &&\n+\t\tls -la . &&\n+\t\tgit show-index <$(ls test2-*.idx) >actual.raw &&\n+\n+\t\tcut -d\" \" -f2 <expect.raw | sort >expect &&\n+\t\tcut -d\" \" -f2 <actual.raw | sort >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_done\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417526","messageId":"6de9f0c52bbde21b478fa5e45743b7e687001744.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 4/8] p5303: add missing &&-chains","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:13Z","receivedAt":"2021-02-23T02:26:24Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nThese are in a helper function, so the usual chain-lint doesn't notice\nthem. This function is still not perfect, as it has some git invocations\non the left-hand-side of the pipe, but it's primary purpose is timing,\nnot finding bugs or correctness issues.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex ce0c42cc9f..d90d714923 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -28,11 +28,11 @@ repack_into_n () {\n \t\t\tpush @commits, $_ if $. % 5 == 1;\n \t\t}\n \t\tprint reverse @commits;\n-\t' \"$1\" >pushes\n+\t' \"$1\" >pushes &&\n \n \t# create base packfile\n \thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack\n+\tgit pack-objects --delta-base-offset --revs staging/pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417527","messageId":"a116587fb2b7f6484b9206de68ff66d10bb2a2a2.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 6/8] builtin/pack-objects.c: rewrite honor-pack-keep logic","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:20Z","receivedAt":"2021-02-23T02:26:26Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nNow that we have find_kept_pack_entry(), we don't have to manually keep\nhunting through every pack to find a possible \"kept\" duplicate of the\nobject. This should be faster, assuming only a portion of your total\npacks are actually kept.\n\nNote that we have to re-order the logic a bit here; we can deal with the\ndisqualifying situations first (e.g., finding the object in a non-local\npack with --local), then \"kept\" situation(s), and then just fall back to\nother \"--local\" conditions.\n\nHere are the results from p5303 (measurements again taken on the\nkernel):\n\n  Test                                        HEAD^                   HEAD\n  -----------------------------------------------------------------------------------------------\n  5303.5: repack (1)                          57.26(54.59+10.84)      57.34(54.66+10.88) +0.1%\n  5303.6: repack with kept (1)                57.33(54.80+10.51)      57.38(54.83+10.49) +0.1%\n  5303.11: repack (50)                        71.54(88.57+4.84)       71.70(88.99+4.74) +0.2%\n  5303.12: repack with kept (50)              85.12(102.05+4.94)      72.58(89.61+4.78) -14.7%\n  5303.17: repack (1000)                      216.87(490.79+14.57)    217.19(491.72+14.25) +0.1%\n  5303.18: repack with kept (1000)            665.63(938.87+15.76)    246.12(520.07+14.93) -63.0%\n\nand the --stdin-packs timings:\n\n  5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)         0.00(0.00+0.00) -100.0%\n  5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)        3.43(11.75+0.24) -2.8%\n  5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)     130.50(307.15+7.66) -33.4%\n\nSo our repack with an empty .keep pack is roughly as fast as one without\na .keep pack up to 50 packs. But the --stdin-packs case scales a little\nbetter, too.\n\nNotably, it is faster than a repack of the same size and a kept pack. It\nlooks at fewer objects, of course, but the penalty for looking at many\npacks isn't as costly.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 131 ++++++++++++++++++++++++-----------------\n 1 file changed, 78 insertions(+), 53 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 6ee8e40665..8cb32763b7 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1188,7 +1188,8 @@ static int have_duplicate_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int want_found_object(int exclude, struct packed_git *p)\n+static int want_found_object(const struct object_id *oid, int exclude,\n+\t\t\t     struct packed_git *p)\n {\n \tif (exclude)\n \t\treturn 1;\n@@ -1204,27 +1205,82 @@ static int want_found_object(int exclude, struct packed_git *p)\n \t * make sure no copy of this object appears in _any_ pack that makes us\n \t * to omit the object, so we need to check all the packs.\n \t *\n-\t * We can however first check whether these options can possible matter;\n+\t * We can however first check whether these options can possibly matter;\n \t * if they do not matter we know we want the object in generated pack.\n \t * Otherwise, we signal \"-1\" at the end to tell the caller that we do\n \t * not know either way, and it needs to check more packs.\n \t */\n-\tif (!ignore_packed_keep_on_disk &&\n-\t    !ignore_packed_keep_in_core &&\n-\t    (!local || !have_non_local_packs))\n-\t\treturn 1;\n \n+\t/*\n+\t * Objects in packs borrowed from elsewhere are discarded regardless of\n+\t * if they appear in other packs that weren't borrowed.\n+\t */\n \tif (local && !p->pack_local)\n \t\treturn 0;\n-\tif (p->pack_local &&\n-\t    ((ignore_packed_keep_on_disk && p->pack_keep) ||\n-\t     (ignore_packed_keep_in_core && p->pack_keep_in_core)))\n-\t\treturn 0;\n+\n+\t/*\n+\t * Then handle .keep first, as we have a fast(er) path there.\n+\t */\n+\tif (ignore_packed_keep_on_disk || ignore_packed_keep_in_core) {\n+\t\t/*\n+\t\t * Set the flags for the kept-pack cache to be the ones we want\n+\t\t * to ignore.\n+\t\t *\n+\t\t * That is, if we are ignoring objects in on-disk keep packs,\n+\t\t * then we want to search through the on-disk keep and ignore\n+\t\t * the in-core ones.\n+\t\t */\n+\t\tunsigned flags = 0;\n+\t\tif (ignore_packed_keep_on_disk)\n+\t\t\tflags |= ON_DISK_KEEP_PACKS;\n+\t\tif (ignore_packed_keep_in_core)\n+\t\t\tflags |= IN_CORE_KEEP_PACKS;\n+\n+\t\tif (ignore_packed_keep_on_disk && p->pack_keep)\n+\t\t\treturn 0;\n+\t\tif (ignore_packed_keep_in_core && p->pack_keep_in_core)\n+\t\t\treturn 0;\n+\t\tif (has_object_kept_pack(oid, flags))\n+\t\t\treturn 0;\n+\t}\n+\n+\t/*\n+\t * At this point we know definitively that either we don't care about\n+\t * keep-packs, or the object is not in one. Keep checking other\n+\t * conditions...\n+\t */\n+\tif (!local || !have_non_local_packs)\n+\t\treturn 1;\n \n \t/* we don't know yet; keep looking for more packs */\n \treturn -1;\n }\n \n+static int want_object_in_pack_one(struct packed_git *p,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   int exclude,\n+\t\t\t\t   struct packed_git **found_pack,\n+\t\t\t\t   off_t *found_offset)\n+{\n+\toff_t offset;\n+\n+\tif (p == *found_pack)\n+\t\toffset = *found_offset;\n+\telse\n+\t\toffset = find_pack_entry_one(oid->hash, p);\n+\n+\tif (offset) {\n+\t\tif (!*found_pack) {\n+\t\t\tif (!is_pack_valid(p))\n+\t\t\t\treturn -1;\n+\t\t\t*found_offset = offset;\n+\t\t\t*found_pack = p;\n+\t\t}\n+\t\treturn want_found_object(oid, exclude, p);\n+\t}\n+\treturn -1;\n+}\n+\n /*\n  * Check whether we want the object in the pack (e.g., we do not want\n  * objects found in non-local stores if the \"--local\" option was used).\n@@ -1252,7 +1308,7 @@ static int want_object_in_pack(const struct object_id *oid,\n \t * are present we will determine the answer right now.\n \t */\n \tif (*found_pack) {\n-\t\twant = want_found_object(exclude, *found_pack);\n+\t\twant = want_found_object(oid, exclude, *found_pack);\n \t\tif (want != -1)\n \t\t\treturn want;\n \t}\n@@ -1260,53 +1316,22 @@ static int want_object_in_pack(const struct object_id *oid,\n \tfor (m = get_multi_pack_index(the_repository); m; m = m->next) {\n \t\tstruct pack_entry e;\n \t\tif (fill_midx_entry(the_repository, oid, &e, m)) {\n-\t\t\tstruct packed_git *p = e.p;\n-\t\t\toff_t offset;\n-\n-\t\t\tif (p == *found_pack)\n-\t\t\t\toffset = *found_offset;\n-\t\t\telse\n-\t\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\t\tif (offset) {\n-\t\t\t\tif (!*found_pack) {\n-\t\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\t\tcontinue;\n-\t\t\t\t\t*found_offset = offset;\n-\t\t\t\t\t*found_pack = p;\n-\t\t\t\t}\n-\t\t\t\twant = want_found_object(exclude, p);\n-\t\t\t\tif (want != -1)\n-\t\t\t\t\treturn want;\n-\t\t\t}\n-\t\t}\n-\t}\n-\n-\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n-\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\toff_t offset;\n-\n-\t\tif (p == *found_pack)\n-\t\t\toffset = *found_offset;\n-\t\telse\n-\t\t\toffset = find_pack_entry_one(oid->hash, p);\n-\n-\t\tif (offset) {\n-\t\t\tif (!*found_pack) {\n-\t\t\t\tif (!is_pack_valid(p))\n-\t\t\t\t\tcontinue;\n-\t\t\t\t*found_offset = offset;\n-\t\t\t\t*found_pack = p;\n-\t\t\t}\n-\t\t\twant = want_found_object(exclude, p);\n-\t\t\tif (!exclude && want > 0)\n-\t\t\t\tlist_move(&p->mru,\n-\t\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\t\twant = want_object_in_pack_one(e.p, oid, exclude, found_pack, found_offset);\n \t\t\tif (want != -1)\n \t\t\t\treturn want;\n \t\t}\n \t}\n \n+\tlist_for_each(pos, get_packed_git_mru(the_repository)) {\n+\t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n+\t\twant = want_object_in_pack_one(p, oid, exclude, found_pack, found_offset);\n+\t\tif (!exclude && want > 0)\n+\t\t\tlist_move(&p->mru,\n+\t\t\t\t  get_packed_git_mru(the_repository));\n+\t\tif (want != -1)\n+\t\t\treturn want;\n+\t}\n+\n \tif (uri_protocols.nr) {\n \t\tstruct configured_exclusion *ex =\n \t\t\toidmap_get(&configured_exclusions, oid);\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417528","messageId":"94e4f3ee3af3181c3805ee397d043e343038005a.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 5/8] p5303: measure time to repack with keep","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:17Z","receivedAt":"2021-02-23T02:26:28Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nAdd two new tests to measure repack performance. Both tests split the\nrepository into synthetic \"pushes\", and then leave the remaining objects\nin a big base pack.\n\nThe first new test marks an empty pack as \"kept\" and then passes\n--honor-pack-keep to avoid including objects in it. That doesn't change\nthe resulting pack, but it does let us compare to the normal repack case\nto see how much overhead we add to check whether objects are kept or\nnot.\n\nThe other test is of --stdin-packs, which gives us a sense of how that\nnumber scales based on the number of packs we provide as input. In each\nof those tests, the empty pack isn't considered, but the residual pack\n(objects that were left over and not included in one of the synthetic\npush packs) is marked as kept.\n\n(Note that in the single-pack case of the --stdin-packs test, there is\nnothing do since there are no non-excluded packs).\n\nHere are some timings on a recent clone of the kernel:\n\n  5303.5: repack (1)                          57.26(54.59+10.84)\n  5303.6: repack with kept (1)                57.33(54.80+10.51)\n\nin the 50-pack case, things start to slow down:\n\n  5303.11: repack (50)                        71.54(88.57+4.84)\n  5303.12: repack with kept (50)              85.12(102.05+4.94)\n\nand by the time we hit 1,000 packs, things are substantially worse, even\nthough the resulting pack produced is the same:\n\n  5303.17: repack (1000)                      216.87(490.79+14.57)\n  5303.18: repack with kept (1000)            665.63(938.87+15.76)\n\nThat's because the code paths around handling .keep files are known to\nscale badly; they look in every single pack file to find each object.\nOur solution to that was to notice that most repos don't have keep\nfiles, and to make that case a fast path. But as soon as you add a\nsingle .keep, that part of pack-objects slows down again (even if we\nhave fewer objects total to look at).\n\nLikewise, the scaling is pretty extreme on --stdin-packs (but each\nsubsequent test is also being asked to do more work):\n\n  5303.7: repack with --stdin-packs (1)       0.01(0.01+0.00)\n  5303.13: repack with --stdin-packs (50)     3.53(12.07+0.24)\n  5303.19: repack with --stdin-packs (1000)   195.83(371.82+8.10)\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n t/perf/p5303-many-packs.sh | 34 ++++++++++++++++++++++++++++++++--\n 1 file changed, 32 insertions(+), 2 deletions(-)\n\ndiff --git a/t/perf/p5303-many-packs.sh b/t/perf/p5303-many-packs.sh\nindex d90d714923..35c0cbdf49 100755\n--- a/t/perf/p5303-many-packs.sh\n+++ b/t/perf/p5303-many-packs.sh\n@@ -31,8 +31,15 @@ repack_into_n () {\n \t' \"$1\" >pushes &&\n \n \t# create base packfile\n-\thead -n 1 pushes |\n-\tgit pack-objects --delta-base-offset --revs staging/pack &&\n+\tbase_pack=$(\n+\t\thead -n 1 pushes |\n+\t\tgit pack-objects --delta-base-offset --revs staging/pack\n+\t) &&\n+\ttest_export base_pack &&\n+\n+\t# create an empty packfile\n+\tempty_pack=$(git pack-objects staging/pack </dev/null) &&\n+\ttest_export empty_pack &&\n \n \t# and then incrementals between each pair of commits\n \tlast= &&\n@@ -49,6 +56,12 @@ repack_into_n () {\n \t\tlast=$rev\n \tdone <pushes &&\n \n+\t(\n+\t\tfind staging -type f -name 'pack-*.pack' |\n+\t\t\txargs -n 1 basename | grep -v \"$base_pack\" &&\n+\t\tprintf \"^pack-%s.pack\\n\" $base_pack\n+\t) >stdin.packs\n+\n \t# and install the whole thing\n \trm -f .git/objects/pack/* &&\n \tmv staging/* .git/objects/pack/\n@@ -91,6 +104,23 @@ do\n \t\t  --reflog --indexed-objects --delta-base-offset \\\n \t\t  --stdout </dev/null >/dev/null\n \t'\n+\n+\ttest_perf \"repack with kept ($nr_packs)\" '\n+\t\tgit pack-objects --keep-true-parents \\\n+\t\t  --keep-pack=pack-$empty_pack.pack \\\n+\t\t  --honor-pack-keep --non-empty --all \\\n+\t\t  --reflog --indexed-objects --delta-base-offset \\\n+\t\t  --stdout </dev/null >/dev/null\n+\t'\n+\n+\ttest_perf \"repack with --stdin-packs ($nr_packs)\" '\n+\t\tgit pack-objects \\\n+\t\t  --keep-true-parents \\\n+\t\t  --stdin-packs \\\n+\t\t  --non-empty \\\n+\t\t  --delta-base-offset \\\n+\t\t  --stdout <stdin.packs >/dev/null\n+\t'\n done\n \n # Measure pack loading with 10,000 packs.\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417529","messageId":"51f57d5da23244ebde27ad6c14cbf4b63da3317d.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 8/8] builtin/repack.c: add '--geometric' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:27Z","receivedAt":"2021-02-23T02:26:38Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Often it is useful to both:\n\n  - have relatively few packfiles in a repository, and\n\n  - avoid having so few packfiles in a repository that we repack its\n    entire contents regularly\n\nThis patch implements a '--geometric=<n>' option in 'git repack'. This\nallows the caller to specify that they would like each pack to be at\nleast a factor times as large as the previous largest pack (by object\ncount).\n\nConcretely, say that a repository has 'n' packfiles, labeled P1, P2,\n..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\nWith a geometric factor of 'r', it should be that:\n\n  objects(Pi) > r*objects(P(i-1))\n\nfor all i in [1, n], where the packs are sorted by\n\n  objects(P1) <= objects(P2) <= ... <= objects(Pn).\n\nSince finding a true optimal repacking is NP-hard, we approximate it\nalong two directions:\n\n  1. We assume that there is a cutoff of packs _before starting the\n     repack_ where everything to the right of that cut-off already forms\n     a geometric progression (or no cutoff exists and everything must be\n     repacked).\n\n  2. We assume that everything smaller than the cutoff count must be\n     repacked. This forms our base assumption, but it can also cause\n     even the \"heavy\" packs to get repacked, for e.g., if we have 6\n     packs containing the following number of objects:\n\n       1, 1, 1, 2, 4, 32\n\n     then we would place the cutoff between '1, 1' and '1, 2, 4, 32',\n     rolling up the first two packs into a pack with 2 objects. That\n     breaks our progression and leaves us:\n\n       2, 1, 2, 4, 32\n         ^\n\n     (where the '^' indicates the position of our split). To restore a\n     progression, we move the split forward (towards larger packs)\n     joining each pack into our new pack until a geometric progression\n     is restored. Here, that looks like:\n\n       2, 1, 2, 4, 32  ~>  3, 2, 4, 32  ~>  5, 4, 32  ~> ... ~> 9, 32\n         ^                   ^                ^                   ^\n\nThis has the advantage of not repacking the heavy-side of packs too\noften while also only creating one new pack at a time. Another wrinkle\nis that we assume that loose, indexed, and reflog'd objects are\ninsignificant, and lump them into any new pack that we create. This can\nlead to non-idempotent results.\n\nSuggested-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt |  23 +++++\n builtin/repack.c             | 187 ++++++++++++++++++++++++++++++++++-\n t/t7703-repack-geometric.sh  | 137 +++++++++++++++++++++++++\n 3 files changed, 343 insertions(+), 4 deletions(-)\n create mode 100755 t/t7703-repack-geometric.sh\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex 92f146d27d..136da9fa0b 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -165,6 +165,29 @@ depth is 4095.\n \tPass the `--delta-islands` option to `git-pack-objects`, see\n \tlinkgit:git-pack-objects[1].\n \n+-g=<factor>::\n+--geometric=<factor>::\n+\tArrange resulting pack structure so that each successive pack\n+\tcontains at least `<factor>` times the number of objects as the\n+\tnext-largest pack.\n++\n+`git repack` ensures this by determining a \"cut\" of packfiles that need\n+to be repacked into one in order to ensure a geometric progression. It\n+picks the smallest set of packfiles such that as many of the larger\n+packfiles (by count of objects contained in that pack) may be left\n+intact.\n++\n+Unlike other repack modes, the set of objects to pack is determined\n+uniquely by the set of packs being \"rolled-up\"; in other words, the\n+packs determined to need to be combined in order to restore a geometric\n+progression.\n++\n+When `--unpacked` is specified, loose objects are implicitly included in\n+this \"roll-up\", without respect to their reachability. This is subject\n+to change in the future. This option (implying a drastically different\n+repack mode) is not guaranteed to work with all other combinations of\n+option to `git repack`).\n+\n Configuration\n -------------\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 01440de2d5..bcf280b10d 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -297,6 +297,124 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n #define ALL_INTO_ONE 1\n #define LOOSEN_UNREACHABLE 2\n \n+struct pack_geometry {\n+\tstruct packed_git **pack;\n+\tuint32_t pack_nr, pack_alloc;\n+\tuint32_t split;\n+};\n+\n+static uint32_t geometry_pack_weight(struct packed_git *p)\n+{\n+\tif (open_pack_index(p))\n+\t\tdie(_(\"cannot open index for %s\"), p->pack_name);\n+\treturn p->num_objects;\n+}\n+\n+static int geometry_cmp(const void *va, const void *vb)\n+{\n+\tuint32_t aw = geometry_pack_weight(*(struct packed_git **)va),\n+\t\t bw = geometry_pack_weight(*(struct packed_git **)vb);\n+\n+\tif (aw < bw)\n+\t\treturn -1;\n+\tif (aw > bw)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static void init_pack_geometry(struct pack_geometry **geometry_p)\n+{\n+\tstruct packed_git *p;\n+\tstruct pack_geometry *geometry;\n+\n+\t*geometry_p = xcalloc(1, sizeof(struct pack_geometry));\n+\tgeometry = *geometry_p;\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tif (!pack_kept_objects && p->pack_keep)\n+\t\t\tcontinue;\n+\n+\t\tALLOC_GROW(geometry->pack,\n+\t\t\t   geometry->pack_nr + 1,\n+\t\t\t   geometry->pack_alloc);\n+\n+\t\tgeometry->pack[geometry->pack_nr] = p;\n+\t\tgeometry->pack_nr++;\n+\t}\n+\n+\tQSORT(geometry->pack, geometry->pack_nr, geometry_cmp);\n+}\n+\n+static void split_pack_geometry(struct pack_geometry *geometry, int factor)\n+{\n+\tuint32_t i;\n+\tuint32_t split;\n+\toff_t total_size = 0;\n+\n+\tif (geometry->pack_nr <= 1) {\n+\t\tgeometry->split = geometry->pack_nr;\n+\t\treturn;\n+\t}\n+\n+\tsplit = geometry->pack_nr - 1;\n+\n+\t/*\n+\t * First, count the number of packs (in descending order of size) which\n+\t * already form a geometric progression.\n+\t */\n+\tfor (i = geometry->pack_nr - 1; i > 0; i--) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tstruct packed_git *prev = geometry->pack[i - 1];\n+\t\tif (geometry_pack_weight(ours) >= factor * geometry_pack_weight(prev))\n+\t\t\tsplit--;\n+\t\telse\n+\t\t\tbreak;\n+\t}\n+\n+\tif (split) {\n+\t\t/*\n+\t\t * Move the split one to the right, since the top element in the\n+\t\t * last-compared pair can't be in the progression. Only do this\n+\t\t * when we split in the middle of the array (otherwise if we got\n+\t\t * to the end, then the split is in the right place).\n+\t\t */\n+\t\tsplit++;\n+\t}\n+\n+\t/*\n+\t * Then, anything to the left of 'split' must be in a new pack. But,\n+\t * creating that new pack may cause packs in the heavy half to no longer\n+\t * form a geometric progression.\n+\t *\n+\t * Compute an expected size of the new pack, and then determine how many\n+\t * packs in the heavy half need to be joined into it (if any) to restore\n+\t * the geometric progression.\n+\t */\n+\tfor (i = 0; i < split; i++)\n+\t\ttotal_size += geometry_pack_weight(geometry->pack[i]);\n+\tfor (i = split; i < geometry->pack_nr; i++) {\n+\t\tstruct packed_git *ours = geometry->pack[i];\n+\t\tif (geometry_pack_weight(ours) < factor * total_size) {\n+\t\t\tsplit++;\n+\t\t\ttotal_size += geometry_pack_weight(ours);\n+\t\t} else\n+\t\t\tbreak;\n+\t}\n+\n+\tgeometry->split = split;\n+}\n+\n+static void clear_pack_geometry(struct pack_geometry *geometry)\n+{\n+\tif (!geometry)\n+\t\treturn;\n+\n+\tfree(geometry->pack);\n+\tgeometry->pack_nr = 0;\n+\tgeometry->pack_alloc = 0;\n+\tgeometry->split = 0;\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -304,6 +422,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list names = STRING_LIST_INIT_DUP;\n \tstruct string_list rollback = STRING_LIST_INIT_NODUP;\n \tstruct string_list existing_packs = STRING_LIST_INIT_DUP;\n+\tstruct pack_geometry *geometry = NULL;\n \tstruct strbuf line = STRBUF_INIT;\n \tint i, ext, ret;\n \tFILE *out;\n@@ -316,6 +435,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tint no_update_server_info = 0;\n \tstruct pack_objects_args po_args = {NULL};\n+\tint geometric_factor = 0;\n \n \tstruct option builtin_repack_options[] = {\n \t\tOPT_BIT('a', NULL, &pack_everything,\n@@ -356,6 +476,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t\tN_(\"repack objects in packs marked with .keep\")),\n \t\tOPT_STRING_LIST(0, \"keep-pack\", &keep_pack_list, N_(\"name\"),\n \t\t\t\tN_(\"do not repack this pack\")),\n+\t\tOPT_INTEGER('g', \"geometric\", &geometric_factor,\n+\t\t\t    N_(\"find a geometric progression with factor <N>\")),\n \t\tOPT_END()\n \t};\n \n@@ -382,6 +504,13 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (write_bitmaps && !(pack_everything & ALL_INTO_ONE))\n \t\tdie(_(incremental_bitmap_conflict_error));\n \n+\tif (geometric_factor) {\n+\t\tif (pack_everything)\n+\t\t\tdie(_(\"--geometric is incompatible with -A, -a\"));\n+\t\tinit_pack_geometry(&geometry);\n+\t\tsplit_pack_geometry(geometry, geometric_factor);\n+\t}\n+\n \tpackdir = mkpathdup(\"%s/pack\", get_object_directory());\n \tpacktmp = mkpathdup(\"%s/.tmp-%d-pack\", packdir, (int)getpid());\n \n@@ -396,9 +525,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s\",\n \t\t\t     keep_pack_list.items[i].string);\n \tstrvec_push(&cmd.args, \"--non-empty\");\n-\tstrvec_push(&cmd.args, \"--all\");\n-\tstrvec_push(&cmd.args, \"--reflog\");\n-\tstrvec_push(&cmd.args, \"--indexed-objects\");\n+\tif (!geometry) {\n+\t\t/*\n+\t\t * 'git pack-objects' will up all objects loose or packed\n+\t\t * (either rolling them up or leaving them alone), so don't pass\n+\t\t * these options.\n+\t\t *\n+\t\t * The implementation of 'git pack-objects --stdin-packs'\n+\t\t * makes them redundant (and the two are incompatible).\n+\t\t */\n+\t\tstrvec_push(&cmd.args, \"--all\");\n+\t\tstrvec_push(&cmd.args, \"--reflog\");\n+\t\tstrvec_push(&cmd.args, \"--indexed-objects\");\n+\t}\n \tif (has_promisor_remote())\n \t\tstrvec_push(&cmd.args, \"--exclude-promisor-objects\");\n \tif (write_bitmaps > 0)\n@@ -429,17 +568,37 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t\tstrvec_push(&cmd.env_array, \"GIT_REF_PARANOIA=1\");\n \t\t\t}\n \t\t}\n+\t} else if (geometry) {\n+\t\tstrvec_push(&cmd.args, \"--stdin-packs\");\n+\t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t} else {\n \t\tstrvec_push(&cmd.args, \"--unpacked\");\n \t\tstrvec_push(&cmd.args, \"--incremental\");\n \t}\n \n-\tcmd.no_stdin = 1;\n+\tif (geometry)\n+\t\tcmd.in = -1;\n+\telse\n+\t\tcmd.no_stdin = 1;\n \n \tret = start_command(&cmd);\n \tif (ret)\n \t\treturn ret;\n \n+\tif (geometry) {\n+\t\tFILE *in = xfdopen(cmd.in, \"w\");\n+\t\t/*\n+\t\t * The resulting pack should contain all objects in packs that\n+\t\t * are going to be rolled up, but exclude objects in packs which\n+\t\t * are being left alone.\n+\t\t */\n+\t\tfor (i = 0; i < geometry->split; i++)\n+\t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n+\t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n+\t\t\tfprintf(in, \"^%s\\n\", pack_basename(geometry->pack[i]));\n+\t\tfclose(in);\n+\t}\n+\n \tout = xfdopen(cmd.out, \"r\");\n \twhile (strbuf_getline_lf(&line, out) != EOF) {\n \t\tif (line.len != the_hash_algo->hexsz)\n@@ -507,6 +666,25 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tif (!string_list_has_string(&names, sha1))\n \t\t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n+\n+\t\tif (geometry) {\n+\t\t\tstruct strbuf buf = STRBUF_INIT;\n+\n+\t\t\tuint32_t i;\n+\t\t\tfor (i = 0; i < geometry->split; i++) {\n+\t\t\t\tstruct packed_git *p = geometry->pack[i];\n+\t\t\t\tif (string_list_has_string(&names,\n+\t\t\t\t\t\t\t   hash_to_hex(p->hash)))\n+\t\t\t\t\tcontinue;\n+\n+\t\t\t\tstrbuf_reset(&buf);\n+\t\t\t\tstrbuf_addstr(&buf, pack_basename(p));\n+\t\t\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\n+\t\t\t\tremove_redundant_pack(packdir, buf.buf);\n+\t\t\t}\n+\t\t\tstrbuf_release(&buf);\n+\t\t}\n \t\tif (!po_args.quiet && isatty(2))\n \t\t\topts |= PRUNE_PACKED_VERBOSE;\n \t\tprune_packed_objects(opts);\n@@ -528,6 +706,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstring_list_clear(&names, 0);\n \tstring_list_clear(&rollback, 0);\n \tstring_list_clear(&existing_packs, 0);\n+\tclear_pack_geometry(geometry);\n \tstrbuf_release(&line);\n \n \treturn 0;\ndiff --git a/t/t7703-repack-geometric.sh b/t/t7703-repack-geometric.sh\nnew file mode 100755\nindex 0000000000..96917fc163\n--- /dev/null\n+++ b/t/t7703-repack-geometric.sh\n@@ -0,0 +1,137 @@\n+#!/bin/sh\n+\n+test_description='git repack --geometric works correctly'\n+\n+. ./test-lib.sh\n+\n+GIT_TEST_MULTI_PACK_INDEX=0\n+\n+objdir=.git/objects\n+midx=$objdir/pack/multi-pack-index\n+\n+test_expect_success '--geometric with no packs' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\tgit repack --geometric 2 >out &&\n+\t\ttest_i18ngrep \"Nothing new to pack\" out\n+\t)\n+'\n+\n+test_expect_success '--geometric with an intact progression' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# These packs already form a geometric progression.\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 2 && # 6 objects\n+\t\ttest_commit_bulk --start=4 4 && # 12 objects\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >expect &&\n+\t\tgit repack --geometric 2 -d &&\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >actual &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success '--geometric with small-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >small &&\n+\t\ttest_commit_bulk --start=3 4 && # 12 objects\n+\t\ttest_commit_bulk --start=7 8 && # 24 objects\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tgit repack --geometric 2 -d &&\n+\n+\t\t# Three packs in total; two of the existing large ones, and one\n+\t\t# new one.\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\ttest_line_count = 3 after &&\n+\t\tcomm -3 small before | tr -d \"\\t\" >large &&\n+\t\tgrep -qFf large after\n+\t)\n+'\n+\n+test_expect_success '--geometric with small- and large-pack rollup' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\t# size(small1) + size(small2) > size(medium) / 2\n+\t\ttest_commit_bulk --start=1 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 1 && # 3 objects\n+\t\ttest_commit_bulk --start=2 3 && # 7 objects\n+\t\ttest_commit_bulk --start=6 9 && # 27 objects &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\n+\t\tgit repack --geometric 2 -d &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\t\tcomm -12 before after >untouched &&\n+\n+\t\t# Two packs in total; the largest pack from before running \"git\n+\t\t# repack\", and one new one.\n+\t\ttest_line_count = 1 untouched &&\n+\t\ttest_line_count = 2 after\n+\t)\n+'\n+\n+test_expect_success '--geometric ignores kept packs' '\n+\tgit init geometric &&\n+\ttest_when_finished \"rm -fr geometric\" &&\n+\t(\n+\t\tcd geometric &&\n+\n+\t\ttest_commit kept && # 3 objects\n+\t\ttest_commit pack && # 3 objects\n+\n+\t\tKEPT=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n+\t\trefs/tags/kept\n+\t\tEOF\n+\t\t) &&\n+\t\tPACK=$(git pack-objects --revs $objdir/pack/pack <<-EOF\n+\t\trefs/tags/pack\n+\t\t^refs/tags/kept\n+\t\tEOF\n+\t\t) &&\n+\n+\t\t# neither pack contains more than twice the number of objects in\n+\t\t# the other, so they should be combined. but, marking one as\n+\t\t# .kept on disk will \"freeze\" it, so the pack structure should\n+\t\t# remain unchanged.\n+\t\ttouch $objdir/pack/pack-$KEPT.keep &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >before &&\n+\t\tgit repack --geometric 2 -d &&\n+\t\tfind $objdir/pack -name \"*.pack\" | sort >after &&\n+\n+\t\t# both packs should still exist\n+\t\ttest_path_is_file $objdir/pack/pack-$KEPT.pack &&\n+\t\ttest_path_is_file $objdir/pack/pack-$PACK.pack &&\n+\n+\t\t# and no new packs should be created\n+\t\ttest_cmp before after &&\n+\n+\t\t# Passing --pack-kept-objects causes packs with a .keep file to\n+\t\t# be repacked, too.\n+\t\tgit repack --geometric 2 -d --pack-kept-objects &&\n+\n+\t\tfind $objdir/pack -name \"*.pack\" >after &&\n+\t\ttest_line_count = 1 after\n+\t)\n+'\n+\n+test_done\n-- \n2.30.0.667.g81c0cbc6fd\n"},{"id":"417530","messageId":"db9f07ec1ae3fe0d9dce922ee0dea831e4a25b13.1614047097.git.me@ttaylorr.com","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"[PATCH v4 7/8] packfile: add kept-pack cache for find_kept_pack_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T02:25:23Z","receivedAt":"2021-02-23T02:26:40Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nIn a recent patch we added a function 'find_kept_pack_entry()' to look\nfor an object only among kept packs.\n\nWhile this function avoids doing any lookup work in non-kept packs, it\nis still linear in the number of packs, since we have to traverse the\nlinked list of packs once per object. Let's cache a reduced version of\nthat list to save us time.\n\nNote that this cache will last the lifetime of the program. We could\ninvalidate it on reprepare_packed_git(), but there's not much point in\nbeing rigorous here:\n\n  - we might already fail to notice new .keep packs showing up after the\n    program starts. We only reprepare_packed_git() when we fail to find\n    an object. But adding a new pack won't cause that to happen.\n    Somebody repacking could add a new pack and delete an old one, but\n    most of the time we'd have a descriptor or mmap open to the old\n    pack anyway, so we might not even notice.\n\n  - in pack-objects we already cache the .keep state at startup, since\n    56dfeb6263 (pack-objects: compute local/ignore_pack_keep early,\n    2016-07-29). So this is just extending that concept further.\n\n  - we don't have to worry about any packed_git being removed; we always\n    keep the old structs around, even after reprepare_packed_git()\n\nWe do defensively invalidate the cache in case the set of kept packs\nbeing asked for changes (e.g., only in-core kept packs were cached, but\nsuddenly the caller also wants on-disk kept packs, too). In theory we\ncould build all three caches and switch between them, but it's not\nnecessary, since this patch (and series) never changes the set of kept\npacks that it wants to inspect from the cache.\n\nSo that \"optimization\" is more about being defensive in the face of\nfuture changes than it is about asking for multiple kinds of kept packs\nin this patch.\n\nHere are p5303 results (as always, measured against the kernel):\n\n  Test                                        HEAD^                   HEAD\n  -----------------------------------------------------------------------------------------------\n  5303.5: repack (1)                          57.34(54.66+10.88)      56.98(54.36+10.98) -0.6%\n  5303.6: repack with kept (1)                57.38(54.83+10.49)      57.17(54.97+10.26) -0.4%\n  5303.11: repack (50)                        71.70(88.99+4.74)       71.62(88.48+5.08) -0.1%\n  5303.12: repack with kept (50)              72.58(89.61+4.78)       71.56(88.80+4.59) -1.4%\n  5303.17: repack (1000)                      217.19(491.72+14.25)    217.31(490.82+14.53) +0.1%\n  5303.18: repack with kept (1000)            246.12(520.07+14.93)    217.08(490.37+15.10) -11.8%\n\nand the --stdin-packs case, which scales a little bit better (although\nnot by that much even at 1,000 packs):\n\n  5303.7: repack with --stdin-packs (1)       0.00(0.00+0.00)         0.00(0.00+0.00) =\n  5303.13: repack with --stdin-packs (50)     3.43(11.75+0.24)        3.43(11.69+0.30) +0.0%\n  5303.19: repack with --stdin-packs (1000)   130.50(307.15+7.66)     125.13(301.36+8.04) -4.1%\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-store.h |  5 +++\n packfile.c     | 99 ++++++++++++++++++++++++++++----------------------\n 2 files changed, 61 insertions(+), 43 deletions(-)\n\ndiff --git a/object-store.h b/object-store.h\nindex 541dab0858..ec32c23dcb 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -153,6 +153,11 @@ struct raw_object_store {\n \t/* A most-recently-used ordered version of the packed_git list. */\n \tstruct list_head packed_git_mru;\n \n+\tstruct {\n+\t\tstruct packed_git **packs;\n+\t\tunsigned flags;\n+\t} kept_pack_cache;\n+\n \t/*\n \t * A map of packfiles to packed_git structs for tracking which\n \t * packs have been loaded already.\ndiff --git a/packfile.c b/packfile.c\nindex 7f84f221ce..57d5b436fb 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2042,10 +2042,7 @@ static int fill_pack_entry(const struct object_id *oid,\n \treturn 1;\n }\n \n-static int find_one_pack_entry(struct repository *r,\n-\t\t\t       const struct object_id *oid,\n-\t\t\t       struct pack_entry *e,\n-\t\t\t       int kept_only)\n+int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n {\n \tstruct list_head *pos;\n \tstruct multi_pack_index *m;\n@@ -2055,49 +2052,63 @@ static int find_one_pack_entry(struct repository *r,\n \t\treturn 0;\n \n \tfor (m = r->objects->multi_pack_index; m; m = m->next) {\n-\t\tif (!fill_midx_entry(r, oid, e, m))\n-\t\t\tcontinue;\n-\n-\t\tif (!kept_only)\n-\t\t\treturn 1;\n-\n-\t\tif (((kept_only & ON_DISK_KEEP_PACKS) && e->p->pack_keep) ||\n-\t\t    ((kept_only & IN_CORE_KEEP_PACKS) && e->p->pack_keep_in_core))\n+\t\tif (fill_midx_entry(r, oid, e, m))\n \t\t\treturn 1;\n \t}\n \n \tlist_for_each(pos, &r->objects->packed_git_mru) {\n \t\tstruct packed_git *p = list_entry(pos, struct packed_git, mru);\n-\t\tif (p->multi_pack_index && !kept_only) {\n-\t\t\t/*\n-\t\t\t * If this pack is covered by the MIDX, we'd have found\n-\t\t\t * the object already in the loop above if it was here,\n-\t\t\t * so don't bother looking.\n-\t\t\t *\n-\t\t\t * The exception is if we are looking only at kept\n-\t\t\t * packs. An object can be present in two packs covered\n-\t\t\t * by the MIDX, one kept and one not-kept. And as the\n-\t\t\t * MIDX points to only one copy of each object, it might\n-\t\t\t * have returned only the non-kept version above. We\n-\t\t\t * have to check again to be thorough.\n-\t\t\t */\n-\t\t\tcontinue;\n-\t\t}\n-\t\tif (!kept_only ||\n-\t\t    (((kept_only & ON_DISK_KEEP_PACKS) && p->pack_keep) ||\n-\t\t     ((kept_only & IN_CORE_KEEP_PACKS) && p->pack_keep_in_core))) {\n-\t\t\tif (fill_pack_entry(oid, e, p)) {\n-\t\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n-\t\t\t\treturn 1;\n-\t\t\t}\n+\t\tif (!p->multi_pack_index && fill_pack_entry(oid, e, p)) {\n+\t\t\tlist_move(&p->mru, &r->objects->packed_git_mru);\n+\t\t\treturn 1;\n \t\t}\n \t}\n \treturn 0;\n }\n \n-int find_pack_entry(struct repository *r, const struct object_id *oid, struct pack_entry *e)\n+static void maybe_invalidate_kept_pack_cache(struct repository *r,\n+\t\t\t\t\t     unsigned flags)\n {\n-\treturn find_one_pack_entry(r, oid, e, 0);\n+\tif (!r->objects->kept_pack_cache.packs)\n+\t\treturn;\n+\tif (r->objects->kept_pack_cache.flags == flags)\n+\t\treturn;\n+\tFREE_AND_NULL(r->objects->kept_pack_cache.packs);\n+\tr->objects->kept_pack_cache.flags = 0;\n+}\n+\n+static struct packed_git **kept_pack_cache(struct repository *r, unsigned flags)\n+{\n+\tmaybe_invalidate_kept_pack_cache(r, flags);\n+\n+\tif (!r->objects->kept_pack_cache.packs) {\n+\t\tstruct packed_git **packs = NULL;\n+\t\tsize_t nr = 0, alloc = 0;\n+\t\tstruct packed_git *p;\n+\n+\t\t/*\n+\t\t * We want \"all\" packs here, because we need to cover ones that\n+\t\t * are used by a midx, as well. We need to look in every one of\n+\t\t * them (instead of the midx itself) to cover duplicates. It's\n+\t\t * possible that an object is found in two packs that the midx\n+\t\t * covers, one kept and one not kept, but the midx returns only\n+\t\t * the non-kept version.\n+\t\t */\n+\t\tfor (p = get_all_packs(r); p; p = p->next) {\n+\t\t\tif ((p->pack_keep && (flags & ON_DISK_KEEP_PACKS)) ||\n+\t\t\t    (p->pack_keep_in_core && (flags & IN_CORE_KEEP_PACKS))) {\n+\t\t\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\t\t\tpacks[nr++] = p;\n+\t\t\t}\n+\t\t}\n+\t\tALLOC_GROW(packs, nr + 1, alloc);\n+\t\tpacks[nr] = NULL;\n+\n+\t\tr->objects->kept_pack_cache.packs = packs;\n+\t\tr->objects->kept_pack_cache.flags = flags;\n+\t}\n+\n+\treturn r->objects->kept_pack_cache.packs;\n }\n \n int find_kept_pack_entry(struct repository *r,\n@@ -2105,13 +2116,15 @@ int find_kept_pack_entry(struct repository *r,\n \t\t\t unsigned flags,\n \t\t\t struct pack_entry *e)\n {\n-\t/*\n-\t * Load all packs, including midx packs, since our \"kept\" strategy\n-\t * relies on that. We're relying on the side effect of it setting up\n-\t * r->objects->packed_git, which is a little ugly.\n-\t */\n-\tget_all_packs(r);\n-\treturn find_one_pack_entry(r, oid, e, flags);\n+\tstruct packed_git **cache;\n+\n+\tfor (cache = kept_pack_cache(r, flags); *cache; cache++) {\n+\t\tstruct packed_git *p = *cache;\n+\t\tif (fill_pack_entry(oid, e, p))\n+\t\t\treturn 1;\n+\t}\n+\n+\treturn 0;\n }\n \n int has_object_pack(const struct object_id *oid)\n-- \n2.30.0.667.g81c0cbc6fd\n\n"},{"id":"417531","messageId":"YDR5CApT3xw4QKwd@coredump.intra.peff.net","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-23T03:39:52Z","receivedAt":"2021-02-23T03:40:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Feb 22, 2021 at 09:24:59PM -0500, Taylor Blau wrote:\n\n> Here's a very lightly modified version on v3 of mine and Peff's series\n> to add a new 'git repack --geometric' mode. Almost nothing has changed\n> since last time, with the exception of:\n> \n>   - Packs listed over standard input to 'git pack-objects --stdin-packs'\n>     are sorted in descending mtime order (and objects are strung\n>     together in pack order as before) so that objects are laid out\n>     roughly newest-to-oldest in the resulting pack.\n> \n>   - Swapped the order of two paragraphs in patch 5 to make the perf\n>     results clearer.\n> \n>   - Mention '--unpacked' specifically in the documentation for 'git\n>     repack --geometric'.\n> \n>   - Typo fixes.\n\nThanks, this all looks great to me.\n\n-Peff\n"},{"id":"417547","messageId":"xmqq7dmz5iw5.fsf@gitster.g","threadId":"55012","inReplyTo":"cover.1614047097.git.me@ttaylorr.com","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-23T07:43:22Z","receivedAt":"2021-02-23T07:44:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n\n>     ++\t/*\n>     ++\t * order packs by descending mtime so that objects are laid out\n>     ++\t * roughly as newest-to-oldest\n>     ++\t */\n>      +\tif (a->mtime < b->mtime)\n>      +\t\treturn 1;\n>     ++\telse if (b->mtime < a->mtime)\n>     ++\t\treturn -1;\n>      +\telse\n>      +\t\treturn 0;\n\nI think this strategy makes sense when this repack using this new\nfeature is run for the first time in a repository that acquired many\npacks over time.  I am not sure what happens after the feature is\nused a few times---it won't always be the newest sets of packs that\nwill be rewritten, but sometimes older ones are also coalesced, and\nwhen that happens the resulting pack that consists primarily of older\nobjects would end up having a more recent timestamp, no?\n\nEven then, I do agree that newer to older would be beneficial most\nof the time, so this is of course not an objection against this\nparticular sort order.\n"},{"id":"417549","messageId":"xmqqsg5n436v.fsf@gitster.g","threadId":"55012","inReplyTo":"649cf9020bfdae9f48e3efbfbc52429cefd31432.1614047097.git.me@ttaylorr.com","subject":"Re: [PATCH v4 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-23T08:07:52Z","receivedAt":"2021-02-23T08:09:45Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> (I found it convenient while developing this patch to have 'git\n> pack-objects' report the number of objects which were visited and got\n> their namehash fields filled in during traversal. This is also included\n> in the below patch via trace2 data lines).\n\nIt does sound like a well thought out strategy to give name-hash to\nentries that we may have to find good delta bases afresh, while\nstopping upon hitting parts of the history we won't have to (either\nbecause they are in \"excluded\" packs, which you did here, or because\nthey can take advantage of the \"reuse existing delta base\" logic [*],\nwhich we may want to look further into in future follow-on topics).\n\n\n[Footnote]\n\n* I presume that such a logic may, instead of stopping at an object\n  that is in an excluded pack, stop at an object that is stored in\n  the current pack as a delta and its base is also going to be\n  packed (and the latter by definition is always true, I presume, as\n  everything in the included pack would be packed)\n\n\n"},{"id":"417577","messageId":"YDVM9U7zLstNBVq2@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqq7dmz5iw5.fsf@gitster.g","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-23T18:44:05Z","receivedAt":"2021-02-23T18:44:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Feb 22, 2021 at 11:43:22PM -0800, Junio C Hamano wrote:\n\n> >     ++\t/*\n> >     ++\t * order packs by descending mtime so that objects are laid out\n> >     ++\t * roughly as newest-to-oldest\n> >     ++\t */\n> >      +\tif (a->mtime < b->mtime)\n> >      +\t\treturn 1;\n> >     ++\telse if (b->mtime < a->mtime)\n> >     ++\t\treturn -1;\n> >      +\telse\n> >      +\t\treturn 0;\n> \n> I think this strategy makes sense when this repack using this new\n> feature is run for the first time in a repository that acquired many\n> packs over time.  I am not sure what happens after the feature is\n> used a few times---it won't always be the newest sets of packs that\n> will be rewritten, but sometimes older ones are also coalesced, and\n> when that happens the resulting pack that consists primarily of older\n> objects would end up having a more recent timestamp, no?\n\nYeah, this is definitely a heuristic that can get out of sync with\nreality. I think in general if you have base pack A and somebody pushes\nup B, C, and D in sequence, we're likely to roll up a single DBC (in\nthat order) pack. Further pushes E, F, G would have newer mtimes. So we\nmight get GFEDBC directly. Or we might get GFE and DBC, but the former\nwould still have a newer mtime, so we'd create GFEDBC on the next run.\n\nThe issues come from:\n\n  - we are deciding what to roll up based on size. A big push might not\n    get rolled up immediately, putting it out-of-sync with the rest of\n    the rollups.\n\n  - we are happy to manipulate pack mtimes under the hood as part of the\n    freshen_*() code.\n\nI think you probably wouldn't want to use this roll-up strategy all the\ntime (even though in theory it would eventually roll up to a single good\npack), just because it is based on heuristics like this. You'd want to\noccasionally run a \"real\" repack that does a full traversal, possibly\npruning objects, etc.\n\nAnd that's how we plan to use it at GitHub. I don't remember how much of\nthe root problem we've discussed on-list, but the crux of it is:\nper-object costs including traversal can get really high on big\nrepositories. Our shared-storage repo for all torvalds/linux forks is on\nthe order of 45M objects, and some companies with large and active\nprivate repositories are close to that. Traversing the object graph\ntakes 15+ minutes (plus another 15 for delta island marking). For busy\nrepositories, by the time you finish repacking, it's time to start\nagain. :)\n\n> Even then, I do agree that newer to older would be beneficial most\n> of the time, so this is of course not an objection against this\n> particular sort order.\n\nSo yeah. I consider this best-effort for sure, and I think this sort\norder is the best we can do without traversing.\n\nOTOH, we _do_ actually do a partial traversal in this latest version of\nthe series. We could use that to impact the final write order. It\ndoesn't necessarily hit every object, though, so we'd still want to fall\nback on this pack ordering heuristic. I'm content to leave punt on that\nwork for now, and leave it for a future series after we see how this\nheuristic performs in practice.\n\n-Peff\n"},{"id":"417578","messageId":"YDVOuwMTBoQ57Omk@coredump.intra.peff.net","threadId":"55012","inReplyTo":"xmqqsg5n436v.fsf@gitster.g","subject":"Re: [PATCH v4 3/8] builtin/pack-objects.c: add '--stdin-packs' option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-23T18:51:39Z","receivedAt":"2021-02-23T18:52:26Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 23, 2021 at 12:07:52AM -0800, Junio C Hamano wrote:\n\n> Taylor Blau <me@ttaylorr.com> writes:\n> \n> > (I found it convenient while developing this patch to have 'git\n> > pack-objects' report the number of objects which were visited and got\n> > their namehash fields filled in during traversal. This is also included\n> > in the below patch via trace2 data lines).\n> \n> It does sound like a well thought out strategy to give name-hash to\n> entries that we may have to find good delta bases afresh, while\n> stopping upon hitting parts of the history we won't have to (either\n> because they are in \"excluded\" packs, which you did here, or because\n> they can take advantage of the \"reuse existing delta base\" logic [*],\n> which we may want to look further into in future follow-on topics).\n> \n> [Footnote]\n> \n> * I presume that such a logic may, instead of stopping at an object\n>   that is in an excluded pack, stop at an object that is stored in\n>   the current pack as a delta and its base is also going to be\n>   packed (and the latter by definition is always true, I presume, as\n>   everything in the included pack would be packed)\n\nI'm not sure if using deltas as a heuristic for stopping traversal makes\nsense. They don't necessarily correspond to the history graph, or to\nwhat was pushed. E.g., if I see that tree X is a delta against tree Y,\nthen we might say: if Y is not excluded by being in one of the base\npacks, then we will reuse the delta. We do not need the namehash of X,\nsince we already know its delta.\n\nBut that does not tell us anything about the subtrees and blobs\ncontained in X. We still want to traverse X in order to find out _their_\nname hashes, because it is likely that we will need to delta some of\nthose.\n\nOf course if you see a blob that is a delta that you plan to reuse, you\nknow you can stop there. But by the time you get to it, you already know\nits namehash, and there is nothing left to traverse. :)\n\n-Peff\n"},{"id":"417592","messageId":"1724378.IzK8VI2DXP@mfick-lnx","threadId":"55012","inReplyTo":"YDVM9U7zLstNBVq2@coredump.intra.peff.net","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-02-23T19:54:56Z","receivedAt":"2021-02-23T19:55:50Z","isPatch":true,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Tuesday, February 23, 2021 1:44:05 PM MST Jeff King wrote:\n> On Mon, Feb 22, 2021 at 11:43:22PM -0800, Junio C Hamano wrote:\n> > >     ++\t/*\n> > >     ++\t * order packs by descending mtime so that objects are laid out\n> > >     ++\t * roughly as newest-to-oldest\n> > >     ++\t */\n> > >     \n> > >      +\tif (a->mtime < b->mtime)\n> > >      +\t\treturn 1;\n> > >     \n> > >     ++\telse if (b->mtime < a->mtime)\n> > >     ++\t\treturn -1;\n> > >     \n> > >      +\telse\n> > >      +\t\treturn 0;\n> > \n> > I think this strategy makes sense when this repack using this new\n> > feature is run for the first time in a repository that acquired many\n> > packs over time.  I am not sure what happens after the feature is\n> > used a few times---it won't always be the newest sets of packs that\n> > will be rewritten, but sometimes older ones are also coalesced, and\n> > when that happens the resulting pack that consists primarily of older\n> > objects would end up having a more recent timestamp, no?\n> \n> Yeah, this is definitely a heuristic that can get out of sync with\n> reality. I think in general if you have base pack A and somebody pushes\n> up B, C, and D in sequence, we're likely to roll up a single DBC (in\n> that order) pack. Further pushes E, F, G would have newer mtimes. So we\n> might get GFEDBC directly. Or we might get GFE and DBC, but the former\n> would still have a newer mtime, so we'd create GFEDBC on the next run.\n> \n> The issues come from:\n> \n>   - we are deciding what to roll up based on size. A big push might not\n>     get rolled up immediately, putting it out-of-sync with the rest of\n>     the rollups.\n\nWould it make sense to somehow detect all new packs since the last rollup and \nalways include them in the rollup no matter what their size? That is one thing \nthat my git-exproll script did. One of the main reasons to do this was because \nnewer packs tended to look big (I was using bytes to determine size), and \nnewer packs were often bigger on disk compared to other packs with similar \nobjects in them (I think you suggested this was due to the thickening of packs \non receipt). Maybe roll up all packs with a timestamp \"new enough\", no matter \nhow big they are?\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"417594","messageId":"YDVgPtkaTb9zNq0/@nand.local","threadId":"55012","inReplyTo":"1724378.IzK8VI2DXP@mfick-lnx","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-23T20:06:22Z","receivedAt":"2021-02-23T20:08:54Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Feb 23, 2021 at 12:54:56PM -0700, Martin Fick wrote:\n> Would it make sense to somehow detect all new packs since the last rollup and\n> always include them in the rollup no matter what their size? That is one thing\n> that my git-exproll script did.\n\nI'm certainly not opposed, and this could certainly be done in an\nadditive way (i.e., after this series). I think the current approach has\nnice properties, but I could also see \"roll-up all packs that have\nmtimes after xyz timestamp\" being useful.\n\nIt would even be possible to reuse a lot of the geometric repack\nmachinery. Having a separate path to arrange packs by their mtimes and\ndetermine the \"split\" at pack whose mtime is nearest the provided one\nwould do exactly what you want.\n\n(As a side-note, reading the original threads about your git-exproll was\nquite humbling, since it turns out all of the problems I thought were\nhard had already been discussed eight years ago!)\n\nThanks,\nTaylor\n"},{"id":"417609","messageId":"YDViUPT4JhGJLjji@coredump.intra.peff.net","threadId":"55012","inReplyTo":"1724378.IzK8VI2DXP@mfick-lnx","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-23T20:15:12Z","receivedAt":"2021-02-23T20:16:48Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 23, 2021 at 12:54:56PM -0700, Martin Fick wrote:\n\n> > Yeah, this is definitely a heuristic that can get out of sync with\n> > reality. I think in general if you have base pack A and somebody pushes\n> > up B, C, and D in sequence, we're likely to roll up a single DBC (in\n> > that order) pack. Further pushes E, F, G would have newer mtimes. So we\n> > might get GFEDBC directly. Or we might get GFE and DBC, but the former\n> > would still have a newer mtime, so we'd create GFEDBC on the next run.\n> > \n> > The issues come from:\n> > \n> >   - we are deciding what to roll up based on size. A big push might not\n> >     get rolled up immediately, putting it out-of-sync with the rest of\n> >     the rollups.\n> \n> Would it make sense to somehow detect all new packs since the last rollup and \n> always include them in the rollup no matter what their size? That is one thing \n> that my git-exproll script did. One of the main reasons to do this was because \n> newer packs tended to look big (I was using bytes to determine size), and \n> newer packs were often bigger on disk compared to other packs with similar \n> objects in them (I think you suggested this was due to the thickening of packs \n> on receipt). Maybe roll up all packs with a timestamp \"new enough\", no matter \n> how big they are?\n\nThat works against the \"geometric\" part of the strategy, which is trying\nto roll up in a sequence that is amortized-linear. I.e., we are not\nalways rolling up everything outside of the base pack, but trying to\nroll up little into medium, and then eventually medium into large. If\nyou roll up things that are \"too big\", then you end up rewriting the\nbytes more often, and your amount of work becomes super-linear.\n\nNow whether that matters all that much or not is perhaps another\ndiscussion. The current strategy is mostly to repack all-into-one with\nno base, which is the worst possible case. So just about any rollup\nstrategy will be an improvement. ;)\n\n-Peff\n"},{"id":"417623","messageId":"8347289.KuArlTWdtP@mfick-lnx","threadId":"55012","inReplyTo":"YDViUPT4JhGJLjji@coredump.intra.peff.net","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-02-23T21:41:09Z","receivedAt":"2021-02-23T21:42:22Z","isPatch":true,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Tuesday, February 23, 2021 3:15:12 PM MST Jeff King wrote:\n> On Tue, Feb 23, 2021 at 12:54:56PM -0700, Martin Fick wrote:\n> > > Yeah, this is definitely a heuristic that can get out of sync with\n> > > reality. I think in general if you have base pack A and somebody pushes\n> > > up B, C, and D in sequence, we're likely to roll up a single DBC (in\n> > > that order) pack. Further pushes E, F, G would have newer mtimes. So we\n> > > might get GFEDBC directly. Or we might get GFE and DBC, but the former\n> > > would still have a newer mtime, so we'd create GFEDBC on the next run.\n> > > \n> > > The issues come from:\n> > >   - we are deciding what to roll up based on size. A big push might not\n> > >   \n> > >     get rolled up immediately, putting it out-of-sync with the rest of\n> > >     the rollups.\n> > \n> > Would it make sense to somehow detect all new packs since the last rollup\n> > and always include them in the rollup no matter what their size? That is\n> > one thing that my git-exproll script did. One of the main reasons to do\n> > this was because newer packs tended to look big (I was using bytes to\n> > determine size), and newer packs were often bigger on disk compared to\n> > other packs with similar objects in them (I think you suggested this was\n> > due to the thickening of packs on receipt). Maybe roll up all packs with\n> > a timestamp \"new enough\", no matter how big they are?\n> \n> That works against the \"geometric\" part of the strategy, which is trying\n> to roll up in a sequence that is amortized-linear. I.e., we are not\n> always rolling up everything outside of the base pack, but trying to\n> roll up little into medium, and then eventually medium into large. If\n> you roll up things that are \"too big\", then you end up rewriting the\n> bytes more often, and your amount of work becomes super-linear.\n\nI'm not sure I follow, it would seem to me that it would stay linear, and be \nat most rewriting each new packfile once more than previously? Are you \nenvisioning more work than that?\n\n> Now whether that matters all that much or not is perhaps another\n> discussion. The current strategy is mostly to repack all-into-one with\n> no base, which is the worst possible case. So just about any rollup\n> strategy will be an improvement. ;)\n\n+1 Yes, while anything would be an improvement, this series' approach is very \ngood! Thanks for doing this!!\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"417625","messageId":"YDV5UX6n2nZb9Tn2@coredump.intra.peff.net","threadId":"55012","inReplyTo":"8347289.KuArlTWdtP@mfick-lnx","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-23T21:53:21Z","receivedAt":"2021-02-23T21:54:21Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 23, 2021 at 02:41:09PM -0700, Martin Fick wrote:\n\n> > > Would it make sense to somehow detect all new packs since the last rollup\n> > > and always include them in the rollup no matter what their size? That is\n> > > one thing that my git-exproll script did. One of the main reasons to do\n> > > this was because newer packs tended to look big (I was using bytes to\n> > > determine size), and newer packs were often bigger on disk compared to\n> > > other packs with similar objects in them (I think you suggested this was\n> > > due to the thickening of packs on receipt). Maybe roll up all packs with\n> > > a timestamp \"new enough\", no matter how big they are?\n> > \n> > That works against the \"geometric\" part of the strategy, which is trying\n> > to roll up in a sequence that is amortized-linear. I.e., we are not\n> > always rolling up everything outside of the base pack, but trying to\n> > roll up little into medium, and then eventually medium into large. If\n> > you roll up things that are \"too big\", then you end up rewriting the\n> > bytes more often, and your amount of work becomes super-linear.\n> \n> I'm not sure I follow, it would seem to me that it would stay linear, and be \n> at most rewriting each new packfile once more than previously? Are you \n> envisioning more work than that?\n\nMaybe I don't understand what you're proposing.\n\nThe idea of the geometric repack is that by sorting by size and then\nfinding a \"cutoff\" within the size array, we can make sure that we roll\nup a sufficiently small number of bytes in each roll-up that it ends up\nlinear in the size of the repo in the long run. But if we roll up\nwithout regard to size, then our worst case is that the biggest pack is\nthe newest (imagine a repo with 10 small pushes and then one gigantic\none). So we roll that up with some small packs, doing effectively\nO(size_of_repo) work. And then in the next roll up we do it again, and\nso on. So we end up with O(size_of_repo * nr_rollups) total work. Which\nis no better than having just done a full repack at each rollup.\n\nNow I don't think we'd see that worst case in practice that much. And\ndepending on your definition of \"new enough\", you might keep nr_rollups\npretty small.\n\n-Peff\n"},{"id":"417626","messageId":"2190798.zr6bNypoxz@mfick-lnx","threadId":"55012","inReplyTo":"YDVgPtkaTb9zNq0/@nand.local","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-02-23T21:57:25Z","receivedAt":"2021-02-23T21:57:58Z","isPatch":true,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Tuesday, February 23, 2021 3:06:22 PM MST Taylor Blau wrote:\n> On Tue, Feb 23, 2021 at 12:54:56PM -0700, Martin Fick wrote:\n> > Would it make sense to somehow detect all new packs since the last rollup\n> > and always include them in the rollup no matter what their size? That is\n> > one thing that my git-exproll script did.\n> \n> I'm certainly not opposed, and this could certainly be done in an\n> additive way (i.e., after this series). I think the current approach has\n> nice properties, but I could also see \"roll-up all packs that have\n> mtimes after xyz timestamp\" being useful.\n\nJust to be clear, I meant to combine the two approaches. And yes, my \nsuggestion would likely make more sense as an additive switch later on.\n \n> It would even be possible to reuse a lot of the geometric repack\n> machinery. Having a separate path to arrange packs by their mtimes and\n> determine the \"split\" at pack whose mtime is nearest the provided one\n> would do exactly what you want.\n\nI was thinking to keep all of your geometric repack machinery and only looking \nfor the split point starting at the right most pack which is newer than the \nprovided mtime, and then possibly enhancing the approach with a clever way to \nuse the mtime of the last consolidation (maybe by touching a pack/.geometric \nfile?).\n\n> (As a side-note, reading the original threads about your git-exproll was\n> quite humbling, since it turns out all of the problems I thought were\n> hard had already been discussed eight years ago!)\n\nThanks, but I think you have likely done a much better job than what I did. \nYour approach of using object counts is likely much better as it should be \nstable, using byte counts is not. You are also solving only one problem at a \ntime, that's probably better than my hodge-podge of at least 3 different \nproblems. And the most important part of your approach as I understand it, is \nthat it actually saves CPU time whereas my approach only saved IO.\n\nCheers,\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"417690","messageId":"4678947.aH8v75nzy7@mfick-lnx","threadId":"55012","inReplyTo":"YDV5UX6n2nZb9Tn2@coredump.intra.peff.net","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-02-24T18:13:53Z","receivedAt":"2021-02-24T18:15:10Z","isPatch":true,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Tuesday, February 23, 2021 4:53:21 PM MST Jeff King wrote:\n> On Tue, Feb 23, 2021 at 02:41:09PM -0700, Martin Fick wrote:\n> > > > Would it make sense to somehow detect all new packs since the last\n> > > > rollup\n> > > > and always include them in the rollup no matter what their size? That\n> > > > is\n> > > > one thing that my git-exproll script did. One of the main reasons to\n> > > > do\n> > > > this was because newer packs tended to look big (I was using bytes to\n> > > > determine size), and newer packs were often bigger on disk compared to\n> > > > other packs with similar objects in them (I think you suggested this\n> > > > was\n> > > > due to the thickening of packs on receipt). Maybe roll up all packs\n> > > > with\n> > > > a timestamp \"new enough\", no matter how big they are?\n> > > \n> > > That works against the \"geometric\" part of the strategy, which is trying\n> > > to roll up in a sequence that is amortized-linear. I.e., we are not\n> > > always rolling up everything outside of the base pack, but trying to\n> > > roll up little into medium, and then eventually medium into large. If\n> > > you roll up things that are \"too big\", then you end up rewriting the\n> > > bytes more often, and your amount of work becomes super-linear.\n> > \n> > I'm not sure I follow, it would seem to me that it would stay linear, and\n> > be at most rewriting each new packfile once more than previously? Are you\n> > envisioning more work than that?\n> \n> Maybe I don't understand what you're proposing.\n> \n> The idea of the geometric repack is that by sorting by size and then\n> finding a \"cutoff\" within the size array, we can make sure that we roll\n> up a sufficiently small number of bytes in each roll-up that it ends up\n> linear in the size of the repo in the long run. But if we roll up\n> without regard to size, then our worst case is that the biggest pack is\n> the newest (imagine a repo with 10 small pushes and then one gigantic\n> one). So we roll that up with some small packs, doing effectively\n> O(size_of_repo) work.\n\nThis isn't quite a fair evaluation, it should be O(size_of_push) I think?\n\n> And then in the next roll up we do it again, and so on. \n \nI should have clarified that the intent is to prevent this by specifying an \nmtime after the last rollup so that this should only ever happen once for new \npackfiles. It also means you probably need special logic to ensure this roll-up \ndoesn't happen if there would only be one file in the rollup, \n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"417764","messageId":"xmqqv9ahxddp.fsf@gitster.g","threadId":"55012","inReplyTo":"51f57d5da23244ebde27ad6c14cbf4b63da3317d.1614047097.git.me@ttaylorr.com","subject":"Re: [PATCH v4 8/8] builtin/repack.c: add '--geometric' option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-24T23:19:30Z","receivedAt":"2021-02-24T23:20:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> Concretely, say that a repository has 'n' packfiles, labeled P1, P2,\n> ..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\n> With a geometric factor of 'r', it should be that:\n>\n>   objects(Pi) > r*objects(P(i-1))\n>\n> for all i in [1, n], where the packs are sorted by\n>\n>   objects(P1) <= objects(P2) <= ... <= objects(Pn).\n>\n> Since finding a true optimal repacking is NP-hard, we approximate it\n> along two directions:\n>\n>   1. We assume that there is a cutoff of packs _before starting the\n>      repack_ where everything to the right of that cut-off already forms\n>      a geometric progression (or no cutoff exists and everything must be\n>      repacked).\n\nWhen you order existing packs like how you explained the next\n\"direction\" below, do we assume loose ones would sit before\n(i.e. \"newer and smaller\" than) all of the packs?\n\n>   2. We assume that everything smaller than the cutoff count must be\n>      repacked. This forms our base assumption, but it can also cause\n>      even the \"heavy\" packs to get repacked, for e.g., if we have 6\n>      packs containing the following number of objects:\n>\n>        1, 1, 1, 2, 4, 32\n>\n>      then we would place the cutoff between '1, 1' and '1, 2, 4, 32',\n>      rolling up the first two packs into a pack with 2 objects. That\n>      breaks our progression and leaves us:\n>\n>        2, 1, 2, 4, 32\n>          ^\n>\n>      (where the '^' indicates the position of our split). To restore a\n>      progression, we move the split forward (towards larger packs)\n>      joining each pack into our new pack until a geometric progression\n>      is restored. Here, that looks like:\n>\n>        2, 1, 2, 4, 32  ~>  3, 2, 4, 32  ~>  5, 4, 32  ~> ... ~> 9, 32\n>          ^                   ^                ^                   ^\n\nThis explanation is very intuitive and easy to understand (I assume\nwe aren't actually repacking 1+1 into 2 and then 2+1 into 3 and then\nchoosing to repack 3+2 to create 5, but we scan before doing any\nrepacking and decide to repack 2+1+2+4 into a single 9).\n\nWhat is not so clear is how this picture changes depending on the\nvalue of 'r'.\n\n> ... Another wrinkle\n> is that we assume that loose, indexed, and reflog'd objects are\n> insignificant, and lump them into any new pack that we create.\n\nIn the example of 2 above, these are treated as insignificant\ncompared to the first '1' in the 1+1+1+2+4+32, so the choice of\nrepacked packs are made by computing 1+1+1+2+4 and noticing that is\nwhere we should stop, but we pack these insignificant ones together\nwith these repacked packs into the new pack that is supposed to\ncontain \"9\" objects?\n\n> This can\n> lead to non-idempotent results.\n\nLet me try to follow aloud to see if I got this right.\n\nIf we start from 1+1+1+2+4+32+... (similarly to the example given to\nexplain 2 above, but with more larger packs---but the assumption\nhere is that everything larger than 32 is already in good\nprogression), depending on how many loose objects we have, the\nresult of packing 1+1+1+2+4+loose might not necessarily be 9 but 100\n(collecting too many loose objects), and the set of packs would be\n32+... (from before the \"repack -g\") plus a 100-object pack, not\n9+32+... as the above explanation for 2 suggested.  Starting from\nthat state, re-running \"repack -g\" again would then have to repack\nthe packs existed before the first repack (i.e. 32+...) into one.\nIn other words, the second \"git repack -g\" in back-to-back \"git\nrepack -g && git repack -g\" may necessarily be a no-op.\n\nIs that what you meant by non-idempotent?\n\n> diff --git a/builtin/repack.c b/builtin/repack.c\n> index 01440de2d5..bcf280b10d 100644\n> --- a/builtin/repack.c\n> +++ b/builtin/repack.c\n> @@ -297,6 +297,124 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n>  #define ALL_INTO_ONE 1\n>  #define LOOSEN_UNREACHABLE 2\n>  \n> +struct pack_geometry {\n> +\tstruct packed_git **pack;\n> +\tuint32_t pack_nr, pack_alloc;\n> +\tuint32_t split;\n> +};\n> +\n> +static uint32_t geometry_pack_weight(struct packed_git *p)\n> +{\n> +\tif (open_pack_index(p))\n> +\t\tdie(_(\"cannot open index for %s\"), p->pack_name);\n> +\treturn p->num_objects;\n> +}\n> +\n> +static int geometry_cmp(const void *va, const void *vb)\n> +{\n> +\tuint32_t aw = geometry_pack_weight(*(struct packed_git **)va),\n> +\t\t bw = geometry_pack_weight(*(struct packed_git **)vb);\n> +\n> +\tif (aw < bw)\n> +\t\treturn -1;\n> +\tif (aw > bw)\n> +\t\treturn 1;\n> +\treturn 0;\n> +}\n> +\n> +static void init_pack_geometry(struct pack_geometry **geometry_p)\n> +{\n> +\tstruct packed_git *p;\n> +\tstruct pack_geometry *geometry;\n> +\n> +\t*geometry_p = xcalloc(1, sizeof(struct pack_geometry));\n> +\tgeometry = *geometry_p;\n> +\n> +\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n> +\t\tif (!pack_kept_objects && p->pack_keep)\n> +\t\t\tcontinue;\n> +\n> +\t\tALLOC_GROW(geometry->pack,\n> +\t\t\t   geometry->pack_nr + 1,\n> +\t\t\t   geometry->pack_alloc);\n> +\n> +\t\tgeometry->pack[geometry->pack_nr] = p;\n> +\t\tgeometry->pack_nr++;\n> +\t}\n> +\n> +\tQSORT(geometry->pack, geometry->pack_nr, geometry_cmp);\n> +}\n\nAfter calling this helper, we get geometry->pack[] that is sorted by\nthe number of objects in each pack, packs with fewer objects sort\nbefore the ones with more objects.  OK.\n\n> +static void split_pack_geometry(struct pack_geometry *geometry, int factor)\n> +{\n> +\tuint32_t i;\n> +\tuint32_t split;\n> +\toff_t total_size = 0;\n> +\n> +\tif (geometry->pack_nr <= 1) {\n> +\t\tgeometry->split = geometry->pack_nr;\n> +\t\treturn;\n> +\t}\n\nWhen there is a single pack (or no pack), we place the split to 1\n(let's keep reading with the need to find out what split means in\nmind; it is not yet clear if it points at the pack that will be part\nof the kept set, or at the pack that is the last one among the\nrepacked set, at this point in the code).\n\n> +\tsplit = geometry->pack_nr - 1;\n> +\n> +\t/*\n> +\t * First, count the number of packs (in descending order of size) which\n> +\t * already form a geometric progression.\n> +\t */\n> +\tfor (i = geometry->pack_nr - 1; i > 0; i--) {\n> +\t\tstruct packed_git *ours = geometry->pack[i];\n> +\t\tstruct packed_git *prev = geometry->pack[i - 1];\n> +\t\tif (geometry_pack_weight(ours) >= factor * geometry_pack_weight(prev))\n> +\t\t\tsplit--;\n> +\t\telse\n> +\t\t\tbreak;\n> +\t}\n\nInstead of rolling up from smaller ones like explained in the log\nmessage, we scan from the larger end and see where the existing\nprogression is broken.  When the loop breaks in the middle, the pack\nat position 'i-1' (prev) is too big.\n\nWhy do we need to initialize 'split' before the loop and decrement\nit?  Wouldn't it be equivalent to assign 'i' after the loop breaks\nto 'split'?\n\nIn any case, after the loop breaks, the packs starting at position\n'i+1' (one after ours when the loop broke) thru to the end of the\ngeometry->pack[] array are in good progression.  We have 'i' in\n'split' at this point, so ...\n\n> +\tif (split) {\n> +\t\t/*\n> +\t\t * Move the split one to the right, since the top element in the\n> +\t\t * last-compared pair can't be in the progression. Only do this\n> +\t\t * when we split in the middle of the array (otherwise if we got\n> +\t\t * to the end, then the split is in the right place).\n> +\t\t */\n> +\t\tsplit++;\n> +\t}\n\n... we increment it.  It means geometry->pack[split] is small enough\nrelative to geometry->pack[split+1] and so on thru to the end of the\narray.\n\nWhat if split==0 when we exited the loop?  That would mean that the\neverything in the array was in good progression, which is in line\nwith the \"in the middle\" case.  Either way, the pack at 'split' and\nlater are in good progression.\n\n> +\t/*\n> +\t * Then, anything to the left of 'split' must be in a new pack. But,\n> +\t * creating that new pack may cause packs in the heavy half to no longer\n> +\t * form a geometric progression.\n> +\t *\n> +\t * Compute an expected size of the new pack, and then determine how many\n> +\t * packs in the heavy half need to be joined into it (if any) to restore\n> +\t * the geometric progression.\n> +\t */\n> +\tfor (i = 0; i < split; i++)\n> +\t\ttotal_size += geometry_pack_weight(geometry->pack[i]);\n\nWe guestimate the number of objects in the rolled-up pack to be\ncreated.  Some objects may appear in multiple packs, but the number\nof them ought to be insignificant.  OK.\n\n> +\tfor (i = split; i < geometry->pack_nr; i++) {\n> +\t\tstruct packed_git *ours = geometry->pack[i];\n> +\t\tif (geometry_pack_weight(ours) < factor * total_size) {\n\nIf the pack at the bottom end of the range we previously thought to\nkeep turns out to be too small, we'd also roll that one in, by\nshifting the split point to the right.  And of course we update the\nexpected size of the new pack.  OK.\n\n> +\t\t\tsplit++;\n> +\t\t\ttotal_size += geometry_pack_weight(ours);\n> +\t\t} else\n> +\t\t\tbreak;\n> +\t}\n> +\n> +\tgeometry->split = split;\n\nThe code makes me wonder if we can compute all of the above in a\nsingle pass, but that is purely an intellectual curiosity.  The\nlogic in the code is crystal clear (the \"what if everything was\nalready in a good progression\" case was the only part that made me\nstop and think about the correctness of the logic) and the\nimplementation looks good, except for a few small nits:\n\n - why initialize 'split' so early before the first loop, which I\n   already mentioned.\n\n - we know many numbers are in uint32_t because that is how\n   packfiles limit their contents, but is it safe to perform the\n   multiplication with factor and comparison in that type?\n\n>  int cmd_repack(int argc, const char **argv, const char *prefix)\n>  {\n>  \tstruct child_process cmd = CHILD_PROCESS_INIT;\n> @@ -304,6 +422,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \tstruct string_list names = STRING_LIST_INIT_DUP;\n>  \tstruct string_list rollback = STRING_LIST_INIT_NODUP;\n>  \tstruct string_list existing_packs = STRING_LIST_INIT_DUP;\n> +\tstruct pack_geometry *geometry = NULL;\n>  \tstruct strbuf line = STRBUF_INIT;\n>  \tint i, ext, ret;\n>  \tFILE *out;\n> @@ -316,6 +435,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n>  \tint no_update_server_info = 0;\n>  \tstruct pack_objects_args po_args = {NULL};\n> +\tint geometric_factor = 0;\n>  \n>  \tstruct option builtin_repack_options[] = {\n>  \t\tOPT_BIT('a', NULL, &pack_everything,\n> @@ -356,6 +476,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \t\t\t\tN_(\"repack objects in packs marked with .keep\")),\n>  \t\tOPT_STRING_LIST(0, \"keep-pack\", &keep_pack_list, N_(\"name\"),\n>  \t\t\t\tN_(\"do not repack this pack\")),\n> +\t\tOPT_INTEGER('g', \"geometric\", &geometric_factor,\n> +\t\t\t    N_(\"find a geometric progression with factor <N>\")),\n>  \t\tOPT_END()\n>  \t};\n>  \n> @@ -382,6 +504,13 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \tif (write_bitmaps && !(pack_everything & ALL_INTO_ONE))\n>  \t\tdie(_(incremental_bitmap_conflict_error));\n>  \n> +\tif (geometric_factor) {\n> +\t\tif (pack_everything)\n> +\t\t\tdie(_(\"--geometric is incompatible with -A, -a\"));\n> +\t\tinit_pack_geometry(&geometry);\n> +\t\tsplit_pack_geometry(geometry, geometric_factor);\n> +\t}\n> +\n>  \tpackdir = mkpathdup(\"%s/pack\", get_object_directory());\n>  \tpacktmp = mkpathdup(\"%s/.tmp-%d-pack\", packdir, (int)getpid());\n>  \n> @@ -396,9 +525,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s\",\n>  \t\t\t     keep_pack_list.items[i].string);\n>  \tstrvec_push(&cmd.args, \"--non-empty\");\n> -\tstrvec_push(&cmd.args, \"--all\");\n> -\tstrvec_push(&cmd.args, \"--reflog\");\n> -\tstrvec_push(&cmd.args, \"--indexed-objects\");\n> +\tif (!geometry) {\n> +\t\t/*\n> +\t\t * 'git pack-objects' will up all objects loose or packed\n\n\"git pack-objects --stdin-packs\" will?\nWhat verb is missing in \"will VERB up all objects\"?\n\n> +\t\t * (either rolling them up or leaving them alone), so don't pass\n> +\t\t * these options.\n> +\t\t *\n> +\t\t * The implementation of 'git pack-objects --stdin-packs'\n> +\t\t * makes them redundant (and the two are incompatible).\n\nI am not sure if that is true.\n\nMore importantly, if you read this comment after you are done with\nthe series and no longer feel that geometric repacking is the most\nimportant thing in the world, you'd realize that an important piece\nof information is missing to help readers.  It talks about what\n\"geometric\" code does (i.e. uses --stdin-packs hence no need to pass\nthese options) in a block that is for !geometric.\n\n\tWe need to grab all reachable objects, including those that\n\tare reachable from reflogs and the index.\n\n\tWhen repacking into a geometric progression of packs,\n\thowever, we ask 'git pack-objects --stdin-packs', and it is\n\tnot about packing objects based on reachability but about\n\trepacking all the objects in specified packs and loose ones\n\t(indeed, --stdin-packs is incompatible with these options).\n\nor something?  I suspect that --stdin-packs does not make --all and\nothers \"redundant\".  The operation is about creating a new pack out\nof the objects contained in these packs, regardless of the objects'\nreachability from the usual \"refs, index and reflogs\" anchor points,\nno?\n\n> +\t\t */\n> +\t\tstrvec_push(&cmd.args, \"--all\");\n> +\t\tstrvec_push(&cmd.args, \"--reflog\");\n> +\t\tstrvec_push(&cmd.args, \"--indexed-objects\");\n> +\t}\n>  \tif (has_promisor_remote())\n>  \t\tstrvec_push(&cmd.args, \"--exclude-promisor-objects\");\n>  \tif (write_bitmaps > 0)\n> @@ -429,17 +568,37 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \t\t\t\tstrvec_push(&cmd.env_array, \"GIT_REF_PARANOIA=1\");\n>  \t\t\t}\n>  \t\t}\n> +\t} else if (geometry) {\n> +\t\tstrvec_push(&cmd.args, \"--stdin-packs\");\n> +\t\tstrvec_push(&cmd.args, \"--unpacked\");\n>  \t} else {\n>  \t\tstrvec_push(&cmd.args, \"--unpacked\");\n>  \t\tstrvec_push(&cmd.args, \"--incremental\");\n>  \t}\n>  \n> -\tcmd.no_stdin = 1;\n> +\tif (geometry)\n> +\t\tcmd.in = -1;\n> +\telse\n> +\t\tcmd.no_stdin = 1;\n\nIt is a bit sad that we need to do this before start_command() in\nthat the code structure does not make it clear why two modes have\ndifferent handling of the standard input stream, but I do not think\nof anything better, so I'll let it pass.\n\n>  \tret = start_command(&cmd);\n>  \tif (ret)\n>  \t\treturn ret;\n>  \n> +\tif (geometry) {\n> +\t\tFILE *in = xfdopen(cmd.in, \"w\");\n> +\t\t/*\n> +\t\t * The resulting pack should contain all objects in packs that\n> +\t\t * are going to be rolled up, but exclude objects in packs which\n> +\t\t * are being left alone.\n> +\t\t */\n> +\t\tfor (i = 0; i < geometry->split; i++)\n> +\t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n> +\t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n> +\t\t\tfprintf(in, \"^%s\\n\", pack_basename(geometry->pack[i]));\n> +\t\tfclose(in);\n> +\t}\n> +\n>  \tout = xfdopen(cmd.out, \"r\");\n>  \twhile (strbuf_getline_lf(&line, out) != EOF) {\n>  \t\tif (line.len != the_hash_algo->hexsz)\n> @@ -507,6 +666,25 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \t\t\tif (!string_list_has_string(&names, sha1))\n>  \t\t\t\tremove_redundant_pack(packdir, item->string);\n>  \t\t}\n> +\n> +\t\tif (geometry) {\n> +\t\t\tstruct strbuf buf = STRBUF_INIT;\n> +\n> +\t\t\tuint32_t i;\n> +\t\t\tfor (i = 0; i < geometry->split; i++) {\n> +\t\t\t\tstruct packed_git *p = geometry->pack[i];\n> +\t\t\t\tif (string_list_has_string(&names,\n> +\t\t\t\t\t\t\t   hash_to_hex(p->hash)))\n> +\t\t\t\t\tcontinue;\n> +\n> +\t\t\t\tstrbuf_reset(&buf);\n> +\t\t\t\tstrbuf_addstr(&buf, pack_basename(p));\n> +\t\t\t\tstrbuf_strip_suffix(&buf, \".pack\");\n> +\n> +\t\t\t\tremove_redundant_pack(packdir, buf.buf);\n> +\t\t\t}\n> +\t\t\tstrbuf_release(&buf);\n> +\t\t}\n\nBefore this new code, we seem to remove all pre-existing packfiles\nthat are not in the output from the pack-objects already.  The only\nreason that code does not harm the geometry case is we assume\nget_non_kept_pack_filenames() call is never made while doing\ngeometric repack (iow, ALL_INTO_ONE is not set) and the list of\npre-existing packfiles &existing_packs is empty.  Am I reading the\ncode correctly?\n\n - It is a bit unnerving to learn (and it will be a maintenance\n   burden in the future) that a variable whose name is\n   existing_packs does not necessarily have a list of existing packs\n   depending on the mode we are operating in.\n\n - The guard to make geometric incompatible with ALL_INTO_ONE does\n   not mention ALL_INTO_ONE, even though that bit is what would\n   corrupt the resulting repository if overlooked.  We should\n   probably need s/pack_everything/& \\& ALL_INTO_ONE/ in the hunk\n    below.\n\n> @@ -382,6 +504,13 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n>  \tif (write_bitmaps && !(pack_everything & ALL_INTO_ONE))\n>  \t\tdie(_(incremental_bitmap_conflict_error));\n>  \n> +\tif (geometric_factor) {\n> +\t\tif (pack_everything)\n> +\t\t\tdie(_(\"--geometric is incompatible with -A, -a\"));\n> +\t\tinit_pack_geometry(&geometry);\n> +\t\tsplit_pack_geometry(geometry, geometric_factor);\n> +\t}\n> +\n>  \tpackdir = mkpathdup(\"%s/pack\", get_object_directory());\n>  \tpacktmp = mkpathdup(\"%s/.tmp-%d-pack\", packdir, (int)getpid());\n>  \n\nOther than that, it was a fun patch to read.\n\nThanks.\n"},{"id":"417766","messageId":"xmqqo8g9xc9l.fsf@gitster.g","threadId":"55012","inReplyTo":"xmqqv9ahxddp.fsf@gitster.g","subject":"Re: [PATCH v4 8/8] builtin/repack.c: add '--geometric' option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-24T23:43:34Z","receivedAt":"2021-02-24T23:44:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Let me try to follow aloud to see if I got this right.\n>\n> If we start from 1+1+1+2+4+32+... (similarly to the example given to\n> explain 2 above, but with more larger packs---but the assumption\n> here is that everything larger than 32 is already in good\n> progression), depending on how many loose objects we have, the\n> result of packing 1+1+1+2+4+loose might not necessarily be 9 but 100\n> (collecting too many loose objects), and the set of packs would be\n> 32+... (from before the \"repack -g\") plus a 100-object pack, not\n> 9+32+... as the above explanation for 2 suggested.  Starting from\n> that state, re-running \"repack -g\" again would then have to repack\n> the packs existed before the first repack (i.e. 32+...) into one.\n> In other words, the second \"git repack -g\" in back-to-back \"git\n> repack -g && git repack -g\" may necessarily be a no-op.\n\n\"... may not necessarily be a no-op\" is what I should have typed here.\n\n> Is that what you meant by non-idempotent?\n\nAnd I think it makes sense for the repack to be non-idempotent.\nOnce we have packs in good progression, it is the only way to make\nprogress by keep rolling loose objects up into the smallest pack\nuntil it grows larger than the geometry factor allows it to be\nrelative to the next smallest pack.\n\n"},{"id":"417872","messageId":"YDiT2KeTYKZPamz8@coredump.intra.peff.net","threadId":"55012","inReplyTo":"4678947.aH8v75nzy7@mfick-lnx","subject":"Re: [PATCH v4 0/8] repack: support repacking into a geometric sequence","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-26T06:23:20Z","receivedAt":"2021-02-26T06:24:04Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 24, 2021 at 11:13:53AM -0700, Martin Fick wrote:\n\n> > The idea of the geometric repack is that by sorting by size and then\n> > finding a \"cutoff\" within the size array, we can make sure that we roll\n> > up a sufficiently small number of bytes in each roll-up that it ends up\n> > linear in the size of the repo in the long run. But if we roll up\n> > without regard to size, then our worst case is that the biggest pack is\n> > the newest (imagine a repo with 10 small pushes and then one gigantic\n> > one). So we roll that up with some small packs, doing effectively\n> > O(size_of_repo) work.\n> \n> This isn't quite a fair evaluation, it should be O(size_of_push) I think?\n\nSorry, I had a longer example, but then cut it down in the name of\nsimplicity. But I think I made it too simple. :)\n\nYou can imagine more pushes after the gigantic one, in which case we'd\nroll them up with the gigantic push. So that gigantic one is part of\nmultiple sequential rollups, until it is itself rolled up further.\n\nBut...\n\n> > And then in the next roll up we do it again, and so on. \n>  \n> I should have clarified that the intent is to prevent this by specifying an \n> mtime after the last rollup so that this should only ever happen once for new \n> packfiles. It also means you probably need special logic to ensure this roll-up \n> doesn't happen if there would only be one file in the rollup, \n\nYes, I agree that if you record a cut point, and then avoid rolling up\nacross it, then you'd only consider the single push once. You probably\nwant to record the actual pack set rather than just an mtime cutoff,\nthough, since Git will update the mtime on packs sometimes (to freshen\nthem whenever it optimizes out an object write for an object in the\npack).\n\nOne of the nice things about looking only at the pack sizes is that you\ndon't have to record that cut point. :) But it's possible you'd want to\nfor other reasons (e.g., you may spend extra work to find good deltas in\nyour on-disk packs, so you want to know what is old and what is new in\norder to discard on-disk deltas from pushed-up packs).\n\n-Peff\n"},{"id":"418263","messageId":"YEFT2P2DxzlT3/+t@nand.local","threadId":"55012","inReplyTo":"xmqqo8g9xc9l.fsf@gitster.g","subject":"Re: [PATCH v4 8/8] builtin/repack.c: add '--geometric' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-03-04T21:40:40Z","receivedAt":"2021-03-04T21:41:59Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 24, 2021 at 03:43:34PM -0800, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > Let me try to follow aloud to see if I got this right.\n> >\n> > If we start from 1+1+1+2+4+32+... (similarly to the example given to\n> > explain 2 above, but with more larger packs---but the assumption\n> > here is that everything larger than 32 is already in good\n> > progression), depending on how many loose objects we have, the\n> > result of packing 1+1+1+2+4+loose might not necessarily be 9 but 100\n> > (collecting too many loose objects), and the set of packs would be\n> > 32+... (from before the \"repack -g\") plus a 100-object pack, not\n> > 9+32+... as the above explanation for 2 suggested.  Starting from\n> > that state, re-running \"repack -g\" again would then have to repack\n> > the packs existed before the first repack (i.e. 32+...) into one.\n> > In other words, the second \"git repack -g\" in back-to-back \"git\n> > repack -g && git repack -g\" may necessarily be a no-op.\n>\n> \"... may not necessarily be a no-op\" is what I should have typed here.\n\nExactly.\n\n> > Is that what you meant by non-idempotent?\n>\n> And I think it makes sense for the repack to be non-idempotent.\n> Once we have packs in good progression, it is the only way to make\n> progress by keep rolling loose objects up into the smallest pack\n> until it grows larger than the geometry factor allows it to be\n> relative to the next smallest pack.\n\nRight again. It *would* be idempotent if we didn't push any new objects\ninto the repository (and repacked it with the same geometric factor once\nmore to clean up any inconsistencies after creating a pack with loose\nobjects), which is what you'd expect.\n\nOf course, pushing new objects into the repository means that the\nprogression will either grow (i.e., because the smallest pack in an\nexisting progression was quite large, and so we have some space to grow\nsmaller packs before rolling up the larger one), or it will get rolled\nup.\n\nThanks,\nTaylor\n"},{"id":"418264","messageId":"YEFXRwyMpyXHgArH@nand.local","threadId":"55012","inReplyTo":"xmqqv9ahxddp.fsf@gitster.g","subject":"Re: [PATCH v4 8/8] builtin/repack.c: add '--geometric' option","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-03-04T21:55:19Z","receivedAt":"2021-03-04T22:01:48Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 24, 2021 at 03:19:30PM -0800, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > Concretely, say that a repository has 'n' packfiles, labeled P1, P2,\n> > ..., up to Pn. Each packfile has an object count equal to 'objects(Pn)'.\n> > With a geometric factor of 'r', it should be that:\n> >\n> >   objects(Pi) > r*objects(P(i-1))\n> >\n> > for all i in [1, n], where the packs are sorted by\n> >\n> >   objects(P1) <= objects(P2) <= ... <= objects(Pn).\n> >\n> > Since finding a true optimal repacking is NP-hard, we approximate it\n> > along two directions:\n> >\n> >   1. We assume that there is a cutoff of packs _before starting the\n> >      repack_ where everything to the right of that cut-off already forms\n> >      a geometric progression (or no cutoff exists and everything must be\n> >      repacked).\n>\n> When you order existing packs like how you explained the next\n> \"direction\" below, do we assume loose ones would sit before\n> (i.e. \"newer and smaller\" than) all of the packs?\n\nKind of. We don't consider them to be part of any pack when deciding\nwhere to place the split (in other words, we don't consider them at\nall until the subsequent repack by which time they are packed).\n\nThat's a fine assumption to make (as you note in the reply below this\none), since we'll eventually reach a geometric progression. This\napproximation can be as wrong as there are loose objects (but hopefully\nthere aren't so many by the time we want to do a geometric repack).\n\n> >   2. We assume that everything smaller than the cutoff count must be\n> >      repacked. This forms our base assumption, but it can also cause\n> >      even the \"heavy\" packs to get repacked, for e.g., if we have 6\n> >      packs containing the following number of objects:\n> >\n> >        1, 1, 1, 2, 4, 32\n> >\n> >      then we would place the cutoff between '1, 1' and '1, 2, 4, 32',\n> >      rolling up the first two packs into a pack with 2 objects. That\n> >      breaks our progression and leaves us:\n> >\n> >        2, 1, 2, 4, 32\n> >          ^\n> >\n> >      (where the '^' indicates the position of our split). To restore a\n> >      progression, we move the split forward (towards larger packs)\n> >      joining each pack into our new pack until a geometric progression\n> >      is restored. Here, that looks like:\n> >\n> >        2, 1, 2, 4, 32  ~>  3, 2, 4, 32  ~>  5, 4, 32  ~> ... ~> 9, 32\n> >          ^                   ^                ^                   ^\n>\n> This explanation is very intuitive and easy to understand (I assume\n> we aren't actually repacking 1+1 into 2 and then 2+1 into 3 and then\n> choosing to repack 3+2 to create 5, but we scan before doing any\n> repacking and decide to repack 2+1+2+4 into a single 9).\n\nCorrect, and thanks. The split is determined ahead of time before we\nactually get to writing any new packs.\n\n> What is not so clear is how this picture changes depending on the\n> value of 'r'.\n\nIt only means that subsequent packs need to contain at least 'r' times\nas many objects as the previous pack does.\n\n> > +static void split_pack_geometry(struct pack_geometry *geometry, int factor)\n> > +{\n> > +\tuint32_t i;\n> > +\tuint32_t split;\n> > +\toff_t total_size = 0;\n> > +\n> > +\tif (geometry->pack_nr <= 1) {\n> > +\t\tgeometry->split = geometry->pack_nr;\n> > +\t\treturn;\n> > +\t}\n>\n> When there is a single pack (or no pack), we place the split to 1\n> (let's keep reading with the need to find out what split means in\n> mind; it is not yet clear if it points at the pack that will be part\n> of the kept set, or at the pack that is the last one among the\n> repacked set, at this point in the code).\n\nEverything that is strictly less than the split will get repacked, which\nupon reading this again means that we'll repack a repository containing\njust a single pack again. That's wasteful, so we may in the future want\nto adjust this to set the split to 0 regardless of whether we have zero\nor one pack here.\n\n> > +\tsplit = geometry->pack_nr - 1;\n> > +\n> > +\t/*\n> > +\t * First, count the number of packs (in descending order of size) which\n> > +\t * already form a geometric progression.\n> > +\t */\n> > +\tfor (i = geometry->pack_nr - 1; i > 0; i--) {\n> > +\t\tstruct packed_git *ours = geometry->pack[i];\n> > +\t\tstruct packed_git *prev = geometry->pack[i - 1];\n> > +\t\tif (geometry_pack_weight(ours) >= factor * geometry_pack_weight(prev))\n> > +\t\t\tsplit--;\n> > +\t\telse\n> > +\t\t\tbreak;\n> > +\t}\n>\n> Instead of rolling up from smaller ones like explained in the log\n> message, we scan from the larger end and see where the existing\n> progression is broken.  When the loop breaks in the middle, the pack\n> at position 'i-1' (prev) is too big.\n>\n> Why do we need to initialize 'split' before the loop and decrement\n> it?  Wouldn't it be equivalent to assign 'i' after the loop breaks\n> to 'split'?\n\nYep, they are equivalent.\n\n> In any case, after the loop breaks, the packs starting at position\n> 'i+1' (one after ours when the loop broke) thru to the end of the\n> geometry->pack[] array are in good progression.  We have 'i' in\n> 'split' at this point, so ...\n>\n> > +\tif (split) {\n> > +\t\t/*\n> > +\t\t * Move the split one to the right, since the top element in the\n> > +\t\t * last-compared pair can't be in the progression. Only do this\n> > +\t\t * when we split in the middle of the array (otherwise if we got\n> > +\t\t * to the end, then the split is in the right place).\n> > +\t\t */\n> > +\t\tsplit++;\n> > +\t}\n>\n> ... we increment it.  It means geometry->pack[split] is small enough\n> relative to geometry->pack[split+1] and so on thru to the end of the\n> array.\n>\n> What if split==0 when we exited the loop?  That would mean that the\n> everything in the array was in good progression, which is in line\n> with the \"in the middle\" case.  Either way, the pack at 'split' and\n> later are in good progression.\n\nRight (and ditto that we wouldn't do anything if split==0 in that case).\n\n>  - we know many numbers are in uint32_t because that is how\n>    packfiles limit their contents, but is it safe to perform the\n>    multiplication with factor and comparison in that type?\n\nWe could arguably be more careful here, yes.\n\n> > @@ -396,9 +525,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n> >  \t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s\",\n> >  \t\t\t     keep_pack_list.items[i].string);\n> >  \tstrvec_push(&cmd.args, \"--non-empty\");\n> > -\tstrvec_push(&cmd.args, \"--all\");\n> > -\tstrvec_push(&cmd.args, \"--reflog\");\n> > -\tstrvec_push(&cmd.args, \"--indexed-objects\");\n> > +\tif (!geometry) {\n> > +\t\t/*\n> > +\t\t * 'git pack-objects' will up all objects loose or packed\n>\n> \"git pack-objects --stdin-packs\" will?\n> What verb is missing in \"will VERB up all objects\"?\n\nLikely I meant to say \"roll\" here before \"up\".\n\n> > +\t\t * (either rolling them up or leaving them alone), so don't pass\n> > +\t\t * these options.\n> > +\t\t *\n> > +\t\t * The implementation of 'git pack-objects --stdin-packs'\n> > +\t\t * makes them redundant (and the two are incompatible).\n>\n> I am not sure if that is true.\n>\n> More importantly, if you read this comment after you are done with\n> the series and no longer feel that geometric repacking is the most\n> important thing in the world, you'd realize that an important piece\n> of information is missing to help readers.  It talks about what\n> \"geometric\" code does (i.e. uses --stdin-packs hence no need to pass\n> these options) in a block that is for !geometric.\n>\n> \tWe need to grab all reachable objects, including those that\n> \tare reachable from reflogs and the index.\n>\n> \tWhen repacking into a geometric progression of packs,\n> \thowever, we ask 'git pack-objects --stdin-packs', and it is\n> \tnot about packing objects based on reachability but about\n> \trepacking all the objects in specified packs and loose ones\n> \t(indeed, --stdin-packs is incompatible with these options).\n>\n> or something?  I suspect that --stdin-packs does not make --all and\n> others \"redundant\".  The operation is about creating a new pack out\n> of the objects contained in these packs, regardless of the objects'\n> reachability from the usual \"refs, index and reflogs\" anchor points,\n> no?\n\nExactly right. And I am certainly in favor of your wording above. Since\nthis series is already on next, I'd be happy to pick this up with the\nfew other minor things above in a separate series to apply on top (but\nsince I don't think any of these are correctness issues, you should feel\nfree to continue merging this down in the meantime).\n\n> > @@ -507,6 +666,25 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n> >  \t\t\tif (!string_list_has_string(&names, sha1))\n> >  \t\t\t\tremove_redundant_pack(packdir, item->string);\n> >  \t\t}\n> > +\n> > +\t\tif (geometry) {\n> > +\t\t\tstruct strbuf buf = STRBUF_INIT;\n> > +\n> > +\t\t\tuint32_t i;\n> > +\t\t\tfor (i = 0; i < geometry->split; i++) {\n> > +\t\t\t\tstruct packed_git *p = geometry->pack[i];\n> > +\t\t\t\tif (string_list_has_string(&names,\n> > +\t\t\t\t\t\t\t   hash_to_hex(p->hash)))\n> > +\t\t\t\t\tcontinue;\n> > +\n> > +\t\t\t\tstrbuf_reset(&buf);\n> > +\t\t\t\tstrbuf_addstr(&buf, pack_basename(p));\n> > +\t\t\t\tstrbuf_strip_suffix(&buf, \".pack\");\n> > +\n> > +\t\t\t\tremove_redundant_pack(packdir, buf.buf);\n> > +\t\t\t}\n> > +\t\t\tstrbuf_release(&buf);\n> > +\t\t}\n>\n> Before this new code, we seem to remove all pre-existing packfiles\n> that are not in the output from the pack-objects already.  The only\n> reason that code does not harm the geometry case is we assume\n> get_non_kept_pack_filenames() call is never made while doing\n> geometric repack (iow, ALL_INTO_ONE is not set) and the list of\n> pre-existing packfiles &existing_packs is empty.  Am I reading the\n> code correctly?\n>\n>  - It is a bit unnerving to learn (and it will be a maintenance\n>    burden in the future) that a variable whose name is\n>    existing_packs does not necessarily have a list of existing packs\n>    depending on the mode we are operating in.\n>\n>  - The guard to make geometric incompatible with ALL_INTO_ONE does\n>    not mention ALL_INTO_ONE, even though that bit is what would\n>    corrupt the resulting repository if overlooked.  We should\n>    probably need s/pack_everything/& \\& ALL_INTO_ONE/ in the hunk\n>     below.\n\nEek, yes. This is because the geometric code takes its own view of the\npack directory when figuring out where to place to split line, and so it\nseemed easier to have separate paths.\n\nI'm not sure whether I maintain that that was a good idea in hindsight\n;). Certainly it does create a little bit of a maintenance burden for\nus. But they really are two different things: the geometric code really\nwants to have the packs laid out in order of object size, while the\n\"existing\" string_list wants packs laid out in lexicographic order of\ntheir filename to check whether certain packs exist or not.\n\n> Other than that, it was a fun patch to read.\n\nThanks, I think the few suggestions you made here are good ones. I'll\nput it on my to-do list of things to clean up in a separate little\nseries.\n\nSince this is already in next, I would suggest continuing to merge it\ndown since none of these suggestions impact the patch's correctness.\n\nThanks,\nTaylor\n"}]}