{"thread":{"id":"64911","subject":"[PATCH 0/5] builtin/repo: include largest object information","startedAt":"2026-02-03T22:18:35Z","lastAt":"2026-03-08T18:44:11Z","messageCount":50,"participants":["Justin Tobler","Junio C Hamano","Kristoffer Haugsbakk","Patrick Steinhardt","Lucas Seiki Oshiro"],"isPatch":true,"patchVersion":1,"patchTotal":5},"messages":[{"id":"535099","messageId":"20260203221758.1164434-1-jltobler@gmail.com","threadId":"64911","inReplyTo":null,"subject":"[PATCH 0/5] builtin/repo: include largest object information","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-03T22:17:53Z","receivedAt":"2026-02-03T22:18:35Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThe \"structure\" output for git-repo(1) currently provides count\ninformation for references/objects as well as total inflated/disk sizes\nof objects by type. Info regarding the largest individual objects in the\nrepository is not yet collected, but would be useful to users wishing to\nidentify such large objects.\n\nThis patch series adds the following data points:\n- The OID and size of the largest objects by object type\n- The OID and parent count of the commit with the most parents\n- The OID and entries count of the tree with the most entries\n\nThanks,\n-Justin\n\nJustin Tobler (5):\n  builtin/repo: update stats for each object\n  builtin/repo: collect largest inflated objects\n  builtin/repo: add OID annotations to table output\n  builtin/repo: find commit with most parents\n  builtin/repo: find tree with most entries\n\n Documentation/git-repo.adoc |   1 +\n builtin/repo.c              | 249 +++++++++++++++++++++++++++++++-----\n t/t1901-repo-structure.sh   | 143 +++++++++++++--------\n 3 files changed, 313 insertions(+), 80 deletions(-)\n\n\nbase-commit: 67ad42147a7acc2af6074753ebd03d904476118f\n-- \n2.53.0\n\n"},{"id":"535100","messageId":"20260203221758.1164434-2-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260203221758.1164434-1-jltobler@gmail.com","subject":"[PATCH 1/5] builtin/repo: update stats for each object","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-03T22:17:54Z","receivedAt":"2026-02-03T22:18:35Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"When walking reachable objects in the repository, `count_objects()`\nprocesses a set of objects and updates the `struct object_stats`. In\npreparation for more granular statistics being collected, update the\n`struct object_stats` for each individual object instead.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 53 +++++++++++++++++++++++---------------------------\n 1 file changed, 24 insertions(+), 29 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 0ea045abc1..c7c9f0f497 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -558,8 +558,6 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n-\tsize_t inflated_total = 0;\n-\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n@@ -575,33 +573,30 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n \t\t\tcontinue;\n \n-\t\tinflated_total += inflated;\n-\t\tdisk_total += disk;\n-\t}\n-\n-\tswitch (type) {\n-\tcase OBJ_TAG:\n-\t\tstats->type_counts.tags += oids->nr;\n-\t\tstats->inflated_sizes.tags += inflated_total;\n-\t\tstats->disk_sizes.tags += disk_total;\n-\t\tbreak;\n-\tcase OBJ_COMMIT:\n-\t\tstats->type_counts.commits += oids->nr;\n-\t\tstats->inflated_sizes.commits += inflated_total;\n-\t\tstats->disk_sizes.commits += disk_total;\n-\t\tbreak;\n-\tcase OBJ_TREE:\n-\t\tstats->type_counts.trees += oids->nr;\n-\t\tstats->inflated_sizes.trees += inflated_total;\n-\t\tstats->disk_sizes.trees += disk_total;\n-\t\tbreak;\n-\tcase OBJ_BLOB:\n-\t\tstats->type_counts.blobs += oids->nr;\n-\t\tstats->inflated_sizes.blobs += inflated_total;\n-\t\tstats->disk_sizes.blobs += disk_total;\n-\t\tbreak;\n-\tdefault:\n-\t\tBUG(\"invalid object type\");\n+\t\tswitch (type) {\n+\t\tcase OBJ_TAG:\n+\t\t\tstats->type_counts.tags++;\n+\t\t\tstats->inflated_sizes.tags += inflated;\n+\t\t\tstats->disk_sizes.tags += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_COMMIT:\n+\t\t\tstats->type_counts.commits++;\n+\t\t\tstats->inflated_sizes.commits += inflated;\n+\t\t\tstats->disk_sizes.commits += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_TREE:\n+\t\t\tstats->type_counts.trees++;\n+\t\t\tstats->inflated_sizes.trees += inflated;\n+\t\t\tstats->disk_sizes.trees += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_BLOB:\n+\t\t\tstats->type_counts.blobs++;\n+\t\t\tstats->inflated_sizes.blobs += inflated;\n+\t\t\tstats->disk_sizes.blobs += disk;\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tBUG(\"invalid object type\");\n+\t\t}\n \t}\n \n \tobject_count = get_total_object_values(&stats->type_counts);\n-- \n2.53.0\n\n"},{"id":"535101","messageId":"20260203221758.1164434-3-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260203221758.1164434-1-jltobler@gmail.com","subject":"[PATCH 2/5] builtin/repo: collect largest inflated objects","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-03T22:17:55Z","receivedAt":"2026-02-03T22:18:36Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The \"structure\" output for git-repo(1) shows the total inflated and disk\nsizes of reachable objects in the repository, but doesn't show the size\nof the largest individual objects. Since an individual object may be a\nlarge contributor to the overall repository size, it is useful for users\nto know the maximum size of individual objects.\n\nWhile interating across objects, record the size and OID of the largest\nobjects encountered for each object type to provide as output. Note that\nthe default \"table\" output format only displays size information and not\nthe corresponding OID. In a subsequent commit, the table format is\nupdated to add table annotations that mention the OID.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 63 +++++++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   | 28 +++++++++++++++++\n 3 files changed, 92 insertions(+)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 7d70270dfa..e812e59158 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -52,6 +52,7 @@ supported:\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n * Total disk size of reachable objects by type\n+* Largest reachable objects in the repository by type\n +\n The output format can be chosen through the flag `--format`. Three formats are\n supported:\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex c7c9f0f497..51a4359685 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,7 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hash.h\"\n #include \"hex.h\"\n #include \"odb.h\"\n #include \"parse-options.h\"\n@@ -197,6 +198,18 @@ static int cmd_repo_info(int argc, const char **argv, const char *prefix,\n \t\treturn print_fields(argc, argv, repo, format);\n }\n \n+struct object_data {\n+\tstruct object_id oid;\n+\tsize_t value;\n+};\n+\n+struct largest_objects {\n+\tstruct object_data tag_size;\n+\tstruct object_data commit_size;\n+\tstruct object_data tree_size;\n+\tstruct object_data blob_size;\n+};\n+\n struct ref_stats {\n \tsize_t branches;\n \tsize_t remotes;\n@@ -215,6 +228,7 @@ struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n \tstruct object_values disk_sizes;\n+\tstruct largest_objects largest;\n };\n \n struct repo_structure {\n@@ -371,6 +385,21 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->disk_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tstats_table_addf(table, \"\");\n+\tstats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->largest.commit_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->largest.tree_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->largest.blob_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Tags\"));\n+\tstats_table_size_addf(table, objects->largest.tag_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -485,6 +514,23 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n \n+\tprintf(\"objects.commits.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.commit_size.value, value_delim);\n+\tprintf(\"objects.commits.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.commit_size.oid), value_delim);\n+\tprintf(\"objects.trees.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.tree_size.value, value_delim);\n+\tprintf(\"objects.trees.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.tree_size.oid), value_delim);\n+\tprintf(\"objects.blobs.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.blob_size.value, value_delim);\n+\tprintf(\"objects.blobs.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.blob_size.oid), value_delim);\n+\tprintf(\"objects.tags.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.tag_size.value, value_delim);\n+\tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -553,6 +599,15 @@ struct count_objects_data {\n \tstruct progress *progress;\n };\n \n+static void check_largest(struct object_data *data, struct object_id *oid,\n+\t\t\t  size_t value)\n+{\n+\tif (value > data->value) {\n+\t\toidcpy(&data->oid, oid);\n+\t\tdata->value = value;\n+\t}\n+}\n+\n static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t enum object_type type, void *cb_data)\n {\n@@ -578,21 +633,29 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tstats->type_counts.tags++;\n \t\t\tstats->inflated_sizes.tags += inflated;\n \t\t\tstats->disk_sizes.tags += disk;\n+\t\t\tcheck_largest(&stats->largest.tag_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_COMMIT:\n \t\t\tstats->type_counts.commits++;\n \t\t\tstats->inflated_sizes.commits += inflated;\n \t\t\tstats->disk_sizes.commits += disk;\n+\t\t\tcheck_largest(&stats->largest.commit_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tstats->type_counts.trees++;\n \t\t\tstats->inflated_sizes.trees += inflated;\n \t\t\tstats->disk_sizes.trees += disk;\n+\t\t\tcheck_largest(&stats->largest.tree_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_BLOB:\n \t\t\tstats->type_counts.blobs++;\n \t\t\tstats->inflated_sizes.blobs += inflated;\n \t\t\tstats->disk_sizes.blobs += disk;\n+\t\t\tcheck_largest(&stats->largest.blob_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 17ff164b05..1999f325d0 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -52,6 +52,16 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|                      |        |\n+\t\t| * Largest objects    |        |\n+\t\t|   * Commits          |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Trees            |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Blobs            |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Tags             |        |\n+\t\t|     * Maximum size   |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -104,6 +114,16 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Trees          | $(object_type_disk_usage tree true) |\n \t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n \t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n+\t\t|                      |            |\n+\t\t| * Largest objects    |            |\n+\t\t|   * Commits          |            |\n+\t\t|     * Maximum size   |    223 B   |\n+\t\t|   * Trees            |            |\n+\t\t|     * Maximum size   |  32.29 KiB |\n+\t\t|   * Blobs            |            |\n+\t\t|     * Maximum size   |     13 B   |\n+\t\t|   * Tags             |            |\n+\t\t|     * Maximum size   |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -138,6 +158,14 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n \t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n \t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n+\t\tobjects.commits.max_size=221\n+\t\tobjects.commits.max_size_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n+\t\tobjects.trees.max_size=1335\n+\t\tobjects.trees.max_size_oid=09931deea9d81ec21300d3e13c74412f32eacec5\n+\t\tobjects.blobs.max_size=11\n+\t\tobjects.blobs.max_size_oid=eaeeedced46482bd4281fda5a5f05ce24854151f\n+\t\tobjects.tags.max_size=132\n+\t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"535102","messageId":"20260203221758.1164434-4-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260203221758.1164434-1-jltobler@gmail.com","subject":"[PATCH 3/5] builtin/repo: add OID annotations to table output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-03T22:17:56Z","receivedAt":"2026-02-03T22:18:37Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The \"structure\" output for git-repo(1) does not show the corresponding\nOIDs for the largest objects in its \"table\" output. Update the output to\ninclude a list of OID annotations with an index to the corresponding row\nin the table.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            |  77 +++++++++++++++++---\n t/t1901-repo-structure.sh | 145 ++++++++++++++++++++------------------\n 2 files changed, 142 insertions(+), 80 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 51a4359685..6fc2d9db12 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -238,6 +238,7 @@ struct repo_structure {\n \n struct stats_table {\n \tstruct string_list rows;\n+\tstruct string_list annotations;\n \n \tint name_col_width;\n \tint value_col_width;\n@@ -250,6 +251,8 @@ struct stats_table {\n struct stats_table_entry {\n \tchar *value;\n \tconst char *unit;\n+\tsize_t index;\n+\tstruct object_id *oid;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -272,6 +275,12 @@ static void stats_table_vaddf(struct stats_table *table,\n \t\ttable->name_col_width = name_width;\n \tif (!entry)\n \t\treturn;\n+\tif (entry->oid) {\n+\t\tentry->index = table->annotations.nr + 1;\n+\t\tstrbuf_addf(&buf, \"[%\" PRIuMAX \"] %s\", (uintmax_t)entry->index,\n+\t\t\t    oid_to_hex(entry->oid));\n+\t\tstring_list_append(&table->annotations, buf.buf);\n+\t}\n \tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n@@ -282,6 +291,8 @@ static void stats_table_vaddf(struct stats_table *table,\n \t\tif (unit_width > table->unit_col_width)\n \t\t\ttable->unit_col_width = unit_width;\n \t}\n+\n+\tstrbuf_release(&buf);\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -321,6 +332,27 @@ static void stats_table_size_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_object_size_addf(struct stats_table *table,\n+\t\t\t\t\t struct object_id *oid, size_t value,\n+\t\t\t\t\t const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n+\n+\t/*\n+\t * A NULL OID should not have a table annotation.\n+\t */\n+\tif (!is_null_oid(oid))\n+\t\tentry->oid = oid;\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -389,19 +421,29 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Commits\"));\n-\tstats_table_size_addf(table, objects->largest.commit_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.commit_size.oid,\n+\t\t\t\t     objects->largest.commit_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n-\tstats_table_size_addf(table, objects->largest.tree_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.tree_size.oid,\n+\t\t\t\t     objects->largest.tree_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n-\tstats_table_size_addf(table, objects->largest.blob_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.blob_size.oid,\n+\t\t\t\t     objects->largest.blob_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Tags\"));\n-\tstats_table_size_addf(table, objects->largest.tag_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.tag_size.oid,\n+\t\t\t\t     objects->largest.tag_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n }\n \n+#define INDEX_WIDTH 4\n+\n static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n@@ -420,7 +462,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width + INDEX_WIDTH,\n+\t\t\t  name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n \t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n@@ -428,7 +471,7 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tprintf(\"%s\\n\", buf.buf);\n \n \tprintf(\"| \");\n-\tfor (int i = 0; i < name_col_width; i++)\n+\tfor (int i = 0; i < name_col_width + INDEX_WIDTH; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n \tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n@@ -450,6 +493,13 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_reset(&buf);\n \t\tstrbuf_addstr(&buf, \"| \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n+\n+\t\tif (entry && entry->oid)\n+\t\t\tstrbuf_addf(&buf, \" [%\" PRIuMAX \"]\",\n+\t\t\t\t    (uintmax_t)entry->index);\n+\t\telse\n+\t\t\tstrbuf_addchars(&buf, ' ', INDEX_WIDTH);\n+\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n \t\tstrbuf_addch(&buf, ' ');\n@@ -458,6 +508,11 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\n \n+\tif (table->annotations.nr)\n+\t\tprintf(\"\\n\");\n+\tfor_each_string_list_item(item, &table->annotations)\n+\t\tprintf(\"%s\\n\", item->string);\n+\n \tstrbuf_release(&buf);\n }\n \n@@ -473,6 +528,7 @@ static void stats_table_clear(struct stats_table *table)\n \t}\n \n \tstring_list_clear(&table->rows, 1);\n+\tstring_list_clear(&table->annotations, 1);\n }\n \n static void structure_keyvalue_print(struct repo_structure *stats,\n@@ -695,6 +751,7 @@ static int cmd_repo_structure(int argc, const char **argv, const char *prefix,\n {\n \tstruct stats_table table = {\n \t\t.rows = STRING_LIST_INIT_DUP,\n+\t\t.annotations = STRING_LIST_INIT_DUP,\n \t};\n \tenum output_format format = FORMAT_TABLE;\n \tstruct repo_structure stats = { 0 };\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 1999f325d0..918af7269f 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -27,41 +27,41 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    0   |\n-\t\t|     * Branches       |    0   |\n-\t\t|     * Tags           |    0   |\n-\t\t|     * Remotes        |    0   |\n-\t\t|     * Others         |    0   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            |    0   |\n-\t\t|     * Commits        |    0   |\n-\t\t|     * Trees          |    0   |\n-\t\t|     * Blobs          |    0   |\n-\t\t|     * Tags           |    0   |\n-\t\t|   * Inflated size    |    0 B |\n-\t\t|     * Commits        |    0 B |\n-\t\t|     * Trees          |    0 B |\n-\t\t|     * Blobs          |    0 B |\n-\t\t|     * Tags           |    0 B |\n-\t\t|   * Disk size        |    0 B |\n-\t\t|     * Commits        |    0 B |\n-\t\t|     * Trees          |    0 B |\n-\t\t|     * Blobs          |    0 B |\n-\t\t|     * Tags           |    0 B |\n-\t\t|                      |        |\n-\t\t| * Largest objects    |        |\n-\t\t|   * Commits          |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Trees            |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Blobs            |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Tags             |        |\n-\t\t|     * Maximum size   |    0 B |\n+\t\t| Repository structure     | Value  |\n+\t\t| ------------------------ | ------ |\n+\t\t| * References             |        |\n+\t\t|   * Count                |    0   |\n+\t\t|     * Branches           |    0   |\n+\t\t|     * Tags               |    0   |\n+\t\t|     * Remotes            |    0   |\n+\t\t|     * Others             |    0   |\n+\t\t|                          |        |\n+\t\t| * Reachable objects      |        |\n+\t\t|   * Count                |    0   |\n+\t\t|     * Commits            |    0   |\n+\t\t|     * Trees              |    0   |\n+\t\t|     * Blobs              |    0   |\n+\t\t|     * Tags               |    0   |\n+\t\t|   * Inflated size        |    0 B |\n+\t\t|     * Commits            |    0 B |\n+\t\t|     * Trees              |    0 B |\n+\t\t|     * Blobs              |    0 B |\n+\t\t|     * Tags               |    0 B |\n+\t\t|   * Disk size            |    0 B |\n+\t\t|     * Commits            |    0 B |\n+\t\t|     * Trees              |    0 B |\n+\t\t|     * Blobs              |    0 B |\n+\t\t|     * Tags               |    0 B |\n+\t\t|                          |        |\n+\t\t| * Largest objects        |        |\n+\t\t|   * Commits              |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Trees                |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Blobs                |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Tags                 |        |\n+\t\t|     * Maximum size       |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -89,41 +89,46 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# git-rev-list(1) --disk-usage=human option printing the full\n \t\t# \"byte/bytes\" unit string instead of just \"B\".\n \t\tcat >expect <<-EOF &&\n-\t\t| Repository structure | Value      |\n-\t\t| -------------------- | ---------- |\n-\t\t| * References         |            |\n-\t\t|   * Count            |      4     |\n-\t\t|     * Branches       |      1     |\n-\t\t|     * Tags           |      1     |\n-\t\t|     * Remotes        |      1     |\n-\t\t|     * Others         |      1     |\n-\t\t|                      |            |\n-\t\t| * Reachable objects  |            |\n-\t\t|   * Count            |   3.02 k   |\n-\t\t|     * Commits        |   1.01 k   |\n-\t\t|     * Trees          |   1.01 k   |\n-\t\t|     * Blobs          |   1.01 k   |\n-\t\t|     * Tags           |      1     |\n-\t\t|   * Inflated size    |  16.03 MiB |\n-\t\t|     * Commits        | 217.92 KiB |\n-\t\t|     * Trees          |  15.81 MiB |\n-\t\t|     * Blobs          |  11.68 KiB |\n-\t\t|     * Tags           |    132 B   |\n-\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n-\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n-\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n-\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n-\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n-\t\t|                      |            |\n-\t\t| * Largest objects    |            |\n-\t\t|   * Commits          |            |\n-\t\t|     * Maximum size   |    223 B   |\n-\t\t|   * Trees            |            |\n-\t\t|     * Maximum size   |  32.29 KiB |\n-\t\t|   * Blobs            |            |\n-\t\t|     * Maximum size   |     13 B   |\n-\t\t|   * Tags             |            |\n-\t\t|     * Maximum size   |    132 B   |\n+\t\t| Repository structure     | Value      |\n+\t\t| ------------------------ | ---------- |\n+\t\t| * References             |            |\n+\t\t|   * Count                |      4     |\n+\t\t|     * Branches           |      1     |\n+\t\t|     * Tags               |      1     |\n+\t\t|     * Remotes            |      1     |\n+\t\t|     * Others             |      1     |\n+\t\t|                          |            |\n+\t\t| * Reachable objects      |            |\n+\t\t|   * Count                |   3.02 k   |\n+\t\t|     * Commits            |   1.01 k   |\n+\t\t|     * Trees              |   1.01 k   |\n+\t\t|     * Blobs              |   1.01 k   |\n+\t\t|     * Tags               |      1     |\n+\t\t|   * Inflated size        |  16.03 MiB |\n+\t\t|     * Commits            | 217.92 KiB |\n+\t\t|     * Trees              |  15.81 MiB |\n+\t\t|     * Blobs              |  11.68 KiB |\n+\t\t|     * Tags               |    132 B   |\n+\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n+\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n+\t\t|                          |            |\n+\t\t| * Largest objects        |            |\n+\t\t|   * Commits              |            |\n+\t\t|     * Maximum size   [1] |    223 B   |\n+\t\t|   * Trees                |            |\n+\t\t|     * Maximum size   [2] |  32.29 KiB |\n+\t\t|   * Blobs                |            |\n+\t\t|     * Maximum size   [3] |     13 B   |\n+\t\t|   * Tags                 |            |\n+\t\t|     * Maximum size   [4] |    132 B   |\n+\n+\t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n+\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"535103","messageId":"20260203221758.1164434-5-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260203221758.1164434-1-jltobler@gmail.com","subject":"[PATCH 4/5] builtin/repo: find commit with most parents","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-03T22:17:57Z","receivedAt":"2026-02-03T22:18:37Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Complex merge events may produce an octopus merge where the resulting\nmerge commit has more than two parents. While iterating through objects\nin the repository for git-repo-structure, identify the commit with the\nmost parents and display it in the output.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            |  47 ++++++++++++\n t/t1901-repo-structure.sh | 151 ++++++++++++++++++++------------------\n 2 files changed, 125 insertions(+), 73 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 6fc2d9db12..dc1ac7ad3b 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -1,6 +1,7 @@\n #define USE_THE_REPOSITORY_VARIABLE\n \n #include \"builtin.h\"\n+#include \"commit.h\"\n #include \"environment.h\"\n #include \"hash.h\"\n #include \"hex.h\"\n@@ -208,6 +209,8 @@ struct largest_objects {\n \tstruct object_data commit_size;\n \tstruct object_data tree_size;\n \tstruct object_data blob_size;\n+\n+\tstruct object_data parent_count;\n };\n \n struct ref_stats {\n@@ -318,6 +321,27 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_object_count_addf(struct stats_table *table,\n+\t\t\t\t\t  struct object_id *oid, size_t value,\n+\t\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_count(value, &entry->value, &entry->unit);\n+\n+\t/*\n+\t * A NULL OID should not have a table annotation.\n+\t */\n+\tif (!is_null_oid(oid))\n+\t\tentry->oid = oid;\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static void stats_table_size_addf(struct stats_table *table, size_t value,\n \t\t\t\t  const char *format, ...)\n {\n@@ -425,6 +449,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t\t     &objects->largest.commit_size.oid,\n \t\t\t\t     objects->largest.commit_size.value,\n \t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_count_addf(table,\n+\t\t\t\t      &objects->largest.parent_count.oid,\n+\t\t\t\t      objects->largest.parent_count.value,\n+\t\t\t\t      \"    * %s\", _(\"Maximum parents\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n \tstats_table_object_size_addf(table,\n \t\t\t\t     &objects->largest.tree_size.oid,\n@@ -587,6 +615,11 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n \t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n \n+\tprintf(\"objects.commits.max_parents%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.parent_count.value, value_delim);\n+\tprintf(\"objects.commits.max_parents_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.parent_count.oid), value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -674,16 +707,24 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\tstruct commit *commit;\n+\t\tstruct object *obj;\n+\t\tvoid *content;\n \t\toff_t disk;\n+\t\tint eaten;\n \n \t\toi.sizep = &inflated;\n \t\toi.disk_sizep = &disk;\n+\t\toi.contentp = &content;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n \t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n \t\t\tcontinue;\n \n+\t\tobj = parse_object_buffer(the_repository, &oids->oid[i], type,\n+\t\t\t\t\t  inflated, content, &eaten);\n+\n \t\tswitch (type) {\n \t\tcase OBJ_TAG:\n \t\t\tstats->type_counts.tags++;\n@@ -693,11 +734,14 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_COMMIT:\n+\t\t\tcommit = object_as_type(obj, OBJ_COMMIT, 0);\n \t\t\tstats->type_counts.commits++;\n \t\t\tstats->inflated_sizes.commits += inflated;\n \t\t\tstats->disk_sizes.commits += disk;\n \t\t\tcheck_largest(&stats->largest.commit_size, &oids->oid[i],\n \t\t\t\t      inflated);\n+\t\t\tcheck_largest(&stats->largest.parent_count, &oids->oid[i],\n+\t\t\t\t      commit_list_count(commit->parents));\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tstats->type_counts.trees++;\n@@ -716,6 +760,9 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\tdefault:\n \t\t\tBUG(\"invalid object type\");\n \t\t}\n+\n+\t\tif (!eaten)\n+\t\t\tfree(content);\n \t}\n \n \tobject_count = get_total_object_values(&stats->type_counts);\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 918af7269f..d003d64a8e 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -27,41 +27,42 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure     | Value  |\n-\t\t| ------------------------ | ------ |\n-\t\t| * References             |        |\n-\t\t|   * Count                |    0   |\n-\t\t|     * Branches           |    0   |\n-\t\t|     * Tags               |    0   |\n-\t\t|     * Remotes            |    0   |\n-\t\t|     * Others             |    0   |\n-\t\t|                          |        |\n-\t\t| * Reachable objects      |        |\n-\t\t|   * Count                |    0   |\n-\t\t|     * Commits            |    0   |\n-\t\t|     * Trees              |    0   |\n-\t\t|     * Blobs              |    0   |\n-\t\t|     * Tags               |    0   |\n-\t\t|   * Inflated size        |    0 B |\n-\t\t|     * Commits            |    0 B |\n-\t\t|     * Trees              |    0 B |\n-\t\t|     * Blobs              |    0 B |\n-\t\t|     * Tags               |    0 B |\n-\t\t|   * Disk size            |    0 B |\n-\t\t|     * Commits            |    0 B |\n-\t\t|     * Trees              |    0 B |\n-\t\t|     * Blobs              |    0 B |\n-\t\t|     * Tags               |    0 B |\n-\t\t|                          |        |\n-\t\t| * Largest objects        |        |\n-\t\t|   * Commits              |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Trees                |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Blobs                |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Tags                 |        |\n-\t\t|     * Maximum size       |    0 B |\n+\t\t| Repository structure      | Value  |\n+\t\t| ------------------------- | ------ |\n+\t\t| * References              |        |\n+\t\t|   * Count                 |    0   |\n+\t\t|     * Branches            |    0   |\n+\t\t|     * Tags                |    0   |\n+\t\t|     * Remotes             |    0   |\n+\t\t|     * Others              |    0   |\n+\t\t|                           |        |\n+\t\t| * Reachable objects       |        |\n+\t\t|   * Count                 |    0   |\n+\t\t|     * Commits             |    0   |\n+\t\t|     * Trees               |    0   |\n+\t\t|     * Blobs               |    0   |\n+\t\t|     * Tags                |    0   |\n+\t\t|   * Inflated size         |    0 B |\n+\t\t|     * Commits             |    0 B |\n+\t\t|     * Trees               |    0 B |\n+\t\t|     * Blobs               |    0 B |\n+\t\t|     * Tags                |    0 B |\n+\t\t|   * Disk size             |    0 B |\n+\t\t|     * Commits             |    0 B |\n+\t\t|     * Trees               |    0 B |\n+\t\t|     * Blobs               |    0 B |\n+\t\t|     * Tags                |    0 B |\n+\t\t|                           |        |\n+\t\t| * Largest objects         |        |\n+\t\t|   * Commits               |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|     * Maximum parents     |    0   |\n+\t\t|   * Trees                 |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|   * Blobs                 |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|   * Tags                  |        |\n+\t\t|     * Maximum size        |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -89,46 +90,48 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# git-rev-list(1) --disk-usage=human option printing the full\n \t\t# \"byte/bytes\" unit string instead of just \"B\".\n \t\tcat >expect <<-EOF &&\n-\t\t| Repository structure     | Value      |\n-\t\t| ------------------------ | ---------- |\n-\t\t| * References             |            |\n-\t\t|   * Count                |      4     |\n-\t\t|     * Branches           |      1     |\n-\t\t|     * Tags               |      1     |\n-\t\t|     * Remotes            |      1     |\n-\t\t|     * Others             |      1     |\n-\t\t|                          |            |\n-\t\t| * Reachable objects      |            |\n-\t\t|   * Count                |   3.02 k   |\n-\t\t|     * Commits            |   1.01 k   |\n-\t\t|     * Trees              |   1.01 k   |\n-\t\t|     * Blobs              |   1.01 k   |\n-\t\t|     * Tags               |      1     |\n-\t\t|   * Inflated size        |  16.03 MiB |\n-\t\t|     * Commits            | 217.92 KiB |\n-\t\t|     * Trees              |  15.81 MiB |\n-\t\t|     * Blobs              |  11.68 KiB |\n-\t\t|     * Tags               |    132 B   |\n-\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n-\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n-\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n-\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n-\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n-\t\t|                          |            |\n-\t\t| * Largest objects        |            |\n-\t\t|   * Commits              |            |\n-\t\t|     * Maximum size   [1] |    223 B   |\n-\t\t|   * Trees                |            |\n-\t\t|     * Maximum size   [2] |  32.29 KiB |\n-\t\t|   * Blobs                |            |\n-\t\t|     * Maximum size   [3] |     13 B   |\n-\t\t|   * Tags                 |            |\n-\t\t|     * Maximum size   [4] |    132 B   |\n+\t\t| Repository structure      | Value      |\n+\t\t| ------------------------- | ---------- |\n+\t\t| * References              |            |\n+\t\t|   * Count                 |      4     |\n+\t\t|     * Branches            |      1     |\n+\t\t|     * Tags                |      1     |\n+\t\t|     * Remotes             |      1     |\n+\t\t|     * Others              |      1     |\n+\t\t|                           |            |\n+\t\t| * Reachable objects       |            |\n+\t\t|   * Count                 |   3.02 k   |\n+\t\t|     * Commits             |   1.01 k   |\n+\t\t|     * Trees               |   1.01 k   |\n+\t\t|     * Blobs               |   1.01 k   |\n+\t\t|     * Tags                |      1     |\n+\t\t|   * Inflated size         |  16.03 MiB |\n+\t\t|     * Commits             | 217.92 KiB |\n+\t\t|     * Trees               |  15.81 MiB |\n+\t\t|     * Blobs               |  11.68 KiB |\n+\t\t|     * Tags                |    132 B   |\n+\t\t|   * Disk size             | $(object_type_disk_usage all true) |\n+\t\t|     * Commits             | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees               | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs               |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags                |    $(object_type_disk_usage tag) B   |\n+\t\t|                           |            |\n+\t\t| * Largest objects         |            |\n+\t\t|   * Commits               |            |\n+\t\t|     * Maximum size    [1] |    223 B   |\n+\t\t|     * Maximum parents [2] |      1     |\n+\t\t|   * Trees                 |            |\n+\t\t|     * Maximum size    [3] |  32.29 KiB |\n+\t\t|   * Blobs                 |            |\n+\t\t|     * Maximum size    [4] |     13 B   |\n+\t\t|   * Tags                  |            |\n+\t\t|     * Maximum size    [5] |    132 B   |\n \n \t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n-\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n-\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n-\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n+\t\t[2] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n+\t\t[3] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[4] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[5] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -171,6 +174,8 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.blobs.max_size_oid=eaeeedced46482bd4281fda5a5f05ce24854151f\n \t\tobjects.tags.max_size=132\n \t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n+\t\tobjects.commits.max_parents=1\n+\t\tobjects.commits.max_parents_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"535104","messageId":"20260203221758.1164434-6-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260203221758.1164434-1-jltobler@gmail.com","subject":"[PATCH 5/5] builtin/repo: find tree with most entries","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-03T22:17:58Z","receivedAt":"2026-02-03T22:18:38Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The size of a tree object usually corresponds with the number of entries\nit has. While iterating through objects in the repository for\ngit-repo-structure, identify the tree with the most entries and display\nit in the output.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 27 +++++++++++++++++++++++++++\n t/t1901-repo-structure.sh | 13 +++++++++----\n 2 files changed, 36 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex dc1ac7ad3b..0f77d8f68f 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -16,6 +16,8 @@\n #include \"strbuf.h\"\n #include \"string-list.h\"\n #include \"shallow.h\"\n+#include \"tree.h\"\n+#include \"tree-walk.h\"\n #include \"utf8.h\"\n \n static const char *const repo_usage[] = {\n@@ -211,6 +213,7 @@ struct largest_objects {\n \tstruct object_data blob_size;\n \n \tstruct object_data parent_count;\n+\tstruct object_data tree_entries;\n };\n \n struct ref_stats {\n@@ -458,6 +461,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t\t     &objects->largest.tree_size.oid,\n \t\t\t\t     objects->largest.tree_size.value,\n \t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_count_addf(table,\n+\t\t\t\t      &objects->largest.tree_entries.oid,\n+\t\t\t\t      objects->largest.tree_entries.value,\n+\t\t\t\t      \"    * %s\", _(\"Maximum entries\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n \tstats_table_object_size_addf(table,\n \t\t\t\t     &objects->largest.blob_size.oid,\n@@ -619,6 +626,10 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \t       (uintmax_t)stats->objects.largest.parent_count.value, value_delim);\n \tprintf(\"objects.commits.max_parents_oid%c%s%c\", key_delim,\n \t       oid_to_hex(&stats->objects.largest.parent_count.oid), value_delim);\n+\tprintf(\"objects.trees.max_entries%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.tree_entries.value, value_delim);\n+\tprintf(\"objects.trees.max_entries_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.tree_entries.oid), value_delim);\n \n \tfflush(stdout);\n }\n@@ -697,6 +708,20 @@ static void check_largest(struct object_data *data, struct object_id *oid,\n \t}\n }\n \n+static size_t count_tree_entries(struct object *obj)\n+{\n+\tstruct tree *t = object_as_type(obj, OBJ_TREE, 0);\n+\tstruct name_entry entry;\n+\tstruct tree_desc desc;\n+\tsize_t count = 0;\n+\n+\tinit_tree_desc(&desc, &t->object.oid, t->buffer, t->size);\n+\twhile (tree_entry(&desc, &entry))\n+\t\tcount++;\n+\n+\treturn count;\n+}\n+\n static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t enum object_type type, void *cb_data)\n {\n@@ -749,6 +774,8 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tstats->disk_sizes.trees += disk;\n \t\t\tcheck_largest(&stats->largest.tree_size, &oids->oid[i],\n \t\t\t\t      inflated);\n+\t\t\tcheck_largest(&stats->largest.tree_entries, &oids->oid[i],\n+\t\t\t\t      count_tree_entries(obj));\n \t\t\tbreak;\n \t\tcase OBJ_BLOB:\n \t\t\tstats->type_counts.blobs++;\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex d003d64a8e..12ed67e846 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -59,6 +59,7 @@ test_expect_success 'empty repository' '\n \t\t|     * Maximum parents     |    0   |\n \t\t|   * Trees                 |        |\n \t\t|     * Maximum size        |    0 B |\n+\t\t|     * Maximum entries     |    0   |\n \t\t|   * Blobs                 |        |\n \t\t|     * Maximum size        |    0 B |\n \t\t|   * Tags                  |        |\n@@ -122,16 +123,18 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Maximum parents [2] |      1     |\n \t\t|   * Trees                 |            |\n \t\t|     * Maximum size    [3] |  32.29 KiB |\n+\t\t|     * Maximum entries [4] |   1.01 k   |\n \t\t|   * Blobs                 |            |\n-\t\t|     * Maximum size    [4] |     13 B   |\n+\t\t|     * Maximum size    [5] |     13 B   |\n \t\t|   * Tags                  |            |\n-\t\t|     * Maximum size    [5] |    132 B   |\n+\t\t|     * Maximum size    [6] |    132 B   |\n \n \t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n \t\t[2] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n \t\t[3] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n-\t\t[4] 97d808e45116bf02103490294d3d46dad7a2ac62\n-\t\t[5] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n+\t\t[4] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[5] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[6] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -176,6 +179,8 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n \t\tobjects.commits.max_parents=1\n \t\tobjects.commits.max_parents_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n+\t\tobjects.trees.max_entries=42\n+\t\tobjects.trees.max_entries_oid=09931deea9d81ec21300d3e13c74412f32eacec5\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"535105","messageId":"xmqqzf5pqwtm.fsf@gitster.g","threadId":"64911","inReplyTo":"20260203221758.1164434-2-jltobler@gmail.com","subject":"Re: [PATCH 1/5] builtin/repo: update stats for each object","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-03T22:36:05Z","receivedAt":"2026-02-03T22:36:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> +\t\tswitch (type) {\n> +\t\tcase OBJ_TAG:\n> +\t\t\tstats->type_counts.tags++;\n> +\t\t\tstats->inflated_sizes.tags += inflated;\n> +\t\t\tstats->disk_sizes.tags += disk;\n> +\t\t\tbreak;\n> +\t\tcase OBJ_COMMIT:\n> +\t\t\tstats->type_counts.commits++;\n> +\t\t\tstats->inflated_sizes.commits += inflated;\n> +\t\t\tstats->disk_sizes.commits += disk;\n> +\t\t\tbreak;\n> +\t\tcase OBJ_TREE:\n> +\t\t\tstats->type_counts.trees++;\n> +\t\t\tstats->inflated_sizes.trees += inflated;\n> +\t\t\tstats->disk_sizes.trees += disk;\n> +\t\t\tbreak;\n> +\t\tcase OBJ_BLOB:\n> +\t\t\tstats->type_counts.blobs++;\n> +\t\t\tstats->inflated_sizes.blobs += inflated;\n> +\t\t\tstats->disk_sizes.blobs += disk;\n> +\t\t\tbreak;\n> +\t\tdefault:\n> +\t\t\tBUG(\"invalid object type\");\n> +\t\t}\n>  \t}\n\nThe repetition above makes me wonder if it might be a better\norganization to have\n\n    struct object_stat {       \n        struct type_stat {\n            size_t count;\n            size_t inflated_size;\n            size_t disk_size;\n\t} tag, commit, tree, blob;\n\t... possibly other members ...\n    } *stats;\n\nor even\n\n    struct object_stat {       \n        struct type_stat {\n            size_t count;\n            size_t inflated_size;\n            size_t disk_size;\n\t} t[4];\n\t... possibly other members ...\n    };\n\nand have this part of the code be\n\n\tstruct type_stat *t;\n\n\tif (OBJ_COMMIT <= type && type <= OBJ_TAG)\n\t\tt = stats->t[type - 1];\n\telse\n\t\tBUG(\"invalid object type\");\n\n\tt->count++;\n\tt->inflated_size += inflated;\n\tt->disk_size += disk;\n\nbut that is probably only because I am looking at this part of the\ncode.  Other parts of the code may have good reasons to have the\nstructure nested the other way around like you have.\n\n"},{"id":"535106","messageId":"xmqqv7gdqwei.fsf@gitster.g","threadId":"64911","inReplyTo":"20260203221758.1164434-3-jltobler@gmail.com","subject":"Re: [PATCH 2/5] builtin/repo: collect largest inflated objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-03T22:45:09Z","receivedAt":"2026-02-03T22:45:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> The \"structure\" output for git-repo(1) shows the total inflated and disk\n> sizes of reachable objects in the repository, but doesn't show the size\n> of the largest individual objects. Since an individual object may be a\n> large contributor to the overall repository size, it is useful for users\n> to know the maximum size of individual objects.\n\nHmph.  It is true that a byte is worth the same amount of money no\nmatter what object it is used to represent, but comparing the size\nof a commit object and the size of a blob object feels inherently\nmeaningless to me.\n\nIt all depends on what you are trying to learn out of the stats, but\nhaving many small blob objects that add up to 1GB and having medium\nnumber of medium sized tree objects that adds up to the same 1GB\nwould give the same number in object_stats.inflated_sizes for both\ntypes, indicating that they are costing you about the same.  But the\nmembers in largest_objects for these types would be different,\nhinting (incorrectly) that one type may be costing more than the\nother.  Would that really tell us something useful, I have to\nwonder?\n\nOne thing that is related to \"largest\" that might be useful is how\nspiky size distribution is.  Among many medium sized blobs, if there\nis only a handful of super huge blobs, that is quite a notable thing\nto know (as opposed to the case where these super huge blobs are\nnot so unusual).\n\n"},{"id":"535107","messageId":"xmqqpl6lqw86.fsf@gitster.g","threadId":"64911","inReplyTo":"20260203221758.1164434-5-jltobler@gmail.com","subject":"Re: [PATCH 4/5] builtin/repo: find commit with most parents","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-03T22:48:57Z","receivedAt":"2026-02-03T22:48:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> Complex merge events may produce an octopus merge where the resulting\n> merge commit has more than two parents. While iterating through objects\n> in the repository for git-repo-structure, identify the commit with the\n> most parents and display it in the output.\n\nDoes the size of octopus have anything more than a curiosity value?\n\nThe opposite, the commit with most direct children, might be even\nmore interesting, but that may be just me.\n"},{"id":"535108","messageId":"xmqqldh9qw5d.fsf@gitster.g","threadId":"64911","inReplyTo":"20260203221758.1164434-6-jltobler@gmail.com","subject":"Re: [PATCH 5/5] builtin/repo: find tree with most entries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-03T22:50:38Z","receivedAt":"2026-02-03T22:50:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> The size of a tree object usually corresponds with the number of entries\n> it has. While iterating through objects in the repository for\n> git-repo-structure, identify the tree with the most entries and display\n> it in the output.\n\nAll of these \"largest\" and \"most\", it would be a lot more\ninteresting if we can give not just these extreme values but\ndistrubution, possibly in a graphical way for bonus points.\n\n;-)\n"},{"id":"535111","messageId":"e48578d5-ec48-4369-901a-597de3be9455@app.fastmail.com","threadId":"64911","inReplyTo":"xmqqpl6lqw86.fsf@gitster.g","subject":"Re: [PATCH 4/5] builtin/repo: find commit with most parents","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2026-02-03T23:14:07Z","receivedAt":"2026-02-03T23:15:10Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"On Tue, Feb 3, 2026, at 23:48, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n>\n>> Complex merge events may produce an octopus merge where the resulting\n>> merge commit has more than two parents. While iterating through objects\n>> in the repository for git-repo-structure, identify the commit with the\n>> most parents and display it in the output.\n>\n> Does the size of octopus have anything more than a curiosity value?\n\nI’m guessing this stat is inspired by git-sizer.[1][2] This is all that\nthe project says about “octopus”:\n\n    * Are there other bizarre and questionable things in your repository?\n\n        * Annotated tags pointing at one another in long chains?\n        * Octopus merges with dozens of parents?\n        * Commits with gigantic log messages?\n\nIt marks the max of 10 in this repo as a “one star” (*) concern\n(lowest). The 66 parent commit in the Linux Kernel gets six stars.\n\nBy the way: why did this project stop doing 3+ parent merges?\n\n🔗 1: https://lore.kernel.org/git/20251021182601.2687284-5-jltobler@gmail.com/\n🔗 2: https://github.com/github/git-sizer\n\n>\n> The opposite, the commit with most direct children, might be even\n> more interesting, but that may be just me.\n"},{"id":"535112","messageId":"xmqqy0l9pflc.fsf@gitster.g","threadId":"64911","inReplyTo":"e48578d5-ec48-4369-901a-597de3be9455@app.fastmail.com","subject":"Re: [PATCH 4/5] builtin/repo: find commit with most parents","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-03T23:33:35Z","receivedAt":"2026-02-03T23:33:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Kristoffer Haugsbakk\" <kristofferhaugsbakk@fastmail.com> writes:\n\n> By the way: why did this project stop doing 3+ parent merges?\n\nIf you mean Git, the primary reason is because I do not see much\nvalue in Octopus merges, which is very hostile to bisection.  It was\n\"interesting\" to view them in gitk while the tool was young and\nnobody has seen such a structure, but curiosity rapidly wanes ;-).\n"},{"id":"535119","messageId":"aYMDL4m7Ceifl1Ja@pks.im","threadId":"64911","inReplyTo":"xmqqldh9qw5d.fsf@gitster.g","subject":"Re: [PATCH 5/5] builtin/repo: find tree with most entries","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-02-04T08:28:31Z","receivedAt":"2026-02-04T08:28:43Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Feb 03, 2026 at 02:50:38PM -0800, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > The size of a tree object usually corresponds with the number of entries\n> > it has. While iterating through objects in the repository for\n> > git-repo-structure, identify the tree with the most entries and display\n> > it in the output.\n> \n> All of these \"largest\" and \"most\", it would be a lot more\n> interesting if we can give not just these extreme values but\n> distrubution, possibly in a graphical way for bonus points.\n\nThat would be amazing indeed! I think having the largest values is still\nvaluable as it allows you to detect weird outliers quite easily. But\nhaving a histogram would of course give the bigger picture.\n\nI guess the challenging part would be to compute the buckets of that\nhistogram in a streaming fashion. But I guess we could:\n\n  1. Pick a target number of buckets.\n\n  2. Track the maximum respective values as we stream.\n\n  3. Merge existing buckets and create new ones in case the maximum\n     value changes.\n\nThe target number of buckets may not necessarily be the same number as\nthe number of buckets that we will eventually print for increased\nresolution.\n\nThe distributions could then be printed as an ASCII bar chart, for\nexample something like:\n\n    0-50   │████████████████████████████████████████ 1,247\n   50-100  │█████████████████████████████████ 812\n  100-150  │█████████████ 401\n  150-200  │████████ 253\n  200-250  │████ 128\n  250-300  │██ 67\n  300+     │▏ 12\n           0        250       500       750      1k     1.2k count\n    bytes / count\n\nFrom my point of view that would be the cherry on top of the new tool :)\nI'd personally still like to learn about maximum values in the table, as\nI've found that info to be useful with some customer incidents in the\npast. It's not giving you a trend, but it immediately gives you some\ngood signal that the repo shape might be weird if you have commits with\nhundreds of parents.\n\nSo maybe this is another step we can do in a subsequent patch series?\n\nThanks!\n\nPatrick\n"},{"id":"535163","messageId":"xmqqtsvwply1.fsf@gitster.g","threadId":"64911","inReplyTo":"aYMDL4m7Ceifl1Ja@pks.im","subject":"Re: [PATCH 5/5] builtin/repo: find tree with most entries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-04T15:28:38Z","receivedAt":"2026-02-04T15:28:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> From my point of view that would be the cherry on top of the new tool :)\n> I'd personally still like to learn about maximum values in the table, as\n> I've found that info to be useful with some customer incidents in the\n> past. It's not giving you a trend, but it immediately gives you some\n> good signal that the repo shape might be weird if you have commits with\n> hundreds of parents.\n>\n> So maybe this is another step we can do in a subsequent patch series?\n\nOh, I didn't mean to say \"the maximum alone is not interesting\nenough for me to bother, come back with histograms.\"  If you already\nhave a good feel for normal range/distribution already, then one\ndata point at the extreme is a sign enough for you to notice when\nthere is something fishy going on.\n\nThanks.\n"},{"id":"535926","messageId":"aY8jt_8LoN1LdboG@pks.im","threadId":"64911","inReplyTo":"20260203221758.1164434-4-jltobler@gmail.com","subject":"Re: [PATCH 3/5] builtin/repo: add OID annotations to table output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-02-13T13:14:31Z","receivedAt":"2026-02-13T13:14:37Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Feb 03, 2026 at 04:17:56PM -0600, Justin Tobler wrote:\n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index 51a4359685..6fc2d9db12 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -272,6 +275,12 @@ static void stats_table_vaddf(struct stats_table *table,\n>  \t\ttable->name_col_width = name_width;\n>  \tif (!entry)\n>  \t\treturn;\n> +\tif (entry->oid) {\n> +\t\tentry->index = table->annotations.nr + 1;\n> +\t\tstrbuf_addf(&buf, \"[%\" PRIuMAX \"] %s\", (uintmax_t)entry->index,\n> +\t\t\t    oid_to_hex(entry->oid));\n> +\t\tstring_list_append(&table->annotations, buf.buf);\n\nCan't we hand over ownership here to avoid the extra string copy?\n\n    string_list_append_nodup(&table->rows, strbuf_detach(&buf, NULL));\n\n> @@ -282,6 +291,8 @@ static void stats_table_vaddf(struct stats_table *table,\n>  \t\tif (unit_width > table->unit_col_width)\n>  \t\t\ttable->unit_col_width = unit_width;\n>  \t}\n> +\n> +\tstrbuf_release(&buf);\n>  }\n>  \n>  static void stats_table_addf(struct stats_table *table, const char *format, ...)\n\nI was wondering why we only start releasing the buffer now. But before\nthese changes we used `strbuf_detach()` on it, so there's been no memory\nleak here.\n\n> @@ -321,6 +332,27 @@ static void stats_table_size_addf(struct stats_table *table, size_t value,\n>  \tva_end(ap);\n>  }\n>  \n> +static void stats_table_object_size_addf(struct stats_table *table,\n> +\t\t\t\t\t struct object_id *oid, size_t value,\n> +\t\t\t\t\t const char *format, ...)\n> +{\n> +\tstruct stats_table_entry *entry;\n> +\tva_list ap;\n> +\n> +\tCALLOC_ARRAY(entry, 1);\n> +\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n> +\n> +\t/*\n> +\t * A NULL OID should not have a table annotation.\n> +\t */\n> +\tif (!is_null_oid(oid))\n> +\t\tentry->oid = oid;\n\nI guess this case could be hit if a certain object type didn't have any\nobjects at all?\n\n> diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> index 1999f325d0..918af7269f 100755\n> --- a/t/t1901-repo-structure.sh\n> +++ b/t/t1901-repo-structure.sh\n> @@ -89,41 +89,46 @@ test_expect_success SHA1 'repository with references and objects' '\n>  \t\t# git-rev-list(1) --disk-usage=human option printing the full\n>  \t\t# \"byte/bytes\" unit string instead of just \"B\".\n>  \t\tcat >expect <<-EOF &&\n> -\t\t| Repository structure | Value      |\n> -\t\t| -------------------- | ---------- |\n> -\t\t| * References         |            |\n> -\t\t|   * Count            |      4     |\n> -\t\t|     * Branches       |      1     |\n> -\t\t|     * Tags           |      1     |\n> -\t\t|     * Remotes        |      1     |\n> -\t\t|     * Others         |      1     |\n> -\t\t|                      |            |\n> -\t\t| * Reachable objects  |            |\n> -\t\t|   * Count            |   3.02 k   |\n> -\t\t|     * Commits        |   1.01 k   |\n> -\t\t|     * Trees          |   1.01 k   |\n> -\t\t|     * Blobs          |   1.01 k   |\n> -\t\t|     * Tags           |      1     |\n> -\t\t|   * Inflated size    |  16.03 MiB |\n> -\t\t|     * Commits        | 217.92 KiB |\n> -\t\t|     * Trees          |  15.81 MiB |\n> -\t\t|     * Blobs          |  11.68 KiB |\n> -\t\t|     * Tags           |    132 B   |\n> -\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n> -\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n> -\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n> -\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n> -\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n> -\t\t|                      |            |\n> -\t\t| * Largest objects    |            |\n> -\t\t|   * Commits          |            |\n> -\t\t|     * Maximum size   |    223 B   |\n> -\t\t|   * Trees            |            |\n> -\t\t|     * Maximum size   |  32.29 KiB |\n> -\t\t|   * Blobs            |            |\n> -\t\t|     * Maximum size   |     13 B   |\n> -\t\t|   * Tags             |            |\n> -\t\t|     * Maximum size   |    132 B   |\n> +\t\t| Repository structure     | Value      |\n> +\t\t| ------------------------ | ---------- |\n> +\t\t| * References             |            |\n> +\t\t|   * Count                |      4     |\n> +\t\t|     * Branches           |      1     |\n> +\t\t|     * Tags               |      1     |\n> +\t\t|     * Remotes            |      1     |\n> +\t\t|     * Others             |      1     |\n> +\t\t|                          |            |\n> +\t\t| * Reachable objects      |            |\n> +\t\t|   * Count                |   3.02 k   |\n> +\t\t|     * Commits            |   1.01 k   |\n> +\t\t|     * Trees              |   1.01 k   |\n> +\t\t|     * Blobs              |   1.01 k   |\n> +\t\t|     * Tags               |      1     |\n> +\t\t|   * Inflated size        |  16.03 MiB |\n> +\t\t|     * Commits            | 217.92 KiB |\n> +\t\t|     * Trees              |  15.81 MiB |\n> +\t\t|     * Blobs              |  11.68 KiB |\n> +\t\t|     * Tags               |    132 B   |\n> +\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n> +\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n> +\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n> +\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n> +\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n> +\t\t|                          |            |\n> +\t\t| * Largest objects        |            |\n> +\t\t|   * Commits              |            |\n> +\t\t|     * Maximum size   [1] |    223 B   |\n> +\t\t|   * Trees                |            |\n> +\t\t|     * Maximum size   [2] |  32.29 KiB |\n> +\t\t|   * Blobs                |            |\n> +\t\t|     * Maximum size   [3] |     13 B   |\n> +\t\t|   * Tags                 |            |\n> +\t\t|     * Maximum size   [4] |    132 B   |\n> +\n> +\t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n> +\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n> +\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n> +\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n>  \t\tEOF\n\nI was briefly wondering whether we can do better here and for example\noutput something like this:\n\n\t| * Largest objects              |            |\n\t|   * Commits                    |            |\n\t|     * Maximum size   [commits] |    223 B   |\n\n\t[commits] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n\nBut I think that becomes quite unwieldy as the column's size is extended\nquite a bit, so numbers it probably the better interface.\n\nPatrick\n"},{"id":"536326","messageId":"aZYUwjSEAcGRSXNa@denethor","threadId":"64911","inReplyTo":"xmqqzf5pqwtm.fsf@gitster.g","subject":"Re: [PATCH 1/5] builtin/repo: update stats for each object","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-18T19:40:21Z","receivedAt":"2026-02-18T19:40:23Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/03 02:36PM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > +\t\tswitch (type) {\n> > +\t\tcase OBJ_TAG:\n> > +\t\t\tstats->type_counts.tags++;\n> > +\t\t\tstats->inflated_sizes.tags += inflated;\n> > +\t\t\tstats->disk_sizes.tags += disk;\n> > +\t\t\tbreak;\n> > +\t\tcase OBJ_COMMIT:\n> > +\t\t\tstats->type_counts.commits++;\n> > +\t\t\tstats->inflated_sizes.commits += inflated;\n> > +\t\t\tstats->disk_sizes.commits += disk;\n> > +\t\t\tbreak;\n> > +\t\tcase OBJ_TREE:\n> > +\t\t\tstats->type_counts.trees++;\n> > +\t\t\tstats->inflated_sizes.trees += inflated;\n> > +\t\t\tstats->disk_sizes.trees += disk;\n> > +\t\t\tbreak;\n> > +\t\tcase OBJ_BLOB:\n> > +\t\t\tstats->type_counts.blobs++;\n> > +\t\t\tstats->inflated_sizes.blobs += inflated;\n> > +\t\t\tstats->disk_sizes.blobs += disk;\n> > +\t\t\tbreak;\n> > +\t\tdefault:\n> > +\t\t\tBUG(\"invalid object type\");\n> > +\t\t}\n> >  \t}\n> \n> The repetition above makes me wonder if it might be a better\n> organization to have\n> \n>     struct object_stat {       \n>         struct type_stat {\n>             size_t count;\n>             size_t inflated_size;\n>             size_t disk_size;\n> \t} tag, commit, tree, blob;\n> \t... possibly other members ...\n>     } *stats;\n> \n> or even\n> \n>     struct object_stat {       \n>         struct type_stat {\n>             size_t count;\n>             size_t inflated_size;\n>             size_t disk_size;\n> \t} t[4];\n> \t... possibly other members ...\n>     };\n> \n> and have this part of the code be\n> \n> \tstruct type_stat *t;\n> \n> \tif (OBJ_COMMIT <= type && type <= OBJ_TAG)\n> \t\tt = stats->t[type - 1];\n> \telse\n> \t\tBUG(\"invalid object type\");\n> \n> \tt->count++;\n> \tt->inflated_size += inflated;\n> \tt->disk_size += disk;\n> \n> but that is probably only because I am looking at this part of the\n> code.  Other parts of the code may have good reasons to have the\n> structure nested the other way around like you have.\n\nGood suggestion. Some of the info added in the following commits is\nobject specific and will need to be handled accordinly, but we could\nprobably still benefit by structuring the data a bit better. Will\nexplore in the next version.\n\n-Justin\n"},{"id":"536327","messageId":"aZYV0o9Xp-v5IPL1@denethor","threadId":"64911","inReplyTo":"xmqqv7gdqwei.fsf@gitster.g","subject":"Re: [PATCH 2/5] builtin/repo: collect largest inflated objects","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-18T20:01:04Z","receivedAt":"2026-02-18T20:01:08Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/03 02:45PM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > The \"structure\" output for git-repo(1) shows the total inflated and disk\n> > sizes of reachable objects in the repository, but doesn't show the size\n> > of the largest individual objects. Since an individual object may be a\n> > large contributor to the overall repository size, it is useful for users\n> > to know the maximum size of individual objects.\n> \n> Hmph.  It is true that a byte is worth the same amount of money no\n> matter what object it is used to represent, but comparing the size\n> of a commit object and the size of a blob object feels inherently\n> meaningless to me.\n\nI certainly agree that comparing max size values between the types\nthemselves is not particularly meaningfull. I do think though the max\nsize values by themselves provide insight into the extremes of the\nrepository.\n\n> It all depends on what you are trying to learn out of the stats, but\n> having many small blob objects that add up to 1GB and having medium\n> number of medium sized tree objects that adds up to the same 1GB\n> would give the same number in object_stats.inflated_sizes for both\n> types, indicating that they are costing you about the same.  But the\n> members in largest_objects for these types would be different,\n> hinting (incorrectly) that one type may be costing more than the\n> other.  Would that really tell us something useful, I have to\n> wonder?\n\nYa the largest objects and inflated sizes you can not really gain any\ninsight regarding the distribution, but I think it still a good idea to\nshowcase the extremes. If I see the max size values are \"normal\", that\nat least gives me some insight into the repository usage patterns.\n\n> One thing that is related to \"largest\" that might be useful is how\n> spiky size distribution is.  Among many medium sized blobs, if there\n> is only a handful of super huge blobs, that is quite a notable thing\n> to know (as opposed to the case where these super huge blobs are\n> not so unusual).\n\nI agree that showing a distribution here would be quite useful. This is\nsomething I plan to explore in a followup series. :)\n\n-Justin\n"},{"id":"536328","messageId":"aZYa8U1hQ1oaeCKn@denethor","threadId":"64911","inReplyTo":"e48578d5-ec48-4369-901a-597de3be9455@app.fastmail.com","subject":"Re: [PATCH 4/5] builtin/repo: find commit with most parents","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-18T20:06:45Z","receivedAt":"2026-02-18T20:06:47Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/04 12:14AM, Kristoffer Haugsbakk wrote:\n> On Tue, Feb 3, 2026, at 23:48, Junio C Hamano wrote:\n> > Justin Tobler <jltobler@gmail.com> writes:\n> >\n> >> Complex merge events may produce an octopus merge where the resulting\n> >> merge commit has more than two parents. While iterating through objects\n> >> in the repository for git-repo-structure, identify the commit with the\n> >> most parents and display it in the output.\n> >\n> > Does the size of octopus have anything more than a curiosity value?\n> \n> I’m guessing this stat is inspired by git-sizer.[1][2] This is all that\n> the project says about “octopus”:\n> \n>     * Are there other bizarre and questionable things in your repository?\n> \n>         * Annotated tags pointing at one another in long chains?\n>         * Octopus merges with dozens of parents?\n>         * Commits with gigantic log messages?\n> \n> It marks the max of 10 in this repo as a “one star” (*) concern\n> (lowest). The 66 parent commit in the Linux Kernel gets six stars.\n\nYup, this is taken from git-sizer. From my perspective the max parents\nvalue largely just provides additional insight into how the repository\nmay have been used/structured.\n\n-Justin\n"},{"id":"536329","messageId":"aZYb9BBBjlviZsnR@denethor","threadId":"64911","inReplyTo":"aY8jt_8LoN1LdboG@pks.im","subject":"Re: [PATCH 3/5] builtin/repo: add OID annotations to table output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-18T20:13:26Z","receivedAt":"2026-02-18T20:13:28Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/13 02:14PM, Patrick Steinhardt wrote:\n> On Tue, Feb 03, 2026 at 04:17:56PM -0600, Justin Tobler wrote:\n> > diff --git a/builtin/repo.c b/builtin/repo.c\n> > index 51a4359685..6fc2d9db12 100644\n> > --- a/builtin/repo.c\n> > +++ b/builtin/repo.c\n> > @@ -272,6 +275,12 @@ static void stats_table_vaddf(struct stats_table *table,\n> >  \t\ttable->name_col_width = name_width;\n> >  \tif (!entry)\n> >  \t\treturn;\n> > +\tif (entry->oid) {\n> > +\t\tentry->index = table->annotations.nr + 1;\n> > +\t\tstrbuf_addf(&buf, \"[%\" PRIuMAX \"] %s\", (uintmax_t)entry->index,\n> > +\t\t\t    oid_to_hex(entry->oid));\n> > +\t\tstring_list_append(&table->annotations, buf.buf);\n> \n> Can't we hand over ownership here to avoid the extra string copy?\n> \n>     string_list_append_nodup(&table->rows, strbuf_detach(&buf, NULL));\n\nGood suggestion. Will adapt in the next version.\n\n> > @@ -321,6 +332,27 @@ static void stats_table_size_addf(struct stats_table *table, size_t value,\n> >  \tva_end(ap);\n> >  }\n> >  \n> > +static void stats_table_object_size_addf(struct stats_table *table,\n> > +\t\t\t\t\t struct object_id *oid, size_t value,\n> > +\t\t\t\t\t const char *format, ...)\n> > +{\n> > +\tstruct stats_table_entry *entry;\n> > +\tva_list ap;\n> > +\n> > +\tCALLOC_ARRAY(entry, 1);\n> > +\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n> > +\n> > +\t/*\n> > +\t * A NULL OID should not have a table annotation.\n> > +\t */\n> > +\tif (!is_null_oid(oid))\n> > +\t\tentry->oid = oid;\n> \n> I guess this case could be hit if a certain object type didn't have any\n> objects at all?\n\nYup. An example here would be running this command on an empty\nrepository.\n\n> > diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> > index 1999f325d0..918af7269f 100755\n> > --- a/t/t1901-repo-structure.sh\n> > +++ b/t/t1901-repo-structure.sh\n> > @@ -89,41 +89,46 @@ test_expect_success SHA1 'repository with references and objects' '\n> >  \t\t# git-rev-list(1) --disk-usage=human option printing the full\n> >  \t\t# \"byte/bytes\" unit string instead of just \"B\".\n> >  \t\tcat >expect <<-EOF &&\n> > -\t\t| Repository structure | Value      |\n> > -\t\t| -------------------- | ---------- |\n> > -\t\t| * References         |            |\n> > -\t\t|   * Count            |      4     |\n> > -\t\t|     * Branches       |      1     |\n> > -\t\t|     * Tags           |      1     |\n> > -\t\t|     * Remotes        |      1     |\n> > -\t\t|     * Others         |      1     |\n> > -\t\t|                      |            |\n> > -\t\t| * Reachable objects  |            |\n> > -\t\t|   * Count            |   3.02 k   |\n> > -\t\t|     * Commits        |   1.01 k   |\n> > -\t\t|     * Trees          |   1.01 k   |\n> > -\t\t|     * Blobs          |   1.01 k   |\n> > -\t\t|     * Tags           |      1     |\n> > -\t\t|   * Inflated size    |  16.03 MiB |\n> > -\t\t|     * Commits        | 217.92 KiB |\n> > -\t\t|     * Trees          |  15.81 MiB |\n> > -\t\t|     * Blobs          |  11.68 KiB |\n> > -\t\t|     * Tags           |    132 B   |\n> > -\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n> > -\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n> > -\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n> > -\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n> > -\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n> > -\t\t|                      |            |\n> > -\t\t| * Largest objects    |            |\n> > -\t\t|   * Commits          |            |\n> > -\t\t|     * Maximum size   |    223 B   |\n> > -\t\t|   * Trees            |            |\n> > -\t\t|     * Maximum size   |  32.29 KiB |\n> > -\t\t|   * Blobs            |            |\n> > -\t\t|     * Maximum size   |     13 B   |\n> > -\t\t|   * Tags             |            |\n> > -\t\t|     * Maximum size   |    132 B   |\n> > +\t\t| Repository structure     | Value      |\n> > +\t\t| ------------------------ | ---------- |\n> > +\t\t| * References             |            |\n> > +\t\t|   * Count                |      4     |\n> > +\t\t|     * Branches           |      1     |\n> > +\t\t|     * Tags               |      1     |\n> > +\t\t|     * Remotes            |      1     |\n> > +\t\t|     * Others             |      1     |\n> > +\t\t|                          |            |\n> > +\t\t| * Reachable objects      |            |\n> > +\t\t|   * Count                |   3.02 k   |\n> > +\t\t|     * Commits            |   1.01 k   |\n> > +\t\t|     * Trees              |   1.01 k   |\n> > +\t\t|     * Blobs              |   1.01 k   |\n> > +\t\t|     * Tags               |      1     |\n> > +\t\t|   * Inflated size        |  16.03 MiB |\n> > +\t\t|     * Commits            | 217.92 KiB |\n> > +\t\t|     * Trees              |  15.81 MiB |\n> > +\t\t|     * Blobs              |  11.68 KiB |\n> > +\t\t|     * Tags               |    132 B   |\n> > +\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n> > +\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n> > +\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n> > +\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n> > +\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n> > +\t\t|                          |            |\n> > +\t\t| * Largest objects        |            |\n> > +\t\t|   * Commits              |            |\n> > +\t\t|     * Maximum size   [1] |    223 B   |\n> > +\t\t|   * Trees                |            |\n> > +\t\t|     * Maximum size   [2] |  32.29 KiB |\n> > +\t\t|   * Blobs                |            |\n> > +\t\t|     * Maximum size   [3] |     13 B   |\n> > +\t\t|   * Tags                 |            |\n> > +\t\t|     * Maximum size   [4] |    132 B   |\n> > +\n> > +\t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n> > +\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n> > +\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n> > +\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n> >  \t\tEOF\n> \n> I was briefly wondering whether we can do better here and for example\n> output something like this:\n> \n> \t| * Largest objects              |            |\n> \t|   * Commits                    |            |\n> \t|     * Maximum size   [commits] |    223 B   |\n> \n> \t[commits] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n> \n> But I think that becomes quite unwieldy as the column's size is extended\n> quite a bit, so numbers it probably the better interface.\n\nIn later commits we also add more annotations. One example is max commit\nsize and max commit parents may be diffent commit OIDs. I think numbers\nmay be the simplest for now.\n\nThanks,\n-Justin\n"},{"id":"536857","messageId":"20260223174120.2356504-1-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260203221758.1164434-1-jltobler@gmail.com","subject":"[PATCH v2 0/5] builtin/repo: include largest object information","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-23T17:41:15Z","receivedAt":"2026-02-23T17:41:28Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThe \"structure\" output for git-repo(1) currently provides count\ninformation for references/objects as well as total inflated/disk sizes\nof objects by type. Info regarding the largest individual objects in the\nrepository is not yet collected, but would be useful to users wishing to\nidentify such large objects.\n\nThis patch series adds the following data points:\n- The OID and size of the largest objects by object type\n- The OID and parent count of the commit with the most parents\n- The OID and entries count of the tree with the most entries\n\nChanges from V1:\n\n- Avoided duplicating the annotation string by handing over ownership.\n- I decided to leave the `struct object_stats` structure alone for now\n  as storing the various object values per-type does make it convenient\n  to calulate the various totals. I may revisit this in a future series\n  though.\n\nThanks,\n-Justin\n\nJustin Tobler (5):\n  builtin/repo: update stats for each object\n  builtin/repo: collect largest inflated objects\n  builtin/repo: add OID annotations to table output\n  builtin/repo: find commit with most parents\n  builtin/repo: find tree with most entries\n\n Documentation/git-repo.adoc |   1 +\n builtin/repo.c              | 249 +++++++++++++++++++++++++++++++-----\n t/t1901-repo-structure.sh   | 143 +++++++++++++--------\n 3 files changed, 313 insertions(+), 80 deletions(-)\n\nRange-diff against v1:\n1:  94a44e0e0f = 1:  94a44e0e0f builtin/repo: update stats for each object\n2:  92dbf34f2c = 2:  92dbf34f2c builtin/repo: collect largest inflated objects\n3:  1811d03afe ! 3:  1457d5d59c builtin/repo: add OID annotations to table output\n    @@ builtin/repo.c: static void stats_table_vaddf(struct stats_table *table,\n     +\t\tentry->index = table->annotations.nr + 1;\n     +\t\tstrbuf_addf(&buf, \"[%\" PRIuMAX \"] %s\", (uintmax_t)entry->index,\n     +\t\t\t    oid_to_hex(entry->oid));\n    -+\t\tstring_list_append(&table->annotations, buf.buf);\n    ++\t\tstring_list_append_nodup(&table->annotations, strbuf_detach(&buf, NULL));\n     +\t}\n      \tif (entry->value) {\n      \t\tint value_width = utf8_strwidth(entry->value);\n4:  471d352cc1 = 4:  f4e92e3f09 builtin/repo: find commit with most parents\n5:  7f1b7f9657 = 5:  af404fcc6c builtin/repo: find tree with most entries\n\nbase-commit: 67ad42147a7acc2af6074753ebd03d904476118f\n-- \n2.53.0\n\n"},{"id":"536858","messageId":"20260223174120.2356504-2-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"[PATCH v2 1/5] builtin/repo: update stats for each object","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-23T17:41:16Z","receivedAt":"2026-02-23T17:41:29Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"When walking reachable objects in the repository, `count_objects()`\nprocesses a set of objects and updates the `struct object_stats`. In\npreparation for more granular statistics being collected, update the\n`struct object_stats` for each individual object instead.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 53 +++++++++++++++++++++++---------------------------\n 1 file changed, 24 insertions(+), 29 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 0ea045abc1..c7c9f0f497 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -558,8 +558,6 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n-\tsize_t inflated_total = 0;\n-\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n@@ -575,33 +573,30 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n \t\t\tcontinue;\n \n-\t\tinflated_total += inflated;\n-\t\tdisk_total += disk;\n-\t}\n-\n-\tswitch (type) {\n-\tcase OBJ_TAG:\n-\t\tstats->type_counts.tags += oids->nr;\n-\t\tstats->inflated_sizes.tags += inflated_total;\n-\t\tstats->disk_sizes.tags += disk_total;\n-\t\tbreak;\n-\tcase OBJ_COMMIT:\n-\t\tstats->type_counts.commits += oids->nr;\n-\t\tstats->inflated_sizes.commits += inflated_total;\n-\t\tstats->disk_sizes.commits += disk_total;\n-\t\tbreak;\n-\tcase OBJ_TREE:\n-\t\tstats->type_counts.trees += oids->nr;\n-\t\tstats->inflated_sizes.trees += inflated_total;\n-\t\tstats->disk_sizes.trees += disk_total;\n-\t\tbreak;\n-\tcase OBJ_BLOB:\n-\t\tstats->type_counts.blobs += oids->nr;\n-\t\tstats->inflated_sizes.blobs += inflated_total;\n-\t\tstats->disk_sizes.blobs += disk_total;\n-\t\tbreak;\n-\tdefault:\n-\t\tBUG(\"invalid object type\");\n+\t\tswitch (type) {\n+\t\tcase OBJ_TAG:\n+\t\t\tstats->type_counts.tags++;\n+\t\t\tstats->inflated_sizes.tags += inflated;\n+\t\t\tstats->disk_sizes.tags += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_COMMIT:\n+\t\t\tstats->type_counts.commits++;\n+\t\t\tstats->inflated_sizes.commits += inflated;\n+\t\t\tstats->disk_sizes.commits += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_TREE:\n+\t\t\tstats->type_counts.trees++;\n+\t\t\tstats->inflated_sizes.trees += inflated;\n+\t\t\tstats->disk_sizes.trees += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_BLOB:\n+\t\t\tstats->type_counts.blobs++;\n+\t\t\tstats->inflated_sizes.blobs += inflated;\n+\t\t\tstats->disk_sizes.blobs += disk;\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tBUG(\"invalid object type\");\n+\t\t}\n \t}\n \n \tobject_count = get_total_object_values(&stats->type_counts);\n-- \n2.53.0\n\n"},{"id":"536859","messageId":"20260223174120.2356504-3-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"[PATCH v2 2/5] builtin/repo: collect largest inflated objects","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-23T17:41:17Z","receivedAt":"2026-02-23T17:41:30Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The \"structure\" output for git-repo(1) shows the total inflated and disk\nsizes of reachable objects in the repository, but doesn't show the size\nof the largest individual objects. Since an individual object may be a\nlarge contributor to the overall repository size, it is useful for users\nto know the maximum size of individual objects.\n\nWhile interating across objects, record the size and OID of the largest\nobjects encountered for each object type to provide as output. Note that\nthe default \"table\" output format only displays size information and not\nthe corresponding OID. In a subsequent commit, the table format is\nupdated to add table annotations that mention the OID.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 63 +++++++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   | 28 +++++++++++++++++\n 3 files changed, 92 insertions(+)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 7d70270dfa..e812e59158 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -52,6 +52,7 @@ supported:\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n * Total disk size of reachable objects by type\n+* Largest reachable objects in the repository by type\n +\n The output format can be chosen through the flag `--format`. Three formats are\n supported:\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex c7c9f0f497..51a4359685 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,7 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hash.h\"\n #include \"hex.h\"\n #include \"odb.h\"\n #include \"parse-options.h\"\n@@ -197,6 +198,18 @@ static int cmd_repo_info(int argc, const char **argv, const char *prefix,\n \t\treturn print_fields(argc, argv, repo, format);\n }\n \n+struct object_data {\n+\tstruct object_id oid;\n+\tsize_t value;\n+};\n+\n+struct largest_objects {\n+\tstruct object_data tag_size;\n+\tstruct object_data commit_size;\n+\tstruct object_data tree_size;\n+\tstruct object_data blob_size;\n+};\n+\n struct ref_stats {\n \tsize_t branches;\n \tsize_t remotes;\n@@ -215,6 +228,7 @@ struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n \tstruct object_values disk_sizes;\n+\tstruct largest_objects largest;\n };\n \n struct repo_structure {\n@@ -371,6 +385,21 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->disk_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tstats_table_addf(table, \"\");\n+\tstats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->largest.commit_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->largest.tree_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->largest.blob_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Tags\"));\n+\tstats_table_size_addf(table, objects->largest.tag_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -485,6 +514,23 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n \n+\tprintf(\"objects.commits.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.commit_size.value, value_delim);\n+\tprintf(\"objects.commits.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.commit_size.oid), value_delim);\n+\tprintf(\"objects.trees.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.tree_size.value, value_delim);\n+\tprintf(\"objects.trees.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.tree_size.oid), value_delim);\n+\tprintf(\"objects.blobs.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.blob_size.value, value_delim);\n+\tprintf(\"objects.blobs.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.blob_size.oid), value_delim);\n+\tprintf(\"objects.tags.max_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.tag_size.value, value_delim);\n+\tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -553,6 +599,15 @@ struct count_objects_data {\n \tstruct progress *progress;\n };\n \n+static void check_largest(struct object_data *data, struct object_id *oid,\n+\t\t\t  size_t value)\n+{\n+\tif (value > data->value) {\n+\t\toidcpy(&data->oid, oid);\n+\t\tdata->value = value;\n+\t}\n+}\n+\n static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t enum object_type type, void *cb_data)\n {\n@@ -578,21 +633,29 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tstats->type_counts.tags++;\n \t\t\tstats->inflated_sizes.tags += inflated;\n \t\t\tstats->disk_sizes.tags += disk;\n+\t\t\tcheck_largest(&stats->largest.tag_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_COMMIT:\n \t\t\tstats->type_counts.commits++;\n \t\t\tstats->inflated_sizes.commits += inflated;\n \t\t\tstats->disk_sizes.commits += disk;\n+\t\t\tcheck_largest(&stats->largest.commit_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tstats->type_counts.trees++;\n \t\t\tstats->inflated_sizes.trees += inflated;\n \t\t\tstats->disk_sizes.trees += disk;\n+\t\t\tcheck_largest(&stats->largest.tree_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_BLOB:\n \t\t\tstats->type_counts.blobs++;\n \t\t\tstats->inflated_sizes.blobs += inflated;\n \t\t\tstats->disk_sizes.blobs += disk;\n+\t\t\tcheck_largest(&stats->largest.blob_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 17ff164b05..1999f325d0 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -52,6 +52,16 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|                      |        |\n+\t\t| * Largest objects    |        |\n+\t\t|   * Commits          |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Trees            |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Blobs            |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Tags             |        |\n+\t\t|     * Maximum size   |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -104,6 +114,16 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Trees          | $(object_type_disk_usage tree true) |\n \t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n \t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n+\t\t|                      |            |\n+\t\t| * Largest objects    |            |\n+\t\t|   * Commits          |            |\n+\t\t|     * Maximum size   |    223 B   |\n+\t\t|   * Trees            |            |\n+\t\t|     * Maximum size   |  32.29 KiB |\n+\t\t|   * Blobs            |            |\n+\t\t|     * Maximum size   |     13 B   |\n+\t\t|   * Tags             |            |\n+\t\t|     * Maximum size   |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -138,6 +158,14 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n \t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n \t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n+\t\tobjects.commits.max_size=221\n+\t\tobjects.commits.max_size_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n+\t\tobjects.trees.max_size=1335\n+\t\tobjects.trees.max_size_oid=09931deea9d81ec21300d3e13c74412f32eacec5\n+\t\tobjects.blobs.max_size=11\n+\t\tobjects.blobs.max_size_oid=eaeeedced46482bd4281fda5a5f05ce24854151f\n+\t\tobjects.tags.max_size=132\n+\t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"536860","messageId":"20260223174120.2356504-4-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"[PATCH v2 3/5] builtin/repo: add OID annotations to table output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-23T17:41:18Z","receivedAt":"2026-02-23T17:41:31Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The \"structure\" output for git-repo(1) does not show the corresponding\nOIDs for the largest objects in its \"table\" output. Update the output to\ninclude a list of OID annotations with an index to the corresponding row\nin the table.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            |  77 +++++++++++++++++---\n t/t1901-repo-structure.sh | 145 ++++++++++++++++++++------------------\n 2 files changed, 142 insertions(+), 80 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 51a4359685..bdf2820463 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -238,6 +238,7 @@ struct repo_structure {\n \n struct stats_table {\n \tstruct string_list rows;\n+\tstruct string_list annotations;\n \n \tint name_col_width;\n \tint value_col_width;\n@@ -250,6 +251,8 @@ struct stats_table {\n struct stats_table_entry {\n \tchar *value;\n \tconst char *unit;\n+\tsize_t index;\n+\tstruct object_id *oid;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -272,6 +275,12 @@ static void stats_table_vaddf(struct stats_table *table,\n \t\ttable->name_col_width = name_width;\n \tif (!entry)\n \t\treturn;\n+\tif (entry->oid) {\n+\t\tentry->index = table->annotations.nr + 1;\n+\t\tstrbuf_addf(&buf, \"[%\" PRIuMAX \"] %s\", (uintmax_t)entry->index,\n+\t\t\t    oid_to_hex(entry->oid));\n+\t\tstring_list_append_nodup(&table->annotations, strbuf_detach(&buf, NULL));\n+\t}\n \tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n@@ -282,6 +291,8 @@ static void stats_table_vaddf(struct stats_table *table,\n \t\tif (unit_width > table->unit_col_width)\n \t\t\ttable->unit_col_width = unit_width;\n \t}\n+\n+\tstrbuf_release(&buf);\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -321,6 +332,27 @@ static void stats_table_size_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_object_size_addf(struct stats_table *table,\n+\t\t\t\t\t struct object_id *oid, size_t value,\n+\t\t\t\t\t const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n+\n+\t/*\n+\t * A NULL OID should not have a table annotation.\n+\t */\n+\tif (!is_null_oid(oid))\n+\t\tentry->oid = oid;\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -389,19 +421,29 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Commits\"));\n-\tstats_table_size_addf(table, objects->largest.commit_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.commit_size.oid,\n+\t\t\t\t     objects->largest.commit_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n-\tstats_table_size_addf(table, objects->largest.tree_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.tree_size.oid,\n+\t\t\t\t     objects->largest.tree_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n-\tstats_table_size_addf(table, objects->largest.blob_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.blob_size.oid,\n+\t\t\t\t     objects->largest.blob_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Tags\"));\n-\tstats_table_size_addf(table, objects->largest.tag_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.tag_size.oid,\n+\t\t\t\t     objects->largest.tag_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n }\n \n+#define INDEX_WIDTH 4\n+\n static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n@@ -420,7 +462,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width + INDEX_WIDTH,\n+\t\t\t  name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n \t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n@@ -428,7 +471,7 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tprintf(\"%s\\n\", buf.buf);\n \n \tprintf(\"| \");\n-\tfor (int i = 0; i < name_col_width; i++)\n+\tfor (int i = 0; i < name_col_width + INDEX_WIDTH; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n \tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n@@ -450,6 +493,13 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_reset(&buf);\n \t\tstrbuf_addstr(&buf, \"| \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n+\n+\t\tif (entry && entry->oid)\n+\t\t\tstrbuf_addf(&buf, \" [%\" PRIuMAX \"]\",\n+\t\t\t\t    (uintmax_t)entry->index);\n+\t\telse\n+\t\t\tstrbuf_addchars(&buf, ' ', INDEX_WIDTH);\n+\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n \t\tstrbuf_addch(&buf, ' ');\n@@ -458,6 +508,11 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\n \n+\tif (table->annotations.nr)\n+\t\tprintf(\"\\n\");\n+\tfor_each_string_list_item(item, &table->annotations)\n+\t\tprintf(\"%s\\n\", item->string);\n+\n \tstrbuf_release(&buf);\n }\n \n@@ -473,6 +528,7 @@ static void stats_table_clear(struct stats_table *table)\n \t}\n \n \tstring_list_clear(&table->rows, 1);\n+\tstring_list_clear(&table->annotations, 1);\n }\n \n static void structure_keyvalue_print(struct repo_structure *stats,\n@@ -695,6 +751,7 @@ static int cmd_repo_structure(int argc, const char **argv, const char *prefix,\n {\n \tstruct stats_table table = {\n \t\t.rows = STRING_LIST_INIT_DUP,\n+\t\t.annotations = STRING_LIST_INIT_DUP,\n \t};\n \tenum output_format format = FORMAT_TABLE;\n \tstruct repo_structure stats = { 0 };\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 1999f325d0..918af7269f 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -27,41 +27,41 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    0   |\n-\t\t|     * Branches       |    0   |\n-\t\t|     * Tags           |    0   |\n-\t\t|     * Remotes        |    0   |\n-\t\t|     * Others         |    0   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            |    0   |\n-\t\t|     * Commits        |    0   |\n-\t\t|     * Trees          |    0   |\n-\t\t|     * Blobs          |    0   |\n-\t\t|     * Tags           |    0   |\n-\t\t|   * Inflated size    |    0 B |\n-\t\t|     * Commits        |    0 B |\n-\t\t|     * Trees          |    0 B |\n-\t\t|     * Blobs          |    0 B |\n-\t\t|     * Tags           |    0 B |\n-\t\t|   * Disk size        |    0 B |\n-\t\t|     * Commits        |    0 B |\n-\t\t|     * Trees          |    0 B |\n-\t\t|     * Blobs          |    0 B |\n-\t\t|     * Tags           |    0 B |\n-\t\t|                      |        |\n-\t\t| * Largest objects    |        |\n-\t\t|   * Commits          |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Trees            |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Blobs            |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Tags             |        |\n-\t\t|     * Maximum size   |    0 B |\n+\t\t| Repository structure     | Value  |\n+\t\t| ------------------------ | ------ |\n+\t\t| * References             |        |\n+\t\t|   * Count                |    0   |\n+\t\t|     * Branches           |    0   |\n+\t\t|     * Tags               |    0   |\n+\t\t|     * Remotes            |    0   |\n+\t\t|     * Others             |    0   |\n+\t\t|                          |        |\n+\t\t| * Reachable objects      |        |\n+\t\t|   * Count                |    0   |\n+\t\t|     * Commits            |    0   |\n+\t\t|     * Trees              |    0   |\n+\t\t|     * Blobs              |    0   |\n+\t\t|     * Tags               |    0   |\n+\t\t|   * Inflated size        |    0 B |\n+\t\t|     * Commits            |    0 B |\n+\t\t|     * Trees              |    0 B |\n+\t\t|     * Blobs              |    0 B |\n+\t\t|     * Tags               |    0 B |\n+\t\t|   * Disk size            |    0 B |\n+\t\t|     * Commits            |    0 B |\n+\t\t|     * Trees              |    0 B |\n+\t\t|     * Blobs              |    0 B |\n+\t\t|     * Tags               |    0 B |\n+\t\t|                          |        |\n+\t\t| * Largest objects        |        |\n+\t\t|   * Commits              |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Trees                |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Blobs                |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Tags                 |        |\n+\t\t|     * Maximum size       |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -89,41 +89,46 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# git-rev-list(1) --disk-usage=human option printing the full\n \t\t# \"byte/bytes\" unit string instead of just \"B\".\n \t\tcat >expect <<-EOF &&\n-\t\t| Repository structure | Value      |\n-\t\t| -------------------- | ---------- |\n-\t\t| * References         |            |\n-\t\t|   * Count            |      4     |\n-\t\t|     * Branches       |      1     |\n-\t\t|     * Tags           |      1     |\n-\t\t|     * Remotes        |      1     |\n-\t\t|     * Others         |      1     |\n-\t\t|                      |            |\n-\t\t| * Reachable objects  |            |\n-\t\t|   * Count            |   3.02 k   |\n-\t\t|     * Commits        |   1.01 k   |\n-\t\t|     * Trees          |   1.01 k   |\n-\t\t|     * Blobs          |   1.01 k   |\n-\t\t|     * Tags           |      1     |\n-\t\t|   * Inflated size    |  16.03 MiB |\n-\t\t|     * Commits        | 217.92 KiB |\n-\t\t|     * Trees          |  15.81 MiB |\n-\t\t|     * Blobs          |  11.68 KiB |\n-\t\t|     * Tags           |    132 B   |\n-\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n-\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n-\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n-\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n-\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n-\t\t|                      |            |\n-\t\t| * Largest objects    |            |\n-\t\t|   * Commits          |            |\n-\t\t|     * Maximum size   |    223 B   |\n-\t\t|   * Trees            |            |\n-\t\t|     * Maximum size   |  32.29 KiB |\n-\t\t|   * Blobs            |            |\n-\t\t|     * Maximum size   |     13 B   |\n-\t\t|   * Tags             |            |\n-\t\t|     * Maximum size   |    132 B   |\n+\t\t| Repository structure     | Value      |\n+\t\t| ------------------------ | ---------- |\n+\t\t| * References             |            |\n+\t\t|   * Count                |      4     |\n+\t\t|     * Branches           |      1     |\n+\t\t|     * Tags               |      1     |\n+\t\t|     * Remotes            |      1     |\n+\t\t|     * Others             |      1     |\n+\t\t|                          |            |\n+\t\t| * Reachable objects      |            |\n+\t\t|   * Count                |   3.02 k   |\n+\t\t|     * Commits            |   1.01 k   |\n+\t\t|     * Trees              |   1.01 k   |\n+\t\t|     * Blobs              |   1.01 k   |\n+\t\t|     * Tags               |      1     |\n+\t\t|   * Inflated size        |  16.03 MiB |\n+\t\t|     * Commits            | 217.92 KiB |\n+\t\t|     * Trees              |  15.81 MiB |\n+\t\t|     * Blobs              |  11.68 KiB |\n+\t\t|     * Tags               |    132 B   |\n+\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n+\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n+\t\t|                          |            |\n+\t\t| * Largest objects        |            |\n+\t\t|   * Commits              |            |\n+\t\t|     * Maximum size   [1] |    223 B   |\n+\t\t|   * Trees                |            |\n+\t\t|     * Maximum size   [2] |  32.29 KiB |\n+\t\t|   * Blobs                |            |\n+\t\t|     * Maximum size   [3] |     13 B   |\n+\t\t|   * Tags                 |            |\n+\t\t|     * Maximum size   [4] |    132 B   |\n+\n+\t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n+\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"536861","messageId":"20260223174120.2356504-5-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"[PATCH v2 4/5] builtin/repo: find commit with most parents","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-23T17:41:19Z","receivedAt":"2026-02-23T17:41:32Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Complex merge events may produce an octopus merge where the resulting\nmerge commit has more than two parents. While iterating through objects\nin the repository for git-repo-structure, identify the commit with the\nmost parents and display it in the output.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            |  47 ++++++++++++\n t/t1901-repo-structure.sh | 151 ++++++++++++++++++++------------------\n 2 files changed, 125 insertions(+), 73 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex bdf2820463..97da147f68 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -1,6 +1,7 @@\n #define USE_THE_REPOSITORY_VARIABLE\n \n #include \"builtin.h\"\n+#include \"commit.h\"\n #include \"environment.h\"\n #include \"hash.h\"\n #include \"hex.h\"\n@@ -208,6 +209,8 @@ struct largest_objects {\n \tstruct object_data commit_size;\n \tstruct object_data tree_size;\n \tstruct object_data blob_size;\n+\n+\tstruct object_data parent_count;\n };\n \n struct ref_stats {\n@@ -318,6 +321,27 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_object_count_addf(struct stats_table *table,\n+\t\t\t\t\t  struct object_id *oid, size_t value,\n+\t\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_count(value, &entry->value, &entry->unit);\n+\n+\t/*\n+\t * A NULL OID should not have a table annotation.\n+\t */\n+\tif (!is_null_oid(oid))\n+\t\tentry->oid = oid;\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static void stats_table_size_addf(struct stats_table *table, size_t value,\n \t\t\t\t  const char *format, ...)\n {\n@@ -425,6 +449,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t\t     &objects->largest.commit_size.oid,\n \t\t\t\t     objects->largest.commit_size.value,\n \t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_count_addf(table,\n+\t\t\t\t      &objects->largest.parent_count.oid,\n+\t\t\t\t      objects->largest.parent_count.value,\n+\t\t\t\t      \"    * %s\", _(\"Maximum parents\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n \tstats_table_object_size_addf(table,\n \t\t\t\t     &objects->largest.tree_size.oid,\n@@ -587,6 +615,11 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n \t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n \n+\tprintf(\"objects.commits.max_parents%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.parent_count.value, value_delim);\n+\tprintf(\"objects.commits.max_parents_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.parent_count.oid), value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -674,16 +707,24 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\tstruct commit *commit;\n+\t\tstruct object *obj;\n+\t\tvoid *content;\n \t\toff_t disk;\n+\t\tint eaten;\n \n \t\toi.sizep = &inflated;\n \t\toi.disk_sizep = &disk;\n+\t\toi.contentp = &content;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n \t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n \t\t\tcontinue;\n \n+\t\tobj = parse_object_buffer(the_repository, &oids->oid[i], type,\n+\t\t\t\t\t  inflated, content, &eaten);\n+\n \t\tswitch (type) {\n \t\tcase OBJ_TAG:\n \t\t\tstats->type_counts.tags++;\n@@ -693,11 +734,14 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_COMMIT:\n+\t\t\tcommit = object_as_type(obj, OBJ_COMMIT, 0);\n \t\t\tstats->type_counts.commits++;\n \t\t\tstats->inflated_sizes.commits += inflated;\n \t\t\tstats->disk_sizes.commits += disk;\n \t\t\tcheck_largest(&stats->largest.commit_size, &oids->oid[i],\n \t\t\t\t      inflated);\n+\t\t\tcheck_largest(&stats->largest.parent_count, &oids->oid[i],\n+\t\t\t\t      commit_list_count(commit->parents));\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tstats->type_counts.trees++;\n@@ -716,6 +760,9 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\tdefault:\n \t\t\tBUG(\"invalid object type\");\n \t\t}\n+\n+\t\tif (!eaten)\n+\t\t\tfree(content);\n \t}\n \n \tobject_count = get_total_object_values(&stats->type_counts);\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 918af7269f..d003d64a8e 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -27,41 +27,42 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure     | Value  |\n-\t\t| ------------------------ | ------ |\n-\t\t| * References             |        |\n-\t\t|   * Count                |    0   |\n-\t\t|     * Branches           |    0   |\n-\t\t|     * Tags               |    0   |\n-\t\t|     * Remotes            |    0   |\n-\t\t|     * Others             |    0   |\n-\t\t|                          |        |\n-\t\t| * Reachable objects      |        |\n-\t\t|   * Count                |    0   |\n-\t\t|     * Commits            |    0   |\n-\t\t|     * Trees              |    0   |\n-\t\t|     * Blobs              |    0   |\n-\t\t|     * Tags               |    0   |\n-\t\t|   * Inflated size        |    0 B |\n-\t\t|     * Commits            |    0 B |\n-\t\t|     * Trees              |    0 B |\n-\t\t|     * Blobs              |    0 B |\n-\t\t|     * Tags               |    0 B |\n-\t\t|   * Disk size            |    0 B |\n-\t\t|     * Commits            |    0 B |\n-\t\t|     * Trees              |    0 B |\n-\t\t|     * Blobs              |    0 B |\n-\t\t|     * Tags               |    0 B |\n-\t\t|                          |        |\n-\t\t| * Largest objects        |        |\n-\t\t|   * Commits              |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Trees                |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Blobs                |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Tags                 |        |\n-\t\t|     * Maximum size       |    0 B |\n+\t\t| Repository structure      | Value  |\n+\t\t| ------------------------- | ------ |\n+\t\t| * References              |        |\n+\t\t|   * Count                 |    0   |\n+\t\t|     * Branches            |    0   |\n+\t\t|     * Tags                |    0   |\n+\t\t|     * Remotes             |    0   |\n+\t\t|     * Others              |    0   |\n+\t\t|                           |        |\n+\t\t| * Reachable objects       |        |\n+\t\t|   * Count                 |    0   |\n+\t\t|     * Commits             |    0   |\n+\t\t|     * Trees               |    0   |\n+\t\t|     * Blobs               |    0   |\n+\t\t|     * Tags                |    0   |\n+\t\t|   * Inflated size         |    0 B |\n+\t\t|     * Commits             |    0 B |\n+\t\t|     * Trees               |    0 B |\n+\t\t|     * Blobs               |    0 B |\n+\t\t|     * Tags                |    0 B |\n+\t\t|   * Disk size             |    0 B |\n+\t\t|     * Commits             |    0 B |\n+\t\t|     * Trees               |    0 B |\n+\t\t|     * Blobs               |    0 B |\n+\t\t|     * Tags                |    0 B |\n+\t\t|                           |        |\n+\t\t| * Largest objects         |        |\n+\t\t|   * Commits               |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|     * Maximum parents     |    0   |\n+\t\t|   * Trees                 |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|   * Blobs                 |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|   * Tags                  |        |\n+\t\t|     * Maximum size        |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -89,46 +90,48 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# git-rev-list(1) --disk-usage=human option printing the full\n \t\t# \"byte/bytes\" unit string instead of just \"B\".\n \t\tcat >expect <<-EOF &&\n-\t\t| Repository structure     | Value      |\n-\t\t| ------------------------ | ---------- |\n-\t\t| * References             |            |\n-\t\t|   * Count                |      4     |\n-\t\t|     * Branches           |      1     |\n-\t\t|     * Tags               |      1     |\n-\t\t|     * Remotes            |      1     |\n-\t\t|     * Others             |      1     |\n-\t\t|                          |            |\n-\t\t| * Reachable objects      |            |\n-\t\t|   * Count                |   3.02 k   |\n-\t\t|     * Commits            |   1.01 k   |\n-\t\t|     * Trees              |   1.01 k   |\n-\t\t|     * Blobs              |   1.01 k   |\n-\t\t|     * Tags               |      1     |\n-\t\t|   * Inflated size        |  16.03 MiB |\n-\t\t|     * Commits            | 217.92 KiB |\n-\t\t|     * Trees              |  15.81 MiB |\n-\t\t|     * Blobs              |  11.68 KiB |\n-\t\t|     * Tags               |    132 B   |\n-\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n-\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n-\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n-\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n-\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n-\t\t|                          |            |\n-\t\t| * Largest objects        |            |\n-\t\t|   * Commits              |            |\n-\t\t|     * Maximum size   [1] |    223 B   |\n-\t\t|   * Trees                |            |\n-\t\t|     * Maximum size   [2] |  32.29 KiB |\n-\t\t|   * Blobs                |            |\n-\t\t|     * Maximum size   [3] |     13 B   |\n-\t\t|   * Tags                 |            |\n-\t\t|     * Maximum size   [4] |    132 B   |\n+\t\t| Repository structure      | Value      |\n+\t\t| ------------------------- | ---------- |\n+\t\t| * References              |            |\n+\t\t|   * Count                 |      4     |\n+\t\t|     * Branches            |      1     |\n+\t\t|     * Tags                |      1     |\n+\t\t|     * Remotes             |      1     |\n+\t\t|     * Others              |      1     |\n+\t\t|                           |            |\n+\t\t| * Reachable objects       |            |\n+\t\t|   * Count                 |   3.02 k   |\n+\t\t|     * Commits             |   1.01 k   |\n+\t\t|     * Trees               |   1.01 k   |\n+\t\t|     * Blobs               |   1.01 k   |\n+\t\t|     * Tags                |      1     |\n+\t\t|   * Inflated size         |  16.03 MiB |\n+\t\t|     * Commits             | 217.92 KiB |\n+\t\t|     * Trees               |  15.81 MiB |\n+\t\t|     * Blobs               |  11.68 KiB |\n+\t\t|     * Tags                |    132 B   |\n+\t\t|   * Disk size             | $(object_type_disk_usage all true) |\n+\t\t|     * Commits             | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees               | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs               |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags                |    $(object_type_disk_usage tag) B   |\n+\t\t|                           |            |\n+\t\t| * Largest objects         |            |\n+\t\t|   * Commits               |            |\n+\t\t|     * Maximum size    [1] |    223 B   |\n+\t\t|     * Maximum parents [2] |      1     |\n+\t\t|   * Trees                 |            |\n+\t\t|     * Maximum size    [3] |  32.29 KiB |\n+\t\t|   * Blobs                 |            |\n+\t\t|     * Maximum size    [4] |     13 B   |\n+\t\t|   * Tags                  |            |\n+\t\t|     * Maximum size    [5] |    132 B   |\n \n \t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n-\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n-\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n-\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n+\t\t[2] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n+\t\t[3] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[4] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[5] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -171,6 +174,8 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.blobs.max_size_oid=eaeeedced46482bd4281fda5a5f05ce24854151f\n \t\tobjects.tags.max_size=132\n \t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n+\t\tobjects.commits.max_parents=1\n+\t\tobjects.commits.max_parents_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"536862","messageId":"20260223174120.2356504-6-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"[PATCH v2 5/5] builtin/repo: find tree with most entries","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-23T17:41:20Z","receivedAt":"2026-02-23T17:41:33Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The size of a tree object usually corresponds with the number of entries\nit has. While iterating through objects in the repository for\ngit-repo-structure, identify the tree with the most entries and display\nit in the output.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 27 +++++++++++++++++++++++++++\n t/t1901-repo-structure.sh | 13 +++++++++----\n 2 files changed, 36 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 97da147f68..349cb27aca 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -16,6 +16,8 @@\n #include \"strbuf.h\"\n #include \"string-list.h\"\n #include \"shallow.h\"\n+#include \"tree.h\"\n+#include \"tree-walk.h\"\n #include \"utf8.h\"\n \n static const char *const repo_usage[] = {\n@@ -211,6 +213,7 @@ struct largest_objects {\n \tstruct object_data blob_size;\n \n \tstruct object_data parent_count;\n+\tstruct object_data tree_entries;\n };\n \n struct ref_stats {\n@@ -458,6 +461,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t\t     &objects->largest.tree_size.oid,\n \t\t\t\t     objects->largest.tree_size.value,\n \t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_count_addf(table,\n+\t\t\t\t      &objects->largest.tree_entries.oid,\n+\t\t\t\t      objects->largest.tree_entries.value,\n+\t\t\t\t      \"    * %s\", _(\"Maximum entries\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n \tstats_table_object_size_addf(table,\n \t\t\t\t     &objects->largest.blob_size.oid,\n@@ -619,6 +626,10 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \t       (uintmax_t)stats->objects.largest.parent_count.value, value_delim);\n \tprintf(\"objects.commits.max_parents_oid%c%s%c\", key_delim,\n \t       oid_to_hex(&stats->objects.largest.parent_count.oid), value_delim);\n+\tprintf(\"objects.trees.max_entries%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.largest.tree_entries.value, value_delim);\n+\tprintf(\"objects.trees.max_entries_oid%c%s%c\", key_delim,\n+\t       oid_to_hex(&stats->objects.largest.tree_entries.oid), value_delim);\n \n \tfflush(stdout);\n }\n@@ -697,6 +708,20 @@ static void check_largest(struct object_data *data, struct object_id *oid,\n \t}\n }\n \n+static size_t count_tree_entries(struct object *obj)\n+{\n+\tstruct tree *t = object_as_type(obj, OBJ_TREE, 0);\n+\tstruct name_entry entry;\n+\tstruct tree_desc desc;\n+\tsize_t count = 0;\n+\n+\tinit_tree_desc(&desc, &t->object.oid, t->buffer, t->size);\n+\twhile (tree_entry(&desc, &entry))\n+\t\tcount++;\n+\n+\treturn count;\n+}\n+\n static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t enum object_type type, void *cb_data)\n {\n@@ -749,6 +774,8 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tstats->disk_sizes.trees += disk;\n \t\t\tcheck_largest(&stats->largest.tree_size, &oids->oid[i],\n \t\t\t\t      inflated);\n+\t\t\tcheck_largest(&stats->largest.tree_entries, &oids->oid[i],\n+\t\t\t\t      count_tree_entries(obj));\n \t\t\tbreak;\n \t\tcase OBJ_BLOB:\n \t\t\tstats->type_counts.blobs++;\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex d003d64a8e..12ed67e846 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -59,6 +59,7 @@ test_expect_success 'empty repository' '\n \t\t|     * Maximum parents     |    0   |\n \t\t|   * Trees                 |        |\n \t\t|     * Maximum size        |    0 B |\n+\t\t|     * Maximum entries     |    0   |\n \t\t|   * Blobs                 |        |\n \t\t|     * Maximum size        |    0 B |\n \t\t|   * Tags                  |        |\n@@ -122,16 +123,18 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Maximum parents [2] |      1     |\n \t\t|   * Trees                 |            |\n \t\t|     * Maximum size    [3] |  32.29 KiB |\n+\t\t|     * Maximum entries [4] |   1.01 k   |\n \t\t|   * Blobs                 |            |\n-\t\t|     * Maximum size    [4] |     13 B   |\n+\t\t|     * Maximum size    [5] |     13 B   |\n \t\t|   * Tags                  |            |\n-\t\t|     * Maximum size    [5] |    132 B   |\n+\t\t|     * Maximum size    [6] |    132 B   |\n \n \t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n \t\t[2] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n \t\t[3] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n-\t\t[4] 97d808e45116bf02103490294d3d46dad7a2ac62\n-\t\t[5] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n+\t\t[4] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[5] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[6] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -176,6 +179,8 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n \t\tobjects.commits.max_parents=1\n \t\tobjects.commits.max_parents_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n+\t\tobjects.trees.max_entries=42\n+\t\tobjects.trees.max_entries_oid=09931deea9d81ec21300d3e13c74412f32eacec5\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"536955","messageId":"aZ1wy55xSaAOye49@pks.im","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"Re: [PATCH v2 0/5] builtin/repo: include largest object information","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-02-24T09:35:07Z","receivedAt":"2026-02-24T09:35:12Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Feb 23, 2026 at 11:41:15AM -0600, Justin Tobler wrote:\n> Range-diff against v1:\n> 1:  94a44e0e0f = 1:  94a44e0e0f builtin/repo: update stats for each object\n> 2:  92dbf34f2c = 2:  92dbf34f2c builtin/repo: collect largest inflated objects\n> 3:  1811d03afe ! 3:  1457d5d59c builtin/repo: add OID annotations to table output\n>     @@ builtin/repo.c: static void stats_table_vaddf(struct stats_table *table,\n>      +\t\tentry->index = table->annotations.nr + 1;\n>      +\t\tstrbuf_addf(&buf, \"[%\" PRIuMAX \"] %s\", (uintmax_t)entry->index,\n>      +\t\t\t    oid_to_hex(entry->oid));\n>     -+\t\tstring_list_append(&table->annotations, buf.buf);\n>     ++\t\tstring_list_append_nodup(&table->annotations, strbuf_detach(&buf, NULL));\n>      +\t}\n>       \tif (entry->value) {\n>       \t\tint value_width = utf8_strwidth(entry->value);\n> 4:  471d352cc1 = 4:  f4e92e3f09 builtin/repo: find commit with most parents\n> 5:  7f1b7f9657 = 5:  af404fcc6c builtin/repo: find tree with most entries\n\nThanks, this addresses my only comment I had on the first version.\n\nPatrick\n"},{"id":"537219","messageId":"xmqqzf4v1fdc.fsf@gitster.g","threadId":"64911","inReplyTo":"aZYUwjSEAcGRSXNa@denethor","subject":"Re: [PATCH 1/5] builtin/repo: update stats for each object","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-26T19:20:15Z","receivedAt":"2026-02-26T19:20:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> Good suggestion. Some of the info added in the following commits is\n> object specific and will need to be handled accordinly, but we could\n> probably still benefit by structuring the data a bit better. Will\n> explore in the next version.\n\nI am looking at v2 patches, but did this happen?\n"},{"id":"537220","messageId":"aaCefUd8JLpKAyPu@denethor","threadId":"64911","inReplyTo":"xmqqzf4v1fdc.fsf@gitster.g","subject":"Re: [PATCH 1/5] builtin/repo: update stats for each object","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-02-26T19:29:02Z","receivedAt":"2026-02-26T19:29:04Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/26 11:20AM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > Good suggestion. Some of the info added in the following commits is\n> > object specific and will need to be handled accordinly, but we could\n> > probably still benefit by structuring the data a bit better. Will\n> > explore in the next version.\n> \n> I am looking at v2 patches, but did this happen?\n\nApologies, I mentioned it in the cover letter, but should have replied\nto this thread also. In version 2 I kept this the same for now.\n\n-Justin\n"},{"id":"537221","messageId":"xmqqv7fj1dzg.fsf@gitster.g","threadId":"64911","inReplyTo":"20260223174120.2356504-3-jltobler@gmail.com","subject":"Re: [PATCH v2 2/5] builtin/repo: collect largest inflated objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-26T19:50:11Z","receivedAt":"2026-02-26T19:50:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> @@ -485,6 +514,23 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n>  \tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n>  \t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n>  \n> +\tprintf(\"objects.commits.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> +\t       (uintmax_t)stats->objects.largest.commit_size.value, value_delim);\n> +\tprintf(\"objects.commits.max_size_oid%c%s%c\", key_delim,\n> +\t       oid_to_hex(&stats->objects.largest.commit_size.oid), value_delim);\n> +\tprintf(\"objects.trees.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> +\t       (uintmax_t)stats->objects.largest.tree_size.value, value_delim);\n> +\tprintf(\"objects.trees.max_size_oid%c%s%c\", key_delim,\n> +\t       oid_to_hex(&stats->objects.largest.tree_size.oid), value_delim);\n> +\tprintf(\"objects.blobs.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> +\t       (uintmax_t)stats->objects.largest.blob_size.value, value_delim);\n> +\tprintf(\"objects.blobs.max_size_oid%c%s%c\", key_delim,\n> +\t       oid_to_hex(&stats->objects.largest.blob_size.oid), value_delim);\n> +\tprintf(\"objects.tags.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> +\t       (uintmax_t)stats->objects.largest.tag_size.value, value_delim);\n> +\tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n> +\t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n\nThe repetition tires reviewers' eyes.  I am reasonably sure if there\nwere an intentional copy-and-paste error, I wouldn't be able to spot\nit.  But I tried to be careful and read it over three times ;-).\n\n> @@ -553,6 +599,15 @@ struct count_objects_data {\n>  \tstruct progress *progress;\n>  };\n>  \n> +static void check_largest(struct object_data *data, struct object_id *oid,\n> +\t\t\t  size_t value)\n> +{\n> +\tif (value > data->value) {\n> +\t\toidcpy(&data->oid, oid);\n> +\t\tdata->value = value;\n> +\t}\n> +}\n\nHow important is it for this application to end up with a valid\nvalue in data->oid?\n\nIf data->value is initialized to a valid value, instead of an\nimpossible sentinel value that is strictly smaller than any valid\nvalues, this can leave data->value to a valid value from an existing\nobject without recording its object name.  Imagine a repository with\na single empty blob, and data->value initialized to zero (it cannot\nbe initialized to a sentinel -1, as use of size_t here makes it\nimpossible to have any reasonable sentinel values).\n\n\n> @@ -138,6 +158,14 @@ test_expect_success SHA1 'keyvalue and nul format' '\n>  \t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n>  \t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n>  \t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n> +\t\tobjects.commits.max_size=221\n> +\t\tobjects.commits.max_size_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n> +\t\tobjects.trees.max_size=1335\n> +\t\tobjects.trees.max_size_oid=09931deea9d81ec21300d3e13c74412f32eacec5\n> +\t\tobjects.blobs.max_size=11\n> +\t\tobjects.blobs.max_size_oid=eaeeedced46482bd4281fda5a5f05ce24854151f\n> +\t\tobjects.tags.max_size=132\n> +\t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n>  \t\tEOF\n>  \n>  \t\tgit repo structure --format=keyvalue >out 2>err &&\n"},{"id":"537222","messageId":"xmqqpl5r1dom.fsf@gitster.g","threadId":"64911","inReplyTo":"20260223174120.2356504-4-jltobler@gmail.com","subject":"Re: [PATCH v2 3/5] builtin/repo: add OID annotations to table output","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-26T19:56:41Z","receivedAt":"2026-02-26T19:56:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> +\tif (table->annotations.nr)\n> +\t\tprintf(\"\\n\");\n> +\tfor_each_string_list_item(item, &table->annotations)\n> +\t\tprintf(\"%s\\n\", item->string);\n> +\n\nIt is minor, but I suspect\n\n\tif (table->annotations.nr) {\n\t\tprintf(\"\\n\");\n\t\tfor_each_string_list_item(...)\n\t\t\tprintf(\"%s\\n\", item->string);\n\t}\n\nwould be easier to reason about.\n\n"},{"id":"537419","messageId":"C99EBCF9-7980-495A-94C5-576AC6D140F3@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-3-jltobler@gmail.com","subject":"Re: [PATCH v2 2/5] builtin/repo: collect largest inflated objects","fromName":"Lucas Seiki Oshiro","fromEmail":"lucasseikioshiro@gmail.com","sentAt":"2026-02-28T23:36:26Z","receivedAt":"2026-02-28T23:36:41Z","isPatch":true,"sender":{"key":"lucasseikioshiro@gmail.com","avatar":"https://avatars.githubusercontent.com/u/12701580?v=4"},"body":"\n> struct repo_structure {\n> @@ -371,6 +385,21 @@ static void stats_table_setup_structure(struct stats_table *table,\n>      \"    * %s\", _(\"Blobs\"));\n> stats_table_size_addf(table, objects->disk_sizes.tags,\n>      \"    * %s\", _(\"Tags\"));\n> +\n> + stats_table_addf(table, \"\");\n> + stats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n> + stats_table_addf(table, \"  * %s\", _(\"Commits\"));\n> + stats_table_size_addf(table, objects->largest.commit_size.value,\n> +      \"    * %s\", _(\"Maximum size\"));\n\nI don't know if it's the best place to comment this, but it would be\nnice if we could find the commit that introduced the largest change,\nin terms of size or number of lines.\n\nThis would be useful for people who are asking \"what's the largest\ncommmit?\" thinking about the introduced changes (like what we see in\nGitLab's interface) instead of the size of the commit object, which\ngenerally is proportional to the message size + the number of\nparents."},{"id":"537420","messageId":"EB04AA40-87BA-41D9-B2DC-92E87FACEB54@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"Re: [PATCH v2 0/5] builtin/repo: include largest object information","fromName":"Lucas Seiki Oshiro","fromEmail":"lucasseikioshiro@gmail.com","sentAt":"2026-02-28T23:43:56Z","receivedAt":"2026-02-28T23:44:12Z","isPatch":true,"sender":{"key":"lucasseikioshiro@gmail.com","avatar":"https://avatars.githubusercontent.com/u/12701580?v=4"},"body":"Hi, Justin!\n\nI was trying this patch series and I noticed that it took\nmore time to run than before. In my machine, I tested it\nwith the Git repository itself and it took 6s to run, while\nit took 3s to run in the current master [1].\n\nI understand the reason and I don't think we could avoid\nthat, but I'm wondering if wouldn't be nice to have some\nway to only retrieve the \"lighter\" data (perhaps a flag,\nor something like the keys in git-repo-info).\n\nThanks!\n\n\n[1] 2cc7191751 (The 8th batch, 2026-02-27)\n"},{"id":"537471","messageId":"aaR6a7o4omOIWJSe@denethor","threadId":"64911","inReplyTo":"EB04AA40-87BA-41D9-B2DC-92E87FACEB54@gmail.com","subject":"Re: [PATCH v2 0/5] builtin/repo: include largest object information","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-01T19:22:33Z","receivedAt":"2026-03-01T19:22:43Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/28 08:43PM, Lucas Seiki Oshiro wrote:\n> I was trying this patch series and I noticed that it took\n> more time to run than before. In my machine, I tested it\n> with the Git repository itself and it took 6s to run, while\n> it took 3s to run in the current master [1].\n\nYes, now that objects are being parsed to fetch additional\ncommit/tree information we incur some additional overhead when\ncollecting metrics.\n\nWith git-repo-structure, the goal is to provide the user with an\noverview of size/structure related statistics that may showcase problems\nfor a given repostiory and is directly inspired by git-sizer [1]. Thus\nas it currently stands, the implementation of git-repo-structure is\nstill incomplete and as we collect additional metrics in subseqent\nseries the performance characteristics may still change.\n\n> I understand the reason and I don't think we could avoid\n> that, but I'm wondering if wouldn't be nice to have some\n> way to only retrieve the \"lighter\" data (perhaps a flag,\n> or something like the keys in git-repo-info).\n\nIf the main motivation is to allow the user to reduce the time spent by\nselecting only a subset of metrics, I don't think using keys like\ngit-repo-info would be a good fit. Most of the collected metrics pull\nfrom the same data sources so including/excluding any given metric may\nnot have any bearing on actual performance. For example: if the user\nwants to collect largest object info which is a more expensive check, we\nstill have to collect the underlying data used by the other metrics\nregardless of if they are shown or not. Furthermore, it would likely not\nbe obvious to users which categories of metrics would be more expensive\nthan others.\n\nI could maybe see something akin to a `--[no-]extended` option that\nbreaks metrics into cheap/expensive categories and computes/displays the\nmetrics accordingly, but it would be important that the default set of\nmetrics collected satisfy the repository overview this command aims to\nprovide.\n\nIf we are more interested in adding a mechanism to filter\ngit-repo-structure results independent of performance considerations,\nmaybe we could eventually explore adding something like the\ngit-repo-info keys or a `--filter` option to restrict the output to a\nspecified subset. At the same time though, it is probably easy enough\nfor git-repo-structure users to filter the machine-parsable output\nthemselves if they wish to do so. For now I think this should be fine,\nbut an included result filtering option is still something we could\nexplore in the future. :)\n\nThanks,\n-Justin\n\n[1]: https://github.com/github/git-sizer\n"},{"id":"537558","messageId":"aaXFcz8AQrwRtr5C@denethor","threadId":"64911","inReplyTo":"xmqqv7fj1dzg.fsf@gitster.g","subject":"Re: [PATCH v2 2/5] builtin/repo: collect largest inflated objects","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T17:28:32Z","receivedAt":"2026-03-02T17:28:36Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/26 11:50AM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > @@ -485,6 +514,23 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n> >  \tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n> >  \t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n> >  \n> > +\tprintf(\"objects.commits.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> > +\t       (uintmax_t)stats->objects.largest.commit_size.value, value_delim);\n> > +\tprintf(\"objects.commits.max_size_oid%c%s%c\", key_delim,\n> > +\t       oid_to_hex(&stats->objects.largest.commit_size.oid), value_delim);\n> > +\tprintf(\"objects.trees.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> > +\t       (uintmax_t)stats->objects.largest.tree_size.value, value_delim);\n> > +\tprintf(\"objects.trees.max_size_oid%c%s%c\", key_delim,\n> > +\t       oid_to_hex(&stats->objects.largest.tree_size.oid), value_delim);\n> > +\tprintf(\"objects.blobs.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> > +\t       (uintmax_t)stats->objects.largest.blob_size.value, value_delim);\n> > +\tprintf(\"objects.blobs.max_size_oid%c%s%c\", key_delim,\n> > +\t       oid_to_hex(&stats->objects.largest.blob_size.oid), value_delim);\n> > +\tprintf(\"objects.tags.max_size%c%\" PRIuMAX \"%c\", key_delim,\n> > +\t       (uintmax_t)stats->objects.largest.tag_size.value, value_delim);\n> > +\tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n> > +\t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n> \n> The repetition tires reviewers' eyes.  I am reasonably sure if there\n> were an intentional copy-and-paste error, I wouldn't be able to spot\n> it.  But I tried to be careful and read it over three times ;-).\n\nYa, I was thinking about adding another patch that reduces the\nduplication for the output here. I'll go ahead and do that in the next\nversion.\n\n> > @@ -553,6 +599,15 @@ struct count_objects_data {\n> >  \tstruct progress *progress;\n> >  };\n> >  \n> > +static void check_largest(struct object_data *data, struct object_id *oid,\n> > +\t\t\t  size_t value)\n> > +{\n> > +\tif (value > data->value) {\n> > +\t\toidcpy(&data->oid, oid);\n> > +\t\tdata->value = value;\n> > +\t}\n> > +}\n> \n> How important is it for this application to end up with a valid\n> value in data->oid?\n> \n> If data->value is initialized to a valid value, instead of an\n> impossible sentinel value that is strictly smaller than any valid\n> values, this can leave data->value to a valid value from an existing\n> object without recording its object name.  Imagine a repository with\n> a single empty blob, and data->value initialized to zero (it cannot\n> be initialized to a sentinel -1, as use of size_t here makes it\n> impossible to have any reasonable sentinel values).\n\nSo in cases where we do not record an OID for an object, the table\noutput format knows not to show any annotations and the machine parsable\nformats display null OIDs. In the example you provided though, this\ntechnically wouldn't be correct though as it possible we could have an\nempty blob.\n\nOne way we could deal which this is have a sentinel value of -1 for the\nsize value as you mentioned. Another option could be to check if the OID\nis a null value and if so record the value regardless. I'll work on this\nin the next version.\n\nThanks,\n-Justin\n"},{"id":"537559","messageId":"aaXI0M7Ztk-Swm18@denethor","threadId":"64911","inReplyTo":"C99EBCF9-7980-495A-94C5-576AC6D140F3@gmail.com","subject":"Re: [PATCH v2 2/5] builtin/repo: collect largest inflated objects","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T17:38:16Z","receivedAt":"2026-03-02T17:38:18Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/28 08:36PM, Lucas Seiki Oshiro wrote:\n> \n> > struct repo_structure {\n> > @@ -371,6 +385,21 @@ static void stats_table_setup_structure(struct stats_table *table,\n> >      \"    * %s\", _(\"Blobs\"));\n> > stats_table_size_addf(table, objects->disk_sizes.tags,\n> >      \"    * %s\", _(\"Tags\"));\n> > +\n> > + stats_table_addf(table, \"\");\n> > + stats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n> > + stats_table_addf(table, \"  * %s\", _(\"Commits\"));\n> > + stats_table_size_addf(table, objects->largest.commit_size.value,\n> > +      \"    * %s\", _(\"Maximum size\"));\n> \n> I don't know if it's the best place to comment this, but it would be\n> nice if we could find the commit that introduced the largest change,\n> in terms of size or number of lines.\n> \n> This would be useful for people who are asking \"what's the largest\n> commmit?\" thinking about the introduced changes (like what we see in\n> GitLab's interface) instead of the size of the commit object, which\n> generally is proportional to the message size + the number of\n> parents.\n\nI assume by largest change we are referring to finding the commit that\nhas the most lines changed between it and its parent. This could be\ninteresting, but I suspect it could be quite costly to compute for large\nrepositories with many commits. This type of information is not actually\nstored in the repository and would have to be computed on the fly. Since\nthis information is not really part of the repository structure, it\nmight not be a great fit for this command either. I'm not quite sure\nabout this one.\n\nThanks,\n-Justin\n"},{"id":"537560","messageId":"aaXLH9Lt-K3_ojyC@denethor","threadId":"64911","inReplyTo":"xmqqpl5r1dom.fsf@gitster.g","subject":"Re: [PATCH v2 3/5] builtin/repo: add OID annotations to table output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T17:39:01Z","receivedAt":"2026-03-02T17:39:03Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/02/26 11:56AM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > +\tif (table->annotations.nr)\n> > +\t\tprintf(\"\\n\");\n> > +\tfor_each_string_list_item(item, &table->annotations)\n> > +\t\tprintf(\"%s\\n\", item->string);\n> > +\n> \n> It is minor, but I suspect\n> \n> \tif (table->annotations.nr) {\n> \t\tprintf(\"\\n\");\n> \t\tfor_each_string_list_item(...)\n> \t\t\tprintf(\"%s\\n\", item->string);\n> \t}\n> \n> would be easier to reason about.\n\nMakes sense, will update. Thanks.\n\n-Justin\n"},{"id":"537609","messageId":"20260302214526.2034279-1-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260223174120.2356504-1-jltobler@gmail.com","subject":"[PATCH v3 0/6] builtin/repo: include largest object information","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T21:45:20Z","receivedAt":"2026-03-02T21:45:32Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThe \"structure\" output for git-repo(1) currently provides count\ninformation for references/objects as well as total inflated/disk sizes\nof objects by type. Info regarding the largest individual objects in the\nrepository is not yet collected, but would be useful to users wishing to\nidentify such large objects.\n\nThis patch series adds the following data points:\n- The OID and size of the largest objects by object type\n- The OID and parent count of the commit with the most parents\n- The OID and entries count of the tree with the most entries\n\nChanges from V2:\n- When checking for largest objects, zero valued objects were not\n  recorded even if they were the \"largest\" object. In this version, if\n  an object ID has not been recorded yet, it is always added even if its\n  value is zero.\n- Added some helper functions for printing keyvalue info to cut down on\n  duplicate code and hopefully make it a bit easier on the eyes.\n- Moved the for-each loop that printed table OID annoations inside the\n  preceding if-block making it a bit easier to reason about.\n\nChanges from V1:\n- Avoided duplicating the annotation string by handing over ownership.\n- I decided to leave the `struct object_stats` structure alone for now\n  as storing the various object values per-type does make it convenient\n  to calulate the various totals. I may revisit this in a future series\n  though.\n\nThanks,\n-Justin\n\nJustin Tobler (6):\n  builtin/repo: update stats for each object\n  builtin/repo: add helper for printing keyvalue output\n  builtin/repo: collect largest inflated objects\n  builtin/repo: add OID annotations to table output\n  builtin/repo: find commit with most parents\n  builtin/repo: find tree with most entries\n\n Documentation/git-repo.adoc |   1 +\n builtin/repo.c              | 323 ++++++++++++++++++++++++++++--------\n t/t1901-repo-structure.sh   | 143 ++++++++++------\n 3 files changed, 352 insertions(+), 115 deletions(-)\n\nRange-diff against v2:\n1:  94a44e0e0f = 1:  94a44e0e0f builtin/repo: update stats for each object\n-:  ---------- > 2:  36c11351ae builtin/repo: add helper for printing keyvalue output\n2:  92dbf34f2c ! 3:  90e71c058d builtin/repo: collect largest inflated objects\n    @@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *tabl\n      }\n      \n      static void stats_table_print_structure(const struct stats_table *table)\n    +@@ builtin/repo.c: static inline void print_keyvalue(const char *key, char key_delim, size_t value,\n    + \t       value_delim);\n    + }\n    + \n    ++static void print_object_data(const char *key, char key_delim,\n    ++\t\t\t      struct object_data *data, char value_delim)\n    ++{\n    ++\tprint_keyvalue(key, key_delim, data->value, value_delim);\n    ++\tprintf(\"%s_oid%c%s%c\", key, key_delim, oid_to_hex(&data->oid),\n    ++\t       value_delim);\n    ++}\n    ++\n    + static void structure_keyvalue_print(struct repo_structure *stats,\n    + \t\t\t\t     char key_delim, char value_delim)\n    + {\n     @@ builtin/repo.c: static void structure_keyvalue_print(struct repo_structure *stats,\n    - \tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n    - \t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n    + \tprint_keyvalue(\"objects.tags.disk_size\", key_delim,\n    + \t\t       stats->objects.disk_sizes.tags, value_delim);\n      \n    -+\tprintf(\"objects.commits.max_size%c%\" PRIuMAX \"%c\", key_delim,\n    -+\t       (uintmax_t)stats->objects.largest.commit_size.value, value_delim);\n    -+\tprintf(\"objects.commits.max_size_oid%c%s%c\", key_delim,\n    -+\t       oid_to_hex(&stats->objects.largest.commit_size.oid), value_delim);\n    -+\tprintf(\"objects.trees.max_size%c%\" PRIuMAX \"%c\", key_delim,\n    -+\t       (uintmax_t)stats->objects.largest.tree_size.value, value_delim);\n    -+\tprintf(\"objects.trees.max_size_oid%c%s%c\", key_delim,\n    -+\t       oid_to_hex(&stats->objects.largest.tree_size.oid), value_delim);\n    -+\tprintf(\"objects.blobs.max_size%c%\" PRIuMAX \"%c\", key_delim,\n    -+\t       (uintmax_t)stats->objects.largest.blob_size.value, value_delim);\n    -+\tprintf(\"objects.blobs.max_size_oid%c%s%c\", key_delim,\n    -+\t       oid_to_hex(&stats->objects.largest.blob_size.oid), value_delim);\n    -+\tprintf(\"objects.tags.max_size%c%\" PRIuMAX \"%c\", key_delim,\n    -+\t       (uintmax_t)stats->objects.largest.tag_size.value, value_delim);\n    -+\tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n    -+\t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n    ++\tprint_object_data(\"objects.commits.max_size\", key_delim,\n    ++\t\t\t  &stats->objects.largest.commit_size, value_delim);\n    ++\tprint_object_data(\"objects.trees.max_size\", key_delim,\n    ++\t\t\t  &stats->objects.largest.tree_size, value_delim);\n    ++\tprint_object_data(\"objects.blobs.max_size\", key_delim,\n    ++\t\t\t  &stats->objects.largest.blob_size, value_delim);\n    ++\tprint_object_data(\"objects.tags.max_size\", key_delim,\n    ++\t\t\t  &stats->objects.largest.tag_size, value_delim);\n     +\n      \tfflush(stdout);\n      }\n    @@ builtin/repo.c: struct count_objects_data {\n     +static void check_largest(struct object_data *data, struct object_id *oid,\n     +\t\t\t  size_t value)\n     +{\n    -+\tif (value > data->value) {\n    ++\tif (value > data->value || is_null_oid(&data->oid)) {\n     +\t\toidcpy(&data->oid, oid);\n     +\t\tdata->value = value;\n     +\t}\n3:  1457d5d59c ! 4:  938c36df91 builtin/repo: add OID annotations to table output\n    @@ builtin/repo.c: static void stats_table_print_structure(const struct stats_table\n      \t\tprintf(\"%s\\n\", buf.buf);\n      \t}\n      \n    -+\tif (table->annotations.nr)\n    ++\tif (table->annotations.nr) {\n     +\t\tprintf(\"\\n\");\n    -+\tfor_each_string_list_item(item, &table->annotations)\n    -+\t\tprintf(\"%s\\n\", item->string);\n    ++\t\tfor_each_string_list_item(item, &table->annotations)\n    ++\t\t\tprintf(\"%s\\n\", item->string);\n    ++\t}\n     +\n      \tstrbuf_release(&buf);\n      }\n    @@ builtin/repo.c: static void stats_table_clear(struct stats_table *table)\n     +\tstring_list_clear(&table->annotations, 1);\n      }\n      \n    - static void structure_keyvalue_print(struct repo_structure *stats,\n    + static inline void print_keyvalue(const char *key, char key_delim, size_t value,\n     @@ builtin/repo.c: static int cmd_repo_structure(int argc, const char **argv, const char *prefix,\n      {\n      \tstruct stats_table table = {\n4:  f4e92e3f09 ! 5:  ab9870f06e builtin/repo: find commit with most parents\n    @@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *tabl\n      \tstats_table_object_size_addf(table,\n      \t\t\t\t     &objects->largest.tree_size.oid,\n     @@ builtin/repo.c: static void structure_keyvalue_print(struct repo_structure *stats,\n    - \tprintf(\"objects.tags.max_size_oid%c%s%c\", key_delim,\n    - \t       oid_to_hex(&stats->objects.largest.tag_size.oid), value_delim);\n    + \tprint_object_data(\"objects.tags.max_size\", key_delim,\n    + \t\t\t  &stats->objects.largest.tag_size, value_delim);\n      \n    -+\tprintf(\"objects.commits.max_parents%c%\" PRIuMAX \"%c\", key_delim,\n    -+\t       (uintmax_t)stats->objects.largest.parent_count.value, value_delim);\n    -+\tprintf(\"objects.commits.max_parents_oid%c%s%c\", key_delim,\n    -+\t       oid_to_hex(&stats->objects.largest.parent_count.oid), value_delim);\n    ++\tprint_object_data(\"objects.commits.max_parents\", key_delim,\n    ++\t\t\t  &stats->objects.largest.parent_count, value_delim);\n     +\n      \tfflush(stdout);\n      }\n5:  af404fcc6c ! 6:  2884cb451c builtin/repo: find tree with most entries\n    @@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *tabl\n      \tstats_table_object_size_addf(table,\n      \t\t\t\t     &objects->largest.blob_size.oid,\n     @@ builtin/repo.c: static void structure_keyvalue_print(struct repo_structure *stats,\n    - \t       (uintmax_t)stats->objects.largest.parent_count.value, value_delim);\n    - \tprintf(\"objects.commits.max_parents_oid%c%s%c\", key_delim,\n    - \t       oid_to_hex(&stats->objects.largest.parent_count.oid), value_delim);\n    -+\tprintf(\"objects.trees.max_entries%c%\" PRIuMAX \"%c\", key_delim,\n    -+\t       (uintmax_t)stats->objects.largest.tree_entries.value, value_delim);\n    -+\tprintf(\"objects.trees.max_entries_oid%c%s%c\", key_delim,\n    -+\t       oid_to_hex(&stats->objects.largest.tree_entries.oid), value_delim);\n    + \n    + \tprint_object_data(\"objects.commits.max_parents\", key_delim,\n    + \t\t\t  &stats->objects.largest.parent_count, value_delim);\n    ++\tprint_object_data(\"objects.trees.max_entries\", key_delim,\n    ++\t\t\t  &stats->objects.largest.tree_entries, value_delim);\n      \n      \tfflush(stdout);\n      }\n\nbase-commit: 67ad42147a7acc2af6074753ebd03d904476118f\n-- \n2.53.0\n\n"},{"id":"537610","messageId":"20260302214526.2034279-2-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260302214526.2034279-1-jltobler@gmail.com","subject":"[PATCH v3 1/6] builtin/repo: update stats for each object","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T21:45:21Z","receivedAt":"2026-03-02T21:45:33Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"When walking reachable objects in the repository, `count_objects()`\nprocesses a set of objects and updates the `struct object_stats`. In\npreparation for more granular statistics being collected, update the\n`struct object_stats` for each individual object instead.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 53 +++++++++++++++++++++++---------------------------\n 1 file changed, 24 insertions(+), 29 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 0ea045abc1..c7c9f0f497 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -558,8 +558,6 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n-\tsize_t inflated_total = 0;\n-\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n@@ -575,33 +573,30 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n \t\t\tcontinue;\n \n-\t\tinflated_total += inflated;\n-\t\tdisk_total += disk;\n-\t}\n-\n-\tswitch (type) {\n-\tcase OBJ_TAG:\n-\t\tstats->type_counts.tags += oids->nr;\n-\t\tstats->inflated_sizes.tags += inflated_total;\n-\t\tstats->disk_sizes.tags += disk_total;\n-\t\tbreak;\n-\tcase OBJ_COMMIT:\n-\t\tstats->type_counts.commits += oids->nr;\n-\t\tstats->inflated_sizes.commits += inflated_total;\n-\t\tstats->disk_sizes.commits += disk_total;\n-\t\tbreak;\n-\tcase OBJ_TREE:\n-\t\tstats->type_counts.trees += oids->nr;\n-\t\tstats->inflated_sizes.trees += inflated_total;\n-\t\tstats->disk_sizes.trees += disk_total;\n-\t\tbreak;\n-\tcase OBJ_BLOB:\n-\t\tstats->type_counts.blobs += oids->nr;\n-\t\tstats->inflated_sizes.blobs += inflated_total;\n-\t\tstats->disk_sizes.blobs += disk_total;\n-\t\tbreak;\n-\tdefault:\n-\t\tBUG(\"invalid object type\");\n+\t\tswitch (type) {\n+\t\tcase OBJ_TAG:\n+\t\t\tstats->type_counts.tags++;\n+\t\t\tstats->inflated_sizes.tags += inflated;\n+\t\t\tstats->disk_sizes.tags += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_COMMIT:\n+\t\t\tstats->type_counts.commits++;\n+\t\t\tstats->inflated_sizes.commits += inflated;\n+\t\t\tstats->disk_sizes.commits += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_TREE:\n+\t\t\tstats->type_counts.trees++;\n+\t\t\tstats->inflated_sizes.trees += inflated;\n+\t\t\tstats->disk_sizes.trees += disk;\n+\t\t\tbreak;\n+\t\tcase OBJ_BLOB:\n+\t\t\tstats->type_counts.blobs++;\n+\t\t\tstats->inflated_sizes.blobs += inflated;\n+\t\t\tstats->disk_sizes.blobs += disk;\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tBUG(\"invalid object type\");\n+\t\t}\n \t}\n \n \tobject_count = get_total_object_values(&stats->type_counts);\n-- \n2.53.0\n\n"},{"id":"537611","messageId":"20260302214526.2034279-3-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260302214526.2034279-1-jltobler@gmail.com","subject":"[PATCH v3 2/6] builtin/repo: add helper for printing keyvalue output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T21:45:22Z","receivedAt":"2026-03-02T21:45:34Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The machine-parsable formats for the git-repo(1) \"structure\" subcommand\nprint output in keyvalue pairs. Introduce the helper function\n`print_keyvalue()` to remove some code duplication and improve\nreadability.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 77 +++++++++++++++++++++++++++-----------------------\n 1 file changed, 42 insertions(+), 35 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex c7c9f0f497..782194cf4c 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -446,44 +446,51 @@ static void stats_table_clear(struct stats_table *table)\n \tstring_list_clear(&table->rows, 1);\n }\n \n+static inline void print_keyvalue(const char *key, char key_delim, size_t value,\n+\t\t\t\t  char value_delim)\n+{\n+\tprintf(\"%s%c%\" PRIuMAX \"%c\", key, key_delim, (uintmax_t)value,\n+\t       value_delim);\n+}\n+\n static void structure_keyvalue_print(struct repo_structure *stats,\n \t\t\t\t     char key_delim, char value_delim)\n {\n-\tprintf(\"references.branches.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->refs.branches, value_delim);\n-\tprintf(\"references.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->refs.tags, value_delim);\n-\tprintf(\"references.remotes.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->refs.remotes, value_delim);\n-\tprintf(\"references.others.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->refs.others, value_delim);\n-\n-\tprintf(\"objects.commits.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.type_counts.commits, value_delim);\n-\tprintf(\"objects.trees.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.type_counts.trees, value_delim);\n-\tprintf(\"objects.blobs.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.type_counts.blobs, value_delim);\n-\tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n-\n-\tprintf(\"objects.commits.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n-\tprintf(\"objects.trees.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n-\tprintf(\"objects.blobs.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n-\tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n-\n-\tprintf(\"objects.commits.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n-\tprintf(\"objects.trees.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n-\tprintf(\"objects.blobs.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n-\tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n+\tprint_keyvalue(\"references.branches.count\", key_delim,\n+\t\t       stats->refs.branches, value_delim);\n+\tprint_keyvalue(\"references.tags.count\", key_delim,\n+\t\t       stats->refs.tags, value_delim);\n+\tprint_keyvalue(\"references.remotes.count\", key_delim,\n+\t\t       stats->refs.remotes, value_delim);\n+\tprint_keyvalue(\"references.others.count\", key_delim,\n+\t\t       stats->refs.others, value_delim);\n+\n+\tprint_keyvalue(\"objects.commits.count\", key_delim,\n+\t\t       stats->objects.type_counts.commits, value_delim);\n+\tprint_keyvalue(\"objects.trees.count\", key_delim,\n+\t\t       stats->objects.type_counts.trees, value_delim);\n+\tprint_keyvalue(\"objects.blobs.count\", key_delim,\n+\t\t       stats->objects.type_counts.blobs, value_delim);\n+\tprint_keyvalue(\"objects.tags.count\", key_delim,\n+\t\t       stats->objects.type_counts.tags, value_delim);\n+\n+\tprint_keyvalue(\"objects.commits.inflated_size\", key_delim,\n+\t\t       stats->objects.inflated_sizes.commits, value_delim);\n+\tprint_keyvalue(\"objects.trees.inflated_size\", key_delim,\n+\t\t       stats->objects.inflated_sizes.trees, value_delim);\n+\tprint_keyvalue(\"objects.blobs.inflated_size\", key_delim,\n+\t\t       stats->objects.inflated_sizes.blobs, value_delim);\n+\tprint_keyvalue(\"objects.tags.inflated_size\", key_delim,\n+\t\t       stats->objects.inflated_sizes.tags, value_delim);\n+\n+\tprint_keyvalue(\"objects.commits.disk_size\", key_delim,\n+\t\t       stats->objects.disk_sizes.commits, value_delim);\n+\tprint_keyvalue(\"objects.trees.disk_size\", key_delim,\n+\t\t       stats->objects.disk_sizes.trees, value_delim);\n+\tprint_keyvalue(\"objects.blobs.disk_size\", key_delim,\n+\t\t       stats->objects.disk_sizes.blobs, value_delim);\n+\tprint_keyvalue(\"objects.tags.disk_size\", key_delim,\n+\t\t       stats->objects.disk_sizes.tags, value_delim);\n \n \tfflush(stdout);\n }\n-- \n2.53.0\n\n"},{"id":"537612","messageId":"20260302214526.2034279-4-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260302214526.2034279-1-jltobler@gmail.com","subject":"[PATCH v3 3/6] builtin/repo: collect largest inflated objects","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T21:45:23Z","receivedAt":"2026-03-02T21:45:34Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The \"structure\" output for git-repo(1) shows the total inflated and disk\nsizes of reachable objects in the repository, but doesn't show the size\nof the largest individual objects. Since an individual object may be a\nlarge contributor to the overall repository size, it is useful for users\nto know the maximum size of individual objects.\n\nWhile interating across objects, record the size and OID of the largest\nobjects encountered for each object type to provide as output. Note that\nthe default \"table\" output format only displays size information and not\nthe corresponding OID. In a subsequent commit, the table format is\nupdated to add table annotations that mention the OID.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 63 +++++++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   | 28 +++++++++++++++++\n 3 files changed, 92 insertions(+)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 7d70270dfa..e812e59158 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -52,6 +52,7 @@ supported:\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n * Total disk size of reachable objects by type\n+* Largest reachable objects in the repository by type\n +\n The output format can be chosen through the flag `--format`. Three formats are\n supported:\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 782194cf4c..59d5cb2551 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,7 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hash.h\"\n #include \"hex.h\"\n #include \"odb.h\"\n #include \"parse-options.h\"\n@@ -197,6 +198,18 @@ static int cmd_repo_info(int argc, const char **argv, const char *prefix,\n \t\treturn print_fields(argc, argv, repo, format);\n }\n \n+struct object_data {\n+\tstruct object_id oid;\n+\tsize_t value;\n+};\n+\n+struct largest_objects {\n+\tstruct object_data tag_size;\n+\tstruct object_data commit_size;\n+\tstruct object_data tree_size;\n+\tstruct object_data blob_size;\n+};\n+\n struct ref_stats {\n \tsize_t branches;\n \tsize_t remotes;\n@@ -215,6 +228,7 @@ struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n \tstruct object_values disk_sizes;\n+\tstruct largest_objects largest;\n };\n \n struct repo_structure {\n@@ -371,6 +385,21 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->disk_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tstats_table_addf(table, \"\");\n+\tstats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->largest.commit_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->largest.tree_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->largest.blob_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_addf(table, \"  * %s\", _(\"Tags\"));\n+\tstats_table_size_addf(table, objects->largest.tag_size.value,\n+\t\t\t      \"    * %s\", _(\"Maximum size\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -453,6 +482,14 @@ static inline void print_keyvalue(const char *key, char key_delim, size_t value,\n \t       value_delim);\n }\n \n+static void print_object_data(const char *key, char key_delim,\n+\t\t\t      struct object_data *data, char value_delim)\n+{\n+\tprint_keyvalue(key, key_delim, data->value, value_delim);\n+\tprintf(\"%s_oid%c%s%c\", key, key_delim, oid_to_hex(&data->oid),\n+\t       value_delim);\n+}\n+\n static void structure_keyvalue_print(struct repo_structure *stats,\n \t\t\t\t     char key_delim, char value_delim)\n {\n@@ -492,6 +529,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprint_keyvalue(\"objects.tags.disk_size\", key_delim,\n \t\t       stats->objects.disk_sizes.tags, value_delim);\n \n+\tprint_object_data(\"objects.commits.max_size\", key_delim,\n+\t\t\t  &stats->objects.largest.commit_size, value_delim);\n+\tprint_object_data(\"objects.trees.max_size\", key_delim,\n+\t\t\t  &stats->objects.largest.tree_size, value_delim);\n+\tprint_object_data(\"objects.blobs.max_size\", key_delim,\n+\t\t\t  &stats->objects.largest.blob_size, value_delim);\n+\tprint_object_data(\"objects.tags.max_size\", key_delim,\n+\t\t\t  &stats->objects.largest.tag_size, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -560,6 +606,15 @@ struct count_objects_data {\n \tstruct progress *progress;\n };\n \n+static void check_largest(struct object_data *data, struct object_id *oid,\n+\t\t\t  size_t value)\n+{\n+\tif (value > data->value || is_null_oid(&data->oid)) {\n+\t\toidcpy(&data->oid, oid);\n+\t\tdata->value = value;\n+\t}\n+}\n+\n static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t enum object_type type, void *cb_data)\n {\n@@ -585,21 +640,29 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tstats->type_counts.tags++;\n \t\t\tstats->inflated_sizes.tags += inflated;\n \t\t\tstats->disk_sizes.tags += disk;\n+\t\t\tcheck_largest(&stats->largest.tag_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_COMMIT:\n \t\t\tstats->type_counts.commits++;\n \t\t\tstats->inflated_sizes.commits += inflated;\n \t\t\tstats->disk_sizes.commits += disk;\n+\t\t\tcheck_largest(&stats->largest.commit_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tstats->type_counts.trees++;\n \t\t\tstats->inflated_sizes.trees += inflated;\n \t\t\tstats->disk_sizes.trees += disk;\n+\t\t\tcheck_largest(&stats->largest.tree_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_BLOB:\n \t\t\tstats->type_counts.blobs++;\n \t\t\tstats->inflated_sizes.blobs += inflated;\n \t\t\tstats->disk_sizes.blobs += disk;\n+\t\t\tcheck_largest(&stats->largest.blob_size, &oids->oid[i],\n+\t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 17ff164b05..1999f325d0 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -52,6 +52,16 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|                      |        |\n+\t\t| * Largest objects    |        |\n+\t\t|   * Commits          |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Trees            |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Blobs            |        |\n+\t\t|     * Maximum size   |    0 B |\n+\t\t|   * Tags             |        |\n+\t\t|     * Maximum size   |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -104,6 +114,16 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Trees          | $(object_type_disk_usage tree true) |\n \t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n \t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n+\t\t|                      |            |\n+\t\t| * Largest objects    |            |\n+\t\t|   * Commits          |            |\n+\t\t|     * Maximum size   |    223 B   |\n+\t\t|   * Trees            |            |\n+\t\t|     * Maximum size   |  32.29 KiB |\n+\t\t|   * Blobs            |            |\n+\t\t|     * Maximum size   |     13 B   |\n+\t\t|   * Tags             |            |\n+\t\t|     * Maximum size   |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -138,6 +158,14 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n \t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n \t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n+\t\tobjects.commits.max_size=221\n+\t\tobjects.commits.max_size_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n+\t\tobjects.trees.max_size=1335\n+\t\tobjects.trees.max_size_oid=09931deea9d81ec21300d3e13c74412f32eacec5\n+\t\tobjects.blobs.max_size=11\n+\t\tobjects.blobs.max_size_oid=eaeeedced46482bd4281fda5a5f05ce24854151f\n+\t\tobjects.tags.max_size=132\n+\t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"537613","messageId":"20260302214526.2034279-5-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260302214526.2034279-1-jltobler@gmail.com","subject":"[PATCH v3 4/6] builtin/repo: add OID annotations to table output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T21:45:24Z","receivedAt":"2026-03-02T21:45:36Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The \"structure\" output for git-repo(1) does not show the corresponding\nOIDs for the largest objects in its \"table\" output. Update the output to\ninclude a list of OID annotations with an index to the corresponding row\nin the table.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            |  78 +++++++++++++++++---\n t/t1901-repo-structure.sh | 145 ++++++++++++++++++++------------------\n 2 files changed, 143 insertions(+), 80 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 59d5cb2551..ea7f5acd3e 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -238,6 +238,7 @@ struct repo_structure {\n \n struct stats_table {\n \tstruct string_list rows;\n+\tstruct string_list annotations;\n \n \tint name_col_width;\n \tint value_col_width;\n@@ -250,6 +251,8 @@ struct stats_table {\n struct stats_table_entry {\n \tchar *value;\n \tconst char *unit;\n+\tsize_t index;\n+\tstruct object_id *oid;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -272,6 +275,12 @@ static void stats_table_vaddf(struct stats_table *table,\n \t\ttable->name_col_width = name_width;\n \tif (!entry)\n \t\treturn;\n+\tif (entry->oid) {\n+\t\tentry->index = table->annotations.nr + 1;\n+\t\tstrbuf_addf(&buf, \"[%\" PRIuMAX \"] %s\", (uintmax_t)entry->index,\n+\t\t\t    oid_to_hex(entry->oid));\n+\t\tstring_list_append_nodup(&table->annotations, strbuf_detach(&buf, NULL));\n+\t}\n \tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n@@ -282,6 +291,8 @@ static void stats_table_vaddf(struct stats_table *table,\n \t\tif (unit_width > table->unit_col_width)\n \t\t\ttable->unit_col_width = unit_width;\n \t}\n+\n+\tstrbuf_release(&buf);\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -321,6 +332,27 @@ static void stats_table_size_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_object_size_addf(struct stats_table *table,\n+\t\t\t\t\t struct object_id *oid, size_t value,\n+\t\t\t\t\t const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n+\n+\t/*\n+\t * A NULL OID should not have a table annotation.\n+\t */\n+\tif (!is_null_oid(oid))\n+\t\tentry->oid = oid;\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -389,19 +421,29 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Largest objects\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Commits\"));\n-\tstats_table_size_addf(table, objects->largest.commit_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.commit_size.oid,\n+\t\t\t\t     objects->largest.commit_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n-\tstats_table_size_addf(table, objects->largest.tree_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.tree_size.oid,\n+\t\t\t\t     objects->largest.tree_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n-\tstats_table_size_addf(table, objects->largest.blob_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.blob_size.oid,\n+\t\t\t\t     objects->largest.blob_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Tags\"));\n-\tstats_table_size_addf(table, objects->largest.tag_size.value,\n-\t\t\t      \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_size_addf(table,\n+\t\t\t\t     &objects->largest.tag_size.oid,\n+\t\t\t\t     objects->largest.tag_size.value,\n+\t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n }\n \n+#define INDEX_WIDTH 4\n+\n static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n@@ -420,7 +462,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width + INDEX_WIDTH,\n+\t\t\t  name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n \t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n@@ -428,7 +471,7 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tprintf(\"%s\\n\", buf.buf);\n \n \tprintf(\"| \");\n-\tfor (int i = 0; i < name_col_width; i++)\n+\tfor (int i = 0; i < name_col_width + INDEX_WIDTH; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n \tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n@@ -450,6 +493,13 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_reset(&buf);\n \t\tstrbuf_addstr(&buf, \"| \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n+\n+\t\tif (entry && entry->oid)\n+\t\t\tstrbuf_addf(&buf, \" [%\" PRIuMAX \"]\",\n+\t\t\t\t    (uintmax_t)entry->index);\n+\t\telse\n+\t\t\tstrbuf_addchars(&buf, ' ', INDEX_WIDTH);\n+\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n \t\tstrbuf_addch(&buf, ' ');\n@@ -458,6 +508,12 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\n \n+\tif (table->annotations.nr) {\n+\t\tprintf(\"\\n\");\n+\t\tfor_each_string_list_item(item, &table->annotations)\n+\t\t\tprintf(\"%s\\n\", item->string);\n+\t}\n+\n \tstrbuf_release(&buf);\n }\n \n@@ -473,6 +529,7 @@ static void stats_table_clear(struct stats_table *table)\n \t}\n \n \tstring_list_clear(&table->rows, 1);\n+\tstring_list_clear(&table->annotations, 1);\n }\n \n static inline void print_keyvalue(const char *key, char key_delim, size_t value,\n@@ -702,6 +759,7 @@ static int cmd_repo_structure(int argc, const char **argv, const char *prefix,\n {\n \tstruct stats_table table = {\n \t\t.rows = STRING_LIST_INIT_DUP,\n+\t\t.annotations = STRING_LIST_INIT_DUP,\n \t};\n \tenum output_format format = FORMAT_TABLE;\n \tstruct repo_structure stats = { 0 };\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 1999f325d0..918af7269f 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -27,41 +27,41 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    0   |\n-\t\t|     * Branches       |    0   |\n-\t\t|     * Tags           |    0   |\n-\t\t|     * Remotes        |    0   |\n-\t\t|     * Others         |    0   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            |    0   |\n-\t\t|     * Commits        |    0   |\n-\t\t|     * Trees          |    0   |\n-\t\t|     * Blobs          |    0   |\n-\t\t|     * Tags           |    0   |\n-\t\t|   * Inflated size    |    0 B |\n-\t\t|     * Commits        |    0 B |\n-\t\t|     * Trees          |    0 B |\n-\t\t|     * Blobs          |    0 B |\n-\t\t|     * Tags           |    0 B |\n-\t\t|   * Disk size        |    0 B |\n-\t\t|     * Commits        |    0 B |\n-\t\t|     * Trees          |    0 B |\n-\t\t|     * Blobs          |    0 B |\n-\t\t|     * Tags           |    0 B |\n-\t\t|                      |        |\n-\t\t| * Largest objects    |        |\n-\t\t|   * Commits          |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Trees            |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Blobs            |        |\n-\t\t|     * Maximum size   |    0 B |\n-\t\t|   * Tags             |        |\n-\t\t|     * Maximum size   |    0 B |\n+\t\t| Repository structure     | Value  |\n+\t\t| ------------------------ | ------ |\n+\t\t| * References             |        |\n+\t\t|   * Count                |    0   |\n+\t\t|     * Branches           |    0   |\n+\t\t|     * Tags               |    0   |\n+\t\t|     * Remotes            |    0   |\n+\t\t|     * Others             |    0   |\n+\t\t|                          |        |\n+\t\t| * Reachable objects      |        |\n+\t\t|   * Count                |    0   |\n+\t\t|     * Commits            |    0   |\n+\t\t|     * Trees              |    0   |\n+\t\t|     * Blobs              |    0   |\n+\t\t|     * Tags               |    0   |\n+\t\t|   * Inflated size        |    0 B |\n+\t\t|     * Commits            |    0 B |\n+\t\t|     * Trees              |    0 B |\n+\t\t|     * Blobs              |    0 B |\n+\t\t|     * Tags               |    0 B |\n+\t\t|   * Disk size            |    0 B |\n+\t\t|     * Commits            |    0 B |\n+\t\t|     * Trees              |    0 B |\n+\t\t|     * Blobs              |    0 B |\n+\t\t|     * Tags               |    0 B |\n+\t\t|                          |        |\n+\t\t| * Largest objects        |        |\n+\t\t|   * Commits              |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Trees                |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Blobs                |        |\n+\t\t|     * Maximum size       |    0 B |\n+\t\t|   * Tags                 |        |\n+\t\t|     * Maximum size       |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -89,41 +89,46 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# git-rev-list(1) --disk-usage=human option printing the full\n \t\t# \"byte/bytes\" unit string instead of just \"B\".\n \t\tcat >expect <<-EOF &&\n-\t\t| Repository structure | Value      |\n-\t\t| -------------------- | ---------- |\n-\t\t| * References         |            |\n-\t\t|   * Count            |      4     |\n-\t\t|     * Branches       |      1     |\n-\t\t|     * Tags           |      1     |\n-\t\t|     * Remotes        |      1     |\n-\t\t|     * Others         |      1     |\n-\t\t|                      |            |\n-\t\t| * Reachable objects  |            |\n-\t\t|   * Count            |   3.02 k   |\n-\t\t|     * Commits        |   1.01 k   |\n-\t\t|     * Trees          |   1.01 k   |\n-\t\t|     * Blobs          |   1.01 k   |\n-\t\t|     * Tags           |      1     |\n-\t\t|   * Inflated size    |  16.03 MiB |\n-\t\t|     * Commits        | 217.92 KiB |\n-\t\t|     * Trees          |  15.81 MiB |\n-\t\t|     * Blobs          |  11.68 KiB |\n-\t\t|     * Tags           |    132 B   |\n-\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n-\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n-\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n-\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n-\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n-\t\t|                      |            |\n-\t\t| * Largest objects    |            |\n-\t\t|   * Commits          |            |\n-\t\t|     * Maximum size   |    223 B   |\n-\t\t|   * Trees            |            |\n-\t\t|     * Maximum size   |  32.29 KiB |\n-\t\t|   * Blobs            |            |\n-\t\t|     * Maximum size   |     13 B   |\n-\t\t|   * Tags             |            |\n-\t\t|     * Maximum size   |    132 B   |\n+\t\t| Repository structure     | Value      |\n+\t\t| ------------------------ | ---------- |\n+\t\t| * References             |            |\n+\t\t|   * Count                |      4     |\n+\t\t|     * Branches           |      1     |\n+\t\t|     * Tags               |      1     |\n+\t\t|     * Remotes            |      1     |\n+\t\t|     * Others             |      1     |\n+\t\t|                          |            |\n+\t\t| * Reachable objects      |            |\n+\t\t|   * Count                |   3.02 k   |\n+\t\t|     * Commits            |   1.01 k   |\n+\t\t|     * Trees              |   1.01 k   |\n+\t\t|     * Blobs              |   1.01 k   |\n+\t\t|     * Tags               |      1     |\n+\t\t|   * Inflated size        |  16.03 MiB |\n+\t\t|     * Commits            | 217.92 KiB |\n+\t\t|     * Trees              |  15.81 MiB |\n+\t\t|     * Blobs              |  11.68 KiB |\n+\t\t|     * Tags               |    132 B   |\n+\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n+\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n+\t\t|                          |            |\n+\t\t| * Largest objects        |            |\n+\t\t|   * Commits              |            |\n+\t\t|     * Maximum size   [1] |    223 B   |\n+\t\t|   * Trees                |            |\n+\t\t|     * Maximum size   [2] |  32.29 KiB |\n+\t\t|   * Blobs                |            |\n+\t\t|     * Maximum size   [3] |     13 B   |\n+\t\t|   * Tags                 |            |\n+\t\t|     * Maximum size   [4] |    132 B   |\n+\n+\t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n+\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"537614","messageId":"20260302214526.2034279-6-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260302214526.2034279-1-jltobler@gmail.com","subject":"[PATCH v3 5/6] builtin/repo: find commit with most parents","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T21:45:25Z","receivedAt":"2026-03-02T21:45:37Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Complex merge events may produce an octopus merge where the resulting\nmerge commit has more than two parents. While iterating through objects\nin the repository for git-repo-structure, identify the commit with the\nmost parents and display it in the output.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            |  45 ++++++++++++\n t/t1901-repo-structure.sh | 151 ++++++++++++++++++++------------------\n 2 files changed, 123 insertions(+), 73 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex ea7f5acd3e..047f5e098d 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -1,6 +1,7 @@\n #define USE_THE_REPOSITORY_VARIABLE\n \n #include \"builtin.h\"\n+#include \"commit.h\"\n #include \"environment.h\"\n #include \"hash.h\"\n #include \"hex.h\"\n@@ -208,6 +209,8 @@ struct largest_objects {\n \tstruct object_data commit_size;\n \tstruct object_data tree_size;\n \tstruct object_data blob_size;\n+\n+\tstruct object_data parent_count;\n };\n \n struct ref_stats {\n@@ -318,6 +321,27 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_object_count_addf(struct stats_table *table,\n+\t\t\t\t\t  struct object_id *oid, size_t value,\n+\t\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_count(value, &entry->value, &entry->unit);\n+\n+\t/*\n+\t * A NULL OID should not have a table annotation.\n+\t */\n+\tif (!is_null_oid(oid))\n+\t\tentry->oid = oid;\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static void stats_table_size_addf(struct stats_table *table, size_t value,\n \t\t\t\t  const char *format, ...)\n {\n@@ -425,6 +449,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t\t     &objects->largest.commit_size.oid,\n \t\t\t\t     objects->largest.commit_size.value,\n \t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_count_addf(table,\n+\t\t\t\t      &objects->largest.parent_count.oid,\n+\t\t\t\t      objects->largest.parent_count.value,\n+\t\t\t\t      \"    * %s\", _(\"Maximum parents\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Trees\"));\n \tstats_table_object_size_addf(table,\n \t\t\t\t     &objects->largest.tree_size.oid,\n@@ -595,6 +623,9 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprint_object_data(\"objects.tags.max_size\", key_delim,\n \t\t\t  &stats->objects.largest.tag_size, value_delim);\n \n+\tprint_object_data(\"objects.commits.max_parents\", key_delim,\n+\t\t\t  &stats->objects.largest.parent_count, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -682,16 +713,24 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\tstruct commit *commit;\n+\t\tstruct object *obj;\n+\t\tvoid *content;\n \t\toff_t disk;\n+\t\tint eaten;\n \n \t\toi.sizep = &inflated;\n \t\toi.disk_sizep = &disk;\n+\t\toi.contentp = &content;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n \t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n \t\t\tcontinue;\n \n+\t\tobj = parse_object_buffer(the_repository, &oids->oid[i], type,\n+\t\t\t\t\t  inflated, content, &eaten);\n+\n \t\tswitch (type) {\n \t\tcase OBJ_TAG:\n \t\t\tstats->type_counts.tags++;\n@@ -701,11 +740,14 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t\t      inflated);\n \t\t\tbreak;\n \t\tcase OBJ_COMMIT:\n+\t\t\tcommit = object_as_type(obj, OBJ_COMMIT, 0);\n \t\t\tstats->type_counts.commits++;\n \t\t\tstats->inflated_sizes.commits += inflated;\n \t\t\tstats->disk_sizes.commits += disk;\n \t\t\tcheck_largest(&stats->largest.commit_size, &oids->oid[i],\n \t\t\t\t      inflated);\n+\t\t\tcheck_largest(&stats->largest.parent_count, &oids->oid[i],\n+\t\t\t\t      commit_list_count(commit->parents));\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tstats->type_counts.trees++;\n@@ -724,6 +766,9 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\tdefault:\n \t\t\tBUG(\"invalid object type\");\n \t\t}\n+\n+\t\tif (!eaten)\n+\t\t\tfree(content);\n \t}\n \n \tobject_count = get_total_object_values(&stats->type_counts);\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 918af7269f..d003d64a8e 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -27,41 +27,42 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure     | Value  |\n-\t\t| ------------------------ | ------ |\n-\t\t| * References             |        |\n-\t\t|   * Count                |    0   |\n-\t\t|     * Branches           |    0   |\n-\t\t|     * Tags               |    0   |\n-\t\t|     * Remotes            |    0   |\n-\t\t|     * Others             |    0   |\n-\t\t|                          |        |\n-\t\t| * Reachable objects      |        |\n-\t\t|   * Count                |    0   |\n-\t\t|     * Commits            |    0   |\n-\t\t|     * Trees              |    0   |\n-\t\t|     * Blobs              |    0   |\n-\t\t|     * Tags               |    0   |\n-\t\t|   * Inflated size        |    0 B |\n-\t\t|     * Commits            |    0 B |\n-\t\t|     * Trees              |    0 B |\n-\t\t|     * Blobs              |    0 B |\n-\t\t|     * Tags               |    0 B |\n-\t\t|   * Disk size            |    0 B |\n-\t\t|     * Commits            |    0 B |\n-\t\t|     * Trees              |    0 B |\n-\t\t|     * Blobs              |    0 B |\n-\t\t|     * Tags               |    0 B |\n-\t\t|                          |        |\n-\t\t| * Largest objects        |        |\n-\t\t|   * Commits              |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Trees                |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Blobs                |        |\n-\t\t|     * Maximum size       |    0 B |\n-\t\t|   * Tags                 |        |\n-\t\t|     * Maximum size       |    0 B |\n+\t\t| Repository structure      | Value  |\n+\t\t| ------------------------- | ------ |\n+\t\t| * References              |        |\n+\t\t|   * Count                 |    0   |\n+\t\t|     * Branches            |    0   |\n+\t\t|     * Tags                |    0   |\n+\t\t|     * Remotes             |    0   |\n+\t\t|     * Others              |    0   |\n+\t\t|                           |        |\n+\t\t| * Reachable objects       |        |\n+\t\t|   * Count                 |    0   |\n+\t\t|     * Commits             |    0   |\n+\t\t|     * Trees               |    0   |\n+\t\t|     * Blobs               |    0   |\n+\t\t|     * Tags                |    0   |\n+\t\t|   * Inflated size         |    0 B |\n+\t\t|     * Commits             |    0 B |\n+\t\t|     * Trees               |    0 B |\n+\t\t|     * Blobs               |    0 B |\n+\t\t|     * Tags                |    0 B |\n+\t\t|   * Disk size             |    0 B |\n+\t\t|     * Commits             |    0 B |\n+\t\t|     * Trees               |    0 B |\n+\t\t|     * Blobs               |    0 B |\n+\t\t|     * Tags                |    0 B |\n+\t\t|                           |        |\n+\t\t| * Largest objects         |        |\n+\t\t|   * Commits               |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|     * Maximum parents     |    0   |\n+\t\t|   * Trees                 |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|   * Blobs                 |        |\n+\t\t|     * Maximum size        |    0 B |\n+\t\t|   * Tags                  |        |\n+\t\t|     * Maximum size        |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -89,46 +90,48 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# git-rev-list(1) --disk-usage=human option printing the full\n \t\t# \"byte/bytes\" unit string instead of just \"B\".\n \t\tcat >expect <<-EOF &&\n-\t\t| Repository structure     | Value      |\n-\t\t| ------------------------ | ---------- |\n-\t\t| * References             |            |\n-\t\t|   * Count                |      4     |\n-\t\t|     * Branches           |      1     |\n-\t\t|     * Tags               |      1     |\n-\t\t|     * Remotes            |      1     |\n-\t\t|     * Others             |      1     |\n-\t\t|                          |            |\n-\t\t| * Reachable objects      |            |\n-\t\t|   * Count                |   3.02 k   |\n-\t\t|     * Commits            |   1.01 k   |\n-\t\t|     * Trees              |   1.01 k   |\n-\t\t|     * Blobs              |   1.01 k   |\n-\t\t|     * Tags               |      1     |\n-\t\t|   * Inflated size        |  16.03 MiB |\n-\t\t|     * Commits            | 217.92 KiB |\n-\t\t|     * Trees              |  15.81 MiB |\n-\t\t|     * Blobs              |  11.68 KiB |\n-\t\t|     * Tags               |    132 B   |\n-\t\t|   * Disk size            | $(object_type_disk_usage all true) |\n-\t\t|     * Commits            | $(object_type_disk_usage commit true) |\n-\t\t|     * Trees              | $(object_type_disk_usage tree true) |\n-\t\t|     * Blobs              |  $(object_type_disk_usage blob true) |\n-\t\t|     * Tags               |    $(object_type_disk_usage tag) B   |\n-\t\t|                          |            |\n-\t\t| * Largest objects        |            |\n-\t\t|   * Commits              |            |\n-\t\t|     * Maximum size   [1] |    223 B   |\n-\t\t|   * Trees                |            |\n-\t\t|     * Maximum size   [2] |  32.29 KiB |\n-\t\t|   * Blobs                |            |\n-\t\t|     * Maximum size   [3] |     13 B   |\n-\t\t|   * Tags                 |            |\n-\t\t|     * Maximum size   [4] |    132 B   |\n+\t\t| Repository structure      | Value      |\n+\t\t| ------------------------- | ---------- |\n+\t\t| * References              |            |\n+\t\t|   * Count                 |      4     |\n+\t\t|     * Branches            |      1     |\n+\t\t|     * Tags                |      1     |\n+\t\t|     * Remotes             |      1     |\n+\t\t|     * Others              |      1     |\n+\t\t|                           |            |\n+\t\t| * Reachable objects       |            |\n+\t\t|   * Count                 |   3.02 k   |\n+\t\t|     * Commits             |   1.01 k   |\n+\t\t|     * Trees               |   1.01 k   |\n+\t\t|     * Blobs               |   1.01 k   |\n+\t\t|     * Tags                |      1     |\n+\t\t|   * Inflated size         |  16.03 MiB |\n+\t\t|     * Commits             | 217.92 KiB |\n+\t\t|     * Trees               |  15.81 MiB |\n+\t\t|     * Blobs               |  11.68 KiB |\n+\t\t|     * Tags                |    132 B   |\n+\t\t|   * Disk size             | $(object_type_disk_usage all true) |\n+\t\t|     * Commits             | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees               | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs               |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags                |    $(object_type_disk_usage tag) B   |\n+\t\t|                           |            |\n+\t\t| * Largest objects         |            |\n+\t\t|   * Commits               |            |\n+\t\t|     * Maximum size    [1] |    223 B   |\n+\t\t|     * Maximum parents [2] |      1     |\n+\t\t|   * Trees                 |            |\n+\t\t|     * Maximum size    [3] |  32.29 KiB |\n+\t\t|   * Blobs                 |            |\n+\t\t|     * Maximum size    [4] |     13 B   |\n+\t\t|   * Tags                  |            |\n+\t\t|     * Maximum size    [5] |    132 B   |\n \n \t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n-\t\t[2] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n-\t\t[3] 97d808e45116bf02103490294d3d46dad7a2ac62\n-\t\t[4] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n+\t\t[2] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n+\t\t[3] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[4] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[5] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -171,6 +174,8 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.blobs.max_size_oid=eaeeedced46482bd4281fda5a5f05ce24854151f\n \t\tobjects.tags.max_size=132\n \t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n+\t\tobjects.commits.max_parents=1\n+\t\tobjects.commits.max_parents_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"537615","messageId":"20260302214526.2034279-7-jltobler@gmail.com","threadId":"64911","inReplyTo":"20260302214526.2034279-1-jltobler@gmail.com","subject":"[PATCH v3 6/6] builtin/repo: find tree with most entries","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-02T21:45:26Z","receivedAt":"2026-03-02T21:45:37Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The size of a tree object usually corresponds with the number of entries\nit has. While iterating through objects in the repository for\ngit-repo-structure, identify the tree with the most entries and display\nit in the output.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 25 +++++++++++++++++++++++++\n t/t1901-repo-structure.sh | 13 +++++++++----\n 2 files changed, 34 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 047f5e098d..e726bb858c 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -16,6 +16,8 @@\n #include \"strbuf.h\"\n #include \"string-list.h\"\n #include \"shallow.h\"\n+#include \"tree.h\"\n+#include \"tree-walk.h\"\n #include \"utf8.h\"\n \n static const char *const repo_usage[] = {\n@@ -211,6 +213,7 @@ struct largest_objects {\n \tstruct object_data blob_size;\n \n \tstruct object_data parent_count;\n+\tstruct object_data tree_entries;\n };\n \n struct ref_stats {\n@@ -458,6 +461,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t\t     &objects->largest.tree_size.oid,\n \t\t\t\t     objects->largest.tree_size.value,\n \t\t\t\t     \"    * %s\", _(\"Maximum size\"));\n+\tstats_table_object_count_addf(table,\n+\t\t\t\t      &objects->largest.tree_entries.oid,\n+\t\t\t\t      objects->largest.tree_entries.value,\n+\t\t\t\t      \"    * %s\", _(\"Maximum entries\"));\n \tstats_table_addf(table, \"  * %s\", _(\"Blobs\"));\n \tstats_table_object_size_addf(table,\n \t\t\t\t     &objects->largest.blob_size.oid,\n@@ -625,6 +632,8 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \n \tprint_object_data(\"objects.commits.max_parents\", key_delim,\n \t\t\t  &stats->objects.largest.parent_count, value_delim);\n+\tprint_object_data(\"objects.trees.max_entries\", key_delim,\n+\t\t\t  &stats->objects.largest.tree_entries, value_delim);\n \n \tfflush(stdout);\n }\n@@ -703,6 +712,20 @@ static void check_largest(struct object_data *data, struct object_id *oid,\n \t}\n }\n \n+static size_t count_tree_entries(struct object *obj)\n+{\n+\tstruct tree *t = object_as_type(obj, OBJ_TREE, 0);\n+\tstruct name_entry entry;\n+\tstruct tree_desc desc;\n+\tsize_t count = 0;\n+\n+\tinit_tree_desc(&desc, &t->object.oid, t->buffer, t->size);\n+\twhile (tree_entry(&desc, &entry))\n+\t\tcount++;\n+\n+\treturn count;\n+}\n+\n static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t enum object_type type, void *cb_data)\n {\n@@ -755,6 +778,8 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tstats->disk_sizes.trees += disk;\n \t\t\tcheck_largest(&stats->largest.tree_size, &oids->oid[i],\n \t\t\t\t      inflated);\n+\t\t\tcheck_largest(&stats->largest.tree_entries, &oids->oid[i],\n+\t\t\t\t      count_tree_entries(obj));\n \t\t\tbreak;\n \t\tcase OBJ_BLOB:\n \t\t\tstats->type_counts.blobs++;\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex d003d64a8e..12ed67e846 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -59,6 +59,7 @@ test_expect_success 'empty repository' '\n \t\t|     * Maximum parents     |    0   |\n \t\t|   * Trees                 |        |\n \t\t|     * Maximum size        |    0 B |\n+\t\t|     * Maximum entries     |    0   |\n \t\t|   * Blobs                 |        |\n \t\t|     * Maximum size        |    0 B |\n \t\t|   * Tags                  |        |\n@@ -122,16 +123,18 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Maximum parents [2] |      1     |\n \t\t|   * Trees                 |            |\n \t\t|     * Maximum size    [3] |  32.29 KiB |\n+\t\t|     * Maximum entries [4] |   1.01 k   |\n \t\t|   * Blobs                 |            |\n-\t\t|     * Maximum size    [4] |     13 B   |\n+\t\t|     * Maximum size    [5] |     13 B   |\n \t\t|   * Tags                  |            |\n-\t\t|     * Maximum size    [5] |    132 B   |\n+\t\t|     * Maximum size    [6] |    132 B   |\n \n \t\t[1] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n \t\t[2] 0dc91eb18580102a3a216c8bfecedeba2b9f9b9a\n \t\t[3] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n-\t\t[4] 97d808e45116bf02103490294d3d46dad7a2ac62\n-\t\t[5] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n+\t\t[4] 60665251ab71dbd8c18d9bf2174f4ee0d58aa06c\n+\t\t[5] 97d808e45116bf02103490294d3d46dad7a2ac62\n+\t\t[6] 4dae4f5954f5e6feb3577cfb1b181daa3fd3afd2\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -176,6 +179,8 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.tags.max_size_oid=1ee0f2b16ea37d895dbe9dbd76cd2ac70446176c\n \t\tobjects.commits.max_parents=1\n \t\tobjects.commits.max_parents_oid=de3508174b5c2ace6993da67cae9be9069e2df39\n+\t\tobjects.trees.max_entries=42\n+\t\tobjects.trees.max_entries_oid=09931deea9d81ec21300d3e13c74412f32eacec5\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.53.0\n\n"},{"id":"537618","messageId":"xmqqqzq1yjcl.fsf@gitster.g","threadId":"64911","inReplyTo":"20260302214526.2034279-1-jltobler@gmail.com","subject":"Re: [PATCH v3 0/6] builtin/repo: include largest object information","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-02T22:09:14Z","receivedAt":"2026-03-02T22:09:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> Changes from V2:\n> - When checking for largest objects, zero valued objects were not\n>   recorded even if they were the \"largest\" object. In this version, if\n>   an object ID has not been recorded yet, it is always added even if its\n>   value is zero.\n> - Added some helper functions for printing keyvalue info to cut down on\n>   duplicate code and hopefully make it a bit easier on the eyes.\n> - Moved the for-each loop that printed table OID annoations inside the\n>   preceding if-block making it a bit easier to reason about.\n\nThe changes I see in the diff relative to the previous iteration all\nlook sane to me.  Will replace.  Thanks.\n"},{"id":"537660","messageId":"aabhtfWZG90YyhQ5@pks.im","threadId":"64911","inReplyTo":"20260302214526.2034279-3-jltobler@gmail.com","subject":"Re: [PATCH v3 2/6] builtin/repo: add helper for printing keyvalue output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-03T13:27:17Z","receivedAt":"2026-03-03T13:27:23Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Mar 02, 2026 at 03:45:22PM -0600, Justin Tobler wrote:\n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index c7c9f0f497..782194cf4c 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -446,44 +446,51 @@ static void stats_table_clear(struct stats_table *table)\n>  \tstring_list_clear(&table->rows, 1);\n>  }\n>  \n> +static inline void print_keyvalue(const char *key, char key_delim, size_t value,\n> +\t\t\t\t  char value_delim)\n> +{\n> +\tprintf(\"%s%c%\" PRIuMAX \"%c\", key, key_delim, (uintmax_t)value,\n> +\t       value_delim);\n> +}\n> +\n>  static void structure_keyvalue_print(struct repo_structure *stats,\n>  \t\t\t\t     char key_delim, char value_delim)\n>  {\n> -\tprintf(\"references.branches.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->refs.branches, value_delim);\n> -\tprintf(\"references.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->refs.tags, value_delim);\n> -\tprintf(\"references.remotes.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->refs.remotes, value_delim);\n> -\tprintf(\"references.others.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->refs.others, value_delim);\n> -\n> -\tprintf(\"objects.commits.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.type_counts.commits, value_delim);\n> -\tprintf(\"objects.trees.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.type_counts.trees, value_delim);\n> -\tprintf(\"objects.blobs.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.type_counts.blobs, value_delim);\n> -\tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n> -\n> -\tprintf(\"objects.commits.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n> -\tprintf(\"objects.trees.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n> -\tprintf(\"objects.blobs.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n> -\tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n> -\n> -\tprintf(\"objects.commits.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n> -\tprintf(\"objects.trees.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n> -\tprintf(\"objects.blobs.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n> -\tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n> -\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n> +\tprint_keyvalue(\"references.branches.count\", key_delim,\n> +\t\t       stats->refs.branches, value_delim);\n> +\tprint_keyvalue(\"references.tags.count\", key_delim,\n> +\t\t       stats->refs.tags, value_delim);\n> +\tprint_keyvalue(\"references.remotes.count\", key_delim,\n> +\t\t       stats->refs.remotes, value_delim);\n> +\tprint_keyvalue(\"references.others.count\", key_delim,\n> +\t\t       stats->refs.others, value_delim);\n> +\n> +\tprint_keyvalue(\"objects.commits.count\", key_delim,\n> +\t\t       stats->objects.type_counts.commits, value_delim);\n> +\tprint_keyvalue(\"objects.trees.count\", key_delim,\n> +\t\t       stats->objects.type_counts.trees, value_delim);\n> +\tprint_keyvalue(\"objects.blobs.count\", key_delim,\n> +\t\t       stats->objects.type_counts.blobs, value_delim);\n> +\tprint_keyvalue(\"objects.tags.count\", key_delim,\n> +\t\t       stats->objects.type_counts.tags, value_delim);\n> +\n> +\tprint_keyvalue(\"objects.commits.inflated_size\", key_delim,\n> +\t\t       stats->objects.inflated_sizes.commits, value_delim);\n> +\tprint_keyvalue(\"objects.trees.inflated_size\", key_delim,\n> +\t\t       stats->objects.inflated_sizes.trees, value_delim);\n> +\tprint_keyvalue(\"objects.blobs.inflated_size\", key_delim,\n> +\t\t       stats->objects.inflated_sizes.blobs, value_delim);\n> +\tprint_keyvalue(\"objects.tags.inflated_size\", key_delim,\n> +\t\t       stats->objects.inflated_sizes.tags, value_delim);\n> +\n> +\tprint_keyvalue(\"objects.commits.disk_size\", key_delim,\n> +\t\t       stats->objects.disk_sizes.commits, value_delim);\n> +\tprint_keyvalue(\"objects.trees.disk_size\", key_delim,\n> +\t\t       stats->objects.disk_sizes.trees, value_delim);\n> +\tprint_keyvalue(\"objects.blobs.disk_size\", key_delim,\n> +\t\t       stats->objects.disk_sizes.blobs, value_delim);\n> +\tprint_keyvalue(\"objects.tags.disk_size\", key_delim,\n> +\t\t       stats->objects.disk_sizes.tags, value_delim);\n\nIt's still easy to miss any mismatch here, but I guess the result is\ndefinitely easier to read regardless of that.\n\nThanks!\n\nPatrick\n"},{"id":"537661","messageId":"aabhurGdiN0jFqkb@pks.im","threadId":"64911","inReplyTo":"20260302214526.2034279-4-jltobler@gmail.com","subject":"Re: [PATCH v3 3/6] builtin/repo: collect largest inflated objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-03T13:27:22Z","receivedAt":"2026-03-03T13:27:27Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Mar 02, 2026 at 03:45:23PM -0600, Justin Tobler wrote:\n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index 782194cf4c..59d5cb2551 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -453,6 +482,14 @@ static inline void print_keyvalue(const char *key, char key_delim, size_t value,\n>  \t       value_delim);\n>  }\n>  \n> +static void print_object_data(const char *key, char key_delim,\n> +\t\t\t      struct object_data *data, char value_delim)\n> +{\n> +\tprint_keyvalue(key, key_delim, data->value, value_delim);\n> +\tprintf(\"%s_oid%c%s%c\", key, key_delim, oid_to_hex(&data->oid),\n> +\t       value_delim);\n> +}\n> +\n>  static void structure_keyvalue_print(struct repo_structure *stats,\n>  \t\t\t\t     char key_delim, char value_delim)\n>  {\n\nAnd this helper is also quite a welcome improvement.\n\n> @@ -492,6 +529,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n>  \tprint_keyvalue(\"objects.tags.disk_size\", key_delim,\n>  \t\t       stats->objects.disk_sizes.tags, value_delim);\n>  \n> +\tprint_object_data(\"objects.commits.max_size\", key_delim,\n> +\t\t\t  &stats->objects.largest.commit_size, value_delim);\n> +\tprint_object_data(\"objects.trees.max_size\", key_delim,\n> +\t\t\t  &stats->objects.largest.tree_size, value_delim);\n> +\tprint_object_data(\"objects.blobs.max_size\", key_delim,\n> +\t\t\t  &stats->objects.largest.blob_size, value_delim);\n> +\tprint_object_data(\"objects.tags.max_size\", key_delim,\n> +\t\t\t  &stats->objects.largest.tag_size, value_delim);\n> +\n>  \tfflush(stdout);\n>  }\n\nCertainly makes this part easier to verify.\n\nPatrick\n"},{"id":"537704","messageId":"xmqq7brserq7.fsf@gitster.g","threadId":"64911","inReplyTo":"aabhtfWZG90YyhQ5@pks.im","subject":"Re: [PATCH v3 2/6] builtin/repo: add helper for printing keyvalue output","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-03T17:40:48Z","receivedAt":"2026-03-03T17:40:51Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> +\tprint_keyvalue(\"references.branches.count\", key_delim,\n>> +\t\t       stats->refs.branches, value_delim);\n>> ...\n>\n> It's still easy to miss any mismatch here, but I guess the result is\n> definitely easier to read regardless of that.\n\nSure, we could further do something silly like\n\n#define P(name, source) print_keyvalue(name, key_delim, source, value_delim)\n\nand reduce the above to\n\n\tP(\"references.branches.count\", stats->refs.branches);\n\nif we wanted to.\n"},{"id":"537710","messageId":"aachrznSGC_gcElv@denethor","threadId":"64911","inReplyTo":"xmqq7brserq7.fsf@gitster.g","subject":"Re: [PATCH v3 2/6] builtin/repo: add helper for printing keyvalue output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-03T18:08:19Z","receivedAt":"2026-03-03T18:08:20Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/03/03 09:40AM, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> >> +\tprint_keyvalue(\"references.branches.count\", key_delim,\n> >> +\t\t       stats->refs.branches, value_delim);\n> >> ...\n> >\n> > It's still easy to miss any mismatch here, but I guess the result is\n> > definitely easier to read regardless of that.\n> \n> Sure, we could further do something silly like\n> \n> #define P(name, source) print_keyvalue(name, key_delim, source, value_delim)\n> \n> and reduce the above to\n> \n> \tP(\"references.branches.count\", stats->refs.branches);\n\nThis does indeed cut down on some of the boilerplate, which maybe would\nmake it a little bit easier to catch any mismatches.\n\n> if we wanted to.\n\nUltimately, I don't feel strongly either way though. I've amended\nlocally, but will hold off on sending another version for now unless\nthere is additional feedback.\n\nThanks,\n-Justin\n"},{"id":"538126","messageId":"xmqq342cy49e.fsf@gitster.g","threadId":"64911","inReplyTo":"xmqqqzq1yjcl.fsf@gitster.g","subject":"Re: [PATCH v3 0/6] builtin/repo: include largest object information","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-06T22:36:29Z","receivedAt":"2026-03-06T22:36:31Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Justin Tobler <jltobler@gmail.com> writes:\n>\n>> Changes from V2:\n>> - When checking for largest objects, zero valued objects were not\n>>   recorded even if they were the \"largest\" object. In this version, if\n>>   an object ID has not been recorded yet, it is always added even if its\n>>   value is zero.\n>> - Added some helper functions for printing keyvalue info to cut down on\n>>   duplicate code and hopefully make it a bit easier on the eyes.\n>> - Moved the for-each loop that printed table OID annoations inside the\n>>   preceding if-block making it a bit easier to reason about.\n>\n> The changes I see in the diff relative to the previous iteration all\n> look sane to me.  Will replace.  Thanks.\n\nIt seems that no further review comments are coming and new\niterations are not happening on this topic, so shall we declare\nvictory and mark the topic for 'next' now?\n\nThanks.\n"},{"id":"538212","messageId":"aa3DNVshSsAjFY1y@denethor","threadId":"64911","inReplyTo":"xmqq342cy49e.fsf@gitster.g","subject":"Re: [PATCH v3 0/6] builtin/repo: include largest object information","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-03-08T18:44:09Z","receivedAt":"2026-03-08T18:44:11Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 26/03/06 02:36PM, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> > Justin Tobler <jltobler@gmail.com> writes:\n> >\n> >> Changes from V2:\n> >> - When checking for largest objects, zero valued objects were not\n> >>   recorded even if they were the \"largest\" object. In this version, if\n> >>   an object ID has not been recorded yet, it is always added even if its\n> >>   value is zero.\n> >> - Added some helper functions for printing keyvalue info to cut down on\n> >>   duplicate code and hopefully make it a bit easier on the eyes.\n> >> - Moved the for-each loop that printed table OID annoations inside the\n> >>   preceding if-block making it a bit easier to reason about.\n> >\n> > The changes I see in the diff relative to the previous iteration all\n> > look sane to me.  Will replace.  Thanks.\n> \n> It seems that no further review comments are coming and new\n> iterations are not happening on this topic, so shall we declare\n> victory and mark the topic for 'next' now?\n\nFrom my perspective, I think this topic is good for 'next' now.\n\nThanks,\n-Justin\n"}]}