{"thread":{"id":"64606","subject":"[PATCH 0/6] builtin/repo: add object size info to structure output","startedAt":"2025-12-09T22:58:27Z","lastAt":"2025-12-18T06:32:28Z","messageCount":80,"participants":["Justin Tobler","Patrick Steinhardt","Junio C Hamano","Lucas Seiki Oshiro","Jiang Xin"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"531926","messageId":"20251209225820.2861276-1-jltobler@gmail.com","threadId":"64606","inReplyTo":null,"subject":"[PATCH 0/6] builtin/repo: add object size info to structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-09T22:58:14Z","receivedAt":"2025-12-09T22:58:27Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThis patch series extends the recently introduced \"structure\" subcommand\nfor git-repo(1) to collect object size information. More specifically,\nit shows total inflated and disk sizes of objects by object type. The\naim to provide additional insight that may be useful to users regarding\nthe structure of a repository.\n\nIn addition to this change, this series also updates the table output\nformat to downscale larger output values along with the appropriate unit\nprefix. This is done to make table output more human friendly. The\nkeyvalue and nul output formats are left the same since they are\nintended more for machine parsing.\n\nThanks,\n-Justin\n\nJustin Tobler (6):\n  builtin/repo: group per-type object values into struct\n  builtin/repo: humanise count values in structure output\n  builtin/repo: add inflated object info to keyvalue structure output\n  builtin/repo: add inflated object info to structure table\n  builtin/repo: add disk size info to keyvalue stucture output\n  builtin/repo: add object disk size info to structure table\n\n Documentation/git-repo.adoc |   2 +\n builtin/repo.c              | 222 +++++++++++++++++++++++++++++++-----\n t/t1901-repo-structure.sh   | 142 ++++++++++++++++-------\n 3 files changed, 295 insertions(+), 71 deletions(-)\n\n\nbase-commit: e85ae279b0d58edc2f4c3fd5ac391b51e1223985\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"531927","messageId":"20251209225820.2861276-2-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251209225820.2861276-1-jltobler@gmail.com","subject":"[PATCH 1/6] builtin/repo: group per-type object values into struct","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-09T22:58:15Z","receivedAt":"2025-12-09T22:58:28Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The `object_stats` structure stores object counts by type. In a\nsubsequent commit, additional per-type object measurements will also be\nstored. Group per-type object values into a new struct to allow better\nreuse.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 42 +++++++++++++++++++++++++-----------------\n 1 file changed, 25 insertions(+), 17 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 2a653bd3ea..a69699857a 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -202,13 +202,17 @@ struct ref_stats {\n \tsize_t others;\n };\n \n-struct object_stats {\n+struct object_values {\n \tsize_t tags;\n \tsize_t commits;\n \tsize_t trees;\n \tsize_t blobs;\n };\n \n+struct object_stats {\n+\tstruct object_values type_counts;\n+};\n+\n struct repo_structure {\n \tstruct ref_stats refs;\n \tstruct object_stats objects;\n@@ -281,9 +285,9 @@ static inline size_t get_total_reference_count(struct ref_stats *stats)\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n }\n \n-static inline size_t get_total_object_count(struct object_stats *stats)\n+static inline size_t get_total_object_values(struct object_values *values)\n {\n-\treturn stats->tags + stats->commits + stats->trees + stats->blobs;\n+\treturn values->tags + values->commits + values->trees + values->blobs;\n }\n \n static void stats_table_setup_structure(struct stats_table *table,\n@@ -302,14 +306,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_count(objects);\n+\tobject_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n \tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n-\tstats_table_count_addf(table, objects->commits, \"    * %s\", _(\"Commits\"));\n-\tstats_table_count_addf(table, objects->trees, \"    * %s\", _(\"Trees\"));\n-\tstats_table_count_addf(table, objects->blobs, \"    * %s\", _(\"Blobs\"));\n-\tstats_table_count_addf(table, objects->tags, \"    * %s\", _(\"Tags\"));\n+\tstats_table_count_addf(table, objects->type_counts.commits,\n+\t\t\t       \"    * %s\", _(\"Commits\"));\n+\tstats_table_count_addf(table, objects->type_counts.trees,\n+\t\t\t       \"    * %s\", _(\"Trees\"));\n+\tstats_table_count_addf(table, objects->type_counts.blobs,\n+\t\t\t       \"    * %s\", _(\"Blobs\"));\n+\tstats_table_count_addf(table, objects->type_counts.tags,\n+\t\t\t       \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -389,13 +397,13 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \t       (uintmax_t)stats->refs.others, value_delim);\n \n \tprintf(\"objects.commits.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.commits, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.commits, value_delim);\n \tprintf(\"objects.trees.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.trees, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.trees, value_delim);\n \tprintf(\"objects.blobs.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.blobs, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.blobs, value_delim);\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.tags, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n \tfflush(stdout);\n }\n@@ -473,22 +481,22 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n-\t\tstats->tags += oids->nr;\n+\t\tstats->type_counts.tags += oids->nr;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n-\t\tstats->commits += oids->nr;\n+\t\tstats->type_counts.commits += oids->nr;\n \t\tbreak;\n \tcase OBJ_TREE:\n-\t\tstats->trees += oids->nr;\n+\t\tstats->type_counts.trees += oids->nr;\n \t\tbreak;\n \tcase OBJ_BLOB:\n-\t\tstats->blobs += oids->nr;\n+\t\tstats->type_counts.blobs += oids->nr;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n \t}\n \n-\tobject_count = get_total_object_count(stats);\n+\tobject_count = get_total_object_values(&stats->type_counts);\n \tdisplay_progress(data->progress, object_count);\n \n \treturn 0;\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"531928","messageId":"20251209225820.2861276-3-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251209225820.2861276-1-jltobler@gmail.com","subject":"[PATCH 2/6] builtin/repo: humanise count values in structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-09T22:58:16Z","receivedAt":"2025-12-09T22:58:29Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The table output format for the git-repo(1) structure subcommand is used\nby default and intended to provide output to users in a human-friendly\nmanner. When the reference/object count values in a repository are\nlarge, it becomes more cumbersome for users to read the values.\n\nFor larger values, update the table output format to instead produce\nmore human-friendly count values that are scaled down with the\nappropriate unit prefix. Output for the keyvalue and nul formats remains\nunchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 61 +++++++++++++++++++++++++++++++-------\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++--------------------\n 2 files changed, 82 insertions(+), 41 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex a69699857a..8fb728b3a5 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -223,6 +223,7 @@ struct stats_table {\n \n \tint name_col_width;\n \tint value_col_width;\n+\tint unit_col_width;\n };\n \n /*\n@@ -230,6 +231,7 @@ struct stats_table {\n  */\n struct stats_table_entry {\n \tchar *value;\n+\tconst char *unit;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -250,11 +252,18 @@ static void stats_table_vaddf(struct stats_table *table,\n \n \tif (name_width > table->name_col_width)\n \t\ttable->name_col_width = name_width;\n-\tif (entry) {\n+\tif (!entry)\n+\t\treturn;\n+\tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n \t\t\ttable->value_col_width = value_width;\n \t}\n+\tif (entry->unit) {\n+\t\tint unit_width = utf8_strwidth(entry->unit);\n+\t\tif (unit_width > table->unit_col_width)\n+\t\t\ttable->unit_col_width = unit_width;\n+\t}\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -266,6 +275,10 @@ static void stats_table_addf(struct stats_table *table, const char *format, ...)\n \tva_end(ap);\n }\n \n+static const char *unit_k = \"k\";\n+static const char *unit_M = \"M\";\n+static const char *unit_G = \"G\";\n+\n static void stats_table_count_addf(struct stats_table *table, size_t value,\n \t\t\t\t   const char *format, ...)\n {\n@@ -273,7 +286,26 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_list ap;\n \n \tCALLOC_ARRAY(entry, 1);\n-\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n+\n+\tif (value >= 1000000000) {\n+\t\tuintmax_t x = (uintmax_t)value + 5000000;\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n+\t\t\t\t       x / 1000000000,\n+\t\t\t\t       x % 1000000000 / 10000000);\n+\t\tentry->unit = unit_G;\n+\t} else if (value >= 1000000) {\n+\t\tuintmax_t x = (uintmax_t)value + 5000;\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n+\t\t\t\t       x / 1000000, x % 1000000 / 10000);\n+\t\tentry->unit = unit_M;\n+\t} else if (value >= 1000) {\n+\t\tuintmax_t x = (uintmax_t)value + 5;\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n+\t\t\t\t       x / 1000, x % 1000 / 10);\n+\t\tentry->unit = unit_k;\n+\t} else {\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n+\t}\n \n \tva_start(ap, format);\n \tstats_table_vaddf(table, entry, format, ap);\n@@ -324,20 +356,24 @@ static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n \tconst char *value_col_title = _(\"Value\");\n-\tint name_col_width = utf8_strwidth(name_col_title);\n-\tint value_col_width = utf8_strwidth(value_col_title);\n+\tint title_name_width = utf8_strwidth(name_col_title);\n+\tint title_value_width = utf8_strwidth(value_col_title);\n+\tint name_col_width = table->name_col_width;\n+\tint value_col_width = table->value_col_width;\n+\tint unit_col_width = table->unit_col_width;\n \tstruct string_list_item *item;\n \tstruct strbuf buf = STRBUF_INIT;\n \n-\tif (table->name_col_width > name_col_width)\n-\t\tname_col_width = table->name_col_width;\n-\tif (table->value_col_width > value_col_width)\n-\t\tvalue_col_width = table->value_col_width;\n+\tif (title_name_width > name_col_width)\n+\t\tname_col_width = title_name_width;\n+\tif (title_value_width > value_col_width + unit_col_width + 1)\n+\t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n+\t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n \tstrbuf_addstr(&buf, \" |\");\n \tprintf(\"%s\\n\", buf.buf);\n \n@@ -345,17 +381,20 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tfor (int i = 0; i < name_col_width; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n-\tfor (int i = 0; i < value_col_width; i++)\n+\tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n \t\tputchar('-');\n \tprintf(\" |\\n\");\n \n \tfor_each_string_list_item(item, &table->rows) {\n \t\tstruct stats_table_entry *entry = item->util;\n \t\tconst char *value = \"\";\n+\t\tconst char *unit = \"\";\n \n \t\tif (entry) {\n \t\t\tstruct stats_table_entry *entry = item->util;\n \t\t\tvalue = entry->value;\n+\t\t\tif (entry->unit)\n+\t\t\t\tunit = entry->unit;\n \t\t}\n \n \t\tstrbuf_reset(&buf);\n@@ -363,6 +402,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n+\t\tstrbuf_addch(&buf, ' ');\n+\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, unit_col_width, unit);\n \t\tstrbuf_addstr(&buf, \" |\");\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 36a71a144e..55fd13ad1b 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -10,21 +10,21 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Branches       |     0 |\n-\t\t|     * Tags           |     0 |\n-\t\t|     * Remotes        |     0 |\n-\t\t|     * Others         |     0 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Commits        |     0 |\n-\t\t|     * Trees          |     0 |\n-\t\t|     * Blobs          |     0 |\n-\t\t|     * Tags           |     0 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Branches       |     0  |\n+\t\t|     * Tags           |     0  |\n+\t\t|     * Remotes        |     0  |\n+\t\t|     * Others         |     0  |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Commits        |     0  |\n+\t\t|     * Trees          |     0  |\n+\t\t|     * Blobs          |     0  |\n+\t\t|     * Tags           |     0  |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -39,7 +39,7 @@ test_expect_success 'repository with references and objects' '\n \tgit init repo &&\n \t(\n \t\tcd repo &&\n-\t\ttest_commit_bulk 42 &&\n+\t\ttest_commit_bulk 1005 &&\n \t\tgit tag -a foo -m bar &&\n \n \t\toid=\"$(git rev-parse HEAD)\" &&\n@@ -49,21 +49,21 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     4 |\n-\t\t|     * Branches       |     1 |\n-\t\t|     * Tags           |     1 |\n-\t\t|     * Remotes        |     1 |\n-\t\t|     * Others         |     1 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |   130 |\n-\t\t|     * Commits        |    43 |\n-\t\t|     * Trees          |    43 |\n-\t\t|     * Blobs          |    43 |\n-\t\t|     * Tags           |     1 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |    4   |\n+\t\t|     * Branches       |    1   |\n+\t\t|     * Tags           |    1   |\n+\t\t|     * Remotes        |    1   |\n+\t\t|     * Others         |    1   |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            | 3.02 k |\n+\t\t|     * Commits        | 1.01 k |\n+\t\t|     * Trees          | 1.01 k |\n+\t\t|     * Blobs          | 1.01 k |\n+\t\t|     * Tags           |    1   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"531929","messageId":"20251209225820.2861276-4-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251209225820.2861276-1-jltobler@gmail.com","subject":"[PATCH 3/6] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-09T22:58:17Z","receivedAt":"2025-12-09T22:58:30Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The structure subcommand for git-repo(1) outputs basic count information\nfor objects and references. Extend this output to also provide\ninformation regarding total size of inflated objects by object type.\n\nFor now, object size by object type info is only added to the keyvalue\nand nul output formats. In a subsequent commit, this info is also added\nto the table format.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 33 +++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   |  6 +++++-\n 3 files changed, 39 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 70f0a6d2e4..287eee4b93 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -50,6 +50,7 @@ supported:\n +\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n+* Total inflated size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 8fb728b3a5..a67215ae31 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,8 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hex.h\"\n+#include \"odb.h\"\n #include \"parse-options.h\"\n #include \"path-walk.h\"\n #include \"progress.h\"\n@@ -211,6 +213,7 @@ struct object_values {\n \n struct object_stats {\n \tstruct object_values type_counts;\n+\tstruct object_values inflated_sizes;\n };\n \n struct repo_structure {\n@@ -446,6 +449,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n+\tprintf(\"objects.commits.inflated%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.inflated%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.inflated%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.inflated%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -509,6 +521,7 @@ static void structure_count_references(struct ref_stats *stats,\n }\n \n struct count_objects_data {\n+\tstruct object_database *odb;\n \tstruct object_stats *stats;\n \tstruct progress *progress;\n };\n@@ -518,20 +531,39 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n+\tsize_t inflated_total = 0;\n \tsize_t object_count;\n \n+\tfor (size_t i = 0; i < oids->nr; i++) {\n+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n+\t\tunsigned long inflated;\n+\n+\t\toi.sizep = &inflated;\n+\n+\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n+\t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n+\t\t\tdie(_(\"cannot read object for %s\"),\n+\t\t\t    oid_to_hex(&oids->oid[i]));\n+\n+\t\tinflated_total += inflated;\n+\t}\n+\n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n+\t\tstats->inflated_sizes.tags += inflated_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n+\t\tstats->inflated_sizes.commits += inflated_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n+\t\tstats->inflated_sizes.trees += inflated_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n+\t\tstats->inflated_sizes.blobs += inflated_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n@@ -549,6 +581,7 @@ static void structure_count_objects(struct object_stats *stats,\n {\n \tstruct path_walk_info info = PATH_WALK_INFO_INIT;\n \tstruct count_objects_data data = {\n+\t\t.odb = repo->objects,\n \t\t.stats = stats,\n \t};\n \ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 55fd13ad1b..cf5e252f10 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -73,7 +73,7 @@ test_expect_success 'repository with references and objects' '\n \t)\n '\n \n-test_expect_success 'keyvalue and nul format' '\n+test_expect_success SHA1 'keyvalue and nul format' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -90,6 +90,10 @@ test_expect_success 'keyvalue and nul format' '\n \t\tobjects.trees.count=42\n \t\tobjects.blobs.count=42\n \t\tobjects.tags.count=1\n+\t\tobjects.commits.inflated=9225\n+\t\tobjects.trees.inflated=28554\n+\t\tobjects.blobs.inflated=453\n+\t\tobjects.tags.inflated=132\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"531930","messageId":"20251209225820.2861276-5-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251209225820.2861276-1-jltobler@gmail.com","subject":"[PATCH 4/6] builtin/repo: add inflated object info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-09T22:58:18Z","receivedAt":"2025-12-09T22:58:31Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Update the table output format for the git-repo(1) structure command to\nbegin printing the total inflated object size info by object type. To be\nmore human-friendly, larger values are scaled down and displayed with\nthe appropriate unit prefix. Output for the keyvalue and nul formats\nremains unchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 57 +++++++++++++++++++++++++++++++++--\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++++++----------------\n 2 files changed, 90 insertions(+), 29 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex a67215ae31..5c37f4116f 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -315,6 +315,44 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static const char *unit_B = \"B\";\n+static const char *unit_KiB = \"KiB\";\n+static const char *unit_MiB = \"MiB\";\n+static const char *unit_GiB = \"GiB\";\n+\n+static void stats_table_size_addf(struct stats_table *table, size_t value,\n+\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\n+\tif (value > 1 << 30) {\n+\t\tuintmax_t x = (uintmax_t)value + 5368709;\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 30,\n+\t\t\t\t       ((x & ((1 << 30) - 1)) * 100) >> 30);\n+\t\tentry->unit = unit_GiB;\n+\t} else if (value > 1 << 20) {\n+\t\tuintmax_t x = (uintmax_t)value + 5243;\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 20,\n+\t\t\t\t       ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\tentry->unit = unit_MiB;\n+\t} else if (value > 1 << 10) {\n+\t\tuintmax_t x = (uintmax_t)value + 5;\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 10,\n+\t\t\t\t       ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\tentry->unit = unit_KiB;\n+\t} else {\n+\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n+\t\tentry->unit = unit_B;\n+\t}\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -330,7 +368,8 @@ static void stats_table_setup_structure(struct stats_table *table,\n {\n \tstruct object_stats *objects = &stats->objects;\n \tstruct ref_stats *refs = &stats->refs;\n-\tsize_t object_total;\n+\tsize_t inflated_object_total;\n+\tsize_t object_count_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -341,10 +380,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_values(&objects->type_counts);\n+\tobject_count_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n-\tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n+\tstats_table_count_addf(table, object_count_total, \"  * %s\", _(\"Count\"));\n \tstats_table_count_addf(table, objects->type_counts.commits,\n \t\t\t       \"    * %s\", _(\"Commits\"));\n \tstats_table_count_addf(table, objects->type_counts.trees,\n@@ -353,6 +392,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t       \"    * %s\", _(\"Blobs\"));\n \tstats_table_count_addf(table, objects->type_counts.tags,\n \t\t\t       \"    * %s\", _(\"Tags\"));\n+\n+\tinflated_object_total = get_total_object_values(&objects->inflated_sizes);\n+\tstats_table_size_addf(table, inflated_object_total,\n+\t\t\t      \"  * %s\", _(\"Inflated size\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex cf5e252f10..0ae96e6bbf 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -13,18 +13,23 @@ test_expect_success 'empty repository' '\n \t\t| Repository structure | Value  |\n \t\t| -------------------- | ------ |\n \t\t| * References         |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Branches       |     0  |\n-\t\t|     * Tags           |     0  |\n-\t\t|     * Remotes        |     0  |\n-\t\t|     * Others         |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Branches       |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|     * Remotes        |    0   |\n+\t\t|     * Others         |    0   |\n \t\t|                      |        |\n \t\t| * Reachable objects  |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Commits        |     0  |\n-\t\t|     * Trees          |     0  |\n-\t\t|     * Blobs          |     0  |\n-\t\t|     * Tags           |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Commits        |    0   |\n+\t\t|     * Trees          |    0   |\n+\t\t|     * Blobs          |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|   * Inflated size    |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -34,7 +39,7 @@ test_expect_success 'empty repository' '\n \t)\n '\n \n-test_expect_success 'repository with references and objects' '\n+test_expect_success SHA1 'repository with references and objects' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -49,21 +54,26 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    4   |\n-\t\t|     * Branches       |    1   |\n-\t\t|     * Tags           |    1   |\n-\t\t|     * Remotes        |    1   |\n-\t\t|     * Others         |    1   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            | 3.02 k |\n-\t\t|     * Commits        | 1.01 k |\n-\t\t|     * Trees          | 1.01 k |\n-\t\t|     * Blobs          | 1.01 k |\n-\t\t|     * Tags           |    1   |\n+\t\t| Repository structure | Value      |\n+\t\t| -------------------- | ---------- |\n+\t\t| * References         |            |\n+\t\t|   * Count            |      4     |\n+\t\t|     * Branches       |      1     |\n+\t\t|     * Tags           |      1     |\n+\t\t|     * Remotes        |      1     |\n+\t\t|     * Others         |      1     |\n+\t\t|                      |            |\n+\t\t| * Reachable objects  |            |\n+\t\t|   * Count            |   3.02 k   |\n+\t\t|     * Commits        |   1.01 k   |\n+\t\t|     * Trees          |   1.01 k   |\n+\t\t|     * Blobs          |   1.01 k   |\n+\t\t|     * Tags           |      1     |\n+\t\t|   * Inflated size    |  16.03 MiB |\n+\t\t|     * Commits        | 217.92 KiB |\n+\t\t|     * Trees          |  15.81 MiB |\n+\t\t|     * Blobs          |  11.68 KiB |\n+\t\t|     * Tags           |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"531931","messageId":"20251209225820.2861276-6-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251209225820.2861276-1-jltobler@gmail.com","subject":"[PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-09T22:58:19Z","receivedAt":"2025-12-09T22:58:32Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, extend the keyvalue and nul output formats of\nthe git-repo(1) structure command to additionally provide info regarding\ntotal object disk sizes by object type.\n\nSince disk size may vary between platforms, tests do not validate actual\nvalues and only check that size info is printed in an empty repository.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 18 +++++++++++++++\n t/t1901-repo-structure.sh   | 45 +++++++++++++++++++++++++++++--------\n 3 files changed, 55 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 287eee4b93..861073f641 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -51,6 +51,7 @@ supported:\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n+* Total disk size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 5c37f4116f..8ea7c9b24f 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -214,6 +214,7 @@ struct object_values {\n struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n+\tstruct object_values disk_sizes;\n };\n \n struct repo_structure {\n@@ -509,6 +510,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.inflated%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n \n+\tprintf(\"objects.commits.disk%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.disk%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.disk%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.disk%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -583,13 +593,16 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n \tsize_t inflated_total = 0;\n+\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\toff_t disk;\n \n \t\toi.sizep = &inflated;\n+\t\toi.disk_sizep = &disk;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n@@ -597,24 +610,29 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\t    oid_to_hex(&oids->oid[i]));\n \n \t\tinflated_total += inflated;\n+\t\tdisk_total += disk;\n \t}\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n \t\tstats->inflated_sizes.tags += inflated_total;\n+\t\tstats->disk_sizes.tags += disk_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n \t\tstats->inflated_sizes.commits += inflated_total;\n+\t\tstats->disk_sizes.commits += disk_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n \t\tstats->inflated_sizes.trees += inflated_total;\n+\t\tstats->disk_sizes.trees += disk_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n \t\tstats->inflated_sizes.blobs += inflated_total;\n+\t\tstats->disk_sizes.blobs += disk_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 0ae96e6bbf..a98c651f1d 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -35,6 +35,37 @@ test_expect_success 'empty repository' '\n \t\tgit repo structure >out 2>err &&\n \n \t\ttest_cmp expect out &&\n+\t\ttest_line_count = 0 err &&\n+\n+\t\tcat >expect <<-\\EOF &&\n+\t\treferences.branches.count=0\n+\t\treferences.tags.count=0\n+\t\treferences.remotes.count=0\n+\t\treferences.others.count=0\n+\t\tobjects.commits.count=0\n+\t\tobjects.trees.count=0\n+\t\tobjects.blobs.count=0\n+\t\tobjects.tags.count=0\n+\t\tobjects.commits.inflated=0\n+\t\tobjects.trees.inflated=0\n+\t\tobjects.blobs.inflated=0\n+\t\tobjects.tags.inflated=0\n+\t\tobjects.commits.disk=0\n+\t\tobjects.trees.disk=0\n+\t\tobjects.blobs.disk=0\n+\t\tobjects.tags.disk=0\n+\t\tEOF\n+\n+\t\tgit repo structure --format=keyvalue >out 2>err &&\n+\n+\t\ttest_cmp expect out &&\n+\t\ttest_line_count = 0 err &&\n+\n+\t\t# Replace key and value delimiters for nul format.\n+\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n+\t\tgit repo structure --format=nul >out 2>err &&\n+\n+\t\ttest_cmp expect_nul out &&\n \t\ttest_line_count = 0 err\n \t)\n '\n@@ -83,7 +114,7 @@ test_expect_success SHA1 'repository with references and objects' '\n \t)\n '\n \n-test_expect_success SHA1 'keyvalue and nul format' '\n+test_expect_success SHA1 'keyvalue format' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -106,16 +137,12 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.tags.inflated=132\n \t\tEOF\n \n-\t\tgit repo structure --format=keyvalue >out 2>err &&\n+\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n \n-\t\ttest_cmp expect out &&\n-\t\ttest_line_count = 0 err &&\n+\t\t# Strip object disk usage from output due to platform variance.\n+\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n \n-\t\t# Replace key and value delimiters for nul format.\n-\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n-\t\tgit repo structure --format=nul >out 2>err &&\n-\n-\t\ttest_cmp expect_nul out &&\n+\t\ttest_cmp expect out &&\n \t\ttest_line_count = 0 err\n \t)\n '\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"531932","messageId":"20251209225820.2861276-7-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251209225820.2861276-1-jltobler@gmail.com","subject":"[PATCH 6/6] builtin/repo: add object disk size info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-09T22:58:20Z","receivedAt":"2025-12-09T22:58:33Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, update the table output format for the\ngit-repo(1) structure commdn to display the total object disk usage by\nobject type.\n\nSince disk size may vary between platforms, tests do not validate actual\nvalues and only check that size info is printed in an empty repository.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 13 +++++++++++++\n t/t1901-repo-structure.sh | 19 ++++++++++++++++++-\n 2 files changed, 31 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 8ea7c9b24f..8ddefd523e 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -371,6 +371,7 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstruct ref_stats *refs = &stats->refs;\n \tsize_t inflated_object_total;\n \tsize_t object_count_total;\n+\tsize_t disk_object_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -405,6 +406,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->inflated_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tdisk_object_total = get_total_object_values(&objects->disk_sizes);\n+\tstats_table_size_addf(table, disk_object_total,\n+\t\t\t      \"  * %s\", _(\"Disk size\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex a98c651f1d..51820cc3f6 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -4,6 +4,15 @@ test_description='test git repo structure'\n \n . ./test-lib.sh\n \n+strip_object_disk_usage() {\n+\tawk '\n+\t\t/^\\|   \\* Disk size/ { skip=1; next }\n+\t\tskip && /^\\|     \\* / { next }\n+\t\tskip && !/^\\|     \\* / { skip=0 }\n+\t\t{ print }\n+\t' $1\n+}\n+\n test_expect_success 'empty repository' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n@@ -30,6 +39,11 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|   * Disk size        |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -107,7 +121,10 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Tags           |    132 B   |\n \t\tEOF\n \n-\t\tgit repo structure >out 2>err &&\n+\t\tgit repo structure >out.raw 2>err &&\n+\n+\t\t# Skip object disk sizes due to platform variance.\n+\t\tstrip_object_disk_usage out.raw >out &&\n \n \t\ttest_cmp expect out &&\n \t\ttest_line_count = 0 err\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"531944","messageId":"aTkS_kBlNsnbPyP5@pks.im","threadId":"64606","inReplyTo":"20251209225820.2861276-3-jltobler@gmail.com","subject":"Re: [PATCH 2/6] builtin/repo: humanise count values in structure output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-10T06:28:14Z","receivedAt":"2025-12-10T06:28:25Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 09, 2025 at 04:58:16PM -0600, Justin Tobler wrote:\n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index a69699857a..8fb728b3a5 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -266,6 +275,10 @@ static void stats_table_addf(struct stats_table *table, const char *format, ...)\n>  \tva_end(ap);\n>  }\n>  \n> +static const char *unit_k = \"k\";\n> +static const char *unit_M = \"M\";\n> +static const char *unit_G = \"G\";\n> +\n>  static void stats_table_count_addf(struct stats_table *table, size_t value,\n>  \t\t\t\t   const char *format, ...)\n>  {\n\nI would assume that these units should be translatable.\n\n> @@ -273,7 +286,26 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n>  \tva_list ap;\n>  \n>  \tCALLOC_ARRAY(entry, 1);\n> -\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n> +\n> +\tif (value >= 1000000000) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5000000;\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n> +\t\t\t\t       x / 1000000000,\n> +\t\t\t\t       x % 1000000000 / 10000000);\n> +\t\tentry->unit = unit_G;\n> +\t} else if (value >= 1000000) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5000;\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n> +\t\t\t\t       x / 1000000, x % 1000000 / 10000);\n> +\t\tentry->unit = unit_M;\n> +\t} else if (value >= 1000) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5;\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n> +\t\t\t\t       x / 1000, x % 1000 / 10);\n> +\t\tentry->unit = unit_k;\n> +\t} else {\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n> +\t}\n>  \n>  \tva_start(ap, format);\n>  \tstats_table_vaddf(table, entry, format, ap);\n\nThese units are decimal-based (1000), whereas in \"parse.c\" we have\n`get_unit_factor()` that is binary-based (1024). Arguably, it's\n\"parse.c\" that is wrong because \"k\" is generally decimal-based whereas\n\"Ki\" would be binary-based.\n\nNot quite sure what to do with this. For counts it _could_ be okay if we\ncontinue to use the wrong unit prefix. But as soon as we get to disk\nsizes we certainly should use the correct units, which would probably be\nKiB.\n\n> diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> index 36a71a144e..55fd13ad1b 100755\n> --- a/t/t1901-repo-structure.sh\n> +++ b/t/t1901-repo-structure.sh\n> @@ -10,21 +10,21 @@ test_expect_success 'empty repository' '\n>  \t(\n>  \t\tcd repo &&\n>  \t\tcat >expect <<-\\EOF &&\n> -\t\t| Repository structure | Value |\n> -\t\t| -------------------- | ----- |\n> -\t\t| * References         |       |\n> -\t\t|   * Count            |     0 |\n> -\t\t|     * Branches       |     0 |\n> -\t\t|     * Tags           |     0 |\n> -\t\t|     * Remotes        |     0 |\n> -\t\t|     * Others         |     0 |\n> -\t\t|                      |       |\n> -\t\t| * Reachable objects  |       |\n> -\t\t|   * Count            |     0 |\n> -\t\t|     * Commits        |     0 |\n> -\t\t|     * Trees          |     0 |\n> -\t\t|     * Blobs          |     0 |\n> -\t\t|     * Tags           |     0 |\n> +\t\t| Repository structure | Value  |\n> +\t\t| -------------------- | ------ |\n> +\t\t| * References         |        |\n> +\t\t|   * Count            |     0  |\n> +\t\t|     * Branches       |     0  |\n> +\t\t|     * Tags           |     0  |\n> +\t\t|     * Remotes        |     0  |\n> +\t\t|     * Others         |     0  |\n> +\t\t|                      |        |\n> +\t\t| * Reachable objects  |        |\n> +\t\t|   * Count            |     0  |\n> +\t\t|     * Commits        |     0  |\n> +\t\t|     * Trees          |     0  |\n> +\t\t|     * Blobs          |     0  |\n> +\t\t|     * Tags           |     0  |\n>  \t\tEOF\n>  \n>  \t\tgit repo structure >out 2>err &&\n\nIt's a bit weird that this test here changes even though we don't even\nuse any units. But I don't mind it too much.\n\nPatrick\n"},{"id":"531945","messageId":"aTkTCplQuSX_Y3oG@pks.im","threadId":"64606","inReplyTo":"20251209225820.2861276-6-jltobler@gmail.com","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-10T06:28:26Z","receivedAt":"2025-12-10T06:28:30Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 09, 2025 at 04:58:19PM -0600, Justin Tobler wrote:\n> diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> index 0ae96e6bbf..a98c651f1d 100755\n> --- a/t/t1901-repo-structure.sh\n> +++ b/t/t1901-repo-structure.sh\n> @@ -35,6 +35,37 @@ test_expect_success 'empty repository' '\n>  \t\tgit repo structure >out 2>err &&\n>  \n>  \t\ttest_cmp expect out &&\n> +\t\ttest_line_count = 0 err &&\n> +\n> +\t\tcat >expect <<-\\EOF &&\n> +\t\treferences.branches.count=0\n> +\t\treferences.tags.count=0\n> +\t\treferences.remotes.count=0\n> +\t\treferences.others.count=0\n> +\t\tobjects.commits.count=0\n> +\t\tobjects.trees.count=0\n> +\t\tobjects.blobs.count=0\n> +\t\tobjects.tags.count=0\n> +\t\tobjects.commits.inflated=0\n> +\t\tobjects.trees.inflated=0\n> +\t\tobjects.blobs.inflated=0\n> +\t\tobjects.tags.inflated=0\n> +\t\tobjects.commits.disk=0\n> +\t\tobjects.trees.disk=0\n> +\t\tobjects.blobs.disk=0\n> +\t\tobjects.tags.disk=0\n> +\t\tEOF\n\nDo we maybe want to adapt the keys to be \"inflated_size\" and\n\"disk_size\"?\n\n> @@ -106,16 +137,12 @@ test_expect_success SHA1 'keyvalue and nul format' '\n>  \t\tobjects.tags.inflated=132\n>  \t\tEOF\n>  \n> -\t\tgit repo structure --format=keyvalue >out 2>err &&\n> +\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n>  \n> -\t\ttest_cmp expect out &&\n> -\t\ttest_line_count = 0 err &&\n> +\t\t# Strip object disk usage from output due to platform variance.\n> +\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n>  \n> -\t\t# Replace key and value delimiters for nul format.\n> -\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n> -\t\tgit repo structure --format=nul >out 2>err &&\n> -\n> -\t\ttest_cmp expect_nul out &&\n> +\t\ttest_cmp expect out &&\n>  \t\ttest_line_count = 0 err\n>  \t)\n>  '\n\nWe could test disk sizes here test if we use git-rev-list(1) to compute\ndisk size by type:\n\n    git rev-list --disk-usage HEAD --objects --filter=object:type=blob\n    git rev-list --disk-usage HEAD --objects --filter=object:type=commit\n    git rev-list --disk-usage HEAD --objects --filter=object:type=tag\n    git rev-list --disk-usage HEAD --objects --filter=object:type=tree\n\nThe `--disk-usage` option also supports `--disk-usage=human`, which we\ncan use in the next commit to verify that our computations are the same\nacross git-rev-list(1) and git-repo(1).\n\nPatrick\n"},{"id":"531946","messageId":"aTkTEselZ4yL11qd@pks.im","threadId":"64606","inReplyTo":"20251209225820.2861276-5-jltobler@gmail.com","subject":"Re: [PATCH 4/6] builtin/repo: add inflated object info to structure table","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-10T06:28:34Z","receivedAt":"2025-12-10T06:28:39Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 09, 2025 at 04:58:18PM -0600, Justin Tobler wrote:\n> Update the table output format for the git-repo(1) structure command to\n> begin printing the total inflated object size info by object type. To be\n> more human-friendly, larger values are scaled down and displayed with\n> the appropriate unit prefix. Output for the keyvalue and nul formats\n> remains unchanged.\n> \n> Signed-off-by: Justin Tobler <jltobler@gmail.com>\n> ---\n>  builtin/repo.c            | 57 +++++++++++++++++++++++++++++++++--\n>  t/t1901-repo-structure.sh | 62 +++++++++++++++++++++++----------------\n>  2 files changed, 90 insertions(+), 29 deletions(-)\n> \n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index a67215ae31..5c37f4116f 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -315,6 +315,44 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n>  \tva_end(ap);\n>  }\n>  \n> +static const char *unit_B = \"B\";\n> +static const char *unit_KiB = \"KiB\";\n> +static const char *unit_MiB = \"MiB\";\n> +static const char *unit_GiB = \"GiB\";\n\nOkay, nice, you already use KiB et al as I suggested in an earlier\ncomment. But I guess these should also be marked as translatable.\n\n> +static void stats_table_size_addf(struct stats_table *table, size_t value,\n> +\t\t\t\t  const char *format, ...)\n> +{\n> +\tstruct stats_table_entry *entry;\n> +\tva_list ap;\n> +\n> +\tCALLOC_ARRAY(entry, 1);\n> +\n> +\tif (value > 1 << 30) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5368709;\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 30,\n> +\t\t\t\t       ((x & ((1 << 30) - 1)) * 100) >> 30);\n> +\t\tentry->unit = unit_GiB;\n> +\t} else if (value > 1 << 20) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5243;\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 20,\n> +\t\t\t\t       ((x & ((1 << 20) - 1)) * 100) >> 20);\n> +\t\tentry->unit = unit_MiB;\n> +\t} else if (value > 1 << 10) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5;\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 10,\n> +\t\t\t\t       ((x & ((1 << 10) - 1)) * 100) >> 10);\n> +\t\tentry->unit = unit_KiB;\n> +\t} else {\n> +\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n> +\t\tentry->unit = unit_B;\n> +\t}\n\nEuh. What kind of black magic is this? This block at least warrants a\ncomment how you came up with these incantations.\n\nAlso, git-rev-list(1) already has logic to output human-formatted disk\nsizes via `git rev-list --disk-usage=human`. Can we share the logic?\n\n> diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> index cf5e252f10..0ae96e6bbf 100755\n> --- a/t/t1901-repo-structure.sh\n> +++ b/t/t1901-repo-structure.sh\n> @@ -49,21 +54,26 @@ test_expect_success 'repository with references and objects' '\n>  \t\tgit notes add -m foo &&\n>  \n>  \t\tcat >expect <<-\\EOF &&\n> -\t\t| Repository structure | Value  |\n> -\t\t| -------------------- | ------ |\n> -\t\t| * References         |        |\n> -\t\t|   * Count            |    4   |\n> -\t\t|     * Branches       |    1   |\n> -\t\t|     * Tags           |    1   |\n> -\t\t|     * Remotes        |    1   |\n> -\t\t|     * Others         |    1   |\n> -\t\t|                      |        |\n> -\t\t| * Reachable objects  |        |\n> -\t\t|   * Count            | 3.02 k |\n> -\t\t|     * Commits        | 1.01 k |\n> -\t\t|     * Trees          | 1.01 k |\n> -\t\t|     * Blobs          | 1.01 k |\n> -\t\t|     * Tags           |    1   |\n> +\t\t| Repository structure | Value      |\n> +\t\t| -------------------- | ---------- |\n> +\t\t| * References         |            |\n> +\t\t|   * Count            |      4     |\n> +\t\t|     * Branches       |      1     |\n> +\t\t|     * Tags           |      1     |\n> +\t\t|     * Remotes        |      1     |\n> +\t\t|     * Others         |      1     |\n> +\t\t|                      |            |\n> +\t\t| * Reachable objects  |            |\n> +\t\t|   * Count            |   3.02 k   |\n> +\t\t|     * Commits        |   1.01 k   |\n> +\t\t|     * Trees          |   1.01 k   |\n> +\t\t|     * Blobs          |   1.01 k   |\n> +\t\t|     * Tags           |      1     |\n> +\t\t|   * Inflated size    |  16.03 MiB |\n> +\t\t|     * Commits        | 217.92 KiB |\n> +\t\t|     * Trees          |  15.81 MiB |\n> +\t\t|     * Blobs          |  11.68 KiB |\n> +\t\t|     * Tags           |    132 B   |\n>  \t\tEOF\n\nNice, I like the end result.\n\nPatrick\n"},{"id":"531947","messageId":"aTkTGilv-xRRQVHA@pks.im","threadId":"64606","inReplyTo":"20251209225820.2861276-7-jltobler@gmail.com","subject":"Re: [PATCH 6/6] builtin/repo: add object disk size info to structure table","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-10T06:28:42Z","receivedAt":"2025-12-10T06:28:47Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 09, 2025 at 04:58:20PM -0600, Justin Tobler wrote:\n> diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> index a98c651f1d..51820cc3f6 100755\n> --- a/t/t1901-repo-structure.sh\n> +++ b/t/t1901-repo-structure.sh\n> @@ -107,7 +121,10 @@ test_expect_success SHA1 'repository with references and objects' '\n>  \t\t|     * Tags           |    132 B   |\n>  \t\tEOF\n>  \n> -\t\tgit repo structure >out 2>err &&\n> +\t\tgit repo structure >out.raw 2>err &&\n> +\n> +\t\t# Skip object disk sizes due to platform variance.\n> +\t\tstrip_object_disk_usage out.raw >out &&\n\nAs mentioned, we can use git-rev-list(1) to compute the expected disk\nsizes.\n\nThanks!\n\nPatrick\n"},{"id":"531977","messageId":"xmqqikeegz8q.fsf@gitster.g","threadId":"64606","inReplyTo":"20251209225820.2861276-6-jltobler@gmail.com","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-10T14:58:29Z","receivedAt":"2025-12-10T14:58:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> -test_expect_success SHA1 'keyvalue and nul format' '\n> +test_expect_success SHA1 'keyvalue format' '\n>  \ttest_when_finished \"rm -rf repo\" &&\n>  \tgit init repo &&\n>  \t(\n> @@ -106,16 +137,12 @@ test_expect_success SHA1 'keyvalue and nul format' '\n>  \t\tobjects.tags.inflated=132\n>  \t\tEOF\n>  \n> -\t\tgit repo structure --format=keyvalue >out 2>err &&\n> +\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n>  \n> -\t\ttest_cmp expect out &&\n> -\t\ttest_line_count = 0 err &&\n> +\t\t# Strip object disk usage from output due to platform variance.\n> +\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n>  \n> -\t\t# Replace key and value delimiters for nul format.\n> -\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n> -\t\tgit repo structure --format=nul >out 2>err &&\n> -\n> -\t\ttest_cmp expect_nul out &&\n> +\t\ttest_cmp expect out &&\n>  \t\ttest_line_count = 0 err\n>  \t)\n>  '\n\nThis part has both textual and semantic conflicts with Lucas's \"-z\nis a synonym for --format=nul\" topic.  I think I resolved it\ncorrectly while improving the \"munge expected output into expected\nNUL-terminated output\" approach to \"munge -z output into textual\noutput and compare with textual expected output\".  Please sanity\ncheck the result after I push it out, merged at 32f8d84b (Merge\nbranch 'jt/repo-struct-more-objinfo' into seen, 2025-12-10)\n\nThanks.\n\ncommit 32f8d84b5cfc5a5704e30fe4abc9d8755893179c\nMerge: 09bd4419e7 b8cacabfa5\nAuthor: Junio C Hamano <gitster@pobox.com>\nDate:   Wed Dec 10 20:41:31 2025 +0900\n\n    Merge branch 'jt/repo-struct-more-objinfo' into seen\n    \n    * jt/repo-struct-more-objinfo:\n      builtin/repo: add object disk size info to structure table\n      builtin/repo: add disk size info to keyvalue stucture output\n      builtin/repo: add inflated object info to structure table\n      builtin/repo: add inflated object info to keyvalue structure output\n      builtin/repo: humanise count values in structure output\n      builtin/repo: group per-type object values into struct\n\ndiff --cc t/t1901-repo-structure.sh\nindex df7d4ea524,51820cc3f6..31c77c4666\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@@ -90,25 -148,18 +148,29 @@@ test_expect_success SHA1 'keyvalue form\n  \t\tobjects.trees.count=42\n  \t\tobjects.blobs.count=42\n  \t\tobjects.tags.count=1\n+ \t\tobjects.commits.inflated=9225\n+ \t\tobjects.trees.inflated=28554\n+ \t\tobjects.blobs.inflated=453\n+ \t\tobjects.tags.inflated=132\n  \t\tEOF\n  \n- \t\tgit repo structure --format=keyvalue >out 2>err &&\n+ \t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n  \n- \t\ttest_cmp expect out &&\n- \t\ttest_line_count = 0 err &&\n+ \t\t# Strip object disk usage from output due to platform variance.\n+ \t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n  \n- \t\t# Replace key and value delimiters for nul format.\n- \t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n- \t\tgit repo structure --format=nul >out 2>err &&\n- \n- \t\ttest_cmp expect_nul out &&\n++\t\ttest_cmp expect out &&\n +\t\ttest_line_count = 0 err &&\n +\n +\t\t# \"-z\", as a synonym to \"--format=nul\", participates in the\n +\t\t# usual \"last one wins\" rule.\n- \t\tgit repo structure --format=table -z >out 2>err &&\n++\t\tgit repo structure --format=table -z >out.raw 2>err &&\n +\n- \t\ttest_cmp expect_nul out &&\n++\t\t# Replace key and value delimiters for nul format.\n++\t\ttr \"\\0\\n\" \"\\n=\" <out.raw |\n++\t\tgrep -v \"objects\\..*\\.disk=\" >out &&\n++\n+ \t\ttest_cmp expect out &&\n  \t\ttest_line_count = 0 err\n  \t)\n  '\n"},{"id":"531979","messageId":"kf7vavs5yetooe6u2ygttzfriul4u5ywdnhtyksh2pbar4mpfz@orlg7ppajd7s","threadId":"64606","inReplyTo":"aTkS_kBlNsnbPyP5@pks.im","subject":"Re: [PATCH 2/6] builtin/repo: humanise count values in structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-10T15:10:29Z","receivedAt":"2025-12-10T15:10:35Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n> On Tue, Dec 09, 2025 at 04:58:16PM -0600, Justin Tobler wrote:\n> > diff --git a/builtin/repo.c b/builtin/repo.c\n> > index a69699857a..8fb728b3a5 100644\n> > --- a/builtin/repo.c\n> > +++ b/builtin/repo.c\n> > @@ -266,6 +275,10 @@ static void stats_table_addf(struct stats_table *table, const char *format, ...)\n> >  \tva_end(ap);\n> >  }\n> >  \n> > +static const char *unit_k = \"k\";\n> > +static const char *unit_M = \"M\";\n> > +static const char *unit_G = \"G\";\n> > +\n> >  static void stats_table_count_addf(struct stats_table *table, size_t value,\n> >  \t\t\t\t   const char *format, ...)\n> >  {\n> \n> I would assume that these units should be translatable.\n\nYa, you are right. I'll make units translatable in the next version.\n\n> > @@ -273,7 +286,26 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n> >  \tva_list ap;\n> >  \n> >  \tCALLOC_ARRAY(entry, 1);\n> > -\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n> > +\n> > +\tif (value >= 1000000000) {\n> > +\t\tuintmax_t x = (uintmax_t)value + 5000000;\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n> > +\t\t\t\t       x / 1000000000,\n> > +\t\t\t\t       x % 1000000000 / 10000000);\n> > +\t\tentry->unit = unit_G;\n> > +\t} else if (value >= 1000000) {\n> > +\t\tuintmax_t x = (uintmax_t)value + 5000;\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n> > +\t\t\t\t       x / 1000000, x % 1000000 / 10000);\n> > +\t\tentry->unit = unit_M;\n> > +\t} else if (value >= 1000) {\n> > +\t\tuintmax_t x = (uintmax_t)value + 5;\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n> > +\t\t\t\t       x / 1000, x % 1000 / 10);\n> > +\t\tentry->unit = unit_k;\n> > +\t} else {\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n> > +\t}\n> >  \n> >  \tva_start(ap, format);\n> >  \tstats_table_vaddf(table, entry, format, ap);\n> \n> These units are decimal-based (1000), whereas in \"parse.c\" we have\n> `get_unit_factor()` that is binary-based (1024). Arguably, it's\n> \"parse.c\" that is wrong because \"k\" is generally decimal-based whereas\n> \"Ki\" would be binary-based.\n> \n> Not quite sure what to do with this. For counts it _could_ be okay if we\n> continue to use the wrong unit prefix. But as soon as we get to disk\n> sizes we certainly should use the correct units, which would probably be\n> KiB.\n\nFor count values, such as number of references/objects, I'm using SI\nunit prefixes which I think is more correct. In a subsequent patch where\nwe start collect size information, I add a separate\n`stats_table_size_addf()` function which uses the IEC unit prefixes.\nThis way we use the most appropriate option for both scenarios.\n\n> > diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> > index 36a71a144e..55fd13ad1b 100755\n> > --- a/t/t1901-repo-structure.sh\n> > +++ b/t/t1901-repo-structure.sh\n> > @@ -10,21 +10,21 @@ test_expect_success 'empty repository' '\n> >  \t(\n> >  \t\tcd repo &&\n> >  \t\tcat >expect <<-\\EOF &&\n> > -\t\t| Repository structure | Value |\n> > -\t\t| -------------------- | ----- |\n> > -\t\t| * References         |       |\n> > -\t\t|   * Count            |     0 |\n> > -\t\t|     * Branches       |     0 |\n> > -\t\t|     * Tags           |     0 |\n> > -\t\t|     * Remotes        |     0 |\n> > -\t\t|     * Others         |     0 |\n> > -\t\t|                      |       |\n> > -\t\t| * Reachable objects  |       |\n> > -\t\t|   * Count            |     0 |\n> > -\t\t|     * Commits        |     0 |\n> > -\t\t|     * Trees          |     0 |\n> > -\t\t|     * Blobs          |     0 |\n> > -\t\t|     * Tags           |     0 |\n> > +\t\t| Repository structure | Value  |\n> > +\t\t| -------------------- | ------ |\n> > +\t\t| * References         |        |\n> > +\t\t|   * Count            |     0  |\n> > +\t\t|     * Branches       |     0  |\n> > +\t\t|     * Tags           |     0  |\n> > +\t\t|     * Remotes        |     0  |\n> > +\t\t|     * Others         |     0  |\n> > +\t\t|                      |        |\n> > +\t\t| * Reachable objects  |        |\n> > +\t\t|   * Count            |     0  |\n> > +\t\t|     * Commits        |     0  |\n> > +\t\t|     * Trees          |     0  |\n> > +\t\t|     * Blobs          |     0  |\n> > +\t\t|     * Tags           |     0  |\n> >  \t\tEOF\n> >  \n> >  \t\tgit repo structure >out 2>err &&\n> \n> It's a bit weird that this test here changes even though we don't even\n> use any units. But I don't mind it too much.\n\nYa, the added space comes from the fixed space character between the\nvalue and unit columns. I didn't think it mattered too much, but I may\ntry to only conditionally add it if needed in the next version.\n\n-Justin\n"},{"id":"531980","messageId":"vrlxdgvibiuohfc6k6nbmloibivntex33ucnmfdqpqu4dparbi@piyq3kckcyje","threadId":"64606","inReplyTo":"aTkTEselZ4yL11qd@pks.im","subject":"Re: [PATCH 4/6] builtin/repo: add inflated object info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-10T15:21:43Z","receivedAt":"2025-12-10T15:21:47Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n> On Tue, Dec 09, 2025 at 04:58:18PM -0600, Justin Tobler wrote:\n> > Update the table output format for the git-repo(1) structure command to\n> > begin printing the total inflated object size info by object type. To be\n> > more human-friendly, larger values are scaled down and displayed with\n> > the appropriate unit prefix. Output for the keyvalue and nul formats\n> > remains unchanged.\n> > \n> > Signed-off-by: Justin Tobler <jltobler@gmail.com>\n> > ---\n> >  builtin/repo.c            | 57 +++++++++++++++++++++++++++++++++--\n> >  t/t1901-repo-structure.sh | 62 +++++++++++++++++++++++----------------\n> >  2 files changed, 90 insertions(+), 29 deletions(-)\n> > \n> > diff --git a/builtin/repo.c b/builtin/repo.c\n> > index a67215ae31..5c37f4116f 100644\n> > --- a/builtin/repo.c\n> > +++ b/builtin/repo.c\n> > @@ -315,6 +315,44 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n> >  \tva_end(ap);\n> >  }\n> >  \n> > +static const char *unit_B = \"B\";\n> > +static const char *unit_KiB = \"KiB\";\n> > +static const char *unit_MiB = \"MiB\";\n> > +static const char *unit_GiB = \"GiB\";\n> \n> Okay, nice, you already use KiB et al as I suggested in an earlier\n> comment. But I guess these should also be marked as translatable.\n\nWill do.\n\n> > +static void stats_table_size_addf(struct stats_table *table, size_t value,\n> > +\t\t\t\t  const char *format, ...)\n> > +{\n> > +\tstruct stats_table_entry *entry;\n> > +\tva_list ap;\n> > +\n> > +\tCALLOC_ARRAY(entry, 1);\n> > +\n> > +\tif (value > 1 << 30) {\n> > +\t\tuintmax_t x = (uintmax_t)value + 5368709;\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 30,\n> > +\t\t\t\t       ((x & ((1 << 30) - 1)) * 100) >> 30);\n> > +\t\tentry->unit = unit_GiB;\n> > +\t} else if (value > 1 << 20) {\n> > +\t\tuintmax_t x = (uintmax_t)value + 5243;\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 20,\n> > +\t\t\t\t       ((x & ((1 << 20) - 1)) * 100) >> 20);\n> > +\t\tentry->unit = unit_MiB;\n> > +\t} else if (value > 1 << 10) {\n> > +\t\tuintmax_t x = (uintmax_t)value + 5;\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 10,\n> > +\t\t\t\t       ((x & ((1 << 10) - 1)) * 100) >> 10);\n> > +\t\tentry->unit = unit_KiB;\n> > +\t} else {\n> > +\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n> > +\t\tentry->unit = unit_B;\n> > +\t}\n> \n> Euh. What kind of black magic is this? This block at least warrants a\n> comment how you came up with these incantations.\n\nYa, I'll add some comments to explain what is going on here. :)\n\n> Also, git-rev-list(1) already has logic to output human-formatted disk\n> sizes via `git rev-list --disk-usage=human`. Can we share the logic?\n\nSo I believe `git rev-list --disk-usage=human` relies on\nstrbuf_humanise_bytes() under the hood. The problem here is that it\ncombines the value and unit prefix together. For alignment purposes in\nthe table output, we need to store the value and unit prefix separately.\nI couldn't immediately think of a great way to share logic here so I\nopted to implement it separately to accommodate this specific use case.\n\n-Justin\n"},{"id":"531981","messageId":"zq7iwyz6jhhj4bf5th2dwoe3ldtmtxeqqgrhx2mc4dgiaujzaa@frsaxmwmdzps","threadId":"64606","inReplyTo":"aTkTCplQuSX_Y3oG@pks.im","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-10T15:24:11Z","receivedAt":"2025-12-10T15:24:13Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n> On Tue, Dec 09, 2025 at 04:58:19PM -0600, Justin Tobler wrote:\n> > diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> > index 0ae96e6bbf..a98c651f1d 100755\n> > --- a/t/t1901-repo-structure.sh\n> > +++ b/t/t1901-repo-structure.sh\n> > @@ -35,6 +35,37 @@ test_expect_success 'empty repository' '\n> >  \t\tgit repo structure >out 2>err &&\n> >  \n> >  \t\ttest_cmp expect out &&\n> > +\t\ttest_line_count = 0 err &&\n> > +\n> > +\t\tcat >expect <<-\\EOF &&\n> > +\t\treferences.branches.count=0\n> > +\t\treferences.tags.count=0\n> > +\t\treferences.remotes.count=0\n> > +\t\treferences.others.count=0\n> > +\t\tobjects.commits.count=0\n> > +\t\tobjects.trees.count=0\n> > +\t\tobjects.blobs.count=0\n> > +\t\tobjects.tags.count=0\n> > +\t\tobjects.commits.inflated=0\n> > +\t\tobjects.trees.inflated=0\n> > +\t\tobjects.blobs.inflated=0\n> > +\t\tobjects.tags.inflated=0\n> > +\t\tobjects.commits.disk=0\n> > +\t\tobjects.trees.disk=0\n> > +\t\tobjects.blobs.disk=0\n> > +\t\tobjects.tags.disk=0\n> > +\t\tEOF\n> \n> Do we maybe want to adapt the keys to be \"inflated_size\" and\n> \"disk_size\"?\n\nGood suggestion. I'll update in the next version.\n\n> > @@ -106,16 +137,12 @@ test_expect_success SHA1 'keyvalue and nul format' '\n> >  \t\tobjects.tags.inflated=132\n> >  \t\tEOF\n> >  \n> > -\t\tgit repo structure --format=keyvalue >out 2>err &&\n> > +\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n> >  \n> > -\t\ttest_cmp expect out &&\n> > -\t\ttest_line_count = 0 err &&\n> > +\t\t# Strip object disk usage from output due to platform variance.\n> > +\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n> >  \n> > -\t\t# Replace key and value delimiters for nul format.\n> > -\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n> > -\t\tgit repo structure --format=nul >out 2>err &&\n> > -\n> > -\t\ttest_cmp expect_nul out &&\n> > +\t\ttest_cmp expect out &&\n> >  \t\ttest_line_count = 0 err\n> >  \t)\n> >  '\n> \n> We could test disk sizes here test if we use git-rev-list(1) to compute\n> disk size by type:\n> \n>     git rev-list --disk-usage HEAD --objects --filter=object:type=blob\n>     git rev-list --disk-usage HEAD --objects --filter=object:type=commit\n>     git rev-list --disk-usage HEAD --objects --filter=object:type=tag\n>     git rev-list --disk-usage HEAD --objects --filter=object:type=tree\n> \n> The `--disk-usage` option also supports `--disk-usage=human`, which we\n> can use in the next commit to verify that our computations are the same\n> across git-rev-list(1) and git-repo(1).\n\nThanks! I hadn't considered this. I'll try to update the tests in the\nnext version using this.\n\n-Justin\n"},{"id":"531982","messageId":"j4pc7xn4jjoyt3ay7clriz4kb2dz7toqluapt3xknm3gzlvitm@ge3ubr4eruw6","threadId":"64606","inReplyTo":"aTkTGilv-xRRQVHA@pks.im","subject":"Re: [PATCH 6/6] builtin/repo: add object disk size info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-10T15:24:36Z","receivedAt":"2025-12-10T15:24:38Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n> On Tue, Dec 09, 2025 at 04:58:20PM -0600, Justin Tobler wrote:\n> > diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> > index a98c651f1d..51820cc3f6 100755\n> > --- a/t/t1901-repo-structure.sh\n> > +++ b/t/t1901-repo-structure.sh\n> > @@ -107,7 +121,10 @@ test_expect_success SHA1 'repository with references and objects' '\n> >  \t\t|     * Tags           |    132 B   |\n> >  \t\tEOF\n> >  \n> > -\t\tgit repo structure >out 2>err &&\n> > +\t\tgit repo structure >out.raw 2>err &&\n> > +\n> > +\t\t# Skip object disk sizes due to platform variance.\n> > +\t\tstrip_object_disk_usage out.raw >out &&\n> \n> As mentioned, we can use git-rev-list(1) to compute the expected disk\n> sizes.\n\nThanks, I'll give this a go. :)\n\n-Justin\n"},{"id":"532000","messageId":"DF127A2A-AC63-4CB8-A405-7932D2A79E2C@gmail.com","threadId":"64606","inReplyTo":"xmqqikeegz8q.fsf@gitster.g","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Lucas Seiki Oshiro","fromEmail":"lucasseikioshiro@gmail.com","sentAt":"2025-12-10T19:09:05Z","receivedAt":"2025-12-10T19:09:19Z","isPatch":true,"sender":{"key":"lucasseikioshiro@gmail.com","avatar":"https://avatars.githubusercontent.com/u/12701580?v=4"},"body":"\n> This part has both textual and semantic conflicts with Lucas's \"-z\n> is a synonym for --format=nul\" topic.\n\nMy only change in this test was:\n\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 36a71a144e..df7d4ea524 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -101,6 +101,13 @@ test_expect_success 'keyvalue and nul format' '\n \t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n \t\tgit repo structure --format=nul >out 2>err &&\n\n+\t\ttest_cmp expect_nul out &&\n+\t\ttest_line_count = 0 err &&\n+\n+\t\t# \"-z\", as a synonym to \"--format=nul\", participates in the\n+\t\t# usual \"last one wins\" rule.\n+\t\tgit repo structure --format=table -z >out 2>err &&\n+\n \t\ttest_cmp expect_nul out &&\n \t\ttest_line_count = 0 err\n \t)\n\nGiven that Justin moved the --format=nul test to\n`test_expect_success 'empty repository'`, it should be ok to move\nmy change together with it. I did it here and everything seems to\nbe working.\n"},{"id":"532010","messageId":"xmqq1pl1hgj4.fsf@gitster.g","threadId":"64606","inReplyTo":"kf7vavs5yetooe6u2ygttzfriul4u5ywdnhtyksh2pbar4mpfz@orlg7ppajd7s","subject":"Re: [PATCH 2/6] builtin/repo: humanise count values in structure output","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-11T02:57:19Z","receivedAt":"2025-12-11T02:57:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n>> On Tue, Dec 09, 2025 at 04:58:16PM -0600, Justin Tobler wrote:\n>> > diff --git a/builtin/repo.c b/builtin/repo.c\n>> > index a69699857a..8fb728b3a5 100644\n>> > --- a/builtin/repo.c\n>> > +++ b/builtin/repo.c\n>> > @@ -266,6 +275,10 @@ static void stats_table_addf(struct stats_table *table, const char *format, ...)\n>> >  \tva_end(ap);\n>> >  }\n>> >  \n>> > +static const char *unit_k = \"k\";\n>> > +static const char *unit_M = \"M\";\n>> > +static const char *unit_G = \"G\";\n>> > +\n>> >  static void stats_table_count_addf(struct stats_table *table, size_t value,\n>> >  \t\t\t\t   const char *format, ...)\n>> >  {\n>> \n>> I would assume that these units should be translatable.\n>\n> Ya, you are right. I'll make units translatable in the next version.\n\nWhatever you do, please first consider reusing existing\n\"human-readable numbers\" helpers, like strbuf_humanise_bytes() used\nby the progress.c for showing throughput, before rolling your own\nvariant like the above.\n\nThanks.\n\n"},{"id":"532085","messageId":"epw5bctnxs7gpcg733qqtf2jxmsknntuylbqxt3xngs5zj7htn@zltxezklckge","threadId":"64606","inReplyTo":"xmqq1pl1hgj4.fsf@gitster.g","subject":"Re: [PATCH 2/6] builtin/repo: humanise count values in structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T16:46:28Z","receivedAt":"2025-12-12T16:46:33Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/11 11:57AM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n> >> On Tue, Dec 09, 2025 at 04:58:16PM -0600, Justin Tobler wrote:\n> >> > diff --git a/builtin/repo.c b/builtin/repo.c\n> >> > index a69699857a..8fb728b3a5 100644\n> >> > --- a/builtin/repo.c\n> >> > +++ b/builtin/repo.c\n> >> > @@ -266,6 +275,10 @@ static void stats_table_addf(struct stats_table *table, const char *format, ...)\n> >> >  \tva_end(ap);\n> >> >  }\n> >> >  \n> >> > +static const char *unit_k = \"k\";\n> >> > +static const char *unit_M = \"M\";\n> >> > +static const char *unit_G = \"G\";\n> >> > +\n> >> >  static void stats_table_count_addf(struct stats_table *table, size_t value,\n> >> >  \t\t\t\t   const char *format, ...)\n> >> >  {\n> >> \n> >> I would assume that these units should be translatable.\n> >\n> > Ya, you are right. I'll make units translatable in the next version.\n> \n> Whatever you do, please first consider reusing existing\n> \"human-readable numbers\" helpers, like strbuf_humanise_bytes() used\n> by the progress.c for showing throughput, before rolling your own\n> variant like the above.\n\nYa, I originally looked into using strbuf_humanise_bytes(), but went a\ndifferent direction due do how I wanted to align the values and unit\nprefixes in the table output. In the next version though, I'm trying to\nsplit out and reuse some of the same logic to avoid the duplication.\n\n-Justin\n"},{"id":"532090","messageId":"54kuvik2ecbkygjp57osmqjxiy7xtyjeffbzavuxbhuvta2oc5@mkqufah7cb3z","threadId":"64606","inReplyTo":"aTkTCplQuSX_Y3oG@pks.im","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T20:40:24Z","receivedAt":"2025-12-12T20:40:29Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n> On Tue, Dec 09, 2025 at 04:58:19PM -0600, Justin Tobler wrote:\n> > @@ -106,16 +137,12 @@ test_expect_success SHA1 'keyvalue and nul format' '\n> >  \t\tobjects.tags.inflated=132\n> >  \t\tEOF\n> >  \n> > -\t\tgit repo structure --format=keyvalue >out 2>err &&\n> > +\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n> >  \n> > -\t\ttest_cmp expect out &&\n> > -\t\ttest_line_count = 0 err &&\n> > +\t\t# Strip object disk usage from output due to platform variance.\n> > +\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n> >  \n> > -\t\t# Replace key and value delimiters for nul format.\n> > -\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n> > -\t\tgit repo structure --format=nul >out 2>err &&\n> > -\n> > -\t\ttest_cmp expect_nul out &&\n> > +\t\ttest_cmp expect out &&\n> >  \t\ttest_line_count = 0 err\n> >  \t)\n> >  '\n> \n> We could test disk sizes here test if we use git-rev-list(1) to compute\n> disk size by type:\n> \n>     git rev-list --disk-usage HEAD --objects --filter=object:type=blob\n>     git rev-list --disk-usage HEAD --objects --filter=object:type=commit\n>     git rev-list --disk-usage HEAD --objects --filter=object:type=tag\n>     git rev-list --disk-usage HEAD --objects --filter=object:type=tree\n> \n> The `--disk-usage` option also supports `--disk-usage=human`, which we\n> can use in the next commit to verify that our computations are the same\n> across git-rev-list(1) and git-repo(1).\n\nSo, I'm not sure we can use git-rev-list(1) in the manner suggested\nabove. It looks like user-specified objects are always included in the\noutput. When using \"HEAD\" this means the referenced object will always\nbe included regardless of the filter used. In practice, this means\nreported disk-usage when filtering by trees or blobs will likely be\ninflated by objects not specified by the filter. As far as I am aware,\nthere is not a way to suppress user-specified objects in git-rev-list(1)\noutput.\n\nI am somewhat curious if always including user-specified objects in\ngit-rev-list(1) output regardless of the specified filter is\nintentional. Looking at git-rev-list(1) --filter documentation:\n\n  The form --filter=object:type=(tag|commit|tree|blob) omits all objects\n  which are not of the requested type.\n\ndoesn't indicate this limitation. From looking at the code in\nlist-objects-filter.c:list_objects_filter__filter_object() though, it\ndoes somewhat seem like this behavior is intentional.\n\nRegardless, in the tests I can hack around this problem by using\nsomething like:\n\n  $ git cat-file --batch-check='$(objectsize:disk)' --batch-all-objects \\\n    --filter=object:type=tree | awk '{ sum += $1 } END { print sum }'\n\nto add up the sizes by object type. This doesn't really leave me a great\nway to verify the human-readable values in the table output though. I\nmay just continue to omit those values from the test like I already do\nin the next patch.\n\n-Justin\n"},{"id":"532096","messageId":"e5hsuevw5t37yt3zgp4hhtunusdyeg2lkph52pj4valpmlyrdt@7teicd67atbj","threadId":"64606","inReplyTo":"xmqqikeegz8q.fsf@gitster.g","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:15Z","receivedAt":"2025-12-12T22:36:20Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/10 11:58PM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > -test_expect_success SHA1 'keyvalue and nul format' '\n> > +test_expect_success SHA1 'keyvalue format' '\n> >  \ttest_when_finished \"rm -rf repo\" &&\n> >  \tgit init repo &&\n> >  \t(\n> > @@ -106,16 +137,12 @@ test_expect_success SHA1 'keyvalue and nul format' '\n> >  \t\tobjects.tags.inflated=132\n> >  \t\tEOF\n> >  \n> > -\t\tgit repo structure --format=keyvalue >out 2>err &&\n> > +\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n> >  \n> > -\t\ttest_cmp expect out &&\n> > -\t\ttest_line_count = 0 err &&\n> > +\t\t# Strip object disk usage from output due to platform variance.\n> > +\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n> >  \n> > -\t\t# Replace key and value delimiters for nul format.\n> > -\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n> > -\t\tgit repo structure --format=nul >out 2>err &&\n> > -\n> > -\t\ttest_cmp expect_nul out &&\n> > +\t\ttest_cmp expect out &&\n> >  \t\ttest_line_count = 0 err\n> >  \t)\n> >  '\n> \n> This part has both textual and semantic conflicts with Lucas's \"-z\n> is a synonym for --format=nul\" topic.  I think I resolved it\n> correctly while improving the \"munge expected output into expected\n> NUL-terminated output\" approach to \"munge -z output into textual\n> output and compare with textual expected output\".  Please sanity\n> check the result after I push it out, merged at 32f8d84b (Merge\n> branch 'jt/repo-struct-more-objinfo' into seen, 2025-12-10)\n\nThanks, this looks correct.\n\nJust FYI, some of the test changes I made here are reverted in the next\nversion since Patrick suggested a better way to test disk usage output.\nThis should allow Lucas's changes to apply a bit more cleanly to this\nfile.\n\nThanks,\n-Justin\n"},{"id":"532097","messageId":"20251212223644.3090879-1-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251209225820.2861276-1-jltobler@gmail.com","subject":"[PATCH v2 0/7] builtin/repo: add object size info to structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:37Z","receivedAt":"2025-12-12T22:36:49Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThis patch series extends the recently introduced \"structure\" subcommand\nfor git-repo(1) to collect object size information. More specifically,\nit shows total inflated and disk sizes of objects by object type. The\naim to provide additional insight that may be useful to users regarding\nthe structure of a repository.\n\nIn addition to this change, this series also updates the table output\nformat to downscale larger output values along with the appropriate unit\nprefix. This is done to make table output more human friendly. The\nkeyvalue and nul output formats are left the same since they are\nintended more for machine parsing.\n\nChanges in V2:\n- Factor out and reuse existing logic from strbuf_humanise() to handle\n  downscaling values and determining the appropriate unit prefix\n  separately. This enables more control over how exactly the values are\n  written to the structure output table which is useful for alignment\n  reasons. I'm not how about the interface used in patch 2. Feedback is\n  most welcome.\n- In the previous version, when checking object size on a missing object\n  we would die. Instead we now ignore missing objects. This allows the\n  structure command to work on partial clones.\n- disk/inflated keyvalue names renamed to disk_size/inflated_size.\n- Unit prefixes are marked for translation.\n- The test for keyvalue disk size values are updated to check against\n  real expected values instead of skipping. Table output tests still\n  skip verifing human-readable values though.\n\nThanks,\n-Justin\n\nJustin Tobler (7):\n  builtin/repo: group per-type object values into struct\n  strbuf: split out logic to humanise byte values\n  builtin/repo: humanise count values in structure output\n  builtin/repo: add inflated object info to keyvalue structure output\n  builtin/repo: add inflated object info to structure table\n  builtin/repo: add disk size info to keyvalue stucture output\n  builtin/repo: add object disk size info to structure table\n\n Documentation/git-repo.adoc |   2 +\n builtin/repo.c              | 185 ++++++++++++++++++++++++++++++------\n strbuf.c                    |  89 ++++++++++-------\n strbuf.h                    |  17 ++++\n t/t1901-repo-structure.sh   | 110 ++++++++++++++-------\n 5 files changed, 304 insertions(+), 99 deletions(-)\n\nRange-diff against v1:\n1:  bd3f1e6ec6 = 1:  be14de68f6 builtin/repo: group per-type object values into struct\n6:  bce4c7b5f1 ! 2:  5ca6f9b708 builtin/repo: add object disk size info to structure table\n    @@ Metadata\n     Author: Justin Tobler <jltobler@gmail.com>\n     \n      ## Commit message ##\n    -    builtin/repo: add object disk size info to structure table\n    +    strbuf: split out logic to humanise byte values\n     \n    -    Similar to a prior commit, update the table output format for the\n    -    git-repo(1) structure commdn to display the total object disk usage by\n    -    object type.\n    -\n    -    Since disk size may vary between platforms, tests do not validate actual\n    -    values and only check that size info is printed in an empty repository.\n    +    In a subsequent commit, byte size values displayed in table output for\n    +    the git-repo(1) \"structure\" subcommand will be shown in a more\n    +    human-readable format with the appropriate unit prefixes. For this\n    +    usecase, the downscaled values and unit prefixes must be handled\n    +    separately to ensure proper column alignment. Refactor strbuf_humanise()\n    +    to instead append the downscaled byte value to the buffer only and\n    +    return the appropriate unit prefix string.\n     \n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n    - ## builtin/repo.c ##\n    -@@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *table,\n    - \tstruct ref_stats *refs = &stats->refs;\n    - \tsize_t inflated_object_total;\n    - \tsize_t object_count_total;\n    -+\tsize_t disk_object_total;\n    - \tsize_t ref_total;\n    + ## strbuf.c ##\n    +@@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n    + \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n    + }\n      \n    - \tref_total = get_total_reference_count(refs);\n    -@@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *table,\n    - \t\t\t      \"    * %s\", _(\"Blobs\"));\n    - \tstats_table_size_addf(table, objects->inflated_sizes.tags,\n    - \t\t\t      \"    * %s\", _(\"Tags\"));\n    +-static void strbuf_humanise(struct strbuf *buf, off_t bytes,\n    +-\t\t\t\t int humanise_rate)\n    ++char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n    + {\n    ++\tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n     +\n    -+\tdisk_object_total = get_total_object_values(&objects->disk_sizes);\n    -+\tstats_table_size_addf(table, disk_object_total,\n    -+\t\t\t      \"  * %s\", _(\"Disk size\"));\n    -+\tstats_table_size_addf(table, objects->disk_sizes.commits,\n    -+\t\t\t      \"    * %s\", _(\"Commits\"));\n    -+\tstats_table_size_addf(table, objects->disk_sizes.trees,\n    -+\t\t\t      \"    * %s\", _(\"Trees\"));\n    -+\tstats_table_size_addf(table, objects->disk_sizes.blobs,\n    -+\t\t\t      \"    * %s\", _(\"Blobs\"));\n    -+\tstats_table_size_addf(table, objects->disk_sizes.tags,\n    -+\t\t\t      \"    * %s\", _(\"Tags\"));\n    + \tif (bytes > 1 << 30) {\n    +-\t\tstrbuf_addf(buf,\n    +-\t\t\t\thumanise_rate == 0 ?\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte */\n    +-\t\t\t\t\t_(\"%u.%2.2u GiB\") :\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second */\n    +-\t\t\t\t\t_(\"%u.%2.2u GiB/s\"),\n    +-\t\t\t    (unsigned)(bytes >> 30),\n    ++\t\tstrbuf_addf(buf, \"%u.%2.2u\", (unsigned)(bytes >> 30),\n    + \t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n    ++\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n    ++\t\treturn humanise_rate ? xstrfmt(_(\"GiB/s\")) : xstrfmt(_(\"GiB\"));\n    + \t} else if (bytes > 1 << 20) {\n    +-\t\tunsigned x = bytes + 5243;  /* for rounding */\n    +-\t\tstrbuf_addf(buf,\n    +-\t\t\t\thumanise_rate == 0 ?\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte */\n    +-\t\t\t\t\t_(\"%u.%2.2u MiB\") :\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second */\n    +-\t\t\t\t\t_(\"%u.%2.2u MiB/s\"),\n    +-\t\t\t    x >> 20, ((x & ((1 << 20) - 1)) * 100) >> 20);\n    ++\t\tunsigned x = bytes + 5243; /* for rounding */\n    ++\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 20,\n    ++\t\t\t    ((x & ((1 << 20) - 1)) * 100) >> 20);\n    ++\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second and mebibyte */\n    ++\t\treturn humanise_rate ? xstrfmt(_(\"MiB/s\")) : xstrfmt(_(\"MiB\"));\n    + \t} else if (bytes > 1 << 10) {\n    +-\t\tunsigned x = bytes + 5;  /* for rounding */\n    +-\t\tstrbuf_addf(buf,\n    +-\t\t\t\thumanise_rate == 0 ?\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte */\n    +-\t\t\t\t\t_(\"%u.%2.2u KiB\") :\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second */\n    +-\t\t\t\t\t_(\"%u.%2.2u KiB/s\"),\n    +-\t\t\t    x >> 10, ((x & ((1 << 10) - 1)) * 100) >> 10);\n    ++\t\tunsigned x = bytes + 5; /* for rounding */\n    ++\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 10,\n    ++\t\t\t    ((x & ((1 << 10) - 1)) * 100) >> 10);\n    ++\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second and kibibyte */\n    ++\t\treturn humanise_rate ? xstrfmt(_(\"KiB/s\")) : xstrfmt(_(\"KiB\"));\n    + \t} else {\n    +-\t\tstrbuf_addf(buf,\n    +-\t\t\t\thumanise_rate == 0 ?\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n    +-\t\t\t\t\tQ_(\"%u byte\", \"%u bytes\", bytes) :\n    +-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n    +-\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n    +-\t\t\t\t(unsigned)bytes);\n    ++\t\tstrbuf_addf(buf, \"%u\", (unsigned)bytes);\n    ++\t\treturn humanise_rate ?\n    ++\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n    ++\t\t\t       xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n    ++\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n    ++\t\t\t       xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n    + \t}\n      }\n      \n    - static void stats_table_print_structure(const struct stats_table *table)\n    -\n    - ## t/t1901-repo-structure.sh ##\n    -@@ t/t1901-repo-structure.sh: test_description='test git repo structure'\n    - \n    - . ./test-lib.sh\n    + void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n    + {\n    +-\tstrbuf_humanise(buf, bytes, 0);\n    ++\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, 0);\n    ++\tstrbuf_addf(buf, \" %s\", unit);\n    ++\tfree(unit);\n    + }\n      \n    -+strip_object_disk_usage() {\n    -+\tawk '\n    -+\t\t/^\\|   \\* Disk size/ { skip=1; next }\n    -+\t\tskip && /^\\|     \\* / { next }\n    -+\t\tskip && !/^\\|     \\* / { skip=0 }\n    -+\t\t{ print }\n    -+\t' $1\n    -+}\n    -+\n    - test_expect_success 'empty repository' '\n    - \ttest_when_finished \"rm -rf repo\" &&\n    - \tgit init repo &&\n    -@@ t/t1901-repo-structure.sh: test_expect_success 'empty repository' '\n    - \t\t|     * Trees          |    0 B |\n    - \t\t|     * Blobs          |    0 B |\n    - \t\t|     * Tags           |    0 B |\n    -+\t\t|   * Disk size        |    0 B |\n    -+\t\t|     * Commits        |    0 B |\n    -+\t\t|     * Trees          |    0 B |\n    -+\t\t|     * Blobs          |    0 B |\n    -+\t\t|     * Tags           |    0 B |\n    - \t\tEOF\n    + void strbuf_humanise_rate(struct strbuf *buf, off_t bytes)\n    + {\n    +-\tstrbuf_humanise(buf, bytes, 1);\n    ++\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, STRBUF_HUMANISE_RATE);\n    ++\tstrbuf_addf(buf, \" %s\", unit);\n    ++\tfree(unit);\n    + }\n      \n    - \t\tgit repo structure >out 2>err &&\n    -@@ t/t1901-repo-structure.sh: test_expect_success SHA1 'repository with references and objects' '\n    - \t\t|     * Tags           |    132 B   |\n    - \t\tEOF\n    + int printf_ln(const char *fmt, ...)\n    +\n    + ## strbuf.h ##\n    +@@ strbuf.h: void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n    +  */\n    + void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n      \n    --\t\tgit repo structure >out 2>err &&\n    -+\t\tgit repo structure >out.raw 2>err &&\n    ++#define STRBUF_HUMANISE_RATE 1 << 0\n     +\n    -+\t\t# Skip object disk sizes due to platform variance.\n    -+\t\tstrip_object_disk_usage out.raw >out &&\n    - \n    - \t\ttest_cmp expect out &&\n    - \t\ttest_line_count = 0 err\n    ++/**\n    ++ * Append the given byte size as a human-readable string that is downscaled by\n    ++ * some factor. A string with the corresponding unit prefix is returned\n    ++ * separately.\n    ++ */\n    ++char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags);\n    ++\n    + /**\n    +  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n    +  * 3.50 MiB).\n2:  3f56d52cd9 ! 3:  2efc3533ef builtin/repo: humanise count values in structure output\n    @@ builtin/repo.c: struct stats_table {\n       */\n      struct stats_table_entry {\n      \tchar *value;\n    -+\tconst char *unit;\n    ++\tchar *unit;\n      };\n      \n      static void stats_table_vaddf(struct stats_table *table,\n    @@ builtin/repo.c: static void stats_table_vaddf(struct stats_table *table,\n      }\n      \n      static void stats_table_addf(struct stats_table *table, const char *format, ...)\n    -@@ builtin/repo.c: static void stats_table_addf(struct stats_table *table, const char *format, ...)\n    - \tva_end(ap);\n    - }\n    - \n    -+static const char *unit_k = \"k\";\n    -+static const char *unit_M = \"M\";\n    -+static const char *unit_G = \"G\";\n    -+\n    - static void stats_table_count_addf(struct stats_table *table, size_t value,\n    +@@ builtin/repo.c: static void stats_table_count_addf(struct stats_table *table, size_t value,\n      \t\t\t\t   const char *format, ...)\n      {\n    -@@ builtin/repo.c: static void stats_table_count_addf(struct stats_table *table, size_t value,\n    + \tstruct stats_table_entry *entry;\n    ++\tstruct strbuf buf = STRBUF_INIT;\n      \tva_list ap;\n      \n      \tCALLOC_ARRAY(entry, 1);\n     -\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n     +\n    -+\tif (value >= 1000000000) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5000000;\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n    -+\t\t\t\t       x / 1000000000,\n    -+\t\t\t\t       x % 1000000000 / 10000000);\n    -+\t\tentry->unit = unit_G;\n    -+\t} else if (value >= 1000000) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5000;\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n    -+\t\t\t\t       x / 1000000, x % 1000000 / 10000);\n    -+\t\tentry->unit = unit_M;\n    -+\t} else if (value >= 1000) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5;\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX,\n    -+\t\t\t\t       x / 1000, x % 1000 / 10);\n    -+\t\tentry->unit = unit_k;\n    -+\t} else {\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n    -+\t}\n    ++\tentry->unit = strbuf_humanise_count_value(&buf, value);\n    ++\tentry->value = strbuf_detach(&buf, NULL);\n      \n      \tva_start(ap, format);\n      \tstats_table_vaddf(table, entry, format, ap);\n    @@ builtin/repo.c: static void stats_table_print_structure(const struct stats_table\n      \t\tstrbuf_addstr(&buf, \" |\");\n      \t\tprintf(\"%s\\n\", buf.buf);\n      \t}\n    +@@ builtin/repo.c: static void stats_table_clear(struct stats_table *table)\n    + \n    + \tfor_each_string_list_item(item, &table->rows) {\n    + \t\tentry = item->util;\n    +-\t\tif (entry)\n    ++\t\tif (entry) {\n    + \t\t\tfree(entry->value);\n    ++\t\t\tfree(entry->unit);\n    ++\t\t}\n    + \t}\n    + \n    + \tstring_list_clear(&table->rows, 1);\n    +\n    + ## strbuf.c ##\n    +@@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n    + \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n    + }\n    + \n    ++char *strbuf_humanise_count_value(struct strbuf *buf, size_t value)\n    ++{\n    ++\tif (value >= 1000000000) {\n    ++\t\tuintmax_t x = (uintmax_t)value + 5000000; /* for rounding */\n    ++\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n    ++\t\t\t    x / 1000000000, x % 1000000000 / 10000000);\n    ++\t\treturn xstrfmt(_(\"G\"));\n    ++\t} else if (value >= 1000000) {\n    ++\t\tuintmax_t x = (uintmax_t)value + 5000; /* for rounding */\n    ++\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n    ++\t\t\t    x / 1000000, x % 1000000 / 10000);\n    ++\t\treturn xstrfmt(_(\"M\"));\n    ++\t} else if (value >= 1000) {\n    ++\t\tuintmax_t x = (uintmax_t)value + 5; /* for rounding */\n    ++\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n    ++\t\t\t    x / 1000, x % 1000 / 10);\n    ++\t\treturn xstrfmt(_(\"k\"));\n    ++\t} else {\n    ++\t\tstrbuf_addf(buf, \"%\" PRIuMAX, (uintmax_t)value);\n    ++\t\treturn NULL;\n    ++\t}\n    ++}\n    ++\n    + char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n    + {\n    + \tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n    +\n    + ## strbuf.h ##\n    +@@ strbuf.h: void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n    +  */\n    + char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags);\n    + \n    ++/**\n    ++ * Append the given count value as a human-readable string that is downsacled by\n    ++ * some factor. A string with the corresponding unit prefix is returned\n    ++ * separately.\n    ++ */\n    ++char *strbuf_humanise_count_value(struct strbuf *buf, size_t value);\n    ++\n    + /**\n    +  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n    +  * 3.50 MiB).\n     \n      ## t/t1901-repo-structure.sh ##\n     @@ t/t1901-repo-structure.sh: test_expect_success 'empty repository' '\n3:  594bd320d1 ! 4:  627b8bf025 builtin/repo: add inflated object info to keyvalue structure output\n    @@ builtin/repo.c: static void structure_keyvalue_print(struct repo_structure *stat\n      \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n      \t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n      \n    -+\tprintf(\"objects.commits.inflated%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.commits.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n    -+\tprintf(\"objects.trees.inflated%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.trees.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n    -+\tprintf(\"objects.blobs.inflated%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.blobs.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n    -+\tprintf(\"objects.tags.inflated%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n     +\n      \tfflush(stdout);\n    @@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_arr\n     +\n     +\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n     +\t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n    -+\t\t\tdie(_(\"cannot read object for %s\"),\n    -+\t\t\t    oid_to_hex(&oids->oid[i]));\n    ++\t\t\tcontinue;\n     +\n     +\t\tinflated_total += inflated;\n     +\t}\n    @@ t/t1901-repo-structure.sh: test_expect_success 'keyvalue and nul format' '\n      \t\tobjects.trees.count=42\n      \t\tobjects.blobs.count=42\n      \t\tobjects.tags.count=1\n    -+\t\tobjects.commits.inflated=9225\n    -+\t\tobjects.trees.inflated=28554\n    -+\t\tobjects.blobs.inflated=453\n    -+\t\tobjects.tags.inflated=132\n    ++\t\tobjects.commits.inflated_size=9225\n    ++\t\tobjects.trees.inflated_size=28554\n    ++\t\tobjects.blobs.inflated_size=453\n    ++\t\tobjects.tags.inflated_size=132\n      \t\tEOF\n      \n      \t\tgit repo structure --format=keyvalue >out 2>err &&\n4:  3406b1ed90 ! 5:  14f4983e1d builtin/repo: add inflated object info to structure table\n    @@ builtin/repo.c: static void stats_table_count_addf(struct stats_table *table, si\n      \tva_end(ap);\n      }\n      \n    -+static const char *unit_B = \"B\";\n    -+static const char *unit_KiB = \"KiB\";\n    -+static const char *unit_MiB = \"MiB\";\n    -+static const char *unit_GiB = \"GiB\";\n    -+\n     +static void stats_table_size_addf(struct stats_table *table, size_t value,\n     +\t\t\t\t  const char *format, ...)\n     +{\n     +\tstruct stats_table_entry *entry;\n    ++\tstruct strbuf buf = STRBUF_INIT;\n     +\tva_list ap;\n     +\n     +\tCALLOC_ARRAY(entry, 1);\n     +\n    -+\tif (value > 1 << 30) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5368709;\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 30,\n    -+\t\t\t\t       ((x & ((1 << 30) - 1)) * 100) >> 30);\n    -+\t\tentry->unit = unit_GiB;\n    -+\t} else if (value > 1 << 20) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5243;\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 20,\n    -+\t\t\t\t       ((x & ((1 << 20) - 1)) * 100) >> 20);\n    -+\t\tentry->unit = unit_MiB;\n    -+\t} else if (value > 1 << 10) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5;\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX \".%02\" PRIuMAX, x >> 10,\n    -+\t\t\t\t       ((x & ((1 << 10) - 1)) * 100) >> 10);\n    -+\t\tentry->unit = unit_KiB;\n    -+\t} else {\n    -+\t\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n    -+\t\tentry->unit = unit_B;\n    -+\t}\n    ++\tentry->unit = strbuf_humanise_bytes_value(&buf, value,\n    ++\t\t\t\t\t\t  STRBUF_HUMANISE_COMPACT);\n    ++\tentry->value = strbuf_detach(&buf, NULL);\n     +\n     +\tva_start(ap, format);\n     +\tstats_table_vaddf(table, entry, format, ap);\n    @@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *tabl\n      \n      static void stats_table_print_structure(const struct stats_table *table)\n     \n    + ## strbuf.c ##\n    +@@ strbuf.c: char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flag\n    + \t\treturn humanise_rate ? xstrfmt(_(\"KiB/s\")) : xstrfmt(_(\"KiB\"));\n    + \t} else {\n    + \t\tstrbuf_addf(buf, \"%u\", (unsigned)bytes);\n    ++\t\tif (flags & STRBUF_HUMANISE_COMPACT)\n    ++\t\t\treturn humanise_rate ?\n    ++\t\t\t\t       xstrfmt(_(\"B/s\")) :\n    ++\t\t\t\t       xstrfmt(_(\"B\"));\n    + \t\treturn humanise_rate ?\n    + \t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n    + \t\t\t       xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n    +\n    + ## strbuf.h ##\n    +@@ strbuf.h: void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n    +  */\n    + void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n    + \n    +-#define STRBUF_HUMANISE_RATE 1 << 0\n    ++#define STRBUF_HUMANISE_RATE\t1 << 0\n    ++#define STRBUF_HUMANISE_COMPACT 1 << 1\n    + \n    + /**\n    +  * Append the given byte size as a human-readable string that is downscaled by\n    +\n      ## t/t1901-repo-structure.sh ##\n     @@ t/t1901-repo-structure.sh: test_expect_success 'empty repository' '\n      \t\t| Repository structure | Value  |\n5:  48461ac6a0 ! 6:  dc9e82889f builtin/repo: add disk size info to keyvalue stucture output\n    @@ Commit message\n         the git-repo(1) structure command to additionally provide info regarding\n         total object disk sizes by object type.\n     \n    -    Since disk size may vary between platforms, tests do not validate actual\n    -    values and only check that size info is printed in an empty repository.\n    -\n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n      ## Documentation/git-repo.adoc ##\n    @@ builtin/repo.c: struct object_values {\n      \n      struct repo_structure {\n     @@ builtin/repo.c: static void structure_keyvalue_print(struct repo_structure *stats,\n    - \tprintf(\"objects.tags.inflated%c%\" PRIuMAX \"%c\", key_delim,\n    + \tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n      \t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n      \n    -+\tprintf(\"objects.commits.disk%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.commits.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n    -+\tprintf(\"objects.trees.disk%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.trees.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n    -+\tprintf(\"objects.blobs.disk%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.blobs.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n    -+\tprintf(\"objects.tags.disk%c%\" PRIuMAX \"%c\", key_delim,\n    ++\tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n     +\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n     +\n      \tfflush(stdout);\n    @@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_arr\n      \n      \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n      \t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n    -@@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_array *oids,\n    - \t\t\t    oid_to_hex(&oids->oid[i]));\n    + \t\t\tcontinue;\n      \n      \t\tinflated_total += inflated;\n     +\t\tdisk_total += disk;\n    @@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_arr\n      \t\tBUG(\"invalid object type\");\n     \n      ## t/t1901-repo-structure.sh ##\n    -@@ t/t1901-repo-structure.sh: test_expect_success 'empty repository' '\n    - \t\tgit repo structure >out 2>err &&\n    +@@ t/t1901-repo-structure.sh: test_description='test git repo structure'\n      \n    - \t\ttest_cmp expect out &&\n    -+\t\ttest_line_count = 0 err &&\n    -+\n    -+\t\tcat >expect <<-\\EOF &&\n    -+\t\treferences.branches.count=0\n    -+\t\treferences.tags.count=0\n    -+\t\treferences.remotes.count=0\n    -+\t\treferences.others.count=0\n    -+\t\tobjects.commits.count=0\n    -+\t\tobjects.trees.count=0\n    -+\t\tobjects.blobs.count=0\n    -+\t\tobjects.tags.count=0\n    -+\t\tobjects.commits.inflated=0\n    -+\t\tobjects.trees.inflated=0\n    -+\t\tobjects.blobs.inflated=0\n    -+\t\tobjects.tags.inflated=0\n    -+\t\tobjects.commits.disk=0\n    -+\t\tobjects.trees.disk=0\n    -+\t\tobjects.blobs.disk=0\n    -+\t\tobjects.tags.disk=0\n    -+\t\tEOF\n    -+\n    -+\t\tgit repo structure --format=keyvalue >out 2>err &&\n    -+\n    -+\t\ttest_cmp expect out &&\n    -+\t\ttest_line_count = 0 err &&\n    -+\n    -+\t\t# Replace key and value delimiters for nul format.\n    -+\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n    -+\t\tgit repo structure --format=nul >out 2>err &&\n    -+\n    -+\t\ttest_cmp expect_nul out &&\n    - \t\ttest_line_count = 0 err\n    - \t)\n    - '\n    -@@ t/t1901-repo-structure.sh: test_expect_success SHA1 'repository with references and objects' '\n    - \t)\n    - '\n    + . ./test-lib.sh\n      \n    --test_expect_success SHA1 'keyvalue and nul format' '\n    -+test_expect_success SHA1 'keyvalue format' '\n    ++object_type_disk_usage() {\n    ++\tgit cat-file --batch-check='%(objectsize:disk)' --batch-all-objects \\\n    ++\t\t--filter=object:type=$1 | awk '{ sum += $1 } END { print sum }'\n    ++}\n    ++\n    + test_expect_success 'empty repository' '\n      \ttest_when_finished \"rm -rf repo\" &&\n      \tgit init repo &&\n    - \t(\n     @@ t/t1901-repo-structure.sh: test_expect_success SHA1 'keyvalue and nul format' '\n    - \t\tobjects.tags.inflated=132\n    - \t\tEOF\n    + \t\ttest_commit_bulk 42 &&\n    + \t\tgit tag -a foo -m bar &&\n      \n    --\t\tgit repo structure --format=keyvalue >out 2>err &&\n    -+\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n    - \n    --\t\ttest_cmp expect out &&\n    --\t\ttest_line_count = 0 err &&\n    -+\t\t# Strip object disk usage from output due to platform variance.\n    -+\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n    +-\t\tcat >expect <<-\\EOF &&\n    ++\t\tcat >expect <<-EOF &&\n    + \t\treferences.branches.count=1\n    + \t\treferences.tags.count=1\n    + \t\treferences.remotes.count=0\n    +@@ t/t1901-repo-structure.sh: test_expect_success SHA1 'keyvalue and nul format' '\n    + \t\tobjects.trees.inflated_size=28554\n    + \t\tobjects.blobs.inflated_size=453\n    + \t\tobjects.tags.inflated_size=132\n    ++\t\tobjects.commits.disk_size=$(object_type_disk_usage commit)\n    ++\t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n    ++\t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n    ++\t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n    + \t\tEOF\n      \n    --\t\t# Replace key and value delimiters for nul format.\n    --\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n    --\t\tgit repo structure --format=nul >out 2>err &&\n    --\n    --\t\ttest_cmp expect_nul out &&\n    -+\t\ttest_cmp expect out &&\n    - \t\ttest_line_count = 0 err\n    - \t)\n    - '\n    + \t\tgit repo structure --format=keyvalue >out 2>err &&\n-:  ---------- > 7:  213b19dc7f builtin/repo: add object disk size info to structure table\n\nbase-commit: e85ae279b0d58edc2f4c3fd5ac391b51e1223985\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532098","messageId":"20251212223644.3090879-2-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v2 1/7] builtin/repo: group per-type object values into struct","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:38Z","receivedAt":"2025-12-12T22:36:50Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The `object_stats` structure stores object counts by type. In a\nsubsequent commit, additional per-type object measurements will also be\nstored. Group per-type object values into a new struct to allow better\nreuse.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 42 +++++++++++++++++++++++++-----------------\n 1 file changed, 25 insertions(+), 17 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 2a653bd3ea..a69699857a 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -202,13 +202,17 @@ struct ref_stats {\n \tsize_t others;\n };\n \n-struct object_stats {\n+struct object_values {\n \tsize_t tags;\n \tsize_t commits;\n \tsize_t trees;\n \tsize_t blobs;\n };\n \n+struct object_stats {\n+\tstruct object_values type_counts;\n+};\n+\n struct repo_structure {\n \tstruct ref_stats refs;\n \tstruct object_stats objects;\n@@ -281,9 +285,9 @@ static inline size_t get_total_reference_count(struct ref_stats *stats)\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n }\n \n-static inline size_t get_total_object_count(struct object_stats *stats)\n+static inline size_t get_total_object_values(struct object_values *values)\n {\n-\treturn stats->tags + stats->commits + stats->trees + stats->blobs;\n+\treturn values->tags + values->commits + values->trees + values->blobs;\n }\n \n static void stats_table_setup_structure(struct stats_table *table,\n@@ -302,14 +306,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_count(objects);\n+\tobject_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n \tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n-\tstats_table_count_addf(table, objects->commits, \"    * %s\", _(\"Commits\"));\n-\tstats_table_count_addf(table, objects->trees, \"    * %s\", _(\"Trees\"));\n-\tstats_table_count_addf(table, objects->blobs, \"    * %s\", _(\"Blobs\"));\n-\tstats_table_count_addf(table, objects->tags, \"    * %s\", _(\"Tags\"));\n+\tstats_table_count_addf(table, objects->type_counts.commits,\n+\t\t\t       \"    * %s\", _(\"Commits\"));\n+\tstats_table_count_addf(table, objects->type_counts.trees,\n+\t\t\t       \"    * %s\", _(\"Trees\"));\n+\tstats_table_count_addf(table, objects->type_counts.blobs,\n+\t\t\t       \"    * %s\", _(\"Blobs\"));\n+\tstats_table_count_addf(table, objects->type_counts.tags,\n+\t\t\t       \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -389,13 +397,13 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \t       (uintmax_t)stats->refs.others, value_delim);\n \n \tprintf(\"objects.commits.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.commits, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.commits, value_delim);\n \tprintf(\"objects.trees.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.trees, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.trees, value_delim);\n \tprintf(\"objects.blobs.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.blobs, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.blobs, value_delim);\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.tags, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n \tfflush(stdout);\n }\n@@ -473,22 +481,22 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n-\t\tstats->tags += oids->nr;\n+\t\tstats->type_counts.tags += oids->nr;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n-\t\tstats->commits += oids->nr;\n+\t\tstats->type_counts.commits += oids->nr;\n \t\tbreak;\n \tcase OBJ_TREE:\n-\t\tstats->trees += oids->nr;\n+\t\tstats->type_counts.trees += oids->nr;\n \t\tbreak;\n \tcase OBJ_BLOB:\n-\t\tstats->blobs += oids->nr;\n+\t\tstats->type_counts.blobs += oids->nr;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n \t}\n \n-\tobject_count = get_total_object_count(stats);\n+\tobject_count = get_total_object_values(&stats->type_counts);\n \tdisplay_progress(data->progress, object_count);\n \n \treturn 0;\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532099","messageId":"20251212223644.3090879-3-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:39Z","receivedAt":"2025-12-12T22:36:51Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"In a subsequent commit, byte size values displayed in table output for\nthe git-repo(1) \"structure\" subcommand will be shown in a more\nhuman-readable format with the appropriate unit prefixes. For this\nusecase, the downscaled values and unit prefixes must be handled\nseparately to ensure proper column alignment. Refactor strbuf_humanise()\nto instead append the downscaled byte value to the buffer only and\nreturn the appropriate unit prefix string.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n strbuf.c | 62 +++++++++++++++++++++++++-------------------------------\n strbuf.h |  9 ++++++++\n 2 files changed, 37 insertions(+), 34 deletions(-)\n\ndiff --git a/strbuf.c b/strbuf.c\nindex 6c3851a7f8..1fb47bf21b 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,55 +836,49 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n-static void strbuf_humanise(struct strbuf *buf, off_t bytes,\n-\t\t\t\t int humanise_rate)\n+char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n {\n+\tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n+\n \tif (bytes > 1 << 30) {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u GiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u GiB/s\"),\n-\t\t\t    (unsigned)(bytes >> 30),\n+\t\tstrbuf_addf(buf, \"%u.%2.2u\", (unsigned)(bytes >> 30),\n \t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n+\t\treturn humanise_rate ? xstrfmt(_(\"GiB/s\")) : xstrfmt(_(\"GiB\"));\n \t} else if (bytes > 1 << 20) {\n-\t\tunsigned x = bytes + 5243;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte */\n-\t\t\t\t\t_(\"%u.%2.2u MiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u MiB/s\"),\n-\t\t\t    x >> 20, ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\tunsigned x = bytes + 5243; /* for rounding */\n+\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 20,\n+\t\t\t    ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second and mebibyte */\n+\t\treturn humanise_rate ? xstrfmt(_(\"MiB/s\")) : xstrfmt(_(\"MiB\"));\n \t} else if (bytes > 1 << 10) {\n-\t\tunsigned x = bytes + 5;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u KiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u KiB/s\"),\n-\t\t\t    x >> 10, ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\tunsigned x = bytes + 5; /* for rounding */\n+\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 10,\n+\t\t\t    ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second and kibibyte */\n+\t\treturn humanise_rate ? xstrfmt(_(\"KiB/s\")) : xstrfmt(_(\"KiB\"));\n \t} else {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n-\t\t\t\t\tQ_(\"%u byte\", \"%u bytes\", bytes) :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n-\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n-\t\t\t\t(unsigned)bytes);\n+\t\tstrbuf_addf(buf, \"%u\", (unsigned)bytes);\n+\t\treturn humanise_rate ?\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n+\t\t\t       xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n+\t\t\t       xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n \t}\n }\n \n void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n {\n-\tstrbuf_humanise(buf, bytes, 0);\n+\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, 0);\n+\tstrbuf_addf(buf, \" %s\", unit);\n+\tfree(unit);\n }\n \n void strbuf_humanise_rate(struct strbuf *buf, off_t bytes)\n {\n-\tstrbuf_humanise(buf, bytes, 1);\n+\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, STRBUF_HUMANISE_RATE);\n+\tstrbuf_addf(buf, \" %s\", unit);\n+\tfree(unit);\n }\n \n int printf_ln(const char *fmt, ...)\ndiff --git a/strbuf.h b/strbuf.h\nindex a580ac6084..a5e3ab0cb4 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -367,6 +367,15 @@ void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n  */\n void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n \n+#define STRBUF_HUMANISE_RATE 1 << 0\n+\n+/**\n+ * Append the given byte size as a human-readable string that is downscaled by\n+ * some factor. A string with the corresponding unit prefix is returned\n+ * separately.\n+ */\n+char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532100","messageId":"20251212223644.3090879-4-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v2 3/7] builtin/repo: humanise count values in structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:40Z","receivedAt":"2025-12-12T22:36:52Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The table output format for the git-repo(1) structure subcommand is used\nby default and intended to provide output to users in a human-friendly\nmanner. When the reference/object count values in a repository are\nlarge, it becomes more cumbersome for users to read the values.\n\nFor larger values, update the table output format to instead produce\nmore human-friendly count values that are scaled down with the\nappropriate unit prefix. Output for the keyvalue and nul formats remains\nunchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 45 +++++++++++++++++++++-------\n strbuf.c                  | 23 +++++++++++++++\n strbuf.h                  |  7 +++++\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++--------------------\n 4 files changed, 95 insertions(+), 42 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex a69699857a..d3dfe416d0 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -223,6 +223,7 @@ struct stats_table {\n \n \tint name_col_width;\n \tint value_col_width;\n+\tint unit_col_width;\n };\n \n /*\n@@ -230,6 +231,7 @@ struct stats_table {\n  */\n struct stats_table_entry {\n \tchar *value;\n+\tchar *unit;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -250,11 +252,18 @@ static void stats_table_vaddf(struct stats_table *table,\n \n \tif (name_width > table->name_col_width)\n \t\ttable->name_col_width = name_width;\n-\tif (entry) {\n+\tif (!entry)\n+\t\treturn;\n+\tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n \t\t\ttable->value_col_width = value_width;\n \t}\n+\tif (entry->unit) {\n+\t\tint unit_width = utf8_strwidth(entry->unit);\n+\t\tif (unit_width > table->unit_col_width)\n+\t\t\ttable->unit_col_width = unit_width;\n+\t}\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -270,10 +279,13 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \t\t\t\t   const char *format, ...)\n {\n \tstruct stats_table_entry *entry;\n+\tstruct strbuf buf = STRBUF_INIT;\n \tva_list ap;\n \n \tCALLOC_ARRAY(entry, 1);\n-\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n+\n+\tentry->unit = strbuf_humanise_count_value(&buf, value);\n+\tentry->value = strbuf_detach(&buf, NULL);\n \n \tva_start(ap, format);\n \tstats_table_vaddf(table, entry, format, ap);\n@@ -324,20 +336,24 @@ static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n \tconst char *value_col_title = _(\"Value\");\n-\tint name_col_width = utf8_strwidth(name_col_title);\n-\tint value_col_width = utf8_strwidth(value_col_title);\n+\tint title_name_width = utf8_strwidth(name_col_title);\n+\tint title_value_width = utf8_strwidth(value_col_title);\n+\tint name_col_width = table->name_col_width;\n+\tint value_col_width = table->value_col_width;\n+\tint unit_col_width = table->unit_col_width;\n \tstruct string_list_item *item;\n \tstruct strbuf buf = STRBUF_INIT;\n \n-\tif (table->name_col_width > name_col_width)\n-\t\tname_col_width = table->name_col_width;\n-\tif (table->value_col_width > value_col_width)\n-\t\tvalue_col_width = table->value_col_width;\n+\tif (title_name_width > name_col_width)\n+\t\tname_col_width = title_name_width;\n+\tif (title_value_width > value_col_width + unit_col_width + 1)\n+\t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n+\t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n \tstrbuf_addstr(&buf, \" |\");\n \tprintf(\"%s\\n\", buf.buf);\n \n@@ -345,17 +361,20 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tfor (int i = 0; i < name_col_width; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n-\tfor (int i = 0; i < value_col_width; i++)\n+\tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n \t\tputchar('-');\n \tprintf(\" |\\n\");\n \n \tfor_each_string_list_item(item, &table->rows) {\n \t\tstruct stats_table_entry *entry = item->util;\n \t\tconst char *value = \"\";\n+\t\tconst char *unit = \"\";\n \n \t\tif (entry) {\n \t\t\tstruct stats_table_entry *entry = item->util;\n \t\t\tvalue = entry->value;\n+\t\t\tif (entry->unit)\n+\t\t\t\tunit = entry->unit;\n \t\t}\n \n \t\tstrbuf_reset(&buf);\n@@ -363,6 +382,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n+\t\tstrbuf_addch(&buf, ' ');\n+\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, unit_col_width, unit);\n \t\tstrbuf_addstr(&buf, \" |\");\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\n@@ -377,8 +398,10 @@ static void stats_table_clear(struct stats_table *table)\n \n \tfor_each_string_list_item(item, &table->rows) {\n \t\tentry = item->util;\n-\t\tif (entry)\n+\t\tif (entry) {\n \t\t\tfree(entry->value);\n+\t\t\tfree(entry->unit);\n+\t\t}\n \t}\n \n \tstring_list_clear(&table->rows, 1);\ndiff --git a/strbuf.c b/strbuf.c\nindex 1fb47bf21b..cebb1593ab 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,6 +836,29 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n+char *strbuf_humanise_count_value(struct strbuf *buf, size_t value)\n+{\n+\tif (value >= 1000000000) {\n+\t\tuintmax_t x = (uintmax_t)value + 5000000; /* for rounding */\n+\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n+\t\t\t    x / 1000000000, x % 1000000000 / 10000000);\n+\t\treturn xstrfmt(_(\"G\"));\n+\t} else if (value >= 1000000) {\n+\t\tuintmax_t x = (uintmax_t)value + 5000; /* for rounding */\n+\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n+\t\t\t    x / 1000000, x % 1000000 / 10000);\n+\t\treturn xstrfmt(_(\"M\"));\n+\t} else if (value >= 1000) {\n+\t\tuintmax_t x = (uintmax_t)value + 5; /* for rounding */\n+\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n+\t\t\t    x / 1000, x % 1000 / 10);\n+\t\treturn xstrfmt(_(\"k\"));\n+\t} else {\n+\t\tstrbuf_addf(buf, \"%\" PRIuMAX, (uintmax_t)value);\n+\t\treturn NULL;\n+\t}\n+}\n+\n char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n {\n \tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\ndiff --git a/strbuf.h b/strbuf.h\nindex a5e3ab0cb4..7532eadd02 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -376,6 +376,13 @@ void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n  */\n char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags);\n \n+/**\n+ * Append the given count value as a human-readable string that is downsacled by\n+ * some factor. A string with the corresponding unit prefix is returned\n+ * separately.\n+ */\n+char *strbuf_humanise_count_value(struct strbuf *buf, size_t value);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 36a71a144e..55fd13ad1b 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -10,21 +10,21 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Branches       |     0 |\n-\t\t|     * Tags           |     0 |\n-\t\t|     * Remotes        |     0 |\n-\t\t|     * Others         |     0 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Commits        |     0 |\n-\t\t|     * Trees          |     0 |\n-\t\t|     * Blobs          |     0 |\n-\t\t|     * Tags           |     0 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Branches       |     0  |\n+\t\t|     * Tags           |     0  |\n+\t\t|     * Remotes        |     0  |\n+\t\t|     * Others         |     0  |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Commits        |     0  |\n+\t\t|     * Trees          |     0  |\n+\t\t|     * Blobs          |     0  |\n+\t\t|     * Tags           |     0  |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -39,7 +39,7 @@ test_expect_success 'repository with references and objects' '\n \tgit init repo &&\n \t(\n \t\tcd repo &&\n-\t\ttest_commit_bulk 42 &&\n+\t\ttest_commit_bulk 1005 &&\n \t\tgit tag -a foo -m bar &&\n \n \t\toid=\"$(git rev-parse HEAD)\" &&\n@@ -49,21 +49,21 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     4 |\n-\t\t|     * Branches       |     1 |\n-\t\t|     * Tags           |     1 |\n-\t\t|     * Remotes        |     1 |\n-\t\t|     * Others         |     1 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |   130 |\n-\t\t|     * Commits        |    43 |\n-\t\t|     * Trees          |    43 |\n-\t\t|     * Blobs          |    43 |\n-\t\t|     * Tags           |     1 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |    4   |\n+\t\t|     * Branches       |    1   |\n+\t\t|     * Tags           |    1   |\n+\t\t|     * Remotes        |    1   |\n+\t\t|     * Others         |    1   |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            | 3.02 k |\n+\t\t|     * Commits        | 1.01 k |\n+\t\t|     * Trees          | 1.01 k |\n+\t\t|     * Blobs          | 1.01 k |\n+\t\t|     * Tags           |    1   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532101","messageId":"20251212223644.3090879-5-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v2 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:41Z","receivedAt":"2025-12-12T22:36:53Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The structure subcommand for git-repo(1) outputs basic count information\nfor objects and references. Extend this output to also provide\ninformation regarding total size of inflated objects by object type.\n\nFor now, object size by object type info is only added to the keyvalue\nand nul output formats. In a subsequent commit, this info is also added\nto the table format.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 32 ++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   |  6 +++++-\n 3 files changed, 38 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 70f0a6d2e4..287eee4b93 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -50,6 +50,7 @@ supported:\n +\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n+* Total inflated size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex d3dfe416d0..3a2d15cec4 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,8 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hex.h\"\n+#include \"odb.h\"\n #include \"parse-options.h\"\n #include \"path-walk.h\"\n #include \"progress.h\"\n@@ -211,6 +213,7 @@ struct object_values {\n \n struct object_stats {\n \tstruct object_values type_counts;\n+\tstruct object_values inflated_sizes;\n };\n \n struct repo_structure {\n@@ -428,6 +431,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n+\tprintf(\"objects.commits.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -491,6 +503,7 @@ static void structure_count_references(struct ref_stats *stats,\n }\n \n struct count_objects_data {\n+\tstruct object_database *odb;\n \tstruct object_stats *stats;\n \tstruct progress *progress;\n };\n@@ -500,20 +513,38 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n+\tsize_t inflated_total = 0;\n \tsize_t object_count;\n \n+\tfor (size_t i = 0; i < oids->nr; i++) {\n+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n+\t\tunsigned long inflated;\n+\n+\t\toi.sizep = &inflated;\n+\n+\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n+\t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n+\t\t\tcontinue;\n+\n+\t\tinflated_total += inflated;\n+\t}\n+\n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n+\t\tstats->inflated_sizes.tags += inflated_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n+\t\tstats->inflated_sizes.commits += inflated_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n+\t\tstats->inflated_sizes.trees += inflated_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n+\t\tstats->inflated_sizes.blobs += inflated_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n@@ -531,6 +562,7 @@ static void structure_count_objects(struct object_stats *stats,\n {\n \tstruct path_walk_info info = PATH_WALK_INFO_INIT;\n \tstruct count_objects_data data = {\n+\t\t.odb = repo->objects,\n \t\t.stats = stats,\n \t};\n \ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 55fd13ad1b..33237822fd 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -73,7 +73,7 @@ test_expect_success 'repository with references and objects' '\n \t)\n '\n \n-test_expect_success 'keyvalue and nul format' '\n+test_expect_success SHA1 'keyvalue and nul format' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -90,6 +90,10 @@ test_expect_success 'keyvalue and nul format' '\n \t\tobjects.trees.count=42\n \t\tobjects.blobs.count=42\n \t\tobjects.tags.count=1\n+\t\tobjects.commits.inflated_size=9225\n+\t\tobjects.trees.inflated_size=28554\n+\t\tobjects.blobs.inflated_size=453\n+\t\tobjects.tags.inflated_size=132\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532102","messageId":"20251212223644.3090879-6-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v2 5/7] builtin/repo: add inflated object info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:42Z","receivedAt":"2025-12-12T22:36:54Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Update the table output format for the git-repo(1) structure command to\nbegin printing the total inflated object size info by object type. To be\nmore human-friendly, larger values are scaled down and displayed with\nthe appropriate unit prefix. Output for the keyvalue and nul formats\nremains unchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 37 +++++++++++++++++++++--\n strbuf.c                  |  4 +++\n strbuf.h                  |  3 +-\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++++++----------------\n 4 files changed, 76 insertions(+), 30 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 3a2d15cec4..b0609cfae5 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -295,6 +295,24 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_size_addf(struct stats_table *table, size_t value,\n+\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\n+\tentry->unit = strbuf_humanise_bytes_value(&buf, value,\n+\t\t\t\t\t\t  STRBUF_HUMANISE_COMPACT);\n+\tentry->value = strbuf_detach(&buf, NULL);\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -310,7 +328,8 @@ static void stats_table_setup_structure(struct stats_table *table,\n {\n \tstruct object_stats *objects = &stats->objects;\n \tstruct ref_stats *refs = &stats->refs;\n-\tsize_t object_total;\n+\tsize_t inflated_object_total;\n+\tsize_t object_count_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -321,10 +340,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_values(&objects->type_counts);\n+\tobject_count_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n-\tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n+\tstats_table_count_addf(table, object_count_total, \"  * %s\", _(\"Count\"));\n \tstats_table_count_addf(table, objects->type_counts.commits,\n \t\t\t       \"    * %s\", _(\"Commits\"));\n \tstats_table_count_addf(table, objects->type_counts.trees,\n@@ -333,6 +352,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t       \"    * %s\", _(\"Blobs\"));\n \tstats_table_count_addf(table, objects->type_counts.tags,\n \t\t\t       \"    * %s\", _(\"Tags\"));\n+\n+\tinflated_object_total = get_total_object_values(&objects->inflated_sizes);\n+\tstats_table_size_addf(table, inflated_object_total,\n+\t\t\t      \"  * %s\", _(\"Inflated size\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/strbuf.c b/strbuf.c\nindex cebb1593ab..eed4e167ca 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -882,6 +882,10 @@ char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flag\n \t\treturn humanise_rate ? xstrfmt(_(\"KiB/s\")) : xstrfmt(_(\"KiB\"));\n \t} else {\n \t\tstrbuf_addf(buf, \"%u\", (unsigned)bytes);\n+\t\tif (flags & STRBUF_HUMANISE_COMPACT)\n+\t\t\treturn humanise_rate ?\n+\t\t\t\t       xstrfmt(_(\"B/s\")) :\n+\t\t\t\t       xstrfmt(_(\"B\"));\n \t\treturn humanise_rate ?\n \t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n \t\t\t       xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\ndiff --git a/strbuf.h b/strbuf.h\nindex 7532eadd02..919527d26b 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -367,7 +367,8 @@ void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n  */\n void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n \n-#define STRBUF_HUMANISE_RATE 1 << 0\n+#define STRBUF_HUMANISE_RATE\t1 << 0\n+#define STRBUF_HUMANISE_COMPACT 1 << 1\n \n /**\n  * Append the given byte size as a human-readable string that is downscaled by\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 33237822fd..b18213c660 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -13,18 +13,23 @@ test_expect_success 'empty repository' '\n \t\t| Repository structure | Value  |\n \t\t| -------------------- | ------ |\n \t\t| * References         |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Branches       |     0  |\n-\t\t|     * Tags           |     0  |\n-\t\t|     * Remotes        |     0  |\n-\t\t|     * Others         |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Branches       |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|     * Remotes        |    0   |\n+\t\t|     * Others         |    0   |\n \t\t|                      |        |\n \t\t| * Reachable objects  |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Commits        |     0  |\n-\t\t|     * Trees          |     0  |\n-\t\t|     * Blobs          |     0  |\n-\t\t|     * Tags           |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Commits        |    0   |\n+\t\t|     * Trees          |    0   |\n+\t\t|     * Blobs          |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|   * Inflated size    |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -34,7 +39,7 @@ test_expect_success 'empty repository' '\n \t)\n '\n \n-test_expect_success 'repository with references and objects' '\n+test_expect_success SHA1 'repository with references and objects' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -49,21 +54,26 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    4   |\n-\t\t|     * Branches       |    1   |\n-\t\t|     * Tags           |    1   |\n-\t\t|     * Remotes        |    1   |\n-\t\t|     * Others         |    1   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            | 3.02 k |\n-\t\t|     * Commits        | 1.01 k |\n-\t\t|     * Trees          | 1.01 k |\n-\t\t|     * Blobs          | 1.01 k |\n-\t\t|     * Tags           |    1   |\n+\t\t| Repository structure | Value      |\n+\t\t| -------------------- | ---------- |\n+\t\t| * References         |            |\n+\t\t|   * Count            |      4     |\n+\t\t|     * Branches       |      1     |\n+\t\t|     * Tags           |      1     |\n+\t\t|     * Remotes        |      1     |\n+\t\t|     * Others         |      1     |\n+\t\t|                      |            |\n+\t\t| * Reachable objects  |            |\n+\t\t|   * Count            |   3.02 k   |\n+\t\t|     * Commits        |   1.01 k   |\n+\t\t|     * Trees          |   1.01 k   |\n+\t\t|     * Blobs          |   1.01 k   |\n+\t\t|     * Tags           |      1     |\n+\t\t|   * Inflated size    |  16.03 MiB |\n+\t\t|     * Commits        | 217.92 KiB |\n+\t\t|     * Trees          |  15.81 MiB |\n+\t\t|     * Blobs          |  11.68 KiB |\n+\t\t|     * Tags           |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532103","messageId":"20251212223644.3090879-7-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v2 6/7] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:43Z","receivedAt":"2025-12-12T22:36:54Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, extend the keyvalue and nul output formats of\nthe git-repo(1) structure command to additionally provide info regarding\ntotal object disk sizes by object type.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 18 ++++++++++++++++++\n t/t1901-repo-structure.sh   | 11 ++++++++++-\n 3 files changed, 29 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 287eee4b93..861073f641 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -51,6 +51,7 @@ supported:\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n+* Total disk size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex b0609cfae5..252a53f452 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -214,6 +214,7 @@ struct object_values {\n struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n+\tstruct object_values disk_sizes;\n };\n \n struct repo_structure {\n@@ -471,6 +472,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n \n+\tprintf(\"objects.commits.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -545,37 +555,45 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n \tsize_t inflated_total = 0;\n+\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\toff_t disk;\n \n \t\toi.sizep = &inflated;\n+\t\toi.disk_sizep = &disk;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n \t\t\tcontinue;\n \n \t\tinflated_total += inflated;\n+\t\tdisk_total += disk;\n \t}\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n \t\tstats->inflated_sizes.tags += inflated_total;\n+\t\tstats->disk_sizes.tags += disk_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n \t\tstats->inflated_sizes.commits += inflated_total;\n+\t\tstats->disk_sizes.commits += disk_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n \t\tstats->inflated_sizes.trees += inflated_total;\n+\t\tstats->disk_sizes.trees += disk_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n \t\tstats->inflated_sizes.blobs += inflated_total;\n+\t\tstats->disk_sizes.blobs += disk_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex b18213c660..1553f3cd32 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -4,6 +4,11 @@ test_description='test git repo structure'\n \n . ./test-lib.sh\n \n+object_type_disk_usage() {\n+\tgit cat-file --batch-check='%(objectsize:disk)' --batch-all-objects \\\n+\t\t--filter=object:type=$1 | awk '{ sum += $1 } END { print sum }'\n+}\n+\n test_expect_success 'empty repository' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n@@ -91,7 +96,7 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\ttest_commit_bulk 42 &&\n \t\tgit tag -a foo -m bar &&\n \n-\t\tcat >expect <<-\\EOF &&\n+\t\tcat >expect <<-EOF &&\n \t\treferences.branches.count=1\n \t\treferences.tags.count=1\n \t\treferences.remotes.count=0\n@@ -104,6 +109,10 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.trees.inflated_size=28554\n \t\tobjects.blobs.inflated_size=453\n \t\tobjects.tags.inflated_size=132\n+\t\tobjects.commits.disk_size=$(object_type_disk_usage commit)\n+\t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n+\t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n+\t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532104","messageId":"20251212223644.3090879-8-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v2 7/7] builtin/repo: add object disk size info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-12T22:36:44Z","receivedAt":"2025-12-12T22:36:55Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, update the table output format for the\ngit-repo(1) structure command to display the total object disk usage by\nobject type.\n\nSince disk size may vary between platforms, tests do not validate actual\nvalues and only check that size info is printed in an empty repository.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 13 +++++++++++++\n t/t1901-repo-structure.sh | 19 ++++++++++++++++++-\n 2 files changed, 31 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 252a53f452..c294fa11d2 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -331,6 +331,7 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstruct ref_stats *refs = &stats->refs;\n \tsize_t inflated_object_total;\n \tsize_t object_count_total;\n+\tsize_t disk_object_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -365,6 +366,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->inflated_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tdisk_object_total = get_total_object_values(&objects->disk_sizes);\n+\tstats_table_size_addf(table, disk_object_total,\n+\t\t\t      \"  * %s\", _(\"Disk size\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 1553f3cd32..6a992222df 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -9,6 +9,15 @@ object_type_disk_usage() {\n \t\t--filter=object:type=$1 | awk '{ sum += $1 } END { print sum }'\n }\n \n+strip_object_disk_usage() {\n+\tawk '\n+\t\t/^\\|   \\* Disk size/ { skip=1; next }\n+\t\tskip && /^\\|     \\* / { next }\n+\t\tskip && !/^\\|     \\* / { skip=0 }\n+\t\t{ print }\n+\t' $1\n+}\n+\n test_expect_success 'empty repository' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n@@ -35,6 +44,11 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|   * Disk size        |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -81,7 +95,10 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Tags           |    132 B   |\n \t\tEOF\n \n-\t\tgit repo structure >out 2>err &&\n+\t\tgit repo structure >out.raw 2>err &&\n+\n+\t\t# Skip object disk sizes due to platform variance.\n+\t\tstrip_object_disk_usage out.raw >out &&\n \n \t\ttest_cmp expect out &&\n \t\ttest_line_count = 0 err\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532107","messageId":"xmqqa4znb6cq.fsf@gitster.g","threadId":"64606","inReplyTo":"e5hsuevw5t37yt3zgp4hhtunusdyeg2lkph52pj4valpmlyrdt@7teicd67atbj","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-12T23:58:13Z","receivedAt":"2025-12-12T23:58:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> Just FYI, some of the test changes I made here are reverted in the next\n> version since Patrick suggested a better way to test disk usage output.\n> This should allow Lucas's changes to apply a bit more cleanly to this\n> file.\n\nGood.  I expect that Lucas's series would also be updated,\nespecially in the way the nul-delimited output is tested, so we'll\nsee what happens ;-).\n"},{"id":"532160","messageId":"aT-djS-TrQJxxV8i@pks.im","threadId":"64606","inReplyTo":"54kuvik2ecbkygjp57osmqjxiy7xtyjeffbzavuxbhuvta2oc5@mkqufah7cb3z","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-15T05:33:01Z","receivedAt":"2025-12-15T05:33:08Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Dec 12, 2025 at 02:40:24PM -0600, Justin Tobler wrote:\n> On 25/12/10 07:28AM, Patrick Steinhardt wrote:\n> > On Tue, Dec 09, 2025 at 04:58:19PM -0600, Justin Tobler wrote:\n> > > @@ -106,16 +137,12 @@ test_expect_success SHA1 'keyvalue and nul format' '\n> > >  \t\tobjects.tags.inflated=132\n> > >  \t\tEOF\n> > >  \n> > > -\t\tgit repo structure --format=keyvalue >out 2>err &&\n> > > +\t\tgit repo structure --format=keyvalue >out.raw 2>err &&\n> > >  \n> > > -\t\ttest_cmp expect out &&\n> > > -\t\ttest_line_count = 0 err &&\n> > > +\t\t# Strip object disk usage from output due to platform variance.\n> > > +\t\tgrep -v \"objects\\..*\\.disk=\" out.raw >out &&\n> > >  \n> > > -\t\t# Replace key and value delimiters for nul format.\n> > > -\t\ttr \"\\n=\" \"\\0\\n\" <expect >expect_nul &&\n> > > -\t\tgit repo structure --format=nul >out 2>err &&\n> > > -\n> > > -\t\ttest_cmp expect_nul out &&\n> > > +\t\ttest_cmp expect out &&\n> > >  \t\ttest_line_count = 0 err\n> > >  \t)\n> > >  '\n> > \n> > We could test disk sizes here test if we use git-rev-list(1) to compute\n> > disk size by type:\n> > \n> >     git rev-list --disk-usage HEAD --objects --filter=object:type=blob\n> >     git rev-list --disk-usage HEAD --objects --filter=object:type=commit\n> >     git rev-list --disk-usage HEAD --objects --filter=object:type=tag\n> >     git rev-list --disk-usage HEAD --objects --filter=object:type=tree\n> > \n> > The `--disk-usage` option also supports `--disk-usage=human`, which we\n> > can use in the next commit to verify that our computations are the same\n> > across git-rev-list(1) and git-repo(1).\n> \n> So, I'm not sure we can use git-rev-list(1) in the manner suggested\n> above. It looks like user-specified objects are always included in the\n> output. When using \"HEAD\" this means the referenced object will always\n> be included regardless of the filter used. In practice, this means\n> reported disk-usage when filtering by trees or blobs will likely be\n> inflated by objects not specified by the filter. As far as I am aware,\n> there is not a way to suppress user-specified objects in git-rev-list(1)\n> output.\n\nThere is, you can use \"--filter-provided-objects\".\n\n> I am somewhat curious if always including user-specified objects in\n> git-rev-list(1) output regardless of the specified filter is\n> intentional. Looking at git-rev-list(1) --filter documentation:\n> \n>   The form --filter=object:type=(tag|commit|tree|blob) omits all objects\n>   which are not of the requested type.\n> \n> doesn't indicate this limitation. From looking at the code in\n> list-objects-filter.c:list_objects_filter__filter_object() though, it\n> does somewhat seem like this behavior is intentional.\n\nIt is intentional, but I've been bitten by it in the past. Hence I\nintroduced the above option in 9cf68b27d5 (rev-list: allow filtering of\nprovided items, 2021-04-19).\n\nPatrick\n"},{"id":"532161","messageId":"aT-dmuOZyMhV0fX6@pks.im","threadId":"64606","inReplyTo":"20251212223644.3090879-4-jltobler@gmail.com","subject":"Re: [PATCH v2 3/7] builtin/repo: humanise count values in structure output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-15T05:33:14Z","receivedAt":"2025-12-15T05:33:20Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Dec 12, 2025 at 04:36:40PM -0600, Justin Tobler wrote:\n> diff --git a/strbuf.c b/strbuf.c\n> index 1fb47bf21b..cebb1593ab 100644\n> --- a/strbuf.c\n> +++ b/strbuf.c\n> @@ -836,6 +836,29 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n>  \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n>  }\n>  \n> +char *strbuf_humanise_count_value(struct strbuf *buf, size_t value)\n> +{\n> +\tif (value >= 1000000000) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5000000; /* for rounding */\n> +\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n> +\t\t\t    x / 1000000000, x % 1000000000 / 10000000);\n> +\t\treturn xstrfmt(_(\"G\"));\n> +\t} else if (value >= 1000000) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5000; /* for rounding */\n> +\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n> +\t\t\t    x / 1000000, x % 1000000 / 10000);\n> +\t\treturn xstrfmt(_(\"M\"));\n> +\t} else if (value >= 1000) {\n> +\t\tuintmax_t x = (uintmax_t)value + 5; /* for rounding */\n> +\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n> +\t\t\t    x / 1000, x % 1000 / 10);\n> +\t\treturn xstrfmt(_(\"k\"));\n> +\t} else {\n> +\t\tstrbuf_addf(buf, \"%\" PRIuMAX, (uintmax_t)value);\n> +\t\treturn NULL;\n> +\t}\n> +}\n\nSame comment here as in the previous patch, can't we return `const char *`\nhere in case we drop all allocations?\n\nPatrick\n"},{"id":"532162","messageId":"aT-doNe94GYmodQl@pks.im","threadId":"64606","inReplyTo":"20251212223644.3090879-5-jltobler@gmail.com","subject":"Re: [PATCH v2 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-15T05:33:20Z","receivedAt":"2025-12-15T05:33:25Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Dec 12, 2025 at 04:36:41PM -0600, Justin Tobler wrote:\n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index d3dfe416d0..3a2d15cec4 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -500,20 +513,38 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n>  {\n>  \tstruct count_objects_data *data = cb_data;\n>  \tstruct object_stats *stats = data->stats;\n> +\tsize_t inflated_total = 0;\n>  \tsize_t object_count;\n>  \n> +\tfor (size_t i = 0; i < oids->nr; i++) {\n> +\t\tstruct object_info oi = OBJECT_INFO_INIT;\n> +\t\tunsigned long inflated;\n> +\n> +\t\toi.sizep = &inflated;\n> +\n> +\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n> +\t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n\nUsing `OBJECT_INFO_FOR_PREFETCH` feels a bit weird to me, as we're not\nin a context where we want to do a prefetch. And if we ever were to\nextend that flag to have more semantics that are relevant to prefetches,\nonly, then this code here might become broken.\n\nUsing `SKIP_FETCH_OBJECT | INFO_QUICK` does make sense though, so I'd\nsuggest to expand the flag here.\n\nPatrick\n"},{"id":"532163","messageId":"aT-dppZm8TsibzyZ@pks.im","threadId":"64606","inReplyTo":"20251212223644.3090879-3-jltobler@gmail.com","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-15T05:33:26Z","receivedAt":"2025-12-15T05:33:31Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Dec 12, 2025 at 04:36:39PM -0600, Justin Tobler wrote:\n> diff --git a/strbuf.c b/strbuf.c\n> index 6c3851a7f8..1fb47bf21b 100644\n> --- a/strbuf.c\n> +++ b/strbuf.c\n> @@ -836,55 +836,49 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n>  \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n>  }\n>  \n> -static void strbuf_humanise(struct strbuf *buf, off_t bytes,\n> -\t\t\t\t int humanise_rate)\n> +char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n>  {\n> +\tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n> +\n>  \tif (bytes > 1 << 30) {\n> -\t\tstrbuf_addf(buf,\n> -\t\t\t\thumanise_rate == 0 ?\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte */\n> -\t\t\t\t\t_(\"%u.%2.2u GiB\") :\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second */\n> -\t\t\t\t\t_(\"%u.%2.2u GiB/s\"),\n> -\t\t\t    (unsigned)(bytes >> 30),\n> +\t\tstrbuf_addf(buf, \"%u.%2.2u\", (unsigned)(bytes >> 30),\n>  \t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n> +\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n> +\t\treturn humanise_rate ? xstrfmt(_(\"GiB/s\")) : xstrfmt(_(\"GiB\"));\n>  \t} else if (bytes > 1 << 20) {\n> -\t\tunsigned x = bytes + 5243;  /* for rounding */\n> -\t\tstrbuf_addf(buf,\n> -\t\t\t\thumanise_rate == 0 ?\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte */\n> -\t\t\t\t\t_(\"%u.%2.2u MiB\") :\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second */\n> -\t\t\t\t\t_(\"%u.%2.2u MiB/s\"),\n> -\t\t\t    x >> 20, ((x & ((1 << 20) - 1)) * 100) >> 20);\n> +\t\tunsigned x = bytes + 5243; /* for rounding */\n> +\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 20,\n> +\t\t\t    ((x & ((1 << 20) - 1)) * 100) >> 20);\n> +\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second and mebibyte */\n> +\t\treturn humanise_rate ? xstrfmt(_(\"MiB/s\")) : xstrfmt(_(\"MiB\"));\n>  \t} else if (bytes > 1 << 10) {\n> -\t\tunsigned x = bytes + 5;  /* for rounding */\n> -\t\tstrbuf_addf(buf,\n> -\t\t\t\thumanise_rate == 0 ?\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte */\n> -\t\t\t\t\t_(\"%u.%2.2u KiB\") :\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second */\n> -\t\t\t\t\t_(\"%u.%2.2u KiB/s\"),\n> -\t\t\t    x >> 10, ((x & ((1 << 10) - 1)) * 100) >> 10);\n> +\t\tunsigned x = bytes + 5; /* for rounding */\n> +\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 10,\n> +\t\t\t    ((x & ((1 << 10) - 1)) * 100) >> 10);\n> +\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second and kibibyte */\n> +\t\treturn humanise_rate ? xstrfmt(_(\"KiB/s\")) : xstrfmt(_(\"KiB\"));\n>  \t} else {\n> -\t\tstrbuf_addf(buf,\n> -\t\t\t\thumanise_rate == 0 ?\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n> -\t\t\t\t\tQ_(\"%u byte\", \"%u bytes\", bytes) :\n> -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n> -\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n> -\t\t\t\t(unsigned)bytes);\n> +\t\tstrbuf_addf(buf, \"%u\", (unsigned)bytes);\n> +\t\treturn humanise_rate ?\n> +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n> +\t\t\t       xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n> +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n> +\t\t\t       xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n>  \t}\n>  }\n\nAll branches use `xstrfmt()` with strings that are essentially\nconstants, except for the translation part. So isn't it possible to drop\nall these allocations and have the function return a `const char *`\ninstead?\n\n> diff --git a/strbuf.h b/strbuf.h\n> index a580ac6084..a5e3ab0cb4 100644\n> --- a/strbuf.h\n> +++ b/strbuf.h\n> @@ -367,6 +367,15 @@ void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n>   */\n>  void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n>  \n> +#define STRBUF_HUMANISE_RATE 1 << 0\n\nI think nowadays it's a bit more common to use an enum, and I think we\nshould also document what the flag does:\n\n    enum strbuf_humanise_flags {\n        /*\n         * Frobnicate the string.\n         */\n        STRBUF_HUMANISE_RATE = (1 << 0),\n    };\n\nPatrick\n"},{"id":"532164","messageId":"aT-drLh5WBX3vMLU@pks.im","threadId":"64606","inReplyTo":"20251212223644.3090879-7-jltobler@gmail.com","subject":"Re: [PATCH v2 6/7] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-15T05:33:32Z","receivedAt":"2025-12-15T05:33:37Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Dec 12, 2025 at 04:36:43PM -0600, Justin Tobler wrote:\n> diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> index b18213c660..1553f3cd32 100755\n> --- a/t/t1901-repo-structure.sh\n> +++ b/t/t1901-repo-structure.sh\n> @@ -4,6 +4,11 @@ test_description='test git repo structure'\n>  \n>  . ./test-lib.sh\n>  \n> +object_type_disk_usage() {\n> +\tgit cat-file --batch-check='%(objectsize:disk)' --batch-all-objects \\\n> +\t\t--filter=object:type=$1 | awk '{ sum += $1 } END { print sum }'\n> +}\n> +\n\nUsing `git rev-list --all --disk-usage --filter=object:type=$1\n--filter-provided-objects` would avoid the separate call to awk(1).\n\nPatrick\n"},{"id":"532178","messageId":"xmqqh5ts88b1.fsf@gitster.g","threadId":"64606","inReplyTo":"20251212223644.3090879-3-jltobler@gmail.com","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-15T08:21:06Z","receivedAt":"2025-12-15T08:21:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> +char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n>  {\n> +\tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n> +\n>  \tif (bytes > 1 << 30) {\n> +\t\tstrbuf_addf(buf, \"%u.%2.2u\", (unsigned)(bytes >> 30),\n>  \t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n> +\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n> +\t\treturn humanise_rate ? xstrfmt(_(\"GiB/s\")) : xstrfmt(_(\"GiB\"));\n> ...\n> }\n>  void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n>  {\n> -\tstrbuf_humanise(buf, bytes, 0);\n> +\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, 0);\n> +\tstrbuf_addf(buf, \" %s\", unit);\n> +\tfree(unit);\n>  }\n\nThe old \"strbuf-humanise\" used to treat the whole \"<number> <unit>\",\ne.g., _(\"%u.%2.2u GiB\"), as a single thing to be translated.\nHowever, the new code requires that in all languages:\n\n - Decimal point in number MUST be \".\" (don't some Europeans prefer\n   comma instead?);\n\n - Number MUST come before the unit;\n\n - Between the number and the unit, there has to be one and only one\n   SP.\n\nAll of which could be a severe regression from localization's point\nof view.\n\nThe first point among the above three can relatively easily\nremedied.  It is a bit more involved, but it is possible to fix the\nother two, too.\n\n\n\n\n\n"},{"id":"532193","messageId":"3uphps6olbz4qphxmivd7iwiwnfvj6mv7su4i4dhaujixf733i@eqybxhvjsznp","threadId":"64606","inReplyTo":"aT-djS-TrQJxxV8i@pks.im","subject":"Re: [PATCH 5/6] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T16:24:55Z","receivedAt":"2025-12-15T16:24:59Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/15 06:33AM, Patrick Steinhardt wrote:\n> On Fri, Dec 12, 2025 at 02:40:24PM -0600, Justin Tobler wrote:\n> > So, I'm not sure we can use git-rev-list(1) in the manner suggested\n> > above. It looks like user-specified objects are always included in the\n> > output. When using \"HEAD\" this means the referenced object will always\n> > be included regardless of the filter used. In practice, this means\n> > reported disk-usage when filtering by trees or blobs will likely be\n> > inflated by objects not specified by the filter. As far as I am aware,\n> > there is not a way to suppress user-specified objects in git-rev-list(1)\n> > output.\n> \n> There is, you can use \"--filter-provided-objects\".\n\nPerfect! I don't know how I missed that option. XD\n\n> > I am somewhat curious if always including user-specified objects in\n> > git-rev-list(1) output regardless of the specified filter is\n> > intentional. Looking at git-rev-list(1) --filter documentation:\n> > \n> >   The form --filter=object:type=(tag|commit|tree|blob) omits all objects\n> >   which are not of the requested type.\n> > \n> > doesn't indicate this limitation. From looking at the code in\n> > list-objects-filter.c:list_objects_filter__filter_object() though, it\n> > does somewhat seem like this behavior is intentional.\n> \n> It is intentional, but I've been bitten by it in the past. Hence I\n> introduced the above option in 9cf68b27d5 (rev-list: allow filtering of\n> provided items, 2021-04-19).\n\nGood to know. I think I'll submit a small patch today to try to clarify\nthe documentation here a little bit. It might be nice to point out this\nbehavior a bit more explictly in the --filter section. :)\n\n-Justin\n"},{"id":"532194","messageId":"qi2ealcgo5lwjjhx3m3acc7mukwaaxeblnbu7fmyw3gqvu3wvt@hqdatbug7s6b","threadId":"64606","inReplyTo":"aT-dppZm8TsibzyZ@pks.im","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T16:26:36Z","receivedAt":"2025-12-15T16:26:38Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/15 06:33AM, Patrick Steinhardt wrote:\n> On Fri, Dec 12, 2025 at 04:36:39PM -0600, Justin Tobler wrote:\n>\n> All branches use `xstrfmt()` with strings that are essentially\n> constants, except for the translation part. So isn't it possible to drop\n> all these allocations and have the function return a `const char *`\n> instead?\n\nYa, that would indeed be better. Will fix.\n\n> > diff --git a/strbuf.h b/strbuf.h\n> > index a580ac6084..a5e3ab0cb4 100644\n> > --- a/strbuf.h\n> > +++ b/strbuf.h\n> > @@ -367,6 +367,15 @@ void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n> >   */\n> >  void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n> >  \n> > +#define STRBUF_HUMANISE_RATE 1 << 0\n> \n> I think nowadays it's a bit more common to use an enum, and I think we\n> should also document what the flag does:\n> \n>     enum strbuf_humanise_flags {\n>         /*\n>          * Frobnicate the string.\n>          */\n>         STRBUF_HUMANISE_RATE = (1 << 0),\n>     };\n\nWill do.\n\n-Justin\n"},{"id":"532196","messageId":"kx3qdkm7rbd23hc66qamhq45agzofoppfhqnbbtw5cmjojevsq@2kkxiaem3fp4","threadId":"64606","inReplyTo":"xmqqh5ts88b1.fsf@gitster.g","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T16:47:27Z","receivedAt":"2025-12-15T16:47:34Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/15 05:21PM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > +char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n> >  {\n> > +\tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n> > +\n> >  \tif (bytes > 1 << 30) {\n> > +\t\tstrbuf_addf(buf, \"%u.%2.2u\", (unsigned)(bytes >> 30),\n> >  \t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n> > +\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n> > +\t\treturn humanise_rate ? xstrfmt(_(\"GiB/s\")) : xstrfmt(_(\"GiB\"));\n> > ...\n> > }\n> >  void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n> >  {\n> > -\tstrbuf_humanise(buf, bytes, 0);\n> > +\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, 0);\n> > +\tstrbuf_addf(buf, \" %s\", unit);\n> > +\tfree(unit);\n> >  }\n> \n> The old \"strbuf-humanise\" used to treat the whole \"<number> <unit>\",\n> e.g., _(\"%u.%2.2u GiB\"), as a single thing to be translated.\n> However, the new code requires that in all languages:\n> \n>  - Decimal point in number MUST be \".\" (don't some Europeans prefer\n>    comma instead?);\n> \n>  - Number MUST come before the unit;\n> \n>  - Between the number and the unit, there has to be one and only one\n>    SP.\n> \n> All of which could be a severe regression from localization's point\n> of view.\n> \n> The first point among the above three can relatively easily\n> remedied.  It is a bit more involved, but it is possible to fix the\n> other two, too.\n\nThe first point could be addressed by just making \"%u.%2.2u\"\ntranslatable. To address the others, we could have\nstrbuf_humanise_bytes_value() output two separate strings (value and\nunit) instead of appending the the value and returning the unit. Maybe\nsomething like:\n\n  void humanise_bytes(off_t bytes, char **value, const char **unit)\n\nWe could then have another translatable string to configure the format:\n\n  void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n  {\n    char *value;\n    const char *unit;\n\n    humanise_bytes(bytes, &value, &unit);\n    strbuf_addf(buf, _(\"%s %s\"), value, unit);\n    free(value);\n  }\n\nThis is certainly a bit more involved setup for translators though. But\nmaybe it's ok? I'll move forward with something like above in the next\nversion for now.\n\nThanks\n-Justin\n"},{"id":"532197","messageId":"vleglpqcjwzse63actqknbwdykvanzszosbflems33ntt3swoa@f7e3zsfatkoe","threadId":"64606","inReplyTo":"aT-doNe94GYmodQl@pks.im","subject":"Re: [PATCH v2 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T16:48:52Z","receivedAt":"2025-12-15T16:48:54Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/15 06:33AM, Patrick Steinhardt wrote:\n> On Fri, Dec 12, 2025 at 04:36:41PM -0600, Justin Tobler wrote:\n> > diff --git a/builtin/repo.c b/builtin/repo.c\n> > index d3dfe416d0..3a2d15cec4 100644\n> > --- a/builtin/repo.c\n> > +++ b/builtin/repo.c\n> > @@ -500,20 +513,38 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n> >  {\n> >  \tstruct count_objects_data *data = cb_data;\n> >  \tstruct object_stats *stats = data->stats;\n> > +\tsize_t inflated_total = 0;\n> >  \tsize_t object_count;\n> >  \n> > +\tfor (size_t i = 0; i < oids->nr; i++) {\n> > +\t\tstruct object_info oi = OBJECT_INFO_INIT;\n> > +\t\tunsigned long inflated;\n> > +\n> > +\t\toi.sizep = &inflated;\n> > +\n> > +\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n> > +\t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n> \n> Using `OBJECT_INFO_FOR_PREFETCH` feels a bit weird to me, as we're not\n> in a context where we want to do a prefetch. And if we ever were to\n> extend that flag to have more semantics that are relevant to prefetches,\n> only, then this code here might become broken.\n> \n> Using `SKIP_FETCH_OBJECT | INFO_QUICK` does make sense though, so I'd\n> suggest to expand the flag here.\n\nGood points. I'll update in the next version.\n\n-Justin\n"},{"id":"532205","messageId":"20251215205639.2700270-1-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-1-jltobler@gmail.com","subject":"[PATCH v3 0/7] builtin/repo: add object size info to structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:32Z","receivedAt":"2025-12-15T20:56:54Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThis patch series extends the recently introduced \"structure\" subcommand\nfor git-repo(1) to collect object size information. More specifically,\nit shows total inflated and disk sizes of objects by object type. The\naim to provide additional insight that may be useful to users regarding\nthe structure of a repository.\n\nIn addition to this change, this series also updates the table output\nformat to downscale larger output values along with the appropriate unit\nprefix. This is done to make table output more human friendly. The\nkeyvalue and nul output formats are left the same since they are\nintended more for machine parsing.\n\nChanges in V3:\n- Address potential localization regression by making the downscaled\n  number format string also translatable. Also make the format string\n  for how the values and unit prefixes are displayed via\n  `strbuf_humanise_{bytes,rate}()` translatable to be more flexible.\n- `strbuf_humanise_{bytes,count}_value()` has been renamed to\n  `humanise_{bytes,count}()` and updated to provide both the value and\n  unit prefix as separate strings.\n- Unit prefix strings are no longer allocated and instead constant.\n- The humanise flags are now defined in an enum.\n- Instead of using `OBJECT_INFO_FOR_PREFETCH`,\n  `OBJECT_INFO_SKIP_FETCH_OBJECT` and `OBJECT_INFO_QUICK` are used\n  explicitly.\n- Tests now use git-rev-list(1) to verify disk size info.\n\nChanges in V2:\n- Factor out and reuse existing logic from strbuf_humanise() to handle\n  downscaling values and determining the appropriate unit prefix\n  separately. This enables more control over how exactly the values are\n  written to the structure output table which is useful for alignment\n  reasons. I'm not how about the interface used in patch 2. Feedback is\n  most welcome.\n- In the previous version, when checking object size on a missing object\n  we would die. Instead we now ignore missing objects. This allows the\n  structure command to work on partial clones.\n- disk/inflated keyvalue names renamed to disk_size/inflated_size.\n- Unit prefixes are marked for translation.\n- The test for keyvalue disk size values are updated to check against\n  real expected values instead of skipping. Table output tests still\n  skip verifing human-readable values though.\n\nThanks,\n-Justin\n\nJustin Tobler (7):\n  builtin/repo: group per-type object values into struct\n  strbuf: split out logic to humanise byte values\n  builtin/repo: humanise count values in structure output\n  builtin/repo: add inflated object info to keyvalue structure output\n  builtin/repo: add inflated object info to structure table\n  builtin/repo: add disk size info to keyvalue stucture output\n  builtin/repo: add object disk size info to structure table\n\n Documentation/git-repo.adoc |   2 +\n builtin/repo.c              | 175 ++++++++++++++++++++++++++++++------\n strbuf.c                    |  93 ++++++++++++-------\n strbuf.h                    |  25 ++++++\n t/t1901-repo-structure.sh   | 113 +++++++++++++++--------\n 5 files changed, 311 insertions(+), 97 deletions(-)\n\nRange-diff against v2:\n1:  be14de68f6 = 1:  be14de68f6 builtin/repo: group per-type object values into struct\n2:  5ca6f9b708 ! 2:  1fa33f5906 strbuf: split out logic to humanise byte values\n    @@ Commit message\n         the git-repo(1) \"structure\" subcommand will be shown in a more\n         human-readable format with the appropriate unit prefixes. For this\n         usecase, the downscaled values and unit prefixes must be handled\n    -    separately to ensure proper column alignment. Refactor strbuf_humanise()\n    -    to instead append the downscaled byte value to the buffer only and\n    -    return the appropriate unit prefix string.\n    +    separately to ensure proper column alignment.\n    +\n    +    Split out logic from strbuf_humanise() to downscale byte values and\n    +    determine the corresponding unit prefix into a separate humanise_bytes()\n    +    function that provides seperate value and unit strings.\n     \n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n      \n     -static void strbuf_humanise(struct strbuf *buf, off_t bytes,\n     -\t\t\t\t int humanise_rate)\n    -+char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n    ++void humanise_bytes(off_t bytes, char **value, const char **unit,\n    ++\t\t    unsigned flags)\n      {\n    -+\tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n    ++\tint humanise_rate = flags & HUMANISE_RATE;\n     +\n      \tif (bytes > 1 << 30) {\n     -\t\tstrbuf_addf(buf,\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second */\n     -\t\t\t\t\t_(\"%u.%2.2u GiB/s\"),\n     -\t\t\t    (unsigned)(bytes >> 30),\n    -+\t\tstrbuf_addf(buf, \"%u.%2.2u\", (unsigned)(bytes >> 30),\n    - \t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n    +-\t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n    ++\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(bytes >> 30),\n    ++\t\t\t\t (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n     +\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n    -+\t\treturn humanise_rate ? xstrfmt(_(\"GiB/s\")) : xstrfmt(_(\"GiB\"));\n    ++\t\t*unit = humanise_rate ? _(\"GiB/s\") : _(\"GiB\");\n      \t} else if (bytes > 1 << 20) {\n     -\t\tunsigned x = bytes + 5243;  /* for rounding */\n     -\t\tstrbuf_addf(buf,\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     -\t\t\t\t\t_(\"%u.%2.2u MiB/s\"),\n     -\t\t\t    x >> 20, ((x & ((1 << 20) - 1)) * 100) >> 20);\n     +\t\tunsigned x = bytes + 5243; /* for rounding */\n    -+\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 20,\n    -+\t\t\t    ((x & ((1 << 20) - 1)) * 100) >> 20);\n    ++\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 20,\n    ++\t\t\t\t ((x & ((1 << 20) - 1)) * 100) >> 20);\n     +\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second and mebibyte */\n    -+\t\treturn humanise_rate ? xstrfmt(_(\"MiB/s\")) : xstrfmt(_(\"MiB\"));\n    ++\t\t*unit = humanise_rate ? _(\"MiB/s\") : _(\"MiB\");\n      \t} else if (bytes > 1 << 10) {\n     -\t\tunsigned x = bytes + 5;  /* for rounding */\n     -\t\tstrbuf_addf(buf,\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     -\t\t\t\t\t_(\"%u.%2.2u KiB/s\"),\n     -\t\t\t    x >> 10, ((x & ((1 << 10) - 1)) * 100) >> 10);\n     +\t\tunsigned x = bytes + 5; /* for rounding */\n    -+\t\tstrbuf_addf(buf, \"%u.%2.2u\", x >> 10,\n    -+\t\t\t    ((x & ((1 << 10) - 1)) * 100) >> 10);\n    ++\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 10,\n    ++\t\t\t\t ((x & ((1 << 10) - 1)) * 100) >> 10);\n     +\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second and kibibyte */\n    -+\t\treturn humanise_rate ? xstrfmt(_(\"KiB/s\")) : xstrfmt(_(\"KiB\"));\n    ++\t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n      \t} else {\n     -\t\tstrbuf_addf(buf,\n     -\t\t\t\thumanise_rate == 0 ?\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n     -\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n     -\t\t\t\t(unsigned)bytes);\n    -+\t\tstrbuf_addf(buf, \"%u\", (unsigned)bytes);\n    -+\t\treturn humanise_rate ?\n    ++\t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n    ++\t\t*unit = humanise_rate ?\n     +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n    -+\t\t\t       xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n    ++\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n     +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n    -+\t\t\t       xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n    ++\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n      \t}\n      }\n      \n    ++static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n    ++{\n    ++\tchar *value;\n    ++\tconst char *unit;\n    ++\n    ++\thumanise_bytes(bytes, &value, &unit, flags);\n    ++\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n    ++\tfree(value);\n    ++}\n    ++\n      void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n      {\n    --\tstrbuf_humanise(buf, bytes, 0);\n    -+\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, 0);\n    -+\tstrbuf_addf(buf, \" %s\", unit);\n    -+\tfree(unit);\n    - }\n    + \tstrbuf_humanise(buf, bytes, 0);\n    +@@ strbuf.c: void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n      \n      void strbuf_humanise_rate(struct strbuf *buf, off_t bytes)\n      {\n     -\tstrbuf_humanise(buf, bytes, 1);\n    -+\tchar *unit = strbuf_humanise_bytes_value(buf, bytes, STRBUF_HUMANISE_RATE);\n    -+\tstrbuf_addf(buf, \" %s\", unit);\n    -+\tfree(unit);\n    ++\tstrbuf_humanise(buf, bytes, HUMANISE_RATE);\n      }\n      \n      int printf_ln(const char *fmt, ...)\n    @@ strbuf.h: void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbu\n       */\n      void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n      \n    -+#define STRBUF_HUMANISE_RATE 1 << 0\n    ++enum humanise_flags {\n    ++\t/*\n    ++\t * Use rate based unit prefixes for humanised values.\n    ++\t */\n    ++\tHUMANISE_RATE = (1 << 0),\n    ++};\n     +\n     +/**\n    -+ * Append the given byte size as a human-readable string that is downscaled by\n    -+ * some factor. A string with the corresponding unit prefix is returned\n    -+ * separately.\n    ++ * Converts the given byte size into a downscaled human-readable value and\n    ++ * corresponding unit prefix as two separate strings.\n     + */\n    -+char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags);\n    ++void humanise_bytes(off_t bytes, char **value, const char **unit,\n    ++\t\t    unsigned flags);\n     +\n      /**\n       * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n3:  2efc3533ef ! 3:  8f09f6358e builtin/repo: humanise count values in structure output\n    @@ builtin/repo.c: struct stats_table {\n       */\n      struct stats_table_entry {\n      \tchar *value;\n    -+\tchar *unit;\n    ++\tconst char *unit;\n      };\n      \n      static void stats_table_vaddf(struct stats_table *table,\n    @@ builtin/repo.c: static void stats_table_vaddf(struct stats_table *table,\n      \n      static void stats_table_addf(struct stats_table *table, const char *format, ...)\n     @@ builtin/repo.c: static void stats_table_count_addf(struct stats_table *table, size_t value,\n    - \t\t\t\t   const char *format, ...)\n    - {\n    - \tstruct stats_table_entry *entry;\n    -+\tstruct strbuf buf = STRBUF_INIT;\n      \tva_list ap;\n      \n      \tCALLOC_ARRAY(entry, 1);\n     -\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n    -+\n    -+\tentry->unit = strbuf_humanise_count_value(&buf, value);\n    -+\tentry->value = strbuf_detach(&buf, NULL);\n    ++\thumanise_count(value, &entry->value, &entry->unit);\n      \n      \tva_start(ap, format);\n      \tstats_table_vaddf(table, entry, format, ap);\n    @@ builtin/repo.c: static void stats_table_print_structure(const struct stats_table\n      \t\tstrbuf_addstr(&buf, \" |\");\n      \t\tprintf(\"%s\\n\", buf.buf);\n      \t}\n    -@@ builtin/repo.c: static void stats_table_clear(struct stats_table *table)\n    - \n    - \tfor_each_string_list_item(item, &table->rows) {\n    - \t\tentry = item->util;\n    --\t\tif (entry)\n    -+\t\tif (entry) {\n    - \t\t\tfree(entry->value);\n    -+\t\t\tfree(entry->unit);\n    -+\t\t}\n    - \t}\n    - \n    - \tstring_list_clear(&table->rows, 1);\n     \n      ## strbuf.c ##\n     @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n      \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n      }\n      \n    -+char *strbuf_humanise_count_value(struct strbuf *buf, size_t value)\n    ++void humanise_count(size_t count, char **value, const char **unit)\n     +{\n    -+\tif (value >= 1000000000) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5000000; /* for rounding */\n    -+\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n    -+\t\t\t    x / 1000000000, x % 1000000000 / 10000000);\n    -+\t\treturn xstrfmt(_(\"G\"));\n    -+\t} else if (value >= 1000000) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5000; /* for rounding */\n    -+\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n    -+\t\t\t    x / 1000000, x % 1000000 / 10000);\n    -+\t\treturn xstrfmt(_(\"M\"));\n    -+\t} else if (value >= 1000) {\n    -+\t\tuintmax_t x = (uintmax_t)value + 5; /* for rounding */\n    -+\t\tstrbuf_addf(buf, \"%\" PRIuMAX \".%02\" PRIuMAX,\n    -+\t\t\t    x / 1000, x % 1000 / 10);\n    -+\t\treturn xstrfmt(_(\"k\"));\n    ++\tif (count >= 1000000000) {\n    ++\t\tsize_t x = count + 5000000; /* for rounding */\n    ++\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000000),\n    ++\t\t\t\t (unsigned)(x % 1000000000 / 10000000));\n    ++\t\t*unit = _(\"G\");\n    ++\t} else if (count >= 1000000) {\n    ++\t\tsize_t x = count + 5000; /* for rounding */\n    ++\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000),\n    ++\t\t\t\t (unsigned)(x % 1000000 / 10000));\n    ++\t\t*unit = _(\"M\");\n    ++\t} else if (count >= 1000) {\n    ++\t\tsize_t x = count + 5; /* for rounding */\n    ++\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000),\n    ++\t\t\t\t (unsigned)(x % 1000 / 10));\n    ++\t\t*unit = _(\"k\");\n     +\t} else {\n    -+\t\tstrbuf_addf(buf, \"%\" PRIuMAX, (uintmax_t)value);\n    -+\t\treturn NULL;\n    ++\t\t*value = xstrfmt(_(\"%u\"), (unsigned)count);\n    ++\t\t*unit = NULL;\n     +\t}\n     +}\n     +\n    - char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags)\n    + void humanise_bytes(off_t bytes, char **value, const char **unit,\n    + \t\t    unsigned flags)\n      {\n    - \tint humanise_rate = flags & STRBUF_HUMANISE_RATE;\n     \n      ## strbuf.h ##\n    -@@ strbuf.h: void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n    -  */\n    - char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flags);\n    +@@ strbuf.h: enum humanise_flags {\n    + void humanise_bytes(off_t bytes, char **value, const char **unit,\n    + \t\t    unsigned flags);\n      \n     +/**\n    -+ * Append the given count value as a human-readable string that is downsacled by\n    -+ * some factor. A string with the corresponding unit prefix is returned\n    -+ * separately.\n    ++ * Converts the given count into a downscaled human-readable value and\n    ++ * corresponding unit prefix as two separate strings.\n     + */\n    -+char *strbuf_humanise_count_value(struct strbuf *buf, size_t value);\n    ++void humanise_count(size_t count, char **value, const char **unit);\n     +\n      /**\n       * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n4:  627b8bf025 ! 4:  3f4eabe94f builtin/repo: add inflated object info to keyvalue structure output\n    @@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_arr\n     +\t\toi.sizep = &inflated;\n     +\n     +\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n    -+\t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n    ++\t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n    ++\t\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n     +\t\t\tcontinue;\n     +\n     +\t\tinflated_total += inflated;\n5:  14f4983e1d ! 5:  85d1052100 builtin/repo: add inflated object info to structure table\n    @@ builtin/repo.c: static void stats_table_count_addf(struct stats_table *table, si\n     +\t\t\t\t  const char *format, ...)\n     +{\n     +\tstruct stats_table_entry *entry;\n    -+\tstruct strbuf buf = STRBUF_INIT;\n     +\tva_list ap;\n     +\n     +\tCALLOC_ARRAY(entry, 1);\n    -+\n    -+\tentry->unit = strbuf_humanise_bytes_value(&buf, value,\n    -+\t\t\t\t\t\t  STRBUF_HUMANISE_COMPACT);\n    -+\tentry->value = strbuf_detach(&buf, NULL);\n    ++\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n     +\n     +\tva_start(ap, format);\n     +\tstats_table_vaddf(table, entry, format, ap);\n    @@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *tabl\n      static void stats_table_print_structure(const struct stats_table *table)\n     \n      ## strbuf.c ##\n    -@@ strbuf.c: char *strbuf_humanise_bytes_value(struct strbuf *buf, off_t bytes, unsigned flag\n    - \t\treturn humanise_rate ? xstrfmt(_(\"KiB/s\")) : xstrfmt(_(\"KiB\"));\n    +@@ strbuf.c: void humanise_bytes(off_t bytes, char **value, const char **unit,\n    + \t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n      \t} else {\n    - \t\tstrbuf_addf(buf, \"%u\", (unsigned)bytes);\n    -+\t\tif (flags & STRBUF_HUMANISE_COMPACT)\n    -+\t\t\treturn humanise_rate ?\n    -+\t\t\t\t       xstrfmt(_(\"B/s\")) :\n    -+\t\t\t\t       xstrfmt(_(\"B\"));\n    - \t\treturn humanise_rate ?\n    - \t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n    - \t\t\t       xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n    + \t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n    +-\t\t*unit = humanise_rate ?\n    +-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n    +-\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n    +-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n    +-\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n    ++\t\tif (flags & HUMANISE_COMPACT)\n    ++\t\t\t*unit = humanise_rate ? _(\"B/s\") : _(\"B\");\n    ++\t\telse\n    ++\t\t\t*unit = humanise_rate ?\n    ++\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n    ++\t\t\t\t\tQ_(\"byte/s\", \"bytes/s\", bytes) :\n    ++\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n    ++\t\t\t\t\tQ_(\"byte\", \"bytes\", bytes);\n    + \t}\n    + }\n    + \n     \n      ## strbuf.h ##\n    -@@ strbuf.h: void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n    -  */\n    - void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n    - \n    --#define STRBUF_HUMANISE_RATE 1 << 0\n    -+#define STRBUF_HUMANISE_RATE\t1 << 0\n    -+#define STRBUF_HUMANISE_COMPACT 1 << 1\n    +@@ strbuf.h: enum humanise_flags {\n    + \t * Use rate based unit prefixes for humanised values.\n    + \t */\n    + \tHUMANISE_RATE = (1 << 0),\n    ++\t/*\n    ++\t * Use compact \"B\" unit prefixes instead of \"byte/bytes\" for humanised\n    ++\t * values.\n    ++\t */\n    ++\tHUMANISE_COMPACT = (1 << 1),\n    + };\n      \n      /**\n    -  * Append the given byte size as a human-readable string that is downscaled by\n     \n      ## t/t1901-repo-structure.sh ##\n     @@ t/t1901-repo-structure.sh: test_expect_success 'empty repository' '\n6:  dc9e82889f ! 6:  e9fa9babec builtin/repo: add disk size info to keyvalue stucture output\n    @@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_arr\n     +\t\toi.disk_sizep = &disk;\n      \n      \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n    - \t\t\t\t\t\t  OBJECT_INFO_FOR_PREFETCH) < 0)\n    + \t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n    +@@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_array *oids,\n      \t\t\tcontinue;\n      \n      \t\tinflated_total += inflated;\n    @@ t/t1901-repo-structure.sh: test_description='test git repo structure'\n      . ./test-lib.sh\n      \n     +object_type_disk_usage() {\n    -+\tgit cat-file --batch-check='%(objectsize:disk)' --batch-all-objects \\\n    -+\t\t--filter=object:type=$1 | awk '{ sum += $1 } END { print sum }'\n    ++\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n    ++\t\t--filter-provided-objects\n     +}\n     +\n      test_expect_success 'empty repository' '\n7:  213b19dc7f ! 7:  df542c7bdf builtin/repo: add object disk size info to structure table\n    @@ Commit message\n         git-repo(1) structure command to display the total object disk usage by\n         object type.\n     \n    -    Since disk size may vary between platforms, tests do not validate actual\n    -    values and only check that size info is printed in an empty repository.\n    -\n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n      ## builtin/repo.c ##\n    @@ builtin/repo.c: static void stats_table_setup_structure(struct stats_table *tabl\n      static void stats_table_print_structure(const struct stats_table *table)\n     \n      ## t/t1901-repo-structure.sh ##\n    -@@ t/t1901-repo-structure.sh: object_type_disk_usage() {\n    - \t\t--filter=object:type=$1 | awk '{ sum += $1 } END { print sum }'\n    - }\n    +@@ t/t1901-repo-structure.sh: test_description='test git repo structure'\n    + . ./test-lib.sh\n      \n    -+strip_object_disk_usage() {\n    -+\tawk '\n    -+\t\t/^\\|   \\* Disk size/ { skip=1; next }\n    -+\t\tskip && /^\\|     \\* / { next }\n    -+\t\tskip && !/^\\|     \\* / { skip=0 }\n    -+\t\t{ print }\n    -+\t' $1\n    -+}\n    + object_type_disk_usage() {\n    +-\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n    +-\t\t--filter-provided-objects\n    ++\tdisk_usage_opt=\"--disk-usage\"\n    ++\n    ++\tif [ \"$2\" = \"true\" ]; then\n    ++\t\tdisk_usage_opt=\"--disk-usage=human\"\n    ++\tfi\n     +\n    ++\tif [ \"$1\" = \"all\" ]; then\n    ++\t\tgit rev-list --all --objects $disk_usage_opt\n    ++\telse\n    ++\t\tgit rev-list --all --objects $disk_usage_opt \\\n    ++\t\t\t--filter=object:type=$1 --filter-provided-objects\n    ++\tfi\n    + }\n    + \n      test_expect_success 'empty repository' '\n    - \ttest_when_finished \"rm -rf repo\" &&\n    - \tgit init repo &&\n     @@ t/t1901-repo-structure.sh: test_expect_success 'empty repository' '\n      \t\t|     * Trees          |    0 B |\n      \t\t|     * Blobs          |    0 B |\n    @@ t/t1901-repo-structure.sh: test_expect_success 'empty repository' '\n      \n      \t\tgit repo structure >out 2>err &&\n     @@ t/t1901-repo-structure.sh: test_expect_success SHA1 'repository with references and objects' '\n    + \t\t# Also creates a commit, tree, and blob.\n    + \t\tgit notes add -m foo &&\n    + \n    +-\t\tcat >expect <<-\\EOF &&\n    ++\t\tcat >expect <<-EOF &&\n    + \t\t| Repository structure | Value      |\n    + \t\t| -------------------- | ---------- |\n    + \t\t| * References         |            |\n    +@@ t/t1901-repo-structure.sh: test_expect_success SHA1 'repository with references and objects' '\n    + \t\t|     * Trees          |  15.81 MiB |\n    + \t\t|     * Blobs          |  11.68 KiB |\n      \t\t|     * Tags           |    132 B   |\n    ++\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n    ++\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n    ++\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n    ++\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n    ++\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n      \t\tEOF\n      \n    --\t\tgit repo structure >out 2>err &&\n    -+\t\tgit repo structure >out.raw 2>err &&\n    -+\n    -+\t\t# Skip object disk sizes due to platform variance.\n    -+\t\tstrip_object_disk_usage out.raw >out &&\n    - \n    - \t\ttest_cmp expect out &&\n    - \t\ttest_line_count = 0 err\n    + \t\tgit repo structure >out 2>err &&\n\nbase-commit: e85ae279b0d58edc2f4c3fd5ac391b51e1223985\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532206","messageId":"20251215205639.2700270-2-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v3 1/7] builtin/repo: group per-type object values into struct","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:33Z","receivedAt":"2025-12-15T20:56:55Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The `object_stats` structure stores object counts by type. In a\nsubsequent commit, additional per-type object measurements will also be\nstored. Group per-type object values into a new struct to allow better\nreuse.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 42 +++++++++++++++++++++++++-----------------\n 1 file changed, 25 insertions(+), 17 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 2a653bd3ea..a69699857a 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -202,13 +202,17 @@ struct ref_stats {\n \tsize_t others;\n };\n \n-struct object_stats {\n+struct object_values {\n \tsize_t tags;\n \tsize_t commits;\n \tsize_t trees;\n \tsize_t blobs;\n };\n \n+struct object_stats {\n+\tstruct object_values type_counts;\n+};\n+\n struct repo_structure {\n \tstruct ref_stats refs;\n \tstruct object_stats objects;\n@@ -281,9 +285,9 @@ static inline size_t get_total_reference_count(struct ref_stats *stats)\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n }\n \n-static inline size_t get_total_object_count(struct object_stats *stats)\n+static inline size_t get_total_object_values(struct object_values *values)\n {\n-\treturn stats->tags + stats->commits + stats->trees + stats->blobs;\n+\treturn values->tags + values->commits + values->trees + values->blobs;\n }\n \n static void stats_table_setup_structure(struct stats_table *table,\n@@ -302,14 +306,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_count(objects);\n+\tobject_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n \tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n-\tstats_table_count_addf(table, objects->commits, \"    * %s\", _(\"Commits\"));\n-\tstats_table_count_addf(table, objects->trees, \"    * %s\", _(\"Trees\"));\n-\tstats_table_count_addf(table, objects->blobs, \"    * %s\", _(\"Blobs\"));\n-\tstats_table_count_addf(table, objects->tags, \"    * %s\", _(\"Tags\"));\n+\tstats_table_count_addf(table, objects->type_counts.commits,\n+\t\t\t       \"    * %s\", _(\"Commits\"));\n+\tstats_table_count_addf(table, objects->type_counts.trees,\n+\t\t\t       \"    * %s\", _(\"Trees\"));\n+\tstats_table_count_addf(table, objects->type_counts.blobs,\n+\t\t\t       \"    * %s\", _(\"Blobs\"));\n+\tstats_table_count_addf(table, objects->type_counts.tags,\n+\t\t\t       \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -389,13 +397,13 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \t       (uintmax_t)stats->refs.others, value_delim);\n \n \tprintf(\"objects.commits.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.commits, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.commits, value_delim);\n \tprintf(\"objects.trees.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.trees, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.trees, value_delim);\n \tprintf(\"objects.blobs.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.blobs, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.blobs, value_delim);\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.tags, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n \tfflush(stdout);\n }\n@@ -473,22 +481,22 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n-\t\tstats->tags += oids->nr;\n+\t\tstats->type_counts.tags += oids->nr;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n-\t\tstats->commits += oids->nr;\n+\t\tstats->type_counts.commits += oids->nr;\n \t\tbreak;\n \tcase OBJ_TREE:\n-\t\tstats->trees += oids->nr;\n+\t\tstats->type_counts.trees += oids->nr;\n \t\tbreak;\n \tcase OBJ_BLOB:\n-\t\tstats->blobs += oids->nr;\n+\t\tstats->type_counts.blobs += oids->nr;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n \t}\n \n-\tobject_count = get_total_object_count(stats);\n+\tobject_count = get_total_object_values(&stats->type_counts);\n \tdisplay_progress(data->progress, object_count);\n \n \treturn 0;\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532207","messageId":"20251215205639.2700270-3-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v3 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:34Z","receivedAt":"2025-12-15T20:56:55Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"In a subsequent commit, byte size values displayed in table output for\nthe git-repo(1) \"structure\" subcommand will be shown in a more\nhuman-readable format with the appropriate unit prefixes. For this\nusecase, the downscaled values and unit prefixes must be handled\nseparately to ensure proper column alignment.\n\nSplit out logic from strbuf_humanise() to downscale byte values and\ndetermine the corresponding unit prefix into a separate humanise_bytes()\nfunction that provides seperate value and unit strings.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n strbuf.c | 69 ++++++++++++++++++++++++++++----------------------------\n strbuf.h | 14 ++++++++++++\n 2 files changed, 49 insertions(+), 34 deletions(-)\n\ndiff --git a/strbuf.c b/strbuf.c\nindex 6c3851a7f8..bb8e98872f 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,47 +836,48 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n-static void strbuf_humanise(struct strbuf *buf, off_t bytes,\n-\t\t\t\t int humanise_rate)\n+void humanise_bytes(off_t bytes, char **value, const char **unit,\n+\t\t    unsigned flags)\n {\n+\tint humanise_rate = flags & HUMANISE_RATE;\n+\n \tif (bytes > 1 << 30) {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u GiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u GiB/s\"),\n-\t\t\t    (unsigned)(bytes >> 30),\n-\t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(bytes >> 30),\n+\t\t\t\t (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n+\t\t*unit = humanise_rate ? _(\"GiB/s\") : _(\"GiB\");\n \t} else if (bytes > 1 << 20) {\n-\t\tunsigned x = bytes + 5243;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte */\n-\t\t\t\t\t_(\"%u.%2.2u MiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u MiB/s\"),\n-\t\t\t    x >> 20, ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\tunsigned x = bytes + 5243; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 20,\n+\t\t\t\t ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second and mebibyte */\n+\t\t*unit = humanise_rate ? _(\"MiB/s\") : _(\"MiB\");\n \t} else if (bytes > 1 << 10) {\n-\t\tunsigned x = bytes + 5;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u KiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u KiB/s\"),\n-\t\t\t    x >> 10, ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\tunsigned x = bytes + 5; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 10,\n+\t\t\t\t ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second and kibibyte */\n+\t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n \t} else {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n-\t\t\t\t\tQ_(\"%u byte\", \"%u bytes\", bytes) :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n-\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n-\t\t\t\t(unsigned)bytes);\n+\t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n+\t\t*unit = humanise_rate ?\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n+\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n+\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n \t}\n }\n \n+static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n+{\n+\tchar *value;\n+\tconst char *unit;\n+\n+\thumanise_bytes(bytes, &value, &unit, flags);\n+\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n+\tfree(value);\n+}\n+\n void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n {\n \tstrbuf_humanise(buf, bytes, 0);\n@@ -884,7 +885,7 @@ void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n \n void strbuf_humanise_rate(struct strbuf *buf, off_t bytes)\n {\n-\tstrbuf_humanise(buf, bytes, 1);\n+\tstrbuf_humanise(buf, bytes, HUMANISE_RATE);\n }\n \n int printf_ln(const char *fmt, ...)\ndiff --git a/strbuf.h b/strbuf.h\nindex a580ac6084..4426163e7e 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -367,6 +367,20 @@ void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n  */\n void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n \n+enum humanise_flags {\n+\t/*\n+\t * Use rate based unit prefixes for humanised values.\n+\t */\n+\tHUMANISE_RATE = (1 << 0),\n+};\n+\n+/**\n+ * Converts the given byte size into a downscaled human-readable value and\n+ * corresponding unit prefix as two separate strings.\n+ */\n+void humanise_bytes(off_t bytes, char **value, const char **unit,\n+\t\t    unsigned flags);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532208","messageId":"20251215205639.2700270-4-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v3 3/7] builtin/repo: humanise count values in structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:35Z","receivedAt":"2025-12-15T20:56:56Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The table output format for the git-repo(1) structure subcommand is used\nby default and intended to provide output to users in a human-friendly\nmanner. When the reference/object count values in a repository are\nlarge, it becomes more cumbersome for users to read the values.\n\nFor larger values, update the table output format to instead produce\nmore human-friendly count values that are scaled down with the\nappropriate unit prefix. Output for the keyvalue and nul formats remains\nunchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 38 +++++++++++++++++-------\n strbuf.c                  | 23 +++++++++++++++\n strbuf.h                  |  6 ++++\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++--------------------\n 4 files changed, 88 insertions(+), 41 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex a69699857a..9c61bc3e17 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -223,6 +223,7 @@ struct stats_table {\n \n \tint name_col_width;\n \tint value_col_width;\n+\tint unit_col_width;\n };\n \n /*\n@@ -230,6 +231,7 @@ struct stats_table {\n  */\n struct stats_table_entry {\n \tchar *value;\n+\tconst char *unit;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -250,11 +252,18 @@ static void stats_table_vaddf(struct stats_table *table,\n \n \tif (name_width > table->name_col_width)\n \t\ttable->name_col_width = name_width;\n-\tif (entry) {\n+\tif (!entry)\n+\t\treturn;\n+\tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n \t\t\ttable->value_col_width = value_width;\n \t}\n+\tif (entry->unit) {\n+\t\tint unit_width = utf8_strwidth(entry->unit);\n+\t\tif (unit_width > table->unit_col_width)\n+\t\t\ttable->unit_col_width = unit_width;\n+\t}\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -273,7 +282,7 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_list ap;\n \n \tCALLOC_ARRAY(entry, 1);\n-\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n+\thumanise_count(value, &entry->value, &entry->unit);\n \n \tva_start(ap, format);\n \tstats_table_vaddf(table, entry, format, ap);\n@@ -324,20 +333,24 @@ static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n \tconst char *value_col_title = _(\"Value\");\n-\tint name_col_width = utf8_strwidth(name_col_title);\n-\tint value_col_width = utf8_strwidth(value_col_title);\n+\tint title_name_width = utf8_strwidth(name_col_title);\n+\tint title_value_width = utf8_strwidth(value_col_title);\n+\tint name_col_width = table->name_col_width;\n+\tint value_col_width = table->value_col_width;\n+\tint unit_col_width = table->unit_col_width;\n \tstruct string_list_item *item;\n \tstruct strbuf buf = STRBUF_INIT;\n \n-\tif (table->name_col_width > name_col_width)\n-\t\tname_col_width = table->name_col_width;\n-\tif (table->value_col_width > value_col_width)\n-\t\tvalue_col_width = table->value_col_width;\n+\tif (title_name_width > name_col_width)\n+\t\tname_col_width = title_name_width;\n+\tif (title_value_width > value_col_width + unit_col_width + 1)\n+\t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n+\t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n \tstrbuf_addstr(&buf, \" |\");\n \tprintf(\"%s\\n\", buf.buf);\n \n@@ -345,17 +358,20 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tfor (int i = 0; i < name_col_width; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n-\tfor (int i = 0; i < value_col_width; i++)\n+\tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n \t\tputchar('-');\n \tprintf(\" |\\n\");\n \n \tfor_each_string_list_item(item, &table->rows) {\n \t\tstruct stats_table_entry *entry = item->util;\n \t\tconst char *value = \"\";\n+\t\tconst char *unit = \"\";\n \n \t\tif (entry) {\n \t\t\tstruct stats_table_entry *entry = item->util;\n \t\t\tvalue = entry->value;\n+\t\t\tif (entry->unit)\n+\t\t\t\tunit = entry->unit;\n \t\t}\n \n \t\tstrbuf_reset(&buf);\n@@ -363,6 +379,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n+\t\tstrbuf_addch(&buf, ' ');\n+\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, unit_col_width, unit);\n \t\tstrbuf_addstr(&buf, \" |\");\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\ndiff --git a/strbuf.c b/strbuf.c\nindex bb8e98872f..662edd4d19 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,6 +836,29 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n+void humanise_count(size_t count, char **value, const char **unit)\n+{\n+\tif (count >= 1000000000) {\n+\t\tsize_t x = count + 5000000; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000000),\n+\t\t\t\t (unsigned)(x % 1000000000 / 10000000));\n+\t\t*unit = _(\"G\");\n+\t} else if (count >= 1000000) {\n+\t\tsize_t x = count + 5000; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000),\n+\t\t\t\t (unsigned)(x % 1000000 / 10000));\n+\t\t*unit = _(\"M\");\n+\t} else if (count >= 1000) {\n+\t\tsize_t x = count + 5; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000),\n+\t\t\t\t (unsigned)(x % 1000 / 10));\n+\t\t*unit = _(\"k\");\n+\t} else {\n+\t\t*value = xstrfmt(_(\"%u\"), (unsigned)count);\n+\t\t*unit = NULL;\n+\t}\n+}\n+\n void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t    unsigned flags)\n {\ndiff --git a/strbuf.h b/strbuf.h\nindex 4426163e7e..571bd889df 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -381,6 +381,12 @@ enum humanise_flags {\n void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t    unsigned flags);\n \n+/**\n+ * Converts the given count into a downscaled human-readable value and\n+ * corresponding unit prefix as two separate strings.\n+ */\n+void humanise_count(size_t count, char **value, const char **unit);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 36a71a144e..55fd13ad1b 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -10,21 +10,21 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Branches       |     0 |\n-\t\t|     * Tags           |     0 |\n-\t\t|     * Remotes        |     0 |\n-\t\t|     * Others         |     0 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Commits        |     0 |\n-\t\t|     * Trees          |     0 |\n-\t\t|     * Blobs          |     0 |\n-\t\t|     * Tags           |     0 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Branches       |     0  |\n+\t\t|     * Tags           |     0  |\n+\t\t|     * Remotes        |     0  |\n+\t\t|     * Others         |     0  |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Commits        |     0  |\n+\t\t|     * Trees          |     0  |\n+\t\t|     * Blobs          |     0  |\n+\t\t|     * Tags           |     0  |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -39,7 +39,7 @@ test_expect_success 'repository with references and objects' '\n \tgit init repo &&\n \t(\n \t\tcd repo &&\n-\t\ttest_commit_bulk 42 &&\n+\t\ttest_commit_bulk 1005 &&\n \t\tgit tag -a foo -m bar &&\n \n \t\toid=\"$(git rev-parse HEAD)\" &&\n@@ -49,21 +49,21 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     4 |\n-\t\t|     * Branches       |     1 |\n-\t\t|     * Tags           |     1 |\n-\t\t|     * Remotes        |     1 |\n-\t\t|     * Others         |     1 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |   130 |\n-\t\t|     * Commits        |    43 |\n-\t\t|     * Trees          |    43 |\n-\t\t|     * Blobs          |    43 |\n-\t\t|     * Tags           |     1 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |    4   |\n+\t\t|     * Branches       |    1   |\n+\t\t|     * Tags           |    1   |\n+\t\t|     * Remotes        |    1   |\n+\t\t|     * Others         |    1   |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            | 3.02 k |\n+\t\t|     * Commits        | 1.01 k |\n+\t\t|     * Trees          | 1.01 k |\n+\t\t|     * Blobs          | 1.01 k |\n+\t\t|     * Tags           |    1   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532209","messageId":"20251215205639.2700270-5-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v3 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:36Z","receivedAt":"2025-12-15T20:56:57Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The structure subcommand for git-repo(1) outputs basic count information\nfor objects and references. Extend this output to also provide\ninformation regarding total size of inflated objects by object type.\n\nFor now, object size by object type info is only added to the keyvalue\nand nul output formats. In a subsequent commit, this info is also added\nto the table format.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 33 +++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   |  6 +++++-\n 3 files changed, 39 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 70f0a6d2e4..287eee4b93 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -50,6 +50,7 @@ supported:\n +\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n+* Total inflated size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 9c61bc3e17..e207108346 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,8 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hex.h\"\n+#include \"odb.h\"\n #include \"parse-options.h\"\n #include \"path-walk.h\"\n #include \"progress.h\"\n@@ -211,6 +213,7 @@ struct object_values {\n \n struct object_stats {\n \tstruct object_values type_counts;\n+\tstruct object_values inflated_sizes;\n };\n \n struct repo_structure {\n@@ -423,6 +426,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n+\tprintf(\"objects.commits.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -486,6 +498,7 @@ static void structure_count_references(struct ref_stats *stats,\n }\n \n struct count_objects_data {\n+\tstruct object_database *odb;\n \tstruct object_stats *stats;\n \tstruct progress *progress;\n };\n@@ -495,20 +508,39 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n+\tsize_t inflated_total = 0;\n \tsize_t object_count;\n \n+\tfor (size_t i = 0; i < oids->nr; i++) {\n+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n+\t\tunsigned long inflated;\n+\n+\t\toi.sizep = &inflated;\n+\n+\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n+\t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n+\t\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n+\t\t\tcontinue;\n+\n+\t\tinflated_total += inflated;\n+\t}\n+\n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n+\t\tstats->inflated_sizes.tags += inflated_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n+\t\tstats->inflated_sizes.commits += inflated_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n+\t\tstats->inflated_sizes.trees += inflated_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n+\t\tstats->inflated_sizes.blobs += inflated_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n@@ -526,6 +558,7 @@ static void structure_count_objects(struct object_stats *stats,\n {\n \tstruct path_walk_info info = PATH_WALK_INFO_INIT;\n \tstruct count_objects_data data = {\n+\t\t.odb = repo->objects,\n \t\t.stats = stats,\n \t};\n \ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 55fd13ad1b..33237822fd 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -73,7 +73,7 @@ test_expect_success 'repository with references and objects' '\n \t)\n '\n \n-test_expect_success 'keyvalue and nul format' '\n+test_expect_success SHA1 'keyvalue and nul format' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -90,6 +90,10 @@ test_expect_success 'keyvalue and nul format' '\n \t\tobjects.trees.count=42\n \t\tobjects.blobs.count=42\n \t\tobjects.tags.count=1\n+\t\tobjects.commits.inflated_size=9225\n+\t\tobjects.trees.inflated_size=28554\n+\t\tobjects.blobs.inflated_size=453\n+\t\tobjects.tags.inflated_size=132\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532210","messageId":"20251215205639.2700270-6-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v3 5/7] builtin/repo: add inflated object info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:37Z","receivedAt":"2025-12-15T20:56:57Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Update the table output format for the git-repo(1) structure command to\nbegin printing the total inflated object size info by object type. To be\nmore human-friendly, larger values are scaled down and displayed with\nthe appropriate unit prefix. Output for the keyvalue and nul formats\nremains unchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 33 +++++++++++++++++++--\n strbuf.c                  | 13 ++++----\n strbuf.h                  |  5 ++++\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++++++----------------\n 4 files changed, 79 insertions(+), 34 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex e207108346..b73cfd975b 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -292,6 +292,20 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_size_addf(struct stats_table *table, size_t value,\n+\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -307,7 +321,8 @@ static void stats_table_setup_structure(struct stats_table *table,\n {\n \tstruct object_stats *objects = &stats->objects;\n \tstruct ref_stats *refs = &stats->refs;\n-\tsize_t object_total;\n+\tsize_t inflated_object_total;\n+\tsize_t object_count_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -318,10 +333,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_values(&objects->type_counts);\n+\tobject_count_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n-\tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n+\tstats_table_count_addf(table, object_count_total, \"  * %s\", _(\"Count\"));\n \tstats_table_count_addf(table, objects->type_counts.commits,\n \t\t\t       \"    * %s\", _(\"Commits\"));\n \tstats_table_count_addf(table, objects->type_counts.trees,\n@@ -330,6 +345,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t       \"    * %s\", _(\"Blobs\"));\n \tstats_table_count_addf(table, objects->type_counts.tags,\n \t\t\t       \"    * %s\", _(\"Tags\"));\n+\n+\tinflated_object_total = get_total_object_values(&objects->inflated_sizes);\n+\tstats_table_size_addf(table, inflated_object_total,\n+\t\t\t      \"  * %s\", _(\"Inflated size\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/strbuf.c b/strbuf.c\nindex 662edd4d19..1e2d1f70a7 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -883,11 +883,14 @@ void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n \t} else {\n \t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n-\t\t*unit = humanise_rate ?\n-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n-\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n-\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n+\t\tif (flags & HUMANISE_COMPACT)\n+\t\t\t*unit = humanise_rate ? _(\"B/s\") : _(\"B\");\n+\t\telse\n+\t\t\t*unit = humanise_rate ?\n+\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n+\t\t\t\t\tQ_(\"byte/s\", \"bytes/s\", bytes) :\n+\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n+\t\t\t\t\tQ_(\"byte\", \"bytes\", bytes);\n \t}\n }\n \ndiff --git a/strbuf.h b/strbuf.h\nindex 571bd889df..005c155808 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -372,6 +372,11 @@ enum humanise_flags {\n \t * Use rate based unit prefixes for humanised values.\n \t */\n \tHUMANISE_RATE = (1 << 0),\n+\t/*\n+\t * Use compact \"B\" unit prefixes instead of \"byte/bytes\" for humanised\n+\t * values.\n+\t */\n+\tHUMANISE_COMPACT = (1 << 1),\n };\n \n /**\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 33237822fd..b18213c660 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -13,18 +13,23 @@ test_expect_success 'empty repository' '\n \t\t| Repository structure | Value  |\n \t\t| -------------------- | ------ |\n \t\t| * References         |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Branches       |     0  |\n-\t\t|     * Tags           |     0  |\n-\t\t|     * Remotes        |     0  |\n-\t\t|     * Others         |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Branches       |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|     * Remotes        |    0   |\n+\t\t|     * Others         |    0   |\n \t\t|                      |        |\n \t\t| * Reachable objects  |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Commits        |     0  |\n-\t\t|     * Trees          |     0  |\n-\t\t|     * Blobs          |     0  |\n-\t\t|     * Tags           |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Commits        |    0   |\n+\t\t|     * Trees          |    0   |\n+\t\t|     * Blobs          |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|   * Inflated size    |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -34,7 +39,7 @@ test_expect_success 'empty repository' '\n \t)\n '\n \n-test_expect_success 'repository with references and objects' '\n+test_expect_success SHA1 'repository with references and objects' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -49,21 +54,26 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    4   |\n-\t\t|     * Branches       |    1   |\n-\t\t|     * Tags           |    1   |\n-\t\t|     * Remotes        |    1   |\n-\t\t|     * Others         |    1   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            | 3.02 k |\n-\t\t|     * Commits        | 1.01 k |\n-\t\t|     * Trees          | 1.01 k |\n-\t\t|     * Blobs          | 1.01 k |\n-\t\t|     * Tags           |    1   |\n+\t\t| Repository structure | Value      |\n+\t\t| -------------------- | ---------- |\n+\t\t| * References         |            |\n+\t\t|   * Count            |      4     |\n+\t\t|     * Branches       |      1     |\n+\t\t|     * Tags           |      1     |\n+\t\t|     * Remotes        |      1     |\n+\t\t|     * Others         |      1     |\n+\t\t|                      |            |\n+\t\t| * Reachable objects  |            |\n+\t\t|   * Count            |   3.02 k   |\n+\t\t|     * Commits        |   1.01 k   |\n+\t\t|     * Trees          |   1.01 k   |\n+\t\t|     * Blobs          |   1.01 k   |\n+\t\t|     * Tags           |      1     |\n+\t\t|   * Inflated size    |  16.03 MiB |\n+\t\t|     * Commits        | 217.92 KiB |\n+\t\t|     * Trees          |  15.81 MiB |\n+\t\t|     * Blobs          |  11.68 KiB |\n+\t\t|     * Tags           |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532211","messageId":"20251215205639.2700270-7-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v3 6/7] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:38Z","receivedAt":"2025-12-15T20:56:58Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, extend the keyvalue and nul output formats of\nthe git-repo(1) structure command to additionally provide info regarding\ntotal object disk sizes by object type.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 18 ++++++++++++++++++\n t/t1901-repo-structure.sh   | 11 ++++++++++-\n 3 files changed, 29 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 287eee4b93..861073f641 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -51,6 +51,7 @@ supported:\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n+* Total disk size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex b73cfd975b..0ed41bf9d4 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -214,6 +214,7 @@ struct object_values {\n struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n+\tstruct object_values disk_sizes;\n };\n \n struct repo_structure {\n@@ -462,6 +463,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n \n+\tprintf(\"objects.commits.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -536,13 +546,16 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n \tsize_t inflated_total = 0;\n+\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\toff_t disk;\n \n \t\toi.sizep = &inflated;\n+\t\toi.disk_sizep = &disk;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n@@ -550,24 +563,29 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tcontinue;\n \n \t\tinflated_total += inflated;\n+\t\tdisk_total += disk;\n \t}\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n \t\tstats->inflated_sizes.tags += inflated_total;\n+\t\tstats->disk_sizes.tags += disk_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n \t\tstats->inflated_sizes.commits += inflated_total;\n+\t\tstats->disk_sizes.commits += disk_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n \t\tstats->inflated_sizes.trees += inflated_total;\n+\t\tstats->disk_sizes.trees += disk_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n \t\tstats->inflated_sizes.blobs += inflated_total;\n+\t\tstats->disk_sizes.blobs += disk_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex b18213c660..dd17caad05 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -4,6 +4,11 @@ test_description='test git repo structure'\n \n . ./test-lib.sh\n \n+object_type_disk_usage() {\n+\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n+\t\t--filter-provided-objects\n+}\n+\n test_expect_success 'empty repository' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n@@ -91,7 +96,7 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\ttest_commit_bulk 42 &&\n \t\tgit tag -a foo -m bar &&\n \n-\t\tcat >expect <<-\\EOF &&\n+\t\tcat >expect <<-EOF &&\n \t\treferences.branches.count=1\n \t\treferences.tags.count=1\n \t\treferences.remotes.count=0\n@@ -104,6 +109,10 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.trees.inflated_size=28554\n \t\tobjects.blobs.inflated_size=453\n \t\tobjects.tags.inflated_size=132\n+\t\tobjects.commits.disk_size=$(object_type_disk_usage commit)\n+\t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n+\t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n+\t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532212","messageId":"20251215205639.2700270-8-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v3 7/7] builtin/repo: add object disk size info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-15T20:56:39Z","receivedAt":"2025-12-15T20:56:59Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, update the table output format for the\ngit-repo(1) structure command to display the total object disk usage by\nobject type.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 13 +++++++++++++\n t/t1901-repo-structure.sh | 26 +++++++++++++++++++++++---\n 2 files changed, 36 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 0ed41bf9d4..a071d2fdfe 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -324,6 +324,7 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstruct ref_stats *refs = &stats->refs;\n \tsize_t inflated_object_total;\n \tsize_t object_count_total;\n+\tsize_t disk_object_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -358,6 +359,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->inflated_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tdisk_object_total = get_total_object_values(&objects->disk_sizes);\n+\tstats_table_size_addf(table, disk_object_total,\n+\t\t\t      \"  * %s\", _(\"Disk size\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex dd17caad05..64db191234 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -5,8 +5,18 @@ test_description='test git repo structure'\n . ./test-lib.sh\n \n object_type_disk_usage() {\n-\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n-\t\t--filter-provided-objects\n+\tdisk_usage_opt=\"--disk-usage\"\n+\n+\tif [ \"$2\" = \"true\" ]; then\n+\t\tdisk_usage_opt=\"--disk-usage=human\"\n+\tfi\n+\n+\tif [ \"$1\" = \"all\" ]; then\n+\t\tgit rev-list --all --objects $disk_usage_opt\n+\telse\n+\t\tgit rev-list --all --objects $disk_usage_opt \\\n+\t\t\t--filter=object:type=$1 --filter-provided-objects\n+\tfi\n }\n \n test_expect_success 'empty repository' '\n@@ -35,6 +45,11 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|   * Disk size        |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -58,7 +73,7 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# Also creates a commit, tree, and blob.\n \t\tgit notes add -m foo &&\n \n-\t\tcat >expect <<-\\EOF &&\n+\t\tcat >expect <<-EOF &&\n \t\t| Repository structure | Value      |\n \t\t| -------------------- | ---------- |\n \t\t| * References         |            |\n@@ -79,6 +94,11 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Trees          |  15.81 MiB |\n \t\t|     * Blobs          |  11.68 KiB |\n \t\t|     * Tags           |    132 B   |\n+\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n+\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532228","messageId":"xmqqms3j5il1.fsf@gitster.g","threadId":"64606","inReplyTo":"20251215205639.2700270-3-jltobler@gmail.com","subject":"Re: [PATCH v3 2/7] strbuf: split out logic to humanise byte values","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-16T01:19:38Z","receivedAt":"2025-12-16T01:19:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> +\t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n\nDoes this \"%u\" need translation?\n\nI very much doubt it, but if it did, this does need TRANSLATORS\ncomment.\n\n> +\t\t*unit = humanise_rate ?\n> +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n> +\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n> +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n> +\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n>  \t}\n>  }\n>  \n> +static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n> +{\n> +\tchar *value;\n> +\tconst char *unit;\n> +\n> +\thumanise_bytes(bytes, &value, &unit, flags);\n> +\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n\nThis definitely needs the TRANSLATORS comment to tell what is going on.\n\n> +\tfree(value);\n> +}\n\n\n"},{"id":"532231","messageId":"lftfcdnv7cn6ajrkjiim3z2ympvlfmlvtfco3x2wwpknytorif@3uutxricxy5d","threadId":"64606","inReplyTo":"xmqqms3j5il1.fsf@gitster.g","subject":"Re: [PATCH v3 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T01:36:10Z","receivedAt":"2025-12-16T01:36:12Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/16 10:19AM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > +\t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n> \n> Does this \"%u\" need translation?\n> \n> I very much doubt it, but if it did, this does need TRANSLATORS\n> comment.\n\nYa, I don't think one should be necessary. Will remove in the next\nversion.\n\nI think I made the same mistake in humanise_count() in a later patch.\nI'll also adjust it there.\n\n> \n> > +\t\t*unit = humanise_rate ?\n> > +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n> > +\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n> > +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n> > +\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n> >  \t}\n> >  }\n> >  \n> > +static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n> > +{\n> > +\tchar *value;\n> > +\tconst char *unit;\n> > +\n> > +\thumanise_bytes(bytes, &value, &unit, flags);\n> > +\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n> \n> This definitely needs the TRANSLATORS comment to tell what is going on.\n\nOk, will do in the next version. Thanks :)\n\n-Justin\n"},{"id":"532234","messageId":"CANYiYbE3Tx6B5L5rEoDue7hTYzFGxw_qA-MRpC9RSxQ7HRczaw@mail.gmail.com","threadId":"64606","inReplyTo":"20251212223644.3090879-3-jltobler@gmail.com","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-12-16T02:26:10Z","receivedAt":"2025-12-16T02:26:22Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Sat, Dec 13, 2025 at 6:37 AM Justin Tobler <jltobler@gmail.com> wrote:\n> +               return humanise_rate ?\n> +                              /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n> +                              xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n> +                              /* TRANSLATORS: IEC 80000-13:2008 byte */\n> +                              xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n\nWe have already defined \"byte\" as a 10n string without plural forms in the\nfile \"t/helper/test-simple-ipc.c\" via commit 36a7eb6876 (t0052: add simple-ipc\ntests and t/helper/test-simple-ipc tool, 2021-03-22 10:29:48 +0000).\n\n    OPT_STRING(0, \"byte\", &bytevalue, N_(\"byte\"), N_(\"ballast character\")),\n\nThe newly introduced usage of \"byte\" is now marked as having a plural form\n(via Q_(\"byte\", \"bytes\", bytes)), which causes a conflict. This results in make\npot failing with the following error:\n\n    msgcat: msgid 'byte' is used without plural and with plural.\n\nThis happens because gettext requires that a given msgid be treated\nconsistently—either exclusively as a singular string or as part of a plural\nconstruct—but not both.\n\nTo resolve this conflict, we can unmark the singular \"byte\" in\nt/helper/test-simple-ipc.c, allowing it to reuse the translation from the\nplural-form definition of \"byte\".\n\n--\nJiang Xin\n"},{"id":"532236","messageId":"xmqqqzsv3uus.fsf@gitster.g","threadId":"64606","inReplyTo":"CANYiYbE3Tx6B5L5rEoDue7hTYzFGxw_qA-MRpC9RSxQ7HRczaw@mail.gmail.com","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-16T04:37:31Z","receivedAt":"2025-12-16T04:37:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\n> On Sat, Dec 13, 2025 at 6:37 AM Justin Tobler <jltobler@gmail.com> wrote:\n>> +               return humanise_rate ?\n>> +                              /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n>> +                              xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n>> +                              /* TRANSLATORS: IEC 80000-13:2008 byte */\n>> +                              xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n>\n> We have already defined \"byte\" as a 10n string without plural forms in the\n> file \"t/helper/test-simple-ipc.c\" via commit 36a7eb6876 (t0052: add simple-ipc\n> tests and t/helper/test-simple-ipc tool, 2021-03-22 10:29:48 +0000).\n>\n>     OPT_STRING(0, \"byte\", &bytevalue, N_(\"byte\"), N_(\"ballast character\")),\n>\n> The newly introduced usage of \"byte\" is now marked as having a plural form\n> (via Q_(\"byte\", \"bytes\", bytes)), which causes a conflict. This results in make\n> pot failing with the following error:\n>\n>     msgcat: msgid 'byte' is used without plural and with plural.\n>\n> This happens because gettext requires that a given msgid be treated\n> consistently—either exclusively as a singular string or as part of a plural\n> construct—but not both.\n>\n> To resolve this conflict, we can unmark the singular \"byte\" in\n> t/helper/test-simple-ipc.c, allowing it to reuse the translation from the\n> plural-form definition of \"byte\".\n\nI learned a new thing today and am happy :).\n\nBut how does one \"unmark\" the singular \"byte\" there, exactly?\n\nWould something like this ...\n\n     OPT_STRING(0, \"byte\", &bytevalue, Q_(\"byte\", \"bytes\", 1), N_(\"ballast character\")),\n\n... a good idea, to \"mark\" it as a countable noun that has a plural\nform?\n\nOr did you mean that we can simply drop N_() around it, i.e.,\nN_(\"byte\") -> \"byte\", to discard the i18n, because it merely is a\ntest helper?\n\nPunting is fine in this case, but in case a similar situation arises\nin real code, it would be better to establish a pattern we can\nfollow.\n\nThanks.\n"},{"id":"532237","messageId":"CANYiYbExjGoCw4n92a75xtREE_EhjEySVSmk=NwJd3GoMAoVLg@mail.gmail.com","threadId":"64606","inReplyTo":"xmqqqzsv3uus.fsf@gitster.g","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-12-16T06:18:06Z","receivedAt":"2025-12-16T06:18:20Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Tue, Dec 16, 2025 at 12:37 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Jiang Xin <worldhello.net@gmail.com> writes:\n>\n> > On Sat, Dec 13, 2025 at 6:37 AM Justin Tobler <jltobler@gmail.com> wrote:\n> >> +               return humanise_rate ?\n> >> +                              /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n> >> +                              xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n> >> +                              /* TRANSLATORS: IEC 80000-13:2008 byte */\n> >> +                              xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n> >\n> > We have already defined \"byte\" as a 10n string without plural forms in the\n> > file \"t/helper/test-simple-ipc.c\" via commit 36a7eb6876 (t0052: add simple-ipc\n> > tests and t/helper/test-simple-ipc tool, 2021-03-22 10:29:48 +0000).\n> >\n> >     OPT_STRING(0, \"byte\", &bytevalue, N_(\"byte\"), N_(\"ballast character\")),\n> >\n> > The newly introduced usage of \"byte\" is now marked as having a plural form\n> > (via Q_(\"byte\", \"bytes\", bytes)), which causes a conflict. This results in make\n> > pot failing with the following error:\n> >\n> >     msgcat: msgid 'byte' is used without plural and with plural.\n> >\n> > This happens because gettext requires that a given msgid be treated\n> > consistently—either exclusively as a singular string or as part of a plural\n> > construct—but not both.\n> >\n> > To resolve this conflict, we can unmark the singular \"byte\" in\n> > t/helper/test-simple-ipc.c, allowing it to reuse the translation from the\n> > plural-form definition of \"byte\".\n>\n> I learned a new thing today and am happy :).\n>\n> But how does one \"unmark\" the singular \"byte\" there, exactly?\n>\n> Would something like this ...\n>\n>      OPT_STRING(0, \"byte\", &bytevalue, Q_(\"byte\", \"bytes\", 1), N_(\"ballast character\")),\n>\n> ... a good idea, to \"mark\" it as a countable noun that has a plural\n> form?\n>\n> Or did you mean that we can simply drop N_() around it, i.e.,\n> N_(\"byte\") -> \"byte\", to discard the i18n, because it merely is a\n> test helper?\n\nI prefer dropping N_() for \"byte\" in \"t/helper/test-simple-ipc.c\", and\nthe i18n for the test helper will continue to work as before if we also\nmark the plural-form of \"byte\" in this patch series. (i.e., drop the N_()\nfor \"byte\" in the test helper in this patch.)\n\nThis is because N_() is a macro that does not invoke any gettext\nfunction, only returns msgid as in gettext.h:\n\n    #define N_(msgid) msgid\n\nAnd the actual translation for the msgid (the argh field of an option)\noccurs later by calling:\n\n    opts->argh ? _(opts->argh) : _(\"...\")\n\nin \"parse-options.c\".\n\nHowever, replacing N_() with Q_() would cause the string to be\nprocessed by gettext twice: once at runtime via Q_(), and again\nwhen _(opts->argh) is evaluated.\n"},{"id":"532245","messageId":"aUEXdE7qxy8TfUJR@pks.im","threadId":"64606","inReplyTo":"20251215205639.2700270-4-jltobler@gmail.com","subject":"Re: [PATCH v3 3/7] builtin/repo: humanise count values in structure output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-16T08:25:24Z","receivedAt":"2025-12-16T08:25:31Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Dec 15, 2025 at 02:56:35PM -0600, Justin Tobler wrote:\n> diff --git a/strbuf.c b/strbuf.c\n> index bb8e98872f..662edd4d19 100644\n> --- a/strbuf.c\n> +++ b/strbuf.c\n> @@ -836,6 +836,29 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n>  \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n>  }\n>  \n> +void humanise_count(size_t count, char **value, const char **unit)\n> +{\n> +\tif (count >= 1000000000) {\n> +\t\tsize_t x = count + 5000000; /* for rounding */\n> +\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000000),\n> +\t\t\t\t (unsigned)(x % 1000000000 / 10000000));\n> +\t\t*unit = _(\"G\");\n> +\t} else if (count >= 1000000) {\n> +\t\tsize_t x = count + 5000; /* for rounding */\n> +\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000),\n> +\t\t\t\t (unsigned)(x % 1000000 / 10000));\n> +\t\t*unit = _(\"M\");\n> +\t} else if (count >= 1000) {\n> +\t\tsize_t x = count + 5; /* for rounding */\n> +\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000),\n> +\t\t\t\t (unsigned)(x % 1000 / 10));\n> +\t\t*unit = _(\"k\");\n> +\t} else {\n> +\t\t*value = xstrfmt(_(\"%u\"), (unsigned)count);\n> +\t\t*unit = NULL;\n> +\t}\n> +}\n\nI guess these here could also all use TRANSLATOR comments.\n\nPatrick\n"},{"id":"532246","messageId":"aUEXeuCkMDWSfwHi@pks.im","threadId":"64606","inReplyTo":"20251215205639.2700270-8-jltobler@gmail.com","subject":"Re: [PATCH v3 7/7] builtin/repo: add object disk size info to structure table","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-16T08:25:30Z","receivedAt":"2025-12-16T08:25:36Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Dec 15, 2025 at 02:56:39PM -0600, Justin Tobler wrote:\n> diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> index dd17caad05..64db191234 100755\n> --- a/t/t1901-repo-structure.sh\n> +++ b/t/t1901-repo-structure.sh\n> @@ -5,8 +5,18 @@ test_description='test git repo structure'\n>  . ./test-lib.sh\n>  \n>  object_type_disk_usage() {\n> -\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n> -\t\t--filter-provided-objects\n> +\tdisk_usage_opt=\"--disk-usage\"\n> +\n> +\tif [ \"$2\" = \"true\" ]; then\n> +\t\tdisk_usage_opt=\"--disk-usage=human\"\n> +\tfi\n> +\n> +\tif [ \"$1\" = \"all\" ]; then\n> +\t\tgit rev-list --all --objects $disk_usage_opt\n> +\telse\n> +\t\tgit rev-list --all --objects $disk_usage_opt \\\n> +\t\t\t--filter=object:type=$1 --filter-provided-objects\n> +\tfi\n>  }\n>  \n>  test_expect_success 'empty repository' '\n\nWe don't use `if [ ... ]` in our codebase, and we typically have the\n`then` on the next line:\n\n    if test \"$2\" = \"true\"\n    then\n        ...\n    fi\n\n    if test \"$1\" = \"all\"\n    then\n        ...\n    else\n        ...\n    fi\n\n> @@ -79,6 +94,11 @@ test_expect_success SHA1 'repository with references and objects' '\n>  \t\t|     * Trees          |  15.81 MiB |\n>  \t\t|     * Blobs          |  11.68 KiB |\n>  \t\t|     * Tags           |    132 B   |\n> +\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n> +\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n> +\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n> +\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n> +\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n>  \t\tEOF\n\nCurious, but why is the last one special here?\n\nPatrick\n"},{"id":"532276","messageId":"z7fuww4wnfpt5m7rojixyp3atejopjr623bi7o7snplas7dgsg@yktwdifek23m","threadId":"64606","inReplyTo":"CANYiYbExjGoCw4n92a75xtREE_EhjEySVSmk=NwJd3GoMAoVLg@mail.gmail.com","subject":"Re: [PATCH v2 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T14:41:48Z","receivedAt":"2025-12-16T14:41:51Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/16 02:18PM, Jiang Xin wrote:\n> On Tue, Dec 16, 2025 at 12:37 PM Junio C Hamano <gitster@pobox.com> wrote:\n> >\n> > Jiang Xin <worldhello.net@gmail.com> writes:\n> >\n> > > On Sat, Dec 13, 2025 at 6:37 AM Justin Tobler <jltobler@gmail.com> wrote:\n> > >> +               return humanise_rate ?\n> > >> +                              /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n> > >> +                              xstrfmt(Q_(\"byte/s\", \"bytes/s\", bytes)) :\n> > >> +                              /* TRANSLATORS: IEC 80000-13:2008 byte */\n> > >> +                              xstrfmt(Q_(\"byte\", \"bytes\", bytes));\n> > >\n> > > We have already defined \"byte\" as a 10n string without plural forms in the\n> > > file \"t/helper/test-simple-ipc.c\" via commit 36a7eb6876 (t0052: add simple-ipc\n> > > tests and t/helper/test-simple-ipc tool, 2021-03-22 10:29:48 +0000).\n> > >\n> > >     OPT_STRING(0, \"byte\", &bytevalue, N_(\"byte\"), N_(\"ballast character\")),\n> > >\n> > > The newly introduced usage of \"byte\" is now marked as having a plural form\n> > > (via Q_(\"byte\", \"bytes\", bytes)), which causes a conflict. This results in make\n> > > pot failing with the following error:\n> > >\n> > >     msgcat: msgid 'byte' is used without plural and with plural.\n> > >\n> > > This happens because gettext requires that a given msgid be treated\n> > > consistently—either exclusively as a singular string or as part of a plural\n> > > construct—but not both.\n> > >\n> > > To resolve this conflict, we can unmark the singular \"byte\" in\n> > > t/helper/test-simple-ipc.c, allowing it to reuse the translation from the\n> > > plural-form definition of \"byte\".\n> >\n> > I learned a new thing today and am happy :).\n> >\n> > But how does one \"unmark\" the singular \"byte\" there, exactly?\n> >\n> > Would something like this ...\n> >\n> >      OPT_STRING(0, \"byte\", &bytevalue, Q_(\"byte\", \"bytes\", 1), N_(\"ballast character\")),\n> >\n> > ... a good idea, to \"mark\" it as a countable noun that has a plural\n> > form?\n> >\n> > Or did you mean that we can simply drop N_() around it, i.e.,\n> > N_(\"byte\") -> \"byte\", to discard the i18n, because it merely is a\n> > test helper?\n> \n> I prefer dropping N_() for \"byte\" in \"t/helper/test-simple-ipc.c\", and\n> the i18n for the test helper will continue to work as before if we also\n> mark the plural-form of \"byte\" in this patch series. (i.e., drop the N_()\n> for \"byte\" in the test helper in this patch.)\n> \n> This is because N_() is a macro that does not invoke any gettext\n> function, only returns msgid as in gettext.h:\n> \n>     #define N_(msgid) msgid\n> \n> And the actual translation for the msgid (the argh field of an option)\n> occurs later by calling:\n> \n>     opts->argh ? _(opts->argh) : _(\"...\")\n> \n> in \"parse-options.c\".\n> \n> However, replacing N_() with Q_() would cause the string to be\n> processed by gettext twice: once at runtime via Q_(), and again\n> when _(opts->argh) is evaluated.\n\nThanks both! This thread has been very informative. In the version I'll\ngo ahead and drop the N_() here for this patch. :)\n\n-Justin\n"},{"id":"532277","messageId":"y7kutectqntle5557tjmta44wwjvk2f4tvsxfuajaktj647275@6kupww6ldexe","threadId":"64606","inReplyTo":"aUEXeuCkMDWSfwHi@pks.im","subject":"Re: [PATCH v3 7/7] builtin/repo: add object disk size info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T14:48:30Z","receivedAt":"2025-12-16T14:48:32Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/16 09:25AM, Patrick Steinhardt wrote:\n> On Mon, Dec 15, 2025 at 02:56:39PM -0600, Justin Tobler wrote:\n> > diff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\n> > index dd17caad05..64db191234 100755\n> > --- a/t/t1901-repo-structure.sh\n> > +++ b/t/t1901-repo-structure.sh\n> > @@ -5,8 +5,18 @@ test_description='test git repo structure'\n> >  . ./test-lib.sh\n> >  \n> >  object_type_disk_usage() {\n> > -\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n> > -\t\t--filter-provided-objects\n> > +\tdisk_usage_opt=\"--disk-usage\"\n> > +\n> > +\tif [ \"$2\" = \"true\" ]; then\n> > +\t\tdisk_usage_opt=\"--disk-usage=human\"\n> > +\tfi\n> > +\n> > +\tif [ \"$1\" = \"all\" ]; then\n> > +\t\tgit rev-list --all --objects $disk_usage_opt\n> > +\telse\n> > +\t\tgit rev-list --all --objects $disk_usage_opt \\\n> > +\t\t\t--filter=object:type=$1 --filter-provided-objects\n> > +\tfi\n> >  }\n> >  \n> >  test_expect_success 'empty repository' '\n> \n> We don't use `if [ ... ]` in our codebase, and we typically have the\n> `then` on the next line:\n> \n>     if test \"$2\" = \"true\"\n>     then\n>         ...\n>     fi\n> \n>     if test \"$1\" = \"all\"\n>     then\n>         ...\n>     else\n>         ...\n>     fi\n\nNoted, will fix.\n\n> > @@ -79,6 +94,11 @@ test_expect_success SHA1 'repository with references and objects' '\n> >  \t\t|     * Trees          |  15.81 MiB |\n> >  \t\t|     * Blobs          |  11.68 KiB |\n> >  \t\t|     * Tags           |    132 B   |\n> > +\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n> > +\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n> > +\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n> > +\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n> > +\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n> >  \t\tEOF\n> \n> Curious, but why is the last one special here?\n\nThe `--disk-usage=human` rev-list option here outputs \"byte/bytes\"\ninstead of \"B\". In patch 5, the HUMANISE_COMPACT flag was added to\nhumanise_bytes() to toggle this behavior. For the git-repo(1) structure\ntable output, I wanted to always use the more compact unit prefix\nrepresentation.\n\nI'll leave a comment here to explain this special case.\n\n-Justin\n"},{"id":"532287","messageId":"20251216173842.3357832-1-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251215205639.2700270-1-jltobler@gmail.com","subject":"[PATCH v4 0/7] builtin/repo: add object size info to structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:35Z","receivedAt":"2025-12-16T17:39:16Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThis patch series extends the recently introduced \"structure\" subcommand\nfor git-repo(1) to collect object size information. More specifically,\nit shows total inflated and disk sizes of objects by object type. The\naim to provide additional insight that may be useful to users regarding\nthe structure of a repository.\n\nIn addition to this change, this series also updates the table output\nformat to downscale larger output values along with the appropriate unit\nprefix. This is done to make table output more human friendly. The\nkeyvalue and nul output formats are left the same since they are\nintended more for machine parsing.\n\nChanges in V4:\n- Unmark \"byte\" string in \"t/helper/test-simple-ipc.c\" for translation\n  to avoid conflict with translated plural \"byte/bytes\" string.\n- Remove some unnecessary translations and add comments to clarify some\n  of the added translations.\n- Some small changes to the tests in patch 7.\n\nChanges in V3:\n- Address potential localization regression by making the downscaled\n  number format string also translatable. Also make the format string\n  for how the values and unit prefixes are displayed via\n  `strbuf_humanise_{bytes,rate}()` translatable to be more flexible.\n- `strbuf_humanise_{bytes,count}_value()` has been renamed to\n  `humanise_{bytes,count}()` and updated to provide both the value and\n  unit prefix as separate strings.\n- Unit prefix strings are no longer allocated and instead constant.\n- The humanise flags are now defined in an enum.\n- Instead of using `OBJECT_INFO_FOR_PREFETCH`,\n  `OBJECT_INFO_SKIP_FETCH_OBJECT` and `OBJECT_INFO_QUICK` are used\n  explicitly.\n- Tests now use git-rev-list(1) to verify disk size info.\n\nChanges in V2:\n- Factor out and reuse existing logic from strbuf_humanise() to handle\n  downscaling values and determining the appropriate unit prefix\n  separately. This enables more control over how exactly the values are\n  written to the structure output table which is useful for alignment\n  reasons. I'm not how about the interface used in patch 2. Feedback is\n  most welcome.\n- In the previous version, when checking object size on a missing object\n  we would die. Instead we now ignore missing objects. This allows the\n  structure command to work on partial clones.\n- disk/inflated keyvalue names renamed to disk_size/inflated_size.\n- Unit prefixes are marked for translation.\n- The test for keyvalue disk size values are updated to check against\n  real expected values instead of skipping. Table output tests still\n  skip verifing human-readable values though.\n\nThanks,\n-Justin\n\nJustin Tobler (7):\n  builtin/repo: group per-type object values into struct\n  strbuf: split out logic to humanise byte values\n  builtin/repo: humanise count values in structure output\n  builtin/repo: add inflated object info to keyvalue structure output\n  builtin/repo: add inflated object info to structure table\n  builtin/repo: add disk size info to keyvalue stucture output\n  builtin/repo: add object disk size info to structure table\n\n Documentation/git-repo.adoc |   2 +\n builtin/repo.c              | 175 ++++++++++++++++++++++++++++++------\n strbuf.c                    | 102 ++++++++++++++-------\n strbuf.h                    |  25 ++++++\n t/helper/test-simple-ipc.c  |   7 +-\n t/t1901-repo-structure.sh   | 118 ++++++++++++++++--------\n 6 files changed, 331 insertions(+), 98 deletions(-)\n\nRange-diff against v3:\n1:  be14de68f6 = 1:  be14de68f6 builtin/repo: group per-type object values into struct\n2:  1fa33f5906 ! 2:  0a145cfeec strbuf: split out logic to humanise byte values\n    @@ Commit message\n         determine the corresponding unit prefix into a separate humanise_bytes()\n         function that provides seperate value and unit strings.\n     \n    +    Note that the \"byte\" string in \"t/helper/test-simple-ipc.c\" is unmarked\n    +    for translation here so that it doesn't conflict with the newly defined\n    +    plural \"byte/bytes\" translation and instead uses it.\n    +\n         Signed-off-by: Justin Tobler <jltobler@gmail.com>\n     \n      ## strbuf.c ##\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     -\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n     -\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n     -\t\t\t\t(unsigned)bytes);\n    -+\t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n    ++\t\t*value = xstrfmt(\"%u\", (unsigned)bytes);\n     +\t\t*unit = humanise_rate ?\n     +\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n     +\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     +\tconst char *unit;\n     +\n     +\thumanise_bytes(bytes, &value, &unit, flags);\n    ++\n    ++\t/*\n    ++\t * TRANSLATORS: The first argument is the number string. The second\n    ++\t * argument is the unit prefix string (i.e. \"12.34 MiB/s\").\n    ++\t */\n     +\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n     +\tfree(value);\n     +}\n    @@ strbuf.h: void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbu\n      /**\n       * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n       * 3.50 MiB).\n    +\n    + ## t/helper/test-simple-ipc.c ##\n    +@@ t/helper/test-simple-ipc.c: int cmd__simple_ipc(int argc, const char **argv)\n    + \t\tOPT_INTEGER(0, \"bytecount\", &cl_args.bytecount, N_(\"number of bytes\")),\n    + \t\tOPT_INTEGER(0, \"batchsize\", &cl_args.batchsize, N_(\"number of requests per thread\")),\n    + \n    +-\t\tOPT_STRING(0, \"byte\", &bytevalue, N_(\"byte\"), N_(\"ballast character\")),\n    ++\t\t/*\n    ++\t\t * The \"byte\" string here is not marked for translation and\n    ++\t\t * instead relies on translation in strbuf.c:humanise_bytes() to\n    ++\t\t * avoid conflict with the plural form.\n    ++\t\t */\n    ++\t\tOPT_STRING(0, \"byte\", &bytevalue, \"byte\", N_(\"ballast character\")),\n    + \t\tOPT_STRING(0, \"token\", &cl_args.token, N_(\"token\"), N_(\"command token to send to the server\")),\n    + \n    + \t\tOPT_END()\n3:  8f09f6358e ! 3:  eebf0d917b builtin/repo: humanise count values in structure output\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     +\t\tsize_t x = count + 5000000; /* for rounding */\n     +\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000000),\n     +\t\t\t\t (unsigned)(x % 1000000000 / 10000000));\n    ++\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^9 */\n     +\t\t*unit = _(\"G\");\n     +\t} else if (count >= 1000000) {\n     +\t\tsize_t x = count + 5000; /* for rounding */\n     +\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000),\n     +\t\t\t\t (unsigned)(x % 1000000 / 10000));\n    ++\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^6 */\n     +\t\t*unit = _(\"M\");\n     +\t} else if (count >= 1000) {\n     +\t\tsize_t x = count + 5; /* for rounding */\n     +\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000),\n     +\t\t\t\t (unsigned)(x % 1000 / 10));\n    ++\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^3 */\n     +\t\t*unit = _(\"k\");\n     +\t} else {\n    -+\t\t*value = xstrfmt(_(\"%u\"), (unsigned)count);\n    ++\t\t*value = xstrfmt(\"%u\", (unsigned)count);\n     +\t\t*unit = NULL;\n     +\t}\n     +}\n4:  3f4eabe94f = 4:  37f71cc1bc builtin/repo: add inflated object info to keyvalue structure output\n5:  85d1052100 ! 5:  40edf4c20b builtin/repo: add inflated object info to structure table\n    @@ strbuf.c\n     @@ strbuf.c: void humanise_bytes(off_t bytes, char **value, const char **unit,\n      \t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n      \t} else {\n    - \t\t*value = xstrfmt(_(\"%u\"), (unsigned)bytes);\n    + \t\t*value = xstrfmt(\"%u\", (unsigned)bytes);\n     -\t\t*unit = humanise_rate ?\n     -\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n     -\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n     -\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n     -\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n     +\t\tif (flags & HUMANISE_COMPACT)\n    ++\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second and byte */\n     +\t\t\t*unit = humanise_rate ? _(\"B/s\") : _(\"B\");\n     +\t\telse\n     +\t\t\t*unit = humanise_rate ?\n6:  e9fa9babec = 6:  ba861f37c9 builtin/repo: add disk size info to keyvalue stucture output\n7:  df542c7bdf ! 7:  3118c17ae3 builtin/repo: add object disk size info to structure table\n    @@ t/t1901-repo-structure.sh: test_description='test git repo structure'\n     -\t\t--filter-provided-objects\n     +\tdisk_usage_opt=\"--disk-usage\"\n     +\n    -+\tif [ \"$2\" = \"true\" ]; then\n    ++\tif test \"$2\" = \"true\"\n    ++\tthen\n     +\t\tdisk_usage_opt=\"--disk-usage=human\"\n     +\tfi\n     +\n    -+\tif [ \"$1\" = \"all\" ]; then\n    ++\tif test \"$1\" = \"all\"\n    ++\tthen\n     +\t\tgit rev-list --all --objects $disk_usage_opt\n     +\telse\n     +\t\tgit rev-list --all --objects $disk_usage_opt \\\n    @@ t/t1901-repo-structure.sh: test_expect_success SHA1 'repository with references\n      \t\tgit notes add -m foo &&\n      \n     -\t\tcat >expect <<-\\EOF &&\n    ++\t\t# The tags disk size is handled specially due to the\n    ++\t\t# git-rev-list(1) --disk-usage=human option printing the full\n    ++\t\t# \"byte/bytes\" unit prefix instead of just \"B\".\n     +\t\tcat >expect <<-EOF &&\n      \t\t| Repository structure | Value      |\n      \t\t| -------------------- | ---------- |\n\nbase-commit: e85ae279b0d58edc2f4c3fd5ac391b51e1223985\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532288","messageId":"20251216173842.3357832-2-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v4 1/7] builtin/repo: group per-type object values into struct","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:36Z","receivedAt":"2025-12-16T17:39:16Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The `object_stats` structure stores object counts by type. In a\nsubsequent commit, additional per-type object measurements will also be\nstored. Group per-type object values into a new struct to allow better\nreuse.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 42 +++++++++++++++++++++++++-----------------\n 1 file changed, 25 insertions(+), 17 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 2a653bd3ea..a69699857a 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -202,13 +202,17 @@ struct ref_stats {\n \tsize_t others;\n };\n \n-struct object_stats {\n+struct object_values {\n \tsize_t tags;\n \tsize_t commits;\n \tsize_t trees;\n \tsize_t blobs;\n };\n \n+struct object_stats {\n+\tstruct object_values type_counts;\n+};\n+\n struct repo_structure {\n \tstruct ref_stats refs;\n \tstruct object_stats objects;\n@@ -281,9 +285,9 @@ static inline size_t get_total_reference_count(struct ref_stats *stats)\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n }\n \n-static inline size_t get_total_object_count(struct object_stats *stats)\n+static inline size_t get_total_object_values(struct object_values *values)\n {\n-\treturn stats->tags + stats->commits + stats->trees + stats->blobs;\n+\treturn values->tags + values->commits + values->trees + values->blobs;\n }\n \n static void stats_table_setup_structure(struct stats_table *table,\n@@ -302,14 +306,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_count(objects);\n+\tobject_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n \tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n-\tstats_table_count_addf(table, objects->commits, \"    * %s\", _(\"Commits\"));\n-\tstats_table_count_addf(table, objects->trees, \"    * %s\", _(\"Trees\"));\n-\tstats_table_count_addf(table, objects->blobs, \"    * %s\", _(\"Blobs\"));\n-\tstats_table_count_addf(table, objects->tags, \"    * %s\", _(\"Tags\"));\n+\tstats_table_count_addf(table, objects->type_counts.commits,\n+\t\t\t       \"    * %s\", _(\"Commits\"));\n+\tstats_table_count_addf(table, objects->type_counts.trees,\n+\t\t\t       \"    * %s\", _(\"Trees\"));\n+\tstats_table_count_addf(table, objects->type_counts.blobs,\n+\t\t\t       \"    * %s\", _(\"Blobs\"));\n+\tstats_table_count_addf(table, objects->type_counts.tags,\n+\t\t\t       \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -389,13 +397,13 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \t       (uintmax_t)stats->refs.others, value_delim);\n \n \tprintf(\"objects.commits.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.commits, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.commits, value_delim);\n \tprintf(\"objects.trees.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.trees, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.trees, value_delim);\n \tprintf(\"objects.blobs.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.blobs, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.blobs, value_delim);\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.tags, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n \tfflush(stdout);\n }\n@@ -473,22 +481,22 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n-\t\tstats->tags += oids->nr;\n+\t\tstats->type_counts.tags += oids->nr;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n-\t\tstats->commits += oids->nr;\n+\t\tstats->type_counts.commits += oids->nr;\n \t\tbreak;\n \tcase OBJ_TREE:\n-\t\tstats->trees += oids->nr;\n+\t\tstats->type_counts.trees += oids->nr;\n \t\tbreak;\n \tcase OBJ_BLOB:\n-\t\tstats->blobs += oids->nr;\n+\t\tstats->type_counts.blobs += oids->nr;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n \t}\n \n-\tobject_count = get_total_object_count(stats);\n+\tobject_count = get_total_object_values(&stats->type_counts);\n \tdisplay_progress(data->progress, object_count);\n \n \treturn 0;\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532289","messageId":"20251216173842.3357832-3-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v4 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:37Z","receivedAt":"2025-12-16T17:39:17Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"In a subsequent commit, byte size values displayed in table output for\nthe git-repo(1) \"structure\" subcommand will be shown in a more\nhuman-readable format with the appropriate unit prefixes. For this\nusecase, the downscaled values and unit prefixes must be handled\nseparately to ensure proper column alignment.\n\nSplit out logic from strbuf_humanise() to downscale byte values and\ndetermine the corresponding unit prefix into a separate humanise_bytes()\nfunction that provides seperate value and unit strings.\n\nNote that the \"byte\" string in \"t/helper/test-simple-ipc.c\" is unmarked\nfor translation here so that it doesn't conflict with the newly defined\nplural \"byte/bytes\" translation and instead uses it.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n strbuf.c                   | 74 ++++++++++++++++++++------------------\n strbuf.h                   | 14 ++++++++\n t/helper/test-simple-ipc.c |  7 +++-\n 3 files changed, 60 insertions(+), 35 deletions(-)\n\ndiff --git a/strbuf.c b/strbuf.c\nindex 6c3851a7f8..3fbd375ad6 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,47 +836,53 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n-static void strbuf_humanise(struct strbuf *buf, off_t bytes,\n-\t\t\t\t int humanise_rate)\n+void humanise_bytes(off_t bytes, char **value, const char **unit,\n+\t\t    unsigned flags)\n {\n+\tint humanise_rate = flags & HUMANISE_RATE;\n+\n \tif (bytes > 1 << 30) {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u GiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u GiB/s\"),\n-\t\t\t    (unsigned)(bytes >> 30),\n-\t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(bytes >> 30),\n+\t\t\t\t (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n+\t\t*unit = humanise_rate ? _(\"GiB/s\") : _(\"GiB\");\n \t} else if (bytes > 1 << 20) {\n-\t\tunsigned x = bytes + 5243;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte */\n-\t\t\t\t\t_(\"%u.%2.2u MiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u MiB/s\"),\n-\t\t\t    x >> 20, ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\tunsigned x = bytes + 5243; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 20,\n+\t\t\t\t ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second and mebibyte */\n+\t\t*unit = humanise_rate ? _(\"MiB/s\") : _(\"MiB\");\n \t} else if (bytes > 1 << 10) {\n-\t\tunsigned x = bytes + 5;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u KiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u KiB/s\"),\n-\t\t\t    x >> 10, ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\tunsigned x = bytes + 5; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 10,\n+\t\t\t\t ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second and kibibyte */\n+\t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n \t} else {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n-\t\t\t\t\tQ_(\"%u byte\", \"%u bytes\", bytes) :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n-\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n-\t\t\t\t(unsigned)bytes);\n+\t\t*value = xstrfmt(\"%u\", (unsigned)bytes);\n+\t\t*unit = humanise_rate ?\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n+\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n+\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n \t}\n }\n \n+static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n+{\n+\tchar *value;\n+\tconst char *unit;\n+\n+\thumanise_bytes(bytes, &value, &unit, flags);\n+\n+\t/*\n+\t * TRANSLATORS: The first argument is the number string. The second\n+\t * argument is the unit prefix string (i.e. \"12.34 MiB/s\").\n+\t */\n+\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n+\tfree(value);\n+}\n+\n void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n {\n \tstrbuf_humanise(buf, bytes, 0);\n@@ -884,7 +890,7 @@ void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n \n void strbuf_humanise_rate(struct strbuf *buf, off_t bytes)\n {\n-\tstrbuf_humanise(buf, bytes, 1);\n+\tstrbuf_humanise(buf, bytes, HUMANISE_RATE);\n }\n \n int printf_ln(const char *fmt, ...)\ndiff --git a/strbuf.h b/strbuf.h\nindex a580ac6084..4426163e7e 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -367,6 +367,20 @@ void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n  */\n void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n \n+enum humanise_flags {\n+\t/*\n+\t * Use rate based unit prefixes for humanised values.\n+\t */\n+\tHUMANISE_RATE = (1 << 0),\n+};\n+\n+/**\n+ * Converts the given byte size into a downscaled human-readable value and\n+ * corresponding unit prefix as two separate strings.\n+ */\n+void humanise_bytes(off_t bytes, char **value, const char **unit,\n+\t\t    unsigned flags);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\ndiff --git a/t/helper/test-simple-ipc.c b/t/helper/test-simple-ipc.c\nindex 03cc5eea2c..442ad6b16f 100644\n--- a/t/helper/test-simple-ipc.c\n+++ b/t/helper/test-simple-ipc.c\n@@ -603,7 +603,12 @@ int cmd__simple_ipc(int argc, const char **argv)\n \t\tOPT_INTEGER(0, \"bytecount\", &cl_args.bytecount, N_(\"number of bytes\")),\n \t\tOPT_INTEGER(0, \"batchsize\", &cl_args.batchsize, N_(\"number of requests per thread\")),\n \n-\t\tOPT_STRING(0, \"byte\", &bytevalue, N_(\"byte\"), N_(\"ballast character\")),\n+\t\t/*\n+\t\t * The \"byte\" string here is not marked for translation and\n+\t\t * instead relies on translation in strbuf.c:humanise_bytes() to\n+\t\t * avoid conflict with the plural form.\n+\t\t */\n+\t\tOPT_STRING(0, \"byte\", &bytevalue, \"byte\", N_(\"ballast character\")),\n \t\tOPT_STRING(0, \"token\", &cl_args.token, N_(\"token\"), N_(\"command token to send to the server\")),\n \n \t\tOPT_END()\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532290","messageId":"20251216173842.3357832-5-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v4 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:39Z","receivedAt":"2025-12-16T17:39:18Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The structure subcommand for git-repo(1) outputs basic count information\nfor objects and references. Extend this output to also provide\ninformation regarding total size of inflated objects by object type.\n\nFor now, object size by object type info is only added to the keyvalue\nand nul output formats. In a subsequent commit, this info is also added\nto the table format.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 33 +++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   |  6 +++++-\n 3 files changed, 39 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 70f0a6d2e4..287eee4b93 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -50,6 +50,7 @@ supported:\n +\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n+* Total inflated size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 9c61bc3e17..e207108346 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,8 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hex.h\"\n+#include \"odb.h\"\n #include \"parse-options.h\"\n #include \"path-walk.h\"\n #include \"progress.h\"\n@@ -211,6 +213,7 @@ struct object_values {\n \n struct object_stats {\n \tstruct object_values type_counts;\n+\tstruct object_values inflated_sizes;\n };\n \n struct repo_structure {\n@@ -423,6 +426,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n+\tprintf(\"objects.commits.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -486,6 +498,7 @@ static void structure_count_references(struct ref_stats *stats,\n }\n \n struct count_objects_data {\n+\tstruct object_database *odb;\n \tstruct object_stats *stats;\n \tstruct progress *progress;\n };\n@@ -495,20 +508,39 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n+\tsize_t inflated_total = 0;\n \tsize_t object_count;\n \n+\tfor (size_t i = 0; i < oids->nr; i++) {\n+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n+\t\tunsigned long inflated;\n+\n+\t\toi.sizep = &inflated;\n+\n+\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n+\t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n+\t\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n+\t\t\tcontinue;\n+\n+\t\tinflated_total += inflated;\n+\t}\n+\n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n+\t\tstats->inflated_sizes.tags += inflated_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n+\t\tstats->inflated_sizes.commits += inflated_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n+\t\tstats->inflated_sizes.trees += inflated_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n+\t\tstats->inflated_sizes.blobs += inflated_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n@@ -526,6 +558,7 @@ static void structure_count_objects(struct object_stats *stats,\n {\n \tstruct path_walk_info info = PATH_WALK_INFO_INIT;\n \tstruct count_objects_data data = {\n+\t\t.odb = repo->objects,\n \t\t.stats = stats,\n \t};\n \ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 55fd13ad1b..33237822fd 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -73,7 +73,7 @@ test_expect_success 'repository with references and objects' '\n \t)\n '\n \n-test_expect_success 'keyvalue and nul format' '\n+test_expect_success SHA1 'keyvalue and nul format' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -90,6 +90,10 @@ test_expect_success 'keyvalue and nul format' '\n \t\tobjects.trees.count=42\n \t\tobjects.blobs.count=42\n \t\tobjects.tags.count=1\n+\t\tobjects.commits.inflated_size=9225\n+\t\tobjects.trees.inflated_size=28554\n+\t\tobjects.blobs.inflated_size=453\n+\t\tobjects.tags.inflated_size=132\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532291","messageId":"20251216173842.3357832-4-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v4 3/7] builtin/repo: humanise count values in structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:38Z","receivedAt":"2025-12-16T17:39:19Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The table output format for the git-repo(1) structure subcommand is used\nby default and intended to provide output to users in a human-friendly\nmanner. When the reference/object count values in a repository are\nlarge, it becomes more cumbersome for users to read the values.\n\nFor larger values, update the table output format to instead produce\nmore human-friendly count values that are scaled down with the\nappropriate unit prefix. Output for the keyvalue and nul formats remains\nunchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 38 +++++++++++++++++-------\n strbuf.c                  | 26 ++++++++++++++++\n strbuf.h                  |  6 ++++\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++--------------------\n 4 files changed, 91 insertions(+), 41 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex a69699857a..9c61bc3e17 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -223,6 +223,7 @@ struct stats_table {\n \n \tint name_col_width;\n \tint value_col_width;\n+\tint unit_col_width;\n };\n \n /*\n@@ -230,6 +231,7 @@ struct stats_table {\n  */\n struct stats_table_entry {\n \tchar *value;\n+\tconst char *unit;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -250,11 +252,18 @@ static void stats_table_vaddf(struct stats_table *table,\n \n \tif (name_width > table->name_col_width)\n \t\ttable->name_col_width = name_width;\n-\tif (entry) {\n+\tif (!entry)\n+\t\treturn;\n+\tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n \t\t\ttable->value_col_width = value_width;\n \t}\n+\tif (entry->unit) {\n+\t\tint unit_width = utf8_strwidth(entry->unit);\n+\t\tif (unit_width > table->unit_col_width)\n+\t\t\ttable->unit_col_width = unit_width;\n+\t}\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -273,7 +282,7 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_list ap;\n \n \tCALLOC_ARRAY(entry, 1);\n-\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n+\thumanise_count(value, &entry->value, &entry->unit);\n \n \tva_start(ap, format);\n \tstats_table_vaddf(table, entry, format, ap);\n@@ -324,20 +333,24 @@ static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n \tconst char *value_col_title = _(\"Value\");\n-\tint name_col_width = utf8_strwidth(name_col_title);\n-\tint value_col_width = utf8_strwidth(value_col_title);\n+\tint title_name_width = utf8_strwidth(name_col_title);\n+\tint title_value_width = utf8_strwidth(value_col_title);\n+\tint name_col_width = table->name_col_width;\n+\tint value_col_width = table->value_col_width;\n+\tint unit_col_width = table->unit_col_width;\n \tstruct string_list_item *item;\n \tstruct strbuf buf = STRBUF_INIT;\n \n-\tif (table->name_col_width > name_col_width)\n-\t\tname_col_width = table->name_col_width;\n-\tif (table->value_col_width > value_col_width)\n-\t\tvalue_col_width = table->value_col_width;\n+\tif (title_name_width > name_col_width)\n+\t\tname_col_width = title_name_width;\n+\tif (title_value_width > value_col_width + unit_col_width + 1)\n+\t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n+\t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n \tstrbuf_addstr(&buf, \" |\");\n \tprintf(\"%s\\n\", buf.buf);\n \n@@ -345,17 +358,20 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tfor (int i = 0; i < name_col_width; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n-\tfor (int i = 0; i < value_col_width; i++)\n+\tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n \t\tputchar('-');\n \tprintf(\" |\\n\");\n \n \tfor_each_string_list_item(item, &table->rows) {\n \t\tstruct stats_table_entry *entry = item->util;\n \t\tconst char *value = \"\";\n+\t\tconst char *unit = \"\";\n \n \t\tif (entry) {\n \t\t\tstruct stats_table_entry *entry = item->util;\n \t\t\tvalue = entry->value;\n+\t\t\tif (entry->unit)\n+\t\t\t\tunit = entry->unit;\n \t\t}\n \n \t\tstrbuf_reset(&buf);\n@@ -363,6 +379,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n+\t\tstrbuf_addch(&buf, ' ');\n+\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, unit_col_width, unit);\n \t\tstrbuf_addstr(&buf, \" |\");\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\ndiff --git a/strbuf.c b/strbuf.c\nindex 3fbd375ad6..9beebad5b9 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,6 +836,32 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n+void humanise_count(size_t count, char **value, const char **unit)\n+{\n+\tif (count >= 1000000000) {\n+\t\tsize_t x = count + 5000000; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000000),\n+\t\t\t\t (unsigned)(x % 1000000000 / 10000000));\n+\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^9 */\n+\t\t*unit = _(\"G\");\n+\t} else if (count >= 1000000) {\n+\t\tsize_t x = count + 5000; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000),\n+\t\t\t\t (unsigned)(x % 1000000 / 10000));\n+\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^6 */\n+\t\t*unit = _(\"M\");\n+\t} else if (count >= 1000) {\n+\t\tsize_t x = count + 5; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000),\n+\t\t\t\t (unsigned)(x % 1000 / 10));\n+\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^3 */\n+\t\t*unit = _(\"k\");\n+\t} else {\n+\t\t*value = xstrfmt(\"%u\", (unsigned)count);\n+\t\t*unit = NULL;\n+\t}\n+}\n+\n void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t    unsigned flags)\n {\ndiff --git a/strbuf.h b/strbuf.h\nindex 4426163e7e..571bd889df 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -381,6 +381,12 @@ enum humanise_flags {\n void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t    unsigned flags);\n \n+/**\n+ * Converts the given count into a downscaled human-readable value and\n+ * corresponding unit prefix as two separate strings.\n+ */\n+void humanise_count(size_t count, char **value, const char **unit);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 36a71a144e..55fd13ad1b 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -10,21 +10,21 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Branches       |     0 |\n-\t\t|     * Tags           |     0 |\n-\t\t|     * Remotes        |     0 |\n-\t\t|     * Others         |     0 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Commits        |     0 |\n-\t\t|     * Trees          |     0 |\n-\t\t|     * Blobs          |     0 |\n-\t\t|     * Tags           |     0 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Branches       |     0  |\n+\t\t|     * Tags           |     0  |\n+\t\t|     * Remotes        |     0  |\n+\t\t|     * Others         |     0  |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Commits        |     0  |\n+\t\t|     * Trees          |     0  |\n+\t\t|     * Blobs          |     0  |\n+\t\t|     * Tags           |     0  |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -39,7 +39,7 @@ test_expect_success 'repository with references and objects' '\n \tgit init repo &&\n \t(\n \t\tcd repo &&\n-\t\ttest_commit_bulk 42 &&\n+\t\ttest_commit_bulk 1005 &&\n \t\tgit tag -a foo -m bar &&\n \n \t\toid=\"$(git rev-parse HEAD)\" &&\n@@ -49,21 +49,21 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     4 |\n-\t\t|     * Branches       |     1 |\n-\t\t|     * Tags           |     1 |\n-\t\t|     * Remotes        |     1 |\n-\t\t|     * Others         |     1 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |   130 |\n-\t\t|     * Commits        |    43 |\n-\t\t|     * Trees          |    43 |\n-\t\t|     * Blobs          |    43 |\n-\t\t|     * Tags           |     1 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |    4   |\n+\t\t|     * Branches       |    1   |\n+\t\t|     * Tags           |    1   |\n+\t\t|     * Remotes        |    1   |\n+\t\t|     * Others         |    1   |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            | 3.02 k |\n+\t\t|     * Commits        | 1.01 k |\n+\t\t|     * Trees          | 1.01 k |\n+\t\t|     * Blobs          | 1.01 k |\n+\t\t|     * Tags           |    1   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532292","messageId":"20251216173842.3357832-6-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v4 5/7] builtin/repo: add inflated object info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:40Z","receivedAt":"2025-12-16T17:39:19Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Update the table output format for the git-repo(1) structure command to\nbegin printing the total inflated object size info by object type. To be\nmore human-friendly, larger values are scaled down and displayed with\nthe appropriate unit prefix. Output for the keyvalue and nul formats\nremains unchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 33 +++++++++++++++++++--\n strbuf.c                  | 14 +++++----\n strbuf.h                  |  5 ++++\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++++++----------------\n 4 files changed, 80 insertions(+), 34 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex e207108346..b73cfd975b 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -292,6 +292,20 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_size_addf(struct stats_table *table, size_t value,\n+\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -307,7 +321,8 @@ static void stats_table_setup_structure(struct stats_table *table,\n {\n \tstruct object_stats *objects = &stats->objects;\n \tstruct ref_stats *refs = &stats->refs;\n-\tsize_t object_total;\n+\tsize_t inflated_object_total;\n+\tsize_t object_count_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -318,10 +333,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_values(&objects->type_counts);\n+\tobject_count_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n-\tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n+\tstats_table_count_addf(table, object_count_total, \"  * %s\", _(\"Count\"));\n \tstats_table_count_addf(table, objects->type_counts.commits,\n \t\t\t       \"    * %s\", _(\"Commits\"));\n \tstats_table_count_addf(table, objects->type_counts.trees,\n@@ -330,6 +345,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t       \"    * %s\", _(\"Blobs\"));\n \tstats_table_count_addf(table, objects->type_counts.tags,\n \t\t\t       \"    * %s\", _(\"Tags\"));\n+\n+\tinflated_object_total = get_total_object_values(&objects->inflated_sizes);\n+\tstats_table_size_addf(table, inflated_object_total,\n+\t\t\t      \"  * %s\", _(\"Inflated size\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/strbuf.c b/strbuf.c\nindex 9beebad5b9..512c7ba680 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -886,11 +886,15 @@ void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n \t} else {\n \t\t*value = xstrfmt(\"%u\", (unsigned)bytes);\n-\t\t*unit = humanise_rate ?\n-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n-\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n-\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n+\t\tif (flags & HUMANISE_COMPACT)\n+\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second and byte */\n+\t\t\t*unit = humanise_rate ? _(\"B/s\") : _(\"B\");\n+\t\telse\n+\t\t\t*unit = humanise_rate ?\n+\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n+\t\t\t\t\tQ_(\"byte/s\", \"bytes/s\", bytes) :\n+\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n+\t\t\t\t\tQ_(\"byte\", \"bytes\", bytes);\n \t}\n }\n \ndiff --git a/strbuf.h b/strbuf.h\nindex 571bd889df..005c155808 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -372,6 +372,11 @@ enum humanise_flags {\n \t * Use rate based unit prefixes for humanised values.\n \t */\n \tHUMANISE_RATE = (1 << 0),\n+\t/*\n+\t * Use compact \"B\" unit prefixes instead of \"byte/bytes\" for humanised\n+\t * values.\n+\t */\n+\tHUMANISE_COMPACT = (1 << 1),\n };\n \n /**\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 33237822fd..b18213c660 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -13,18 +13,23 @@ test_expect_success 'empty repository' '\n \t\t| Repository structure | Value  |\n \t\t| -------------------- | ------ |\n \t\t| * References         |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Branches       |     0  |\n-\t\t|     * Tags           |     0  |\n-\t\t|     * Remotes        |     0  |\n-\t\t|     * Others         |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Branches       |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|     * Remotes        |    0   |\n+\t\t|     * Others         |    0   |\n \t\t|                      |        |\n \t\t| * Reachable objects  |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Commits        |     0  |\n-\t\t|     * Trees          |     0  |\n-\t\t|     * Blobs          |     0  |\n-\t\t|     * Tags           |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Commits        |    0   |\n+\t\t|     * Trees          |    0   |\n+\t\t|     * Blobs          |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|   * Inflated size    |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -34,7 +39,7 @@ test_expect_success 'empty repository' '\n \t)\n '\n \n-test_expect_success 'repository with references and objects' '\n+test_expect_success SHA1 'repository with references and objects' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -49,21 +54,26 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    4   |\n-\t\t|     * Branches       |    1   |\n-\t\t|     * Tags           |    1   |\n-\t\t|     * Remotes        |    1   |\n-\t\t|     * Others         |    1   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            | 3.02 k |\n-\t\t|     * Commits        | 1.01 k |\n-\t\t|     * Trees          | 1.01 k |\n-\t\t|     * Blobs          | 1.01 k |\n-\t\t|     * Tags           |    1   |\n+\t\t| Repository structure | Value      |\n+\t\t| -------------------- | ---------- |\n+\t\t| * References         |            |\n+\t\t|   * Count            |      4     |\n+\t\t|     * Branches       |      1     |\n+\t\t|     * Tags           |      1     |\n+\t\t|     * Remotes        |      1     |\n+\t\t|     * Others         |      1     |\n+\t\t|                      |            |\n+\t\t| * Reachable objects  |            |\n+\t\t|   * Count            |   3.02 k   |\n+\t\t|     * Commits        |   1.01 k   |\n+\t\t|     * Trees          |   1.01 k   |\n+\t\t|     * Blobs          |   1.01 k   |\n+\t\t|     * Tags           |      1     |\n+\t\t|   * Inflated size    |  16.03 MiB |\n+\t\t|     * Commits        | 217.92 KiB |\n+\t\t|     * Trees          |  15.81 MiB |\n+\t\t|     * Blobs          |  11.68 KiB |\n+\t\t|     * Tags           |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532293","messageId":"20251216173842.3357832-7-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v4 6/7] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:41Z","receivedAt":"2025-12-16T17:39:20Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, extend the keyvalue and nul output formats of\nthe git-repo(1) structure command to additionally provide info regarding\ntotal object disk sizes by object type.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 18 ++++++++++++++++++\n t/t1901-repo-structure.sh   | 11 ++++++++++-\n 3 files changed, 29 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 287eee4b93..861073f641 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -51,6 +51,7 @@ supported:\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n+* Total disk size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex b73cfd975b..0ed41bf9d4 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -214,6 +214,7 @@ struct object_values {\n struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n+\tstruct object_values disk_sizes;\n };\n \n struct repo_structure {\n@@ -462,6 +463,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n \n+\tprintf(\"objects.commits.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -536,13 +546,16 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n \tsize_t inflated_total = 0;\n+\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\toff_t disk;\n \n \t\toi.sizep = &inflated;\n+\t\toi.disk_sizep = &disk;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n@@ -550,24 +563,29 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tcontinue;\n \n \t\tinflated_total += inflated;\n+\t\tdisk_total += disk;\n \t}\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n \t\tstats->inflated_sizes.tags += inflated_total;\n+\t\tstats->disk_sizes.tags += disk_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n \t\tstats->inflated_sizes.commits += inflated_total;\n+\t\tstats->disk_sizes.commits += disk_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n \t\tstats->inflated_sizes.trees += inflated_total;\n+\t\tstats->disk_sizes.trees += disk_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n \t\tstats->inflated_sizes.blobs += inflated_total;\n+\t\tstats->disk_sizes.blobs += disk_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex b18213c660..dd17caad05 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -4,6 +4,11 @@ test_description='test git repo structure'\n \n . ./test-lib.sh\n \n+object_type_disk_usage() {\n+\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n+\t\t--filter-provided-objects\n+}\n+\n test_expect_success 'empty repository' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n@@ -91,7 +96,7 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\ttest_commit_bulk 42 &&\n \t\tgit tag -a foo -m bar &&\n \n-\t\tcat >expect <<-\\EOF &&\n+\t\tcat >expect <<-EOF &&\n \t\treferences.branches.count=1\n \t\treferences.tags.count=1\n \t\treferences.remotes.count=0\n@@ -104,6 +109,10 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.trees.inflated_size=28554\n \t\tobjects.blobs.inflated_size=453\n \t\tobjects.tags.inflated_size=132\n+\t\tobjects.commits.disk_size=$(object_type_disk_usage commit)\n+\t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n+\t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n+\t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532294","messageId":"20251216173842.3357832-8-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v4 7/7] builtin/repo: add object disk size info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T17:38:42Z","receivedAt":"2025-12-16T17:39:21Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, update the table output format for the\ngit-repo(1) structure command to display the total object disk usage by\nobject type.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 13 +++++++++++++\n t/t1901-repo-structure.sh | 31 ++++++++++++++++++++++++++++---\n 2 files changed, 41 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 0ed41bf9d4..a071d2fdfe 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -324,6 +324,7 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstruct ref_stats *refs = &stats->refs;\n \tsize_t inflated_object_total;\n \tsize_t object_count_total;\n+\tsize_t disk_object_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -358,6 +359,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->inflated_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tdisk_object_total = get_total_object_values(&objects->disk_sizes);\n+\tstats_table_size_addf(table, disk_object_total,\n+\t\t\t      \"  * %s\", _(\"Disk size\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex dd17caad05..1b68525079 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -5,8 +5,20 @@ test_description='test git repo structure'\n . ./test-lib.sh\n \n object_type_disk_usage() {\n-\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n-\t\t--filter-provided-objects\n+\tdisk_usage_opt=\"--disk-usage\"\n+\n+\tif test \"$2\" = \"true\"\n+\tthen\n+\t\tdisk_usage_opt=\"--disk-usage=human\"\n+\tfi\n+\n+\tif test \"$1\" = \"all\"\n+\tthen\n+\t\tgit rev-list --all --objects $disk_usage_opt\n+\telse\n+\t\tgit rev-list --all --objects $disk_usage_opt \\\n+\t\t\t--filter=object:type=$1 --filter-provided-objects\n+\tfi\n }\n \n test_expect_success 'empty repository' '\n@@ -35,6 +47,11 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|   * Disk size        |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -58,7 +75,10 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# Also creates a commit, tree, and blob.\n \t\tgit notes add -m foo &&\n \n-\t\tcat >expect <<-\\EOF &&\n+\t\t# The tags disk size is handled specially due to the\n+\t\t# git-rev-list(1) --disk-usage=human option printing the full\n+\t\t# \"byte/bytes\" unit prefix instead of just \"B\".\n+\t\tcat >expect <<-EOF &&\n \t\t| Repository structure | Value      |\n \t\t| -------------------- | ---------- |\n \t\t| * References         |            |\n@@ -79,6 +99,11 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Trees          |  15.81 MiB |\n \t\t|     * Blobs          |  11.68 KiB |\n \t\t|     * Tags           |    132 B   |\n+\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n+\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532303","messageId":"xmqqqzsu2qxy.fsf@gitster.g","threadId":"64606","inReplyTo":"20251216173842.3357832-3-jltobler@gmail.com","subject":"Re: [PATCH v4 2/7] strbuf: split out logic to humanise byte values","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-16T18:59:37Z","receivedAt":"2025-12-16T18:59:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> +static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n> +{\n> +\tchar *value;\n> +\tconst char *unit;\n> +\n> +\thumanise_bytes(bytes, &value, &unit, flags);\n> +\n> +\t/*\n> +\t * TRANSLATORS: The first argument is the number string. The second\n> +\t * argument is the unit prefix string (i.e. \"12.34 MiB/s\").\n> +\t */\n> +\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n\n\"unit prefix string\"?  Prefix is something that comes before\nsomething else, but this one is at the end.  Simply saying a \"unit\nstring\" would probably be a sufficient fix, perhaps?\n\nI read the changes since the last round, and other than this part,\neverything looked good.\n\nThanks.\n"},{"id":"532306","messageId":"uyuorzpq6mqr2icszhzxswdyxpr3de4762yt5fynlpgmymovje@zzix54kgnwwm","threadId":"64606","inReplyTo":"xmqqqzsu2qxy.fsf@gitster.g","subject":"Re: [PATCH v4 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-16T19:39:39Z","receivedAt":"2025-12-16T19:39:41Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/17 03:59AM, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > +static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n> > +{\n> > +\tchar *value;\n> > +\tconst char *unit;\n> > +\n> > +\thumanise_bytes(bytes, &value, &unit, flags);\n> > +\n> > +\t/*\n> > +\t * TRANSLATORS: The first argument is the number string. The second\n> > +\t * argument is the unit prefix string (i.e. \"12.34 MiB/s\").\n> > +\t */\n> > +\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n> \n> \"unit prefix string\"?  Prefix is something that comes before\n> something else, but this one is at the end.  Simply saying a \"unit\n> string\" would probably be a sufficient fix, perhaps?\n\nYa my bad, the prefix part would be just the Ki, Mi, etc. In this case\nit is the whole unit string. Saying \"unit string\" would be correct. I\ncan send another version fixing this it if you would like.\n\nThanks for the review,\n-Justin\n"},{"id":"532322","messageId":"aUJVyHOCsCjjazB-@pks.im","threadId":"64606","inReplyTo":"20251216173842.3357832-5-jltobler@gmail.com","subject":"Re: [PATCH v4 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-17T07:03:36Z","receivedAt":"2025-12-17T07:03:43Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 16, 2025 at 11:38:39AM -0600, Justin Tobler wrote:\n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index 9c61bc3e17..e207108346 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -495,20 +508,39 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n>  {\n>  \tstruct count_objects_data *data = cb_data;\n>  \tstruct object_stats *stats = data->stats;\n> +\tsize_t inflated_total = 0;\n>  \tsize_t object_count;\n>  \n> +\tfor (size_t i = 0; i < oids->nr; i++) {\n> +\t\tstruct object_info oi = OBJECT_INFO_INIT;\n> +\t\tunsigned long inflated;\n> +\n> +\t\toi.sizep = &inflated;\n> +\n> +\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n> +\t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n> +\t\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n\nTiny nit: there seems to be an extra tab here. This really is only worth\nfixing if you intend to reroll anyway.\n\nPatrick\n"},{"id":"532323","messageId":"aUJVzVp9VB7tDfA-@pks.im","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"Re: [PATCH v4 0/7] builtin/repo: add object size info to structure output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-17T07:03:41Z","receivedAt":"2025-12-17T07:03:47Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 16, 2025 at 11:38:35AM -0600, Justin Tobler wrote:\n> Changes in V4:\n> - Unmark \"byte\" string in \"t/helper/test-simple-ipc.c\" for translation\n>   to avoid conflict with translated plural \"byte/bytes\" string.\n> - Remove some unnecessary translations and add comments to clarify some\n>   of the added translations.\n> - Some small changes to the tests in patch 7.\n\nI had a last tiny nit that doesn't warrant a reroll on its own. Other\nthan that this series looks great to me now. Thanks!\n\nPatrick\n"},{"id":"532372","messageId":"ygljaf4o7mgsvzz6upybtj3fslpdk7a5j3jz3lxjhho4is5cjf@o22or2lcvhep","threadId":"64606","inReplyTo":"aUJVyHOCsCjjazB-@pks.im","subject":"Re: [PATCH v4 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T16:10:43Z","receivedAt":"2025-12-17T16:10:45Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/17 08:03AM, Patrick Steinhardt wrote:\n> On Tue, Dec 16, 2025 at 11:38:39AM -0600, Justin Tobler wrote:\n> > diff --git a/builtin/repo.c b/builtin/repo.c\n> > index 9c61bc3e17..e207108346 100644\n> > --- a/builtin/repo.c\n> > +++ b/builtin/repo.c\n> > @@ -495,20 +508,39 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n> >  {\n> >  \tstruct count_objects_data *data = cb_data;\n> >  \tstruct object_stats *stats = data->stats;\n> > +\tsize_t inflated_total = 0;\n> >  \tsize_t object_count;\n> >  \n> > +\tfor (size_t i = 0; i < oids->nr; i++) {\n> > +\t\tstruct object_info oi = OBJECT_INFO_INIT;\n> > +\t\tunsigned long inflated;\n> > +\n> > +\t\toi.sizep = &inflated;\n> > +\n> > +\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n> > +\t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n> > +\t\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n> \n> Tiny nit: there seems to be an extra tab here. This really is only worth\n> fixing if you intend to reroll anyway.\n\nI had that initially, but it was failing the check_style CI job so I\njust opted to what clang format wanted. I can change it though if I sent\nanother version. I haven't quite figured out the best way to wrap long\nlines.\n\n-Justin\n"},{"id":"532373","messageId":"4zhiuhpvik5w2vgawepbvqfsfukkixjk7ht6ixvzhjfbu4wtzz@tqljiqujmy47","threadId":"64606","inReplyTo":"aUJVzVp9VB7tDfA-@pks.im","subject":"Re: [PATCH v4 0/7] builtin/repo: add object size info to structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:49:59Z","receivedAt":"2025-12-17T17:50:04Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/12/17 08:03AM, Patrick Steinhardt wrote:\n> On Tue, Dec 16, 2025 at 11:38:35AM -0600, Justin Tobler wrote:\n> > Changes in V4:\n> > - Unmark \"byte\" string in \"t/helper/test-simple-ipc.c\" for translation\n> >   to avoid conflict with translated plural \"byte/bytes\" string.\n> > - Remove some unnecessary translations and add comments to clarify some\n> >   of the added translations.\n> > - Some small changes to the tests in patch 7.\n> \n> I had a last tiny nit that doesn't warrant a reroll on its own. Other\n> than that this series looks great to me now. Thanks!\n\nJunio also had some small comments. I'll go ahead a send another\nversion. Thanks for the review. :)\n\n-Justin\n"},{"id":"532374","messageId":"20251217175404.37963-1-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251216173842.3357832-1-jltobler@gmail.com","subject":"[PATCH v5 0/7] builtin/repo: add object size info to structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:53:57Z","receivedAt":"2025-12-17T17:54:09Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Greetings,\n\nThis patch series extends the recently introduced \"structure\" subcommand\nfor git-repo(1) to collect object size information. More specifically,\nit shows total inflated and disk sizes of objects by object type. The\naim to provide additional insight that may be useful to users regarding\nthe structure of a repository.\n\nIn addition to this change, this series also updates the table output\nformat to downscale larger output values along with the appropriate unit\nprefix. This is done to make table output more human friendly. The\nkeyvalue and nul output formats are left the same since they are\nintended more for machine parsing.\n\nChanges in V5:\n- Small updates to some comments and log messages to improve\n  correctness.\n- Adjusted spacing in builtin/repo.c:count_objects().\n\nChanges in V4:\n- Unmark \"byte\" string in \"t/helper/test-simple-ipc.c\" for translation\n  to avoid conflict with translated plural \"byte/bytes\" string.\n- Remove some unnecessary translations and add comments to clarify some\n  of the added translations.\n- Some small changes to the tests in patch 7.\n\nChanges in V3:\n- Address potential localization regression by making the downscaled\n  number format string also translatable. Also make the format string\n  for how the values and unit prefixes are displayed via\n  `strbuf_humanise_{bytes,rate}()` translatable to be more flexible.\n- `strbuf_humanise_{bytes,count}_value()` has been renamed to\n  `humanise_{bytes,count}()` and updated to provide both the value and\n  unit prefix as separate strings.\n- Unit prefix strings are no longer allocated and instead constant.\n- The humanise flags are now defined in an enum.\n- Instead of using `OBJECT_INFO_FOR_PREFETCH`,\n  `OBJECT_INFO_SKIP_FETCH_OBJECT` and `OBJECT_INFO_QUICK` are used\n  explicitly.\n- Tests now use git-rev-list(1) to verify disk size info.\n\nChanges in V2:\n- Factor out and reuse existing logic from strbuf_humanise() to handle\n  downscaling values and determining the appropriate unit prefix\n  separately. This enables more control over how exactly the values are\n  written to the structure output table which is useful for alignment\n  reasons. I'm not how about the interface used in patch 2. Feedback is\n  most welcome.\n- In the previous version, when checking object size on a missing object\n  we would die. Instead we now ignore missing objects. This allows the\n  structure command to work on partial clones.\n- disk/inflated keyvalue names renamed to disk_size/inflated_size.\n- Unit prefixes are marked for translation.\n- The test for keyvalue disk size values are updated to check against\n  real expected values instead of skipping. Table output tests still\n  skip verifing human-readable values though.\n\nThanks,\n-Justin\n\nJustin Tobler (7):\n  builtin/repo: group per-type object values into struct\n  strbuf: split out logic to humanise byte values\n  builtin/repo: humanise count values in structure output\n  builtin/repo: add inflated object info to keyvalue structure output\n  builtin/repo: add inflated object info to structure table\n  builtin/repo: add disk size info to keyvalue stucture output\n  builtin/repo: add object disk size info to structure table\n\n Documentation/git-repo.adoc |   2 +\n builtin/repo.c              | 175 ++++++++++++++++++++++++++++++------\n strbuf.c                    | 102 ++++++++++++++-------\n strbuf.h                    |  25 ++++++\n t/helper/test-simple-ipc.c  |   7 +-\n t/t1901-repo-structure.sh   | 118 ++++++++++++++++--------\n 6 files changed, 331 insertions(+), 98 deletions(-)\n\nRange-diff against v4:\n1:  be14de68f6 = 1:  be14de68f6 builtin/repo: group per-type object values into struct\n2:  0a145cfeec ! 2:  61cff22afa strbuf: split out logic to humanise byte values\n    @@ Commit message\n         In a subsequent commit, byte size values displayed in table output for\n         the git-repo(1) \"structure\" subcommand will be shown in a more\n         human-readable format with the appropriate unit prefixes. For this\n    -    usecase, the downscaled values and unit prefixes must be handled\n    +    usecase, the downscaled values and unit strings must be handled\n         separately to ensure proper column alignment.\n     \n         Split out logic from strbuf_humanise() to downscale byte values and\n    @@ strbuf.c: void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n     +\n     +\t/*\n     +\t * TRANSLATORS: The first argument is the number string. The second\n    -+\t * argument is the unit prefix string (i.e. \"12.34 MiB/s\").\n    ++\t * argument is the unit string (i.e. \"12.34 MiB/s\").\n     +\t */\n     +\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n     +\tfree(value);\n    @@ strbuf.h: void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbu\n      \n     +enum humanise_flags {\n     +\t/*\n    -+\t * Use rate based unit prefixes for humanised values.\n    ++\t * Use rate based units for humanised values.\n     +\t */\n     +\tHUMANISE_RATE = (1 << 0),\n     +};\n     +\n     +/**\n     + * Converts the given byte size into a downscaled human-readable value and\n    -+ * corresponding unit prefix as two separate strings.\n    ++ * corresponding unit as two separate strings.\n     + */\n     +void humanise_bytes(off_t bytes, char **value, const char **unit,\n     +\t\t    unsigned flags);\n3:  eebf0d917b ! 3:  0b575738c2 builtin/repo: humanise count values in structure output\n    @@ strbuf.h: enum humanise_flags {\n      \n     +/**\n     + * Converts the given count into a downscaled human-readable value and\n    -+ * corresponding unit prefix as two separate strings.\n    ++ * corresponding unit as two separate strings.\n     + */\n     +void humanise_count(size_t count, char **value, const char **unit);\n     +\n4:  37f71cc1bc ! 4:  e2c79c8759 builtin/repo: add inflated object info to keyvalue structure output\n    @@ builtin/repo.c: static int count_objects(const char *path UNUSED, struct oid_arr\n     +\n     +\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n     +\t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n    -+\t\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n    ++\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n     +\t\t\tcontinue;\n     +\n     +\t\tinflated_total += inflated;\n5:  40edf4c20b ! 5:  03219630cc builtin/repo: add inflated object info to structure table\n    @@ strbuf.c: void humanise_bytes(off_t bytes, char **value, const char **unit,\n     \n      ## strbuf.h ##\n     @@ strbuf.h: enum humanise_flags {\n    - \t * Use rate based unit prefixes for humanised values.\n    + \t * Use rate based units for humanised values.\n      \t */\n      \tHUMANISE_RATE = (1 << 0),\n     +\t/*\n    -+\t * Use compact \"B\" unit prefixes instead of \"byte/bytes\" for humanised\n    ++\t * Use compact \"B\" unit symbol instead of \"byte/bytes\" for humanised\n     +\t * values.\n     +\t */\n     +\tHUMANISE_COMPACT = (1 << 1),\n6:  ba861f37c9 = 6:  7d8862a064 builtin/repo: add disk size info to keyvalue stucture output\n7:  3118c17ae3 ! 7:  3e2d5c20f8 builtin/repo: add object disk size info to structure table\n    @@ t/t1901-repo-structure.sh: test_expect_success SHA1 'repository with references\n     -\t\tcat >expect <<-\\EOF &&\n     +\t\t# The tags disk size is handled specially due to the\n     +\t\t# git-rev-list(1) --disk-usage=human option printing the full\n    -+\t\t# \"byte/bytes\" unit prefix instead of just \"B\".\n    ++\t\t# \"byte/bytes\" unit string instead of just \"B\".\n     +\t\tcat >expect <<-EOF &&\n      \t\t| Repository structure | Value      |\n      \t\t| -------------------- | ---------- |\n\nbase-commit: e85ae279b0d58edc2f4c3fd5ac391b51e1223985\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532375","messageId":"20251217175404.37963-2-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"[PATCH v5 1/7] builtin/repo: group per-type object values into struct","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:53:58Z","receivedAt":"2025-12-17T17:54:10Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The `object_stats` structure stores object counts by type. In a\nsubsequent commit, additional per-type object measurements will also be\nstored. Group per-type object values into a new struct to allow better\nreuse.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c | 42 +++++++++++++++++++++++++-----------------\n 1 file changed, 25 insertions(+), 17 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 2a653bd3ea..a69699857a 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -202,13 +202,17 @@ struct ref_stats {\n \tsize_t others;\n };\n \n-struct object_stats {\n+struct object_values {\n \tsize_t tags;\n \tsize_t commits;\n \tsize_t trees;\n \tsize_t blobs;\n };\n \n+struct object_stats {\n+\tstruct object_values type_counts;\n+};\n+\n struct repo_structure {\n \tstruct ref_stats refs;\n \tstruct object_stats objects;\n@@ -281,9 +285,9 @@ static inline size_t get_total_reference_count(struct ref_stats *stats)\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n }\n \n-static inline size_t get_total_object_count(struct object_stats *stats)\n+static inline size_t get_total_object_values(struct object_values *values)\n {\n-\treturn stats->tags + stats->commits + stats->trees + stats->blobs;\n+\treturn values->tags + values->commits + values->trees + values->blobs;\n }\n \n static void stats_table_setup_structure(struct stats_table *table,\n@@ -302,14 +306,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_count(objects);\n+\tobject_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n \tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n-\tstats_table_count_addf(table, objects->commits, \"    * %s\", _(\"Commits\"));\n-\tstats_table_count_addf(table, objects->trees, \"    * %s\", _(\"Trees\"));\n-\tstats_table_count_addf(table, objects->blobs, \"    * %s\", _(\"Blobs\"));\n-\tstats_table_count_addf(table, objects->tags, \"    * %s\", _(\"Tags\"));\n+\tstats_table_count_addf(table, objects->type_counts.commits,\n+\t\t\t       \"    * %s\", _(\"Commits\"));\n+\tstats_table_count_addf(table, objects->type_counts.trees,\n+\t\t\t       \"    * %s\", _(\"Trees\"));\n+\tstats_table_count_addf(table, objects->type_counts.blobs,\n+\t\t\t       \"    * %s\", _(\"Blobs\"));\n+\tstats_table_count_addf(table, objects->type_counts.tags,\n+\t\t\t       \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\n@@ -389,13 +397,13 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \t       (uintmax_t)stats->refs.others, value_delim);\n \n \tprintf(\"objects.commits.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.commits, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.commits, value_delim);\n \tprintf(\"objects.trees.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.trees, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.trees, value_delim);\n \tprintf(\"objects.blobs.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.blobs, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.blobs, value_delim);\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n-\t       (uintmax_t)stats->objects.tags, value_delim);\n+\t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n \tfflush(stdout);\n }\n@@ -473,22 +481,22 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n-\t\tstats->tags += oids->nr;\n+\t\tstats->type_counts.tags += oids->nr;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n-\t\tstats->commits += oids->nr;\n+\t\tstats->type_counts.commits += oids->nr;\n \t\tbreak;\n \tcase OBJ_TREE:\n-\t\tstats->trees += oids->nr;\n+\t\tstats->type_counts.trees += oids->nr;\n \t\tbreak;\n \tcase OBJ_BLOB:\n-\t\tstats->blobs += oids->nr;\n+\t\tstats->type_counts.blobs += oids->nr;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n \t}\n \n-\tobject_count = get_total_object_count(stats);\n+\tobject_count = get_total_object_values(&stats->type_counts);\n \tdisplay_progress(data->progress, object_count);\n \n \treturn 0;\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532376","messageId":"20251217175404.37963-3-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"[PATCH v5 2/7] strbuf: split out logic to humanise byte values","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:53:59Z","receivedAt":"2025-12-17T17:54:11Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"In a subsequent commit, byte size values displayed in table output for\nthe git-repo(1) \"structure\" subcommand will be shown in a more\nhuman-readable format with the appropriate unit prefixes. For this\nusecase, the downscaled values and unit strings must be handled\nseparately to ensure proper column alignment.\n\nSplit out logic from strbuf_humanise() to downscale byte values and\ndetermine the corresponding unit prefix into a separate humanise_bytes()\nfunction that provides seperate value and unit strings.\n\nNote that the \"byte\" string in \"t/helper/test-simple-ipc.c\" is unmarked\nfor translation here so that it doesn't conflict with the newly defined\nplural \"byte/bytes\" translation and instead uses it.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n strbuf.c                   | 74 ++++++++++++++++++++------------------\n strbuf.h                   | 14 ++++++++\n t/helper/test-simple-ipc.c |  7 +++-\n 3 files changed, 60 insertions(+), 35 deletions(-)\n\ndiff --git a/strbuf.c b/strbuf.c\nindex 6c3851a7f8..349ee9727a 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,47 +836,53 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n-static void strbuf_humanise(struct strbuf *buf, off_t bytes,\n-\t\t\t\t int humanise_rate)\n+void humanise_bytes(off_t bytes, char **value, const char **unit,\n+\t\t    unsigned flags)\n {\n+\tint humanise_rate = flags & HUMANISE_RATE;\n+\n \tif (bytes > 1 << 30) {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u GiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u GiB/s\"),\n-\t\t\t    (unsigned)(bytes >> 30),\n-\t\t\t    (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(bytes >> 30),\n+\t\t\t\t (unsigned)(bytes & ((1 << 30) - 1)) / 10737419);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 gibibyte/second and gibibyte */\n+\t\t*unit = humanise_rate ? _(\"GiB/s\") : _(\"GiB\");\n \t} else if (bytes > 1 << 20) {\n-\t\tunsigned x = bytes + 5243;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte */\n-\t\t\t\t\t_(\"%u.%2.2u MiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u MiB/s\"),\n-\t\t\t    x >> 20, ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\tunsigned x = bytes + 5243; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 20,\n+\t\t\t\t ((x & ((1 << 20) - 1)) * 100) >> 20);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 mebibyte/second and mebibyte */\n+\t\t*unit = humanise_rate ? _(\"MiB/s\") : _(\"MiB\");\n \t} else if (bytes > 1 << 10) {\n-\t\tunsigned x = bytes + 5;  /* for rounding */\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte */\n-\t\t\t\t\t_(\"%u.%2.2u KiB\") :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second */\n-\t\t\t\t\t_(\"%u.%2.2u KiB/s\"),\n-\t\t\t    x >> 10, ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\tunsigned x = bytes + 5; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), x >> 10,\n+\t\t\t\t ((x & ((1 << 10) - 1)) * 100) >> 10);\n+\t\t/* TRANSLATORS: IEC 80000-13:2008 kibibyte/second and kibibyte */\n+\t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n \t} else {\n-\t\tstrbuf_addf(buf,\n-\t\t\t\thumanise_rate == 0 ?\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n-\t\t\t\t\tQ_(\"%u byte\", \"%u bytes\", bytes) :\n-\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n-\t\t\t\t\tQ_(\"%u byte/s\", \"%u bytes/s\", bytes),\n-\t\t\t\t(unsigned)bytes);\n+\t\t*value = xstrfmt(\"%u\", (unsigned)bytes);\n+\t\t*unit = humanise_rate ?\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n+\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n+\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n+\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n \t}\n }\n \n+static void strbuf_humanise(struct strbuf *buf, off_t bytes, unsigned flags)\n+{\n+\tchar *value;\n+\tconst char *unit;\n+\n+\thumanise_bytes(bytes, &value, &unit, flags);\n+\n+\t/*\n+\t * TRANSLATORS: The first argument is the number string. The second\n+\t * argument is the unit string (i.e. \"12.34 MiB/s\").\n+\t */\n+\tstrbuf_addf(buf, _(\"%s %s\"), value, unit);\n+\tfree(value);\n+}\n+\n void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n {\n \tstrbuf_humanise(buf, bytes, 0);\n@@ -884,7 +890,7 @@ void strbuf_humanise_bytes(struct strbuf *buf, off_t bytes)\n \n void strbuf_humanise_rate(struct strbuf *buf, off_t bytes)\n {\n-\tstrbuf_humanise(buf, bytes, 1);\n+\tstrbuf_humanise(buf, bytes, HUMANISE_RATE);\n }\n \n int printf_ln(const char *fmt, ...)\ndiff --git a/strbuf.h b/strbuf.h\nindex a580ac6084..698b3cc4a5 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -367,6 +367,20 @@ void strbuf_addbuf_percentquote(struct strbuf *dst, const struct strbuf *src);\n  */\n void strbuf_add_percentencode(struct strbuf *dst, const char *src, int flags);\n \n+enum humanise_flags {\n+\t/*\n+\t * Use rate based units for humanised values.\n+\t */\n+\tHUMANISE_RATE = (1 << 0),\n+};\n+\n+/**\n+ * Converts the given byte size into a downscaled human-readable value and\n+ * corresponding unit as two separate strings.\n+ */\n+void humanise_bytes(off_t bytes, char **value, const char **unit,\n+\t\t    unsigned flags);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\ndiff --git a/t/helper/test-simple-ipc.c b/t/helper/test-simple-ipc.c\nindex 03cc5eea2c..442ad6b16f 100644\n--- a/t/helper/test-simple-ipc.c\n+++ b/t/helper/test-simple-ipc.c\n@@ -603,7 +603,12 @@ int cmd__simple_ipc(int argc, const char **argv)\n \t\tOPT_INTEGER(0, \"bytecount\", &cl_args.bytecount, N_(\"number of bytes\")),\n \t\tOPT_INTEGER(0, \"batchsize\", &cl_args.batchsize, N_(\"number of requests per thread\")),\n \n-\t\tOPT_STRING(0, \"byte\", &bytevalue, N_(\"byte\"), N_(\"ballast character\")),\n+\t\t/*\n+\t\t * The \"byte\" string here is not marked for translation and\n+\t\t * instead relies on translation in strbuf.c:humanise_bytes() to\n+\t\t * avoid conflict with the plural form.\n+\t\t */\n+\t\tOPT_STRING(0, \"byte\", &bytevalue, \"byte\", N_(\"ballast character\")),\n \t\tOPT_STRING(0, \"token\", &cl_args.token, N_(\"token\"), N_(\"command token to send to the server\")),\n \n \t\tOPT_END()\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532377","messageId":"20251217175404.37963-5-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"[PATCH v5 4/7] builtin/repo: add inflated object info to keyvalue structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:54:01Z","receivedAt":"2025-12-17T17:54:12Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The structure subcommand for git-repo(1) outputs basic count information\nfor objects and references. Extend this output to also provide\ninformation regarding total size of inflated objects by object type.\n\nFor now, object size by object type info is only added to the keyvalue\nand nul output formats. In a subsequent commit, this info is also added\nto the table format.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 33 +++++++++++++++++++++++++++++++++\n t/t1901-repo-structure.sh   |  6 +++++-\n 3 files changed, 39 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 70f0a6d2e4..287eee4b93 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -50,6 +50,7 @@ supported:\n +\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n+* Total inflated size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 9c61bc3e17..8da321a386 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -2,6 +2,8 @@\n \n #include \"builtin.h\"\n #include \"environment.h\"\n+#include \"hex.h\"\n+#include \"odb.h\"\n #include \"parse-options.h\"\n #include \"path-walk.h\"\n #include \"progress.h\"\n@@ -211,6 +213,7 @@ struct object_values {\n \n struct object_stats {\n \tstruct object_values type_counts;\n+\tstruct object_values inflated_sizes;\n };\n \n struct repo_structure {\n@@ -423,6 +426,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.count%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.type_counts.tags, value_delim);\n \n+\tprintf(\"objects.commits.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -486,6 +498,7 @@ static void structure_count_references(struct ref_stats *stats,\n }\n \n struct count_objects_data {\n+\tstruct object_database *odb;\n \tstruct object_stats *stats;\n \tstruct progress *progress;\n };\n@@ -495,20 +508,39 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n {\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n+\tsize_t inflated_total = 0;\n \tsize_t object_count;\n \n+\tfor (size_t i = 0; i < oids->nr; i++) {\n+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n+\t\tunsigned long inflated;\n+\n+\t\toi.sizep = &inflated;\n+\n+\t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n+\t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n+\t\t\t\t\t\t  OBJECT_INFO_QUICK) < 0)\n+\t\t\tcontinue;\n+\n+\t\tinflated_total += inflated;\n+\t}\n+\n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n+\t\tstats->inflated_sizes.tags += inflated_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n+\t\tstats->inflated_sizes.commits += inflated_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n+\t\tstats->inflated_sizes.trees += inflated_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n+\t\tstats->inflated_sizes.blobs += inflated_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\n@@ -526,6 +558,7 @@ static void structure_count_objects(struct object_stats *stats,\n {\n \tstruct path_walk_info info = PATH_WALK_INFO_INIT;\n \tstruct count_objects_data data = {\n+\t\t.odb = repo->objects,\n \t\t.stats = stats,\n \t};\n \ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 55fd13ad1b..33237822fd 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -73,7 +73,7 @@ test_expect_success 'repository with references and objects' '\n \t)\n '\n \n-test_expect_success 'keyvalue and nul format' '\n+test_expect_success SHA1 'keyvalue and nul format' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -90,6 +90,10 @@ test_expect_success 'keyvalue and nul format' '\n \t\tobjects.trees.count=42\n \t\tobjects.blobs.count=42\n \t\tobjects.tags.count=1\n+\t\tobjects.commits.inflated_size=9225\n+\t\tobjects.trees.inflated_size=28554\n+\t\tobjects.blobs.inflated_size=453\n+\t\tobjects.tags.inflated_size=132\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532378","messageId":"20251217175404.37963-4-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"[PATCH v5 3/7] builtin/repo: humanise count values in structure output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:54:00Z","receivedAt":"2025-12-17T17:54:12Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"The table output format for the git-repo(1) structure subcommand is used\nby default and intended to provide output to users in a human-friendly\nmanner. When the reference/object count values in a repository are\nlarge, it becomes more cumbersome for users to read the values.\n\nFor larger values, update the table output format to instead produce\nmore human-friendly count values that are scaled down with the\nappropriate unit prefix. Output for the keyvalue and nul formats remains\nunchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 38 +++++++++++++++++-------\n strbuf.c                  | 26 ++++++++++++++++\n strbuf.h                  |  6 ++++\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++--------------------\n 4 files changed, 91 insertions(+), 41 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex a69699857a..9c61bc3e17 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -223,6 +223,7 @@ struct stats_table {\n \n \tint name_col_width;\n \tint value_col_width;\n+\tint unit_col_width;\n };\n \n /*\n@@ -230,6 +231,7 @@ struct stats_table {\n  */\n struct stats_table_entry {\n \tchar *value;\n+\tconst char *unit;\n };\n \n static void stats_table_vaddf(struct stats_table *table,\n@@ -250,11 +252,18 @@ static void stats_table_vaddf(struct stats_table *table,\n \n \tif (name_width > table->name_col_width)\n \t\ttable->name_col_width = name_width;\n-\tif (entry) {\n+\tif (!entry)\n+\t\treturn;\n+\tif (entry->value) {\n \t\tint value_width = utf8_strwidth(entry->value);\n \t\tif (value_width > table->value_col_width)\n \t\t\ttable->value_col_width = value_width;\n \t}\n+\tif (entry->unit) {\n+\t\tint unit_width = utf8_strwidth(entry->unit);\n+\t\tif (unit_width > table->unit_col_width)\n+\t\t\ttable->unit_col_width = unit_width;\n+\t}\n }\n \n static void stats_table_addf(struct stats_table *table, const char *format, ...)\n@@ -273,7 +282,7 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_list ap;\n \n \tCALLOC_ARRAY(entry, 1);\n-\tentry->value = xstrfmt(\"%\" PRIuMAX, (uintmax_t)value);\n+\thumanise_count(value, &entry->value, &entry->unit);\n \n \tva_start(ap, format);\n \tstats_table_vaddf(table, entry, format, ap);\n@@ -324,20 +333,24 @@ static void stats_table_print_structure(const struct stats_table *table)\n {\n \tconst char *name_col_title = _(\"Repository structure\");\n \tconst char *value_col_title = _(\"Value\");\n-\tint name_col_width = utf8_strwidth(name_col_title);\n-\tint value_col_width = utf8_strwidth(value_col_title);\n+\tint title_name_width = utf8_strwidth(name_col_title);\n+\tint title_value_width = utf8_strwidth(value_col_title);\n+\tint name_col_width = table->name_col_width;\n+\tint value_col_width = table->value_col_width;\n+\tint unit_col_width = table->unit_col_width;\n \tstruct string_list_item *item;\n \tstruct strbuf buf = STRBUF_INIT;\n \n-\tif (table->name_col_width > name_col_width)\n-\t\tname_col_width = table->name_col_width;\n-\tif (table->value_col_width > value_col_width)\n-\t\tvalue_col_width = table->value_col_width;\n+\tif (title_name_width > name_col_width)\n+\t\tname_col_width = title_name_width;\n+\tif (title_value_width > value_col_width + unit_col_width + 1)\n+\t\tvalue_col_width = title_value_width - unit_col_width;\n \n \tstrbuf_addstr(&buf, \"| \");\n \tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n \tstrbuf_addstr(&buf, \" | \");\n-\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT,\n+\t\t\t  value_col_width + unit_col_width + 1, value_col_title);\n \tstrbuf_addstr(&buf, \" |\");\n \tprintf(\"%s\\n\", buf.buf);\n \n@@ -345,17 +358,20 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tfor (int i = 0; i < name_col_width; i++)\n \t\tputchar('-');\n \tprintf(\" | \");\n-\tfor (int i = 0; i < value_col_width; i++)\n+\tfor (int i = 0; i < value_col_width + unit_col_width + 1; i++)\n \t\tputchar('-');\n \tprintf(\" |\\n\");\n \n \tfor_each_string_list_item(item, &table->rows) {\n \t\tstruct stats_table_entry *entry = item->util;\n \t\tconst char *value = \"\";\n+\t\tconst char *unit = \"\";\n \n \t\tif (entry) {\n \t\t\tstruct stats_table_entry *entry = item->util;\n \t\t\tvalue = entry->value;\n+\t\t\tif (entry->unit)\n+\t\t\t\tunit = entry->unit;\n \t\t}\n \n \t\tstrbuf_reset(&buf);\n@@ -363,6 +379,8 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n \t\tstrbuf_addstr(&buf, \" | \");\n \t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n+\t\tstrbuf_addch(&buf, ' ');\n+\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, unit_col_width, unit);\n \t\tstrbuf_addstr(&buf, \" |\");\n \t\tprintf(\"%s\\n\", buf.buf);\n \t}\ndiff --git a/strbuf.c b/strbuf.c\nindex 349ee9727a..995ff15169 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -836,6 +836,32 @@ void strbuf_addstr_urlencode(struct strbuf *sb, const char *s,\n \tstrbuf_add_urlencode(sb, s, strlen(s), allow_unencoded_fn);\n }\n \n+void humanise_count(size_t count, char **value, const char **unit)\n+{\n+\tif (count >= 1000000000) {\n+\t\tsize_t x = count + 5000000; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000000),\n+\t\t\t\t (unsigned)(x % 1000000000 / 10000000));\n+\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^9 */\n+\t\t*unit = _(\"G\");\n+\t} else if (count >= 1000000) {\n+\t\tsize_t x = count + 5000; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000000),\n+\t\t\t\t (unsigned)(x % 1000000 / 10000));\n+\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^6 */\n+\t\t*unit = _(\"M\");\n+\t} else if (count >= 1000) {\n+\t\tsize_t x = count + 5; /* for rounding */\n+\t\t*value = xstrfmt(_(\"%u.%2.2u\"), (unsigned)(x / 1000),\n+\t\t\t\t (unsigned)(x % 1000 / 10));\n+\t\t/* TRANSLATORS: SI decimal prefix symbol for 10^3 */\n+\t\t*unit = _(\"k\");\n+\t} else {\n+\t\t*value = xstrfmt(\"%u\", (unsigned)count);\n+\t\t*unit = NULL;\n+\t}\n+}\n+\n void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t    unsigned flags)\n {\ndiff --git a/strbuf.h b/strbuf.h\nindex 698b3cc4a5..52feef4c1b 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -381,6 +381,12 @@ enum humanise_flags {\n void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t    unsigned flags);\n \n+/**\n+ * Converts the given count into a downscaled human-readable value and\n+ * corresponding unit as two separate strings.\n+ */\n+void humanise_count(size_t count, char **value, const char **unit);\n+\n /**\n  * Append the given byte size as a human-readable string (i.e. 12.23 KiB,\n  * 3.50 MiB).\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 36a71a144e..55fd13ad1b 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -10,21 +10,21 @@ test_expect_success 'empty repository' '\n \t(\n \t\tcd repo &&\n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Branches       |     0 |\n-\t\t|     * Tags           |     0 |\n-\t\t|     * Remotes        |     0 |\n-\t\t|     * Others         |     0 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |     0 |\n-\t\t|     * Commits        |     0 |\n-\t\t|     * Trees          |     0 |\n-\t\t|     * Blobs          |     0 |\n-\t\t|     * Tags           |     0 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Branches       |     0  |\n+\t\t|     * Tags           |     0  |\n+\t\t|     * Remotes        |     0  |\n+\t\t|     * Others         |     0  |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            |     0  |\n+\t\t|     * Commits        |     0  |\n+\t\t|     * Trees          |     0  |\n+\t\t|     * Blobs          |     0  |\n+\t\t|     * Tags           |     0  |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -39,7 +39,7 @@ test_expect_success 'repository with references and objects' '\n \tgit init repo &&\n \t(\n \t\tcd repo &&\n-\t\ttest_commit_bulk 42 &&\n+\t\ttest_commit_bulk 1005 &&\n \t\tgit tag -a foo -m bar &&\n \n \t\toid=\"$(git rev-parse HEAD)\" &&\n@@ -49,21 +49,21 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value |\n-\t\t| -------------------- | ----- |\n-\t\t| * References         |       |\n-\t\t|   * Count            |     4 |\n-\t\t|     * Branches       |     1 |\n-\t\t|     * Tags           |     1 |\n-\t\t|     * Remotes        |     1 |\n-\t\t|     * Others         |     1 |\n-\t\t|                      |       |\n-\t\t| * Reachable objects  |       |\n-\t\t|   * Count            |   130 |\n-\t\t|     * Commits        |    43 |\n-\t\t|     * Trees          |    43 |\n-\t\t|     * Blobs          |    43 |\n-\t\t|     * Tags           |     1 |\n+\t\t| Repository structure | Value  |\n+\t\t| -------------------- | ------ |\n+\t\t| * References         |        |\n+\t\t|   * Count            |    4   |\n+\t\t|     * Branches       |    1   |\n+\t\t|     * Tags           |    1   |\n+\t\t|     * Remotes        |    1   |\n+\t\t|     * Others         |    1   |\n+\t\t|                      |        |\n+\t\t| * Reachable objects  |        |\n+\t\t|   * Count            | 3.02 k |\n+\t\t|     * Commits        | 1.01 k |\n+\t\t|     * Trees          | 1.01 k |\n+\t\t|     * Blobs          | 1.01 k |\n+\t\t|     * Tags           |    1   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532379","messageId":"20251217175404.37963-6-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"[PATCH v5 5/7] builtin/repo: add inflated object info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:54:02Z","receivedAt":"2025-12-17T17:54:13Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Update the table output format for the git-repo(1) structure command to\nbegin printing the total inflated object size info by object type. To be\nmore human-friendly, larger values are scaled down and displayed with\nthe appropriate unit prefix. Output for the keyvalue and nul formats\nremains unchanged.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 33 +++++++++++++++++++--\n strbuf.c                  | 14 +++++----\n strbuf.h                  |  5 ++++\n t/t1901-repo-structure.sh | 62 +++++++++++++++++++++++----------------\n 4 files changed, 80 insertions(+), 34 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 8da321a386..67d7548b88 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -292,6 +292,20 @@ static void stats_table_count_addf(struct stats_table *table, size_t value,\n \tva_end(ap);\n }\n \n+static void stats_table_size_addf(struct stats_table *table, size_t value,\n+\t\t\t\t  const char *format, ...)\n+{\n+\tstruct stats_table_entry *entry;\n+\tva_list ap;\n+\n+\tCALLOC_ARRAY(entry, 1);\n+\thumanise_bytes(value, &entry->value, &entry->unit, HUMANISE_COMPACT);\n+\n+\tva_start(ap, format);\n+\tstats_table_vaddf(table, entry, format, ap);\n+\tva_end(ap);\n+}\n+\n static inline size_t get_total_reference_count(struct ref_stats *stats)\n {\n \treturn stats->branches + stats->remotes + stats->tags + stats->others;\n@@ -307,7 +321,8 @@ static void stats_table_setup_structure(struct stats_table *table,\n {\n \tstruct object_stats *objects = &stats->objects;\n \tstruct ref_stats *refs = &stats->refs;\n-\tsize_t object_total;\n+\tsize_t inflated_object_total;\n+\tsize_t object_count_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -318,10 +333,10 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstats_table_count_addf(table, refs->remotes, \"    * %s\", _(\"Remotes\"));\n \tstats_table_count_addf(table, refs->others, \"    * %s\", _(\"Others\"));\n \n-\tobject_total = get_total_object_values(&objects->type_counts);\n+\tobject_count_total = get_total_object_values(&objects->type_counts);\n \tstats_table_addf(table, \"\");\n \tstats_table_addf(table, \"* %s\", _(\"Reachable objects\"));\n-\tstats_table_count_addf(table, object_total, \"  * %s\", _(\"Count\"));\n+\tstats_table_count_addf(table, object_count_total, \"  * %s\", _(\"Count\"));\n \tstats_table_count_addf(table, objects->type_counts.commits,\n \t\t\t       \"    * %s\", _(\"Commits\"));\n \tstats_table_count_addf(table, objects->type_counts.trees,\n@@ -330,6 +345,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t       \"    * %s\", _(\"Blobs\"));\n \tstats_table_count_addf(table, objects->type_counts.tags,\n \t\t\t       \"    * %s\", _(\"Tags\"));\n+\n+\tinflated_object_total = get_total_object_values(&objects->inflated_sizes);\n+\tstats_table_size_addf(table, inflated_object_total,\n+\t\t\t      \"  * %s\", _(\"Inflated size\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->inflated_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/strbuf.c b/strbuf.c\nindex 995ff15169..7fb7d12ac0 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -886,11 +886,15 @@ void humanise_bytes(off_t bytes, char **value, const char **unit,\n \t\t*unit = humanise_rate ? _(\"KiB/s\") : _(\"KiB\");\n \t} else {\n \t\t*value = xstrfmt(\"%u\", (unsigned)bytes);\n-\t\t*unit = humanise_rate ?\n-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte/second */\n-\t\t\t       Q_(\"byte/s\", \"bytes/s\", bytes) :\n-\t\t\t       /* TRANSLATORS: IEC 80000-13:2008 byte */\n-\t\t\t       Q_(\"byte\", \"bytes\", bytes);\n+\t\tif (flags & HUMANISE_COMPACT)\n+\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second and byte */\n+\t\t\t*unit = humanise_rate ? _(\"B/s\") : _(\"B\");\n+\t\telse\n+\t\t\t*unit = humanise_rate ?\n+\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte/second */\n+\t\t\t\t\tQ_(\"byte/s\", \"bytes/s\", bytes) :\n+\t\t\t\t\t/* TRANSLATORS: IEC 80000-13:2008 byte */\n+\t\t\t\t\tQ_(\"byte\", \"bytes\", bytes);\n \t}\n }\n \ndiff --git a/strbuf.h b/strbuf.h\nindex 52feef4c1b..06e284f9cc 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -372,6 +372,11 @@ enum humanise_flags {\n \t * Use rate based units for humanised values.\n \t */\n \tHUMANISE_RATE = (1 << 0),\n+\t/*\n+\t * Use compact \"B\" unit symbol instead of \"byte/bytes\" for humanised\n+\t * values.\n+\t */\n+\tHUMANISE_COMPACT = (1 << 1),\n };\n \n /**\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 33237822fd..b18213c660 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -13,18 +13,23 @@ test_expect_success 'empty repository' '\n \t\t| Repository structure | Value  |\n \t\t| -------------------- | ------ |\n \t\t| * References         |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Branches       |     0  |\n-\t\t|     * Tags           |     0  |\n-\t\t|     * Remotes        |     0  |\n-\t\t|     * Others         |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Branches       |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|     * Remotes        |    0   |\n+\t\t|     * Others         |    0   |\n \t\t|                      |        |\n \t\t| * Reachable objects  |        |\n-\t\t|   * Count            |     0  |\n-\t\t|     * Commits        |     0  |\n-\t\t|     * Trees          |     0  |\n-\t\t|     * Blobs          |     0  |\n-\t\t|     * Tags           |     0  |\n+\t\t|   * Count            |    0   |\n+\t\t|     * Commits        |    0   |\n+\t\t|     * Trees          |    0   |\n+\t\t|     * Blobs          |    0   |\n+\t\t|     * Tags           |    0   |\n+\t\t|   * Inflated size    |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -34,7 +39,7 @@ test_expect_success 'empty repository' '\n \t)\n '\n \n-test_expect_success 'repository with references and objects' '\n+test_expect_success SHA1 'repository with references and objects' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n \t(\n@@ -49,21 +54,26 @@ test_expect_success 'repository with references and objects' '\n \t\tgit notes add -m foo &&\n \n \t\tcat >expect <<-\\EOF &&\n-\t\t| Repository structure | Value  |\n-\t\t| -------------------- | ------ |\n-\t\t| * References         |        |\n-\t\t|   * Count            |    4   |\n-\t\t|     * Branches       |    1   |\n-\t\t|     * Tags           |    1   |\n-\t\t|     * Remotes        |    1   |\n-\t\t|     * Others         |    1   |\n-\t\t|                      |        |\n-\t\t| * Reachable objects  |        |\n-\t\t|   * Count            | 3.02 k |\n-\t\t|     * Commits        | 1.01 k |\n-\t\t|     * Trees          | 1.01 k |\n-\t\t|     * Blobs          | 1.01 k |\n-\t\t|     * Tags           |    1   |\n+\t\t| Repository structure | Value      |\n+\t\t| -------------------- | ---------- |\n+\t\t| * References         |            |\n+\t\t|   * Count            |      4     |\n+\t\t|     * Branches       |      1     |\n+\t\t|     * Tags           |      1     |\n+\t\t|     * Remotes        |      1     |\n+\t\t|     * Others         |      1     |\n+\t\t|                      |            |\n+\t\t| * Reachable objects  |            |\n+\t\t|   * Count            |   3.02 k   |\n+\t\t|     * Commits        |   1.01 k   |\n+\t\t|     * Trees          |   1.01 k   |\n+\t\t|     * Blobs          |   1.01 k   |\n+\t\t|     * Tags           |      1     |\n+\t\t|   * Inflated size    |  16.03 MiB |\n+\t\t|     * Commits        | 217.92 KiB |\n+\t\t|     * Trees          |  15.81 MiB |\n+\t\t|     * Blobs          |  11.68 KiB |\n+\t\t|     * Tags           |    132 B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532380","messageId":"20251217175404.37963-7-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"[PATCH v5 6/7] builtin/repo: add disk size info to keyvalue stucture output","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:54:03Z","receivedAt":"2025-12-17T17:54:14Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, extend the keyvalue and nul output formats of\nthe git-repo(1) structure command to additionally provide info regarding\ntotal object disk sizes by object type.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n Documentation/git-repo.adoc |  1 +\n builtin/repo.c              | 18 ++++++++++++++++++\n t/t1901-repo-structure.sh   | 11 ++++++++++-\n 3 files changed, 29 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-repo.adoc b/Documentation/git-repo.adoc\nindex 287eee4b93..861073f641 100644\n--- a/Documentation/git-repo.adoc\n+++ b/Documentation/git-repo.adoc\n@@ -51,6 +51,7 @@ supported:\n * Reference counts categorized by type\n * Reachable object counts categorized by type\n * Total inflated size of reachable objects by type\n+* Total disk size of reachable objects by type\n \n +\n The output format can be chosen through the flag `--format`. Three formats are\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 67d7548b88..7ea051f3af 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -214,6 +214,7 @@ struct object_values {\n struct object_stats {\n \tstruct object_values type_counts;\n \tstruct object_values inflated_sizes;\n+\tstruct object_values disk_sizes;\n };\n \n struct repo_structure {\n@@ -462,6 +463,15 @@ static void structure_keyvalue_print(struct repo_structure *stats,\n \tprintf(\"objects.tags.inflated_size%c%\" PRIuMAX \"%c\", key_delim,\n \t       (uintmax_t)stats->objects.inflated_sizes.tags, value_delim);\n \n+\tprintf(\"objects.commits.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.commits, value_delim);\n+\tprintf(\"objects.trees.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.trees, value_delim);\n+\tprintf(\"objects.blobs.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.blobs, value_delim);\n+\tprintf(\"objects.tags.disk_size%c%\" PRIuMAX \"%c\", key_delim,\n+\t       (uintmax_t)stats->objects.disk_sizes.tags, value_delim);\n+\n \tfflush(stdout);\n }\n \n@@ -536,13 +546,16 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \tstruct count_objects_data *data = cb_data;\n \tstruct object_stats *stats = data->stats;\n \tsize_t inflated_total = 0;\n+\tsize_t disk_total = 0;\n \tsize_t object_count;\n \n \tfor (size_t i = 0; i < oids->nr; i++) {\n \t\tstruct object_info oi = OBJECT_INFO_INIT;\n \t\tunsigned long inflated;\n+\t\toff_t disk;\n \n \t\toi.sizep = &inflated;\n+\t\toi.disk_sizep = &disk;\n \n \t\tif (odb_read_object_info_extended(data->odb, &oids->oid[i], &oi,\n \t\t\t\t\t\t  OBJECT_INFO_SKIP_FETCH_OBJECT |\n@@ -550,24 +563,29 @@ static int count_objects(const char *path UNUSED, struct oid_array *oids,\n \t\t\tcontinue;\n \n \t\tinflated_total += inflated;\n+\t\tdisk_total += disk;\n \t}\n \n \tswitch (type) {\n \tcase OBJ_TAG:\n \t\tstats->type_counts.tags += oids->nr;\n \t\tstats->inflated_sizes.tags += inflated_total;\n+\t\tstats->disk_sizes.tags += disk_total;\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tstats->type_counts.commits += oids->nr;\n \t\tstats->inflated_sizes.commits += inflated_total;\n+\t\tstats->disk_sizes.commits += disk_total;\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tstats->type_counts.trees += oids->nr;\n \t\tstats->inflated_sizes.trees += inflated_total;\n+\t\tstats->disk_sizes.trees += disk_total;\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\tstats->type_counts.blobs += oids->nr;\n \t\tstats->inflated_sizes.blobs += inflated_total;\n+\t\tstats->disk_sizes.blobs += disk_total;\n \t\tbreak;\n \tdefault:\n \t\tBUG(\"invalid object type\");\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex b18213c660..dd17caad05 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -4,6 +4,11 @@ test_description='test git repo structure'\n \n . ./test-lib.sh\n \n+object_type_disk_usage() {\n+\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n+\t\t--filter-provided-objects\n+}\n+\n test_expect_success 'empty repository' '\n \ttest_when_finished \"rm -rf repo\" &&\n \tgit init repo &&\n@@ -91,7 +96,7 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\ttest_commit_bulk 42 &&\n \t\tgit tag -a foo -m bar &&\n \n-\t\tcat >expect <<-\\EOF &&\n+\t\tcat >expect <<-EOF &&\n \t\treferences.branches.count=1\n \t\treferences.tags.count=1\n \t\treferences.remotes.count=0\n@@ -104,6 +109,10 @@ test_expect_success SHA1 'keyvalue and nul format' '\n \t\tobjects.trees.inflated_size=28554\n \t\tobjects.blobs.inflated_size=453\n \t\tobjects.tags.inflated_size=132\n+\t\tobjects.commits.disk_size=$(object_type_disk_usage commit)\n+\t\tobjects.trees.disk_size=$(object_type_disk_usage tree)\n+\t\tobjects.blobs.disk_size=$(object_type_disk_usage blob)\n+\t\tobjects.tags.disk_size=$(object_type_disk_usage tag)\n \t\tEOF\n \n \t\tgit repo structure --format=keyvalue >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532381","messageId":"20251217175404.37963-8-jltobler@gmail.com","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"[PATCH v5 7/7] builtin/repo: add object disk size info to structure table","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-12-17T17:54:04Z","receivedAt":"2025-12-17T17:54:14Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"Similar to a prior commit, update the table output format for the\ngit-repo(1) structure command to display the total object disk usage by\nobject type.\n\nSigned-off-by: Justin Tobler <jltobler@gmail.com>\n---\n builtin/repo.c            | 13 +++++++++++++\n t/t1901-repo-structure.sh | 31 ++++++++++++++++++++++++++++---\n 2 files changed, 41 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 7ea051f3af..09bc8fccfd 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -324,6 +324,7 @@ static void stats_table_setup_structure(struct stats_table *table,\n \tstruct ref_stats *refs = &stats->refs;\n \tsize_t inflated_object_total;\n \tsize_t object_count_total;\n+\tsize_t disk_object_total;\n \tsize_t ref_total;\n \n \tref_total = get_total_reference_count(refs);\n@@ -358,6 +359,18 @@ static void stats_table_setup_structure(struct stats_table *table,\n \t\t\t      \"    * %s\", _(\"Blobs\"));\n \tstats_table_size_addf(table, objects->inflated_sizes.tags,\n \t\t\t      \"    * %s\", _(\"Tags\"));\n+\n+\tdisk_object_total = get_total_object_values(&objects->disk_sizes);\n+\tstats_table_size_addf(table, disk_object_total,\n+\t\t\t      \"  * %s\", _(\"Disk size\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.commits,\n+\t\t\t      \"    * %s\", _(\"Commits\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.trees,\n+\t\t\t      \"    * %s\", _(\"Trees\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.blobs,\n+\t\t\t      \"    * %s\", _(\"Blobs\"));\n+\tstats_table_size_addf(table, objects->disk_sizes.tags,\n+\t\t\t      \"    * %s\", _(\"Tags\"));\n }\n \n static void stats_table_print_structure(const struct stats_table *table)\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex dd17caad05..435fd979fa 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -5,8 +5,20 @@ test_description='test git repo structure'\n . ./test-lib.sh\n \n object_type_disk_usage() {\n-\tgit rev-list --all --objects --disk-usage --filter=object:type=$1 \\\n-\t\t--filter-provided-objects\n+\tdisk_usage_opt=\"--disk-usage\"\n+\n+\tif test \"$2\" = \"true\"\n+\tthen\n+\t\tdisk_usage_opt=\"--disk-usage=human\"\n+\tfi\n+\n+\tif test \"$1\" = \"all\"\n+\tthen\n+\t\tgit rev-list --all --objects $disk_usage_opt\n+\telse\n+\t\tgit rev-list --all --objects $disk_usage_opt \\\n+\t\t\t--filter=object:type=$1 --filter-provided-objects\n+\tfi\n }\n \n test_expect_success 'empty repository' '\n@@ -35,6 +47,11 @@ test_expect_success 'empty repository' '\n \t\t|     * Trees          |    0 B |\n \t\t|     * Blobs          |    0 B |\n \t\t|     * Tags           |    0 B |\n+\t\t|   * Disk size        |    0 B |\n+\t\t|     * Commits        |    0 B |\n+\t\t|     * Trees          |    0 B |\n+\t\t|     * Blobs          |    0 B |\n+\t\t|     * Tags           |    0 B |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n@@ -58,7 +75,10 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t# Also creates a commit, tree, and blob.\n \t\tgit notes add -m foo &&\n \n-\t\tcat >expect <<-\\EOF &&\n+\t\t# The tags disk size is handled specially due to the\n+\t\t# git-rev-list(1) --disk-usage=human option printing the full\n+\t\t# \"byte/bytes\" unit string instead of just \"B\".\n+\t\tcat >expect <<-EOF &&\n \t\t| Repository structure | Value      |\n \t\t| -------------------- | ---------- |\n \t\t| * References         |            |\n@@ -79,6 +99,11 @@ test_expect_success SHA1 'repository with references and objects' '\n \t\t|     * Trees          |  15.81 MiB |\n \t\t|     * Blobs          |  11.68 KiB |\n \t\t|     * Tags           |    132 B   |\n+\t\t|   * Disk size        | $(object_type_disk_usage all true) |\n+\t\t|     * Commits        | $(object_type_disk_usage commit true) |\n+\t\t|     * Trees          | $(object_type_disk_usage tree true) |\n+\t\t|     * Blobs          |  $(object_type_disk_usage blob true) |\n+\t\t|     * Tags           |    $(object_type_disk_usage tag) B   |\n \t\tEOF\n \n \t\tgit repo structure >out 2>err &&\n-- \n2.52.0.209.ge85ae279b0\n\n"},{"id":"532423","messageId":"aUOf9hiIWYXgWJ1o@pks.im","threadId":"64606","inReplyTo":"20251217175404.37963-1-jltobler@gmail.com","subject":"Re: [PATCH v5 0/7] builtin/repo: add object size info to structure output","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-18T06:32:22Z","receivedAt":"2025-12-18T06:32:28Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Dec 17, 2025 at 11:53:57AM -0600, Justin Tobler wrote:\n> Greetings,\n> \n> This patch series extends the recently introduced \"structure\" subcommand\n> for git-repo(1) to collect object size information. More specifically,\n> it shows total inflated and disk sizes of objects by object type. The\n> aim to provide additional insight that may be useful to users regarding\n> the structure of a repository.\n> \n> In addition to this change, this series also updates the table output\n> format to downscale larger output values along with the appropriate unit\n> prefix. This is done to make table output more human friendly. The\n> keyvalue and nul output formats are left the same since they are\n> intended more for machine parsing.\n> \n> Changes in V5:\n> - Small updates to some comments and log messages to improve\n>   correctness.\n> - Adjusted spacing in builtin/repo.c:count_objects().\n\nI'm happy with this version, thanks!\n\nPatrick\n"}]}