{"thread":{"id":"55062","subject":"[PATCH 0/2] rev-list --disk-usage","startedAt":"2021-01-27T22:12:32Z","lastAt":"2021-02-17T23:44:59Z","messageCount":30,"participants":["Jeff King","Taylor Blau","Eric Sunshine","Kyle Meyer","Junio C Hamano","Ævar Arnfjörð Bjarmason"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"415459","messageId":"YBHlGPBSJC++CnPy@coredump.intra.peff.net","threadId":"55062","inReplyTo":null,"subject":"[PATCH 0/2] rev-list --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-27T22:11:36Z","receivedAt":"2021-01-27T22:12:32Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This series teaches rev-list to compute the on-disk size used by a set\nof objects. You can do the same thing with cat-file, but this is much\nfaster (see the timings in the second commit).\n\nWe've been running it for about 5 years at GitHub. I hesitated sending\nit upstream because it's a bit weird and special-purpose. But it does\ncome in handy for debugging, analyzing repos, etc. So maybe others will\nfind it useful.\n\nThe first patch is just a test-script enhancement to let test_commit\navoid creating tags. During some recent refactoring, we actually broke\nthe --disk-usage feature but the test script didn't catch it because the\ntags were being picked up by \"--all\". Since this is at least the third\ntime I've run into that in our test suite, I thought I'd make it a\nlittle more convenient to avoid. :)\n\n  [1/2]: t: add --no-tag option to test_commit\n  [2/2]: rev-list: add --disk-usage option for calculating disk usage\n\n Documentation/rev-list-options.txt |  9 ++++++\n builtin/rev-list.c                 | 49 ++++++++++++++++++++++++++++\n pack-bitmap.c                      | 50 +++++++++++++++++++++++++++++\n pack-bitmap.h                      |  2 ++\n t/t4208-log-magic-pathspec.sh      |  9 ++----\n t/t6114-rev-list-du.sh             | 51 ++++++++++++++++++++++++++++++\n t/test-lib-functions.sh            |  9 +++++-\n 7 files changed, 171 insertions(+), 8 deletions(-)\n create mode 100755 t/t6114-rev-list-du.sh\n\n-Peff\n"},{"id":"415460","messageId":"YBHlSQ6cSJWHWeWo@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YBHlGPBSJC++CnPy@coredump.intra.peff.net","subject":"[PATCH 1/2] t: add --no-tag option to test_commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-27T22:12:25Z","receivedAt":"2021-01-27T22:13:13Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"One of the conveniences that test_commit offers is making a tag for each\ncommit. This makes it easy to refer to the commits in subsequent\ncommands. But it can also be a pain if you care about reachability,\nbecause those tags keep the commits reachable even if they are rewound\nfrom the branch they're made on.\n\nThe alternative is that scripts have to call test_tick, git-add, and\ngit-commit themselves. Let's add a --no-tag option to give them the\none-liner convenience of using test_commit.\n\nThis is in preparation for the next patch, which will add some more\ncalls. But I cleaned up an existing site to show off the feature. There\nare probably more cleanups possible.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n t/t4208-log-magic-pathspec.sh | 9 ++-------\n t/test-lib-functions.sh       | 9 ++++++++-\n 2 files changed, 10 insertions(+), 8 deletions(-)\n\ndiff --git a/t/t4208-log-magic-pathspec.sh b/t/t4208-log-magic-pathspec.sh\nindex 5e10136e9a..7f0c1dcc0f 100755\n--- a/t/t4208-log-magic-pathspec.sh\n+++ b/t/t4208-log-magic-pathspec.sh\n@@ -31,13 +31,8 @@ test_expect_success '\"git log :/a -- \" should not be ambiguous' '\n test_expect_success '\"git log :/detached -- \" should find a commit only in HEAD' '\n \ttest_when_finished \"git checkout main\" &&\n \tgit checkout --detach &&\n-\t# Must manually call `test_tick` instead of using `test_commit`,\n-\t# because the latter additionally creates a tag, which would make\n-\t# the commit reachable not only via HEAD.\n-\ttest_tick &&\n-\tgit commit --allow-empty -m detached &&\n-\ttest_tick &&\n-\tgit commit --allow-empty -m something-else &&\n+\ttest_commit --no-tag detached &&\n+\ttest_commit --no-tag something-else &&\n \tgit log :/detached --\n '\n \ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 6bca002316..1587241ba0 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -202,6 +202,7 @@ test_commit () {\n \tauthor= &&\n \tsignoff= &&\n \tindir= &&\n+\tno_tag= &&\n \twhile test $# != 0\n \tdo\n \t\tcase \"$1\" in\n@@ -222,6 +223,9 @@ test_commit () {\n \t\t\tindir=\"$2\"\n \t\t\tshift\n \t\t\t;;\n+\t\t--no-tag)\n+\t\t\tno_tag=yes\n+\t\t\t;;\n \t\t*)\n \t\t\tbreak\n \t\t\t;;\n@@ -244,7 +248,10 @@ test_commit () {\n \tgit ${indir:+ -C \"$indir\"} commit \\\n \t    ${author:+ --author \"$author\"} \\\n \t    $signoff -m \"$1\" &&\n-\tgit ${indir:+ -C \"$indir\"} tag \"${4:-$1}\"\n+\tif test -z \"$no_tag\"\n+\tthen\n+\t\tgit ${indir:+ -C \"$indir\"} tag \"${4:-$1}\"\n+\tfi\n }\n \n # Call test_merge with the arguments \"<message> <commit>\", where <commit>\n-- \n2.30.0.758.g9692d13bf2\n\n"},{"id":"415461","messageId":"YBHmY7vNxu2hqOa/@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YBHlGPBSJC++CnPy@coredump.intra.peff.net","subject":"[PATCH 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-27T22:17:07Z","receivedAt":"2021-01-27T22:17:56Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"It can sometimes be useful to see which refs are contributing to the\noverall repository size (e.g., does some branch have a bunch of objects\nnot found elsewhere in history, which indicates that deleting it would\nshrink the size of a clone).\n\nYou can find that out by generating a list of objects, getting their\nsizes from cat-file, and then summing them, like:\n\n    git rev-list --objects main..branch\n    cut -d' ' -f1 |\n    git cat-file --batch-check='%(objectsize:disk)' |\n    perl -lne '$total += $_; END { print $total }'\n\nThough note that the caveats from git-cat-file(1) apply here. We \"blame\"\nbase objects more than their deltas, even though the relationship could\neasily be flipped. Still, it can be a useful rough measure.\n\nBut one problem is that it's slow to run. Teaching rev-list to sum up\nthe sizes can be much faster for two reasons:\n\n  1. It skips all of the piping of object names and sizes.\n\n  2. If bitmaps are in use, for objects that are in the\n     bitmapped packfile we can skip the oid_object_info()\n     lookup entirely, and just ask the revindex for the\n     on-disk size.\n\nThis patch implements a --disk-usage option which produces the same\nanswer in a fraction of the time. Here are some timings using a clone of\ntorvalds/linux:\n\n  [rev-list piped to cat-file, no bitmaps]\n  $ time git rev-list --objects --all |\n    cut -d' ' -f1 |\n    git cat-file --buffer --batch-check='%(objectsize:disk)' |\n    perl -lne '$total += $_; END { print $total }'\n  1455691059\n  real\t0m34.336s\n  user\t0m46.533s\n  sys\t0m2.953s\n\n  [internal, no bitmaps]\n  $ time git rev-list --disk-usage --all\n  1455691059\n  real\t0m32.662s\n  user\t0m32.306s\n  sys\t0m0.353s\n\nThe wall-clock times aren't that different because of parallelism, but\nnotice the CPU savings between the two. We saved 35% of the CPU just by\navoiding the pipes.\n\nBut the real win is with bitmaps. If we use them without the new option:\n\n  [rev-list piped to cat-file, bitmaps]\n  $ time git rev-list --objects --all --use-bitmap-index |\n    cut -d' ' -f1 |\n    git cat-file --batch-check='%(objectsize:disk)' |\n    perl -lne '$total += $_; END { print $total }'\n  real\t0m9.954s\n  user\t0m11.234s\n  sys\t0m8.522s\n\nthen we're faster to generate the list of objects, but we still spend a\nlot of time piping and looking things up. But if we do both together:\n\n  [internal, bitmaps]\n  $ time git rev-list --disk-usage --all --use-bitmap-index\n  1455691059\n  real\t0m0.235s\n  user\t0m0.186s\n  sys\t0m0.049s\n\nthen we get the same answer much faster.\n\nFor \"--all\", that answer will correspond closely to \"du objects/pack\",\nof course. But we're actually checking reachability here, so we're still\nfast when we ask for more interesting things:\n\n  $ time git rev-list --disk-usage --all --use-bitmap-index v5.0..v5.10\n  374798628\n  real\t0m0.429s\n  user\t0m0.356s\n  sys\t0m0.072s\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\nThis _could_ be made more flexible, but I didn't think it was worth the\ncomplexity. Some obvious things one might want are:\n\n  - not counting up all reachable objects (i.e., requiring --objects for\n    this output, and omitting it just counts up commits). This could be\n    handled in the bitmap case with some extra code (OR-ing with the\n    type bitmaps).\n\n    But after 5 years of this patch, I've never wanted that once. The\n    disk usage of just some of the objects isn't really that useful (and\n    of course you can still get it by piping to cat-file).\n\n  - an option to output the sizes of specific objects along with their\n    oids. But if you want to get to this level of flexibility, I think\n    you're better off just using cat-file (and if we are concerned about\n    the pipe costs, we should teach rev-list to understand cat-file's\n    custom formats).\n\n Documentation/rev-list-options.txt |  9 ++++++\n builtin/rev-list.c                 | 49 ++++++++++++++++++++++++++++\n pack-bitmap.c                      | 50 +++++++++++++++++++++++++++++\n pack-bitmap.h                      |  2 ++\n t/t6114-rev-list-du.sh             | 51 ++++++++++++++++++++++++++++++\n 5 files changed, 161 insertions(+)\n create mode 100755 t/t6114-rev-list-du.sh\n\ndiff --git a/Documentation/rev-list-options.txt b/Documentation/rev-list-options.txt\nindex 002379056a..1e5826f26d 100644\n--- a/Documentation/rev-list-options.txt\n+++ b/Documentation/rev-list-options.txt\n@@ -222,6 +222,15 @@ ifdef::git-rev-list[]\n \ttest the exit status to see if a range of objects is fully\n \tconnected (or not).  It is faster than redirecting stdout\n \tto `/dev/null` as the output does not have to be formatted.\n+\n+--disk-usage::\n+\tSuppress normal output; instead, print the sum of the bytes used\n+\tfor on-disk storage by the selected objects. This is equivalent\n+\tto piping the output of `rev-list --objects` into\n+\t`git cat-file --batch-check='%(objectsize:disk)', except that it\n+\truns much faster (especially with `--use-bitmap-index`). See the\n+\t`CAVEATS` section in linkgit:git-cat-file[1] for the limitations\n+\tof what \"on-disk storage\" means.\n endif::git-rev-list[]\n \n --cherry-mark::\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex 25c6c3b38d..2262b613dd 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -80,6 +80,19 @@ static int arg_show_object_names = 1;\n \n #define DEFAULT_OIDSET_SIZE     (16*1024)\n \n+static int show_disk_usage;\n+static off_t total_disk_usage;\n+\n+static off_t get_object_disk_usage(struct object *obj)\n+{\n+\toff_t size;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\toi.disk_sizep = &size;\n+\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n+\t\tdie(_(\"unable to get disk usage of %s\"), oid_to_hex(&obj->oid));\n+\treturn size;\n+}\n+\n static void finish_commit(struct commit *commit);\n static void show_commit(struct commit *commit, void *data)\n {\n@@ -88,6 +101,9 @@ static void show_commit(struct commit *commit, void *data)\n \n \tdisplay_progress(progress, ++progress_counter);\n \n+\tif (show_disk_usage)\n+\t\ttotal_disk_usage += get_object_disk_usage(&commit->object);\n+\n \tif (info->flags & REV_LIST_QUIET) {\n \t\tfinish_commit(commit);\n \t\treturn;\n@@ -258,6 +274,8 @@ static void show_object(struct object *obj, const char *name, void *cb_data)\n \tif (finish_object(obj, name, cb_data))\n \t\treturn;\n \tdisplay_progress(progress, ++progress_counter);\n+\tif (show_disk_usage)\n+\t\ttotal_disk_usage += get_object_disk_usage(obj);\n \tif (info->flags & REV_LIST_QUIET)\n \t\treturn;\n \n@@ -452,6 +470,23 @@ static int try_bitmap_traversal(struct rev_info *revs,\n \treturn 0;\n }\n \n+static int try_bitmap_disk_usage(struct rev_info *revs,\n+\t\t\t\t struct list_objects_filter_options *filter)\n+{\n+\tstruct bitmap_index *bitmap_git;\n+\n+\tif (!show_disk_usage)\n+\t\treturn -1;\n+\n+\tbitmap_git = prepare_bitmap_walk(revs, filter);\n+\tif (!bitmap_git)\n+\t\treturn -1;\n+\n+\tprintf(\"%\"PRIuMAX\"\\n\",\n+\t       (uintmax_t)get_disk_usage_from_bitmap(bitmap_git));\n+\treturn 0;\n+}\n+\n int cmd_rev_list(int argc, const char **argv, const char *prefix)\n {\n \tstruct rev_info revs;\n@@ -584,6 +619,15 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\tcontinue;\n \t\t}\n \n+\t\tif (!strcmp(arg, \"--disk-usage\")) {\n+\t\t\tshow_disk_usage = 1;\n+\t\t\trevs.tag_objects = 1;\n+\t\t\trevs.tree_objects = 1;\n+\t\t\trevs.blob_objects = 1;\n+\t\t\tinfo.flags |= REV_LIST_QUIET;\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tusage(rev_list_usage);\n \n \t}\n@@ -626,6 +670,8 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \tif (use_bitmap_index) {\n \t\tif (!try_bitmap_count(&revs, &filter_options))\n \t\t\treturn 0;\n+\t\tif (!try_bitmap_disk_usage(&revs, &filter_options))\n+\t\t\treturn 0;\n \t\tif (!try_bitmap_traversal(&revs, &filter_options))\n \t\t\treturn 0;\n \t}\n@@ -690,5 +736,8 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\tprintf(\"%d\\n\", revs.count_left + revs.count_right);\n \t}\n \n+\tif (show_disk_usage)\n+\t\tprintf(\"%\"PRIuMAX\"\\n\", (uintmax_t)total_disk_usage);\n+\n \treturn 0;\n }\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 60fe20fb87..ba36b9c6a0 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -1430,3 +1430,53 @@ int bitmap_has_oid_in_uninteresting(struct bitmap_index *bitmap_git,\n \treturn bitmap_git &&\n \t\tbitmap_walk_contains(bitmap_git, bitmap_git->haves, oid);\n }\n+\n+off_t get_disk_usage_from_bitmap(struct bitmap_index *bitmap_git)\n+{\n+\tstruct bitmap *result = bitmap_git->result;\n+\tstruct packed_git *pack = bitmap_git->pack;\n+\tstruct eindex *eindex = &bitmap_git->ext_index;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\toff_t object_size;\n+\toff_t total = 0;\n+\tsize_t i;\n+\n+\toi.disk_sizep = &object_size;\n+\n+\tfor (i = 0; i < result->word_alloc; i++) {\n+\t\teword_t word = result->words[i];\n+\t\tsize_t base = (i * BITS_IN_EWORD);\n+\t\tunsigned offset;\n+\n+\t\tfor (offset = 0; offset < BITS_IN_EWORD; offset++) {\n+\t\t\tsize_t pos;\n+\n+\t\t\tif ((word >> offset) == 0)\n+\t\t\t\tbreak;\n+\n+\t\t\toffset += ewah_bit_ctz64(word >> offset);\n+\t\t\tpos = base + offset;\n+\n+\t\t\t/*\n+\t\t\t * If it's in the pack, we can use the fast path\n+\t\t\t * and just check the revindex. Otherwise, we\n+\t\t\t * fall back to looking it up.\n+\t\t\t */\n+\t\t\tif (pos < pack->num_objects) {\n+\t\t\t\tobject_size =\n+\t\t\t\t\tpack_pos_to_offset(pack, pos + 1) -\n+\t\t\t\t\tpack_pos_to_offset(pack, pos);\n+\t\t\t} else {\n+\t\t\t\tstruct object *obj;\n+\t\t\t\tobj = eindex->objects[pos - pack->num_objects];\n+\t\t\t\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n+\t\t\t\t\tdie(_(\"unable to get disk usage of %s\"),\n+\t\t\t\t\t      oid_to_hex(&obj->oid));\n+\t\t\t}\n+\n+\t\t\ttotal += object_size;\n+\t\t}\n+\t}\n+\n+\treturn total;\n+}\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nindex 25dfcf5615..c8070606b7 100644\n--- a/pack-bitmap.h\n+++ b/pack-bitmap.h\n@@ -68,6 +68,8 @@ int bitmap_walk_contains(struct bitmap_index *,\n  */\n int bitmap_has_oid_in_uninteresting(struct bitmap_index *, const struct object_id *oid);\n \n+off_t get_disk_usage_from_bitmap(struct bitmap_index *);\n+\n void bitmap_writer_show_progress(int show);\n void bitmap_writer_set_checksum(unsigned char *sha1);\n void bitmap_writer_build_type_index(struct packing_data *to_pack,\ndiff --git a/t/t6114-rev-list-du.sh b/t/t6114-rev-list-du.sh\nnew file mode 100755\nindex 0000000000..1fadbcaded\n--- /dev/null\n+++ b/t/t6114-rev-list-du.sh\n@@ -0,0 +1,51 @@\n+#!/bin/sh\n+\n+test_description='basic tests of rev-list --disk-usage'\n+. ./test-lib.sh\n+\n+# we want a mix of reachable and unreachable, as well as\n+# objects in the bitmapped pack and some outside of it\n+test_expect_success 'set up repository' '\n+\ttest_commit --no-tag one &&\n+\ttest_commit --no-tag two &&\n+\tgit repack -adb &&\n+\tgit reset --hard HEAD^ &&\n+\ttest_commit --no-tag three &&\n+\ttest_commit --no-tag four &&\n+\tgit reset --hard HEAD^\n+'\n+\n+# We don't want to hardcode sizes, because they depend on the exact details of\n+# packing, zlib, etc. We'll assume that the regular rev-list and cat-file\n+# machinery works and compare the --disk-usage output to that.\n+disk_usage_slow () {\n+\tgit rev-list --objects \"$@\" |\n+\tcut -d' ' -f1 |\n+\tgit cat-file --batch-check=\"%(objectsize:disk)\" |\n+\tperl -lne '$total += $_; END { print $total}'\n+}\n+\n+# check behavior with given rev-list options; note that\n+# whitespace is not preserved in args\n+check_du () {\n+\targs=$*\n+\n+\ttest_expect_success \"generate expected size ($args)\" \"\n+\t\tdisk_usage_slow $args >expect\n+\t\"\n+\n+\ttest_expect_success \"rev-list --disk-usage without bitmaps ($args)\" \"\n+\t\tgit rev-list --disk-usage $args >actual &&\n+\t\ttest_cmp expect actual\n+\t\"\n+\n+\ttest_expect_success \"rev-list --disk-usage with bitmaps ($args)\" \"\n+\t\tgit rev-list --disk-usage --use-bitmap-index $args >actual &&\n+\t\ttest_cmp expect actual\n+\t\"\n+}\n+\n+check_du HEAD\n+check_du HEAD^..HEAD\n+\n+test_done\n-- \n2.30.0.758.g9692d13bf2\n"},{"id":"415462","messageId":"YBHtztEsqqvz5zUK@nand.local","threadId":"55062","inReplyTo":"YBHlSQ6cSJWHWeWo@coredump.intra.peff.net","subject":"Re: [PATCH 1/2] t: add --no-tag option to test_commit","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-27T22:48:46Z","receivedAt":"2021-01-27T22:57:30Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 27, 2021 at 05:12:25PM -0500, Jeff King wrote:\n> The alternative is that scripts have to call test_tick, git-add, and\n> git-commit themselves. Let's add a --no-tag option to give them the\n> one-liner convenience of using test_commit.\n\nThanks for finding a spot that does this and making it more readable\nwith the new --no-tag option.\n\nThis patch looks obviously correct. I'm sure that (as you note) there\nare more cleanups possible, but I'm happy to just grab an easy one and\nlet future refactorings clean up the remaining ones.\n\nThanks,\nTaylor\n"},{"id":"415463","messageId":"YBHv0ZHZD4VMHLYR@nand.local","threadId":"55062","inReplyTo":"YBHmY7vNxu2hqOa/@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-27T22:57:21Z","receivedAt":"2021-01-27T23:01:02Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 27, 2021 at 05:17:07PM -0500, Jeff King wrote:\n> It can sometimes be useful to see which refs are contributing to the\n> overall repository size (e.g., does some branch have a bunch of objects\n> not found elsewhere in history, which indicates that deleting it would\n> shrink the size of a clone).\n>\n> You can find that out by generating a list of objects, getting their\n> sizes from cat-file, and then summing them, like:\n>\n>     git rev-list --objects main..branch\n>     cut -d' ' -f1 |\n\nI suspect that this is from the original commit message that you wrote a\nhalf-decade ago. Not that it really means much, but you could shave one\nprocess off of this example by passing '--no-object-names' to 'git\nrev-list'.\n\nThe whole point is that we can avoid having to do this, so I don't think\nit really matters, anyway.\n\n> [...]\n> then we're faster to generate the list of objects, but we still spend a\n> lot of time piping and looking things up. But if we do both together:\n>\n>   [internal, bitmaps]\n>   $ time git rev-list --disk-usage --all --use-bitmap-index\n>   1455691059\n>   real\t0m0.235s\n>   user\t0m0.186s\n>   sys\t0m0.049s\n>\n> then we get the same answer much faster.\n\nVery nice.\n\n> This _could_ be made more flexible, but I didn't think it was worth the\n> complexity. Some obvious things one might want are:\n>\n>   - not counting up all reachable objects (i.e., requiring --objects for\n>     this output, and omitting it just counts up commits). This could be\n>     handled in the bitmap case with some extra code (OR-ing with the\n>     type bitmaps).\n>\n>     But after 5 years of this patch, I've never wanted that once. The\n>     disk usage of just some of the objects isn't really that useful (and\n>     of course you can still get it by piping to cat-file).\n\nYeah. I think it's trivial to support it, but I'm in favor of a simpler\ninterface.\n\nThat said, I worry about painting ourselves into a corner if the default\nimplies --objects. If we wanted to change that, I'm pretty sure you'd\nhave to write a rule that says \"imply objects, unless --tags, --blobs or\netc. are specified, and then only do that\".\n\nMaybe we'll never have to address that, but it's worth thinking about\nbefore committing to implying '--objects'.\n\n>   - an option to output the sizes of specific objects along with their\n>     oids. But if you want to get to this level of flexibility, I think\n>     you're better off just using cat-file (and if we are concerned about\n>     the pipe costs, we should teach rev-list to understand cat-file's\n>     custom formats).\n\nThis I agree with completely. Any caller who wants that level of\nflexibility shouldn't mind the piping.\n\nI have no comments on the patch itself, which looks fine to me (and I\nhave seen over and over again as it seems to regularly cause conflicts\nwhen merging new releases into GitHub's fork :-)).\n\nThanks,\nTaylor\n"},{"id":"415464","messageId":"CAPig+cQTV6ACiOj+GKoBwj15TZBr5craVPT6dYzzSDfrX9a3YA@mail.gmail.com","threadId":"55062","inReplyTo":"YBHmY7vNxu2hqOa/@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2021-01-27T23:07:57Z","receivedAt":"2021-01-27T23:14:31Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Jan 27, 2021 at 5:20 PM Jeff King <peff@peff.net> wrote:\n> This patch implements a --disk-usage option which produces the same\n> answer in a fraction of the time. Here are some timings using a clone of\n> torvalds/linux:\n>\n>   [rev-list piped to cat-file, no bitmaps]\n>   $ time git rev-list --objects --all |\n>     cut -d' ' -f1 |\n>     git cat-file --buffer --batch-check='%(objectsize:disk)' |\n>     perl -lne '$total += $_; END { print $total }'\n>   1455691059\n>   real  0m34.336s\n>   user  0m46.533s\n>   sys   0m2.953s\n\nThis example shows the computed size (1455691059)...\n\n> But the real win is with bitmaps. If we use them without the new option:\n>\n>   [rev-list piped to cat-file, bitmaps]\n>   $ time git rev-list --objects --all --use-bitmap-index |\n>     cut -d' ' -f1 |\n>     git cat-file --batch-check='%(objectsize:disk)' |\n>     perl -lne '$total += $_; END { print $total }'\n>   real  0m9.954s\n>   user  0m11.234s\n>   sys   0m8.522s\n\n...however, this example does not (but all the others do). Simple\ncopy/paste error?\n\nNot worth a re-roll, of course.\n"},{"id":"415465","messageId":"87mtwuvva8.fsf@kyleam.com","threadId":"55062","inReplyTo":"YBHmY7vNxu2hqOa/@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Kyle Meyer","fromEmail":"kyle@kyleam.com","sentAt":"2021-01-27T23:01:51Z","receivedAt":"2021-01-27T23:17:35Z","isPatch":true,"sender":{"key":"kyle@kyleam.com","avatar":"https://avatars.githubusercontent.com/u/1297788?v=4"},"body":"Jeff King writes:\n\n> diff --git a/Documentation/rev-list-options.txt b/Documentation/rev-list-options.txt\n> index 002379056a..1e5826f26d 100644\n> --- a/Documentation/rev-list-options.txt\n> +++ b/Documentation/rev-list-options.txt\n> @@ -222,6 +222,15 @@ ifdef::git-rev-list[]\n>  \ttest the exit status to see if a range of objects is fully\n>  \tconnected (or not).  It is faster than redirecting stdout\n>  \tto `/dev/null` as the output does not have to be formatted.\n> +\n> +--disk-usage::\n> +\tSuppress normal output; instead, print the sum of the bytes used\n> +\tfor on-disk storage by the selected objects. This is equivalent\n> +\tto piping the output of `rev-list --objects` into\n> +\t`git cat-file --batch-check='%(objectsize:disk)', except that it\n\n[ Just a drive-by typo comment from a reader not knowledgeable enough to\n  review the code change :) ]\n\nThe cat-file command is missing its closing quote.\n"},{"id":"415467","messageId":"YBHtS4bJQxHFU3DM@nand.local","threadId":"55062","inReplyTo":"YBHlGPBSJC++CnPy@coredump.intra.peff.net","subject":"Re: [PATCH 0/2] rev-list --disk-usage","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-27T22:46:35Z","receivedAt":"2021-01-27T23:26:06Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 27, 2021 at 05:11:36PM -0500, Jeff King wrote:\n> The first patch is just a test-script enhancement to let test_commit\n> avoid creating tags. During some recent refactoring, we actually broke\n> the --disk-usage feature but the test script didn't catch it because the\n> tags were being picked up by \"--all\". Since this is at least the third\n> time I've run into that in our test suite, I thought I'd make it a\n> little more convenient to avoid. :)\n\nI appreciate the non-incriminating \"we\", but the person who caused the\nregression was most certainly me ;-).\n\nThis happened while cherry-picking Junio's recent merge of\ntb/revindex-api, which obviously did not cause a merge conflict with\nthis new caller. The remaining details are boring, but they definitely\nweren't Peff's fault :-).\n\nThanks,\nTaylor\n"},{"id":"415468","messageId":"YBH4b04kZL8V6GFe@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YBHv0ZHZD4VMHLYR@nand.local","subject":"Re: [PATCH 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-27T23:34:07Z","receivedAt":"2021-01-27T23:36:36Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jan 27, 2021 at 05:57:21PM -0500, Taylor Blau wrote:\n\n> > You can find that out by generating a list of objects, getting their\n> > sizes from cat-file, and then summing them, like:\n> >\n> >     git rev-list --objects main..branch\n> >     cut -d' ' -f1 |\n> \n> I suspect that this is from the original commit message that you wrote a\n> half-decade ago. Not that it really means much, but you could shave one\n> process off of this example by passing '--no-object-names' to 'git\n> rev-list'.\n\nThat, plus my muscle memory to do the cut. We should probably model the\nbetter form here, and use it in the test, though (not worth a re-roll on\nits own, but it looks like there are a few other minor bits).\n\n> >   - not counting up all reachable objects (i.e., requiring --objects for\n> >     this output, and omitting it just counts up commits). This could be\n> >     handled in the bitmap case with some extra code (OR-ing with the\n> >     type bitmaps).\n> >\n> >     But after 5 years of this patch, I've never wanted that once. The\n> >     disk usage of just some of the objects isn't really that useful (and\n> >     of course you can still get it by piping to cat-file).\n> \n> Yeah. I think it's trivial to support it, but I'm in favor of a simpler\n> interface.\n> \n> That said, I worry about painting ourselves into a corner if the default\n> implies --objects. If we wanted to change that, I'm pretty sure you'd\n> have to write a rule that says \"imply objects, unless --tags, --blobs or\n> etc. are specified, and then only do that\".\n> \n> Maybe we'll never have to address that, but it's worth thinking about\n> before committing to implying '--objects'.\n\nYeah, the one thing that gives me pause is that it would be hard to undo\nlater. I didn't write the code to handle it in the bitmap case, but I\ndon't think it would be _too_ bad. It is slightly annoying for the\nall-objects case, because the existing code isn't set up well to iterate\neither a specific type, or all types.\n\n> I have no comments on the patch itself, which looks fine to me (and I\n> have seen over and over again as it seems to regularly cause conflicts\n> when merging new releases into GitHub's fork :-)).\n\nYou are exposing my ulterior motive. :)\n\n-Peff\n"},{"id":"415469","messageId":"YBH5DarZvg/phGjP@coredump.intra.peff.net","threadId":"55062","inReplyTo":"87mtwuvva8.fsf@kyleam.com","subject":"Re: [PATCH 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-27T23:36:45Z","receivedAt":"2021-01-27T23:38:29Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jan 27, 2021 at 06:01:51PM -0500, Kyle Meyer wrote:\n\n> > +--disk-usage::\n> > +\tSuppress normal output; instead, print the sum of the bytes used\n> > +\tfor on-disk storage by the selected objects. This is equivalent\n> > +\tto piping the output of `rev-list --objects` into\n> > +\t`git cat-file --batch-check='%(objectsize:disk)', except that it\n> \n> [ Just a drive-by typo comment from a reader not knowledgeable enough to\n>   review the code change :) ]\n> \n> The cat-file command is missing its closing quote.\n\nThanks for catching that. I should have looked at the output of\ndoc-diff, which does reveal it.\n\n-Peff\n"},{"id":"415470","messageId":"YBH5qFon3qP2vnuZ@coredump.intra.peff.net","threadId":"55062","inReplyTo":"CAPig+cQTV6ACiOj+GKoBwj15TZBr5craVPT6dYzzSDfrX9a3YA@mail.gmail.com","subject":"Re: [PATCH 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-01-27T23:39:20Z","receivedAt":"2021-01-27T23:42:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jan 27, 2021 at 06:07:57PM -0500, Eric Sunshine wrote:\n\n> This example shows the computed size (1455691059)...\n> [...]\n> ...however, this example does not (but all the others do). Simple\n> copy/paste error?\n\nYep, thanks for catching. (Of course I have since repacked my linux.git,\nso now it produces a different answer! It does match the current value\nof the other techniques, though).\n\n> Not worth a re-roll, of course.\n\nAgreed, but it looks like there are a few other minor bits, so I'll\ndefinitely fix it up at the same time. I'll give a little more time\nbefore re-rolling in case there are any other comments.\n\n-Peff\n"},{"id":"416490","messageId":"YCJpbPIlSpCAKSBF@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YBHlGPBSJC++CnPy@coredump.intra.peff.net","subject":"[PATCH v2] rev-list --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-09T10:52:28Z","receivedAt":"2021-02-09T10:55:42Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"Here's a re-roll of my series to add \"rev-list --disk-usage\", for\ncounting up object storage used for various slices of history.\n\nThis fixes the minor bits mentioned in review for v1, but the big change\nis that \"--disk-usage\" no longer implies \"--objects\". I think you\ngenerally would want to use it with that option, but it really seemed to\nviolate the principle of least surprise for the user.\n\nThat requires handling each object type independently, but the code for\nthat turned out to be not too bad (and is modeled after the similar\nlogic in traverse_bitmap_commit_list()). I was slightly concerned that\nit would slow things down to walk over the bitmap multiple times, but it\ndoesn't seem to make much of a difference in practice.\n\nThere's a range-diff below, but it's not really worth looking at. All of\nthe interesting parts were rewritten completely, so you're better off to\njust read patch 2 again (and patch 1 did not change at all).\n\n  [1/2]: t: add --no-tag option to test_commit\n  [2/2]: rev-list: add --disk-usage option for calculating disk usage\n\n Documentation/rev-list-options.txt |  9 ++++\n builtin/rev-list.c                 | 46 +++++++++++++++++\n pack-bitmap.c                      | 81 ++++++++++++++++++++++++++++++\n pack-bitmap.h                      |  2 +\n t/t4208-log-magic-pathspec.sh      |  9 +---\n t/t6114-rev-list-du.sh             | 51 +++++++++++++++++++\n t/test-lib-functions.sh            |  9 +++-\n 7 files changed, 199 insertions(+), 8 deletions(-)\n create mode 100755 t/t6114-rev-list-du.sh\n\n1:  20f8edeff1 = 1:  6365cd94bd t: add --no-tag option to test_commit\n2:  64e28cb6c9 ! 2:  8a93583dee rev-list: add --disk-usage option for calculating disk usage\n    @@ Commit message\n         You can find that out by generating a list of objects, getting their\n         sizes from cat-file, and then summing them, like:\n     \n    -        git rev-list --objects main..branch\n    -        cut -d' ' -f1 |\n    +        git rev-list --objects --no-object-names main..branch\n             git cat-file --batch-check='%(objectsize:disk)' |\n             perl -lne '$total += $_; END { print $total }'\n     \n    @@ Commit message\n         torvalds/linux:\n     \n           [rev-list piped to cat-file, no bitmaps]\n    -      $ time git rev-list --objects --all |\n    -        cut -d' ' -f1 |\n    +      $ time git rev-list --objects --no-object-names --all |\n             git cat-file --buffer --batch-check='%(objectsize:disk)' |\n             perl -lne '$total += $_; END { print $total }'\n    -      1455691059\n    -      real  0m34.336s\n    -      user  0m46.533s\n    -      sys   0m2.953s\n    +      1459938510\n    +      real  0m29.635s\n    +      user  0m38.003s\n    +      sys   0m1.093s\n     \n           [internal, no bitmaps]\n    -      $ time git rev-list --disk-usage --all\n    -      1455691059\n    -      real  0m32.662s\n    -      user  0m32.306s\n    -      sys   0m0.353s\n    +      $ time git rev-list --disk-usage --objects --all\n    +      1459938510\n    +      real  0m31.262s\n    +      user  0m30.885s\n    +      sys   0m0.376s\n     \n    -    The wall-clock times aren't that different because of parallelism, but\n    -    notice the CPU savings between the two. We saved 35% of the CPU just by\n    +    Even though the wall-clock time is slightly worse due to parallelism,\n    +    notice the CPU savings between the two. We saved 21% of the CPU just by\n         avoiding the pipes.\n     \n         But the real win is with bitmaps. If we use them without the new option:\n     \n           [rev-list piped to cat-file, bitmaps]\n    -      $ time git rev-list --objects --all --use-bitmap-index |\n    -        cut -d' ' -f1 |\n    +      $ time git rev-list --objects --no-object-names --all --use-bitmap-index |\n             git cat-file --batch-check='%(objectsize:disk)' |\n             perl -lne '$total += $_; END { print $total }'\n    -      real  0m9.954s\n    -      user  0m11.234s\n    -      sys   0m8.522s\n    +      1459938510\n    +      real  0m6.244s\n    +      user  0m8.452s\n    +      sys   0m0.311s\n     \n         then we're faster to generate the list of objects, but we still spend a\n         lot of time piping and looking things up. But if we do both together:\n     \n           [internal, bitmaps]\n    -      $ time git rev-list --disk-usage --all --use-bitmap-index\n    -      1455691059\n    -      real  0m0.235s\n    -      user  0m0.186s\n    +      $ time git rev-list --disk-usage --objects --all --use-bitmap-index\n    +      1459938510\n    +      real  0m0.219s\n    +      user  0m0.169s\n           sys   0m0.049s\n     \n         then we get the same answer much faster.\n    @@ Commit message\n         of course. But we're actually checking reachability here, so we're still\n         fast when we ask for more interesting things:\n     \n    -      $ time git rev-list --disk-usage --all --use-bitmap-index v5.0..v5.10\n    +      $ time git rev-list --disk-usage --use-bitmap-index v5.0..v5.10\n           374798628\n           real  0m0.429s\n           user  0m0.356s\n    @@ Documentation/rev-list-options.txt: ifdef::git-rev-list[]\n     +\n     +--disk-usage::\n     +\tSuppress normal output; instead, print the sum of the bytes used\n    -+\tfor on-disk storage by the selected objects. This is equivalent\n    -+\tto piping the output of `rev-list --objects` into\n    -+\t`git cat-file --batch-check='%(objectsize:disk)', except that it\n    -+\truns much faster (especially with `--use-bitmap-index`). See the\n    -+\t`CAVEATS` section in linkgit:git-cat-file[1] for the limitations\n    -+\tof what \"on-disk storage\" means.\n    ++\tfor on-disk storage by the selected commits or objects. This is\n    ++\tequivalent to piping the output into `git cat-file\n    ++\t--batch-check='%(objectsize:disk)'`, except that it runs much\n    ++\tfaster (especially with `--use-bitmap-index`). See the `CAVEATS`\n    ++\tsection in linkgit:git-cat-file[1] for the limitations of what\n    ++\t\"on-disk storage\" means.\n      endif::git-rev-list[]\n      \n      --cherry-mark::\n    @@ builtin/rev-list.c: static int try_bitmap_traversal(struct rev_info *revs,\n     +\t\treturn -1;\n     +\n     +\tprintf(\"%\"PRIuMAX\"\\n\",\n    -+\t       (uintmax_t)get_disk_usage_from_bitmap(bitmap_git));\n    ++\t       (uintmax_t)get_disk_usage_from_bitmap(bitmap_git, revs));\n     +\treturn 0;\n     +}\n     +\n    @@ builtin/rev-list.c: int cmd_rev_list(int argc, const char **argv, const char *pr\n      \n     +\t\tif (!strcmp(arg, \"--disk-usage\")) {\n     +\t\t\tshow_disk_usage = 1;\n    -+\t\t\trevs.tag_objects = 1;\n    -+\t\t\trevs.tree_objects = 1;\n    -+\t\t\trevs.blob_objects = 1;\n     +\t\t\tinfo.flags |= REV_LIST_QUIET;\n     +\t\t\tcontinue;\n     +\t\t}\n    @@ pack-bitmap.c: int bitmap_has_oid_in_uninteresting(struct bitmap_index *bitmap_g\n      \t\tbitmap_walk_contains(bitmap_git, bitmap_git->haves, oid);\n      }\n     +\n    -+off_t get_disk_usage_from_bitmap(struct bitmap_index *bitmap_git)\n    ++static off_t get_disk_usage_for_type(struct bitmap_index *bitmap_git,\n    ++\t\t\t\t     enum object_type object_type)\n     +{\n     +\tstruct bitmap *result = bitmap_git->result;\n     +\tstruct packed_git *pack = bitmap_git->pack;\n    -+\tstruct eindex *eindex = &bitmap_git->ext_index;\n    -+\tstruct object_info oi = OBJECT_INFO_INIT;\n    -+\toff_t object_size;\n     +\toff_t total = 0;\n    ++\tstruct ewah_iterator it;\n    ++\teword_t filter;\n     +\tsize_t i;\n     +\n    -+\toi.disk_sizep = &object_size;\n    -+\n    -+\tfor (i = 0; i < result->word_alloc; i++) {\n    -+\t\teword_t word = result->words[i];\n    ++\tinit_type_iterator(&it, bitmap_git, object_type);\n    ++\tfor (i = 0; i < result->word_alloc &&\n    ++\t\t\tewah_iterator_next(&filter, &it); i++) {\n    ++\t\teword_t word = result->words[i] & filter;\n     +\t\tsize_t base = (i * BITS_IN_EWORD);\n     +\t\tunsigned offset;\n     +\n    ++\t\tif (!word)\n    ++\t\t\tcontinue;\n    ++\n     +\t\tfor (offset = 0; offset < BITS_IN_EWORD; offset++) {\n     +\t\t\tsize_t pos;\n     +\n    @@ pack-bitmap.c: int bitmap_has_oid_in_uninteresting(struct bitmap_index *bitmap_g\n     +\n     +\t\t\toffset += ewah_bit_ctz64(word >> offset);\n     +\t\t\tpos = base + offset;\n    -+\n    -+\t\t\t/*\n    -+\t\t\t * If it's in the pack, we can use the fast path\n    -+\t\t\t * and just check the revindex. Otherwise, we\n    -+\t\t\t * fall back to looking it up.\n    -+\t\t\t */\n    -+\t\t\tif (pos < pack->num_objects) {\n    -+\t\t\t\tobject_size =\n    -+\t\t\t\t\tpack_pos_to_offset(pack, pos + 1) -\n    -+\t\t\t\t\tpack_pos_to_offset(pack, pos);\n    -+\t\t\t} else {\n    -+\t\t\t\tstruct object *obj;\n    -+\t\t\t\tobj = eindex->objects[pos - pack->num_objects];\n    -+\t\t\t\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n    -+\t\t\t\t\tdie(_(\"unable to get disk usage of %s\"),\n    -+\t\t\t\t\t      oid_to_hex(&obj->oid));\n    -+\t\t\t}\n    -+\n    -+\t\t\ttotal += object_size;\n    ++\t\t\ttotal += pack_pos_to_offset(pack, pos + 1) -\n    ++\t\t\t\t pack_pos_to_offset(pack, pos);\n     +\t\t}\n     +\t}\n     +\n     +\treturn total;\n    ++}\n    ++\n    ++static off_t get_disk_usage_for_extended(struct bitmap_index *bitmap_git)\n    ++{\n    ++\tstruct bitmap *result = bitmap_git->result;\n    ++\tstruct packed_git *pack = bitmap_git->pack;\n    ++\tstruct eindex *eindex = &bitmap_git->ext_index;\n    ++\toff_t total = 0;\n    ++\tstruct object_info oi = OBJECT_INFO_INIT;\n    ++\toff_t object_size;\n    ++\tsize_t i;\n    ++\n    ++\toi.disk_sizep = &object_size;\n    ++\n    ++\tfor (i = 0; i < eindex->count; i++) {\n    ++\t\tstruct object *obj = eindex->objects[i];\n    ++\n    ++\t\tif (!bitmap_get(result, pack->num_objects + i))\n    ++\t\t\tcontinue;\n    ++\n    ++\t\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n    ++\t\t\tdie(_(\"unable to get disk usage of %s\"),\n    ++\t\t\t    oid_to_hex(&obj->oid));\n    ++\n    ++\t\ttotal += object_size;\n    ++\t}\n    ++\treturn total;\n    ++}\n    ++\n    ++off_t get_disk_usage_from_bitmap(struct bitmap_index *bitmap_git,\n    ++\t\t\t\t struct rev_info *revs)\n    ++{\n    ++\toff_t total = 0;\n    ++\n    ++\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_COMMIT);\n    ++\tif (revs->tree_objects)\n    ++\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_TREE);\n    ++\tif (revs->blob_objects)\n    ++\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_BLOB);\n    ++\tif (revs->tag_objects)\n    ++\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_TAG);\n    ++\n    ++\ttotal += get_disk_usage_for_extended(bitmap_git);\n    ++\n    ++\treturn total;\n     +}\n     \n      ## pack-bitmap.h ##\n     @@ pack-bitmap.h: int bitmap_walk_contains(struct bitmap_index *,\n       */\n      int bitmap_has_oid_in_uninteresting(struct bitmap_index *, const struct object_id *oid);\n      \n    -+off_t get_disk_usage_from_bitmap(struct bitmap_index *);\n    ++off_t get_disk_usage_from_bitmap(struct bitmap_index *, struct rev_info *);\n     +\n      void bitmap_writer_show_progress(int show);\n      void bitmap_writer_set_checksum(unsigned char *sha1);\n    @@ t/t6114-rev-list-du.sh (new)\n     +# packing, zlib, etc. We'll assume that the regular rev-list and cat-file\n     +# machinery works and compare the --disk-usage output to that.\n     +disk_usage_slow () {\n    -+\tgit rev-list --objects \"$@\" |\n    -+\tcut -d' ' -f1 |\n    ++\tgit rev-list --no-object-names \"$@\" |\n     +\tgit cat-file --batch-check=\"%(objectsize:disk)\" |\n     +\tperl -lne '$total += $_; END { print $total}'\n     +}\n    @@ t/t6114-rev-list-du.sh (new)\n     +}\n     +\n     +check_du HEAD\n    -+check_du HEAD^..HEAD\n    ++check_du --objects HEAD\n    ++check_du --objects HEAD^..HEAD\n     +\n     +test_done\n"},{"id":"416491","messageId":"YCJpfYJqevvqBj1D@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YCJpbPIlSpCAKSBF@coredump.intra.peff.net","subject":"[PATCH v2 1/2] t: add --no-tag option to test_commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-09T10:52:45Z","receivedAt":"2021-02-09T10:56:15Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"One of the conveniences that test_commit offers is making a tag for each\ncommit. This makes it easy to refer to the commits in subsequent\ncommands. But it can also be a pain if you care about reachability,\nbecause those tags keep the commits reachable even if they are rewound\nfrom the branch they're made on.\n\nThe alternative is that scripts have to call test_tick, git-add, and\ngit-commit themselves. Let's add a --no-tag option to give them the\none-liner convenience of using test_commit.\n\nThis is in preparation for the next patch, which will add some more\ncalls. But I cleaned up an existing site to show off the feature. There\nare probably more cleanups possible.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n t/t4208-log-magic-pathspec.sh | 9 ++-------\n t/test-lib-functions.sh       | 9 ++++++++-\n 2 files changed, 10 insertions(+), 8 deletions(-)\n\ndiff --git a/t/t4208-log-magic-pathspec.sh b/t/t4208-log-magic-pathspec.sh\nindex 5e10136e9a..7f0c1dcc0f 100755\n--- a/t/t4208-log-magic-pathspec.sh\n+++ b/t/t4208-log-magic-pathspec.sh\n@@ -31,13 +31,8 @@ test_expect_success '\"git log :/a -- \" should not be ambiguous' '\n test_expect_success '\"git log :/detached -- \" should find a commit only in HEAD' '\n \ttest_when_finished \"git checkout main\" &&\n \tgit checkout --detach &&\n-\t# Must manually call `test_tick` instead of using `test_commit`,\n-\t# because the latter additionally creates a tag, which would make\n-\t# the commit reachable not only via HEAD.\n-\ttest_tick &&\n-\tgit commit --allow-empty -m detached &&\n-\ttest_tick &&\n-\tgit commit --allow-empty -m something-else &&\n+\ttest_commit --no-tag detached &&\n+\ttest_commit --no-tag something-else &&\n \tgit log :/detached --\n '\n \ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 6bca002316..1587241ba0 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -202,6 +202,7 @@ test_commit () {\n \tauthor= &&\n \tsignoff= &&\n \tindir= &&\n+\tno_tag= &&\n \twhile test $# != 0\n \tdo\n \t\tcase \"$1\" in\n@@ -222,6 +223,9 @@ test_commit () {\n \t\t\tindir=\"$2\"\n \t\t\tshift\n \t\t\t;;\n+\t\t--no-tag)\n+\t\t\tno_tag=yes\n+\t\t\t;;\n \t\t*)\n \t\t\tbreak\n \t\t\t;;\n@@ -244,7 +248,10 @@ test_commit () {\n \tgit ${indir:+ -C \"$indir\"} commit \\\n \t    ${author:+ --author \"$author\"} \\\n \t    $signoff -m \"$1\" &&\n-\tgit ${indir:+ -C \"$indir\"} tag \"${4:-$1}\"\n+\tif test -z \"$no_tag\"\n+\tthen\n+\t\tgit ${indir:+ -C \"$indir\"} tag \"${4:-$1}\"\n+\tfi\n }\n \n # Call test_merge with the arguments \"<message> <commit>\", where <commit>\n-- \n2.30.1.887.ge7d57fcab0\n\n"},{"id":"416492","messageId":"YCJpvi0V045PkpJ2@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YCJpbPIlSpCAKSBF@coredump.intra.peff.net","subject":"[PATCH v2 2/2] rev-list: add --disk-usage option for calculating disk usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-09T10:53:50Z","receivedAt":"2021-02-09T10:57:17Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"It can sometimes be useful to see which refs are contributing to the\noverall repository size (e.g., does some branch have a bunch of objects\nnot found elsewhere in history, which indicates that deleting it would\nshrink the size of a clone).\n\nYou can find that out by generating a list of objects, getting their\nsizes from cat-file, and then summing them, like:\n\n    git rev-list --objects --no-object-names main..branch\n    git cat-file --batch-check='%(objectsize:disk)' |\n    perl -lne '$total += $_; END { print $total }'\n\nThough note that the caveats from git-cat-file(1) apply here. We \"blame\"\nbase objects more than their deltas, even though the relationship could\neasily be flipped. Still, it can be a useful rough measure.\n\nBut one problem is that it's slow to run. Teaching rev-list to sum up\nthe sizes can be much faster for two reasons:\n\n  1. It skips all of the piping of object names and sizes.\n\n  2. If bitmaps are in use, for objects that are in the\n     bitmapped packfile we can skip the oid_object_info()\n     lookup entirely, and just ask the revindex for the\n     on-disk size.\n\nThis patch implements a --disk-usage option which produces the same\nanswer in a fraction of the time. Here are some timings using a clone of\ntorvalds/linux:\n\n  [rev-list piped to cat-file, no bitmaps]\n  $ time git rev-list --objects --no-object-names --all |\n    git cat-file --buffer --batch-check='%(objectsize:disk)' |\n    perl -lne '$total += $_; END { print $total }'\n  1459938510\n  real\t0m29.635s\n  user\t0m38.003s\n  sys\t0m1.093s\n\n  [internal, no bitmaps]\n  $ time git rev-list --disk-usage --objects --all\n  1459938510\n  real\t0m31.262s\n  user\t0m30.885s\n  sys\t0m0.376s\n\nEven though the wall-clock time is slightly worse due to parallelism,\nnotice the CPU savings between the two. We saved 21% of the CPU just by\navoiding the pipes.\n\nBut the real win is with bitmaps. If we use them without the new option:\n\n  [rev-list piped to cat-file, bitmaps]\n  $ time git rev-list --objects --no-object-names --all --use-bitmap-index |\n    git cat-file --batch-check='%(objectsize:disk)' |\n    perl -lne '$total += $_; END { print $total }'\n  1459938510\n  real\t0m6.244s\n  user\t0m8.452s\n  sys\t0m0.311s\n\nthen we're faster to generate the list of objects, but we still spend a\nlot of time piping and looking things up. But if we do both together:\n\n  [internal, bitmaps]\n  $ time git rev-list --disk-usage --objects --all --use-bitmap-index\n  1459938510\n  real\t0m0.219s\n  user\t0m0.169s\n  sys\t0m0.049s\n\nthen we get the same answer much faster.\n\nFor \"--all\", that answer will correspond closely to \"du objects/pack\",\nof course. But we're actually checking reachability here, so we're still\nfast when we ask for more interesting things:\n\n  $ time git rev-list --disk-usage --use-bitmap-index v5.0..v5.10\n  374798628\n  real\t0m0.429s\n  user\t0m0.356s\n  sys\t0m0.072s\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/rev-list-options.txt |  9 ++++\n builtin/rev-list.c                 | 46 +++++++++++++++++\n pack-bitmap.c                      | 81 ++++++++++++++++++++++++++++++\n pack-bitmap.h                      |  2 +\n t/t6114-rev-list-du.sh             | 51 +++++++++++++++++++\n 5 files changed, 189 insertions(+)\n create mode 100755 t/t6114-rev-list-du.sh\n\ndiff --git a/Documentation/rev-list-options.txt b/Documentation/rev-list-options.txt\nindex 96cc89d157..1238bfd915 100644\n--- a/Documentation/rev-list-options.txt\n+++ b/Documentation/rev-list-options.txt\n@@ -227,6 +227,15 @@ ifdef::git-rev-list[]\n \ttest the exit status to see if a range of objects is fully\n \tconnected (or not).  It is faster than redirecting stdout\n \tto `/dev/null` as the output does not have to be formatted.\n+\n+--disk-usage::\n+\tSuppress normal output; instead, print the sum of the bytes used\n+\tfor on-disk storage by the selected commits or objects. This is\n+\tequivalent to piping the output into `git cat-file\n+\t--batch-check='%(objectsize:disk)'`, except that it runs much\n+\tfaster (especially with `--use-bitmap-index`). See the `CAVEATS`\n+\tsection in linkgit:git-cat-file[1] for the limitations of what\n+\t\"on-disk storage\" means.\n endif::git-rev-list[]\n \n --cherry-mark::\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex 25c6c3b38d..b4d8ea0a35 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -80,6 +80,19 @@ static int arg_show_object_names = 1;\n \n #define DEFAULT_OIDSET_SIZE     (16*1024)\n \n+static int show_disk_usage;\n+static off_t total_disk_usage;\n+\n+static off_t get_object_disk_usage(struct object *obj)\n+{\n+\toff_t size;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\toi.disk_sizep = &size;\n+\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n+\t\tdie(_(\"unable to get disk usage of %s\"), oid_to_hex(&obj->oid));\n+\treturn size;\n+}\n+\n static void finish_commit(struct commit *commit);\n static void show_commit(struct commit *commit, void *data)\n {\n@@ -88,6 +101,9 @@ static void show_commit(struct commit *commit, void *data)\n \n \tdisplay_progress(progress, ++progress_counter);\n \n+\tif (show_disk_usage)\n+\t\ttotal_disk_usage += get_object_disk_usage(&commit->object);\n+\n \tif (info->flags & REV_LIST_QUIET) {\n \t\tfinish_commit(commit);\n \t\treturn;\n@@ -258,6 +274,8 @@ static void show_object(struct object *obj, const char *name, void *cb_data)\n \tif (finish_object(obj, name, cb_data))\n \t\treturn;\n \tdisplay_progress(progress, ++progress_counter);\n+\tif (show_disk_usage)\n+\t\ttotal_disk_usage += get_object_disk_usage(obj);\n \tif (info->flags & REV_LIST_QUIET)\n \t\treturn;\n \n@@ -452,6 +470,23 @@ static int try_bitmap_traversal(struct rev_info *revs,\n \treturn 0;\n }\n \n+static int try_bitmap_disk_usage(struct rev_info *revs,\n+\t\t\t\t struct list_objects_filter_options *filter)\n+{\n+\tstruct bitmap_index *bitmap_git;\n+\n+\tif (!show_disk_usage)\n+\t\treturn -1;\n+\n+\tbitmap_git = prepare_bitmap_walk(revs, filter);\n+\tif (!bitmap_git)\n+\t\treturn -1;\n+\n+\tprintf(\"%\"PRIuMAX\"\\n\",\n+\t       (uintmax_t)get_disk_usage_from_bitmap(bitmap_git, revs));\n+\treturn 0;\n+}\n+\n int cmd_rev_list(int argc, const char **argv, const char *prefix)\n {\n \tstruct rev_info revs;\n@@ -584,6 +619,12 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\tcontinue;\n \t\t}\n \n+\t\tif (!strcmp(arg, \"--disk-usage\")) {\n+\t\t\tshow_disk_usage = 1;\n+\t\t\tinfo.flags |= REV_LIST_QUIET;\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tusage(rev_list_usage);\n \n \t}\n@@ -626,6 +667,8 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \tif (use_bitmap_index) {\n \t\tif (!try_bitmap_count(&revs, &filter_options))\n \t\t\treturn 0;\n+\t\tif (!try_bitmap_disk_usage(&revs, &filter_options))\n+\t\t\treturn 0;\n \t\tif (!try_bitmap_traversal(&revs, &filter_options))\n \t\t\treturn 0;\n \t}\n@@ -690,5 +733,8 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\tprintf(\"%d\\n\", revs.count_left + revs.count_right);\n \t}\n \n+\tif (show_disk_usage)\n+\t\tprintf(\"%\"PRIuMAX\"\\n\", (uintmax_t)total_disk_usage);\n+\n \treturn 0;\n }\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 60fe20fb87..1f69b5fa85 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -1430,3 +1430,84 @@ int bitmap_has_oid_in_uninteresting(struct bitmap_index *bitmap_git,\n \treturn bitmap_git &&\n \t\tbitmap_walk_contains(bitmap_git, bitmap_git->haves, oid);\n }\n+\n+static off_t get_disk_usage_for_type(struct bitmap_index *bitmap_git,\n+\t\t\t\t     enum object_type object_type)\n+{\n+\tstruct bitmap *result = bitmap_git->result;\n+\tstruct packed_git *pack = bitmap_git->pack;\n+\toff_t total = 0;\n+\tstruct ewah_iterator it;\n+\teword_t filter;\n+\tsize_t i;\n+\n+\tinit_type_iterator(&it, bitmap_git, object_type);\n+\tfor (i = 0; i < result->word_alloc &&\n+\t\t\tewah_iterator_next(&filter, &it); i++) {\n+\t\teword_t word = result->words[i] & filter;\n+\t\tsize_t base = (i * BITS_IN_EWORD);\n+\t\tunsigned offset;\n+\n+\t\tif (!word)\n+\t\t\tcontinue;\n+\n+\t\tfor (offset = 0; offset < BITS_IN_EWORD; offset++) {\n+\t\t\tsize_t pos;\n+\n+\t\t\tif ((word >> offset) == 0)\n+\t\t\t\tbreak;\n+\n+\t\t\toffset += ewah_bit_ctz64(word >> offset);\n+\t\t\tpos = base + offset;\n+\t\t\ttotal += pack_pos_to_offset(pack, pos + 1) -\n+\t\t\t\t pack_pos_to_offset(pack, pos);\n+\t\t}\n+\t}\n+\n+\treturn total;\n+}\n+\n+static off_t get_disk_usage_for_extended(struct bitmap_index *bitmap_git)\n+{\n+\tstruct bitmap *result = bitmap_git->result;\n+\tstruct packed_git *pack = bitmap_git->pack;\n+\tstruct eindex *eindex = &bitmap_git->ext_index;\n+\toff_t total = 0;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\toff_t object_size;\n+\tsize_t i;\n+\n+\toi.disk_sizep = &object_size;\n+\n+\tfor (i = 0; i < eindex->count; i++) {\n+\t\tstruct object *obj = eindex->objects[i];\n+\n+\t\tif (!bitmap_get(result, pack->num_objects + i))\n+\t\t\tcontinue;\n+\n+\t\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n+\t\t\tdie(_(\"unable to get disk usage of %s\"),\n+\t\t\t    oid_to_hex(&obj->oid));\n+\n+\t\ttotal += object_size;\n+\t}\n+\treturn total;\n+}\n+\n+off_t get_disk_usage_from_bitmap(struct bitmap_index *bitmap_git,\n+\t\t\t\t struct rev_info *revs)\n+{\n+\toff_t total = 0;\n+\n+\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_COMMIT);\n+\tif (revs->tree_objects)\n+\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_TREE);\n+\tif (revs->blob_objects)\n+\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_BLOB);\n+\tif (revs->tag_objects)\n+\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_TAG);\n+\n+\ttotal += get_disk_usage_for_extended(bitmap_git);\n+\n+\treturn total;\n+}\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nindex 25dfcf5615..36d99930d8 100644\n--- a/pack-bitmap.h\n+++ b/pack-bitmap.h\n@@ -68,6 +68,8 @@ int bitmap_walk_contains(struct bitmap_index *,\n  */\n int bitmap_has_oid_in_uninteresting(struct bitmap_index *, const struct object_id *oid);\n \n+off_t get_disk_usage_from_bitmap(struct bitmap_index *, struct rev_info *);\n+\n void bitmap_writer_show_progress(int show);\n void bitmap_writer_set_checksum(unsigned char *sha1);\n void bitmap_writer_build_type_index(struct packing_data *to_pack,\ndiff --git a/t/t6114-rev-list-du.sh b/t/t6114-rev-list-du.sh\nnew file mode 100755\nindex 0000000000..b4aef32b71\n--- /dev/null\n+++ b/t/t6114-rev-list-du.sh\n@@ -0,0 +1,51 @@\n+#!/bin/sh\n+\n+test_description='basic tests of rev-list --disk-usage'\n+. ./test-lib.sh\n+\n+# we want a mix of reachable and unreachable, as well as\n+# objects in the bitmapped pack and some outside of it\n+test_expect_success 'set up repository' '\n+\ttest_commit --no-tag one &&\n+\ttest_commit --no-tag two &&\n+\tgit repack -adb &&\n+\tgit reset --hard HEAD^ &&\n+\ttest_commit --no-tag three &&\n+\ttest_commit --no-tag four &&\n+\tgit reset --hard HEAD^\n+'\n+\n+# We don't want to hardcode sizes, because they depend on the exact details of\n+# packing, zlib, etc. We'll assume that the regular rev-list and cat-file\n+# machinery works and compare the --disk-usage output to that.\n+disk_usage_slow () {\n+\tgit rev-list --no-object-names \"$@\" |\n+\tgit cat-file --batch-check=\"%(objectsize:disk)\" |\n+\tperl -lne '$total += $_; END { print $total}'\n+}\n+\n+# check behavior with given rev-list options; note that\n+# whitespace is not preserved in args\n+check_du () {\n+\targs=$*\n+\n+\ttest_expect_success \"generate expected size ($args)\" \"\n+\t\tdisk_usage_slow $args >expect\n+\t\"\n+\n+\ttest_expect_success \"rev-list --disk-usage without bitmaps ($args)\" \"\n+\t\tgit rev-list --disk-usage $args >actual &&\n+\t\ttest_cmp expect actual\n+\t\"\n+\n+\ttest_expect_success \"rev-list --disk-usage with bitmaps ($args)\" \"\n+\t\tgit rev-list --disk-usage --use-bitmap-index $args >actual &&\n+\t\ttest_cmp expect actual\n+\t\"\n+}\n+\n+check_du HEAD\n+check_du --objects HEAD\n+check_du --objects HEAD^..HEAD\n+\n+test_done\n-- \n2.30.1.887.ge7d57fcab0\n"},{"id":"416493","messageId":"YCJtbmaguIW+YeAs@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YCJpbPIlSpCAKSBF@coredump.intra.peff.net","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-09T11:09:34Z","receivedAt":"2021-02-09T11:11:24Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 09, 2021 at 05:52:28AM -0500, Jeff King wrote:\n\n> This fixes the minor bits mentioned in review for v1, but the big change\n> is that \"--disk-usage\" no longer implies \"--objects\". I think you\n> generally would want to use it with that option, but it really seemed to\n> violate the principle of least surprise for the user.\n> \n> That requires handling each object type independently, but the code for\n> that turned out to be not too bad (and is modeled after the similar\n> logic in traverse_bitmap_commit_list()). I was slightly concerned that\n> it would slow things down to walk over the bitmap multiple times, but it\n> doesn't seem to make much of a difference in practice.\n\nYou might reasonably ask whether we could just directly use\ntraverse_bitmap_commit_list(), since after all it takes a callback. And\nindeed, doing so reduces the size of the code (see the patch below).\n\nBut it's shockingly slower! It takes consistently 2-3x longer to produce\nthe same answer on linux.git with bitmaps. The problem is that we give\nmore information to the callback than the disk-usage computation needs.\n\nIn particular, finding nth_packed_object_id() is a big killer. Which\nkind of makes sense. We memcpy() the oids out of the .idx file into a\n\"struct object_id\" on the stack. And linux.git has ~200MB of oids to\ncopy (and I'm sure doing it 20 bytes at a time isn't quite optimal).\nThat adds several hundred milliseconds. Not a lot in absolute terms, but\nwe're able to do the whole computation in ~200ms to start with, so it's\nrelatively a big change.\n\nThis could be solved by having a more \"bare\" callback that just passes\nthe pack position, and not the oid (and then the callback is responsible\nfor looking it up if they care). But it gets pretty awkward when we have\nto complete the bitmap traversal with non-bitmap objects (for those we\n_do_ have an oid to pass, but no pack position). I think the\nimplementation in my 2/2 isn't so bad in comparison (and we can always\nswap it out later; these are all just implementation details).\n\nI did find it a bit interesting, though. When we moved to \"struct\nobject_id\" and started copying bits out with nth_packed_object_id(),\nrather than just pointing to the mmap'd .idx bytes, we wondered whether\nthere would be any measurable difference. Likewise when we extended it\nto handle the oid size changing at runtime. At the time, I wasn't able\nto measure any impact for real operations, but I guess we just needed a\ncase that highlighted it more.\n\nI don't know that it's really worth digging into that much, though it's\nquite possible there may be some easy wins by optimizing those memcpy\ncalls. E.g., I'm not sure if the compiler ends up inlining them or not.\nIf it doesn't realize that the_hash_algo->rawsz is only ever \"20\" or\n\"32\", we could perhaps help it along with specialized versions of\nhashcpy(). If somebody does want to play with it, this patch may make a\ngood testbed. :)\n\n-- >8 --\n builtin/pack-objects.c |  3 +-\n builtin/rev-list.c     | 40 +++++++++++------------\n pack-bitmap.c          | 86 ++------------------------------------------------\n pack-bitmap.h          |  1 +\n reachable.c            |  1 +\n 5 files changed, 25 insertions(+), 106 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 13cde5896a..33f7d19eb3 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1388,7 +1388,8 @@ static int add_object_entry(const struct object_id *oid, enum object_type type,\n static int add_object_entry_from_bitmap(const struct object_id *oid,\n \t\t\t\t\tenum object_type type,\n \t\t\t\t\tint flags, uint32_t name_hash,\n-\t\t\t\t\tstruct packed_git *pack, off_t offset)\n+\t\t\t\t\tstruct packed_git *pack,\n+\t\t\t\t\tuint32_t pack_pos, off_t offset)\n {\n \tdisplay_progress(progress_state, ++nr_seen);\n \ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex b4d8ea0a35..cc96b4c854 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -83,13 +83,13 @@ static int arg_show_object_names = 1;\n static int show_disk_usage;\n static off_t total_disk_usage;\n \n-static off_t get_object_disk_usage(struct object *obj)\n+static off_t get_object_disk_usage(const struct object_id *oid)\n {\n \toff_t size;\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \toi.disk_sizep = &size;\n-\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n-\t\tdie(_(\"unable to get disk usage of %s\"), oid_to_hex(&obj->oid));\n+\tif (oid_object_info_extended(the_repository, oid, &oi, 0) < 0)\n+\t\tdie(_(\"unable to get disk usage of %s\"), oid_to_hex(oid));\n \treturn size;\n }\n \n@@ -102,7 +102,7 @@ static void show_commit(struct commit *commit, void *data)\n \tdisplay_progress(progress, ++progress_counter);\n \n \tif (show_disk_usage)\n-\t\ttotal_disk_usage += get_object_disk_usage(&commit->object);\n+\t\ttotal_disk_usage += get_object_disk_usage(&commit->object.oid);\n \n \tif (info->flags & REV_LIST_QUIET) {\n \t\tfinish_commit(commit);\n@@ -275,7 +275,7 @@ static void show_object(struct object *obj, const char *name, void *cb_data)\n \t\treturn;\n \tdisplay_progress(progress, ++progress_counter);\n \tif (show_disk_usage)\n-\t\ttotal_disk_usage += get_object_disk_usage(obj);\n+\t\ttotal_disk_usage += get_object_disk_usage(&obj->oid);\n \tif (info->flags & REV_LIST_QUIET)\n \t\treturn;\n \n@@ -363,8 +363,19 @@ static int show_object_fast(\n \tint exclude,\n \tuint32_t name_hash,\n \tstruct packed_git *found_pack,\n+\tuint32_t pack_pos,\n \toff_t found_offset)\n {\n+\tif (show_disk_usage) {\n+\t\tif (found_pack) {\n+\t\t\ttotal_disk_usage +=\n+\t\t\t\tpack_pos_to_offset(found_pack, pack_pos + 1) -\n+\t\t\t\tfound_offset;\n+\t\t} else {\n+\t\t\ttotal_disk_usage += get_object_disk_usage(oid);\n+\t\t}\n+\t\treturn 1;\n+\t}\n \tfprintf(stdout, \"%s\\n\", oid_to_hex(oid));\n \treturn 1;\n }\n@@ -467,23 +478,10 @@ static int try_bitmap_traversal(struct rev_info *revs,\n \n \ttraverse_bitmap_commit_list(bitmap_git, revs, &show_object_fast);\n \tfree_bitmap_index(bitmap_git);\n-\treturn 0;\n-}\n-\n-static int try_bitmap_disk_usage(struct rev_info *revs,\n-\t\t\t\t struct list_objects_filter_options *filter)\n-{\n-\tstruct bitmap_index *bitmap_git;\n \n-\tif (!show_disk_usage)\n-\t\treturn -1;\n-\n-\tbitmap_git = prepare_bitmap_walk(revs, filter);\n-\tif (!bitmap_git)\n-\t\treturn -1;\n+\tif (show_disk_usage)\n+\t\tprintf(\"%\"PRIuMAX\"\\n\", (uintmax_t)total_disk_usage);\n \n-\tprintf(\"%\"PRIuMAX\"\\n\",\n-\t       (uintmax_t)get_disk_usage_from_bitmap(bitmap_git, revs));\n \treturn 0;\n }\n \n@@ -667,8 +665,6 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \tif (use_bitmap_index) {\n \t\tif (!try_bitmap_count(&revs, &filter_options))\n \t\t\treturn 0;\n-\t\tif (!try_bitmap_disk_usage(&revs, &filter_options))\n-\t\t\treturn 0;\n \t\tif (!try_bitmap_traversal(&revs, &filter_options))\n \t\t\treturn 0;\n \t}\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 1f69b5fa85..f118a365e1 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -655,7 +655,7 @@ static void show_extended_objects(struct bitmap_index *bitmap_git,\n \t\t    (obj->type == OBJ_TAG && !revs->tag_objects))\n \t\t\tcontinue;\n \n-\t\tshow_reach(&obj->oid, obj->type, 0, eindex->hashes[i], NULL, 0);\n+\t\tshow_reach(&obj->oid, obj->type, 0, eindex->hashes[i], NULL, 0, 0);\n \t}\n }\n \n@@ -726,7 +726,8 @@ static void show_objects_for_type(\n \t\t\tif (bitmap_git->hashes)\n \t\t\t\thash = get_be32(bitmap_git->hashes + index_pos);\n \n-\t\t\tshow_reach(&oid, object_type, 0, hash, bitmap_git->pack, ofs);\n+\t\t\tshow_reach(&oid, object_type, 0, hash,\n+\t\t\t\t   bitmap_git->pack, pos + offset, ofs);\n \t\t}\n \t}\n }\n@@ -1430,84 +1431,3 @@ int bitmap_has_oid_in_uninteresting(struct bitmap_index *bitmap_git,\n \treturn bitmap_git &&\n \t\tbitmap_walk_contains(bitmap_git, bitmap_git->haves, oid);\n }\n-\n-static off_t get_disk_usage_for_type(struct bitmap_index *bitmap_git,\n-\t\t\t\t     enum object_type object_type)\n-{\n-\tstruct bitmap *result = bitmap_git->result;\n-\tstruct packed_git *pack = bitmap_git->pack;\n-\toff_t total = 0;\n-\tstruct ewah_iterator it;\n-\teword_t filter;\n-\tsize_t i;\n-\n-\tinit_type_iterator(&it, bitmap_git, object_type);\n-\tfor (i = 0; i < result->word_alloc &&\n-\t\t\tewah_iterator_next(&filter, &it); i++) {\n-\t\teword_t word = result->words[i] & filter;\n-\t\tsize_t base = (i * BITS_IN_EWORD);\n-\t\tunsigned offset;\n-\n-\t\tif (!word)\n-\t\t\tcontinue;\n-\n-\t\tfor (offset = 0; offset < BITS_IN_EWORD; offset++) {\n-\t\t\tsize_t pos;\n-\n-\t\t\tif ((word >> offset) == 0)\n-\t\t\t\tbreak;\n-\n-\t\t\toffset += ewah_bit_ctz64(word >> offset);\n-\t\t\tpos = base + offset;\n-\t\t\ttotal += pack_pos_to_offset(pack, pos + 1) -\n-\t\t\t\t pack_pos_to_offset(pack, pos);\n-\t\t}\n-\t}\n-\n-\treturn total;\n-}\n-\n-static off_t get_disk_usage_for_extended(struct bitmap_index *bitmap_git)\n-{\n-\tstruct bitmap *result = bitmap_git->result;\n-\tstruct packed_git *pack = bitmap_git->pack;\n-\tstruct eindex *eindex = &bitmap_git->ext_index;\n-\toff_t total = 0;\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\toff_t object_size;\n-\tsize_t i;\n-\n-\toi.disk_sizep = &object_size;\n-\n-\tfor (i = 0; i < eindex->count; i++) {\n-\t\tstruct object *obj = eindex->objects[i];\n-\n-\t\tif (!bitmap_get(result, pack->num_objects + i))\n-\t\t\tcontinue;\n-\n-\t\tif (oid_object_info_extended(the_repository, &obj->oid, &oi, 0) < 0)\n-\t\t\tdie(_(\"unable to get disk usage of %s\"),\n-\t\t\t    oid_to_hex(&obj->oid));\n-\n-\t\ttotal += object_size;\n-\t}\n-\treturn total;\n-}\n-\n-off_t get_disk_usage_from_bitmap(struct bitmap_index *bitmap_git,\n-\t\t\t\t struct rev_info *revs)\n-{\n-\toff_t total = 0;\n-\n-\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_COMMIT);\n-\tif (revs->tree_objects)\n-\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_TREE);\n-\tif (revs->blob_objects)\n-\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_BLOB);\n-\tif (revs->tag_objects)\n-\t\ttotal += get_disk_usage_for_type(bitmap_git, OBJ_TAG);\n-\n-\ttotal += get_disk_usage_for_extended(bitmap_git);\n-\n-\treturn total;\n-}\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nindex 36d99930d8..ba71a9f5c6 100644\n--- a/pack-bitmap.h\n+++ b/pack-bitmap.h\n@@ -38,6 +38,7 @@ typedef int (*show_reachable_fn)(\n \tint flags,\n \tuint32_t hash,\n \tstruct packed_git *found_pack,\n+\tuint32_t pack_pos,\n \toff_t found_offset);\n \n struct bitmap_index;\ndiff --git a/reachable.c b/reachable.c\nindex 77a60c70a5..79ebe8f940 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -182,6 +182,7 @@ static int mark_object_seen(const struct object_id *oid,\n \t\t\t     int exclude,\n \t\t\t     uint32_t name_hash,\n \t\t\t     struct packed_git *found_pack,\n+\t\t\t     uint32_t pack_pos,\n \t\t\t     off_t found_offset)\n {\n \tstruct object *obj = lookup_object_by_type(the_repository, oid, type);\n"},{"id":"416563","messageId":"xmqq8s7x0wra.fsf@gitster.c.googlers.com","threadId":"55062","inReplyTo":"YCJtbmaguIW+YeAs@coredump.intra.peff.net","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-09T21:14:17Z","receivedAt":"2021-02-09T21:46:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I don't know that it's really worth digging into that much, though it's\n> quite possible there may be some easy wins by optimizing those memcpy\n> calls. E.g., I'm not sure if the compiler ends up inlining them or not.\n> If it doesn't realize that the_hash_algo->rawsz is only ever \"20\" or\n> \"32\", we could perhaps help it along with specialized versions of\n> hashcpy(). If somebody does want to play with it, this patch may make a\n> good testbed. :)\n\nYuck.  That reminds me of the adventure Shawn he made in the Java\nland benchmarking which one among int[5], int a,b,c,d,e, char[40] is\nthe most efficient way (both storage-wise and performance-wise) to\nstore SHA-1 hash.  I wish we didn't have to go there.\n\nIt indeed is an interesting, despite a bit sad, observation that\neven with a good precomputed information, an overly heavy interface\ncan kill potential performance benefit.\n\nThanks.\n"},{"id":"416597","messageId":"xmqqh7mkycno.fsf@gitster.c.googlers.com","threadId":"55062","inReplyTo":"YCJpbPIlSpCAKSBF@coredump.intra.peff.net","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-10T00:44:27Z","receivedAt":"2021-02-10T00:47:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Here's a re-roll of my series to add \"rev-list --disk-usage\", for\n> counting up object storage used for various slices of history.\n> ...\n>  t/t6114-rev-list-du.sh             | 51 +++++++++++++++++++\n>  t/test-lib-functions.sh            |  9 +++-\n>  7 files changed, 199 insertions(+), 8 deletions(-)\n>  create mode 100755 t/t6114-rev-list-du.sh\n\nI relocated 6114 to 6115 to avoid tests sharing the same number.\n\nI am getting these numbers from random ranges I am interested in,\nbut do they say what I think they mean?  Was the development effort\nwent into the v2.28 release almost half the size of v2.29, and have\nwe already done about the same amont of work for this cycle?\n\n: gitster git.git/seen; rungit seen rev-list --disk-usage master..next\n83105\n: gitster git.git/seen; rungit seen rev-list --disk-usage v2.30.0..master\n183463\n: gitster git.git/seen; rungit seen rev-list --disk-usage v2.29.0..v2.30.0\n231640\n: gitster git.git/seen; rungit seen rev-list --disk-usage v2.28.0..v2.29.0\n334355\n: gitster git.git/seen; rungit seen rev-list --disk-usage v2.27.0..v2.28.0\n182298\n"},{"id":"416599","messageId":"YCM7t3buBR6sL/lh@nand.local","threadId":"55062","inReplyTo":"xmqqh7mkycno.fsf@gitster.c.googlers.com","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-10T01:49:43Z","receivedAt":"2021-02-10T01:52:32Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Feb 09, 2021 at 04:44:27PM -0800, Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n>\n> > Here's a re-roll of my series to add \"rev-list --disk-usage\", for\n> > counting up object storage used for various slices of history.\n> > ...\n> >  t/t6114-rev-list-du.sh             | 51 +++++++++++++++++++\n> >  t/test-lib-functions.sh            |  9 +++-\n> >  7 files changed, 199 insertions(+), 8 deletions(-)\n> >  create mode 100755 t/t6114-rev-list-du.sh\n>\n> I relocated 6114 to 6115 to avoid tests sharing the same number.\n\nThanks.\n\n> I am getting these numbers from random ranges I am interested in,\n> but do they say what I think they mean?  Was the development effort\n> went into the v2.28 release almost half the size of v2.29, and have\n> we already done about the same amont of work for this cycle?\n>\n> : gitster git.git/seen; rungit seen rev-list --disk-usage master..next\n> 83105\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.30.0..master\n> 183463\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.29.0..v2.30.0\n> 231640\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.28.0..v2.29.0\n> 334355\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.27.0..v2.28.0\n> 182298\n\nI think you are surprised by these numbers because you're only counting\ndisk usage of commit objects in those ranges. v1 of this series implied\n--objects by default, but this changed in v2 due to my suggestion.\n\nPassing --objects to count the disk-usage of all objects in those ranges\ngives more reasonable numbers (and match my rough guesses, i.e., that\n2.29 was busier than 2.30, and so on):\n\n    $ for range in origin/master..origin/next v2.30.0..origin/master \\\n        v2.29.0..v2.30.0 v2.28.0..v2.29.0 v2.27.0..v2.28.0\n    do\n      printf \"%s %d vs. %d\\n\" $range \\\n        \"$(git rev-list --objects --no-object-names $range |\n           git cat-file --batch-check='%(objectsize:disk)' |\n           paste -sd+ | bc)\" \\\n        \"$(git.seen rev-list --objects --disk-usage $range)\"\n    done\n    origin/master..origin/next 671380 vs. 671380\n    v2.30.0..origin/master 1618815 vs. 1618815\n    v2.29.0..v2.30.0 3308295 vs. 3308295\n    v2.28.0..v2.29.0 4080789 vs. 4080789\n    v2.27.0..v2.28.0 2846196 vs. 2846196\n\nThanks,\nTaylor\n"},{"id":"416610","messageId":"YCOpq5fDYp+YEzEu@coredump.intra.peff.net","threadId":"55062","inReplyTo":"xmqq8s7x0wra.fsf@gitster.c.googlers.com","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-10T09:38:51Z","receivedAt":"2021-02-10T09:42:15Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 09, 2021 at 01:14:17PM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > I don't know that it's really worth digging into that much, though it's\n> > quite possible there may be some easy wins by optimizing those memcpy\n> > calls. E.g., I'm not sure if the compiler ends up inlining them or not.\n> > If it doesn't realize that the_hash_algo->rawsz is only ever \"20\" or\n> > \"32\", we could perhaps help it along with specialized versions of\n> > hashcpy(). If somebody does want to play with it, this patch may make a\n> > good testbed. :)\n> \n> Yuck.  That reminds me of the adventure Shawn he made in the Java\n> land benchmarking which one among int[5], int a,b,c,d,e, char[40] is\n> the most efficient way (both storage-wise and performance-wise) to\n> store SHA-1 hash.  I wish we didn't have to go there.\n> \n> It indeed is an interesting, despite a bit sad, observation that\n> even with a good precomputed information, an overly heavy interface\n> can kill potential performance benefit.\n\nAgreed. But I'm hoping we can continue to mostly ignore it. I suspect\nthis finding means we are wasting a few hundred milliseconds copying\noids around during a clone of torvalds/linux. But overall that is a\npretty heavy-weight operation, and I doubt anybody really notices. And\nfor something as lightweight as --disk-usage, it was easy enough to\noptimize around it.\n\nIt probably does have a more measurable impact in something like:\n\n  git rev-list --use-bitmap-index --objects HEAD >/dev/null\n\nwhere we really do need those oids, and the extra copying might add up.\nI guess if somebody is interested in micro-optimizing, that is probably\na good command to look at.\n\n-Peff\n"},{"id":"416611","messageId":"YCOu70m5SKU7L4CS@coredump.intra.peff.net","threadId":"55062","inReplyTo":"xmqqh7mkycno.fsf@gitster.c.googlers.com","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-10T10:01:19Z","receivedAt":"2021-02-10T10:04:38Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 09, 2021 at 04:44:27PM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > Here's a re-roll of my series to add \"rev-list --disk-usage\", for\n> > counting up object storage used for various slices of history.\n> > ...\n> >  t/t6114-rev-list-du.sh             | 51 +++++++++++++++++++\n> >  t/test-lib-functions.sh            |  9 +++-\n> >  7 files changed, 199 insertions(+), 8 deletions(-)\n> >  create mode 100755 t/t6114-rev-list-du.sh\n> \n> I relocated 6114 to 6115 to avoid tests sharing the same number.\n\nThanks. I wondered why I didn't notice, but it's because the other 6114\nalso just made it into \"seen\". :)\n\n> I am getting these numbers from random ranges I am interested in,\n> but do they say what I think they mean?  Was the development effort\n> went into the v2.28 release almost half the size of v2.29, and have\n> we already done about the same amont of work for this cycle?\n> \n> : gitster git.git/seen; rungit seen rev-list --disk-usage master..next\n> 83105\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.30.0..master\n> 183463\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.29.0..v2.30.0\n> 231640\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.28.0..v2.29.0\n> 334355\n> : gitster git.git/seen; rungit seen rev-list --disk-usage v2.27.0..v2.28.0\n> 182298\n\nAs Taylor mentioned, this is only hitting the commits. So you might as\nwell just be looking at commit counts as a measure of work, I'd think\n(and indeed v2.28 has about half as many commits as v2.29!).\n\nAdding --objects gets you a rougher estimate of \"bytes changed\", which\nhelps accounts for commits of different sizes. But there I think you'd\ndo just as well to look at the actual number of lines changed with \"git\ndiff --numstat\".\n\nI'd expect the number of on-disk bytes to _roughly_ correspond to the\nsize of the changes. But you are working against the heuristics of the\ndelta chains there. It may well be that we would store a base object in\nthe v2.28..v2.29 range, and a delta against it in v2.27..v2.28. And that\nwould attribute most of the bytes to v2.29, even though they should be\nshared roughly with v2.28.\n\nI'm sure one could devise a scheme for \"sharing\" the bytes from a delta\nfamily across all of its objects. That might even be worth implementing\non top (I don't even think it would be too expensive; you just have to\ncollect the delta chains for any objects you're reporting, and then\naverage the total size among a chain).\n\nBut in practice, we've found this kind of naive --disk-usage useful for\nanswering questions like:\n\n  - do I need all of these objects? Comparing \"rev-list --disk-usage\n    --objects --all\", \"rev-list --disk-usage --objects --all --reflog\",\n    and \"du objects/pack/*.pack\" will tell you if a prune/repack might\n    help, and whether expiring reflogs makes a difference.\n\n  - the size of the shared alternates repo for a set of forks has\n    jumped. Comparing \"rev-list --disk-usage --objects --remotes=$base\n    --not --remotes=$fork\" will tell you what's reachable from a fork\n    but not from the base (we use \"refs/remotes/$id/*\" to keep track of\n    fork refs in our alternates repo). This can be junk like somebody\n    forking git/git and then uploading a bunch of pirated video files.\n    :)\n\n  - likewise, the size of cloning a single repo may jump. Comparing\n    \"rev-list --disk-usage --objects HEAD..$branch\" for each branch\n    might show that one branch is an outlier (e.g., because somebody\n    accidentally committed a bunch of build artifacts).\n\nIn those kinds of cases, it's not usually \"oh, this version is twice as\nbig as this other one\". It's more like \"wow, this branch is 100x as big\nas the other branches\", and little decisions like delta direction are\njust noise. I imagine that in those cases the uncompressed object sizes\nwould probably produce similar patterns and answers. But it's actually\nfaster to produce the on-disk sizes. :)\n\n-Peff\n"},{"id":"416642","messageId":"xmqq1rdn51gz.fsf@gitster.c.googlers.com","threadId":"55062","inReplyTo":"YCOu70m5SKU7L4CS@coredump.intra.peff.net","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-10T16:31:08Z","receivedAt":"2021-02-10T16:33:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> But in practice, we've found this kind of naive --disk-usage useful for\n> answering questions like:\n>\n>   - do I need all of these objects? Comparing \"rev-list --disk-usage\n>     --objects --all\", \"rev-list --disk-usage --objects --all --reflog\",\n>     and \"du objects/pack/*.pack\" will tell you if a prune/repack might\n>     help, and whether expiring reflogs makes a difference.\n>\n>   - the size of the shared alternates repo for a set of forks has\n>     jumped. Comparing \"rev-list --disk-usage --objects --remotes=$base\n>     --not --remotes=$fork\" will tell you what's reachable from a fork\n>     but not from the base (we use \"refs/remotes/$id/*\" to keep track of\n>     fork refs in our alternates repo). This can be junk like somebody\n>     forking git/git and then uploading a bunch of pirated video files.\n>     :)\n>\n>   - likewise, the size of cloning a single repo may jump. Comparing\n>     \"rev-list --disk-usage --objects HEAD..$branch\" for each branch\n>     might show that one branch is an outlier (e.g., because somebody\n>     accidentally committed a bunch of build artifacts).\n>\n> In those kinds of cases, it's not usually \"oh, this version is twice as\n> big as this other one\". It's more like \"wow, this branch is 100x as big\n> as the other branches\", and little decisions like delta direction are\n> just noise. I imagine that in those cases the uncompressed object sizes\n> would probably produce similar patterns and answers. But it's actually\n> faster to produce the on-disk sizes. :)\n\nThanks.\n\nI kind of feel sad to have a nice write-up like this only in the\nlist archive.  Is there a section in our documentation set to keep\ncollection of such a real-life use cases?  Perhaps the examples\nsection of manpages is the closest thing, but it looks a bit too\nnarrowly scoped for the example section of \"rev-list\" manpage.\n\nTHanks.\n\n"},{"id":"416661","messageId":"YCREYmBsnv2wgvXZ@coredump.intra.peff.net","threadId":"55062","inReplyTo":"xmqq1rdn51gz.fsf@gitster.c.googlers.com","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-10T20:38:58Z","receivedAt":"2021-02-10T20:39:54Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 10, 2021 at 08:31:08AM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > But in practice, we've found this kind of naive --disk-usage useful for\n> > answering questions like:\n> [...]\n>\n> I kind of feel sad to have a nice write-up like this only in the\n> list archive.  Is there a section in our documentation set to keep\n> collection of such a real-life use cases?  Perhaps the examples\n> section of manpages is the closest thing, but it looks a bit too\n> narrowly scoped for the example section of \"rev-list\" manpage.\n\nAgreed on both counts. If this gets put into a release, I suspect Taylor\nwould cover it in a release blog post. That is not quite the same thing\nas having it in the documentation, but it may provide more search engine\nboost than the list archive. I dunno.\n\n-Peff\n"},{"id":"416681","messageId":"YCRpBCNJ2yNTbc2i@nand.local","threadId":"55062","inReplyTo":"YCREYmBsnv2wgvXZ@coredump.intra.peff.net","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-10T23:15:16Z","receivedAt":"2021-02-10T23:18:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 10, 2021 at 03:38:58PM -0500, Jeff King wrote:\n> On Wed, Feb 10, 2021 at 08:31:08AM -0800, Junio C Hamano wrote:\n>\n> > Jeff King <peff@peff.net> writes:\n> >\n> > > But in practice, we've found this kind of naive --disk-usage useful for\n> > > answering questions like:\n> > [...]\n> >\n> > I kind of feel sad to have a nice write-up like this only in the\n> > list archive.  Is there a section in our documentation set to keep\n> > collection of such a real-life use cases?  Perhaps the examples\n> > section of manpages is the closest thing, but it looks a bit too\n> > narrowly scoped for the example section of \"rev-list\" manpage.\n>\n> Agreed on both counts. If this gets put into a release, I suspect Taylor\n> would cover it in a release blog post. That is not quite the same thing\n> as having it in the documentation, but it may provide more search engine\n> boost than the list archive. I dunno.\n\nYeah, this is the perfect sort of thing for those blog posts.\n\nBut it makes sense to include some of these examples in our own\ndocumentation here, too. git-rev-list(1) doesn't have an EXAMPLES\nsection, but maybe it should.\n\n> -Peff\n\nThanks,\nTaylor\n"},{"id":"416740","messageId":"YCUOSmnsJ4LLPFgK@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YCRpBCNJ2yNTbc2i@nand.local","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-11T11:00:26Z","receivedAt":"2021-02-11T11:06:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 10, 2021 at 06:15:16PM -0500, Taylor Blau wrote:\n\n> > > I kind of feel sad to have a nice write-up like this only in the\n> > > list archive.  Is there a section in our documentation set to keep\n> > > collection of such a real-life use cases?  Perhaps the examples\n> > > section of manpages is the closest thing, but it looks a bit too\n> > > narrowly scoped for the example section of \"rev-list\" manpage.\n> >\n> > Agreed on both counts. If this gets put into a release, I suspect Taylor\n> > would cover it in a release blog post. That is not quite the same thing\n> > as having it in the documentation, but it may provide more search engine\n> > boost than the list archive. I dunno.\n> \n> Yeah, this is the perfect sort of thing for those blog posts.\n> \n> But it makes sense to include some of these examples in our own\n> documentation here, too. git-rev-list(1) doesn't have an EXAMPLES\n> section, but maybe it should.\n\nI think this is the \"narrowly scoped\" bit from Junio's response above.\nIt would be a bit weird to have an examples section for rev-list that\nonly mentions this rather obscure feature.\n\n-Peff\n"},{"id":"416743","messageId":"875z2ydd4l.fsf@evledraar.gmail.com","threadId":"55062","inReplyTo":"YCUOSmnsJ4LLPFgK@coredump.intra.peff.net","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-02-11T12:04:26Z","receivedAt":"2021-02-11T12:08:19Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Feb 11 2021, Jeff King wrote:\n\n> On Wed, Feb 10, 2021 at 06:15:16PM -0500, Taylor Blau wrote:\n>\n>> > > I kind of feel sad to have a nice write-up like this only in the\n>> > > list archive.  Is there a section in our documentation set to keep\n>> > > collection of such a real-life use cases?  Perhaps the examples\n>> > > section of manpages is the closest thing, but it looks a bit too\n>> > > narrowly scoped for the example section of \"rev-list\" manpage.\n>> >\n>> > Agreed on both counts. If this gets put into a release, I suspect Taylor\n>> > would cover it in a release blog post. That is not quite the same thing\n>> > as having it in the documentation, but it may provide more search engine\n>> > boost than the list archive. I dunno.\n>> \n>> Yeah, this is the perfect sort of thing for those blog posts.\n>> \n>> But it makes sense to include some of these examples in our own\n>> documentation here, too. git-rev-list(1) doesn't have an EXAMPLES\n>> section, but maybe it should.\n>\n> I think this is the \"narrowly scoped\" bit from Junio's response above.\n> It would be a bit weird to have an examples section for rev-list that\n> only mentions this rather obscure feature.\n\nI don't think the lack of an EXAMPLES section or the relative obscurity\nof the switch should preclude us from adding useful documentation.\n\nYes it would feel a bit out of place, but we can always have a\nsub-section of EXAMPLES, and we've got to start somewhere.\n\nIn this case I don't see why it couldn't be added to OPTIONS, we've got\nsome very long discussion there already, and as long as there's a clear\nseparation in prose from an initial brief discussion of the switch and\nfurther prose it won't be confusing for readers, they can just page past\nthe details.\n"},{"id":"416751","messageId":"xmqq1rdmxzbb.fsf@gitster.c.googlers.com","threadId":"55062","inReplyTo":"875z2ydd4l.fsf@evledraar.gmail.com","subject":"Re: [PATCH v2] rev-list --disk-usage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-11T17:57:12Z","receivedAt":"2021-02-11T18:09:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> I think this is the \"narrowly scoped\" bit from Junio's response above.\n>> It would be a bit weird to have an examples section for rev-list that\n>> only mentions this rather obscure feature.\n>\n> I don't think the lack of an EXAMPLES section or the relative obscurity\n> of the switch should preclude us from adding useful documentation.\n>\n> Yes it would feel a bit out of place, but we can always have a\n> sub-section of EXAMPLES, and we've got to start somewhere.\n>\n> In this case I don't see why it couldn't be added to OPTIONS, we've got\n> some very long discussion there already, and as long as there's a clear\n> separation in prose from an initial brief discussion of the switch and\n> further prose it won't be confusing for readers, they can just page past\n> the details.\n\nOK.\n\nIn any case, [v2] as we have it (with test number relocation) should\nbe good as-is, so I'd start preparing to merge it down to 'next'\nsoonish.\n\nThanks.\n"},{"id":"417244","messageId":"YC2nOxPP3SAY2g1I@coredump.intra.peff.net","threadId":"55062","inReplyTo":"875z2ydd4l.fsf@evledraar.gmail.com","subject":"[PATCH 0/2] rev-list --disk-usage example docs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T23:31:07Z","receivedAt":"2021-02-17T23:33:10Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 11, 2021 at 01:04:26PM +0100, Ævar Arnfjörð Bjarmason wrote:\n\n> > I think this is the \"narrowly scoped\" bit from Junio's response above.\n> > It would be a bit weird to have an examples section for rev-list that\n> > only mentions this rather obscure feature.\n> \n> I don't think the lack of an EXAMPLES section or the relative obscurity\n> of the switch should preclude us from adding useful documentation.\n> \n> Yes it would feel a bit out of place, but we can always have a\n> sub-section of EXAMPLES, and we've got to start somewhere.\n\nFair enough. Here are some patches (to go on top of jk/rev-list-disk-usage,\nthough obviously the first one could be applied independently).\n\n> In this case I don't see why it couldn't be added to OPTIONS, we've got\n> some very long discussion there already, and as long as there's a clear\n> separation in prose from an initial brief discussion of the switch and\n> further prose it won't be confusing for readers, they can just page past\n> the details.\n\nIt's already big and scary enough that I prefer starting an EXAMPLES\nsection. :)\n\nBy the way, there's one other finishing touch we might consider:\nenabling --use-bitmap-index automatically when bitmaps are present, for\nrequests that produce the identical answer (so _not_ a regular\ntraversal, because the output order and presence of pathnames are\ndifferent there). I'd prefer to do that as a separate series, though,\nsince there are multiple arguments that might benefit (like --count).\n\n  [1/2]: docs/rev-list: add an examples section\n  [2/2]: docs/rev-list: add some examples of --disk-usage\n\n Documentation/git-rev-list.txt | 93 ++++++++++++++++++++++++++++++++++\n 1 file changed, 93 insertions(+)\n\n-Peff\n"},{"id":"417246","messageId":"YC2n/R1O77sRICSQ@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YC2nOxPP3SAY2g1I@coredump.intra.peff.net","subject":"[PATCH 1/2] docs/rev-list: add an examples section","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T23:34:21Z","receivedAt":"2021-02-17T23:35:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"We currently don't show any examples of using git-rev-list at all. Let's\nadd some pretty elementary examples. They likely seem obvious to anybody\nwho has worked with the tool for a while, but my purpose here is\ntwo-fold:\n\n  - they may be enlightening to people who haven't used the tool a lot\n    to give a general flavor of how it is meant to be used\n\n  - they can serve as a starting point for adding more interesting\n    examples (we can do that without the basic ones, of course, but I\n    think it makes sense to show off the building blocks)\n\nThis set is far from exhaustive, but again, the purpose is to be a\nstarting point for further additions.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\nI'm open to feedback on these. But please, if you have suggestions for\nadding more, do it in the form of a patch on top. :)\n\n Documentation/git-rev-list.txt | 52 ++++++++++++++++++++++++++++++++++\n 1 file changed, 52 insertions(+)\n\ndiff --git a/Documentation/git-rev-list.txt b/Documentation/git-rev-list.txt\nindex 5da66232dc..d7ff519b90 100644\n--- a/Documentation/git-rev-list.txt\n+++ b/Documentation/git-rev-list.txt\n@@ -31,6 +31,58 @@ include::rev-list-options.txt[]\n \n include::pretty-formats.txt[]\n \n+EXAMPLES\n+--------\n+\n+* Print the list of commits reachable from the current branch.\n++\n+----------\n+git rev-list HEAD\n+----------\n+\n+* Print the list of commits on this branch, but not present in the\n+  upstream branch.\n++\n+----------\n+git rev-list @{upstream}..HEAD\n+----------\n+\n+* Format commits with their author and commit message (see also the\n+  porcelain linkgit:git-log[1]).\n++\n+----------\n+git rev-list --format=medium HEAD\n+----------\n+\n+* Format commits along with their diffs (see also the porcelain\n+  linkgit:git-log[1], which can do this in a single process).\n++\n+----------\n+git rev-list HEAD |\n+git diff-tree --stdin --format=medium -p\n+----------\n+\n+* Print the list of commits on the current branch that touched any\n+  file in the `Documentation` directory.\n++\n+----------\n+git rev-list HEAD -- Documentation/\n+----------\n+\n+* Print the list of commits authored by you in the past year, on\n+  any branch, tag, or other ref.\n++\n+----------\n+git rev-list --author=you@example.com --since=1.year.ago --all\n+----------\n+\n+* Print the list of objects reachable from the current branch (i.e., all\n+  commits and the blobs and trees they contain).\n++\n+----------\n+git rev-list --objects HEAD\n+----------\n+\n GIT\n ---\n Part of the linkgit:git[1] suite\n-- \n2.30.1.989.g5e01c2f281\n\n"},{"id":"417247","messageId":"YC2oRaMtBo/zMBmi@coredump.intra.peff.net","threadId":"55062","inReplyTo":"YC2nOxPP3SAY2g1I@coredump.intra.peff.net","subject":"[PATCH 2/2] docs/rev-list: add some examples of --disk-usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-02-17T23:35:33Z","receivedAt":"2021-02-17T23:36:29Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"It's not immediately obvious why --disk-usage might be a useful thing.\nThese examples show off a few of the real-world cases I've used it for.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/git-rev-list.txt | 41 ++++++++++++++++++++++++++++++++++\n 1 file changed, 41 insertions(+)\n\ndiff --git a/Documentation/git-rev-list.txt b/Documentation/git-rev-list.txt\nindex d7ff519b90..20bb8e8217 100644\n--- a/Documentation/git-rev-list.txt\n+++ b/Documentation/git-rev-list.txt\n@@ -83,6 +83,47 @@ git rev-list --author=you@example.com --since=1.year.ago --all\n git rev-list --objects HEAD\n ----------\n \n+* Compare the disk size of all reachable objects, versus those\n+  reachable from reflogs, versus the total packed size. This can tell\n+  you whether running `git repack -ad` might reduce the repository size\n+  (by dropping unreachable objects), and whether expiring reflogs might\n+  help.\n++\n+----------\n+# reachable objects\n+git rev-list --disk-usage --objects --all\n+# plus reflogs\n+git rev-list --disk-usage --objects --all --reflog\n+# total disk size used\n+du -c .git/objects/pack/*.pack .git/objects/??/*\n+# alternative to du: add up \"size\" and \"size-pack\" fields\n+git count-objects -v\n+----------\n+\n+* Report the disk size of each branch, not including objects used by the\n+  current branch. This can find outliers that are contributing to a\n+  bloated repository size (e.g., because somebody accidentally committed\n+  large build artifacts).\n++\n+----------\n+git for-each-ref --format='%(refname)' |\n+while read branch\n+do\n+\tsize=$(git rev-list --disk-usage --objects HEAD..$branch)\n+\techo \"$size $branch\"\n+done |\n+sort -n\n+----------\n+\n+* Compare the on-disk size of branches in one group of refs, excluding\n+  another. If you co-mingle objects from multiple remotes in a single\n+  repository, this can show which remotes are contributing to the\n+  repository size (taking the size of `origin` as a baseline).\n++\n+----------\n+git rev-list --disk-usage --objects --remotes=$suspect --not --remotes=origin\n+----------\n+\n GIT\n ---\n Part of the linkgit:git[1] suite\n-- \n2.30.1.989.g5e01c2f281\n"},{"id":"417249","messageId":"YC2qTedr8agOpQxy@nand.local","threadId":"55062","inReplyTo":"YC2nOxPP3SAY2g1I@coredump.intra.peff.net","subject":"Re: [PATCH 0/2] rev-list --disk-usage example docs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-02-17T23:44:13Z","receivedAt":"2021-02-17T23:44:59Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Feb 17, 2021 at 06:31:07PM -0500, Jeff King wrote:\n> It's already big and scary enough that I prefer starting an EXAMPLES\n> section. :)\n\nThe patches you sent below are great. I think that it's easy to nitpick\nand say \"oh, you should have added this or that example, too\", but I\nthink you gave a great set of starting examples.\n\nI'd be happy to see this merged so that others can add more examples on\ntop.\n\n> By the way, there's one other finishing touch we might consider:\n> enabling --use-bitmap-index automatically when bitmaps are present, for\n> requests that produce the identical answer (so _not_ a regular\n> traversal, because the output order and presence of pathnames are\n> different there). I'd prefer to do that as a separate series, though,\n> since there are multiple arguments that might benefit (like --count).\n\nThis would be really neat. I look forward to it.\n\nThanks,\nTaylor\n"}]}