Re: [PATCH v2 0/5] builtin/repo: include largest object information
- From
Justin Tobler <jltobler@gmail.com>
- Date
- Mar 1, 2026, 19:22 UTC
- Message-ID
- <aaR6a7o4omOIWJSe@denethor>
- In-Reply-To
- <EB04AA40-87BA-41D9-B2DC-92E87FACEB54@gmail.com>
On 26/02/28 08:43PM, Lucas Seiki Oshiro wrote:
> I was trying this patch series and I noticed that it took > more time to run than before. In my machine, I tested it > with the Git repository itself and it took 6s to run, while > it took 3s to run in the current master [1].
Yes, now that objects are being parsed to fetch additional commit/tree information we incur some additional overhead when collecting metrics.
With git-repo-structure, the goal is to provide the user with an overview of size/structure related statistics that may showcase problems for a given repostiory and is directly inspired by git-sizer [1]. Thus as it currently stands, the implementation of git-repo-structure is still incomplete and as we collect additional metrics in subseqent series the performance characteristics may still change.
> I understand the reason and I don't think we could avoid > that, but I'm wondering if wouldn't be nice to have some > way to only retrieve the "lighter" data (perhaps a flag, > or something like the keys in git-repo-info).
If the main motivation is to allow the user to reduce the time spent by selecting only a subset of metrics, I don't think using keys like git-repo-info would be a good fit. Most of the collected metrics pull from the same data sources so including/excluding any given metric may not have any bearing on actual performance. For example: if the user wants to collect largest object info which is a more expensive check, we still have to collect the underlying data used by the other metrics regardless of if they are shown or not. Furthermore, it would likely not be obvious to users which categories of metrics would be more expensive than others.
I could maybe see something akin to a `--[no-]extended` option that breaks metrics into cheap/expensive categories and computes/displays the metrics accordingly, but it would be important that the default set of metrics collected satisfy the repository overview this command aims to provide.
If we are more interested in adding a mechanism to filter git-repo-structure results independent of performance considerations, maybe we could eventually explore adding something like the git-repo-info keys or a `--filter` option to restrict the output to a specified subset. At the same time though, it is probably easy enough for git-repo-structure users to filter the machine-parsable output themselves if they wish to do so. For now I think this should be fine, but an included result filtering option is still something we could explore in the future. :)
Thanks, -Justin
[1]: https://github.com/github/git-sizer