Re: [PATCH 5/5] builtin/repo: find tree with most entries
- From
Patrick Steinhardt <ps@pks.im>
- Date
- Feb 4, 2026, 08:28 UTC
- Message-ID
- <aYMDL4m7Ceifl1Ja@pks.im>
- In-Reply-To
- <xmqqldh9qw5d.fsf@gitster.g>
On Tue, Feb 03, 2026 at 02:50:38PM -0800, Junio C Hamano wrote:
Show 10 quoted lines
> Justin Tobler <jltobler@gmail.com> writes: > > > The size of a tree object usually corresponds with the number of entries > > it has. While iterating through objects in the repository for > > git-repo-structure, identify the tree with the most entries and display > > it in the output. > > All of these "largest" and "most", it would be a lot more > interesting if we can give not just these extreme values but > distrubution, possibly in a graphical way for bonus points.
That would be amazing indeed! I think having the largest values is still valuable as it allows you to detect weird outliers quite easily. But having a histogram would of course give the bigger picture.
I guess the challenging part would be to compute the buckets of that histogram in a streaming fashion. But I guess we could:
1. Pick a target number of buckets.
2. Track the maximum respective values as we stream.
3. Merge existing buckets and create new ones in case the maximum
value changes.The target number of buckets may not necessarily be the same number as the number of buckets that we will eventually print for increased resolution.
The distributions could then be printed as an ASCII bar chart, for example something like:
0-50 │████████████████████████████████████████ 1,247
50-100 │█████████████████████████████████ 812
100-150 │█████████████ 401
150-200 │████████ 253
200-250 │████ 128
250-300 │██ 67
300+ │▏ 12
0 250 500 750 1k 1.2k count
bytes / countFrom my point of view that would be the cherry on top of the new tool :) I'd personally still like to learn about maximum values in the table, as I've found that info to be useful with some customer incidents in the past. It's not giving you a trend, but it immediately gives you some good signal that the repo shape might be weird if you have commits with hundreds of parents.
So maybe this is another step we can do in a subsequent patch series?
Thanks!
Patrick