Re: [PATCH v6] doc: add an explanation of Git's data model
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Nov 9, 2025, 04:59 UTC
- Message-ID
- <xmqqa50v4x8n.fsf@gitster.g>
- In-Reply-To
- <D50AB3E0-E41C-49CD-9407-AB60331A6A43@gmail.com>
Ben Knoble <ben.knoble@gmail.com> writes:
> My only other opinion on the matter is: what does making this > distinction clear do to benefit readers of this document?
I care about teaching people not just _what_ but _why_, because with vague distinction, many tend to memorize _what_ without understanding the reasoning behind it. "Our object names are computed as a hash of the contents in it formatted in a canonical way" is "what we do to compute an object name", but the reason behind the design is because we want to be able to dedup the same thing cheaply, detect two objects that are different cheaply, which is "why" in this example and it is equally, if not more, important.
The refs and objects record object names, and that is "what"; the reason why they do so is to refer to these objects. If somebody comes up with other ways to uniquely refer to these objects, their implementation of git-compatible system does not have to make their refs record object names---they can draw a line from a circle to a rectangle instead of writing the object name of that rectangle in the circle---and their system is still compatible with the Git data model at the higher/conceptual level. IOW, what exactly is done at the byte level (like file format) is lower part of the "data model", but what these byte level details wants to achieve is the other, higher half of the "data model". A data model documentation should teach both levels.