Re: [PATCH v6] doc: add an explanation of Git's data model
- From
Julia Evans <julia@jvns.ca>
- Date
- Nov 10, 2025, 15:56 UTC
- Message-ID
- <150f3442-93a6-4469-9c25-5bca24accc80@app.fastmail.com>
- In-Reply-To
- <xmqqa50v4x8n.fsf@gitster.g>
On Sat, Nov 8, 2025, at 11:59 PM, Junio C Hamano wrote:
Show 26 quoted lines
> Ben Knoble <ben.knoble@gmail.com> writes: > >> My only other opinion on the matter is: what does making this >> distinction clear do to benefit readers of this document? > > I care about teaching people not just _what_ but _why_, because with > vague distinction, many tend to memorize _what_ without > understanding the reasoning behind it. "Our object names are > computed as a hash of the contents in it formatted in a canonical > way" is "what we do to compute an object name", but the reason > behind the design is because we want to be able to dedup the same > thing cheaply, detect two objects that are different cheaply, which > is "why" in this example and it is equally, if not more, important. > > The refs and objects record object names, and that is "what"; the > reason why they do so is to refer to these objects. If somebody > comes up with other ways to uniquely refer to these objects, their > implementation of git-compatible system does not have to make their > refs record object names---they can draw a line from a circle to a > rectangle instead of writing the object name of that rectangle in > the circle---and their system is still compatible with the Git data > model at the higher/conceptual level. IOW, what exactly is done at > the byte level (like file format) is lower part of the "data model", > but what these byte level details wants to achieve is the other, > higher half of the "data model". A data model documentation should > teach both levels.
Thanks, this is exactly what I was looking for when I asked in what way this rephrasing helps the reader. I agree that explaining the "why" is very important.
It sounds like there are 2 "whats" and "whys" here:
#1:
what: object IDs are hashes of the contents
why: this makes it very fast to avoid storing duplicate information,
and it's extremely fast to check if 2 objects are the same or notI love the idea of explaining this. I think we could incorporate it very easily by adding this paragraph in the "Objects" introduction, right before "Here's how each type of object is structured":
The reason the ID is a cryptographic hash is that it makes it extremely
fast for Git to tell if 2 objects have the same contents or not
(if they have the same ID, they have the same contents!),
and it means Git will never store duplicate objects.Will add that unless there are any objections.
#2: what: The refs contain object IDs why: to refer to the object
I think this is so obvious that going out of our way to explain it risks confusing the reader. Spending too much time explaining something obvious can make the reader feel like they're missing something.
I can't imagine any purpose for the refs containing object IDs other than to refer to the object?
Like you noticed in the tag object section, I think saying that the tag object "refers to an object" works well in that context, but in the context of explaining what a branch is it makes the text more confusing.