Re: [PATCH] doc: add a explanation of Git's data model
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Oct 7, 2025, 17:02 UTC
- Message-ID
- <xmqq4isalk5g.fsf@gitster.g>
- In-Reply-To
- <aOUkZa4_fq1hho7Q@pks.im>
Patrick Steinhardt <ps@pks.im> writes:
Show 16 quoted lines
>> +Git's core operations use 4 kinds of data: >> + >> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects >> +2. <<references,References>>: branches, tags, >> + remote-tracking branches, etc >> +3. <<index,The index>>, also known as the staging area >> +4. <<reflogs,Reflogs>> > > This list makes sense to me. There's of course more data structures in > Git, but all the other data structures shouldn't really matter to users > at all as they are mostly caches or internal details of the on-disk > format. > > There's potentially one exception though, namely the Git configuration. > I'd claim that Git "uses" the Git configuration similarly to how it uses > the others, but I get why it's not explicitly mentioned here.
The core operations do not use Git configuration any more than they use what is specified by the command line arguments.
Show 12 quoted lines
>> +[[objects]] >> +OBJECTS >> +------- >> + >> +Commits, trees, blobs, and tag objects are all stored in Git's object database. >> +Every object has: >> + >> +1. an *ID*, which is the SHA-1 hash of its contents. > > I think this needs to be adapted to not single out SHA-1 as the only > hashing algorithm. We already support SHA-256, so we should definitely > say that the algorithm can be swapped. Maybe something like:
Good point. Also officially they are called "object name".
> An *object ID*, which is the cryptographic hash of its contents. By > default, Git uses SHA-1 as object hash, but alternative hashes like > SHA-256 are supported.
I'd avoid "object name is the result of hashing X" which historically was a source of question: "why does 'sha1sum README.md' give different hash from 'git add README.md && git ls-files -s README.md'?"
It is an irrelevant implementation detail (and you'd eventually end up having to say "X is <type> SP <length> NUL <contents>").
An object name, which is derived cryptographically from its
type, size and contents. All versions of Git can use SHA-1 hash
function, but more recent versions of Git can also use SHA-256
hash function.Show 7 quoted lines
>> +commits:: >> + A commit contains: >> ++ >> +1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents, >> + regular commits have 1 parent, merge commits have 2+ parents > > I'd say "at least two parents" instead of "2+ parents".
Yup, that reads much better.
Show 11 quoted lines
>> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a >> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647 >> +author Maya <maya@example.com> 1759173425 -0400 >> +committer Maya <maya@example.com> 1759173425 -0400 >> + >> +Add README >> +---- > > In practice, commits can have other headers that are ignored by Git. But > that's certainly not part of Git's core data model, so I don't think we > should mention that here.
Third-party software can add truly garbage ones that do not have any meaning, and Git tolerates by ignoring them. But there are others that Git does pay attention to, like encoding, gpgsig, etc., which may worth mention (in the form that "these four are what you typically see, but there may be others" without even naming any).