Re: [PATCH] doc: add a explanation of Git's data model
- From
Julia Evans <julia@jvns.ca>
- Date
- Oct 6, 2025, 19:36 UTC
- Message-ID
- <51e0a55c-1f1d-4cae-9459-8c2b9220e52d@app.fastmail.com>
- In-Reply-To
- <8df4c59c-4d27-4f36-a231-f7af32ddf149@app.fastmail.com>
Thanks for the review!
Show 6 quoted lines
>> 2. Don't mention that the full name of the branch `main` is >> technically `refs/heads/main`. This should likely change but I >> haven't worked out how to do it in a clear way yet. > > I think this is worth getting into. This is a pretty > user-facing concept.
I think I'll see if I can figure out a way to mention this and at the same time remove most of the rest of the references to the `.git` directory when explaining references (which you talked about further down), including packed refs.
Show 9 quoted lines
>> + >> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects >> +2. <<references,References>>: branches, tags, >> + remote-tracking branches, etc >> +3. <<index,The index>>, also known as the staging area >> +4. <<reflogs,Reflogs>> > > Reflogs is certainly auxiliary ref data. What makes it qualify as > one-of-the-four? I am open to it being both, to be clear.
The reason I like to talk about reflogs is that it gives you a way to "undo" Git operations that can be really useful. And any Git command that updates refs can updates that ref's reflog.
Understanding how reflogs work helps to understand what the limitations of using reflogs to undo mistakes is: for example the index is not a ref, so you can't use the reflog to undo changes to the index.
Show 37 quoted lines
>> +2. A *commit message* >> +3. All the *files* in the commit, stored as a *<<tree,tree>>* >> +4. An *author* and the time the commit was authored >> +5. A *committer* and the time the commit was committed >> ++ >> +Here's how an example commit is stored: >> ++ >> +---- >> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a >> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647 >> +author Maya <maya@example.com> 1759173425 -0400 >> +committer Maya <maya@example.com> 1759173425 -0400 >> + >> +Add README >> +---- >> ++ >> +Like all other objects, commits can never be changed after they're >> created. >> +For example, "amending" a commit with `git commit --amend` creates a >> new commit. > >> +The old commit will eventually be deleted by `git gc`. > > Maybe this could be moved to a part about what happens (eventually) to > unreachable objects? > > Mentioning `git gc` and how things will get deleted raises > questions naturally. Like why would they be deleted? Okay > that’s clear: the previous commit will be replaced by the > amended one. Then when it is not reachable by anything > (even the reflog) it will get garbage collected. > > It all follows. But is the reader necessarily mature enough > in their understanding to make the inference? > > This is a long-winded way of saying: if you’re gonna discuss > `git gc` you might need to go into all of these concepts.
If folks here think this is a reasonable document to add to Git I'll try get some beta readers to read this, see which parts folks find confusing, and address those, keeping the `git gc` stuff in mind.
Similarly for the style comments.
Show 10 quoted lines
>> +blobs:: >> + A blob is how Git represents a file. A blob object contains the >> + file's contents. >> ++ >> +Storing a new blob for every new version of a file can get big, so >> +`git gc` periodically compresses objects for efficiency in >> `.git/objects/pack`. > > This gets into mentioning implementation files(?) like you mentioned in > the commit message.
That's true! The reason I think this is important to mention is that I find that people often "reject" information that they find implausible, even if it comes from a credible source. ("that can't be true! I must be not understanding correctly. Oh well, I'll just ignore that!")
I sometimes hear from users that "commits can't be snapshots", because it would take up too much disk space to store every version of every commit. So I find that sometimes explaining a little bit about the implementation can make the information more memorable.
Certainly I'm not able to remember details that don't make sense with my mental model of how computers work and I don't expect other people to either, so I think it's important to give an explanation that handles the biggest "objections".
Show 5 quoted lines
> 1. That it’s a packfile and where it is might be too much detail for > this doc > 2. I vaguely recall documents discussing what happens to “storing every > version” discussing deltas instead of packs? Again, I am not a Git > developer though.
I could be wrong about the details here, I'm not a Git developer either. From https://git-scm.com/book/en/v2/Git-Internals-Packfiles it looks like packfiles are implemented using deltas.
Show 10 quoted lines
>> + >> +References can either be: >> + >> +1. References to an object ID, usually a <<commit,commit>> ID >> +2. References to another reference. This is called a "symbolic >> reference". > > You seem to have used `**` when introducing terms: > > This is a *symbolic reference*
Thanks, will take a look at that.
Show 8 quoted lines
>> +[[reflogs]] >> +REFLOGS >> +------- >> + >> +Git stores the history of branch, tag, and HEAD refs in a reflog >> +(you should read "reflog" as "ref log"). Not every ref is logged by > > You’ve heard of the re-flog too?
haha exactly, I just want folks to understand why it's called that :)
> I appreciate that this is the first version and you might have plans > after this one. But I wonder if this doc could use a fair number of > `gitlink` to branch out to all the other parts. Like git-reflog(1), > gitglossary(7).
That's reasonable. Do you often use the "See also" section of man pages? I've never looked at them so I'm always curious about how people are actually using them in practice.
I also need to think about what else could link *to* this, because without attention to discoverability probably nobody will find it. My main idea so far is actually to add it to https://git-scm.com/learn but I wanted to send it here instead of adding it to the website directly because I thought it could benefit from a more detailed review.
> Thanks for starting on a whole new doc. That must take quite > some effort.
All the work on documentation takes a lot of effort, in some ways it's easier to write something new than to edit something existing :)