From: Kristoffer Haugsbakk Date: Fri, 03 Oct 2025 21:46:00 GMT Subject: Re: [PATCH] doc: add a explanation of Git's data model Message-ID: <8df4c59c-4d27-4f36-a231-f7af32ddf149@app.fastmail.com> In-Reply-To: On Fri, Oct 3, 2025, at 19:34, Julia Evans via GitGitGadget wrote: > From: Julia Evans > > Git very often uses the terms "object", "reference", or "index" in its > documentation. > > However, it's hard to find a clear explanation of these terms and how > they relate to each other in the documentation. The closest candidates > currently are: > > 1. `gitglossary`. This makes a good effort, but it's an alphabetically > ordered dictionary and a dictionary is not a good way to learn > concepts. You have to jump around too much and it's not possible to > present the concepts in the order that they should be explained. > 2. `gitcore-tutorial`. This explains how to use the "core" Git commands. > This is a nice document to have, but it's not necessary to learn how > `update-index` works to understand Git's data model, and we should > not be requiring users to learn how to use the "plumbing" commands > if they want to learn what the term "index" or "object" means. > 3. `gitrepository-layout`. This is a great resource, but it includes a > lot of information about configuration and internal implementation > details which are not related to the data model. It also does > not explain how commits work. > > The result of this is that Git users (even users who have been using > Git for 15+ years) struggle to read the documentation because they don't > know what the core terms mean, and it's not possible to add links > to help them learn more. > > Add an explanation of Git's data model. Some choices I've made in > deciding what "core data model" means: > > 1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me > if those are intended to be user facing or if they're more like > internal implementation details. > 2. Don't talk about submodules other than by mentioning how they > relate to trees. This is because Git has a lot of special features, > and explaining how they all work exhaustively could quickly go > down a rabbit hole which would make this document less useful for > understanding Git's core behaviour. > 3. Don't discuss the structure of a commit message > (first line, trailers, GPG signatures, etc). > Perhaps this should change. > > Some other choices I've made: > > 1. Mention packed refs only in a note. I don’t think it’s worth mentioning this at all. More on that later. > 2. Don't mention that the full name of the branch `main` is > technically `refs/heads/main`. This should likely change but I > haven't worked out how to do it in a clear way yet. I think this is worth getting into. This is a pretty user-facing concept. > 3. Mostly avoid referring to the `.git` directory, because the exact > details of how things are stored change over time. > This should perhaps change from "mostly" to "entirely" > but I haven't worked out how to do that in a clear way yet. I think that’s good. I mean, I think us users don’t need that level of detail and shouldn’t be “inspired” to muck with the internals. If that makes sense. (See later) > > Signed-off-by: Julia Evans > --- > doc: Add a explanation of Git's data model >[snip] > diff --git a/Documentation/Makefile b/Documentation/Makefile >[snip] > diff --git a/Documentation/gitdatamodel.adoc > b/Documentation/gitdatamodel.adoc > new file mode 100644 > index 0000000000..4b2cb167dc > --- /dev/null > +++ b/Documentation/gitdatamodel.adoc > @@ -0,0 +1,226 @@ > +gitdatamodel(7) > +=============== > + > +NAME > +---- > +gitdatamodel - Git's core data model > + > +DESCRIPTION > +----------- > + > +It's not necessary to understand Git's data model to use Git, but it's > +very helpful when reading Git's documentation so that you know what it > +means when the documentation says "object" "reference" or "index". I haven’t gone hunting through the docs to see if this is covered elsewhere. But the thrust of all the things here definitely feel to me like something that should be presented and documented in such a way. > + > +Git's core operations use 4 kinds of data: Maybe small numerals should be spelled as words in running text? > + > +1. <>: commits, trees, blobs, and tag objects > +2. <>: branches, tags, > + remote-tracking branches, etc > +3. <>, also known as the staging area > +4. <> Reflogs is certainly auxiliary ref data. What makes it qualify as one-of-the-four? I am open to it being both, to be clear. > + > +[[objects]] > +OBJECTS > +------- > + > +Commits, trees, blobs, and tag objects are all stored in Git's object > database. > +Every object has: > + > +1. an *ID*, which is the SHA-1 hash of its contents. > + It's fast to look up a Git object using its ID. > + The ID is usually represented in hexadecimal, like > + `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`. > +2. a *type*. There are 4 types of objects: > + <>, <>, <>, > + and <>. > +3. *contents*. The structure of the contents depends on the type. > + > +Once an object is created, it can never be changed. > +Here are the 4 types of objects: As a curious Git user this seems correct. > + > +[[commit]] > +commits:: > + A commit contains: > ++ > +1. Its *parent commit ID(s)*. The first commit in a repository has 0 > parents, Maybe this is a subjective style thing but is it necessary to use “(s)” when the context makes clear that it could be zero to many? Its *parent commit IDs. ... > + regular commits have 1 parent, merge commits have 2+ parents s/2+/two or more/ ? Same point as the “numeral” one above. > +2. A *commit message* > +3. All the *files* in the commit, stored as a *<>* > +4. An *author* and the time the commit was authored > +5. A *committer* and the time the commit was committed > ++ > +Here's how an example commit is stored: > ++ > +---- > +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a > +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647 > +author Maya 1759173425 -0400 > +committer Maya 1759173425 -0400 > + > +Add README > +---- > ++ > +Like all other objects, commits can never be changed after they're > created. > +For example, "amending" a commit with `git commit --amend` creates a > new commit. > +The old commit will eventually be deleted by `git gc`. Maybe this could be moved to a part about what happens (eventually) to unreachable objects? Mentioning `git gc` and how things will get deleted raises questions naturally. Like why would they be deleted? Okay that’s clear: the previous commit will be replaced by the amended one. Then when it is not reachable by anything (even the reflog) it will get garbage collected. It all follows. But is the reader necessarily mature enough in their understanding to make the inference? This is a long-winded way of saying: if you’re gonna discuss `git gc` you might need to go into all of these concepts. > + > +[[tree]] > +trees:: > + A tree is how Git represents a directory. It lists, for each item > in > + the tree: > ++ > +1. The *permissions*, for example `100644` > +2. The *type*: either <> (a file), `tree` (a directory), > + or <> (a Git submodule) > +3. The *object ID* > +4. The *filename* > ++ > +For example, this is how a tree containing one directory (`src`) and > one file > +(`README.md`) is stored: > ++ > +---- > +100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md > +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src > +---- > ++ > +*NOTE:* The permissions are in the same format as UNIX permissions, but > +the only allowed permissions for files (blobs) are 644 and 755. > + Makes sense. > +[[blob]] > +blobs:: > + A blob is how Git represents a file. A blob object contains the > + file's contents. > ++ > +Storing a new blob for every new version of a file can get big, so > +`git gc` periodically compresses objects for efficiency in > `.git/objects/pack`. This gets into mentioning implementation files(?) like you mentioned in the commit message. 1. That it’s a packfile and where it is might be too much detail for this doc 2. I vaguely recall documents discussing what happens to “storing every version” discussing deltas instead of packs? Again, I am not a Git developer though. > + > +[[tag-object]] > +tag objects:: > + Tag objects (also known as "annotated tags") contain: > ++ > +1. The *tagger* and tag date > +2. A *tag message*, similar to a commit message > +3. The *ID* of the object (often a commit) that they reference s/often/typically/ ? I know it can get tedious to caveat the 99% cases with things that are technically possible. Maybe if it gets “bad enough” there could be a part that explains/distinguishes the high-level/porcelain Git use and what is technically possible: you make a `git tag -a`, which is on a commit... except if you accidentally run it on top of an existing tag. Then even the porcelain won’t protect you from making a tag-on-tag. (But it will issue a warning I guess.) Hmm. Now I don’t know. > + > +[[references]] > +REFERENCES > +---------- > + > +References are a way to give a name to a commit. > +It's easier to remember "the changes I'm working on are on the `turtle` > +branch" than "the changes are in commit bb69721404348e". > +Git often uses "ref" as shorthand for "reference". Good. > + > +References that you create are stored in the `.git/refs` directory, > +and Git has a few special internal references like `HEAD` that are > stored > +in the base `.git` directory. Implementation file details. You also mention `.git/refs/heads/` below. But refs aren’t stored as files if you are using the *reftable* backend. And that backend will become the default for new repositories in Git 3.0, I think. How does reftable work? I don’t know. But I don’t think we need to know after reading this doc. :) To be clear: how files are stored might not matter here. > + > +References can either be: > + > +1. References to an object ID, usually a <> ID > +2. References to another reference. This is called a "symbolic > reference". You seem to have used `**` when introducing terms: This is a *symbolic reference* >[snip ref stuff] > + > +[[HEAD]] > +HEAD: `.git/HEAD`:: > + `HEAD` is where Git stores your current <>. > + `HEAD` is normally a symbolic reference to your current branch, for > + example `ref: refs/heads/main` if your current branch is `main`. > + `HEAD` can also be a direct reference to a commit ID, > + that's called "detached HEAD state". > + > +[[remote-tracking-branch]] > +remote tracking branches: `.git/refs/remotes//`:: > + A remote-tracking branch is a name for a commit ID. > + It's how Git stores the last-known state of a branch in a remote > + repository. `git fetch` updates remote-tracking branches. When > + `git status` says "you're up to date with origin/main", it's looking at > + this. Looks good. > + > +[[other-refs]] > +Other references:: > + Git tools may create references in any subdirectory of `.git/refs`. > + For example, linkgit:git-stash[1], linkgit:git-bisect[1], > + and linkgit:git-notes[1] all create their own references > + in `.git/refs/stash`, `.git/refs/bisect`, etc. > + Third-party Git tools may also create their own references. > ++ > +Git may also create references in the base `.git` directory > +other than `HEAD`, like `ORIG_HEAD`. > + > +*NOTE:* As an optimization, references may be stored as packed > +refs instead of in `.git/refs`. See linkgit:git-pack-refs[1]. I don’t know if this is relevant for both ref backends. And does it matter? > + > +[[index]] > +THE INDEX > +--------- > + > +The index, also known as the "staging area", contains the current > staged > +version of every file in your Git repository. When you commit, the > files > +in the index are used as the files in the next commit. > + > +Unlike a tree, the index is a flat list of files. > +Each index entry has 4 fields: > + > +1. The *permissions* > +2. The *<> ID* of the file > +3. The *filename* > +4. The *number*. This is normally 0, but if there's a merge conflict > + there can be multiple versions (with numbers 0, 1, 2, ..) > + of the same filename in the index. > + > +It's extremely uncommon to look at the index directly: normally you'd > +run `git status` to see a list of changes between the index and > <>. > +But you can use `git ls-files --stage` to see the index. > +Here's the output of `git ls-files --stage` in a repository with 2 > files: > + > +---- > +100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md > +100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py > +---- > + > +[[reflogs]] > +REFLOGS > +------- > + > +Git stores the history of branch, tag, and HEAD refs in a reflog > +(you should read "reflog" as "ref log"). Not every ref is logged by You’ve heard of the re-flog too? > +default, but any ref can be logged. > + > +Each reflog entry has: > + > +1. *Before/after *commit IDs* > +2. *User* who made the change, for example `Maya ` > +3. *Timestamp* > +4. *Log message*, for example `pull: Fast-forward` > + > +Reflogs only log changes made in your local repository. > +They are not shared with remotes. Makes sense. > + > +GIT > +--- > +Part of the linkgit:git[1] suite I appreciate that this is the first version and you might have plans after this one. But I wonder if this doc could use a fair number of `gitlink` to branch out to all the other parts. Like git-reflog(1), gitglossary(7). Thanks for starting on a whole new doc. That must take quite some effort.