git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v3] doc: add a explanation of Git's data model

From
Junio C Hamano <gitster@pobox.com>
Date
Oct 15, 2025, 19:58 UTC
Message-ID
<xmqqv7kgszr1.fsf@gitster.g>
In-Reply-To
<pull.1981.v3.git.1760476346040.gitgitgadget@gmail.com>
"Julia Evans via GitGitGadget" <gitgitgadget@gmail.com> writes:
Show 7 quoted lines
> +[[commit]]
> +commits::
> +    A commit contains these required fields
> +    (though there are other optional fields):
> ++
> +1. All the *files* in the commit, stored as the *<<tree,tree>>* ID of
> +   the commit's base directory.

"all the files' exact contents at the time of the commit" is what we mean here, and once readers know what a tree is, the above sentence would be understood as such, but "All the files" felt somewhat fuzzy. I wonder if presenting objects in bottom-up fashion makes it easier to see? Learn that a blob records exact content of a file, then learn that a tree records the set of paths with exact contents stored at these paths, and after that, learn that a commit records a tree, hence a snapshot of the whole set of contents. I dunno...

Show 6 quoted lines
> +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,
> +  regular commits have 1 parent, merge commits have 2 or more parents
> +3. An *author* and the time the commit was authored
> +4. A *committer* and the time the commit was committed.
> +   If you cherry-pick (linkgit:git-cherry-pick[1]) someone else's commit,
> +   then they will be the author and you'll be the committer.
It felt a bit odd to single-out cherry-pick here.

I think the important thing to become aware of for the readers at this point is that the author and committer can be different people, and it does not matter how one commits somebody else's patch at the mechanical level.

Perhaps replace "If you cherry-pick..." with something like "note: a change authored by a person at some point in time can be committed by another person at a different time, and these fields are to record both persons' contributions separately", perhaps, if we really want to say more.

> +Git does not store the diff for a commit: when you ask Git for a
> +diff it calculates it on the fly.

I think this is an attempt to demystify "are we really storing snapshot for each commit?" thing, but then "when you ask Git to show the commit, it calculates the diff from its parent on the fly" might achieve that better, perhaps?

Show 14 quoted lines
> +[[tree]]
> +trees::
> +    A tree is how Git represents a directory. It lists, for each item in
> +    the tree:
> ++
> +[[file-mode]]
> +1. The *file mode*, for example `100644`. The format is inspired by Unix
> +   permissions, but Git's modes are much more limited. Git only supports these file modes:
> ++
> +  - `100644`: regular file (with type `blob`)
> +  - `100755`: executable file (with type `blob`)
> +  - `120000`: symbolic link (with type `blob`)
> +  - `040000`: directory (with type `tree`)
> +  - `160000`: gitlink, for use with submodules (with type `commit`)

It is not really "supporting" file modes. Rather, Git only records 5 kinds of entities associated with each path in a tree object, and uses numbers taht remotely resemble POSIX file modes to represent these 5 kinds.

Perhaps "supports" -> "uses"?
Show 5 quoted lines
> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),
> +  or <<commit,`commit`>> (a Git submodule, which is a
> +  commit from a different Git repository)
> +3. The <<object-id,*object ID*>>
> +4. The *filename*

Here it may be worth noting that this "filename" is a single pathname component (roughly, what you would see in non-recursive "ls"). In other words, it may be a directory name.

I wonder if we need to say "<blob> (a file, or a symbolic link)"?
> +[[blob]]
> +blobs::
> +    A blob is how Git represents a file. A blob object contains the
> +    file's contents.

"represents a file" hints as if the thing may know its name, but that is not the case (its name is given only by surrounding tree).

"A blob is how Git represents uninterpreted series of bytes, and most commonly used to store file's contents." or something, perhaps?

> +When you make a new commit, Git only needs to store new versions of
> +files which were changed in that commit. This means that commits
> +can use relatively little disk space even in a very large repository.

That invites the "aren't we storing a delta after all, then?" confusion.

"Git only needs to newly store new versions of files and directories. Files and directories that were not modified by the commit are shared with its parent commit".

> +NOTE: All of the examples in this section were generated with
> +`git cat-file -p <object-id>`, which shows the contents of a Git object.

Was this necessary to say this? Blobs, Commits, and Tags are textual, so "-p" does very minimum thing, but Trees are binary garbage, so "-p" output is heavily massaged version of the contents.

> +[[branch]]
> +branches: `refs/heads/<name>`::
> +    A branch is a name for a commit ID.

Well a commit ID is an alternative way to refer to a commit object *name*, so it is a bit strange to say "a name for a commit ID".

Perhaps "A branch ref stores a commit ID." is better?
> +[[tag]]
> +tags: `refs/tags/<name>`::
> +    A tag is a name for a commit ID, tag object ID, or other object ID.

Likewise. "A tag ref stores any kind of object ID, but commonly they are commit objects or tag objects"

Show 9 quoted lines
> +    Tags that reference a tag object ID are called "annotated tags",
> +    because the tag object contains a tag message.
> +    Tags that reference a commit, blob, or tree ID are
> +    called "lightweight tags".
> ++
> +Even though branches and tags are both "a name for a commit ID", Git
> +treats them very differently.
> +Branches are expected to change over time: when you make a commit, Git
> +will update your <<HEAD,current branch>> to reference the new changes.

This sentence talks about branch moving because it advances with more commits. Did we want to say "HEAD" here before we explain what it is? "HEAD" can move for another reason (i.e. branch switching) and using "HEAD" in the context of talking about growing history might invite confusion. I dunno.

> +Tags are usually not changed after they're created.
> +[[HEAD]]
> +HEAD: `HEAD`::
> +    `HEAD` is where Git stores your current <<branch,branch>>.
Hmm...
Show 5 quoted lines
> +    `HEAD` can either be:
> +    1. A symbolic reference to your current branch, for example `ref:
> +       refs/heads/main` if your current branch is `main`.
> +    2. A direct reference to a commit ID. This is called "detached HEAD
> +	   state", see the DETACHED HEAD section of linkgit:git-checkout[1] for more.

These two are very reasonable. But "your current <<branch>>" refers only to #1.

    `HEAD` refers to the commit your current work is based on, and
    it is the commit that will become the first parent of the commit
    once your current work is concluded.  It can either be ...
perhaps.
> +[[remote-tracking-branch]]
> +remote tracking branches: `refs/remotes/<remote>/<branch>`::
Please always write "remote-tracking" with a hyphen (see glossary).
> +    A remote-tracking branch is a name for a commit ID.

Either "A remote-tracking branch stores a commit object name" or "A remote-tracking branch points at a commit object", followed by "in order to keep track of the last-nown state of ..." in a single sentence.

Show 7 quoted lines
> +[[index]]
> +THE INDEX
> +---------
> +
> +The index, also known as the "staging area", contains a list of every
> +file in the repository and its contents. When you commit, the files in
> +the index are used as the files in the next commit.

It is hard to define what "every file in the repository" really is. Files that you removed last week do not count. Files added in your wip branch elsewhere are obviously not yet in the index when you are working on your primary branch.

> +You can add files to the index or update the version in the index with
> +linkgit:git-add[1]. Adding a file to the index or updating its version
> +is called "staging" the file for commit.

It may be worth to clarify by saying "staging the contents of the file" (you can edit the file further after you "git add") that you are taking a snapshot at the time you ran "git add", instead of giving a general instruction to "keey an eye on this file" to Git (if it were, then the next "git commit" would behave more like "git add -u && git commit").

Show 18 quoted lines
> +[[reflogs]]
> +REFLOGS
> +-------
> +
> +Git stores a history called a "reflog" for every branch, remote-tracking
> +branch, and HEAD. This means that if you make a mistake and "lose" a
> +commit, you can generally recover the commit ID by running
> +`git reflog <reference>`.
> +
> +Each reflog entry has:
> +
> +1. Before/after *commit IDs*
> +2. *User* who made the change, for example `Maya <maya@example.com>`
> +3. *Timestamp* when the change was made
> +4. *Log message*, for example `pull: Fast-forward`
> +
> +Reflogs only log changes made in your local repository.
> +They are not shared with remotes.

Technically it is correct that before/after are recorded, but there is no way for the end-user to interact with them. "git reflog" walking these entries will only give you a single commit object. The username is also recorded, but I do not think of a way to view the information, let alone using it for querying.

Especially when the reftable backend is in use, you cannot even read the raw representation like you can do with files backend (where something like "cat .git/logs/HEAD" would let you peek into the details). I am not sure if we want to go into this detail.

Perhaps drop everything after "Each reflog entry has:"?
Show 40 quoted lines
> +For example, here's how the reflog for `HEAD` in a repository with 2
> +commits is stored:
> +
> +----
> +0000000000000000000000000000000000000000 4ccb6d7b8869a86aae2e84c56523f8705b50c647 Maya <maya@example.com> 1759173408 -0400      commit (initial): Initial commit
> +4ccb6d7b8869a86aae2e84c56523f8705b50c647 750b4ead9c87ceb3ddb7a390e6c7074521797fb3 Maya <maya@example.com> 1759173425 -0400      commit: Add README
> +----
> +
> +GIT
> +---
> +Part of the linkgit:git[1] suite
> diff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc
> index e423e4765b..20ba121314 100644
> --- a/Documentation/glossary-content.adoc
> +++ b/Documentation/glossary-content.adoc
> @@ -297,8 +297,8 @@ This commit is referred to as a "merge commit", or sometimes just a
>  	identified by its <<def_object_name,object name>>. The objects usually
>  	live in `$GIT_DIR/objects/`.
>  
> -[[def_object_identifier]]object identifier (oid)::
> -	Synonym for <<def_object_name,object name>>.
> +[[def_object_identifier]]object identifier, object ID, oid::
> +	Synonyms for <<def_object_name,object name>>.
>  
>  [[def_object_name]]object name::
>  	The unique identifier of an <<def_object,object>>.  The
> diff --git a/Documentation/meson.build b/Documentation/meson.build
> index e34965c5b0..ace0573e82 100644
> --- a/Documentation/meson.build
> +++ b/Documentation/meson.build
> @@ -192,6 +192,7 @@ manpages = {
>    'gitcore-tutorial.adoc' : 7,
>    'gitcredentials.adoc' : 7,
>    'gitcvs-migration.adoc' : 7,
> +  'gitdatamodel.adoc' : 7,
>    'gitdiffcore.adoc' : 7,
>    'giteveryday.adoc' : 7,
>    'gitfaq.adoc' : 7,
>
> base-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8
Previous: Julia EvansNext: Julia Evans
Message 35 of 89 in “doc: add a explanation of Git's data model”
  1. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 3, 2025
  2. Kristoffer HaugsbakkOct 3, 2025
  3. Julia EvansOct 6, 2025
  4. D. Ben KnobleOct 6, 2025
  5. Julia EvansOct 6, 2025
  6. D. Ben KnobleOct 6, 2025
  7. Julia EvansOct 9, 2025
  8. Kristoffer HaugsbakkOct 8, 2025
  9. Junio C HamanoOct 6, 2025
  10. Julia EvansOct 6, 2025
  11. Kristoffer HaugsbakkOct 7, 2025
  12. Junio C HamanoOct 7, 2025
  13. Patrick SteinhardtOct 7, 2025
  14. Junio C HamanoOct 7, 2025
  15. Julia EvansOct 7, 2025
  16. Junio C HamanoOct 7, 2025
  17. D. Ben KnobleOct 7, 2025
  18. Julia EvansOct 7, 2025
  19. Patrick SteinhardtOct 8, 2025
  20. Junio C HamanoOct 8, 2025
  21. Julia EvansOct 8, 2025
  22. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 8, 2025
  23. Patrick SteinhardtOct 10, 2025
  24. Junio C HamanoOct 13, 2025
  25. Patrick SteinhardtOct 14, 2025
  26. Julia EvansOct 14, 2025
  27. Patrick SteinhardtOct 14, 2025
  28. Junio C HamanoOct 14, 2025
  29. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 14, 2025
  30. Patrick SteinhardtOct 15, 2025
  31. Junio C HamanoOct 15, 2025
  32. Julia EvansOct 15, 2025
  33. Junio C HamanoOct 15, 2025
  34. Julia EvansOct 16, 2025
  35. Junio C HamanoOct 15, 2025
  36. Julia EvansOct 16, 2025
  37. Junio C HamanoOct 16, 2025
  38. Julia EvansOct 16, 2025
  39. Junio C HamanoOct 16, 2025
  40. Kristoffer HaugsbakkOct 16, 2025
  41. Kristoffer HaugsbakkOct 20, 2025
  42. Junio C HamanoOct 20, 2025
  43. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Oct 27, 2025
  44. Junio C HamanoOct 27, 2025
  45. Julia EvansOct 28, 2025
  46. Junio C HamanoOct 28, 2025
  47. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Oct 30, 2025
  48. Junio C HamanoOct 31, 2025
  49. Patrick SteinhardtNov 3, 2025
  50. Junio C HamanoNov 3, 2025
  51. Julia EvansNov 3, 2025
  52. Junio C HamanoNov 4, 2025
  53. Julia EvansNov 4, 2025
  54. Junio C HamanoNov 4, 2025
  55. Julia EvansNov 4, 2025
  56. Junio C HamanoNov 4, 2025
  57. Julia EvansNov 5, 2025
  58. Ben KnobleNov 5, 2025
  59. Julia EvansNov 5, 2025
  60. Ben KnobleNov 6, 2025
  61. Junio C HamanoOct 31, 2025
  62. Patrick SteinhardtNov 3, 2025
  63. Julia EvansNov 3, 2025
  64. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Nov 7, 2025
  65. Junio C HamanoNov 7, 2025
  66. Junio C HamanoNov 7, 2025
  67. Julia EvansNov 7, 2025
  68. Junio C HamanoNov 7, 2025
  69. Junio C HamanoNov 8, 2025
  70. Ben KnobleNov 9, 2025
  71. Junio C HamanoNov 9, 2025
  72. Julia EvansNov 10, 2025
  73. Junio C HamanoNov 11, 2025
  74. Ben KnobleNov 11, 2025
  75. Julia EvansNov 11, 2025
  76. Junio C HamanoNov 12, 2025
  77. Junio C HamanoNov 12, 2025
  78. Julia EvansNov 13, 2025
  79. Junio C HamanoNov 13, 2025
  80. Julia EvansNov 13, 2025
  81. Chris TorekNov 13, 2025
  82. Junio C HamanoNov 13, 2025
  83. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Nov 12, 2025
  84. Junio C HamanoNov 12, 2025
  85. Junio C HamanoNov 23, 2025
  86. Patrick SteinhardtDec 1, 2025
  87. Junio C HamanoDec 2, 2025
  88. Julia EvansOct 9, 2025
  89. Ben KnobleOct 10, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.