git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH] doc: add a explanation of Git's data model

From
Julia Evans via GitGitGadget <gitgitgadget@gmail.com>
Date
Oct 3, 2025, 17:34 UTC
Message-ID
<pull.1981.git.1759512876284.gitgitgadget@gmail.com>
From: Julia Evans <julia@jvns.ca>

Git very often uses the terms "object", "reference", or "index" in its documentation.

However, it's hard to find a clear explanation of these terms and how they relate to each other in the documentation. The closest candidates currently are:

1. `gitglossary`. This makes a good effort, but it's an alphabetically
    ordered dictionary and a dictionary is not a good way to learn
    concepts. You have to jump around too much and it's not possible to
    present the concepts in the order that they should be explained.
2. `gitcore-tutorial`. This explains how to use the "core" Git commands.
   This is a nice document to have, but it's not necessary to learn how
   `update-index` works to understand Git's data model, and we should
   not be requiring users to learn how to use the "plumbing" commands
   if they want to learn what the term "index" or "object" means.
3. `gitrepository-layout`. This is a great resource, but it includes a
   lot of information about configuration and internal implementation
   details which are not related to the data model. It also does
   not explain how commits work.

The result of this is that Git users (even users who have been using Git for 15+ years) struggle to read the documentation because they don't know what the core terms mean, and it's not possible to add links to help them learn more.

Add an explanation of Git's data model. Some choices I've made in deciding what "core data model" means:

1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me
   if those are intended to be user facing or if they're more like
   internal implementation details.
2. Don't talk about submodules other than by mentioning how they
   relate to trees. This is because Git has a lot of special features,
   and explaining how they all work exhaustively could quickly go
   down a rabbit hole which would make this document less useful for
   understanding Git's core behaviour.
3. Don't discuss the structure of a commit message
   (first line, trailers, GPG signatures, etc).
   Perhaps this should change.
Some other choices I've made:
1. Mention packed refs only in a note.
2. Don't mention that the full name of the branch `main` is
   technically `refs/heads/main`. This should likely change but I
   haven't worked out how to do it in a clear way yet.
3. Mostly avoid referring to the `.git` directory, because the exact
   details of how things are stored change over time.
   This should perhaps change from "mostly" to "entirely"
   but I haven't worked out how to do that in a clear way yet.
Signed-off-by: Julia Evans <julia@jvns.ca>
---
    doc: Add a explanation of Git's data model
Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v1
Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v1
Pull-Request: https://github.com/gitgitgadget/git/pull/1981
 Documentation/Makefile          |   1 +
 Documentation/gitdatamodel.adoc | 226 ++++++++++++++++++++++++++++++++
 2 files changed, 227 insertions(+)
 create mode 100644 Documentation/gitdatamodel.adoc
diff --git a/Documentation/Makefile b/Documentation/Makefile
index 6fb83d0c6e..5f4acfacbd 100644
--- a/Documentation/Makefile
+++ b/Documentation/Makefile
@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc
 MAN7_TXT += gitcore-tutorial.adoc
 MAN7_TXT += gitcredentials.adoc
 MAN7_TXT += gitcvs-migration.adoc
+MAN7_TXT += gitdatamodel.adoc
 MAN7_TXT += gitdiffcore.adoc
 MAN7_TXT += giteveryday.adoc
 MAN7_TXT += gitfaq.adoc
diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc
new file mode 100644
index 0000000000..4b2cb167dc
--- /dev/null
+++ b/Documentation/gitdatamodel.adoc
@@ -0,0 +1,226 @@
+gitdatamodel(7)
+===============
+
+NAME
+----
+gitdatamodel - Git's core data model
+
+DESCRIPTION
+-----------
+
+It's not necessary to understand Git's data model to use Git, but it's
+very helpful when reading Git's documentation so that you know what it
+means when the documentation says "object" "reference" or "index".
+
+Git's core operations use 4 kinds of data:
+
+1. <<objects,Objects>>: commits, trees, blobs, and tag objects
+2. <<references,References>>: branches, tags,
+   remote-tracking branches, etc
+3. <<index,The index>>, also known as the staging area
+4. <<reflogs,Reflogs>>
+
+[[objects]]
+OBJECTS
+-------
+
+Commits, trees, blobs, and tag objects are all stored in Git's object database.
+Every object has:
+
+1. an *ID*, which is the SHA-1 hash of its contents.
+  It's fast to look up a Git object using its ID.
+  The ID is usually represented in hexadecimal, like
+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.
+2. a *type*. There are 4 types of objects:
+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,
+   and <<tag-object,tag objects>>.
+3. *contents*. The structure of the contents depends on the type.
+
+Once an object is created, it can never be changed.
+Here are the 4 types of objects:
+
+[[commit]]
+commits::
+    A commit contains:
++
+1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,
+  regular commits have 1 parent, merge commits have 2+ parents
+2. A *commit message*
+3. All the *files* in the commit, stored as a *<<tree,tree>>*
+4. An *author* and the time the commit was authored
+5. A *committer* and the time the commit was committed
++
+Here's how an example commit is stored:
++
+----
+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a
+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647
+author Maya <maya@example.com> 1759173425 -0400
+committer Maya <maya@example.com> 1759173425 -0400
+
+Add README
+----
++
+Like all other objects, commits can never be changed after they're created.
+For example, "amending" a commit with `git commit --amend` creates a new commit.
+The old commit will eventually be deleted by `git gc`.
+
+[[tree]]
+trees::
+    A tree is how Git represents a directory. It lists, for each item in
+    the tree:
++
+1. The *permissions*, for example `100644`
+2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),
+  or <<commit,`commit`>> (a Git submodule)
+3. The *object ID*
+4. The *filename*
++
+For example, this is how a tree containing one directory (`src`) and one file
+(`README.md`) is stored:
++
+----
+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md
+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src
+----
++
+*NOTE:* The permissions are in the same format as UNIX permissions, but
+the only allowed permissions for files (blobs) are 644 and 755.
+
+[[blob]]
+blobs::
+    A blob is how Git represents a file. A blob object contains the
+    file's contents.
++
+Storing a new blob for every new version of a file can get big, so
+`git gc` periodically compresses objects for efficiency in `.git/objects/pack`.
+
+[[tag-object]]
+tag objects::
+    Tag objects (also known as "annotated tags") contain:
++
+1. The *tagger* and tag date
+2. A *tag message*, similar to a commit message
+3. The *ID* of the object (often a commit) that they reference
+
+[[references]]
+REFERENCES
+----------
+
+References are a way to give a name to a commit.
+It's easier to remember "the changes I'm working on are on the `turtle`
+branch" than "the changes are in commit bb69721404348e".
+Git often uses "ref" as shorthand for "reference".
+
+References that you create are stored in the `.git/refs` directory,
+and Git has a few special internal references like `HEAD` that are stored
+in the base `.git` directory.
+
+References can either be:
+
+1. References to an object ID, usually a <<commit,commit>> ID
+2. References to another reference. This is called a "symbolic reference".
+
+Git handles references differently based on which subdirectory of
+`.git/refs` they're stored in.
+Here are the main types:
+
+[[branch]]
+branches: `.git/refs/heads/<name>`::
+    A branch is a name for a commit ID.
+    That commit is the latest commit on the branch.
+    Branches are stored in the `.git/refs/heads/` directory.
++
+To get the history of commits on a branch, Git will start at the commit
+ID the branch references, and then look at the commit's parent(s),
+the parent's parent, etc.
+
+[[tag]]
+tags: `.git/refs/tags/<name>`::
+    A tag is a name for a commit ID, tag object ID, or other object ID.
+    Tags are stored in the `refs/tags/` directory.
++
+Even though branches and commits are both "a name for a commit ID", Git
+treats them very differently.
+Branches are expected to be regularly updated as you work on the branch,
+but it's expected that a tag will never change after you create it.
+
+[[HEAD]]
+HEAD: `.git/HEAD`::
+    `HEAD` is where Git stores your current <<branch,branch>>.
+    `HEAD` is normally a symbolic reference to your current branch, for
+    example `ref: refs/heads/main` if your current branch is `main`.
+    `HEAD` can also be a direct reference to a commit ID,
+    that's called "detached HEAD state".
+
+[[remote-tracking-branch]]
+remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::
+    A remote-tracking branch is a name for a commit ID.
+    It's how Git stores the last-known state of a branch in a remote
+    repository. `git fetch` updates remote-tracking branches. When
+    `git status` says "you're up to date with origin/main", it's looking at
+    this.
+
+[[other-refs]]
+Other references::
+    Git tools may create references in any subdirectory of `.git/refs`.
+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],
+    and linkgit:git-notes[1] all create their own references
+    in `.git/refs/stash`, `.git/refs/bisect`, etc.
+    Third-party Git tools may also create their own references.
++
+Git may also create references in the base `.git` directory
+other than `HEAD`, like `ORIG_HEAD`.
+
+*NOTE:* As an optimization, references may be stored as packed
+refs instead of in `.git/refs`. See linkgit:git-pack-refs[1].
+
+[[index]]
+THE INDEX
+---------
+
+The index, also known as the "staging area", contains the current staged
+version of every file in your Git repository. When you commit, the files
+in the index are used as the files in the next commit.
+
+Unlike a tree, the index is a flat list of files.
+Each index entry has 4 fields:
+
+1. The *permissions*
+2. The *<<blob,blob>> ID* of the file
+3. The *filename*
+4. The *number*. This is normally 0, but if there's a merge conflict
+   there can be multiple versions (with numbers 0, 1, 2, ..)
+   of the same filename in the index.
+
+It's extremely uncommon to look at the index directly: normally you'd
+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.
+But you can use `git ls-files --stage` to see the index.
+Here's the output of `git ls-files --stage` in a repository with 2 files:
+
+----
+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md
+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py
+----
+
+[[reflogs]]
+REFLOGS
+-------
+
+Git stores the history of branch, tag, and HEAD refs in a reflog
+(you should read "reflog" as "ref log"). Not every ref is logged by
+default, but any ref can be logged.
+
+Each reflog entry has:
+
+1. *Before/after *commit IDs*
+2. *User* who made the change, for example `Maya <maya@example.com>`
+3. *Timestamp*
+4. *Log message*, for example `pull: Fast-forward`
+
+Reflogs only log changes made in your local repository.
+They are not shared with remotes.
+
+GIT
+---
+Part of the linkgit:git[1] suite

base-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8
-- 
gitgitgadget
Next: Kristoffer Haugsbakk
Message 1 of 89 in “doc: add a explanation of Git's data model”
  1. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 3, 2025
  2. Kristoffer HaugsbakkOct 3, 2025
  3. Julia EvansOct 6, 2025
  4. D. Ben KnobleOct 6, 2025
  5. Julia EvansOct 6, 2025
  6. D. Ben KnobleOct 6, 2025
  7. Julia EvansOct 9, 2025
  8. Kristoffer HaugsbakkOct 8, 2025
  9. Junio C HamanoOct 6, 2025
  10. Julia EvansOct 6, 2025
  11. Kristoffer HaugsbakkOct 7, 2025
  12. Junio C HamanoOct 7, 2025
  13. Patrick SteinhardtOct 7, 2025
  14. Junio C HamanoOct 7, 2025
  15. Julia EvansOct 7, 2025
  16. Junio C HamanoOct 7, 2025
  17. D. Ben KnobleOct 7, 2025
  18. Julia EvansOct 7, 2025
  19. Patrick SteinhardtOct 8, 2025
  20. Junio C HamanoOct 8, 2025
  21. Julia EvansOct 8, 2025
  22. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 8, 2025
  23. Patrick SteinhardtOct 10, 2025
  24. Junio C HamanoOct 13, 2025
  25. Patrick SteinhardtOct 14, 2025
  26. Julia EvansOct 14, 2025
  27. Patrick SteinhardtOct 14, 2025
  28. Junio C HamanoOct 14, 2025
  29. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 14, 2025
  30. Patrick SteinhardtOct 15, 2025
  31. Junio C HamanoOct 15, 2025
  32. Julia EvansOct 15, 2025
  33. Junio C HamanoOct 15, 2025
  34. Julia EvansOct 16, 2025
  35. Junio C HamanoOct 15, 2025
  36. Julia EvansOct 16, 2025
  37. Junio C HamanoOct 16, 2025
  38. Julia EvansOct 16, 2025
  39. Junio C HamanoOct 16, 2025
  40. Kristoffer HaugsbakkOct 16, 2025
  41. Kristoffer HaugsbakkOct 20, 2025
  42. Junio C HamanoOct 20, 2025
  43. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Oct 27, 2025
  44. Junio C HamanoOct 27, 2025
  45. Julia EvansOct 28, 2025
  46. Junio C HamanoOct 28, 2025
  47. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Oct 30, 2025
  48. Junio C HamanoOct 31, 2025
  49. Patrick SteinhardtNov 3, 2025
  50. Junio C HamanoNov 3, 2025
  51. Julia EvansNov 3, 2025
  52. Junio C HamanoNov 4, 2025
  53. Julia EvansNov 4, 2025
  54. Junio C HamanoNov 4, 2025
  55. Julia EvansNov 4, 2025
  56. Junio C HamanoNov 4, 2025
  57. Julia EvansNov 5, 2025
  58. Ben KnobleNov 5, 2025
  59. Julia EvansNov 5, 2025
  60. Ben KnobleNov 6, 2025
  61. Junio C HamanoOct 31, 2025
  62. Patrick SteinhardtNov 3, 2025
  63. Julia EvansNov 3, 2025
  64. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Nov 7, 2025
  65. Junio C HamanoNov 7, 2025
  66. Junio C HamanoNov 7, 2025
  67. Julia EvansNov 7, 2025
  68. Junio C HamanoNov 7, 2025
  69. Junio C HamanoNov 8, 2025
  70. Ben KnobleNov 9, 2025
  71. Junio C HamanoNov 9, 2025
  72. Julia EvansNov 10, 2025
  73. Junio C HamanoNov 11, 2025
  74. Ben KnobleNov 11, 2025
  75. Julia EvansNov 11, 2025
  76. Junio C HamanoNov 12, 2025
  77. Junio C HamanoNov 12, 2025
  78. Julia EvansNov 13, 2025
  79. Junio C HamanoNov 13, 2025
  80. Julia EvansNov 13, 2025
  81. Chris TorekNov 13, 2025
  82. Junio C HamanoNov 13, 2025
  83. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Nov 12, 2025
  84. Junio C HamanoNov 12, 2025
  85. Junio C HamanoNov 23, 2025
  86. Patrick SteinhardtDec 1, 2025
  87. Junio C HamanoDec 2, 2025
  88. Julia EvansOct 9, 2025
  89. Ben KnobleOct 10, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.