{"thread":{"id":"64244","subject":"[PATCH] doc: add a explanation of Git's data model","startedAt":"2025-10-03T17:34:39Z","lastAt":"2025-12-02T12:25:43Z","messageCount":89,"participants":["Julia Evans via GitGitGadget","Kristoffer Haugsbakk","Junio C Hamano","Julia Evans","D. Ben Knoble","Patrick Steinhardt","Ben Knoble","Chris Torek"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"527894","messageId":"pull.1981.git.1759512876284.gitgitgadget@gmail.com","threadId":"64244","inReplyTo":null,"subject":"[PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-03T17:34:36Z","receivedAt":"2025-10-03T17:34:39Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"From: Julia Evans <julia@jvns.ca>\n\nGit very often uses the terms \"object\", \"reference\", or \"index\" in its\ndocumentation.\n\nHowever, it's hard to find a clear explanation of these terms and how\nthey relate to each other in the documentation. The closest candidates\ncurrently are:\n\n1. `gitglossary`. This makes a good effort, but it's an alphabetically\n    ordered dictionary and a dictionary is not a good way to learn\n    concepts. You have to jump around too much and it's not possible to\n    present the concepts in the order that they should be explained.\n2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n   This is a nice document to have, but it's not necessary to learn how\n   `update-index` works to understand Git's data model, and we should\n   not be requiring users to learn how to use the \"plumbing\" commands\n   if they want to learn what the term \"index\" or \"object\" means.\n3. `gitrepository-layout`. This is a great resource, but it includes a\n   lot of information about configuration and internal implementation\n   details which are not related to the data model. It also does\n   not explain how commits work.\n\nThe result of this is that Git users (even users who have been using\nGit for 15+ years) struggle to read the documentation because they don't\nknow what the core terms mean, and it's not possible to add links\nto help them learn more.\n\nAdd an explanation of Git's data model. Some choices I've made in\ndeciding what \"core data model\" means:\n\n1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n   if those are intended to be user facing or if they're more like\n   internal implementation details.\n2. Don't talk about submodules other than by mentioning how they\n   relate to trees. This is because Git has a lot of special features,\n   and explaining how they all work exhaustively could quickly go\n   down a rabbit hole which would make this document less useful for\n   understanding Git's core behaviour.\n3. Don't discuss the structure of a commit message\n   (first line, trailers, GPG signatures, etc).\n   Perhaps this should change.\n\nSome other choices I've made:\n\n1. Mention packed refs only in a note.\n2. Don't mention that the full name of the branch `main` is\n   technically `refs/heads/main`. This should likely change but I\n   haven't worked out how to do it in a clear way yet.\n3. Mostly avoid referring to the `.git` directory, because the exact\n   details of how things are stored change over time.\n   This should perhaps change from \"mostly\" to \"entirely\"\n   but I haven't worked out how to do that in a clear way yet.\n\nSigned-off-by: Julia Evans <julia@jvns.ca>\n---\n    doc: Add a explanation of Git's data model\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/1981\n\n Documentation/Makefile          |   1 +\n Documentation/gitdatamodel.adoc | 226 ++++++++++++++++++++++++++++++++\n 2 files changed, 227 insertions(+)\n create mode 100644 Documentation/gitdatamodel.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 6fb83d0c6e..5f4acfacbd 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n MAN7_TXT += gitcore-tutorial.adoc\n MAN7_TXT += gitcredentials.adoc\n MAN7_TXT += gitcvs-migration.adoc\n+MAN7_TXT += gitdatamodel.adoc\n MAN7_TXT += gitdiffcore.adoc\n MAN7_TXT += giteveryday.adoc\n MAN7_TXT += gitfaq.adoc\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nnew file mode 100644\nindex 0000000000..4b2cb167dc\n--- /dev/null\n+++ b/Documentation/gitdatamodel.adoc\n@@ -0,0 +1,226 @@\n+gitdatamodel(7)\n+===============\n+\n+NAME\n+----\n+gitdatamodel - Git's core data model\n+\n+DESCRIPTION\n+-----------\n+\n+It's not necessary to understand Git's data model to use Git, but it's\n+very helpful when reading Git's documentation so that you know what it\n+means when the documentation says \"object\" \"reference\" or \"index\".\n+\n+Git's core operations use 4 kinds of data:\n+\n+1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+2. <<references,References>>: branches, tags,\n+   remote-tracking branches, etc\n+3. <<index,The index>>, also known as the staging area\n+4. <<reflogs,Reflogs>>\n+\n+[[objects]]\n+OBJECTS\n+-------\n+\n+Commits, trees, blobs, and tag objects are all stored in Git's object database.\n+Every object has:\n+\n+1. an *ID*, which is the SHA-1 hash of its contents.\n+  It's fast to look up a Git object using its ID.\n+  The ID is usually represented in hexadecimal, like\n+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+2. a *type*. There are 4 types of objects:\n+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n+   and <<tag-object,tag objects>>.\n+3. *contents*. The structure of the contents depends on the type.\n+\n+Once an object is created, it can never be changed.\n+Here are the 4 types of objects:\n+\n+[[commit]]\n+commits::\n+    A commit contains:\n++\n+1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n+  regular commits have 1 parent, merge commits have 2+ parents\n+2. A *commit message*\n+3. All the *files* in the commit, stored as a *<<tree,tree>>*\n+4. An *author* and the time the commit was authored\n+5. A *committer* and the time the commit was committed\n++\n+Here's how an example commit is stored:\n++\n+----\n+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n+author Maya <maya@example.com> 1759173425 -0400\n+committer Maya <maya@example.com> 1759173425 -0400\n+\n+Add README\n+----\n++\n+Like all other objects, commits can never be changed after they're created.\n+For example, \"amending\" a commit with `git commit --amend` creates a new commit.\n+The old commit will eventually be deleted by `git gc`.\n+\n+[[tree]]\n+trees::\n+    A tree is how Git represents a directory. It lists, for each item in\n+    the tree:\n++\n+1. The *permissions*, for example `100644`\n+2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n+  or <<commit,`commit`>> (a Git submodule)\n+3. The *object ID*\n+4. The *filename*\n++\n+For example, this is how a tree containing one directory (`src`) and one file\n+(`README.md`) is stored:\n++\n+----\n+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n+----\n++\n+*NOTE:* The permissions are in the same format as UNIX permissions, but\n+the only allowed permissions for files (blobs) are 644 and 755.\n+\n+[[blob]]\n+blobs::\n+    A blob is how Git represents a file. A blob object contains the\n+    file's contents.\n++\n+Storing a new blob for every new version of a file can get big, so\n+`git gc` periodically compresses objects for efficiency in `.git/objects/pack`.\n+\n+[[tag-object]]\n+tag objects::\n+    Tag objects (also known as \"annotated tags\") contain:\n++\n+1. The *tagger* and tag date\n+2. A *tag message*, similar to a commit message\n+3. The *ID* of the object (often a commit) that they reference\n+\n+[[references]]\n+REFERENCES\n+----------\n+\n+References are a way to give a name to a commit.\n+It's easier to remember \"the changes I'm working on are on the `turtle`\n+branch\" than \"the changes are in commit bb69721404348e\".\n+Git often uses \"ref\" as shorthand for \"reference\".\n+\n+References that you create are stored in the `.git/refs` directory,\n+and Git has a few special internal references like `HEAD` that are stored\n+in the base `.git` directory.\n+\n+References can either be:\n+\n+1. References to an object ID, usually a <<commit,commit>> ID\n+2. References to another reference. This is called a \"symbolic reference\".\n+\n+Git handles references differently based on which subdirectory of\n+`.git/refs` they're stored in.\n+Here are the main types:\n+\n+[[branch]]\n+branches: `.git/refs/heads/<name>`::\n+    A branch is a name for a commit ID.\n+    That commit is the latest commit on the branch.\n+    Branches are stored in the `.git/refs/heads/` directory.\n++\n+To get the history of commits on a branch, Git will start at the commit\n+ID the branch references, and then look at the commit's parent(s),\n+the parent's parent, etc.\n+\n+[[tag]]\n+tags: `.git/refs/tags/<name>`::\n+    A tag is a name for a commit ID, tag object ID, or other object ID.\n+    Tags are stored in the `refs/tags/` directory.\n++\n+Even though branches and commits are both \"a name for a commit ID\", Git\n+treats them very differently.\n+Branches are expected to be regularly updated as you work on the branch,\n+but it's expected that a tag will never change after you create it.\n+\n+[[HEAD]]\n+HEAD: `.git/HEAD`::\n+    `HEAD` is where Git stores your current <<branch,branch>>.\n+    `HEAD` is normally a symbolic reference to your current branch, for\n+    example `ref: refs/heads/main` if your current branch is `main`.\n+    `HEAD` can also be a direct reference to a commit ID,\n+    that's called \"detached HEAD state\".\n+\n+[[remote-tracking-branch]]\n+remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::\n+    A remote-tracking branch is a name for a commit ID.\n+    It's how Git stores the last-known state of a branch in a remote\n+    repository. `git fetch` updates remote-tracking branches. When\n+    `git status` says \"you're up to date with origin/main\", it's looking at\n+    this.\n+\n+[[other-refs]]\n+Other references::\n+    Git tools may create references in any subdirectory of `.git/refs`.\n+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n+    and linkgit:git-notes[1] all create their own references\n+    in `.git/refs/stash`, `.git/refs/bisect`, etc.\n+    Third-party Git tools may also create their own references.\n++\n+Git may also create references in the base `.git` directory\n+other than `HEAD`, like `ORIG_HEAD`.\n+\n+*NOTE:* As an optimization, references may be stored as packed\n+refs instead of in `.git/refs`. See linkgit:git-pack-refs[1].\n+\n+[[index]]\n+THE INDEX\n+---------\n+\n+The index, also known as the \"staging area\", contains the current staged\n+version of every file in your Git repository. When you commit, the files\n+in the index are used as the files in the next commit.\n+\n+Unlike a tree, the index is a flat list of files.\n+Each index entry has 4 fields:\n+\n+1. The *permissions*\n+2. The *<<blob,blob>> ID* of the file\n+3. The *filename*\n+4. The *number*. This is normally 0, but if there's a merge conflict\n+   there can be multiple versions (with numbers 0, 1, 2, ..)\n+   of the same filename in the index.\n+\n+It's extremely uncommon to look at the index directly: normally you'd\n+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n+But you can use `git ls-files --stage` to see the index.\n+Here's the output of `git ls-files --stage` in a repository with 2 files:\n+\n+----\n+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n+----\n+\n+[[reflogs]]\n+REFLOGS\n+-------\n+\n+Git stores the history of branch, tag, and HEAD refs in a reflog\n+(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n+default, but any ref can be logged.\n+\n+Each reflog entry has:\n+\n+1. *Before/after *commit IDs*\n+2. *User* who made the change, for example `Maya <maya@example.com>`\n+3. *Timestamp*\n+4. *Log message*, for example `pull: Fast-forward`\n+\n+Reflogs only log changes made in your local repository.\n+They are not shared with remotes.\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\n\nbase-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n-- \ngitgitgadget\n"},{"id":"527911","messageId":"8df4c59c-4d27-4f36-a231-f7af32ddf149@app.fastmail.com","threadId":"64244","inReplyTo":"pull.1981.git.1759512876284.gitgitgadget@gmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-10-03T21:46:00Z","receivedAt":"2025-10-03T21:47:16Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"On Fri, Oct 3, 2025, at 19:34, Julia Evans via GitGitGadget wrote:\n> From: Julia Evans <julia@jvns.ca>\n>\n> Git very often uses the terms \"object\", \"reference\", or \"index\" in its\n> documentation.\n>\n> However, it's hard to find a clear explanation of these terms and how\n> they relate to each other in the documentation. The closest candidates\n> currently are:\n>\n> 1. `gitglossary`. This makes a good effort, but it's an alphabetically\n>     ordered dictionary and a dictionary is not a good way to learn\n>     concepts. You have to jump around too much and it's not possible to\n>     present the concepts in the order that they should be explained.\n> 2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n>    This is a nice document to have, but it's not necessary to learn how\n>    `update-index` works to understand Git's data model, and we should\n>    not be requiring users to learn how to use the \"plumbing\" commands\n>    if they want to learn what the term \"index\" or \"object\" means.\n> 3. `gitrepository-layout`. This is a great resource, but it includes a\n>    lot of information about configuration and internal implementation\n>    details which are not related to the data model. It also does\n>    not explain how commits work.\n>\n> The result of this is that Git users (even users who have been using\n> Git for 15+ years) struggle to read the documentation because they don't\n> know what the core terms mean, and it's not possible to add links\n> to help them learn more.\n>\n> Add an explanation of Git's data model. Some choices I've made in\n> deciding what \"core data model\" means:\n>\n> 1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n>    if those are intended to be user facing or if they're more like\n>    internal implementation details.\n> 2. Don't talk about submodules other than by mentioning how they\n>    relate to trees. This is because Git has a lot of special features,\n>    and explaining how they all work exhaustively could quickly go\n>    down a rabbit hole which would make this document less useful for\n>    understanding Git's core behaviour.\n> 3. Don't discuss the structure of a commit message\n>    (first line, trailers, GPG signatures, etc).\n>    Perhaps this should change.\n>\n> Some other choices I've made:\n>\n> 1. Mention packed refs only in a note.\n\nI don’t think it’s worth mentioning this at all.  More on that later.\n\n> 2. Don't mention that the full name of the branch `main` is\n>    technically `refs/heads/main`. This should likely change but I\n>    haven't worked out how to do it in a clear way yet.\n\nI think this is worth getting into.  This is a pretty\nuser-facing concept.\n\n> 3. Mostly avoid referring to the `.git` directory, because the exact\n>    details of how things are stored change over time.\n>    This should perhaps change from \"mostly\" to \"entirely\"\n>    but I haven't worked out how to do that in a clear way yet.\n\nI think that’s good.  I mean, I think us users don’t need that level of\ndetail and shouldn’t be “inspired” to muck with the internals.  If that\nmakes sense.  (See later)\n\n>\n> Signed-off-by: Julia Evans <julia@jvns.ca>\n> ---\n>     doc: Add a explanation of Git's data model\n>[snip]\n> diff --git a/Documentation/Makefile b/Documentation/Makefile\n>[snip]\n> diff --git a/Documentation/gitdatamodel.adoc\n> b/Documentation/gitdatamodel.adoc\n> new file mode 100644\n> index 0000000000..4b2cb167dc\n> --- /dev/null\n> +++ b/Documentation/gitdatamodel.adoc\n> @@ -0,0 +1,226 @@\n> +gitdatamodel(7)\n> +===============\n> +\n> +NAME\n> +----\n> +gitdatamodel - Git's core data model\n> +\n> +DESCRIPTION\n> +-----------\n> +\n> +It's not necessary to understand Git's data model to use Git, but it's\n> +very helpful when reading Git's documentation so that you know what it\n> +means when the documentation says \"object\" \"reference\" or \"index\".\n\nI haven’t gone hunting through the docs to see if this is covered\nelsewhere.  But the thrust of all the things here definitely feel to me\nlike something that should be presented and documented in such a way.\n\n> +\n> +Git's core operations use 4 kinds of data:\n\nMaybe small numerals should be spelled as words in running text?\n\n> +\n> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n> +2. <<references,References>>: branches, tags,\n> +   remote-tracking branches, etc\n> +3. <<index,The index>>, also known as the staging area\n> +4. <<reflogs,Reflogs>>\n\nReflogs is certainly auxiliary ref data. What makes it qualify as\none-of-the-four?  I am open to it being both, to be clear.\n\n> +\n> +[[objects]]\n> +OBJECTS\n> +-------\n> +\n> +Commits, trees, blobs, and tag objects are all stored in Git's object\n> database.\n> +Every object has:\n> +\n> +1. an *ID*, which is the SHA-1 hash of its contents.\n> +  It's fast to look up a Git object using its ID.\n> +  The ID is usually represented in hexadecimal, like\n> +  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n> +2. a *type*. There are 4 types of objects:\n> +   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n> +   and <<tag-object,tag objects>>.\n> +3. *contents*. The structure of the contents depends on the type.\n> +\n> +Once an object is created, it can never be changed.\n> +Here are the 4 types of objects:\n\nAs a curious Git user this seems correct.\n\n> +\n> +[[commit]]\n> +commits::\n> +    A commit contains:\n> ++\n> +1. Its *parent commit ID(s)*. The first commit in a repository has 0\n> parents,\n\nMaybe this is a subjective style thing but is it necessary to use “(s)”\nwhen the context makes clear that it could be zero to many?\n\n    Its *parent commit IDs. ...\n\n> +  regular commits have 1 parent, merge commits have 2+ parents\n\ns/2+/two or more/ ?\n\nSame point as the “numeral” one above.\n\n> +2. A *commit message*\n> +3. All the *files* in the commit, stored as a *<<tree,tree>>*\n> +4. An *author* and the time the commit was authored\n> +5. A *committer* and the time the commit was committed\n> ++\n> +Here's how an example commit is stored:\n> ++\n> +----\n> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n> +author Maya <maya@example.com> 1759173425 -0400\n> +committer Maya <maya@example.com> 1759173425 -0400\n> +\n> +Add README\n> +----\n> ++\n> +Like all other objects, commits can never be changed after they're\n> created.\n> +For example, \"amending\" a commit with `git commit --amend` creates a\n> new commit.\n\n> +The old commit will eventually be deleted by `git gc`.\n\nMaybe this could be moved to a part about what happens (eventually) to\nunreachable objects?\n\nMentioning `git gc` and how things will get deleted raises\nquestions naturally. Like why would they be deleted? Okay\nthat’s clear: the previous commit will be replaced by the\namended one. Then when it is not reachable by anything\n(even the reflog) it will get garbage collected.\n\nIt all follows. But is the reader necessarily mature enough\nin their understanding to make the inference?\n\nThis is a long-winded way of saying: if you’re gonna discuss\n`git gc` you might need to go into all of these concepts.\n\n> +\n> +[[tree]]\n> +trees::\n> +    A tree is how Git represents a directory. It lists, for each item\n> in\n> +    the tree:\n> ++\n> +1. The *permissions*, for example `100644`\n> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n> +  or <<commit,`commit`>> (a Git submodule)\n> +3. The *object ID*\n> +4. The *filename*\n> ++\n> +For example, this is how a tree containing one directory (`src`) and\n> one file\n> +(`README.md`) is stored:\n> ++\n> +----\n> +100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n> +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n> +----\n> ++\n> +*NOTE:* The permissions are in the same format as UNIX permissions, but\n> +the only allowed permissions for files (blobs) are 644 and 755.\n> +\n\nMakes sense.\n\n> +[[blob]]\n> +blobs::\n> +    A blob is how Git represents a file. A blob object contains the\n> +    file's contents.\n> ++\n> +Storing a new blob for every new version of a file can get big, so\n> +`git gc` periodically compresses objects for efficiency in\n> `.git/objects/pack`.\n\nThis gets into mentioning implementation files(?) like you mentioned in\nthe commit message.\n\n1. That it’s a packfile and where it is might be too much detail for\n   this doc\n2. I vaguely recall documents discussing what happens to “storing every\n   version” discussing deltas instead of packs? Again, I am not a Git\n   developer though.\n\n> +\n> +[[tag-object]]\n> +tag objects::\n> +    Tag objects (also known as \"annotated tags\") contain:\n> ++\n> +1. The *tagger* and tag date\n> +2. A *tag message*, similar to a commit message\n> +3. The *ID* of the object (often a commit) that they reference\n\ns/often/typically/ ?\n\nI know it can get tedious to caveat the 99% cases with things that are\ntechnically possible.  Maybe if it gets “bad enough” there could be a\npart that explains/distinguishes the high-level/porcelain Git use and\nwhat is technically possible: you make a `git tag -a`, which is on a\ncommit... except if you accidentally run it on top of an existing\ntag. Then even the porcelain won’t protect you from making a \ntag-on-tag. (But it will issue a warning I guess.) Hmm. Now I don’t know.\n\n> +\n> +[[references]]\n> +REFERENCES\n> +----------\n> +\n> +References are a way to give a name to a commit.\n> +It's easier to remember \"the changes I'm working on are on the `turtle`\n> +branch\" than \"the changes are in commit bb69721404348e\".\n> +Git often uses \"ref\" as shorthand for \"reference\".\n\nGood.\n\n> +\n> +References that you create are stored in the `.git/refs` directory,\n> +and Git has a few special internal references like `HEAD` that are\n> stored\n> +in the base `.git` directory.\n\nImplementation file details.\n\nYou also mention `.git/refs/heads/<name>` below.  But refs aren’t stored\nas files if you are using the *reftable* backend.  And that backend will\nbecome the default for new repositories in Git 3.0, I think.\n\nHow does reftable work?  I don’t know.  But I don’t think we need to\nknow after reading this doc. :)\n\nTo be clear: how files are stored might not matter here.\n\n> +\n> +References can either be:\n> +\n> +1. References to an object ID, usually a <<commit,commit>> ID\n> +2. References to another reference. This is called a \"symbolic\n> reference\".\n\nYou seem to have used `**` when introducing terms:\n\n    This is a *symbolic reference*\n\n>[snip ref stuff]\n> +\n> +[[HEAD]]\n> +HEAD: `.git/HEAD`::\n> +    `HEAD` is where Git stores your current <<branch,branch>>.\n> +    `HEAD` is normally a symbolic reference to your current branch, for\n> +    example `ref: refs/heads/main` if your current branch is `main`.\n> +    `HEAD` can also be a direct reference to a commit ID,\n> +    that's called \"detached HEAD state\".\n> +\n> +[[remote-tracking-branch]]\n> +remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::\n> +    A remote-tracking branch is a name for a commit ID.\n> +    It's how Git stores the last-known state of a branch in a remote\n> +    repository. `git fetch` updates remote-tracking branches. When\n> +    `git status` says \"you're up to date with origin/main\", it's looking at\n> +    this.\n\nLooks good.\n\n> +\n> +[[other-refs]]\n> +Other references::\n> +    Git tools may create references in any subdirectory of `.git/refs`.\n> +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n> +    and linkgit:git-notes[1] all create their own references\n> +    in `.git/refs/stash`, `.git/refs/bisect`, etc.\n> +    Third-party Git tools may also create their own references.\n> ++\n> +Git may also create references in the base `.git` directory\n> +other than `HEAD`, like `ORIG_HEAD`.\n> +\n\n> +*NOTE:* As an optimization, references may be stored as packed\n> +refs instead of in `.git/refs`. See linkgit:git-pack-refs[1].\n\nI don’t know if this is relevant for both ref backends. And does it\nmatter?\n\n> +\n> +[[index]]\n> +THE INDEX\n> +---------\n> +\n> +The index, also known as the \"staging area\", contains the current\n> staged\n> +version of every file in your Git repository. When you commit, the\n> files\n> +in the index are used as the files in the next commit.\n> +\n> +Unlike a tree, the index is a flat list of files.\n> +Each index entry has 4 fields:\n> +\n> +1. The *permissions*\n> +2. The *<<blob,blob>> ID* of the file\n> +3. The *filename*\n> +4. The *number*. This is normally 0, but if there's a merge conflict\n> +   there can be multiple versions (with numbers 0, 1, 2, ..)\n> +   of the same filename in the index.\n> +\n> +It's extremely uncommon to look at the index directly: normally you'd\n> +run `git status` to see a list of changes between the index and\n> <<HEAD,HEAD>>.\n> +But you can use `git ls-files --stage` to see the index.\n> +Here's the output of `git ls-files --stage` in a repository with 2\n> files:\n> +\n> +----\n> +100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n> +100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n> +----\n> +\n> +[[reflogs]]\n> +REFLOGS\n> +-------\n> +\n> +Git stores the history of branch, tag, and HEAD refs in a reflog\n> +(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n\nYou’ve heard of the re-flog too?\n\n> +default, but any ref can be logged.\n> +\n> +Each reflog entry has:\n> +\n> +1. *Before/after *commit IDs*\n> +2. *User* who made the change, for example `Maya <maya@example.com>`\n> +3. *Timestamp*\n> +4. *Log message*, for example `pull: Fast-forward`\n> +\n> +Reflogs only log changes made in your local repository.\n> +They are not shared with remotes.\n\nMakes sense.\n\n> +\n> +GIT\n> +---\n> +Part of the linkgit:git[1] suite\n\nI appreciate that this is the first version and you might have plans\nafter this one. But I wonder if this doc could use a fair number of\n`gitlink` to branch out to all the other parts. Like git-reflog(1),\ngitglossary(7).\n\nThanks for starting on a whole new doc. That must take quite\nsome effort.\n"},{"id":"527958","messageId":"xmqqy0por9g7.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.git.1759512876284.gitgitgadget@gmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-06T03:32:56Z","receivedAt":"2025-10-06T03:32:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +MAN7_TXT += gitdatamodel.adoc\n>  MAN7_TXT += gitdiffcore.adoc\n> ...\n> +gitdatamodel(7)\n> +===============\n> +\n> +NAME\n> +----\n> +gitdatamodel - Git's core data model\n> +\n> +DESCRIPTION\n> +-----------\n\nThe above causes doc-lint to barf.\n\nhttps://github.com/git/git/actions/runs/18265502271/job/51999236907#step:4:655\n\ngitdatamodel.adoc:226: has no required 'SYNOPSIS' section!\n    LINT MAN SEC giteveryday.adoc\nmake[1]: *** [Makefile:498: .build/lint-docs/man-section-order/gitdatamodel.ok] Error 1\n\n\nYou can check locally with \"make check-docs\" without waiting for my\nintegration cycle to push to GitHub CI.\n\nThanks.\n"},{"id":"528021","messageId":"9d02c334-b5dc-4faa-8bd4-75c344bf5576@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqy0por9g7.fsf@gitster.g","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-06T19:03:26Z","receivedAt":"2025-10-06T19:03:47Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"> The above causes doc-lint to barf.\n>\n> https://github.com/git/git/actions/runs/18265502271/job/51999236907#step:4:655\n>\n> gitdatamodel.adoc:226: has no required 'SYNOPSIS' section!\n>     LINT MAN SEC giteveryday.adoc\n> make[1]: *** [Makefile:498: \n> .build/lint-docs/man-section-order/gitdatamodel.ok] Error 1\n>\n>\n> You can check locally with \"make check-docs\" without waiting for my\n> integration cycle to push to GitHub CI.\n\n\nThanks, will fix.\n"},{"id":"528034","messageId":"51e0a55c-1f1d-4cae-9459-8c2b9220e52d@app.fastmail.com","threadId":"64244","inReplyTo":"8df4c59c-4d27-4f36-a231-f7af32ddf149@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-06T19:36:54Z","receivedAt":"2025-10-06T19:37:19Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"Thanks for the review!\n\n>> 2. Don't mention that the full name of the branch `main` is\n>>    technically `refs/heads/main`. This should likely change but I\n>>    haven't worked out how to do it in a clear way yet.\n>\n> I think this is worth getting into.  This is a pretty\n> user-facing concept.\n\nI think I'll see if I can figure out a way to mention this and at the\nsame time remove most of the rest of the references to the `.git`\ndirectory when explaining references (which you talked about\nfurther down), including packed refs.\n\n>> +\n>> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n>> +2. <<references,References>>: branches, tags,\n>> +   remote-tracking branches, etc\n>> +3. <<index,The index>>, also known as the staging area\n>> +4. <<reflogs,Reflogs>>\n>\n> Reflogs is certainly auxiliary ref data. What makes it qualify as\n> one-of-the-four?  I am open to it being both, to be clear.\n\nThe reason I like to talk about reflogs is that it gives you a\nway to \"undo\" Git operations that can be really useful. \nAnd any Git command that updates refs can updates that\nref's reflog.\n\nUnderstanding how reflogs work helps to understand what the\nlimitations of using reflogs to undo mistakes is: for example\nthe index is not a ref, so you can't use the reflog to undo\nchanges to the index.\n\n>> +2. A *commit message*\n>> +3. All the *files* in the commit, stored as a *<<tree,tree>>*\n>> +4. An *author* and the time the commit was authored\n>> +5. A *committer* and the time the commit was committed\n>> ++\n>> +Here's how an example commit is stored:\n>> ++\n>> +----\n>> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n>> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n>> +author Maya <maya@example.com> 1759173425 -0400\n>> +committer Maya <maya@example.com> 1759173425 -0400\n>> +\n>> +Add README\n>> +----\n>> ++\n>> +Like all other objects, commits can never be changed after they're\n>> created.\n>> +For example, \"amending\" a commit with `git commit --amend` creates a\n>> new commit.\n>\n>> +The old commit will eventually be deleted by `git gc`.\n>\n> Maybe this could be moved to a part about what happens (eventually) to\n> unreachable objects?\n>\n> Mentioning `git gc` and how things will get deleted raises\n> questions naturally. Like why would they be deleted? Okay\n> that’s clear: the previous commit will be replaced by the\n> amended one. Then when it is not reachable by anything\n> (even the reflog) it will get garbage collected.\n>\n> It all follows. But is the reader necessarily mature enough\n> in their understanding to make the inference?\n>\n> This is a long-winded way of saying: if you’re gonna discuss\n> `git gc` you might need to go into all of these concepts.\n\nIf folks here think this is a reasonable document to add to\nGit I'll try get some beta readers to read this, see which parts\nfolks find confusing, and address those, keeping the `git gc`\nstuff in mind.\n\nSimilarly for the style comments.\n\n>> +blobs::\n>> +    A blob is how Git represents a file. A blob object contains the\n>> +    file's contents.\n>> ++\n>> +Storing a new blob for every new version of a file can get big, so\n>> +`git gc` periodically compresses objects for efficiency in\n>> `.git/objects/pack`.\n>\n> This gets into mentioning implementation files(?) like you mentioned in\n> the commit message.\n\nThat's true! The reason I think this is important to mention is that I find\nthat people often \"reject\" information that they find implausible, even\nif it comes from a credible source. (\"that can't be true! I must be\nnot understanding correctly. Oh well, I'll just ignore that!\")\n\nI sometimes hear from users that \"commits can't be snapshots\", because\nit would take up too much disk space to store every version of\nevery commit. So I find that sometimes explaining a little bit about the\nimplementation can make the information more memorable.\n\nCertainly I'm not able to remember details that don't make sense\nwith my mental model of how computers work and I don't expect other\npeople to either, so I think it's important to give an explanation that\nhandles the biggest \"objections\".\n\n> 1. That it’s a packfile and where it is might be too much detail for\n>    this doc\n> 2. I vaguely recall documents discussing what happens to “storing every\n>    version” discussing deltas instead of packs? Again, I am not a Git\n>    developer though.\n\nI could be wrong about the details here, I'm not a Git developer either.\nFrom https://git-scm.com/book/en/v2/Git-Internals-Packfiles\nit looks like packfiles are implemented using deltas.\n\n>> +\n>> +References can either be:\n>> +\n>> +1. References to an object ID, usually a <<commit,commit>> ID\n>> +2. References to another reference. This is called a \"symbolic\n>> reference\".\n>\n> You seem to have used `**` when introducing terms:\n>\n>     This is a *symbolic reference*\n\nThanks, will take a look at that.\n\n>> +[[reflogs]]\n>> +REFLOGS\n>> +-------\n>> +\n>> +Git stores the history of branch, tag, and HEAD refs in a reflog\n>> +(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n>\n> You’ve heard of the re-flog too?\n\nhaha exactly, I just want folks to understand why it's called that :)\n\n> I appreciate that this is the first version and you might have plans\n> after this one. But I wonder if this doc could use a fair number of\n> `gitlink` to branch out to all the other parts. Like git-reflog(1),\n> gitglossary(7).\n\nThat's reasonable. Do you often use the \"See also\" section of\nman pages? I've never looked at them so I'm always curious about\nhow people are actually using them in practice.\n\nI also need to think about what else could link *to* this, because\nwithout attention to discoverability probably nobody will find it.\nMy main idea so far is actually to add it to\nhttps://git-scm.com/learn\nbut I wanted to send it here instead of adding it to the website\ndirectly because I thought it could benefit from a more detailed\nreview.\n\n> Thanks for starting on a whole new doc. That must take quite\n> some effort.\n\nAll the work on documentation takes a lot of effort, in some\nways it's easier to write something new than to edit something\nexisting :)\n"},{"id":"528042","messageId":"CALnO6CA29HA_FOQAJp_bkskKF-6Vy0_SKVL_OyJASByvKEZTqQ@mail.gmail.com","threadId":"64244","inReplyTo":"51e0a55c-1f1d-4cae-9459-8c2b9220e52d@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-10-06T21:44:47Z","receivedAt":"2025-10-06T21:45:01Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Mon, Oct 6, 2025 at 3:37 PM Julia Evans <julia@jvns.ca> wrote:\n>\n> Thanks for the review!\n>\n> >> 2. Don't mention that the full name of the branch `main` is\n> >>    technically `refs/heads/main`. This should likely change but I\n> >>    haven't worked out how to do it in a clear way yet.\n> >\n> > I think this is worth getting into.  This is a pretty\n> > user-facing concept.\n>\n> I think I'll see if I can figure out a way to mention this and at the\n> same time remove most of the rest of the references to the `.git`\n> directory when explaining references (which you talked about\n> further down), including packed refs.\n\nA colleague will be explaining reflog for an audience tomorrow, and\ndecided to briefly explain refs, too—which tells me this is\nmuch-needed.\n\nFor refs themselves, perhaps \"git for-each-ref\" is a reasonable place\nto start? Since it tells you the refs you have and how to spell them\nexplicitly regardless of how they are stored?\n\n-- \nD. Ben Knoble\n"},{"id":"528043","messageId":"1241cb86-9adf-4c52-87fb-028406ccd8f0@app.fastmail.com","threadId":"64244","inReplyTo":"CALnO6CA29HA_FOQAJp_bkskKF-6Vy0_SKVL_OyJASByvKEZTqQ@mail.gmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-06T21:46:46Z","receivedAt":"2025-10-06T21:47:15Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Mon, Oct 6, 2025, at 5:44 PM, D. Ben Knoble wrote:\n> On Mon, Oct 6, 2025 at 3:37 PM Julia Evans <julia@jvns.ca> wrote:\n>>\n>> Thanks for the review!\n>>\n>> >> 2. Don't mention that the full name of the branch `main` is\n>> >>    technically `refs/heads/main`. This should likely change but I\n>> >>    haven't worked out how to do it in a clear way yet.\n>> >\n>> > I think this is worth getting into.  This is a pretty\n>> > user-facing concept.\n>>\n>> I think I'll see if I can figure out a way to mention this and at the\n>> same time remove most of the rest of the references to the `.git`\n>> directory when explaining references (which you talked about\n>> further down), including packed refs.\n>\n> A colleague will be explaining reflog for an audience tomorrow, and\n> decided to briefly explain refs, too—which tells me this is\n> much-needed.\n>\n> For refs themselves, perhaps \"git for-each-ref\" is a reasonable place\n> to start? Since it tells you the refs you have and how to spell them\n> explicitly regardless of how they are stored?\n\nInteresting, do you use git for-each-ref? \nWhat do you use it for?\n\n> -- \n> D. Ben Knoble\n"},{"id":"528045","messageId":"CALnO6CCsGtjcWBkjV0vsJHDCwiwt9eO2CsA1zFgwFiwJ-KLhew@mail.gmail.com","threadId":"64244","inReplyTo":"1241cb86-9adf-4c52-87fb-028406ccd8f0@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-10-06T21:55:20Z","receivedAt":"2025-10-06T21:55:33Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Mon, Oct 6, 2025 at 5:47 PM Julia Evans <julia@jvns.ca> wrote:\n>\n>\n>\n> On Mon, Oct 6, 2025, at 5:44 PM, D. Ben Knoble wrote:\n> > On Mon, Oct 6, 2025 at 3:37 PM Julia Evans <julia@jvns.ca> wrote:\n> >>\n> >> Thanks for the review!\n> >>\n> >> >> 2. Don't mention that the full name of the branch `main` is\n> >> >>    technically `refs/heads/main`. This should likely change but I\n> >> >>    haven't worked out how to do it in a clear way yet.\n> >> >\n> >> > I think this is worth getting into.  This is a pretty\n> >> > user-facing concept.\n> >>\n> >> I think I'll see if I can figure out a way to mention this and at the\n> >> same time remove most of the rest of the references to the `.git`\n> >> directory when explaining references (which you talked about\n> >> further down), including packed refs.\n> >\n> > A colleague will be explaining reflog for an audience tomorrow, and\n> > decided to briefly explain refs, too—which tells me this is\n> > much-needed.\n> >\n> > For refs themselves, perhaps \"git for-each-ref\" is a reasonable place\n> > to start? Since it tells you the refs you have and how to spell them\n> > explicitly regardless of how they are stored?\n>\n> Interesting, do you use git for-each-ref?\n> What do you use it for?\n\nAh, yes, but primarily for scripting.\n\nWhat I should have clarified is that \"the tool (I know of) to\ninterrogate the refs you currently have is git-for-each-ref\" (like how\ngit-ls-remote is the tool to interrogate a remote's refs). It avoids\nthe issues with assuming \"tree .git/refs\" or similar will capture the\nactual data.\n\n-- \nD. Ben Knoble\n"},{"id":"528104","messageId":"93b30d1e-7d49-44cf-b29b-69e8055bccbc@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqy0por9g7.fsf@gitster.g","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-10-07T12:37:26Z","receivedAt":"2025-10-07T12:37:47Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"On Mon, Oct 6, 2025, at 05:32, Junio C Hamano wrote:\n> \"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> +MAN7_TXT += gitdatamodel.adoc\n>>  MAN7_TXT += gitdiffcore.adoc\n>> ...\n>> +gitdatamodel(7)\n>> +===============\n>> +\n>> +NAME\n>> +----\n>> +gitdatamodel - Git's core data model\n>> +\n>> +DESCRIPTION\n>> +-----------\n>\n> The above causes doc-lint to barf.\n>[snip]\n> You can check locally with \"make check-docs\" without waiting for my\n> integration cycle to push to GitHub CI.\n\nI think you meant `make lint-docs` for both of these.\n"},{"id":"528123","messageId":"aOUkZa4_fq1hho7Q@pks.im","threadId":"64244","inReplyTo":"pull.1981.git.1759512876284.gitgitgadget@gmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-07T14:32:05Z","receivedAt":"2025-10-07T14:32:17Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Oct 03, 2025 at 05:34:36PM +0000, Julia Evans via GitGitGadget wrote:\n> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n> new file mode 100644\n> index 0000000000..4b2cb167dc\n> --- /dev/null\n> +++ b/Documentation/gitdatamodel.adoc\n> @@ -0,0 +1,226 @@\n> +gitdatamodel(7)\n> +===============\n> +\n> +NAME\n> +----\n> +gitdatamodel - Git's core data model\n> +\n> +DESCRIPTION\n> +-----------\n> +\n> +It's not necessary to understand Git's data model to use Git, but it's\n> +very helpful when reading Git's documentation so that you know what it\n> +means when the documentation says \"object\" \"reference\" or \"index\".\n\nThere's a missing comma after \"object\".\n\n> +\n> +Git's core operations use 4 kinds of data:\n> +\n> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n> +2. <<references,References>>: branches, tags,\n> +   remote-tracking branches, etc\n> +3. <<index,The index>>, also known as the staging area\n> +4. <<reflogs,Reflogs>>\n\nThis list makes sense to me. There's of course more data structures in\nGit, but all the other data structures shouldn't really matter to users\nat all as they are mostly caches or internal details of the on-disk\nformat.\n\nThere's potentially one exception though, namely the Git configuration.\nI'd claim that Git \"uses\" the Git configuration similarly to how it uses\nthe others, but I get why it's not explicitly mentioned here.\n\n> +[[objects]]\n> +OBJECTS\n> +-------\n> +\n> +Commits, trees, blobs, and tag objects are all stored in Git's object database.\n> +Every object has:\n> +\n> +1. an *ID*, which is the SHA-1 hash of its contents.\n\nI think this needs to be adapted to not single out SHA-1 as the only\nhashing algorithm. We already support SHA-256, so we should definitely\nsay that the algorithm can be swapped. Maybe something like:\n\n  An *object ID*, which is the cryptographic hash of its contents. By\n  default, Git uses SHA-1 as object hash, but alternative hashes like\n  SHA-256 are supported.\n\n> +  It's fast to look up a Git object using its ID.\n> +  The ID is usually represented in hexadecimal, like\n> +  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n> +2. a *type*. There are 4 types of objects:\n> +   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n> +   and <<tag-object,tag objects>>.\n> +3. *contents*. The structure of the contents depends on the type.\n\nNit: every object also has an object size. Not sure though whether it's\nfine to imply that with \"contents\".\n\n> +Once an object is created, it can never be changed.\n> +Here are the 4 types of objects:\n> +\n> +[[commit]]\n> +commits::\n> +    A commit contains:\n> ++\n> +1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n> +  regular commits have 1 parent, merge commits have 2+ parents\n\nI'd say \"at least two parents\" instead of \"2+ parents\".\n\n> +2. A *commit message*\n> +3. All the *files* in the commit, stored as a *<<tree,tree>>*\n> +4. An *author* and the time the commit was authored\n> +5. A *committer* and the time the commit was committed\n> ++\n> +Here's how an example commit is stored:\n> ++\n> +----\n> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n> +author Maya <maya@example.com> 1759173425 -0400\n> +committer Maya <maya@example.com> 1759173425 -0400\n> +\n> +Add README\n> +----\n\nIn practice, commits can have other headers that are ignored by Git. But\nthat's certainly not part of Git's core data model, so I don't think we\nshould mention that here.\n\n> +Like all other objects, commits can never be changed after they're created.\n> +For example, \"amending\" a commit with `git commit --amend` creates a new commit.\n> +The old commit will eventually be deleted by `git gc`.\n\nIf we mention git-gc(1) I think it would make sense to use\n`linkgit:git-gc[1]` instead to provide a link to its man page.\n\n> +[[tree]]\n> +trees::\n> +    A tree is how Git represents a directory. It lists, for each item in\n> +    the tree:\n> ++\n> +1. The *permissions*, for example `100644`\n\nI think we should rather call these \"mode bits\". These bits are\npermissions indeed when you have a blob, but for subtrees, symlinks and\nsubmodules they aren't.\n\n> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n> +  or <<commit,`commit`>> (a Git submodule)\n\nThere's also symlinks.\n\n> +3. The *object ID*\n> +4. The *filename*\n> ++\n> +For example, this is how a tree containing one directory (`src`) and one file\n> +(`README.md`) is stored:\n> ++\n> +----\n> +100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n> +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n> +----\n> ++\n> +*NOTE:* The permissions are in the same format as UNIX permissions, but\n> +the only allowed permissions for files (blobs) are 644 and 755.\n> +\n> +[[blob]]\n> +blobs::\n> +    A blob is how Git represents a file. A blob object contains the\n> +    file's contents.\n> ++\n> +Storing a new blob for every new version of a file can get big, so\n> +`git gc` periodically compresses objects for efficiency in `.git/objects/pack`.\n\nI would claim that it's not necessary to mention object compression.\nThis should be a low-level detail that users don't ever have to worry\nabout. Furthermore, packing objects isn't only relevant in the context\nof blobs: trees for example also tend to compress very well as there\ntypically is only small incremental updates to trees.\n\n> +[[tag-object]]\n> +tag objects::\n> +    Tag objects (also known as \"annotated tags\") contain:\n> ++\n> +1. The *tagger* and tag date\n> +2. A *tag message*, similar to a commit message\n> +3. The *ID* of the object (often a commit) that they reference\n\nThey can also be signed, if we want to mention that.\n\n> +[[references]]\n> +REFERENCES\n> +----------\n> +\n> +References are a way to give a name to a commit.\n> +It's easier to remember \"the changes I'm working on are on the `turtle`\n> +branch\" than \"the changes are in commit bb69721404348e\".\n> +Git often uses \"ref\" as shorthand for \"reference\".\n> +\n> +References that you create are stored in the `.git/refs` directory,\n> +and Git has a few special internal references like `HEAD` that are stored\n> +in the base `.git` directory.\n\nThis isn't true anymore with the introduction of the reftable backend,\nwhich is slated to become the default backend. I'd argue that this is\nanother implementation detail that the user shouldn't have to worry\nabout.\n\n> +References can either be:\n> +\n> +1. References to an object ID, usually a <<commit,commit>> ID\n> +2. References to another reference. This is called a \"symbolic reference\".\n> +\n> +Git handles references differently based on which subdirectory of\n> +`.git/refs` they're stored in.\n\nSo instead of saying \"subdirectory\", I'd rather say \"reference\nhierarchy\".\n\nIn general, I think we should explain that references are layed out\nin a hierarchy. This is somewhat obvious with the \"files\" backend, as we\nuse directories there. But as we move on to the \"reftable\" backend this\nmay become less obvious over time.\n\n> +Here are the main types:\n> +\n> +[[branch]]\n> +branches: `.git/refs/heads/<name>`::\n\nHere and in the other cases we should then strip the `.git/` prefix.\n\n> +    A branch is a name for a commit ID.\n> +    That commit is the latest commit on the branch.\n> +    Branches are stored in the `.git/refs/heads/` directory.\n> ++\n> +To get the history of commits on a branch, Git will start at the commit\n> +ID the branch references, and then look at the commit's parent(s),\n> +the parent's parent, etc.\n> +\n> +[[tag]]\n> +tags: `.git/refs/tags/<name>`::\n> +    A tag is a name for a commit ID, tag object ID, or other object ID.\n> +    Tags are stored in the `refs/tags/` directory.\n> ++\n> +Even though branches and commits are both \"a name for a commit ID\", Git\n> +treats them very differently.\n> +Branches are expected to be regularly updated as you work on the branch,\n> +but it's expected that a tag will never change after you create it.\n\nThis sounds a bit like the user itself needs to update the branch. How\nabout this instead:\n\n    Even though branches and commits are both \"a name for a commit ID\", Git\n    treats them very differently:\n\n        - Branches can be checked out directly. If so, creating a new\n          commit will automatically update the checked-out branch to\n          point to the new commit.\n\n        - Tags cannot be checked out directly and don't move when\n          creating a new commit. Instead, one can only check out the\n          commit that a branch points to. This is called \"detached\n          HEAD\", and the effect is that a new commit will not update \n\n> +[[HEAD]]\n> +HEAD: `.git/HEAD`::\n> +    `HEAD` is where Git stores your current <<branch,branch>>.\n> +    `HEAD` is normally a symbolic reference to your current branch, for\n> +    example `ref: refs/heads/main` if your current branch is `main`.\n> +    `HEAD` can also be a direct reference to a commit ID,\n> +    that's called \"detached HEAD state\".\n> +\n> +[[remote-tracking-branch]]\n> +remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::\n> +    A remote-tracking branch is a name for a commit ID.\n> +    It's how Git stores the last-known state of a branch in a remote\n> +    repository. `git fetch` updates remote-tracking branches. When\n> +    `git status` says \"you're up to date with origin/main\", it's looking at\n> +    this.\n\nThis misses \"refs/remotes/<remote>/HEAD\". This reference is a symbolic\nreference that indicates the default branch on the remote side.\n\n> +[[other-refs]]\n> +Other references::\n> +    Git tools may create references in any subdirectory of `.git/refs`.\n> +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n> +    and linkgit:git-notes[1] all create their own references\n> +    in `.git/refs/stash`, `.git/refs/bisect`, etc.\n> +    Third-party Git tools may also create their own references.\n> ++\n> +Git may also create references in the base `.git` directory\n> +other than `HEAD`, like `ORIG_HEAD`.\n\nLet's mention that such references are typically spelt all-uppercase\nwith underscores between. You shouldn't ever create a reference that is\nfor example called \".git/foo\".\n\nWe enforce this restriction inconsistently, only, but I don't think that\nshould keep us from spelling out the common rule.\n\n> +*NOTE:* As an optimization, references may be stored as packed\n> +refs instead of in `.git/refs`. See linkgit:git-pack-refs[1].\n\nI'd drop this note. It's an internal implementation detail and only true\nfor the \"files\" backend. The \"reftable\" backend stores references quite\ndifferently and doesn't really \"pack\" references.\n\n> +[[index]]\n> +THE INDEX\n> +---------\n> +\n> +The index, also known as the \"staging area\", contains the current staged\n\nHonestly, I always forget which of these two nouns we are supposed to\nuse nowadays. I think consensus was to use \"index\" and avoid using\n\"staging area\"? Not sure though, but I think we should only mention\none of these.\n\n> +version of every file in your Git repository. When you commit, the files\n> +in the index are used as the files in the next commit.\n> +\n> +Unlike a tree, the index is a flat list of files.\n> +Each index entry has 4 fields:\n> +\n> +1. The *permissions*\n> +2. The *<<blob,blob>> ID* of the file\n> +3. The *filename*\n> +4. The *number*. This is normally 0, but if there's a merge conflict\n\nI think we don't call this \"number\", but \"stage\".\n\n> +   there can be multiple versions (with numbers 0, 1, 2, ..)\n> +   of the same filename in the index.\n> +\n> +It's extremely uncommon to look at the index directly: normally you'd\n> +run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n> +But you can use `git ls-files --stage` to see the index.\n> +Here's the output of `git ls-files --stage` in a repository with 2 files:\n> +\n> +----\n> +100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n> +100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n> +----\n> +\n> +[[reflogs]]\n> +REFLOGS\n> +-------\n> +\n> +Git stores the history of branch, tag, and HEAD refs in a reflog\n> +(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n> +default, but any ref can be logged.\n\nIf we mention this here, do we maybe want to mention how the user can\ndecide which references are logged?\n\n> +Each reflog entry has:\n> +\n> +1. *Before/after *commit IDs*\n\nThis will probably misformat as we have three asterisks here, not two.\n\n> +2. *User* who made the change, for example `Maya <maya@example.com>`\n> +3. *Timestamp*\n\nSuggestion: \"*Timestamp* when that change has been made\".\n\n> +4. *Log message*, for example `pull: Fast-forward`\n> +\n> +Reflogs only log changes made in your local repository.\n> +They are not shared with remotes.\n\nWe may want ot mention that you can reference reflog entries via\n`refs/heads/<branch>@{<reflog-nr>}`.\n\nIn general, one thing that I think would be important to highlight in\nthis document is revisions. Most of the commands tend to not accept\nreferences, but revisions instead, which are a lot more flexible. They\nuse our do-what-I-mean mechanism to resolve, but also allow the user to\nspecify commits relative to one another. It's probably sufficient though\nto mention them briefly and then redirect to girevisions(7).\n\nThanks for working on this!\n\nPatrick\n"},{"id":"528132","messageId":"xmqqbjmillai.fsf@gitster.g","threadId":"64244","inReplyTo":"93b30d1e-7d49-44cf-b29b-69e8055bccbc@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-07T16:38:13Z","receivedAt":"2025-10-07T16:38:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Kristoffer Haugsbakk\" <kristofferhaugsbakk@fastmail.com> writes:\n\n> On Mon, Oct 6, 2025, at 05:32, Junio C Hamano wrote:\n>> \"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>>> +MAN7_TXT += gitdatamodel.adoc\n>>>  MAN7_TXT += gitdiffcore.adoc\n>>> ...\n>>> +gitdatamodel(7)\n>>> +===============\n>>> +\n>>> +NAME\n>>> +----\n>>> +gitdatamodel - Git's core data model\n>>> +\n>>> +DESCRIPTION\n>>> +-----------\n>>\n>> The above causes doc-lint to barf.\n>>[snip]\n>> You can check locally with \"make check-docs\" without waiting for my\n>> integration cycle to push to GitHub CI.\n>\n> I think you meant `make lint-docs` for both of these.\n\nThe former is a typo for \"causes lint-docs to barf\", but I did mean\n\"make check-docs\" as the recipe for local checking.\n\nYou could also do \"make -C Documentation lint-docs\", but that is a\nlot more to type ;-).\n\nThanks.\n"},{"id":"528134","messageId":"xmqq4isalk5g.fsf@gitster.g","threadId":"64244","inReplyTo":"aOUkZa4_fq1hho7Q@pks.im","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-07T17:02:51Z","receivedAt":"2025-10-07T17:02:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> +Git's core operations use 4 kinds of data:\n>> +\n>> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n>> +2. <<references,References>>: branches, tags,\n>> +   remote-tracking branches, etc\n>> +3. <<index,The index>>, also known as the staging area\n>> +4. <<reflogs,Reflogs>>\n>\n> This list makes sense to me. There's of course more data structures in\n> Git, but all the other data structures shouldn't really matter to users\n> at all as they are mostly caches or internal details of the on-disk\n> format.\n>\n> There's potentially one exception though, namely the Git configuration.\n> I'd claim that Git \"uses\" the Git configuration similarly to how it uses\n> the others, but I get why it's not explicitly mentioned here.\n\nThe core operations do not use Git configuration any more than they\nuse what is specified by the command line arguments.\n\n>> +[[objects]]\n>> +OBJECTS\n>> +-------\n>> +\n>> +Commits, trees, blobs, and tag objects are all stored in Git's object database.\n>> +Every object has:\n>> +\n>> +1. an *ID*, which is the SHA-1 hash of its contents.\n>\n> I think this needs to be adapted to not single out SHA-1 as the only\n> hashing algorithm. We already support SHA-256, so we should definitely\n> say that the algorithm can be swapped. Maybe something like:\n\nGood point.  Also officially they are called \"object name\".\n\n>   An *object ID*, which is the cryptographic hash of its contents. By\n>   default, Git uses SHA-1 as object hash, but alternative hashes like\n>   SHA-256 are supported.\n\nI'd avoid \"object name is the result of hashing X\" which historically\nwas a source of question: \"why does 'sha1sum README.md' give different\nhash from 'git add README.md && git ls-files -s README.md'?\"\n\nIt is an irrelevant implementation detail (and you'd eventually end\nup having to say \"X is <type> SP <length> NUL <contents>\").\n\n    An object name, which is derived cryptographically from its\n    type, size and contents.  All versions of Git can use SHA-1 hash\n    function, but more recent versions of Git can also use SHA-256\n    hash function.\n\n>> +commits::\n>> +    A commit contains:\n>> ++\n>> +1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n>> +  regular commits have 1 parent, merge commits have 2+ parents\n>\n> I'd say \"at least two parents\" instead of \"2+ parents\".\n\nYup, that reads much better.\n\n>> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n>> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n>> +author Maya <maya@example.com> 1759173425 -0400\n>> +committer Maya <maya@example.com> 1759173425 -0400\n>> +\n>> +Add README\n>> +----\n>\n> In practice, commits can have other headers that are ignored by Git. But\n> that's certainly not part of Git's core data model, so I don't think we\n> should mention that here.\n\nThird-party software can add truly garbage ones that do not have any\nmeaning, and Git tolerates by ignoring them.  But there are others\nthat Git does pay attention to, like encoding, gpgsig, etc., which\nmay worth mention (in the form that \"these four are what you typically\nsee, but there may be others\" without even naming any).\n\n\n"},{"id":"528150","messageId":"CALnO6CDgB+yWoVv+eP4eNhVkVLw7hXb==1q3Ve+OnkZuERiYYw@mail.gmail.com","threadId":"64244","inReplyTo":"aOUkZa4_fq1hho7Q@pks.im","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-10-07T18:39:52Z","receivedAt":"2025-10-07T18:40:06Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Tue, Oct 7, 2025 at 11:51 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Fri, Oct 03, 2025 at 05:34:36PM +0000, Julia Evans via GitGitGadget wrote:\n[snip]\n> > +    A branch is a name for a commit ID.\n> > +    That commit is the latest commit on the branch.\n> > +    Branches are stored in the `.git/refs/heads/` directory.\n> > ++\n> > +To get the history of commits on a branch, Git will start at the commit\n> > +ID the branch references, and then look at the commit's parent(s),\n> > +the parent's parent, etc.\n> > +\n> > +[[tag]]\n> > +tags: `.git/refs/tags/<name>`::\n> > +    A tag is a name for a commit ID, tag object ID, or other object ID.\n> > +    Tags are stored in the `refs/tags/` directory.\n> > ++\n> > +Even though branches and commits are both \"a name for a commit ID\", Git\n> > +treats them very differently.\n> > +Branches are expected to be regularly updated as you work on the branch,\n> > +but it's expected that a tag will never change after you create it.\n>\n> This sounds a bit like the user itself needs to update the branch. How\n> about this instead:\n>\n>     Even though branches and commits are both \"a name for a commit ID\", Git\n>     treats them very differently:\n>\n>         - Branches can be checked out directly. If so, creating a new\n>           commit will automatically update the checked-out branch to\n>           point to the new commit.\n>\n>         - Tags cannot be checked out directly and don't move when\n>           creating a new commit. Instead, one can only check out the\n>           commit that a branch points to. This is called \"detached\n>           HEAD\", and the effect is that a new commit will not update\n\nmissing \"the tag.\" ?\n"},{"id":"528151","messageId":"dbf0727d-66bf-4698-aa21-d69da86027c3@app.fastmail.com","threadId":"64244","inReplyTo":"aOUkZa4_fq1hho7Q@pks.im","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-07T18:55:37Z","receivedAt":"2025-10-07T18:55:58Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Tue, Oct 7, 2025, at 10:32 AM, Patrick Steinhardt wrote:\n> On Fri, Oct 03, 2025 at 05:34:36PM +0000, Julia Evans via GitGitGadget wrote:\n>> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n>> new file mode 100644\n>> index 0000000000..4b2cb167dc\n>> --- /dev/null\n>> +++ b/Documentation/gitdatamodel.adoc\n>> @@ -0,0 +1,226 @@\n>> +gitdatamodel(7)\n>> +===============\n>> +\n>> +NAME\n>> +----\n>> +gitdatamodel - Git's core data model\n>> +\n>> +DESCRIPTION\n>> +-----------\n>> +\n>> +It's not necessary to understand Git's data model to use Git, but it's\n>> +very helpful when reading Git's documentation so that you know what it\n>> +means when the documentation says \"object\" \"reference\" or \"index\".\n>\n> There's a missing comma after \"object\".\n\nWill fix.\n\n>> +\n>> +Git's core operations use 4 kinds of data:\n>> +\n>> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n>> +2. <<references,References>>: branches, tags,\n>> +   remote-tracking branches, etc\n>> +3. <<index,The index>>, also known as the staging area\n>> +4. <<reflogs,Reflogs>>\n>\n> This list makes sense to me. There's of course more data structures in\n> Git, but all the other data structures shouldn't really matter to users\n> at all as they are mostly caches or internal details of the on-disk\n> format.\n>\n> There's potentially one exception though, namely the Git configuration.\n> I'd claim that Git \"uses\" the Git configuration similarly to how it uses\n> the others, but I get why it's not explicitly mentioned here.\n>\n>> +[[objects]]\n>> +OBJECTS\n>> +-------\n>> +\n>> +Commits, trees, blobs, and tag objects are all stored in Git's object database.\n>> +Every object has:\n>> +\n>> +1. an *ID*, which is the SHA-1 hash of its contents.\n>\n> I think this needs to be adapted to not single out SHA-1 as the only\n> hashing algorithm. We already support SHA-256, so we should definitely\n> say that the algorithm can be swapped. Maybe something like:\n>\n>   An *object ID*, which is the cryptographic hash of its contents. By\n>   default, Git uses SHA-1 as object hash, but alternative hashes like\n>   SHA-256 are supported.\n\nMakes sense. I might just say \"cryptographic hash of its type and contents\"\nand leave it that. I'm not sure it's worth getting into details\nof the exact hash function.\n\n>> +  It's fast to look up a Git object using its ID.\n>> +  The ID is usually represented in hexadecimal, like\n>> +  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n>> +2. a *type*. There are 4 types of objects:\n>> +   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n>> +   and <<tag-object,tag objects>>.\n>> +3. *contents*. The structure of the contents depends on the type.\n>\n> Nit: every object also has an object size. Not sure though whether it's\n> fine to imply that with \"contents\".\n\nI think it is.\n\n>> +Once an object is created, it can never be changed.\n>> +Here are the 4 types of objects:\n>> +\n>> +[[commit]]\n>> +commits::\n>> +    A commit contains:\n>> ++\n>> +1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n>> +  regular commits have 1 parent, merge commits have 2+ parents\n>\n> I'd say \"at least two parents\" instead of \"2+ parents\".\n>\n>> +2. A *commit message*\n>> +3. All the *files* in the commit, stored as a *<<tree,tree>>*\n>> +4. An *author* and the time the commit was authored\n>> +5. A *committer* and the time the commit was committed\n>> ++\n>> +Here's how an example commit is stored:\n>> ++\n>> +----\n>> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n>> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n>> +author Maya <maya@example.com> 1759173425 -0400\n>> +committer Maya <maya@example.com> 1759173425 -0400\n>> +\n>> +Add README\n>> +----\n>\n> In practice, commits can have other headers that are ignored by Git. But\n> that's certainly not part of Git's core data model, so I don't think we\n> should mention that here.\n>\n>> +Like all other objects, commits can never be changed after they're created.\n>> +For example, \"amending\" a commit with `git commit --amend` creates a new commit.\n>> +The old commit will eventually be deleted by `git gc`.\n>\n> If we mention git-gc(1) I think it would make sense to use\n> `linkgit:git-gc[1]` instead to provide a link to its man page.\n\nAgreed.\n\n>> +[[tree]]\n>> +trees::\n>> +    A tree is how Git represents a directory. It lists, for each item in\n>> +    the tree:\n>> ++\n>> +1. The *permissions*, for example `100644`\n>\n> I think we should rather call these \"mode bits\". These bits are\n> permissions indeed when you have a blob, but for subtrees, symlinks and\n> submodules they aren't.\n\nI think it's a bit strange to call them mode bits since I thought they were stored\nas ASCII strings and it's basically an enum of 5 options, but I see your point.\nI think \"file mode\" will work and that's used elsewhere.\n\nI wonder if it would make sense to list all of the possible file modes if\nthis isn't documented anywhere else, my impression is that it's a short\nlist and that it's unlikely to change much in the future.\n\nAnd listing them all might make it more clear that Git's file modes don't\nhave much in common with Unix file modes.\nI looked for where this is documented and it looks like the only place is\nin `man git-fast-import` . That man page says that there are just 5 options\n(040000, 160000, 100644, 100755, 120000)\n\n>> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n>> +  or <<commit,`commit`>> (a Git submodule)\n>\n> There's also symlinks.\n\nI created a test symlink and it looks like symlinks are stored as type \"blob\".\nI might say which type corresponds to which file mode,\nthough I'm not sure what type corresponds to the \"gitlink\" mode (commit?).\n\nI think these are the 5 modes and what they mean / what type they\nshould have. Not sure about the gitlink mode though.\n\n  - `100644`: regular file (with type `blob`)\n  - `100755`: executable file (with type `blob`)\n  - `120000`: symbolic link (with type `blob`)\n  - `040000`: directory (with type `tree`)\n  - `160000`: gitlink, for use with submodules (with type `commit`)\n\n>> +3. The *object ID*\n>> +4. The *filename*\n>> ++\n>> +For example, this is how a tree containing one directory (`src`) and one file\n>> +(`README.md`) is stored:\n>> ++\n>> +----\n>> +100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n>> +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n>> +----\n>> ++\n>> +*NOTE:* The permissions are in the same format as UNIX permissions, but\n>> +the only allowed permissions for files (blobs) are 644 and 755.\n>> +\n>> +[[blob]]\n>> +blobs::\n>> +    A blob is how Git represents a file. A blob object contains the\n>> +    file's contents.\n>> ++\n>> +Storing a new blob for every new version of a file can get big, so\n>> +`git gc` periodically compresses objects for efficiency in `.git/objects/pack`.\n>\n> I would claim that it's not necessary to mention object compression.\n> This should be a low-level detail that users don't ever have to worry\n> about. Furthermore, packing objects isn't only relevant in the context\n> of blobs: trees for example also tend to compress very well as there\n> typically is only small incremental updates to trees.\n\nI discussed why I think this important in another reply,\nhttps://lore.kernel.org/all/51e0a55c-1f1d-4cae-9459-8c2b9220e52d@app.fastmail.com/,\nwill paste what I said here. I'll think about this more though.\n\npaste follows:\n\nThat's true! The reason I think this is important to mention is that I find\nthat people often \"reject\" information that they find implausible, even\nif it comes from a credible source. (\"that can't be true! I must be\nnot understanding correctly. Oh well, I'll just ignore that!\")\n\nI sometimes hear from users that \"commits can't be snapshots\", because\nit would take up too much disk space to store every version of\nevery commit. So I find that sometimes explaining a little bit about the\nimplementation can make the information more memorable.\n\nCertainly I'm not able to remember details that don't make sense\nwith my mental model of how computers work and I don't expect other\npeople to either, so I think it's important to give an explanation that\nhandles the biggest \"objections\".\n\n>> +[[tag-object]]\n>> +tag objects::\n>> +    Tag objects (also known as \"annotated tags\") contain:\n>> ++\n>> +1. The *tagger* and tag date\n>> +2. A *tag message*, similar to a commit message\n>> +3. The *ID* of the object (often a commit) that they reference\n>\n> They can also be signed, if we want to mention that.\n\nI guess that's true for commit objects too. Not sure whether to\nmention it either, can add it if others think it's important.\n\n>> +[[references]]\n>> +REFERENCES\n>> +----------\n>> +\n>> +References are a way to give a name to a commit.\n>> +It's easier to remember \"the changes I'm working on are on the `turtle`\n>> +branch\" than \"the changes are in commit bb69721404348e\".\n>> +Git often uses \"ref\" as shorthand for \"reference\".\n>> +\n>> +References that you create are stored in the `.git/refs` directory,\n>> +and Git has a few special internal references like `HEAD` that are stored\n>> +in the base `.git` directory.\n>\n> This isn't true anymore with the introduction of the reftable backend,\n> which is slated to become the default backend. I'd argue that this is\n> another implementation detail that the user shouldn't have to worry\n> about.\n\nMakes sense, will fix. (as well as other references to the .git prefix and\n\"subdirectories\").\n\n>> +References can either be:\n>> +\n>> +1. References to an object ID, usually a <<commit,commit>> ID\n>> +2. References to another reference. This is called a \"symbolic reference\".\n>> +\n>> +Git handles references differently based on which subdirectory of\n>> +`.git/refs` they're stored in.\n>\n> So instead of saying \"subdirectory\", I'd rather say \"reference\n> hierarchy\".\n>\n> In general, I think we should explain that references are layed out\n> in a hierarchy. This is somewhat obvious with the \"files\" backend, as we\n> use directories there. But as we move on to the \"reftable\" backend this\n> may become less obvious over time.\n\nThat makes sense.\n\n>> +[[tag]]\n>> +tags: `.git/refs/tags/<name>`::\n>> +    A tag is a name for a commit ID, tag object ID, or other object ID.\n>> +    Tags are stored in the `refs/tags/` directory.\n>> ++\n>> +Even though branches and commits are both \"a name for a commit ID\", Git\n>> +treats them very differently.\n>> +Branches are expected to be regularly updated as you work on the branch,\n>> +but it's expected that a tag will never change after you create it.\n>\n> This sounds a bit like the user itself needs to update the branch. How\n> about this instead:\n>\n>     Even though branches and commits are both \"a name for a commit ID\", Git\n>     treats them very differently:\n>\n>         - Branches can be checked out directly. If so, creating a new\n>           commit will automatically update the checked-out branch to\n>           point to the new commit.\n>\n>         - Tags cannot be checked out directly and don't move when\n>           creating a new commit. Instead, one can only check out the\n>           commit that a branch points to. This is called \"detached\n>           HEAD\", and the effect is that a new commit will not update \n\nI think mentioning that branches can be checked out and that tags can't\nis a good idea.\n\n>> +[[HEAD]]\n>> +HEAD: `.git/HEAD`::\n>> +    `HEAD` is where Git stores your current <<branch,branch>>.\n>> +    `HEAD` is normally a symbolic reference to your current branch, for\n>> +    example `ref: refs/heads/main` if your current branch is `main`.\n>> +    `HEAD` can also be a direct reference to a commit ID,\n>> +    that's called \"detached HEAD state\".\n>> +\n>> +[[remote-tracking-branch]]\n>> +remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::\n>> +    A remote-tracking branch is a name for a commit ID.\n>> +    It's how Git stores the last-known state of a branch in a remote\n>> +    repository. `git fetch` updates remote-tracking branches. When\n>> +    `git status` says \"you're up to date with origin/main\", it's looking at\n>> +    this.\n>\n> This misses \"refs/remotes/<remote>/HEAD\". This reference is a symbolic\n> reference that indicates the default branch on the remote side.\n\nIs \"refs/remotes/<remote>/HEAD\" a remote-tracking branch?\nI've never thought about that reference and I'm not sure what to call it.\n\n>> +[[other-refs]]\n>> +Other references::\n>> +    Git tools may create references in any subdirectory of `.git/refs`.\n>> +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n>> +    and linkgit:git-notes[1] all create their own references\n>> +    in `.git/refs/stash`, `.git/refs/bisect`, etc.\n>> +    Third-party Git tools may also create their own references.\n>> ++\n>> +Git may also create references in the base `.git` directory\n>> +other than `HEAD`, like `ORIG_HEAD`.\n>\n> Let's mention that such references are typically spelt all-uppercase\n> with underscores between. You shouldn't ever create a reference that is\n> for example called \".git/foo\".\n>\n> We enforce this restriction inconsistently, only, but I don't think that\n> should keep us from spelling out the common rule.\n\nThat makes sense. I'm also not sure whether third-party\nGit tools are \"supposed\" to create references outside of \"refs/\",\nor whether that's common. \n\n>> +*NOTE:* As an optimization, references may be stored as packed\n>> +refs instead of in `.git/refs`. See linkgit:git-pack-refs[1].\n>\n> I'd drop this note. It's an internal implementation detail and only true\n> for the \"files\" backend. The \"reftable\" backend stores references quite\n> differently and doesn't really \"pack\" references.\n>\n>> +[[index]]\n>> +THE INDEX\n>> +---------\n>> +\n>> +The index, also known as the \"staging area\", contains the current staged\n>\n> Honestly, I always forget which of these two nouns we are supposed to\n> use nowadays. I think consensus was to use \"index\" and avoid using\n> \"staging area\"? Not sure though, but I think we should only mention\n> one of these.\n>\n>> +version of every file in your Git repository. When you commit, the files\n>> +in the index are used as the files in the next commit.\n>> +\n>> +Unlike a tree, the index is a flat list of files.\n>> +Each index entry has 4 fields:\n>> +\n>> +1. The *permissions*\n>> +2. The *<<blob,blob>> ID* of the file\n>> +3. The *filename*\n>> +4. The *number*. This is normally 0, but if there's a merge conflict\n>\n> I think we don't call this \"number\", but \"stage\".\n\nThanks, I see that it's sometimes called \"stage number\" which is a little\neasier to search for so I'll call it that.\n\n>> +   there can be multiple versions (with numbers 0, 1, 2, ..)\n>> +   of the same filename in the index.\n>> +\n>> +It's extremely uncommon to look at the index directly: normally you'd\n>> +run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n>> +But you can use `git ls-files --stage` to see the index.\n>> +Here's the output of `git ls-files --stage` in a repository with 2 files:\n>> +\n>> +----\n>> +100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n>> +100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n>> +----\n>> +\n>> +[[reflogs]]\n>> +REFLOGS\n>> +-------\n>> +\n>> +Git stores the history of branch, tag, and HEAD refs in a reflog\n>> +(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n>> +default, but any ref can be logged.\n>\n> If we mention this here, do we maybe want to mention how the user can\n> decide which references are logged?\n\nDo you mean by using the setting `core.logAllRefUpdates`?\n\n>> +Each reflog entry has:\n>> +\n>> +1. *Before/after *commit IDs*\n>\n> This will probably misformat as we have three asterisks here, not two.\n>\n>> +2. *User* who made the change, for example `Maya <maya@example.com>`\n>> +3. *Timestamp*\n>\n> Suggestion: \"*Timestamp* when that change has been made\".\n\nMakes sense.\n\n>> +4. *Log message*, for example `pull: Fast-forward`\n>> +\n>> +Reflogs only log changes made in your local repository.\n>> +They are not shared with remotes.\n>\n> We may want ot mention that you can reference reflog entries via\n> `refs/heads/<branch>@{<reflog-nr>}`.\n>\n> In general, one thing that I think would be important to highlight in\n> this document is revisions. Most of the commands tend to not accept\n> references, but revisions instead, which are a lot more flexible. They\n> use our do-what-I-mean mechanism to resolve, but also allow the user to\n> specify commits relative to one another. It's probably sufficient though\n> to mention them briefly and then redirect to girevisions(7).\n\nWill think about this, I'm not sure how to best incorporate that.\nMaybe under the commits section.\n\n> Thanks for working on this!\n\nThanks for the review!\n\n- Julia\n"},{"id":"528152","messageId":"ede082ad-5031-4b55-8576-0a6315f16b70@app.fastmail.com","threadId":"64244","inReplyTo":"xmqq4isalk5g.fsf@gitster.g","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-07T19:30:47Z","receivedAt":"2025-10-07T19:31:14Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":">> I think this needs to be adapted to not single out SHA-1 as the only\n>> hashing algorithm. We already support SHA-256, so we should definitely\n>> say that the algorithm can be swapped. Maybe something like:\n>\n> Good point.  Also officially they are called \"object name\".\n\nI hadn't realized that \"object name\" was the official name, it does\nseem to be used a lot in the docs. I'm going to try something like this:\n\n1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n  type and contents.\n\nI think it's useful to refer this as an \"ID\", because usually we call it a\n\"commit ID\" or \"tag ID\" and not a \"commit name\" or \"tag name\"\nand it makes it more clear that \"object name\" and \"commit ID\"\nrefer to the same identifier.\n\n>>> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n>>> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n>>> +author Maya <maya@example.com> 1759173425 -0400\n>>> +committer Maya <maya@example.com> 1759173425 -0400\n>>> +\n>>> +Add README\n>>> +----\n>>\n>> In practice, commits can have other headers that are ignored by Git. But\n>> that's certainly not part of Git's core data model, so I don't think we\n>> should mention that here.\n>\n> Third-party software can add truly garbage ones that do not have any\n> meaning, and Git tolerates by ignoring them.  But there are others\n> that Git does pay attention to, like encoding, gpgsig, etc., which\n> may worth mention (in the form that \"these four are what you typically\n> see, but there may be others\" without even naming any).\n\nI didn't realize that there were other optional fields,\nwill try to communicate this somehow.\n"},{"id":"528154","messageId":"xmqqwm56iiq1.fsf@gitster.g","threadId":"64244","inReplyTo":"ede082ad-5031-4b55-8576-0a6315f16b70@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-07T20:01:58Z","receivedAt":"2025-10-07T20:02:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> I think it's useful to refer this as an \"ID\", because usually we call it a\n> \"commit ID\" or \"tag ID\" and not a \"commit name\" or \"tag name\"\n> and it makes it more clear that \"object name\" and \"commit ID\"\n> refer to the same identifier.\n\nIt is a bit funny that they do not exactly align.\n\n    \"object name\" aka \"object ID\"\n    \"$type object name\" aka \"$type ID\" for type in (commit, blob, tree, tag)\n\nIn any case, we should add \"object ID\" and other \"$type ID\" to the\nglossary, if you are going to use it very often.  We have entries\nfor spelled out \"identifier\" but I do not think \"ID\" is there yet.\n\nThanks.\n"},{"id":"528194","messageId":"aOXmA5L5LsUuXWEh@pks.im","threadId":"64244","inReplyTo":"dbf0727d-66bf-4698-aa21-d69da86027c3@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-08T04:18:11Z","receivedAt":"2025-10-08T04:18:18Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Oct 07, 2025 at 02:55:37PM -0400, Julia Evans wrote:\n> On Tue, Oct 7, 2025, at 10:32 AM, Patrick Steinhardt wrote:\n> > On Fri, Oct 03, 2025 at 05:34:36PM +0000, Julia Evans via GitGitGadget wrote:\n> >> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n> >> new file mode 100644\n> >> index 0000000000..4b2cb167dc\n> >> --- /dev/null\n> >> +++ b/Documentation/gitdatamodel.adoc\n[snip]\n> >> +[[tree]]\n> >> +trees::\n> >> +    A tree is how Git represents a directory. It lists, for each item in\n> >> +    the tree:\n> >> ++\n> >> +1. The *permissions*, for example `100644`\n> >\n> > I think we should rather call these \"mode bits\". These bits are\n> > permissions indeed when you have a blob, but for subtrees, symlinks and\n> > submodules they aren't.\n> \n> I think it's a bit strange to call them mode bits since I thought they were stored\n> as ASCII strings and it's basically an enum of 5 options, but I see your point.\n> I think \"file mode\" will work and that's used elsewhere.\n> \n> I wonder if it would make sense to list all of the possible file modes if\n> this isn't documented anywhere else, my impression is that it's a short\n> list and that it's unlikely to change much in the future.\n\nAgreed, that seems reasonable to me.\n\n> And listing them all might make it more clear that Git's file modes don't\n> have much in common with Unix file modes.\n> I looked for where this is documented and it looks like the only place is\n> in `man git-fast-import` . That man page says that there are just 5 options\n> (040000, 160000, 100644, 100755, 120000)\n> \n> >> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n> >> +  or <<commit,`commit`>> (a Git submodule)\n> >\n> > There's also symlinks.\n> \n> I created a test symlink and it looks like symlinks are stored as type \"blob\".\n> I might say which type corresponds to which file mode,\n> though I'm not sure what type corresponds to the \"gitlink\" mode (commit?).\n\nYeah, gitlinks are used for submodules. They point to an object ID that\nrefers to a commit in the submodule itself.\n\n> I think these are the 5 modes and what they mean / what type they\n> should have. Not sure about the gitlink mode though.\n> \n>   - `100644`: regular file (with type `blob`)\n>   - `100755`: executable file (with type `blob`)\n>   - `120000`: symbolic link (with type `blob`)\n>   - `040000`: directory (with type `tree`)\n>   - `160000`: gitlink, for use with submodules (with type `commit`)\n\nThis list looks good to me. gitlinks are somewhat special given that\nthey refer to a commit stored in the submodule repository, not in the\nrepository that has the gitlink. But the expectation is that the object\nname should always resolve to a commit indeed.\n\n[snip]\n> >> +[[blob]]\n> >> +blobs::\n> >> +    A blob is how Git represents a file. A blob object contains the\n> >> +    file's contents.\n> >> ++\n> >> +Storing a new blob for every new version of a file can get big, so\n> >> +`git gc` periodically compresses objects for efficiency in `.git/objects/pack`.\n> >\n> > I would claim that it's not necessary to mention object compression.\n> > This should be a low-level detail that users don't ever have to worry\n> > about. Furthermore, packing objects isn't only relevant in the context\n> > of blobs: trees for example also tend to compress very well as there\n> > typically is only small incremental updates to trees.\n> \n> I discussed why I think this important in another reply,\n> https://lore.kernel.org/all/51e0a55c-1f1d-4cae-9459-8c2b9220e52d@app.fastmail.com/,\n> will paste what I said here. I'll think about this more though.\n> \n> paste follows:\n> \n> That's true! The reason I think this is important to mention is that I find\n> that people often \"reject\" information that they find implausible, even\n> if it comes from a credible source. (\"that can't be true! I must be\n> not understanding correctly. Oh well, I'll just ignore that!\")\n> \n> I sometimes hear from users that \"commits can't be snapshots\", because\n> it would take up too much disk space to store every version of\n> every commit. So I find that sometimes explaining a little bit about the\n> implementation can make the information more memorable.\n> \n> Certainly I'm not able to remember details that don't make sense\n> with my mental model of how computers work and I don't expect other\n> people to either, so I think it's important to give an explanation that\n> handles the biggest \"objections\".\n\nHm, fair I guess. In any case, if we want to mention this I'd leave away\nthe details how exactly Git achieves this. E.g. we could say something\nlike:\n\n    Storing a new blob for every new version of a file can result to a\n    lot of duplication. Git regularly runs repository maintenance to\n    optimize to counteract this. Part of the maintenance involves\n    compression of objects, where incremental changes to the same object\n    are optimized to be stored as deltas, only.\n\nWe skip over the details, but this should give enough pointers to an\ninterested reader to go dig deeper. We could also generalize this to\nobjects in general, not only blobs.\n\n[snip]\n> >> +[[HEAD]]\n> >> +HEAD: `.git/HEAD`::\n> >> +    `HEAD` is where Git stores your current <<branch,branch>>.\n> >> +    `HEAD` is normally a symbolic reference to your current branch, for\n> >> +    example `ref: refs/heads/main` if your current branch is `main`.\n> >> +    `HEAD` can also be a direct reference to a commit ID,\n> >> +    that's called \"detached HEAD state\".\n> >> +\n> >> +[[remote-tracking-branch]]\n> >> +remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::\n> >> +    A remote-tracking branch is a name for a commit ID.\n> >> +    It's how Git stores the last-known state of a branch in a remote\n> >> +    repository. `git fetch` updates remote-tracking branches. When\n> >> +    `git status` says \"you're up to date with origin/main\", it's looking at\n> >> +    this.\n> >\n> > This misses \"refs/remotes/<remote>/HEAD\". This reference is a symbolic\n> > reference that indicates the default branch on the remote side.\n> \n> Is \"refs/remotes/<remote>/HEAD\" a remote-tracking branch?\n> I've never thought about that reference and I'm not sure what to call it.\n\nNo, it's not. I think the term we use is \"remote reference\".\n\n> >> +[[other-refs]]\n> >> +Other references::\n> >> +    Git tools may create references in any subdirectory of `.git/refs`.\n> >> +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n> >> +    and linkgit:git-notes[1] all create their own references\n> >> +    in `.git/refs/stash`, `.git/refs/bisect`, etc.\n> >> +    Third-party Git tools may also create their own references.\n> >> ++\n> >> +Git may also create references in the base `.git` directory\n> >> +other than `HEAD`, like `ORIG_HEAD`.\n> >\n> > Let's mention that such references are typically spelt all-uppercase\n> > with underscores between. You shouldn't ever create a reference that is\n> > for example called \".git/foo\".\n> >\n> > We enforce this restriction inconsistently, only, but I don't think that\n> > should keep us from spelling out the common rule.\n> \n> That makes sense. I'm also not sure whether third-party\n> Git tools are \"supposed\" to create references outside of \"refs/\",\n> or whether that's common. \n\nThey really shouldn't, and to the best of my knowledge they don't. There\nis only a rather limited number of root references with very specific\nuse cases. And nowadays we have also tightened the meaning of pseudo\nrefs, of which there are only two (\"FETCH_HEAD\" and \"MERGE_HEAD\").\n\n[snip]\n> >> +[[reflogs]]\n> >> +REFLOGS\n> >> +-------\n> >> +\n> >> +Git stores the history of branch, tag, and HEAD refs in a reflog\n> >> +(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n> >> +default, but any ref can be logged.\n> >\n> > If we mention this here, do we maybe want to mention how the user can\n> > decide which references are logged?\n> \n> Do you mean by using the setting `core.logAllRefUpdates`?\n\nYeah. Otherwise the reader won't have any pointers to figure out _how_\nthey can change this. I don't think we have a man page that provides a\nbetter overview than this configuration.\n\nThanks!\n\nPatrick\n"},{"id":"528230","messageId":"b0220dca-d324-49c8-af80-7b19f3b70691@app.fastmail.com","threadId":"64244","inReplyTo":"51e0a55c-1f1d-4cae-9459-8c2b9220e52d@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-10-08T09:59:39Z","receivedAt":"2025-10-08T10:00:04Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"On Mon, Oct 6, 2025, at 21:36, Julia Evans wrote:\n>[snip]\n>>> +blobs::\n>>> +    A blob is how Git represents a file. A blob object contains the\n>>> +    file's contents.\n>>> ++\n>>> +Storing a new blob for every new version of a file can get big, so\n>>> +`git gc` periodically compresses objects for efficiency in\n>>> `.git/objects/pack`.\n>>\n>> This gets into mentioning implementation files(?) like you mentioned in\n>> the commit message.\n>\n> That's true! The reason I think this is important to mention is that I find\n> that people often \"reject\" information that they find implausible, even\n> if it comes from a credible source. (\"that can't be true! I must be\n> not understanding correctly. Oh well, I'll just ignore that!\")\n>\n> I sometimes hear from users that \"commits can't be snapshots\", because\n> it would take up too much disk space to store every version of\n> every commit. So I find that sometimes explaining a little bit about the\n> implementation can make the information more memorable.\n>\n> Certainly I'm not able to remember details that don't make sense\n> with my mental model of how computers work and I don't expect other\n> people to either, so I think it's important to give an explanation that\n> handles the biggest \"objections\".\n\nThat’s very intresting. Yes, maybe people need to be told/taught to a\nlevel which might be considered “just implementation details” or else\nboth neither their curiosity won’t be satisfied *nor* will their own\nsense of error-correction for the seemingly implausible.\n\n>[snip]\n>> I appreciate that this is the first version and you might have plans\n>> after this one. But I wonder if this doc could use a fair number of\n>> `gitlink` to branch out to all the other parts. Like git-reflog(1),\n>> gitglossary(7).\n>\n> That's reasonable. Do you often use the \"See also\" section of\n> man pages? I've never looked at them so I'm always curious about\n> how people are actually using them in practice.\n\nI don’t really use See Also when looking things up. But I notice all the\nmentions of other docs in running text.\n\n>[snip]\n"},{"id":"528251","messageId":"pull.1981.v2.git.1759931621272.gitgitgadget@gmail.com","threadId":"64244","inReplyTo":"pull.1981.git.1759512876284.gitgitgadget@gmail.com","subject":"[PATCH v2] doc: add a explanation of Git's data model","fromName":"Julia Evans via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-08T13:53:41Z","receivedAt":"2025-10-08T13:53:44Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"From: Julia Evans <julia@jvns.ca>\n\nGit very often uses the terms \"object\", \"reference\", or \"index\" in its\ndocumentation.\n\nHowever, it's hard to find a clear explanation of these terms and how\nthey relate to each other in the documentation. The closest candidates\ncurrently are:\n\n1. `gitglossary`. This makes a good effort, but it's an alphabetically\n    ordered dictionary and a dictionary is not a good way to learn\n    concepts. You have to jump around too much and it's not possible to\n    present the concepts in the order that they should be explained.\n2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n   This is a nice document to have, but it's not necessary to learn how\n   `update-index` works to understand Git's data model, and we should\n   not be requiring users to learn how to use the \"plumbing\" commands\n   if they want to learn what the term \"index\" or \"object\" means.\n3. `gitrepository-layout`. This is a great resource, but it includes a\n   lot of information about configuration and internal implementation\n   details which are not related to the data model. It also does\n   not explain how commits work.\n\nThe result of this is that Git users (even users who have been using\nGit for 15+ years) struggle to read the documentation because they don't\nknow what the core terms mean, and it's not possible to add links\nto help them learn more.\n\nAdd an explanation of Git's data model. Some choices I've made in\ndeciding what \"core data model\" means:\n\n1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n   if those are intended to be user facing or if they're more like\n   internal implementation details.\n2. Don't talk about submodules other than by mentioning how they\n   relate to trees. This is because Git has a lot of special features,\n   and explaining how they all work exhaustively could quickly go\n   down a rabbit hole which would make this document less useful for\n   understanding Git's core behaviour.\n3. Don't discuss the structure of a commit message\n   (first line, trailers etc).\n4. Don't mention configuration.\n5. Don't mention the `.git` directory, to avoid getting too much into\n   implementation details\n\nSigned-off-by: Julia Evans <julia@jvns.ca>\n---\n    doc: Add a explanation of Git's data model\n    \n    Changes in v2:\n    \n    The biggest change is to remove all mentions of the .git directory, and\n    explain references in a way that doesn't refer to \"directories\" at all,\n    and instead talks about the \"hierarchy\" (from Kristoffer and Patrick's\n    reviews).\n    \n    Also:\n    \n     * objects: Mention that an object ID is called an \"object name\", and\n       update the glossary to include the term \"object ID\" (from Junio's\n       review)\n     * objects: Replace \"SHA-1 hash\" with \"cryptographic hash\" which is more\n       accurate (from Patrick's review)\n     * blobs: Made the explanation of git gc a little higher level and took\n       some ideas from Patrick's suggested wording (from Patrick's and\n       Kroftoffer's reviews)\n     * commits: Mention that tag objects and commits can optionally have\n       other fields. I didn't mention the GPG signature specifically, but\n       don't have any objections to adding it. (from Patrick and Junio's\n       reviews)\n     * commits: Remove one of the mentions of git gc, since it perhaps opens\n       up too much of a rabbit hole: \"how does git gc decide which commits\n       to clean up?\". (from Kristoffer's review)\n     * tag objects: Add an example of how a tag object is represented (from\n       user feedback on the draft)\n     * index: Use the term \"file mode\" instead of \"permissions\", and list\n       all allowed file modes (from Patrick's review)\n     * index: Use \"stage number\" instead of \"number\" for index entries (from\n       Patrick's review)\n     * reflogs: Remove \"any ref can be logged\", it raises some questions of\n       \"how do you tell Git to log a ref that it isn't normally logging?\"\n       and my guess is that it's uncommon to ask Git to log more refs. I\n       don't think it's a \"lie\" to omit this but I can bring it back if\n       folks disagree. (from Patrick's review)\n     * reflogs: Fix an error I noticed in the explanation of reflogs: tags\n       aren't logged by default and remote-tracking branches are, according\n       to man git-config\n     * branches and tags: Be clearer about how branches are usually updated\n       (by committing), and make it a little more obvious that only branches\n       can be checked out. This is a bit tricky because using the word\n       \"check out\" introduces a rabbit hole that I want to avoid (what does\n       \"check out\" mean?). I've dealt this by just talking about the\n       \"current branch\" (HEAD) since that is defined here, and making it\n       more explicit that HEAD must either be a branch or a commit, there's\n       no \"HEAD is a tag\" option. (from Patrick's review)\n     * tags: Explain the differences between annotated and lightweight tags\n       (this is the main piece of user feedback I've gotten on the draft so\n       far)\n     * Various style/typo changes (\"2 or more\", linkgit:git-gc[1], removed\n       extra asterisks, added empty SYNOPSIS, \"commits -> tags\" typo fix,\n       add to meson build)\n    \n    non-changes:\n    \n     * I still haven't mentioned things that aren't part of the \"data\n       model\", like revision params and configuration. I think there could\n       be a place for them but I haven't found it yet.\n     * tag objects: I noticed that there's a \"tag\" header field in tag\n       objects (like tag v1.0.0) but I didn't mention it yet because I\n       couldn't figure out what the purpose of that field is (I thought the\n       tag name was stored in the reference, why is it duplicated in the tag\n       object?)\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/1981\n\nRange-diff vs v1:\n\n 1:  fcbd21b6da ! 1:  3b38a88dc7 doc: add a explanation of Git's data model\n     @@ Commit message\n             down a rabbit hole which would make this document less useful for\n             understanding Git's core behaviour.\n          3. Don't discuss the structure of a commit message\n     -       (first line, trailers, GPG signatures, etc).\n     -       Perhaps this should change.\n     -\n     -    Some other choices I've made:\n     -\n     -    1. Mention packed refs only in a note.\n     -    2. Don't mention that the full name of the branch `main` is\n     -       technically `refs/heads/main`. This should likely change but I\n     -       haven't worked out how to do it in a clear way yet.\n     -    3. Mostly avoid referring to the `.git` directory, because the exact\n     -       details of how things are stored change over time.\n     -       This should perhaps change from \"mostly\" to \"entirely\"\n     -       but I haven't worked out how to do that in a clear way yet.\n     +       (first line, trailers etc).\n     +    4. Don't mention configuration.\n     +    5. Don't mention the `.git` directory, to avoid getting too much into\n     +       implementation details\n      \n          Signed-off-by: Julia Evans <julia@jvns.ca>\n      \n     @@ Documentation/gitdatamodel.adoc (new)\n      +----\n      +gitdatamodel - Git's core data model\n      +\n     ++SYNOPSIS\n     ++--------\n     ++gitdatamodel\n     ++\n      +DESCRIPTION\n      +-----------\n      +\n      +It's not necessary to understand Git's data model to use Git, but it's\n      +very helpful when reading Git's documentation so that you know what it\n     -+means when the documentation says \"object\" \"reference\" or \"index\".\n     ++means when the documentation says \"object\", \"reference\" or \"index\".\n      +\n      +Git's core operations use 4 kinds of data:\n      +\n     @@ Documentation/gitdatamodel.adoc (new)\n      +Commits, trees, blobs, and tag objects are all stored in Git's object database.\n      +Every object has:\n      +\n     -+1. an *ID*, which is the SHA-1 hash of its contents.\n     ++1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n     ++  type and contents.\n      +  It's fast to look up a Git object using its ID.\n     -+  The ID is usually represented in hexadecimal, like\n     ++  This is usually represented in hexadecimal, like\n      +  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n      +2. a *type*. There are 4 types of objects:\n      +   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n     @@ Documentation/gitdatamodel.adoc (new)\n      +\n      +[[commit]]\n      +commits::\n     -+    A commit contains:\n     ++    A commit contains these required fields\n     ++    (though there are other optional fields):\n      ++\n      +1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n     -+  regular commits have 1 parent, merge commits have 2+ parents\n     ++  regular commits have 1 parent, merge commits have 2 or more parents\n      +2. A *commit message*\n      +3. All the *files* in the commit, stored as a *<<tree,tree>>*\n      +4. An *author* and the time the commit was authored\n     @@ Documentation/gitdatamodel.adoc (new)\n      ++\n      +Like all other objects, commits can never be changed after they're created.\n      +For example, \"amending\" a commit with `git commit --amend` creates a new commit.\n     -+The old commit will eventually be deleted by `git gc`.\n      +\n      +[[tree]]\n      +trees::\n      +    A tree is how Git represents a directory. It lists, for each item in\n      +    the tree:\n      ++\n     -+1. The *permissions*, for example `100644`\n     ++1. The *file mode*, for example `100644`\n      +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n      +  or <<commit,`commit`>> (a Git submodule)\n      +3. The *object ID*\n     @@ Documentation/gitdatamodel.adoc (new)\n      +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n      +----\n      ++\n     -+*NOTE:* The permissions are in the same format as UNIX permissions, but\n     -+the only allowed permissions for files (blobs) are 644 and 755.\n     ++Git only supports these file modes:\n     +++\n     ++  - `100644`: regular file (with type `blob`)\n     ++  - `100755`: executable file (with type `blob`)\n     ++  - `120000`: symbolic link (with type `blob`)\n     ++  - `040000`: directory (with type `tree`)\n     ++  - `160000`: gitlink, for use with submodules (with type `commit`)\n      +\n      +[[blob]]\n      +blobs::\n      +    A blob is how Git represents a file. A blob object contains the\n      +    file's contents.\n      ++\n     -+Storing a new blob for every new version of a file can get big, so\n     -+`git gc` periodically compresses objects for efficiency in `.git/objects/pack`.\n     ++\n     ++NOTE: Storing a new blob for every new version of a file can use a\n     ++lot of disk space. To handle this, Git periodically runs repository\n     ++maintenance with linkgit:git-gc[1]. Part of this maintenance is\n     ++compressing objects so that if a small part of a file was changed, only\n     ++the change is stored instead of the whole file.\n      +\n      +[[tag-object]]\n      +tag objects::\n     -+    Tag objects (also known as \"annotated tags\") contain:\n     ++    Tag objects (also known as \"annotated tags\") contain these required fields\n     ++    (though there are other optional fields):\n      ++\n      +1. The *tagger* and tag date\n      +2. A *tag message*, similar to a commit message\n     -+3. The *ID* of the object (often a commit) that they reference\n     ++3. The *ID* and *type* of the object (often a commit) that they reference\n     ++\n     ++Here's how an example tag object is stored:\n     ++\n     ++----\n     ++object 750b4ead9c87ceb3ddb7a390e6c7074521797fb3\n     ++type commit\n     ++tag v1.0.0\n     ++tagger Maya <maya@example.com> 1759927359 -0400\n     ++\n     ++Release version 1.0.0\n     ++----\n      +\n      +[[references]]\n      +REFERENCES\n     @@ Documentation/gitdatamodel.adoc (new)\n      +branch\" than \"the changes are in commit bb69721404348e\".\n      +Git often uses \"ref\" as shorthand for \"reference\".\n      +\n     -+References that you create are stored in the `.git/refs` directory,\n     -+and Git has a few special internal references like `HEAD` that are stored\n     -+in the base `.git` directory.\n     -+\n      +References can either be:\n      +\n      +1. References to an object ID, usually a <<commit,commit>> ID\n      +2. References to another reference. This is called a \"symbolic reference\".\n      +\n     -+Git handles references differently based on which subdirectory of\n     -+`.git/refs` they're stored in.\n     -+Here are the main types:\n     ++References are stored in a hierarchy, and Git handles references\n     ++differently based on where they are in the hierarchy.\n     ++Most references are under `refs/`. Here are the main types:\n      +\n      +[[branch]]\n     -+branches: `.git/refs/heads/<name>`::\n     ++branches: `refs/heads/<name>`::\n      +    A branch is a name for a commit ID.\n      +    That commit is the latest commit on the branch.\n     -+    Branches are stored in the `.git/refs/heads/` directory.\n      ++\n      +To get the history of commits on a branch, Git will start at the commit\n      +ID the branch references, and then look at the commit's parent(s),\n      +the parent's parent, etc.\n      +\n      +[[tag]]\n     -+tags: `.git/refs/tags/<name>`::\n     ++tags: `refs/tags/<name>`::\n      +    A tag is a name for a commit ID, tag object ID, or other object ID.\n     -+    Tags are stored in the `refs/tags/` directory.\n     ++    Tags that reference a tag object ID are called \"annotated tags\",\n     ++    because the tag object contains a tag message.\n     ++    Tags that reference a commit ID, blob ID, or tree ID are\n     ++    called \"lightweight tags\".\n      ++\n     -+Even though branches and commits are both \"a name for a commit ID\", Git\n     ++Even though branches and tags are both \"a name for a commit ID\", Git\n      +treats them very differently.\n     -+Branches are expected to be regularly updated as you work on the branch,\n     -+but it's expected that a tag will never change after you create it.\n     ++Branches are expected to change over time: when you make a commit, Git\n     ++will update your <<HEAD,current branch>> to reference the new changes.\n     ++It's expected that a tag will never change after you create it.\n      +\n      +[[HEAD]]\n     -+HEAD: `.git/HEAD`::\n     ++HEAD: `HEAD`::\n      +    `HEAD` is where Git stores your current <<branch,branch>>.\n     -+    `HEAD` is normally a symbolic reference to your current branch, for\n     -+    example `ref: refs/heads/main` if your current branch is `main`.\n     -+    `HEAD` can also be a direct reference to a commit ID,\n     -+    that's called \"detached HEAD state\".\n     ++    `HEAD` can either be:\n     ++    1. A symbolic reference to your current branch, for example `ref:\n     ++       refs/heads/main` if your current branch is `main`.\n     ++    2. A direct reference to a commit ID.\n     ++        This is called \"detached HEAD state\".\n      +\n      +[[remote-tracking-branch]]\n     -+remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::\n     ++remote tracking branches: `refs/remotes/<remote>/<branch>`::\n      +    A remote-tracking branch is a name for a commit ID.\n      +    It's how Git stores the last-known state of a branch in a remote\n      +    repository. `git fetch` updates remote-tracking branches. When\n     @@ Documentation/gitdatamodel.adoc (new)\n      +\n      +[[other-refs]]\n      +Other references::\n     -+    Git tools may create references in any subdirectory of `.git/refs`.\n     ++    Git tools may create references anywhere under `refs/`.\n      +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n      +    and linkgit:git-notes[1] all create their own references\n     -+    in `.git/refs/stash`, `.git/refs/bisect`, etc.\n     ++    in `refs/stash`, `refs/bisect`, etc.\n      +    Third-party Git tools may also create their own references.\n      ++\n     -+Git may also create references in the base `.git` directory\n     -+other than `HEAD`, like `ORIG_HEAD`.\n     -+\n     -+*NOTE:* As an optimization, references may be stored as packed\n     -+refs instead of in `.git/refs`. See linkgit:git-pack-refs[1].\n     ++Git may also create references other than `HEAD` at the base of the\n     ++hierarchy, like `ORIG_HEAD`.\n      +\n      +[[index]]\n      +THE INDEX\n     @@ Documentation/gitdatamodel.adoc (new)\n      +1. The *permissions*\n      +2. The *<<blob,blob>> ID* of the file\n      +3. The *filename*\n     -+4. The *number*. This is normally 0, but if there's a merge conflict\n     ++4. The *stage number*. This is normally 0, but if there's a merge conflict\n      +   there can be multiple versions (with numbers 0, 1, 2, ..)\n      +   of the same filename in the index.\n      +\n     @@ Documentation/gitdatamodel.adoc (new)\n      +REFLOGS\n      +-------\n      +\n     -+Git stores the history of branch, tag, and HEAD refs in a reflog\n     -+(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n     -+default, but any ref can be logged.\n     ++Git stores the history of your branch, remote-tracking branch, and HEAD refs\n     ++in a reflog (you should read \"reflog\" as \"ref log\").\n      +\n      +Each reflog entry has:\n      +\n     -+1. *Before/after *commit IDs*\n     ++1. Before/after *commit IDs*\n      +2. *User* who made the change, for example `Maya <maya@example.com>`\n     -+3. *Timestamp*\n     ++3. *Timestamp* when the change was made\n      +4. *Log message*, for example `pull: Fast-forward`\n      +\n      +Reflogs only log changes made in your local repository.\n     @@ Documentation/gitdatamodel.adoc (new)\n      +GIT\n      +---\n      +Part of the linkgit:git[1] suite\n     +\n     + ## Documentation/glossary-content.adoc ##\n     +@@ Documentation/glossary-content.adoc: This commit is referred to as a \"merge commit\", or sometimes just a\n     + \tidentified by its <<def_object_name,object name>>. The objects usually\n     + \tlive in `$GIT_DIR/objects/`.\n     + \n     +-[[def_object_identifier]]object identifier (oid)::\n     +-\tSynonym for <<def_object_name,object name>>.\n     ++[[def_object_identifier]]object identifier, object ID, oid::\n     ++\tSynonyms for <<def_object_name,object name>>.\n     + \n     + [[def_object_name]]object name::\n     + \tThe unique identifier of an <<def_object,object>>.  The\n     +\n     + ## Documentation/meson.build ##\n     +@@ Documentation/meson.build: manpages = {\n     +   'gitcore-tutorial.adoc' : 7,\n     +   'gitcredentials.adoc' : 7,\n     +   'gitcvs-migration.adoc' : 7,\n     ++  'gitdatamodel.adoc' : 7,\n     +   'gitdiffcore.adoc' : 7,\n     +   'giteveryday.adoc' : 7,\n     +   'gitfaq.adoc' : 7,\n\n\n Documentation/Makefile              |   1 +\n Documentation/gitdatamodel.adoc     | 248 ++++++++++++++++++++++++++++\n Documentation/glossary-content.adoc |   4 +-\n Documentation/meson.build           |   1 +\n 4 files changed, 252 insertions(+), 2 deletions(-)\n create mode 100644 Documentation/gitdatamodel.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 6fb83d0c6e..5f4acfacbd 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n MAN7_TXT += gitcore-tutorial.adoc\n MAN7_TXT += gitcredentials.adoc\n MAN7_TXT += gitcvs-migration.adoc\n+MAN7_TXT += gitdatamodel.adoc\n MAN7_TXT += gitdiffcore.adoc\n MAN7_TXT += giteveryday.adoc\n MAN7_TXT += gitfaq.adoc\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nnew file mode 100644\nindex 0000000000..c3a25ea8d2\n--- /dev/null\n+++ b/Documentation/gitdatamodel.adoc\n@@ -0,0 +1,248 @@\n+gitdatamodel(7)\n+===============\n+\n+NAME\n+----\n+gitdatamodel - Git's core data model\n+\n+SYNOPSIS\n+--------\n+gitdatamodel\n+\n+DESCRIPTION\n+-----------\n+\n+It's not necessary to understand Git's data model to use Git, but it's\n+very helpful when reading Git's documentation so that you know what it\n+means when the documentation says \"object\", \"reference\" or \"index\".\n+\n+Git's core operations use 4 kinds of data:\n+\n+1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+2. <<references,References>>: branches, tags,\n+   remote-tracking branches, etc\n+3. <<index,The index>>, also known as the staging area\n+4. <<reflogs,Reflogs>>\n+\n+[[objects]]\n+OBJECTS\n+-------\n+\n+Commits, trees, blobs, and tag objects are all stored in Git's object database.\n+Every object has:\n+\n+1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n+  type and contents.\n+  It's fast to look up a Git object using its ID.\n+  This is usually represented in hexadecimal, like\n+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+2. a *type*. There are 4 types of objects:\n+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n+   and <<tag-object,tag objects>>.\n+3. *contents*. The structure of the contents depends on the type.\n+\n+Once an object is created, it can never be changed.\n+Here are the 4 types of objects:\n+\n+[[commit]]\n+commits::\n+    A commit contains these required fields\n+    (though there are other optional fields):\n++\n+1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n+  regular commits have 1 parent, merge commits have 2 or more parents\n+2. A *commit message*\n+3. All the *files* in the commit, stored as a *<<tree,tree>>*\n+4. An *author* and the time the commit was authored\n+5. A *committer* and the time the commit was committed\n++\n+Here's how an example commit is stored:\n++\n+----\n+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n+author Maya <maya@example.com> 1759173425 -0400\n+committer Maya <maya@example.com> 1759173425 -0400\n+\n+Add README\n+----\n++\n+Like all other objects, commits can never be changed after they're created.\n+For example, \"amending\" a commit with `git commit --amend` creates a new commit.\n+\n+[[tree]]\n+trees::\n+    A tree is how Git represents a directory. It lists, for each item in\n+    the tree:\n++\n+1. The *file mode*, for example `100644`\n+2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n+  or <<commit,`commit`>> (a Git submodule)\n+3. The *object ID*\n+4. The *filename*\n++\n+For example, this is how a tree containing one directory (`src`) and one file\n+(`README.md`) is stored:\n++\n+----\n+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n+----\n++\n+Git only supports these file modes:\n++\n+  - `100644`: regular file (with type `blob`)\n+  - `100755`: executable file (with type `blob`)\n+  - `120000`: symbolic link (with type `blob`)\n+  - `040000`: directory (with type `tree`)\n+  - `160000`: gitlink, for use with submodules (with type `commit`)\n+\n+[[blob]]\n+blobs::\n+    A blob is how Git represents a file. A blob object contains the\n+    file's contents.\n++\n+\n+NOTE: Storing a new blob for every new version of a file can use a\n+lot of disk space. To handle this, Git periodically runs repository\n+maintenance with linkgit:git-gc[1]. Part of this maintenance is\n+compressing objects so that if a small part of a file was changed, only\n+the change is stored instead of the whole file.\n+\n+[[tag-object]]\n+tag objects::\n+    Tag objects (also known as \"annotated tags\") contain these required fields\n+    (though there are other optional fields):\n++\n+1. The *tagger* and tag date\n+2. A *tag message*, similar to a commit message\n+3. The *ID* and *type* of the object (often a commit) that they reference\n+\n+Here's how an example tag object is stored:\n+\n+----\n+object 750b4ead9c87ceb3ddb7a390e6c7074521797fb3\n+type commit\n+tag v1.0.0\n+tagger Maya <maya@example.com> 1759927359 -0400\n+\n+Release version 1.0.0\n+----\n+\n+[[references]]\n+REFERENCES\n+----------\n+\n+References are a way to give a name to a commit.\n+It's easier to remember \"the changes I'm working on are on the `turtle`\n+branch\" than \"the changes are in commit bb69721404348e\".\n+Git often uses \"ref\" as shorthand for \"reference\".\n+\n+References can either be:\n+\n+1. References to an object ID, usually a <<commit,commit>> ID\n+2. References to another reference. This is called a \"symbolic reference\".\n+\n+References are stored in a hierarchy, and Git handles references\n+differently based on where they are in the hierarchy.\n+Most references are under `refs/`. Here are the main types:\n+\n+[[branch]]\n+branches: `refs/heads/<name>`::\n+    A branch is a name for a commit ID.\n+    That commit is the latest commit on the branch.\n++\n+To get the history of commits on a branch, Git will start at the commit\n+ID the branch references, and then look at the commit's parent(s),\n+the parent's parent, etc.\n+\n+[[tag]]\n+tags: `refs/tags/<name>`::\n+    A tag is a name for a commit ID, tag object ID, or other object ID.\n+    Tags that reference a tag object ID are called \"annotated tags\",\n+    because the tag object contains a tag message.\n+    Tags that reference a commit ID, blob ID, or tree ID are\n+    called \"lightweight tags\".\n++\n+Even though branches and tags are both \"a name for a commit ID\", Git\n+treats them very differently.\n+Branches are expected to change over time: when you make a commit, Git\n+will update your <<HEAD,current branch>> to reference the new changes.\n+It's expected that a tag will never change after you create it.\n+\n+[[HEAD]]\n+HEAD: `HEAD`::\n+    `HEAD` is where Git stores your current <<branch,branch>>.\n+    `HEAD` can either be:\n+    1. A symbolic reference to your current branch, for example `ref:\n+       refs/heads/main` if your current branch is `main`.\n+    2. A direct reference to a commit ID.\n+        This is called \"detached HEAD state\".\n+\n+[[remote-tracking-branch]]\n+remote tracking branches: `refs/remotes/<remote>/<branch>`::\n+    A remote-tracking branch is a name for a commit ID.\n+    It's how Git stores the last-known state of a branch in a remote\n+    repository. `git fetch` updates remote-tracking branches. When\n+    `git status` says \"you're up to date with origin/main\", it's looking at\n+    this.\n+\n+[[other-refs]]\n+Other references::\n+    Git tools may create references anywhere under `refs/`.\n+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n+    and linkgit:git-notes[1] all create their own references\n+    in `refs/stash`, `refs/bisect`, etc.\n+    Third-party Git tools may also create their own references.\n++\n+Git may also create references other than `HEAD` at the base of the\n+hierarchy, like `ORIG_HEAD`.\n+\n+[[index]]\n+THE INDEX\n+---------\n+\n+The index, also known as the \"staging area\", contains the current staged\n+version of every file in your Git repository. When you commit, the files\n+in the index are used as the files in the next commit.\n+\n+Unlike a tree, the index is a flat list of files.\n+Each index entry has 4 fields:\n+\n+1. The *permissions*\n+2. The *<<blob,blob>> ID* of the file\n+3. The *filename*\n+4. The *stage number*. This is normally 0, but if there's a merge conflict\n+   there can be multiple versions (with numbers 0, 1, 2, ..)\n+   of the same filename in the index.\n+\n+It's extremely uncommon to look at the index directly: normally you'd\n+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n+But you can use `git ls-files --stage` to see the index.\n+Here's the output of `git ls-files --stage` in a repository with 2 files:\n+\n+----\n+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n+----\n+\n+[[reflogs]]\n+REFLOGS\n+-------\n+\n+Git stores the history of your branch, remote-tracking branch, and HEAD refs\n+in a reflog (you should read \"reflog\" as \"ref log\").\n+\n+Each reflog entry has:\n+\n+1. Before/after *commit IDs*\n+2. *User* who made the change, for example `Maya <maya@example.com>`\n+3. *Timestamp* when the change was made\n+4. *Log message*, for example `pull: Fast-forward`\n+\n+Reflogs only log changes made in your local repository.\n+They are not shared with remotes.\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\ndiff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc\nindex e423e4765b..20ba121314 100644\n--- a/Documentation/glossary-content.adoc\n+++ b/Documentation/glossary-content.adoc\n@@ -297,8 +297,8 @@ This commit is referred to as a \"merge commit\", or sometimes just a\n \tidentified by its <<def_object_name,object name>>. The objects usually\n \tlive in `$GIT_DIR/objects/`.\n \n-[[def_object_identifier]]object identifier (oid)::\n-\tSynonym for <<def_object_name,object name>>.\n+[[def_object_identifier]]object identifier, object ID, oid::\n+\tSynonyms for <<def_object_name,object name>>.\n \n [[def_object_name]]object name::\n \tThe unique identifier of an <<def_object,object>>.  The\ndiff --git a/Documentation/meson.build b/Documentation/meson.build\nindex e34965c5b0..ace0573e82 100644\n--- a/Documentation/meson.build\n+++ b/Documentation/meson.build\n@@ -192,6 +192,7 @@ manpages = {\n   'gitcore-tutorial.adoc' : 7,\n   'gitcredentials.adoc' : 7,\n   'gitcvs-migration.adoc' : 7,\n+  'gitdatamodel.adoc' : 7,\n   'gitdiffcore.adoc' : 7,\n   'giteveryday.adoc' : 7,\n   'gitfaq.adoc' : 7,\n\nbase-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n-- \ngitgitgadget\n"},{"id":"528275","messageId":"xmqqecrdgzk8.fsf@gitster.g","threadId":"64244","inReplyTo":"aOXmA5L5LsUuXWEh@pks.im","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-08T15:53:27Z","receivedAt":"2025-10-08T15:53:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> I sometimes hear from users that \"commits can't be snapshots\", because\n>> it would take up too much disk space to store every version of\n>> every commit. So I find that sometimes explaining a little bit about the\n>> implementation can make the information more memorable.\n>> \n>> Certainly I'm not able to remember details that don't make sense\n>> with my mental model of how computers work and I don't expect other\n>> people to either, so I think it's important to give an explanation that\n>> handles the biggest \"objections\".\n>\n> Hm, fair I guess. In any case, if we want to mention this I'd leave away\n> the details how exactly Git achieves this. E.g. we could say something\n> like:\n>\n>     Storing a new blob for every new version of a file can result to a\n>     lot of duplication. Git regularly runs repository maintenance to\n>     optimize to counteract this. Part of the maintenance involves\n>     compression of objects, where incremental changes to the same object\n>     are optimized to be stored as deltas, only.\n>\n> We skip over the details, but this should give enough pointers to an\n> interested reader to go dig deeper. We could also generalize this to\n> objects in general, not only blobs.\n\nInteresting.  It is of course not wrong at all, but it was not what\nI would have expected for the first explanation to help confused\nfolks who say \"commits cannot be snapshots as they take too much\nspace\".\n\nTo me, it was a realization that even in a project whose tree (think\nof \"du -s .\")  is huge, each of its commits touches only a handful\nof paths, hence a large portion of that huge tree would be shared\nwith the previous snapshot.\n\n>> > This misses \"refs/remotes/<remote>/HEAD\". This reference is a symbolic\n>> > reference that indicates the default branch on the remote side.\n>> \n>> Is \"refs/remotes/<remote>/HEAD\" a remote-tracking branch?\n>> I've never thought about that reference and I'm not sure what to call it.\n>\n> No, it's not. I think the term we use is \"remote reference\".\n\nHonestly I didn't know/think we have any special terminology for the\nrefs/remotes/*/HEAD symref.\n\nHistorically HEAD did not \"track\" the remote state, and we did take\nadvantage of that fact to use it as a place to record the preference\nwith respect to which remote-tracking branch we would want to\nprimarily interact with.\n\nBut these days because the protocol is capable of expressing where\nthe symrefs point at, the users can make it track just like all\nother refs inside refs/remotes/*/ hiearchy.  So I personally think\nit is OK to call it in remote-tracking branch.\n\nThanks.\n"},{"id":"528290","messageId":"395232da-4ac6-4311-ae44-2bbf92fa6d2f@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqecrdgzk8.fsf@gitster.g","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-08T19:06:29Z","receivedAt":"2025-10-08T19:06:54Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Wed, Oct 8, 2025, at 11:53 AM, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n>\n>>> I sometimes hear from users that \"commits can't be snapshots\", because\n>>> it would take up too much disk space to store every version of\n>>> every commit. So I find that sometimes explaining a little bit about the\n>>> implementation can make the information more memorable.\n>>> \n>>> Certainly I'm not able to remember details that don't make sense\n>>> with my mental model of how computers work and I don't expect other\n>>> people to either, so I think it's important to give an explanation that\n>>> handles the biggest \"objections\".\n>>\n>> Hm, fair I guess. In any case, if we want to mention this I'd leave away\n>> the details how exactly Git achieves this. E.g. we could say something\n>> like:\n>>\n>>     Storing a new blob for every new version of a file can result to a\n>>     lot of duplication. Git regularly runs repository maintenance to\n>>     optimize to counteract this. Part of the maintenance involves\n>>     compression of objects, where incremental changes to the same object\n>>     are optimized to be stored as deltas, only.\n>>\n>> We skip over the details, but this should give enough pointers to an\n>> interested reader to go dig deeper. We could also generalize this to\n>> objects in general, not only blobs.\n>\n> Interesting.  It is of course not wrong at all, but it was not what\n> I would have expected for the first explanation to help confused\n> folks who say \"commits cannot be snapshots as they take too much\n> space\".\n>\n> To me, it was a realization that even in a project whose tree (think\n> of \"du -s .\")  is huge, each of its commits touches only a handful\n> of paths, hence a large portion of that huge tree would be shared\n> with the previous snapshot.\n\nThat's a good point, I forgot that I've explained it that way too.\nI might change it to that instead. \n\n>>> > This misses \"refs/remotes/<remote>/HEAD\". This reference is a symbolic\n>>> > reference that indicates the default branch on the remote side.\n>>> \n>>> Is \"refs/remotes/<remote>/HEAD\" a remote-tracking branch?\n>>> I've never thought about that reference and I'm not sure what to call it.\n>>\n>> No, it's not. I think the term we use is \"remote reference\".\n>\n> Honestly I didn't know/think we have any special terminology for the\n> refs/remotes/*/HEAD symref.\n>\n> Historically HEAD did not \"track\" the remote state, and we did take\n> advantage of that fact to use it as a place to record the preference\n> with respect to which remote-tracking branch we would want to\n> primarily interact with.\n>\n> But these days because the protocol is capable of expressing where\n> the symrefs point at, the users can make it track just like all\n> other refs inside refs/remotes/*/ hiearchy.  So I personally think\n> it is OK to call it in remote-tracking branch.\n\nI may just add this to the remote-tracking branch sentence then,\nwhich is hopefully correct:\n\n`refs/remotes/<remote>/HEAD` is a symbolic reference to the remote's\ndefault branch. This is the branch that `git clone` checks out by default.\n"},{"id":"528390","messageId":"d1bd63d4-3ac7-458c-86e0-a4f7aac300d9@app.fastmail.com","threadId":"64244","inReplyTo":"CALnO6CCsGtjcWBkjV0vsJHDCwiwt9eO2CsA1zFgwFiwJ-KLhew@mail.gmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-09T13:20:53Z","receivedAt":"2025-10-09T13:21:14Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":">> >> I think I'll see if I can figure out a way to mention this and at the\n>> >> same time remove most of the rest of the references to the `.git`\n>> >> directory when explaining references (which you talked about\n>> >> further down), including packed refs.\n>> >\n>> > A colleague will be explaining reflog for an audience tomorrow, and\n>> > decided to briefly explain refs, too—which tells me this is\n>> > much-needed.\n>> >\n>> > For refs themselves, perhaps \"git for-each-ref\" is a reasonable place\n>> > to start? Since it tells you the refs you have and how to spell them\n>> > explicitly regardless of how they are stored?\n>>\n>> Interesting, do you use git for-each-ref?\n>> What do you use it for?\n>\n> Ah, yes, but primarily for scripting.\n>\n> What I should have clarified is that \"the tool (I know of) to\n> interrogate the refs you currently have is git-for-each-ref\" (like how\n> git-ls-remote is the tool to interrogate a remote's refs). It avoids\n> the issues with assuming \"tree .git/refs\" or similar will capture the\n> actual data.\n\nAh, that makes sense! I spent a little while trying to come up with\nsomething that would give a \"similar result\" to running\n`cat .git/<refname>` and I came up with this:\n\ngit for-each-ref <ref-name> --include-root-refs  --format=\"%(refname)   %(if)%(symref)%(then)%(symref)%(else)%(objectname:short)%(end)\"\n\nI hoped to find a simple equivalent to that `cat` command\n(kind of the equivalent of `git cat-file -p`) that would work with\nother ref backends but couldn't find one.\n"},{"id":"528392","messageId":"b2a9b8ca-8f2a-40f0-a724-0da707902985@app.fastmail.com","threadId":"64244","inReplyTo":"pull.1981.git.1759512876284.gitgitgadget@gmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-09T14:20:14Z","receivedAt":"2025-10-09T14:20:46Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"I collected some feedback from Git users on this v2 document. I'm expecting more\nfeedback, but here's an initial brain dump of my notes. I mostly wrote this for\nmy own use but I thought it might be interesting to other folks too.\n\nintro:\n- Say that we're going to explain what \"objects\", \"references\" etc are.\n  (so that readers know they're not expected to know what those words mean yet)\n- It's confusing that tags are both \"an object\" and \"a reference\".\n  Need to think about whether there's a way to address this, I was hoping\n  that using the terms \"tag object\" and \"tag\" would be enough but maybe not.\n- Give a 1-sentence intro to \"reflog\" (easy)\n\ncommits:\n- The order the fields are given in don't match the order in the\n  example, maybe they should.\n- \"All the files in the commit, stored as a tree\" is throwing a few\n  people off. I think we should communicate something like \"a tree hash\n  that describes the root of the project and then by extension the whole\n  project\", but phrased more clearly. Will figure that out.\n- \"Are commits stored as a diff?\" (2 people asked where diffs come\n  from, I think we need to add a note saying that diffs are calculated\n  at runtime, it's a very common misconception and I think it should\n  be easy to clear up)\n- \"What's the difference between an author and committer?\n  (I actually don't know either, will try to find out and see if it's\n  straightforward to add a short note explaining it)\n- In the note about commits being amended: one person suggested saying\n  \"creates a new commit with the same parent\" which I think might be clearer.\n\ntrees:\n- One person asks what a \"working tree\" is.\n  I don't think this is a good place for that, but it made me wonder if\n  \"the current working directory\" has a place in this document.\n  I feel like no but not 100% sure.\n- 2 people want to know more about \"The file mode, for example 100644\".\n  Moving \"Git only supports these file modes...\" further up so that folks\n  can immediately see what the options are here should help with this.\n- On \"so git-gc(1) periodically compresses objects to save disk space\",\n  there are a few follow up comments wondering about more, which makes me\n  think the comment about compression is actually a distraction.\n  I'll say something simpler instead, from Junio's suggestion.\n\n\ntag objects:\n- Requests for an example, will add one.\n- Requests to explain the difference between\n  \"lightweight\" and \"annotated\" tags, will add.\n\nreferences:\n\n- Two people pointed out that because references are often stored as files,\n  you can't have two references named `julia/ticket-number` and\n  `julia/ticket-number/task-name`.\n  I'm not sure if this is a fundamental limit of the refs data model\n  (does the reftable backend have the same limitation?), but it could be\n  a good reason to mention that refs are often stored as files, because\n  it makes it obvious that you can't have a file and a directory with\n  the same name.\n  Obviously this is an issue that is affecting people relatively often\n  in practice though so I think it's worth mentioning in some way.\n\n\nbranches:\n\nHEAD:\n- One person asks if there are any other symbolic references other than HEAD,\n  or if they can make their own symbolic references. I don't know and I don't\n  know if this is worth mentioning.\n- `HEAD: HEAD` looks weird, it made sense when it was `HEAD: .git/HEAD`.\n  Will think about how to fix this.\n- Several people are asking for more detail about detached HEAD state.\n  My current idea here is to just give an example of a way you can end up\n  in detached HEAD state. (\"by checking out a tag\"), but in an ideal world\n  it would be easy to find out what it means, how it happens, what it implies,\n  and how you might adjust your workflow to avoid it (by using `git switch`).\n  But we can't get into all of that here.\n  I'd love to just link to a more detailed explanation of detached HEAD state\n  but I'm not totally satisfied with the one that's currently in `man git-checkout`.\n  It may be best to just leave this in a slightly suboptimal state, write\n  a really clear explanation of detached HEAD state somewhere else, and\n  then link to it.\n\nthe index:\n- \"permissions\" should be \"file mode\" (like with trees)\n- \"filename\" should be \"file path\"\n- The index can also be locked. Might be worth mentioning.\n- This doesn't explain what \"staged\" means, perhaps mention\n  the relationship to `git add`\n\nreflogs\n- Mention the role of the reflog in retrieving \"lost\" commits or\n  undoing bad rebases.\n- How can you see the full data in the reflog?\n  `git reflog show` doesn't list the user who made the change\n  git reflog show <refname> --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso\n  works but it's really a mouthful, not sure I want to include all that\n\nOverall: several people suggested mentioning more about where things\nare stored in the `.git` directory, which I just removed.\n\nI think I want to avoid this (not sure yet), but I'm going to think\nabout the underlying motivation for this suggestion and see if it can be\naddressed in a different way.\n\nSome ideas for what functions discussing the `.git` directory has:\n\n1. Like I mentioned above with branches, sometimes the implementation causes\n   some extra constraints like \"you can't have branches `julia/ticket`\n   and `julia/ticket/task`\". So often people like to know a little\n   about the implementation because it can help predict some of the\n   holes in the abstractions you're using.\n2. It lets you view the \"raw\" data, so you can be totally sure about\n   what Git is storing. This is nice because Git's UI can be very\n   inconsistent sometimes, so looking at the raw data gives a sense of\n   certainty about what's actually there.\n\nI tried to put together a list of ways to look at the \"raw\" data without\nlooking in the `.git` directory. The ways for objects and the index are great,\nbut for references and the reflog they involve these pretty complex format\nstrings, I'm not confident I've gotten the format strings right and IMO\nthey don't inspire a lot of confidence.\n\nView an object with:\n----\ngit cat-file -p <object-id>\n----\n\nView a reference with:\n\n----\ngit for-each-ref <ref-name> --include-root-refs  --format=\"%(refname) %(if)%(symref)%(then)%(symref)%(else)%(objectname:short)%(end)\"\n----\n\nView the index with:\n\n----\ngit ls-files --stage\n----\n\nView the reflog for a reference with:\n\n----\ngit reflog show <refname> --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso\n----\n\n\nOn Fri, Oct 3, 2025, at 1:34 PM, Julia Evans via GitGitGadget wrote:\n> From: Julia Evans <julia@jvns.ca>\n>\n> Git very often uses the terms \"object\", \"reference\", or \"index\" in its\n> documentation.\n>\n> However, it's hard to find a clear explanation of these terms and how\n> they relate to each other in the documentation. The closest candidates\n> currently are:\n>\n> 1. `gitglossary`. This makes a good effort, but it's an alphabetically\n>     ordered dictionary and a dictionary is not a good way to learn\n>     concepts. You have to jump around too much and it's not possible to\n>     present the concepts in the order that they should be explained.\n> 2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n>    This is a nice document to have, but it's not necessary to learn how\n>    `update-index` works to understand Git's data model, and we should\n>    not be requiring users to learn how to use the \"plumbing\" commands\n>    if they want to learn what the term \"index\" or \"object\" means.\n> 3. `gitrepository-layout`. This is a great resource, but it includes a\n>    lot of information about configuration and internal implementation\n>    details which are not related to the data model. It also does\n>    not explain how commits work.\n>\n> The result of this is that Git users (even users who have been using\n> Git for 15+ years) struggle to read the documentation because they don't\n> know what the core terms mean, and it's not possible to add links\n> to help them learn more.\n>\n> Add an explanation of Git's data model. Some choices I've made in\n> deciding what \"core data model\" means:\n>\n> 1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n>    if those are intended to be user facing or if they're more like\n>    internal implementation details.\n> 2. Don't talk about submodules other than by mentioning how they\n>    relate to trees. This is because Git has a lot of special features,\n>    and explaining how they all work exhaustively could quickly go\n>    down a rabbit hole which would make this document less useful for\n>    understanding Git's core behaviour.\n> 3. Don't discuss the structure of a commit message\n>    (first line, trailers, GPG signatures, etc).\n>    Perhaps this should change.\n>\n> Some other choices I've made:\n>\n> 1. Mention packed refs only in a note.\n> 2. Don't mention that the full name of the branch `main` is\n>    technically `refs/heads/main`. This should likely change but I\n>    haven't worked out how to do it in a clear way yet.\n> 3. Mostly avoid referring to the `.git` directory, because the exact\n>    details of how things are stored change over time.\n>    This should perhaps change from \"mostly\" to \"entirely\"\n>    but I haven't worked out how to do that in a clear way yet.\n>\n> Signed-off-by: Julia Evans <julia@jvns.ca>\n> ---\n>     doc: Add a explanation of Git's data model\n>\n> Published-As: \n> https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git \n> pr-1981/jvns/gitdatamodel-v1\n> Pull-Request: https://github.com/gitgitgadget/git/pull/1981\n>\n>  Documentation/Makefile          |   1 +\n>  Documentation/gitdatamodel.adoc | 226 ++++++++++++++++++++++++++++++++\n>  2 files changed, 227 insertions(+)\n>  create mode 100644 Documentation/gitdatamodel.adoc\n>\n> diff --git a/Documentation/Makefile b/Documentation/Makefile\n> index 6fb83d0c6e..5f4acfacbd 100644\n> --- a/Documentation/Makefile\n> +++ b/Documentation/Makefile\n> @@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n>  MAN7_TXT += gitcore-tutorial.adoc\n>  MAN7_TXT += gitcredentials.adoc\n>  MAN7_TXT += gitcvs-migration.adoc\n> +MAN7_TXT += gitdatamodel.adoc\n>  MAN7_TXT += gitdiffcore.adoc\n>  MAN7_TXT += giteveryday.adoc\n>  MAN7_TXT += gitfaq.adoc\n> diff --git a/Documentation/gitdatamodel.adoc \n> b/Documentation/gitdatamodel.adoc\n> new file mode 100644\n> index 0000000000..4b2cb167dc\n> --- /dev/null\n> +++ b/Documentation/gitdatamodel.adoc\n> @@ -0,0 +1,226 @@\n> +gitdatamodel(7)\n> +===============\n> +\n> +NAME\n> +----\n> +gitdatamodel - Git's core data model\n> +\n> +DESCRIPTION\n> +-----------\n> +\n> +It's not necessary to understand Git's data model to use Git, but it's\n> +very helpful when reading Git's documentation so that you know what it\n> +means when the documentation says \"object\" \"reference\" or \"index\".\n> +\n> +Git's core operations use 4 kinds of data:\n> +\n> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n> +2. <<references,References>>: branches, tags,\n> +   remote-tracking branches, etc\n> +3. <<index,The index>>, also known as the staging area\n> +4. <<reflogs,Reflogs>>\n> +\n> +[[objects]]\n> +OBJECTS\n> +-------\n> +\n> +Commits, trees, blobs, and tag objects are all stored in Git's object \n> database.\n> +Every object has:\n> +\n> +1. an *ID*, which is the SHA-1 hash of its contents.\n> +  It's fast to look up a Git object using its ID.\n> +  The ID is usually represented in hexadecimal, like\n> +  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n> +2. a *type*. There are 4 types of objects:\n> +   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n> +   and <<tag-object,tag objects>>.\n> +3. *contents*. The structure of the contents depends on the type.\n> +\n> +Once an object is created, it can never be changed.\n> +Here are the 4 types of objects:\n> +\n> +[[commit]]\n> +commits::\n> +    A commit contains:\n> ++\n> +1. Its *parent commit ID(s)*. The first commit in a repository has 0 \n> parents,\n> +  regular commits have 1 parent, merge commits have 2+ parents\n> +2. A *commit message*\n> +3. All the *files* in the commit, stored as a *<<tree,tree>>*\n> +4. An *author* and the time the commit was authored\n> +5. A *committer* and the time the commit was committed\n> ++\n> +Here's how an example commit is stored:\n> ++\n> +----\n> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n> +author Maya <maya@example.com> 1759173425 -0400\n> +committer Maya <maya@example.com> 1759173425 -0400\n> +\n> +Add README\n> +----\n> ++\n> +Like all other objects, commits can never be changed after they're \n> created.\n> +For example, \"amending\" a commit with `git commit --amend` creates a \n> new commit.\n> +The old commit will eventually be deleted by `git gc`.\n> +\n> +[[tree]]\n> +trees::\n> +    A tree is how Git represents a directory. It lists, for each item \n> in\n> +    the tree:\n> ++\n> +1. The *permissions*, for example `100644`\n> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n> +  or <<commit,`commit`>> (a Git submodule)\n> +3. The *object ID*\n> +4. The *filename*\n> ++\n> +For example, this is how a tree containing one directory (`src`) and \n> one file\n> +(`README.md`) is stored:\n> ++\n> +----\n> +100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n> +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n> +----\n> ++\n> +*NOTE:* The permissions are in the same format as UNIX permissions, but\n> +the only allowed permissions for files (blobs) are 644 and 755.\n> +\n> +[[blob]]\n> +blobs::\n> +    A blob is how Git represents a file. A blob object contains the\n> +    file's contents.\n> ++\n> +Storing a new blob for every new version of a file can get big, so\n> +`git gc` periodically compresses objects for efficiency in \n> `.git/objects/pack`.\n> +\n> +[[tag-object]]\n> +tag objects::\n> +    Tag objects (also known as \"annotated tags\") contain:\n> ++\n> +1. The *tagger* and tag date\n> +2. A *tag message*, similar to a commit message\n> +3. The *ID* of the object (often a commit) that they reference\n> +\n> +[[references]]\n> +REFERENCES\n> +----------\n> +\n> +References are a way to give a name to a commit.\n> +It's easier to remember \"the changes I'm working on are on the `turtle`\n> +branch\" than \"the changes are in commit bb69721404348e\".\n> +Git often uses \"ref\" as shorthand for \"reference\".\n> +\n> +References that you create are stored in the `.git/refs` directory,\n> +and Git has a few special internal references like `HEAD` that are \n> stored\n> +in the base `.git` directory.\n> +\n> +References can either be:\n> +\n> +1. References to an object ID, usually a <<commit,commit>> ID\n> +2. References to another reference. This is called a \"symbolic \n> reference\".\n> +\n> +Git handles references differently based on which subdirectory of\n> +`.git/refs` they're stored in.\n> +Here are the main types:\n> +\n> +[[branch]]\n> +branches: `.git/refs/heads/<name>`::\n> +    A branch is a name for a commit ID.\n> +    That commit is the latest commit on the branch.\n> +    Branches are stored in the `.git/refs/heads/` directory.\n> ++\n> +To get the history of commits on a branch, Git will start at the commit\n> +ID the branch references, and then look at the commit's parent(s),\n> +the parent's parent, etc.\n> +\n> +[[tag]]\n> +tags: `.git/refs/tags/<name>`::\n> +    A tag is a name for a commit ID, tag object ID, or other object ID.\n> +    Tags are stored in the `refs/tags/` directory.\n> ++\n> +Even though branches and commits are both \"a name for a commit ID\", Git\n> +treats them very differently.\n> +Branches are expected to be regularly updated as you work on the \n> branch,\n> +but it's expected that a tag will never change after you create it.\n> +\n> +[[HEAD]]\n> +HEAD: `.git/HEAD`::\n> +    `HEAD` is where Git stores your current <<branch,branch>>.\n> +    `HEAD` is normally a symbolic reference to your current branch, for\n> +    example `ref: refs/heads/main` if your current branch is `main`.\n> +    `HEAD` can also be a direct reference to a commit ID,\n> +    that's called \"detached HEAD state\".\n> +\n> +[[remote-tracking-branch]]\n> +remote tracking branches: `.git/refs/remotes/<remote>/<branch>`::\n> +    A remote-tracking branch is a name for a commit ID.\n> +    It's how Git stores the last-known state of a branch in a remote\n> +    repository. `git fetch` updates remote-tracking branches. When\n> +    `git status` says \"you're up to date with origin/main\", it's \n> looking at\n> +    this.\n> +\n> +[[other-refs]]\n> +Other references::\n> +    Git tools may create references in any subdirectory of `.git/refs`.\n> +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n> +    and linkgit:git-notes[1] all create their own references\n> +    in `.git/refs/stash`, `.git/refs/bisect`, etc.\n> +    Third-party Git tools may also create their own references.\n> ++\n> +Git may also create references in the base `.git` directory\n> +other than `HEAD`, like `ORIG_HEAD`.\n> +\n> +*NOTE:* As an optimization, references may be stored as packed\n> +refs instead of in `.git/refs`. See linkgit:git-pack-refs[1].\n> +\n> +[[index]]\n> +THE INDEX\n> +---------\n> +\n> +The index, also known as the \"staging area\", contains the current \n> staged\n> +version of every file in your Git repository. When you commit, the \n> files\n> +in the index are used as the files in the next commit.\n> +\n> +Unlike a tree, the index is a flat list of files.\n> +Each index entry has 4 fields:\n> +\n> +1. The *permissions*\n> +2. The *<<blob,blob>> ID* of the file\n> +3. The *filename*\n> +4. The *number*. This is normally 0, but if there's a merge conflict\n> +   there can be multiple versions (with numbers 0, 1, 2, ..)\n> +   of the same filename in the index.\n> +\n> +It's extremely uncommon to look at the index directly: normally you'd\n> +run `git status` to see a list of changes between the index and \n> <<HEAD,HEAD>>.\n> +But you can use `git ls-files --stage` to see the index.\n> +Here's the output of `git ls-files --stage` in a repository with 2 \n> files:\n> +\n> +----\n> +100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n> +100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n> +----\n> +\n> +[[reflogs]]\n> +REFLOGS\n> +-------\n> +\n> +Git stores the history of branch, tag, and HEAD refs in a reflog\n> +(you should read \"reflog\" as \"ref log\"). Not every ref is logged by\n> +default, but any ref can be logged.\n> +\n> +Each reflog entry has:\n> +\n> +1. *Before/after *commit IDs*\n> +2. *User* who made the change, for example `Maya <maya@example.com>`\n> +3. *Timestamp*\n> +4. *Log message*, for example `pull: Fast-forward`\n> +\n> +Reflogs only log changes made in your local repository.\n> +They are not shared with remotes.\n> +\n> +GIT\n> +---\n> +Part of the linkgit:git[1] suite\n>\n> base-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n> -- \n> gitgitgadget\n"},{"id":"528441","messageId":"F485A91C-2F1E-44DD-9179-0B47426DA3B7@gmail.com","threadId":"64244","inReplyTo":"b2a9b8ca-8f2a-40f0-a724-0da707902985@app.fastmail.com","subject":"Re: [PATCH] doc: add a explanation of Git's data model","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-10-10T00:42:43Z","receivedAt":"2025-10-10T00:42:55Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 9 oct. 2025 à 10:21, Julia Evans <julia@jvns.ca> a écrit :\n> \n> ﻿I collected some feedback from Git users on this v2 document. I'm expecting more\n> feedback, but here's an initial brain dump of my notes. I mostly wrote this for\n> my own use but I thought it might be interesting to other folks too.\n> \n[snip]\n> references:\n> \n> - Two people pointed out that because references are often stored as files,\n>  you can't have two references named `julia/ticket-number` and\n>  `julia/ticket-number/task-name`.\n>  I'm not sure if this is a fundamental limit of the refs data model\n>  (does the reftable backend have the same limitation?), but it could be\n>  a good reason to mention that refs are often stored as files, because\n>  it makes it obvious that you can't have a file and a directory with\n>  the same name.\n>  Obviously this is an issue that is affecting people relatively often\n>  in practice though so I think it's worth mentioning in some way.\n\nI don’t think the reftable backend has this limitation (?), but it reminded me of another important one: on case-insensitive filesystems you cannot have both « julia » and « JULIA » branches!\n\nThis occasionally creates problems where someone cannot fetch/clone what has been pushed.\n\nAnyway: it’s worth mentioning the files for that purpose. It would be nice to improve the UI as you describe below to continue to be able to naturally interrogate Git without needing to know about all the storage formats (recall that cat-file works just fine with packs and MIDXs!). \n\n> Overall: several people suggested mentioning more about where things\n> are stored in the `.git` directory, which I just removed.\n> \n> I think I want to avoid this (not sure yet), but I'm going to think\n> about the underlying motivation for this suggestion and see if it can be\n> addressed in a different way.\n> \n> Some ideas for what functions discussing the `.git` directory has:\n> \n> 1. Like I mentioned above with branches, sometimes the implementation causes\n>   some extra constraints like \"you can't have branches `julia/ticket`\n>   and `julia/ticket/task`\". So often people like to know a little\n>   about the implementation because it can help predict some of the\n>   holes in the abstractions you're using.\n> 2. It lets you view the \"raw\" data, so you can be totally sure about\n>   what Git is storing. This is nice because Git's UI can be very\n>   inconsistent sometimes, so looking at the raw data gives a sense of\n>   certainty about what's actually there.\n> \n> I tried to put together a list of ways to look at the \"raw\" data without\n> looking in the `.git` directory. The ways for objects and the index are great,\n> but for references and the reflog they involve these pretty complex format\n> strings, I'm not confident I've gotten the format strings right and IMO\n> they don't inspire a lot of confidence.\n> \n> View an object with:\n> ----\n> git cat-file -p <object-id>\n> ----\n> \n> View a reference with:\n> \n> ----\n> git for-each-ref <ref-name> --include-root-refs  --format=\"%(refname) %(if)%(symref)%(then)%(symref)%(else)%(objectname:short)%(end)\"\n> ----\n> \n> View the index with:\n> \n> ----\n> git ls-files --stage\n> ----\n> \n> View the reflog for a reference with:\n> \n> ----\n> git reflog show <refname> --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso\n> ----\n\n[kept for context]"},{"id":"528499","messageId":"aOjzQ7-88m5e_YJl@pks.im","threadId":"64244","inReplyTo":"pull.1981.v2.git.1759931621272.gitgitgadget@gmail.com","subject":"Re: [PATCH v2] doc: add a explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-10T11:51:31Z","receivedAt":"2025-10-10T11:51:38Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Oct 08, 2025 at 01:53:41PM +0000, Julia Evans via GitGitGadget wrote:\n[snip]\n> +[[blob]]\n> +blobs::\n> +    A blob is how Git represents a file. A blob object contains the\n> +    file's contents.\n> ++\n> +\n> +NOTE: Storing a new blob for every new version of a file can use a\n> +lot of disk space. To handle this, Git periodically runs repository\n> +maintenance with linkgit:git-gc[1]. Part of this maintenance is\n\nBy the way, this isn't true nowadays: Git does not use `git gc --auto`\nanymore, but instead `git maintenance run --auto`. So we really should\nbe linking to \"linkgit:git-maintenance[1]\".\n\nThis tool _by default_ executes git-gc(1). But it can be configured to\nuse alternative strategies, and when using scalar(1) we actually use a\ndifferent strategy.\n\n[snip]\n> +[[references]]\n> +REFERENCES\n> +----------\n> +\n> +References are a way to give a name to a commit.\n> +It's easier to remember \"the changes I'm working on are on the `turtle`\n> +branch\" than \"the changes are in commit bb69721404348e\".\n> +Git often uses \"ref\" as shorthand for \"reference\".\n> +\n> +References can either be:\n> +\n> +1. References to an object ID, usually a <<commit,commit>> ID\n> +2. References to another reference. This is called a \"symbolic reference\".\n> +\n> +References are stored in a hierarchy, and Git handles references\n> +differently based on where they are in the hierarchy.\n> +Most references are under `refs/`. Here are the main types:\n\nNot quite true. Pseudo refs are outside the hierarchy and are in fact\ntreated differently. But root refs are treated the same as any other\nreference.\n\n    References are stored in a hierarchy. While most references are\n    stored in the \"refs/\" hierarchy, some references with special\n    meaning like for example \"HEAD\" are stored directly in the root of\n    the hierarchy.\n\nI don't really think we should get into root refs vs pseudo refs here,\nso maybe this is sufficient?\n\n[snip]\n> +[[other-refs]]\n> +Other references::\n> +    Git tools may create references anywhere under `refs/`.\n> +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n> +    and linkgit:git-notes[1] all create their own references\n> +    in `refs/stash`, `refs/bisect`, etc.\n> +    Third-party Git tools may also create their own references.\n> ++\n> +Git may also create references other than `HEAD` at the base of the\n> +hierarchy, like `ORIG_HEAD`.\n\nMaybe append: \"These references are called root refs (see\nlinkgit:gitglossary[7]).\"\n\nPatrick\n"},{"id":"528628","messageId":"xmqq8qhe5040.fsf@gitster.g","threadId":"64244","inReplyTo":"aOjzQ7-88m5e_YJl@pks.im","subject":"Re: [PATCH v2] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-13T14:48:15Z","receivedAt":"2025-10-13T14:48:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> On Wed, Oct 08, 2025 at 01:53:41PM +0000, Julia Evans via GitGitGadget wrote:\n> [snip]\n>> +[[blob]]\n>> +blobs::\n>> +    A blob is how Git represents a file. A blob object contains the\n>> +    file's contents.\n>> ++\n>> +\n>> +NOTE: Storing a new blob for every new version of a file can use a\n>> +lot of disk space. To handle this, Git periodically runs repository\n>> +maintenance with linkgit:git-gc[1]. Part of this maintenance is\n>\n> By the way, this isn't true nowadays: Git does not use `git gc --auto`\n> anymore, but instead `git maintenance run --auto`. So we really should\n> be linking to \"linkgit:git-maintenance[1]\".\n>\n> This tool _by default_ executes git-gc(1). But it can be configured to\n> use alternative strategies, and when using scalar(1) we actually use a\n> different strategy.\n\nFor the curious, this happened around a95ce124 (maintenance: replace\nrun_auto_gc(), 2020-09-17).\n\n> Not quite true. Pseudo refs are outside the hierarchy and are in fact\n> treated differently. But root refs are treated the same as any other\n> reference.\n>\n>     References are stored in a hierarchy. While most references are\n>     stored in the \"refs/\" hierarchy, some references with special\n>     meaning like for example \"HEAD\" are stored directly in the root of\n>     the hierarchy.\n>\n> I don't really think we should get into root refs vs pseudo refs here,\n> so maybe this is sufficient?\n\nI do not think \"root ref\" (or pseudo for that matter) is a concept\nthat has no use in this context.  If this is really about data\nmodel, where you find refs (or what the \"pathname looking\" thing\nexactly look like that names your refs) should be immaterial.  It\ndoes help to know that HEAD is just a ref.  It also would help to\nknow there are symbolic refs that point at other refs, which is much\nmore relevant to the data model.\n\nThanks.\n"},{"id":"528685","messageId":"aO3jbnXRI67JsAx7@pks.im","threadId":"64244","inReplyTo":"xmqq8qhe5040.fsf@gitster.g","subject":"Re: [PATCH v2] doc: add a explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-14T05:45:18Z","receivedAt":"2025-10-14T05:45:25Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Oct 13, 2025 at 07:48:15AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > On Wed, Oct 08, 2025 at 01:53:41PM +0000, Julia Evans via GitGitGadget wrote:\n> > [snip]\n> > Not quite true. Pseudo refs are outside the hierarchy and are in fact\n> > treated differently. But root refs are treated the same as any other\n> > reference.\n> >\n> >     References are stored in a hierarchy. While most references are\n> >     stored in the \"refs/\" hierarchy, some references with special\n> >     meaning like for example \"HEAD\" are stored directly in the root of\n> >     the hierarchy.\n> >\n> > I don't really think we should get into root refs vs pseudo refs here,\n> > so maybe this is sufficient?\n> \n> I do not think \"root ref\" (or pseudo for that matter) is a concept\n> that has no use in this context.  If this is really about data\n> model, where you find refs (or what the \"pathname looking\" thing\n> exactly look like that names your refs) should be immaterial.  It\n> does help to know that HEAD is just a ref.  It also would help to\n> know there are symbolic refs that point at other refs, which is much\n> more relevant to the data model.\n\nYeah, I don't necessarily think that we need to mention root refs here.\nBut what I think we need to avoid is the following sentence, as it is\nmisleading:\n\n    References are stored in a hierarchy, and Git handles references\n    differently based on where they are in the hierarchy.\n\nPseudo refs are stored outside of the hierarchy and are indeed handled\ndifferently. But root refs are stored outside of the hierarchy and are\ntreated the same as any other ref, even though they of course have\nspecial meaning to some commands.\n\nSo maybe something like this would be preferable:\n\n    References are stored in a hierarchy. References that sit at the\n    root of the hierarchy often have special meaning to Git commands,\n    like for example \"HEAD\" or \"REBASE_HEAD\".\n\nIt hints at the fact that these references are special, but not in how\nthey are handled but rather in what they mean. It doesn't go into our\ntwo pseudo refs at all, but given that there's only FETCH_HEAD and\nMERGE_HEAD I don't think we should explain them. The water is getting\nsomewhat murky around pseudorefs anyway, so it probably only causes more\nconfusion.\n\nPatrick\n"},{"id":"528698","messageId":"46c6ca15-c1d2-4dd9-a6d3-2538f482b475@app.fastmail.com","threadId":"64244","inReplyTo":"aO3jbnXRI67JsAx7@pks.im","subject":"Re: [PATCH v2] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-14T09:18:58Z","receivedAt":"2025-10-14T09:19:19Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Tue, Oct 14, 2025, at 1:45 AM, Patrick Steinhardt wrote:\n> On Mon, Oct 13, 2025 at 07:48:15AM -0700, Junio C Hamano wrote:\n>> Patrick Steinhardt <ps@pks.im> writes:\n>> > On Wed, Oct 08, 2025 at 01:53:41PM +0000, Julia Evans via GitGitGadget wrote:\n>> > [snip]\n>> > Not quite true. Pseudo refs are outside the hierarchy and are in fact\n>> > treated differently. But root refs are treated the same as any other\n>> > reference.\n>> >\n>> >     References are stored in a hierarchy. While most references are\n>> >     stored in the \"refs/\" hierarchy, some references with special\n>> >     meaning like for example \"HEAD\" are stored directly in the root of\n>> >     the hierarchy.\n>> >\n>> > I don't really think we should get into root refs vs pseudo refs here,\n>> > so maybe this is sufficient?\n>> \n>> I do not think \"root ref\" (or pseudo for that matter) is a concept\n>> that has no use in this context.  If this is really about data\n>> model, where you find refs (or what the \"pathname looking\" thing\n>> exactly look like that names your refs) should be immaterial.  It\n>> does help to know that HEAD is just a ref.  It also would help to\n>> know there are symbolic refs that point at other refs, which is much\n>> more relevant to the data model.\n>\n> Yeah, I don't necessarily think that we need to mention root refs here.\n> But what I think we need to avoid is the following sentence, as it is\n> misleading:\n>\n>     References are stored in a hierarchy, and Git handles references\n>     differently based on where they are in the hierarchy.\n>\n\nWhy do you say that it’s misleading? (what do you think it’s implying that is not true?)\n\nWhat i’m trying to communicate is that branches, tags, etc are treated differently from each other and that Git knows how to handle them based on where they are in the hierarchy.\n\n> Pseudo refs are stored outside of the hierarchy and are indeed handled\n> differently. But root refs are stored outside of the hierarchy and are\n> treated the same as any other ref, even though they of course have\n> special meaning to some commands.\n>\n> So maybe something like this would be preferable:\n>\n>     References are stored in a hierarchy. References that sit at the\n>     root of the hierarchy often have special meaning to Git commands,\n>     like for example \"HEAD\" or \"REBASE_HEAD\".\n>\n> It hints at the fact that these references are special, but not in how\n> they are handled but rather in what they mean. It doesn't go into our\n> two pseudo refs at all, but given that there's only FETCH_HEAD and\n> MERGE_HEAD I don't think we should explain them. The water is getting\n> somewhat murky around pseudorefs anyway, so it probably only causes more\n> confusion.\n> Patrick\n"},{"id":"528701","messageId":"aO43ywdCKXWchRqG@pks.im","threadId":"64244","inReplyTo":"46c6ca15-c1d2-4dd9-a6d3-2538f482b475@app.fastmail.com","subject":"Re: [PATCH v2] doc: add a explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-14T11:45:15Z","receivedAt":"2025-10-14T11:45:22Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Oct 14, 2025 at 05:18:58AM -0400, Julia Evans wrote:\n> \n> \n> On Tue, Oct 14, 2025, at 1:45 AM, Patrick Steinhardt wrote:\n> > On Mon, Oct 13, 2025 at 07:48:15AM -0700, Junio C Hamano wrote:\n> >> Patrick Steinhardt <ps@pks.im> writes:\n> >> > On Wed, Oct 08, 2025 at 01:53:41PM +0000, Julia Evans via GitGitGadget wrote:\n> >> > [snip]\n> >> > Not quite true. Pseudo refs are outside the hierarchy and are in fact\n> >> > treated differently. But root refs are treated the same as any other\n> >> > reference.\n> >> >\n> >> >     References are stored in a hierarchy. While most references are\n> >> >     stored in the \"refs/\" hierarchy, some references with special\n> >> >     meaning like for example \"HEAD\" are stored directly in the root of\n> >> >     the hierarchy.\n> >> >\n> >> > I don't really think we should get into root refs vs pseudo refs here,\n> >> > so maybe this is sufficient?\n> >> \n> >> I do not think \"root ref\" (or pseudo for that matter) is a concept\n> >> that has no use in this context.  If this is really about data\n> >> model, where you find refs (or what the \"pathname looking\" thing\n> >> exactly look like that names your refs) should be immaterial.  It\n> >> does help to know that HEAD is just a ref.  It also would help to\n> >> know there are symbolic refs that point at other refs, which is much\n> >> more relevant to the data model.\n> >\n> > Yeah, I don't necessarily think that we need to mention root refs here.\n> > But what I think we need to avoid is the following sentence, as it is\n> > misleading:\n> >\n> >     References are stored in a hierarchy, and Git handles references\n> >     differently based on where they are in the hierarchy.\n> >\n> \n> Why do you say that it’s misleading? (what do you think it’s implying\n> that is not true?)\n> \n> What i’m trying to communicate is that branches, tags, etc are treated\n> differently from each other and that Git knows how to handle them\n> based on where they are in the hierarchy.\n\nOh, I think I managed to repeatedly misread this sentence! I was\nbasically s/where/whether/ and thought that this was saying that refs\nare handled differently depending on whether they are stored _in_ that\nhierarchy or _outside_ of it. And that would have been misleading\nindeed.\n\nBut that's not what this sentence says at all. So please ignore this\ntangent, sorry :)\n\nPatrick\n"},{"id":"528738","messageId":"xmqqa51tzjpm.fsf@gitster.g","threadId":"64244","inReplyTo":"46c6ca15-c1d2-4dd9-a6d3-2538f482b475@app.fastmail.com","subject":"Re: [PATCH v2] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-14T13:39:01Z","receivedAt":"2025-10-14T13:39:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n>> Yeah, I don't necessarily think that we need to mention root refs here.\n>> But what I think we need to avoid is the following sentence, as it is\n>> misleading:\n>>\n>>     References are stored in a hierarchy, and Git handles references\n>>     differently based on where they are in the hierarchy.\n>>\n>\n> Why do you say that it’s misleading? (what do you think it’s\n> implying that is not true?)\n>\n> What i’m trying to communicate is that branches, tags, etc are\n> treated differently from each other and that Git knows how to\n> handle them based on where they are in the hierarchy.\n\nFWIW, I had the same reaction to what response you are responding to\nsaid.  I think Patrick assumes that our target audiences would\nassume the \"hierarchy\" begins at \"refs/\" and \"root\" things are\noutside the hierarchy, but my mental model saw that the hierarchy\nbegan at the root level, most of things are in \"refs/\", but one\nlevel above it lives things like HEAD and ORIG_HEAD.  I think both\ncan be valid, but I do not know which views are more common.\n\nThanks.\n\n"},{"id":"528765","messageId":"pull.1981.v3.git.1760476346040.gitgitgadget@gmail.com","threadId":"64244","inReplyTo":"pull.1981.v2.git.1759931621272.gitgitgadget@gmail.com","subject":"[PATCH v3] doc: add a explanation of Git's data model","fromName":"Julia Evans via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-14T21:12:26Z","receivedAt":"2025-10-14T21:12:29Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"From: Julia Evans <julia@jvns.ca>\n\nGit very often uses the terms \"object\", \"reference\", or \"index\" in its\ndocumentation.\n\nHowever, it's hard to find a clear explanation of these terms and how\nthey relate to each other in the documentation. The closest candidates\ncurrently are:\n\n1. `gitglossary`. This makes a good effort, but it's an alphabetically\n    ordered dictionary and a dictionary is not a good way to learn\n    concepts. You have to jump around too much and it's not possible to\n    present the concepts in the order that they should be explained.\n2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n   This is a nice document to have, but it's not necessary to learn how\n   `update-index` works to understand Git's data model, and we should\n   not be requiring users to learn how to use the \"plumbing\" commands\n   if they want to learn what the term \"index\" or \"object\" means.\n3. `gitrepository-layout`. This is a great resource, but it includes a\n   lot of information about configuration and internal implementation\n   details which are not related to the data model. It also does\n   not explain how commits work.\n\nThe result of this is that Git users (even users who have been using\nGit for 15+ years) struggle to read the documentation because they don't\nknow what the core terms mean, and it's not possible to add links\nto help them learn more.\n\nAdd an explanation of Git's data model. Some choices I've made in\ndeciding what \"core data model\" means:\n\n1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n   if those are intended to be user facing or if they're more like\n   internal implementation details.\n2. Don't talk about submodules other than by mentioning how they\n   relate to trees. This is because Git has a lot of special features,\n   and explaining how they all work exhaustively could quickly go\n   down a rabbit hole which would make this document less useful for\n   understanding Git's core behaviour.\n3. Don't discuss the structure of a commit message\n   (first line, trailers etc).\n4. Don't mention configuration.\n5. Don't mention the `.git` directory, to avoid getting too much into\n   implementation details\n\nSigned-off-by: Julia Evans <julia@jvns.ca>\n---\n    doc: Add a explanation of Git's data model\n    \n    Changes in v2:\n    \n    The biggest change is to remove all mentions of the .git directory, and\n    explain references in a way that doesn't refer to \"directories\" at all,\n    and instead talks about the \"hierarchy\" (from Kristoffer and Patrick's\n    reviews).\n    \n    Also:\n    \n     * objects: Mention that an object ID is called an \"object name\", and\n       update the glossary to include the term \"object ID\" (from Junio's\n       review)\n     * objects: Replace \"SHA-1 hash\" with \"cryptographic hash\" which is more\n       accurate (from Patrick's review)\n     * blobs: Made the explanation of git gc a little higher level and took\n       some ideas from Patrick's suggested wording (from Patrick's and\n       Kroftoffer's reviews)\n     * commits: Mention that tag objects and commits can optionally have\n       other fields. I didn't mention the GPG signature specifically, but\n       don't have any objections to adding it. (from Patrick and Junio's\n       reviews)\n     * commits: Remove one of the mentions of git gc, since it perhaps opens\n       up too much of a rabbit hole: \"how does git gc decide which commits\n       to clean up?\". (from Kristoffer's review)\n     * tag objects: Add an example of how a tag object is represented (from\n       user feedback on the draft)\n     * index: Use the term \"file mode\" instead of \"permissions\", and list\n       all allowed file modes (from Patrick's review)\n     * index: Use \"stage number\" instead of \"number\" for index entries (from\n       Patrick's review)\n     * reflogs: Remove \"any ref can be logged\", it raises some questions of\n       \"how do you tell Git to log a ref that it isn't normally logging?\"\n       and my guess is that it's uncommon to ask Git to log more refs. I\n       don't think it's a \"lie\" to omit this but I can bring it back if\n       folks disagree. (from Patrick's review)\n     * reflogs: Fix an error I noticed in the explanation of reflogs: tags\n       aren't logged by default and remote-tracking branches are, according\n       to man git-config\n     * branches and tags: Be clearer about how branches are usually updated\n       (by committing), and make it a little more obvious that only branches\n       can be checked out. This is a bit tricky because using the word\n       \"check out\" introduces a rabbit hole that I want to avoid (what does\n       \"check out\" mean?). I've dealt this by just talking about the\n       \"current branch\" (HEAD) since that is defined here, and making it\n       more explicit that HEAD must either be a branch or a commit, there's\n       no \"HEAD is a tag\" option. (from Patrick's review)\n     * tags: Explain the differences between annotated and lightweight tags\n       (this is the main piece of user feedback I've gotten on the draft so\n       far)\n     * Various style/typo changes (\"2 or more\", linkgit:git-gc[1], removed\n       extra asterisks, added empty SYNOPSIS, \"commits -> tags\" typo fix,\n       add to meson build)\n    \n    non-changes:\n    \n     * I still haven't mentioned things that aren't part of the \"data\n       model\", like revision params and configuration. I think there could\n       be a place for them but I haven't found it yet.\n     * tag objects: I noticed that there's a \"tag\" header field in tag\n       objects (like tag v1.0.0) but I didn't mention it yet because I\n       couldn't figure out what the purpose of that field is (I thought the\n       tag name was stored in the reference, why is it duplicated in the tag\n       object?)\n    \n    Changes in v3:\n    \n    I asked for feedback from Git users on Mastodon and got 220 pieces of\n    feedback from 48 different users. People seemed very excited to read\n    about Git's data model. Usually I judge explanations by what folks\n    report learning from them. Here people reported learning:\n    \n     * how branches are stored (that a branch is \"a name for a commit\")\n     * how objects work\n     * that Git has separate \"author\" and \"committer\" fields\n     * that amending a commit does not change it\n     * that a tree is \"just a directory\" (not something more complicated),\n       and how trees are stored\n     * that Git repos can contain symlinks\n     * that Git saves modes separately from the OS.\n     * how the stage number works\n     * that when you git add a file, Git will create an object\n     * that third-party tools can create their own refs.\n     * that the reflog stores the history of branches (not just HEAD), and\n       what reflogs are for\n    \n    Also (of course) there were quite a few points of confusion! The main 4\n    pieces of feedback were\n    \n     1. The index section doesn't explain what the word \"staged\" means, and\n        one person says that it makes it sounds like only files that you\n        \"git add\"ed are in the index. Rewrite the explanation to avoid using\n        the word \"staged\" to define the index and instead define the word\n        \"staging\".\n     2. Explain the difference between \"annotated tags\" and \"lightweight\n        tags\" (done)\n     3. Add examples for tag objects and reflogs (done)\n     4. Mention a little more about where things are stored in the .git\n        directory, which I'd removed in v2. This seems most important for\n        .git/refs, so I added a hopefully accurate note about how refs are\n        stored by default, with a comment about one of the major\n        implications. I did not discuss where objects or the index are\n        stored, because I don't think the implementation details of how\n        objects are stored are as important, and there are better tools for\n        viewing the \"raw\" state of objects and the index (with git cat-file\n        -p or git ls-files --staged).\n    \n    Here's every other change I made in response to the feedback, as well as\n    a few comments that I did not address.\n    \n    intro:\n    \n     * Give a 1-sentence intro to \"reflog\"\n    \n    objects:\n    \n     * people really like having git ls-files --stage as a way to view the\n       index, so add git cat-file -p as well in a note\n    \n    commits:\n    \n     * 2 people asked \"Are commits stored as a diff?\". Say that diffs are\n       calculated at runtime, this is very important.\n     * The order the fields are given in don't match the order in the\n       example. Make them match.\n     * \"All the files in the commit, stored as a tree\" is throwing a few\n       people off. Be clearer that it's the tree ID of the base directory.\n     * Several people asked \"What's the difference between an author and\n       committer? I added an example using git cherry-pick that I'm not 100%\n       happy with (what if the reader doesn't know what cherry-pick does?).\n       There might be a better example to give here.\n     * In the note about commits being amended: one person suggested saying\n       \"creates a new commit with the same parent\" to make it clearer what\n       the relationship between the new and old commit are. I liked that\n       idea so I did it.\n    \n    trees:\n    \n     * file modes. 2 people want to know more about \"The file mode, for\n       example 100644\". Also 2 people are curious about what relationship\n       these have to Unix permissions. Say that they're inspired by Unix\n       permissions, and move the list of possible file modes up to make the\n       relationship clearer\n     * On \"so git-gc(1) periodically compresses objects to save disk space\",\n       there are a few follow up comments wondering about more, which makes\n       me think the comment about compression is actually a distraction. Say\n       something simpler instead, (\"Git only needs to store new versions of\n       files which were changed in that commit\"), from Junio's suggestion\n     * Re \"commit (a Git submodule)\": 2 people say it's not clear how trees\n       relate to submodules. Say that it refers to a commit in a different\n       repository.\n     * One person says they're not sure if the \"object ID\" is a hash. Link\n       it to the definition of \"object ID\".\n    \n    tag objects:\n    \n     * Requests for an example, added one.\n     * Requests to explain the difference between \"lightweight\" and\n       \"annotated\" tags, added it.\n    \n    tags:\n    \n     * one person thinks \"It’s expected that a tag will never change after\n       you create it.\" is too strong (since of course you can change it with\n       git tag -f). Say instead that tags are \"usually\" not changed.\n    \n    HEAD:\n    \n     * Several people are asking for more detail about detached HEAD state.\n       There's actually quite a lot to talk about here (what it means, how\n       it happens, what it implies, and how you might adjust your workflow\n       to avoid it by using git switch). I don't think we can get into all\n       of that here, so refer to the DETACHED HEAD section of git-checkout\n       instead. I'm not totally happy with the current version of that\n       section but that seems like the most practical solution right now.\n    \n    remote-tracking branches:\n    \n     * discuss refs/remotes/<remote>/HEAD.\n    \n    the index:\n    \n     * \"permissions\" should be \"file mode\" (like with trees). Changed.\n     * \"filename\" should be \"file path\". Changed.\n     * the stage number can only be 0, 1, 2, or 3, since it's 2 bits. Also\n       maybe say that the numbers have specific meanings. Said it can only\n       be 0/1/2/3 but did not give the specific meanings.\n    \n    reflogs\n    \n     * Request for an example. Added one.\n     * It's not clear if there's one reflog per branch/tag/HEAD, or if\n       there's one universal reflog. Make this clearer.\n     * Mention the role of the reflog in retrieving \"lost\" commits or\n       undoing bad rebases.\n    \n    Not fixed:\n    \n     * intro: A couple of people say that it's confusing that tags are both\n       \"an object\" and \"a reference\". Handled this by just explaining the\n       difference between an annotated and a lightweight tag further down.\n       I'd like to make this clearer in the intro but not sure if there's a\n       way to do it.\n     * commits and tag objects: one person asks if there's a reference for\n       the other \"optional fields\", like \"encoding\" and \"gpgsig\". I couldn't\n       find one, so left this as is.\n     * HEAD: A couple of people ask if there are any other symbolic\n       references other than HEAD, or if they can make their own symbolic\n       references. I don't know the answer to this.\n     * HEAD: the HEAD: HEAD thing looks weird, it made more sense when it\n       was HEAD: .git/HEAD. Will think about this.\n     * reflogs: One person asks: if reflogs only store local changes, why\n       does it track the user who made the change? Is that for remote\n       operations like fetches and pulls? Or for cases where more than one\n       user is using the same repo on a system? I don't know the answer to\n       this.\n     * reflogs: How can you see the full data in the reflog? git reflog show\n       doesn't list the user who made the change. git reflog show <refname>\n       --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso seems to work but\n       it's really a mouthful, not sure it's useful to include all that.\n     * index: Is it worth mentioning that the index can be locked? I don't\n       have an opinion about this.\n     * other: One person asks what a \"working tree\" is. It made me wonder if\n       \"the current working directory\" has a place in Git's data model. My\n       feeling is \"no\" but I could be convinced otherwise.\n     * overall: \"How can Git be so fast? If I switch branches, how does it\n       figure out what to add, remove or replace?\". I don't think this is\n       the right place for that discussion but it would\n     * there are some docs CI errors I haven't figured out yet (IDREF\n       attribute linkend references an unknown ID \"tree\")\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/1981\n\nRange-diff vs v2:\n\n 1:  3b38a88dc7 ! 1:  39da4e04cf doc: add a explanation of Git's data model\n     @@ Documentation/gitdatamodel.adoc (new)\n      +2. <<references,References>>: branches, tags,\n      +   remote-tracking branches, etc\n      +3. <<index,The index>>, also known as the staging area\n     -+4. <<reflogs,Reflogs>>\n     ++4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n      +\n      +[[objects]]\n      +OBJECTS\n     @@ Documentation/gitdatamodel.adoc (new)\n      +Commits, trees, blobs, and tag objects are all stored in Git's object database.\n      +Every object has:\n      +\n     ++[[object-id]]\n      +1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n      +  type and contents.\n      +  It's fast to look up a Git object using its ID.\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    A commit contains these required fields\n      +    (though there are other optional fields):\n      ++\n     -+1. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n     ++1. All the *files* in the commit, stored as the *<<tree,tree>>* ID of\n     ++   the commit's base directory.\n     ++2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n      +  regular commits have 1 parent, merge commits have 2 or more parents\n     -+2. A *commit message*\n     -+3. All the *files* in the commit, stored as a *<<tree,tree>>*\n     -+4. An *author* and the time the commit was authored\n     -+5. A *committer* and the time the commit was committed\n     ++3. An *author* and the time the commit was authored\n     ++4. A *committer* and the time the commit was committed.\n     ++   If you cherry-pick (linkgit:git-cherry-pick[1]) someone else's commit,\n     ++   then they will be the author and you'll be the committer.\n     ++5. A *commit message*\n      ++\n      +Here's how an example commit is stored:\n      ++\n     @@ Documentation/gitdatamodel.adoc (new)\n      +----\n      ++\n      +Like all other objects, commits can never be changed after they're created.\n     -+For example, \"amending\" a commit with `git commit --amend` creates a new commit.\n     ++For example, \"amending\" a commit with `git commit --amend` creates a new\n     ++commit with the same parent.\n     +++\n     ++Git does not store the diff for a commit: when you ask Git for a\n     ++diff it calculates it on the fly.\n      +\n      +[[tree]]\n      +trees::\n      +    A tree is how Git represents a directory. It lists, for each item in\n      +    the tree:\n      ++\n     -+1. The *file mode*, for example `100644`\n     ++[[file-mode]]\n     ++1. The *file mode*, for example `100644`. The format is inspired by Unix\n     ++   permissions, but Git's modes are much more limited. Git only supports these file modes:\n     +++\n     ++  - `100644`: regular file (with type `blob`)\n     ++  - `100755`: executable file (with type `blob`)\n     ++  - `120000`: symbolic link (with type `blob`)\n     ++  - `040000`: directory (with type `tree`)\n     ++  - `160000`: gitlink, for use with submodules (with type `commit`)\n     ++\n      +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n     -+  or <<commit,`commit`>> (a Git submodule)\n     -+3. The *object ID*\n     ++  or <<commit,`commit`>> (a Git submodule, which is a\n     ++  commit from a different Git repository)\n     ++3. The <<object-id,*object ID*>>\n      +4. The *filename*\n      ++\n      +For example, this is how a tree containing one directory (`src`) and one file\n     @@ Documentation/gitdatamodel.adoc (new)\n      +100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n      +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n      +----\n     -++\n     -+Git only supports these file modes:\n     -++\n     -+  - `100644`: regular file (with type `blob`)\n     -+  - `100755`: executable file (with type `blob`)\n     -+  - `120000`: symbolic link (with type `blob`)\n     -+  - `040000`: directory (with type `tree`)\n     -+  - `160000`: gitlink, for use with submodules (with type `commit`)\n     ++\n      +\n      +[[blob]]\n      +blobs::\n      +    A blob is how Git represents a file. A blob object contains the\n      +    file's contents.\n      ++\n     -+\n     -+NOTE: Storing a new blob for every new version of a file can use a\n     -+lot of disk space. To handle this, Git periodically runs repository\n     -+maintenance with linkgit:git-gc[1]. Part of this maintenance is\n     -+compressing objects so that if a small part of a file was changed, only\n     -+the change is stored instead of the whole file.\n     ++When you make a new commit, Git only needs to store new versions of\n     ++files which were changed in that commit. This means that commits\n     ++can use relatively little disk space even in a very large repository.\n      +\n      +[[tag-object]]\n      +tag objects::\n     -+    Tag objects (also known as \"annotated tags\") contain these required fields\n     ++    Tag objects contain these required fields\n      +    (though there are other optional fields):\n      ++\n     -+1. The *tagger* and tag date\n     -+2. A *tag message*, similar to a commit message\n     -+3. The *ID* and *type* of the object (often a commit) that they reference\n     ++1. The *ID* and *type* of the object (often a commit) that they reference\n     ++2. The *tagger* and tag date\n     ++3. A *tag message*, similar to a commit message\n      +\n      +Here's how an example tag object is stored:\n      +\n     @@ Documentation/gitdatamodel.adoc (new)\n      +Release version 1.0.0\n      +----\n      +\n     ++NOTE: All of the examples in this section were generated with\n     ++`git cat-file -p <object-id>`, which shows the contents of a Git object.\n     ++\n      +[[references]]\n      +REFERENCES\n      +----------\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    A tag is a name for a commit ID, tag object ID, or other object ID.\n      +    Tags that reference a tag object ID are called \"annotated tags\",\n      +    because the tag object contains a tag message.\n     -+    Tags that reference a commit ID, blob ID, or tree ID are\n     ++    Tags that reference a commit, blob, or tree ID are\n      +    called \"lightweight tags\".\n      ++\n      +Even though branches and tags are both \"a name for a commit ID\", Git\n      +treats them very differently.\n      +Branches are expected to change over time: when you make a commit, Git\n      +will update your <<HEAD,current branch>> to reference the new changes.\n     -+It's expected that a tag will never change after you create it.\n     ++Tags are usually not changed after they're created.\n      +\n      +[[HEAD]]\n      +HEAD: `HEAD`::\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    `HEAD` can either be:\n      +    1. A symbolic reference to your current branch, for example `ref:\n      +       refs/heads/main` if your current branch is `main`.\n     -+    2. A direct reference to a commit ID.\n     -+        This is called \"detached HEAD state\".\n     ++    2. A direct reference to a commit ID. This is called \"detached HEAD\n     ++\t   state\", see the DETACHED HEAD section of linkgit:git-checkout[1] for more.\n      +\n      +[[remote-tracking-branch]]\n      +remote tracking branches: `refs/remotes/<remote>/<branch>`::\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    repository. `git fetch` updates remote-tracking branches. When\n      +    `git status` says \"you're up to date with origin/main\", it's looking at\n      +    this.\n     +++\n     ++`refs/remotes/<remote>/HEAD` is a symbolic reference to the remote's\n     ++default branch. This is the branch that `git clone` checks out by default.\n      +\n      +[[other-refs]]\n      +Other references::\n     @@ Documentation/gitdatamodel.adoc (new)\n      ++\n      +Git may also create references other than `HEAD` at the base of the\n      +hierarchy, like `ORIG_HEAD`.\n     +++\n     ++NOTE: By default, Git references are stored as files in the `.git` directory.\n     ++For example, the branch `main` is stored in `.git/refs/heads/main`.\n     ++This means that you can't have branches named both `maya` and `maya/some-task`,\n     ++because there can't be a file and a directory with the same name.\n      +\n      +[[index]]\n      +THE INDEX\n      +---------\n      +\n     -+The index, also known as the \"staging area\", contains the current staged\n     -+version of every file in your Git repository. When you commit, the files\n     -+in the index are used as the files in the next commit.\n     ++The index, also known as the \"staging area\", contains a list of every\n     ++file in the repository and its contents. When you commit, the files in\n     ++the index are used as the files in the next commit.\n     ++\n     ++You can add files to the index or update the version in the index with\n     ++linkgit:git-add[1]. Adding a file to the index or updating its version\n     ++is called \"staging\" the file for commit.\n      +\n     -+Unlike a tree, the index is a flat list of files.\n     ++Unlike a <<tree,tree>>, the index is a flat list of files.\n      +Each index entry has 4 fields:\n      +\n     -+1. The *permissions*\n     ++1. The *<<file-mode,file mode>>*\n      +2. The *<<blob,blob>> ID* of the file\n     -+3. The *filename*\n     -+4. The *stage number*. This is normally 0, but if there's a merge conflict\n     -+   there can be multiple versions (with numbers 0, 1, 2, ..)\n     -+   of the same filename in the index.\n     ++3. The *file path*, for example `src/hello.py`\n     ++4. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n     ++   there's a merge conflict there can be multiple versions of the same\n     ++   filename in the index.\n      +\n      +It's extremely uncommon to look at the index directly: normally you'd\n      +run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n     @@ Documentation/gitdatamodel.adoc (new)\n      +REFLOGS\n      +-------\n      +\n     -+Git stores the history of your branch, remote-tracking branch, and HEAD refs\n     -+in a reflog (you should read \"reflog\" as \"ref log\").\n     ++Git stores a history called a \"reflog\" for every branch, remote-tracking\n     ++branch, and HEAD. This means that if you make a mistake and \"lose\" a\n     ++commit, you can generally recover the commit ID by running\n     ++`git reflog <reference>`.\n      +\n      +Each reflog entry has:\n      +\n     @@ Documentation/gitdatamodel.adoc (new)\n      +Reflogs only log changes made in your local repository.\n      +They are not shared with remotes.\n      +\n     ++For example, here's how the reflog for `HEAD` in a repository with 2\n     ++commits is stored:\n     ++\n     ++----\n     ++0000000000000000000000000000000000000000 4ccb6d7b8869a86aae2e84c56523f8705b50c647 Maya <maya@example.com> 1759173408 -0400      commit (initial): Initial commit\n     ++4ccb6d7b8869a86aae2e84c56523f8705b50c647 750b4ead9c87ceb3ddb7a390e6c7074521797fb3 Maya <maya@example.com> 1759173425 -0400      commit: Add README\n     ++----\n     ++\n      +GIT\n      +---\n      +Part of the linkgit:git[1] suite\n\n\n Documentation/Makefile              |   1 +\n Documentation/gitdatamodel.adoc     | 281 ++++++++++++++++++++++++++++\n Documentation/glossary-content.adoc |   4 +-\n Documentation/meson.build           |   1 +\n 4 files changed, 285 insertions(+), 2 deletions(-)\n create mode 100644 Documentation/gitdatamodel.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 6fb83d0c6e..5f4acfacbd 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n MAN7_TXT += gitcore-tutorial.adoc\n MAN7_TXT += gitcredentials.adoc\n MAN7_TXT += gitcvs-migration.adoc\n+MAN7_TXT += gitdatamodel.adoc\n MAN7_TXT += gitdiffcore.adoc\n MAN7_TXT += giteveryday.adoc\n MAN7_TXT += gitfaq.adoc\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nnew file mode 100644\nindex 0000000000..f49574dfae\n--- /dev/null\n+++ b/Documentation/gitdatamodel.adoc\n@@ -0,0 +1,281 @@\n+gitdatamodel(7)\n+===============\n+\n+NAME\n+----\n+gitdatamodel - Git's core data model\n+\n+SYNOPSIS\n+--------\n+gitdatamodel\n+\n+DESCRIPTION\n+-----------\n+\n+It's not necessary to understand Git's data model to use Git, but it's\n+very helpful when reading Git's documentation so that you know what it\n+means when the documentation says \"object\", \"reference\" or \"index\".\n+\n+Git's core operations use 4 kinds of data:\n+\n+1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+2. <<references,References>>: branches, tags,\n+   remote-tracking branches, etc\n+3. <<index,The index>>, also known as the staging area\n+4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n+\n+[[objects]]\n+OBJECTS\n+-------\n+\n+Commits, trees, blobs, and tag objects are all stored in Git's object database.\n+Every object has:\n+\n+[[object-id]]\n+1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n+  type and contents.\n+  It's fast to look up a Git object using its ID.\n+  This is usually represented in hexadecimal, like\n+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+2. a *type*. There are 4 types of objects:\n+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n+   and <<tag-object,tag objects>>.\n+3. *contents*. The structure of the contents depends on the type.\n+\n+Once an object is created, it can never be changed.\n+Here are the 4 types of objects:\n+\n+[[commit]]\n+commits::\n+    A commit contains these required fields\n+    (though there are other optional fields):\n++\n+1. All the *files* in the commit, stored as the *<<tree,tree>>* ID of\n+   the commit's base directory.\n+2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n+  regular commits have 1 parent, merge commits have 2 or more parents\n+3. An *author* and the time the commit was authored\n+4. A *committer* and the time the commit was committed.\n+   If you cherry-pick (linkgit:git-cherry-pick[1]) someone else's commit,\n+   then they will be the author and you'll be the committer.\n+5. A *commit message*\n++\n+Here's how an example commit is stored:\n++\n+----\n+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n+author Maya <maya@example.com> 1759173425 -0400\n+committer Maya <maya@example.com> 1759173425 -0400\n+\n+Add README\n+----\n++\n+Like all other objects, commits can never be changed after they're created.\n+For example, \"amending\" a commit with `git commit --amend` creates a new\n+commit with the same parent.\n++\n+Git does not store the diff for a commit: when you ask Git for a\n+diff it calculates it on the fly.\n+\n+[[tree]]\n+trees::\n+    A tree is how Git represents a directory. It lists, for each item in\n+    the tree:\n++\n+[[file-mode]]\n+1. The *file mode*, for example `100644`. The format is inspired by Unix\n+   permissions, but Git's modes are much more limited. Git only supports these file modes:\n++\n+  - `100644`: regular file (with type `blob`)\n+  - `100755`: executable file (with type `blob`)\n+  - `120000`: symbolic link (with type `blob`)\n+  - `040000`: directory (with type `tree`)\n+  - `160000`: gitlink, for use with submodules (with type `commit`)\n+\n+2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n+  or <<commit,`commit`>> (a Git submodule, which is a\n+  commit from a different Git repository)\n+3. The <<object-id,*object ID*>>\n+4. The *filename*\n++\n+For example, this is how a tree containing one directory (`src`) and one file\n+(`README.md`) is stored:\n++\n+----\n+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n+----\n+\n+\n+[[blob]]\n+blobs::\n+    A blob is how Git represents a file. A blob object contains the\n+    file's contents.\n++\n+When you make a new commit, Git only needs to store new versions of\n+files which were changed in that commit. This means that commits\n+can use relatively little disk space even in a very large repository.\n+\n+[[tag-object]]\n+tag objects::\n+    Tag objects contain these required fields\n+    (though there are other optional fields):\n++\n+1. The *ID* and *type* of the object (often a commit) that they reference\n+2. The *tagger* and tag date\n+3. A *tag message*, similar to a commit message\n+\n+Here's how an example tag object is stored:\n+\n+----\n+object 750b4ead9c87ceb3ddb7a390e6c7074521797fb3\n+type commit\n+tag v1.0.0\n+tagger Maya <maya@example.com> 1759927359 -0400\n+\n+Release version 1.0.0\n+----\n+\n+NOTE: All of the examples in this section were generated with\n+`git cat-file -p <object-id>`, which shows the contents of a Git object.\n+\n+[[references]]\n+REFERENCES\n+----------\n+\n+References are a way to give a name to a commit.\n+It's easier to remember \"the changes I'm working on are on the `turtle`\n+branch\" than \"the changes are in commit bb69721404348e\".\n+Git often uses \"ref\" as shorthand for \"reference\".\n+\n+References can either be:\n+\n+1. References to an object ID, usually a <<commit,commit>> ID\n+2. References to another reference. This is called a \"symbolic reference\".\n+\n+References are stored in a hierarchy, and Git handles references\n+differently based on where they are in the hierarchy.\n+Most references are under `refs/`. Here are the main types:\n+\n+[[branch]]\n+branches: `refs/heads/<name>`::\n+    A branch is a name for a commit ID.\n+    That commit is the latest commit on the branch.\n++\n+To get the history of commits on a branch, Git will start at the commit\n+ID the branch references, and then look at the commit's parent(s),\n+the parent's parent, etc.\n+\n+[[tag]]\n+tags: `refs/tags/<name>`::\n+    A tag is a name for a commit ID, tag object ID, or other object ID.\n+    Tags that reference a tag object ID are called \"annotated tags\",\n+    because the tag object contains a tag message.\n+    Tags that reference a commit, blob, or tree ID are\n+    called \"lightweight tags\".\n++\n+Even though branches and tags are both \"a name for a commit ID\", Git\n+treats them very differently.\n+Branches are expected to change over time: when you make a commit, Git\n+will update your <<HEAD,current branch>> to reference the new changes.\n+Tags are usually not changed after they're created.\n+\n+[[HEAD]]\n+HEAD: `HEAD`::\n+    `HEAD` is where Git stores your current <<branch,branch>>.\n+    `HEAD` can either be:\n+    1. A symbolic reference to your current branch, for example `ref:\n+       refs/heads/main` if your current branch is `main`.\n+    2. A direct reference to a commit ID. This is called \"detached HEAD\n+\t   state\", see the DETACHED HEAD section of linkgit:git-checkout[1] for more.\n+\n+[[remote-tracking-branch]]\n+remote tracking branches: `refs/remotes/<remote>/<branch>`::\n+    A remote-tracking branch is a name for a commit ID.\n+    It's how Git stores the last-known state of a branch in a remote\n+    repository. `git fetch` updates remote-tracking branches. When\n+    `git status` says \"you're up to date with origin/main\", it's looking at\n+    this.\n++\n+`refs/remotes/<remote>/HEAD` is a symbolic reference to the remote's\n+default branch. This is the branch that `git clone` checks out by default.\n+\n+[[other-refs]]\n+Other references::\n+    Git tools may create references anywhere under `refs/`.\n+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n+    and linkgit:git-notes[1] all create their own references\n+    in `refs/stash`, `refs/bisect`, etc.\n+    Third-party Git tools may also create their own references.\n++\n+Git may also create references other than `HEAD` at the base of the\n+hierarchy, like `ORIG_HEAD`.\n++\n+NOTE: By default, Git references are stored as files in the `.git` directory.\n+For example, the branch `main` is stored in `.git/refs/heads/main`.\n+This means that you can't have branches named both `maya` and `maya/some-task`,\n+because there can't be a file and a directory with the same name.\n+\n+[[index]]\n+THE INDEX\n+---------\n+\n+The index, also known as the \"staging area\", contains a list of every\n+file in the repository and its contents. When you commit, the files in\n+the index are used as the files in the next commit.\n+\n+You can add files to the index or update the version in the index with\n+linkgit:git-add[1]. Adding a file to the index or updating its version\n+is called \"staging\" the file for commit.\n+\n+Unlike a <<tree,tree>>, the index is a flat list of files.\n+Each index entry has 4 fields:\n+\n+1. The *<<file-mode,file mode>>*\n+2. The *<<blob,blob>> ID* of the file\n+3. The *file path*, for example `src/hello.py`\n+4. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n+   there's a merge conflict there can be multiple versions of the same\n+   filename in the index.\n+\n+It's extremely uncommon to look at the index directly: normally you'd\n+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n+But you can use `git ls-files --stage` to see the index.\n+Here's the output of `git ls-files --stage` in a repository with 2 files:\n+\n+----\n+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n+----\n+\n+[[reflogs]]\n+REFLOGS\n+-------\n+\n+Git stores a history called a \"reflog\" for every branch, remote-tracking\n+branch, and HEAD. This means that if you make a mistake and \"lose\" a\n+commit, you can generally recover the commit ID by running\n+`git reflog <reference>`.\n+\n+Each reflog entry has:\n+\n+1. Before/after *commit IDs*\n+2. *User* who made the change, for example `Maya <maya@example.com>`\n+3. *Timestamp* when the change was made\n+4. *Log message*, for example `pull: Fast-forward`\n+\n+Reflogs only log changes made in your local repository.\n+They are not shared with remotes.\n+\n+For example, here's how the reflog for `HEAD` in a repository with 2\n+commits is stored:\n+\n+----\n+0000000000000000000000000000000000000000 4ccb6d7b8869a86aae2e84c56523f8705b50c647 Maya <maya@example.com> 1759173408 -0400      commit (initial): Initial commit\n+4ccb6d7b8869a86aae2e84c56523f8705b50c647 750b4ead9c87ceb3ddb7a390e6c7074521797fb3 Maya <maya@example.com> 1759173425 -0400      commit: Add README\n+----\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\ndiff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc\nindex e423e4765b..20ba121314 100644\n--- a/Documentation/glossary-content.adoc\n+++ b/Documentation/glossary-content.adoc\n@@ -297,8 +297,8 @@ This commit is referred to as a \"merge commit\", or sometimes just a\n \tidentified by its <<def_object_name,object name>>. The objects usually\n \tlive in `$GIT_DIR/objects/`.\n \n-[[def_object_identifier]]object identifier (oid)::\n-\tSynonym for <<def_object_name,object name>>.\n+[[def_object_identifier]]object identifier, object ID, oid::\n+\tSynonyms for <<def_object_name,object name>>.\n \n [[def_object_name]]object name::\n \tThe unique identifier of an <<def_object,object>>.  The\ndiff --git a/Documentation/meson.build b/Documentation/meson.build\nindex e34965c5b0..ace0573e82 100644\n--- a/Documentation/meson.build\n+++ b/Documentation/meson.build\n@@ -192,6 +192,7 @@ manpages = {\n   'gitcore-tutorial.adoc' : 7,\n   'gitcredentials.adoc' : 7,\n   'gitcvs-migration.adoc' : 7,\n+  'gitdatamodel.adoc' : 7,\n   'gitdiffcore.adoc' : 7,\n   'giteveryday.adoc' : 7,\n   'gitfaq.adoc' : 7,\n\nbase-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n-- \ngitgitgadget\n"},{"id":"528788","messageId":"aO8-NtJPNBAM2tVn@pks.im","threadId":"64244","inReplyTo":"pull.1981.v3.git.1760476346040.gitgitgadget@gmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-15T06:24:54Z","receivedAt":"2025-10-15T06:25:03Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Oct 14, 2025 at 09:12:26PM +0000, Julia Evans via GitGitGadget wrote:\n[snip]\n> +[[commit]]\n> +commits::\n> +    A commit contains these required fields\n> +    (though there are other optional fields):\n> ++\n> +1. All the *files* in the commit, stored as the *<<tree,tree>>* ID of\n> +   the commit's base directory.\n> +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n> +  regular commits have 1 parent, merge commits have 2 or more parents\n> +3. An *author* and the time the commit was authored\n> +4. A *committer* and the time the commit was committed.\n> +   If you cherry-pick (linkgit:git-cherry-pick[1]) someone else's commit,\n> +   then they will be the author and you'll be the committer.\n> +5. A *commit message*\n> ++\n> +Here's how an example commit is stored:\n> ++\n> +----\n> +tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n> +parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n> +author Maya <maya@example.com> 1759173425 -0400\n> +committer Maya <maya@example.com> 1759173425 -0400\n> +\n> +Add README\n> +----\n> ++\n> +Like all other objects, commits can never be changed after they're created.\n> +For example, \"amending\" a commit with `git commit --amend` creates a new\n> +commit with the same parent.\n\nLet's say \"parents\" instead of \"parent\" here so that it also works for\nroot and merge commits.\n\n[snip]\n> +[[other-refs]]\n> +Other references::\n> +    Git tools may create references anywhere under `refs/`.\n> +    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n> +    and linkgit:git-notes[1] all create their own references\n> +    in `refs/stash`, `refs/bisect`, etc.\n> +    Third-party Git tools may also create their own references.\n> ++\n> +Git may also create references other than `HEAD` at the base of the\n> +hierarchy, like `ORIG_HEAD`.\n> ++\n> +NOTE: By default, Git references are stored as files in the `.git` directory.\n> +For example, the branch `main` is stored in `.git/refs/heads/main`.\n> +This means that you can't have branches named both `maya` and `maya/some-task`,\n> +because there can't be a file and a directory with the same name.\n\nHm. I think mentioning this can help, but it may also creates questions\nwhen someone has a \"main\" branch but is unable find it in\n\".git/refs/heads/main\" because it has either been packed, or because the\nrepository uses reftables.\n\nI don't really know what to do about this. I think the most sensible\nthing would be to introduce two man pages gitformat-reffiles(5) and\ngitformat-reftables(5) that we can reference here for further reading.\n\n[snip]\n> +[[reflogs]]\n> +REFLOGS\n> +-------\n> +\n> +Git stores a history called a \"reflog\" for every branch, remote-tracking\n\nI think it's a bit unclear what \"history\" means here. Maybe:\n\n    Git stores a \"reflog\" for every branch, remote-tracking branch and\n    \"HEAD\" that contains the annotated history of all updates for a\n    particular reference. This means...\n\nOther than those handful of comments I'm happy with the current version,\nthanks!\n\nPatrick\n"},{"id":"528822","messageId":"xmqqsefkuqkv.fsf@gitster.g","threadId":"64244","inReplyTo":"aO8-NtJPNBAM2tVn@pks.im","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-15T15:34:08Z","receivedAt":"2025-10-15T15:34:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> +Like all other objects, commits can never be changed after they're created.\n>> +For example, \"amending\" a commit with `git commit --amend` creates a new\n>> +commit with the same parent.\n>\n> Let's say \"parents\" instead of \"parent\" here so that it also works for\n> root and merge commits.\n\nI just found it amusing that parents can be 0 ;-)\n\n>> +NOTE: By default, Git references are stored as files in the `.git` directory.\n>> +For example, the branch `main` is stored in `.git/refs/heads/main`.\n>> +This means that you can't have branches named both `maya` and `maya/some-task`,\n>> +because there can't be a file and a directory with the same name.\n>\n> Hm. I think mentioning this can help, but it may also creates questions\n> when someone has a \"main\" branch but is unable find it in\n> \".git/refs/heads/main\" because it has either been packed, or because the\n> repository uses reftables.\n\nI had the same thought.  The only thing we want to stress here is\nthat the names of refs _behave_ like filesystem entities.  So how\nabout saying just\n\n    Note: when you have a branch with <name>, you cannot have any\n    branch whose name begins with \"<name>/\".\n\nand stop at it?  It may look like an arbitrary limitation, and once\nin a distant future ref-files gets retired, it will become one (as\nthere is no inherent reason why reftable backend must retain it; it\nonly enforces the same limitation to ensure that the names it stores\ninteroperate with another clone that uses ref-files backend).  At\nthe data-model level (which is the theme of this document), it is\njust as immaterial as refnames may be case insensitive on some\nsystems.\n\nMentioning the limitation may be good, but the data model document\nis not the right place to explain where this limitation comes from\n(i.e. to be compatible with and expressible in ref-files backend).\nWe do not say \"you may not be able to have 'maya' branch and 'mAYa'\nbranch at the same time on some systems\", either ;-).\n\n>> +Git stores a history called a \"reflog\" for every branch, remote-tracking\n>\n> I think it's a bit unclear what \"history\" means here. Maybe:\n\n\"records of updates\", perhaps?\n"},{"id":"528827","messageId":"353916d3-c977-40e5-9251-1535b226cc9e@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqsefkuqkv.fsf@gitster.g","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-15T17:20:30Z","receivedAt":"2025-10-15T17:22:18Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"On Wed, Oct 15, 2025, at 11:34 AM, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n>\n>>> +Like all other objects, commits can never be changed after they're created.\n>>> +For example, \"amending\" a commit with `git commit --amend` creates a new\n>>> +commit with the same parent.\n>>\n>> Let's say \"parents\" instead of \"parent\" here so that it also works for\n>> root and merge commits.\n>\n> I just found it amusing that parents can be 0 ;-)\n\nWill change to parent(s).\n\n>>> +NOTE: By default, Git references are stored as files in the `.git` directory.\n>>> +For example, the branch `main` is stored in `.git/refs/heads/main`.\n>>> +This means that you can't have branches named both `maya` and `maya/some-task`,\n>>> +because there can't be a file and a directory with the same name.\n>>\n>> Hm. I think mentioning this can help, but it may also creates questions\n>> when someone has a \"main\" branch but is unable find it in\n>> \".git/refs/heads/main\" because it has either been packed, or because the\n>> repository uses reftables.\n>\n> I had the same thought.  The only thing we want to stress here is\n> that the names of refs _behave_ like filesystem entities.  So how\n> about saying just\n>\n>     Note: when you have a branch with <name>, you cannot have any\n>     branch whose name begins with \"<name>/\".\n>\n> and stop at it?  It may look like an arbitrary limitation, and once\n> in a distant future ref-files gets retired, it will become one (as\n> there is no inherent reason why reftable backend must retain it; it\n> only enforces the same limitation to ensure that the names it stores\n> interoperate with another clone that uses ref-files backend).  At\n> the data-model level (which is the theme of this document), it is\n> just as immaterial as refnames may be case insensitive on some\n> systems.\n>\n> Mentioning the limitation may be good, but the data model document\n> is not the right place to explain where this limitation comes from\n> (i.e. to be compatible with and expressible in ref-files backend).\n\nI'm still not clear on why you think we shouldn't mention that how\nreferences behave depends on which filesystem you're using.\n\nIs it because the fact that how references behave depends on which FS\nyou're using is considered a \"bug\", Git is working on eventually fixing\nthat bug via the reftable backend, and we don't want to document\n\"bugs\" as an expected part of the data model?\n\nI do think it's important to tell users where the data model has \"weak points\"\nwhere the abstraction leaks through to the implementation, pretending that\nabstractions are stronger than they are leads to unnecessary confusion.\n\n> We do not say \"you may not be able to have 'maya' branch and 'mAYa'\n> branch at the same time on some systems\", either ;-).\n\nSpeaking of case-insensitive filesystems, I wonder if we should add a\nshort note about the rules for filenames in Git. I ran into an issue\nrecently where I had a filename with a colon in it, and my\ncollaborator (who was using Windows) could not check out the branch\nbecause of that, and I saw another similar issue recently where one\ncollaborator was using a case-insensitive filesystem and the other wasn't.\n\nMy guess is that Git does not enforce any rules about filenames (?),\nand it's up to the user to make sure that the filenames in the repository\nwill work well for everyone collaborating on the repository.\n\n>>> +Git stores a history called a \"reflog\" for every branch, remote-tracking\n>>\n>> I think it's a bit unclear what \"history\" means here. Maybe:\n>\n> \"records of updates\", perhaps?\n\nAgreed. Perhaps this instead:\n\nEvery time a branch, remote-tracking branch, or HEAD is updated, Git\nupdates a log called a \"reflog\" for that <<reference,reference>>.\n"},{"id":"528829","messageId":"xmqqv7kgszr1.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v3.git.1760476346040.gitgitgadget@gmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-15T19:58:58Z","receivedAt":"2025-10-15T19:59:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +[[commit]]\n> +commits::\n> +    A commit contains these required fields\n> +    (though there are other optional fields):\n> ++\n> +1. All the *files* in the commit, stored as the *<<tree,tree>>* ID of\n> +   the commit's base directory.\n\n\"all the files' exact contents at the time of the commit\" is what we\nmean here, and once readers know what a tree is, the above sentence\nwould be understood as such, but \"All the files\" felt somewhat\nfuzzy.  I wonder if presenting objects in bottom-up fashion makes it\neasier to see?  Learn that a blob records exact content of a file,\nthen learn that a tree records the set of paths with exact contents\nstored at these paths, and after that, learn that a commit records a\ntree, hence a snapshot of the whole set of contents.  I dunno...\n\n> +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n> +  regular commits have 1 parent, merge commits have 2 or more parents\n> +3. An *author* and the time the commit was authored\n> +4. A *committer* and the time the commit was committed.\n> +   If you cherry-pick (linkgit:git-cherry-pick[1]) someone else's commit,\n> +   then they will be the author and you'll be the committer.\n\nIt felt a bit odd to single-out cherry-pick here.\n\nI think the important thing to become aware of for the readers at\nthis point is that the author and committer can be different people,\nand it does not matter how one commits somebody else's patch at the\nmechanical level.\n\nPerhaps replace \"If you cherry-pick...\" with something like \"note: a\nchange authored by a person at some point in time can be committed\nby another person at a different time, and these fields are to\nrecord both persons' contributions separately\", perhaps, if we\nreally want to say more.\n\n> +Git does not store the diff for a commit: when you ask Git for a\n> +diff it calculates it on the fly.\n\nI think this is an attempt to demystify \"are we really storing\nsnapshot for each commit?\" thing, but then \"when you ask Git to show\nthe commit, it calculates the diff from its parent on the fly\" might\nachieve that better, perhaps?\n\n> +[[tree]]\n> +trees::\n> +    A tree is how Git represents a directory. It lists, for each item in\n> +    the tree:\n> ++\n> +[[file-mode]]\n> +1. The *file mode*, for example `100644`. The format is inspired by Unix\n> +   permissions, but Git's modes are much more limited. Git only supports these file modes:\n> ++\n> +  - `100644`: regular file (with type `blob`)\n> +  - `100755`: executable file (with type `blob`)\n> +  - `120000`: symbolic link (with type `blob`)\n> +  - `040000`: directory (with type `tree`)\n> +  - `160000`: gitlink, for use with submodules (with type `commit`)\n\nIt is not really \"supporting\" file modes.  Rather, Git only records\n5 kinds of entities associated with each path in a tree object, and\nuses numbers taht remotely resemble POSIX file modes to represent\nthese 5 kinds.\n\nPerhaps \"supports\" -> \"uses\"?\n\n> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n> +  or <<commit,`commit`>> (a Git submodule, which is a\n> +  commit from a different Git repository)\n> +3. The <<object-id,*object ID*>>\n> +4. The *filename*\n\nHere it may be worth noting that this \"filename\" is a single\npathname component (roughly, what you would see in non-recursive\n\"ls\").  In other words, it may be a directory name.\n\nI wonder if we need to say \"<blob> (a file, or a symbolic link)\"?\n\n> +[[blob]]\n> +blobs::\n> +    A blob is how Git represents a file. A blob object contains the\n> +    file's contents.\n\n\"represents a file\" hints as if the thing may know its name, but\nthat is not the case (its name is given only by surrounding tree).\n\n\"A blob is how Git represents uninterpreted series of bytes, and\nmost commonly used to store file's contents.\" or something, perhaps?\n\n> +When you make a new commit, Git only needs to store new versions of\n> +files which were changed in that commit. This means that commits\n> +can use relatively little disk space even in a very large repository.\n\nThat invites the \"aren't we storing a delta after all, then?\"\nconfusion.\n\n\"Git only needs to newly store new versions of files and\ndirectories.  Files and directories that were not modified by the\ncommit are shared with its parent commit\".\n\n> +NOTE: All of the examples in this section were generated with\n> +`git cat-file -p <object-id>`, which shows the contents of a Git object.\n\nWas this necessary to say this?  Blobs, Commits, and Tags are\ntextual, so \"-p\" does very minimum thing, but Trees are binary\ngarbage, so \"-p\" output is heavily massaged version of the contents.\n\n> +[[branch]]\n> +branches: `refs/heads/<name>`::\n> +    A branch is a name for a commit ID.\n\nWell a commit ID is an alternative way to refer to a commit object\n*name*, so it is a bit strange to say \"a name for a commit ID\".\n\nPerhaps \"A branch ref stores a commit ID.\" is better?\n\n> +[[tag]]\n> +tags: `refs/tags/<name>`::\n> +    A tag is a name for a commit ID, tag object ID, or other object ID.\n\nLikewise.  \"A tag ref stores any kind of object ID, but commonly\nthey are commit objects or tag objects\"\n\n> +    Tags that reference a tag object ID are called \"annotated tags\",\n> +    because the tag object contains a tag message.\n> +    Tags that reference a commit, blob, or tree ID are\n> +    called \"lightweight tags\".\n> ++\n> +Even though branches and tags are both \"a name for a commit ID\", Git\n> +treats them very differently.\n> +Branches are expected to change over time: when you make a commit, Git\n> +will update your <<HEAD,current branch>> to reference the new changes.\n\nThis sentence talks about branch moving because it advances with\nmore commits.  Did we want to say \"HEAD\" here before we explain what\nit is?  \"HEAD\" can move for another reason (i.e. branch switching)\nand using \"HEAD\" in the context of talking about growing history\nmight invite confusion.  I dunno.\n\n> +Tags are usually not changed after they're created.\n\n> +[[HEAD]]\n> +HEAD: `HEAD`::\n> +    `HEAD` is where Git stores your current <<branch,branch>>.\n\nHmm...\n\n> +    `HEAD` can either be:\n> +    1. A symbolic reference to your current branch, for example `ref:\n> +       refs/heads/main` if your current branch is `main`.\n> +    2. A direct reference to a commit ID. This is called \"detached HEAD\n> +\t   state\", see the DETACHED HEAD section of linkgit:git-checkout[1] for more.\n\nThese two are very reasonable.  But \"your current <<branch>>\" refers\nonly to #1.\n\n    `HEAD` refers to the commit your current work is based on, and\n    it is the commit that will become the first parent of the commit\n    once your current work is concluded.  It can either be ...\n\nperhaps.\n\n> +[[remote-tracking-branch]]\n> +remote tracking branches: `refs/remotes/<remote>/<branch>`::\n\nPlease always write \"remote-tracking\" with a hyphen (see glossary).\n\n> +    A remote-tracking branch is a name for a commit ID.\n\nEither \"A remote-tracking branch stores a commit object name\" or \"A\nremote-tracking branch points at a commit object\", followed by \"in\norder to keep track of the last-nown state of ...\" in a single\nsentence.\n\n> +[[index]]\n> +THE INDEX\n> +---------\n> +\n> +The index, also known as the \"staging area\", contains a list of every\n> +file in the repository and its contents. When you commit, the files in\n> +the index are used as the files in the next commit.\n\nIt is hard to define what \"every file in the repository\" really is.\nFiles that you removed last week do not count.  Files added in your\nwip branch elsewhere are obviously not yet in the index when you are\nworking on your primary branch.\n\n> +You can add files to the index or update the version in the index with\n> +linkgit:git-add[1]. Adding a file to the index or updating its version\n> +is called \"staging\" the file for commit.\n\nIt may be worth to clarify by saying \"staging the contents of the\nfile\" (you can edit the file further after you \"git add\") that you\nare taking a snapshot at the time you ran \"git add\", instead of\ngiving a general instruction to \"keey an eye on this file\" to Git\n(if it were, then the next \"git commit\" would behave more like \"git\nadd -u && git commit\").\n\n> +[[reflogs]]\n> +REFLOGS\n> +-------\n> +\n> +Git stores a history called a \"reflog\" for every branch, remote-tracking\n> +branch, and HEAD. This means that if you make a mistake and \"lose\" a\n> +commit, you can generally recover the commit ID by running\n> +`git reflog <reference>`.\n> +\n> +Each reflog entry has:\n> +\n> +1. Before/after *commit IDs*\n> +2. *User* who made the change, for example `Maya <maya@example.com>`\n> +3. *Timestamp* when the change was made\n> +4. *Log message*, for example `pull: Fast-forward`\n> +\n> +Reflogs only log changes made in your local repository.\n> +They are not shared with remotes.\n\nTechnically it is correct that before/after are recorded, but there\nis no way for the end-user to interact with them.  \"git reflog\"\nwalking these entries will only give you a single commit object.\nThe username is also recorded, but I do not think of a way to view\nthe information, let alone using it for querying.\n\nEspecially when the reftable backend is in use, you cannot even read\nthe raw representation like you can do with files backend (where\nsomething like \"cat .git/logs/HEAD\" would let you peek into the\ndetails).  I am not sure if we want to go into this detail.\n\nPerhaps drop everything after \"Each reflog entry has:\"?\n\n> +For example, here's how the reflog for `HEAD` in a repository with 2\n> +commits is stored:\n> +\n> +----\n> +0000000000000000000000000000000000000000 4ccb6d7b8869a86aae2e84c56523f8705b50c647 Maya <maya@example.com> 1759173408 -0400      commit (initial): Initial commit\n> +4ccb6d7b8869a86aae2e84c56523f8705b50c647 750b4ead9c87ceb3ddb7a390e6c7074521797fb3 Maya <maya@example.com> 1759173425 -0400      commit: Add README\n> +----\n> +\n> +GIT\n> +---\n> +Part of the linkgit:git[1] suite\n> diff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc\n> index e423e4765b..20ba121314 100644\n> --- a/Documentation/glossary-content.adoc\n> +++ b/Documentation/glossary-content.adoc\n> @@ -297,8 +297,8 @@ This commit is referred to as a \"merge commit\", or sometimes just a\n>  \tidentified by its <<def_object_name,object name>>. The objects usually\n>  \tlive in `$GIT_DIR/objects/`.\n>  \n> -[[def_object_identifier]]object identifier (oid)::\n> -\tSynonym for <<def_object_name,object name>>.\n> +[[def_object_identifier]]object identifier, object ID, oid::\n> +\tSynonyms for <<def_object_name,object name>>.\n>  \n>  [[def_object_name]]object name::\n>  \tThe unique identifier of an <<def_object,object>>.  The\n> diff --git a/Documentation/meson.build b/Documentation/meson.build\n> index e34965c5b0..ace0573e82 100644\n> --- a/Documentation/meson.build\n> +++ b/Documentation/meson.build\n> @@ -192,6 +192,7 @@ manpages = {\n>    'gitcore-tutorial.adoc' : 7,\n>    'gitcredentials.adoc' : 7,\n>    'gitcvs-migration.adoc' : 7,\n> +  'gitdatamodel.adoc' : 7,\n>    'gitdiffcore.adoc' : 7,\n>    'giteveryday.adoc' : 7,\n>    'gitfaq.adoc' : 7,\n>\n> base-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n"},{"id":"528833","messageId":"xmqqecr3ucba.fsf@gitster.g","threadId":"64244","inReplyTo":"353916d3-c977-40e5-9251-1535b226cc9e@app.fastmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-15T20:42:17Z","receivedAt":"2025-10-15T20:42:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> I'm still not clear on why you think we shouldn't mention that how\n> references behave depends on which filesystem you're using.\n\nSimply because the main purpose of this document is to give a\ndata-model.  A case insensitive filesystem limiting the set of names\nyou can use depending on what other names are in use is a quality of\nimplementation issue, which I view as a mere distraction when we are\ngiving overview at the conceptual level.\n\nThanks.\n"},{"id":"528957","messageId":"52e9036f-6432-46ed-b606-056e5cafe3b9@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqecr3ucba.fsf@gitster.g","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-16T14:21:45Z","receivedAt":"2025-10-16T14:22:06Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"On Wed, Oct 15, 2025, at 4:42 PM, Junio C Hamano wrote:\n> \"Julia Evans\" <julia@jvns.ca> writes:\n>\n>> I'm still not clear on why you think we shouldn't mention that how\n>> references behave depends on which filesystem you're using.\n>\n> Simply because the main purpose of this document is to give a\n> data-model.  A case insensitive filesystem limiting the set of names\n> you can use depending on what other names are in use is a quality of\n> implementation issue, which I view as a mere distraction when we are\n> giving overview at the conceptual level.\n\nOkay, I'll delete the note.\n"},{"id":"528960","messageId":"0eb276ef-7b1a-4e79-93da-13a83226aa01@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqv7kgszr1.fsf@gitster.g","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-16T15:19:46Z","receivedAt":"2025-10-16T15:20:15Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"On Wed, Oct 15, 2025, at 3:58 PM, Junio C Hamano wrote:\n> \"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> +[[commit]]\n>> +commits::\n>> +    A commit contains these required fields\n>> +    (though there are other optional fields):\n>> ++\n>> +1. All the *files* in the commit, stored as the *<<tree,tree>>* ID of\n>> +   the commit's base directory.\n>\n> \"all the files' exact contents at the time of the commit\" is what we\n> mean here, and once readers know what a tree is, the above sentence\n> would be understood as such, but \"All the files\" felt somewhat\n> fuzzy.  I wonder if presenting objects in bottom-up fashion makes it\n> easier to see?  Learn that a blob records exact content of a file,\n> then learn that a tree records the set of paths with exact contents\n> stored at these paths, and after that, learn that a commit records a\n> tree, hence a snapshot of the whole set of contents.  I dunno...\n\nWill try \"The contents of all the *files* in the commit...\" to make it a little\nmore explicit that it's a snapshot.\n\n>> +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n>> +  regular commits have 1 parent, merge commits have 2 or more parents\n>> +3. An *author* and the time the commit was authored\n>> +4. A *committer* and the time the commit was committed.\n>> +   If you cherry-pick (linkgit:git-cherry-pick[1]) someone else's commit,\n>> +   then they will be the author and you'll be the committer.\n>\n> It felt a bit odd to single-out cherry-pick here.\n>\n> I think the important thing to become aware of for the readers at\n> this point is that the author and committer can be different people,\n> and it does not matter how one commits somebody else's patch at the\n> mechanical level.\n>\n> Perhaps replace \"If you cherry-pick...\" with something like \"note: a\n> change authored by a person at some point in time can be committed\n> by another person at a different time, and these fields are to\n> record both persons' contributions separately\", perhaps, if we\n> really want to say more.\n\nI'll just delete the comment about cherry-pick.\nI think it's already obvious (from the fact that are two different fields)\nthat the author and committer can be different (and happen at\ndifferent times), and if we don't want to explain why that might\nhappen there's no need to say more.\n\n>> +Git does not store the diff for a commit: when you ask Git for a\n>> +diff it calculates it on the fly.\n>\n> I think this is an attempt to demystify \"are we really storing\n> snapshot for each commit?\" thing, but then \"when you ask Git to show\n> the commit, it calculates the diff from its parent on the fly\" might\n> achieve that better, perhaps?\n\nSure, can change it to that.\n\n>> +[[tree]]\n>> +trees::\n>> +    A tree is how Git represents a directory. It lists, for each item in\n>> +    the tree:\n>> ++\n>> +[[file-mode]]\n>> +1. The *file mode*, for example `100644`. The format is inspired by Unix\n>> +   permissions, but Git's modes are much more limited. Git only supports these file modes:\n>> ++\n>> +  - `100644`: regular file (with type `blob`)\n>> +  - `100755`: executable file (with type `blob`)\n>> +  - `120000`: symbolic link (with type `blob`)\n>> +  - `040000`: directory (with type `tree`)\n>> +  - `160000`: gitlink, for use with submodules (with type `commit`)\n>\n> It is not really \"supporting\" file modes.  Rather, Git only records\n> 5 kinds of entities associated with each path in a tree object, and\n> uses numbers taht remotely resemble POSIX file modes to represent\n> these 5 kinds.\n>\n> Perhaps \"supports\" -> \"uses\"?\n\n\"Uses\" sounds good to me.\n\n>> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n>> +  or <<commit,`commit`>> (a Git submodule, which is a\n>> +  commit from a different Git repository)\n>> +3. The <<object-id,*object ID*>>\n>> +4. The *filename*\n>\n> Here it may be worth noting that this \"filename\" is a single\n> pathname component (roughly, what you would see in non-recursive\n> \"ls\").  In other words, it may be a directory name.\n>\n> I wonder if we need to say \"<blob> (a file, or a symbolic link)\"?\n\nI'm inclined to leave this alone because arguably a symbolic link is\na file but I don't feel strongly about this.\n\n>> +[[blob]]\n>> +blobs::\n>> +    A blob is how Git represents a file. A blob object contains the\n>> +    file's contents.\n>\n> \"represents a file\" hints as if the thing may know its name, but\n> that is not the case (its name is given only by surrounding tree).\n>\n> \"A blob is how Git represents uninterpreted series of bytes, and\n> most commonly used to store file's contents.\" or something, perhaps?\n\nI'll say \"A blob is how Git represents a file's contents\", unless Git has\nanother use for blobs that I don't know about (I think it's not\nthat much of a stretch to say that a symbolic link is a special kind\nof file where the \"contents\" are the the link destination).\n\nI think it's always clearer to be more specific when possible, if there's only\none purpose for blobs it's unnecessary (and IMO a bit misleading, because\nit makes the reader wonder if there are other purposes that they should\nknow about) to say that blobs can be used to store any arbitrary bytes for\nany purpose.\n\nIf there is another purpose I think we should give an example.\n\n>> +When you make a new commit, Git only needs to store new versions of\n>> +files which were changed in that commit. This means that commits\n>> +can use relatively little disk space even in a very large repository.\n>\n> That invites the \"aren't we storing a delta after all, then?\"\n> confusion.\n>\n> \"Git only needs to newly store new versions of files and\n> directories.  Files and directories that were not modified by the\n> commit are shared with its parent commit\".\n\nI agree it makes it sound a little bit like we're storing a delta.\nWill think about how to phrase this differently.\n\n>> +NOTE: All of the examples in this section were generated with\n>> +`git cat-file -p <object-id>`, which shows the contents of a Git object.\n>\n> Was this necessary to say this?  Blobs, Commits, and Tags are\n> textual, so \"-p\" does very minimum thing, but Trees are binary\n> garbage, so \"-p\" output is heavily massaged version of the contents.\n\nAh, I didn't know how trees were stored, thanks. \nI can remove \"which shows the contents of a Git object\", people\ncan read the man page for `git cat-file` if they want details.\n\n>> +[[branch]]\n>> +branches: `refs/heads/<name>`::\n>> +    A branch is a name for a commit ID.\n>\n> Well a commit ID is an alternative way to refer to a commit object\n> *name*, so it is a bit strange to say \"a name for a commit ID\".\n>\n> Perhaps \"A branch ref stores a commit ID.\" is better?\n\nI think I'll leave this alone, none of the many test readers reported\nbeing confused by it.\n\n>> +[[tag]]\n>> +tags: `refs/tags/<name>`::\n>> +    A tag is a name for a commit ID, tag object ID, or other object ID.\n>\n> Likewise.  \"A tag ref stores any kind of object ID, but commonly\n> they are commit objects or tag objects\"\n>\n>> +    Tags that reference a tag object ID are called \"annotated tags\",\n>> +    because the tag object contains a tag message.\n>> +    Tags that reference a commit, blob, or tree ID are\n>> +    called \"lightweight tags\".\n>> ++\n>> +Even though branches and tags are both \"a name for a commit ID\", Git\n>> +treats them very differently.\n>> +Branches are expected to change over time: when you make a commit, Git\n>> +will update your <<HEAD,current branch>> to reference the new changes.\n>\n> This sentence talks about branch moving because it advances with\n> more commits.  Did we want to say \"HEAD\" here before we explain what\n> it is?  \"HEAD\" can move for another reason (i.e. branch switching)\n> and using \"HEAD\" in the context of talking about growing history\n> might invite confusion.  I dunno.\n\nThe text says \"current branch\", it just cross-references the \"HEAD\" section in the\nHTML version if someone wants to read about what is meant by \"current branch\".\n\n>> +Tags are usually not changed after they're created.\n>\n>> +[[HEAD]]\n>> +HEAD: `HEAD`::\n>> +    `HEAD` is where Git stores your current <<branch,branch>>.\n>\n> Hmm...\n>\n>> +    `HEAD` can either be:\n>> +    1. A symbolic reference to your current branch, for example `ref:\n>> +       refs/heads/main` if your current branch is `main`.\n>> +    2. A direct reference to a commit ID. This is called \"detached HEAD\n>> +\t   state\", see the DETACHED HEAD section of linkgit:git-checkout[1] for more.\n>\n> These two are very reasonable.  But \"your current <<branch>>\" refers\n> only to #1.\n>\n>     `HEAD` refers to the commit your current work is based on, and\n>     it is the commit that will become the first parent of the commit\n>     once your current work is concluded.  It can either be ...\n>\n> perhaps.\n\nI like the idea of mentioning that HEAD will be the parent commit\nof any commit that you make. Will think about how to incorporate\nthat, and about how to resolve \" `HEAD` is where Git stores your\ncurrent <<branch,branch>>.\" being not exactly true.\n\n>> +[[remote-tracking-branch]]\n>> +remote tracking branches: `refs/remotes/<remote>/<branch>`::\n>\n> Please always write \"remote-tracking\" with a hyphen (see glossary).\n\nWill fix.\n\n>> +    A remote-tracking branch is a name for a commit ID.\n>\n> Either \"A remote-tracking branch stores a commit object name\" or \"A\n> remote-tracking branch points at a commit object\", followed by \"in\n> order to keep track of the last-nown state of ...\" in a single\n> sentence.\n\nI see that you don't like the \"name for a commit ID\" phrasing :)\nMaybe there's another way to say it, though again none of the test\nreaders said they were confused by this or disagreed with the phrasing.\n\n>> +[[index]]\n>> +THE INDEX\n>> +---------\n>> +\n>> +The index, also known as the \"staging area\", contains a list of every\n>> +file in the repository and its contents. When you commit, the files in\n>> +the index are used as the files in the next commit.\n>\n> It is hard to define what \"every file in the repository\" really is.\n> Files that you removed last week do not count.  Files added in your\n> wip branch elsewhere are obviously not yet in the index when you are\n> working on your primary branch.\n\nAgreed, I'm not so happy with \"every file in the repository\" either.\nMy intent was to make it clear that it's not \"just the files you `git add`ed\".\nI'll think about a different phrasing that communicates the same thing.\nPerhaps mentioning how it relates to the HEAD commit would help.\n\n>> +You can add files to the index or update the version in the index with\n>> +linkgit:git-add[1]. Adding a file to the index or updating its version\n>> +is called \"staging\" the file for commit.\n>\n> It may be worth to clarify by saying \"staging the contents of the\n> file\" (you can edit the file further after you \"git add\") that you\n> are taking a snapshot at the time you ran \"git add\", instead of\n> giving a general instruction to \"keey an eye on this file\" to Git\n> (if it were, then the next \"git commit\" would behave more like \"git\n> add -u && git commit\").\n\nMaybe, will think about this too.\n\n>> +[[reflogs]]\n>> +REFLOGS\n>> +-------\n>> +\n>> +Git stores a history called a \"reflog\" for every branch, remote-tracking\n>> +branch, and HEAD. This means that if you make a mistake and \"lose\" a\n>> +commit, you can generally recover the commit ID by running\n>> +`git reflog <reference>`.\n>> +\n>> +Each reflog entry has:\n>> +\n>> +1. Before/after *commit IDs*\n>> +2. *User* who made the change, for example `Maya <maya@example.com>`\n>> +3. *Timestamp* when the change was made\n>> +4. *Log message*, for example `pull: Fast-forward`\n>> +\n>> +Reflogs only log changes made in your local repository.\n>> +They are not shared with remotes.\n>\n> Technically it is correct that before/after are recorded, but there\n> is no way for the end-user to interact with them.  \"git reflog\"\n> walking these entries will only give you a single commit object.\n> The username is also recorded, but I do not think of a way to view\n> the information, let alone using it for querying.\n\nYou can view the username with git reflog --format=\"%gn <%ge>\".\n(according to `man git-log`). I don't see a way to view the old commit ID.\n\nPerhaps we should include the username but not the old commit ID then.\nI'm not sure.\n\n> Especially when the reftable backend is in use, you cannot even read\n> the raw representation like you can do with files backend (where\n> something like \"cat .git/logs/HEAD\" would let you peek into the\n> details).  I am not sure if we want to go into this detail.\n>\n> Perhaps drop everything after \"Each reflog entry has:\"?\n\nPerhaps we could give a stripped down list, like\n\n1. The new *commit ID* the reference points to\n2. *Timestamp* when the change was made\n3. *Log message*, for example `pull: Fast-forward`\n\nAnd then instead of giving the contents of `.git/logs/HEAD`\n(which as you say includes some fields that there's no way\nfor the user to interact with), instead we could just show the\noutput of `git reflog main`, like this:\n\n    You can view the reflog for `git reflog`, for example here's the reflog\n    for a `main` branch which has changed twice:\n\n    $ git reflog main --date=iso --no-decorate\n    750b4ea main@{2025-09-29 15:17:05 -0400}: commit: Add README\n    4ccb6d7 main@{2025-09-29 15:16:48 -0400}: commit (initial): Initial commit\n\nI added `--no-decorate`  there because the decorations are a distraction\nwhen talking about the data model.\n\nThis version omits the username which is a little weird (it is possible to\naccess the username) but mentioning the username is a little weird\ntoo because it raises some questions that are hard to answer about\nwhat that field is for, and you have to pass an obscure format string\nto view it. Not sure what's best here.\n\n>> +For example, here's how the reflog for `HEAD` in a repository with 2\n>> +commits is stored:\n>> +\n>> +----\n>> +0000000000000000000000000000000000000000 4ccb6d7b8869a86aae2e84c56523f8705b50c647 Maya <maya@example.com> 1759173408 -0400      commit (initial): Initial commit\n>> +4ccb6d7b8869a86aae2e84c56523f8705b50c647 750b4ead9c87ceb3ddb7a390e6c7074521797fb3 Maya <maya@example.com> 1759173425 -0400      commit: Add README\n>> +----\n\nThanks for the review.\n- Julia\n"},{"id":"528961","messageId":"0ec8192d-558c-4caa-9d18-0e0c1e1203ca@app.fastmail.com","threadId":"64244","inReplyTo":"pull.1981.v3.git.1760476346040.gitgitgadget@gmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-10-16T15:24:12Z","receivedAt":"2025-10-16T15:25:30Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"> [PATCH v3] doc: add a explanation of Git's data model\n\ns/a explanation/an explanation/\n\nOn Tue, Oct 14, 2025, at 23:12, Julia Evans via GitGitGadget wrote:\n> From: Julia Evans <julia@jvns.ca>\n>\n> Git very often uses the terms \"object\", \"reference\", or \"index\" in its\n> documentation.\n>[snip]\n"},{"id":"528972","messageId":"xmqq347i948a.fsf@gitster.g","threadId":"64244","inReplyTo":"0eb276ef-7b1a-4e79-93da-13a83226aa01@app.fastmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-16T16:54:45Z","receivedAt":"2025-10-16T16:54:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n>>> +[[tree]]\n>>> +trees::\n>>> +    A tree is how Git represents a directory. It lists, for each item in\n>>> +    the tree:\n>>> ++\n>>> +[[file-mode]]\n>>> +1. The *file mode*, for example `100644`. The format is inspired by Unix\n>>> +   permissions, but Git's modes are much more limited. Git only supports these file modes:\n>>> ++\n>>> +  - `100644`: regular file (with type `blob`)\n>>> +  - `100755`: executable file (with type `blob`)\n>>> +  - `120000`: symbolic link (with type `blob`)\n>>> +  - `040000`: directory (with type `tree`)\n>>> +  - `160000`: gitlink, for use with submodules (with type `commit`)\n>>\n>> It is not really \"supporting\" file modes.  Rather, Git only records\n>> 5 kinds of entities associated with each path in a tree object, and\n>> uses numbers taht remotely resemble POSIX file modes to represent\n>> these 5 kinds.\n>>\n>> Perhaps \"supports\" -> \"uses\"?\n>\n> \"Uses\" sounds good to me.\n\nAlso \"much more limited\" is misleading.  We only represent 5 kinds\nof things, so we use only 5 mode-bits-looking numbers.\n\n>>> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n>>> +  or <<commit,`commit`>> (a Git submodule, which is a\n>>> +  commit from a different Git repository)\n>>> +3. The <<object-id,*object ID*>>\n>>> +4. The *filename*\n>>\n>> Here it may be worth noting that this \"filename\" is a single\n>> pathname component (roughly, what you would see in non-recursive\n>> \"ls\").  In other words, it may be a directory name.\n\nComments?\n\n>>> +[[blob]]\n>>> +blobs::\n>>> +    A blob is how Git represents a file. A blob object contains the\n>>> +    file's contents.\n>>\n>> \"represents a file\" hints as if the thing may know its name, but\n>> that is not the case (its name is given only by surrounding tree).\n>>\n>> \"A blob is how Git represents uninterpreted series of bytes, and\n>> most commonly used to store file's contents.\" or something, perhaps?\n>\n> I'll say \"A blob is how Git represents a file's contents\", unless Git has\n> another use for blobs that I don't know about (I think it's not\n> that much of a stretch to say that a symbolic link is a special kind\n> of file where the \"contents\" are the the link destination).\n\nA few configuration variables like mailmap.blob name a blob object,\nfor which _only_ its contents, i.e., the sequence of bytes, matter\nand where they originally were stored does not matter.\n\nBut we are falling into the area of tautology, as any sequence of\nbytes can be stored in a file so they can be called \"contents of a\nfile\".  But the point is that these bytes do not have to be stored\nto become a blob (think: \"git cat-file -t blob -w --stdin\").\n\n> I think it's always clearer to be more specific when possible, if there's only\n> one purpose for blobs it's unnecessary (and IMO a bit misleading, because\n> it makes the reader wonder if there are other purposes that they should\n> know about) to say that blobs can be used to store any arbitrary bytes for\n> any purpose.\n\nI do not think describing other use cases is unnecessary.  Even if\nwe limit ourselves to discuss a single purpose for blob, i.e. to\nrepresent the contents of a file, we should stress that blob is to\nstore _only_ contents, and not other aspects of the file (e.g., in\nwhat paths with what mode), and that is where my reaction to \"how\nGit reprsents a file\" comes from.\n\n>>> +[[branch]]\n>>> +branches: `refs/heads/<name>`::\n>>> +    A branch is a name for a commit ID.\n>>\n>> Well a commit ID is an alternative way to refer to a commit object\n>> *name*, so it is a bit strange to say \"a name for a commit ID\".\n>>\n>> Perhaps \"A branch ref stores a commit ID.\" is better?\n>\n> I think I'll leave this alone, none of the many test readers reported\n> being confused by it.\n\nWould a confused person report that they are confused? ;-)\n\n> I see that you don't like the \"name for a commit ID\" phrasing :)\n> Maybe there's another way to say it, though again none of the test\n> readers said they were confused by this or disagreed with the phrasing.\n\nYes, I get that given \"refs/heads/main\", you want to say \"main\" is\none of the ways to have repo_get_oid() to yield the commit object,\nand you are using \"name\" in that sense, but it is more like a ref\ncan be used to name an object.  It is *not* the name of the object,\nbecause the object can have other names, and more importantly, it\n(i.e., to give a name for an object) is not the only thing that a\nref can do.  And that is why I do not like that phrasing, combined\nwith the target of giving that name is spelled \"a commit ID\".  The\ncommit ID is already another way to name the thing the refname can\nbe also used to name: a commit object.  A commit object and a commit\nobject name are different things.  The latter is a name that can\nrefer to the former.  And a ref can be used just like the latter to\nrefer to the former (i.e. \"commit object\").\n\nBy the way, I do like the way many of your responses are \"will think\nabout it more\", not \"I'll take your version\".\n\nVery much appreciated.\n\nThanks.\n"},{"id":"528989","messageId":"03db91a6-148b-436f-8afa-0273a1f5d508@app.fastmail.com","threadId":"64244","inReplyTo":"xmqq347i948a.fsf@gitster.g","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-16T18:59:01Z","receivedAt":"2025-10-16T19:00:15Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"On Thu, Oct 16, 2025, at 12:54 PM, Junio C Hamano wrote:\n> \"Julia Evans\" <julia@jvns.ca> writes:\n>\n>>>> +[[tree]]\n>>>> +trees::\n>>>> +    A tree is how Git represents a directory. It lists, for each item in\n>>>> +    the tree:\n>>>> ++\n>>>> +[[file-mode]]\n>>>> +1. The *file mode*, for example `100644`. The format is inspired by Unix\n>>>> +   permissions, but Git's modes are much more limited. Git only supports these file modes:\n>>>> ++\n>>>> +  - `100644`: regular file (with type `blob`)\n>>>> +  - `100755`: executable file (with type `blob`)\n>>>> +  - `120000`: symbolic link (with type `blob`)\n>>>> +  - `040000`: directory (with type `tree`)\n>>>> +  - `160000`: gitlink, for use with submodules (with type `commit`)\n>>>\n>>> It is not really \"supporting\" file modes.  Rather, Git only records\n>>> 5 kinds of entities associated with each path in a tree object, and\n>>> uses numbers taht remotely resemble POSIX file modes to represent\n>>> these 5 kinds.\n>>>\n>>> Perhaps \"supports\" -> \"uses\"?\n>>\n>> \"Uses\" sounds good to me.\n>\n> Also \"much more limited\" is misleading.  We only represent 5 kinds\n> of things, so we use only 5 mode-bits-looking numbers.\n\nWhat does it mislead the reader to think? My goal is to communicate that\nif you want to tell Git to remember that a file's Unix permissions were\n700, that's not possible.\n\n>>>> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n>>>> +  or <<commit,`commit`>> (a Git submodule, which is a\n>>>> +  commit from a different Git repository)\n>>>> +3. The <<object-id,*object ID*>>\n>>>> +4. The *filename*\n>>>\n>>> Here it may be worth noting that this \"filename\" is a single\n>>> pathname component (roughly, what you would see in non-recursive\n>>> \"ls\").  In other words, it may be a directory name.\n>\n> Comments?\n\nOops, missed this in my first pass.\n\nI looked at them man pages for a couple of commands (\"mv\", \"cp\")\nand it looks like it's normal to refer to files and directories jointly\nas \"files\", or refer to them as having a \"file name\". So I think it's okay\nto call it a \"file name\" even if the \"file\" may be a directory.\n\n>>>> +[[blob]]\n>>>> +blobs::\n>>>> +    A blob is how Git represents a file. A blob object contains the\n>>>> +    file's contents.\n>>>\n>>> \"represents a file\" hints as if the thing may know its name, but\n>>> that is not the case (its name is given only by surrounding tree).\n>>>\n>>> \"A blob is how Git represents uninterpreted series of bytes, and\n>>> most commonly used to store file's contents.\" or something, perhaps?\n>>\n>> I'll say \"A blob is how Git represents a file's contents\", unless Git has\n>> another use for blobs that I don't know about (I think it's not\n>> that much of a stretch to say that a symbolic link is a special kind\n>> of file where the \"contents\" are the the link destination).\n>\n> A few configuration variables like mailmap.blob name a blob object,\n> for which _only_ its contents, i.e., the sequence of bytes, matter\n> and where they originally were stored does not matter.\n>\n> But we are falling into the area of tautology, as any sequence of\n> bytes can be stored in a file so they can be called \"contents of a\n> file\".  But the point is that these bytes do not have to be stored\n> to become a blob (think: \"git cat-file -t blob -w --stdin\").\n\nI'm trying to think through what the goal of explaining the nature of\na \"blob\" is.\n\nTo me describing blobs primarily as \"bytes\" makes it sound a bit like\n\"Git will treat this as opaque binary data, Git will not attempt to\ninterpret the contents of a blob in any way\" (which is certainly true\nfor many blob storage systems!).\n\nBut it's not true that Git treats blobs as opaque binary data, unlike\nother blob storage systems, Git has diff and merge algorithms to\ninterpret the contents of the file to some extent and try to do useful\nthings with them.\n\nAnother goal we could have is to be clear that there are no limits to\nwhat kind of files you can store in Git: you can equally well store text\nfiles and binary files.\n\n>> I think it's always clearer to be more specific when possible, if there's only\n>> one purpose for blobs it's unnecessary (and IMO a bit misleading, because\n>> it makes the reader wonder if there are other purposes that they should\n>> know about) to say that blobs can be used to store any arbitrary bytes for\n>> any purpose.\n>\n> I do not think describing other use cases is unnecessary.  Even if\n> we limit ourselves to discuss a single purpose for blob, i.e. to\n> represent the contents of a file, we should stress that blob is to\n> store _only_ contents, and not other aspects of the file (e.g., in\n> what paths with what mode), and that is where my reaction to \"how\n> Git reprsents a file\" comes from.\n\nI think it does make sense to say the blob stores only the contents,\nthough IMO that's fairly clear already since we've already explained\nwhere the other parts of the file are stored by the time we get to\nexplaining \"blob\".\n\n>>>> +[[branch]]\n>>>> +branches: `refs/heads/<name>`::\n>>>> +    A branch is a name for a commit ID.\n>>>\n>>> Well a commit ID is an alternative way to refer to a commit object\n>>> *name*, so it is a bit strange to say \"a name for a commit ID\".\n>>>\n>>> Perhaps \"A branch ref stores a commit ID.\" is better?\n>>\n>> I think I'll leave this alone, none of the many test readers reported\n>> being confused by it.\n>\n> Would a confused person report that they are confused? ;-)\n\nEveryone leaving feedback gets a prompt something like this\nasking them to categorize their feedback,\nand \"I'm confused\" is one of the options.\nhttps://jvns.ca/images/feedback-categories.png\n\nI definitely got many \"I'm confused\" and \"I have a question\"\ncomments about other things that were confusing to readers.\n\n>> I see that you don't like the \"name for a commit ID\" phrasing :)\n>> Maybe there's another way to say it, though again none of the test\n>> readers said they were confused by this or disagreed with the phrasing.\n>\n> Yes, I get that given \"refs/heads/main\", you want to say \"main\" is\n> one of the ways to have repo_get_oid() to yield the commit object,\n> and you are using \"name\" in that sense, but it is more like a ref\n> can be used to name an object.  It is *not* the name of the object,\n> because the object can have other names, and more importantly, it\n> (i.e., to give a name for an object) is not the only thing that a\n> ref can do.  \n\nThat's interesting,  what else can a ref do other than to give a name to\nan object?\n\n> And that is why I do not like that phrasing, combined\n> with the target of giving that name is spelled \"a commit ID\".  The\n> commit ID is already another way to name the thing the refname can\n> be also used to name: a commit object.  A commit object and a commit\n> object name are different things.  The latter is a name that can\n> refer to the former.\n\nI'm curious about why it's important to you to make this distinction\nbetween a commit ID and a commit object. To me the commit ID and the\ncommit object come as a package, since the commit ID is calculated from\nthe commit object.\n\n>  And a ref can be used just like the latter to\n> refer to the former (i.e. \"commit object\").\n\n> By the way, I do like the way many of your responses are \"will think\n> about it more\", not \"I'll take your version\".\n>\n> Very much appreciated.\n\nI'm glad to hear that! It's a fun puzzle to figure out how to express\nthings clearly and accurately and concisely.\n\n- Julia\n"},{"id":"529009","messageId":"xmqqbjm63751.fsf@gitster.g","threadId":"64244","inReplyTo":"03db91a6-148b-436f-8afa-0273a1f5d508@app.fastmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-16T20:48:26Z","receivedAt":"2025-10-16T20:48:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n>>>> It is not really \"supporting\" file modes.  Rather, Git only records\n>>>> 5 kinds of entities associated with each path in a tree object, and\n>>>> uses numbers taht remotely resemble POSIX file modes to represent\n>>>> these 5 kinds.\n>>>>\n>>>> Perhaps \"supports\" -> \"uses\"?\n>>>\n>>> \"Uses\" sounds good to me.\n>>\n>> Also \"much more limited\" is misleading.  We only represent 5 kinds\n>> of things, so we use only 5 mode-bits-looking numbers.\n>\n> What does it mislead the reader to think? My goal is to communicate that\n> if you want to tell Git to remember that a file's Unix permissions were\n> 700, that's not possible.\n\nYes, rewording \"support\" to \"use\" is one good way to do so.  But\n\"limited\" implies that lifting the limitation would allow you to\nstore more.  That is the misguided thinking I want to avoid here.\nThere is no limitations to lift.  We only differentiate 5 kinds\nhence we only use 5 permission-bit-looking numbers.  We do not\ndifferenciate a file with permission 0600 from aother with 0644.\n\n>>>> Here it may be worth noting that this \"filename\" is a single\n>>>> pathname component (roughly, what you would see in non-recursive\n>>>> \"ls\").  In other words, it may be a directory name.\n>>\n>> Comments?\n>\n> Oops, missed this in my first pass.\n>\n> I looked at them man pages for a couple of commands (\"mv\", \"cp\")\n> and it looks like it's normal to refer to files and directories jointly\n> as \"files\", or refer to them as having a \"file name\". So I think it's okay\n> to call it a \"file name\" even if the \"file\" may be a directory.\n\nAh, not that part.  I was more interested in seeing how we express\n\"in these names, there won't be any slashes\".\n\n>>>>> +[[blob]]\n>>>>> +blobs::\n\nBy the way, I kept forgetting to mention, but why are all of these\nlisted terms plural (not just object types but also \"branches\" and\n\"tags\"?\n\n> But it's not true that Git treats blobs as opaque binary data, unlike\n> other blob storage systems, Git has diff and merge algorithms to\n> interpret the contents of the file to some extent and try to do useful\n> things with them.\n\nYes, but diff and merge happens way above the object layer, where\nthe question \"what is blob\" has a meaning.  And these \"blobs are\nrecorded in a tree together with other blobs and trees recursively,\nand the single top-level tree describes a snapshot of a single\nstate, which is recorded in a commit\" data model descriptions is\nexactly about the lower-level object layer.\n\n> Another goal we could have is to be clear that there are no limits to\n> what kind of files you can store in Git: you can equally well store text\n> files and binary files.\n\nThat is a natural consequence of blobs being nothing more than\nuninterpreted sequence of bytes.\n\n>>> I see that you don't like the \"name for a commit ID\" phrasing :)\n>>> Maybe there's another way to say it, though again none of the test\n>>> readers said they were confused by this or disagreed with the phrasing.\n>>\n>> Yes, I get that given \"refs/heads/main\", you want to say \"main\" is\n>> one of the ways to have repo_get_oid() to yield the commit object,\n>> and you are using \"name\" in that sense, but it is more like a ref\n>> can be used to name an object.  It is *not* the name of the object,\n>> because the object can have other names, and more importantly, it\n>> (i.e., to give a name for an object) is not the only thing that a\n>> ref can do.  \n>\n> That's interesting,  what else can a ref do other than to give a name to\n> an object?\n\nFor example, a ref is a key to reflog, so obvoiusly it is more than\njust a single commit.  If you say \"git checkout main\" and \"git\ncheckout main^{commit}\", they refer to the same commit, but the\nformer is a sign that you want the next commit you make from that\nstate to grow that branch (and not any other branch you may have\nthat happen to be pointing at the same commit), while the other one\nis not.\n\n>> And that is why I do not like that phrasing, combined\n>> with the target of giving that name is spelled \"a commit ID\".  The\n>> commit ID is already another way to name the thing the refname can\n>> be also used to name: a commit object.  A commit object and a commit\n>> object name are different things.  The latter is a name that can\n>> refer to the former.\n>\n> I'm curious about why it's important to you to make this distinction\n> between a commit ID and a commit object. To me the commit ID and the\n> commit object come as a package, since the commit ID is calculated from\n> the commit object.\n\nIt may be the most natural name for the commit object, but that does\nnot mean the name is the object.  Let's not go phylosophical.\n"},{"id":"529178","messageId":"c1c456b5-aca7-4b24-a4a2-558405214f24@app.fastmail.com","threadId":"64244","inReplyTo":"pull.1981.v3.git.1760476346040.gitgitgadget@gmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-10-20T16:37:30Z","receivedAt":"2025-10-20T16:37:51Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"On Tue, Oct 14, 2025, at 23:12, Julia Evans via GitGitGadget wrote:\n> From: Julia Evans <julia@jvns.ca>\n>\n> Git very often uses the terms \"object\", \"reference\", or \"index\" in its\n> documentation.\n>\n> However, it's hard to find a clear explanation of these terms and how\n> they relate to each other in the documentation. The closest candidates\n> currently are:\n>[snip]\n\nFor some reason I get an error with `Documentation/doc-diff` when run\nagainst 446c8a72 (Merge branch 'je/doc-data-model' into seen,\n2025-10-16).  Here I’m comparing with `master`.\n\n    $ ./doc-diff 4253630c6f07a4bdcc9aa62a50e26a4d466219d1 446c8a72be6cf1b6121e643590a9acacfc21c5fb\n    Previous HEAD position was b20e48e0232 doc: add a explanation of Git's data model\n    HEAD is now at 446c8a72be6 Merge branch 'je/doc-data-model' into seen\n    make: Entering directory '<git repo>/Documentation/tmp-doc-diff/worktree'\n    install -d -m 755 '<git repo>/Documentation/tmp-doc-diff/installed/446c8a72be6cf1b6121e643590a9acacfc21c5fb+/home/kristoffer/share/man/man3'\n    (cd perl/build/man/man3 && tar cf - .) | \\\n    (cd '<git repo>/Documentation/tmp-doc-diff/installed/446c8a72be6cf1b6121e643590a9acacfc21c5fb+/home/kristoffer/share/man/man3' && umask 022 && tar xof -)\n    make -C Documentation install-man\n    make[1]: Entering directory '<git repo>/Documentation/tmp-doc-diff/worktree/Documentation'\n        GEN cmd-list.made\n        GEN doc.dep\n        GEN asciidoc.conf\n        ASCIIDOC git-add.xml\n        ASCIIDOC git-config.xml\n        ASCIIDOC git-diff-tree.xml\n        ASCIIDOC git-fast-import.xml\n        ASCIIDOC git-fetch.xml\n        ASCIIDOC git-fsck.xml\n        ASCIIDOC git-log.xml\n        ASCIIDOC git-merge-tree.xml\n        ASCIIDOC git-patch-id.xml\n        ASCIIDOC git-pull.xml\n        ASCIIDOC git-push.xml\n        ASCIIDOC git-replay.xml\n        ASCIIDOC git-repo.xml\n        ASCIIDOC git-rev-list.xml\n        ASCIIDOC git-rev-parse.xml\n        ASCIIDOC git-shortlog.xml\n        ASCIIDOC git-show.xml\n        ASCIIDOC git-sparse-checkout.xml\n        ASCIIDOC git-stash.xml\n        ASCIIDOC git-tag.xml\n        ASCIIDOC git-worktree.xml\n        ASCIIDOC git.xml\n        ASCIIDOC gitformat-loose.xml\n        ASCIIDOC gitformat-pack.xml\n        ASCIIDOC gitcli.xml\n        ASCIIDOC gitcredentials.xml\n        XMLTO gitdatamodel.7\n        XMLTO git-add.1\n        XMLTO git-diff-tree.1\n        XMLTO git-fast-import.1\n        XMLTO git-fetch.1\n        XMLTO git-fsck.1\n    xmlto: <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml does not validate (status 3)\n    xmlto: Fix document syntax or use --skip-validation option\n    <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml:71: element link: validity error : IDREF attribute linkend references an unknown ID \"tree\"\n    <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml:96: element link: validity error : IDREF attribute linkend references an unknown ID \"tree\"\n    <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml:397: element link: validity error : IDREF attribute linkend references an unknown ID \"tree\"\n    Document <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml does not validate\n    make[1]: *** [Makefile:380: gitdatamodel.7] Error 13\n    make[1]: *** Waiting for unfinished jobs....\n    make[1]: Leaving directory '<git repo>/Documentation/tmp-doc-diff/worktree/Documentation'\n    make: *** [Makefile:3676: install-man] Error 2\n    make: Leaving directory '<git repo>/Documentation/tmp-doc-diff/worktree'\n\nThe syntax looks correct.  So I don’t know what is wrong.  `make html`\nworks *and* makes the link.\n\nAt first look it might be to do with the anchor on a definition list but\nI tried removing the anchors and expected to get an error for `blob`\nnext.  But that didn’t happen.\n\nIn short I don’t see what is special about `tree`.\n"},{"id":"529186","messageId":"xmqqy0p5zc3n.fsf@gitster.g","threadId":"64244","inReplyTo":"c1c456b5-aca7-4b24-a4a2-558405214f24@app.fastmail.com","subject":"Re: [PATCH v3] doc: add a explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-20T18:01:32Z","receivedAt":"2025-10-20T18:01:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Kristoffer Haugsbakk\" <kristofferhaugsbakk@fastmail.com> writes:\n\n>     xmlto: <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml does not validate (status 3)\n>     xmlto: Fix document syntax or use --skip-validation option\n>     <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml:71: element link: validity error : IDREF attribute linkend references an unknown ID \"tree\"\n>     <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml:96: element link: validity error : IDREF attribute linkend references an unknown ID \"tree\"\n>     <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml:397: element link: validity error : IDREF attribute linkend references an unknown ID \"tree\"\n>     Document <git repo>/Documentation/tmp-doc-diff/worktree/Documentation/gitdatamodel.xml does not validate\n>     make[1]: *** [Makefile:380: gitdatamodel.7] Error 13\n>     make[1]: *** Waiting for unfinished jobs....\n>     make[1]: Leaving directory '<git repo>/Documentation/tmp-doc-diff/worktree/Documentation'\n>     make: *** [Makefile:3676: install-man] Error 2\n>     make: Leaving directory '<git repo>/Documentation/tmp-doc-diff/worktree'\n>\n> The syntax looks correct.  So I don’t know what is wrong.  `make html`\n> works *and* makes the link.\n>\n> At first look it might be to do with the anchor on a definition list but\n> I tried removing the anchors and expected to get an error for `blob`\n> next.  But that didn’t happen.\n>\n> In short I don’t see what is special about `tree`.\n\nThis seems to work it around without breaking .html generation too\nbadly for AsciiDoc and without breaking .7/.html generation for\nAsciidoctor.  Generation of .7 were broken with AsciiDoc so we\ncannot complain even if the result is suboptimal, but the generated\nmanpage with this patch using AsciiDoc did not look too bad, either.\n\nI do not know AsciiDoc internals (and I am not particularly\ninterested to learn it now), but I am guessing that the bug is that\nwhen it sees [[tree]], it tries to find an element to put id=\"tree\",\nbut before it finds any approprifate one, it sees [[filemode]] and\nuses the element it finds to hold id=\"filemode\", losing sight of the\nneed to add id=\"tree\" somewhere.\n\n\n\n Documentation/gitdatamodel.adoc | 4 +++-\n 1 file changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nindex f49574dfae..7232fe3861 100644\n--- a/Documentation/gitdatamodel.adoc\n+++ b/Documentation/gitdatamodel.adoc\n@@ -83,8 +83,10 @@ trees::\n     A tree is how Git represents a directory. It lists, for each item in\n     the tree:\n +\n+1. The *file mode*, for example `100644`.\n++\n [[file-mode]]\n-1. The *file mode*, for example `100644`. The format is inspired by Unix\n+The format is inspired by Unix\n    permissions, but Git's modes are much more limited. Git only supports these file modes:\n +\n   - `100644`: regular file (with type `blob`)\n-- \n2.51.1-556-g06b2a500e9\n\n"},{"id":"529751","messageId":"pull.1981.v4.git.1761593537924.gitgitgadget@gmail.com","threadId":"64244","inReplyTo":"pull.1981.v3.git.1760476346040.gitgitgadget@gmail.com","subject":"[PATCH v4] doc: add an explanation of Git's data model","fromName":"Julia Evans via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-27T19:32:17Z","receivedAt":"2025-10-27T19:32:23Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"From: Julia Evans <julia@jvns.ca>\n\nGit very often uses the terms \"object\", \"reference\", or \"index\" in its\ndocumentation.\n\nHowever, it's hard to find a clear explanation of these terms and how\nthey relate to each other in the documentation. The closest candidates\ncurrently are:\n\n1. `gitglossary`. This makes a good effort, but it's an alphabetically\n    ordered dictionary and a dictionary is not a good way to learn\n    concepts. You have to jump around too much and it's not possible to\n    present the concepts in the order that they should be explained.\n2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n   This is a nice document to have, but it's not necessary to learn how\n   `update-index` works to understand Git's data model, and we should\n   not be requiring users to learn how to use the \"plumbing\" commands\n   if they want to learn what the term \"index\" or \"object\" means.\n3. `gitrepository-layout`. This is a great resource, but it includes a\n   lot of information about configuration and internal implementation\n   details which are not related to the data model. It also does\n   not explain how commits work.\n\nThe result of this is that Git users (even users who have been using\nGit for 15+ years) struggle to read the documentation because they don't\nknow what the core terms mean, and it's not possible to add links\nto help them learn more.\n\nAdd an explanation of Git's data model. Some choices I've made in\ndeciding what \"core data model\" means:\n\n1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n   if those are intended to be user facing or if they're more like\n   internal implementation details.\n2. Don't talk about submodules other than by mentioning how they\n   relate to trees. This is because Git has a lot of special features,\n   and explaining how they all work exhaustively could quickly go\n   down a rabbit hole which would make this document less useful for\n   understanding Git's core behaviour.\n3. Don't discuss the structure of a commit message\n   (first line, trailers etc).\n4. Don't mention configuration.\n5. Don't mention the `.git` directory, to avoid getting too much into\n   implementation details\n\nSigned-off-by: Julia Evans <julia@jvns.ca>\n---\n    doc: Add a explanation of Git's data model\n    \n    Changes in v2:\n    \n    The biggest change is to remove all mentions of the .git directory, and\n    explain references in a way that doesn't refer to \"directories\" at all,\n    and instead talks about the \"hierarchy\" (from Kristoffer and Patrick's\n    reviews).\n    \n    Also:\n    \n     * objects: Mention that an object ID is called an \"object name\", and\n       update the glossary to include the term \"object ID\" (from Junio's\n       review)\n     * objects: Replace \"SHA-1 hash\" with \"cryptographic hash\" which is more\n       accurate (from Patrick's review)\n     * blobs: Made the explanation of git gc a little higher level and took\n       some ideas from Patrick's suggested wording (from Patrick's and\n       Kroftoffer's reviews)\n     * commits: Mention that tag objects and commits can optionally have\n       other fields. I didn't mention the GPG signature specifically, but\n       don't have any objections to adding it. (from Patrick and Junio's\n       reviews)\n     * commits: Remove one of the mentions of git gc, since it perhaps opens\n       up too much of a rabbit hole: \"how does git gc decide which commits\n       to clean up?\". (from Kristoffer's review)\n     * tag objects: Add an example of how a tag object is represented (from\n       user feedback on the draft)\n     * index: Use the term \"file mode\" instead of \"permissions\", and list\n       all allowed file modes (from Patrick's review)\n     * index: Use \"stage number\" instead of \"number\" for index entries (from\n       Patrick's review)\n     * reflogs: Remove \"any ref can be logged\", it raises some questions of\n       \"how do you tell Git to log a ref that it isn't normally logging?\"\n       and my guess is that it's uncommon to ask Git to log more refs. I\n       don't think it's a \"lie\" to omit this but I can bring it back if\n       folks disagree. (from Patrick's review)\n     * reflogs: Fix an error I noticed in the explanation of reflogs: tags\n       aren't logged by default and remote-tracking branches are, according\n       to man git-config\n     * branches and tags: Be clearer about how branches are usually updated\n       (by committing), and make it a little more obvious that only branches\n       can be checked out. This is a bit tricky because using the word\n       \"check out\" introduces a rabbit hole that I want to avoid (what does\n       \"check out\" mean?). I've dealt this by just talking about the\n       \"current branch\" (HEAD) since that is defined here, and making it\n       more explicit that HEAD must either be a branch or a commit, there's\n       no \"HEAD is a tag\" option. (from Patrick's review)\n     * tags: Explain the differences between annotated and lightweight tags\n       (this is the main piece of user feedback I've gotten on the draft so\n       far)\n     * Various style/typo changes (\"2 or more\", linkgit:git-gc[1], removed\n       extra asterisks, added empty SYNOPSIS, \"commits -> tags\" typo fix,\n       add to meson build)\n    \n    non-changes:\n    \n     * I still haven't mentioned things that aren't part of the \"data\n       model\", like revision params and configuration. I think there could\n       be a place for them but I haven't found it yet.\n     * tag objects: I noticed that there's a \"tag\" header field in tag\n       objects (like tag v1.0.0) but I didn't mention it yet because I\n       couldn't figure out what the purpose of that field is (I thought the\n       tag name was stored in the reference, why is it duplicated in the tag\n       object?)\n    \n    Changes in v3:\n    \n    I asked for feedback from Git users on Mastodon and got 220 pieces of\n    feedback from 48 different users. People seemed very excited to read\n    about Git's data model. Usually I judge explanations by what folks\n    report learning from them. Here people reported learning:\n    \n     * how branches are stored (that a branch is \"a name for a commit\")\n     * how objects work\n     * that Git has separate \"author\" and \"committer\" fields\n     * that amending a commit does not change it\n     * that a tree is \"just a directory\" (not something more complicated),\n       and how trees are stored\n     * that Git repos can contain symlinks\n     * that Git saves modes separately from the OS.\n     * how the stage number works\n     * that when you git add a file, Git will create an object\n     * that third-party tools can create their own refs.\n     * that the reflog stores the history of branches (not just HEAD), and\n       what reflogs are for\n    \n    Also (of course) there were quite a few points of confusion! The main 4\n    pieces of feedback were\n    \n     1. The index section doesn't explain what the word \"staged\" means, and\n        one person says that it makes it sounds like only files that you\n        \"git add\"ed are in the index. Rewrite the explanation to avoid using\n        the word \"staged\" to define the index and instead define the word\n        \"staging\".\n     2. Explain the difference between \"annotated tags\" and \"lightweight\n        tags\" (done)\n     3. Add examples for tag objects and reflogs (done)\n     4. Mention a little more about where things are stored in the .git\n        directory, which I'd removed in v2. This seems most important for\n        .git/refs, so I added a hopefully accurate note about how refs are\n        stored by default, with a comment about one of the major\n        implications. I did not discuss where objects or the index are\n        stored, because I don't think the implementation details of how\n        objects are stored are as important, and there are better tools for\n        viewing the \"raw\" state of objects and the index (with git cat-file\n        -p or git ls-files --staged).\n    \n    Here's every other change I made in response to the feedback, as well as\n    a few comments that I did not address.\n    \n    intro:\n    \n     * Give a 1-sentence intro to \"reflog\"\n    \n    objects:\n    \n     * people really like having git ls-files --stage as a way to view the\n       index, so add git cat-file -p as well in a note\n    \n    commits:\n    \n     * 2 people asked \"Are commits stored as a diff?\". Say that diffs are\n       calculated at runtime, this is very important.\n     * The order the fields are given in don't match the order in the\n       example. Make them match.\n     * \"All the files in the commit, stored as a tree\" is throwing a few\n       people off. Be clearer that it's the tree ID of the base directory.\n     * Several people asked \"What's the difference between an author and\n       committer? I added an example using git cherry-pick that I'm not 100%\n       happy with (what if the reader doesn't know what cherry-pick does?).\n       There might be a better example to give here.\n     * In the note about commits being amended: one person suggested saying\n       \"creates a new commit with the same parent\" to make it clearer what\n       the relationship between the new and old commit are. I liked that\n       idea so I did it.\n    \n    trees:\n    \n     * file modes. 2 people want to know more about \"The file mode, for\n       example 100644\". Also 2 people are curious about what relationship\n       these have to Unix permissions. Say that they're inspired by Unix\n       permissions, and move the list of possible file modes up to make the\n       relationship clearer\n     * On \"so git-gc(1) periodically compresses objects to save disk space\",\n       there are a few follow up comments wondering about more, which makes\n       me think the comment about compression is actually a distraction. Say\n       something simpler instead, (\"Git only needs to store new versions of\n       files which were changed in that commit\"), from Junio's suggestion\n     * Re \"commit (a Git submodule)\": 2 people say it's not clear how trees\n       relate to submodules. Say that it refers to a commit in a different\n       repository.\n     * One person says they're not sure if the \"object ID\" is a hash. Link\n       it to the definition of \"object ID\".\n    \n    tag objects:\n    \n     * Requests for an example, added one.\n     * Requests to explain the difference between \"lightweight\" and\n       \"annotated\" tags, added it.\n    \n    tags:\n    \n     * one person thinks \"It’s expected that a tag will never change after\n       you create it.\" is too strong (since of course you can change it with\n       git tag -f). Say instead that tags are \"usually\" not changed.\n    \n    HEAD:\n    \n     * Several people are asking for more detail about detached HEAD state.\n       There's actually quite a lot to talk about here (what it means, how\n       it happens, what it implies, and how you might adjust your workflow\n       to avoid it by using git switch). I don't think we can get into all\n       of that here, so refer to the DETACHED HEAD section of git-checkout\n       instead. I'm not totally happy with the current version of that\n       section but that seems like the most practical solution right now.\n    \n    remote-tracking branches:\n    \n     * discuss refs/remotes/<remote>/HEAD.\n    \n    the index:\n    \n     * \"permissions\" should be \"file mode\" (like with trees). Changed.\n     * \"filename\" should be \"file path\". Changed.\n     * the stage number can only be 0, 1, 2, or 3, since it's 2 bits. Also\n       maybe say that the numbers have specific meanings. Said it can only\n       be 0/1/2/3 but did not give the specific meanings.\n    \n    reflogs\n    \n     * Request for an example. Added one.\n     * It's not clear if there's one reflog per branch/tag/HEAD, or if\n       there's one universal reflog. Make this clearer.\n     * Mention the role of the reflog in retrieving \"lost\" commits or\n       undoing bad rebases.\n    \n    Not fixed:\n    \n     * intro: A couple of people say that it's confusing that tags are both\n       \"an object\" and \"a reference\". Handled this by just explaining the\n       difference between an annotated and a lightweight tag further down.\n       I'd like to make this clearer in the intro but not sure if there's a\n       way to do it.\n     * commits and tag objects: one person asks if there's a reference for\n       the other \"optional fields\", like \"encoding\" and \"gpgsig\". I couldn't\n       find one, so left this as is.\n     * HEAD: A couple of people ask if there are any other symbolic\n       references other than HEAD, or if they can make their own symbolic\n       references. I don't know the answer to this.\n     * HEAD: the HEAD: HEAD thing looks weird, it made more sense when it\n       was HEAD: .git/HEAD. Will think about this.\n     * reflogs: One person asks: if reflogs only store local changes, why\n       does it track the user who made the change? Is that for remote\n       operations like fetches and pulls? Or for cases where more than one\n       user is using the same repo on a system? I don't know the answer to\n       this.\n     * reflogs: How can you see the full data in the reflog? git reflog show\n       doesn't list the user who made the change. git reflog show <refname>\n       --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso seems to work but\n       it's really a mouthful, not sure it's useful to include all that.\n     * index: Is it worth mentioning that the index can be locked? I don't\n       have an opinion about this.\n     * other: One person asks what a \"working tree\" is. It made me wonder if\n       \"the current working directory\" has a place in Git's data model. My\n       feeling is \"no\" but I could be convinced otherwise.\n     * overall: \"How can Git be so fast? If I switch branches, how does it\n       figure out what to add, remove or replace?\". I don't think this is\n       the right place for that discussion but it would\n     * there are some docs CI errors I haven't figured out yet (IDREF\n       attribute linkend references an unknown ID \"tree\")\n    \n    changes in v4:\n    \n    This is a combination of trying to make some of the intro text a little\n    more \"friendly\" for someone new to Git's data model, avoiding implying\n    things that are false, and removing information that isn't relevant to\n    the data model.\n    \n    intro:\n    \n     * Add a 1-line description of what a \"reflog\" is (from user feedback)\n    \n    objects:\n    \n     * Start with a \"friendly\" description of what an object is, similar to\n       what we do for references and the reflog\n     * Rename \"commits\" to \"commit\" and similarly for trees etc (from\n       Junio's review)\n     * Remove the explanation of what git cat-file -p does, since it might\n       be misleading and if people want to know they can read the man page\n       (from Junio's review)\n    \n    commits:\n    \n     * Start by saying that the commit contains the full directory structure\n       of all the files (from Junio's comment about how it may not be clear\n       that the commit contains all the files' exact contents at the time of\n       the commit)\n     * Remove the comment about cherry-pick (from Junio's review)\n     * Replace \"ask Git for a diff\" with \"ask Git to show the commit with\n       git show\" (from Junio's review)\n    \n    trees:\n    \n     * Make the description a little more friendly\n     * Reorder so that \"type\" is defined before we refer to the \"type\"\n     * Say that file modes are \"only spiritually related\" to Unix\n       permissions instead of talking about what Git \"supports\" (from\n       Junio's review)\n    \n    blobs:\n    \n     * Try to make it clearer how \"commits use relatively little disk space\"\n       is true while not implying that commits are diffs, by using an\n       example (from Junio's review)\n    \n    branches:\n    \n     * Replace \"a branch is a name for a commit ID\" with \"a branch refers to\n       a commit ID\" (except in the intro sentence for the \"references\"\n       section). Similarly for tags etc. (from Junio's review)\n     * Remove the note about how branches are stored in .git (from Junio's\n       review)\n    \n    HEAD:\n    \n     * Be clearer that HEAD is not always the current branch, because there\n       may not be a current branch (from Junio's review)\n    \n    index:\n    \n     * Be a little more specific about how exactly the index is converted\n       into a commit. (from Junio's comment about how it's not clear what\n       \"every file in the repository\" means)\n    \n    reflog:\n    \n     * Be clearer that there are many reflogs (one for each reference with a\n       log), not just one reflog (from Junio and Patrick's reviews)\n     * Omit the user and \"Before\" commit IDs from the list of fields,\n       because you usually don't see them (from Junio's review)\n     * Show the output of git reflog main in the example instead of the\n       contents of the reflog file, to avoid showing the user and before\n       commit ID\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/1981\n\nRange-diff vs v3:\n\n 1:  39da4e04cf ! 1:  92249b5b08 doc: add a explanation of Git's data model\n     @@ Metadata\n      Author: Julia Evans <julia@jvns.ca>\n      \n       ## Commit message ##\n     -    doc: add a explanation of Git's data model\n     +    doc: add an explanation of Git's data model\n      \n          Git very often uses the terms \"object\", \"reference\", or \"index\" in its\n          documentation.\n     @@ Documentation/gitdatamodel.adoc (new)\n      +OBJECTS\n      +-------\n      +\n     -+Commits, trees, blobs, and tag objects are all stored in Git's object database.\n     ++All of the commits and files in a Git repository are stored as \"Git objects\".\n     ++Git objects never change after they're created, and every object has an ID,\n     ++like `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n     ++\n     ++This means that if you have an object's ID, you can always recover its\n     ++exact contents as long as the object hasn't been deleted.\n     ++\n      +Every object has:\n      +\n      +[[object-id]]\n     @@ Documentation/gitdatamodel.adoc (new)\n      +   and <<tag-object,tag objects>>.\n      +3. *contents*. The structure of the contents depends on the type.\n      +\n     -+Once an object is created, it can never be changed.\n     -+Here are the 4 types of objects:\n     ++Here's how each type of object is structured:\n      +\n      +[[commit]]\n     -+commits::\n     -+    A commit contains these required fields\n     ++commit::\n     ++    A commit contains the full directory structure of every file\n     ++    in that version of the repository and each file's contents.\n     ++    It has these these required fields\n      +    (though there are other optional fields):\n      ++\n     -+1. All the *files* in the commit, stored as the *<<tree,tree>>* ID of\n     -+   the commit's base directory.\n     ++1. The *files* in the commit, stored as the *<<tree,tree>>* ID\n     ++   of the commit's base directory.\n      +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n      +  regular commits have 1 parent, merge commits have 2 or more parents\n      +3. An *author* and the time the commit was authored\n      +4. A *committer* and the time the commit was committed.\n     -+   If you cherry-pick (linkgit:git-cherry-pick[1]) someone else's commit,\n     -+   then they will be the author and you'll be the committer.\n      +5. A *commit message*\n      ++\n      +Here's how an example commit is stored:\n     @@ Documentation/gitdatamodel.adoc (new)\n      +For example, \"amending\" a commit with `git commit --amend` creates a new\n      +commit with the same parent.\n      ++\n     -+Git does not store the diff for a commit: when you ask Git for a\n     -+diff it calculates it on the fly.\n     ++Git does not store the diff for a commit: when you ask Git to show\n     ++the commit with linkgit:git-show[1], it calculates the diff from its\n     ++parent on the fly.\n      +\n      +[[tree]]\n     -+trees::\n     -+    A tree is how Git represents a directory. It lists, for each item in\n     -+    the tree:\n     ++tree::\n     ++    A tree is how Git represents a directory.\n     ++    It can contain files or other trees (which are subdirectories).\n     ++    It lists, for each item in the tree:\n      ++\n     -+[[file-mode]]\n     -+1. The *file mode*, for example `100644`. The format is inspired by Unix\n     -+   permissions, but Git's modes are much more limited. Git only supports these file modes:\n     ++1. The *filename*, for example `hello.py`\n     ++2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n     ++  or <<commit,`commit`>> (a Git submodule, which is a\n     ++  commit from a different Git repository)\n     ++3. The *file mode*. Git has these file modes. which are only\n     ++   spiritually related to Unix permissions:\n      ++\n      +  - `100644`: regular file (with type `blob`)\n      +  - `100755`: executable file (with type `blob`)\n     @@ Documentation/gitdatamodel.adoc (new)\n      +  - `040000`: directory (with type `tree`)\n      +  - `160000`: gitlink, for use with submodules (with type `commit`)\n      +\n     -+2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n     -+  or <<commit,`commit`>> (a Git submodule, which is a\n     -+  commit from a different Git repository)\n     -+3. The <<object-id,*object ID*>>\n     -+4. The *filename*\n     ++4. The <<object-id,*object ID*>> with the contents of the file or directory\n      ++\n      +For example, this is how a tree containing one directory (`src`) and one file\n      +(`README.md`) is stored:\n     @@ Documentation/gitdatamodel.adoc (new)\n      +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n      +----\n      +\n     -+\n      +[[blob]]\n     -+blobs::\n     -+    A blob is how Git represents a file. A blob object contains the\n     -+    file's contents.\n     ++blob::\n     ++    A blob object contains a file's contents.\n      ++\n     -+When you make a new commit, Git only needs to store new versions of\n     -+files which were changed in that commit. This means that commits\n     -+can use relatively little disk space even in a very large repository.\n     ++When you make a commit, Git stores the full contents of each file that\n     ++you changed as a blob.\n     ++For example, if you have a commit that changes 2 files in a repository\n     ++with 1000 files, that commit will create 2 new blobs, and use the\n     ++previous blob ID for the other 998 files.\n     ++This means that commits can use relatively little disk space even in a\n     ++very large repository.\n      +\n      +[[tag-object]]\n     -+tag objects::\n     ++tag object::\n      +    Tag objects contain these required fields\n      +    (though there are other optional fields):\n      ++\n     @@ Documentation/gitdatamodel.adoc (new)\n      +----\n      +\n      +NOTE: All of the examples in this section were generated with\n     -+`git cat-file -p <object-id>`, which shows the contents of a Git object.\n     ++`git cat-file -p <object-id>`.\n      +\n      +[[references]]\n      +REFERENCES\n     @@ Documentation/gitdatamodel.adoc (new)\n      +branch\" than \"the changes are in commit bb69721404348e\".\n      +Git often uses \"ref\" as shorthand for \"reference\".\n      +\n     -+References can either be:\n     ++References can either refer to:\n      +\n     -+1. References to an object ID, usually a <<commit,commit>> ID\n     -+2. References to another reference. This is called a \"symbolic reference\".\n     ++1. An object ID, usually a <<commit,commit>> ID\n     ++2. Another reference. This is called a \"symbolic reference\".\n      +\n      +References are stored in a hierarchy, and Git handles references\n      +differently based on where they are in the hierarchy.\n     @@ Documentation/gitdatamodel.adoc (new)\n      +\n      +[[branch]]\n      +branches: `refs/heads/<name>`::\n     -+    A branch is a name for a commit ID.\n     ++    A branch refers to a commit ID.\n      +    That commit is the latest commit on the branch.\n      ++\n      +To get the history of commits on a branch, Git will start at the commit\n     @@ Documentation/gitdatamodel.adoc (new)\n      +\n      +[[tag]]\n      +tags: `refs/tags/<name>`::\n     -+    A tag is a name for a commit ID, tag object ID, or other object ID.\n     -+    Tags that reference a tag object ID are called \"annotated tags\",\n     -+    because the tag object contains a tag message.\n     -+    Tags that reference a commit, blob, or tree ID are\n     -+    called \"lightweight tags\".\n     ++    A tag refers to a commit ID, tag object ID, or other object ID.\n     ++    There are two types of tags:\n     ++    1. \"Annotated tags\", which reference a <<tag-object,tag object>> ID\n     ++       which contains a tag message\n     ++    2. \"Lightweight tags\", which reference a commit, blob, or tree ID\n     ++       directly\n      ++\n     -+Even though branches and tags are both \"a name for a commit ID\", Git\n     ++Even though branches and tags both refer to a commit ID, Git\n      +treats them very differently.\n      +Branches are expected to change over time: when you make a commit, Git\n     -+will update your <<HEAD,current branch>> to reference the new changes.\n     ++will update your <<HEAD,current branch>> to point to the new commit.\n      +Tags are usually not changed after they're created.\n      +\n      +[[HEAD]]\n      +HEAD: `HEAD`::\n     -+    `HEAD` is where Git stores your current <<branch,branch>>.\n     -+    `HEAD` can either be:\n     -+    1. A symbolic reference to your current branch, for example `ref:\n     -+       refs/heads/main` if your current branch is `main`.\n     -+    2. A direct reference to a commit ID. This is called \"detached HEAD\n     -+\t   state\", see the DETACHED HEAD section of linkgit:git-checkout[1] for more.\n     ++    `HEAD` is where Git stores your current <<branch,branch>>,\n     ++    if there is a current branch. `HEAD` can either be:\n     +++\n     ++1. A symbolic reference to your current branch, for example `ref:\n     ++   refs/heads/main` if your current branch is `main`.\n     ++2. A direct reference to a commit ID. In this case there is no current branch.\n     ++   This is called \"detached HEAD state\", see the DETACHED HEAD section\n     ++   of linkgit:git-checkout[1] for more.\n      +\n      +[[remote-tracking-branch]]\n     -+remote tracking branches: `refs/remotes/<remote>/<branch>`::\n     -+    A remote-tracking branch is a name for a commit ID.\n     ++remote-tracking branches: `refs/remotes/<remote>/<branch>`::\n     ++    A remote-tracking branch refers to a commit ID.\n      +    It's how Git stores the last-known state of a branch in a remote\n      +    repository. `git fetch` updates remote-tracking branches. When\n      +    `git status` says \"you're up to date with origin/main\", it's looking at\n     @@ Documentation/gitdatamodel.adoc (new)\n      ++\n      +Git may also create references other than `HEAD` at the base of the\n      +hierarchy, like `ORIG_HEAD`.\n     -++\n     -+NOTE: By default, Git references are stored as files in the `.git` directory.\n     -+For example, the branch `main` is stored in `.git/refs/heads/main`.\n     -+This means that you can't have branches named both `maya` and `maya/some-task`,\n     -+because there can't be a file and a directory with the same name.\n      +\n      +[[index]]\n      +THE INDEX\n      +---------\n     -+\n     -+The index, also known as the \"staging area\", contains a list of every\n     -+file in the repository and its contents. When you commit, the files in\n     -+the index are used as the files in the next commit.\n     -+\n     -+You can add files to the index or update the version in the index with\n     -+linkgit:git-add[1]. Adding a file to the index or updating its version\n     -+is called \"staging\" the file for commit.\n     ++The index, also known as the \"staging area\", is a list of files and\n     ++the contents of each file, stored as a <<blob,blob>>.\n     ++You can add files to the index or update the contents of a file in the\n     ++index with linkgit:git-add[1]. This is called \"staging\" the file for commit.\n      +\n      +Unlike a <<tree,tree>>, the index is a flat list of files.\n     ++When you commit, Git converts the list of files in the index to a\n     ++directory <<tree,tree>> and uses that tree in the new <<commit,commit>>.\n     ++\n      +Each index entry has 4 fields:\n      +\n     -+1. The *<<file-mode,file mode>>*\n     ++1. The *<<tree,file mode>>*\n      +2. The *<<blob,blob>> ID* of the file\n      +3. The *file path*, for example `src/hello.py`\n      +4. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n     @@ Documentation/gitdatamodel.adoc (new)\n      +REFLOGS\n      +-------\n      +\n     -+Git stores a history called a \"reflog\" for every branch, remote-tracking\n     -+branch, and HEAD. This means that if you make a mistake and \"lose\" a\n     -+commit, you can generally recover the commit ID by running\n     -+`git reflog <reference>`.\n     ++Every time a branch, remote-tracking branch, or HEAD is updated, Git\n     ++updates a log called a \"reflog\" for that <<references,reference>>.\n     ++This means that if you make a mistake and \"lose\" a commit, you can\n     ++generally recover the commit ID by running `git reflog <reference>`.\n      +\n     -+Each reflog entry has:\n     ++A reflog is a list of log entries. Each entry has:\n      +\n     -+1. Before/after *commit IDs*\n     -+2. *User* who made the change, for example `Maya <maya@example.com>`\n     -+3. *Timestamp* when the change was made\n     -+4. *Log message*, for example `pull: Fast-forward`\n     ++1. The *commit ID*\n     ++2. *Timestamp* when the change was made\n     ++3. *Log message*, for example `pull: Fast-forward`\n      +\n      +Reflogs only log changes made in your local repository.\n      +They are not shared with remotes.\n      +\n     -+For example, here's how the reflog for `HEAD` in a repository with 2\n     -+commits is stored:\n     ++You can view a reflog with `git reflog <reference>`.\n     ++For example, here's the reflog for a `main` branch which has changed twice:\n      +\n      +----\n     -+0000000000000000000000000000000000000000 4ccb6d7b8869a86aae2e84c56523f8705b50c647 Maya <maya@example.com> 1759173408 -0400      commit (initial): Initial commit\n     -+4ccb6d7b8869a86aae2e84c56523f8705b50c647 750b4ead9c87ceb3ddb7a390e6c7074521797fb3 Maya <maya@example.com> 1759173425 -0400      commit: Add README\n     ++$ git reflog main --date=iso --no-decorate\n     ++750b4ea main@{2025-09-29 15:17:05 -0400}: commit: Add README\n     ++4ccb6d7 main@{2025-09-29 15:16:48 -0400}: commit (initial): Initial commit\n      +----\n      +\n      +GIT\n\n\n Documentation/Makefile              |   1 +\n Documentation/gitdatamodel.adoc     | 286 ++++++++++++++++++++++++++++\n Documentation/glossary-content.adoc |   4 +-\n Documentation/meson.build           |   1 +\n 4 files changed, 290 insertions(+), 2 deletions(-)\n create mode 100644 Documentation/gitdatamodel.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 6fb83d0c6e..5f4acfacbd 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n MAN7_TXT += gitcore-tutorial.adoc\n MAN7_TXT += gitcredentials.adoc\n MAN7_TXT += gitcvs-migration.adoc\n+MAN7_TXT += gitdatamodel.adoc\n MAN7_TXT += gitdiffcore.adoc\n MAN7_TXT += giteveryday.adoc\n MAN7_TXT += gitfaq.adoc\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nnew file mode 100644\nindex 0000000000..e36e833f66\n--- /dev/null\n+++ b/Documentation/gitdatamodel.adoc\n@@ -0,0 +1,286 @@\n+gitdatamodel(7)\n+===============\n+\n+NAME\n+----\n+gitdatamodel - Git's core data model\n+\n+SYNOPSIS\n+--------\n+gitdatamodel\n+\n+DESCRIPTION\n+-----------\n+\n+It's not necessary to understand Git's data model to use Git, but it's\n+very helpful when reading Git's documentation so that you know what it\n+means when the documentation says \"object\", \"reference\" or \"index\".\n+\n+Git's core operations use 4 kinds of data:\n+\n+1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+2. <<references,References>>: branches, tags,\n+   remote-tracking branches, etc\n+3. <<index,The index>>, also known as the staging area\n+4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n+\n+[[objects]]\n+OBJECTS\n+-------\n+\n+All of the commits and files in a Git repository are stored as \"Git objects\".\n+Git objects never change after they're created, and every object has an ID,\n+like `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+\n+This means that if you have an object's ID, you can always recover its\n+exact contents as long as the object hasn't been deleted.\n+\n+Every object has:\n+\n+[[object-id]]\n+1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n+  type and contents.\n+  It's fast to look up a Git object using its ID.\n+  This is usually represented in hexadecimal, like\n+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+2. a *type*. There are 4 types of objects:\n+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n+   and <<tag-object,tag objects>>.\n+3. *contents*. The structure of the contents depends on the type.\n+\n+Here's how each type of object is structured:\n+\n+[[commit]]\n+commit::\n+    A commit contains the full directory structure of every file\n+    in that version of the repository and each file's contents.\n+    It has these these required fields\n+    (though there are other optional fields):\n++\n+1. The *files* in the commit, stored as the *<<tree,tree>>* ID\n+   of the commit's base directory.\n+2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n+  regular commits have 1 parent, merge commits have 2 or more parents\n+3. An *author* and the time the commit was authored\n+4. A *committer* and the time the commit was committed.\n+5. A *commit message*\n++\n+Here's how an example commit is stored:\n++\n+----\n+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n+author Maya <maya@example.com> 1759173425 -0400\n+committer Maya <maya@example.com> 1759173425 -0400\n+\n+Add README\n+----\n++\n+Like all other objects, commits can never be changed after they're created.\n+For example, \"amending\" a commit with `git commit --amend` creates a new\n+commit with the same parent.\n++\n+Git does not store the diff for a commit: when you ask Git to show\n+the commit with linkgit:git-show[1], it calculates the diff from its\n+parent on the fly.\n+\n+[[tree]]\n+tree::\n+    A tree is how Git represents a directory.\n+    It can contain files or other trees (which are subdirectories).\n+    It lists, for each item in the tree:\n++\n+1. The *filename*, for example `hello.py`\n+2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n+  or <<commit,`commit`>> (a Git submodule, which is a\n+  commit from a different Git repository)\n+3. The *file mode*. Git has these file modes. which are only\n+   spiritually related to Unix permissions:\n++\n+  - `100644`: regular file (with type `blob`)\n+  - `100755`: executable file (with type `blob`)\n+  - `120000`: symbolic link (with type `blob`)\n+  - `040000`: directory (with type `tree`)\n+  - `160000`: gitlink, for use with submodules (with type `commit`)\n+\n+4. The <<object-id,*object ID*>> with the contents of the file or directory\n++\n+For example, this is how a tree containing one directory (`src`) and one file\n+(`README.md`) is stored:\n++\n+----\n+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n+----\n+\n+[[blob]]\n+blob::\n+    A blob object contains a file's contents.\n++\n+When you make a commit, Git stores the full contents of each file that\n+you changed as a blob.\n+For example, if you have a commit that changes 2 files in a repository\n+with 1000 files, that commit will create 2 new blobs, and use the\n+previous blob ID for the other 998 files.\n+This means that commits can use relatively little disk space even in a\n+very large repository.\n+\n+[[tag-object]]\n+tag object::\n+    Tag objects contain these required fields\n+    (though there are other optional fields):\n++\n+1. The *ID* and *type* of the object (often a commit) that they reference\n+2. The *tagger* and tag date\n+3. A *tag message*, similar to a commit message\n+\n+Here's how an example tag object is stored:\n+\n+----\n+object 750b4ead9c87ceb3ddb7a390e6c7074521797fb3\n+type commit\n+tag v1.0.0\n+tagger Maya <maya@example.com> 1759927359 -0400\n+\n+Release version 1.0.0\n+----\n+\n+NOTE: All of the examples in this section were generated with\n+`git cat-file -p <object-id>`.\n+\n+[[references]]\n+REFERENCES\n+----------\n+\n+References are a way to give a name to a commit.\n+It's easier to remember \"the changes I'm working on are on the `turtle`\n+branch\" than \"the changes are in commit bb69721404348e\".\n+Git often uses \"ref\" as shorthand for \"reference\".\n+\n+References can either refer to:\n+\n+1. An object ID, usually a <<commit,commit>> ID\n+2. Another reference. This is called a \"symbolic reference\".\n+\n+References are stored in a hierarchy, and Git handles references\n+differently based on where they are in the hierarchy.\n+Most references are under `refs/`. Here are the main types:\n+\n+[[branch]]\n+branches: `refs/heads/<name>`::\n+    A branch refers to a commit ID.\n+    That commit is the latest commit on the branch.\n++\n+To get the history of commits on a branch, Git will start at the commit\n+ID the branch references, and then look at the commit's parent(s),\n+the parent's parent, etc.\n+\n+[[tag]]\n+tags: `refs/tags/<name>`::\n+    A tag refers to a commit ID, tag object ID, or other object ID.\n+    There are two types of tags:\n+    1. \"Annotated tags\", which reference a <<tag-object,tag object>> ID\n+       which contains a tag message\n+    2. \"Lightweight tags\", which reference a commit, blob, or tree ID\n+       directly\n++\n+Even though branches and tags both refer to a commit ID, Git\n+treats them very differently.\n+Branches are expected to change over time: when you make a commit, Git\n+will update your <<HEAD,current branch>> to point to the new commit.\n+Tags are usually not changed after they're created.\n+\n+[[HEAD]]\n+HEAD: `HEAD`::\n+    `HEAD` is where Git stores your current <<branch,branch>>,\n+    if there is a current branch. `HEAD` can either be:\n++\n+1. A symbolic reference to your current branch, for example `ref:\n+   refs/heads/main` if your current branch is `main`.\n+2. A direct reference to a commit ID. In this case there is no current branch.\n+   This is called \"detached HEAD state\", see the DETACHED HEAD section\n+   of linkgit:git-checkout[1] for more.\n+\n+[[remote-tracking-branch]]\n+remote-tracking branches: `refs/remotes/<remote>/<branch>`::\n+    A remote-tracking branch refers to a commit ID.\n+    It's how Git stores the last-known state of a branch in a remote\n+    repository. `git fetch` updates remote-tracking branches. When\n+    `git status` says \"you're up to date with origin/main\", it's looking at\n+    this.\n++\n+`refs/remotes/<remote>/HEAD` is a symbolic reference to the remote's\n+default branch. This is the branch that `git clone` checks out by default.\n+\n+[[other-refs]]\n+Other references::\n+    Git tools may create references anywhere under `refs/`.\n+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n+    and linkgit:git-notes[1] all create their own references\n+    in `refs/stash`, `refs/bisect`, etc.\n+    Third-party Git tools may also create their own references.\n++\n+Git may also create references other than `HEAD` at the base of the\n+hierarchy, like `ORIG_HEAD`.\n+\n+[[index]]\n+THE INDEX\n+---------\n+The index, also known as the \"staging area\", is a list of files and\n+the contents of each file, stored as a <<blob,blob>>.\n+You can add files to the index or update the contents of a file in the\n+index with linkgit:git-add[1]. This is called \"staging\" the file for commit.\n+\n+Unlike a <<tree,tree>>, the index is a flat list of files.\n+When you commit, Git converts the list of files in the index to a\n+directory <<tree,tree>> and uses that tree in the new <<commit,commit>>.\n+\n+Each index entry has 4 fields:\n+\n+1. The *<<tree,file mode>>*\n+2. The *<<blob,blob>> ID* of the file\n+3. The *file path*, for example `src/hello.py`\n+4. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n+   there's a merge conflict there can be multiple versions of the same\n+   filename in the index.\n+\n+It's extremely uncommon to look at the index directly: normally you'd\n+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n+But you can use `git ls-files --stage` to see the index.\n+Here's the output of `git ls-files --stage` in a repository with 2 files:\n+\n+----\n+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n+----\n+\n+[[reflogs]]\n+REFLOGS\n+-------\n+\n+Every time a branch, remote-tracking branch, or HEAD is updated, Git\n+updates a log called a \"reflog\" for that <<references,reference>>.\n+This means that if you make a mistake and \"lose\" a commit, you can\n+generally recover the commit ID by running `git reflog <reference>`.\n+\n+A reflog is a list of log entries. Each entry has:\n+\n+1. The *commit ID*\n+2. *Timestamp* when the change was made\n+3. *Log message*, for example `pull: Fast-forward`\n+\n+Reflogs only log changes made in your local repository.\n+They are not shared with remotes.\n+\n+You can view a reflog with `git reflog <reference>`.\n+For example, here's the reflog for a `main` branch which has changed twice:\n+\n+----\n+$ git reflog main --date=iso --no-decorate\n+750b4ea main@{2025-09-29 15:17:05 -0400}: commit: Add README\n+4ccb6d7 main@{2025-09-29 15:16:48 -0400}: commit (initial): Initial commit\n+----\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\ndiff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc\nindex e423e4765b..20ba121314 100644\n--- a/Documentation/glossary-content.adoc\n+++ b/Documentation/glossary-content.adoc\n@@ -297,8 +297,8 @@ This commit is referred to as a \"merge commit\", or sometimes just a\n \tidentified by its <<def_object_name,object name>>. The objects usually\n \tlive in `$GIT_DIR/objects/`.\n \n-[[def_object_identifier]]object identifier (oid)::\n-\tSynonym for <<def_object_name,object name>>.\n+[[def_object_identifier]]object identifier, object ID, oid::\n+\tSynonyms for <<def_object_name,object name>>.\n \n [[def_object_name]]object name::\n \tThe unique identifier of an <<def_object,object>>.  The\ndiff --git a/Documentation/meson.build b/Documentation/meson.build\nindex e34965c5b0..ace0573e82 100644\n--- a/Documentation/meson.build\n+++ b/Documentation/meson.build\n@@ -192,6 +192,7 @@ manpages = {\n   'gitcore-tutorial.adoc' : 7,\n   'gitcredentials.adoc' : 7,\n   'gitcvs-migration.adoc' : 7,\n+  'gitdatamodel.adoc' : 7,\n   'gitdiffcore.adoc' : 7,\n   'giteveryday.adoc' : 7,\n   'gitfaq.adoc' : 7,\n\nbase-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n-- \ngitgitgadget\n"},{"id":"529760","messageId":"xmqqikg0f1tk.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v4.git.1761593537924.gitgitgadget@gmail.com","subject":"Re: [PATCH v4] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-27T21:54:15Z","receivedAt":"2025-10-27T21:54:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n> new file mode 100644\n> index 0000000000..e36e833f66\n> --- /dev/null\n> +++ b/Documentation/gitdatamodel.adoc\n> @@ -0,0 +1,286 @@\n> +gitdatamodel(7)\n> +===============\n> +\n> +NAME\n> +----\n> +gitdatamodel - Git's core data model\n> +\n> +SYNOPSIS\n> +--------\n> +gitdatamodel\n> +\n> +DESCRIPTION\n> +-----------\n> +\n> +It's not necessary to understand Git's data model to use Git, but it's\n> +very helpful when reading Git's documentation so that you know what it\n> +means when the documentation says \"object\", \"reference\" or \"index\".\n\n\"While it is not necessary ..., it is helpful ...\" may flow better\nthan \"It is not necesary ..., but it is very helpful\".\n\n> +This means that if you have an object's ID, you can always recover its\n> +exact contents as long as the object hasn't been deleted.\n\nSomewhere in distant footnote, we may want to mention that objects\nthat are in use are never deleted, and when they get removed (i.e.,\ngarbage collection).  As part of the data model, \"everything is\nretained by default, until we can prove it is no longer reachable\"\nprobably belongs somewhere.\n\n> +Here's how each type of object is structured:\n> +\n> +[[commit]]\n> +commit::\n> +    A commit contains the full directory structure of every file\n> +    in that version of the repository and each file's contents.\n\nWhat you are describing here is more of the property of a tree; a\ncommit is a bit richer.\n\n    A commit records a snapshot of the every file in the project at\n    one point in time, records who contributed to create such a\n    snapshot and why, and how that particular snapshot relates to\n    other snapshots in the history.\n\n> +    It has these these required fields\n\n\"these these\".\n\n> +Like all other objects, commits can never be changed after they're created.\n> +For example, \"amending\" a commit with `git commit --amend` creates a new\n> +commit with the same parent.\n\n\"same parent.\" -> \"same parent, without modifying the original\ncommit object at all\"?  Maybe redundant?  I dunno.\n\n> +[[tree]]\n> +tree::\n> +    A tree is how Git represents a directory.\n\n\"a directory\" -> \"contents in a directory\"?  I dunno.\n\n> +    It can contain files or other trees (which are subdirectories).\n> +    It lists, for each item in the tree:\n> ++\n> +1. The *filename*, for example `hello.py`\n> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n> +  or <<commit,`commit`>> (a Git submodule, which is a\n> +  commit from a different Git repository)\n\nThis is a bit of white lie.  A tree object entry never stores the\ntype of the object.  It records <mode, object name, path component>.\n\nThe second field you see in git ls-tree output is computed from the\nobject name (when the object is available) or inferred from the mode\nbits.\n\n> +3. The *file mode*. Git has these file modes. which are only\n> +   spiritually related to Unix permissions:\n\nIn the cover letter part of the message I am responding to, I saw\nrepeated mention of \"permissions should be \"file mode\"; let's be\nconsistent.\n\n\"Git has these file modes, which are ...\" -> \n\n    Git uses the following file mode to represent what each tree\n    entry is (because an object of the same type, e.g. \"blob\", is\n    used to represent more than one kind of things).  The file mode\n    are assigned to resemble Unix file mode.\n\n    Note that Git does not _store_ permissions, and there are only\n    two kinds of regular files; non-executable (100644) or\n    executable (100755).  To Git, there are no files that are\n    \"readable only by the owner\" etc., so file mode bits like\n    100600, 100400, etc., are never used.\n\n> +[[tag-object]]\n> +tag object::\n> +    Tag objects contain these required fields\n> +    (though there are other optional fields):\n> ++\n> +1. The *ID* and *type* of the object (often a commit) that they reference\n\nNot wrong per-se, but it is a bit curious to lump these two into a\nsingle enumerated item here, unlike \"author\" and \"committer\" were\nenumerated separately for commit objects.  If you are going to show\n\"cat-file -p\" output for illustration, it may be help readers\nunderstand them if you had them separately listed here.\n\n> +2. The *tagger* and tag date\n> +3. A *tag message*, similar to a commit message\n\n> +[[index]]\n> +THE INDEX\n> +---------\n> +The index, also known as the \"staging area\", is a list of files and\n> +the contents of each file, stored as a <<blob,blob>>.\n> +You can add files to the index or update the contents of a file in the\n> +index with linkgit:git-add[1]. This is called \"staging\" the file for commit.\n> +\n> +Unlike a <<tree,tree>>, the index is a flat list of files.\n\nThis is a bit of white lie, as modern versions of Git could be\ncollapsing uninteresting parts of the directory structure as a\nsingle tree in an index entry (this is called \"sparse index\"), and\ncan expand such collapsed \"tree\" in the index on-demand into its\nconstituent files and directories.  But I do not mind presenting the\ntraditional world model for conceptual simplicity.\n\n> +When you commit, Git converts the list of files in the index to a\n> +directory <<tree,tree>> and uses that tree in the new <<commit,commit>>.\n> +\n> +Each index entry has 4 fields:\n> +\n> +1. The *<<tree,file mode>>*\n> +2. The *<<blob,blob>> ID* of the file\n\nIf you were to collapse descriptions like you did for tag objects\nwhere ID and TYPE were treated as a unit, here is the place to do\nso.  With the mode bits and object ID, we can represent regular\nfiles that are non-executable, regular files that are executable,  \nsymbolic links, and submodules (if a sparse-index is in use, an\nindex entry could be a subdirectory, but I suggested above that we\ncan ignore them for simplicity).\n\nBut <<blob,blob>> is highly misleading.  Even if we ignore\nsparse-index, we may see a commit object there.\n\n    Each index entry records\n\n    1. The object that occupies the path, as (file mode, object\n       name) tuple.  Most often, it is a regular file whose contents\n       are stored in a blob object, that is either non-executable\n       (100644), executable (100755), or a symbolic link (120000),\n       but the object can be a commit in another repository if it\n       represents a submodule.\n\n    2. The stage number, which is normally 0, but entries with\n       higher stages for the same path are used during a conflicted\n       merge.\n\n    3. The path name for the index entry.\n\n> +3. The *file path*, for example `src/hello.py`\n> +4. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n> +   there's a merge conflict there can be multiple versions of the same\n> +   filename in the index.\n\nIf you are going by \"ls-files -s\" output, it may be better to swap 3\nand 4 above for ease of understanding.\n\n> +It's extremely uncommon to look at the index directly: normally you'd\n> +run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n> +But you can use `git ls-files --stage` to see the index.\n> +Here's the output of `git ls-files --stage` in a repository with 2 files:\n> +\n> +----\n> +100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n> +100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n> +----\n> +\n> +[[reflogs]]\n> +REFLOGS\n> +-------\n> +\n> +Every time a branch, remote-tracking branch, or HEAD is updated, Git\n> +updates a log called a \"reflog\" for that <<references,reference>>.\n\nIf we want to avoid using word X while explaining X, then we can\nrephrase it as \"Git updates a record in the reflog for that\nreference\".\n"},{"id":"529829","messageId":"5b078fae-6fe9-4fde-ba84-1070761c168b@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqikg0f1tk.fsf@gitster.g","subject":"Re: [PATCH v4] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-10-28T20:10:52Z","receivedAt":"2025-10-28T20:11:21Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":">> +\n>> +It's not necessary to understand Git's data model to use Git, but it's\n>> +very helpful when reading Git's documentation so that you know what it\n>> +means when the documentation says \"object\", \"reference\" or \"index\".\n>\n> \"While it is not necessary ..., it is helpful ...\" may flow better\n> than \"It is not necesary ..., but it is very helpful\".\n>\n>> +This means that if you have an object's ID, you can always recover its\n>> +exact contents as long as the object hasn't been deleted.\n>\n> Somewhere in distant footnote, we may want to mention that objects\n> that are in use are never deleted, and when they get removed (i.e.,\n> garbage collection).  As part of the data model, \"everything is\n> retained by default, until we can prove it is no longer reachable\"\n> probably belongs somewhere.\n\nAgreed, I really like this idea. Came up with the following, which I'll put at\nthe bottom of the \"References\" section if I don't come up with a better idea.\n(I don't feel strongly about where exactly it should go):\n\nNOTE: Objects will only be deleted if they aren't \"reachable\" from any reference.\nAn object is \"reachable\" if we can find it by following tags to whatever\nthey tag, commits to their parents or trees, and trees to the trees or\nblobs that they contain.\nFor example, if you amend a commit, with `git commit --amend`,\nthe old commit will usually not be reachable, so it may be deleted eventually.\n\n>> +Here's how each type of object is structured:\n>> +\n>> +[[commit]]\n>> +commit::\n>> +    A commit contains the full directory structure of every file\n>> +    in that version of the repository and each file's contents.\n>\n> What you are describing here is more of the property of a tree; a\n> commit is a bit richer.\n>\n>     A commit records a snapshot of the every file in the project at\n>     one point in time, records who contributed to create such a\n>     snapshot and why, and how that particular snapshot relates to\n>     other snapshots in the history.\n\nI don't understand the goal of explaining a commit in detail in\nparagraph form when we already explain everything in a commit right\nbelow this.\n\nMy goal of this intro sentence is just to emphasize what I think is the\nleast obvious point in that list, which is that commits contain every file. \n\nHappy to change it to something shorter like\n\"A commit records a snapshot of the every file in the project\" if you\nprefer that wording.\n\n>> +    It has these these required fields\n>\n> \"these these\".\n\nOops, will fix\n\n>> +Like all other objects, commits can never be changed after they're created.\n>> +For example, \"amending\" a commit with `git commit --amend` creates a new\n>> +commit with the same parent.\n>\n> \"same parent.\" -> \"same parent, without modifying the original\n> commit object at all\"?  Maybe redundant?  I dunno.\n>\n>> +[[tree]]\n>> +tree::\n>> +    A tree is how Git represents a directory.\n>\n> \"a directory\" -> \"contents in a directory\"?  I dunno.\n>\n>> +    It can contain files or other trees (which are subdirectories).\n>> +    It lists, for each item in the tree:\n>> ++\n>> +1. The *filename*, for example `hello.py`\n>> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n>> +  or <<commit,`commit`>> (a Git submodule, which is a\n>> +  commit from a different Git repository)\n>\n> This is a bit of white lie.  A tree object entry never stores the\n> type of the object.  It records <mode, object name, path component>.\n>\n> The second field you see in git ls-tree output is computed from the\n> object name (when the object is available) or inferred from the mode\n> bits.\n\nThanks, I didn't realize how tree object entries were stored.\nWill remove \"type\".\n\n>> +3. The *file mode*. Git has these file modes. which are only\n>> +   spiritually related to Unix permissions:\n>\n> In the cover letter part of the message I am responding to, I saw\n> repeated mention of \"permissions should be \"file mode\"; let's be\n> consistent.\n>\n> \"Git has these file modes, which are ...\" -> \n\nMakes sense. Will change to \"Unix file modes\" from \"Unix permissions\".\nI don't think this needs a more dramatic rewrite though.\n\n>     Git uses the following file mode to represent what each tree\n>     entry is (because an object of the same type, e.g. \"blob\", is\n>     used to represent more than one kind of things).  The file mode\n>     are assigned to resemble Unix file mode.\n>\n>     Note that Git does not _store_ permissions, and there are only\n>     two kinds of regular files; non-executable (100644) or\n>     executable (100755).  To Git, there are no files that are\n>     \"readable only by the owner\" etc., so file mode bits like\n>     100600, 100400, etc., are never used.\n>\n>> +[[tag-object]]\n>> +tag object::\n>> +    Tag objects contain these required fields\n>> +    (though there are other optional fields):\n>> ++\n>> +1. The *ID* and *type* of the object (often a commit) that they reference\n>\n> Not wrong per-se, but it is a bit curious to lump these two into a\n> single enumerated item here, unlike \"author\" and \"committer\" were\n> enumerated separately for commit objects.  If you are going to show\n> \"cat-file -p\" output for illustration, it may be help readers\n> understand them if you had them separately listed here.\n\nAgreed, I'll split them into two items.\n\n>> +2. The *tagger* and tag date\n>> +3. A *tag message*, similar to a commit message\n>\n>> +[[index]]\n>> +THE INDEX\n>> +---------\n>> +The index, also known as the \"staging area\", is a list of files and\n>> +the contents of each file, stored as a <<blob,blob>>.\n>> +You can add files to the index or update the contents of a file in the\n>> +index with linkgit:git-add[1]. This is called \"staging\" the file for commit.\n>> +\n>> +Unlike a <<tree,tree>>, the index is a flat list of files.\n>\n> This is a bit of white lie, as modern versions of Git could be\n> collapsing uninteresting parts of the directory structure as a\n> single tree in an index entry (this is called \"sparse index\"), and\n> can expand such collapsed \"tree\" in the index on-demand into its\n> constituent files and directories.  But I do not mind presenting the\n> traditional world model for conceptual simplicity.\n\nI didn't know that, thanks. I guess I'll leave it the way it is for now.\nIt could be good to add a footnote, but I don't actually know how\nto add footnotes in this document format.\n\n>> +When you commit, Git converts the list of files in the index to a\n>> +directory <<tree,tree>> and uses that tree in the new <<commit,commit>>.\n>> +\n>> +Each index entry has 4 fields:\n>> +\n>> +1. The *<<tree,file mode>>*\n>> +2. The *<<blob,blob>> ID* of the file\n>\n> If you were to collapse descriptions like you did for tag objects\n> where ID and TYPE were treated as a unit, here is the place to do\n> so.  With the mode bits and object ID, we can represent regular\n> files that are non-executable, regular files that are executable,  \n> symbolic links, and submodules (if a sparse-index is in use, an\n> index entry could be a subdirectory, but I suggested above that we\n> can ignore them for simplicity).\n>\n> But <<blob,blob>> is highly misleading.  Even if we ignore\n> sparse-index, we may see a commit object there.\n\nThanks, I didn't realize that. Will change to say that it can be a blob\nor commit ID. I don't think that collapsing will help, IMO it's\nimportant to keep a consistent format.\n\n>     Each index entry records\n>\n>     1. The object that occupies the path, as (file mode, object\n>        name) tuple.  Most often, it is a regular file whose contents\n>        are stored in a blob object, that is either non-executable\n>        (100644), executable (100755), or a symbolic link (120000),\n>        but the object can be a commit in another repository if it\n>        represents a submodule.\n>\n>     2. The stage number, which is normally 0, but entries with\n>        higher stages for the same path are used during a conflicted\n>        merge.\n>\n>     3. The path name for the index entry.\n>\n>> +3. The *file path*, for example `src/hello.py`\n>> +4. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n>> +   there's a merge conflict there can be multiple versions of the same\n>> +   filename in the index.\n>\n> If you are going by \"ls-files -s\" output, it may be better to swap 3\n> and 4 above for ease of understanding.\n\nGood point, will do.\n\n>> +It's extremely uncommon to look at the index directly: normally you'd\n>> +run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n>> +But you can use `git ls-files --stage` to see the index.\n>> +Here's the output of `git ls-files --stage` in a repository with 2 files:\n>> +\n>> +----\n>> +100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n>> +100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n>> +----\n>> +\n>> +[[reflogs]]\n>> +REFLOGS\n>> +-------\n>> +\n>> +Every time a branch, remote-tracking branch, or HEAD is updated, Git\n>> +updates a log called a \"reflog\" for that <<references,reference>>.\n>\n> If we want to avoid using word X while explaining X, then we can\n> rephrase it as \"Git updates a record in the reflog for that\n> reference\".\n\nI think the current phrasing is okay. I also didn't respond to some of the\nphrasing suggestions above if I didn't understand the goal of them.\nHope that's okay.\n"},{"id":"529834","messageId":"xmqqldkubwfp.fsf@gitster.g","threadId":"64244","inReplyTo":"5b078fae-6fe9-4fde-ba84-1070761c168b@app.fastmail.com","subject":"Re: [PATCH v4] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-28T20:31:06Z","receivedAt":"2025-10-28T20:31:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> Agreed, I really like this idea. Came up with the following, which I'll put at\n> the bottom of the \"References\" section if I don't come up with a better idea.\n> (I don't feel strongly about where exactly it should go):\n>\n> NOTE: Objects will only be deleted if they aren't \"reachable\" from any reference.\n> An object is \"reachable\" if we can find it by following tags to whatever\n> they tag, commits to their parents or trees, and trees to the trees or\n> blobs that they contain.\n> For example, if you amend a commit, with `git commit --amend`,\n> the old commit will usually not be reachable, so it may be deleted eventually.\n\nOther reachability anchors exist, like the index and reflog entries,\nbut we have to stop at somewhere.  I am fine if we do not mention\nthem explicitly for the sake of simplicity.\n\n>>> +Here's how each type of object is structured:\n>>> +\n>>> +[[commit]]\n>>> +commit::\n>>> +    A commit contains the full directory structure of every file\n>>> +    in that version of the repository and each file's contents.\n>>\n>> What you are describing here is more of the property of a tree; a\n>> commit is a bit richer.\n>>\n>>     A commit records a snapshot of the every file in the project at\n>>     one point in time, records who contributed to create such a\n>>     snapshot and why, and how that particular snapshot relates to\n>>     other snapshots in the history.\n>\n> I don't understand the goal of explaining a commit in detail in\n> paragraph form when we already explain everything in a commit right\n> below this.\n>\n> My goal of this intro sentence is just to emphasize what I think is the\n> least obvious point in that list, which is that commits contain every file. \n>\n> Happy to change it to something shorter like\n> \"A commit records a snapshot of the every file in the project\" if you\n> prefer that wording.\n\nNot really.  Somebody who is skimming, who reads just the headline\nwithout reading enumeration, would not be able to tell differenes\nbetween a tree and a commit.  Your enumeration lists _what_ is\nrecorded, the headline I gave you above explains _what_ they are\nrecorded _for_.\n\n> I think the current phrasing is okay. I also didn't respond to some of the\n> phrasing suggestions above if I didn't understand the goal of them.\n> Hope that's okay.\n\nIf you do not understand, please ask ;-)\n"},{"id":"529986","messageId":"pull.1981.v5.git.1761856336360.gitgitgadget@gmail.com","threadId":"64244","inReplyTo":"pull.1981.v4.git.1761593537924.gitgitgadget@gmail.com","subject":"[PATCH v5] doc: add an explanation of Git's data model","fromName":"Julia Evans via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-30T20:32:16Z","receivedAt":"2025-10-30T20:32:19Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"From: Julia Evans <julia@jvns.ca>\n\nGit very often uses the terms \"object\", \"reference\", or \"index\" in its\ndocumentation.\n\nHowever, it's hard to find a clear explanation of these terms and how\nthey relate to each other in the documentation. The closest candidates\ncurrently are:\n\n1. `gitglossary`. This makes a good effort, but it's an alphabetically\n    ordered dictionary and a dictionary is not a good way to learn\n    concepts. You have to jump around too much and it's not possible to\n    present the concepts in the order that they should be explained.\n2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n   This is a nice document to have, but it's not necessary to learn how\n   `update-index` works to understand Git's data model, and we should\n   not be requiring users to learn how to use the \"plumbing\" commands\n   if they want to learn what the term \"index\" or \"object\" means.\n3. `gitrepository-layout`. This is a great resource, but it includes a\n   lot of information about configuration and internal implementation\n   details which are not related to the data model. It also does\n   not explain how commits work.\n\nThe result of this is that Git users (even users who have been using\nGit for 15+ years) struggle to read the documentation because they don't\nknow what the core terms mean, and it's not possible to add links\nto help them learn more.\n\nAdd an explanation of Git's data model. Some choices I've made in\ndeciding what \"core data model\" means:\n\n1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n   if those are intended to be user facing or if they're more like\n   internal implementation details.\n2. Don't talk about submodules other than by mentioning how they\n   relate to trees. This is because Git has a lot of special features,\n   and explaining how they all work exhaustively could quickly go\n   down a rabbit hole which would make this document less useful for\n   understanding Git's core behaviour.\n3. Don't discuss the structure of a commit message\n   (first line, trailers etc).\n4. Don't mention configuration.\n5. Don't mention the `.git` directory, to avoid getting too much into\n   implementation details\n\nSigned-off-by: Julia Evans <julia@jvns.ca>\n---\n    doc: Add a explanation of Git's data model\n    \n    Changes in v2:\n    \n    The biggest change is to remove all mentions of the .git directory, and\n    explain references in a way that doesn't refer to \"directories\" at all,\n    and instead talks about the \"hierarchy\" (from Kristoffer and Patrick's\n    reviews).\n    \n    Also:\n    \n     * objects: Mention that an object ID is called an \"object name\", and\n       update the glossary to include the term \"object ID\" (from Junio's\n       review)\n     * objects: Replace \"SHA-1 hash\" with \"cryptographic hash\" which is more\n       accurate (from Patrick's review)\n     * blobs: Made the explanation of git gc a little higher level and took\n       some ideas from Patrick's suggested wording (from Patrick's and\n       Kroftoffer's reviews)\n     * commits: Mention that tag objects and commits can optionally have\n       other fields. I didn't mention the GPG signature specifically, but\n       don't have any objections to adding it. (from Patrick and Junio's\n       reviews)\n     * commits: Remove one of the mentions of git gc, since it perhaps opens\n       up too much of a rabbit hole: \"how does git gc decide which commits\n       to clean up?\". (from Kristoffer's review)\n     * tag objects: Add an example of how a tag object is represented (from\n       user feedback on the draft)\n     * index: Use the term \"file mode\" instead of \"permissions\", and list\n       all allowed file modes (from Patrick's review)\n     * index: Use \"stage number\" instead of \"number\" for index entries (from\n       Patrick's review)\n     * reflogs: Remove \"any ref can be logged\", it raises some questions of\n       \"how do you tell Git to log a ref that it isn't normally logging?\"\n       and my guess is that it's uncommon to ask Git to log more refs. I\n       don't think it's a \"lie\" to omit this but I can bring it back if\n       folks disagree. (from Patrick's review)\n     * reflogs: Fix an error I noticed in the explanation of reflogs: tags\n       aren't logged by default and remote-tracking branches are, according\n       to man git-config\n     * branches and tags: Be clearer about how branches are usually updated\n       (by committing), and make it a little more obvious that only branches\n       can be checked out. This is a bit tricky because using the word\n       \"check out\" introduces a rabbit hole that I want to avoid (what does\n       \"check out\" mean?). I've dealt this by just talking about the\n       \"current branch\" (HEAD) since that is defined here, and making it\n       more explicit that HEAD must either be a branch or a commit, there's\n       no \"HEAD is a tag\" option. (from Patrick's review)\n     * tags: Explain the differences between annotated and lightweight tags\n       (this is the main piece of user feedback I've gotten on the draft so\n       far)\n     * Various style/typo changes (\"2 or more\", linkgit:git-gc[1], removed\n       extra asterisks, added empty SYNOPSIS, \"commits -> tags\" typo fix,\n       add to meson build)\n    \n    non-changes:\n    \n     * I still haven't mentioned things that aren't part of the \"data\n       model\", like revision params and configuration. I think there could\n       be a place for them but I haven't found it yet.\n     * tag objects: I noticed that there's a \"tag\" header field in tag\n       objects (like tag v1.0.0) but I didn't mention it yet because I\n       couldn't figure out what the purpose of that field is (I thought the\n       tag name was stored in the reference, why is it duplicated in the tag\n       object?)\n    \n    Changes in v3:\n    \n    I asked for feedback from Git users on Mastodon and got 220 pieces of\n    feedback from 48 different users. People seemed very excited to read\n    about Git's data model. Usually I judge explanations by what folks\n    report learning from them. Here people reported learning:\n    \n     * how branches are stored (that a branch is \"a name for a commit\")\n     * how objects work\n     * that Git has separate \"author\" and \"committer\" fields\n     * that amending a commit does not change it\n     * that a tree is \"just a directory\" (not something more complicated),\n       and how trees are stored\n     * that Git repos can contain symlinks\n     * that Git saves modes separately from the OS.\n     * how the stage number works\n     * that when you git add a file, Git will create an object\n     * that third-party tools can create their own refs.\n     * that the reflog stores the history of branches (not just HEAD), and\n       what reflogs are for\n    \n    Also (of course) there were quite a few points of confusion! The main 4\n    pieces of feedback were\n    \n     1. The index section doesn't explain what the word \"staged\" means, and\n        one person says that it makes it sounds like only files that you\n        \"git add\"ed are in the index. Rewrite the explanation to avoid using\n        the word \"staged\" to define the index and instead define the word\n        \"staging\".\n     2. Explain the difference between \"annotated tags\" and \"lightweight\n        tags\" (done)\n     3. Add examples for tag objects and reflogs (done)\n     4. Mention a little more about where things are stored in the .git\n        directory, which I'd removed in v2. This seems most important for\n        .git/refs, so I added a hopefully accurate note about how refs are\n        stored by default, with a comment about one of the major\n        implications. I did not discuss where objects or the index are\n        stored, because I don't think the implementation details of how\n        objects are stored are as important, and there are better tools for\n        viewing the \"raw\" state of objects and the index (with git cat-file\n        -p or git ls-files --staged).\n    \n    Here's every other change I made in response to the feedback, as well as\n    a few comments that I did not address.\n    \n    intro:\n    \n     * Give a 1-sentence intro to \"reflog\"\n    \n    objects:\n    \n     * people really like having git ls-files --stage as a way to view the\n       index, so add git cat-file -p as well in a note\n    \n    commits:\n    \n     * 2 people asked \"Are commits stored as a diff?\". Say that diffs are\n       calculated at runtime, this is very important.\n     * The order the fields are given in don't match the order in the\n       example. Make them match.\n     * \"All the files in the commit, stored as a tree\" is throwing a few\n       people off. Be clearer that it's the tree ID of the base directory.\n     * Several people asked \"What's the difference between an author and\n       committer? I added an example using git cherry-pick that I'm not 100%\n       happy with (what if the reader doesn't know what cherry-pick does?).\n       There might be a better example to give here.\n     * In the note about commits being amended: one person suggested saying\n       \"creates a new commit with the same parent\" to make it clearer what\n       the relationship between the new and old commit are. I liked that\n       idea so I did it.\n    \n    trees:\n    \n     * file modes. 2 people want to know more about \"The file mode, for\n       example 100644\". Also 2 people are curious about what relationship\n       these have to Unix permissions. Say that they're inspired by Unix\n       permissions, and move the list of possible file modes up to make the\n       relationship clearer\n     * On \"so git-gc(1) periodically compresses objects to save disk space\",\n       there are a few follow up comments wondering about more, which makes\n       me think the comment about compression is actually a distraction. Say\n       something simpler instead, (\"Git only needs to store new versions of\n       files which were changed in that commit\"), from Junio's suggestion\n     * Re \"commit (a Git submodule)\": 2 people say it's not clear how trees\n       relate to submodules. Say that it refers to a commit in a different\n       repository.\n     * One person says they're not sure if the \"object ID\" is a hash. Link\n       it to the definition of \"object ID\".\n    \n    tag objects:\n    \n     * Requests for an example, added one.\n     * Requests to explain the difference between \"lightweight\" and\n       \"annotated\" tags, added it.\n    \n    tags:\n    \n     * one person thinks \"It’s expected that a tag will never change after\n       you create it.\" is too strong (since of course you can change it with\n       git tag -f). Say instead that tags are \"usually\" not changed.\n    \n    HEAD:\n    \n     * Several people are asking for more detail about detached HEAD state.\n       There's actually quite a lot to talk about here (what it means, how\n       it happens, what it implies, and how you might adjust your workflow\n       to avoid it by using git switch). I don't think we can get into all\n       of that here, so refer to the DETACHED HEAD section of git-checkout\n       instead. I'm not totally happy with the current version of that\n       section but that seems like the most practical solution right now.\n    \n    remote-tracking branches:\n    \n     * discuss refs/remotes/<remote>/HEAD.\n    \n    the index:\n    \n     * \"permissions\" should be \"file mode\" (like with trees). Changed.\n     * \"filename\" should be \"file path\". Changed.\n     * the stage number can only be 0, 1, 2, or 3, since it's 2 bits. Also\n       maybe say that the numbers have specific meanings. Said it can only\n       be 0/1/2/3 but did not give the specific meanings.\n    \n    reflogs\n    \n     * Request for an example. Added one.\n     * It's not clear if there's one reflog per branch/tag/HEAD, or if\n       there's one universal reflog. Make this clearer.\n     * Mention the role of the reflog in retrieving \"lost\" commits or\n       undoing bad rebases.\n    \n    Not fixed:\n    \n     * intro: A couple of people say that it's confusing that tags are both\n       \"an object\" and \"a reference\". Handled this by just explaining the\n       difference between an annotated and a lightweight tag further down.\n       I'd like to make this clearer in the intro but not sure if there's a\n       way to do it.\n     * commits and tag objects: one person asks if there's a reference for\n       the other \"optional fields\", like \"encoding\" and \"gpgsig\". I couldn't\n       find one, so left this as is.\n     * HEAD: A couple of people ask if there are any other symbolic\n       references other than HEAD, or if they can make their own symbolic\n       references. I don't know the answer to this.\n     * HEAD: the HEAD: HEAD thing looks weird, it made more sense when it\n       was HEAD: .git/HEAD. Will think about this.\n     * reflogs: One person asks: if reflogs only store local changes, why\n       does it track the user who made the change? Is that for remote\n       operations like fetches and pulls? Or for cases where more than one\n       user is using the same repo on a system? I don't know the answer to\n       this.\n     * reflogs: How can you see the full data in the reflog? git reflog show\n       doesn't list the user who made the change. git reflog show <refname>\n       --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso seems to work but\n       it's really a mouthful, not sure it's useful to include all that.\n     * index: Is it worth mentioning that the index can be locked? I don't\n       have an opinion about this.\n     * other: One person asks what a \"working tree\" is. It made me wonder if\n       \"the current working directory\" has a place in Git's data model. My\n       feeling is \"no\" but I could be convinced otherwise.\n     * overall: \"How can Git be so fast? If I switch branches, how does it\n       figure out what to add, remove or replace?\". I don't think this is\n       the right place for that discussion but it would\n     * there are some docs CI errors I haven't figured out yet (IDREF\n       attribute linkend references an unknown ID \"tree\")\n    \n    changes in v4:\n    \n    This is a combination of trying to make some of the intro text a little\n    more \"friendly\" for someone new to Git's data model, avoiding implying\n    things that are false, and removing information that isn't relevant to\n    the data model.\n    \n    intro:\n    \n     * Add a 1-line description of what a \"reflog\" is (from user feedback)\n    \n    objects:\n    \n     * Start with a \"friendly\" description of what an object is, similar to\n       what we do for references and the reflog\n     * Rename \"commits\" to \"commit\" and similarly for trees etc (from\n       Junio's review)\n     * Remove the explanation of what git cat-file -p does, since it might\n       be misleading and if people want to know they can read the man page\n       (from Junio's review)\n    \n    commits:\n    \n     * Start by saying that the commit contains the full directory structure\n       of all the files (from Junio's comment about how it may not be clear\n       that the commit contains all the files' exact contents at the time of\n       the commit)\n     * Remove the comment about cherry-pick (from Junio's review)\n     * Replace \"ask Git for a diff\" with \"ask Git to show the commit with\n       git show\" (from Junio's review)\n    \n    trees:\n    \n     * Make the description a little more friendly\n     * Reorder so that \"type\" is defined before we refer to the \"type\"\n     * Say that file modes are \"only spiritually related\" to Unix\n       permissions instead of talking about what Git \"supports\" (from\n       Junio's review)\n    \n    blobs:\n    \n     * Try to make it clearer how \"commits use relatively little disk space\"\n       is true while not implying that commits are diffs, by using an\n       example (from Junio's review)\n    \n    branches:\n    \n     * Replace \"a branch is a name for a commit ID\" with \"a branch refers to\n       a commit ID\" (except in the intro sentence for the \"references\"\n       section). Similarly for tags etc. (from Junio's review)\n     * Remove the note about how branches are stored in .git (from Junio's\n       review)\n    \n    HEAD:\n    \n     * Be clearer that HEAD is not always the current branch, because there\n       may not be a current branch (from Junio's review)\n    \n    index:\n    \n     * Be a little more specific about how exactly the index is converted\n       into a commit. (from Junio's comment about how it's not clear what\n       \"every file in the repository\" means)\n    \n    reflog:\n    \n     * Be clearer that there are many reflogs (one for each reference with a\n       log), not just one reflog (from Junio and Patrick's reviews)\n     * Omit the user and \"Before\" commit IDs from the list of fields,\n       because you usually don't see them (from Junio's review)\n     * Show the output of git reflog main in the example instead of the\n       contents of the reflog file, to avoid showing the user and before\n       commit ID\n    \n    changes in v5:\n    \n    Mostly smaller tweaks this time. The only major addition is to add a\n    note about how unreachable objects may be deleted.\n    \n    From Junio's review:\n    \n     * Remove \"type\" in the description of what's in a tree (since I have\n       learned that is not a separate field, it's part of the file mode)\n     * Fix a typo (\"these these\")\n     * Remove the intro sentence about what a \"commit\" is and instead only\n       describe its contents in the list of fields, to avoid implying that a\n       commit is the same as a tree\n     * Say \"Unix file modes\" instead of \"Unix permissions\"\n     * In the tag objects contents: make \"ID\" and \"type\" separate list items\n       since they're separate fields\n     * in the index section:\n       * list all of the possible file modes (since from my understanding\n         there are fewer allowed file modes here than in a tree)\n       * mention that the object can be either a commit or blob\n       * make the order match the order in git ls-files\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/1981\n\nRange-diff vs v4:\n\n 1:  92249b5b08 ! 1:  d342255dad doc: add an explanation of Git's data model\n     @@ Documentation/gitdatamodel.adoc (new)\n      +\n      +[[commit]]\n      +commit::\n     -+    A commit contains the full directory structure of every file\n     -+    in that version of the repository and each file's contents.\n     -+    It has these these required fields\n     ++    A commit contains these required fields\n      +    (though there are other optional fields):\n      ++\n     -+1. The *files* in the commit, stored as the *<<tree,tree>>* ID\n     ++1. The full directory structure of all the files in that version of the\n     ++   repository and each file's contents, stored as the *<<tree,tree>>* ID\n      +   of the commit's base directory.\n      +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n      +  regular commits have 1 parent, merge commits have 2 or more parents\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    It lists, for each item in the tree:\n      ++\n      +1. The *filename*, for example `hello.py`\n     -+2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),\n     -+  or <<commit,`commit`>> (a Git submodule, which is a\n     -+  commit from a different Git repository)\n     -+3. The *file mode*. Git has these file modes. which are only\n     -+   spiritually related to Unix permissions:\n     ++2. The *file mode*. Git has these file modes. which are only\n     ++   spiritually related to Unix file modes:\n      ++\n     -+  - `100644`: regular file (with type `blob`)\n     ++  - `100644`: regular file (with <<object,object type>> `blob`)\n      +  - `100755`: executable file (with type `blob`)\n      +  - `120000`: symbolic link (with type `blob`)\n      +  - `040000`: directory (with type `tree`)\n      +  - `160000`: gitlink, for use with submodules (with type `commit`)\n      +\n     -+4. The <<object-id,*object ID*>> with the contents of the file or directory\n     ++3. The <<object-id,*object ID*>> with the contents of the file or directory\n      ++\n      +For example, this is how a tree containing one directory (`src`) and one file\n      +(`README.md`) is stored:\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    Tag objects contain these required fields\n      +    (though there are other optional fields):\n      ++\n     -+1. The *ID* and *type* of the object (often a commit) that they reference\n     -+2. The *tagger* and tag date\n     -+3. A *tag message*, similar to a commit message\n     ++1. The object *ID* it references\n     ++2. The object *type*\n     ++3. The *tagger* and tag date\n     ++4. A *tag message*, similar to a commit message\n      +\n      +Here's how an example tag object is stored:\n      +\n     @@ Documentation/gitdatamodel.adoc (new)\n      +Git may also create references other than `HEAD` at the base of the\n      +hierarchy, like `ORIG_HEAD`.\n      +\n     ++NOTE: Git may delete objects that aren't \"reachable\" from any reference.\n     ++An object is \"reachable\" if we can find it by following tags to whatever\n     ++they tag, commits to their parents or trees, and trees to the trees or\n     ++blobs that they contain.\n     ++For example, if you amend a commit, with `git commit --amend`,\n     ++the old commit will usually not be reachable, so it may be deleted eventually.\n     ++Reachable objects will never be deleted.\n     ++\n      +[[index]]\n      +THE INDEX\n      +---------\n     @@ Documentation/gitdatamodel.adoc (new)\n      +\n      +Each index entry has 4 fields:\n      +\n     -+1. The *<<tree,file mode>>*\n     -+2. The *<<blob,blob>> ID* of the file\n     -+3. The *file path*, for example `src/hello.py`\n     -+4. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n     ++1. The *file mode*, which must be one of:\n     ++  - `100644`: regular file (with <<object,object type>> `blob`)\n     ++  - `100755`: executable file (with type `blob`)\n     ++  - `120000`: symbolic link (with type `blob`)\n     ++  - `160000`: gitlink, for use with submodules (with type `commit`)\n     ++2. The *<<blob,blob>>* ID of the file,\n     ++   or (rarely) the *<<commit,commit>>* ID of the submodule\n     ++3. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n      +   there's a merge conflict there can be multiple versions of the same\n      +   filename in the index.\n     ++4. The *file path*, for example `src/hello.py`\n      +\n      +It's extremely uncommon to look at the index directly: normally you'd\n      +run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n\n\n Documentation/Makefile              |   1 +\n Documentation/gitdatamodel.adoc     | 296 ++++++++++++++++++++++++++++\n Documentation/glossary-content.adoc |   4 +-\n Documentation/meson.build           |   1 +\n 4 files changed, 300 insertions(+), 2 deletions(-)\n create mode 100644 Documentation/gitdatamodel.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 6fb83d0c6e..5f4acfacbd 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n MAN7_TXT += gitcore-tutorial.adoc\n MAN7_TXT += gitcredentials.adoc\n MAN7_TXT += gitcvs-migration.adoc\n+MAN7_TXT += gitdatamodel.adoc\n MAN7_TXT += gitdiffcore.adoc\n MAN7_TXT += giteveryday.adoc\n MAN7_TXT += gitfaq.adoc\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nnew file mode 100644\nindex 0000000000..1cefbb4833\n--- /dev/null\n+++ b/Documentation/gitdatamodel.adoc\n@@ -0,0 +1,296 @@\n+gitdatamodel(7)\n+===============\n+\n+NAME\n+----\n+gitdatamodel - Git's core data model\n+\n+SYNOPSIS\n+--------\n+gitdatamodel\n+\n+DESCRIPTION\n+-----------\n+\n+It's not necessary to understand Git's data model to use Git, but it's\n+very helpful when reading Git's documentation so that you know what it\n+means when the documentation says \"object\", \"reference\" or \"index\".\n+\n+Git's core operations use 4 kinds of data:\n+\n+1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+2. <<references,References>>: branches, tags,\n+   remote-tracking branches, etc\n+3. <<index,The index>>, also known as the staging area\n+4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n+\n+[[objects]]\n+OBJECTS\n+-------\n+\n+All of the commits and files in a Git repository are stored as \"Git objects\".\n+Git objects never change after they're created, and every object has an ID,\n+like `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+\n+This means that if you have an object's ID, you can always recover its\n+exact contents as long as the object hasn't been deleted.\n+\n+Every object has:\n+\n+[[object-id]]\n+1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n+  type and contents.\n+  It's fast to look up a Git object using its ID.\n+  This is usually represented in hexadecimal, like\n+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+2. a *type*. There are 4 types of objects:\n+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n+   and <<tag-object,tag objects>>.\n+3. *contents*. The structure of the contents depends on the type.\n+\n+Here's how each type of object is structured:\n+\n+[[commit]]\n+commit::\n+    A commit contains these required fields\n+    (though there are other optional fields):\n++\n+1. The full directory structure of all the files in that version of the\n+   repository and each file's contents, stored as the *<<tree,tree>>* ID\n+   of the commit's base directory.\n+2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n+  regular commits have 1 parent, merge commits have 2 or more parents\n+3. An *author* and the time the commit was authored\n+4. A *committer* and the time the commit was committed.\n+5. A *commit message*\n++\n+Here's how an example commit is stored:\n++\n+----\n+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n+author Maya <maya@example.com> 1759173425 -0400\n+committer Maya <maya@example.com> 1759173425 -0400\n+\n+Add README\n+----\n++\n+Like all other objects, commits can never be changed after they're created.\n+For example, \"amending\" a commit with `git commit --amend` creates a new\n+commit with the same parent.\n++\n+Git does not store the diff for a commit: when you ask Git to show\n+the commit with linkgit:git-show[1], it calculates the diff from its\n+parent on the fly.\n+\n+[[tree]]\n+tree::\n+    A tree is how Git represents a directory.\n+    It can contain files or other trees (which are subdirectories).\n+    It lists, for each item in the tree:\n++\n+1. The *filename*, for example `hello.py`\n+2. The *file mode*. Git has these file modes. which are only\n+   spiritually related to Unix file modes:\n++\n+  - `100644`: regular file (with <<object,object type>> `blob`)\n+  - `100755`: executable file (with type `blob`)\n+  - `120000`: symbolic link (with type `blob`)\n+  - `040000`: directory (with type `tree`)\n+  - `160000`: gitlink, for use with submodules (with type `commit`)\n+\n+3. The <<object-id,*object ID*>> with the contents of the file or directory\n++\n+For example, this is how a tree containing one directory (`src`) and one file\n+(`README.md`) is stored:\n++\n+----\n+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n+----\n+\n+[[blob]]\n+blob::\n+    A blob object contains a file's contents.\n++\n+When you make a commit, Git stores the full contents of each file that\n+you changed as a blob.\n+For example, if you have a commit that changes 2 files in a repository\n+with 1000 files, that commit will create 2 new blobs, and use the\n+previous blob ID for the other 998 files.\n+This means that commits can use relatively little disk space even in a\n+very large repository.\n+\n+[[tag-object]]\n+tag object::\n+    Tag objects contain these required fields\n+    (though there are other optional fields):\n++\n+1. The object *ID* it references\n+2. The object *type*\n+3. The *tagger* and tag date\n+4. A *tag message*, similar to a commit message\n+\n+Here's how an example tag object is stored:\n+\n+----\n+object 750b4ead9c87ceb3ddb7a390e6c7074521797fb3\n+type commit\n+tag v1.0.0\n+tagger Maya <maya@example.com> 1759927359 -0400\n+\n+Release version 1.0.0\n+----\n+\n+NOTE: All of the examples in this section were generated with\n+`git cat-file -p <object-id>`.\n+\n+[[references]]\n+REFERENCES\n+----------\n+\n+References are a way to give a name to a commit.\n+It's easier to remember \"the changes I'm working on are on the `turtle`\n+branch\" than \"the changes are in commit bb69721404348e\".\n+Git often uses \"ref\" as shorthand for \"reference\".\n+\n+References can either refer to:\n+\n+1. An object ID, usually a <<commit,commit>> ID\n+2. Another reference. This is called a \"symbolic reference\".\n+\n+References are stored in a hierarchy, and Git handles references\n+differently based on where they are in the hierarchy.\n+Most references are under `refs/`. Here are the main types:\n+\n+[[branch]]\n+branches: `refs/heads/<name>`::\n+    A branch refers to a commit ID.\n+    That commit is the latest commit on the branch.\n++\n+To get the history of commits on a branch, Git will start at the commit\n+ID the branch references, and then look at the commit's parent(s),\n+the parent's parent, etc.\n+\n+[[tag]]\n+tags: `refs/tags/<name>`::\n+    A tag refers to a commit ID, tag object ID, or other object ID.\n+    There are two types of tags:\n+    1. \"Annotated tags\", which reference a <<tag-object,tag object>> ID\n+       which contains a tag message\n+    2. \"Lightweight tags\", which reference a commit, blob, or tree ID\n+       directly\n++\n+Even though branches and tags both refer to a commit ID, Git\n+treats them very differently.\n+Branches are expected to change over time: when you make a commit, Git\n+will update your <<HEAD,current branch>> to point to the new commit.\n+Tags are usually not changed after they're created.\n+\n+[[HEAD]]\n+HEAD: `HEAD`::\n+    `HEAD` is where Git stores your current <<branch,branch>>,\n+    if there is a current branch. `HEAD` can either be:\n++\n+1. A symbolic reference to your current branch, for example `ref:\n+   refs/heads/main` if your current branch is `main`.\n+2. A direct reference to a commit ID. In this case there is no current branch.\n+   This is called \"detached HEAD state\", see the DETACHED HEAD section\n+   of linkgit:git-checkout[1] for more.\n+\n+[[remote-tracking-branch]]\n+remote-tracking branches: `refs/remotes/<remote>/<branch>`::\n+    A remote-tracking branch refers to a commit ID.\n+    It's how Git stores the last-known state of a branch in a remote\n+    repository. `git fetch` updates remote-tracking branches. When\n+    `git status` says \"you're up to date with origin/main\", it's looking at\n+    this.\n++\n+`refs/remotes/<remote>/HEAD` is a symbolic reference to the remote's\n+default branch. This is the branch that `git clone` checks out by default.\n+\n+[[other-refs]]\n+Other references::\n+    Git tools may create references anywhere under `refs/`.\n+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n+    and linkgit:git-notes[1] all create their own references\n+    in `refs/stash`, `refs/bisect`, etc.\n+    Third-party Git tools may also create their own references.\n++\n+Git may also create references other than `HEAD` at the base of the\n+hierarchy, like `ORIG_HEAD`.\n+\n+NOTE: Git may delete objects that aren't \"reachable\" from any reference.\n+An object is \"reachable\" if we can find it by following tags to whatever\n+they tag, commits to their parents or trees, and trees to the trees or\n+blobs that they contain.\n+For example, if you amend a commit, with `git commit --amend`,\n+the old commit will usually not be reachable, so it may be deleted eventually.\n+Reachable objects will never be deleted.\n+\n+[[index]]\n+THE INDEX\n+---------\n+The index, also known as the \"staging area\", is a list of files and\n+the contents of each file, stored as a <<blob,blob>>.\n+You can add files to the index or update the contents of a file in the\n+index with linkgit:git-add[1]. This is called \"staging\" the file for commit.\n+\n+Unlike a <<tree,tree>>, the index is a flat list of files.\n+When you commit, Git converts the list of files in the index to a\n+directory <<tree,tree>> and uses that tree in the new <<commit,commit>>.\n+\n+Each index entry has 4 fields:\n+\n+1. The *file mode*, which must be one of:\n+  - `100644`: regular file (with <<object,object type>> `blob`)\n+  - `100755`: executable file (with type `blob`)\n+  - `120000`: symbolic link (with type `blob`)\n+  - `160000`: gitlink, for use with submodules (with type `commit`)\n+2. The *<<blob,blob>>* ID of the file,\n+   or (rarely) the *<<commit,commit>>* ID of the submodule\n+3. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n+   there's a merge conflict there can be multiple versions of the same\n+   filename in the index.\n+4. The *file path*, for example `src/hello.py`\n+\n+It's extremely uncommon to look at the index directly: normally you'd\n+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n+But you can use `git ls-files --stage` to see the index.\n+Here's the output of `git ls-files --stage` in a repository with 2 files:\n+\n+----\n+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n+----\n+\n+[[reflogs]]\n+REFLOGS\n+-------\n+\n+Every time a branch, remote-tracking branch, or HEAD is updated, Git\n+updates a log called a \"reflog\" for that <<references,reference>>.\n+This means that if you make a mistake and \"lose\" a commit, you can\n+generally recover the commit ID by running `git reflog <reference>`.\n+\n+A reflog is a list of log entries. Each entry has:\n+\n+1. The *commit ID*\n+2. *Timestamp* when the change was made\n+3. *Log message*, for example `pull: Fast-forward`\n+\n+Reflogs only log changes made in your local repository.\n+They are not shared with remotes.\n+\n+You can view a reflog with `git reflog <reference>`.\n+For example, here's the reflog for a `main` branch which has changed twice:\n+\n+----\n+$ git reflog main --date=iso --no-decorate\n+750b4ea main@{2025-09-29 15:17:05 -0400}: commit: Add README\n+4ccb6d7 main@{2025-09-29 15:16:48 -0400}: commit (initial): Initial commit\n+----\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\ndiff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc\nindex e423e4765b..20ba121314 100644\n--- a/Documentation/glossary-content.adoc\n+++ b/Documentation/glossary-content.adoc\n@@ -297,8 +297,8 @@ This commit is referred to as a \"merge commit\", or sometimes just a\n \tidentified by its <<def_object_name,object name>>. The objects usually\n \tlive in `$GIT_DIR/objects/`.\n \n-[[def_object_identifier]]object identifier (oid)::\n-\tSynonym for <<def_object_name,object name>>.\n+[[def_object_identifier]]object identifier, object ID, oid::\n+\tSynonyms for <<def_object_name,object name>>.\n \n [[def_object_name]]object name::\n \tThe unique identifier of an <<def_object,object>>.  The\ndiff --git a/Documentation/meson.build b/Documentation/meson.build\nindex e34965c5b0..ace0573e82 100644\n--- a/Documentation/meson.build\n+++ b/Documentation/meson.build\n@@ -192,6 +192,7 @@ manpages = {\n   'gitcore-tutorial.adoc' : 7,\n   'gitcredentials.adoc' : 7,\n   'gitcvs-migration.adoc' : 7,\n+  'gitdatamodel.adoc' : 7,\n   'gitdiffcore.adoc' : 7,\n   'giteveryday.adoc' : 7,\n   'gitfaq.adoc' : 7,\n\nbase-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n-- \ngitgitgadget\n"},{"id":"530035","messageId":"xmqqtszf2kro.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v5.git.1761856336360.gitgitgadget@gmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-31T14:44:43Z","receivedAt":"2025-10-31T14:44:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n> new file mode 100644\n> index 0000000000..1cefbb4833\n> --- /dev/null\n> +++ b/Documentation/gitdatamodel.adoc\n> @@ -0,0 +1,296 @@\n> +gitdatamodel(7)\n> +===============\n> +\n> +NAME\n> +----\n> +gitdatamodel - Git's core data model\n> +\n> +SYNOPSIS\n> +--------\n> +gitdatamodel\n> +\n> +DESCRIPTION\n> +-----------\n> +\n> +It's not necessary to understand Git's data model to use Git, but it's\n> +very helpful when reading Git's documentation so that you know what it\n> +means when the documentation says \"object\", \"reference\" or \"index\".\n> +\n> +Git's core operations use 4 kinds of data:\n> +\n> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n> +2. <<references,References>>: branches, tags,\n> +   remote-tracking branches, etc\n> +3. <<index,The index>>, also known as the staging area\n> +4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n> +\n> +[[objects]]\n> +OBJECTS\n> +-------\n> +\n> +All of the commits and files in a Git repository are stored as \"Git objects\".\n> +Git objects never change after they're created, and every object has an ID,\n> +like `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n> +\n> +This means that if you have an object's ID, you can always recover its\n> +exact contents as long as the object hasn't been deleted.\n> +\n> +Every object has:\n> +\n> +[[object-id]]\n> +1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n> +  type and contents.\n> +  It's fast to look up a Git object using its ID.\n> +  This is usually represented in hexadecimal, like\n> +  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n> +2. a *type*. There are 4 types of objects:\n> +   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n> +   and <<tag-object,tag objects>>.\n> +3. *contents*. The structure of the contents depends on the type.\n> +\n> +Here's how each type of object is structured:\n> +\n> +[[commit]]\n> +commit::\n> +    A commit contains these required fields\n> +    (though there are other optional fields):\n> ++\n> +1. The full directory structure of all the files in that version of the\n> +   repository and each file's contents, stored as the *<<tree,tree>>* ID\n> +   of the commit's base directory.\n\n\"base directory\" is a new term; I think we most often use\n\"top-level\" directory (in various spellings).\n\n$ git grep -e 'base directory' -e 'level directory' Documentation/\n\n> +[[tree]]\n> +tree::\n> +    A tree is how Git represents a directory.\n> +    It can contain files or other trees (which are subdirectories).\n> +    It lists, for each item in the tree:\n> ++\n> +1. The *filename*, for example `hello.py`\n> +2. The *file mode*. Git has these file modes. which are only\n\n\"has these\" -> \"uses only these\" to clarify that this is an\nexhaustive enumeration and users cannot invent 100664 and others,\nwhich is a mistake Git itself used to make/allow.\n\n> +[[tag-object]]\n> +tag object::\n> +    Tag objects contain these required fields\n> +    (though there are other optional fields):\n> ++\n> +1. The object *ID* it references\n> +2. The object *type*\n\nI would rephrase these to\n\n    1. The *ID* of the object it references\n    2. The *type* of the object it references\n\nbecause (1) a tag object references another object, not ID.  To name\nthe object it reference, it uses the object name of it, but just\nlike your name is not you, object name is not the object (it merely\nis *one* way to refer to it). (2) unless it is very clear to readers\nthat \"The object\" in 1. and 2. refer to the same object, 2. invites\na question \"type of which object?\".\n\n> +[[branch]]\n> +branches: `refs/heads/<name>`::\n> +    A branch refers to a commit ID.\n\nA branch refers to a commit object (by its ID).  Ditto for tags.\n\n> +NOTE: Git may delete objects that aren't \"reachable\" from any reference.\n> +An object is \"reachable\" if we can find it by following tags to whatever\n> +they tag, commits to their parents or trees, and trees to the trees or\n> +blobs that they contain.\n> +For example, if you amend a commit, with `git commit --amend`,\n> +the old commit will usually not be reachable, so it may be deleted eventually.\n> +Reachable objects will never be deleted.\n\nVery good write-up.  As we would touch upon reflog later in the same\ndocument, we may want to extend the \"amend\" example a bit, perhaps\nlike\n\n    Note: Git never deletes objects that are \"reachable\".  An object\n    is \"reachable\" if ....  An unreachable object may be deleted.\n\n    For example, ... a newly created commit will replace the old\n    commit and the current branch ref points at the new commit.  The\n    old commit is recorded in the <<reflogs,reflog>> of the current\n    branch, so it is still \"reachable\", but sufficiently old reflog\n    entries are expired away, the old commit may become unreachable\n    at that point, and would get deleted.\n\nOther than the above, I found everything very nicely written.\n\nThanks.\n"},{"id":"530052","messageId":"xmqqcy62213a.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v5.git.1761856336360.gitgitgadget@gmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-31T21:49:45Z","receivedAt":"2025-10-31T21:49:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The document refers to <<object,object type>> but the id to refer\nthe descripion of the object is defined as [[objects]]; we need a\nband-aid like this one to pass GitHub Actions CI.\n\nAs description for individual object types are titled singular like\n[[commit]], [[blob]], etc., this band-aid drops the plural 's' from\nthe tail of [[objects]], but as long as we are consistent, of course,\nwe could go the other direction.\n\n\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nindex 1cefbb4833..eaab3f800b 100644\n--- a/Documentation/gitdatamodel.adoc\n+++ b/Documentation/gitdatamodel.adoc\n@@ -18,13 +18,13 @@ means when the documentation says \"object\", \"reference\" or \"index\".\n \n Git's core operations use 4 kinds of data:\n \n-1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+1. <<object,Objects>>: commits, trees, blobs, and tag objects\n 2. <<references,References>>: branches, tags,\n    remote-tracking branches, etc\n 3. <<index,The index>>, also known as the staging area\n 4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n \n-[[objects]]\n+[[object]]\n OBJECTS\n -------\n \n-- \n2.51.2-719-gbbf487eab4\n\n"},{"id":"530094","messageId":"aQhcZwv0PdwNc6RW@pks.im","threadId":"64244","inReplyTo":"pull.1981.v5.git.1761856336360.gitgitgadget@gmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-03T07:40:23Z","receivedAt":"2025-11-03T07:40:30Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Thu, Oct 30, 2025 at 08:32:16PM +0000, Julia Evans via GitGitGadget wrote:\n> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n> new file mode 100644\n> index 0000000000..1cefbb4833\n> --- /dev/null\n> +++ b/Documentation/gitdatamodel.adoc\n[snip]\n> +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n> +  regular commits have 1 parent, merge commits have 2 or more parents\n> +3. An *author* and the time the commit was authored\n> +4. A *committer* and the time the commit was committed.\n> +5. A *commit message*\n\nNit: The punctuation is a bit inconsistent here, as some list items have\na trailing dot while others don't.\n\n> +[[references]]\n> +REFERENCES\n> +----------\n> +\n> +References are a way to give a name to a commit.\n> +It's easier to remember \"the changes I'm working on are on the `turtle`\n> +branch\" than \"the changes are in commit bb69721404348e\".\n> +Git often uses \"ref\" as shorthand for \"reference\".\n> +\n> +References can either refer to:\n> +\n> +1. An object ID, usually a <<commit,commit>> ID\n> +2. Another reference. This is called a \"symbolic reference\".\n\nSame here.\n\nOther than these two nits and Junio's comments I think this is in a good\nenough shape. Thanks for working on this!\n\nPatrick\n"},{"id":"530095","messageId":"aQhcbHJjiI5GtV6Y@pks.im","threadId":"64244","inReplyTo":"xmqqtszf2kro.fsf@gitster.g","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-11-03T07:40:28Z","receivedAt":"2025-11-03T07:40:33Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Oct 31, 2025 at 07:44:43AM -0700, Junio C Hamano wrote:\n> \"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n[snip]\n> > +[[commit]]\n> > +commit::\n> > +    A commit contains these required fields\n> > +    (though there are other optional fields):\n> > ++\n> > +1. The full directory structure of all the files in that version of the\n> > +   repository and each file's contents, stored as the *<<tree,tree>>* ID\n> > +   of the commit's base directory.\n> \n> \"base directory\" is a new term; I think we most often use\n> \"top-level\" directory (in various spellings).\n> \n> $ git grep -e 'base directory' -e 'level directory' Documentation/\n\nWe'd refer to the top-level directory when talking about the worktree.\nBut what's referenced here is not referring to the worktree, but to the\ncommit's tree. And here I think we rather consistently use \"root tree\",\ndon't we? Our docs already mention \"root tree\" in several contexts.\n\nPatrick\n"},{"id":"530126","messageId":"xmqqwm47unw3.fsf@gitster.g","threadId":"64244","inReplyTo":"aQhcbHJjiI5GtV6Y@pks.im","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-03T15:38:52Z","receivedAt":"2025-11-03T15:38:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> We'd refer to the top-level directory when talking about the worktree.\n> But what's referenced here is not referring to the worktree, but to the\n> commit's tree. And here I think we rather consistently use \"root tree\",\n> don't we? Our docs already mention \"root tree\" in several contexts.\n\nAh, thanks.  I wasn't aware that we use the phrase \"root tree\"; I\nrecall that I've always said something awkward like \"the tree that\ncorresponds to the top-level of your working tree\", due to lack of\nthat exact word.\n\nIt would be nice to add it to Documentation/glossary-content.adoc,\nperhaps?  Here is my attempt (I am not committing this, and I won't\nbe polishing it myself, but recording it as #leftoverbit material\nfor somebody else to polish and make it a part of our documentation\nset).\n\n Documentation/glossary-content.adoc | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git c/Documentation/glossary-content.adoc w/Documentation/glossary-content.adoc\nindex e423e4765b..bdf469f137 100644\n--- c/Documentation/glossary-content.adoc\n+++ w/Documentation/glossary-content.adoc\n@@ -627,6 +627,12 @@ the `refs/tags/` hierarchy is used to represent local tags..\n \tTo throw away part of the development, i.e. to assign the\n \t<<def_head,head>> to an earlier <<def_revision,revision>>.\n \n+[[def_root_tree]]root tree::\n+\tThe tree objct that corresponds to the top-level directory\n+\tof a checkout of the project.  A <<def_commit,commit>> object\n+\tholds a snapshot of the project state by recording the object\n+\tname of its root tree.\n+\n [[def_SCM]]SCM::\n \tSource code management (tool).\n \n"},{"id":"530149","messageId":"8b70796e-b5a4-4f70-8b27-c0ed80d1fc4d@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqtszf2kro.fsf@gitster.g","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-03T19:43:39Z","receivedAt":"2025-11-03T19:44:01Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Fri, Oct 31, 2025, at 10:44 AM, Junio C Hamano wrote:\n> \"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n>> new file mode 100644\n>> index 0000000000..1cefbb4833\n>> --- /dev/null\n>> +++ b/Documentation/gitdatamodel.adoc\n>> @@ -0,0 +1,296 @@\n>> +gitdatamodel(7)\n>> +===============\n>> +\n>> +NAME\n>> +----\n>> +gitdatamodel - Git's core data model\n>> +\n>> +SYNOPSIS\n>> +--------\n>> +gitdatamodel\n>> +\n>> +DESCRIPTION\n>> +-----------\n>> +\n>> +It's not necessary to understand Git's data model to use Git, but it's\n>> +very helpful when reading Git's documentation so that you know what it\n>> +means when the documentation says \"object\", \"reference\" or \"index\".\n>> +\n>> +Git's core operations use 4 kinds of data:\n>> +\n>> +1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n>> +2. <<references,References>>: branches, tags,\n>> +   remote-tracking branches, etc\n>> +3. <<index,The index>>, also known as the staging area\n>> +4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n>> +\n>> +[[objects]]\n>> +OBJECTS\n>> +-------\n>> +\n>> +All of the commits and files in a Git repository are stored as \"Git objects\".\n>> +Git objects never change after they're created, and every object has an ID,\n>> +like `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n>> +\n>> +This means that if you have an object's ID, you can always recover its\n>> +exact contents as long as the object hasn't been deleted.\n>> +\n>> +Every object has:\n>> +\n>> +[[object-id]]\n>> +1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n>> +  type and contents.\n>> +  It's fast to look up a Git object using its ID.\n>> +  This is usually represented in hexadecimal, like\n>> +  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n>> +2. a *type*. There are 4 types of objects:\n>> +   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n>> +   and <<tag-object,tag objects>>.\n>> +3. *contents*. The structure of the contents depends on the type.\n>> +\n>> +Here's how each type of object is structured:\n>> +\n>> +[[commit]]\n>> +commit::\n>> +    A commit contains these required fields\n>> +    (though there are other optional fields):\n>> ++\n>> +1. The full directory structure of all the files in that version of the\n>> +   repository and each file's contents, stored as the *<<tree,tree>>* ID\n>> +   of the commit's base directory.\n>\n> \"base directory\" is a new term; I think we most often use\n> \"top-level\" directory (in various spellings).\n>\n> $ git grep -e 'base directory' -e 'level directory' Documentation/\n>\n>> +[[tree]]\n>> +tree::\n>> +    A tree is how Git represents a directory.\n>> +    It can contain files or other trees (which are subdirectories).\n>> +    It lists, for each item in the tree:\n>> ++\n>> +1. The *filename*, for example `hello.py`\n>> +2. The *file mode*. Git has these file modes. which are only\n>\n> \"has these\" -> \"uses only these\" to clarify that this is an\n> exhaustive enumeration and users cannot invent 100664 and others,\n> which is a mistake Git itself used to make/allow.\n\nI like the idea to make it more explicit that this is an exhaustive\nenumeration. I'll try changing it to this instead: \"These are all of the file\nmodes in Git (which are only spiritually related to Unix file modes):\"\n\n>> +[[tag-object]]\n>> +tag object::\n>> +    Tag objects contain these required fields\n>> +    (though there are other optional fields):\n>> ++\n>> +1. The object *ID* it references\n>> +2. The object *type*\n>\n> I would rephrase these to\n>\n>     1. The *ID* of the object it references\n>     2. The *type* of the object it references\n>\n> because (1) a tag object references another object, not ID.  To name\n> the object it reference, it uses the object name of it, but just\n> like your name is not you, object name is not the object (it merely\n> is *one* way to refer to it). (2) unless it is very clear to readers\n> that \"The object\" in 1. and 2. refer to the same object, 2. invites\n> a question \"type of which object?\".\n\nThat makes sense to me, will change it to that.\n\n>> +[[branch]]\n>> +branches: `refs/heads/<name>`::\n>> +    A branch refers to a commit ID.\n>\n> A branch refers to a commit object (by its ID).  Ditto for tags.\n\nWhat's the goal of this? I can't tell what misconception you're\ntrying to avoid here.\n\n>> +NOTE: Git may delete objects that aren't \"reachable\" from any reference.\n>> +An object is \"reachable\" if we can find it by following tags to whatever\n>> +they tag, commits to their parents or trees, and trees to the trees or\n>> +blobs that they contain.\n>> +For example, if you amend a commit, with `git commit --amend`,\n>> +the old commit will usually not be reachable, so it may be deleted eventually.\n>> +Reachable objects will never be deleted.\n>\n> Very good write-up.  As we would touch upon reflog later in the same\n> document, we may want to extend the \"amend\" example a bit, perhaps\n> like\n>\n>     Note: Git never deletes objects that are \"reachable\".  An object\n>     is \"reachable\" if ....  An unreachable object may be deleted.\n>\n>     For example, ... a newly created commit will replace the old\n>     commit and the current branch ref points at the new commit.  The\n>     old commit is recorded in the <<reflogs,reflog>> of the current\n>     branch, so it is still \"reachable\", but sufficiently old reflog\n>     entries are expired away, the old commit may become unreachable\n>     at that point, and would get deleted.\n\nI like that, will include something similar, lightly reworded.\n\n> Other than the above, I found everything very nicely written.\n>\n> Thanks.\n"},{"id":"530150","messageId":"9ec9a2ec-d141-4c0b-8fd7-bfcd6a3f8249@app.fastmail.com","threadId":"64244","inReplyTo":"aQhcZwv0PdwNc6RW@pks.im","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-03T19:52:37Z","receivedAt":"2025-11-03T19:53:53Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Mon, Nov 3, 2025, at 2:40 AM, Patrick Steinhardt wrote:\n> On Thu, Oct 30, 2025 at 08:32:16PM +0000, Julia Evans via GitGitGadget wrote:\n>> diff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\n>> new file mode 100644\n>> index 0000000000..1cefbb4833\n>> --- /dev/null\n>> +++ b/Documentation/gitdatamodel.adoc\n> [snip]\n>> +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n>> +  regular commits have 1 parent, merge commits have 2 or more parents\n>> +3. An *author* and the time the commit was authored\n>> +4. A *committer* and the time the commit was committed.\n>> +5. A *commit message*\n>\n> Nit: The punctuation is a bit inconsistent here, as some list items have\n> a trailing dot while others don't.\n\nThanks, will fix.\n\n>> +[[references]]\n>> +REFERENCES\n>> +----------\n>> +\n>> +References are a way to give a name to a commit.\n>> +It's easier to remember \"the changes I'm working on are on the `turtle`\n>> +branch\" than \"the changes are in commit bb69721404348e\".\n>> +Git often uses \"ref\" as shorthand for \"reference\".\n>> +\n>> +References can either refer to:\n>> +\n>> +1. An object ID, usually a <<commit,commit>> ID\n>> +2. Another reference. This is called a \"symbolic reference\".\n>\n> Same here.\n>\n> Other than these two nits and Junio's comments I think this is in a good\n> enough shape. Thanks for working on this!\n>\n> Patrick\n"},{"id":"530156","messageId":"xmqqpl9yshrr.fsf@gitster.g","threadId":"64244","inReplyTo":"8b70796e-b5a4-4f70-8b27-c0ed80d1fc4d@app.fastmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-04T01:34:00Z","receivedAt":"2025-11-04T01:34:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n>>> +tree::\n>>> +    A tree is how Git represents a directory.\n>>> +    It can contain files or other trees (which are subdirectories).\n>>> +    It lists, for each item in the tree:\n>>> ++\n>>> +1. The *filename*, for example `hello.py`\n>>> +2. The *file mode*. Git has these file modes. which are only\n>>\n>> \"has these\" -> \"uses only these\" to clarify that this is an\n>> exhaustive enumeration and users cannot invent 100664 and others,\n>> which is a mistake Git itself used to make/allow.\n>\n> I like the idea to make it more explicit that this is an exhaustive\n> enumeration. I'll try changing it to this instead: \"These are all of the file\n> modes in Git (which are only spiritually related to Unix file modes):\"\n\nThe primary reason why I suggested \"uses only these\" was because I\nthought it would strongly hint that random additions beyond the set\nis unwelcome.  As long as that implication is not lost, I do not\nhave strong preference between \"we only use these and nothing else\"\nand your \"these are all that we use\".\n\n>>> +[[tag-object]]\n>>> +tag object::\n>>> +    Tag objects contain these required fields\n>>> +    (though there are other optional fields):\n>>> ++\n>>> +1. The object *ID* it references\n>>> +2. The object *type*\n>>\n>> I would rephrase these to\n>>\n>>     1. The *ID* of the object it references\n>>     2. The *type* of the object it references\n>>\n>> because (1) a tag object references another object, not ID.  To name\n>> the object it reference, it uses the object name of it, but just\n>> like your name is not you, object name is not the object (it merely\n>> is *one* way to refer to it). (2) unless it is very clear to readers\n>> that \"The object\" in 1. and 2. refer to the same object, 2. invites\n>> a question \"type of which object?\".\n>\n> That makes sense to me, will change it to that.\n>\n>>> +[[branch]]\n>>> +branches: `refs/heads/<name>`::\n>>> +    A branch refers to a commit ID.\n>>\n>> A branch refers to a commit object (by its ID).  Ditto for tags.\n>\n> What's the goal of this? I can't tell what misconception you're\n> trying to avoid here.\n\nThis comes from the same place as the suggestion for the tag object\nabove, i.e. \"a tag object references another object, not ID.\".\n\nExactly the same reasoning applies here.  A branch refers to a\ncommit, and to name the object it references, it uses the object\nname of it, but just like your name is not you, object name is not\nthe object itself.\n\nThanks.\n"},{"id":"530199","messageId":"9ff9d97e-2fae-488c-990b-cb574fbe8c71@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqpl9yshrr.fsf@gitster.g","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-04T15:45:25Z","receivedAt":"2025-11-04T15:45:52Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Mon, Nov 3, 2025, at 8:34 PM, Junio C Hamano wrote:\n> \"Julia Evans\" <julia@jvns.ca> writes:\n>\n>>>> +tree::\n>>>> +    A tree is how Git represents a directory.\n>>>> +    It can contain files or other trees (which are subdirectories).\n>>>> +    It lists, for each item in the tree:\n>>>> ++\n>>>> +1. The *filename*, for example `hello.py`\n>>>> +2. The *file mode*. Git has these file modes. which are only\n>>>\n>>> \"has these\" -> \"uses only these\" to clarify that this is an\n>>> exhaustive enumeration and users cannot invent 100664 and others,\n>>> which is a mistake Git itself used to make/allow.\n>>\n>> I like the idea to make it more explicit that this is an exhaustive\n>> enumeration. I'll try changing it to this instead: \"These are all of the file\n>> modes in Git (which are only spiritually related to Unix file modes):\"\n>\n> The primary reason why I suggested \"uses only these\" was because I\n> thought it would strongly hint that random additions beyond the set\n> is unwelcome.  As long as that implication is not lost, I do not\n> have strong preference between \"we only use these and nothing else\"\n> and your \"these are all that we use\".\n>\n>>>> +[[tag-object]]\n>>>> +tag object::\n>>>> +    Tag objects contain these required fields\n>>>> +    (though there are other optional fields):\n>>>> ++\n>>>> +1. The object *ID* it references\n>>>> +2. The object *type*\n>>>\n>>> I would rephrase these to\n>>>\n>>>     1. The *ID* of the object it references\n>>>     2. The *type* of the object it references\n>>>\n>>> because (1) a tag object references another object, not ID.  To name\n>>> the object it reference, it uses the object name of it, but just\n>>> like your name is not you, object name is not the object (it merely\n>>> is *one* way to refer to it). (2) unless it is very clear to readers\n>>> that \"The object\" in 1. and 2. refer to the same object, 2. invites\n>>> a question \"type of which object?\".\n>>\n>> That makes sense to me, will change it to that.\n>>\n>>>> +[[branch]]\n>>>> +branches: `refs/heads/<name>`::\n>>>> +    A branch refers to a commit ID.\n>>>\n>>> A branch refers to a commit object (by its ID).  Ditto for tags.\n>>\n>> What's the goal of this? I can't tell what misconception you're\n>> trying to avoid here.\n>\n> This comes from the same place as the suggestion for the tag object\n> above, i.e. \"a tag object references another object, not ID.\".\n>\n> Exactly the same reasoning applies here.  A branch refers to a\n> commit, and to name the object it references, it uses the object\n> name of it, but just like your name is not you, object name is not\n> the object itself.\n\nI agree the ID of a commit is not the same as the commit itself.\nThe reason I said \"refers to a commit ID\" is that it's a very concise\nexplanation and  I don't see any risk that the reader will be\nconfused by it.\n\nUnlike with my name, commit IDs uniquely identify commits, so\nI think it will be clear to the reader that the commit ID is going to\nbe used to retrieve the commit object.\n\nThe problem with \"A branch refers to a commit object (by its ID).\" is\nthat it introduces some more potential for confusion: it makes it\nsound like there might be other ways to refer to a commit object\nthan by its ID.\n\nMaybe there's another option? To me this introduces the potential\nfor more confusion and does not solve any specific problem.\n"},{"id":"530217","messageId":"xmqq346tpliw.fsf@gitster.g","threadId":"64244","inReplyTo":"9ff9d97e-2fae-488c-990b-cb574fbe8c71@app.fastmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-04T20:53:27Z","receivedAt":"2025-11-04T20:53:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> The problem with \"A branch refers to a commit object (by its ID).\" is\n\nAh, I didn't mean to say \"you must use exactly that phrase\".\n\nBut branch refers to a commit object, it does not refer to the name\nof a commit object.\n\nPerhaps \"a branch ref records the object name of a commit object\",\nwould be better?  The untold implication of the phrasing is that\nanybody who reads what is recorded by that ref can then use the\nresult to refer to (find) the commit object.\n\n> it introduces some more potential for confusion: it makes it\n> sound like there might be other ways to refer to a commit object\n> than by its ID.\n\nYes, there are unbound number of ways to refer to a commit object.\n\n $ git show-ref refs/heads/maint\n bb5c624209fcaebd60b9572b2cc8c61086e39b57 refs/heads/maint\n\nThe branch ref let you refer to a commit object by recording its\ncommit object name bb5c6242, but for humans, it is much easier to\nrefer to the same commit as \"v2.51.2^{commit}\", which is far more\nmemorable.  Of course I can use master~32^2 to call the same commit\nobject, which is less memorable gives us a hint that the tip of\nmaster fully contains that maintenance release.  What's more useful\ndepends on how the name will be used, and the hexadecimal object\nnames happen to be how refs record the objects they refer to.\n\n\n\n\n"},{"id":"530219","messageId":"5ac4f09e-927c-4125-adea-f7d5ed3d1caf@app.fastmail.com","threadId":"64244","inReplyTo":"xmqq346tpliw.fsf@gitster.g","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-04T21:24:48Z","receivedAt":"2025-11-04T21:25:09Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Tue, Nov 4, 2025, at 3:53 PM, Junio C Hamano wrote:\n> \"Julia Evans\" <julia@jvns.ca> writes:\n>\n>> The problem with \"A branch refers to a commit object (by its ID).\" is\n>\n> Ah, I didn't mean to say \"you must use exactly that phrase\".\n>\n> But branch refers to a commit object, it does not refer to the name\n> of a commit object.\n>\n> Perhaps \"a branch ref records the object name of a commit object\",\n> would be better?  The untold implication of the phrasing is that\n> anybody who reads what is recorded by that ref can then use the\n> result to refer to (find) the commit object.\n>\n>> it introduces some more potential for confusion: it makes it\n>> sound like there might be other ways to refer to a commit object\n>> than by its ID.\n>\n> Yes, there are unbound number of ways to refer to a commit object.\n>\n>  $ git show-ref refs/heads/maint\n>  bb5c624209fcaebd60b9572b2cc8c61086e39b57 refs/heads/maint\n>\n> The branch ref let you refer to a commit object by recording its\n> commit object name bb5c6242, but for humans, it is much easier to\n> refer to the same commit as \"v2.51.2^{commit}\", which is far more\n> memorable.  Of course I can use master~32^2 to call the same commit\n> object, which is less memorable gives us a hint that the tip of\n> master fully contains that maintenance release.  What's more useful\n> depends on how the name will be used, and the hexadecimal object\n> names happen to be how refs record the objects they refer to.\n\nI'm aware that there are other ways to refer to a commit other than its ID, but\nas far as I know literally every other way to refer to a commit eventually ends\nup going through the commit ID to retrieve the commit.\n\nFor example you could use `master^32`. but presumably what that does is\nto find `master`, look up the commit ID for `master`, and then go through 32\nparents until it finds the appropriate commit ID and then looks up the object\ncorresponding to that ID\n\nI do not see the point of implying that the commit ID is not \"special\", or that\nit's only one of many ways to find a commit because to me it seems very special,\nsince there is no way I know of to retrieve a commit that doesn't ultimately\nend up using the commit ID at some point. (though that ID might not be encoded\nin hexadecimal)\n"},{"id":"530226","messageId":"xmqq8qglnyzj.fsf@gitster.g","threadId":"64244","inReplyTo":"5ac4f09e-927c-4125-adea-f7d5ed3d1caf@app.fastmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-04T23:45:36Z","receivedAt":"2025-11-04T23:45:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> I do not see the point of implying that the commit ID is not \"special\", or that\n> it's only one of many ways to find a commit because to me it seems very special,\n> since there is no way I know of to retrieve a commit that doesn't ultimately\n> end up using the commit ID at some point. (though that ID might not be encoded\n> in hexadecimal)\n\nThat is not what I am trying to say.  The hexadecimal name is the\nmost neutral way to refer to a commit object, and in that sense it\nis special.  It is the way ref subsystem uses to record the name of\nobjects, and that makes it special enough.\n\nBut that does not mean that the name _is_ the object.  The\nhexadecimal name is a way you use to name the object, but is not the\nobject itself, and the special-ness of that name does not change it.\n"},{"id":"530229","messageId":"c268c98d-0a8d-48fe-99dd-b4a2fdcd0fb9@app.fastmail.com","threadId":"64244","inReplyTo":"xmqq8qglnyzj.fsf@gitster.g","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-05T00:02:38Z","receivedAt":"2025-11-05T00:03:00Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Tue, Nov 4, 2025, at 6:45 PM, Junio C Hamano wrote:\n> \"Julia Evans\" <julia@jvns.ca> writes:\n>\n>> I do not see the point of implying that the commit ID is not \"special\", or that\n>> it's only one of many ways to find a commit because to me it seems very special,\n>> since there is no way I know of to retrieve a commit that doesn't ultimately\n>> end up using the commit ID at some point. (though that ID might not be encoded\n>> in hexadecimal)\n>\n> That is not what I am trying to say.  The hexadecimal name is the\n> most neutral way to refer to a commit object, and in that sense it\n> is special.  It is the way ref subsystem uses to record the name of\n> objects, and that makes it special enough.\n>\n> But that does not mean that the name _is_ the object.  The\n> hexadecimal name is a way you use to name the object, but is not the\n> object itself, and the special-ness of that name does not change it.\n\nOkay. I still do not understand at all why this is so important to you\n(for the reasons I mentioned before) but I'll see if there's anything I can do.\n"},{"id":"530232","messageId":"7E8706EF-6C15-43AD-A847-5C896D9235AA@gmail.com","threadId":"64244","inReplyTo":"c268c98d-0a8d-48fe-99dd-b4a2fdcd0fb9@app.fastmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-11-05T03:21:23Z","receivedAt":"2025-11-05T03:21:35Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 4 nov. 2025 à 19:02, Julia Evans <julia@jvns.ca> a écrit :\n> \n> ﻿\n> \n>> On Tue, Nov 4, 2025, at 6:45 PM, Junio C Hamano wrote:\n>> \"Julia Evans\" <julia@jvns.ca> writes:\n>>> I do not see the point of implying that the commit ID is not \"special\", or that\n>>> it's only one of many ways to find a commit because to me it seems very special,\n>>> since there is no way I know of to retrieve a commit that doesn't ultimately\n>>> end up using the commit ID at some point. (though that ID might not be encoded\n>>> in hexadecimal)\n>> That is not what I am trying to say.  The hexadecimal name is the\n>> most neutral way to refer to a commit object, and in that sense it\n>> is special.  It is the way ref subsystem uses to record the name of\n>> objects, and that makes it special enough.\n>> But that does not mean that the name _is_ the object.  The\n>> hexadecimal name is a way you use to name the object, but is not the\n>> object itself, and the special-ness of that name does not change it.\n> \n> Okay. I still do not understand at all why this is so important to you\n> (for the reasons I mentioned before) but I'll see if there's anything I can do.\n\nPerhaps one way to look at is, what diagram would I draw given different textual explanations?\n\nThe diagram we _want_ folks to draw (?) is the one where a branch points at a commit [a circle, perhaps], which points to a tree [triangle] and recursively blobs [squares], like I’ve seen Stolee draw for GitHub blogs.\n\nWe might also want folks to label the arrows with names, or not.\n\nOne way to interpret the “branch refers to a commit ID” might be to draw a diagram where the branch points to an ID label, and to find the circle you have to separately consult a different part of the diagram.\n\nBoth seem useful to me, though as the former has fewer moving pieces might be better for the model this document describes? I dunno. "},{"id":"530253","messageId":"7217ae44-5ad5-468c-b76b-c485247fb2f4@app.fastmail.com","threadId":"64244","inReplyTo":"7E8706EF-6C15-43AD-A847-5C896D9235AA@gmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-05T16:26:42Z","receivedAt":"2025-11-05T16:27:03Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Tue, Nov 4, 2025, at 10:21 PM, Ben Knoble wrote:\n>> Le 4 nov. 2025 à 19:02, Julia Evans <julia@jvns.ca> a écrit :\n>> \n>> ﻿\n>> \n>>> On Tue, Nov 4, 2025, at 6:45 PM, Junio C Hamano wrote:\n>>> \"Julia Evans\" <julia@jvns.ca> writes:\n>>>> I do not see the point of implying that the commit ID is not \"special\", or that\n>>>> it's only one of many ways to find a commit because to me it seems very special,\n>>>> since there is no way I know of to retrieve a commit that doesn't ultimately\n>>>> end up using the commit ID at some point. (though that ID might not be encoded\n>>>> in hexadecimal)\n>>> That is not what I am trying to say.  The hexadecimal name is the\n>>> most neutral way to refer to a commit object, and in that sense it\n>>> is special.  It is the way ref subsystem uses to record the name of\n>>> objects, and that makes it special enough.\n>>> But that does not mean that the name _is_ the object.  The\n>>> hexadecimal name is a way you use to name the object, but is not the\n>>> object itself, and the special-ness of that name does not change it.\n>> \n>> Okay. I still do not understand at all why this is so important to you\n>> (for the reasons I mentioned before) but I'll see if there's anything I can do.\n>\n> Perhaps one way to look at is, what diagram would I draw given \n> different textual explanations?\n>\n> The diagram we _want_ folks to draw (?) is the one where a branch \n> points at a commit [a circle, perhaps], which points to a tree \n> [triangle] and recursively blobs [squares], like I’ve seen Stolee draw \n> for GitHub blogs.\n>\n> We might also want folks to label the arrows with names, or not.\n>\n> One way to interpret the “branch refers to a commit ID” might be to \n> draw a diagram where the branch points to an ID label, and to find the \n> circle you have to separately consult a different part of the diagram.\n\nYes, the most common type of Git diagram I see is something like this:\nhttps://git-scm.com/book/en/v2/images/head-to-master.png\nwhich only includes references, commits, and HEAD. \n\nThat's the diagram I have in mind when writing this text, and I think it's\na useful and accurate diagram to keep in mind, and it's one that you see\nvery often when using Git tools, including in `git log --graph`. (it's not\na _complete_ diagram of every type of object, but diagrams do not need to be\ncomplete to be accurate)\n\nI personally would not use a graph diagram to explain how commits relate to\ntrees and blobs (normally I use `git cat-file -p` instead, like I did in this\n`gitdatamodel` document. You can see this comic for a \"visual\" example of how\nI've approached discussing trees and blobs in the past with `git cat-file -p`\nhttps://wizardzines.com/comics/explore-a-commit/).\n\n> Both seem useful to me, though as the former has fewer moving pieces \n> might be better for the model this document describes? I dunno.\n"},{"id":"530286","messageId":"875BAB6E-724A-4AB3-85D7-8750667949E5@gmail.com","threadId":"64244","inReplyTo":"7217ae44-5ad5-468c-b76b-c485247fb2f4@app.fastmail.com","subject":"Re: [PATCH v5] doc: add an explanation of Git's data model","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-11-06T03:07:57Z","receivedAt":"2025-11-06T03:08:10Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 5 nov. 2025 à 11:27, Julia Evans <julia@jvns.ca> a écrit :\n> \n> ﻿\n> \n> On Tue, Nov 4, 2025, at 10:21 PM, Ben Knoble wrote:\n>>>> Le 4 nov. 2025 à 19:02, Julia Evans <julia@jvns.ca> a écrit :\n>>> \n>>> ﻿\n>>> \n>>>> On Tue, Nov 4, 2025, at 6:45 PM, Junio C Hamano wrote:\n>>>> \"Julia Evans\" <julia@jvns.ca> writes:\n>>>>> I do not see the point of implying that the commit ID is not \"special\", or that\n>>>>> it's only one of many ways to find a commit because to me it seems very special,\n>>>>> since there is no way I know of to retrieve a commit that doesn't ultimately\n>>>>> end up using the commit ID at some point. (though that ID might not be encoded\n>>>>> in hexadecimal)\n>>>> That is not what I am trying to say.  The hexadecimal name is the\n>>>> most neutral way to refer to a commit object, and in that sense it\n>>>> is special.  It is the way ref subsystem uses to record the name of\n>>>> objects, and that makes it special enough.\n>>>> But that does not mean that the name _is_ the object.  The\n>>>> hexadecimal name is a way you use to name the object, but is not the\n>>>> object itself, and the special-ness of that name does not change it.\n>>> \n>>> Okay. I still do not understand at all why this is so important to you\n>>> (for the reasons I mentioned before) but I'll see if there's anything I can do.\n>> \n>> Perhaps one way to look at is, what diagram would I draw given\n>> different textual explanations?\n>> \n>> The diagram we _want_ folks to draw (?) is the one where a branch\n>> points at a commit [a circle, perhaps], which points to a tree\n>> [triangle] and recursively blobs [squares], like I’ve seen Stolee draw\n>> for GitHub blogs.\n>> \n>> We might also want folks to label the arrows with names, or not.\n>> \n>> One way to interpret the “branch refers to a commit ID” might be to\n>> draw a diagram where the branch points to an ID label, and to find the\n>> circle you have to separately consult a different part of the diagram.\n> \n> Yes, the most common type of Git diagram I see is something like this:\n> https://git-scm.com/book/en/v2/images/head-to-master.png\n> which only includes references, commits, and HEAD.\n> \n> That's the diagram I have in mind when writing this text, and I think it's\n> a useful and accurate diagram to keep in mind, and it's one that you see\n> very often when using Git tools, including in `git log --graph`. (it's not\n> a _complete_ diagram of every type of object, but diagrams do not need to be\n> complete to be accurate)\n> \n> I personally would not use a graph diagram to explain how commits relate to\n> trees and blobs (normally I use `git cat-file -p` instead, like I did in this\n> `gitdatamodel` document. You can see this comic for a \"visual\" example of how\n> I've approached discussing trees and blobs in the past with `git cat-file -p`\n> https://wizardzines.com/comics/explore-a-commit/).\n\nFair enough. Here’s a post I “stole” ;) the shapes from, for posterity:\n\nhttps://github.blog/open-source/git/commits-are-snapshots-not-diffs \n\nMy larger point was: since these are the diagrams I’m imagining we want to convey to a reader, perhaps ID can be omitted for brevity? IOW, the relationship between objects is the thing to highlight.\n\nOTOH, when exploring the data, especially at the plumbing level it seems we have to do the “pointer-chasing” ourselves (see cat-file).\n\nSo idk. "},{"id":"530389","messageId":"pull.1981.v6.git.1762545177204.gitgitgadget@gmail.com","threadId":"64244","inReplyTo":"pull.1981.v5.git.1761856336360.gitgitgadget@gmail.com","subject":"[PATCH v6] doc: add an explanation of Git's data model","fromName":"Julia Evans via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-07T19:52:57Z","receivedAt":"2025-11-07T19:53:00Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"From: Julia Evans <julia@jvns.ca>\n\nGit very often uses the terms \"object\", \"reference\", or \"index\" in its\ndocumentation.\n\nHowever, it's hard to find a clear explanation of these terms and how\nthey relate to each other in the documentation. The closest candidates\ncurrently are:\n\n1. `gitglossary`. This makes a good effort, but it's an alphabetically\n    ordered dictionary and a dictionary is not a good way to learn\n    concepts. You have to jump around too much and it's not possible to\n    present the concepts in the order that they should be explained.\n2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n   This is a nice document to have, but it's not necessary to learn how\n   `update-index` works to understand Git's data model, and we should\n   not be requiring users to learn how to use the \"plumbing\" commands\n   if they want to learn what the term \"index\" or \"object\" means.\n3. `gitrepository-layout`. This is a great resource, but it includes a\n   lot of information about configuration and internal implementation\n   details which are not related to the data model. It also does\n   not explain how commits work.\n\nThe result of this is that Git users (even users who have been using\nGit for 15+ years) struggle to read the documentation because they don't\nknow what the core terms mean, and it's not possible to add links\nto help them learn more.\n\nAdd an explanation of Git's data model. Some choices I've made in\ndeciding what \"core data model\" means:\n\n1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n   if those are intended to be user facing or if they're more like\n   internal implementation details.\n2. Don't talk about submodules other than by mentioning how they\n   relate to trees. This is because Git has a lot of special features,\n   and explaining how they all work exhaustively could quickly go\n   down a rabbit hole which would make this document less useful for\n   understanding Git's core behaviour.\n3. Don't discuss the structure of a commit message\n   (first line, trailers etc).\n4. Don't mention configuration.\n5. Don't mention the `.git` directory, to avoid getting too much into\n   implementation details\n\nSigned-off-by: Julia Evans <julia@jvns.ca>\n---\n    doc: Add a explanation of Git's data model\n    \n    Changes in v2:\n    \n    The biggest change is to remove all mentions of the .git directory, and\n    explain references in a way that doesn't refer to \"directories\" at all,\n    and instead talks about the \"hierarchy\" (from Kristoffer and Patrick's\n    reviews).\n    \n    Also:\n    \n     * objects: Mention that an object ID is called an \"object name\", and\n       update the glossary to include the term \"object ID\" (from Junio's\n       review)\n     * objects: Replace \"SHA-1 hash\" with \"cryptographic hash\" which is more\n       accurate (from Patrick's review)\n     * blobs: Made the explanation of git gc a little higher level and took\n       some ideas from Patrick's suggested wording (from Patrick's and\n       Kroftoffer's reviews)\n     * commits: Mention that tag objects and commits can optionally have\n       other fields. I didn't mention the GPG signature specifically, but\n       don't have any objections to adding it. (from Patrick and Junio's\n       reviews)\n     * commits: Remove one of the mentions of git gc, since it perhaps opens\n       up too much of a rabbit hole: \"how does git gc decide which commits\n       to clean up?\". (from Kristoffer's review)\n     * tag objects: Add an example of how a tag object is represented (from\n       user feedback on the draft)\n     * index: Use the term \"file mode\" instead of \"permissions\", and list\n       all allowed file modes (from Patrick's review)\n     * index: Use \"stage number\" instead of \"number\" for index entries (from\n       Patrick's review)\n     * reflogs: Remove \"any ref can be logged\", it raises some questions of\n       \"how do you tell Git to log a ref that it isn't normally logging?\"\n       and my guess is that it's uncommon to ask Git to log more refs. I\n       don't think it's a \"lie\" to omit this but I can bring it back if\n       folks disagree. (from Patrick's review)\n     * reflogs: Fix an error I noticed in the explanation of reflogs: tags\n       aren't logged by default and remote-tracking branches are, according\n       to man git-config\n     * branches and tags: Be clearer about how branches are usually updated\n       (by committing), and make it a little more obvious that only branches\n       can be checked out. This is a bit tricky because using the word\n       \"check out\" introduces a rabbit hole that I want to avoid (what does\n       \"check out\" mean?). I've dealt this by just talking about the\n       \"current branch\" (HEAD) since that is defined here, and making it\n       more explicit that HEAD must either be a branch or a commit, there's\n       no \"HEAD is a tag\" option. (from Patrick's review)\n     * tags: Explain the differences between annotated and lightweight tags\n       (this is the main piece of user feedback I've gotten on the draft so\n       far)\n     * Various style/typo changes (\"2 or more\", linkgit:git-gc[1], removed\n       extra asterisks, added empty SYNOPSIS, \"commits -> tags\" typo fix,\n       add to meson build)\n    \n    non-changes:\n    \n     * I still haven't mentioned things that aren't part of the \"data\n       model\", like revision params and configuration. I think there could\n       be a place for them but I haven't found it yet.\n     * tag objects: I noticed that there's a \"tag\" header field in tag\n       objects (like tag v1.0.0) but I didn't mention it yet because I\n       couldn't figure out what the purpose of that field is (I thought the\n       tag name was stored in the reference, why is it duplicated in the tag\n       object?)\n    \n    Changes in v3:\n    \n    I asked for feedback from Git users on Mastodon and got 220 pieces of\n    feedback from 48 different users. People seemed very excited to read\n    about Git's data model. Usually I judge explanations by what folks\n    report learning from them. Here people reported learning:\n    \n     * how branches are stored (that a branch is \"a name for a commit\")\n     * how objects work\n     * that Git has separate \"author\" and \"committer\" fields\n     * that amending a commit does not change it\n     * that a tree is \"just a directory\" (not something more complicated),\n       and how trees are stored\n     * that Git repos can contain symlinks\n     * that Git saves modes separately from the OS.\n     * how the stage number works\n     * that when you git add a file, Git will create an object\n     * that third-party tools can create their own refs.\n     * that the reflog stores the history of branches (not just HEAD), and\n       what reflogs are for\n    \n    Also (of course) there were quite a few points of confusion! The main 4\n    pieces of feedback were\n    \n     1. The index section doesn't explain what the word \"staged\" means, and\n        one person says that it makes it sounds like only files that you\n        \"git add\"ed are in the index. Rewrite the explanation to avoid using\n        the word \"staged\" to define the index and instead define the word\n        \"staging\".\n     2. Explain the difference between \"annotated tags\" and \"lightweight\n        tags\" (done)\n     3. Add examples for tag objects and reflogs (done)\n     4. Mention a little more about where things are stored in the .git\n        directory, which I'd removed in v2. This seems most important for\n        .git/refs, so I added a hopefully accurate note about how refs are\n        stored by default, with a comment about one of the major\n        implications. I did not discuss where objects or the index are\n        stored, because I don't think the implementation details of how\n        objects are stored are as important, and there are better tools for\n        viewing the \"raw\" state of objects and the index (with git cat-file\n        -p or git ls-files --staged).\n    \n    Here's every other change I made in response to the feedback, as well as\n    a few comments that I did not address.\n    \n    intro:\n    \n     * Give a 1-sentence intro to \"reflog\"\n    \n    objects:\n    \n     * people really like having git ls-files --stage as a way to view the\n       index, so add git cat-file -p as well in a note\n    \n    commits:\n    \n     * 2 people asked \"Are commits stored as a diff?\". Say that diffs are\n       calculated at runtime, this is very important.\n     * The order the fields are given in don't match the order in the\n       example. Make them match.\n     * \"All the files in the commit, stored as a tree\" is throwing a few\n       people off. Be clearer that it's the tree ID of the base directory.\n     * Several people asked \"What's the difference between an author and\n       committer? I added an example using git cherry-pick that I'm not 100%\n       happy with (what if the reader doesn't know what cherry-pick does?).\n       There might be a better example to give here.\n     * In the note about commits being amended: one person suggested saying\n       \"creates a new commit with the same parent\" to make it clearer what\n       the relationship between the new and old commit are. I liked that\n       idea so I did it.\n    \n    trees:\n    \n     * file modes. 2 people want to know more about \"The file mode, for\n       example 100644\". Also 2 people are curious about what relationship\n       these have to Unix permissions. Say that they're inspired by Unix\n       permissions, and move the list of possible file modes up to make the\n       relationship clearer\n     * On \"so git-gc(1) periodically compresses objects to save disk space\",\n       there are a few follow up comments wondering about more, which makes\n       me think the comment about compression is actually a distraction. Say\n       something simpler instead, (\"Git only needs to store new versions of\n       files which were changed in that commit\"), from Junio's suggestion\n     * Re \"commit (a Git submodule)\": 2 people say it's not clear how trees\n       relate to submodules. Say that it refers to a commit in a different\n       repository.\n     * One person says they're not sure if the \"object ID\" is a hash. Link\n       it to the definition of \"object ID\".\n    \n    tag objects:\n    \n     * Requests for an example, added one.\n     * Requests to explain the difference between \"lightweight\" and\n       \"annotated\" tags, added it.\n    \n    tags:\n    \n     * one person thinks \"It’s expected that a tag will never change after\n       you create it.\" is too strong (since of course you can change it with\n       git tag -f). Say instead that tags are \"usually\" not changed.\n    \n    HEAD:\n    \n     * Several people are asking for more detail about detached HEAD state.\n       There's actually quite a lot to talk about here (what it means, how\n       it happens, what it implies, and how you might adjust your workflow\n       to avoid it by using git switch). I don't think we can get into all\n       of that here, so refer to the DETACHED HEAD section of git-checkout\n       instead. I'm not totally happy with the current version of that\n       section but that seems like the most practical solution right now.\n    \n    remote-tracking branches:\n    \n     * discuss refs/remotes/<remote>/HEAD.\n    \n    the index:\n    \n     * \"permissions\" should be \"file mode\" (like with trees). Changed.\n     * \"filename\" should be \"file path\". Changed.\n     * the stage number can only be 0, 1, 2, or 3, since it's 2 bits. Also\n       maybe say that the numbers have specific meanings. Said it can only\n       be 0/1/2/3 but did not give the specific meanings.\n    \n    reflogs\n    \n     * Request for an example. Added one.\n     * It's not clear if there's one reflog per branch/tag/HEAD, or if\n       there's one universal reflog. Make this clearer.\n     * Mention the role of the reflog in retrieving \"lost\" commits or\n       undoing bad rebases.\n    \n    Not fixed:\n    \n     * intro: A couple of people say that it's confusing that tags are both\n       \"an object\" and \"a reference\". Handled this by just explaining the\n       difference between an annotated and a lightweight tag further down.\n       I'd like to make this clearer in the intro but not sure if there's a\n       way to do it.\n     * commits and tag objects: one person asks if there's a reference for\n       the other \"optional fields\", like \"encoding\" and \"gpgsig\". I couldn't\n       find one, so left this as is.\n     * HEAD: A couple of people ask if there are any other symbolic\n       references other than HEAD, or if they can make their own symbolic\n       references. I don't know the answer to this.\n     * HEAD: the HEAD: HEAD thing looks weird, it made more sense when it\n       was HEAD: .git/HEAD. Will think about this.\n     * reflogs: One person asks: if reflogs only store local changes, why\n       does it track the user who made the change? Is that for remote\n       operations like fetches and pulls? Or for cases where more than one\n       user is using the same repo on a system? I don't know the answer to\n       this.\n     * reflogs: How can you see the full data in the reflog? git reflog show\n       doesn't list the user who made the change. git reflog show <refname>\n       --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso seems to work but\n       it's really a mouthful, not sure it's useful to include all that.\n     * index: Is it worth mentioning that the index can be locked? I don't\n       have an opinion about this.\n     * other: One person asks what a \"working tree\" is. It made me wonder if\n       \"the current working directory\" has a place in Git's data model. My\n       feeling is \"no\" but I could be convinced otherwise.\n     * overall: \"How can Git be so fast? If I switch branches, how does it\n       figure out what to add, remove or replace?\". I don't think this is\n       the right place for that discussion but it would\n     * there are some docs CI errors I haven't figured out yet (IDREF\n       attribute linkend references an unknown ID \"tree\")\n    \n    changes in v4:\n    \n    This is a combination of trying to make some of the intro text a little\n    more \"friendly\" for someone new to Git's data model, avoiding implying\n    things that are false, and removing information that isn't relevant to\n    the data model.\n    \n    intro:\n    \n     * Add a 1-line description of what a \"reflog\" is (from user feedback)\n    \n    objects:\n    \n     * Start with a \"friendly\" description of what an object is, similar to\n       what we do for references and the reflog\n     * Rename \"commits\" to \"commit\" and similarly for trees etc (from\n       Junio's review)\n     * Remove the explanation of what git cat-file -p does, since it might\n       be misleading and if people want to know they can read the man page\n       (from Junio's review)\n    \n    commits:\n    \n     * Start by saying that the commit contains the full directory structure\n       of all the files (from Junio's comment about how it may not be clear\n       that the commit contains all the files' exact contents at the time of\n       the commit)\n     * Remove the comment about cherry-pick (from Junio's review)\n     * Replace \"ask Git for a diff\" with \"ask Git to show the commit with\n       git show\" (from Junio's review)\n    \n    trees:\n    \n     * Make the description a little more friendly\n     * Reorder so that \"type\" is defined before we refer to the \"type\"\n     * Say that file modes are \"only spiritually related\" to Unix\n       permissions instead of talking about what Git \"supports\" (from\n       Junio's review)\n    \n    blobs:\n    \n     * Try to make it clearer how \"commits use relatively little disk space\"\n       is true while not implying that commits are diffs, by using an\n       example (from Junio's review)\n    \n    branches:\n    \n     * Replace \"a branch is a name for a commit ID\" with \"a branch refers to\n       a commit ID\" (except in the intro sentence for the \"references\"\n       section). Similarly for tags etc. (from Junio's review)\n     * Remove the note about how branches are stored in .git (from Junio's\n       review)\n    \n    HEAD:\n    \n     * Be clearer that HEAD is not always the current branch, because there\n       may not be a current branch (from Junio's review)\n    \n    index:\n    \n     * Be a little more specific about how exactly the index is converted\n       into a commit. (from Junio's comment about how it's not clear what\n       \"every file in the repository\" means)\n    \n    reflog:\n    \n     * Be clearer that there are many reflogs (one for each reference with a\n       log), not just one reflog (from Junio and Patrick's reviews)\n     * Omit the user and \"Before\" commit IDs from the list of fields,\n       because you usually don't see them (from Junio's review)\n     * Show the output of git reflog main in the example instead of the\n       contents of the reflog file, to avoid showing the user and before\n       commit ID\n    \n    changes in v5:\n    \n    Mostly smaller tweaks this time. The only major addition is to add a\n    note about how unreachable objects may be deleted.\n    \n    From Junio's review:\n    \n     * Remove \"type\" in the description of what's in a tree (since I have\n       learned that is not a separate field, it's part of the file mode)\n     * Fix a typo (\"these these\")\n     * Remove the intro sentence about what a \"commit\" is and instead only\n       describe its contents in the list of fields, to avoid implying that a\n       commit is the same as a tree\n     * Say \"Unix file modes\" instead of \"Unix permissions\"\n     * In the tag objects contents: make \"ID\" and \"type\" separate list items\n       since they're separate fields\n     * in the index section:\n       * list all of the possible file modes (since from my understanding\n         there are fewer allowed file modes here than in a tree)\n       * mention that the object can be either a commit or blob\n       * make the order match the order in git ls-files\n    \n    changes in v6:\n    \n     * Make punctuation more consistent (from Patrick's review)\n     * Explain more about when exactly amended commits will get deleted\n       (when their reflog entry expires), from Junio's review\n     * Be more explicit that there are only 5 file modes in Git (from\n       Junio's review)\n     * Make tag object description clearer (from Junio's review)\n     * We had a long discussion about the phrasing of \"A branch refers to a\n       commit ID\" but I didn't come up with any ideas for how to improve the\n       phrasing so I left it as is.\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v6\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v6\nPull-Request: https://github.com/gitgitgadget/git/pull/1981\n\nRange-diff vs v5:\n\n 1:  d342255dad ! 1:  6e2a7bbe6b doc: add an explanation of Git's data model\n     @@ Documentation/gitdatamodel.adoc (new)\n      ++\n      +1. The full directory structure of all the files in that version of the\n      +   repository and each file's contents, stored as the *<<tree,tree>>* ID\n     -+   of the commit's base directory.\n     ++   of the commit's base directory\n      +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n      +  regular commits have 1 parent, merge commits have 2 or more parents\n      +3. An *author* and the time the commit was authored\n     -+4. A *committer* and the time the commit was committed.\n     ++4. A *committer* and the time the commit was committed\n      +5. A *commit message*\n      ++\n      +Here's how an example commit is stored:\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    It lists, for each item in the tree:\n      ++\n      +1. The *filename*, for example `hello.py`\n     -+2. The *file mode*. Git has these file modes. which are only\n     -+   spiritually related to Unix file modes:\n     ++2. The *file mode*. These are all of the file modes in Git.\n     ++   They're only spiritually related to Unix file modes.\n      ++\n      +  - `100644`: regular file (with <<object,object type>> `blob`)\n      +  - `100755`: executable file (with type `blob`)\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    Tag objects contain these required fields\n      +    (though there are other optional fields):\n      ++\n     -+1. The object *ID* it references\n     -+2. The object *type*\n     ++1. The *ID* of the object it references\n     ++2. The *type* of the object it references\n      +3. The *tagger* and tag date\n      +4. A *tag message*, similar to a commit message\n      +\n     @@ Documentation/gitdatamodel.adoc (new)\n      +References can either refer to:\n      +\n      +1. An object ID, usually a <<commit,commit>> ID\n     -+2. Another reference. This is called a \"symbolic reference\".\n     ++2. Another reference. This is called a \"symbolic reference\"\n      +\n      +References are stored in a hierarchy, and Git handles references\n      +differently based on where they are in the hierarchy.\n     @@ Documentation/gitdatamodel.adoc (new)\n      +Git may also create references other than `HEAD` at the base of the\n      +hierarchy, like `ORIG_HEAD`.\n      +\n     -+NOTE: Git may delete objects that aren't \"reachable\" from any reference.\n     ++NOTE: Git may delete objects that aren't \"reachable\" from any reference\n     ++or <<reflogs,reflog>>.\n      +An object is \"reachable\" if we can find it by following tags to whatever\n      +they tag, commits to their parents or trees, and trees to the trees or\n      +blobs that they contain.\n     -+For example, if you amend a commit, with `git commit --amend`,\n     ++For example, if you amend a commit with `git commit --amend`,\n     ++there will no longer be a branch that points at the old commit.\n     ++The old commit is recorded in the current branch's <<reflogs,reflog>>,\n     ++so it is still \"reachable\", but when the reflog entry expires it may\n     ++become unreachable and get deleted.\n     ++\n      +the old commit will usually not be reachable, so it may be deleted eventually.\n      +Reachable objects will never be deleted.\n      +\n\n\n Documentation/Makefile              |   1 +\n Documentation/gitdatamodel.adoc     | 302 ++++++++++++++++++++++++++++\n Documentation/glossary-content.adoc |   4 +-\n Documentation/meson.build           |   1 +\n 4 files changed, 306 insertions(+), 2 deletions(-)\n create mode 100644 Documentation/gitdatamodel.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 6fb83d0c6e..5f4acfacbd 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n MAN7_TXT += gitcore-tutorial.adoc\n MAN7_TXT += gitcredentials.adoc\n MAN7_TXT += gitcvs-migration.adoc\n+MAN7_TXT += gitdatamodel.adoc\n MAN7_TXT += gitdiffcore.adoc\n MAN7_TXT += giteveryday.adoc\n MAN7_TXT += gitfaq.adoc\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nnew file mode 100644\nindex 0000000000..b54ff0e52b\n--- /dev/null\n+++ b/Documentation/gitdatamodel.adoc\n@@ -0,0 +1,302 @@\n+gitdatamodel(7)\n+===============\n+\n+NAME\n+----\n+gitdatamodel - Git's core data model\n+\n+SYNOPSIS\n+--------\n+gitdatamodel\n+\n+DESCRIPTION\n+-----------\n+\n+It's not necessary to understand Git's data model to use Git, but it's\n+very helpful when reading Git's documentation so that you know what it\n+means when the documentation says \"object\", \"reference\" or \"index\".\n+\n+Git's core operations use 4 kinds of data:\n+\n+1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+2. <<references,References>>: branches, tags,\n+   remote-tracking branches, etc\n+3. <<index,The index>>, also known as the staging area\n+4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n+\n+[[objects]]\n+OBJECTS\n+-------\n+\n+All of the commits and files in a Git repository are stored as \"Git objects\".\n+Git objects never change after they're created, and every object has an ID,\n+like `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+\n+This means that if you have an object's ID, you can always recover its\n+exact contents as long as the object hasn't been deleted.\n+\n+Every object has:\n+\n+[[object-id]]\n+1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n+  type and contents.\n+  It's fast to look up a Git object using its ID.\n+  This is usually represented in hexadecimal, like\n+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+2. a *type*. There are 4 types of objects:\n+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n+   and <<tag-object,tag objects>>.\n+3. *contents*. The structure of the contents depends on the type.\n+\n+Here's how each type of object is structured:\n+\n+[[commit]]\n+commit::\n+    A commit contains these required fields\n+    (though there are other optional fields):\n++\n+1. The full directory structure of all the files in that version of the\n+   repository and each file's contents, stored as the *<<tree,tree>>* ID\n+   of the commit's base directory\n+2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n+  regular commits have 1 parent, merge commits have 2 or more parents\n+3. An *author* and the time the commit was authored\n+4. A *committer* and the time the commit was committed\n+5. A *commit message*\n++\n+Here's how an example commit is stored:\n++\n+----\n+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n+author Maya <maya@example.com> 1759173425 -0400\n+committer Maya <maya@example.com> 1759173425 -0400\n+\n+Add README\n+----\n++\n+Like all other objects, commits can never be changed after they're created.\n+For example, \"amending\" a commit with `git commit --amend` creates a new\n+commit with the same parent.\n++\n+Git does not store the diff for a commit: when you ask Git to show\n+the commit with linkgit:git-show[1], it calculates the diff from its\n+parent on the fly.\n+\n+[[tree]]\n+tree::\n+    A tree is how Git represents a directory.\n+    It can contain files or other trees (which are subdirectories).\n+    It lists, for each item in the tree:\n++\n+1. The *filename*, for example `hello.py`\n+2. The *file mode*. These are all of the file modes in Git.\n+   They're only spiritually related to Unix file modes.\n++\n+  - `100644`: regular file (with <<object,object type>> `blob`)\n+  - `100755`: executable file (with type `blob`)\n+  - `120000`: symbolic link (with type `blob`)\n+  - `040000`: directory (with type `tree`)\n+  - `160000`: gitlink, for use with submodules (with type `commit`)\n+\n+3. The <<object-id,*object ID*>> with the contents of the file or directory\n++\n+For example, this is how a tree containing one directory (`src`) and one file\n+(`README.md`) is stored:\n++\n+----\n+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n+----\n+\n+[[blob]]\n+blob::\n+    A blob object contains a file's contents.\n++\n+When you make a commit, Git stores the full contents of each file that\n+you changed as a blob.\n+For example, if you have a commit that changes 2 files in a repository\n+with 1000 files, that commit will create 2 new blobs, and use the\n+previous blob ID for the other 998 files.\n+This means that commits can use relatively little disk space even in a\n+very large repository.\n+\n+[[tag-object]]\n+tag object::\n+    Tag objects contain these required fields\n+    (though there are other optional fields):\n++\n+1. The *ID* of the object it references\n+2. The *type* of the object it references\n+3. The *tagger* and tag date\n+4. A *tag message*, similar to a commit message\n+\n+Here's how an example tag object is stored:\n+\n+----\n+object 750b4ead9c87ceb3ddb7a390e6c7074521797fb3\n+type commit\n+tag v1.0.0\n+tagger Maya <maya@example.com> 1759927359 -0400\n+\n+Release version 1.0.0\n+----\n+\n+NOTE: All of the examples in this section were generated with\n+`git cat-file -p <object-id>`.\n+\n+[[references]]\n+REFERENCES\n+----------\n+\n+References are a way to give a name to a commit.\n+It's easier to remember \"the changes I'm working on are on the `turtle`\n+branch\" than \"the changes are in commit bb69721404348e\".\n+Git often uses \"ref\" as shorthand for \"reference\".\n+\n+References can either refer to:\n+\n+1. An object ID, usually a <<commit,commit>> ID\n+2. Another reference. This is called a \"symbolic reference\"\n+\n+References are stored in a hierarchy, and Git handles references\n+differently based on where they are in the hierarchy.\n+Most references are under `refs/`. Here are the main types:\n+\n+[[branch]]\n+branches: `refs/heads/<name>`::\n+    A branch refers to a commit ID.\n+    That commit is the latest commit on the branch.\n++\n+To get the history of commits on a branch, Git will start at the commit\n+ID the branch references, and then look at the commit's parent(s),\n+the parent's parent, etc.\n+\n+[[tag]]\n+tags: `refs/tags/<name>`::\n+    A tag refers to a commit ID, tag object ID, or other object ID.\n+    There are two types of tags:\n+    1. \"Annotated tags\", which reference a <<tag-object,tag object>> ID\n+       which contains a tag message\n+    2. \"Lightweight tags\", which reference a commit, blob, or tree ID\n+       directly\n++\n+Even though branches and tags both refer to a commit ID, Git\n+treats them very differently.\n+Branches are expected to change over time: when you make a commit, Git\n+will update your <<HEAD,current branch>> to point to the new commit.\n+Tags are usually not changed after they're created.\n+\n+[[HEAD]]\n+HEAD: `HEAD`::\n+    `HEAD` is where Git stores your current <<branch,branch>>,\n+    if there is a current branch. `HEAD` can either be:\n++\n+1. A symbolic reference to your current branch, for example `ref:\n+   refs/heads/main` if your current branch is `main`.\n+2. A direct reference to a commit ID. In this case there is no current branch.\n+   This is called \"detached HEAD state\", see the DETACHED HEAD section\n+   of linkgit:git-checkout[1] for more.\n+\n+[[remote-tracking-branch]]\n+remote-tracking branches: `refs/remotes/<remote>/<branch>`::\n+    A remote-tracking branch refers to a commit ID.\n+    It's how Git stores the last-known state of a branch in a remote\n+    repository. `git fetch` updates remote-tracking branches. When\n+    `git status` says \"you're up to date with origin/main\", it's looking at\n+    this.\n++\n+`refs/remotes/<remote>/HEAD` is a symbolic reference to the remote's\n+default branch. This is the branch that `git clone` checks out by default.\n+\n+[[other-refs]]\n+Other references::\n+    Git tools may create references anywhere under `refs/`.\n+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n+    and linkgit:git-notes[1] all create their own references\n+    in `refs/stash`, `refs/bisect`, etc.\n+    Third-party Git tools may also create their own references.\n++\n+Git may also create references other than `HEAD` at the base of the\n+hierarchy, like `ORIG_HEAD`.\n+\n+NOTE: Git may delete objects that aren't \"reachable\" from any reference\n+or <<reflogs,reflog>>.\n+An object is \"reachable\" if we can find it by following tags to whatever\n+they tag, commits to their parents or trees, and trees to the trees or\n+blobs that they contain.\n+For example, if you amend a commit with `git commit --amend`,\n+there will no longer be a branch that points at the old commit.\n+The old commit is recorded in the current branch's <<reflogs,reflog>>,\n+so it is still \"reachable\", but when the reflog entry expires it may\n+become unreachable and get deleted.\n+\n+the old commit will usually not be reachable, so it may be deleted eventually.\n+Reachable objects will never be deleted.\n+\n+[[index]]\n+THE INDEX\n+---------\n+The index, also known as the \"staging area\", is a list of files and\n+the contents of each file, stored as a <<blob,blob>>.\n+You can add files to the index or update the contents of a file in the\n+index with linkgit:git-add[1]. This is called \"staging\" the file for commit.\n+\n+Unlike a <<tree,tree>>, the index is a flat list of files.\n+When you commit, Git converts the list of files in the index to a\n+directory <<tree,tree>> and uses that tree in the new <<commit,commit>>.\n+\n+Each index entry has 4 fields:\n+\n+1. The *file mode*, which must be one of:\n+  - `100644`: regular file (with <<object,object type>> `blob`)\n+  - `100755`: executable file (with type `blob`)\n+  - `120000`: symbolic link (with type `blob`)\n+  - `160000`: gitlink, for use with submodules (with type `commit`)\n+2. The *<<blob,blob>>* ID of the file,\n+   or (rarely) the *<<commit,commit>>* ID of the submodule\n+3. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n+   there's a merge conflict there can be multiple versions of the same\n+   filename in the index.\n+4. The *file path*, for example `src/hello.py`\n+\n+It's extremely uncommon to look at the index directly: normally you'd\n+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n+But you can use `git ls-files --stage` to see the index.\n+Here's the output of `git ls-files --stage` in a repository with 2 files:\n+\n+----\n+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n+----\n+\n+[[reflogs]]\n+REFLOGS\n+-------\n+\n+Every time a branch, remote-tracking branch, or HEAD is updated, Git\n+updates a log called a \"reflog\" for that <<references,reference>>.\n+This means that if you make a mistake and \"lose\" a commit, you can\n+generally recover the commit ID by running `git reflog <reference>`.\n+\n+A reflog is a list of log entries. Each entry has:\n+\n+1. The *commit ID*\n+2. *Timestamp* when the change was made\n+3. *Log message*, for example `pull: Fast-forward`\n+\n+Reflogs only log changes made in your local repository.\n+They are not shared with remotes.\n+\n+You can view a reflog with `git reflog <reference>`.\n+For example, here's the reflog for a `main` branch which has changed twice:\n+\n+----\n+$ git reflog main --date=iso --no-decorate\n+750b4ea main@{2025-09-29 15:17:05 -0400}: commit: Add README\n+4ccb6d7 main@{2025-09-29 15:16:48 -0400}: commit (initial): Initial commit\n+----\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\ndiff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc\nindex e423e4765b..20ba121314 100644\n--- a/Documentation/glossary-content.adoc\n+++ b/Documentation/glossary-content.adoc\n@@ -297,8 +297,8 @@ This commit is referred to as a \"merge commit\", or sometimes just a\n \tidentified by its <<def_object_name,object name>>. The objects usually\n \tlive in `$GIT_DIR/objects/`.\n \n-[[def_object_identifier]]object identifier (oid)::\n-\tSynonym for <<def_object_name,object name>>.\n+[[def_object_identifier]]object identifier, object ID, oid::\n+\tSynonyms for <<def_object_name,object name>>.\n \n [[def_object_name]]object name::\n \tThe unique identifier of an <<def_object,object>>.  The\ndiff --git a/Documentation/meson.build b/Documentation/meson.build\nindex e34965c5b0..ace0573e82 100644\n--- a/Documentation/meson.build\n+++ b/Documentation/meson.build\n@@ -192,6 +192,7 @@ manpages = {\n   'gitcore-tutorial.adoc' : 7,\n   'gitcredentials.adoc' : 7,\n   'gitcvs-migration.adoc' : 7,\n+  'gitdatamodel.adoc' : 7,\n   'gitdiffcore.adoc' : 7,\n   'giteveryday.adoc' : 7,\n   'gitfaq.adoc' : 7,\n\nbase-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n-- \ngitgitgadget\n"},{"id":"530390","messageId":"xmqq4ir5ftcv.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v6.git.1762545177204.gitgitgadget@gmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-07T21:03:28Z","receivedAt":"2025-11-07T21:03:31Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Julia Evans <julia@jvns.ca>\n>\n> Git very often uses the terms \"object\", \"reference\", or \"index\" in its\n> documentation.\n\nNot about the updated text (which I haven't carefully read yet), but\nwe'd need this squashed in to avoid xml that does not validate when\nusing AsciiDoc (not Asciidoctor) to format gitdatamode.7\ndocumentation.\n\n    XMLTO gitdatamodel.7\nxmlto: /home/gitster/w/git.git/Documentation/gitdatamodel.xml does not validate (status 3)\nxmlto: Fix document syntax or use --skip-validation option\nDocument /home/gitster/w/git.git/Documentation/gitdatamodel.xml does not validate\n\nPerhaps I forgot to send this after queuing the previous round, even\nthough it was queued on top of the previous round in 'seen'.  The\npatch still applies cleanly to this version, and seems to fix the\nbreakage for me.\n\n   ... goes and looks ...\n\nAh, no, I did not forget.  The same patch is in the review thread of\nthe previous round:\n\n    https://lore.kernel.org/git/xmqqcy62213a.fsf@gitster.g/\n\n\n\n Documentation/gitdatamodel.adoc | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nindex 1cefbb4833..eaab3f800b 100644\n--- a/Documentation/gitdatamodel.adoc\n+++ b/Documentation/gitdatamodel.adoc\n@@ -18,13 +18,13 @@ means when the documentation says \"object\", \"reference\" or \"index\".\n \n Git's core operations use 4 kinds of data:\n \n-1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+1. <<object,Objects>>: commits, trees, blobs, and tag objects\n 2. <<references,References>>: branches, tags,\n    remote-tracking branches, etc\n 3. <<index,The index>>, also known as the staging area\n 4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n \n-[[objects]]\n+[[object]]\n OBJECTS\n -------\n \n-- \n2.52.0-rc1-455-g30608eb744\n\n"},{"id":"530391","messageId":"xmqqseepedue.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v6.git.1762545177204.gitgitgadget@gmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-07T21:23:53Z","receivedAt":"2025-11-07T21:23:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n>     changes in v6:\n>     \n>      * Make punctuation more consistent (from Patrick's review)\n\nGood.\n\n>      * Explain more about when exactly amended commits will get deleted\n>        (when their reflog entry expires), from Junio's review\n\nLooked good.\n\n>      * Be more explicit that there are only 5 file modes in Git (from\n>        Junio's review)\n\nI find \"These are all of the file modes in Git\" hard to read and\nunderstand, and more importantly, does not imply that we won't be\nadding any others strongly enough, than something like \"Git uses\nonly the following modes to represent the objects it stores\".\n\n>      * Make tag object description clearer (from Junio's review)\n\nOK.\n\n>      * We had a long discussion about the phrasing of \"A branch refers to a\n>        commit ID\" but I didn't come up with any ideas for how to improve the\n>        phrasing so I left it as is.\n\nI gave you something that is clearly an improvement there, though.\nJust like a tag object records \"the ID of the object it references\",\na branch records \"the ID of the commit it references\".\n\nAnother thing we discussed and a better alternative offered during\nthe last round was \"base directory\", to which Patrick mentioned \n\"we rather consistently use 'root tree'\"\n\n cf. https://lore.kernel.org/git/aQhcbHJjiI5GtV6Y@pks.im/\n\nOther than a few minor points I pointed out above, and the broken\nxml id/idref that does not validate, this round looks good to me.\n\nThanks.\n\n\n"},{"id":"530392","messageId":"07cca81a-10fd-49aa-b175-17b49e4f1116@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqseepedue.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-07T21:40:08Z","receivedAt":"2025-11-07T21:40:29Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Fri, Nov 7, 2025, at 4:23 PM, Junio C Hamano wrote:\n> \"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>>     changes in v6:\n>>     \n>>      * Make punctuation more consistent (from Patrick's review)\n>\n> Good.\n>\n>>      * Explain more about when exactly amended commits will get deleted\n>>        (when their reflog entry expires), from Junio's review\n>\n> Looked good.\n>\n>>      * Be more explicit that there are only 5 file modes in Git (from\n>>        Junio's review)\n>\n> I find \"These are all of the file modes in Git\" hard to read and\n> understand, and more importantly, does not imply that we won't be\n> adding any others strongly enough, than something like \"Git uses\n> only the following modes to represent the objects it stores\".\n>\n>>      * Make tag object description clearer (from Junio's review)\n\nI wonder if it would help to de-emphasize the octal representation\nof the file modes, and instead give them names since (from a\ndata model section Git's file modes are really more like an enum with\n5 values than )\n\nSomething like this:\n\n\tGit has 5 file modes:\n\n\t  - *regular file* (with <<object,object type>> `blob`)\n\t  - *executable file* (with type `blob`)\n\t  - *symbolic link* (with type `blob`)\n\t  - *directory* (with type `tree`)\n\t  - *gitlink*, for use with submodules (with type `commit`)\n\n\tNOTE: Git normally displays file modes in the same format as Unix file modes\n\t(100644, 100755, 120000, 040000, and 160000 respectively), but file modes are\n\tonly spiritually related to Unix file modes.\n\n> OK.\n>\n>>      * We had a long discussion about the phrasing of \"A branch refers to a\n>>        commit ID\" but I didn't come up with any ideas for how to improve the\n>>        phrasing so I left it as is.\n>\n> I gave you something that is clearly an improvement there, though.\n> Just like a tag object records \"the ID of the object it references\",\n> a branch records \"the ID of the commit it references\".\n\nTo me an \"improvement\" is something that helps the reader understand how Git's\ndata model, and I do not understand in what way this rephrasing helps the\nreader, or how you think the current phrasing might cause confusion for the\nreader.\n\nFrom my point of view \"a branch refers to a commit ID\" clearly means the exact\nsame thing as \"a branch records the ID of the commit it references\" and \n\"a branch records the ID of the commit it references\" is just a less clear and\nmore indirect way to communicate that.\n\n> Another thing we discussed and a better alternative offered during\n> the last round was \"base directory\", to which Patrick mentioned \n> \"we rather consistently use 'root tree'\"\n>\n>  cf. https://lore.kernel.org/git/aQhcbHJjiI5GtV6Y@pks.im/\n\nI think it would be better to stick with \"directory\" here, because I've gotten\nseveral reader comments saying that they do not understand the\nterm \"tree\" when it is used as a synonym for \"directory\".\n\nMaybe \"root directory\"?\n\n> Other than a few minor points I pointed out above, and the broken\n> xml id/idref that does not validate, this round looks good to me.\n\nWill fix the broken XML.\n"},{"id":"530394","messageId":"xmqqo6pde90w.fsf@gitster.g","threadId":"64244","inReplyTo":"07cca81a-10fd-49aa-b175-17b49e4f1116@app.fastmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-07T23:07:59Z","receivedAt":"2025-11-07T23:08:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> I wonder if it would help to de-emphasize the octal representation\n> of the file modes, and instead give them names since (from a\n> data model section Git's file modes are really more like an enum with\n> 5 values than )\n>\n> Something like this:\n>\n> \tGit has 5 file modes:\n>\n> \t  - *regular file* (with <<object,object type>> `blob`)\n> \t  - *executable file* (with type `blob`)\n> \t  - *symbolic link* (with type `blob`)\n> \t  - *directory* (with type `tree`)\n> \t  - *gitlink*, for use with submodules (with type `commit`)\n>\n> \tNOTE: Git normally displays file modes in the same format as Unix file modes\n> \t(100644, 100755, 120000, 040000, and 160000 respectively), but file modes are\n> \tonly spiritually related to Unix file modes.\n\nThen, I would suggest further deemphasize the \"file modes\" even\nmore.  \n\n    * Git stores/tracks 5 different file types, which are\n      non-executable files, executable files, symbolic links,\n      directories, and gitlinks.\n\n    * Git uses one bitpattern each to mark these 5 different kinds\n      of things in tree objects.  These bitpatterns were loosely\n      modelled after UNIX file mode bits.\n\nThe first half entirely avoids saying \"mode\" and that is very\ndeliberate.\n\n> ... I do not understand in what way this rephrasing helps the\n> reader, or how you think the current phrasing might cause confusion for the\n> reader.\n\nA branch (or any ref) does *not* *REFERENCE* an ID.  They refer to\nobjects by *recording* an ID.  The distinction is not clear with\nyour wording.\n\n>> Another thing we discussed and a better alternative offered during\n>> the last round was \"base directory\", to which Patrick mentioned \n>> \"we rather consistently use 'root tree'\"\n>>\n>>  cf. https://lore.kernel.org/git/aQhcbHJjiI5GtV6Y@pks.im/\n>\n> I think it would be better to stick with \"directory\" here, because I've gotten\n> several reader comments saying that they do not understand the\n> term \"tree\" when it is used as a synonym for \"directory\".\n>\n> Maybe \"root directory\"?\n\nI am OK with \"root\" but that is conditional; only if it is not used\ntogether with the word \"directory\".  We are not talking about \"root\ndirectory\" where common directories like /usr, /etc, /dev and /tmp\nhang immediately below.  If we use the word \"directory\", I'd\nstrongly prefer to see it with adjective like \"top-level\" that\nimplies that it is something different from \"root directory\" but is\nrelative to the project in question.\n\nThanks.\n"},{"id":"530408","messageId":"xmqqh5v448fr.fsf@gitster.g","threadId":"64244","inReplyTo":"xmqqo6pde90w.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-08T19:43:04Z","receivedAt":"2025-11-08T19:43:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Julia Evans\" <julia@jvns.ca> writes:\n>\n>> I wonder if it would help to de-emphasize the octal representation\n>> of the file modes, and instead give them names since (from a\n>> data model section Git's file modes are really more like an enum with\n>> 5 values than )\n>>\n>> Something like this:\n>>\n>> \tGit has 5 file modes:\n>>\n>> \t  - *regular file* (with <<object,object type>> `blob`)\n>> \t  - *executable file* (with type `blob`)\n>> \t  - *symbolic link* (with type `blob`)\n>> \t  - *directory* (with type `tree`)\n>> \t  - *gitlink*, for use with submodules (with type `commit`)\n>>\n>> \tNOTE: Git normally displays file modes in the same format as Unix file modes\n>> \t(100644, 100755, 120000, 040000, and 160000 respectively), but file modes are\n>> \tonly spiritually related to Unix file modes.\n>\n> Then, I would suggest further deemphasize the \"file modes\" even\n> more.  \n>\n>     * Git stores/tracks 5 different file types, which are\n>       non-executable files, executable files, symbolic links,\n>       directories, and gitlinks.\n>\n>     * Git uses one bitpattern each to mark these 5 different kinds\n>       of things in tree objects.  These bitpatterns were loosely\n>       modelled after UNIX file mode bits.\n>\n> The first half entirely avoids saying \"mode\" and that is very\n> deliberate.\n>\n>>> Another thing we discussed and a better alternative offered during\n>>> the last round was \"base directory\", to which Patrick mentioned \n>>> \"we rather consistently use 'root tree'\"\n>>>\n>>>  cf. https://lore.kernel.org/git/aQhcbHJjiI5GtV6Y@pks.im/\n>>\n>> I think it would be better to stick with \"directory\" here, because I've gotten\n>> several reader comments saying that they do not understand the\n>> term \"tree\" when it is used as a synonym for \"directory\".\n>>\n>> Maybe \"root directory\"?\n>\n> I am OK with \"root\" but that is conditional; only if it is not used\n> together with the word \"directory\".  We are not talking about \"root\n> directory\" where common directories like /usr, /etc, /dev and /tmp\n> hang immediately below.  If we use the word \"directory\", I'd\n> strongly prefer to see it with adjective like \"top-level\" that\n> implies that it is something different from \"root directory\" but is\n> relative to the project in question.\n\nThe above two points should probably be trivial to address.  I've\nalready squashed in the xml validation fixes to [v6], so let's\nfinish the rest quickly.\n\nI have no more words to offer somebody, who says she does not know\nwhy saying \"branch records ID of the commit it refers to\" is an\nimprovement over \"branch refers to ID of the commit\", when she\nalready accepts that \"The *ID* of the object it references\" is a\nbetter way than \"The object *ID* it references\" to describe one of\nthe fields in an annotated tag object.  So I wouldn't mind if v7\nstill said \"branch refers to commit id\".  We can update it with\nfollow-up series as needed, and it is not worth blocking the rest of\nthe document.\n\nRefs (including branches), refer to objects exactly the same way an\nannotated tag refers to another object, or a tree entry in a tree\nobject refers to a blob, tree, or a commit object.  Recording the\nhexadecimal hash is an implementation detail of the way how they\nreference the object, and the phrasing used for the tag field in an\nannotated tag reflects that by clearly distinguishing \n\n - recording the ID \n - referring to the object\n\nas two separate things.  The former is merely a means to the end\nwhich is the latter, i.e. the purpose of refs, tree-entry in a tree,\ntag field in a tag object, and all other things that refer to an\nobject by recording its ID.\n\n"},{"id":"530419","messageId":"D50AB3E0-E41C-49CD-9407-AB60331A6A43@gmail.com","threadId":"64244","inReplyTo":"xmqqo6pde90w.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-11-09T00:48:56Z","receivedAt":"2025-11-09T00:49:09Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 7 nov. 2025 à 18:08, Junio C Hamano <gitster@pobox.com> a écrit :\n> \n> ﻿\"Julia Evans\" <julia@jvns.ca> writes:\n> \n>> ... I do not understand in what way this rephrasing helps the\n>> reader, or how you think the current phrasing might cause confusion for the\n>> reader.\n> \n> A branch (or any ref) does *not* *REFERENCE* an ID.  They refer to\n> objects by *recording* an ID.  The distinction is not clear with\n> your wording.\n\nI concur with your later email that this is not worth delaying the rest of the document for.\n\nMy only other opinion on the matter is: what does making this distinction clear do to benefit readers of this document? I cannot come up with one, and I suspect Julia cannot either. \n\nClearly you feel strongly about it, though, given the shouty caps and “I have no more words” phrasing, which I find convey a tone that is… less than welcoming. Perhaps it’s simply time to move on? And someone motivated can propose improvements.  "},{"id":"530420","messageId":"xmqqa50v4x8n.fsf@gitster.g","threadId":"64244","inReplyTo":"D50AB3E0-E41C-49CD-9407-AB60331A6A43@gmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-09T04:59:36Z","receivedAt":"2025-11-09T04:59:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Knoble <ben.knoble@gmail.com> writes:\n\n> My only other opinion on the matter is: what does making this\n> distinction clear do to benefit readers of this document?\n\nI care about teaching people not just _what_ but _why_, because with\nvague distinction, many tend to memorize _what_ without\nunderstanding the reasoning behind it.  \"Our object names are\ncomputed as a hash of the contents in it formatted in a canonical\nway\" is \"what we do to compute an object name\", but the reason\nbehind the design is because we want to be able to dedup the same\nthing cheaply, detect two objects that are different cheaply, which\nis \"why\" in this example and it is equally, if not more, important.\n\nThe refs and objects record object names, and that is \"what\"; the\nreason why they do so is to refer to these objects.  If somebody\ncomes up with other ways to uniquely refer to these objects, their\nimplementation of git-compatible system does not have to make their\nrefs record object names---they can draw a line from a circle to a\nrectangle instead of writing the object name of that rectangle in\nthe circle---and their system is still compatible with the Git data\nmodel at the higher/conceptual level.  IOW, what exactly is done at\nthe byte level (like file format) is lower part of the \"data model\",\nbut what these byte level details wants to achieve is the other,\nhigher half of the \"data model\".  A data model documentation should\nteach both levels.\n"},{"id":"530448","messageId":"150f3442-93a6-4469-9c25-5bca24accc80@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqa50v4x8n.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-10T15:56:03Z","receivedAt":"2025-11-10T15:56:47Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"On Sat, Nov 8, 2025, at 11:59 PM, Junio C Hamano wrote:\n> Ben Knoble <ben.knoble@gmail.com> writes:\n>\n>> My only other opinion on the matter is: what does making this\n>> distinction clear do to benefit readers of this document?\n>\n> I care about teaching people not just _what_ but _why_, because with\n> vague distinction, many tend to memorize _what_ without\n> understanding the reasoning behind it.  \"Our object names are\n> computed as a hash of the contents in it formatted in a canonical\n> way\" is \"what we do to compute an object name\", but the reason\n> behind the design is because we want to be able to dedup the same\n> thing cheaply, detect two objects that are different cheaply, which\n> is \"why\" in this example and it is equally, if not more, important.\n>\n> The refs and objects record object names, and that is \"what\"; the\n> reason why they do so is to refer to these objects.  If somebody\n> comes up with other ways to uniquely refer to these objects, their\n> implementation of git-compatible system does not have to make their\n> refs record object names---they can draw a line from a circle to a\n> rectangle instead of writing the object name of that rectangle in\n> the circle---and their system is still compatible with the Git data\n> model at the higher/conceptual level.  IOW, what exactly is done at\n> the byte level (like file format) is lower part of the \"data model\",\n> but what these byte level details wants to achieve is the other,\n> higher half of the \"data model\".  A data model documentation should\n> teach both levels.\n\nThanks, this is exactly what I was looking for when I asked in what way\nthis rephrasing helps the reader. I agree that explaining the \"why\" is\nvery important.\n\nIt sounds like there are 2 \"whats\" and \"whys\" here:\n\n#1:\nwhat: object IDs are hashes of the contents\nwhy: this makes it very fast to avoid storing duplicate information,\n     and it's extremely fast to check if 2 objects are the same or not\n\nI love the idea of explaining this. I think we could incorporate it very easily\nby adding this paragraph in the \"Objects\" introduction, right before\n\"Here's how each type of object is structured\":\n\n    The reason the ID is a cryptographic hash is that it makes it extremely\n    fast for Git to tell if 2 objects have the same contents or not\n    (if they have the same ID, they have the same contents!),\n    and it means Git will never store duplicate objects.\n\nWill add that unless there are any objections.\n\n#2:\nwhat: The refs contain object IDs\nwhy: to refer to the object\n\nI think this is so obvious that going out of our way to explain it\nrisks confusing the reader. Spending too much time explaining something\nobvious can make the reader feel like they're missing something.\n\nI can't imagine any purpose for the refs containing object IDs other\nthan to refer to the object?\n\nLike you noticed in the tag object section, I think saying that the tag\nobject \"refers to an object\" works well in that context, but in the context\nof explaining what a branch is it makes the text more confusing.\n"},{"id":"530500","messageId":"xmqqfrakyj0w.fsf@gitster.g","threadId":"64244","inReplyTo":"150f3442-93a6-4469-9c25-5bca24accc80@app.fastmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T10:13:03Z","receivedAt":"2025-11-11T10:13:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> Like you noticed in the tag object section, I think saying that the tag\n> object \"refers to an object\" works well in that context, but in the context\n> of explaining what a branch is it makes the text more confusing.\n\nSorry, but I do not understand your objection, as I cannot see what\nconfusion it would bring in in saying \"a ref refers to an object\"\n(or \"a branch refers to a commit object\").  A ref refers to an\nobject, just like a tag field in a tag object or a tree-entry in a\ntree object refer to another object.  They do so by recording the\nname of the object they refer to.  So what's so confusing if we said\nthat straight?\n\nAre you saying that the noun \"reference\" (or \"ref\") is a sufficient\nclue to readers that their objective is to \"refer to\" something, so\n\"refers to\" is a redundant thing to say?\n\nMaybe its just me, but I find it a quite roundabout thing to say\nthat a ref refers to an object name (or \"ID\" if you like), simply\nbecause name or ID *is* a way to refer to the thing that is assigned\nthat name, so you are making a ref to refer to something (\"name\")\nthat refers to what it (\"ref\") originally wanted to refer to\n(\"object\").\n\nThat is what I find the most strange in the construction \"A branch\nrefers to ID\" at the conceptual level.  I am much less unhappy with\n\"A branch records an ID\", but stopping at that may make readers ask\nthe obvious question \"what goal does that design aim to achieve?\"\n(whose answer is of course \"to refer to the object that is assigned\nthat ID\").\n\n\"A branch refers to a commit object by recording its object name\",\n\"A branch records the ID of a commit it refers to\", \"A branch\nrecords the ID of the commit at the tip of its history\".  Any of the\nphrasing that does not make \"ID\" the object/target of the verb\n\"refer to\" would work to avoid that strange construction.\n\nBy the way, Ben used a word \"unwelcome\", but the words that are more\nappropriate to describe my reaction were \"frustrated\" (for not being\nable to explain what I know to be true clearly to make others\nunderstand) and \"disappointed\".\n"},{"id":"530505","messageId":"1348B322-2DE0-4D5A-9B87-888884555FD2@gmail.com","threadId":"64244","inReplyTo":"xmqqfrakyj0w.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-11-11T13:07:27Z","receivedAt":"2025-11-11T13:07:39Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 11 nov. 2025 à 05:13, Junio C Hamano <gitster@pobox.com> a écrit :\n> \n> By the way, Ben used a word \"unwelcome\", but the words that are more\n> appropriate to describe my reaction were \"frustrated\" (for not being\n> able to explain what I know to be true clearly to make others\n> understand) and \"disappointed\".\n\nWhile that isn’t precisely my phrasing, I certainly had a similar impact. Either way, thanks for clarifying: my intent (not always the same as impact; see prior) was to point out the way such things may be perceived. \n\nI deeply appreciate, by way of having been in similar situations, that you were frustrated/disappointed with the conversation or yourself—language is hard! I’m glad that frustration was not directed at any contributor, and I am also hopeful that future contributors will see a maintainer who cares about contributors and clear communication and decide they should spend time on this project. That is the source of my use of the word “welcome.” :)"},{"id":"530518","messageId":"2474339d-67bc-4a68-9f26-fe7edd172ec4@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqfrakyj0w.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-11T15:24:38Z","receivedAt":"2025-11-11T15:25:10Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"(this message got a bit long but the tl;dr is: maybe\n\"a branch is a label for a commit ID\" would work?)\n\nOn Tue, Nov 11, 2025, at 5:13 AM, Junio C Hamano wrote:\n> \"Julia Evans\" <julia@jvns.ca> writes:\n>\n>> Like you noticed in the tag object section, I think saying that the tag\n>> object \"refers to an object\" works well in that context, but in the context\n>> of explaining what a branch is it makes the text more confusing.\n>\n> Sorry, but I do not understand your objection, as I cannot see what\n> confusion it would bring in in saying \"a ref refers to an object\"\n> (or \"a branch refers to a commit object\"). \n\n> A ref refers to an\n> object, just like a tag field in a tag object or a tree-entry in a\n> tree object refer to another object.  They do so by recording the\n> name of the object they refer to.  So what's so confusing if we said\n> that straight?\n\nMy main strategy for figuring out if something is confusing or not is\nto talk to a few different users of the software and to ask them what\nthey think. I have a pretty empirical approach to figuring out if an\nexplanation is clear or not, if people think it's clear, then it's clear.\n(the question of \"accuracy\" is separate of course)\n\nFrom experience talking to people about references in Git I know\nthat this particular thing is extremely easy to get wrong, I used to\noften try to explain branches by saying something like \"A branch\nto a commit\" and I would get kind of a blank stare, which is why\nI'm so cautious about the phrasing here.\n\nThe reason I started with \"a branch is a name for a commit ID\"\ninitially is that I've found that people respond well to that phrasing\nin the past, and I don't think it gives a misleading impression about\nwhat a branch is. But I thought your point that (in the context\nof this document) the term \"name\" could perhaps be confused\nwith \"object name\" was reasonable, so I've been trying to\ncome up with an alternative.\n\nIt's always a little tricky to explain from first principles _why_\nsomething is confusing, when I started working on explaining Git a\ncouple of years ago I would have thought that many of your suggested\nphrasings would be an effective way to explain how Git branches work to\npeople and I was very surprised to see how careful I had to be around\nthe phrasing to get folks to understand how branches work.\n\n\n> Are you saying that the noun \"reference\" (or \"ref\") is a sufficient\n> clue to readers that their objective is to \"refer to\" something, so\n> \"refers to\" is a redundant thing to say?\n>\n> Maybe its just me, but I find it a quite roundabout thing to say\n> that a ref refers to an object name (or \"ID\" if you like), simply\n> because name or ID *is* a way to refer to the thing that is assigned\n> that name, so you are making a ref to refer to something (\"name\")\n> that refers to what it (\"ref\") originally wanted to refer to\n> (\"object\").\n\nMy thought process is sort of like this: I have two descriptions of\n\"a Git branch\" that people have responded well to in the past in\npractice:\n\n1. a branch is a name for a commit ID (you said that the use of\n   \"name\" could be confused with \"object name\", which\n   I thought was fair)\n2. a branch is a file that contains a commit ID (people often\n   respond very well to how concrete this is, but it refers to\n   Git's implementation which we're trying to avoid in this context)\n\nSo I'm trying to find a different wording that's similar to one of these\ntwo phrasings that I know are effective, but that doesn't have those\nproblems.\n\nSome of the options we've discussed are:\n\n- \"a branch refers to a commit ID\" (which as you've said has kind of a\n  \"type\" issue since technically the branch refers to a commit, though\n  when I've discussed it people they don't seem to think it's a problem\n  in practice)\n- \"a branch refers to a commit, using its ID\" (we had a long discussion\n  about how \"using its ID\" can lead the reader to think \"wait, how else\n  could you refer to a commit\", which in the context of trying to learn\n  what a branch is an unproductive distraction)\n- \"a branch records a commit ID\" (from my discussions I'm pretty sure the\n  word \"records\" does not work, I think it's because introducing a new\n  verb like \"records\" is always a bit dangerous)\n\nOne idea I just had is \"a branch is a label for a commit ID\", which\nI think avoids the issue with \"name\" from earlier.\n\n> That is what I find the most strange in the construction \"A branch\n> refers to ID\" at the conceptual level.  I am much less unhappy with\n> \"A branch records an ID\", but stopping at that may make readers ask\n> the obvious question \"what goal does that design aim to achieve?\"\n> (whose answer is of course \"to refer to the object that is assigned\n> that ID\").\n>\n> \"A branch refers to a commit object by recording its object name\",\n> \"A branch records the ID of a commit it refers to\", \"A branch\n> records the ID of the commit at the tip of its history\".  Any of the\n> phrasing that does not make \"ID\" the object/target of the verb\n> \"refer to\" would work to avoid that strange construction.\n>\n> By the way, Ben used a word \"unwelcome\", but the words that are more\n> appropriate to describe my reaction were \"frustrated\" (for not being\n> able to explain what I know to be true clearly to make others\n> understand) and \"disappointed\".\n"},{"id":"530612","messageId":"xmqqa50rqcy1.fsf@gitster.g","threadId":"64244","inReplyTo":"2474339d-67bc-4a68-9f26-fe7edd172ec4@app.fastmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-12T19:16:06Z","receivedAt":"2025-11-12T19:16:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n>> Maybe its just me, but I find it a quite roundabout thing to say\n>> that a ref refers to an object name (or \"ID\" if you like), simply\n>> because name or ID *is* a way to refer to the thing that is assigned\n>> that name, so you are making a ref to refer to something (\"name\")\n>> that refers to what it (\"ref\") originally wanted to refer to\n>> (\"object\").\n> ...\n> One idea I just had is \"a branch is a label for a commit ID\", which\n> I think avoids the issue with \"name\" from earlier.\n>\n>> That is what I find the most strange in the construction \"A branch\n>> refers to ID\" at the conceptual level.  I am much less unhappy with\n>> \"A branch records an ID\", but stopping at that may make readers ask\n>> the obvious question \"what goal does that design aim to achieve?\"\n>> (whose answer is of course \"to refer to the object that is assigned\n>> that ID\").\n>>\n>> \"A branch refers to a commit object by recording its object name\",\n>> \"A branch records the ID of a commit it refers to\", \"A branch\n>> records the ID of the commit at the tip of its history\".  Any of the\n>> phrasing that does not make \"ID\" the object/target of the verb\n>> \"refer to\" would work to avoid that strange construction.\n\nSorry, but I am having a hard time to come up with something that I\ncan give to help somebody who rejects \"record\", saying that it is a\nnew verb, and in the same message introduces \"label\" as a better\nalternative, as we haven't seen \"label\" used in this context,\neither.\n\nBesides, a label, a name, or an ID are all that are used to refer to\nsomething (in this context, \"a commit object\"), so I find the newly\nproposed one just as roundabout as \"a branch refers to ID\" in the\nsame way.  The use of *ID* is a low-level implementation detail to\nmake the ref work as a label for, or make the ref refer to, an\nobject, so \"is a label for ID\" is just as bad as \"refers to ID\".\n\nIf we do not hesitate using a new word and introduce \"label\", \"a\nbranch works as a label for a commit object\" may probably work,\nprobably.\n\nThanks.\n"},{"id":"530614","messageId":"pull.1981.v7.git.1762977200244.gitgitgadget@gmail.com","threadId":"64244","inReplyTo":"pull.1981.v6.git.1762545177204.gitgitgadget@gmail.com","subject":"[PATCH v7] doc: add an explanation of Git's data model","fromName":"Julia Evans via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-12T19:53:20Z","receivedAt":"2025-11-12T19:53:24Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"From: Julia Evans <julia@jvns.ca>\n\nGit very often uses the terms \"object\", \"reference\", or \"index\" in its\ndocumentation.\n\nHowever, it's hard to find a clear explanation of these terms and how\nthey relate to each other in the documentation. The closest candidates\ncurrently are:\n\n1. `gitglossary`. This makes a good effort, but it's an alphabetically\n    ordered dictionary and a dictionary is not a good way to learn\n    concepts. You have to jump around too much and it's not possible to\n    present the concepts in the order that they should be explained.\n2. `gitcore-tutorial`. This explains how to use the \"core\" Git commands.\n   This is a nice document to have, but it's not necessary to learn how\n   `update-index` works to understand Git's data model, and we should\n   not be requiring users to learn how to use the \"plumbing\" commands\n   if they want to learn what the term \"index\" or \"object\" means.\n3. `gitrepository-layout`. This is a great resource, but it includes a\n   lot of information about configuration and internal implementation\n   details which are not related to the data model. It also does\n   not explain how commits work.\n\nThe result of this is that Git users (even users who have been using\nGit for 15+ years) struggle to read the documentation because they don't\nknow what the core terms mean, and it's not possible to add links\nto help them learn more.\n\nAdd an explanation of Git's data model. Some choices I've made in\ndeciding what \"core data model\" means:\n\n1. Omit pseudorefs like `FETCH_HEAD`, because it's not clear to me\n   if those are intended to be user facing or if they're more like\n   internal implementation details.\n2. Don't talk about submodules other than by mentioning how they\n   relate to trees. This is because Git has a lot of special features,\n   and explaining how they all work exhaustively could quickly go\n   down a rabbit hole which would make this document less useful for\n   understanding Git's core behaviour.\n3. Don't discuss the structure of a commit message\n   (first line, trailers etc).\n4. Don't mention configuration.\n5. Don't mention the `.git` directory, to avoid getting too much into\n   implementation details\n\nSigned-off-by: Julia Evans <julia@jvns.ca>\n---\n    doc: Add a explanation of Git's data model\n    \n    Changes in v2:\n    \n    The biggest change is to remove all mentions of the .git directory, and\n    explain references in a way that doesn't refer to \"directories\" at all,\n    and instead talks about the \"hierarchy\" (from Kristoffer and Patrick's\n    reviews).\n    \n    Also:\n    \n     * objects: Mention that an object ID is called an \"object name\", and\n       update the glossary to include the term \"object ID\" (from Junio's\n       review)\n     * objects: Replace \"SHA-1 hash\" with \"cryptographic hash\" which is more\n       accurate (from Patrick's review)\n     * blobs: Made the explanation of git gc a little higher level and took\n       some ideas from Patrick's suggested wording (from Patrick's and\n       Kroftoffer's reviews)\n     * commits: Mention that tag objects and commits can optionally have\n       other fields. I didn't mention the GPG signature specifically, but\n       don't have any objections to adding it. (from Patrick and Junio's\n       reviews)\n     * commits: Remove one of the mentions of git gc, since it perhaps opens\n       up too much of a rabbit hole: \"how does git gc decide which commits\n       to clean up?\". (from Kristoffer's review)\n     * tag objects: Add an example of how a tag object is represented (from\n       user feedback on the draft)\n     * index: Use the term \"file mode\" instead of \"permissions\", and list\n       all allowed file modes (from Patrick's review)\n     * index: Use \"stage number\" instead of \"number\" for index entries (from\n       Patrick's review)\n     * reflogs: Remove \"any ref can be logged\", it raises some questions of\n       \"how do you tell Git to log a ref that it isn't normally logging?\"\n       and my guess is that it's uncommon to ask Git to log more refs. I\n       don't think it's a \"lie\" to omit this but I can bring it back if\n       folks disagree. (from Patrick's review)\n     * reflogs: Fix an error I noticed in the explanation of reflogs: tags\n       aren't logged by default and remote-tracking branches are, according\n       to man git-config\n     * branches and tags: Be clearer about how branches are usually updated\n       (by committing), and make it a little more obvious that only branches\n       can be checked out. This is a bit tricky because using the word\n       \"check out\" introduces a rabbit hole that I want to avoid (what does\n       \"check out\" mean?). I've dealt this by just talking about the\n       \"current branch\" (HEAD) since that is defined here, and making it\n       more explicit that HEAD must either be a branch or a commit, there's\n       no \"HEAD is a tag\" option. (from Patrick's review)\n     * tags: Explain the differences between annotated and lightweight tags\n       (this is the main piece of user feedback I've gotten on the draft so\n       far)\n     * Various style/typo changes (\"2 or more\", linkgit:git-gc[1], removed\n       extra asterisks, added empty SYNOPSIS, \"commits -> tags\" typo fix,\n       add to meson build)\n    \n    non-changes:\n    \n     * I still haven't mentioned things that aren't part of the \"data\n       model\", like revision params and configuration. I think there could\n       be a place for them but I haven't found it yet.\n     * tag objects: I noticed that there's a \"tag\" header field in tag\n       objects (like tag v1.0.0) but I didn't mention it yet because I\n       couldn't figure out what the purpose of that field is (I thought the\n       tag name was stored in the reference, why is it duplicated in the tag\n       object?)\n    \n    Changes in v3:\n    \n    I asked for feedback from Git users on Mastodon and got 220 pieces of\n    feedback from 48 different users. People seemed very excited to read\n    about Git's data model. Usually I judge explanations by what folks\n    report learning from them. Here people reported learning:\n    \n     * how branches are stored (that a branch is \"a name for a commit\")\n     * how objects work\n     * that Git has separate \"author\" and \"committer\" fields\n     * that amending a commit does not change it\n     * that a tree is \"just a directory\" (not something more complicated),\n       and how trees are stored\n     * that Git repos can contain symlinks\n     * that Git saves modes separately from the OS.\n     * how the stage number works\n     * that when you git add a file, Git will create an object\n     * that third-party tools can create their own refs.\n     * that the reflog stores the history of branches (not just HEAD), and\n       what reflogs are for\n    \n    Also (of course) there were quite a few points of confusion! The main 4\n    pieces of feedback were\n    \n     1. The index section doesn't explain what the word \"staged\" means, and\n        one person says that it makes it sounds like only files that you\n        \"git add\"ed are in the index. Rewrite the explanation to avoid using\n        the word \"staged\" to define the index and instead define the word\n        \"staging\".\n     2. Explain the difference between \"annotated tags\" and \"lightweight\n        tags\" (done)\n     3. Add examples for tag objects and reflogs (done)\n     4. Mention a little more about where things are stored in the .git\n        directory, which I'd removed in v2. This seems most important for\n        .git/refs, so I added a hopefully accurate note about how refs are\n        stored by default, with a comment about one of the major\n        implications. I did not discuss where objects or the index are\n        stored, because I don't think the implementation details of how\n        objects are stored are as important, and there are better tools for\n        viewing the \"raw\" state of objects and the index (with git cat-file\n        -p or git ls-files --staged).\n    \n    Here's every other change I made in response to the feedback, as well as\n    a few comments that I did not address.\n    \n    intro:\n    \n     * Give a 1-sentence intro to \"reflog\"\n    \n    objects:\n    \n     * people really like having git ls-files --stage as a way to view the\n       index, so add git cat-file -p as well in a note\n    \n    commits:\n    \n     * 2 people asked \"Are commits stored as a diff?\". Say that diffs are\n       calculated at runtime, this is very important.\n     * The order the fields are given in don't match the order in the\n       example. Make them match.\n     * \"All the files in the commit, stored as a tree\" is throwing a few\n       people off. Be clearer that it's the tree ID of the base directory.\n     * Several people asked \"What's the difference between an author and\n       committer? I added an example using git cherry-pick that I'm not 100%\n       happy with (what if the reader doesn't know what cherry-pick does?).\n       There might be a better example to give here.\n     * In the note about commits being amended: one person suggested saying\n       \"creates a new commit with the same parent\" to make it clearer what\n       the relationship between the new and old commit are. I liked that\n       idea so I did it.\n    \n    trees:\n    \n     * file modes. 2 people want to know more about \"The file mode, for\n       example 100644\". Also 2 people are curious about what relationship\n       these have to Unix permissions. Say that they're inspired by Unix\n       permissions, and move the list of possible file modes up to make the\n       relationship clearer\n     * On \"so git-gc(1) periodically compresses objects to save disk space\",\n       there are a few follow up comments wondering about more, which makes\n       me think the comment about compression is actually a distraction. Say\n       something simpler instead, (\"Git only needs to store new versions of\n       files which were changed in that commit\"), from Junio's suggestion\n     * Re \"commit (a Git submodule)\": 2 people say it's not clear how trees\n       relate to submodules. Say that it refers to a commit in a different\n       repository.\n     * One person says they're not sure if the \"object ID\" is a hash. Link\n       it to the definition of \"object ID\".\n    \n    tag objects:\n    \n     * Requests for an example, added one.\n     * Requests to explain the difference between \"lightweight\" and\n       \"annotated\" tags, added it.\n    \n    tags:\n    \n     * one person thinks \"It’s expected that a tag will never change after\n       you create it.\" is too strong (since of course you can change it with\n       git tag -f). Say instead that tags are \"usually\" not changed.\n    \n    HEAD:\n    \n     * Several people are asking for more detail about detached HEAD state.\n       There's actually quite a lot to talk about here (what it means, how\n       it happens, what it implies, and how you might adjust your workflow\n       to avoid it by using git switch). I don't think we can get into all\n       of that here, so refer to the DETACHED HEAD section of git-checkout\n       instead. I'm not totally happy with the current version of that\n       section but that seems like the most practical solution right now.\n    \n    remote-tracking branches:\n    \n     * discuss refs/remotes/<remote>/HEAD.\n    \n    the index:\n    \n     * \"permissions\" should be \"file mode\" (like with trees). Changed.\n     * \"filename\" should be \"file path\". Changed.\n     * the stage number can only be 0, 1, 2, or 3, since it's 2 bits. Also\n       maybe say that the numbers have specific meanings. Said it can only\n       be 0/1/2/3 but did not give the specific meanings.\n    \n    reflogs\n    \n     * Request for an example. Added one.\n     * It's not clear if there's one reflog per branch/tag/HEAD, or if\n       there's one universal reflog. Make this clearer.\n     * Mention the role of the reflog in retrieving \"lost\" commits or\n       undoing bad rebases.\n    \n    Not fixed:\n    \n     * intro: A couple of people say that it's confusing that tags are both\n       \"an object\" and \"a reference\". Handled this by just explaining the\n       difference between an annotated and a lightweight tag further down.\n       I'd like to make this clearer in the intro but not sure if there's a\n       way to do it.\n     * commits and tag objects: one person asks if there's a reference for\n       the other \"optional fields\", like \"encoding\" and \"gpgsig\". I couldn't\n       find one, so left this as is.\n     * HEAD: A couple of people ask if there are any other symbolic\n       references other than HEAD, or if they can make their own symbolic\n       references. I don't know the answer to this.\n     * HEAD: the HEAD: HEAD thing looks weird, it made more sense when it\n       was HEAD: .git/HEAD. Will think about this.\n     * reflogs: One person asks: if reflogs only store local changes, why\n       does it track the user who made the change? Is that for remote\n       operations like fetches and pulls? Or for cases where more than one\n       user is using the same repo on a system? I don't know the answer to\n       this.\n     * reflogs: How can you see the full data in the reflog? git reflog show\n       doesn't list the user who made the change. git reflog show <refname>\n       --format=\"%h | %gd | %gn <%ge> | %gs\" --date=iso seems to work but\n       it's really a mouthful, not sure it's useful to include all that.\n     * index: Is it worth mentioning that the index can be locked? I don't\n       have an opinion about this.\n     * other: One person asks what a \"working tree\" is. It made me wonder if\n       \"the current working directory\" has a place in Git's data model. My\n       feeling is \"no\" but I could be convinced otherwise.\n     * overall: \"How can Git be so fast? If I switch branches, how does it\n       figure out what to add, remove or replace?\". I don't think this is\n       the right place for that discussion but it would\n     * there are some docs CI errors I haven't figured out yet (IDREF\n       attribute linkend references an unknown ID \"tree\")\n    \n    changes in v4:\n    \n    This is a combination of trying to make some of the intro text a little\n    more \"friendly\" for someone new to Git's data model, avoiding implying\n    things that are false, and removing information that isn't relevant to\n    the data model.\n    \n    intro:\n    \n     * Add a 1-line description of what a \"reflog\" is (from user feedback)\n    \n    objects:\n    \n     * Start with a \"friendly\" description of what an object is, similar to\n       what we do for references and the reflog\n     * Rename \"commits\" to \"commit\" and similarly for trees etc (from\n       Junio's review)\n     * Remove the explanation of what git cat-file -p does, since it might\n       be misleading and if people want to know they can read the man page\n       (from Junio's review)\n    \n    commits:\n    \n     * Start by saying that the commit contains the full directory structure\n       of all the files (from Junio's comment about how it may not be clear\n       that the commit contains all the files' exact contents at the time of\n       the commit)\n     * Remove the comment about cherry-pick (from Junio's review)\n     * Replace \"ask Git for a diff\" with \"ask Git to show the commit with\n       git show\" (from Junio's review)\n    \n    trees:\n    \n     * Make the description a little more friendly\n     * Reorder so that \"type\" is defined before we refer to the \"type\"\n     * Say that file modes are \"only spiritually related\" to Unix\n       permissions instead of talking about what Git \"supports\" (from\n       Junio's review)\n    \n    blobs:\n    \n     * Try to make it clearer how \"commits use relatively little disk space\"\n       is true while not implying that commits are diffs, by using an\n       example (from Junio's review)\n    \n    branches:\n    \n     * Replace \"a branch is a name for a commit ID\" with \"a branch refers to\n       a commit ID\" (except in the intro sentence for the \"references\"\n       section). Similarly for tags etc. (from Junio's review)\n     * Remove the note about how branches are stored in .git (from Junio's\n       review)\n    \n    HEAD:\n    \n     * Be clearer that HEAD is not always the current branch, because there\n       may not be a current branch (from Junio's review)\n    \n    index:\n    \n     * Be a little more specific about how exactly the index is converted\n       into a commit. (from Junio's comment about how it's not clear what\n       \"every file in the repository\" means)\n    \n    reflog:\n    \n     * Be clearer that there are many reflogs (one for each reference with a\n       log), not just one reflog (from Junio and Patrick's reviews)\n     * Omit the user and \"Before\" commit IDs from the list of fields,\n       because you usually don't see them (from Junio's review)\n     * Show the output of git reflog main in the example instead of the\n       contents of the reflog file, to avoid showing the user and before\n       commit ID\n    \n    changes in v5:\n    \n    Mostly smaller tweaks this time. The only major addition is to add a\n    note about how unreachable objects may be deleted.\n    \n    From Junio's review:\n    \n     * Remove \"type\" in the description of what's in a tree (since I have\n       learned that is not a separate field, it's part of the file mode)\n     * Fix a typo (\"these these\")\n     * Remove the intro sentence about what a \"commit\" is and instead only\n       describe its contents in the list of fields, to avoid implying that a\n       commit is the same as a tree\n     * Say \"Unix file modes\" instead of \"Unix permissions\"\n     * In the tag objects contents: make \"ID\" and \"type\" separate list items\n       since they're separate fields\n     * in the index section:\n       * list all of the possible file modes (since from my understanding\n         there are fewer allowed file modes here than in a tree)\n       * mention that the object can be either a commit or blob\n       * make the order match the order in git ls-files\n    \n    changes in v6:\n    \n     * Make punctuation more consistent (from Patrick's review)\n     * Explain more about when exactly amended commits will get deleted\n       (when their reflog entry expires), from Junio's review\n     * Be more explicit that there are only 5 file modes in Git (from\n       Junio's review)\n     * Make tag object description clearer (from Junio's review)\n     * We had a long discussion about the phrasing of \"A branch refers to a\n       commit ID\" but I didn't come up with any ideas for how to improve the\n       phrasing so I left it as is.\n    \n    changes in v7:\n    \n     * Replace \"file mode\" with \"file type\", to make it more obvious that\n       Git does not support general Unix file modes. Remove a broken XML\n       link as a side effect.\n     * Use \"top-level directory\" instead of \"base directory\"\n     * Like last time, I still don't have any better ideas for \"A branch\n       refers to a commit ID\"\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1981%2Fjvns%2Fgitdatamodel-v7\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1981/jvns/gitdatamodel-v7\nPull-Request: https://github.com/gitgitgadget/git/pull/1981\n\nRange-diff vs v6:\n\n 1:  6e2a7bbe6b ! 1:  22a1b32017 doc: add an explanation of Git's data model\n     @@ Documentation/gitdatamodel.adoc (new)\n      ++\n      +1. The full directory structure of all the files in that version of the\n      +   repository and each file's contents, stored as the *<<tree,tree>>* ID\n     -+   of the commit's base directory\n     ++   of the commit's top-level directory\n      +2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n      +  regular commits have 1 parent, merge commits have 2 or more parents\n      +3. An *author* and the time the commit was authored\n     @@ Documentation/gitdatamodel.adoc (new)\n      +    It lists, for each item in the tree:\n      ++\n      +1. The *filename*, for example `hello.py`\n     -+2. The *file mode*. These are all of the file modes in Git.\n     -+   They're only spiritually related to Unix file modes.\n     -++\n     -+  - `100644`: regular file (with <<object,object type>> `blob`)\n     -+  - `100755`: executable file (with type `blob`)\n     -+  - `120000`: symbolic link (with type `blob`)\n     -+  - `040000`: directory (with type `tree`)\n     -+  - `160000`: gitlink, for use with submodules (with type `commit`)\n     -+\n     -+3. The <<object-id,*object ID*>> with the contents of the file or directory\n     ++2. The *file type*, which must be one of these five types:\n     ++  - *regular file*\n     ++  - *executable file*\n     ++  - *symbolic link*\n     ++  - *directory*\n     ++  - *gitlink* (for use with submodules)\n     ++3. The <<object-id,*object ID*>> with the contents of the file, directory,\n     ++   or gitlink.\n      ++\n      +For example, this is how a tree containing one directory (`src`) and one file\n      +(`README.md`) is stored:\n     @@ Documentation/gitdatamodel.adoc (new)\n      +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n      +----\n      +\n     ++NOTE: In the output above, Git displays the file type of each tree entry\n     ++using a format that's loosely modelled on Unix file modes (`100644` is\n     ++\"regular file\", `100755` is \"executable file\", `120000` is \"symbolic\n     ++link\", `040000` is \"directory\", and `160000` is \"gitlink\"). It also\n     ++displays the object's type: `blob` for files and symlinks, `tree` for\n     ++directories, and `commit` for gitlinks.\n     ++\n      +[[blob]]\n      +blob::\n      +    A blob object contains a file's contents.\n     @@ Documentation/gitdatamodel.adoc (new)\n      +\n      +Each index entry has 4 fields:\n      +\n     -+1. The *file mode*, which must be one of:\n     -+  - `100644`: regular file (with <<object,object type>> `blob`)\n     -+  - `100755`: executable file (with type `blob`)\n     -+  - `120000`: symbolic link (with type `blob`)\n     -+  - `160000`: gitlink, for use with submodules (with type `commit`)\n     ++1. The *file type*, which must be one of:\n     ++  - *regular file*\n     ++  - *executable file*\n     ++  - *symbolic link*\n     ++  - *gitlink* (for use with submodules)\n      +2. The *<<blob,blob>>* ID of the file,\n      +   or (rarely) the *<<commit,commit>>* ID of the submodule\n      +3. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n\n\n Documentation/Makefile              |   1 +\n Documentation/gitdatamodel.adoc     | 307 ++++++++++++++++++++++++++++\n Documentation/glossary-content.adoc |   4 +-\n Documentation/meson.build           |   1 +\n 4 files changed, 311 insertions(+), 2 deletions(-)\n create mode 100644 Documentation/gitdatamodel.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 6fb83d0c6e..5f4acfacbd 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -52,6 +52,7 @@ MAN7_TXT += gitcli.adoc\n MAN7_TXT += gitcore-tutorial.adoc\n MAN7_TXT += gitcredentials.adoc\n MAN7_TXT += gitcvs-migration.adoc\n+MAN7_TXT += gitdatamodel.adoc\n MAN7_TXT += gitdiffcore.adoc\n MAN7_TXT += giteveryday.adoc\n MAN7_TXT += gitfaq.adoc\ndiff --git a/Documentation/gitdatamodel.adoc b/Documentation/gitdatamodel.adoc\nnew file mode 100644\nindex 0000000000..3614f5960e\n--- /dev/null\n+++ b/Documentation/gitdatamodel.adoc\n@@ -0,0 +1,307 @@\n+gitdatamodel(7)\n+===============\n+\n+NAME\n+----\n+gitdatamodel - Git's core data model\n+\n+SYNOPSIS\n+--------\n+gitdatamodel\n+\n+DESCRIPTION\n+-----------\n+\n+It's not necessary to understand Git's data model to use Git, but it's\n+very helpful when reading Git's documentation so that you know what it\n+means when the documentation says \"object\", \"reference\" or \"index\".\n+\n+Git's core operations use 4 kinds of data:\n+\n+1. <<objects,Objects>>: commits, trees, blobs, and tag objects\n+2. <<references,References>>: branches, tags,\n+   remote-tracking branches, etc\n+3. <<index,The index>>, also known as the staging area\n+4. <<reflogs,Reflogs>>: logs of changes to references (\"ref log\")\n+\n+[[objects]]\n+OBJECTS\n+-------\n+\n+All of the commits and files in a Git repository are stored as \"Git objects\".\n+Git objects never change after they're created, and every object has an ID,\n+like `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+\n+This means that if you have an object's ID, you can always recover its\n+exact contents as long as the object hasn't been deleted.\n+\n+Every object has:\n+\n+[[object-id]]\n+1. an *ID* (aka \"object name\"), which is a cryptographic hash of its\n+  type and contents.\n+  It's fast to look up a Git object using its ID.\n+  This is usually represented in hexadecimal, like\n+  `1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a`.\n+2. a *type*. There are 4 types of objects:\n+   <<commit,commits>>, <<tree,trees>>, <<blob,blobs>>,\n+   and <<tag-object,tag objects>>.\n+3. *contents*. The structure of the contents depends on the type.\n+\n+Here's how each type of object is structured:\n+\n+[[commit]]\n+commit::\n+    A commit contains these required fields\n+    (though there are other optional fields):\n++\n+1. The full directory structure of all the files in that version of the\n+   repository and each file's contents, stored as the *<<tree,tree>>* ID\n+   of the commit's top-level directory\n+2. Its *parent commit ID(s)*. The first commit in a repository has 0 parents,\n+  regular commits have 1 parent, merge commits have 2 or more parents\n+3. An *author* and the time the commit was authored\n+4. A *committer* and the time the commit was committed\n+5. A *commit message*\n++\n+Here's how an example commit is stored:\n++\n+----\n+tree 1b61de420a21a2f1aaef93e38ecd0e45e8bc9f0a\n+parent 4ccb6d7b8869a86aae2e84c56523f8705b50c647\n+author Maya <maya@example.com> 1759173425 -0400\n+committer Maya <maya@example.com> 1759173425 -0400\n+\n+Add README\n+----\n++\n+Like all other objects, commits can never be changed after they're created.\n+For example, \"amending\" a commit with `git commit --amend` creates a new\n+commit with the same parent.\n++\n+Git does not store the diff for a commit: when you ask Git to show\n+the commit with linkgit:git-show[1], it calculates the diff from its\n+parent on the fly.\n+\n+[[tree]]\n+tree::\n+    A tree is how Git represents a directory.\n+    It can contain files or other trees (which are subdirectories).\n+    It lists, for each item in the tree:\n++\n+1. The *filename*, for example `hello.py`\n+2. The *file type*, which must be one of these five types:\n+  - *regular file*\n+  - *executable file*\n+  - *symbolic link*\n+  - *directory*\n+  - *gitlink* (for use with submodules)\n+3. The <<object-id,*object ID*>> with the contents of the file, directory,\n+   or gitlink.\n++\n+For example, this is how a tree containing one directory (`src`) and one file\n+(`README.md`) is stored:\n++\n+----\n+100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n+040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n+----\n+\n+NOTE: In the output above, Git displays the file type of each tree entry\n+using a format that's loosely modelled on Unix file modes (`100644` is\n+\"regular file\", `100755` is \"executable file\", `120000` is \"symbolic\n+link\", `040000` is \"directory\", and `160000` is \"gitlink\"). It also\n+displays the object's type: `blob` for files and symlinks, `tree` for\n+directories, and `commit` for gitlinks.\n+\n+[[blob]]\n+blob::\n+    A blob object contains a file's contents.\n++\n+When you make a commit, Git stores the full contents of each file that\n+you changed as a blob.\n+For example, if you have a commit that changes 2 files in a repository\n+with 1000 files, that commit will create 2 new blobs, and use the\n+previous blob ID for the other 998 files.\n+This means that commits can use relatively little disk space even in a\n+very large repository.\n+\n+[[tag-object]]\n+tag object::\n+    Tag objects contain these required fields\n+    (though there are other optional fields):\n++\n+1. The *ID* of the object it references\n+2. The *type* of the object it references\n+3. The *tagger* and tag date\n+4. A *tag message*, similar to a commit message\n+\n+Here's how an example tag object is stored:\n+\n+----\n+object 750b4ead9c87ceb3ddb7a390e6c7074521797fb3\n+type commit\n+tag v1.0.0\n+tagger Maya <maya@example.com> 1759927359 -0400\n+\n+Release version 1.0.0\n+----\n+\n+NOTE: All of the examples in this section were generated with\n+`git cat-file -p <object-id>`.\n+\n+[[references]]\n+REFERENCES\n+----------\n+\n+References are a way to give a name to a commit.\n+It's easier to remember \"the changes I'm working on are on the `turtle`\n+branch\" than \"the changes are in commit bb69721404348e\".\n+Git often uses \"ref\" as shorthand for \"reference\".\n+\n+References can either refer to:\n+\n+1. An object ID, usually a <<commit,commit>> ID\n+2. Another reference. This is called a \"symbolic reference\"\n+\n+References are stored in a hierarchy, and Git handles references\n+differently based on where they are in the hierarchy.\n+Most references are under `refs/`. Here are the main types:\n+\n+[[branch]]\n+branches: `refs/heads/<name>`::\n+    A branch refers to a commit ID.\n+    That commit is the latest commit on the branch.\n++\n+To get the history of commits on a branch, Git will start at the commit\n+ID the branch references, and then look at the commit's parent(s),\n+the parent's parent, etc.\n+\n+[[tag]]\n+tags: `refs/tags/<name>`::\n+    A tag refers to a commit ID, tag object ID, or other object ID.\n+    There are two types of tags:\n+    1. \"Annotated tags\", which reference a <<tag-object,tag object>> ID\n+       which contains a tag message\n+    2. \"Lightweight tags\", which reference a commit, blob, or tree ID\n+       directly\n++\n+Even though branches and tags both refer to a commit ID, Git\n+treats them very differently.\n+Branches are expected to change over time: when you make a commit, Git\n+will update your <<HEAD,current branch>> to point to the new commit.\n+Tags are usually not changed after they're created.\n+\n+[[HEAD]]\n+HEAD: `HEAD`::\n+    `HEAD` is where Git stores your current <<branch,branch>>,\n+    if there is a current branch. `HEAD` can either be:\n++\n+1. A symbolic reference to your current branch, for example `ref:\n+   refs/heads/main` if your current branch is `main`.\n+2. A direct reference to a commit ID. In this case there is no current branch.\n+   This is called \"detached HEAD state\", see the DETACHED HEAD section\n+   of linkgit:git-checkout[1] for more.\n+\n+[[remote-tracking-branch]]\n+remote-tracking branches: `refs/remotes/<remote>/<branch>`::\n+    A remote-tracking branch refers to a commit ID.\n+    It's how Git stores the last-known state of a branch in a remote\n+    repository. `git fetch` updates remote-tracking branches. When\n+    `git status` says \"you're up to date with origin/main\", it's looking at\n+    this.\n++\n+`refs/remotes/<remote>/HEAD` is a symbolic reference to the remote's\n+default branch. This is the branch that `git clone` checks out by default.\n+\n+[[other-refs]]\n+Other references::\n+    Git tools may create references anywhere under `refs/`.\n+    For example, linkgit:git-stash[1], linkgit:git-bisect[1],\n+    and linkgit:git-notes[1] all create their own references\n+    in `refs/stash`, `refs/bisect`, etc.\n+    Third-party Git tools may also create their own references.\n++\n+Git may also create references other than `HEAD` at the base of the\n+hierarchy, like `ORIG_HEAD`.\n+\n+NOTE: Git may delete objects that aren't \"reachable\" from any reference\n+or <<reflogs,reflog>>.\n+An object is \"reachable\" if we can find it by following tags to whatever\n+they tag, commits to their parents or trees, and trees to the trees or\n+blobs that they contain.\n+For example, if you amend a commit with `git commit --amend`,\n+there will no longer be a branch that points at the old commit.\n+The old commit is recorded in the current branch's <<reflogs,reflog>>,\n+so it is still \"reachable\", but when the reflog entry expires it may\n+become unreachable and get deleted.\n+\n+the old commit will usually not be reachable, so it may be deleted eventually.\n+Reachable objects will never be deleted.\n+\n+[[index]]\n+THE INDEX\n+---------\n+The index, also known as the \"staging area\", is a list of files and\n+the contents of each file, stored as a <<blob,blob>>.\n+You can add files to the index or update the contents of a file in the\n+index with linkgit:git-add[1]. This is called \"staging\" the file for commit.\n+\n+Unlike a <<tree,tree>>, the index is a flat list of files.\n+When you commit, Git converts the list of files in the index to a\n+directory <<tree,tree>> and uses that tree in the new <<commit,commit>>.\n+\n+Each index entry has 4 fields:\n+\n+1. The *file type*, which must be one of:\n+  - *regular file*\n+  - *executable file*\n+  - *symbolic link*\n+  - *gitlink* (for use with submodules)\n+2. The *<<blob,blob>>* ID of the file,\n+   or (rarely) the *<<commit,commit>>* ID of the submodule\n+3. The *stage number*, either 0, 1, 2, or 3. This is normally 0, but if\n+   there's a merge conflict there can be multiple versions of the same\n+   filename in the index.\n+4. The *file path*, for example `src/hello.py`\n+\n+It's extremely uncommon to look at the index directly: normally you'd\n+run `git status` to see a list of changes between the index and <<HEAD,HEAD>>.\n+But you can use `git ls-files --stage` to see the index.\n+Here's the output of `git ls-files --stage` in a repository with 2 files:\n+\n+----\n+100644 8728a858d9d21a8c78488c8b4e70e531b659141f 0 README.md\n+100644 665c637a360874ce43bf74018768a96d2d4d219a 0 src/hello.py\n+----\n+\n+[[reflogs]]\n+REFLOGS\n+-------\n+\n+Every time a branch, remote-tracking branch, or HEAD is updated, Git\n+updates a log called a \"reflog\" for that <<references,reference>>.\n+This means that if you make a mistake and \"lose\" a commit, you can\n+generally recover the commit ID by running `git reflog <reference>`.\n+\n+A reflog is a list of log entries. Each entry has:\n+\n+1. The *commit ID*\n+2. *Timestamp* when the change was made\n+3. *Log message*, for example `pull: Fast-forward`\n+\n+Reflogs only log changes made in your local repository.\n+They are not shared with remotes.\n+\n+You can view a reflog with `git reflog <reference>`.\n+For example, here's the reflog for a `main` branch which has changed twice:\n+\n+----\n+$ git reflog main --date=iso --no-decorate\n+750b4ea main@{2025-09-29 15:17:05 -0400}: commit: Add README\n+4ccb6d7 main@{2025-09-29 15:16:48 -0400}: commit (initial): Initial commit\n+----\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\ndiff --git a/Documentation/glossary-content.adoc b/Documentation/glossary-content.adoc\nindex e423e4765b..20ba121314 100644\n--- a/Documentation/glossary-content.adoc\n+++ b/Documentation/glossary-content.adoc\n@@ -297,8 +297,8 @@ This commit is referred to as a \"merge commit\", or sometimes just a\n \tidentified by its <<def_object_name,object name>>. The objects usually\n \tlive in `$GIT_DIR/objects/`.\n \n-[[def_object_identifier]]object identifier (oid)::\n-\tSynonym for <<def_object_name,object name>>.\n+[[def_object_identifier]]object identifier, object ID, oid::\n+\tSynonyms for <<def_object_name,object name>>.\n \n [[def_object_name]]object name::\n \tThe unique identifier of an <<def_object,object>>.  The\ndiff --git a/Documentation/meson.build b/Documentation/meson.build\nindex e34965c5b0..ace0573e82 100644\n--- a/Documentation/meson.build\n+++ b/Documentation/meson.build\n@@ -192,6 +192,7 @@ manpages = {\n   'gitcore-tutorial.adoc' : 7,\n   'gitcredentials.adoc' : 7,\n   'gitcvs-migration.adoc' : 7,\n+  'gitdatamodel.adoc' : 7,\n   'gitdiffcore.adoc' : 7,\n   'giteveryday.adoc' : 7,\n   'gitfaq.adoc' : 7,\n\nbase-commit: bb69721404348ea2db0a081c41ab6ebfe75bdec8\n-- \ngitgitgadget\n"},{"id":"530616","messageId":"xmqqy0obov4d.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v7.git.1762977200244.gitgitgadget@gmail.com","subject":"Re: [PATCH v7] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-12T20:26:26Z","receivedAt":"2025-11-12T20:26:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +2. The *file type*, which must be one of these five types:\n> +  - *regular file*\n> +  - *executable file*\n> +  - *symbolic link*\n> +  - *directory*\n> +  - *gitlink* (for use with submodules)\n> +3. The <<object-id,*object ID*>> with the contents of the file, directory,\n> +   or gitlink.\n> ++\n> +For example, this is how a tree containing one directory (`src`) and one file\n> +(`README.md`) is stored:\n> ++\n> +----\n> +100644 blob 8728a858d9d21a8c78488c8b4e70e531b659141f README.md\n> +040000 tree 89b1d2e0495f66d6929f4ff76ff1bb07fc41947d src\n> +----\n> +\n> +NOTE: In the output above, Git displays the file type of each tree entry\n> +using a format that's loosely modelled on Unix file modes (`100644` is\n> +\"regular file\", `100755` is \"executable file\", `120000` is \"symbolic\n> +link\", `040000` is \"directory\", and `160000` is \"gitlink\"). It also\n> +displays the object's type: `blob` for files and symlinks, `tree` for\n> +directories, and `commit` for gitlinks.\n\nAs a description of the data model, moving the exact bit assignment\nto a side note like the above hunk (relative to the previous\niteration) does make the body text less cluttered, which I think is\na welcome change.\n"},{"id":"530631","messageId":"xmqqo6p6q32v.fsf@gitster.g","threadId":"64244","inReplyTo":"xmqqa50rqcy1.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-12T22:49:12Z","receivedAt":"2025-11-12T22:49:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> If we do not hesitate using a new word and introduce \"label\", \"a\n> branch works as a label for a commit object\" may probably work,\n> probably.\n\nAnother thing.\n\nDo we want to limit the definition of \"branch\" very narrowly, i.e.,\n\"subset of refs whose refname begins with refs/heads/\"?  \n\nOr do we want to give a description at a bit higher conceptual\nlevel, something like:\n\n  A branch is a mechanism to help you grow one line of history (in\n  the sea/cloud of commits) by (1) keeping track of the commit it\n  currently is at (by recording its ID in the ref used to implement\n  the branch), (2) allowing you easily record a new commit you\n  create while you are on it as a child of the current commit (by\n  allowing the symbolic ref \"HEAD\" to point the ref used to\n  implement the branch), (3) keeping the description of the theme of\n  the particular line of history being developed there (by using\n  \"branch.<name>.description\" configuration variable for the branch)\n  which is incorporated when the branch gets merged to an\n  integration branch, and (4) keeping track of how the branch has\n  grown over time (in the reflog for the ref used to implement the\n  branch).\n\nWe can limit ourselves to view a \"branch\" as a narrow subset of a\nref that can point at a single commit in the dag of commits, and it\ncan be updated at any time to point another different commit that\nhas no relation to the previous commit.\n\nOnce we stop limiting ourselves and explain the purpose of using a\n\"branch\", \"it can be updated to point any random commit\" stops being\nentirely true.  While the \"git branch -f\" command can be used to do\nso, doing so all the time would go against what makes a branch a\nbranch, i.e. to keep track of the process of growing the history,\nand it is expected that it would be a lot more common for the commit\npointed at by the branch ref to move by growing the history with\n\"git commit\", refining the history with \"git rebase\", etc.  But that\ncan only follow if readers understand the branch as more than \"just\na ref whose name begins with refs/heads/\".\n\nI am not sure what level the data model description you are writing\nshould be at.  The current description seems to concentrate too\nnarrowly on \"a branch is a specialization of a ref\" aspect, and\nwhile it is not incorrect as a description of a building block of a\ntool set to implement a workflow, it might be too limiting to form\na proper mental model.  I dunno.\n\n"},{"id":"530659","messageId":"2265ecb5-b0ba-4a28-904f-186ef5318562@app.fastmail.com","threadId":"64244","inReplyTo":"xmqqo6p6q32v.fsf@gitster.g","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-13T19:50:13Z","receivedAt":"2025-11-13T19:50:35Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"\n\nOn Wed, Nov 12, 2025, at 5:49 PM, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> If we do not hesitate using a new word and introduce \"label\", \"a\n>> branch works as a label for a commit object\" may probably work,\n>> probably.\n>\n> Another thing.\n>\n> Do we want to limit the definition of \"branch\" very narrowly, i.e.,\n> \"subset of refs whose refname begins with refs/heads/\"?  \n>\n> Or do we want to give a description at a bit higher conceptual\n> level, something like:\n>\n>   A branch is a mechanism to help you grow one line of history (in\n>   the sea/cloud of commits) by (1) keeping track of the commit it\n>   currently is at (by recording its ID in the ref used to implement\n>   the branch), (2) allowing you easily record a new commit you\n>   create while you are on it as a child of the current commit (by\n>   allowing the symbolic ref \"HEAD\" to point the ref used to\n>   implement the branch), (3) keeping the description of the theme of\n>   the particular line of history being developed there (by using\n>   \"branch.<name>.description\" configuration variable for the branch)\n>   which is incorporated when the branch gets merged to an\n>   integration branch, and (4) keeping track of how the branch has\n>   grown over time (in the reflog for the ref used to implement the\n>   branch).\n>\n> We can limit ourselves to view a \"branch\" as a narrow subset of a\n> ref that can point at a single commit in the dag of commits, and it\n> can be updated at any time to point another different commit that\n> has no relation to the previous commit.\n\n﻿﻿﻿I think talking too much about the intentions behind branches runs\nthe risk of getting into a discussion from Git workflows which IMO\nis definitely out of scope for this document. For example \"which is\nincorporated when the branch gets merged to an integration branch\" is\ntalking about a specific Git workflow.\n\nFrom my point of view as a Git user one of Git's biggest strengths is its\nflexibility; because branches _can_ be moved to point at a different\ncommit at any time in various ways (via `git reset --hard`, `git rebase`, or\n`git commit --amend`), there's a lot of flexibility in how someone can\nchoose to use Git, including never using branches at all. \n(the flexibility is also one of the things that makes Git hard of course :) )\n\nSo I'd prefer to keep editorializing about what a branch \"means\"\nto a minimum.\n\nRight now we have this, which tries to explain a very small amount\nabout how branches are used that should apply to almost\nall Git workflows:\n\n\"Even though branches and tags both refer to a commit ID, Git treats\nthem very differently. Branches are expected to change over time: when\nyou make a commit, Git will update your current branch to point to the\nnew commit. \"\n\n> Once we stop limiting ourselves and explain the purpose of using a\n> \"branch\", \"it can be updated to point any random commit\" stops being\n> entirely true.  While the \"git branch -f\" command can be used to do\n> so, doing so all the time would go against what makes a branch a\n> branch, i.e. to keep track of the process of growing the history,\n> and it is expected that it would be a lot more common for the commit\n> pointed at by the branch ref to move by growing the history with\n> \"git commit\", refining the history with \"git rebase\", etc.  But that\n> can only follow if readers understand the branch as more than \"just\n> a ref whose name begins with refs/heads/\".\n>\n> I am not sure what level the data model description you are writing\n> should be at.  The current description seems to concentrate too\n> narrowly on \"a branch is a specialization of a ref\" aspect, and\n> while it is not incorrect as a description of a building block of a\n> tool set to implement a workflow, it might be too limiting to form\n> a proper mental model.  I dunno.\n"},{"id":"530661","messageId":"xmqqv7jdlmr7.fsf@gitster.g","threadId":"64244","inReplyTo":"2265ecb5-b0ba-4a28-904f-186ef5318562@app.fastmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-13T20:07:40Z","receivedAt":"2025-11-13T20:07:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n> From my point of view as a Git user one of Git's biggest strengths is its\n> flexibility; because branches _can_ be moved to point at a different\n> commit at any time in various ways (via `git reset --hard`, `git rebase`, or\n> `git commit --amend`), there's a lot of flexibility in how someone can\n> choose to use Git, including never using branches at all. \n> (the flexibility is also one of the things that makes Git hard of course :) )\n>\n> So I'd prefer to keep editorializing about what a branch \"means\"\n> to a minimum.\n\nOK.\n\nIf you go to such an extreme and make readers oblivious to what a\nbranch means, some of the sanity measures we have (e.g., a branch\nref will never point at anything but a commit object, not even a\ncommit-ish tag is allowed) would become \"unnecessary nuisance\" to\nthem.  In other words, as I hinted, it takes a delicate balancing\nact.\n"},{"id":"530662","messageId":"160ef4a8-8e9c-4034-9607-2f268fdbf29d@app.fastmail.com","threadId":"64244","inReplyTo":"2265ecb5-b0ba-4a28-904f-186ef5318562@app.fastmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Julia Evans","fromEmail":"julia@jvns.ca","sentAt":"2025-11-13T20:18:25Z","receivedAt":"2025-11-13T20:18:48Z","isPatch":true,"sender":{"key":"julia@jvns.ca","avatar":"https://avatars.githubusercontent.com/u/817739?v=4"},"body":"On Thu, Nov 13, 2025, at 2:50 PM, Julia Evans wrote:\n> On Wed, Nov 12, 2025, at 5:49 PM, Junio C Hamano wrote:\n>> Junio C Hamano <gitster@pobox.com> writes:\n>>\n>>> If we do not hesitate using a new word and introduce \"label\", \"a\n>>> branch works as a label for a commit object\" may probably work,\n>>> probably.\n>>\n>> Another thing.\n>>\n>> Do we want to limit the definition of \"branch\" very narrowly, i.e.,\n>> \"subset of refs whose refname begins with refs/heads/\"?  \n>>\n>> Or do we want to give a description at a bit higher conceptual\n>> level, something like:\n>>\n>>   A branch is a mechanism to help you grow one line of history (in\n>>   the sea/cloud of commits) by (1) keeping track of the commit it\n>>   currently is at (by recording its ID in the ref used to implement\n>>   the branch), (2) allowing you easily record a new commit you\n>>   create while you are on it as a child of the current commit (by\n>>   allowing the symbolic ref \"HEAD\" to point the ref used to\n>>   implement the branch), (3) keeping the description of the theme of\n>>   the particular line of history being developed there (by using\n>>   \"branch.<name>.description\" configuration variable for the branch)\n>>   which is incorporated when the branch gets merged to an\n>>   integration branch, and (4) keeping track of how the branch has\n>>   grown over time (in the reflog for the ref used to implement the\n>>   branch).\n>>\n>> We can limit ourselves to view a \"branch\" as a narrow subset of a\n>> ref that can point at a single commit in the dag of commits, and it\n>> can be updated at any time to point another different commit that\n>> has no relation to the previous commit.\n>\n> ﻿﻿﻿I think talking too much about the intentions behind branches runs\n> the risk of getting into a discussion from Git workflows which IMO\n> is definitely out of scope for this document. For example \"which is\n> incorporated when the branch gets merged to an integration branch\" is\n> talking about a specific Git workflow.\n>\n> From my point of view as a Git user one of Git's biggest strengths is its\n> flexibility; because branches _can_ be moved to point at a different\n> commit at any time in various ways (via `git reset --hard`, `git rebase`, or\n> `git commit --amend`), there's a lot of flexibility in how someone can\n> choose to use Git, including never using branches at all. \n> (the flexibility is also one of the things that makes Git hard of course :) )\n>\n> So I'd prefer to keep editorializing about what a branch \"means\"\n> to a minimum.\n\nTo immediately contradict myself a bit: after sending this I thought to\nlook through Mark Dominus's great blog posts about Git to see if\nhe has anything to say about this, and I came across this article:\nhttps://blog.plover.com/prog/git/branches.html, called \"I wish people\nwould stop insisting that Git branches are nothing but refs\".\n\nIt reminded me that of course in Git the word \"branch\" often is used\nto mean \"a sequence of commits\", for example if I make a branch called\n`topic` and add 2 commits to it I might say that that \"branch\" is that\nsequence of two commits. I think the way Dominus talks about this is\nvery interesting:\n\n\tThe reason people say this, the disconnection is that the Git software\n\tdoesn't have any formal representation of branches. Conceptually, the\n\tbranch is there; the git commands just don't understand it. This is the\n\tmost important mismatch between the conceptual model and what the Git\n\tsoftware actually does.\n\nTo me the sticky point is that \"the branch is these two commits\" is an important\nand useful concept in Git, but it doesn't really _exist_ in Git's data model,\nbecause Git only stores a branch as a reference to a commit.\n\nOne way I've resolved this in the past is to say something like\n\"you can think about a branch in 3 different ways!\"\nhttps://wizardzines.com/comics/whats-a-branch/\n\nThe idea there is to talk about how a branch might be _conceptually_\n\"a line of development\", but that Git doesn't have anything in its data\nmodel to track what the \"base\" of the line of development is, so any\ntime you want Git to think of a branch as \"these 2 commits\" you need to\ngive it a way to determine the base.\n\n> Right now we have this, which tries to explain a very small amount\n> about how branches are used that should apply to almost\n> all Git workflows:\n>\n> \"Even though branches and tags both refer to a commit ID, Git treats\n> them very differently. Branches are expected to change over time: when\n> you make a commit, Git will update your current branch to point to the\n> new commit. \"\n>\n>> Once we stop limiting ourselves and explain the purpose of using a\n>> \"branch\", \"it can be updated to point any random commit\" stops being\n>> entirely true.  While the \"git branch -f\" command can be used to do\n>> so, doing so all the time would go against what makes a branch a\n>> branch, i.e. to keep track of the process of growing the history,\n>> and it is expected that it would be a lot more common for the commit\n>> pointed at by the branch ref to move by growing the history with\n>> \"git commit\", refining the history with \"git rebase\", etc.  But that\n>> can only follow if readers understand the branch as more than \"just\n>> a ref whose name begins with refs/heads/\".\n>>\n>> I am not sure what level the data model description you are writing\n>> should be at.  The current description seems to concentrate too\n>> narrowly on \"a branch is a specialization of a ref\" aspect, and\n>> while it is not incorrect as a description of a building block of a\n>> tool set to implement a workflow, it might be too limiting to form\n>> a proper mental model.  I dunno.\n"},{"id":"530664","messageId":"CAPx1Gvcf5=nBg9=AakfF=2tXVakBfddt7vTj+Wy9-497OcjviQ@mail.gmail.com","threadId":"64244","inReplyTo":"160ef4a8-8e9c-4034-9607-2f268fdbf29d@app.fastmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Chris Torek","fromEmail":"chris.torek@gmail.com","sentAt":"2025-11-13T20:34:35Z","receivedAt":"2025-11-13T20:34:48Z","isPatch":true,"sender":{"key":"chris.torek@gmail.com","avatar":"https://avatars.githubusercontent.com/u/16826774?v=4"},"body":"On Thu, Nov 13, 2025 at 12:19 PM Julia Evans <julia@jvns.ca> wrote:\n> To immediately contradict myself a bit: after sending this I thought to\n> look through Mark Dominus's great blog posts about Git to see if\n> he has anything to say about this, and I came across this article:\n> https://blog.plover.com/prog/git/branches.html, called \"I wish people\n> would stop insisting that Git branches are nothing but refs\".\n>\n> It reminded me that of course in Git the word \"branch\" often is used\n> to mean \"a sequence of commits\" ...\n\nYes, this is the crux of the issue: The word \"branch\" is ambiguous.\n\nIn Git, the *branch name* is the `refs/heads/whatever` name, and\nwe also have remote-tracking branch names under `refs/remotes/`.\nThe *branch*, however, is some ill-defined set of commits starting\nfrom the specific commit identified by a branch name *or* any other\nunique identifier, and then working backwards for some unspecified\nnumber of steps with unspecified constraints.\n\nSometimes the bare term \"branch\" means one or another of\nthese various things, and sometimes it's meant to encompass\nall of them...\n\nChris\n"},{"id":"530671","messageId":"xmqqwm3tjzoj.fsf@gitster.g","threadId":"64244","inReplyTo":"160ef4a8-8e9c-4034-9607-2f268fdbf29d@app.fastmail.com","subject":"Re: [PATCH v6] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-13T23:11:24Z","receivedAt":"2025-11-13T23:11:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans\" <julia@jvns.ca> writes:\n\n>> So I'd prefer to keep editorializing about what a branch \"means\"\n>> to a minimum.\n>\n> To immediately contradict myself a bit: after sending this I thought to\n> look through Mark Dominus's great blog posts about Git to see if\n> he has anything to say about this ...\n> The idea there is to talk about how a branch might be _conceptually_\n> \"a line of development\", but that Git doesn't have anything in its data\n> model to track what the \"base\" of the line of development is, so any\n> time you want Git to think of a branch as \"these 2 commits\" you need to\n> give it a way to determine the base.\n\nYup, that is why you need to walk a fine line between what is hard\nand mechanical \"bits in the system\" data model, and the conceptual\ngoal human users build using the bits as building blocks.\n\nAside from that \"branch\" description, the rest of the document has\nbeen polished well enough that we are quickly approaching the point\nof diminishing returns, I would think.  Should we declare victory\nand mark the topic for 'next' by now?\n\nThanks.\n"},{"id":"531168","messageId":"xmqqv7j11nkc.fsf@gitster.g","threadId":"64244","inReplyTo":"pull.1981.v7.git.1762977200244.gitgitgadget@gmail.com","subject":"Re: [PATCH v7] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-23T02:37:39Z","receivedAt":"2025-11-23T02:37:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n>     changes in v7:\n>     \n>      * Replace \"file mode\" with \"file type\", to make it more obvious that\n>        Git does not support general Unix file modes. Remove a broken XML\n>        link as a side effect.\n>      * Use \"top-level directory\" instead of \"base directory\"\n>      * Like last time, I still don't have any better ideas for \"A branch\n>        refers to a commit ID\"\n\nWe haven't seen much comment on this iteration, and hopefully that\nis not showing the lack of interest ;-)  Shall we mark the topic for\n'next' now?\n\nThanks.\n"},{"id":"531486","messageId":"aS1OXhBcx0IegwRw@pks.im","threadId":"64244","inReplyTo":"xmqqv7j11nkc.fsf@gitster.g","subject":"Re: [PATCH v7] doc: add an explanation of Git's data model","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-12-01T08:14:22Z","receivedAt":"2025-12-01T08:14:29Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sat, Nov 22, 2025 at 06:37:39PM -0800, Junio C Hamano wrote:\n> \"Julia Evans via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> >     changes in v7:\n> >     \n> >      * Replace \"file mode\" with \"file type\", to make it more obvious that\n> >        Git does not support general Unix file modes. Remove a broken XML\n> >        link as a side effect.\n> >      * Use \"top-level directory\" instead of \"base directory\"\n> >      * Like last time, I still don't have any better ideas for \"A branch\n> >        refers to a commit ID\"\n> \n> We haven't seen much comment on this iteration, and hopefully that\n> is not showing the lack of interest ;-)  Shall we mark the topic for\n> 'next' now?\n\nI've been out of office, but I certainly think that this version is more\nthan \"good enough\", and I have a lot of interest in these topics. I've\nseen you already merged it to 'master' -- yay!\n\nBy the way, thanks a ton Julia for all these improvements to our docs. I\nhighly appreciate them and think that this is sorely needed. Our users\nwill certainly appreciate your work!\n\nPatrick\n"},{"id":"531553","messageId":"xmqq345t840s.fsf@gitster.g","threadId":"64244","inReplyTo":"aS1OXhBcx0IegwRw@pks.im","subject":"Re: [PATCH v7] doc: add an explanation of Git's data model","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-12-02T12:25:07Z","receivedAt":"2025-12-02T12:25:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> By the way, thanks a ton Julia for all these improvements to our docs. I\n> highly appreciate them and think that this is sorely needed. Our users\n> will certainly appreciate your work!\n\nSame here.\n"}]}