git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v3] doc: add a explanation of Git's data model

From
Julia Evans <julia@jvns.ca>
Date
Oct 16, 2025, 18:59 UTC
Message-ID
<03db91a6-148b-436f-8afa-0273a1f5d508@app.fastmail.com>
In-Reply-To
<xmqq347i948a.fsf@gitster.g>
On Thu, Oct 16, 2025, at 12:54 PM, Junio C Hamano wrote:
Show 28 quoted lines
> "Julia Evans" <julia@jvns.ca> writes:
>
>>>> +[[tree]]
>>>> +trees::
>>>> +    A tree is how Git represents a directory. It lists, for each item in
>>>> +    the tree:
>>>> ++
>>>> +[[file-mode]]
>>>> +1. The *file mode*, for example `100644`. The format is inspired by Unix
>>>> +   permissions, but Git's modes are much more limited. Git only supports these file modes:
>>>> ++
>>>> +  - `100644`: regular file (with type `blob`)
>>>> +  - `100755`: executable file (with type `blob`)
>>>> +  - `120000`: symbolic link (with type `blob`)
>>>> +  - `040000`: directory (with type `tree`)
>>>> +  - `160000`: gitlink, for use with submodules (with type `commit`)
>>>
>>> It is not really "supporting" file modes.  Rather, Git only records
>>> 5 kinds of entities associated with each path in a tree object, and
>>> uses numbers taht remotely resemble POSIX file modes to represent
>>> these 5 kinds.
>>>
>>> Perhaps "supports" -> "uses"?
>>
>> "Uses" sounds good to me.
>
> Also "much more limited" is misleading.  We only represent 5 kinds
> of things, so we use only 5 mode-bits-looking numbers.

What does it mislead the reader to think? My goal is to communicate that if you want to tell Git to remember that a file's Unix permissions were 700, that's not possible.

Show 11 quoted lines
>>>> +2. The *type*: either <<blob,`blob`>> (a file), `tree` (a directory),
>>>> +  or <<commit,`commit`>> (a Git submodule, which is a
>>>> +  commit from a different Git repository)
>>>> +3. The <<object-id,*object ID*>>
>>>> +4. The *filename*
>>>
>>> Here it may be worth noting that this "filename" is a single
>>> pathname component (roughly, what you would see in non-recursive
>>> "ls").  In other words, it may be a directory name.
>
> Comments?
Oops, missed this in my first pass.

I looked at them man pages for a couple of commands ("mv", "cp") and it looks like it's normal to refer to files and directories jointly as "files", or refer to them as having a "file name". So I think it's okay to call it a "file name" even if the "file" may be a directory.

Show 24 quoted lines
>>>> +[[blob]]
>>>> +blobs::
>>>> +    A blob is how Git represents a file. A blob object contains the
>>>> +    file's contents.
>>>
>>> "represents a file" hints as if the thing may know its name, but
>>> that is not the case (its name is given only by surrounding tree).
>>>
>>> "A blob is how Git represents uninterpreted series of bytes, and
>>> most commonly used to store file's contents." or something, perhaps?
>>
>> I'll say "A blob is how Git represents a file's contents", unless Git has
>> another use for blobs that I don't know about (I think it's not
>> that much of a stretch to say that a symbolic link is a special kind
>> of file where the "contents" are the the link destination).
>
> A few configuration variables like mailmap.blob name a blob object,
> for which _only_ its contents, i.e., the sequence of bytes, matter
> and where they originally were stored does not matter.
>
> But we are falling into the area of tautology, as any sequence of
> bytes can be stored in a file so they can be called "contents of a
> file".  But the point is that these bytes do not have to be stored
> to become a blob (think: "git cat-file -t blob -w --stdin").

I'm trying to think through what the goal of explaining the nature of a "blob" is.

To me describing blobs primarily as "bytes" makes it sound a bit like "Git will treat this as opaque binary data, Git will not attempt to interpret the contents of a blob in any way" (which is certainly true for many blob storage systems!).

But it's not true that Git treats blobs as opaque binary data, unlike other blob storage systems, Git has diff and merge algorithms to interpret the contents of the file to some extent and try to do useful things with them.

Another goal we could have is to be clear that there are no limits to what kind of files you can store in Git: you can equally well store text files and binary files.

Show 12 quoted lines
>> I think it's always clearer to be more specific when possible, if there's only
>> one purpose for blobs it's unnecessary (and IMO a bit misleading, because
>> it makes the reader wonder if there are other purposes that they should
>> know about) to say that blobs can be used to store any arbitrary bytes for
>> any purpose.
>
> I do not think describing other use cases is unnecessary.  Even if
> we limit ourselves to discuss a single purpose for blob, i.e. to
> represent the contents of a file, we should stress that blob is to
> store _only_ contents, and not other aspects of the file (e.g., in
> what paths with what mode), and that is where my reaction to "how
> Git reprsents a file" comes from.

I think it does make sense to say the blob stores only the contents, though IMO that's fairly clear already since we've already explained where the other parts of the file are stored by the time we get to explaining "blob".

Show 13 quoted lines
>>>> +[[branch]]
>>>> +branches: `refs/heads/<name>`::
>>>> +    A branch is a name for a commit ID.
>>>
>>> Well a commit ID is an alternative way to refer to a commit object
>>> *name*, so it is a bit strange to say "a name for a commit ID".
>>>
>>> Perhaps "A branch ref stores a commit ID." is better?
>>
>> I think I'll leave this alone, none of the many test readers reported
>> being confused by it.
>
> Would a confused person report that they are confused? ;-)

Everyone leaving feedback gets a prompt something like this asking them to categorize their feedback, and "I'm confused" is one of the options. https://jvns.ca/images/feedback-categories.png

I definitely got many "I'm confused" and "I have a question" comments about other things that were confusing to readers.

Show 11 quoted lines
>> I see that you don't like the "name for a commit ID" phrasing :)
>> Maybe there's another way to say it, though again none of the test
>> readers said they were confused by this or disagreed with the phrasing.
>
> Yes, I get that given "refs/heads/main", you want to say "main" is
> one of the ways to have repo_get_oid() to yield the commit object,
> and you are using "name" in that sense, but it is more like a ref
> can be used to name an object.  It is *not* the name of the object,
> because the object can have other names, and more importantly, it
> (i.e., to give a name for an object) is not the only thing that a
> ref can do.  

That's interesting, what else can a ref do other than to give a name to an object?

Show 6 quoted lines
> And that is why I do not like that phrasing, combined
> with the target of giving that name is spelled "a commit ID".  The
> commit ID is already another way to name the thing the refname can
> be also used to name: a commit object.  A commit object and a commit
> object name are different things.  The latter is a name that can
> refer to the former.

I'm curious about why it's important to you to make this distinction between a commit ID and a commit object. To me the commit ID and the commit object come as a package, since the commit ID is calculated from the commit object.

>  And a ref can be used just like the latter to
> refer to the former (i.e. "commit object").
> By the way, I do like the way many of your responses are "will think
> about it more", not "I'll take your version".
>
> Very much appreciated.

I'm glad to hear that! It's a fun puzzle to figure out how to express things clearly and accurately and concisely.

- Julia
Previous: Junio C HamanoNext: Junio C Hamano
Message 38 of 89 in “doc: add a explanation of Git's data model”
  1. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 3, 2025
  2. Kristoffer HaugsbakkOct 3, 2025
  3. Julia EvansOct 6, 2025
  4. D. Ben KnobleOct 6, 2025
  5. Julia EvansOct 6, 2025
  6. D. Ben KnobleOct 6, 2025
  7. Julia EvansOct 9, 2025
  8. Kristoffer HaugsbakkOct 8, 2025
  9. Junio C HamanoOct 6, 2025
  10. Julia EvansOct 6, 2025
  11. Kristoffer HaugsbakkOct 7, 2025
  12. Junio C HamanoOct 7, 2025
  13. Patrick SteinhardtOct 7, 2025
  14. Junio C HamanoOct 7, 2025
  15. Julia EvansOct 7, 2025
  16. Junio C HamanoOct 7, 2025
  17. D. Ben KnobleOct 7, 2025
  18. Julia EvansOct 7, 2025
  19. Patrick SteinhardtOct 8, 2025
  20. Junio C HamanoOct 8, 2025
  21. Julia EvansOct 8, 2025
  22. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 8, 2025
  23. Patrick SteinhardtOct 10, 2025
  24. Junio C HamanoOct 13, 2025
  25. Patrick SteinhardtOct 14, 2025
  26. Julia EvansOct 14, 2025
  27. Patrick SteinhardtOct 14, 2025
  28. Junio C HamanoOct 14, 2025
  29. doc: add a explanation of Git's data modelJulia Evans via GitGitGadget, Oct 14, 2025
  30. Patrick SteinhardtOct 15, 2025
  31. Junio C HamanoOct 15, 2025
  32. Julia EvansOct 15, 2025
  33. Junio C HamanoOct 15, 2025
  34. Julia EvansOct 16, 2025
  35. Junio C HamanoOct 15, 2025
  36. Julia EvansOct 16, 2025
  37. Junio C HamanoOct 16, 2025
  38. Julia EvansOct 16, 2025
  39. Junio C HamanoOct 16, 2025
  40. Kristoffer HaugsbakkOct 16, 2025
  41. Kristoffer HaugsbakkOct 20, 2025
  42. Junio C HamanoOct 20, 2025
  43. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Oct 27, 2025
  44. Junio C HamanoOct 27, 2025
  45. Julia EvansOct 28, 2025
  46. Junio C HamanoOct 28, 2025
  47. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Oct 30, 2025
  48. Junio C HamanoOct 31, 2025
  49. Patrick SteinhardtNov 3, 2025
  50. Junio C HamanoNov 3, 2025
  51. Julia EvansNov 3, 2025
  52. Junio C HamanoNov 4, 2025
  53. Julia EvansNov 4, 2025
  54. Junio C HamanoNov 4, 2025
  55. Julia EvansNov 4, 2025
  56. Junio C HamanoNov 4, 2025
  57. Julia EvansNov 5, 2025
  58. Ben KnobleNov 5, 2025
  59. Julia EvansNov 5, 2025
  60. Ben KnobleNov 6, 2025
  61. Junio C HamanoOct 31, 2025
  62. Patrick SteinhardtNov 3, 2025
  63. Julia EvansNov 3, 2025
  64. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Nov 7, 2025
  65. Junio C HamanoNov 7, 2025
  66. Junio C HamanoNov 7, 2025
  67. Julia EvansNov 7, 2025
  68. Junio C HamanoNov 7, 2025
  69. Junio C HamanoNov 8, 2025
  70. Ben KnobleNov 9, 2025
  71. Junio C HamanoNov 9, 2025
  72. Julia EvansNov 10, 2025
  73. Junio C HamanoNov 11, 2025
  74. Ben KnobleNov 11, 2025
  75. Julia EvansNov 11, 2025
  76. Junio C HamanoNov 12, 2025
  77. Junio C HamanoNov 12, 2025
  78. Julia EvansNov 13, 2025
  79. Junio C HamanoNov 13, 2025
  80. Julia EvansNov 13, 2025
  81. Chris TorekNov 13, 2025
  82. Junio C HamanoNov 13, 2025
  83. doc: add an explanation of Git's data modelJulia Evans via GitGitGadget, Nov 12, 2025
  84. Junio C HamanoNov 12, 2025
  85. Junio C HamanoNov 23, 2025
  86. Patrick SteinhardtDec 1, 2025
  87. Junio C HamanoDec 2, 2025
  88. Julia EvansOct 9, 2025
  89. Ben KnobleOct 10, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.