Re: [PATCH v3] doc: add a explanation of Git's data model
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Oct 16, 2025, 20:48 UTC
- Message-ID
- <xmqqbjm63751.fsf@gitster.g>
- In-Reply-To
- <03db91a6-148b-436f-8afa-0273a1f5d508@app.fastmail.com>
"Julia Evans" <julia@jvns.ca> writes:
Show 15 quoted lines
>>>> It is not really "supporting" file modes. Rather, Git only records >>>> 5 kinds of entities associated with each path in a tree object, and >>>> uses numbers taht remotely resemble POSIX file modes to represent >>>> these 5 kinds. >>>> >>>> Perhaps "supports" -> "uses"? >>> >>> "Uses" sounds good to me. >> >> Also "much more limited" is misleading. We only represent 5 kinds >> of things, so we use only 5 mode-bits-looking numbers. > > What does it mislead the reader to think? My goal is to communicate that > if you want to tell Git to remember that a file's Unix permissions were > 700, that's not possible.
Yes, rewording "support" to "use" is one good way to do so. But "limited" implies that lifting the limitation would allow you to store more. That is the misguided thinking I want to avoid here. There is no limitations to lift. We only differentiate 5 kinds hence we only use 5 permission-bit-looking numbers. We do not differenciate a file with permission 0600 from aother with 0644.
Show 12 quoted lines
>>>> Here it may be worth noting that this "filename" is a single
>>>> pathname component (roughly, what you would see in non-recursive
>>>> "ls"). In other words, it may be a directory name.
>>
>> Comments?
>
> Oops, missed this in my first pass.
>
> I looked at them man pages for a couple of commands ("mv", "cp")
> and it looks like it's normal to refer to files and directories jointly
> as "files", or refer to them as having a "file name". So I think it's okay
> to call it a "file name" even if the "file" may be a directory.Ah, not that part. I was more interested in seeing how we express "in these names, there won't be any slashes".
>>>>> +[[blob]] >>>>> +blobs::
By the way, I kept forgetting to mention, but why are all of these listed terms plural (not just object types but also "branches" and "tags"?
> But it's not true that Git treats blobs as opaque binary data, unlike > other blob storage systems, Git has diff and merge algorithms to > interpret the contents of the file to some extent and try to do useful > things with them.
Yes, but diff and merge happens way above the object layer, where the question "what is blob" has a meaning. And these "blobs are recorded in a tree together with other blobs and trees recursively, and the single top-level tree describes a snapshot of a single state, which is recorded in a commit" data model descriptions is exactly about the lower-level object layer.
> Another goal we could have is to be clear that there are no limits to > what kind of files you can store in Git: you can equally well store text > files and binary files.
That is a natural consequence of blobs being nothing more than uninterpreted sequence of bytes.
Show 14 quoted lines
>>> I see that you don't like the "name for a commit ID" phrasing :) >>> Maybe there's another way to say it, though again none of the test >>> readers said they were confused by this or disagreed with the phrasing. >> >> Yes, I get that given "refs/heads/main", you want to say "main" is >> one of the ways to have repo_get_oid() to yield the commit object, >> and you are using "name" in that sense, but it is more like a ref >> can be used to name an object. It is *not* the name of the object, >> because the object can have other names, and more importantly, it >> (i.e., to give a name for an object) is not the only thing that a >> ref can do. > > That's interesting, what else can a ref do other than to give a name to > an object?
For example, a ref is a key to reflog, so obvoiusly it is more than just a single commit. If you say "git checkout main" and "git checkout main^{commit}", they refer to the same commit, but the former is a sign that you want the next commit you make from that state to grow that branch (and not any other branch you may have that happen to be pointing at the same commit), while the other one is not.
Show 11 quoted lines
>> And that is why I do not like that phrasing, combined >> with the target of giving that name is spelled "a commit ID". The >> commit ID is already another way to name the thing the refname can >> be also used to name: a commit object. A commit object and a commit >> object name are different things. The latter is a name that can >> refer to the former. > > I'm curious about why it's important to you to make this distinction > between a commit ID and a commit object. To me the commit ID and the > commit object come as a package, since the commit ID is calculated from > the commit object.
It may be the most natural name for the commit object, but that does not mean the name is the object. Let's not go phylosophical.