Re: [PATCH v6] doc: add an explanation of Git's data model
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Nov 8, 2025, 19:43 UTC
- Message-ID
- <xmqqh5v448fr.fsf@gitster.g>
- In-Reply-To
- <xmqqo6pde90w.fsf@gitster.g>
Junio C Hamano <gitster@pobox.com> writes:
Show 54 quoted lines
> "Julia Evans" <julia@jvns.ca> writes: > >> I wonder if it would help to de-emphasize the octal representation >> of the file modes, and instead give them names since (from a >> data model section Git's file modes are really more like an enum with >> 5 values than ) >> >> Something like this: >> >> Git has 5 file modes: >> >> - *regular file* (with <<object,object type>> `blob`) >> - *executable file* (with type `blob`) >> - *symbolic link* (with type `blob`) >> - *directory* (with type `tree`) >> - *gitlink*, for use with submodules (with type `commit`) >> >> NOTE: Git normally displays file modes in the same format as Unix file modes >> (100644, 100755, 120000, 040000, and 160000 respectively), but file modes are >> only spiritually related to Unix file modes. > > Then, I would suggest further deemphasize the "file modes" even > more. > > * Git stores/tracks 5 different file types, which are > non-executable files, executable files, symbolic links, > directories, and gitlinks. > > * Git uses one bitpattern each to mark these 5 different kinds > of things in tree objects. These bitpatterns were loosely > modelled after UNIX file mode bits. > > The first half entirely avoids saying "mode" and that is very > deliberate. > >>> Another thing we discussed and a better alternative offered during >>> the last round was "base directory", to which Patrick mentioned >>> "we rather consistently use 'root tree'" >>> >>> cf. https://lore.kernel.org/git/aQhcbHJjiI5GtV6Y@pks.im/ >> >> I think it would be better to stick with "directory" here, because I've gotten >> several reader comments saying that they do not understand the >> term "tree" when it is used as a synonym for "directory". >> >> Maybe "root directory"? > > I am OK with "root" but that is conditional; only if it is not used > together with the word "directory". We are not talking about "root > directory" where common directories like /usr, /etc, /dev and /tmp > hang immediately below. If we use the word "directory", I'd > strongly prefer to see it with adjective like "top-level" that > implies that it is something different from "root directory" but is > relative to the project in question.
The above two points should probably be trivial to address. I've already squashed in the xml validation fixes to [v6], so let's finish the rest quickly.
I have no more words to offer somebody, who says she does not know why saying "branch records ID of the commit it refers to" is an improvement over "branch refers to ID of the commit", when she already accepts that "The *ID* of the object it references" is a better way than "The object *ID* it references" to describe one of the fields in an annotated tag object. So I wouldn't mind if v7 still said "branch refers to commit id". We can update it with follow-up series as needed, and it is not worth blocking the rest of the document.
Refs (including branches), refer to objects exactly the same way an annotated tag refers to another object, or a tree entry in a tree object refers to a blob, tree, or a commit object. Recording the hexadecimal hash is an implementation detail of the way how they reference the object, and the phrasing used for the tag field in an annotated tag reflects that by clearly distinguishing
- recording the ID - referring to the object
as two separate things. The former is merely a means to the end which is the latter, i.e. the purpose of refs, tree-entry in a tree, tag field in a tag object, and all other things that refer to an object by recording its ID.