Re: [PATCH v3] doc: add a explanation of Git's data model
- From
Julia Evans <julia@jvns.ca>
- Date
- Oct 15, 2025, 17:20 UTC
- Message-ID
- <353916d3-c977-40e5-9251-1535b226cc9e@app.fastmail.com>
- In-Reply-To
- <xmqqsefkuqkv.fsf@gitster.g>
On Wed, Oct 15, 2025, at 11:34 AM, Junio C Hamano wrote:
Show 10 quoted lines
> Patrick Steinhardt <ps@pks.im> writes: > >>> +Like all other objects, commits can never be changed after they're created. >>> +For example, "amending" a commit with `git commit --amend` creates a new >>> +commit with the same parent. >> >> Let's say "parents" instead of "parent" here so that it also works for >> root and merge commits. > > I just found it amusing that parents can be 0 ;-)
Will change to parent(s).
Show 29 quoted lines
>>> +NOTE: By default, Git references are stored as files in the `.git` directory. >>> +For example, the branch `main` is stored in `.git/refs/heads/main`. >>> +This means that you can't have branches named both `maya` and `maya/some-task`, >>> +because there can't be a file and a directory with the same name. >> >> Hm. I think mentioning this can help, but it may also creates questions >> when someone has a "main" branch but is unable find it in >> ".git/refs/heads/main" because it has either been packed, or because the >> repository uses reftables. > > I had the same thought. The only thing we want to stress here is > that the names of refs _behave_ like filesystem entities. So how > about saying just > > Note: when you have a branch with <name>, you cannot have any > branch whose name begins with "<name>/". > > and stop at it? It may look like an arbitrary limitation, and once > in a distant future ref-files gets retired, it will become one (as > there is no inherent reason why reftable backend must retain it; it > only enforces the same limitation to ensure that the names it stores > interoperate with another clone that uses ref-files backend). At > the data-model level (which is the theme of this document), it is > just as immaterial as refnames may be case insensitive on some > systems. > > Mentioning the limitation may be good, but the data model document > is not the right place to explain where this limitation comes from > (i.e. to be compatible with and expressible in ref-files backend).
I'm still not clear on why you think we shouldn't mention that how references behave depends on which filesystem you're using.
Is it because the fact that how references behave depends on which FS you're using is considered a "bug", Git is working on eventually fixing that bug via the reftable backend, and we don't want to document "bugs" as an expected part of the data model?
I do think it's important to tell users where the data model has "weak points" where the abstraction leaks through to the implementation, pretending that abstractions are stronger than they are leads to unnecessary confusion.
> We do not say "you may not be able to have 'maya' branch and 'mAYa' > branch at the same time on some systems", either ;-).
Speaking of case-insensitive filesystems, I wonder if we should add a short note about the rules for filenames in Git. I ran into an issue recently where I had a filename with a colon in it, and my collaborator (who was using Windows) could not check out the branch because of that, and I saw another similar issue recently where one collaborator was using a case-insensitive filesystem and the other wasn't.
My guess is that Git does not enforce any rules about filenames (?), and it's up to the user to make sure that the filenames in the repository will work well for everyone collaborating on the repository.
Show 5 quoted lines
>>> +Git stores a history called a "reflog" for every branch, remote-tracking >> >> I think it's a bit unclear what "history" means here. Maybe: > > "records of updates", perhaps?
Agreed. Perhaps this instead:
Every time a branch, remote-tracking branch, or HEAD is updated, Git updates a log called a "reflog" for that <<reference,reference>>.