Re: [PATCH v6] doc: add an explanation of Git's data model
- From
Julia Evans <julia@jvns.ca>
- Date
- Nov 13, 2025, 19:50 UTC
- Message-ID
- <2265ecb5-b0ba-4a28-904f-186ef5318562@app.fastmail.com>
- In-Reply-To
- <xmqqo6p6q32v.fsf@gitster.g>
On Wed, Nov 12, 2025, at 5:49 PM, Junio C Hamano wrote:
Show 32 quoted lines
> Junio C Hamano <gitster@pobox.com> writes: > >> If we do not hesitate using a new word and introduce "label", "a >> branch works as a label for a commit object" may probably work, >> probably. > > Another thing. > > Do we want to limit the definition of "branch" very narrowly, i.e., > "subset of refs whose refname begins with refs/heads/"? > > Or do we want to give a description at a bit higher conceptual > level, something like: > > A branch is a mechanism to help you grow one line of history (in > the sea/cloud of commits) by (1) keeping track of the commit it > currently is at (by recording its ID in the ref used to implement > the branch), (2) allowing you easily record a new commit you > create while you are on it as a child of the current commit (by > allowing the symbolic ref "HEAD" to point the ref used to > implement the branch), (3) keeping the description of the theme of > the particular line of history being developed there (by using > "branch.<name>.description" configuration variable for the branch) > which is incorporated when the branch gets merged to an > integration branch, and (4) keeping track of how the branch has > grown over time (in the reflog for the ref used to implement the > branch). > > We can limit ourselves to view a "branch" as a narrow subset of a > ref that can point at a single commit in the dag of commits, and it > can be updated at any time to point another different commit that > has no relation to the previous commit.
I think talking too much about the intentions behind branches runs the risk of getting into a discussion from Git workflows which IMO is definitely out of scope for this document. For example "which is incorporated when the branch gets merged to an integration branch" is talking about a specific Git workflow.
From my point of view as a Git user one of Git's biggest strengths is its flexibility; because branches _can_ be moved to point at a different commit at any time in various ways (via `git reset --hard`, `git rebase`, or `git commit --amend`), there's a lot of flexibility in how someone can choose to use Git, including never using branches at all. (the flexibility is also one of the things that makes Git hard of course :) )
So I'd prefer to keep editorializing about what a branch "means" to a minimum.
Right now we have this, which tries to explain a very small amount about how branches are used that should apply to almost all Git workflows:
"Even though branches and tags both refer to a commit ID, Git treats them very differently. Branches are expected to change over time: when you make a commit, Git will update your current branch to point to the new commit. "
Show 17 quoted lines
> Once we stop limiting ourselves and explain the purpose of using a > "branch", "it can be updated to point any random commit" stops being > entirely true. While the "git branch -f" command can be used to do > so, doing so all the time would go against what makes a branch a > branch, i.e. to keep track of the process of growing the history, > and it is expected that it would be a lot more common for the commit > pointed at by the branch ref to move by growing the history with > "git commit", refining the history with "git rebase", etc. But that > can only follow if readers understand the branch as more than "just > a ref whose name begins with refs/heads/". > > I am not sure what level the data model description you are writing > should be at. The current description seems to concentrate too > narrowly on "a branch is a specialization of a ref" aspect, and > while it is not incorrect as a description of a building block of a > tool set to implement a workflow, it might be too limiting to form > a proper mental model. I dunno.