From: Julia Evans Date: Wed, 08 Oct 2025 19:06:29 GMT Subject: Re: [PATCH] doc: add a explanation of Git's data model Message-ID: <395232da-4ac6-4311-ae44-2bbf92fa6d2f@app.fastmail.com> In-Reply-To: On Wed, Oct 8, 2025, at 11:53 AM, Junio C Hamano wrote: > Patrick Steinhardt writes: > >>> I sometimes hear from users that "commits can't be snapshots", because >>> it would take up too much disk space to store every version of >>> every commit. So I find that sometimes explaining a little bit about the >>> implementation can make the information more memorable. >>> >>> Certainly I'm not able to remember details that don't make sense >>> with my mental model of how computers work and I don't expect other >>> people to either, so I think it's important to give an explanation that >>> handles the biggest "objections". >> >> Hm, fair I guess. In any case, if we want to mention this I'd leave away >> the details how exactly Git achieves this. E.g. we could say something >> like: >> >> Storing a new blob for every new version of a file can result to a >> lot of duplication. Git regularly runs repository maintenance to >> optimize to counteract this. Part of the maintenance involves >> compression of objects, where incremental changes to the same object >> are optimized to be stored as deltas, only. >> >> We skip over the details, but this should give enough pointers to an >> interested reader to go dig deeper. We could also generalize this to >> objects in general, not only blobs. > > Interesting. It is of course not wrong at all, but it was not what > I would have expected for the first explanation to help confused > folks who say "commits cannot be snapshots as they take too much > space". > > To me, it was a realization that even in a project whose tree (think > of "du -s .") is huge, each of its commits touches only a handful > of paths, hence a large portion of that huge tree would be shared > with the previous snapshot. That's a good point, I forgot that I've explained it that way too. I might change it to that instead. >>> > This misses "refs/remotes//HEAD". This reference is a symbolic >>> > reference that indicates the default branch on the remote side. >>> >>> Is "refs/remotes//HEAD" a remote-tracking branch? >>> I've never thought about that reference and I'm not sure what to call it. >> >> No, it's not. I think the term we use is "remote reference". > > Honestly I didn't know/think we have any special terminology for the > refs/remotes/*/HEAD symref. > > Historically HEAD did not "track" the remote state, and we did take > advantage of that fact to use it as a place to record the preference > with respect to which remote-tracking branch we would want to > primarily interact with. > > But these days because the protocol is capable of expressing where > the symrefs point at, the users can make it track just like all > other refs inside refs/remotes/*/ hiearchy. So I personally think > it is OK to call it in remote-tracking branch. I may just add this to the remote-tracking branch sentence then, which is hopefully correct: `refs/remotes//HEAD` is a symbolic reference to the remote's default branch. This is the branch that `git clone` checks out by default.