Re: [PATCH] doc: add a explanation of Git's data model
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Oct 8, 2025, 15:53 UTC
- Message-ID
- <xmqqecrdgzk8.fsf@gitster.g>
- In-Reply-To
- <aOXmA5L5LsUuXWEh@pks.im>
Patrick Steinhardt <ps@pks.im> writes:
Show 23 quoted lines
>> I sometimes hear from users that "commits can't be snapshots", because >> it would take up too much disk space to store every version of >> every commit. So I find that sometimes explaining a little bit about the >> implementation can make the information more memorable. >> >> Certainly I'm not able to remember details that don't make sense >> with my mental model of how computers work and I don't expect other >> people to either, so I think it's important to give an explanation that >> handles the biggest "objections". > > Hm, fair I guess. In any case, if we want to mention this I'd leave away > the details how exactly Git achieves this. E.g. we could say something > like: > > Storing a new blob for every new version of a file can result to a > lot of duplication. Git regularly runs repository maintenance to > optimize to counteract this. Part of the maintenance involves > compression of objects, where incremental changes to the same object > are optimized to be stored as deltas, only. > > We skip over the details, but this should give enough pointers to an > interested reader to go dig deeper. We could also generalize this to > objects in general, not only blobs.
Interesting. It is of course not wrong at all, but it was not what I would have expected for the first explanation to help confused folks who say "commits cannot be snapshots as they take too much space".
To me, it was a realization that even in a project whose tree (think of "du -s .") is huge, each of its commits touches only a handful of paths, hence a large portion of that huge tree would be shared with the previous snapshot.
Show 7 quoted lines
>> > This misses "refs/remotes/<remote>/HEAD". This reference is a symbolic >> > reference that indicates the default branch on the remote side. >> >> Is "refs/remotes/<remote>/HEAD" a remote-tracking branch? >> I've never thought about that reference and I'm not sure what to call it. > > No, it's not. I think the term we use is "remote reference".
Honestly I didn't know/think we have any special terminology for the refs/remotes/*/HEAD symref.
Historically HEAD did not "track" the remote state, and we did take advantage of that fact to use it as a place to record the preference with respect to which remote-tracking branch we would want to primarily interact with.
But these days because the protocol is capable of expressing where the symrefs point at, the users can make it track just like all other refs inside refs/remotes/*/ hiearchy. So I personally think it is OK to call it in remote-tracking branch.
Thanks.