Volume XXII, number 279Tuesday, October 6, 2026Latest message 39 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

Persistent shallow + fake-linearizing a whole mainline

1 messages between Jul 14, 2026 and Jul 14, 2026, from Richard Fine.

Plain Markdown or JSON for tools and agents.

Richard FineJul 14, 2026, 20:04 UTC on lore
Hi,

The repository at my company uses a standard branch-and-pull-request model for developers to make changes. Pull requests are integrated as 2-parent merge commits. I'm trying to find ways to optimise working with the repository, particularly by reducing the Git database size to make git operations faster. I see two clear opportunities, though I'm struggling to get engineering alignment on each:

* Squash-merging. If we switched to squash-merging pull requests into
our mainline branch, developers wouldn't have to carry the added load
of the individual commits from the branches used to construct the PR.
However, two objections arise: first, landing a sequence of stacked
PRs becomes painful because Git can no longer accurately identify
merge-bases; and second, in-branch history is sometimes useful for
code archaeology. Suggestions that the in-branch history would still
be available in our source-of-truth repository are met with complaints
that looking at a second repo for history is inconvenient :)
* Shallow cloning. Setting a shallow boundary with something like
--shallow-since="3 months ago" gives us an intuitive way to trade off
repository size/speed against history availability. The problem with
this is that we sometimes have PRs which initially branched off old
revisions - earlier than our shallow-since point - and then get landed
with merge commits. When people pull those merge commits, git follows
the second parent of the merge, pulls the history of the branch, and
ends up bypassing the 'firebreak' revisions defined in `.git/shallow`,
pulling large amounts of history. People can avoid this by specifying
the --shallow options when running `git fetch`, but very often, this
is not something they are running manually: UI tools are running it
for them, or they run `git remote update`, or an AI agent is doing it
for them, etc. Suggestions that we should block people from landing
PRs with branches based on ancient revisions, and that people should
instead rebase the work on a newer revision, are met with the
objection that if the branch is based on that old of a revision it's
typically because it's a long-running branch which accumulated a lot
of work, and rebasing that work on a more recent revision of mainline
is painful.

I've not yet given up trying to get my colleagues to change their workflows (and I welcome advice on how others approach these engineering-culture problems). In the meantime, I have a couple of thoughts for possible Git improvements that might help, which I figured I'd raise here.

* The biggest issue with the shallow clone solution is the possibility
that someone fetches one of these 'based on ancient history' merges
without passing --shallow-since, causing Git to end up pulling huge
amounts of history. What if one could set the shallow options
persistently? For example, a "remote.origin.shallowsince" in the
.git/config. If set, it would make fetches from that remote behave as
if --shallow-since was specified on the command-line, regardless of
how the fetch was triggered.
* This is more complicated, but... I did wonder if there is some way
to use .git/shallow (or something similar, like replace refs) to make
Git pretend that the merge commits on our mainline are actually linear
commits (i.e. pretend they only have their first parent). Then when a
developer actually wants to delve into the commits that made up a
specific branch, they could make Git stop pretending for that specific
merge commit, and Git would then fetch the commits needed to fill in
the missing second parent and its ancestors. When they're done, they
flip it back to being a fake linear commit, and the branch commits
would no longer be reachable, eventually being cleaned up by `git gc`.
I think I could probably write a script to convert an existing
repository into this 'fake linear mainline' mode, but I'm not sure how
I'd then make incremental fetches of the mainline continue to keep up
the masquerade without pulling all the commits first and then running
the script. I'd like to avoid pulling the extra commits if possible.

What do you think? I could probably take on at least the first idea if there is interest in it.

- Richard

Back to recent threads