{"thread":{"id":"65997","subject":"Persistent shallow + fake-linearizing a whole mainline","startedAt":"2026-07-14T20:05:10Z","lastAt":"2026-07-14T20:05:10Z","messageCount":1,"participants":["Richard Fine"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"548163","messageId":"CAAFFBtiOBsbHuKh0J2ytUXOB1PpzMhT6LvyjGozH5g5ygYZjHw@mail.gmail.com","threadId":"65997","inReplyTo":null,"subject":"Persistent shallow + fake-linearizing a whole mainline","fromName":"Richard Fine","fromEmail":"richardf@unity3d.com","sentAt":"2026-07-14T20:04:58Z","receivedAt":"2026-07-14T20:05:10Z","isPatch":false,"body":"Hi,\n\nThe repository at my company uses a standard branch-and-pull-request\nmodel for developers to make changes. Pull requests are integrated as\n2-parent merge commits. I'm trying to find ways to optimise working\nwith the repository, particularly by reducing the Git database size to\nmake git operations faster. I see two clear opportunities, though I'm\nstruggling to get engineering alignment on each:\n\n* Squash-merging. If we switched to squash-merging pull requests into\nour mainline branch, developers wouldn't have to carry the added load\nof the individual commits from the branches used to construct the PR.\nHowever, two objections arise: first, landing a sequence of stacked\nPRs becomes painful because Git can no longer accurately identify\nmerge-bases; and second, in-branch history is sometimes useful for\ncode archaeology. Suggestions that the in-branch history would still\nbe available in our source-of-truth repository are met with complaints\nthat looking at a second repo for history is inconvenient :)\n\n* Shallow cloning. Setting a shallow boundary with something like\n--shallow-since=\"3 months ago\" gives us an intuitive way to trade off\nrepository size/speed against history availability. The problem with\nthis is that we sometimes have PRs which initially branched off old\nrevisions - earlier than our shallow-since point - and then get landed\nwith merge commits. When people pull those merge commits, git follows\nthe second parent of the merge, pulls the history of the branch, and\nends up bypassing the 'firebreak' revisions defined in `.git/shallow`,\npulling large amounts of history. People can avoid this by specifying\nthe --shallow options when running `git fetch`, but very often, this\nis not something they are running manually: UI tools are running it\nfor them, or they run `git remote update`, or an AI agent is doing it\nfor them, etc. Suggestions that we should block people from landing\nPRs with branches based on ancient revisions, and that people should\ninstead rebase the work on a newer revision, are met with the\nobjection that if the branch is based on that old of a revision it's\ntypically because it's a long-running branch which accumulated a lot\nof work, and rebasing that work on a more recent revision of mainline\nis painful.\n\nI've not yet given up trying to get my colleagues to change their\nworkflows (and I welcome advice on how others approach these\nengineering-culture problems). In the meantime, I have a couple of\nthoughts for possible Git improvements that might help, which I\nfigured I'd raise here.\n\n* The biggest issue with the shallow clone solution is the possibility\nthat someone fetches one of these 'based on ancient history' merges\nwithout passing --shallow-since, causing Git to end up pulling huge\namounts of history. What if one could set the shallow options\npersistently? For example, a \"remote.origin.shallowsince\" in the\n.git/config. If set, it would make fetches from that remote behave as\nif --shallow-since was specified on the command-line, regardless of\nhow the fetch was triggered.\n\n* This is more complicated, but... I did wonder if there is some way\nto use .git/shallow (or something similar, like replace refs) to make\nGit pretend that the merge commits on our mainline are actually linear\ncommits (i.e. pretend they only have their first parent). Then when a\ndeveloper actually wants to delve into the commits that made up a\nspecific branch, they could make Git stop pretending for that specific\nmerge commit, and Git would then fetch the commits needed to fill in\nthe missing second parent and its ancestors. When they're done, they\nflip it back to being a fake linear commit, and the branch commits\nwould no longer be reachable, eventually being cleaned up by `git gc`.\nI think I could probably write a script to convert an existing\nrepository into this 'fake linear mainline' mode, but I'm not sure how\nI'd then make incremental fetches of the mainline continue to keep up\nthe masquerade without pulling all the commits first and then running\nthe script. I'd like to avoid pulling the extra commits if possible.\n\nWhat do you think? I could probably take on at least the first idea if\nthere is interest in it.\n\n- Richard\n"}]}