threads / discuss / 64180

Could Git be smarter about object reuse?

Subject: Could Git be smarter about object reuse?

## tl;dr

8 messages between Sep 22, 2025 and Oct 3, 2025.

replies: 7people: 3as markdown or json

Sainan· Sep 22, 2025, 10:06 UTC · lore
Hello, I'm not entirely sure about the details of pushing, but I've noticed that basically it just uploads every object the server might need in relationship to that commit which can be a huge amount of data and sometimes even exceed the server-defined request timeout, causing the push to fail.
It's especially annoying because I know the server already has basically all the blobs needed and hence would only need to receive the commit and tree objects.
Are there any hidden flags or future cosiderations that could be made to reduce the bandwidth needed for such pushes?
-- Sainan
Simon Richter· Sep 22, 2025, 10:41 UTC · re: Sainan · lore

Re: Could Git be smarter about object reuse?

Hi,
On 9/22/25 7:06 PM, Sainan wrote:
> It's especially annoying because I know the server already has basically all the blobs needed and hence would only need to receive the commit and tree objects.

Git already does this. The receiver sends a list of commits it has, and the sender omits all objects (of any kind) that are reachable from any of these.

For this to work, the sender needs to be able to follow the commits from these references, so this does not work properly if the sender is operating from a shallow clone, or is missing branches, because the receiver only sends a list of branch tips (and, if the receiver is shallow, missing commits), not a full list of objects present, because that would be a lot.

    Simon
Sainan· Sep 22, 2025, 10:49 UTC · re: Simon Richter · lore

Re: Could Git be smarter about object reuse?

> The receiver sends a list of commits it has
This alone is not enough because if I'm amending a commit, it doesn't have the new commit(s), but it does have the previous commit(s), so the fact of blobs/trees being reusable is missed.
> For this to work, the sender needs to be able to follow the commits from these references
I think the issue here is more fundamental than this, because even just taking commits from one branch to the other is causing all objects to be resent even if again only the commit objects are new, not any attached trees or blobs.
Of course, it gets even more complicated when there's forks involved which are usually stored in the same repo by the server, but the client might not be entirely aware that different remotes are related. But that's a different issue that I could work around if that were all that were missing.
-- Sainan
Jeff King· Sep 22, 2025, 20:05 UTC · re: Sainan · lore

Re: Could Git be smarter about object reuse?

On Mon, Sep 22, 2025 at 10:49:07AM +0000, Sainan wrote:
Show 5 quoted lines
> > The receiver sends a list of commits it has
> 
> This alone is not enough because if I'm amending a commit, it doesn't
> have the new commit(s), but it does have the previous commit(s), so
> the fact of blobs/trees being reusable is missed.

Pushing doesn't dig into every possible blob/tree within each commit to look for duplicates. Doing that is very expensive in the most general case (you'd have to walk the entire object graph to check if some old commit mentions the blob you are about to send). So there are some heuristics about how much to dig.

We can simulate this case in a single repo like this:
  git init
  # or any big file; we want it to be obvious when it is sent
  dd if=/dev/urandom bs=1M count=10 >rand.bin
  git add rand.bin
  # now make one commit
  git commit -m one
  one=$(git rev-parse HEAD)
  # and an amended one with the same tree
  git commit --amend -m two
  two=$(git rev-parse HEAD)

If we pushed $one to a server, and then tried to push $two the server will tell us it has $one already. And push will feed this to pack-objects:

  echo ^$one >input
  echo $two >>input
And now we can run that same pack-objects locally to see the output:
  $ git pack-objects --stdout --revs --thin --no-progress <input | wc -c
  10489164

So that demonstrates the issue. Interestingly, we used to suppress the duplicate long ago. If I use Git v2.0.5, for example, we send only 147 bytes. Bisecting turns up the culprit as 2dacf26d09 (pack-objects: use --objects-edge-aggressive for shallow repos, 2014-12-24). The subject is a bit misleading there. It is enabling the "aggressive" form _only_ for shallow repos, whereas it had been used for both before that.

And the reasoning there is better explained by 1684c1b219 (rev-list: add an option to mark fewer edges as uninteresting, 2014-12-24), which says:

    In commit fbd4a70 (list-objects: mark more commits as edges in
    mark_edges_uninteresting - 2013-08-16), we marked an increasing number
    of edges uninteresting.  This change, and the subsequent change to make
    this conditional on --objects-edge, are used by --thin to make much
    smaller packs for shallow clones.
    Unfortunately, they cause a significant performance regression when
    pushing non-shallow clones with lots of refs (23.322 seconds vs.
    4.785 seconds with 22400 refs).  Add an option to git rev-list,
    --objects-edge-aggressive, that preserves this more aggressive behavior,
    while leaving --objects-edge to provide more performant behavior.
    Preserve the current behavior for the moment by using the aggressive
    option.

Under the hood this is being handled by calls to rev-list. So we could see the objects more directly like this:

  # this shows the blob; we are not doing any edge reporting at all
  git rev-list --objects ^$one $two
  # this is what pack-objects does by default; it also shows the blob
  git rev-list --objects-edge ^$one $two
  # and this is the more aggressive form that does suppress the blob
  git rev-list --objects-edge-aggressive ^$one $two
So I think there are a few things to ponder here:
  1. Possibly our heuristics could be smarter.
     This case is easy because it's the tree of a commit we know the
     other side has. We could detect it without digging into any trees
     by just marking the tree pointer of each uninteresting commit as
     also uninteresting. I'm actually a little surprised we don't do
     that already.
     But there are more complex --amend cases, too. E.g., you might have
     changed a nearby file, and the trees would be different (but the
     blob may still be unchanged). To detect that we'd have to walk the
     whole tree of the commit that the other side claims not to have.
     And I suspect that's what --object-edge-aggressive is doing, and
     why it would be expensive if the other side has a lot of refs.
     But possibly we could be do the aggressive thing on just the tip of
     a server-side ref when we are force-pushing over it. That would
     help with amends, rebases, and so forth.
  2. It would be nice if there was a knob for the user to turn, so they
     can spend more CPU time to find duplicates that might make the push
     smaller. There is a knob for rev-list, as shown above. But I don't
     think you can control how pack-objects behaves (aside from lying
     to it by passing --shallow), nor can you convince git-push itself
     to trigger pack-objects with specific options. But you could
     imagine a config option that would you do:
       git -c pack.aggressiveEdges=true push ...
     or something. It might be reasonable to turn on all the time in
     repos with few refs, or you could do a one-off like the command
     above if you saw that a push was going to be big.

And finally, there is one more trick up our sleeve: reachability bitmaps. The idea there is that we store bitmaps of which objects are reachable from which commit, which lets us answer object-graph questions quickly. And in particular it lets us produce a full set difference between the reachable objects in two commits.

So doing:
  git repack -adb

before running pack-objects (or git-push) will also produce the desired pack. The downside is that generating bitmaps is relatively expensive (much more CPU than the push would have used in the first place). In theory the results can then be amortized across many pushes, but the tradeoff isn't always great for a local repository which mostly packs to push (it's much better on a server that will serve many clones and fetches).

-Peff
Sainan· Sep 22, 2025, 22:09 UTC · re: Jeff King · lore

Re: Could Git be smarter about object reuse?

> git repack -adb
Certainly got my fans spinning for a bit. :)
But I can indeed confirm that it does solve the issue at least when amending a commit (will need to do further testing and generally get a feel for it).
However, one issue that I'm immediately noticing is that if I do 'git pull', I get a new pack that is bitmap-less, once again likely exposing me to the same problems.
Repacking before pushing would be okay for me as long as it guarantees a successful push, but it also seems a bit inefficient?
-- Sainan
Jeff King· Sep 23, 2025, 00:54 UTC · re: Sainan · lore

Re: Could Git be smarter about object reuse?

On Mon, Sep 22, 2025 at 10:09:41PM +0000, Sainan wrote:
Show 11 quoted lines
> > git repack -adb
> 
> Certainly got my fans spinning for a bit. :)
> 
> But I can indeed confirm that it does solve the issue at least when
> amending a commit (will need to do further testing and generally get a
> feel for it).
> 
> However, one issue that I'm immediately noticing is that if I do 'git
> pull', I get a new pack that is bitmap-less, once again likely
> exposing me to the same problems.

It should still work. The bitmap format is meant to degrade progressively. So if if you have a history like this:

   ROOT--...--A--B--C
               \
		D

and we have a bitmap for commit "B", then asking about "C" will let us traverse backwards until we hit "B", when we can fill in everything down to the root from the bitmap. We just have to traverse C's tree (but we can even avoid going into subtrees that are already mentioned in the bitmap).

For D it's a little trickier. We can't use the bitmap for B, but the idea is that we sprinkle them throughout history so that we'll eventually hit one and stop traversing.

So the big thing for making your case work is deciding whether we should use bitmaps at all. And if we have some already, then generally Git will try to use them, even if it means doing some fill-in traversal (because we really don't know how much fill-in traversal there will be ahead of time).

-Peff
Sainan· Oct 3, 2025, 11:39 UTC · re: Jeff King · lore

Re: Could Git be smarter about object reuse?

Yeah, I have to be honest, I'm not sure 'git repack -abd' is such a magic solution. I have a branch that diverged from main like 1000 commits ago and added roughly 30 commits of its own (they're all small commits, at most 1 blob), and yet even with a bitmap it's pushing over 475000 objects when my diverged branch could be expressed in about 100.
-- Sainan
Sainan· Oct 3, 2025, 11:51 UTC · re: Sainan · lore

Re: Could Git be smarter about object reuse?

The push didn't fail! git fetch on another clone gives me:

remote: Enumerating objects: 46, done. remote: Counting objects: 100% (46/46), done. remote: Compressing objects: 100% (16/16), done. remote: Total 40 (delta 30), reused 34 (delta 24), pack-reused 0 (from 0) Unpacking objects: 100% (40/40), 4.29 KiB | 28.00 KiB/s, done.

-- Sainan

← back to recent threads