{"thread":{"id":"64180","subject":"Could Git be smarter about object reuse?","startedAt":"2025-09-22T10:06:46Z","lastAt":"2025-10-03T11:52:07Z","messageCount":8,"participants":["Sainan","Simon Richter","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"526919","messageId":"pmKix6R7b3WVLrcK6ig1Lh7RhrB5G4Hm5yam_fEoC839aatB-OjJEmSJJ-weErGEnt4Mvgf5slxgu6Pm1xlGZ4mr_i4MIAAEMYy8DjJnWgk=@calamity.inc","threadId":"64180","inReplyTo":null,"subject":"Could Git be smarter about object reuse?","fromName":"Sainan","fromEmail":"sainan@calamity.inc","sentAt":"2025-09-22T10:06:31Z","receivedAt":"2025-09-22T10:06:46Z","isPatch":false,"sender":{"key":"sainan@calamity.inc","avatar":null},"body":"Hello, I'm not entirely sure about the details of pushing, but I've noticed that basically it just uploads every object the server might need in relationship to that commit which can be a huge amount of data and sometimes even exceed the server-defined request timeout, causing the push to fail.\n\nIt's especially annoying because I know the server already has basically all the blobs needed and hence would only need to receive the commit and tree objects.\n\nAre there any hidden flags or future cosiderations that could be made to reduce the bandwidth needed for such pushes?\n\n-- Sainan\n"},{"id":"526921","messageId":"2RWL_muy24EPDZ9wWFx-WZfu4Br_F2LenvcVJbKewfSVYipYM3qmeEIgV-6o4EbL39ZjMXtLHbVFOCPcBdHHVAU-0BrgBtuQ9BdRjS_2niE=@calamity.inc","threadId":"64180","inReplyTo":"f478fc6f-77ab-4d4e-a8d9-2d44622ba8dd@hogyros.de","subject":"Re: Could Git be smarter about object reuse?","fromName":"Sainan","fromEmail":"sainan@calamity.inc","sentAt":"2025-09-22T10:49:07Z","receivedAt":"2025-09-22T10:49:19Z","isPatch":false,"sender":{"key":"sainan@calamity.inc","avatar":null},"body":"> The receiver sends a list of commits it has\n\nThis alone is not enough because if I'm amending a commit, it doesn't have the new commit(s), but it does have the previous commit(s), so the fact of blobs/trees being reusable is missed.\n\n> For this to work, the sender needs to be able to follow the commits from these references\n\nI think the issue here is more fundamental than this, because even just taking commits from one branch to the other is causing all objects to be resent even if again only the commit objects are new, not any attached trees or blobs.\n\nOf course, it gets even more complicated when there's forks involved which are usually stored in the same repo by the server, but the client might not be entirely aware that different remotes are related. But that's a different issue that I could work around if that were all that were missing.\n\n-- Sainan\n"},{"id":"526922","messageId":"f478fc6f-77ab-4d4e-a8d9-2d44622ba8dd@hogyros.de","threadId":"64180","inReplyTo":"pmKix6R7b3WVLrcK6ig1Lh7RhrB5G4Hm5yam_fEoC839aatB-OjJEmSJJ-weErGEnt4Mvgf5slxgu6Pm1xlGZ4mr_i4MIAAEMYy8DjJnWgk=@calamity.inc","subject":"Re: Could Git be smarter about object reuse?","fromName":"Simon Richter","fromEmail":"simon.richter@hogyros.de","sentAt":"2025-09-22T10:41:23Z","receivedAt":"2025-09-22T10:51:36Z","isPatch":false,"sender":{"key":"simon.richter@hogyros.de","avatar":"https://gravatar.com/avatar/1192aa9fa5dd19ce258b12b044cc27111dd7cdb58d920dc2123a02a24f55b5c5?d=mp&s=160"},"body":"Hi,\n\nOn 9/22/25 7:06 PM, Sainan wrote:\n\n> It's especially annoying because I know the server already has basically all the blobs needed and hence would only need to receive the commit and tree objects.\n\nGit already does this. The receiver sends a list of commits it has, and \nthe sender omits all objects (of any kind) that are reachable from any \nof these.\n\nFor this to work, the sender needs to be able to follow the commits from \nthese references, so this does not work properly if the sender is \noperating from a shallow clone, or is missing branches, because the \nreceiver only sends a list of branch tips (and, if the receiver is \nshallow, missing commits), not a full list of objects present, because \nthat would be a lot.\n\n    Simon\n"},{"id":"526980","messageId":"20250922200510.GC2205919@coredump.intra.peff.net","threadId":"64180","inReplyTo":"2RWL_muy24EPDZ9wWFx-WZfu4Br_F2LenvcVJbKewfSVYipYM3qmeEIgV-6o4EbL39ZjMXtLHbVFOCPcBdHHVAU-0BrgBtuQ9BdRjS_2niE=@calamity.inc","subject":"Re: Could Git be smarter about object reuse?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2025-09-22T20:05:10Z","receivedAt":"2025-09-22T20:05:11Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 22, 2025 at 10:49:07AM +0000, Sainan wrote:\n\n> > The receiver sends a list of commits it has\n> \n> This alone is not enough because if I'm amending a commit, it doesn't\n> have the new commit(s), but it does have the previous commit(s), so\n> the fact of blobs/trees being reusable is missed.\n\nPushing doesn't dig into every possible blob/tree within each commit to\nlook for duplicates. Doing that is very expensive in the most general\ncase (you'd have to walk the entire object graph to check if some old\ncommit mentions the blob you are about to send). So there are some\nheuristics about how much to dig.\n\nWe can simulate this case in a single repo like this:\n\n  git init\n  # or any big file; we want it to be obvious when it is sent\n  dd if=/dev/urandom bs=1M count=10 >rand.bin\n  git add rand.bin\n\n  # now make one commit\n  git commit -m one\n  one=$(git rev-parse HEAD)\n\n  # and an amended one with the same tree\n  git commit --amend -m two\n  two=$(git rev-parse HEAD)\n\nIf we pushed $one to a server, and then tried to push $two the server\nwill tell us it has $one already. And push will feed this to\npack-objects:\n\n  echo ^$one >input\n  echo $two >>input\n\nAnd now we can run that same pack-objects locally to see the output:\n\n  $ git pack-objects --stdout --revs --thin --no-progress <input | wc -c\n  10489164\n\nSo that demonstrates the issue. Interestingly, we used to suppress the\nduplicate long ago. If I use Git v2.0.5, for example, we send only 147\nbytes. Bisecting turns up the culprit as 2dacf26d09 (pack-objects: use\n--objects-edge-aggressive for shallow repos, 2014-12-24). The subject is\na bit misleading there. It is enabling the \"aggressive\" form _only_ for\nshallow repos, whereas it had been used for both before that. \n\nAnd the reasoning there is better explained by 1684c1b219 (rev-list: add\nan option to mark fewer edges as uninteresting, 2014-12-24), which says:\n\n    In commit fbd4a70 (list-objects: mark more commits as edges in\n    mark_edges_uninteresting - 2013-08-16), we marked an increasing number\n    of edges uninteresting.  This change, and the subsequent change to make\n    this conditional on --objects-edge, are used by --thin to make much\n    smaller packs for shallow clones.\n\n    Unfortunately, they cause a significant performance regression when\n    pushing non-shallow clones with lots of refs (23.322 seconds vs.\n    4.785 seconds with 22400 refs).  Add an option to git rev-list,\n    --objects-edge-aggressive, that preserves this more aggressive behavior,\n    while leaving --objects-edge to provide more performant behavior.\n    Preserve the current behavior for the moment by using the aggressive\n    option.\n\nUnder the hood this is being handled by calls to rev-list. So we could\nsee the objects more directly like this:\n\n  # this shows the blob; we are not doing any edge reporting at all\n  git rev-list --objects ^$one $two\n\n  # this is what pack-objects does by default; it also shows the blob\n  git rev-list --objects-edge ^$one $two\n\n  # and this is the more aggressive form that does suppress the blob\n  git rev-list --objects-edge-aggressive ^$one $two\n\nSo I think there are a few things to ponder here:\n\n  1. Possibly our heuristics could be smarter.\n\n     This case is easy because it's the tree of a commit we know the\n     other side has. We could detect it without digging into any trees\n     by just marking the tree pointer of each uninteresting commit as\n     also uninteresting. I'm actually a little surprised we don't do\n     that already.\n\n     But there are more complex --amend cases, too. E.g., you might have\n     changed a nearby file, and the trees would be different (but the\n     blob may still be unchanged). To detect that we'd have to walk the\n     whole tree of the commit that the other side claims not to have.\n     And I suspect that's what --object-edge-aggressive is doing, and\n     why it would be expensive if the other side has a lot of refs.\n\n     But possibly we could be do the aggressive thing on just the tip of\n     a server-side ref when we are force-pushing over it. That would\n     help with amends, rebases, and so forth.\n\n  2. It would be nice if there was a knob for the user to turn, so they\n     can spend more CPU time to find duplicates that might make the push\n     smaller. There is a knob for rev-list, as shown above. But I don't\n     think you can control how pack-objects behaves (aside from lying\n     to it by passing --shallow), nor can you convince git-push itself\n     to trigger pack-objects with specific options. But you could\n     imagine a config option that would you do:\n\n       git -c pack.aggressiveEdges=true push ...\n\n     or something. It might be reasonable to turn on all the time in\n     repos with few refs, or you could do a one-off like the command\n     above if you saw that a push was going to be big.\n\nAnd finally, there is one more trick up our sleeve: reachability\nbitmaps. The idea there is that we store bitmaps of which objects are\nreachable from which commit, which lets us answer object-graph questions\nquickly. And in particular it lets us produce a full set difference\nbetween the reachable objects in two commits.\n\nSo doing:\n\n  git repack -adb\n\nbefore running pack-objects (or git-push) will also produce the desired\npack. The downside is that generating bitmaps is relatively expensive\n(much more CPU than the push would have used in the first place). In\ntheory the results can then be amortized across many pushes, but the\ntradeoff isn't always great for a local repository which mostly packs to\npush (it's much better on a server that will serve many clones and\nfetches).\n\n-Peff\n"},{"id":"526998","messageId":"ZURUr5sfXi0wsjBeXiwAxyNgalVa2ZveXDgoTcexUNOAgcP_JscHvFFDIss4stpsiB2MzUQ_Z30tFrPSgr8W8V02ecfCj4BFFwQqWwJpba4=@calamity.inc","threadId":"64180","inReplyTo":"20250922200510.GC2205919@coredump.intra.peff.net","subject":"Re: Could Git be smarter about object reuse?","fromName":"Sainan","fromEmail":"sainan@calamity.inc","sentAt":"2025-09-22T22:09:41Z","receivedAt":"2025-09-22T22:10:01Z","isPatch":false,"sender":{"key":"sainan@calamity.inc","avatar":null},"body":"> git repack -adb\n\nCertainly got my fans spinning for a bit. :)\n\nBut I can indeed confirm that it does solve the issue at least when amending a commit (will need to do further testing and generally get a feel for it).\n\nHowever, one issue that I'm immediately noticing is that if I do 'git pull', I get a new pack that is bitmap-less, once again likely exposing me to the same problems.\n\nRepacking before pushing would be okay for me as long as it guarantees a successful push, but it also seems a bit inefficient?\n\n-- Sainan\n"},{"id":"527010","messageId":"20250923005421.GB2271307@coredump.intra.peff.net","threadId":"64180","inReplyTo":"ZURUr5sfXi0wsjBeXiwAxyNgalVa2ZveXDgoTcexUNOAgcP_JscHvFFDIss4stpsiB2MzUQ_Z30tFrPSgr8W8V02ecfCj4BFFwQqWwJpba4=@calamity.inc","subject":"Re: Could Git be smarter about object reuse?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2025-09-23T00:54:21Z","receivedAt":"2025-09-23T00:54:22Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 22, 2025 at 10:09:41PM +0000, Sainan wrote:\n\n> > git repack -adb\n> \n> Certainly got my fans spinning for a bit. :)\n> \n> But I can indeed confirm that it does solve the issue at least when\n> amending a commit (will need to do further testing and generally get a\n> feel for it).\n> \n> However, one issue that I'm immediately noticing is that if I do 'git\n> pull', I get a new pack that is bitmap-less, once again likely\n> exposing me to the same problems.\n\nIt should still work. The bitmap format is meant to degrade\nprogressively. So if if you have a history like this:\n\n   ROOT--...--A--B--C\n               \\\n\t\tD\n\nand we have a bitmap for commit \"B\", then asking about \"C\" will let us\ntraverse backwards until we hit \"B\", when we can fill in everything down\nto the root from the bitmap. We just have to traverse C's tree (but we\ncan even avoid going into subtrees that are already mentioned in the\nbitmap).\n\nFor D it's a little trickier. We can't use the bitmap for B, but the\nidea is that we sprinkle them throughout history so that we'll\neventually hit one and stop traversing.\n\nSo the big thing for making your case work is deciding whether we should\nuse bitmaps at all. And if we have some already, then generally Git will\ntry to use them, even if it means doing some fill-in traversal (because\nwe really don't know how much fill-in traversal there will be ahead of\ntime).\n\n-Peff\n"},{"id":"527874","messageId":"B0Y9iigwIf1VSJpVtY_IzINon0LTimi0sIg9B4j8pDJt2FoxHmQ-Gn5C0s0l-GhsHMP2ZptgNbm879BqQgpDOo-CFEOhh-nQqhlISosKoWY=@calamity.inc","threadId":"64180","inReplyTo":"20250923005421.GB2271307@coredump.intra.peff.net","subject":"Re: Could Git be smarter about object reuse?","fromName":"Sainan","fromEmail":"sainan@calamity.inc","sentAt":"2025-10-03T11:39:13Z","receivedAt":"2025-10-03T11:41:00Z","isPatch":false,"sender":{"key":"sainan@calamity.inc","avatar":null},"body":"Yeah, I have to be honest, I'm not sure 'git repack -abd' is such a magic solution. I have a branch that diverged from main like 1000 commits ago and added roughly 30 commits of its own (they're all small commits, at most 1 blob), and yet even with a bitmap it's pushing over 475000 objects when my diverged branch could be expressed in about 100.\n\n-- Sainan\n"},{"id":"527876","messageId":"mVMA1eOYhWQp-1-EXnXsp31DUwjoqszklcsfGiT15qy_QKmQWj6Z8PrDLMoIQeAvEvxrRhW3dXblWbtSKLuYycnaU5x_DM06A_XxWH_lWBk=@calamity.inc","threadId":"64180","inReplyTo":"B0Y9iigwIf1VSJpVtY_IzINon0LTimi0sIg9B4j8pDJt2FoxHmQ-Gn5C0s0l-GhsHMP2ZptgNbm879BqQgpDOo-CFEOhh-nQqhlISosKoWY=@calamity.inc","subject":"Re: Could Git be smarter about object reuse?","fromName":"Sainan","fromEmail":"sainan@calamity.inc","sentAt":"2025-10-03T11:51:53Z","receivedAt":"2025-10-03T11:52:07Z","isPatch":false,"sender":{"key":"sainan@calamity.inc","avatar":null},"body":"The push didn't fail! git fetch on another clone gives me:\n\nremote: Enumerating objects: 46, done.\nremote: Counting objects: 100% (46/46), done.\nremote: Compressing objects: 100% (16/16), done.\nremote: Total 40 (delta 30), reused 34 (delta 24), pack-reused 0 (from 0)\nUnpacking objects: 100% (40/40), 4.29 KiB | 28.00 KiB/s, done.\n\n-- Sainan\n"}]}