{"thread":{"id":"64801","subject":"[BUG] Git push sends too much data unnecessarily","startedAt":"2026-01-14T12:41:23Z","lastAt":"2026-01-15T09:43:27Z","messageCount":7,"participants":["Rajiv Sharma","Karthik Nayak","Junio C Hamano","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"533821","messageId":"CAGe2LO0nxXuNNRYS0fk0JuPBDa3UCT8EDJ6G1u4GNW1d9rzRgA@mail.gmail.com","threadId":"64801","inReplyTo":null,"subject":"[BUG] Git push sends too much data unnecessarily","fromName":"Rajiv Sharma","fromEmail":"rajiv.tilakraj.sharma@gmail.com","sentAt":"2026-01-14T12:41:07Z","receivedAt":"2026-01-14T12:41:23Z","isPatch":false,"sender":{"key":"rajiv.tilakraj.sharma@gmail.com","avatar":"https://gravatar.com/avatar/37a35fa524c1baf12341e58bd7688e26ee84ee0a1e44c81d100f671dcdb51654?d=mp&s=160"},"body":"Thank you for filling out a Git bug report!\nPlease answer the following questions to help us understand your issue.\n\nWhat did you do before the bug happened? (Steps to reproduce your issue)\n\nI tried to create a new branch pointing to the commit which was the\nancestor of the current branch (i.e. HEAD~1) and pushing it to the\nremote. Since the commit was already known to the server, I expected\nthe push to be kind of no-op since it's simply creating a new pointer.\nHowever the push ended up taking 10+ minutes. Since I was running with\nthe `--verbose` flag, I realised that the push ended up sending\nmultiple GBs worth of data just for creating a new branch on an\nexisting commit already known to the remote. After some\nexperimentation, I managed to find an easy repro for this issue:\n\nClone a non-empty repo from some remote (e.g. git clone\nhttps://SERVER_HOSTNAME/repo_name.git) in two locations, `primary` and\n`secondary` and ensure that both have the same branch checked out.\nNavigate to the `primary` location and create a local commit for repo\n`repo_name`. Push this commit C1 to the remote server\nNavigate to the `secondary` location and try to create a new branch by\nrunning `git push origin HEAD:refs/heads/shiny_new_branch --verbose`\n(or by checking out that branch and pushing it). Note that `HEAD` here\nrefers to the `HEAD` commit as seen by `secondary` which in reality is\n`HEAD~1` compared to the remote\nIf the repo had some commits on the checked out branch, you will\nnotice the verbose output highlighting objects being sent to the\nserver where there was no need to do so\n\n\nTo understand more about exactly how much data is sent, I ran a few\nmore experiments and came to the conclusion that the git client sends\nHEAD commit + all ancestors of HEAD commit except the commits which\nare also ancestors of some other branch / ref known to Git.\nPictorially, it can be represented as:\n\nB1  B2       <-- HEAD\n*      *         (sent)\n|       |\n*       *         (sent)\n|        |\n*        *        (sent)\n|      /\n|    /\n*                  (NOT sent)\n|\n*                  (NOT sent)\n\nThis explains the multi GB push in my case because I was working on a\nlong standing branch with lots of commits. Initially I assumed this\nwas a server problem but then realised that in the push path the\nserver just advertises refs and where they point and it's the client\nthat does the negotiation. I think the bug exists somewhere in the\nnegotiation logic but I am not sure.\n\nWhat did you expect to happen? (Expected behavior)\n\nI would have expected the push to be extremely lightweight without\nsending any objects to the server.\n\n\nWhat happened instead? (Actual behavior)\n\nAlready detailed in the first section above.\n\n\nWhat's different between what you expected and what actually happened?\n\nThe git client sends loads of data to the server when it shouldn't\nhave had to send anything at all.\n\n\nAnything else you want to add:\n\nNote that there are workarounds for this problem. If I do a `git pull`\nand get the latest state of the repo before performing any push, this\nproblem doesn't occur. Nevertheless, I think it might be worthwhile to\nfix this. I managed to repro this across OS (Linux, MacOS) and across\nversions.\n\n\n\n[System Info]\ngit version:\ngit version 2.47.3\ncpu: x86_64\nno commit associated with this build\nsizeof-long: 8\nsizeof-size_t: 8\nshell-path: /bin/sh\nlibcurl: 7.76.1\nOpenSSL: OpenSSL 3.5.1 1 Jul 2025\nzlib: 1.2.11\nuname: Linux 6.9.0-0_fbk12_0_g28f2d09ad102 #1 SMP Thu Nov  6 08:05:52\nPST 2025 x86_64\ncompiler info: gnuc: 11.5\nlibc info: glibc: 2.34\n$SHELL (typically, interactive shell): /bin/bash\n\n\n[Enabled Hooks]\n"},{"id":"533838","messageId":"CAOLa=ZT4fQdHqG+1AeviYuLUR5VG33voJk_DU1y0MzhUKBQvvw@mail.gmail.com","threadId":"64801","inReplyTo":"CAGe2LO0nxXuNNRYS0fk0JuPBDa3UCT8EDJ6G1u4GNW1d9rzRgA@mail.gmail.com","subject":"Re: [BUG] Git push sends too much data unnecessarily","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-01-14T16:27:21Z","receivedAt":"2026-01-14T16:27:24Z","isPatch":false,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Rajiv Sharma <rajiv.tilakraj.sharma@gmail.com> writes:\n\n> Thank you for filling out a Git bug report!\n> Please answer the following questions to help us understand your issue.\n>\n> What did you do before the bug happened? (Steps to reproduce your issue)\n>\n> I tried to create a new branch pointing to the commit which was the\n> ancestor of the current branch (i.e. HEAD~1) and pushing it to the\n> remote. Since the commit was already known to the server, I expected\n> the push to be kind of no-op since it's simply creating a new pointer.\n> However the push ended up taking 10+ minutes. Since I was running with\n> the `--verbose` flag, I realised that the push ended up sending\n> multiple GBs worth of data just for creating a new branch on an\n> existing commit already known to the remote. After some\n> experimentation, I managed to find an easy repro for this issue:\n>\n> Clone a non-empty repo from some remote (e.g. git clone\n> https://SERVER_HOSTNAME/repo_name.git) in two locations, `primary` and\n> `secondary` and ensure that both have the same branch checked out.\n> Navigate to the `primary` location and create a local commit for repo\n> `repo_name`. Push this commit C1 to the remote server\n> Navigate to the `secondary` location and try to create a new branch by\n> running `git push origin HEAD:refs/heads/shiny_new_branch --verbose`\n> (or by checking out that branch and pushing it). Note that `HEAD` here\n> refers to the `HEAD` commit as seen by `secondary` which in reality is\n> `HEAD~1` compared to the remote\n> If the repo had some commits on the checked out branch, you will\n> notice the verbose output highlighting objects being sent to the\n> server where there was no need to do so\n>\n>\n> To understand more about exactly how much data is sent, I ran a few\n> more experiments and came to the conclusion that the git client sends\n> HEAD commit + all ancestors of HEAD commit except the commits which\n> are also ancestors of some other branch / ref known to Git.\n> Pictorially, it can be represented as:\n>\n> B1  B2       <-- HEAD\n> *      *         (sent)\n> |       |\n> *       *         (sent)\n> |        |\n> *        *        (sent)\n> |      /\n> |    /\n> *                  (NOT sent)\n> |\n> *                  (NOT sent)\n>\n> This explains the multi GB push in my case because I was working on a\n> long standing branch with lots of commits. Initially I assumed this\n> was a server problem but then realised that in the push path the\n> server just advertises refs and where they point and it's the client\n> that does the negotiation. I think the bug exists somewhere in the\n> negotiation logic but I am not sure.\n>\n\nThanks for the detailed explanation. I don't think this is a bug per-se,\nbut that doesn't mean this isn't something we can't discuss and\npotentiall optimize\n\nTo reiterate my understanding, I did a quick local PoC:\n\n$ git init remote\n$ git -C remote config set receive.denyCurrentBranch ignore\n$ git -C remote commit --allow-empty -m \"C1\"\n$ git -C remote commit --allow-empty -m \"C2\"\n$ git -C remote commit --allow-empty -m \"C3\"\n\n$ git clone remote/ base1\n$ git clone remote/ base2\n\n$ git -C base1 commit --allow-empty -m \"C4\"\n$ git -C base1 push -f --verbose\nPushing to /tmp/remote/\nEnumerating objects: 1, done.\nCounting objects: 100% (1/1), done.\nWriting objects: 100% (1/1), 704 bytes | 704.00 KiB/s, done.\nTotal 1 (delta 0), reused 0 (delta 0), pack-reused 0 (from 0)\nTo /tmp/remote/\n   78c400c..affbad8  master -> master\nupdating local tracking ref 'refs/remotes/origin/master'\n\n$ git -C base2 push -f --verbose origin HEAD:refs/heads/fun\nPushing to /tmp/remote/\nEnumerating objects: 4, done.\nCounting objects: 100% (4/4), done.\nDelta compression using up to 16 threads\nCompressing objects: 100% (3/3), done.\nWriting objects: 100% (4/4), 1.98 KiB | 1.98 MiB/s, done.\nTotal 4 (delta 0), reused 0 (delta 0), pack-reused 0 (from 0)\nTo /tmp/remote/\n * [new branch]      HEAD -> fun\nupdating local tracking ref 'refs/remotes/origin/fun'\n\nWhat you're stating about and can be easily seen here is that while\npushing C4 from base1 only transferred one object, pushing HEAD from\nbase2 (which is C4~1), pushes 4 objects.\n\nAfter base1 creates C4 and pushes:\n==================================\nremote:     C1 --- C2 --- C3 --- C4 (master)\n\nbase1:      C1 --- C2 --- C3 --- C4 (master, origin/master)\n                                  ^\n                                  |\n                         (transfers only C4)\n\nbase2:      C1 --- C2 --- C3 (master, origin/master)\n\n\nWhen base2 pushes HEAD (=C3) to refs/heads/fun:\n================================================\nremote:     C1 --- C2 --- C3 --- C4 (master)\n                            \\\n                             fun\n\nbase2:      C1 --- C2 --- C3 (master, origin/master)\n                        ^\n                        |\n              (transfers C1, C2, C3, + tree object)\n              (4 objects total)\n\nThis boils down to how Git negotiates between the client <> server.\nIn our case, remote will list the references it already contains. So in\nour experiment, that'd be:\n\n - C4: affbad8\n\nWith this information, the client should find all the objects the remote\nwould need to satisfy the new references being pushed.\n\nSince C4 is a reference the client (base2) knows nothing about, it\ncannot find a common ancestor between the provided commit vs all commits\npresent within the repository itself. This is seems obvious to us, since\nC4~1 is the common ancestor here, but base2 doesn't have sufficient\ninformation to come to that conclusion.\n\nSo it sends all objects required to create the reference, in our case 4\nobjects, in your case GBs of data.\n\n> What did you expect to happen? (Expected behavior)\n>\n> I would have expected the push to be extremely lightweight without\n> sending any objects to the server.\n>\n>\n> What happened instead? (Actual behavior)\n>\n> Already detailed in the first section above.\n>\n>\n> What's different between what you expected and what actually happened?\n>\n> The git client sends loads of data to the server when it shouldn't\n> have had to send anything at all.\n>\n>\n> Anything else you want to add:\n>\n> Note that there are workarounds for this problem. If I do a `git pull`\n> and get the latest state of the repo before performing any push, this\n> problem doesn't occur. Nevertheless, I think it might be worthwhile to\n> fix this. I managed to repro this across OS (Linux, MacOS) and across\n> versions.\n>\n\nThat said, I do think we can potentially optimize this, AFAIK the\nnegotiation phase has the server listing its refs and this is compared\nto the list of refs locally present to determine all missing objects.\n\nSo any commits which are not represented by a ref, would be missed. One\nway to reduce this would be for the server to also provide additional\ninformation such as commits which are not represented by any refs. But\nhow many such commits? What about sampling? Finally we'd have to\nconsider if it is worth it.\n\nThanks,\nKarthik\n"},{"id":"533850","messageId":"xmqqh5sof61i.fsf@gitster.g","threadId":"64801","inReplyTo":"CAOLa=ZT4fQdHqG+1AeviYuLUR5VG33voJk_DU1y0MzhUKBQvvw@mail.gmail.com","subject":"Re: [BUG] Git push sends too much data unnecessarily","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-01-14T17:38:49Z","receivedAt":"2026-01-14T17:38:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Karthik Nayak <karthik.188@gmail.com> writes:\n\n> So it sends all objects required to create the reference, in our case 4\n> objects, in your case GBs of data.\n\n\"push.negotiate\"?\n"},{"id":"533851","messageId":"CAGe2LO3t3B1g1ARH-LQ9V0UoGmToO-Z9XYpeMOTKkaSQvCpaRA@mail.gmail.com","threadId":"64801","inReplyTo":"xmqqh5sof61i.fsf@gitster.g","subject":"Re: [BUG] Git push sends too much data unnecessarily","fromName":"Rajiv Sharma","fromEmail":"rajiv.tilakraj.sharma@gmail.com","sentAt":"2026-01-14T17:39:43Z","receivedAt":"2026-01-14T17:39:58Z","isPatch":false,"sender":{"key":"rajiv.tilakraj.sharma@gmail.com","avatar":"https://gravatar.com/avatar/37a35fa524c1baf12341e58bd7688e26ee84ee0a1e44c81d100f671dcdb51654?d=mp&s=160"},"body":"Thanks for the great explanation! You are right, it's not really a bug\n(because there is no correctness problem here) but it surely is\nsuboptimal behavior.\n\n> This boils down to how Git negotiates between the client <> server\n\nI think that's the crux of the problem here. I don't think git\nnegotiates in the push path the way it does in the read path, i.e.\nthere is no process of client-server communication that involves\ngradually arriving at the common base (in this case it would be C3).\nThe read path does this quite well (using something akin to a skiplist\nIIRC?) and the common base is found in a couple iterations in most\ncases. I am unaware of the historical context behind this difference\nbut I assume the server sending unnecessary extra data during the read\npath would be much more expensive than the client doing it hence the\npush protocol is kept simpler.\n\nThis kind of negotiation _could_ be added to the push path but it\nwould be a breaking change. I read somewhere that there were plans for\nPush Protocol V2 (in the same vein as Read Protocol V2) so it would be\ngreat to see this improvement making its way there!\n\nThanks\nRajiv Sharma\n\nOn Wed, Jan 14, 2026 at 5:38 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Karthik Nayak <karthik.188@gmail.com> writes:\n>\n> > So it sends all objects required to create the reference, in our case 4\n> > objects, in your case GBs of data.\n>\n> \"push.negotiate\"?\n"},{"id":"533890","messageId":"20260114211115.GC1008851@coredump.intra.peff.net","threadId":"64801","inReplyTo":"CAGe2LO3t3B1g1ARH-LQ9V0UoGmToO-Z9XYpeMOTKkaSQvCpaRA@mail.gmail.com","subject":"Re: [BUG] Git push sends too much data unnecessarily","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-01-14T21:11:15Z","receivedAt":"2026-01-14T21:11:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jan 14, 2026 at 05:39:43PM +0000, Rajiv Sharma wrote:\n\n> > This boils down to how Git negotiates between the client <> server\n> \n> I think that's the crux of the problem here. I don't think git\n> negotiates in the push path the way it does in the read path, i.e.\n> there is no process of client-server communication that involves\n> gradually arriving at the common base (in this case it would be C3).\n> The read path does this quite well (using something akin to a skiplist\n> IIRC?) and the common base is found in a couple iterations in most\n> cases. I am unaware of the historical context behind this difference\n> but I assume the server sending unnecessary extra data during the read\n> path would be much more expensive than the client doing it hence the\n> push protocol is kept simpler.\n> \n> This kind of negotiation _could_ be added to the push path but it\n> would be a breaking change. I read somewhere that there were plans for\n> Push Protocol V2 (in the same vein as Read Protocol V2) so it would be\n> great to see this improvement making its way there!\n\nI think you may have misunderstood Junio's response. We do have\npush.negotiate already. It's just not the default.\n\nDid you try your example with \"git -c push.negotiate=true push ...\"?\n\n-Peff\n"},{"id":"533902","messageId":"CAGe2LO188CuDetOKRQZs8MNw3Fq9LxpAwM8HMEP2AMHAB_g0_A@mail.gmail.com","threadId":"64801","inReplyTo":"20260114211115.GC1008851@coredump.intra.peff.net","subject":"Re: [BUG] Git push sends too much data unnecessarily","fromName":"Rajiv Sharma","fromEmail":"rajiv.tilakraj.sharma@gmail.com","sentAt":"2026-01-14T21:48:22Z","receivedAt":"2026-01-14T21:48:40Z","isPatch":false,"sender":{"key":"rajiv.tilakraj.sharma@gmail.com","avatar":"https://gravatar.com/avatar/37a35fa524c1baf12341e58bd7688e26ee84ee0a1e44c81d100f671dcdb51654?d=mp&s=160"},"body":"Ah you are right, \"push.negotiate\" is exactly what is needed here. I\ntried this out and it works like a charm. Thanks for sorting this out.\n\n- Rajiv Sharma\n\nOn Wed, Jan 14, 2026 at 9:11 PM Jeff King <peff@peff.net> wrote:\n>\n> On Wed, Jan 14, 2026 at 05:39:43PM +0000, Rajiv Sharma wrote:\n>\n> > > This boils down to how Git negotiates between the client <> server\n> >\n> > I think that's the crux of the problem here. I don't think git\n> > negotiates in the push path the way it does in the read path, i.e.\n> > there is no process of client-server communication that involves\n> > gradually arriving at the common base (in this case it would be C3).\n> > The read path does this quite well (using something akin to a skiplist\n> > IIRC?) and the common base is found in a couple iterations in most\n> > cases. I am unaware of the historical context behind this difference\n> > but I assume the server sending unnecessary extra data during the read\n> > path would be much more expensive than the client doing it hence the\n> > push protocol is kept simpler.\n> >\n> > This kind of negotiation _could_ be added to the push path but it\n> > would be a breaking change. I read somewhere that there were plans for\n> > Push Protocol V2 (in the same vein as Read Protocol V2) so it would be\n> > great to see this improvement making its way there!\n>\n> I think you may have misunderstood Junio's response. We do have\n> push.negotiate already. It's just not the default.\n>\n> Did you try your example with \"git -c push.negotiate=true push ...\"?\n>\n> -Peff\n"},{"id":"533921","messageId":"CAOLa=ZTZhWscU=4mAb=FhMSNS2r1S0stE9NpCQBKpioTudhfXw@mail.gmail.com","threadId":"64801","inReplyTo":"xmqqh5sof61i.fsf@gitster.g","subject":"Re: [BUG] Git push sends too much data unnecessarily","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-01-15T09:43:25Z","receivedAt":"2026-01-15T09:43:27Z","isPatch":false,"sender":{"key":"karthik.188@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1786334?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Karthik Nayak <karthik.188@gmail.com> writes:\n>\n>> So it sends all objects required to create the reference, in our case 4\n>> objects, in your case GBs of data.\n>\n> \"push.negotiate\"?\n\nNeat. Everyday there is something new to know! Thanks.\n"}]}