git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git 2.26 fetches many times more objects than it should, wasting gigabytes

From
Jeff King <peff@peff.net>
Date
Apr 24, 2020, 05:32 UTC
Message-ID
<20200424053204.GD1648190@coredump.intra.peff.net>
In-Reply-To
<20200423213735.242662-1-jonathantanmy@google.com>
On Thu, Apr 23, 2020 at 02:37:35PM -0700, Jonathan Tan wrote:
Show 11 quoted lines
> Thanks for the reproduction recipe (in [1]) and your analysis. I took a
> look, and it's because the check for in_vain is done differently. In v0:
> 
>   if (got_continue && MAX_IN_VAIN < in_vain) {
> 
> reflecting the documentation in pack-protocol.txt:
> 
>   However, the 256 limit *only* turns on in the canonical client
>   implementation if we have received at least one "ACK %s continue"
>   during a prior round.  This helps to ensure that at least one common
>   ancestor is found before we give up entirely.

Ah, thanks for that; I hadn't though to look in that file for more clues.

Show 14 quoted lines
> When debugging, I noticed that in_vain was increasing far in excess of
> MAX_IN_VAIN, but because got_continue was false, the client did not give
> up.
> 
> But in v2:
> 
>   if (!haves_added || *in_vain >= MAX_IN_VAIN) {
> 
> ("haves_added" is irrelevant to this discussion. It is another
> termination condition - when we have run out of "have"s to send.)
> 
> So there is no check that "continue" was sent. We probably should change
> v2 to match v0. I can start writing a patch unless someone else would
> like to take a further look at it.
Yeah, this fills in the final pieces of the puzzle I was chasing in:
 https://lore.kernel.org/git/20200422193324.GB558336@coredump.intra.peff.net/
And the patch you suggest sounds like the best solution.

I think there's some room for discussion about what the optimal strategies are (e.g., v0 does send a lot more haves than v2 in this instance, and it wouldn't always be helpful). But it makes sense to me to put v2 and v0 on the same footing for now, especially given the regressions people have mentioned, and then we can explore new options at our convenience (like switching on the skipping negotiation algorithm).

-Peff
Previous: Junio C HamanoNext: Jonathan Nieder
Message 9 of 18 in “Git 2.26 fetches many times more objects than it should, wasting gigabytes”
  1. Lubomir RintelApr 22, 2020
  2. Jeff KingApr 22, 2020
  3. Jeff KingApr 22, 2020
  4. Jeff KingApr 22, 2020
  5. Junio C HamanoApr 22, 2020
  6. Jeff KingApr 22, 2020
  7. Jonathan TanApr 23, 2020
  8. Junio C HamanoApr 23, 2020
  9. Jeff KingApr 24, 2020
  10. Jonathan NiederApr 22, 2020
  11. Jeff KingApr 22, 2020
  12. Revert "fetch: default to protocol version 2"Jonathan Nieder, Apr 22, 2020
  13. Junio C HamanoApr 22, 2020
  14. Jeff KingApr 22, 2020
  15. Jeff KingApr 22, 2020
  16. Jonathan NiederApr 22, 2020
  17. Junio C HamanoApr 22, 2020
  18. Jeff KingApr 22, 2020

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.