git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git-fetch per-repository speed issues

From
Linus Torvalds <torvalds@osdl.org>
Date
Jul 4, 2006, 05:36 UTC
Message-ID
<Pine.LNX.4.64.0607032213030.12404@g5.osdl.org>
In-Reply-To
<1151989503.4723.126.camel@neko.keithp.com>
On Mon, 3 Jul 2006, Keith Packard wrote:
Show 14 quoted lines
> 
> 5 Start:                             21:59:01.584648000
> 66 After args:                       21:59:01.605987000
> 248 fetch_main() start:              21:59:02.408559000
> 339 fetch_main() before fetch-pack:  21:59:03.293228000
> 387 fetch_main() done:               21:59:04.784388000
> 422 After tag following:             21:59:05.311439000
> 438 All done:                        21:59:05.315338000
> 
> fetch-pack itself took 0.421 seconds (measured with time(1)).
> 
> Looks like the bulk of the time here is caused by simple shell
> processing overhead, some of which scales with the number of heads and
> tags to track.
Ahh.. Do you have tons of tags at the other end?
Looking closer, I suspect a big part of it is that
	git-ls-remote $upload_pack --tags "$remote" |
	sed -ne 's|^\([0-9a-f]*\)[      ]\(refs/tags/.*\)^{}$|\1 \2|p' |
	while read sha1 name
	do
		..
	done
loop.

With a lot of tags, the shell overhead there can indeed be pretty disgusting. And I was wrong - I thought it would do that git-ls-remote only if the first time around we noticed that we would need to, but we do actually do it all the time that we're fetching any new branches.

The sad part is that we really already got the list once, we just never saved it away (ie "git-fetch-pack" actually _knows_ what the tags at the other end are, and also knows which tags we already have, so if we made git-fetch-pack just create that list and save it off, all the overhead would just go away).

And yes, the shell script loops are really really simple, but some of them are actually quadratic in the number of refs (O(local*remote)). If this was a C program, we'd never even care, but with shell, the thing is slow enough that having even a modest amount of tags and refs is going to just make it waste a lot of time in shell scripting.

We already do a lot of the infrastructure for "git fetch" in C - the remotes parsing etc is all things that "git fetch" used to share with "git push", but "git push" has been a builtin C program for a while now. I suspect we should just do the same to "git fetch", which would make all these issues just totally go away.

			Linus
Previous: Keith PackardNext: Junio C Hamano
Message 27 of 30 in “git-fetch per-repository speed issues”
  1. Keith PackardJul 3, 2006
  2. Linus TorvaldsJul 3, 2006
  3. Jeff KingJul 4, 2006
  4. Ryan AndersonJul 4, 2006
  5. Jeff KingJul 4, 2006
  6. Ryan AndersonJul 4, 2006
  7. Linus TorvaldsJul 4, 2006
  8. Jeff KingJul 5, 2006
  9. Linus TorvaldsJul 5, 2006
  10. Jakub NarebskiJul 4, 2006
  11. Jakub NarebskiJul 4, 2006
  12. Thomas GlanzmannJul 4, 2006
  13. Junio C HamanoJul 4, 2006
  14. Linus TorvaldsJul 4, 2006
  15. Junio C HamanoJul 4, 2006
  16. David WoodhouseJul 6, 2006
  17. Linus TorvaldsJul 4, 2006
  18. Junio C HamanoJul 4, 2006
  19. Linus TorvaldsJul 4, 2006
  20. Keith PackardJul 4, 2006
  21. Andreas EricssonJul 4, 2006
  22. Matthias KestenholzJul 4, 2006
  23. Andreas EricssonJul 4, 2006
  24. Keith PackardJul 4, 2006
  25. Linus TorvaldsJul 4, 2006
  26. Keith PackardJul 4, 2006
  27. Linus TorvaldsJul 4, 2006
  28. Junio C HamanoJul 4, 2006
  29. Keith PackardJul 4, 2006
  30. Linus TorvaldsJul 4, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.