git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git clone sending unneeded objects

From
Nicolas Pitre <nico@fluxnic.net>
Date
Sep 27, 2009, 02:26 UTC
Message-ID
<alpine.LFD.2.00.0909262140280.4997@xanadu.home>
In-Reply-To
<4ABE1818.6010007@redhat.com>
On Sat, 26 Sep 2009, Jason Merrill wrote:
Show 13 quoted lines
> On 09/26/2009 12:44 AM, Jason Merrill wrote:
> > git config remote.origin.fetch 'refs/remotes/*:refs/remotes/origin/*'
> > git fetch
> 
> git count-objects -v before:
> 
> count: 44
> size: 1768
> in-pack: 1399509
> packs: 1
> size-pack: 600456
> prune-packable: 0
> garbage: 0

I'm sure if you had done 'git rev-list --all --objects | wc -l' at that point, the result would have been something around 900000. That's the actual number of objects git had a reference to, compared to the total objects contained in the object store.

Show 9 quoted lines
> and after (transferred 278MB):
> 
> count: 44
> size: 1768
> in-pack: 1947339
> packs: 2
> size-pack: 1178408
> prune-packable: 8
> garbage: 0

And those 500000 extra objects or so (minus a couple dozens which were probably used to "complete" the fetched thin pack and are duplicates of local objects -- the fetch progress message gave the exact number) were obtained from the remote repository because git has no way to tell the remote it already had them. That's what I was explaining in my previous email.

Show 12 quoted lines
> and then after git gc --prune=now:
> 
> count: 0
> size: 0
> in-pack: 1399613
> packs: 1
> size-pack: 839900
> prune-packable: 0
> garbage: 0
> 
> So I only actually needed 104 more objects, but fetch wasn't clever enough to
> see that, and my new pack is much less efficient.

Like I said, it's not that the fetch wasn't clever enough. Rather that your initial clone asked for way too many objects in the first place. That's what my patch fixed.

Now the pack efficiency can be explained as well. A single pack is always going to be more efficient than 2 packs. Problem is when you do a gc, by default git does the least costly operation which consists of copying as much data from existing packs without extra processing. That means that many objects were copied from the second (newly received) pack although a better delta representation was most probably available in the other larger pack (remember that most objects from that second pack already existed in the first pack). Git do select the second pack in preference to the other pack because it is more recent, and normally more recent packs contains more recent objects which is a good heuristic to optimizes the object enumeration. In this case this didn't produce a good result, but again we're talking about a scenario which is bogus from the start and shouldn't be.

So if you do a 'git gc --aggressive' and let it run for a while, you should get back a smaller pack, possibly even much smaller than the original one.

Nicolas
Previous: Jason MerrillNext: Nicolas Pitre
Message 20 of 26 in “Re: git gc expanding packed data?”
  1. Andreas SchwabAug 8, 2009
  2. Hin-Tak LeungAug 8, 2009
  3. Andreas SchwabAug 8, 2009
  4. Nicolas PitreAug 9, 2009
  5. Andreas SchwabAug 9, 2009
  6. Jason MerrillSep 25, 2009
  7. Matthieu MoySep 25, 2009
  8. Jason MerrillSep 25, 2009
  9. Nicolas PitreSep 25, 2009
  10. Jason MerrillSep 25, 2009
  11. Nicolas PitreSep 25, 2009
  12. Jason MerrillSep 25, 2009
  13. Nicolas PitreSep 26, 2009
  14. make 'git clone' ask the remote only for objects it cares aboutNicolas Pitre, Sep 26, 2009
  15. Andreas SchwabSep 26, 2009
  16. Shawn O. PearceSep 26, 2009
  17. Nicolas PitreSep 27, 2009
  18. Jason MerrillSep 26, 2009
  19. Jason MerrillSep 26, 2009
  20. Nicolas PitreSep 27, 2009
  21. Nicolas PitreSep 27, 2009
  22. Shawn O. PearceSep 27, 2009
  23. Nicolas PitreSep 27, 2009
  24. Jason MerrillSep 27, 2009
  25. Nicolas PitreSep 28, 2009
  26. Hin-Tak LeungSep 26, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.