git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git-index-pack really does suck..

From
Nicolas Pitre <nico@cam.org>
Date
Apr 3, 2007, 16:33 UTC
Message-ID
<alpine.LFD.0.98.0704031220470.28181@xanadu.home>
In-Reply-To
<Pine.LNX.4.64.0704030754020.6730@woody.linux-foundation.org>
On Tue, 3 Apr 2007, Linus Torvalds wrote:
> 
> Junio, Nico,
>  I think we need to do something about it.
Sure.  Mea culpa.
Show 9 quoted lines
> CLee was complaining about git-index-pack on #irc with the partial KDE 
> repo, and while I don't have the KDE repo, I decided to investigate a bit.
> 
> Even with just the kernel repo (with a single 170MB pack-file), I can do
> 
> 	git index-pack --stdin --fix-thin new.pack < .git/objects/pack/pack-*.pack
> 
> and it uses 52s of CPU-time, and on my 4GB machine it actually started 
> doing IO and swapping, because git-index-pack grew to 4.8GB in size.
Right.
Show 22 quoted lines
> Two suggestion for other ways:
> 
>  - simple one: don't keep unexploded objects around, just keep the deltas, 
>    and spend tons of CPU-time just re-expanding them if required.
> 
>    We *should* be able to do it with just keeping the original 170MB 
>    pack-file in memory, not expanding it to 3.8GB! 
> 
>    Still, even this will be painful once you have a big pack-file, and the 
>    CPU waste is nasty (although a delta-base cache like we do in 
>    sha1_file.c would probably fix it 99% - at that point it's getting 
>    less simple, and the "best" solution below looks more palatable)
> 
>  - best one: when writing out the pack-file, we incrementally keep a 
>    "struct packed_git" around, and update the index for it dynamically, 
>    and totally get rid of all objects that we've written out, because we 
>    can re-create them.
> 
>    This means that we should have _zero_ memory footprint except for the 
>    one object that we're working on right then and there, and any 
>    unresolved deltas where we've not seen the base at all (and the latter 
>    generally shouldn't happen any more with most pack-files)
Even better:
  - Fix my own stupid mistake with a _single_ line of code:
diff --git a/index-pack.c b/index-pack.c
index 6284fe3..3c768fb 100644
--- a/index-pack.c
+++ b/index-pack.c
@@ -358,6 +358,7 @@ static void sha1_object(const void *data, unsigned long size,
 		if (size != has_size || type != has_type ||
 		    memcmp(data, has_data, size) != 0)
 			die("SHA1 COLLISION FOUND WITH %s !", sha1_to_hex(sha1));
+		free(has_data);
 	}
 }

The thing is, that code path is executed _only_ when index-pack is 
encountering an object already in the repository in order to protect 
against possible SHA1 collision attacks.  See commit 8685da42561d log 
for the full story.

Normally this should not happen in normal usage scenarios because the 
objects you fetch are those that you don't already have.  But if you 
manually run index-pack inside an existing repository then you'll 
already have _all_ those objects already explaining the high CPU usage.

But this is no excuse for not freeing the data though.


Nicolas
Previous: Nicolas PitreNext: Chris Lee
Message 4 of 58 in “git-index-pack really does suck..”
  1. Linus TorvaldsApr 3, 2007
  2. Linus TorvaldsApr 3, 2007
  3. Nicolas PitreApr 3, 2007
  4. Nicolas PitreApr 3, 2007
  5. Chris LeeApr 3, 2007
  6. Nicolas PitreApr 3, 2007
  7. Chris LeeApr 3, 2007
  8. Linus TorvaldsApr 3, 2007
  9. Nicolas PitreApr 3, 2007
  10. Junio C HamanoApr 3, 2007
  11. Linus TorvaldsApr 3, 2007
  12. Nicolas PitreApr 3, 2007
  13. Chris LeeApr 3, 2007
  14. Linus TorvaldsApr 3, 2007
  15. Linus TorvaldsApr 3, 2007
  16. Shawn O. PearceApr 3, 2007
  17. Linus TorvaldsApr 3, 2007
  18. Shawn O. PearceApr 3, 2007
  19. Linus TorvaldsApr 3, 2007
  20. Linus TorvaldsApr 3, 2007
  21. Junio C HamanoApr 3, 2007
  22. Shawn O. PearceApr 3, 2007
  23. Junio C HamanoApr 3, 2007
  24. 1/2 git-fetch--tool pick-rrefJunio C Hamano, Apr 5, 2007
  25. 2/2 git-fetch: use fetch--tool pick-rref to avoid local fetch from alternateJunio C Hamano, Apr 5, 2007
  26. Shawn O. PearceApr 5, 2007
  27. Junio C HamanoApr 5, 2007
  28. Nicolas PitreApr 3, 2007
  29. Shawn O. PearceApr 3, 2007
  30. Junio C HamanoApr 3, 2007
  31. Shawn O. PearceApr 3, 2007
  32. Jeff KingApr 3, 2007
  33. Dana HowApr 3, 2007
  34. Linus TorvaldsApr 3, 2007
  35. David LangApr 3, 2007
  36. Nicolas PitreApr 3, 2007
  37. Nicolas PitreApr 3, 2007
  38. Linus TorvaldsApr 3, 2007
  39. Nicolas PitreApr 3, 2007
  40. Shawn O. PearceApr 3, 2007
  41. Linus TorvaldsApr 3, 2007
  42. Nicolas PitreApr 3, 2007
  43. Junio C HamanoApr 3, 2007
  44. Shawn O. PearceApr 3, 2007
  45. Nicolas PitreApr 3, 2007
  46. Linus TorvaldsApr 3, 2007
  47. Nicolas PitreApr 3, 2007
  48. David LangApr 3, 2007
  49. Alex RiesenApr 4, 2007
  50. David LangApr 6, 2007
  51. Junio C HamanoApr 6, 2007
  52. Junio C HamanoApr 6, 2007
  53. David LangApr 6, 2007
  54. Junio C HamanoApr 6, 2007
  55. David LangApr 6, 2007
  56. Linus TorvaldsApr 3, 2007
  57. Junio C HamanoApr 3, 2007
  58. Nicolas PitreApr 3, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.