git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 0/5] handling 4GB .idx files

From
Jeff King <peff@peff.net>
Date
Nov 16, 2020, 04:10 UTC
Message-ID
<20201116041051.GA883199@coredump.intra.peff.net>
In-Reply-To
<323fd904-a7ee-061d-d846-5da5afbc88b2@virtuell-zuhause.de>
On Sun, Nov 15, 2020 at 03:43:39PM +0100, Thomas Braun wrote:
Show 13 quoted lines
> On 13.11.2020 06:06, Jeff King wrote:
> > I recently ran into a case where Git could not read the pack it had
> > produced via running "git repack". The culprit turned out to be an .idx
> > file which crossed the 4GB barrier (in bytes, not number of objects).
> > This series fixes the problems I saw, along with similar ones I couldn't
> > trigger in practice, and protects the .idx loading code against integer
> > overflows that would fool the size checks.
> 
> Would it be feasible to have a test case for this large index case? This
> should very certainly have an EXPENSIVE tag, or might even not yet work
> on windows. But hopefully someday I'll find some more time to push large
> object support on windows forward, and these kind of tests would really
> help then.

I think it would be a level beyond what we usually consider even for EXPENSIVE. The cheapest I could come up with to generate the case is:

  perl -e '
	for (0..154_000_000) {
		print "blob\n";
		print "data <<EOF\n";
		print "$_\n";
		print "EOF\n";
	}
  ' |
  git fast-import

which took almost 13 minutes of CPU to run, and peaked around 15GB of RAM (and takes about 6.7GB on disk).

In the resulting repo, the old code barfed on lookups:
  $ blob=$(echo 0 | git hash-object --stdin)
  $ git cat-file blob $blob
  error: wrong index v2 file size in .git/objects/pack/pack-f8f43ae56c25c1c8ff49ad6320df6efb393f551e.idx
  error: wrong index v2 file size in .git/objects/pack/pack-f8f43ae56c25c1c8ff49ad6320df6efb393f551e.idx
  error: wrong index v2 file size in .git/objects/pack/pack-f8f43ae56c25c1c8ff49ad6320df6efb393f551e.idx
  error: wrong index v2 file size in .git/objects/pack/pack-f8f43ae56c25c1c8ff49ad6320df6efb393f551e.idx
  error: wrong index v2 file size in .git/objects/pack/pack-f8f43ae56c25c1c8ff49ad6320df6efb393f551e.idx
  error: wrong index v2 file size in .git/objects/pack/pack-f8f43ae56c25c1c8ff49ad6320df6efb393f551e.idx
  fatal: git cat-file 573541ac9702dd3969c9bc859d2b91ec1f7e6e56: bad file
whereas now it works:
  $ git cat-file blob $blob
  0

That's the most basic test I think you could do. More interesting is looking at entries that are actually after the 4GB mark. That requires dumping the whole index:

  final=$(git show-index <.git/objects/pack/*.idx | tail -1 | awk '{print $2}')
  git cat-file blob $final

That takes ~35s to run. Curiously, it also allocates 5GB of heap. For some reason it decides to make an internal copy of the entries table. I guess because it reads the file sequentially rather than mmap-ing it, and 64-bit offsets in v2 idx files can't be resolved until we've read the whole entry table (and it wants to output the entries in sha1 order).

The checksum bug requires running git-fsck on the repo. That's another 5 minutes of CPU (and even higher peak memory; I think we create a "struct blob" for each one, and it seems to hit 20GB).

Hitting the other cases that I fixed but never triggered in practice would need a repo about 4x as large. So figure an hour of CPU and 60GB of RAM.

So I dunno. I wouldn't be opposed to codifying some of that in a script, but I can't imagine anybody ever running it unless they were working on this specific problem.

-Peff
Previous: Thomas BraunNext: Derrick Stolee
Message 9 of 19 in “handling 4GB .idx files”
  1. 0/5 handling 4GB .idx filesJeff King, Nov 13, 2020
  2. 1/5 compute pack .idx byte offsets using size_tJeff King, Nov 13, 2020
  3. 2/5 use size_t to store pack .idx byte offsetsJeff King, Nov 13, 2020
  4. 3/5 fsck: correctly compute checksums on idx files larger than 4GBJeff King, Nov 13, 2020
  5. 4/5 block-sha1: take a size_t length parameterJeff King, Nov 13, 2020
  6. 5/5 packfile: detect overflow in .idx file size checksJeff King, Nov 13, 2020
  7. Johannes SchindelinNov 13, 2020
  8. Thomas BraunNov 15, 2020
  9. Jeff KingNov 16, 2020
  10. Derrick StoleeNov 16, 2020
  11. Jeff KingNov 16, 2020
  12. Thomas BraunNov 30, 2020
  13. Jeff KingDec 1, 2020
  14. t7900's new expensive testJeff King, Dec 1, 2020
  15. Derrick StoleeDec 1, 2020
  16. t7900: speed up expensive testJeff King, Dec 2, 2020
  17. Derrick StoleeDec 3, 2020
  18. Taylor BlauDec 1, 2020
  19. Jeff KingDec 2, 2020

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.