git/list[1] front-page[2] threads[3] people[4] search[5] about
 

propagating repo corruption across clone

From
Jeff King <peff@peff.net>
Date
Mar 24, 2013, 18:31 UTC
Message-ID
<20130324183133.GA11200@sigill.intra.peff.net>
I saw this post-mortem on recent disk corruption seen on git.kde.org:
  http://jefferai.org/2013/03/24/too-perfect-a-mirror/

The interesting bit to me is that object corruption propagated across a clone (and oddly, that --mirror made complaints about corruption go away). I did a little testing and found some curious results (this ended up long; skip to the bottom for my conclusions).

Here's a fairly straight-forward corruption recipe:
-- >8 --
obj_to_file() {
  echo ".git/objects/$(echo $1 | sed 's,..,&/,')"
}
# corrupt a single byte inside the object
corrupt_object() {
  fn=$(obj_to_file "$1") &&
  chmod +w "$fn" &&
  printf '\0' | dd of="$fn" bs=1 conv=notrunc seek=10
}

git init repo && cd repo && echo content >file && git add file && git commit -m one && corrupt_object $(git rev-parse HEAD:file) -- 8< --

report git clone . fast-local report git clone --no-local . no-local report git -c transfer.unpackLimit=1 clone --no-local . index-pack report git -c fetch.fsckObjects=1 clone --no-local . fsck

and here is how clone reacts in a few situations:
  $ git clone --bare . local-bare && echo WORKED
  Cloning into bare repository 'local-bare'...
  done.
  WORKED

We don't notice the problem during the transport phase, which is to be expected; we're using the fast "just hardlink it" code path. So that's OK.

  $ git clone . local-tree && echo WORKED
  Cloning into 'local-tree'...
  done.
  error: inflate: data stream error (invalid distance too far back)
  error: unable to unpack d95f3ad14dee633a758d2e331151e950dd13e4ed header
  WORKED

We _do_ see a problem during the checkout phase, but we don't propagate a checkout failure to the exit code from clone. That is bad in general, and should probably be fixed. Though it would never find corruption of older objects in the history, anyway, so checkout should not be relied on for robustness.

  $ git clone --no-local . non-local && echo WORKED
  Cloning into 'non-local'...
  remote: Counting objects: 3, done.
  remote: error: inflate: data stream error (invalid distance too far back)
  remote: error: unable to unpack d95f3ad14dee633a758d2e331151e950dd13e4ed header
  remote: error: inflate: data stream error (invalid distance too far back)
  remote: fatal: loose object d95f3ad14dee633a758d2e331151e950dd13e4ed (stored in ./objects/d9/5f3ad14dee633a758d2e331151e950dd13e4ed) is corrupt
  error: git upload-pack: git-pack-objects died with error.
  fatal: git upload-pack: aborting due to possible repository corruption on the remote side.
  remote: aborting due to possible repository corruption on the remote side.
  fatal: early EOF
  fatal: index-pack failed

Here we detect the error. It's noticed by pack-objects on the remote side as it tries to put the bogus object into a pack. But what if we already have a pack that's been corrupted, and pack-objects is just pushing out entries without doing any recompression?

Let's change our corrupt_object to:
  corrupt_object() {
    git repack -ad &&
    pack=`echo .git/objects/pack/*.pack` &&
    chmod +w "$pack" &&
    printf '\0' | dd of="$pack" bs=1 conv=notrunc seek=175
  }
and try again:
  $ git clone --no-local . non-local && echo WORKED
  Cloning into 'non-local'...
  remote: Counting objects: 3, done.
  remote: Total 3 (delta 0), reused 3 (delta 0)
  error: inflate: data stream error (invalid distance too far back)
  fatal: pack has bad object at offset 169: inflate returned -3
  fatal: index-pack failed

Great, we still notice the problem in unpack-objects on the receiving end. But what if there's a more subtle corruption, where filesystem corruption points the directory entry for one object at the inode of another. Like:

  corrupt_object() {
    corrupt=$(echo corrupted | git hash-object -w --stdin) &&
    mv -f $(obj_to_file $corrupt) $(obj_to_file $1)
  }

This is going to be more subtle, because the object in the packfile is self-consistent but the object graph as a whole is broken.

  $ git clone --no-local . non-local && echo WORKED
  Cloning into 'non-local'...
  remote: Counting objects: 3, done.
  remote: Total 3 (delta 0), reused 0 (delta 0)
  Receiving objects: 100% (3/3), done.
  error: unable to find d95f3ad14dee633a758d2e331151e950dd13e4ed
  WORKED

Like the --local cases earlier, we notice the missing object during the checkout phase, but do not correctly propagate the error.

We do not notice the sha1 mis-match on the sending side (which we could, if we checked the sha1 as we were sending). We do not notice the broken object graph during the receive process either. I would have expected check_everything_connected to handle this, but we don't actually call it during clone! If you do this:

  $ git init non-local && cd non-local && git fetch ..
  remote: Counting objects: 3, done.
  remote: Total 3 (delta 0), reused 3 (delta 0)
  Unpacking objects: 100% (3/3), done.
  fatal: missing blob object 'd95f3ad14dee633a758d2e331151e950dd13e4ed'
  error: .. did not send all necessary objects
we do notice.
And one final check:
  $ git -c transfer.fsckobjects=1 clone --no-local . fsck
  Cloning into 'fsck'...
  remote: Counting objects: 3, done.
  remote: Total 3 (delta 0), reused 3 (delta 0)
  Receiving objects: 100% (3/3), done.
  error: unable to find d95f3ad14dee633a758d2e331151e950dd13e4ed
  fatal: object of unexpected type
  fatal: index-pack failed

Fscking the incoming objects does work, but of course it comes at a cost in the normal case (for linux-2.6, I measured an increase in CPU time with "index-pack --strict" from ~2.5 minutes to ~4 minutes). And I think it is probably overkill for finding corruption; index-pack already recognizes bit corruption inside an object, and check_everything_connected can detect object graph problems much more cheaply.

One thing I didn't check is bit corruption inside a packed object that still correctly zlib inflates. check_everything_connected will end up reading all of the commits and trees (to walk them), but not the blobs. And I don't think that we explicitly re-sha1 every incoming object (only if we detect a possible collision). So it may be that transfer.fsckObjects would save us there (it also introduces new problems if there are ignorable warnings in the objects you receive, like zero-padded trees).

So I think at the very least we should:
  1. Make sure clone propagates errors from checkout to the final exit
     code.
  2. Teach clone to run check_everything_connected.

I don't have details on the KDE corruption, or why it wasn't detected (if it was one of the cases I mentioned above, or a more subtle issue).

-Peff
Next: Ævar Arnfjörð Bjarmason
Message 1 of 60 in “propagating repo corruption across clone”
  1. Jeff KingMar 24, 2013
  2. Ævar Arnfjörð BjarmasonMar 24, 2013
  3. Jeff KingMar 24, 2013
  4. Jeff MitchellMar 25, 2013
  5. Jeff KingMar 25, 2013
  6. Duy NguyenMar 25, 2013
  7. Jeff KingMar 25, 2013
  8. Jeff MitchellMar 25, 2013
  9. Jeff KingMar 25, 2013
  10. Jeff MitchellMar 26, 2013
  11. Jeff KingMar 26, 2013
  12. Philip OakleyMar 26, 2013
  13. Jeff KingMar 26, 2013
  14. Rich FrommMar 26, 2013
  15. Jonathan NiederMar 27, 2013
  16. Rich FrommMar 27, 2013
  17. Jeff KingMar 27, 2013
  18. Jeff KingMar 27, 2013
  19. Junio C HamanoMar 27, 2013
  20. Sitaram ChamartyMar 27, 2013
  21. Junio C HamanoMar 27, 2013
  22. Sitaram ChamartyMar 27, 2013
  23. Rich FrommMar 27, 2013
  24. Junio C HamanoMar 27, 2013
  25. Jeff MitchellMar 28, 2013
  26. Jeff MitchellMar 28, 2013
  27. Duy NguyenMar 26, 2013
  28. Ilari LiusvaaraMar 24, 2013
  29. Junio C HamanoMar 25, 2013
  30. Jeff KingMar 25, 2013
  31. 0/9 corrupt object potpourriJeff King, Mar 25, 2013
  32. 1/9 stream_blob_to_fd: detect errors reading from streamJeff King, Mar 25, 2013
  33. Junio C HamanoMar 26, 2013
  34. 2/9 check_sha1_signature: check return value from read_istreamJeff King, Mar 25, 2013
  35. 3/9 read_istream_filtered: propagate read error from upstreamJeff King, Mar 25, 2013
  36. 4/9 avoid infinite loop in read_istream_looseJeff King, Mar 25, 2013
  37. 5/9 add test for streaming corrupt blobsJeff King, Mar 25, 2013
  38. Jonathan NiederMar 25, 2013
  39. Jeff KingMar 25, 2013
  40. Jeff KingMar 27, 2013
  41. Junio C HamanoMar 27, 2013
  42. 6/9 streaming_write_entry: propagate streaming errorsJeff King, Mar 25, 2013
  43. Eric SunshineMar 25, 2013
  44. Jeff KingMar 25, 2013
  45. Jonathan NiederMar 25, 2013
  46. 6/9 streaming_write_entry: propagate streaming errorsJeff King, Mar 25, 2013
  47. Jonathan NiederMar 25, 2013
  48. Junio C HamanoMar 26, 2013
  49. 7/9 add tests for cloning corrupted repositoriesJeff King, Mar 25, 2013
  50. 8/9 clone: die on errors from unpack_treesJeff King, Mar 25, 2013
  51. Junio C HamanoMar 26, 2013
  52. 10/9 clone: leave repo in place after checkout errorsJeff King, Mar 26, 2013
  53. Jonathan NiederMar 26, 2013
  54. Jeff KingMar 27, 2013
  55. 9/9 clone: run check_everything_connectedJeff King, Mar 25, 2013
  56. Duy NguyenMar 26, 2013
  57. Jeff KingMar 26, 2013
  58. Junio C HamanoMar 26, 2013
  59. Duy NguyenMar 28, 2013
  60. Duy NguyenMar 31, 2013

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.