{"thread":{"id":"15196","subject":"\"failed to read delta base object at...\"","startedAt":"2008-08-25T16:46:02Z","lastAt":"2008-08-27T20:46:38Z","messageCount":16,"participants":["J. Bruce Fields","Nicolas Pitre","Linus Torvalds","Jason McMullan","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"88473","messageId":"20080825164602.GA2213@fieldses.org","threadId":"15196","inReplyTo":null,"subject":"\"failed to read delta base object at...\"","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2008-08-25T16:46:02Z","receivedAt":"2008-08-25T16:46:02Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"Today I got this:\n\nfatal: failed to read delta base object at 3025976 from\n/home/bfields/local/linux-2.6/.git/objects/pack/pack-f7261d96cf1161b1b0a1593f673a67d0f2469e9b.pack\n\nThis has happened once before recently, I believe with a pack that had\njust been created on a recent fetch.  (If I remember correctly, this was\nsoon after a failed suspend/resume cycle that might have interrupted an\nin-progress fetch; could that possible explain the error?)  In that case\nI reset origin/master, deleted a tag or two, and fetched, and the\nproblem seemed to be fixed.\n\n--b.\n"},{"id":"88487","messageId":"alpine.LFD.1.10.0808251445090.1624@xanadu.home","threadId":"15196","inReplyTo":"20080825164602.GA2213@fieldses.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-08-25T18:58:05Z","receivedAt":"2008-08-25T18:58:05Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 25 Aug 2008, J. Bruce Fields wrote:\n\n> Today I got this:\n> \n> fatal: failed to read delta base object at 3025976 from\n> /home/bfields/local/linux-2.6/.git/objects/pack/pack-f7261d96cf1161b1b0a1593f673a67d0f2469e9b.pack\n> \n> This has happened once before recently, I believe with a pack that had\n> just been created on a recent fetch.  (If I remember correctly, this was\n> soon after a failed suspend/resume cycle that might have interrupted an\n> in-progress fetch; could that possible explain the error?)  In that case\n> I reset origin/master, deleted a tag or two, and fetched, and the\n> problem seemed to be fixed.\n\nThe above error is indicative of a corrupted pack on disk.  To confirm \nit you could use 'git verify-pack' with the given pack file.\n\nWith a sufficiently recent git, you only need to copy over another pack \ncontaining the corrupted object, or the object itself in loose form, \ninto your object store to \"fix\" it.\n\nAs to the source of disk corruptions... that's up to you to find the \ncause amongst many (including a failed suspend).\n\n\nNicolas\n"},{"id":"88488","messageId":"alpine.LFD.1.10.0808251153210.3363@nehalem.linux-foundation.org","threadId":"15196","inReplyTo":"20080825164602.GA2213@fieldses.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-08-25T19:01:42Z","receivedAt":"2008-08-25T19:01:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Aug 2008, J. Bruce Fields wrote:\n>\n> Today I got this:\n> \n> fatal: failed to read delta base object at 3025976 from\n> /home/bfields/local/linux-2.6/.git/objects/pack/pack-f7261d96cf1161b1b0a1593f673a67d0f2469e9b.pack\n\nThis is almost certainly due to some corruption. Basically, the call to \n\"cache_or_unpack_entry()\" failed, which in turn is because \n'unpack_entry()' will have failed. \n\nAnd since you didn't see any other error, that failure is almost certainly \ndue to unpack_compressed_entry() having failed. We don't print out _why_ \n(which is a bit sad), but the only thing that unpack_compressed_entry() \ndoes is to just \"inflate()\" the data at that offset.\n\nSo it probably got a zlib data error, or an adler32 crc failure.\n\n> This has happened once before recently, I believe with a pack that had\n> just been created on a recent fetch.  (If I remember correctly, this was\n> soon after a failed suspend/resume cycle that might have interrupted an\n> in-progress fetch; could that possible explain the error?)  In that case\n> I reset origin/master, deleted a tag or two, and fetched, and the\n> problem seemed to be fixed.\n\nAn interrupted fetch shouldn't have caused this, it really should only \nhappen if you have some actual filesystem data error. Something didn't get \nwritten back correctly, or the page cache isn't coherent (or it got \ncorrupted by something else like a wild kernel pointer, of course).\n\n\t\t\tLinus\n"},{"id":"88512","messageId":"20080825211830.GH2213@fieldses.org","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808251445090.1624@xanadu.home","subject":"Re: \"failed to read delta base object at...\"","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2008-08-25T21:18:30Z","receivedAt":"2008-08-25T21:18:30Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Mon, Aug 25, 2008 at 02:58:05PM -0400, Nicolas Pitre wrote:\n> On Mon, 25 Aug 2008, J. Bruce Fields wrote:\n> \n> > Today I got this:\n> > \n> > fatal: failed to read delta base object at 3025976 from\n> > /home/bfields/local/linux-2.6/.git/objects/pack/pack-f7261d96cf1161b1b0a1593f673a67d0f2469e9b.pack\n> > \n> > This has happened once before recently, I believe with a pack that had\n> > just been created on a recent fetch.  (If I remember correctly, this was\n> > soon after a failed suspend/resume cycle that might have interrupted an\n> > in-progress fetch; could that possible explain the error?)  In that case\n> > I reset origin/master, deleted a tag or two, and fetched, and the\n> > problem seemed to be fixed.\n> \n> The above error is indicative of a corrupted pack on disk.  To confirm \n> it you could use 'git verify-pack' with the given pack file.\n\nThat gives:\n\n\terror: Packfile .git/objects/pack/pack-f7261d96cf1161b1b0a1593f673a67d0f2469e9b.pack SHA1 mismatch with itself\n\nNo surprise, I assume.\n\n> With a sufficiently recent git, you only need to copy over another pack \n> containing the corrupted object, or the object itself in loose form, \n> into your object store to \"fix\" it.\n\nYeah, I just did a git repack -a -d in a known good repository and\ncopied the resulting pack over.  Seems OK.  Thanks.\n\n--b.\n\n> As to the source of disk corruptions... that's up to you to find the \n> cause amongst many (including a failed suspend).\n"},{"id":"88515","messageId":"20080825213104.GI2213@fieldses.org","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808251153210.3363@nehalem.linux-foundation.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2008-08-25T21:31:04Z","receivedAt":"2008-08-25T21:31:04Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"Thanks to you and Nicolas for the responses.\n\nOn Mon, Aug 25, 2008 at 12:01:42PM -0700, Linus Torvalds wrote:\n> On Mon, 25 Aug 2008, J. Bruce Fields wrote:\n> >\n> > Today I got this:\n> > \n> > fatal: failed to read delta base object at 3025976 from\n> > /home/bfields/local/linux-2.6/.git/objects/pack/pack-f7261d96cf1161b1b0a1593f673a67d0f2469e9b.pack\n> \n> This is almost certainly due to some corruption. Basically, the call to \n> \"cache_or_unpack_entry()\" failed, which in turn is because \n> 'unpack_entry()' will have failed. \n> \n> And since you didn't see any other error, that failure is almost certainly \n> due to unpack_compressed_entry() having failed. We don't print out _why_ \n> (which is a bit sad), but the only thing that unpack_compressed_entry() \n> does is to just \"inflate()\" the data at that offset.\n> \n> So it probably got a zlib data error, or an adler32 crc failure.\n> \n> > This has happened once before recently, I believe with a pack that had\n> > just been created on a recent fetch.  (If I remember correctly, this was\n> > soon after a failed suspend/resume cycle that might have interrupted an\n> > in-progress fetch; could that possible explain the error?)  In that case\n> > I reset origin/master, deleted a tag or two, and fetched, and the\n> > problem seemed to be fixed.\n> \n> An interrupted fetch shouldn't have caused this, it really should only \n> happen if you have some actual filesystem data error. Something didn't get \n> written back correctly, or the page cache isn't coherent (or it got \n> corrupted by something else like a wild kernel pointer, of course).\n\nOK.  I seem to recall these pack files are created with something like\n\n\topen\n\twrite\n\tsync\n\tclose\n\trename\n\n?  This is just ext3 with data=writeback on a local laptop disk,\nubuntu's 2.6.24-21-generic.  Would it be any use trying to look more\nclosely at the pack in connection for any hints?\n\n(But with my git repo back I'm happy enough to just forget this for now\nif there's not anything obvious to try.)\n\n--b.\n"},{"id":"88518","messageId":"alpine.LFD.1.10.0808251435540.3363@nehalem.linux-foundation.org","threadId":"15196","inReplyTo":"20080825213104.GI2213@fieldses.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-08-25T21:37:39Z","receivedAt":"2008-08-25T21:37:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Aug 2008, J. Bruce Fields wrote:\n> \n> OK.  I seem to recall these pack files are created with something like\n> \n> \topen\n> \twrite\n> \tsync\n> \tclose\n> \trename\n> \n> ? \n\nYes. We're trying to be _extremely_ safe and only do things that should \nwork for everything.\n\n> This is just ext3 with data=writeback on a local laptop disk,\n> ubuntu's 2.6.24-21-generic.  Would it be any use trying to look more\n> closely at the pack in connection for any hints?\n\nYou still have the packfile that caused problems available somewhere? If \nso, absolutely yes. If you have the corrupt pack, please make it \navailable.\n\n> (But with my git repo back I'm happy enough to just forget this for now\n> if there's not anything obvious to try.)\n\nWith the actual corrupt pack, we can make a fairly intelligent guess about \nexactly what the corruption was. Was it a flipped bit, or what? So if you \nhave it, please do send it over.\n\n\t\t\tLinus\n"},{"id":"88523","messageId":"20080825221321.GL2213@fieldses.org","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808251435540.3363@nehalem.linux-foundation.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2008-08-25T22:13:21Z","receivedAt":"2008-08-25T22:13:21Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Mon, Aug 25, 2008 at 02:37:39PM -0700, Linus Torvalds wrote:\n> \n> \n> On Mon, 25 Aug 2008, J. Bruce Fields wrote:\n> > \n> > OK.  I seem to recall these pack files are created with something like\n> > \n> > \topen\n> > \twrite\n> > \tsync\n> > \tclose\n> > \trename\n> > \n> > ? \n> \n> Yes. We're trying to be _extremely_ safe and only do things that should \n> work for everything.\n> \n> > This is just ext3 with data=writeback on a local laptop disk,\n> > ubuntu's 2.6.24-21-generic.  Would it be any use trying to look more\n> > closely at the pack in connection for any hints?\n> \n> You still have the packfile that caused problems available somewhere? If \n> so, absolutely yes. If you have the corrupt pack, please make it \n> available.\n> \n> > (But with my git repo back I'm happy enough to just forget this for now\n> > if there's not anything obvious to try.)\n> \n> With the actual corrupt pack, we can make a fairly intelligent guess about \n> exactly what the corruption was. Was it a flipped bit, or what? So if you \n> have it, please do send it over.\n\nOK!  It's in:\n\n\thttp://www.citi.umich.edu/u/bfields/bad-pack/\n\nI assume the .idx file isn't interesting, but it's there anyway in case.\n\n--b.\n"},{"id":"88527","messageId":"alpine.LFD.1.10.0808251616240.3363@nehalem.linux-foundation.org","threadId":"15196","inReplyTo":"20080825221321.GL2213@fieldses.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-08-25T23:59:26Z","receivedAt":"2008-08-25T23:59:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Aug 2008, J. Bruce Fields wrote:\n> \n> OK!  It's in:\n> \n> \thttp://www.citi.umich.edu/u/bfields/bad-pack/\n> \n> I assume the .idx file isn't interesting, but it's there anyway in case.\n\nIt makes things slightly easier.\n\nBut yes, the file is corrupt. The most obvious corruption is simply the \ncomparison that git-fsck does with the SHA1 that the pack-file itself \ncontains: all pack-files have at the end the SHA1 of the contents of the \npack-file, and git fsck says:\n\n\terror: .git/objects/pack/pack-f7261d96cf1161b1b0a1593f673a67d0f2469e9b.pack SHA1 checksum mismatch\n\nso it's definitely not a valid pack-file, and perhaps more importantly, \nsince git calculates the SHA1 as it writes the file out, the data really \ndoesn't match what git wrote, and it got corrupted at some point in \nbetween the generation of the data and the reading back.\n\n(It _could_ still be some git corruption where git itself corrupted it \nafter it had done the SHA1 writing, but I doubt it).\n\nAs to the exact error: zlib reports \"invalid distance too far back\" as the \nerror message when unpacking, which doesn't really tell me much. It's just \na common data error. But since I actually _have_ that corrupt object in my \npack too, I can actually look at what it should be.\n\nIt's object 784716eb7fc5735a3e6b6209223508984013819f (it's a blob encoding \nfor arch/arm/mach-pxa/sleep.S from earlier this year), and in my pack-file \nit's at byte offset 48094431-48098049.\n\nSo I can actually extract just those bytes from my pack, and compare them \nagainst the bytes in your corrupt pack. It's interesting, because the \ndifferences aren't all that big, but they aren't a single word either.\n\nHere's a \"diff -u\" of \"od -t x1\" output:\n\n\t--- good-object.hex     2008-08-25 16:33:41.000000000 -0700\n\t+++ bad-object.hex      2008-08-25 16:33:50.000000000 -0700\n\t@@ -113,16 +113,16 @@\n\t 0003400 e1 b1 dc ef 45 fa b1 ec 0f a0 2c 72 1b 25 55 f7\n\t 0003420 5a d6 f7 5c 72 dd 8b b2 44 dd 25 e6 c6 33 09 98\n\t 0003440 03 9b d1 49 8e 30 e6 7a df 76 76 bb cc c9 fb 44\n\t-0003460 a0 14 8d 7d cd 12 75 81 2b c4 d9 71 27 62 ae b5\n\t-0003500 87 a9 4f 76 c5 90 82 04 2b 15 53 47 ad 97 d1 02\n\t-0003520 98 14 df 45 7c 5f 93 b1 08 fd b6 bf e1 62 1b 00\n\t-0003540 b5 f6 ee 6c 70 30 78 73 74 38 3c 7f f3 e3 e5 90\n\t-0003560 39 7e 2e df e5 e1 c5 c1 29 62 02 fd 61 36 61 7d\n\t-0003600 c0 83 a0 4f 17 04 02 c6 ec 59 2e c1 7d 20 41 b1\n\t-0003620 fd 51 56 d0 b1 74 b9 76 81 4c 47 63 f0 10 76 15\n\t-0003640 db a7 0b b8 d1 6a f8 59 73 8c c4 e5 c9 ab 8b cb\n\t-0003660 37 32 5d 4a 91 c3 c1 c9 77 33 b4 2c b7 61 a0 37\n\t-0003700 47 f7 9b 3d b6 a5 99 5a cf 49 ef e0 3f 54 08 f4\n\t+0003460 a0 14 8d 7d 22 00 00 00 6b 57 fe ff 55 57 fe ff\n\t+0003500 aa 57 fe ff b0 00 0e 00 d9 66 22 00 00 00 00 e0\n\t+0003520 19 9f fe ff ff ff d1 43 fe ff d1 43 fe ff 00 c0\n\t+0003540 bf fe f0 33 69 bf b2 33 69 bf 00 00 00 10 01 fe\n\t+0003560 bf fe 66 01 00 00 01 1f 37 17 2e 59 fe ff 00 00\n\t+0003600 00 00 c2 41 fe ff 00 00 01 1f 37 17 44 42 fe ff\n\t+0003620 00 01 fe ff 01 1f 37 07 d4 42 fe ff 02 01 a8 44\n\t+0003640 00 f8 bf fe 6b 57 fe ff 55 57 fe ff 97 57 fe ff\n\t+0003660 00 f8 bf fe 04 00 fe ff ab 45 fe ff 02 00 fe ff\n\t+0003700 ae 41 49 50 aa 40 fe ff cf 49 ef e0 3f 54 08 f4\n\t 0003720 1e 2b 6f 59 e5 47 06 f4 80 a7 76 3a 38 a7 88 0a\n\t 0003740 30 72 a2 0a ed e1 b0 d5 6d 6f 9e ca 96 c7 28 d0\n\t 0003760 2c 19 b6 db 51 97 a9 f6 c3 31 8a 50 22 87 7a 97\n\nand the difference literally seems to be just a block of less than 160 \nbytes. But it's certainly not a single-bit error, and it's also not a disk \nblock or anything like that.\n\nLooking at the first line that differs, they are:\n\n\t-0003460 a0 14 8d 7d cd 12 75 81 2b c4 d9 71 27 62 ae b5\n\t+0003460 a0 14 8d 7d 22 00 00 00 6b 57 fe ff 55 57 fe ff\n\nso the differences start at offset 03464. The last line is:\n\n\t-0003700 47 f7 9b 3d b6 a5 99 5a cf 49 ef e0 3f 54 08 f4\n\t+0003700 ae 41 49 50 aa 40 fe ff cf 49 ef e0 3f 54 08 f4\n\nso they re-join at 03708 (octal offsets because of 'od' behavior).  But in \nbetween there seems to be nothing in common.\n\nSo the corrupt data looks like\n\n\t                    22 00 00 00 6b 57 fe ff 55 57 fe ff\n        0003500 aa 57 fe ff b0 00 0e 00 d9 66 22 00 00 00 00 e0\n        0003520 19 9f fe ff ff ff d1 43 fe ff d1 43 fe ff 00 c0\n        0003540 bf fe f0 33 69 bf b2 33 69 bf 00 00 00 10 01 fe\n        0003560 bf fe 66 01 00 00 01 1f 37 17 2e 59 fe ff 00 00\n        0003600 00 00 c2 41 fe ff 00 00 01 1f 37 17 44 42 fe ff\n        0003620 00 01 fe ff 01 1f 37 07 d4 42 fe ff 02 01 a8 44\n        0003640 00 f8 bf fe 6b 57 fe ff 55 57 fe ff 97 57 fe ff\n        0003660 00 f8 bf fe 04 00 fe ff ab 45 fe ff 02 00 fe ff\n        0003700 ae 41 49 50 aa 40 fe ff \n\nand I don't see what the patern is, except that there's a lot of \"fe ff\" \nin there. 148 bytes, no obvious source. It definitely does _not_ look like \ncompressed input (it's too regular to be something zlib spits out), nor is \nit text. Nor is it valid x86 assembly, so it's not some code-segment data \neither.\n\nAnd as far as I can tell, that's the _only_ corruption in the whole file, \nbut I didn't really double-check.\n\nDoes anybody see a pattern?\n\n\t\t\tLinus\n"},{"id":"88648","messageId":"48B46B04.70102@gmail.com","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808251616240.3363@nehalem.linux-foundation.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Jason McMullan","fromEmail":"jason.mcmullan@gmail.com","sentAt":"2008-08-26T20:43:48Z","receivedAt":"2008-08-26T20:43:48Z","isPatch":false,"sender":{"key":"jason.mcmullan@gmail.com","avatar":null},"body":"Linus Torvalds wrote:\n> So the corrupt data looks like\n> \n> [snip]\n> \n> And as far as I can tell, that's the _only_ corruption in the whole file, \n> but I didn't really double-check.\n> \n> Does anybody see a pattern?\n> \n\nWas this pack created on a journaled file system? Reiserfs? Ext3?\n\nIf there's journal corruption in a commonly used filesystem,\nThat Would Be Bad.\n\nI would suspect the Reiserfs 'file tail' behavour and journalling\nhave something to do with it, only because I still use and like\nReiserV3+LVM (dynamic grow and offline shrink baby!).\n\nOnly something like silently knifing files in the back would cause me\nto leave my beloved RedrumFS.\n\nWhich I probably should. Eventually. Once ZFS changes it's license. HA!\n\n- Jason McMullan\n"},{"id":"88649","messageId":"20080826205552.GR4380@fieldses.org","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808251616240.3363@nehalem.linux-foundation.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2008-08-26T20:55:52Z","receivedAt":"2008-08-26T20:55:52Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Mon, Aug 25, 2008 at 04:59:26PM -0700, Linus Torvalds wrote:\n> and the difference literally seems to be just a block of less than 160 \n> bytes. But it's certainly not a single-bit error, and it's also not a disk \n> block or anything like that.\n> \n> Looking at the first line that differs, they are:\n> \n> \t-0003460 a0 14 8d 7d cd 12 75 81 2b c4 d9 71 27 62 ae b5\n> \t+0003460 a0 14 8d 7d 22 00 00 00 6b 57 fe ff 55 57 fe ff\n> \n> so the differences start at offset 03464. The last line is:\n> \n> \t-0003700 47 f7 9b 3d b6 a5 99 5a cf 49 ef e0 3f 54 08 f4\n> \t+0003700 ae 41 49 50 aa 40 fe ff cf 49 ef e0 3f 54 08 f4\n> \n> so they re-join at 03708 (octal offsets because of 'od' behavior).  But in \n> between there seems to be nothing in common.\n> \n> So the corrupt data looks like\n> \n>                             22 00 00 00 6b 57 fe ff 55 57 fe ff\n>         0003500 aa 57 fe ff b0 00 0e 00 d9 66 22 00 00 00 00 e0\n>         0003520 19 9f fe ff ff ff d1 43 fe ff d1 43 fe ff 00 c0\n>         0003540 bf fe f0 33 69 bf b2 33 69 bf 00 00 00 10 01 fe\n>         0003560 bf fe 66 01 00 00 01 1f 37 17 2e 59 fe ff 00 00\n>         0003600 00 00 c2 41 fe ff 00 00 01 1f 37 17 44 42 fe ff\n>         0003620 00 01 fe ff 01 1f 37 07 d4 42 fe ff 02 01 a8 44\n>         0003640 00 f8 bf fe 6b 57 fe ff 55 57 fe ff 97 57 fe ff\n>         0003660 00 f8 bf fe 04 00 fe ff ab 45 fe ff 02 00 fe ff\n>         0003700 ae 41 49 50 aa 40 fe ff \n> \n> and I don't see what the patern is, except that there's a lot of \"fe ff\" \n> in there.\n\nand all aligned on 2-byte boundaries?  I don't see anything suggestive\nthere either.\n\n> 148 bytes, no obvious source. It definitely does _not_ look like \n> compressed input (it's too regular to be something zlib spits out), nor is \n> it text. Nor is it valid x86 assembly, so it's not some code-segment data \n> either.\n> \n> And as far as I can tell, that's the _only_ corruption in the whole file, \n> but I didn't really double-check.\n> \n> Does anybody see a pattern?\n\nGot me.  Thanks for taking a look at it.\n\n--b.\n"},{"id":"88652","messageId":"48B46F46.9090302@gmail.com","threadId":"15196","inReplyTo":"48B46B04.70102@gmail.com","subject":"Re: \"failed to read delta base object at...\"","fromName":"Jason McMullan","fromEmail":"jason.mcmullan@gmail.com","sentAt":"2008-08-26T21:01:58Z","receivedAt":"2008-08-26T21:01:58Z","isPatch":false,"sender":{"key":"jason.mcmullan@gmail.com","avatar":null},"body":"Jason McMullan wrote:\n> \n> Was this pack created on a journaled file system? Reiserfs? Ext3?\n> \n> If there's journal corruption in a commonly used filesystem,\n> That Would Be Bad.\n\nIn private mail, it was indicated that this was an ext3 with the\ndata=writeback option. From mount(1), man page, ext3 section:\n\n    writeback\n       Data ordering is not preserved - data may be written into\n       the main file system after its metadata has been  commit‐\n       ted  to the journal.  This is rumoured to be the highest-\n       throughput option.  It guarantees  internal  file  system\n       integrity,  however  it  can  allow old data to appear in\n       files after a crash and journal recovery.\n\nAll bets are off when data=writeback.\n\n\n- Jason McMullan\n"},{"id":"88732","messageId":"alpine.LFD.1.10.0808270937340.3363@nehalem.linux-foundation.org","threadId":"15196","inReplyTo":"48B46F46.9090302@gmail.com","subject":"Re: \"failed to read delta base object at...\"","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-08-27T17:05:54Z","receivedAt":"2008-08-27T17:05:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 26 Aug 2008, Jason McMullan wrote:\n> \n> All bets are off when data=writeback.\n\nNot the way git writes pack-files. It does a fsync() before moving them \ninto place (at least newer git versions do), so the data is stable.\n\nI do worry about wild pointers. I can't recognize the data, and it \ndefinitely doesn't look like any git internal data structures, but 16-bit \ndata _is_ what zlib internally uses for things like the decoding tables. \n\nSo if there is some use-after-free issue, I could imagine things like this \nhappening inside of git. People do occasionally run valgrind on git, \nthough, and it's been clean in the past, but I don't know if that has ever \nbeen done on the threaded packing, for example.\n\nFor example, the corrupting data had patterns like this:\n\n\t00 f8 bf fe 6b 57 fe ff 55 57 fe ff 97 57 fe ff\n\nwhere the pattern _could_ be something like\n\n\t{ 00 f8 febf },\n\t{ 6b 57 fffe },\n\t{ 55 57 fffe },\n\t{ 97 57 fffe },\n\nassuming that the \"fe ff\" pattern really is meaningful and is a 16-bit \nlittle-endian word.\n\nAnd the thign is, zlib \"code\" tables look exactly like that:\n\n\ttypedef struct {\n\t    unsigned char op;           /* operation, extra bits, table bits */\n\t    unsigned char bits;         /* bits in this part of the code */\n\t    unsigned short val;         /* offset in table or code value */\n\t} code;\n\n\t/* op values as set by inflate_table():\n\t    00000000 - literal\n\t    0000tttt - table link, tttt != 0 is the number of table index bits\n\t    0001eeee - length or distance, eeee is the number of extra bits\n\t    01100000 - end of block\n\t    01000000 - invalid code\n\t */\n\nbut those particular op/val things don't make sense in that context \neither. But I don't know zlib that well, maybe the deflate routines use \nsome other model.\n\n\t\t\tLinus\n"},{"id":"88758","messageId":"alpine.LFD.1.10.0808271458320.1624@xanadu.home","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808270937340.3363@nehalem.linux-foundation.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-08-27T19:17:05Z","receivedAt":"2008-08-27T19:17:05Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 27 Aug 2008, Linus Torvalds wrote:\n\n> \n> \n> On Tue, 26 Aug 2008, Jason McMullan wrote:\n> > \n> > All bets are off when data=writeback.\n> \n> Not the way git writes pack-files. It does a fsync() before moving them \n> into place (at least newer git versions do), so the data is stable.\n\nAnd isn't the bad data block size and alignment a bit odd for a \nfilesystem crash corruption?\n\n> I do worry about wild pointers. I can't recognize the data, and it \n> definitely doesn't look like any git internal data structures, but 16-bit \n> data _is_ what zlib internally uses for things like the decoding tables. \n> \n> \n> So if there is some use-after-free issue, I could imagine things like this \n> happening inside of git. People do occasionally run valgrind on git, \n> though, and it's been clean in the past, but I don't know if that has ever \n> been done on the threaded packing, for example.\n\nHowever, in the pack-objects case, it is almost impossible to have such \na corruption since the data is SHA1 summed immediately before being \nwritten out.  In the index-pack case, it is even less likely since the \ndata is SHA1 summed immediately after being written out.  So there is no \nwindow for random pointer access corrupting the data without also \ninfluencing the pack checksum outcome.  Therefore, if the pack content \ndoesn't match its checksum, then the corruption must have occurred \noutside of git, otherwise the checksum would have included corrupted \ndata already and the pack checksum would match.\n\n\nNicolas\n"},{"id":"88764","messageId":"alpine.LFD.1.10.0808271222250.3363@nehalem.linux-foundation.org","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808271458320.1624@xanadu.home","subject":"Re: \"failed to read delta base object at...\"","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-08-27T19:48:00Z","receivedAt":"2008-08-27T19:48:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Aug 2008, Nicolas Pitre wrote:\n> \n> And isn't the bad data block size and alignment a bit odd for a \n> filesystem crash corruption?\n\nYes. If it was a filesystem issue, I'd expect it to be at least disk block \naligned (512 bytes, most of the time) and more likely filesystem block \naligned (ie mostly 4kB).\n\nHowever, if we were to re-write the file afterwards, it could still get \nnon-block-aligned corruption - simply because there was a \nnon-block-aligned rewrite that got lost. But we don't actually ever do \nthat, except for the header and the SHA1 at the end in some unusual cases.\n\n> However, in the pack-objects case, it is almost impossible to have such \n> a corruption since the data is SHA1 summed immediately before being \n> written out.\n\nYes. Anything that uses the \"sha1write()\" model (which includes the \nregular pack-file _and_ the index) should generally be pretty safe. \n\nHowever, we do have this odd case of fixing up the pack after-the-fact \nwhen we receive it from somebody else (because we get a thin pack and \ndon't know how many objects the final result will have). And that case \nseems to be not as safe, because it\n\n - re-reads the file to recompute the SHA1\n\n   This is understandable, and it's fairly ok, but it does mean that there \n   is a bigger chance of the SHA1 matching if something has corrupted the \n   file in the meantime!\n\n   (That was not the case of this corruption, obviously, since the SHA1 \n   didn't match)\n\n - but it also forgets to fsync the result, because it only did that in \n   one path rather in all cases of fixup.\n\n   Again, this wasn't actually the cause of this corruption, because the \n   corruption wasn't near the header or tail, so if it had been due to a \n   missed write due to missing an fsync, the pattern would have been \n   different.\n\nAnyway, we should fix the latter problem regardless, even if it's (a) damn \nunlikely and (b) definietly not the case in this thing.\n\nThe fix is trivial - just move the \"fsync_or_die()\" into the fixup routine \nrather than doing it in one of the callers.\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n\n\t\tLinus\n\n---\n builtin-pack-objects.c |    1 -\n pack-write.c           |    1 +\n 2 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 2dadec1..d394c49 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -499,7 +499,6 @@ static void write_pack_file(void)\n \t\t} else {\n \t\t\tint fd = sha1close(f, NULL, 0);\n \t\t\tfixup_pack_header_footer(fd, sha1, pack_tmp_name, nr_written);\n-\t\t\tfsync_or_die(fd, pack_tmp_name);\n \t\t\tclose(fd);\n \t\t}\n \ndiff --git a/pack-write.c b/pack-write.c\nindex a8f0269..ddcfd37 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -179,6 +179,7 @@ void fixup_pack_header_footer(int pack_fd,\n \n \tSHA1_Final(pack_file_sha1, &c);\n \twrite_or_die(pack_fd, pack_file_sha1, 20);\n+\tfsync_or_die(pack_fd, pack_name);\n }\n \n char *index_pack_lockfile(int ip_out)\n"},{"id":"88771","messageId":"7vy72inrlj.fsf@gitster.siamese.dyndns.org","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808251616240.3363@nehalem.linux-foundation.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-27T20:14:00Z","receivedAt":"2008-08-27T20:14:00Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> And as far as I can tell, that's the _only_ corruption in the whole file, \n> but I didn't really double-check.\n\nJust FYI, replacing these 3619 bytes in JBF's packfile with the good\nobject with\n\n\tdd conv=notrunc bs=1 seek=3025976 count=3619\n\nmakes the fixed pack pass verify-pack, so this part seems to be the only\ncorruption.\n"},{"id":"88779","messageId":"alpine.LFD.1.10.0808271627540.1624@xanadu.home","threadId":"15196","inReplyTo":"alpine.LFD.1.10.0808271222250.3363@nehalem.linux-foundation.org","subject":"Re: \"failed to read delta base object at...\"","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-08-27T20:46:38Z","receivedAt":"2008-08-27T20:46:38Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 27 Aug 2008, Linus Torvalds wrote:\n\n> On Wed, 27 Aug 2008, Nicolas Pitre wrote:\n> \n> > However, in the pack-objects case, it is almost impossible to have such \n> > a corruption since the data is SHA1 summed immediately before being \n> > written out.\n> \n> Yes. Anything that uses the \"sha1write()\" model (which includes the \n> regular pack-file _and_ the index) should generally be pretty safe. \n\nWhat that means is that if git was the cause of the corruption itself \nthen the pack would still match its checksum (verify-pâck would still \nfail nevertheless).\n\n> However, we do have this odd case of fixing up the pack after-the-fact \n> when we receive it from somebody else (because we get a thin pack and \n> don't know how many objects the final result will have). And that case \n> seems to be not as safe, because it\n> \n>  - re-reads the file to recompute the SHA1\n> \n>    This is understandable, and it's fairly ok, but it does mean that there \n>    is a bigger chance of the SHA1 matching if something has corrupted the \n>    file in the meantime!\n\nI think that can be fixed.  When reading the file back, it is possible \nto compute 2 sha1s: one to compare with the recieved one using original \npack header, and the second which would be the final one.  FRom a \ncertain offset, new objects were added, so that first sha1 is validated \nagainst the received one and reset, and at the end, it should correspond \nto the sha1 of added objects that we should compute when writing them.\n\n\nNicolas\n"}]}