{"thread":{"id":"24628","subject":"Inspecting a corrupt git object","startedAt":"2010-08-04T09:25:30Z","lastAt":"2010-08-04T13:09:57Z","messageCount":6,"participants":["Magnus Bäck","Alejandro Riveira Fernández","Thomas Rast","Holger Hellmuth"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"147099","messageId":"20100804092530.GA30070@jpl.local","threadId":"24628","inReplyTo":null,"subject":"Inspecting a corrupt git object","fromName":"Magnus Bäck","fromEmail":"magnus.back@sonyericsson.com","sentAt":"2010-08-04T09:25:30Z","receivedAt":"2010-08-04T09:25:30Z","isPatch":false,"sender":{"key":"magnus.back@sonyericsson.com","avatar":null},"body":"We recently discovered a git tree object corruption in one of our\nbusiest gits on the master server. From what I can tell \"git cat-file -p\"\noutput looked just fine, but \"git gc\" complained loudly about the object\nbeing corrupt. I had the same git cloned on my machine and found (after\nunpacking the packfiles) that my object was different from the one on\nthe server. Same size and everything, but the second byte (and only the\nsecond byte) differed between good and bad object.\n\n$ head -n 5 /tmp/hexdump_corrupt.txt\n00000000  78 9c 2b 29 4a 4d 55 30  32 36 62 30 34 30 30 33 |x.+)JMU026b04003|\n00000010  31 51 70 cc 4b 29 ca cf  4c d1 cb cd 66 a8 38 dd |1Qp.K)..L...f.8.|\n00000020  76 77 82 ba af da a1 66  06 b9 b4 03 66 9d 27 18 |vw.....f....f.'.|\n00000030  93 ec 50 55 f9 26 e6 65  a6 a5 16 97 e8 55 e4 e6 |..PU.&.e.....U..|\n00000040  30 d8 98 fe a9 93 98 cc  be 24 a4 ac 93 3b 43 b7 |0........$...;C.|\n$ head -n 5 /tmp/hexdump_okay.txt\n00000000  78 01 2b 29 4a 4d 55 30  32 36 62 30 34 30 30 33 |x.+)JMU026b04003|\n00000010  31 51 70 cc 4b 29 ca cf  4c d1 cb cd 66 a8 38 dd |1Qp.K)..L...f.8.|\n00000020  76 77 82 ba af da a1 66  06 b9 b4 03 66 9d 27 18 |vw.....f....f.'.|\n00000030  93 ec 50 55 f9 26 e6 65  a6 a5 16 97 e8 55 e4 e6 |..PU.&.e.....U..|\n00000040  30 d8 98 fe a9 93 98 cc  be 24 a4 ac 93 3b 43 b7 |0........$...;C.|\n\nFrom what I gather from the community book and Pro Git, a git object\nfile is a deflated representation of the object type as a string, the\npayload size, a null byte, and the payload. Is there a standard tool for\ninflating the file back so that I can inspect what the actual difference\nbetween these two are? Short of writing a tool utilizing zlib, at least.\n\nAny other ideas why we would see such a difference? Hardware\nmalfunction or memory corruption I guess, but something else?\nI can supply the actual object files if necessary.\n\n-- \nMagnus Bäck                      Opinions are my own and do not necessarily\nSW Configuration Manager         represent the ones of my employer, etc.\nSony Ericsson\n"},{"id":"147100","messageId":"i3bd0r$g2l$1@dough.gmane.org","threadId":"24628","inReplyTo":"20100804092530.GA30070@jpl.local","subject":"Re: Inspecting a corrupt git object","fromName":"Alejandro Riveira Fernández","fromEmail":"ariveira@gmail.com","sentAt":"2010-08-04T09:48:11Z","receivedAt":"2010-08-04T09:48:11Z","isPatch":false,"sender":{"key":"ariveira@gmail.com","avatar":null},"body":"On Wed, 04 Aug 2010 11:25:30 +0200, Magnus Bäck wrote:\n\n[ ... ]\n> \n> From what I gather from the community book and Pro Git, a git object\n> file is a deflated representation of the object type as a string, the\n> payload size, a null byte, and the payload. Is there a standard tool for\n> inflating the file back so that I can inspect what the actual difference\n> between these two are? Short of writing a tool utilizing zlib, at least.\n\n Maybe\n\n git cat-file -p <sha1>\n \n ?\n\n> \n> Any other ideas why we would see such a difference? Hardware malfunction\n> or memory corruption I guess, but something else? I can supply the\n> actual object files if necessary.\n\nAlejandro\n"},{"id":"147101","messageId":"201008041148.49668.trast@student.ethz.ch","threadId":"24628","inReplyTo":"20100804092530.GA30070@jpl.local","subject":"Re: Inspecting a corrupt git object","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2010-08-04T09:48:49Z","receivedAt":"2010-08-04T09:48:49Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Magnus Bäck wrote:\n> \n> $ head -n 1 /tmp/hexdump_corrupt.txt\n> 00000000  78 9c 2b 29 4a 4d 55 30  32 36 62 30 34 30 30 33 |x.+)JMU026b04003|\n> $ head -n 1 /tmp/hexdump_okay.txt\n> 00000000  78 01 2b 29 4a 4d 55 30  32 36 62 30 34 30 30 33 |x.+)JMU026b04003|\n> \n> From what I gather from the community book and Pro Git, a git object\n> file is a deflated representation of the object type as a string, the\n> payload size, a null byte, and the payload. Is there a standard tool for\n> inflating the file back so that I can inspect what the actual difference\n> between these two are? Short of writing a tool utilizing zlib, at least.\n\nI'm sure it's a one-liner in almost any scripting language, e.g. you\ncan use\n\n  python -c 'import sys,zlib; sys.stdout.write(zlib.decompress(open(sys.argv[1]).read()))'\n\nwith a filename argument if you have Python at hand.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"147105","messageId":"4C594AE3.4000708@ira.uka.de","threadId":"24628","inReplyTo":"20100804092530.GA30070@jpl.local","subject":"Re: Inspecting a corrupt git object","fromName":"Holger Hellmuth","fromEmail":"hellmuth@ira.uka.de","sentAt":"2010-08-04T11:11:31Z","receivedAt":"2010-08-04T11:11:31Z","isPatch":false,"sender":{"key":"hellmuth@ira.uka.de","avatar":null},"body":"Magnus Bäck schrieb:\n> Any other ideas why we would see such a difference? Hardware\n> malfunction or memory corruption I guess, but something else?\n> I can supply the actual object files if necessary.\n> \n\nI checked with a repository here and all objects seem to start with 78\n01. That means it is a common prefix. Ergo no malicious tampering, as\nthat would make only sense if the contents of the blob had changed.\n\nSo a random hardware or software malfunction is left as explanation IMHO\n\nHolger\n"},{"id":"147111","messageId":"20100804130229.GA1537@jpl.local","threadId":"24628","inReplyTo":"201008041148.49668.trast@student.ethz.ch","subject":"Re: Inspecting a corrupt git object","fromName":"Magnus Bäck","fromEmail":"magnus.back@sonyericsson.com","sentAt":"2010-08-04T13:02:29Z","receivedAt":"2010-08-04T13:02:29Z","isPatch":false,"sender":{"key":"magnus.back@sonyericsson.com","avatar":null},"body":"On Wednesday, August 04, 2010 at 11:48 CEST,\n     Thomas Rast <trast@student.ethz.ch> wrote:\n\n> Magnus Bäck wrote:\n>\n> > From what I gather from the community book and Pro Git, a git object\n> > file is a deflated representation of the object type as a string,\n> > the payload size, a null byte, and the payload. Is there a standard\n> > tool for inflating the file back so that I can inspect what the\n> > actual difference between these two are? Short of writing a tool\n> > utilizing zlib, at least.\n> \n> I'm sure it's a one-liner in almost any scripting language, e.g. you\n> can use\n> \n>   python -c 'import sys,zlib; sys.stdout.write(zlib.decompress(open(sys.argv[1]).read()))'\n> \n> with a filename argument if you have Python at hand.\n\nThat worked fine, thanks. Apparently this difference in the second byte\nof the compressed data makes no difference for the end result -- the two\ninflated files are identical.\n\nInterestingly, just as we were about to transplant the loose object from\nmy working repository to the server where \"git gc\" failed and the object\nwas seemingly corrupt, the person doing the actual work (I don't have\naccess to the server) ran \"git gc\" to find the id of the bad object, and\nsuddenly it completed without errors. The object in question had now\nbeen included in a packfile, and upon unpacking that packfile to inspect\nthe object it was identical to the file I had, i.e. the new loose object\nwas different from the original loose object. I had expected a loose\nobject -> packfile -> loose object cycle to not change anything.\nEverything seems to be back to normal now, which is good, but I prefer\nI understand why things get fixed.\n\nWe did have some initial problems with reaching the per-process limit\nfor open files (as no repack had been done for an extended time and 5000\npackfiles were lingering), but it seems weird for such a problem to be\nrelated to the possible corruptness of a single tree object.\n\n-- \nMagnus Bäck                      Opinions are my own and do not necessarily\nSW Configuration Manager         represent the ones of my employer, etc.\nSony Ericsson\n"},{"id":"147112","messageId":"20100804130957.GB1537@jpl.local","threadId":"24628","inReplyTo":"i3bd0r$g2l$1@dough.gmane.org","subject":"Re: Inspecting a corrupt git object","fromName":"Magnus Bäck","fromEmail":"magnus.back@sonyericsson.com","sentAt":"2010-08-04T13:09:57Z","receivedAt":"2010-08-04T13:09:57Z","isPatch":false,"sender":{"key":"magnus.back@sonyericsson.com","avatar":null},"body":"On Wednesday, August 04, 2010 at 11:48 CEST,\n     Alejandro Riveira Fernández <ariveira@gmail.com> wrote:\n\n> On Wed, 04 Aug 2010 11:25:30 +0200, Magnus Bäck wrote:\n>\n> > From what I gather from the community book and Pro Git, a git object\n> > file is a deflated representation of the object type as a string,\n> > the payload size, a null byte, and the payload. Is there a standard\n> > tool for inflating the file back so that I can inspect what the\n> > actual difference between these two are? Short of writing a tool\n> > utilizing zlib, at least.\n>\n>  Maybe\n>\n>  git cat-file -p <sha1>\n>\n>  ?\n\nSorry, I should've been more clear here. I know about cat-file's\npretty-printing abilities, but I just wanted to inflate the loose\nobject data and see *exactly* where the differing byte ended up.\n\n-- \nMagnus Bäck                      Opinions are my own and do not necessarily\nSW Configuration Manager         represent the ones of my employer, etc.\nSony Ericsson\n"}]}