{"thread":{"id":"7009","subject":"My git repo is broken, how to fix it ?","startedAt":"2007-02-28T04:36:30Z","lastAt":"2007-03-23T03:55:22Z","messageCount":28,"participants":["Alexander Litvinov","Linus Torvalds","Alex Riesen","Junio C Hamano","Nicolas Pitre","Johannes Sixt","Jeff King","Bill Lear"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"35773","messageId":"200702281036.30539.litvinov2004@gmail.com","threadId":"7009","inReplyTo":null,"subject":"My git repo is broken, how to fix it ?","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-02-28T04:36:30Z","receivedAt":"2007-02-28T04:36:30Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"Hello,\n\nI use manualy compiled git under cygwin. Some time ago I have imported project \nfrom CVS and start to us it under git. To emulate 'separate remotes' schema \nfor branches from CVS I clone imported git repo and work with it. From time \nto time I incrementaly update imported repo from cvs and sometimes use \ngit-cvsexportcommit (from work repo) to export my changes and then get them \nusing git-cvsimport.\n\nSome times ago I descide to run fsck and found that by working repo is broken, \nwhile imported repo is correct. Is there way to fix it ? \n\n>git version\ngit version 1.5.0.GIT\n(It was actualy compiled from e86d552 commit)\n> git fsck\n> git fsck --full\nerror: packed 7f5fed8131fb32972c602dede29b9257a053ba67 \nfrom .git/objects/pack/pack-c4554978bbe079c9a43d6a13546a2fa314fe0884.pack is \ncorrupt\nsha1 mismatch 7f5fed8131fb32972c602dede29b9257a053ba67\n\n(This is a blob, git cat-file blob 7f5fed813 shows me my c++ header file that \nis partialy broken with ^@ symbols)\n\nThe repo I get using git-cvsimport is correct and does not contains that blob. \nI also tried git-log -p for all by branches to force git to show me what is \nthe commit was broken but git-log finished without errors.\n\nBy the way, several times I interrupt git's commands like commit and pull \nusing Ctrl-C.\n\nI tried to unpack all objects:\n> git-unpack-objects -r \n< .git/objects/pack/pack-c4554978bbe079c9a43d6a13546a2fa314fe0884.pack; echo \n$?\nUnpacking 12868 objects\n 100% (12868/12868) done\n0\n\nNo erorts here. But fsck find that broken blob:\n> git fsck \ndangling blob beb992198d4d8813ea51fd1cbbf38313ef490c22\n\ngit-cat-file shows me this this is a broken object with correct sha1 sum.\n\n\nAs a cunclusion: my repo has broken file and I don't see there is the brakage. \nCan I reconstruct file by sha1 sum :-) or can I do something to stop fsck \nwarn me ?\n"},{"id":"35776","messageId":"Pine.LNX.4.64.0702272039540.12485@woody.linux-foundation.org","threadId":"7009","inReplyTo":"200702281036.30539.litvinov2004@gmail.com","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-28T04:57:28Z","receivedAt":"2007-02-28T04:57:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 28 Feb 2007, Alexander Litvinov wrote:\n> \n> Some times ago I descide to run fsck and found that by working repo is broken, \n> while imported repo is correct. Is there way to fix it ? \n\nGenerally, the best way to fix things is (I've written this up at \nsomewhat more length before, but I'm too lazy to find it):\n\n - back up all your state so that anything you do is re-doable if you \n   corrupt things more!\n\n - explode any corrupt pack-files\n\n   See \"man git-unpack-objects\", and in particular the \"-r\" flag. Also, \n   please realize that it only unpacks objects that aren't already \n   available, so you need to move the pack-file away from its normal \n   location first (otherwise git-unpack-objects will find all objects \n   that are in the pack-file in the pack-file itself, and not unpack \n   anything at all)\n\n - replace any broken and/or missing objects\n\n   This is the challenging part. Sometimes (hopefully often!) you can find \n   the missing objects in other copies of the repositories. At other \n   times, you may need to try to find the data some other way (for \n   example, maybe your checked-out copy contains the file content that \n   when hashed will be the missing object?).\n\n - make sure everything is happy with \"git-fsck --full\"\n\n - repack everything to get back to an efficient state again.\n\nAnd remember: git does _not_ make backups pointless. It hopefully makes \nbackups *easy* (since cloning and pulling is easy), but the basic need for \nbackups does not go away!\n\n> By the way, several times I interrupt git's commands like commit and pull \n> using Ctrl-C.\n\nShouldn't matter, at least as long as you are using the native git \nprotocol: git will create objects fully under a temporary name, and then \natomically rename things to their right names. \n\nUsing rsync and/or http may not be as safe.\n\nHOWEVER! I do not know how well Windows and/or cygwin does file renames. \nIf cygwin does a rename as a copy + delete, a lot of the safety \nassumptions just fly out the window.\n\n> I tried to unpack all objects:\n>\n> > git-unpack-objects -r < .git/objects/pack/pack-c4554978bbe079c9a43d6a13546a2fa314fe0884.pack; echo  $?\n> Unpacking 12868 objects\n>  100% (12868/12868) done\n\nOk, that's a good thing, but see above: I don't think anything should have \ngotten unpacked, because it found all objects already existing in the very \npack-file you tried to unpack.\n\nSo you might well need to do\n\n\tmv .git/objects/pack/pack-c4554978bbe079c9a43d6a13546a2fa314fe0884.pack oldpack\n\tgit-unpack-objects -r < oldpack\n\n(or rename the .idx file instead).\n\nAlternatively (and in many ways this migth be better when you're trying to \nrecover something) just create a totally *new* git repo, by doing\n\n\tmkdir new-repo\n\tcd new-repo\n\tgit init\n\tgit unpack-objects -r < ../other-repo/.git/pack/pack-.....pack\n\nand re-create the objects somewhere else - you can do all of this without \nat all disturbing the old repository (but you'd need to copy all the refs \nand all the loose objects by hand, of course!)\n\n> No erorts here. But fsck find that broken blob:\n> > git fsck \n> dangling blob beb992198d4d8813ea51fd1cbbf38313ef490c22\n> \n> git-cat-file shows me this this is a broken object with correct sha1 sum.\n> \n> As a cunclusion: my repo has broken file and I don't see there is the brakage. \n> Can I reconstruct file by sha1 sum :-) or can I do something to stop fsck \n> warn me ?\n\nYou didn't do \"--full\", so it's not looking inside your pack, so the fsck \nwasn't very interesting in this case.\n\nAnd no, you cannot reconstruct the file by sha1 sum, although you may be \nable to reconstruct the file some *other* way (by looking at the other \nblobs and remembering what the missing case is), and then you can \nobviously use the sha1-sum to *confirm* that you reconstructed the file \nexactly as it was!\n\nSo yes, reconstruction of missing objects is possible, but no, you can't \ndo it based purely based on SHA1, you need to base reconstruction on some \nother information. That's kind of what \"cryptographically secure hash\" \nmeans ;^p\n\n\t\t\tLinus\n"},{"id":"35796","messageId":"200702281754.42383.litvinov2004@gmail.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0702272039540.12485@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-02-28T11:54:42Z","receivedAt":"2007-02-28T11:54:42Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Wednesday 28 February 2007 10:57 Linus Torvalds написал(a):\n>  - replace any broken and/or missing objects\n>\n>    This is the challenging part. Sometimes (hopefully often!) you can find\n>    the missing objects in other copies of the repositories. At other\n>    times, you may need to try to find the data some other way (for\n>    example, maybe your checked-out copy contains the file content that\n>    when hashed will be the missing object?).\n\nThanks for answer. I have found this blob in cloned repo. I just copy it into \nobjects subdir and repack repo again. fsck works without any errors.\n\nThanks again.\n"},{"id":"35835","messageId":"Pine.LNX.4.64.0702280802150.12485@woody.linux-foundation.org","threadId":"7009","inReplyTo":"200702281754.42383.litvinov2004@gmail.com","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-28T16:19:50Z","receivedAt":"2007-02-28T16:19:50Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 28 Feb 2007, Alexander Litvinov wrote:\n>\n> >  - replace any broken and/or missing objects\n> >\n> >    This is the challenging part. Sometimes (hopefully often!) you can find\n> >    the missing objects in other copies of the repositories. At other\n> >    times, you may need to try to find the data some other way (for\n> >    example, maybe your checked-out copy contains the file content that\n> >    when hashed will be the missing object?).\n> \n> Thanks for answer. I have found this blob in cloned repo. I just copy it into \n> objects subdir and repack repo again. fsck works without any errors.\n\nGood to hear.\n\nIt would probably be good to\n\n - try to figure out why things got corrupted in the first place.\n\n   In particular, we should probably check our (my) assumption that the \n   file rename on cygwin is atomic. Your comment that you use ^C a lot \n   makes me worry that something basically caused an incomplete write or \n   other thing to happen..\n\n   To be more precise, git will actually start off trying to do a \"link + \n   unlink\" pair, because that is the safest thing to do on a UNIX \n   filesystem: if the linked target already exists, we won't overwrite it \n   (and the git data consistency rules have always been: honor existing \n   data over new one, and *never* change anything that has already been \n   written).\n\n   But I would not be shocked to hear that the \"link + unlink\" sequence \n   ends up being emulated under cygwin as a \"copy + delete\" due to lack of \n   hardlinks or something.\n\n   Also, even if the link fails, and git then falls back to \"rename()\" \n   (since some filesystems don't do hardlinks at all, or limit them to one \n   particular directory), I would _still_ not be totally surprised if the \n   rename got emulated as a copy/delete for some strange Windows reason.\n\n   There are other possibilities for corruption, of course: just plain \n   disk corruption, or (again) some other subtle cygwin emulation or \n   Windows issue could bite us. \n\n - Even under UNIX, I'm not entirely sure about http/ftp/rsync transfers. \n   rsync in particular doesn't check anything at all, but last I looked, \n   the http fetcher was also doing things like checking the integrity of \n   the object *after* it had already moved it to its final resting place \n   (which is again unsafe with ^C).\n\n   In general, I strongly suggest that people use the \"native git\" \n   pack-transfers. The \"dumb protocol\" transfers are called \"dumb\" for a \n   reason..\n\n - It would probably be good to write up the \"How to recover\" thing, \n   regardless of why any corruption happens. It doesn't matter if you're \n   under UNIX and using native protocols, and just being careful as hell: \n   disks get corrupted, sh*t happens, alpha-particles in the wrong place \n   do bad things to memory cells. And bugs _are_ inevitable, even if \n   we've been pretty damn good about these things.\n\n   So it's important for people to know what the limits on corruption are, \n   and tell people that regardless of how stable the git data structures \n   are, if you care about your data, you need to have things in multiple \n   places (and no, RAID is _not_ the answer either, even if it can be a \n   small _part_ of doing things well).\n\nAnybody?\n\n\t\tLinus\n"},{"id":"35870","messageId":"81b0412b0702281112y714c3b0eread09ee9440d0537@mail.gmail.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0702280802150.12485@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-02-28T19:12:03Z","receivedAt":"2007-02-28T19:12:03Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 2/28/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>  - try to figure out why things got corrupted in the first place.\n> ...\n>    But I would not be shocked to hear that the \"link + unlink\" sequence\n>    ends up being emulated under cygwin as a \"copy + delete\" due to lack of\n>    hardlinks or something.\n\nWell, cygwin has a whole rename implementation using wincalls (which\nmay or maybe not syscalls). It's hard to tell whether micros$#%^ used\nlink+unlink: windows traditionally had no link(2). It has it since W2K,\nallowed only on ntfs, and even there it spent some time undocumented :)\n\n>    Also, even if the link fails, and git then falls back to \"rename()\"\n>    (since some filesystems don't do hardlinks at all, or limit them to one\n>    particular directory), I would _still_ not be totally surprised if the\n>    rename got emulated as a copy/delete for some strange Windows reason.\n>\n>    There are other possibilities for corruption, of course: just plain\n>    disk corruption, or (again) some other subtle cygwin emulation or\n>    Windows issue could bite us.\n\nIt is very hard to tell: the rename function in cygwin is over 200 lines long,\nhas multiple calls down to the kernel (or whatever it is), and there are two\ndifferent wincalls used. It is hard to predict what path will be taken or\nwhether there were actually any data copying happened.\nOf course, windows known to corrupt data just so...\n\n>  - It would probably be good to write up the \"How to recover\" thing,\n>    regardless of why any corruption happens. It doesn't matter if you're\n>    under UNIX and using native protocols, and just being careful as hell:\n>    disks get corrupted, sh*t happens, alpha-particles in the wrong place\n>    do bad things to memory cells. And bugs _are_ inevitable, even if\n>    we've been pretty damn good about these things.\n\nIt'd be a good start to link the recent disaster recoveries from\ngit.or.cz wiki from the title page (unless it is already).\n"},{"id":"37496","messageId":"200703191932.26856.litvinov2004@gmail.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0702280802150.12485@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-03-19T13:32:26Z","receivedAt":"2007-03-19T13:32:26Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Wednesday 28 February 2007 22:19 Linus Torvalds написал(a):\n> On Wed, 28 Feb 2007, Alexander Litvinov wrote:\n> > Thanks for answer. I have found this blob in cloned repo. I just copy it\n> > into objects subdir and repack repo again. fsck works without any errors.\n>\n> Good to hear.\n\nHello, its me again.\n\nIt is pity but my repo was corrupted again. I have WinXP + cygwin + \ngit-1.5.0-572-ge86d552. I was doing \ngit-apply/git-am/git-reset/git-cvsexportcommit and broke repo somehow. I have \ntwo broken blobs that should be done by my recent patches.\n\nWill try to recover them and report the result.\n\nIs there any way to catch and solve the problem ?\nThanks for help.\n"},{"id":"37506","messageId":"Pine.LNX.4.64.0703190804350.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"200703191932.26856.litvinov2004@gmail.com","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-19T15:20:23Z","receivedAt":"2007-03-19T15:20:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 19 Mar 2007, Alexander Litvinov wrote:\n> \n> It is pity but my repo was corrupted again. I have WinXP + cygwin + \n> git-1.5.0-572-ge86d552. I was doing \n> git-apply/git-am/git-reset/git-cvsexportcommit and broke repo somehow. I have \n> two broken blobs that should be done by my recent patches.\n\nOk, can you send me just the two broken blobs? I assume that they are \nloose objects, so when fsck complaines about a corrupt object xyzzy..., \njust take those objects from\n\n\t.git/objects/xy/zzy..\n\nand tar the two broken ones up, and send it to me by email. I'll keep it \nprivate if need be, but if you don't care about that, it would be even \nbetter if you can send them to the list publicly so that others can see \nwhat the corruption looks like.\n\n> Is there any way to catch and solve the problem ?\n\nI'd like to see the objects to look at what the corruption looks like, but \nI suspect that it's cygwin and/or WinXP. I'm not at all convinced that \nWindows is all that safe in general when it comes to data consistency, and \nI suspect cygwin makes it much worse by making operations that *should* be \natomic be non-atomic.\n\nHere's a patch that is probably a good idea to try. It disables the \nhardlinking code for CYGWIN, and it also checks for errors from \"close()\". \nThose are the two most obvious issues that I could imagine causing \nproblems..\n\n\t\tLinus\n---\n sha1_file.c |   11 ++++++++++-\n 1 files changed, 10 insertions(+), 1 deletions(-)\n\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 372af60..d829dc7 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1760,6 +1760,14 @@ static void write_sha1_file_prepare(void *buf, unsigned long len,\n \tSHA1_Final(sha1, &c);\n }\n \n+#ifdef __CYGWIN__\n+static int link(const char *old, const char *new)\n+{\n+\terrno = ENOSYS;\n+\treturn -1;\n+}\n+#endif\n+\n /*\n  * Link the tempfile to the final place, possibly creating the\n  * last directory level as you do so.\n@@ -1951,7 +1959,8 @@ int write_sha1_file(void *buf, unsigned long len, const char *type, unsigned cha\n \tif (write_buffer(fd, compressed, size) < 0)\n \t\tdie(\"unable to write sha1 file\");\n \tfchmod(fd, 0444);\n-\tclose(fd);\n+\tif (close(fd))\n+\t\tdie(\"error closing sha1 file (%s)\", strerror(errno));\n \tfree(compressed);\n \n \treturn move_temp_to_file(tmpfile, filename);\n"},{"id":"37590","messageId":"Pine.LNX.4.64.0703192212280.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"200703201013.39169.litvinov2004@gmail.com","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-20T05:34:10Z","receivedAt":"2007-03-20T05:34:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 20 Mar 2007, Alexander Litvinov wrote:\n> \n> Actualy, I have packed that objects already, so fsck warn me:\n> $ git fsck --full\n> error: packed 8edc906985f00cf27180b1d9d4c3217ffd1896f8 from .git/objects/pack/pack-abc5cbabfc05c213e50c43ea07f43158bf1de236.pack is corrupt\n> error: packed f6aca57bb30a12e9ac5d71558e0b6052d6fb67a8 from .git/objects/pack/pack-abc5cbabfc05c213e50c43ea07f43158bf1de236.pack is corrupt\n> sha1 mismatch 8edc906985f00cf27180b1d9d4c3217ffd1896f8\n> sha1 mismatch f6aca57bb30a12e9ac5d71558e0b6052d6fb67a8\n\nOk, this is different from what I expected. \n\nSince your pack-file seems to pass its own internal SHA1 checks, it means \nthat it was likely corrupt already when it was written out in the pack. \nWhat's interesting is that it seems to unpack, but then the SHA1 of the \nunpacked object doesn't match.\n\nThe reason I say that's interesting is that it would seem to mean that the \nzlib CRC/adler check didn't trigger - which probably means that the object \nwas corrupted *before* it was compressed (but after it was originally \nSHA1-summed), or the compression itself was corrupting (eg a libz \nproblem).\n\nAnd since the SHA1 of the pack-file matches, the thing was apparently \nalso written out \"correctly\" after compression (but by that \"correctly\" I \nobviously mean that the *corrupted* data was written out). \n\nSadly, by the time it's in a pack-file, it is *really* hard to figure out \nwhat went wrong: I see your unpacked data, but it's really the packed raw \nobjects that I wanted to look at, in case there would be some pattern in \nthe actual corruption (the corruption will then result in random crud when \nactually unpacking, which is why the unpacked data isn't that interesting, \nsimple because there's no pattern left to analyze - it got inflated to \nbogus \"data\").\n\n> I also use autocrlf feature:\n> $ git config core.autocrlf\n> true\n\nI doubt autocrlf affects anything here, it's only used at checkin and \ncheckout time, and it wouldn't affect the raw internal git objects.\n\nMore interesting might be if you might be using any of the other flags \nthat actually affect internal git object packing: \"use_legacy_headers\" in \nparticular? If we have a bug there, that could be nasty.\n\nBut to really look at this we should probably add a \"really_careful\" flag \nthat actually re-verifies the SHA1 on read so that we'd catch these kinds \nof corruptions early. \n\n> This files are cpp code from our project and tham need to be private. Really.\n\nOk, no problem. I added back the git list (but not your attachments, \nobviously) but as explained above, there is not a lot I can do with the \nunpacked data, I'd like to see the actual \"raw\" stuff.\n\nI'm hoping somebody has any ideas. We really *could* check the SHA1 on \neach read (and slow down git a lot) and that would catch corruption much \nfaster and hopefully pinpoint it more quickly where exactly it happens. \nBut maybe somebody has some other smart idea?\n\n\t\tLinus\n"},{"id":"37602","messageId":"200703201255.22701.litvinov2004@gmail.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703192212280.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-03-20T06:55:22Z","receivedAt":"2007-03-20T06:55:22Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Tuesday 20 March 2007 11:34 Linus Torvalds написал:\n> Ok, this is different from what I expected.\n\nI will try to stop using git-gc for some time to find out broken loose \nobjects.\n\n> > I also use autocrlf feature:\n> > $ git config core.autocrlf\n> > true\n>\n> More interesting might be if you might be using any of the other flags\n> that actually affect internal git object packing: \"use_legacy_headers\" in\n> particular? If we have a bug there, that could be nasty.\nThis is the all my config options:\n$ git config -l\nuser.name=Alexander Litvinov\nuser.email=XXX\ncore.logallrefupdates=true\ncore.filemode=false\ncore.autocrlf=true\ndiff.color=auto\nstatus.color=auto\napply.whitespace=strip\ncore.repositoryformatversion=0\ncore.filemode=false\ncore.bare=false\nremote.origin.url=/home/lan/src/XXX\nremote.origin.fetch=+refs/heads/*:refs/remotes/origin/*\nbranch.master.remote=origin\nbranch.master.merge=refs/heads/master\nbranch.XXX.remote=origin\nbranch.XXX.merge=refs/heads/XXX\n\n> Ok, no problem. I added back the git list (but not your attachments,\n> obviously) but as explained above, there is not a lot I can do with the\n> unpacked data, I'd like to see the actual \"raw\" stuff.\n\nI undertand your wish.\n\n> I'm hoping somebody has any ideas. We really *could* check the SHA1 on\n> each read (and slow down git a lot) and that would catch corruption much\n> faster and hopefully pinpoint it more quickly where exactly it happens.\n\nI can live with such slowdown as far as cygwin not fast and I am ready to wait \nright now. I don't think the situation become realy worser than now :-)\n"},{"id":"37605","messageId":"7vd5349x97.fsf@assigned-by-dhcp.cox.net","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703192212280.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-03-20T07:42:28Z","receivedAt":"2007-03-20T07:42:28Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> But to really look at this we should probably add a \"really_careful\" flag \n> that actually re-verifies the SHA1 on read so that we'd catch these kinds \n> of corruptions early. \n> ...\n> I'm hoping somebody has any ideas. We really *could* check the SHA1 on \n> each read (and slow down git a lot) and that would catch corruption much \n> faster and hopefully pinpoint it more quickly where exactly it happens. \n\nAt least, we could do something like this to catch the breakage\nwhen we (re)pack, to prevent damage from propagating.\n\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 73d448b..5d0692a 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -65,6 +65,7 @@ static int no_reuse_delta;\n static int local;\n static int incremental;\n static int allow_ofs_delta;\n+static int revalidate_sha1;\n \n static struct object_entry **sorted_by_sha, **sorted_by_type;\n static struct object_entry *objects;\n@@ -974,8 +975,31 @@ static void add_preferred_base(unsigned char *sha1)\n \tit->pcache.tree_size = size;\n }\n \n-static void check_object(struct object_entry *entry)\n+static void check_object(struct object_entry *entry, int ith, unsigned *last)\n {\n+\tif (revalidate_sha1) {\n+\t\tunsigned char sha1[20];\n+\t\tenum object_type type;\n+\t\tunsigned long size;\n+\t\tvoid *buf;\n+\n+\t\tbuf = read_sha1_file(entry->sha1, &type, &size);\n+\t\thash_sha1_file(buf, size, typename(type), sha1);\n+\t\tif (hashcmp(sha1, entry->sha1))\n+\t\t\tdie(\"'%s': hash mismatch\", sha1_to_hex(entry->sha1));\n+\t\tfree(buf);\n+\n+\t\tif (progress) {\n+\t\t\tunsigned percent = ith * 100 / nr_objects;\n+\t\t\tif (percent != *last || progress_update) {\n+\t\t\t\tfprintf(stderr, \"%4u%% (%u/%u) done\\r\",\n+\t\t\t\t\tpercent, ith, nr_objects);\n+\t\t\t\tprogress_update = 0;\n+\t\t\t\t*last = percent;\n+\t\t\t}\n+\t\t}\n+\t}\n+\n \tif (entry->in_pack && !entry->preferred_base) {\n \t\tstruct packed_git *p = entry->in_pack;\n \t\tstruct pack_window *w_curs = NULL;\n@@ -1082,10 +1106,16 @@ static void get_object_details(void)\n {\n \tuint32_t i;\n \tstruct object_entry *entry;\n+\tunsigned last_percent = 999;\n+\n+\tif (progress && revalidate_sha1)\n+\t\tfprintf(stderr, \"Revalidating %u objects.\\n\", nr_objects);\n \n \tprepare_pack_ix();\n \tfor (i = 0, entry = objects; i < nr_objects; i++, entry++)\n-\t\tcheck_object(entry);\n+\t\tcheck_object(entry, i+1, &last_percent);\n+\tif (progress && revalidate_sha1)\n+\t\tfputc('\\n', stderr);\n \n \tif (nr_objects == nr_result) {\n \t\t/*\n@@ -1629,6 +1659,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\trp_av[1] = \"--objects-edge\";\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(\"--revalidate\", arg)) {\n+\t\t\trevalidate_sha1 = 1;\n+\t\t\tcontinue;\n+\t\t}\n \t\tusage(pack_usage);\n \t}\n \n"},{"id":"37622","messageId":"alpine.LFD.0.83.0703201113180.18328@xanadu.home","threadId":"7009","inReplyTo":"7vd5349x97.fsf@assigned-by-dhcp.cox.net","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-20T15:23:01Z","receivedAt":"2007-03-20T15:23:01Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 20 Mar 2007, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > But to really look at this we should probably add a \"really_careful\" flag \n> > that actually re-verifies the SHA1 on read so that we'd catch these kinds \n> > of corruptions early. \n> > ...\n> > I'm hoping somebody has any ideas. We really *could* check the SHA1 on \n> > each read (and slow down git a lot) and that would catch corruption much \n> > faster and hopefully pinpoint it more quickly where exactly it happens. \n> \n> At least, we could do something like this to catch the breakage\n> when we (re)pack, to prevent damage from propagating.\n\nI think it would be better to retest the SHA1 when we're about to \n_write_ the object out to the pack, replacing check_pack_inflate() and \nrevalidate_loose_object() with the full SHA1 check, and testing objects \nwhich data isn't reused from a pack too.  And make it conditional on \n!pack_to_stdout like we already do of course.\n\n\nNicolas\n"},{"id":"37740","messageId":"Pine.LNX.4.64.0703220847540.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"200703210956.50018.litvinov2004@gmail.com","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-22T15:58:36Z","receivedAt":"2007-03-22T15:58:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n[ Sorry, I got sidetracked, have been looking at kernel bugs and git \n  optimizations. I added back the git mailing list, since there is no \n  private file data here any more, and I'd like others to follow this\n  saga too ]\n\nOn Wed, 21 Mar 2007, Alexander Litvinov wrote:\n>\n> > Oh, btw, before I do that - do you by any chance have the *uncorrupt*\n> > version of the file that should be this object? In other words, do you\n> > have object 03312463e194d68d0d677b51e09b47cb29ca926a in another\n> > repository? It should be a version of your file.\n>\n> It is pity, but I don't have that version. It was broken at my last commit and \n> I will redo it again. I remeber the changes I have done.\n> \n> The new file has different sha1 sum but I attach it to show you the real file \n> content.\n\nSadly, I actually would need to compare it to the exact object it *should* \nhave generated, and that means that I'm not actually all that interested \nin the \"real file content\" per se, I really would need to see the exact \nblob, so that I can generate the object it should have been, and then \ncompare that against the corrupt one... \n\nIt's the binary data I'd like to compare, so that I can tell (for example) \nif there is just a chunk missing in the middle, or something like that. \nBut since even slight differences in the source data will lead to \ndifferent binary data, and since the compressed and corrupt data I have \nfrom you earlier doesn't make sense on its own, I do care about the exact \nobject that got corrupted.\n\n> By the way, I now I am using git (3ba7a10) taken from next with tree your \n> patches:\n> 1. disables the hardlinking code for CYGWIN, and it also checks for errors \n> from \"close()\"\n> 2. Don't ever return corrupt objects from \"parse_object()\"\n> 3. Be more careful about zlib return values.\n\nOk, apart from #1, those should be in current -git now, along with better \nvalidation checks (by Nico) when packing. So hopefully at least when there \nis corruption in a loose object, we will now always notice when we do a \n\"git repack\", and will never generate a broken pack-file. Knock wood.\n\nOf course, I actually wonder if the bug might be in your version of zlib \n(miscompiled or some other thing), in which case *any* amount of \npre-validation won't really help, because it will become corrupted when we \ndeflate it prior to writing. For example, if \"deflateBound()\" sometimes \ndoesn't give a valid upper bound and we allocate too little space..\n\n> Yesterday brakage was made by git with only first patch.\n\nI see that you seem to be able to reproduce it again - I'll answer that \nemail separately just so that the git list sees that message too. But it \nboils down to: if it's reproducible, I'd *really* like to see the \ncorrupted object and the exact file that it should have been generated \nfrom..\n\n\t\tLinus\n"},{"id":"37741","messageId":"Pine.LNX.4.64.0703220858400.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"200703211024.04740.litvinov2004@gmail.com","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-22T16:17:11Z","receivedAt":"2007-03-22T16:17:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n[ Git list cc'd again - I modified your branch names and commit header \n  line just in case you care about those ]\n\nOn Wed, 21 Mar 2007, Alexander Litvinov wrote:\n> \n> I have a good news : I got the breakage again. And I can reproduce it :-)\n> \n> This is a version of git with three your patches.\n> \n> Here is the steps to broke my repo:\n\nSo does it break every time if you do this particular sequence with the \nparticular state that it has?  If so, wonderful, since it should mean that \nyou can also recreate the file that got corrupted as a blob.\n\n> $ git prune\n> $ git fsck --full\n> dangling commit 50267ccaa820c456bd361db808f99d81714cbce8\n> $ git rebase fix-autoxyzzy                                    \n> First, rewinding head to replay your work on top of it...\n> HEAD is now at 42af3b2... Replace ...\n> Applying 'Show ...'\n> \n> Wrote tree 851c5d8d2213c60efc1bd081b0012bfcc9e558b5\n> Committed: e7117e5637e881368ff04e94a27dca2abdb12d38\n\nAnd then..\n\n> [lan@ac-7923bb4c6c14 navitel (debug-autoxyzzy)]$ git fsck --full\n> error: corrupt loose object 'c01848491b53c3dcfd738149193a14d3c9abe107'\n> error: c01848491b53c3dcfd738149193a14d3c9abe107: object corrupt or missing\n> missing blob c01848491b53c3dcfd738149193a14d3c9abe107\n> dangling commit 50267ccaa820c456bd361db808f99d81714cbce8\n> \n> What can I do to debug this ?\n\nEvery time there's a corrupt object, if you can send it to me, that would \nbe good. If you can tell the source for the corrupt object and can send \nthat to me too, that's always even better, but even in the absense of \nthat, the more corrupt objects I have, the better the chances that I see \nsome pattern. And if it's always the same object that gets corrupted the \nsame way when you start from a particular starting point, that would also \nbe very interesting to know.\n\nConsidering that the \"don't use hardlinks on cygwin\" thing didn't matter \nfor you (and really, I would have only expected it to matter if you used \n^C to kill a process in the middle or something), you migth also be better \noff just trackng the standard git, since it now has Nicos extra \nconsistency checks over and beyond those I send you. \n\nIt's also possible that the real bug is that we have some memory scribble \ninternally in git, and that it shows up for you just because Cygwin and/or \nWinXp has different allocation patterns than other platforms. Do you know \nif there are any \"debugging malloc\" libraries for Cygwin? Something like \nElectricFence/dmalloc under Linux, or running with valgrind.\n\nSince it happens after a single rebase, if it's a git bug (as opposed \nto,for example, a zlib problem or simply a problem in your combination of \nvmware/winxp/cygwin), it would be the recursive merge that screws up. It \n*is* one of the more complex operations (especially if it also ends up \ndoing file-level merging, which I assume it does), so some memory \nallocation problem there is not out of the question, although it's strange \nthat you see it but the (many more) users on UNIX never seem to - it's not \nlike rebase is an uncommon operation!\n\n\t\tLinus\n"},{"id":"37743","messageId":"Pine.LNX.4.64.0703220924590.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703220858400.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-22T16:29:11Z","receivedAt":"2007-03-22T16:29:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 22 Mar 2007, Linus Torvalds wrote:\n> \n> It's also possible that the real bug is that we have some memory scribble \n> internally in git, and that it shows up for you just because Cygwin and/or \n> WinXp has different allocation patterns than other platforms. Do you know \n> if there are any \"debugging malloc\" libraries for Cygwin? Something like \n> ElectricFence/dmalloc under Linux, or running with valgrind.\n\nYeehaa! I think I'm on the right trail.\n\nGit people: do this:\n\n\tyum install ElectricFence\n\n(or similar, apt-get, whatever), and then apply this patch, and do\n\n\tmake test\n\nand it will fail in \"git-apply\"! Which (having read Alexander's corruption \nsequence once more) must have been what corrupted things for Alexander \ntoo!\n\nI've not debugged it any more, but gdb on the core-dump shows:\n\n\t(gdb) where\n\t#0  0x0000003768462331 in SHA1_Update () from /lib64/libcrypto.so.6\n\t#1  0x0000000000461b0e in write_sha1_file_prepare (buf=0x2ba371737fd0, len=46, type=0x49a50b \"blob\", sha1=0x7fff395c84a0 \"\",\n\t    hdr=0x7fff395c8480 \"blob 46\", hdrlen=0x7fff395c847c) at sha1_file.c:1823\n\t#2  0x0000000000461e6f in write_sha1_file (buf=0x2ba371737fd0, len=46, type=0x49a50b \"blob\", returnsha1=0x2ba371733fe0 \"\")\n\t    at sha1_file.c:1962\n\t#3  0x000000000040aa9a in add_index_file (path=0x2ba371719ffc \"one\", mode=33188, buf=0x2ba371737fd0, size=46) at builtin-apply.c:2350\n\t#4  0x000000000040aeb5 in create_file (patch=0x2ba37170cf40) at builtin-apply.c:2451\n\t#5  0x000000000040af45 in write_out_one_result (patch=0x2ba37170cf40, phase=1) at builtin-apply.c:2475\n\t#6  0x000000000040b291 in write_out_results (list=0x2ba37170cf40, skipped_patch=0) at builtin-apply.c:2560\n\t#7  0x000000000040b71c in apply_patch (fd=6, filename=0x7fff395c96ec \"patch.file\", inaccurate_eof=0) at builtin-apply.c:2676\n\t#8  0x000000000040bd1b in cmd_apply (argc=3, argv=0x7fff395c8990, unused_prefix=0x0) at builtin-apply.c:2836\n\t#9  0x0000000000403fbb in handle_internal_command (argc=3, argv=0x7fff395c8990, envp=0x7fff395c89b0) at git.c:322\n\t#10 0x0000000000404193 in main (argc=3, argv=0x7fff395c8990, envp=0x7fff395c89b0) at git.c:391\n\nso I thought I'd send out this email asap, in case somebody else finds the \nbug before I do.\n\nAnyway, this looks like a real smoking gun..\n\n\t\tLinus\n----\ndiff --git a/Makefile b/Makefile\nindex 51c1fed..7e20410 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -123,7 +123,7 @@ uname_P := $(shell sh -c 'uname -p 2>/dev/null || echo not')\n \n # CFLAGS and LDFLAGS are for the users to override from the command line.\n \n-CFLAGS = -g -O2 -Wall\n+CFLAGS = -g -Wall\n LDFLAGS =\n ALL_CFLAGS = $(CFLAGS)\n ALL_LDFLAGS = $(LDFLAGS)\n@@ -647,7 +647,7 @@ prefix_SQ = $(subst ','\\'',$(prefix))\n SHELL_PATH_SQ = $(subst ','\\'',$(SHELL_PATH))\n PERL_PATH_SQ = $(subst ','\\'',$(PERL_PATH))\n \n-LIBS = $(GITLIBS) $(EXTLIBS)\n+LIBS = $(GITLIBS) $(EXTLIBS) -lefence\n \n BASIC_CFLAGS += -DSHA1_HEADER='$(SHA1_HEADER_SQ)' \\\n \t-DETC_GITCONFIG='\"$(ETC_GITCONFIG_SQ)\"' $(COMPAT_CFLAGS)\n"},{"id":"37744","messageId":"alpine.LFD.0.83.0703221215150.18328@xanadu.home","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703220847540.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-22T16:34:02Z","receivedAt":"2007-03-22T16:34:02Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 22 Mar 2007, Linus Torvalds wrote:\n\n> Ok, apart from #1, those should be in current -git now, along with better \n> validation checks (by Nico) when packing. So hopefully at least when there \n> is corruption in a loose object, we will now always notice when we do a \n> \"git repack\", and will never generate a broken pack-file. Knock wood.\n\nNot yet actually.  What I did do is to make index-pack perform more \nvalidation and ensure it never accept SHA1 collisions.\n\nFor the repack case... I think there should be a better way.  Either we \nrevalidate the full SHA1 which would be expensive as we'd basically lose \nmost advantages of direct pack data copy.\n\nWhat I'm pondering is some sort of lightweight checksum like adler32 for \nobject data in the pack but stored in the index.  Since index-pack \nalready perform the full SHA1 already, it could as well provide a \nchecksum for the raw pack object data for the repack case.  Currently we \ntry to validate reused pack data by attempting an inflate pass on the \nobject payload, but that doesn't validate the object type nor the \nreference SHA1 to delta base objects which could get corrupted and \ncopied without noticing into another pack.\n\n> Of course, I actually wonder if the bug might be in your version of zlib \n> (miscompiled or some other thing), in which case *any* amount of \n> pre-validation won't really help, because it will become corrupted when we \n> deflate it prior to writing. For example, if \"deflateBound()\" sometimes \n> doesn't give a valid upper bound and we allocate too little space..\n\nWell, since we provide the size of the allocated output buffer to zlib \nit would be seriously broken if it overflowed it.  Also zlib perform a \nchecksum verification of the deflated data if I remember correctly.  So \nit seems to me that zlib should be quite self validating already.\n\n\nNicolas\n"},{"id":"37745","messageId":"Pine.LNX.4.64.0703220931120.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703220924590.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-22T16:48:30Z","receivedAt":"2007-03-22T16:48:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 22 Mar 2007, Linus Torvalds wrote:\n> \n> Yeehaa! I think I'm on the right trail.\n\n.. and the reason only Alexander sees it, and nobody else does, is that \nthis one is a bug in the CR/LF creation. \n\nJunio: I think it's your git-apply commit 67160271.\n\nIn \"try_create_file()\", we do:\n\n\t...\n        if (convert_to_working_tree(path, &nbuf, &nsize)) {\n                free((char *) buf);\n                buf = nbuf;\n                size = nsize;\n        }\n\t...\n\nbut the thing is, the *caller* still uses the old \"buf/nsize\", so when you \nfree it, the caller will now use the free'd data structure, and if it gets \nre-used by - for example - the zlib deflate() buffers, you'll get a \ncorrupt object (if it gets re-used *before*, you'll get the *wrong* \nobject!). Exactly Alexander's patterns.\n\nAlexander - sorry for all the trouble, this was definitely our bad.\n\nI think the easy temporary fix is to just remove that \"free()\" and leak a \nbit of memory. That gets it through that test with efence for me.\n\nDoes that fix it for you, Alexander?\n\nI can't really say whether there are other problems too - electric fence \nhas a few bugs in that it considers zero-length allocations to be \n\"probably a bug\" and aborts. This makes some of the tests fail with \nefence, when re_compile_internal wants to allocate a zero-length object.\n\n(It also writes crap to stderr, which could make others fail, I didn't \ncheck).\n\n\t\tLinus\n"},{"id":"37746","messageId":"alpine.LFD.0.83.0703221257020.18328@xanadu.home","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703220931120.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-22T17:01:05Z","receivedAt":"2007-03-22T17:01:05Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 22 Mar 2007, Linus Torvalds wrote:\n\n> I can't really say whether there are other problems too - electric fence \n> has a few bugs in that it considers zero-length allocations to be \n> \"probably a bug\" and aborts. This makes some of the tests fail with \n> efence, when re_compile_internal wants to allocate a zero-length object.\n\nYou can tell it not to abort on zero-length allocations by setting the \nEF_ALLOW_MALLOC_0 environment variable to 1.\n\n\nNicolas\n"},{"id":"37748","messageId":"Pine.LNX.4.64.0703221006360.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"alpine.LFD.0.83.0703221257020.18328@xanadu.home","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-22T17:10:24Z","receivedAt":"2007-03-22T17:10:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 22 Mar 2007, Nicolas Pitre wrote:\n\n> On Thu, 22 Mar 2007, Linus Torvalds wrote:\n> \n> > I can't really say whether there are other problems too - electric fence \n> > has a few bugs in that it considers zero-length allocations to be \n> > \"probably a bug\" and aborts. This makes some of the tests fail with \n> > efence, when re_compile_internal wants to allocate a zero-length object.\n> \n> You can tell it not to abort on zero-length allocations by setting the \n> EF_ALLOW_MALLOC_0 environment variable to 1.\n\nAhh,that gets me further, but then it bombs out on the added error \nmessages. Is there something for that braindamage too?\n\n(if it at least tested it with \"isatty()\" or something I'd understand it, \nas it is, it has made my life miserable in the past too.. I like efence, \nbut bruce seems to think his ego is more important than the program \nyou're debugging.)\n\n\t\tLinus\n"},{"id":"37749","messageId":"4602B917.5C2912F4@eudaptics.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703220924590.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Johannes Sixt","fromEmail":"j.sixt@eudaptics.com","sentAt":"2007-03-22T17:12:55Z","receivedAt":"2007-03-22T17:12:55Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Linus Torvalds wrote:\n> Yeehaa! I think I'm on the right trail.\n> \n> Git people: do this:\n> \n>         yum install ElectricFence\n> \n> (or similar, apt-get, whatever), and then apply this patch, and do\n> \n>         make test\n> \n> and it will fail in \"git-apply\"! Which (having read Alexander's corruption\n> sequence once more) must have been what corrupted things for Alexander\n> too!\n\nHere is a repo, created with the MinGW port. It has a corrupted object,\nthat you can see even on Linux with just\n\n$ git show :one\n\nIt was created with an obviously bogus \"git apply\" on Windows.\nTo restore the correct object, just do this on Linux:\n\n$ git read-tree --reset -u HEAD\n$ git apply --index patch.file\n\nNow \"git show :one\" gives you the correct one.\n\nHTH,\n-- Hannes"},{"id":"37750","messageId":"alpine.LFD.0.83.0703221327370.18328@xanadu.home","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703221006360.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-22T17:28:55Z","receivedAt":"2007-03-22T17:28:55Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 22 Mar 2007, Linus Torvalds wrote:\n\n> On Thu, 22 Mar 2007, Nicolas Pitre wrote:\n> \n> > On Thu, 22 Mar 2007, Linus Torvalds wrote:\n> > \n> > > I can't really say whether there are other problems too - electric fence \n> > > has a few bugs in that it considers zero-length allocations to be \n> > > \"probably a bug\" and aborts. This makes some of the tests fail with \n> > > efence, when re_compile_internal wants to allocate a zero-length object.\n> > \n> > You can tell it not to abort on zero-length allocations by setting the \n> > EF_ALLOW_MALLOC_0 environment variable to 1.\n> \n> Ahh,that gets me further, but then it bombs out on the added error \n> messages. Is there something for that braindamage too?\n\nThe efence man page doesn't mention any.\n\n\nNicolas\n"},{"id":"37756","messageId":"7virctt3yi.fsf_-_@assigned-by-dhcp.cox.net","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703220931120.6730@woody.linux-foundation.org","subject":"[PATCH] git-apply: Do not free the wrong buffer when we convert the data for writeout","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-03-22T20:31:49Z","receivedAt":"2007-03-22T20:31:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"When we write out the result of patch application, we sometimes\nneed to munge the data (e.g. under core.autocrlf).  After doing\nso, what we should free is the temporary buffer that holds the\nconverted data returned from convert_to_working_tree(), not the\noriginal one.\n\nThis patch also moves the call to open() up in the function, as\nthe caller expects us to fail cheaply if leading directories\nneed to be created (and then the caller creates them and calls\nus again).  For that calling pattern, attempting conversion\nbefore opening the file adds unnecessary overhead.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n  Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n  > In \"try_create_file()\", we do:\n  >\n  > \t...\n  >         if (convert_to_working_tree(path, &nbuf, &nsize)) {\n  >                 free((char *) buf);\n  >                 buf = nbuf;\n  >                 size = nsize;\n  >         }\n  > \t...\n  >\n  > but the thing is, the *caller* still uses the old \"buf/nsize\", so when you \n  > free it, the caller will now use the free'd data structure, and if it gets \n  > re-used by - for example - the zlib deflate() buffers, you'll get a \n  > corrupt object (if it gets re-used *before*, you'll get the *wrong* \n  > object!). Exactly Alexander's patterns.\n  >\n  > Alexander - sorry for all the trouble, this was definitely our bad.\n\n  I concur.  Sorry for this, Alexander.\n\n  git-apply in general is quite leaky (e.g. it never frees a\n  finished patch, and freeing a patch is quite difficult as some\n  strings like filenames are shared without refcounting), but\n  this part deals with a large amount of data and I would rather\n  not add a new one.\n\n builtin-apply.c |   17 ++++++++++-------\n 1 files changed, 10 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin-apply.c b/builtin-apply.c\nindex dfa1716..27a182b 100644\n--- a/builtin-apply.c\n+++ b/builtin-apply.c\n@@ -2355,7 +2355,7 @@ static void add_index_file(const char *path, unsigned mode, void *buf, unsigned\n \n static int try_create_file(const char *path, unsigned int mode, const char *buf, unsigned long size)\n {\n-\tint fd;\n+\tint fd, converted;\n \tchar *nbuf;\n \tunsigned long nsize;\n \n@@ -2364,17 +2364,18 @@ static int try_create_file(const char *path, unsigned int mode, const char *buf,\n \t\t * terminated.\n \t\t */\n \t\treturn symlink(buf, path);\n+\n+\tfd = open(path, O_CREAT | O_EXCL | O_WRONLY, (mode & 0100) ? 0777 : 0666);\n+\tif (fd < 0)\n+\t\treturn -1;\n+\n \tnsize = size;\n \tnbuf = (char *) buf;\n-\tif (convert_to_working_tree(path, &nbuf, &nsize)) {\n-\t\tfree((char *) buf);\n+\tconverted = convert_to_working_tree(path, &nbuf, &nsize);\n+\tif (converted) {\n \t\tbuf = nbuf;\n \t\tsize = nsize;\n \t}\n-\n-\tfd = open(path, O_CREAT | O_EXCL | O_WRONLY, (mode & 0100) ? 0777 : 0666);\n-\tif (fd < 0)\n-\t\treturn -1;\n \twhile (size) {\n \t\tint written = xwrite(fd, buf, size);\n \t\tif (written < 0)\n@@ -2386,6 +2387,8 @@ static int try_create_file(const char *path, unsigned int mode, const char *buf,\n \t}\n \tif (close(fd) < 0)\n \t\tdie(\"closing file %s: %s\", path, strerror(errno));\n+\tif (converted)\n+\t\tfree(nbuf);\n \treturn 0;\n }\n \n"},{"id":"37759","messageId":"Pine.LNX.4.64.0703221355110.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"7virctt3yi.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] git-apply: Do not free the wrong buffer when we convert the data for writeout","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-22T20:55:29Z","receivedAt":"2007-03-22T20:55:29Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 22 Mar 2007, Junio C Hamano wrote:\n> \n> This patch also moves the call to open() up in the function, as\n> the caller expects us to fail cheaply if leading directories\n> need to be created (and then the caller creates them and calls\n> us again).  For that calling pattern, attempting conversion\n> before opening the file adds unnecessary overhead.\n> \n> Signed-off-by: Junio C Hamano <junkio@cox.net>\n\nAck, looks good.\n\n\t\tLinus\n"},{"id":"37765","messageId":"20070322221340.GA13867@segfault.peff.net","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703221006360.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-03-22T22:13:41Z","receivedAt":"2007-03-22T22:13:41Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 22, 2007 at 10:10:24AM -0700, Linus Torvalds wrote:\n\n> Ahh,that gets me further, but then it bombs out on the added error \n> messages. Is there something for that braindamage too?\n\nTry EF_DISABLE_BANNER=1\n\n-Peff\n"},{"id":"37767","messageId":"Pine.LNX.4.64.0703221720481.6730@woody.linux-foundation.org","threadId":"7009","inReplyTo":"20070322221340.GA13867@segfault.peff.net","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-23T00:25:37Z","receivedAt":"2007-03-23T00:25:37Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 22 Mar 2007, Jeff King wrote:\n> \n> Try EF_DISABLE_BANNER=1\n\nThat does nothing for me. Nor does\n\n\tstrings -- /usr/lib64/libefence.so | grep EF_\n\nshow that string or anything else half-way promising..\n\nGoogling for that shows that some versions of efence have had that flag \n(not necessarily as a environment variable, though). But certainly not the \nversion I have.\n\n\t\tLinus\n"},{"id":"37768","messageId":"17923.8813.162118.405908@lisa.zopyra.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703221720481.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Bill Lear","fromEmail":"rael@zopyra.com","sentAt":"2007-03-23T00:42:21Z","receivedAt":"2007-03-23T00:42:21Z","isPatch":false,"sender":{"key":"rael@zopyra.com","avatar":"https://gravatar.com/avatar/c4f2d2790ca3828d3b4e7dfebabf61d2fe94fd82fa49cdac2a5295dd2d46a874?d=mp&s=160"},"body":"On Thursday, March 22, 2007 at 17:25:37 (-0700) Linus Torvalds writes:\n>\n>\n>On Thu, 22 Mar 2007, Jeff King wrote:\n>> \n>> Try EF_DISABLE_BANNER=1\n>\n>That does nothing for me. Nor does\n>\n>\tstrings -- /usr/lib64/libefence.so | grep EF_\n>\n>show that string or anything else half-way promising..\n>\n>Googling for that shows that some versions of efence have had that flag \n>(not necessarily as a environment variable, though). But certainly not the \n>version I have.\n\nI just downloaded and installed the latest version (2.1.13):\n\n% strings /usr/lib/libefence.a | grep BANNER\nLLEF_DISABLE_BANNER\n EF_DISABLE_BANNER\nEF_DISABLE_BANNER\nEF_DISABLE_BANNER\nEF_DISABLE_BANNER\n\nhttp://perens.com/FreeSoftware/ElectricFence/electric-fence_2.1.13-0.1.tar.gz\n\n\nBill\n"},{"id":"37770","messageId":"20070323005116.GA29901@segfault.peff.net","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703221720481.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-03-23T00:51:16Z","receivedAt":"2007-03-23T00:51:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 22, 2007 at 05:25:37PM -0700, Linus Torvalds wrote:\n\n> That does nothing for me. Nor does\n> \n> \tstrings -- /usr/lib64/libefence.so | grep EF_\n> \n> show that string or anything else half-way promising..\n> \n> Googling for that shows that some versions of efence have had that flag \n> (not necessarily as a environment variable, though). But certainly not the \n> version I have.\n\nHmm. It's in the latest debian package (2.1.14.1) and works as\nadvertised. I just poked at the FC6 version (2.2.2 -- but Bruce's last\nversion seemed to be the 2.1 series, so no idea who is responsible for\nthis brain damage), and it now unconditionally prints the banner.\nHuzzah.\n\nIt at least goes to stderr, which might be redirectable; otherwise\nyou're stuck editing the source (see efence.c:initialize).\n\n-Peff\n"},{"id":"37779","messageId":"200703230940.25103.litvinov2004@gmail.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703220931120.6730@woody.linux-foundation.org","subject":"Re: My git repo is broken, how to fix it ?","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-03-23T03:40:24Z","receivedAt":"2007-03-23T03:40:24Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Thursday 22 March 2007 22:48 Linus Torvalds написал(a):\n> In \"try_create_file()\", we do:\n>\n> \t...\n>         if (convert_to_working_tree(path, &nbuf, &nsize)) {\n>                 free((char *) buf);\n>                 buf = nbuf;\n>                 size = nsize;\n>         }\n> \t...\n>\n> I think the easy temporary fix is to just remove that \"free()\" and leak a\n> bit of memory. That gets it through that test with efence for me.\n>\n> Does that fix it for you, Alexander?\n\nYes, commenting out free fix repo breakage.\n\nThanks for help !\n"},{"id":"37780","messageId":"200703230955.22801.litvinov2004@gmail.com","threadId":"7009","inReplyTo":"Pine.LNX.4.64.0703221355110.6730@woody.linux-foundation.org","subject":"Re: [PATCH] git-apply: Do not free the wrong buffer when we convert the data for writeout","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-03-23T03:55:22Z","receivedAt":"2007-03-23T03:55:22Z","isPatch":true,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Friday 23 March 2007 02:55 Linus Torvalds написал(a):\n> On Thu, 22 Mar 2007, Junio C Hamano wrote:\n> > This patch also moves the call to open() up in the function, as\n> > the caller expects us to fail cheaply if leading directories\n> > need to be created (and then the caller creates them and calls\n> > us again).  For that calling pattern, attempting conversion\n> > before opening the file adds unnecessary overhead.\n\nI have applied this patch ontop of next (d06644b) and it also fix by repo \nbreakage. \n\nThanks for help !\nAlexander Litvinov.\n"}]}