{"thread":{"id":"34443","subject":"repo consistency under crashes and power failures?","startedAt":"2013-07-15T17:48:23Z","lastAt":"2013-07-27T03:10:17Z","messageCount":4,"participants":["Greg Troxel","Jonathan Nieder","Johannes Sixt","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"223458","messageId":"rmiy597iujc.fsf@fnord.ir.bbn.com","threadId":"34443","inReplyTo":null,"subject":"repo consistency under crashes and power failures?","fromName":"Greg Troxel","fromEmail":"gdt@ir.bbn.com","sentAt":"2013-07-15T17:48:23Z","receivedAt":"2013-07-15T17:48:23Z","isPatch":false,"sender":{"key":"gdt@ir.bbn.com","avatar":null},"body":"\nClearly there is the possibility of creating a corrupt repository when\nreceiving objects and updating refs, if a crash or power failure causes\ndata not to get written to disk but that data is pointed to.  Journaling\nmitigates this, but I'd argue that programs should function safely with\nonly the guarantees from POSIX.\n\nI am curious if anyone has actual experiences to share, either\n\n  a report of corruption after a crash (where corruption means that\n  either 1) git fsck reports worse than dangling objects or 2) some ref\n  did not either point to the old place or the new place)\n\n  experiments intended to provoke corruption, like dropping power during\n  pushes, or forced panics in the kernel due to timers, etc.\n\nAlternatively, is there somewhere a first-principles analysis vs POSIX\nspecs (such as fsyncing object files before updating refs to point to\nthem, which I realize has performance negatives)?\n\n(I have not done experiments, but have observed no corruption.)\n\n    Thanks,\n    Greg\n"},{"id":"223459","messageId":"20130715175142.GC14690@google.com","threadId":"34443","inReplyTo":"rmiy597iujc.fsf@fnord.ir.bbn.com","subject":"Re: repo consistency under crashes and power failures?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2013-07-15T17:51:42Z","receivedAt":"2013-07-15T17:51:42Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Greg Troxel wrote:\n\n> Alternatively, is there somewhere a first-principles analysis vs POSIX\n> specs (such as fsyncing object files before updating refs to point to\n> them, which I realize has performance negatives)?\n\nYou might be interested in the 'core.fsyncobjectfiles' setting.\ngit-config(1) has details.\n\nThanks and hope that helps,\nJonathan\n"},{"id":"223501","messageId":"51E4E570.1060403@viscovery.net","threadId":"34443","inReplyTo":"rmiy597iujc.fsf@fnord.ir.bbn.com","subject":"Re: repo consistency under crashes and power failures?","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2013-07-16T06:17:20Z","receivedAt":"2013-07-16T06:17:20Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 7/15/2013 19:48, schrieb Greg Troxel:\n> Clearly there is the possibility of creating a corrupt repository when\n> receiving objects and updating refs, if a crash or power failure causes\n> data not to get written to disk but that data is pointed to.  Journaling\n> mitigates this, but I'd argue that programs should function safely with\n> only the guarantees from POSIX.\n\nEven under POSIX, \"guarantees\" and \"crash/power failure\" do not mesh well.\nThis has been under dispute recently, for example:\n\nhttp://thread.gmane.org/gmane.comp.standards.posix.austin.general/7456/focus=7487\n\nThe best we can achieve with POSIX alone is \"to make bad consequences less\nlikely\".\n\nJonathan already mentioned the knob that allows you to trade performance\nfor more safety.\n\n-- Hannes\n"},{"id":"224166","messageId":"20130727031017.GA20207@sigill.intra.peff.net","threadId":"34443","inReplyTo":"rmiy597iujc.fsf@fnord.ir.bbn.com","subject":"Re: repo consistency under crashes and power failures?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-07-27T03:10:17Z","receivedAt":"2013-07-27T03:10:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 15, 2013 at 01:48:23PM -0400, Greg Troxel wrote:\n\n> I am curious if anyone has actual experiences to share, either\n> \n>   a report of corruption after a crash (where corruption means that\n>   either 1) git fsck reports worse than dangling objects or 2) some ref\n>   did not either point to the old place or the new place)\n> \n>   experiments intended to provoke corruption, like dropping power during\n>   pushes, or forced panics in the kernel due to timers, etc.\n\nI have quite a bit of experience with this, as I investigate all repo\ncorruption that we see on github.com, and have run experiments to try to\nreproduce such corruption.\n\nOur backend git systems are ext3 with journaling and data=ordered. We\nrun that on top of drbd, with two redundant machines sharing the block\ndevice. If one dies, we fail over to the spare. Writes to the block\ndevice are not considered committed until they are written to both\nmachines.\n\nGit's scheme is to write objects (both loose and when receiving packs\nover the wire) via tempfile, with an atomic link-into-place after close.\nWe do not fsync object files by default, but we do fsync packs. However,\nit shouldn't matter as long as your filesystem orders data and metadata\nwrites (if it doesn't, you probably want to turn on object fsyncing).\nSo for our data=ordered filesystems, that's fine.\n\nRef writes have a similar fsync situation to loose object files. We\nwrite the new ref to a tempfile, close, and then rename into place. If\nthe data and metadata writes are out of order, one could have problems\n(but again, not a problem with data=ordered).\n\nMost of the corruption we have seen at GitHub has been one of:\n\n  1. Buggy non-core-git implementations that do not properly use\n     tempfiles to create objects (Grit used to have this problem, but it\n     is now fixed).\n\n  2. Race conditions in examining ref state that can cause refs to be\n     missed when determining reachability (thus you might prune objects\n     that should be left). The worst of these is fixed in the current\n     \"master\" and will be part of git v1.8.4. There are still ways that\n     we can prune too much, but they are reasonably unlikely unless you\n     are pruning constantly.\n\nWe did once experience some lost objects after a server failover.  After\nmuch experimentation, we finally found out that the machine in question\nhad a RAID card with bad memory which would drop some writes which it\nclaimed to have committed after a power failure (so even fsync did not\nhelp).\n\nSo for ordered data and metadata writes, in my experience git is quite\nsolid against power failures and crashes. For systems without that\nguarantee, you should turn on core.fsyncobjectfiles, but I suspect you\ncould also see some ref corruption (and possibly index corruption, too,\nas it does not fsync either).\n\n-Peff\n"}]}