git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Recovering from repository corruption

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Jun 10, 2008, 22:45 UTC
Message-ID
<alpine.LFD.1.10.0806101518590.3101@woody.linux-foundation.org>
In-Reply-To
<6dbd4d000806101509l516cf467me06fadee6ead0964@mail.gmail.com>
On Tue, 10 Jun 2008, Denis Bueno wrote:
Show 10 quoted lines
> 
> > Do you have some odd filesystem in play? Was the current corruption in a
> > similar environment as the old one? IOW, I'm trying to find a pattern
> > here, to see if there might be something we can do about it..
> 
> I can't remember if the old one happened after a panic or not, but I'd
> bet it did.  The filesystem is HFS+, as indeed most OS X 10.4
> installations are.  Maybe the HD has been going south?  However, that
> doesn't seem likely, since when I got the computer it was new, and
> that was around Jun 2007.

Yeah, it's almost certainly not the disk. Disks do go bad, but the behavior tends to be rather different when they do (usually you will get read errors with uncorrectably CRC failures, and you'd know that _very_ clearly).

Sure, I could imagine something like the sector remapping could be flaking out on you, but that sounds really unlikely. Especially since:

Show 8 quoted lines
> > But it *sounds* like the objects you lost were literally old ones, no? Ie
> > the lost stuff wasn't something you had committed in the last five minutes
> > or so? If so, then you really do seem to have a filesystem that corrupts
> > *old* files when it crashes. That's fairly scary. What FS is it?
> 
> No, in fact I had just committed those changes not 10 minutes before
> the panic.  Last time they were also fresh changes, although perhaps
> older than 10 minutes.  I can't remember.

Oh, ok. If so, then this is much less worrisome, and is in fact almost "normal" HFS+ behaviour. It is a journaling filesystem, but it only journals metadata, so the filenames and inodes will be fine after a crash, but the contents will be random.

[ Yeah, yeah, I know - it sounds rather stupid, but it's a common kind of 
  stupidity. The journaling essentially protects the only thing that fsck 
  can find. Ext3 does similar things in "writeback" mode - but you should 
  use "data=ordered" which writes out the data before metadata.
  Basically, such journaling doesn't help data integrity per se, but it 
  does mean that the metadata is ok, and that in turn means that while the 
  file contents won't be dependable, at least things like free block 
  bitmaps etc hopefully are.
  That in turn hopefully means that new file allocations won't be 
  crapping out all over old ones etc due to bad resource allocations, so 
  while it doesn't mean that the data is trust-worthy, it at least means 
  that you can trust _some_ things ]

If your machine crashes often, you could trivially add a "sync" to your commit hook. That would make things better. And maybe we should have a "safe mode" that does these things more carefully. You would definitely want to turn it on on that machine.

Are you doing something special to make the machine crash so much? Or do OS X machines always crash, and Apple PR is just so good that people aren't aware of it?

Anyway, I'll think about sane ways to add a "safe" mode without making it _too_ painful. In the meantime, here's a trial patch that you should probably use. It does slow things down, but hopefully not too much.

(I really don't much like it - but I think this is a good change, and I just need to come up with a better way to do the fsync() than to be totally synchronous about it.)

It's going to make big "git add" calls *much* slower, so I'm not very happy about it (especially since we don't actually care that deeply about the files really being there until much later, so doing something asynchronous would be perfectly acceptable), but for you this is definitely worth-while.

			Linus
---
 sha1_file.c |   17 +++++++++++------
 1 files changed, 11 insertions(+), 6 deletions(-)
diff --git a/sha1_file.c b/sha1_file.c
index adcf37c..86a653b 100644
--- a/sha1_file.c
+++ b/sha1_file.c
@@ -2105,6 +2105,15 @@ int hash_sha1_file(const void *buf, unsigned long len, const char *type,
 	return 0;
 }
 
+/* Finalize a file on disk, and close it. */
+static void close_sha1_file(int fd)
+{
+	fsync_or_die(fd, "sha1 file");
+	fchmod(fd, 0444);
+	if (close(fd) != 0)
+		die("unable to write sha1 file");
+}
+
 static int write_loose_object(const unsigned char *sha1, char *hdr, int hdrlen,
 			      void *buf, unsigned long len, time_t mtime)
 {
@@ -2170,9 +2179,7 @@ static int write_loose_object(const unsigned char *sha1, char *hdr, int hdrlen,
 
 	if (write_buffer(fd, compressed, size) < 0)
 		die("unable to write sha1 file");
-	fchmod(fd, 0444);
-	if (close(fd))
-		die("unable to write sha1 file");
+	close_sha1_file(fd);
 	free(compressed);
 
 	if (mtime) {
@@ -2350,9 +2357,7 @@ int write_sha1_from_fd(const unsigned char *sha1, int fd, char *buffer,
 	} while (1);
 	inflateEnd(&stream);
 
-	fchmod(local, 0444);
-	if (close(local) != 0)
-		die("unable to write sha1 file");
+	close_sha1_file(local);
 	SHA1_Final(real_sha1, &c);
 	if (ret != Z_STREAM_END) {
 		unlink(tmpfile);
Previous: Denis BuenoNext: Linus Torvalds
Message 16 of 31 in “Recovering from repository corruption”
  1. Denis BuenoJun 10, 2008
  2. Jakub NarebskiJun 10, 2008
  3. Denis BuenoJun 10, 2008
  4. Jakub NarebskiJun 10, 2008
  5. Denis BuenoJun 10, 2008
  6. Jakub NarebskiJun 10, 2008
  7. Denis BuenoJun 10, 2008
  8. Linus TorvaldsJun 10, 2008
  9. Denis BuenoJun 10, 2008
  10. Linus TorvaldsJun 10, 2008
  11. Denis BuenoJun 10, 2008
  12. Linus TorvaldsJun 10, 2008
  13. Denis BuenoJun 10, 2008
  14. TarmiganJun 10, 2008
  15. Denis BuenoJun 10, 2008
  16. Linus TorvaldsJun 10, 2008
  17. Linus TorvaldsJun 10, 2008
  18. Nicolas PitreJun 11, 2008
  19. Linus TorvaldsJun 11, 2008
  20. Nicolas PitreJun 11, 2008
  21. Denis BuenoJun 10, 2008
  22. Junio C HamanoJun 10, 2008
  23. To graft or not to graft... (Re: Recovering from repository corruption)Stephen R. van den Berg, Jun 11, 2008
  24. Jakub NarebskiJun 11, 2008
  25. Linus TorvaldsJun 11, 2008
  26. Johan HerlandJun 12, 2008
  27. Jeff KingJun 12, 2008
  28. Johan HerlandJun 12, 2008
  29. Stephen R. van den BergJun 12, 2008
  30. Nicolas PitreJun 10, 2008
  31. Denis BuenoJun 10, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.