{"thread":{"id":"32452","subject":"GIT get corrupted on lustre","startedAt":"2012-12-24T14:08:46Z","lastAt":"2013-02-04T13:58:40Z","messageCount":36,"participants":["Eric Chamberland","Andreas Schwab","Brian J. Murrell","Greg Troxel","Jeff King","Philippe Vaucher","Pyeron, Jason J CTR (US)","Maxime Boissonneault","Erik Faye-Lund","Thomas Rast","Junio C Hamano","Sébastien Boisvert","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"205470","messageId":"50D861EE.6020105@giref.ulaval.ca","threadId":"32452","inReplyTo":null,"subject":"GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2012-12-24T14:08:46Z","receivedAt":"2012-12-24T14:08:46Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"Hi,\n\nwe are using git since may and all is working fine for all of us (almost \n20 people) on our workstations.  However, when we clone our repositories \nto the cluster, only and only there\nwe are having many problems similiar to this post:\n\nhttp://thread.gmane.org/gmane.comp.file-systems.lustre.user/12093\n\nDoing a \"git clone\" always work fine, but when we \"git pull\" or \"git gc\" \nor \"git fsck\", often (1/5) the local repository get corrupted.\nfor example, I got this error two days ago while doing \"git gc\":\n\nerror: index file .git/objects/pack/pack-7b43b1c613a851392aaf4f66916dff2577931576.idx is too small\nerror: refs/heads/mail_seekable does not point to a valid object!\n\nalso, I got this error 5 days ago:\n\nerror: index file .git/objects/pack/pack-ef9b5bbff1ebc1af63ef4262ade3e18b439c58af.idx is too small\nerror: refs/heads/mail_seekable does not point to a valid object!\nRemoving stale temporary file .git/objects/pack/tmp_pack_lO7aw2\n\nand this one some time ago:\n\nRemoving stale temporary file .git/objects/pack/tmp_pack_5CHb2F\nRemoving stale temporary file .git/objects/pack/tmp_pack_GY159g\nRemoving stale temporary file .git/objects/pack/tmp_pack_aKkXTS\n\nWe are using git 1.8.0.1 on CentOS release 5.8 (Final).\n\nWe think it could be related to the fact that we are on a *Lustre* \nfilesystem, which I think doesn't fully support file locking.\n\nQuestions:\n\n#1) However, how can we *test* the filesystem (lustre) compatibility \nwith git? (Is there a unit test we can run?)\n\n#2) Is there a way to compile GIT to be compatible with lustre? (ex: no \nthreads?)\n\n#3) If you *know* your filesystem doesn't allow file locking, how would \nyou configure/compile GIT to work on it?\n\n#4) Anyone has another idea on how to solve this?\n\nThanks,\n\nEric\n"},{"id":"205471","messageId":"m2bodjv74i.fsf@igel.home","threadId":"32452","inReplyTo":"50D861EE.6020105@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Andreas Schwab","fromEmail":"schwab@linux-m68k.org","sentAt":"2012-12-24T14:48:13Z","receivedAt":"2012-12-24T14:48:13Z","isPatch":false,"sender":{"key":"schwab@linux-m68k.org","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n\n> #1) However, how can we *test* the filesystem (lustre) compatibility with\n> git? (Is there a unit test we can run?)\n\nHave you considered running git's testsuite?\n\nAndreas.\n\n-- \nAndreas Schwab, schwab@linux-m68k.org\nGPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5\n\"And now for something completely different.\"\n"},{"id":"205472","messageId":"50D870A0.90205@interlinx.bc.ca","threadId":"32452","inReplyTo":"50D861EE.6020105@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Brian J. Murrell","fromEmail":"brian@interlinx.bc.ca","sentAt":"2012-12-24T15:11:28Z","receivedAt":"2012-12-24T15:11:28Z","isPatch":false,"sender":{"key":"brian@interlinx.bc.ca","avatar":null},"body":"On 12-12-24 09:08 AM, Eric Chamberland wrote:\n> Hi,\n\nHi,\n\n> Doing a \"git clone\" always work fine, but when we \"git pull\" or \"git gc\"\n> or \"git fsck\", often (1/5) the local repository get corrupted.\n\nHave you tried adding a \"-q\" to the git command line to quiet down git's\n\"feedback\" messages?\n\nI discovered other oddities with using git on Lustre which I described\nin this thread:\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/208886\n\nI found that by simply disabling the feedback (which disables the\ncopious SIGALRM processing) I could alleviate the issue.\n\nI wonder if your issues are more of the same.\n\nI filed Lustre bug LU-2276 about it at:\n\nhttp://jira.whamcloud.com/browse/LU-2276\n\nCheers,\nb.\n\n\n"},{"id":"205482","messageId":"rmilicnlyvi.fsf@fnord.ir.bbn.com","threadId":"32452","inReplyTo":"50D861EE.6020105@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Greg Troxel","fromEmail":"gdt@ir.bbn.com","sentAt":"2012-12-25T01:11:13Z","receivedAt":"2012-12-25T01:11:13Z","isPatch":false,"sender":{"key":"gdt@ir.bbn.com","avatar":null},"body":"\n  we are using git since may and all is working fine for all of us\n  (almost 20 people) on our workstations.  However, when we clone our\n  repositories to the cluster, only and only there\n  we are having many problems similiar to this post:\n\nWhat filesystem tests have you run on lustre?  I would run every test\nyou can find, and lustre should have a robust test suite.  It's really\nhard to be certain, but given how many filesystems git is used with,\nyour experience points to a lustre bug.\n\nI would also suggest using ktrace/ktruss/strace and perhaps poring over\nthe logs to see if you can spot any bad behavior.\n\n"},{"id":"205517","messageId":"20121226225152.GB11491@sigill.intra.peff.net","threadId":"32452","inReplyTo":"50D861EE.6020105@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-12-26T22:51:52Z","receivedAt":"2012-12-26T22:51:52Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Dec 24, 2012 at 09:08:46AM -0500, Eric Chamberland wrote:\n\n> Doing a \"git clone\" always work fine, but when we \"git pull\" or \"git\n> gc\" or \"git fsck\", often (1/5) the local repository get corrupted.\n> for example, I got this error two days ago while doing \"git gc\":\n> \n> error: index file .git/objects/pack/pack-7b43b1c613a851392aaf4f66916dff2577931576.idx is too small\n> error: refs/heads/mail_seekable does not point to a valid object!\n> [...]\n> We think it could be related to the fact that we are on a *Lustre*\n> filesystem, which I think doesn't fully support file locking.\n\nI don't think locking is a problem here. The problem is that you have a\ncorrupt .idx file (the second error is almost certainly an effect of the\nfirst one; git cannot look in the packfile, and therefore cannot find\nthe object the ref points to). But we do not ever lock the .idx files.\nThey are generated in tmpfiles and then atomically moved into place\nusing a hard link.\n\nSo if anything, I would suspect that lustre has trouble with the\nwrite/fsync/close/link sequence. Is it possible that it does not keep\nthe ordering, and readers might see a linked file that is missing some\ndata? If you wait (or do some synchronizing operation on the filesystem,\nlike \"sync\", or an unmount/mount), does the repo later work, or is it\nbroken forever?\n\n> #1) However, how can we *test* the filesystem (lustre) compatibility\n> with git? (Is there a unit test we can run?)\n\nRunning \"make test\" in git.git would be a good start. You could also try\nrunning the C program I'm including below. It repeatedly runs a\nwrite/close/fsync/link sequence like the one that index-pack runs, and\nthen verifies the result. If it does not run forever without error, that\nwould be a sign of the possible ordering problem I mentioned above.\n\n> #2) Is there a way to compile GIT to be compatible with lustre? (ex:\n> no threads?)\n\nThis isn't a known issue, so I don't know offhand what compile flags\nmight help. The complete list is at the top of Makefile. You might try\nwith NO_PTHREADS=Yes, but I kind of doubt that threads are at work here.\n\n> #3) If you *know* your filesystem doesn't allow file locking, how\n> would you configure/compile GIT to work on it?\n\nI think locking is a red herring here, as it is not used to create the\n.idx files at all (and we don't do flock locking anyway; everything\nhappens via O_EXCL creation).\n\n-Peff\n\n-- >8 --\n#include <fcntl.h>\n#include <unistd.h>\n#include <stdlib.h>\n#include <stdio.h>\n#include <string.h>\n\nstatic int randomize(unsigned char *buf, int len)\n{\n  int i;\n  len = rand() % len;\n  for (i = 0; i < len; i++)\n    buf[i] = rand() & 0xff;\n  return len;\n}\n\nstatic int check_eof(int fd)\n{\n  int ch;\n  int r = read(fd, &ch, 1);\n  if (r < 0) {\n    perror(\"read error after expected EOF\");\n    return -1;\n  }\n  if (r > 0) {\n    fprintf(stderr, \"extra byte after expected EOF\");\n    return -1;\n  }\n  return 0;\n}\n\nstatic int verify(int fd, const unsigned char *buf, int len)\n{\n  while (len) {\n    char to_check[4096];\n    int got = read(fd, to_check,\n                   len < sizeof(to_check) ? len : sizeof(to_check));\n\n    if (got < 0) {\n      perror(\"unable to read\");\n      return -1;\n    }\n    if (got == 0) {\n      fprintf(stderr, \"premature EOF (%d bytes remaining)\", len);\n      return -1;\n    }\n    if (memcmp(buf, to_check, got)) {\n      fprintf(stderr, \"bytes differ\");\n      return -1;\n    }\n\n    buf += got;\n    len -= got;\n  }\n\n  return check_eof(fd);\n}\n\nint write_in_full(int fd, const unsigned char *buf, int len)\n{\n  while (len) {\n    int r = write(fd, buf, len);\n    if (r < 0)\n      return -1;\n    buf += r;\n    len -= r;\n  }\n  return 0;\n}\n\nint move_into_place(const char *old, const char *new)\n{\n  if (link(old, new) < 0) {\n    perror(\"unable to create hard link\");\n    return 1;\n  }\n  unlink(old);\n  return 0;\n}\n\nint main(void)\n{\n  while (1) {\n    static unsigned char junk[1024*1024];\n    int len = randomize(junk, sizeof(junk));\n    int fd;\n\n    /* clean up from any previous round */\n    unlink(\"tmpfile\");\n    unlink(\"final.idx\");\n\n    fd = open(\"tmpfile\", O_WRONLY|O_CREAT, 0666);\n    if (fd < 0) {\n      perror(\"unable to open tmpfile\");\n      return 1;\n    }\n    if (write_in_full(fd, junk, len) < 0 ||\n        fsync(fd) < 0 ||\n        close(fd) < 0) {\n      perror(\"unable to write\");\n      return 1;\n    }\n\n    if (move_into_place(\"tmpfile\", \"final.idx\") < 0)\n      return 1;\n\n    fd = open(\"final.idx\", O_RDONLY);\n    if (fd < 0) {\n      perror(\"unable to open index file\");\n      return 1;\n    }\n    if (verify(fd, junk, len) < 0)\n      return 1;\n    close(fd);\n  }\n}\n"},{"id":"206300","messageId":"50EC453A.2060306@giref.ulaval.ca","threadId":"32452","inReplyTo":"50D870A0.90205@interlinx.bc.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-08T16:11:38Z","receivedAt":"2013-01-08T16:11:38Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"On 12/24/2012 10:11 AM, Brian J. Murrell wrote:\n> Have you tried adding a \"-q\" to the git command line to quiet down git's\n> \"feedback\" messages?\n>\n\nOk, I have modified my crontab to use \"-q\" and I will wait to see if the \nproblem occurs from now.\n\n> I discovered other oddities with using git on Lustre which I described\n> in this thread:\n>\n> http://thread.gmane.org/gmane.comp.version-control.git/208886\n>\n> I found that by simply disabling the feedback (which disables the\n> copious SIGALRM processing) I could alleviate the issue.\n>\n> I wonder if your issues are more of the same.\n>\n> I filed Lustre bug LU-2276 about it at:\n>\n> http://jira.whamcloud.com/browse/LU-2276\n\nThank you for these informations.  I see the bug is unresolved!...\n\nEric\n"},{"id":"206427","messageId":"50EDDF12.3080800@giref.ulaval.ca","threadId":"32452","inReplyTo":"50EC453A.2060306@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-09T21:20:18Z","receivedAt":"2013-01-09T21:20:18Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"Hi Brian,\n\nOn 01/08/2013 11:11 AM, Eric Chamberland wrote:\n> On 12/24/2012 10:11 AM, Brian J. Murrell wrote:\n>> Have you tried adding a \"-q\" to the git command line to quiet down git's\n>> \"feedback\" messages?\n>>\n>\n\nI moved to git 1.8.1 and added the \"-q\" to the command \"git gc\" but it \noccured to return an error, so the \"-q\" option is not avoiding the \nproblem here... :-/\n\ncommand in crontab:\n\ncd /rap/jsf-051-aa/ericc/tests_git_clones/GIREF && for i in seq 10; do \n/software/apps/git/1.8.1/bin/git gc -q || true;done\n\nresults:\nerror: index file \n.git/objects/pack/pack-1f09879c88cd71a15dcc891713cf038d249830ad.idx is \ntoo small\nerror: refs/remotes/origin/BIB_Branche_1_4_x does not point to a valid \nobject!\n\nand this clone was a \"clean\" clone in which only \"git qc -q\" has been \nrun on....\n\nI still have a doubt on threads....\n\nEric\n"},{"id":"207143","messageId":"50F7F793.80507@giref.ulaval.ca","threadId":"32452","inReplyTo":"50EDDF12.3080800@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-17T13:07:31Z","receivedAt":"2013-01-17T13:07:31Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"Hi!\n\nI still have the corruption problems....\n\nWe just compiled a git without threads to try... (by the way, \n--without-pthreads doesn't work, you have to do a --disable-pthreads \ninstead).\n\nAnd to remove the warnings about threads at \"git gc\" execution, I did a:\n\ngit config --local pack.threads 1\n\nand cloned a repository and started to do:\n\ngit gc\n\nonce every hour.\n\nThen this night (at 05:35:02 exactly), the same error as usual occurred:\n\nerror: index file \n.git/objects/pack/pack-bf0748cee64a1964be0a1061c82aca51c993b825.idx is \ntoo small\nerror: refs/heads/master does not point to a valid object!\n\nSo now I am convinced that it is not a thread problem....\n\nI am kind of discouraged, we like to use git, but in this case we have \nthis error which seems unsolvable!\n\nAnyone has a new idea?\n\nThanks,\n\nEric\n\n\nOn 01/09/2013 04:20 PM, Eric Chamberland wrote:\n> Hi Brian,\n>\n> On 01/08/2013 11:11 AM, Eric Chamberland wrote:\n>> On 12/24/2012 10:11 AM, Brian J. Murrell wrote:\n>>> Have you tried adding a \"-q\" to the git command line to quiet down git's\n>>> \"feedback\" messages?\n>>>\n>>\n>\n> I moved to git 1.8.1 and added the \"-q\" to the command \"git gc\" but it\n> occured to return an error, so the \"-q\" option is not avoiding the\n> problem here... :-/\n>\n> command in crontab:\n>\n> cd /rap/jsf-051-aa/ericc/tests_git_clones/GIREF && for i in seq 10; do\n> /software/apps/git/1.8.1/bin/git gc -q || true;done\n>\n> results:\n> error: index file\n> .git/objects/pack/pack-1f09879c88cd71a15dcc891713cf038d249830ad.idx is\n> too small\n> error: refs/remotes/origin/BIB_Branche_1_4_x does not point to a valid\n> object!\n>\n> and this clone was a \"clean\" clone in which only \"git qc -q\" has been\n> run on....\n>\n> I still have a doubt on threads....\n>\n> Eric\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"207146","messageId":"CAGK7Mr4R=OwfWt4Kat75C8YDi3iLTavMLxeoLxkf1-gKhxrucg@mail.gmail.com","threadId":"32452","inReplyTo":"50F7F793.80507@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Philippe Vaucher","fromEmail":"philippe.vaucher@gmail.com","sentAt":"2013-01-17T14:23:29Z","receivedAt":"2013-01-17T14:23:29Z","isPatch":false,"sender":{"key":"philippe.vaucher@gmail.com","avatar":null},"body":"> Anyone has a new idea?\n\nDid you try Jeff King's code to confirm his idea?\n\nPhilippe\n"},{"id":"207147","messageId":"50F8273E.5050803@giref.ulaval.ca","threadId":"32452","inReplyTo":"CAGK7Mr4R=OwfWt4Kat75C8YDi3iLTavMLxeoLxkf1-gKhxrucg@mail.gmail.com","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-17T16:30:54Z","receivedAt":"2013-01-17T16:30:54Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"On 01/17/2013 09:23 AM, Philippe Vaucher wrote:\n>> Anyone has a new idea?\n>\n> Did you try Jeff King's code to confirm his idea?\n>\n> Philippe\n>\n\nYes I did, but it was running without any problem....\n\nI find that my test case is \"simple\" (fresh git clone then \"git gc\" in a \ncrontab), I bet anyone who has access to a Lustre filesystem can \nreproduce the problem...  The problem is to have such a filesystem to do \nthe tests....\n\nBut I am available to do it...\n\nThanks,\n\nEric\n"},{"id":"207148","messageId":"871B6C10EBEFE342A772D1159D1320853A042AD7@umechphj.easf.csd.disa.mil","threadId":"32452","inReplyTo":"50F8273E.5050803@giref.ulaval.ca","subject":"RE: GIT get corrupted on lustre","fromName":"Pyeron, Jason J CTR (US)","fromEmail":"jason.j.pyeron.ctr@mail.mil","sentAt":"2013-01-17T16:40:30Z","receivedAt":"2013-01-17T16:40:30Z","isPatch":false,"sender":{"key":"jason.j.pyeron.ctr@mail.mil","avatar":null},"body":"> -----Original Message-----\n> From: Eric Chamberland\n> Sent: Thursday, January 17, 2013 11:31 AM\n> \n> On 01/17/2013 09:23 AM, Philippe Vaucher wrote:\n> >> Anyone has a new idea?\n> >\n> > Did you try Jeff King's code to confirm his idea?\n> >\n> > Philippe\n> >\n> \n> Yes I did, but it was running without any problem....\n> \n> I find that my test case is \"simple\" (fresh git clone then \"git gc\" in\n> a\n> crontab), I bet anyone who has access to a Lustre filesystem can\n> reproduce the problem...  The problem is to have such a filesystem to\n> do\n> the tests....\n\nStabbing in the dark, but can you log the details with ProcessMon?\n\nhttp://technet.microsoft.com/en-us/sysinternals/bb896645\n\n> \n> But I am available to do it...\n\n-Jason\n"},{"id":"207149","messageId":"50F829A9.7090606@calculquebec.ca","threadId":"32452","inReplyTo":"871B6C10EBEFE342A772D1159D1320853A042AD7@umechphj.easf.csd.disa.mil","subject":"Re: GIT get corrupted on lustre","fromName":"Maxime Boissonneault","fromEmail":"maxime.boissonneault@calculquebec.ca","sentAt":"2013-01-17T16:41:13Z","receivedAt":"2013-01-17T16:41:13Z","isPatch":false,"sender":{"key":"maxime.boissonneault@calculquebec.ca","avatar":null},"body":"I don't know of any lustre filesystem that is used on Windows. Barely \nanybody uses Windows in the HPC industry.\nThis is a Linux cluster.\n\nMaxime Boissonneault\n\nLe 2013-01-17 11:40, Pyeron, Jason J CTR (US) a écrit :\n>> -----Original Message-----\n>> From: Eric Chamberland\n>> Sent: Thursday, January 17, 2013 11:31 AM\n>>\n>> On 01/17/2013 09:23 AM, Philippe Vaucher wrote:\n>>>> Anyone has a new idea?\n>>> Did you try Jeff King's code to confirm his idea?\n>>>\n>>> Philippe\n>>>\n>> Yes I did, but it was running without any problem....\n>>\n>> I find that my test case is \"simple\" (fresh git clone then \"git gc\" in\n>> a\n>> crontab), I bet anyone who has access to a Lustre filesystem can\n>> reproduce the problem...  The problem is to have such a filesystem to\n>> do\n>> the tests....\n> Stabbing in the dark, but can you log the details with ProcessMon?\n>\n> http://technet.microsoft.com/en-us/sysinternals/bb896645\n>\n>> But I am available to do it...\n> -Jason\n\n\n-- \n---------------------------------\nMaxime Boissonneault\nAnalyste de calcul - Calcul Québec, Université Laval\nPh. D. en physique\n"},{"id":"207153","messageId":"871B6C10EBEFE342A772D1159D1320853A044B42@umechphj.easf.csd.disa.mil","threadId":"32452","inReplyTo":"50F829A9.7090606@calculquebec.ca","subject":"RE: GIT get corrupted on lustre","fromName":"Pyeron, Jason J CTR (US)","fromEmail":"jason.j.pyeron.ctr@mail.mil","sentAt":"2013-01-17T17:17:11Z","receivedAt":"2013-01-17T17:17:11Z","isPatch":false,"sender":{"key":"jason.j.pyeron.ctr@mail.mil","avatar":null},"body":"Sorry, I am in cygwin mode, and I had crossed wires in my head. s/ProcessMon/strace/\n\n> -----Original Message-----\n> From: git-owner@vger.kernel.org [mailto:git-owner@vger.kernel.org] On\n> Behalf Of Maxime Boissonneault\n> Sent: Thursday, January 17, 2013 11:41 AM\n> To: Pyeron, Jason J CTR (US)\n> Cc: Eric Chamberland; Philippe Vaucher; git@vger.kernel.org; Sébastien\n> Boisvert\n> Subject: Re: GIT get corrupted on lustre\n> \n> I don't know of any lustre filesystem that is used on Windows. Barely\n> anybody uses Windows in the HPC industry.\n> This is a Linux cluster.\n> \n> Maxime Boissonneault\n> \n> Le 2013-01-17 11:40, Pyeron, Jason J CTR (US) a écrit :\n> >> -----Original Message-----\n> >> From: Eric Chamberland\n> >> Sent: Thursday, January 17, 2013 11:31 AM\n> >>\n> >> On 01/17/2013 09:23 AM, Philippe Vaucher wrote:\n> >>>> Anyone has a new idea?\n> >>> Did you try Jeff King's code to confirm his idea?\n> >>>\n> >>> Philippe\n> >>>\n> >> Yes I did, but it was running without any problem....\n> >>\n> >> I find that my test case is \"simple\" (fresh git clone then \"git gc\"\n> in\n> >> a\n> >> crontab), I bet anyone who has access to a Lustre filesystem can\n> >> reproduce the problem...  The problem is to have such a filesystem\n> to\n> >> do\n> >> the tests....\n> > Stabbing in the dark, but can you log the details with ProcessMon?\n> >\n> > http://technet.microsoft.com/en-us/sysinternals/bb896645\n> >\n> >> But I am available to do it...\n> > -Jason\n> \n> \n> --\n> ---------------------------------\n> Maxime Boissonneault\n> Analyste de calcul - Calcul Québec, Université Laval\n> Ph. D. en physique\n> \n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"207207","messageId":"50F98B53.9080109@giref.ulaval.ca","threadId":"32452","inReplyTo":"871B6C10EBEFE342A772D1159D1320853A044B42@umechphj.easf.csd.disa.mil","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-18T17:50:11Z","receivedAt":"2013-01-18T17:50:11Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"Good idea!\n\nI did a strace and here is the output with the error:\n\nhttp://www.giref.ulaval.ca/~ericc/strace_git_error.txt\n\nHope it will be insightful!\n\nEric\n\n\nOn 01/17/2013 12:17 PM, Pyeron, Jason J CTR (US) wrote:\n> Sorry, I am in cygwin mode, and I had crossed wires in my head. s/ProcessMon/strace/\n>\n>> -----Original Message-----\n>> From: git-owner@vger.kernel.org [mailto:git-owner@vger.kernel.org] On\n>> Behalf Of Maxime Boissonneault\n>> Sent: Thursday, January 17, 2013 11:41 AM\n>> To: Pyeron, Jason J CTR (US)\n>> Cc: Eric Chamberland; Philippe Vaucher; git@vger.kernel.org; Sébastien\n>> Boisvert\n>> Subject: Re: GIT get corrupted on lustre\n>>\n>> I don't know of any lustre filesystem that is used on Windows. Barely\n>> anybody uses Windows in the HPC industry.\n>> This is a Linux cluster.\n>>\n>> Maxime Boissonneault\n>>\n>> Le 2013-01-17 11:40, Pyeron, Jason J CTR (US) a écrit :\n>>>> -----Original Message-----\n>>>> From: Eric Chamberland\n>>>> Sent: Thursday, January 17, 2013 11:31 AM\n>>>>\n>>>> On 01/17/2013 09:23 AM, Philippe Vaucher wrote:\n>>>>>> Anyone has a new idea?\n>>>>> Did you try Jeff King's code to confirm his idea?\n>>>>>\n>>>>> Philippe\n>>>>>\n>>>> Yes I did, but it was running without any problem....\n>>>>\n>>>> I find that my test case is \"simple\" (fresh git clone then \"git gc\"\n>> in\n>>>> a\n>>>> crontab), I bet anyone who has access to a Lustre filesystem can\n>>>> reproduce the problem...  The problem is to have such a filesystem\n>> to\n>>>> do\n>>>> the tests....\n>>> Stabbing in the dark, but can you log the details with ProcessMon?\n>>>\n>>> http://technet.microsoft.com/en-us/sysinternals/bb896645\n>>>\n>>>> But I am available to do it...\n>>> -Jason\n>>\n>>\n>> --\n>> ---------------------------------\n>> Maxime Boissonneault\n>> Analyste de calcul - Calcul Québec, Université Laval\n>> Ph. D. en physique\n>>\n>> --\n>> To unsubscribe from this list: send the line \"unsubscribe git\" in\n>> the body of a message to majordomo@vger.kernel.org\n>> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"207397","messageId":"CABPQNSbJr4dR9mq+kCwGe-RKb9PA7q=SKzbFW+=md_PLzZh=nQ@mail.gmail.com","threadId":"32452","inReplyTo":"50F98B53.9080109@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2013-01-21T13:29:14Z","receivedAt":"2013-01-21T13:29:14Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Fri, Jan 18, 2013 at 6:50 PM, Eric Chamberland\n<Eric.Chamberland@giref.ulaval.ca> wrote:\n> Good idea!\n>\n> I did a strace and here is the output with the error:\n>\n> http://www.giref.ulaval.ca/~ericc/strace_git_error.txt\n>\n> Hope it will be insightful!\n\nThis trace doesn't seem to contain child-processes, but instead having\ntheir stderr inlined into the log. Try using \"strace -f\" instead...\n"},{"id":"207404","messageId":"87a9s2o6ri.fsf@pctrast.inf.ethz.ch","threadId":"32452","inReplyTo":"CABPQNSbJr4dR9mq+kCwGe-RKb9PA7q=SKzbFW+=md_PLzZh=nQ@mail.gmail.com","subject":"Re: GIT get corrupted on lustre","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2013-01-21T16:11:45Z","receivedAt":"2013-01-21T16:11:45Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Erik Faye-Lund <kusmabite@gmail.com> writes:\n\n> On Fri, Jan 18, 2013 at 6:50 PM, Eric Chamberland\n> <Eric.Chamberland@giref.ulaval.ca> wrote:\n>> Good idea!\n>>\n>> I did a strace and here is the output with the error:\n>>\n>> http://www.giref.ulaval.ca/~ericc/strace_git_error.txt\n>>\n>> Hope it will be insightful!\n>\n> This trace doesn't seem to contain child-processes, but instead having\n> their stderr inlined into the log. Try using \"strace -f\" instead...\n\nI happen to have access to a lustre FS on the brutus cluster of ETH\nZurich, so I figured I could give it a shot.\n\nWhat's odd is that while I cannot reproduce the original problem, there\nseems to be another issue/bug with utime():\n\n  $ strace -f -o ~/gc.trace git gc\n  warning: failed utime() on .git/objects/68/tmp_obj_sCAEVc: Interrupted system call\n  warning: failed utime() on .git/objects/a6/tmp_obj_3cdB2c: Interrupted system call\n  warning: failed utime() on .git/objects/69/tmp_obj_lbU3Xc: Interrupted system call\n  warning: failed utime() on .git/objects/c3/tmp_obj_EU97Wc: Interrupted system call\n  warning: failed utime() on .git/objects/3e/tmp_obj_tb2j3c: Interrupted system call\n  warning: failed utime() on .git/objects/15/tmp_obj_e6zMXc: Interrupted system call\n  warning: failed utime() on .git/objects/54/tmp_obj_ExOJVc: Interrupted system call\n  warning: failed utime() on .git/objects/e3/tmp_obj_GtPw4c: Interrupted system call\n  warning: failed utime() on .git/objects/21/tmp_obj_Xex32c: Interrupted system call\n  warning: failed utime() on .git/objects/1a/tmp_obj_CzwsZc: Interrupted system call\n  warning: failed utime() on .git/objects/18/tmp_obj_o6fp3c: Interrupted system call\n  warning: failed utime() on .git/objects/32/tmp_obj_Ih0G4c: Interrupted system call\n  warning: failed utime() on .git/objects/41/tmp_obj_1RXV1c: Interrupted system call\n  Counting objects: 137744, done.\n  Delta compression using up to 48 threads.\n  Compressing objects: 100% (36510/36510), done.\n  Writing objects: 100% (137744/137744), done.\n  Total 137744 (delta 101591), reused 135512 (delta 99472)\n\nThe trace is here (2.1MB compressed):\n\n  http://thomasrast.ch/download/gc.trace.bz2\n\nFor the test I used a clone of another git.git I had around.  I think\nthe error is from sha1_file.c:2564.  While that doesn't look too\nimportant (see ca11b212 for context), it does raise the question: what\nother system calls that we never expect to EINTR can do this on\nsufficiently arcane system/FS combinations?\n\n\nPeff's test ran without any apparent issue for a few minutes.  I also\nran an extended version (at the end) that sets alarms, so as to actually\nget interrupted.  That proved more interesting.  I had to fix verify()\nand write_in_full() to account for EINTR in read()/write(), as those\nseem likely to fail.  I also got link() to fail:\n\n  $ ~/lustre-peff-reproducer \n  unable to create hard link: Interrupted system call\n  unable to open index file: No such file or directory\n\nbut it took a long time.  Unfortunately, when running it with strace I\nmanaged to panic the host I ran it on:\n\n  $ strace -o ~/peff-reproducer.trace ~/lustre-peff-reproducer \n\n  Message from syslogd@brutus1 at Jan 21 17:09:43 ...                                                    \n   kernel:LustreError: 37417:0:(osc_lock.c:1182:osc_lock_enqueue()) ASSERTION( ols->ols_state == OLS_NEW ) failed: Impossible state: 4\n\n  Message from syslogd@brutus1 at Jan 21 17:09:43 ...\n   kernel:LustreError: 37417:0:(osc_lock.c:1182:osc_lock_enqueue()) LBUG\n\n  Message from syslogd@brutus1 at Jan 21 17:09:43 ...\n   kernel:Kernel panic - not syncing: LBUG\n\nYay for now having to explain this to the cluster team.\n\n\nI tried finding a standard that limits the syscalls to which EINTR\napplies, without too much success.  I'm not sure how far I should trust\nmy manpages, but while some of them explicitly list EINTR as a possible\nerror (read, write, etc.) link() does not.  (And the linux manpages\nagree with the POSOIX ones for once.)\n\nIf somebody finds such a standard, we could of course use it to blame\nlustre instead :-)\n\nIn the absence of it, wouldn't we in theory have to write a simple\nloop-on-EINTR wrapper for *all* syscalls?\n\nOf course there's the added problem that when open(O_CREAT|O_EXCL) fails\nwith EINTR, it's hard to tell whether a file that may now exist is\nindeed yours or some other process's.\n\n--- 8< ----\n#include <fcntl.h>\n#include <unistd.h>\n#include <stdlib.h>\n#include <stdio.h>\n#include <string.h>\n#include <sys/time.h>\n#include <signal.h>\n#include <errno.h>\n\nstruct itimerval itv;\n\nstatic int randomize(unsigned char *buf, int len)\n{\n  int i;\n  len = rand() % len;\n  for (i = 0; i < len; i++)\n    buf[i] = rand() & 0xff;\n  return len;\n}\n\nstatic int check_eof(int fd)\n{\n  int ch;\n  int r = read(fd, &ch, 1);\n  if (r < 0) {\n    perror(\"read error after expected EOF\");\n    return -1;\n  }\n  if (r > 0) {\n    fprintf(stderr, \"extra byte after expected EOF\");\n    return -1;\n  }\n  return 0;\n}\n\nstatic int verify(int fd, const unsigned char *buf, int len)\n{\n  while (len) {\n    char to_check[4096];\n    int got = read(fd, to_check,\n                   len < sizeof(to_check) ? len : sizeof(to_check));\n\n    if (got < 0 && errno == EINTR)\n      continue;\n    if (got < 0) {\n      perror(\"unable to read\");\n      return -1;\n    }\n    if (got == 0) {\n      fprintf(stderr, \"premature EOF (%d bytes remaining)\", len);\n      return -1;\n    }\n    if (memcmp(buf, to_check, got)) {\n      fprintf(stderr, \"bytes differ\");\n      return -1;\n    }\n\n    buf += got;\n    len -= got;\n  }\n\n  return check_eof(fd);\n}\n\nint write_in_full(int fd, const unsigned char *buf, int len)\n{\n  while (len) {\n    int r = write(fd, buf, len);\n    if (r < 0 && errno == EINTR)\n      continue;\n    if (r < 0)\n      return -1;\n    buf += r;\n    len -= r;\n  }\n  return 0;\n}\n\nint move_into_place(const char *old, const char *new)\n{\n  if (link(old, new) < 0) {\n    perror(\"unable to create hard link\");\n    return 1;\n  }\n  unlink(old);\n  return 0;\n}\n\nvoid handle_alarm(int signal)\n{\n}\n\nint main(void)\n{\n  struct sigaction sa;\n\n  sa.sa_handler = handle_alarm;\n  sa.sa_flags = SA_RESTART;\n  sigaction(SIGALRM, &sa, NULL);\n\n  itv.it_interval.tv_sec = 0;\n  itv.it_interval.tv_usec = 10000;\n  itv.it_value.tv_sec = 0;\n  itv.it_value.tv_usec = 100000;\n  setitimer(ITIMER_REAL, &itv, NULL);\n\n  while (1) {\n    static unsigned char junk[1024*1024];\n    int len = randomize(junk, sizeof(junk));\n    int fd;\n\n    /* clean up from any previous round */\n    unlink(\"tmpfile\");\n    unlink(\"final.idx\");\n\n    fd = open(\"tmpfile\", O_WRONLY|O_CREAT, 0666);\n    if (fd < 0) {\n      perror(\"unable to open tmpfile\");\n      return 1;\n    }\n    if (write_in_full(fd, junk, len) < 0 ||\n        fsync(fd) < 0 ||\n        close(fd) < 0) {\n      perror(\"unable to write\");\n      return 1;\n    }\n\n    if (move_into_place(\"tmpfile\", \"final.idx\") < 0)\n      return 1;\n\n    fd = open(\"final.idx\", O_RDONLY);\n    if (fd < 0) {\n      perror(\"unable to open index file\");\n      return 1;\n    }\n    if (verify(fd, junk, len) < 0)\n      return 1;\n    close(fd);\n  }\n}\n"},{"id":"207405","messageId":"50FD696B.5000205@calculquebec.ca","threadId":"32452","inReplyTo":"87a9s2o6ri.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Maxime Boissonneault","fromEmail":"maxime.boissonneault@calculquebec.ca","sentAt":"2013-01-21T16:14:35Z","receivedAt":"2013-01-21T16:14:35Z","isPatch":false,"sender":{"key":"maxime.boissonneault@calculquebec.ca","avatar":null},"body":"Hi Thomas,\nCan you tell me what is the version of the lustre servers and the lustre \nclients ?\n\nThanks,\n\nMaxime Boissonneault\nHPC specialist @ Calcul Québec.\n\nLe 2013-01-21 11:11, Thomas Rast a écrit :\n> Erik Faye-Lund <kusmabite@gmail.com> writes:\n>\n>> On Fri, Jan 18, 2013 at 6:50 PM, Eric Chamberland\n>> <Eric.Chamberland@giref.ulaval.ca> wrote:\n>>> Good idea!\n>>>\n>>> I did a strace and here is the output with the error:\n>>>\n>>> http://www.giref.ulaval.ca/~ericc/strace_git_error.txt\n>>>\n>>> Hope it will be insightful!\n>> This trace doesn't seem to contain child-processes, but instead having\n>> their stderr inlined into the log. Try using \"strace -f\" instead...\n> I happen to have access to a lustre FS on the brutus cluster of ETH\n> Zurich, so I figured I could give it a shot.\n>\n> What's odd is that while I cannot reproduce the original problem, there\n> seems to be another issue/bug with utime():\n>\n>    $ strace -f -o ~/gc.trace git gc\n>    warning: failed utime() on .git/objects/68/tmp_obj_sCAEVc: Interrupted system call\n>    warning: failed utime() on .git/objects/a6/tmp_obj_3cdB2c: Interrupted system call\n>    warning: failed utime() on .git/objects/69/tmp_obj_lbU3Xc: Interrupted system call\n>    warning: failed utime() on .git/objects/c3/tmp_obj_EU97Wc: Interrupted system call\n>    warning: failed utime() on .git/objects/3e/tmp_obj_tb2j3c: Interrupted system call\n>    warning: failed utime() on .git/objects/15/tmp_obj_e6zMXc: Interrupted system call\n>    warning: failed utime() on .git/objects/54/tmp_obj_ExOJVc: Interrupted system call\n>    warning: failed utime() on .git/objects/e3/tmp_obj_GtPw4c: Interrupted system call\n>    warning: failed utime() on .git/objects/21/tmp_obj_Xex32c: Interrupted system call\n>    warning: failed utime() on .git/objects/1a/tmp_obj_CzwsZc: Interrupted system call\n>    warning: failed utime() on .git/objects/18/tmp_obj_o6fp3c: Interrupted system call\n>    warning: failed utime() on .git/objects/32/tmp_obj_Ih0G4c: Interrupted system call\n>    warning: failed utime() on .git/objects/41/tmp_obj_1RXV1c: Interrupted system call\n>    Counting objects: 137744, done.\n>    Delta compression using up to 48 threads.\n>    Compressing objects: 100% (36510/36510), done.\n>    Writing objects: 100% (137744/137744), done.\n>    Total 137744 (delta 101591), reused 135512 (delta 99472)\n>\n> The trace is here (2.1MB compressed):\n>\n>    http://thomasrast.ch/download/gc.trace.bz2\n>\n> For the test I used a clone of another git.git I had around.  I think\n> the error is from sha1_file.c:2564.  While that doesn't look too\n> important (see ca11b212 for context), it does raise the question: what\n> other system calls that we never expect to EINTR can do this on\n> sufficiently arcane system/FS combinations?\n>\n>\n> Peff's test ran without any apparent issue for a few minutes.  I also\n> ran an extended version (at the end) that sets alarms, so as to actually\n> get interrupted.  That proved more interesting.  I had to fix verify()\n> and write_in_full() to account for EINTR in read()/write(), as those\n> seem likely to fail.  I also got link() to fail:\n>\n>    $ ~/lustre-peff-reproducer\n>    unable to create hard link: Interrupted system call\n>    unable to open index file: No such file or directory\n>\n> but it took a long time.  Unfortunately, when running it with strace I\n> managed to panic the host I ran it on:\n>\n>    $ strace -o ~/peff-reproducer.trace ~/lustre-peff-reproducer\n>\n>    Message from syslogd@brutus1 at Jan 21 17:09:43 ...\n>     kernel:LustreError: 37417:0:(osc_lock.c:1182:osc_lock_enqueue()) ASSERTION( ols->ols_state == OLS_NEW ) failed: Impossible state: 4\n>\n>    Message from syslogd@brutus1 at Jan 21 17:09:43 ...\n>     kernel:LustreError: 37417:0:(osc_lock.c:1182:osc_lock_enqueue()) LBUG\n>\n>    Message from syslogd@brutus1 at Jan 21 17:09:43 ...\n>     kernel:Kernel panic - not syncing: LBUG\n>\n> Yay for now having to explain this to the cluster team.\n>\n>\n> I tried finding a standard that limits the syscalls to which EINTR\n> applies, without too much success.  I'm not sure how far I should trust\n> my manpages, but while some of them explicitly list EINTR as a possible\n> error (read, write, etc.) link() does not.  (And the linux manpages\n> agree with the POSOIX ones for once.)\n>\n> If somebody finds such a standard, we could of course use it to blame\n> lustre instead :-)\n>\n> In the absence of it, wouldn't we in theory have to write a simple\n> loop-on-EINTR wrapper for *all* syscalls?\n>\n> Of course there's the added problem that when open(O_CREAT|O_EXCL) fails\n> with EINTR, it's hard to tell whether a file that may now exist is\n> indeed yours or some other process's.\n>\n> --- 8< ----\n> #include <fcntl.h>\n> #include <unistd.h>\n> #include <stdlib.h>\n> #include <stdio.h>\n> #include <string.h>\n> #include <sys/time.h>\n> #include <signal.h>\n> #include <errno.h>\n>\n> struct itimerval itv;\n>\n> static int randomize(unsigned char *buf, int len)\n> {\n>    int i;\n>    len = rand() % len;\n>    for (i = 0; i < len; i++)\n>      buf[i] = rand() & 0xff;\n>    return len;\n> }\n>\n> static int check_eof(int fd)\n> {\n>    int ch;\n>    int r = read(fd, &ch, 1);\n>    if (r < 0) {\n>      perror(\"read error after expected EOF\");\n>      return -1;\n>    }\n>    if (r > 0) {\n>      fprintf(stderr, \"extra byte after expected EOF\");\n>      return -1;\n>    }\n>    return 0;\n> }\n>\n> static int verify(int fd, const unsigned char *buf, int len)\n> {\n>    while (len) {\n>      char to_check[4096];\n>      int got = read(fd, to_check,\n>                     len < sizeof(to_check) ? len : sizeof(to_check));\n>\n>      if (got < 0 && errno == EINTR)\n>        continue;\n>      if (got < 0) {\n>        perror(\"unable to read\");\n>        return -1;\n>      }\n>      if (got == 0) {\n>        fprintf(stderr, \"premature EOF (%d bytes remaining)\", len);\n>        return -1;\n>      }\n>      if (memcmp(buf, to_check, got)) {\n>        fprintf(stderr, \"bytes differ\");\n>        return -1;\n>      }\n>\n>      buf += got;\n>      len -= got;\n>    }\n>\n>    return check_eof(fd);\n> }\n>\n> int write_in_full(int fd, const unsigned char *buf, int len)\n> {\n>    while (len) {\n>      int r = write(fd, buf, len);\n>      if (r < 0 && errno == EINTR)\n>        continue;\n>      if (r < 0)\n>        return -1;\n>      buf += r;\n>      len -= r;\n>    }\n>    return 0;\n> }\n>\n> int move_into_place(const char *old, const char *new)\n> {\n>    if (link(old, new) < 0) {\n>      perror(\"unable to create hard link\");\n>      return 1;\n>    }\n>    unlink(old);\n>    return 0;\n> }\n>\n> void handle_alarm(int signal)\n> {\n> }\n>\n> int main(void)\n> {\n>    struct sigaction sa;\n>\n>    sa.sa_handler = handle_alarm;\n>    sa.sa_flags = SA_RESTART;\n>    sigaction(SIGALRM, &sa, NULL);\n>\n>    itv.it_interval.tv_sec = 0;\n>    itv.it_interval.tv_usec = 10000;\n>    itv.it_value.tv_sec = 0;\n>    itv.it_value.tv_usec = 100000;\n>    setitimer(ITIMER_REAL, &itv, NULL);\n>\n>    while (1) {\n>      static unsigned char junk[1024*1024];\n>      int len = randomize(junk, sizeof(junk));\n>      int fd;\n>\n>      /* clean up from any previous round */\n>      unlink(\"tmpfile\");\n>      unlink(\"final.idx\");\n>\n>      fd = open(\"tmpfile\", O_WRONLY|O_CREAT, 0666);\n>      if (fd < 0) {\n>        perror(\"unable to open tmpfile\");\n>        return 1;\n>      }\n>      if (write_in_full(fd, junk, len) < 0 ||\n>          fsync(fd) < 0 ||\n>          close(fd) < 0) {\n>        perror(\"unable to write\");\n>        return 1;\n>      }\n>\n>      if (move_into_place(\"tmpfile\", \"final.idx\") < 0)\n>        return 1;\n>\n>      fd = open(\"final.idx\", O_RDONLY);\n>      if (fd < 0) {\n>        perror(\"unable to open index file\");\n>        return 1;\n>      }\n>      if (verify(fd, junk, len) < 0)\n>        return 1;\n>      close(fd);\n>    }\n> }\n\n\n-- \n---------------------------------\nMaxime Boissonneault\nAnalyste de calcul - Calcul Québec, Université Laval\nPh. D. en physique\n"},{"id":"207406","messageId":"8738xuo6c7.fsf@pctrast.inf.ethz.ch","threadId":"32452","inReplyTo":"50FD696B.5000205@calculquebec.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2013-01-21T16:20:56Z","receivedAt":"2013-01-21T16:20:56Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Maxime Boissonneault <maxime.boissonneault@calculquebec.ca> writes:\n\n> Hi Thomas,\n> Can you tell me what is the version of the lustre servers and the\n> lustre clients ?\n\n$ uname -a\nLinux brutus4.ethz.ch 2.6.32-279.14.1.el6.x86_64 #1 SMP Tue Nov 6 23:43:09 UTC 2012 x86_64 x86_64 x86_64 GNU/Linux\n$ cat /proc/fs/lustre/version\nlustre: 2.3.0\nkernel: patchless_client\nbuild:  2.3.0-RC6--PRISTINE-2.6.32-279.14.1.el6.x86_64\n\nI have no idea what the servers are running, I only have client access.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"207412","messageId":"50FD75BE.1030504@giref.ulaval.ca","threadId":"32452","inReplyTo":"CABPQNSbJr4dR9mq+kCwGe-RKb9PA7q=SKzbFW+=md_PLzZh=nQ@mail.gmail.com","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-21T17:07:10Z","receivedAt":"2013-01-21T17:07:10Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"Hi,\n\nIt just happened again.  Now I have the \"strace -f\" output gzipped here:\n\nhttp://www.giref.ulaval.ca/~ericc/strace-f_git_error.txt.gz\n\nthanks,\n\nEric\n\nOn 01/21/2013 08:29 AM, Erik Faye-Lund wrote:\n> On Fri, Jan 18, 2013 at 6:50 PM, Eric Chamberland\n> <Eric.Chamberland@giref.ulaval.ca> wrote:\n>> Good idea!\n>>\n>> I did a strace and here is the output with the error:\n>>\n>> http://www.giref.ulaval.ca/~ericc/strace_git_error.txt\n>>\n>> Hope it will be insightful!\n>\n> This trace doesn't seem to contain child-processes, but instead having\n> their stderr inlined into the log. Try using \"strace -f\" instead...\n>\n"},{"id":"207413","messageId":"50FD88CE.5030508@giref.ulaval.ca","threadId":"32452","inReplyTo":"50FD75BE.1030504@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-21T18:28:30Z","receivedAt":"2013-01-21T18:28:30Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"On 01/21/2013 12:07 PM, Eric Chamberland wrote:\n> Hi,\n>\n> It just happened again.  Now I have the \"strace -f\" output gzipped here:\n>\n> http://www.giref.ulaval.ca/~ericc/strace-f_git_error.txt.gz\n>\n\nI added the \"strace -f\" output when non error occurs...\n\nhttp://www.giref.ulaval.ca/~ericc/strace-f_git_no_error.txt.gz\n\na \"kdiff3\" can show the differences just before the error...\n\nEric\n"},{"id":"207418","messageId":"kdk2ss$498$1@ger.gmane.org","threadId":"32452","inReplyTo":"87a9s2o6ri.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Brian J. Murrell","fromEmail":"brian@interlinx.bc.ca","sentAt":"2013-01-21T18:54:23Z","receivedAt":"2013-01-21T18:54:23Z","isPatch":false,"sender":{"key":"brian@interlinx.bc.ca","avatar":null},"body":"On 13-01-21 11:11 AM, Thomas Rast wrote:\n> \n> What's odd is that while I cannot reproduce the original problem, there\n> seems to be another issue/bug with utime():\n\nI wonder if this is related to http://jira.whamcloud.com/browse/LU-305.\n That was reported as fixed in Lustre 2.0.0 and 2.1.0 but I thought I\nsaw it on 2.1.1 and added a comment to the above ticket about that.\n\n> In the absence of it, wouldn't we in theory have to write a simple\n> loop-on-EINTR wrapper for *all* syscalls?\n\nIIUC, that's what SA_RESTART is all about.\n\n> Of course there's the added problem that when open(O_CREAT|O_EXCL) fails\n> with EINTR, it's hard to tell whether a file that may now exist is\n> indeed yours or some other process's.\n\nOr whether it's in a \"half created\" state such as I hypothesize in\nhttp://jira.whamcloud.com/browse/LU-2276.\n\nb.\n\n\n"},{"id":"207433","messageId":"87r4lejpx8.fsf@pctrast.inf.ethz.ch","threadId":"32452","inReplyTo":"kdk2ss$498$1@ger.gmane.org","subject":"Re: GIT get corrupted on lustre","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2013-01-21T19:29:07Z","receivedAt":"2013-01-21T19:29:07Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Please don't drop the Cc list!\n\n\"Brian J. Murrell\" <brian@interlinx.bc.ca> writes:\n\n>> What's odd is that while I cannot reproduce the original problem, there\n>> seems to be another issue/bug with utime():\n>\n> I wonder if this is related to http://jira.whamcloud.com/browse/LU-305.\n>  That was reported as fixed in Lustre 2.0.0 and 2.1.0 but I thought I\n> saw it on 2.1.1 and added a comment to the above ticket about that.\n\nAha, that's a very interesting bug report.  My observations support\nyours: I managed to get EINTR during utime().\n\n>> In the absence of it, wouldn't we in theory have to write a simple\n>> loop-on-EINTR wrapper for *all* syscalls?\n>\n> IIUC, that's what SA_RESTART is all about.\n\nYes, but there's precious little clear language on when SA_RESTART is\nsupposed to act.  In all cases?\n\nThe wording on\n\n  http://www.delorie.com/gnu/docs/glibc/libc_485.html\n  http://www.delorie.com/gnu/docs/glibc/libc_498.html\n\nleads me to believe that SA_RESTART is actually used on the glibc side\nof things, so that any glibc syscall wrapper not specifically equipped\nwith the restarting behavior would return EINTR unmodified.  This might\nexplain why utime() doesn't restart like it should (assuming we work on\nthe theory that POSIX doesn't allow an EINTR from utime() to begin\nwith).\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"207527","messageId":"50FF051D.5090804@giref.ulaval.ca","threadId":"32452","inReplyTo":"87r4lejpx8.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-22T21:31:09Z","receivedAt":"2013-01-22T21:31:09Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"So, hum, do we have some sort of conclusion?\n\nShall it be a fix for git to get around that lustre \"behavior\"?\n\nIf something can be done in git it would be great: it is a *lot* easier \nto change git than the lustre filesystem software for a cluster in \nrunning in production mode... (words from cluster team) :-/\n\nI hope this subject will not die in the list... :-/\n\nThanks,\n\nEric\n\n\n\nOn 01/21/2013 02:29 PM, Thomas Rast wrote:\n> Please don't drop the Cc list!\n>\n> \"Brian J. Murrell\" <brian@interlinx.bc.ca> writes:\n>\n>>> What's odd is that while I cannot reproduce the original problem, there\n>>> seems to be another issue/bug with utime():\n>>\n>> I wonder if this is related to http://jira.whamcloud.com/browse/LU-305.\n>>   That was reported as fixed in Lustre 2.0.0 and 2.1.0 but I thought I\n>> saw it on 2.1.1 and added a comment to the above ticket about that.\n>\n> Aha, that's a very interesting bug report.  My observations support\n> yours: I managed to get EINTR during utime().\n>\n>>> In the absence of it, wouldn't we in theory have to write a simple\n>>> loop-on-EINTR wrapper for *all* syscalls?\n>>\n>> IIUC, that's what SA_RESTART is all about.\n>\n> Yes, but there's precious little clear language on when SA_RESTART is\n> supposed to act.  In all cases?\n>\n> The wording on\n>\n>    http://www.delorie.com/gnu/docs/glibc/libc_485.html\n>    http://www.delorie.com/gnu/docs/glibc/libc_498.html\n>\n> leads me to believe that SA_RESTART is actually used on the glibc side\n> of things, so that any glibc syscall wrapper not specifically equipped\n> with the restarting behavior would return EINTR unmodified.  This might\n> explain why utime() doesn't restart like it should (assuming we work on\n> the theory that POSIX doesn't allow an EINTR from utime() to begin\n> with).\n>\n"},{"id":"207530","messageId":"7vip6onadi.fsf@alter.siamese.dyndns.org","threadId":"32452","inReplyTo":"50FF051D.5090804@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-01-22T22:03:37Z","receivedAt":"2013-01-22T22:03:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n\n> So, hum, do we have some sort of conclusion?\n>\n> Shall it be a fix for git to get around that lustre \"behavior\"?\n>\n> If something can be done in git it would be great: it is a *lot*\n> easier to change git than the lustre filesystem software for a cluster\n> in running in production mode... (words from cluster team) :-/\n\nDo we know the real cause of the symptom?  I did not follow the\nthread carefully, but the impression I was getting was that the\nfilesystem is broken around EINTR, and even if you \"fix\"ed Git,\nyour other more mission critical applications will be broken by\nthe same filesystem behaviour, no?\n"},{"id":"207531","messageId":"878v7keuh3.fsf@pctrast.inf.ethz.ch","threadId":"32452","inReplyTo":"50FF051D.5090804@giref.ulaval.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2013-01-22T22:14:16Z","receivedAt":"2013-01-22T22:14:16Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n\n> So, hum, do we have some sort of conclusion?\n>\n> Shall it be a fix for git to get around that lustre \"behavior\"?\n>\n> If something can be done in git it would be great: it is a *lot*\n> easier to change git than the lustre filesystem software for a cluster\n> in running in production mode... (words from cluster team) :-/\n\nI thought you already established that simply disabling the progress\ndisplay is a sufficient workaround?  If that doesn't help, you can try\npatching out all use of SIGALRM within git.\n\nOther than that I agree with Junio, from what we've seen so far, Lustre\nreturns EINTR on all sorts of calls that simply aren't allowed to do so.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"207534","messageId":"50FF16B0.8010206@giref.ulaval.ca","threadId":"32452","inReplyTo":"878v7keuh3.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-01-22T22:46:08Z","receivedAt":"2013-01-22T22:46:08Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"On 01/22/2013 05:14 PM, Thomas Rast wrote:\n> Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n>\n>> So, hum, do we have some sort of conclusion?\n>>\n>> Shall it be a fix for git to get around that lustre \"behavior\"?\n>>\n>> If something can be done in git it would be great: it is a *lot*\n>> easier to change git than the lustre filesystem software for a cluster\n>> in running in production mode... (words from cluster team) :-/\n>\n> I thought you already established that simply disabling the progress\n> display is a sufficient workaround?  If that doesn't help, you can try\n> patching out all use of SIGALRM within git.\n>\n\nI tried that solution after Brian told me to try it, but it didn't \nsolved the problem for me! :-(\n\n> Other than that I agree with Junio, from what we've seen so far, Lustre\n> returns EINTR on all sorts of calls that simply aren't allowed to do so.\n>\n\nOk, so now the \"good\" move would be to have all this reported to lustre \ndevelopment team?  Brian, have you seen anything new from what you have \nalready reported?  I have to admit that I am not a fs expert...\n\nAnd I also agree with Junio point of view: The problem may impact \nmission critical applications....\n\nEric\n"},{"id":"207596","messageId":"50FFF79F.6080103@calculquebec.ca","threadId":"32452","inReplyTo":"878v7keuh3.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Sébastien Boisvert","fromEmail":"sebastien.boisvert@calculquebec.ca","sentAt":"2013-01-23T14:45:51Z","receivedAt":"2013-01-23T14:45:51Z","isPatch":false,"sender":{"key":"sebastien.boisvert@calculquebec.ca","avatar":null},"body":"On 01/22/2013 05:14 PM, Thomas Rast wrote:\n> Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n>\n>> So, hum, do we have some sort of conclusion?\n>>\n>> Shall it be a fix for git to get around that lustre \"behavior\"?\n>>\n>> If something can be done in git it would be great: it is a *lot*\n>> easier to change git than the lustre filesystem software for a cluster\n>> in running in production mode... (words from cluster team) :-/\n>\n> I thought you already established that simply disabling the progress\n> display is a sufficient workaround?  If that doesn't help, you can try\n> patching out all use of SIGALRM within git.\n>\n\nIn git (9591fcc6d66), I have found these SIGALRM signal handling:\n\nbuiltin/log.c:268:\tsigaction(SIGALRM, &sa, NULL);\nbuiltin/log.c:285:\tsignal(SIGALRM, SIG_IGN);\ncompat/mingw.c:1590:\t\tmingw_raise(SIGALRM);\ncompat/mingw.c:1666:\tif (sig != SIGALRM)\ncompat/mingw.c:1668:\t\t\terror(\"sigaction only implemented for SIGALRM\");\ncompat/mingw.c:1683:\tcase SIGALRM:\ncompat/mingw.c:1702:\tcase SIGALRM:\ncompat/mingw.c:1706:\t\t\texit(128 + SIGALRM);\ncompat/mingw.c:1708:\t\t\ttimer_fn(SIGALRM);\ncompat/mingw.h:42:#define SIGALRM 14\nperl/Git/SVN.pm:2121:\t\t\tSIGALRM, SIGUSR1, SIGUSR2);\nprogress.c:56:\tsigaction(SIGALRM, &sa, NULL);\nprogress.c:68:\tsignal(SIGALRM, SIG_IGN);\n\n\nI suppose that compat/mingw.{h,c} and SVN.pm can be ignored as our patch to work\naround this problem won't be pushed upstream because the real problem is not in git, right ?\n\nIf I understand correctly, some VFS system calls get interrupted by SIGALRM, but when\nthey resume (via SA_RESTART) they return EINTR. Thomas said that these failed calls may need to be retried,\nbut that open(O_CREAT|O_EXCL) is still tricky around this case.\n\n\nprogress.c SIGALRM code paths are for progress and therefore are required, right ?\n\nbuiltin/log.c SIGALRM code paths are for early output, and the comments in the code say that\n\n    \"If we can get the whole output in less than a tenth of a second, don't even bother doing the\n     early-output thing.\"\n\n\nSo where do I start for the patch ?\n\n> Other than that I agree with Junio, from what we've seen so far, Lustre\n> returns EINTR on all sorts of calls that simply aren't allowed to do so.\n>\n\n\n-- \n---\nSpécialiste en granularité (1 journée / semaine)\nCalcul Québec / Calcul Canada\nPavillon Adrien-Pouliot, Université Laval, Québec (Québec), Canada\n"},{"id":"207598","messageId":"50FFF8B3.4070909@calculquebec.ca","threadId":"32452","inReplyTo":"878v7keuh3.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Sébastien Boisvert","fromEmail":"sebastien.boisvert@calculquebec.ca","sentAt":"2013-01-23T14:50:27Z","receivedAt":"2013-01-23T14:50:27Z","isPatch":false,"sender":{"key":"sebastien.boisvert@calculquebec.ca","avatar":null},"body":"[I forgot to subscribe to the git mailing list, sorry for that]\n\nOn 01/22/2013 05:14 PM, Thomas Rast wrote:\n> Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n>\n>> So, hum, do we have some sort of conclusion?\n>>\n>> Shall it be a fix for git to get around that lustre \"behavior\"?\n>>\n>> If something can be done in git it would be great: it is a *lot*\n>> easier to change git than the lustre filesystem software for a cluster\n>> in running in production mode... (words from cluster team) :-/\n>\n> I thought you already established that simply disabling the progress\n> display is a sufficient workaround?  If that doesn't help, you can try\n> patching out all use of SIGALRM within git.\n>\n\nIn git (9591fcc6d66), I have found these SIGALRM signal handling:\n\nbuiltin/log.c:268:    sigaction(SIGALRM, &sa, NULL);\nbuiltin/log.c:285:    signal(SIGALRM, SIG_IGN);\ncompat/mingw.c:1590:        mingw_raise(SIGALRM);\ncompat/mingw.c:1666:    if (sig != SIGALRM)\ncompat/mingw.c:1668:            error(\"sigaction only implemented for SIGALRM\");\ncompat/mingw.c:1683:    case SIGALRM:\ncompat/mingw.c:1702:    case SIGALRM:\ncompat/mingw.c:1706:            exit(128 + SIGALRM);\ncompat/mingw.c:1708:            timer_fn(SIGALRM);\ncompat/mingw.h:42:#define SIGALRM 14\nperl/Git/SVN.pm:2121:            SIGALRM, SIGUSR1, SIGUSR2);\nprogress.c:56:    sigaction(SIGALRM, &sa, NULL);\nprogress.c:68:    signal(SIGALRM, SIG_IGN);\n\n\nI suppose that compat/mingw.{h,c} and SVN.pm can be ignored as our patch to work\naround this problem won't be pushed upstream because the real problem is not in git, right ?\n\nIf I understand correctly, some VFS system calls get interrupted by SIGALRM, but when\nthey resume (via SA_RESTART) they return EINTR. Thomas said that these failed calls may need to be retried,\nbut that open(O_CREAT|O_EXCL) is still tricky around this case.\n\n\nprogress.c SIGALRM code paths are for progress and therefore are required, right ?\n\nbuiltin/log.c SIGALRM code paths are for early output, and the comments in the code say that\n\n    \"If we can get the whole output in less than a tenth of a second, don't even bother doing the\n     early-output thing.\"\n\n\nSo where do I start for the patch ?\n\n> Other than that I agree with Junio, from what we've seen so far, Lustre\n> returns EINTR on all sorts of calls that simply aren't allowed to do so.\n>\n\n\n-- \n---\nSpécialiste en granularité (1 journée / semaine)\nCalcul Québec / Calcul Canada\nPavillon Adrien-Pouliot, Université Laval, Québec (Québec), Canada\n"},{"id":"207599","messageId":"CABPQNSad1EKbmt3Gjs+uB9fud4YBqmk++5GMqF2s047Lcc8GwQ@mail.gmail.com","threadId":"32452","inReplyTo":"878v7keuh3.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2013-01-23T15:23:14Z","receivedAt":"2013-01-23T15:23:14Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Tue, Jan 22, 2013 at 11:14 PM, Thomas Rast <trast@student.ethz.ch> wrote:\n> Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n>\n> Other than that I agree with Junio, from what we've seen so far, Lustre\n> returns EINTR on all sorts of calls that simply aren't allowed to do so.\n\nI don't think this analysis is 100% accurate, POSIX allows error codes\nto be generated other than those defined. From\nhttp://pubs.opengroup.org/onlinepubs/009695399/functions/xsh_chap02_03.html:\n\n\"Implementations may support additional errors not included in this\nlist, *may generate errors included in this list under circumstances\nother than those described here*, or may contain extensions or\nlimitations that prevent some errors from occurring.\"\n\nSo I don't think Lustre violates POSIX by erroring with errno=EINTR,\nbut I also think guarding every single function call for EINTR just to\nbe safe to be insane :)\n\nHowever, looking at Eric's log, I can't see that being what has\nhappened here - grepping it for EINTR does not produce a single match.\n"},{"id":"207600","messageId":"87d2wvc3v0.fsf@pctrast.inf.ethz.ch","threadId":"32452","inReplyTo":"CABPQNSad1EKbmt3Gjs+uB9fud4YBqmk++5GMqF2s047Lcc8GwQ@mail.gmail.com","subject":"Re: GIT get corrupted on lustre","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2013-01-23T15:32:03Z","receivedAt":"2013-01-23T15:32:03Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Erik Faye-Lund <kusmabite@gmail.com> writes:\n\n> On Tue, Jan 22, 2013 at 11:14 PM, Thomas Rast <trast@student.ethz.ch> wrote:\n>> Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n>>\n>> Other than that I agree with Junio, from what we've seen so far, Lustre\n>> returns EINTR on all sorts of calls that simply aren't allowed to do so.\n>\n> I don't think this analysis is 100% accurate, POSIX allows error codes\n> to be generated other than those defined. From\n> http://pubs.opengroup.org/onlinepubs/009695399/functions/xsh_chap02_03.html:\n>\n> \"Implementations may support additional errors not included in this\n> list, *may generate errors included in this list under circumstances\n> other than those described here*, or may contain extensions or\n> limitations that prevent some errors from occurring.\"\n\nThat same page says, however:\n\n  For functions under the Threads option for which [EINTR] is not listed\n  as a possible error condition in this volume of IEEE Std 1003.1-2001,\n  an implementation shall not return an error code of [EINTR].\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"207601","messageId":"CABPQNSb89h28O_a3uVoVrNisZqPcHHVFm8nP7GdFGCb=PVdcsQ@mail.gmail.com","threadId":"32452","inReplyTo":"87d2wvc3v0.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2013-01-23T15:32:54Z","receivedAt":"2013-01-23T15:32:54Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Wed, Jan 23, 2013 at 4:32 PM, Thomas Rast <trast@student.ethz.ch> wrote:\n> Erik Faye-Lund <kusmabite@gmail.com> writes:\n>\n>> On Tue, Jan 22, 2013 at 11:14 PM, Thomas Rast <trast@student.ethz.ch> wrote:\n>>> Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n>>>\n>>> Other than that I agree with Junio, from what we've seen so far, Lustre\n>>> returns EINTR on all sorts of calls that simply aren't allowed to do so.\n>>\n>> I don't think this analysis is 100% accurate, POSIX allows error codes\n>> to be generated other than those defined. From\n>> http://pubs.opengroup.org/onlinepubs/009695399/functions/xsh_chap02_03.html:\n>>\n>> \"Implementations may support additional errors not included in this\n>> list, *may generate errors included in this list under circumstances\n>> other than those described here*, or may contain extensions or\n>> limitations that prevent some errors from occurring.\"\n>\n> That same page says, however:\n>\n>   For functions under the Threads option for which [EINTR] is not listed\n>   as a possible error condition in this volume of IEEE Std 1003.1-2001,\n>   an implementation shall not return an error code of [EINTR].\n\nYes, but surely that's for pthreads functions, no? utime is not one of\nthose functions...\n"},{"id":"207603","messageId":"871udbc3af.fsf@pctrast.inf.ethz.ch","threadId":"32452","inReplyTo":"CABPQNSb89h28O_a3uVoVrNisZqPcHHVFm8nP7GdFGCb=PVdcsQ@mail.gmail.com","subject":"Re: GIT get corrupted on lustre","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-01-23T15:44:24Z","receivedAt":"2013-01-23T15:44:24Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Erik Faye-Lund <kusmabite@gmail.com> writes:\n\n> On Wed, Jan 23, 2013 at 4:32 PM, Thomas Rast <trast@student.ethz.ch> wrote:\n>> Erik Faye-Lund <kusmabite@gmail.com> writes:\n>>\n>>> POSIX allows error codes\n>>> to be generated other than those defined. From\n>>> http://pubs.opengroup.org/onlinepubs/009695399/functions/xsh_chap02_03.html:\n>>>\n>>> \"Implementations may support additional errors not included in this\n>>> list, *may generate errors included in this list under circumstances\n>>> other than those described here*, or may contain extensions or\n>>> limitations that prevent some errors from occurring.\"\n>>\n>> That same page says, however:\n>>\n>>   For functions under the Threads option for which [EINTR] is not listed\n>>   as a possible error condition in this volume of IEEE Std 1003.1-2001,\n>>   an implementation shall not return an error code of [EINTR].\n>\n> Yes, but surely that's for pthreads functions, no? utime is not one of\n> those functions...\n\nAh, my bad.  In fact in\n\n  http://pubs.opengroup.org/onlinepubs/9699919799/xrat/V4_xsh_chap02.html\n\nthere is a paragraph \"Signal Effects on Other Functions\", which says\n\n  The most common behavior of an interrupted function after a\n  signal-catching function returns is for the interrupted function to\n  give an [EINTR] error unless the SA_RESTART flag is in effect for the\n  signal. However, there are a number of specific exceptions, including\n  sleep() and certain situations with read() and write().\n\n  The historical implementations of many functions defined by IEEE Std\n  1003.1-2001 are not interruptible[...]\n\n  Functions not mentioned explicitly as interruptible may be so on some\n  implementations, possibly as an extension where the function gives an\n  [EINTR] error. There are several functions (for example, getpid(),\n  getuid()) that are specified as never returning an error, which can\n  thus never be extended in this way.\n\n  If a signal-catching function returns while the SA_RESTART flag is in\n  effect, an interrupted function is restarted at the point it was\n  interrupted. Conforming applications cannot make assumptions about the\n  internal behavior of interrupted functions, even if the functions are\n  async-signal-safe. For example, suppose the read() function is\n  interrupted with SA_RESTART in effect, the signal-catching function\n  closes the file descriptor being read from and returns, and the read()\n  function is then restarted; in this case the application cannot assume\n  that the read() function will give an [EBADF] error, since read()\n  might have checked the file descriptor for validity before being\n  interrupted.\n\nTaken together this should mean that the bug is in fact simply that the\ncalls do not *restart*.  They are (like you say) allowed to return EINTR\ndespite not being specified to, *but* SA_RESTART should restart it.\n\nNow, does that make it a lustre bug or a glibc bug? :-)\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"207604","messageId":"CABPQNSaNmuXEn6hNd_xH01oCoykBs_y85=0bigmDBDH3Aazj2g@mail.gmail.com","threadId":"32452","inReplyTo":"871udbc3af.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2013-01-23T15:54:44Z","receivedAt":"2013-01-23T15:54:44Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Wed, Jan 23, 2013 at 4:44 PM, Thomas Rast <trast@inf.ethz.ch> wrote:\n> Erik Faye-Lund <kusmabite@gmail.com> writes:\n>\n>> On Wed, Jan 23, 2013 at 4:32 PM, Thomas Rast <trast@student.ethz.ch> wrote:\n>>> Erik Faye-Lund <kusmabite@gmail.com> writes:\n>>>\n>>>> POSIX allows error codes\n>>>> to be generated other than those defined. From\n>>>> http://pubs.opengroup.org/onlinepubs/009695399/functions/xsh_chap02_03.html:\n>>>>\n>>>> \"Implementations may support additional errors not included in this\n>>>> list, *may generate errors included in this list under circumstances\n>>>> other than those described here*, or may contain extensions or\n>>>> limitations that prevent some errors from occurring.\"\n>>>\n>>> That same page says, however:\n>>>\n>>>   For functions under the Threads option for which [EINTR] is not listed\n>>>   as a possible error condition in this volume of IEEE Std 1003.1-2001,\n>>>   an implementation shall not return an error code of [EINTR].\n>>\n>> Yes, but surely that's for pthreads functions, no? utime is not one of\n>> those functions...\n>\n> Ah, my bad.  In fact in\n>\n>   http://pubs.opengroup.org/onlinepubs/9699919799/xrat/V4_xsh_chap02.html\n>\n> there is a paragraph \"Signal Effects on Other Functions\", which says\n>\n> <snip>\n>\n> Taken together this should mean that the bug is in fact simply that the\n> calls do not *restart*.  They are (like you say) allowed to return EINTR\n> despite not being specified to, *but* SA_RESTART should restart it.\n>\n\nRight, thanks for clearing that up.\n\n> Now, does that make it a lustre bug or a glibc bug? :-)\n\nThat's kind of uninteresting, the important bit is that it is indeed a\nbug (outside of Git).\n"},{"id":"207611","messageId":"20130123172316.GA3238@elie.Belkin","threadId":"32452","inReplyTo":"871udbc3af.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2013-01-23T17:23:16Z","receivedAt":"2013-01-23T17:23:16Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Thomas Rast wrote:\n\n> Taken together this should mean that the bug is in fact simply that the\n> calls do not *restart*.  They are (like you say) allowed to return EINTR\n> despite not being specified to, *but* SA_RESTART should restart it.\n>\n> Now, does that make it a lustre bug or a glibc bug? :-)\n\nThe kernel takes care of SA_RESTART, if I remember correctly.  (See\narch/x86/kernel/signal.c::handle_signal() case -ERESTARTSYS.)\n"},{"id":"207615","messageId":"51002D18.4080201@calculquebec.ca","threadId":"32452","inReplyTo":"878v7keuh3.fsf@pctrast.inf.ethz.ch","subject":"Re: GIT get corrupted on lustre","fromName":"Sébastien Boisvert","fromEmail":"sebastien.boisvert@calculquebec.ca","sentAt":"2013-01-23T18:34:00Z","receivedAt":"2013-01-23T18:34:00Z","isPatch":false,"sender":{"key":"sebastien.boisvert@calculquebec.ca","avatar":null},"body":"Hello,\n\nHere is a patch (with git format-patch) that removes any timer if NO_SETITIMER is set.\n\n\nÉric:\n\nTo test it with your workflow:\n\n$ module load apps/git/1.8.1.1.348.g78eb407-NO_SETITIMER-patch\n\n$ git clone ...\n\n\n                               Sébastien\n\n\nOn 01/22/2013 05:14 PM, Thomas Rast wrote:\n> Eric Chamberland <Eric.Chamberland@giref.ulaval.ca> writes:\n>\n>> So, hum, do we have some sort of conclusion?\n>>\n>> Shall it be a fix for git to get around that lustre \"behavior\"?\n>>\n>> If something can be done in git it would be great: it is a *lot*\n>> easier to change git than the lustre filesystem software for a cluster\n>> in running in production mode... (words from cluster team) :-/\n>\n> I thought you already established that simply disabling the progress\n> display is a sufficient workaround?  If that doesn't help, you can try\n> patching out all use of SIGALRM within git.\n>\n> Other than that I agree with Junio, from what we've seen so far, Lustre\n> returns EINTR on all sorts of calls that simply aren't allowed to do so.\n>\n\n\n-- \n---\nSpécialiste en granularité (1 journée / semaine)\nCalcul Québec / Calcul Canada\nPavillon Adrien-Pouliot, Université Laval, Québec (Québec), Canada\n\n\n>From 78eb4075d98eb9cdc57210c63b8d8de8a3d0cd9e Mon Sep 17 00:00:00 2001\nFrom: =?UTF-8?q?S=C3=A9bastien=20Boisvert?= <sebastien.boisvert@calculquebec.ca>\nDate: Wed, 23 Jan 2013 13:10:57 -0500\nSubject: [PATCH] don't use timers if NO_SETITIMER is set\nMIME-Version: 1.0\nContent-Type: text/plain; charset=UTF-8\nContent-Transfer-Encoding: 8bit\n\nWith NO_SETITIMER, the user experience on legacy Lustre is fixed,\nbut there is no early progress.\n\nThe patch has no effect on the resulting git executable if NO_SETITIMER is\nnot set (the default). So by default this patch has no effect at all, which\nis good.\n\ngit tests:\n\n$ make clean\n$ make NO_SETITIMER=YesPlease\n$ make test NO_SETITIMER=YesPlease &> make-test.log\n\n$ grep \"^not ok\" make-test.log |grep -v \"# TODO known breakage\"|wc -l\n0\n$ grep \"^ok\" make-test.log |wc -l\n9531\n$ grep \"^not ok\" make-test.log |wc -l\n65\n\nNo timers with NO_SETITIMER:\n\n$ objdump -d ./git|grep setitimer|wc -l\n0\n$ objdump -d ./git|grep alarm|wc -l\n0\n\nTimers without NO_SETITIMER:\n\n$ objdump -d /software/apps/git/1.8.1/bin/git|grep setitimer|wc -l\n5\n$ objdump -d /software/apps/git/1.8.1/bin/git|grep alarm|wc -l\n0\n\nSigned-off-by: Sébastien Boisvert <sebastien.boisvert@calculquebec.ca>\n---\n builtin/log.c |    7 +++++++\n daemon.c      |    6 ++++++\n progress.c    |    8 ++++++++\n upload-pack.c |    2 ++\n 4 files changed, 23 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin/log.c b/builtin/log.c\nindex 8f0b2e8..f8321c7 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -198,7 +198,9 @@ static void show_early_header(struct rev_info *rev, const char *stage, int nr)\n \tprintf(_(\"Final output: %d %s\\n\"), nr, stage);\n }\n \n+#ifndef NO_SETITIMER\n static struct itimerval early_output_timer;\n+#endif\n \n static void log_show_early(struct rev_info *revs, struct commit_list *list)\n {\n@@ -240,9 +242,12 @@ static void log_show_early(struct rev_info *revs, struct commit_list *list)\n \t * trigger every second even if we're blocked on a\n \t * reader!\n \t */\n+\n+\t#ifndef NO_SETITIMER\n \tearly_output_timer.it_value.tv_sec = 0;\n \tearly_output_timer.it_value.tv_usec = 500000;\n \tsetitimer(ITIMER_REAL, &early_output_timer, NULL);\n+\t#endif\n }\n \n static void early_output(int signal)\n@@ -274,9 +279,11 @@ static void setup_early_output(struct rev_info *rev)\n \t *\n \t * This is a one-time-only trigger.\n \t */\n+\t#ifndef NO_SETITIMER\n \tearly_output_timer.it_value.tv_sec = 0;\n \tearly_output_timer.it_value.tv_usec = 100000;\n \tsetitimer(ITIMER_REAL, &early_output_timer, NULL);\n+\t#endif\n }\n \n static void finish_early_output(struct rev_info *rev)\ndiff --git a/daemon.c b/daemon.c\nindex 4602b46..eb82c19 100644\n--- a/daemon.c\n+++ b/daemon.c\n@@ -611,9 +611,15 @@ static int execute(void)\n \tif (addr)\n \t\tloginfo(\"Connection from %s:%s\", addr, port);\n \n+\t#ifndef NO_SETITIMER\n \talarm(init_timeout ? init_timeout : timeout);\n+\t#endif\n+\n \tpktlen = packet_read_line(0, line, sizeof(line));\n+\n+\t#ifndef NO_SETITIMER\n \talarm(0);\n+\t#endif\n \n \tlen = strlen(line);\n \tif (pktlen != len)\ndiff --git a/progress.c b/progress.c\nindex 3971f49..b84ccc7 100644\n--- a/progress.c\n+++ b/progress.c\n@@ -45,7 +45,10 @@ static void progress_interval(int signum)\n static void set_progress_signal(void)\n {\n \tstruct sigaction sa;\n+\n+\t#ifndef NO_SETITIMER\n \tstruct itimerval v;\n+\t#endif\n \n \tprogress_update = 0;\n \n@@ -55,16 +58,21 @@ static void set_progress_signal(void)\n \tsa.sa_flags = SA_RESTART;\n \tsigaction(SIGALRM, &sa, NULL);\n \n+\t#ifndef NO_SETITIMER\n \tv.it_interval.tv_sec = 1;\n \tv.it_interval.tv_usec = 0;\n \tv.it_value = v.it_interval;\n \tsetitimer(ITIMER_REAL, &v, NULL);\n+\t#endif\n }\n \n static void clear_progress_signal(void)\n {\n+\t#ifndef NO_SETITIMER\n \tstruct itimerval v = {{0,},};\n \tsetitimer(ITIMER_REAL, &v, NULL);\n+\t#endif\n+\n \tsignal(SIGALRM, SIG_IGN);\n \tprogress_update = 0;\n }\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 95d8313..e0b8b32 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -47,7 +47,9 @@ static int stateless_rpc;\n \n static void reset_timeout(void)\n {\n+\t#ifndef NO_SETITIMER\n \talarm(timeout);\n+\t#endif\n }\n \n static int strip(char *line, int len)\n-- \n1.7.4.1\n\n"},{"id":"208611","messageId":"510FBE90.6000105@giref.ulaval.ca","threadId":"32452","inReplyTo":"51002D18.4080201@calculquebec.ca","subject":"Re: GIT get corrupted on lustre","fromName":"Eric Chamberland","fromEmail":"eric.chamberland@giref.ulaval.ca","sentAt":"2013-02-04T13:58:40Z","receivedAt":"2013-02-04T13:58:40Z","isPatch":false,"sender":{"key":"eric.chamberland@giref.ulaval.ca","avatar":null},"body":"Hi,\n\nOn 01/23/2013 01:34 PM, Sébastien Boisvert wrote:\n> Hello,\n>\n> Here is a patch (with git format-patch) that removes any timer if\n> NO_SETITIMER is set.\n>\n\nEven with the patch, I finally got an error... :-/\n\nHere are the log (strace -f) of a clean execution and one with the error:\n\nhttp://www.giref.ulaval.ca/~ericc/NO_SETITIMER-patch_bin_git_noerror.txt.gz\n\nhttp://www.giref.ulaval.ca/~ericc/NO_SETITIMER-patch_bin_git_with_error.txt.gz\n\nEric\n"}]}