{"thread":{"id":"53782","subject":"Interrupted system call","startedAt":"2020-07-01T09:53:10Z","lastAt":"2020-07-15T16:09:13Z","messageCount":7,"participants":["R. Diez","Santiago Torres Arias","Jeff King","Chris Torek"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"400880","messageId":"c8061cce-71e4-17bd-a56a-a5fed93804da@neanderfunk.de","threadId":"53782","inReplyTo":"14b3d372-f3fe-c06c-dd56-1d9799a12632@yahoo.de","subject":"Interrupted system call","fromName":"R. Diez","fromEmail":"rdiez@neanderfunk.de","sentAt":"2020-07-01T09:43:15Z","receivedAt":"2020-07-01T09:53:10Z","isPatch":false,"sender":{"key":"rdiez@neanderfunk.de","avatar":null},"body":"Hi all:\n\nFirst of all, many thanks for Git.\n\nAfter a 3-month pause, I recently updated my Ubuntu 18.04.4. I am using a PPA to keep Git more up to date, so I have now \"git version 2.27.0\".\n\nI am now getting this kind of errors:\n\nfatal: failed to read object cf965547a433493caa80e84d7a2b78b32a26ee35: Interrupted system call\n\nerror: unable to mmap /home/rdiez/[blah blah]/SrcRepo.git/objects/2e/f96ffba4c0d60f36c8779758f82752be380689: Interrupted system call\n\nI am using a mount point for a network share. Keep in mind that Git thinks it is working on a local directory, so there should be no sockets \nor non-blocking I/O involved.\n\nThe problem is probably caused by using SMB to connect to an outdated Windows server. It has been working for years, but at some point in \ntime it is bound to fail. The Linux kernel itself seems to introduce bugs in the SMB/CIFS code every now and then.\n\nNevertheless, I am surprised to get such an \"Interrupted system call\" from Git. A long time ago I learnt that it is OK for many syscalls to \nget interrupted, so you have to loop around them. See here for more information:\n\nhttp://250bpm.com/blog:12\n\nAs a result, users should never actually get an \"Interrupted system call\" error from any software, at least when no sockets or non-blocking \nI/O is involved.\n\nHow can I pin-point this problem? I would like to know where Git is encountering this error, so that I can troubleshoot it, and maybe report \nyet another bug to the Linux SMB/CIFS maintainer.\n\nThanks in advance,\n   rdiez\n"},{"id":"400920","messageId":"20200701142205.3gwltulmkrd6iljk@LykOS.localdomain","threadId":"53782","inReplyTo":"c8061cce-71e4-17bd-a56a-a5fed93804da@neanderfunk.de","subject":"Re: Interrupted system call","fromName":"Santiago Torres Arias","fromEmail":"santiago@nyu.edu","sentAt":"2020-07-01T14:22:05Z","receivedAt":"2020-07-01T14:22:12Z","isPatch":false,"sender":{"key":"santiago@nyu.edu","avatar":"https://avatars.githubusercontent.com/u/3579933?v=4"},"body":"Hi, \n\n> Nevertheless, I am surprised to get such an \"Interrupted system call\" from\n> Git. A long time ago I learnt that it is OK for many syscalls to get\n> interrupted, so you have to loop around them. See here for more information:\n> \n> https://urldefense.proofpoint.com/v2/url?u=http-3A__250bpm.com_blog-3A12&d=DwICaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=yZMPY-APGKyVIX7HgQFZJA&m=JwtG1XJ8aqvchYKsbjW23-PqEl4qm4xuOrYLaF8MOK4&s=k58MMdPdIRPl0kpuTohwZo_3GbW7elvojU1wjTil2GY&e=\n> \n> As a result, users should never actually get an \"Interrupted system call\"\n> error from any software, at least when no sockets or non-blocking I/O is\n> involved.\n\nI'm not sure if you can blame git right away (it could be an underlying\nlibrary), and I'm also not convinced that \"interrupted system call\" is\nan error that should never exist for users (error handling is generally\nvery nuanced).\n\nI'd advice to use GIT_TRACE_FSMONITOR or just GIT_TRACE to figure out\nwhat component is the last one in place before things failed. You can\nread about these on the manpage on the \"other\" subsection of the\n\"ENVIRONMENT VARIABLES\" section.\n\nI hope this helps!\n-Santiago\n"},{"id":"400933","messageId":"20200701162111.GA934052@coredump.intra.peff.net","threadId":"53782","inReplyTo":"c8061cce-71e4-17bd-a56a-a5fed93804da@neanderfunk.de","subject":"Re: Interrupted system call","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-07-01T16:21:11Z","receivedAt":"2020-07-01T16:21:13Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jul 01, 2020 at 11:43:15AM +0200, R. Diez wrote:\n\n> After a 3-month pause, I recently updated my Ubuntu 18.04.4. I am\n> using a PPA to keep Git more up to date, so I have now \"git version\n> 2.27.0\".\n> \n> I am now getting this kind of errors:\n> \n> fatal: failed to read object cf965547a433493caa80e84d7a2b78b32a26ee35: Interrupted system call\n> \n> error: unable to mmap /home/rdiez/[blah blah]/SrcRepo.git/objects/2e/f96ffba4c0d60f36c8779758f82752be380689: Interrupted system call\n> \n> I am using a mount point for a network share. Keep in mind that Git thinks\n> it is working on a local directory, so there should be no sockets or\n> non-blocking I/O involved.\n\nLooking at the code, that message is slightly deceptive. It's reporting\na failure from map_loose_object_1(), which calls both open() and mmap(),\nas well as fstat().  It would be interesting to know which syscall is\nactually failing. Running the failure case under \"strace\" would be\ninteresting (likewise to see which signal is causing the interruption).\n\n> The problem is probably caused by using SMB to connect to an outdated\n> Windows server. It has been working for years, but at some point in time it\n> is bound to fail. The Linux kernel itself seems to introduce bugs in the\n> SMB/CIFS code every now and then.\n> \n> Nevertheless, I am surprised to get such an \"Interrupted system call\" from\n> Git. A long time ago I learnt that it is OK for many syscalls to get\n> interrupted, so you have to loop around them. See here for more information:\n\nWe do check for signals and re-start read() and write() calls as\nappropriate. We don't for open(), and nobody has ever complained (though\nit definitely is documented to result in EINTR, I'd imagine it's\nrelatively rare). I'm not excited about the prospect of adding retry\ncode to every open(), though perhaps doing it with our git_open()\nwrapper would be sufficient (it's unclear how stdio fopen() behaves).\n\n> How can I pin-point this problem? I would like to know where Git is\n> encountering this error, so that I can troubleshoot it, and maybe report yet\n> another bug to the Linux SMB/CIFS maintainer.\n\nI think the first step is using strace to record the system call\nreturning EINTR (and the signal that interrupted it). I suspect it's in\nopen(), though, and probably not a bug: opening network files may take a\nwhile and need to be interruptable.\n\n-Peff\n"},{"id":"400961","messageId":"11754dcc-3c88-04dd-d009-89da01881e5d@neanderfunk.de","threadId":"53782","inReplyTo":"20200701162111.GA934052@coredump.intra.peff.net","subject":"Re: Interrupted system call","fromName":"R. Diez","fromEmail":"rdiez@neanderfunk.de","sentAt":"2020-07-02T07:07:46Z","receivedAt":"2020-07-02T07:07:52Z","isPatch":false,"sender":{"key":"rdiez@neanderfunk.de","avatar":null},"body":"\n > [...]\n> It would be interesting to know which syscall is\n> actually failing. Running the failure case under \"strace\" would be\n> interesting (likewise to see which signal is causing the interruption).\n > [...]\n\n\nFirst of all, thanks for your help.\n\nGIT_TRACE alone does not tell me anything useful:\n\n$ GIT_TRACE=true git fsck\n07:58:47.229138 git.c:442               trace: built-in: git fsck\nerror: unable to mmap ./objects/cb/fec04963c1090535d2670b741912e17fd27b27: Interrupted system call\nerror: cbfec04963c1090535d2670b741912e17fd27b27: object corrupt or missing: ./objects/cb/fec04963c1090535d2670b741912e17fd27b27\nChecking object directories: 100% (256/256), done.\nChecking objects: 100% (70229/70229), done.\nChecking connectivity: 75316, done.\nmissing commit cbfec04963c1090535d2670b741912e17fd27b27\ndangling commit 6835e962b227e957520addbc5c28aedc97b253f3\ndangling tree a9d1a1321066d8a8402f1c9e584675146d250952\n\n\nGIT_TRACE_FSMONITOR does not either:\n\n$ GIT_TRACE_FSMONITOR=true git fsck \n\nerror: unable to mmap ./objects/56/af267465e7cdb7ccebe8242e55c03d4b675684: Interrupted system call\nerror: 56af267465e7cdb7ccebe8242e55c03d4b675684: object corrupt or missing: ./objects/56/af267465e7cdb7ccebe8242e55c03d4b675684\nChecking object directories: 100% (256/256), done.\nChecking objects: 100% (70229/70229), done.\nChecking connectivity: 75666, done.\nmissing tree 56af267465e7cdb7ccebe8242e55c03d4b675684\n\nIt is the same Git repository, so it looks like every time a different, random file fails.\n\n\nI managed to make it fail once with:\n\n   strace -f -- git fsck --progress\n\nThe signal involved is SIGALRM. I am guessing that Git is setting it up in order to display its progress messages. This is one of the few \ncalls to rt_sigaction(SIGALRM):\n\nrt_sigaction(SIGALRM, {sa_handler=0x556c8ac0fe80, sa_mask=[], sa_flags=SA_RESTORER|SA_RESTART, sa_restorer=0x7fbdca7da890}, NULL, 8) = 0\n\n\nThis is the first failure:\n\nopenat(AT_FDCWD, \"./objects/11/a327f469cc40015d6d873f6eed328e977c4234\", O_RDONLY|O_CLOEXEC) = -1 EINTR (Interrupted system call)\n--- SIGALRM {si_signo=SIGALRM, si_code=SI_KERNEL} ---\nrt_sigreturn({mask=[]})                 = -1 EINTR (Interrupted system call)\nopenat(AT_FDCWD, \"/usr/share/locale/en_US/LC_MESSAGES/libc.mo\", O_RDONLY) = -1 ENOENT (No such file or directory)\nopenat(AT_FDCWD, \"/usr/share/locale/en/LC_MESSAGES/libc.mo\", O_RDONLY) = -1 ENOENT (No such file or directory)\nopenat(AT_FDCWD, \"/usr/share/locale-langpack/en_US/LC_MESSAGES/libc.mo\", O_RDONLY) = -1 ENOENT (No such file or directory)\nopenat(AT_FDCWD, \"/usr/share/locale-langpack/en/LC_MESSAGES/libc.mo\", O_RDONLY) = -1 ENOENT (No such file or directory)\nwrite(2, \"error: unable to mmap ./objects/\"..., 99error: unable to mmap ./objects/11/a327f469cc40015d6d873f6eed328e977c4234: Interrupted \nsystem call\n) = 99\nwrite(2, \"error: 11a327f469cc40015d6d873f6\"..., 128error: 11a327f469cc40015d6d873f6eed328e977c4234: object corrupt or missing: \n./objects/11/a327f469cc40015d6d873f6eed328e977c4234\n) = 128\n\n\nThis is the second one:\n\nopenat(AT_FDCWD, \"./objects/18/5b82729943708795b635899348ecca97aa7804\", O_RDONLY|O_CLOEXEC) = -1 EINTR (Interrupted system call)\n--- SIGALRM {si_signo=SIGALRM, si_code=SI_KERNEL} ---\nrt_sigreturn({mask=[]})                 = -1 EINTR (Interrupted system call)\nwrite(2, \"error: unable to mmap ./objects/\"..., 99error: unable to mmap ./objects/18/5b82729943708795b635899348ecca97aa7804: Interrupted \nsystem call\n) = 99\nwrite(2, \"error: 185b82729943708795b635899\"..., 128error: 185b82729943708795b635899348ecca97aa7804: object corrupt or missing: \n./objects/18/5b82729943708795b635899348ecca97aa7804\n) = 128\n\nThere are a few more failures.\n\nThis is the last one. Afterwards, Git exited:\n\nopenat(AT_FDCWD, \"./objects/f4/56439700761946c57ef467a8a125a80f0304bd\", O_RDONLY|O_CLOEXEC) = -1 EINTR (Interrupted system call)\n--- SIGALRM {si_signo=SIGALRM, si_code=SI_KERNEL} ---\nrt_sigreturn({mask=[]})                 = -1 EINTR (Interrupted system call)\nopenat(AT_FDCWD, \"./objects/pack\", O_RDONLY|O_NONBLOCK|O_CLOEXEC|O_DIRECTORY) = 3\nfstat(3, {st_mode=S_IFDIR|0755, st_size=0, ...}) = 0\nbrk(0x556c934af000)                     = 0x556c934af000\ngetdents(3, /* 19 entries */, 1048576)  = 1272\ngetdents(3, /* 0 entries */, 1048576)   = 0\nclose(3)                                = 0\nwrite(2, \"fatal: failed to read object f45\"..., 95fatal: failed to read object f456439700761946c57ef467a8a125a80f0304bd: Interrupted system call\n) = 95\nexit_group(128)                         = ?\n+++ exited with 128 +++\n\n\nI am not an expert in Unix signals, but I'll do my best here.\n\nI do not understand why Git is getting these interruptions due to SIGALRM, because SA_RESTART is in place.\n\nInterestingly, the man page signal(7) does list open() under that flag, but not openat().\n\nThe description for open() under SA_RESTART is also interesting:\n\n* open(2), if it can block (e.g., when opening a FIFO; see fifo(7)).\n\nI am not sure that opening a normal disk file may qualify as \"can block\" with that definition though.\n\nBest regards,\n   rdiez\n"},{"id":"401404","messageId":"39d520fd-5b46-0163-e7af-ff13d647b133@neanderfunk.de","threadId":"53782","inReplyTo":"20200701162111.GA934052@coredump.intra.peff.net","subject":"Re: Interrupted system call","fromName":"R. Diez","fromEmail":"rdiez@neanderfunk.de","sentAt":"2020-07-12T08:41:54Z","receivedAt":"2020-07-12T08:43:01Z","isPatch":false,"sender":{"key":"rdiez@neanderfunk.de","avatar":null},"body":"\n> fatal: failed to read object cf965547a433493caa80e84d7a2b78b32a26ee35: Interrupted system call\n > [...]\n\nIn case anybody else has the same problem and finds this thread in the future, the workaround I am using is to disable progress messages.\n\nFor \"git pull\", \"git gc\" and \"git push\" the appropriate option is \"--quiet\", but for \"git fsck\" it is \"--no-progress\".\n\nBest regards,\n   rdiez\n"},{"id":"401581","messageId":"20200715093831.GA3259535@coredump.intra.peff.net","threadId":"53782","inReplyTo":"11754dcc-3c88-04dd-d009-89da01881e5d@neanderfunk.de","subject":"Re: Interrupted system call","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-07-15T09:38:31Z","receivedAt":"2020-07-15T09:38:33Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 02, 2020 at 09:07:46AM +0200, R. Diez wrote:\n\n> I managed to make it fail once with:\n> \n>   strace -f -- git fsck --progress\n> \n> The signal involved is SIGALRM. I am guessing that Git is setting it up in\n> order to display its progress messages. This is one of the few calls to\n> rt_sigaction(SIGALRM):\n> \n> rt_sigaction(SIGALRM, {sa_handler=0x556c8ac0fe80, sa_mask=[], sa_flags=SA_RESTORER|SA_RESTART, sa_restorer=0x7fbdca7da890}, NULL, 8) = 0\n\nThat makes sense (and likewise your \"--quiet\" workaround seems\nreasonable).\n\n> I am not an expert in Unix signals, but I'll do my best here.\n> \n> I do not understand why Git is getting these interruptions due to SIGALRM, because SA_RESTART is in place.\n> \n> Interestingly, the man page signal(7) does list open() under that flag, but not openat().\n\nYes, though since open(2) says:\n\n The openat() system call operates in exactly the same way as open(),\n except for the differences described here.\n\nI'd expect that would include any SA_RESTART handling. Peeking at the\nLinux implementation in fs/open.c, it looks like both syscalls quickly\nend up in the same do_sys_open().\n\n> The description for open() under SA_RESTART is also interesting:\n> \n> * open(2), if it can block (e.g., when opening a FIFO; see fifo(7)).\n> \n> I am not sure that opening a normal disk file may qualify as \"can block\" with that definition though.\n\nDelivering EINTR on a non-blocking call seems even more confusing,\nthough. I think the \"if it can block\" is just \"you won't even get a\nsignal if it's not blocking\".\n\nThis really _seems_ like a kernel bug, either:\n\n  - openat() does not get the same SA_RESTART treatment as open(); or\n\n  - open() on a network file can get EINTR even with SA_RESTART\n\nBut it's quite possible that I'm missing some corner case or historical\nreason that it would need to behave the way you're seeing. It might be\nworth reporting to kernel folks.\n\n-Peff\n"},{"id":"401592","messageId":"CAPx1Gve48S+6VmpWD4FoYJ0MzMUVBZB_=Mu2bXy3RoLoZsJBMA@mail.gmail.com","threadId":"53782","inReplyTo":"20200715093831.GA3259535@coredump.intra.peff.net","subject":"Re: Interrupted system call","fromName":"Chris Torek","fromEmail":"chris.torek@gmail.com","sentAt":"2020-07-15T16:06:48Z","receivedAt":"2020-07-15T16:09:13Z","isPatch":false,"sender":{"key":"chris.torek@gmail.com","avatar":"https://avatars.githubusercontent.com/u/16826774?v=4"},"body":"On Wed, Jul 15, 2020 at 2:45 AM Jeff King <peff@peff.net> wrote:\n> On Thu, Jul 02, 2020 at 09:07:46AM +0200, R. Diez wrote:\n> > I do not understand why Git is getting these interruptions due to SIGALRM, because SA_RESTART is in place.\n\nIt really shouldn't -- that's the whole point of SA_RESTART.\n\n> Delivering EINTR on a non-blocking call seems even more confusing,\n> though. I think the \"if it can block\" is just \"you won't even get a\n> signal if it's not blocking\".\n>\n> This really _seems_ like a kernel bug, either:\n>\n>   - openat() does not get the same SA_RESTART treatment as open(); or\n>\n>   - open() on a network file can get EINTR even with SA_RESTART\n>\n> But it's quite possible that I'm missing some corner case or historical\n> reason that it would need to behave the way you're seeing. It might be\n> worth reporting to kernel folks.\n>\n> -Peff\n\nRight.  This goes way back to pre-v7-Unix signals, as a sort of a\nside effect of the implementation.  In ancient times, the kernel\ncode for the internal wait-for-some-event took a priority number,\nand anything below a cutoff value meant \"not interrupted by\nsignals\" while anything above it meant \"interrupted by signals\".\nDisk operations were all at PRIBIO which was never interrupted.\n\nThis is all quite different in modern systems and hence it's all\nadjustable, but in general we like to distinguish between\n\"operations that will definitely complete fairly quickly\"\n(normally not interrupted) and \"operations that might take\nsignificant amounts of time\" (normally interrupted with the option\nof restarting the system call).\n\n*Restarting*, though, means exactly that: not *resuming*, but\n*restarting*.  So whatever system call is to be interrupted by the\nsignal *must* be one that can simply be started over from the\nbeginning.  That means, for instance, that read() or write() can\nonly be restarted if no data have yet moved.  So if you're in a\nread() on a device (e.g., serial port, or tape drive, or whatever)\nand have gotten a few bytes, but not yet all you wanted, and then\nthe system call is to be interrupted by a signal, the read() must\nreturn with a short count.\n\nAn open() can be restarted on the assumption that no path names\nhave been changed.  That's not necessarily a good assumption,\nbut it's traditional.  The openat() can be restarted for the same\nreason (and in fact correct use of openat() can protect against\nsome pathname issues).  It's up to the programmer to decide\nwhether to use SA_RESTART, and hence allow this, or not.\n\nChris\n"}]}