git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: pread() over NFS (again) [1.5.5.4]

From
Shawn O. Pearce <spearce@spearce.org>
Date
Jun 30, 2008, 00:32 UTC
Message-ID
<20080630003203.GJ11793@spearce.org>
In-Reply-To
<1214578229.7437.14.camel@localhost>
Trond Myklebust <Trond.Myklebust@netapp.com> wrote:
Show 57 quoted lines
> On Thu, 2008-06-26 at 22:57 -0400, J. Bruce Fields wrote:
> > On Thu, Jun 26, 2008 at 04:38:40PM -0700, Junio C Hamano wrote:
> > > logank@sent.com writes:
> > > 
> > > > On Jun 26, 2008, at 1:56 PM, Junio C Hamano wrote:
> > > >
> > > >>> "The file shouldn't be short unless someone truncated it, or there
> > > >>> is a bug in index-pack.  Neither is very likely, but I don't think
> > > >>> we would want to retry pread'ing the same block forever.
> > > >>
> > > >> I don't think we would want to retry even once.  Return value of 0
> > > >> from
> > > >> pread is defined to be an EOF, isn't it?
> > > >
> > > > No, it seems to be a simple error-out in this case. We have 2.4.20
> > > > systems with nfs-utils 0.3.3 and used to frequently get the same error
> > > > while pushing. I made a similar change back in February and haven't
> > > > had a problem since:
> > > >
> > > > diff --git a/index-pack.c b/index-pack.c
> > > > index 5ac91ba..85c8bdb 100644
> > > > --- a/index-pack.c
> > > > +++ b/index-pack.c
> > > > @@ -313,7 +313,14 @@ static void *get_data_from_pack(struct
> > > > object_entry *obj)
> > > > 	src = xmalloc(len);
> > > > 	data = src;
> > > > 	do {
> > > > +		// It appears that if multiple threads read across NFS, the
> > > > +		// second read will fail. I know this is awful, but we wait for
> > > > +		// a little bit and try again.
> > > > 		ssize_t n = pread(pack_fd, data + rdy, len - rdy, from + rdy);
> > > > +		if (n <= 0) {
> > > > +			sleep(1);
> > > > +			n = pread(pack_fd, data + rdy, len - rdy, from + rdy);
> > > > +		}
> > > > 		if (n <= 0)
> > > > 			die("cannot pread pack file: %s", strerror(errno));
> > > > 		rdy += n;
> > > >
> > > > I use a sleep request since it seems less likely that the other thread
> > > > will have an outstanding request after a second of waiting.
> > > 
> > > Gaah.  Don't we have NFS experts in house?  Bruce, perhaps?
> > 
> > Trond, you don't have any idea why a 2.6.9-42.0.8.ELsmp client (2.4.28
> > server) might be returning spurious 0's from pread()?
> > 
> > Seems like everything is happening from that one client--the file isn't
> > being simultaneously accessed from the server or from another client.
> 
> Is the file only being read, or could there be a simultaneous write to
> the same file? I'm surmising this could be an effect resulting from
> simultaneous cache invalidations: prior to Linux 2.6.20 or so, we
> weren't rigorously following the VFS/VM rules for page locking, and so
> page cache invalidation in particular could have some curious
> side-effects.

The file was created and opened O_CREAT|O_EXCL|O_RDWR, by this process, written linearly using write(2), without any lseeks. We kept the file descriptor open and starting issuing pread(2) calls for earlier offsets we had alread written. One of those kicks back EOF far too early (and results in this bug report).

Note the only accesses we are using is write(2) and pread(2), and once we start reading we don't ever go back to writing. The pread(2) calls are typically issued in ascending offsets, and we read each position only once. This is to try and take advantage of any read-ahead the kernel may be able to do. The pread(2) calls are rarely (if ever) on a block/page boundary.

Nobody else should know about this file. Its written to a temporary name and no other well behaved processes would attempt to read the file until it gets closed and renamed to its final destination. We haven't reached that far in the processing when we get this error, so there should be only one file descriptor open on the inode, and its the same one that wrote the data.

-- 
Shawn.
Previous: Trond MyklebustNext: Nicolas Pitre
Message 12 of 16 in “pread() over NFS (again) [1.5.5.4]”
  1. Christian HoltjeJun 26, 2008
  2. Shawn O. PearceJun 26, 2008
  3. Junio C HamanoJun 26, 2008
  4. Shawn O. PearceJun 26, 2008
  5. Christian HoltjeJun 26, 2008
  6. Junio C HamanoJun 26, 2008
  7. Shawn O. PearceJun 26, 2008
  8. logank@sent.comJun 26, 2008
  9. Junio C HamanoJun 26, 2008
  10. J. Bruce FieldsJun 27, 2008
  11. Trond MyklebustJun 27, 2008
  12. Shawn O. PearceJun 30, 2008
  13. Nicolas PitreJun 30, 2008
  14. J. Bruce FieldsJun 27, 2008
  15. Christian HoltjeJun 27, 2008
  16. Christian HoltjeJun 27, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.