git/list[1] front-page[2] threads[3] people[4] search[5] about
 

RE: reftable [v5]: new ref storage format

From
DTDavid Turner <david.turner@twosigma.com>
Date
Aug 14, 2017, 16:05 UTC
Message-ID
<4c1c1fc9904f4678823b6c3054c02b4d@exmbdft7.ad.twosigma.com>
In-Reply-To
<576a2361-1a3d-4bb2-1d31-f095f9e3c708@symas.com>
Show 32 quoted lines
> -----Original Message-----
> From: Howard Chu [mailto:hyc@symas.com]
> Sent: Monday, August 14, 2017 8:31 AM
> To: spearce@spearce.org
> Cc: David Turner <David.Turner@twosigma.com>; avarab@gmail.com;
> ben.alex@acegi.com.au; dborowitz@google.com; git@vger.kernel.org;
> gitster@pobox.com; mhagger@alum.mit.edu; peff@peff.net;
> sbeller@google.com; stoffe@gmail.com
> Subject: Re: reftable [v5]: new ref storage format
> 
> Howard Chu wrote:
> > The primary issue with using LMDB over NFS is with performance. All
> > reads are performed thru accesses of mapped memory, and in general,
> > NFS implementations don't cache mmap'd pages. I believe this is a
> > consequence of the fact that they also can't guarantee cache
> > coherence, so the only way for an NFS client to see a write from
> > another NFS client is by always refetching pages whenever they're accessed.
> 
> > LMDB's read lock management also wouldn't perform well over NFS; it
> > also uses an mmap'd file. On a local filesystem LMDB read locks are
> > zero cost since they just atomically update a word in the mmap. Over
> > NFS, each update to the mmap would also require an msync() to
> > propagate the change back to the server. This would seriously limit
> > the speed with which read transactions may be opened and closed.
> > (Ordinarily opening and closing a read txn can be done with zero
> > system calls.)
> 
> All that aside, we could simply add an EXCLUSIVE open-flag to LMDB, and
> prevent multiple processes from using the DB concurrently. In that case,
> maintaining coherence with other NFS clients is a non-issue. It strikes me that git
> doesn't require concurrent multi-process access anyway, and any particular
> process would only use the DB for a short time before closing it and going away.

Git, in general, does require concurrent multi-process access, depending on what that means.

For example, a post-receive hook might call some git command which opens the ref database. This means that git receive-pack would have to close and re-open the ref database. More generally, a fair number of git commands are implemented in terms of other git commands, and might need the same treatment. We could, in general, close and re-open the database around fork/exec, but I am not sure that this solves the general problem -- by mere happenstance, one might be e.g. pushing in one terminal while running git checkout in another. This is especially true with git worktrees, which share one ref database across multiple working directories.

Previous: Howard ChuNext: Jeff King
Message 11 of 12 in “Re: reftable [v5]: new ref storage format”
  1. Shawn PearceAug 6, 2017
  2. Ævar Arnfjörð BjarmasonAug 6, 2017
  3. Shawn PearceAug 6, 2017
  4. Shawn PearceAug 7, 2017
  5. David TurnerAug 7, 2017
  6. Jeff KingAug 8, 2017
  7. Shawn PearceAug 8, 2017
  8. Jeff KingAug 8, 2017
  9. Howard ChuAug 9, 2017
  10. Howard ChuAug 14, 2017
  11. David TurnerAug 14, 2017
  12. Jeff KingAug 15, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.