git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Gitweb caching: Google Summer of Code project

From
Jakub Narebski <jnareb@gmail.com>
Date
May 29, 2008, 23:27 UTC
Message-ID
<200805300127.10454.jnareb@gmail.com>
In-Reply-To
<483DA594.5040803@gmail.com>
On Wed, 28 May 2008, Lea Wiemann wrote:
Show 12 quoted lines
> Jakub Narebski wrote:
> >
> > 1. Caching data
> >  * disadvantages:
> >    - more CPU
> >    - need to serialize and deserialize (parse) data
> >    - more complicated
> 
> CPU: John told me that so far CPU has *never* been an issue on k.org. 
> Unless someone tells me they've had CPU problems, I'll assume that CPU 
> is a non-issue until I actually run into it (and then I can optimize the 
> particular pieces where CPU is actually an issue).
True.

What you have to care about (although I don't think it would be partilcularly difficult) is to not repeat bad I/O patterns with cache...

> Serialization: I was planning to use Storable (memcached's Perl API uses 
> it transparently I think).  I'm hoping that this'll just solve it.
While Storable is part of, I think, any modern Perl installation, there
might be problem with memcached API, and memcached API wrappers such as
CHI one.  Namely you cannot assume that memcached API is installed, so
you have to provide some kind of fallback.
 
> It's true that it's more complicated.  It'll require quite a bit of 
> refactoring, and maybe I'll just back off if I find that it's too hard.

What's more, if you want to implement If-Modified-Since and If-None-Match, you would have to implement it by yourself, while for static pages (cahing HTML output) web server would do this for us "for free".

Show 8 quoted lines
> > I'm afraid that implementing kernel.org caching in mainline in
> > a generic way would be enough work for a whole GSoC 2008.
> 
> I probably won't reimplement the current caching mechanism.  Do you 
> think that a solution using memcached is generic enough?  I'll still 
> need to add some abstraction layer in the code, but when I'm finished 
> the user will either get the normal uncached gitweb, or activate 
> memcached caching with some configuration setting.

Thats good enough, although I think that current caching mechanism in kernel.org's gitweb (your implementation follows more what repo.or.cz's gitweb does) has some good ideas, like for example adaptive (depending on load) expiry time.

By the way what do you think about adding (as an option) information
about gitweb performance to the output, in the form of
  "Site generated in 0.01 seconds, 2 calls to git commands"
or
  "Site generated in 0.0023 seconds, cached output, 1m31s old"
line somewhere in the page footer?

I hope you have some ideas in gitweb access statistics from kernel.org, repo.or.cz, and perhaps other large git hosting sites (e.g. freedesktop.org), and you plan on benchamrking gitweb caching using average / amortized time to generate page, ApacheBench or equivalent, load average on server depending on number of requests, I/O load (using fio tool, for example) depending on number of requests etc.

> By the way, I'll be posting about gitweb on this mailing list 
> occasionally.  If any of you would like to receive CC's on such 
> messages, please let me know, otherwise I'll assume you get them through 
> the mailing list.

I read git mailing list via Usenet / news interface (NNTP gateway) from GMane.

-- 
Jakub Narebski
Poland
Previous: Lea WiemannNext: Lea Wiemann
Message 6 of 18 in “Gitweb caching: Google Summer of Code project”
  1. Lea WiemannMay 27, 2008
  2. Jakub NarebskiMay 27, 2008
  3. Lea WiemannMay 27, 2008
  4. Jakub NarebskiMay 28, 2008
  5. Lea WiemannMay 28, 2008
  6. Jakub NarebskiMay 29, 2008
  7. Lea WiemannMay 30, 2008
  8. Jakub NarebskiMay 30, 2008
  9. Lea WiemannMay 30, 2008
  10. Petr BaudisMay 30, 2008
  11. Lea WiemannMay 30, 2008
  12. Petr BaudisMay 30, 2008
  13. Rafael Garcia-SuarezMay 30, 2008
  14. J.H.May 30, 2008
  15. Junio C HamanoMay 30, 2008
  16. Lea WiemannMay 30, 2008
  17. Lea WiemannMay 30, 2008
  18. Jakub NarebskiMay 31, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.