git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: kernel.org mirroring (Re: [GIT PULL] MMC update)

From
Martin Langhoff <martin.langhoff@gmail.com>
Date
Dec 9, 2006, 01:56 UTC
Message-ID
<46a038f90612081756w1ab4609epcb4a2cbd9f4d8205@mail.gmail.com>
In-Reply-To
<Pine.LNX.4.64.0612081453430.3516@woody.osdl.org>
On 12/9/06, Linus Torvalds <torvalds@osdl.org> wrote:
> Actually, just looking at the examples, it looks like memcached is
> fundamentally flawed, exactly the same way Apache mod_cache is
> fundamentally flawed.

I don't know if fundamentally flawed but (having used memcached) I don't think it's a big win for this at all.

We can make gitweb to detect mod_perl and a few smarter things if it is running inside of it. In fact, we can (ab)use mod_perl and perl facilities a bit to do some serialization which will be a big win for some pages. What we need for that is to set a sensible the ETag and use some IPC to announce/check if other apache/modperl processes are preparing content for the same ETag. The first-process-to-announce a given ETag can then write it to a common temp directory (atomically - write to a temp-name and move to the expected name) while other processes wait, polling for the file. Once the file is in place the latecomers can just serve the content of the file and exit.

(I am calling the "state we are serving" identifier ETag because I think we should also set it as the ETag in the HTTP headers, so well be able to check the ETag of future requests for staleness - all we need is a ref lookup, and if the SHA1 matches, we are sorted). So having this 'unique request identifier' doubles up nicely...

The ETag should probably be:
 - SHA1+displaytype+args for pages that display an object identified by SHA1
 - refname+SHA!+displaytype+args for pages that display something
identified by a ref
 - SHA1(names and sha1s of all refs) for the summary page
> You can't have a cache architecture where the client just does a "get",
> like memcached does. You need to have a "read-for-fill" operation, which
> says:

You _could_ make do with a convention of polling for "entryname" and "workingon-entryname" and if "workingon-entryname" is set to 1, you can expect entryname to be filled real soon now. However, memcached is completely memorybound, so it is only nice for really small stuff or for a large server farm which has gobs of spare ram.

(Note that memcached does have timeouts which means that the 'workingon' value could have a short timeout in case the request is cancelled or the process dies - the nasty bit in the above plan would be the polling.)

> I still don't understand why apache doesn't do it. I guess it wants to be
> stateless or something.

Apache doesn't do it because most web applications don't use the HTTP procol correctly - specially when it comes to the idempotency of GET. So in 99% of the cases, web apps serve truly different pages for the same GET request, depending on your cookie, IP address, time-of-day, etc.

Most websites deal with very little traffic, so this isn't a problem. And many large sites that serve a lot of traffic from a dynamic web app want to be serving custom ads, let you login and see your personalised toolbar, etc,etc, so this wouldn't work for them either.

So in practice, serialising speculatively on GET requests for the same URL has very little payoff except for static content. And that's quite fast anyway.... specially if the underlying OS is smokin' fast ;-)

cheers,
Previous: Jeff GarzikNext: Jakub Narebski
Message 38 of 80 in “Re: kernel.org mirroring (Re: [GIT PULL] MMC update)”
  1. Linus TorvaldsDec 7, 2006
  2. H. Peter AnvinDec 7, 2006
  3. Olivier GalibertDec 7, 2006
  4. H. Peter AnvinDec 7, 2006
  5. Olivier GalibertDec 7, 2006
  6. H. Peter AnvinDec 7, 2006
  7. Jakub NarebskiDec 8, 2006
  8. Rogan DawesDec 8, 2006
  9. Jakub NarebskiDec 8, 2006
  10. Rogan DawesDec 8, 2006
  11. Jonas FonsecaDec 8, 2006
  12. Martin LanghoffDec 9, 2006
  13. H. Peter AnvinDec 9, 2006
  14. Martin LanghoffDec 9, 2006
  15. H. Peter AnvinDec 9, 2006
  16. Martin LanghoffDec 9, 2006
  17. H. Peter AnvinDec 9, 2006
  18. H. Peter AnvinDec 8, 2006
  19. Linus TorvaldsDec 8, 2006
  20. H. Peter AnvinDec 8, 2006
  21. Lars HjemliDec 8, 2006
  22. H. Peter AnvinDec 8, 2006
  23. Lars HjemliDec 8, 2006
  24. H. Peter AnvinDec 8, 2006
  25. rdaDec 10, 2006
  26. Jeff GarzikDec 8, 2006
  27. H. Peter AnvinDec 8, 2006
  28. Jeff GarzikDec 8, 2006
  29. Linus TorvaldsDec 8, 2006
  30. Michael K. EdwardsDec 8, 2006
  31. H. Peter AnvinDec 8, 2006
  32. Michael K. EdwardsDec 9, 2006
  33. H. Peter AnvinDec 9, 2006
  34. Linus TorvaldsDec 9, 2006
  35. H. Peter AnvinDec 9, 2006
  36. Michael K. EdwardsDec 9, 2006
  37. Jeff GarzikDec 9, 2006
  38. Martin LanghoffDec 9, 2006
  39. Jakub NarebskiDec 9, 2006
  40. Jeff GarzikDec 9, 2006
  41. Jakub NarebskiDec 9, 2006
  42. Jeff GarzikDec 9, 2006
  43. Jakub NarebskiDec 9, 2006
  44. Jeff GarzikDec 9, 2006
  45. Martin LanghoffDec 10, 2006
  46. Jakub NarebskiDec 10, 2006
  47. Jeff GarzikDec 10, 2006
  48. Jakub NarebskiDec 10, 2006
  49. Jeff GarzikDec 10, 2006
  50. Jakub NarebskiDec 10, 2006
  51. Linus TorvaldsDec 10, 2006
  52. Jakub NarebskiDec 10, 2006
  53. Linus TorvaldsDec 10, 2006
  54. Martin LanghoffDec 10, 2006
  55. Jeff GarzikDec 10, 2006
  56. Jeff GarzikDec 10, 2006
  57. H. Peter AnvinDec 10, 2006
  58. Jeff GarzikDec 10, 2006
  59. Jakub NarebskiDec 10, 2006
  60. Martin LanghoffDec 11, 2006
  61. Jakub NarebskiDec 11, 2006
  62. Martin LanghoffDec 11, 2006
  63. Linus TorvaldsDec 9, 2006
  64. H. Peter AnvinDec 9, 2006
  65. Martin LanghoffDec 10, 2006
  66. H. Peter AnvinDec 10, 2006
  67. Jakub NarebskiDec 12, 2006
  68. Steven GrimmDec 9, 2006
  69. Linus TorvaldsDec 7, 2006
  70. Shawn PearceDec 7, 2006
  71. Linus TorvaldsDec 7, 2006
  72. Michael K. EdwardsDec 7, 2006
  73. H. Peter AnvinDec 7, 2006
  74. Junio C HamanoDec 7, 2006
  75. H. Peter AnvinDec 7, 2006
  76. Junio C HamanoDec 7, 2006
  77. Jakub NarebskiDec 8, 2006
  78. Linus TorvaldsDec 9, 2006
  79. H. Peter AnvinDec 9, 2006
  80. Jeff GarzikDec 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.