git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: kernel.org mirroring (Re: [GIT PULL] MMC update)

From
RDRogan Dawes <discard@dawes.za.net>
Date
Dec 8, 2006, 12:57 UTC
Message-ID
<4579611F.5010303@dawes.za.net>
In-Reply-To
<4578722E.9030402@zytor.com>
H. Peter Anvin wrote:
Show 11 quoted lines
> Olivier Galibert wrote:
>> On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:
>>> Unfortunately, the most common queries are also extremely expensive.
>>
>> Do you have a top-ten of queries ?  That would be the ones to optimize
>> for.
> 
> The front page, summary page of each project, and the RSS feed for each 
> project.
> 
>     -hpa

How about extending gitweb to check to see if there already exists a cached version of these pages, before recreating them?

e.g. structure the temp dir in such a way that each project has a place for cached pages. Then, before performing expensive operations, check to see if a file corresponding to the requested page already exists. If it does, simply return the contents of the file, otherwise go ahead and create the page dynamically, and return it to the user. Do not create cached pages in gitweb dynamically.

Then, in a post-update hook, for each of the expensive pages, invoke something like:

# delete the cached copy of the file, to force gitweb to recreate it
rm -f $git_temp/$project/rss
# get gitweb to recreate the page appropriately
# use a tmp file to prevent gitweb from getting confused
wget -O $git_temp/$project/rss.tmp \
   http://kernel.org/gitweb.cgi?p=$project;a=rss
# move the tmp file into place
mv $git_temp/$project/rss.tmp $git_temp/$project/rss

This way, we get the exact output returned from the usual gitweb invocation, but we can now cache the result, and only update it when there is a new commit that would affect the page output.

This would also not affect those who do not wish to use this mechanism. If the file does not exist, gitweb.cgi will simply revert to its usual behaviour.

Possible complications are the content-type headers, etc, but you could use the -s flag to wget, and store the server headers as well in the file, and get the necessary headers from the file as you stream it.

i.e. read the headers looking for ones that are "interesting" (Content-Type, charset, expires) until you get a blank line, print out the interesting headers using $cgi->header(), then just dump the remainder of the file to the caller via stdout.

Previous: Jakub NarebskiNext: Jakub Narebski
Message 8 of 80 in “Re: kernel.org mirroring (Re: [GIT PULL] MMC update)”
  1. Linus TorvaldsDec 7, 2006
  2. H. Peter AnvinDec 7, 2006
  3. Olivier GalibertDec 7, 2006
  4. H. Peter AnvinDec 7, 2006
  5. Olivier GalibertDec 7, 2006
  6. H. Peter AnvinDec 7, 2006
  7. Jakub NarebskiDec 8, 2006
  8. Rogan DawesDec 8, 2006
  9. Jakub NarebskiDec 8, 2006
  10. Rogan DawesDec 8, 2006
  11. Jonas FonsecaDec 8, 2006
  12. Martin LanghoffDec 9, 2006
  13. H. Peter AnvinDec 9, 2006
  14. Martin LanghoffDec 9, 2006
  15. H. Peter AnvinDec 9, 2006
  16. Martin LanghoffDec 9, 2006
  17. H. Peter AnvinDec 9, 2006
  18. H. Peter AnvinDec 8, 2006
  19. Linus TorvaldsDec 8, 2006
  20. H. Peter AnvinDec 8, 2006
  21. Lars HjemliDec 8, 2006
  22. H. Peter AnvinDec 8, 2006
  23. Lars HjemliDec 8, 2006
  24. H. Peter AnvinDec 8, 2006
  25. rdaDec 10, 2006
  26. Jeff GarzikDec 8, 2006
  27. H. Peter AnvinDec 8, 2006
  28. Jeff GarzikDec 8, 2006
  29. Linus TorvaldsDec 8, 2006
  30. Michael K. EdwardsDec 8, 2006
  31. H. Peter AnvinDec 8, 2006
  32. Michael K. EdwardsDec 9, 2006
  33. H. Peter AnvinDec 9, 2006
  34. Linus TorvaldsDec 9, 2006
  35. H. Peter AnvinDec 9, 2006
  36. Michael K. EdwardsDec 9, 2006
  37. Jeff GarzikDec 9, 2006
  38. Martin LanghoffDec 9, 2006
  39. Jakub NarebskiDec 9, 2006
  40. Jeff GarzikDec 9, 2006
  41. Jakub NarebskiDec 9, 2006
  42. Jeff GarzikDec 9, 2006
  43. Jakub NarebskiDec 9, 2006
  44. Jeff GarzikDec 9, 2006
  45. Martin LanghoffDec 10, 2006
  46. Jakub NarebskiDec 10, 2006
  47. Jeff GarzikDec 10, 2006
  48. Jakub NarebskiDec 10, 2006
  49. Jeff GarzikDec 10, 2006
  50. Jakub NarebskiDec 10, 2006
  51. Linus TorvaldsDec 10, 2006
  52. Jakub NarebskiDec 10, 2006
  53. Linus TorvaldsDec 10, 2006
  54. Martin LanghoffDec 10, 2006
  55. Jeff GarzikDec 10, 2006
  56. Jeff GarzikDec 10, 2006
  57. H. Peter AnvinDec 10, 2006
  58. Jeff GarzikDec 10, 2006
  59. Jakub NarebskiDec 10, 2006
  60. Martin LanghoffDec 11, 2006
  61. Jakub NarebskiDec 11, 2006
  62. Martin LanghoffDec 11, 2006
  63. Linus TorvaldsDec 9, 2006
  64. H. Peter AnvinDec 9, 2006
  65. Martin LanghoffDec 10, 2006
  66. H. Peter AnvinDec 10, 2006
  67. Jakub NarebskiDec 12, 2006
  68. Steven GrimmDec 9, 2006
  69. Linus TorvaldsDec 7, 2006
  70. Shawn PearceDec 7, 2006
  71. Linus TorvaldsDec 7, 2006
  72. Michael K. EdwardsDec 7, 2006
  73. H. Peter AnvinDec 7, 2006
  74. Junio C HamanoDec 7, 2006
  75. H. Peter AnvinDec 7, 2006
  76. Junio C HamanoDec 7, 2006
  77. Jakub NarebskiDec 8, 2006
  78. Linus TorvaldsDec 9, 2006
  79. H. Peter AnvinDec 9, 2006
  80. Jeff GarzikDec 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.