{"thread":{"id":"43081","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","startedAt":"2006-12-07T19:05:59Z","lastAt":"2006-12-12T21:19:28Z","messageCount":80,"participants":["Jeff Garzik","H. Peter Anvin","Linus Torvalds","Martin Langhoff","Jakub Narebski","Olivier Galibert","Jonas Fonseca","Shawn Pearce","rda","Michael K. Edwards","Junio C Hamano","Lars Hjemli","Steven Grimm","Rogan Dawes"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"294556","messageId":"Pine.LNX.4.64.0612071052560.3615@woody.osdl.org","threadId":"43081","inReplyTo":"45785697.1060001@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-07T19:05:59Z","receivedAt":"2006-12-07T19:05:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 7 Dec 2006, H. Peter Anvin wrote:\n> \n> That all being said, the lack of intrinsic caching in gitweb continues to be a\n> major problem for us.  Under high load, it makes all the problems worse.\n\nI really don't see what gitweb could do that would be somehow better than \napache doing the caching in front of it.. Is there some apache reason why \nthat isn't sufficient (ie limitations on its cache size or timeouts?)\n\nMaybe the cacheability hints from gitweb could be tweaked (a lot of it \nshould be \"infinitely cacheable\", but the stuff that depends on refs and \nthus can change, could be set to some fixed host-wide value - preferably \nsome that depends on how old the ref is).\n\nHaving gitweb be potentially up to an hour out of date is better than \ncausing mirroring problems due to excessive load.\n\nFor example, if the git \"refs/heads/\" (or tags) directory hasn't changed \nin the last two months, we should probably set any ref-relative gitweb \npages to have a caching timeout of a day or two. In contrast, if it's \nchanged in the last hour, maybe we should only cache it for five minutes.\n\nJakub: any way to make gitweb set the \"expires\" fields _much_ more \naggressively. I think we should at least have the ability to set a basic \nrules like\n\n - a _minimum_ of five minutes regardless of anything else\n\n   We might even tweak this based on loadaverage, and it might be \n   worthwhile to add a randomization, to make sure that you don't get into \n   situations where everything webpage needs to be recalculated at once.\n\n - if refs/ directories are old, raise the minimum by the age of the refs\n\n   If it's more than an hour old, raise it to ten minutes. If it's more \n   than a day, raise it to an hour. If it's more than a month old, raise \n   it to a day. And if it's more than half a year, it's some historical \n   archive like linux-history, and should probably default to a week or \n   more.\n\n - infinite for stuff that isn't ref-related.\n\nHmm?\n\n"},{"id":"295126","messageId":"457868AA.2030605@zytor.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612071052560.3615@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-07T19:16:58Z","receivedAt":"2006-12-07T19:16:58Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 7 Dec 2006, H. Peter Anvin wrote:\n>> That all being said, the lack of intrinsic caching in gitweb continues to be a\n>> major problem for us.  Under high load, it makes all the problems worse.\n> \n> I really don't see what gitweb could do that would be somehow better than \n> apache doing the caching in front of it.. Is there some apache reason why \n> that isn't sufficient (ie limitations on its cache size or timeouts?)\n> \n\nWhat it could do better is it could prevent multiple identical queries \nfrom being launched in parallel.  That's the real problem we see; under \nhigh load, Apache times out so the git query never gets into the cache; \nbut in the meantime, the common queries might easily have been launched \n20 times in parallel.  Unfortunately, the most common queries are also \nextremely expensive.\n\n"},{"id":"294350","messageId":"20061207193012.GA84678@dspnet.fr.eu.org","threadId":"43081","inReplyTo":"457868AA.2030605@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Olivier Galibert","fromEmail":"galibert@pobox.com","sentAt":"2006-12-07T19:30:12Z","receivedAt":"2006-12-07T19:30:12Z","isPatch":false,"sender":{"key":"galibert@pobox.com","avatar":null},"body":"On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:\n> Unfortunately, the most common queries are also extremely expensive.\n\nDo you have a top-ten of queries ?  That would be the ones to optimize\nfor.\n\n"},{"id":"294074","messageId":"Pine.LNX.4.64.0612071121410.3615@woody.osdl.org","threadId":"43081","inReplyTo":"457868AA.2030605@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-07T19:30:29Z","receivedAt":"2006-12-07T19:30:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 7 Dec 2006, H. Peter Anvin wrote:\n> \n> What it could do better is it could prevent multiple identical queries from\n> being launched in parallel.  That's the real problem we see; under high load,\n> Apache times out so the git query never gets into the cache; but in the\n> meantime, the common queries might easily have been launched 20 times in\n> parallel.  Unfortunately, the most common queries are also extremely\n> expensive.\n\nAhh. I'd have expected that apache itself had some serialization facility, \nthat would kind of go hand-in-hand with any caching.\n\nIt really would make more sense to have anything that does caching \nserialize the address that gets cached (think \"page cache\" layer in the \nkernel: the _cache_ is also the serialization point, and is what \nguarantees that we don't do stupid multiple reads to the same address).\n\nI'm surprised that Apache can't do that. Or maybe it can, and it just \nneeds some configuration entry? I don't know apache.. I realize that \nbecause Apache doesn't know before-hand whether something is cacheable or \nnot, it must probably _default_ to running the CGI scripts to the same \naddress in parallel, but it would be stupid to not have the option to \nserialize.\n\nThat said, from some of the other horrors I've heard about, \"stupid\" may \nbe just scratching at the surface.\n\n"},{"id":"295355","messageId":"20061207193903.GE12143@spearce.org","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612071121410.3615@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-07T19:39:03Z","receivedAt":"2006-12-07T19:39:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> wrote:\n> I'm surprised that Apache can't do that. Or maybe it can, and it just \n> needs some configuration entry? I don't know apache.. I realize that \n> because Apache doesn't know before-hand whether something is cacheable or \n> not, it must probably _default_ to running the CGI scripts to the same \n> address in parallel, but it would be stupid to not have the option to \n> serialize.\n\nAFAIK it doesn't have such an option, for basically the reason\nyou describe.  I worked on a project which had much more difficult\nto answer queries than gitweb and were also very popular.  Yes,\nthe system died under any load, no matter how much money was thrown\nat it.  :-)\n\n> That said, from some of the other horrors I've heard about, \"stupid\" may \n> be just scratching at the surface.\n\nIt is.  :-)\n\n-- \n"},{"id":"296282","messageId":"4578722E.9030402@zytor.com","threadId":"43081","inReplyTo":"20061207193012.GA84678@dspnet.fr.eu.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-07T19:57:34Z","receivedAt":"2006-12-07T19:57:34Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Olivier Galibert wrote:\n> On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:\n>> Unfortunately, the most common queries are also extremely expensive.\n> \n> Do you have a top-ten of queries ?  That would be the ones to optimize\n> for.\n\nThe front page, summary page of each project, and the RSS feed for each \nproject.\n\n"},{"id":"294795","messageId":"Pine.LNX.4.64.0612071152410.3615@woody.osdl.org","threadId":"43081","inReplyTo":"20061207193903.GE12143@spearce.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-07T19:58:23Z","receivedAt":"2006-12-07T19:58:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 7 Dec 2006, Shawn Pearce wrote:\n> \n> AFAIK it doesn't have such an option, for basically the reason\n> you describe.  I worked on a project which had much more difficult\n> to answer queries than gitweb and were also very popular.  Yes,\n> the system died under any load, no matter how much money was thrown\n> at it.  :-)\n> \n> > That said, from some of the other horrors I've heard about, \"stupid\" may \n> > be just scratching at the surface.\n> \n> It is.  :-)\n\nGaah. That's just stupid. This is such a _basic_ issue for caching (\"if \nconcurrent requests come in, only handle _one_ and give everybody the same \nresult\") that I claim that any cache that doesn't handle it isn't a cache \nat all, but a total disaster written by incompetent people.\n\nSure, you may want to disable it for certain kinds of truly dynamic \ncontent, but that doesn't mean you shouldn't be able to do it at all.\n\nDoes anybody who is web-server clueful know if there is some simple \nfront-end (squid?) that is easy to set up and can just act as a caching \nproxy in front of such an incompetent server?\n\nOr maybe there is some competent Apache module, not just the default \nmod_cache (which is what I assume kernel.org uses now)?\n\n"},{"id":"295812","messageId":"45787268.1030902@zytor.com","threadId":"43081","inReplyTo":"20061207193903.GE12143@spearce.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-07T19:58:32Z","receivedAt":"2006-12-07T19:58:32Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn Pearce wrote:\n> Linus Torvalds <torvalds@osdl.org> wrote:\n>> I'm surprised that Apache can't do that. Or maybe it can, and it just \n>> needs some configuration entry? I don't know apache.. I realize that \n>> because Apache doesn't know before-hand whether something is cacheable or \n>> not, it must probably _default_ to running the CGI scripts to the same \n>> address in parallel, but it would be stupid to not have the option to \n>> serialize.\n> \n> AFAIK it doesn't have such an option, for basically the reason\n> you describe.  I worked on a project which had much more difficult\n> to answer queries than gitweb and were also very popular.  Yes,\n> the system died under any load, no matter how much money was thrown\n> at it.  :-)\n\nYou certainly can be smarter about it when you know the nature of the \nquery, though.  I do that with the patch viewer scripts.\n\n"},{"id":"297616","messageId":"7vac1zsc8z.fsf@assigned-by-dhcp.cox.net","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612071121410.3615@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-07T20:05:16Z","receivedAt":"2006-12-07T20:05:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"If I understand correctly, kernel.org is still running the\nversion of gitweb Kay last installed there (I am too busy to\ntake over the gitweb installation maintenance at kernel.org, and\nI did not ask the $DOCUMENTROOT/git/ directory to be transferred\nto me when I rolled gitweb into the git.git repository).\n\nI do not know what queries are most popular, but I think a newer\ngitweb is more efficient in the summary page (getting list of\nbranches and tags).  It might be worth a try.\n"},{"id":"294347","messageId":"45787515.4010200@zytor.com","threadId":"43081","inReplyTo":"7vac1zsc8z.fsf@assigned-by-dhcp.cox.net","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-07T20:09:57Z","receivedAt":"2006-12-07T20:09:57Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Junio C Hamano wrote:\n> If I understand correctly, kernel.org is still running the\n> version of gitweb Kay last installed there (I am too busy to\n> take over the gitweb installation maintenance at kernel.org, and\n> I did not ask the $DOCUMENTROOT/git/ directory to be transferred\n> to me when I rolled gitweb into the git.git repository).\n\nThat's correct.  I can transfer that directory to you if you want; I \ncan't realistically track gitweb well enough to do this myself (in fact, \nit was pretty much a condition of having it up there that Kay would keep \nmaintaining it.)\n\n> I do not know what queries are most popular, but I think a newer\n> gitweb is more efficient in the summary page (getting list of\n> branches and tags).  It might be worth a try.\n\nHow do you want to handle it?\n\n\t-hpa\n\n"},{"id":"296857","messageId":"7vveknqruo.fsf@assigned-by-dhcp.cox.net","threadId":"43081","inReplyTo":"45787515.4010200@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-07T22:11:11Z","receivedAt":"2006-12-07T22:11:11Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> Junio C Hamano wrote:\n>> If I understand correctly, kernel.org is still running the\n>> version of gitweb Kay last installed there (I am too busy to\n>> take over the gitweb installation maintenance at kernel.org, and\n>> I did not ask the $DOCUMENTROOT/git/ directory to be transferred\n>> to me when I rolled gitweb into the git.git repository).\n>\n> That's correct.  I can transfer that directory to you if you want; I\n> can't realistically track gitweb well enough to do this myself...\n\nWell, the reason I haven't asked to is because I don't have\nenough time myself, so....\n"},{"id":"298011","messageId":"f2b55d220612071533l332d107t5852121eab53e5d8@mail.gmail.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612071152410.3615@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Michael K. Edwards","fromEmail":"medwards.linux@gmail.com","sentAt":"2006-12-07T23:33:48Z","receivedAt":"2006-12-07T23:33:48Z","isPatch":false,"sender":{"key":"medwards.linux@gmail.com","avatar":null},"body":"On 12/7/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Does anybody who is web-server clueful know if there is some simple\n> front-end (squid?) that is easy to set up and can just act as a caching\n> proxy in front of such an incompetent server?\n\nSquid in \"transparent reverse proxy\" mode isn't a bad choice, although\nI don't know offhand whether it queues/clusters concurrent requests\nfor the same URL in the way you want.  I suggest the \"transparent\"\ndeployment (netfilter/netlink integration) because you can slap it in\nwith no changes to the origin server and yank it out again if you have\na problem.  The challenge is in getting conntrack to scale to a\nzillion concurrent sessions, but you could probably find someone in\nyour crowd who knows something about that.  :-)\n\nIgnore any documentation that talks about httpd_accel_*.  Configuring\ntransparent mode is a great deal simpler and saner in squid 2.6 than\nit used to be; you just add a \"transparent\" parameter to the http_port\ntag.  With or without this tag, you set up what used to be called\n\"accelerator mode\" using some parameters to http_port and cache_peer,\nas described in\nhttp://www.squid-cache.org/mail-archive/squid-users/200607/0162.html.\n\nIf transparent mode looks like the right thing for kernel.org, you\nmight be interested in some netfilter hackery to offload part of the\nconntrack session lookup load to a front-end box that blocks DDoS and\nacts more or less as an L4 switch plus session context cache.  I've\nbeen banging on a proof of concept implementation for a while, and am\ncurrently working on integrating against 2.6.19 by splitting\nnf_conntrack into front and back halves that interact via a sort of\nLayer 2+ header.  I have no idea yet whether it will have any\nscalability benefit on dual-x86_64 class hardware (it was originally\nconceived for rigid cache architectures where the random access\npatterns of session lookups have drastic cache effects).\n\nCheers,\n"},{"id":"295761","messageId":"20061207235039.GA423@dspnet.fr.eu.org","threadId":"43081","inReplyTo":"4578722E.9030402@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Olivier Galibert","fromEmail":"galibert@pobox.com","sentAt":"2006-12-07T23:50:39Z","receivedAt":"2006-12-07T23:50:39Z","isPatch":false,"sender":{"key":"galibert@pobox.com","avatar":null},"body":"On Thu, Dec 07, 2006 at 11:57:34AM -0800, H. Peter Anvin wrote:\n> Olivier Galibert wrote:\n> >On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:\n> >>Unfortunately, the most common queries are also extremely expensive.\n> >\n> >Do you have a top-ten of queries ?  That would be the ones to optimize\n> >for.\n> \n> The front page, summary page of each project, and the RSS feed for each \n> project.\n\nHmmm, maybe you could have the summaries and rss feed generated on\npush, which could also generate elementary files with lines of the\nfront page.  That would make these top offenders static page serving.\n\n  OG.\n"},{"id":"297354","messageId":"4578AA3B.7040401@zytor.com","threadId":"43081","inReplyTo":"20061207235039.GA423@dspnet.fr.eu.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-07T23:56:43Z","receivedAt":"2006-12-07T23:56:43Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Olivier Galibert wrote:\n> On Thu, Dec 07, 2006 at 11:57:34AM -0800, H. Peter Anvin wrote:\n>> Olivier Galibert wrote:\n>>> On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:\n>>>> Unfortunately, the most common queries are also extremely expensive.\n>>> Do you have a top-ten of queries ?  That would be the ones to optimize\n>>> for.\n>> The front page, summary page of each project, and the RSS feed for each \n>> project.\n> \n> Hmmm, maybe you could have the summaries and rss feed generated on\n> push, which could also generate elementary files with lines of the\n> front page.  That would make these top offenders static page serving.\n> \n\nThere are a lot of things which \"could be done\" given the proper cache \ninfrastructure and gitweb support.\n\n"},{"id":"294395","messageId":"200612081043.07192.jnareb@gmail.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612071052560.3615@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-08T09:43:06Z","receivedAt":"2006-12-08T09:43:06Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n[...] \n> For example, if the git \"refs/heads/\" (or tags) directory hasn't changed \n> in the last two months, we should probably set any ref-relative gitweb \n> pages to have a caching timeout of a day or two. In contrast, if it's \n> changed in the last hour, maybe we should only cache it for five minutes.\n> \n> Jakub: any way to make gitweb set the \"expires\" fields _much_ more \n> aggressively. I think we should at least have the ability to set a basic \n> rules like\n> \n>  - a _minimum_ of five minutes regardless of anything else\n> \n>    We might even tweak this based on loadaverage, and it might be \n>    worthwhile to add a randomization, to make sure that you don't get into \n>    situations where everything webpage needs to be recalculated at once.\n\nI think the minimum expires (or minimum _additional_ expires: as of now\ngiweb only does expires +1d for explicit hash requests) should depend on\nhow often project changes. How often there are pushes to kernel.org?\n \n>  - if refs/ directories are old, raise the minimum by the age of the refs\n> \n>    If it's more than an hour old, raise it to ten minutes. If it's more \n>    than a day, raise it to an hour. If it's more than a month old, raise \n>    it to a day. And if it's more than half a year, it's some historical \n>    archive like linux-history, and should probably default to a week or \n>    more.\n\nWhat about packed refs?\n\nWe can certainly raise expires for tags (tags objects), as they should not\nusually change.\n \n>  - infinite for stuff that isn't ref-related.\n\nAs sha1 is not changeable, everything that is accessed by explicit \nsha1 (hash), or by explicit sha1 (hash_base) plus pathname (file_name)\nshould have effectively infinite expires.\n\n\nEvery caching would need some temporary memory, or temporary disk space.\nAnd perhaps mod_perl specific caching would be useful here...\n\nP.S. I have added Pasky to Cc:, as he manages http://repo.or.cz public\ngit repository hosting (much smaller than kernel.org and I think under less\nload: but also I think withour kernel.org resources).\n-- \nJakub Narebski\n"},{"id":"296740","messageId":"elbhu9$ke2$1@sea.gmane.org","threadId":"43081","inReplyTo":"20061207235039.GA423@dspnet.fr.eu.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-08T11:25:05Z","receivedAt":"2006-12-08T11:25:05Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Olivier Galibert wrote:\n\n> On Thu, Dec 07, 2006 at 11:57:34AM -0800, H. Peter Anvin wrote:\n>> Olivier Galibert wrote:\n>>>On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:\n>>>>\n>>>>Unfortunately, the most common queries are also extremely expensive.\n>>>\n>>>Do you have a top-ten of queries ?  That would be the ones to optimize\n>>>for.\n>> \n>> The front page, summary page of each project, and the RSS feed for each \n>> project.\n> \n> Hmmm, maybe you could have the summaries and rss feed generated on\n> push, which could also generate elementary files with lines of the\n> front page.  That would make these top offenders static page serving.\n\nThe \"extremely aggresive caching solution\" could be as follows: cache\neverything, invalidate (remove) on push caches of variable variety related\nto push (list of projects and OPML on any push; summary page and every\npage without h=<hash> or hb=<hash>;f=<filename> for a given project).\n\nThe most important problem is that kernel.org uses old gitweb, the last\nversion before incorporating gitweb into git (and also reducing\nsignificantly the time needed for summary, heads and tags pages).\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"298267","messageId":"4579611F.5010303@dawes.za.net","threadId":"43081","inReplyTo":"4578722E.9030402@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Rogan Dawes","fromEmail":"discard@dawes.za.net","sentAt":"2006-12-08T12:57:03Z","receivedAt":"2006-12-08T12:57:03Z","isPatch":false,"sender":{"key":"discard@dawes.za.net","avatar":null},"body":"H. Peter Anvin wrote:\n> Olivier Galibert wrote:\n>> On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:\n>>> Unfortunately, the most common queries are also extremely expensive.\n>>\n>> Do you have a top-ten of queries ?  That would be the ones to optimize\n>> for.\n> \n> The front page, summary page of each project, and the RSS feed for each \n> project.\n> \n>     -hpa\n\nHow about extending gitweb to check to see if there already exists a \ncached version of these pages, before recreating them?\n\ne.g. structure the temp dir in such a way that each project has a place \nfor cached pages. Then, before performing expensive operations, check to \nsee if a file corresponding to the requested page already exists. If it \ndoes, simply return the contents of the file, otherwise go ahead and \ncreate the page dynamically, and return it to the user. Do not create \ncached pages in gitweb dynamically.\n\nThen, in a post-update hook, for each of the expensive pages, invoke \nsomething like:\n\n# delete the cached copy of the file, to force gitweb to recreate it\nrm -f $git_temp/$project/rss\n# get gitweb to recreate the page appropriately\n# use a tmp file to prevent gitweb from getting confused\nwget -O $git_temp/$project/rss.tmp \\\n   http://kernel.org/gitweb.cgi?p=$project;a=rss\n# move the tmp file into place\nmv $git_temp/$project/rss.tmp $git_temp/$project/rss\n\nThis way, we get the exact output returned from the usual gitweb \ninvocation, but we can now cache the result, and only update it when \nthere is a new commit that would affect the page output.\n\nThis would also not affect those who do not wish to use this mechanism. \nIf the file does not exist, gitweb.cgi will simply revert to its usual \nbehaviour.\n\nPossible complications are the content-type headers, etc, but you could \nuse the -s flag to wget, and store the server headers as well in the \nfile, and get the necessary headers from the file as you stream it.\n\ni.e. read the headers looking for ones that are \"interesting\" \n(Content-Type, charset, expires) until you get a blank line, print out \nthe interesting headers using $cgi->header(), then just dump the \nremainder of the file to the caller via stdout.\n\n"},{"id":"295137","messageId":"200612081438.25493.jnareb@gmail.com","threadId":"43081","inReplyTo":"4579611F.5010303@dawes.za.net","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-08T13:38:24Z","receivedAt":"2006-12-08T13:38:24Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dnia piątek 8. grudnia 2006 13:57, Rogan Dawes napisał:\n> H. Peter Anvin wrote:\n>> Olivier Galibert wrote:\n>>> On Thu, Dec 07, 2006 at 11:16:58AM -0800, H. Peter Anvin wrote:\n>>>> Unfortunately, the most common queries are also extremely expensive.\n\nWith newer gitweb, which tries to do the same using less git commands,\nsome of queries (summary, heads, tags pages) should be less expensive.\n\n>>> Do you have a top-ten of queries ?  That would be the ones to optimize\n>>> for.\n>> \n>> The front page, summary page of each project, and the RSS feed for each \n>> project.\n> \n> How about extending gitweb to check to see if there already exists a \n> cached version of these pages, before recreating them?\n> \n> e.g. structure the temp dir in such a way that each project has a place \n> for cached pages. Then, before performing expensive operations, check to \n> see if a file corresponding to the requested page already exists. If it \n> does, simply return the contents of the file, otherwise go ahead and \n> create the page dynamically, and return it to the user. Do not create \n> cached pages in gitweb dynamically.\n\nThis would add the need for directory for temporary files... well,\nit would be optional now...\n\n> Then, in a post-update hook, for each of the expensive pages, invoke \n> something like:\n> \n> # delete the cached copy of the file, to force gitweb to recreate it\n> rm -f $git_temp/$project/rss\n> # get gitweb to recreate the page appropriately\n> # use a tmp file to prevent gitweb from getting confused\n> wget -O $git_temp/$project/rss.tmp \\\n>    http://kernel.org/gitweb.cgi?p=$project;a=rss\n> # move the tmp file into place\n> mv $git_temp/$project/rss.tmp $git_temp/$project/rss\n\nGood idea... although there are some page views which shouldn't change\nat all... well, with the possible exception of changes in gitweb output,\nand even then there are some (blob_plain and snapshot views) which\ndoesn't change at all.\n\nIt would be good to avoid removing them on push, and only remove\nthem using some tmpwatch-like removal.\n \n> This way, we get the exact output returned from the usual gitweb \n> invocation, but we can now cache the result, and only update it when \n> there is a new commit that would affect the page output.\n> \n> This would also not affect those who do not wish to use this mechanism. \n> If the file does not exist, gitweb.cgi will simply revert to its usual \n> behaviour.\n\nGood idea. Perhaps I should add it to gitweb TODO file.\n\nHmmm... perhaps it is time for next \"[RFC] gitweb wishlist and TODO list\"\nthread?\n \n> Possible complications are the content-type headers, etc, but you could \n> use the -s flag to wget, and store the server headers as well in the \n> file, and get the necessary headers from the file as you stream it.\n> \n> i.e. read the headers looking for ones that are \"interesting\" \n> (Content-Type, charset, expires) until you get a blank line, print out \n> the interesting headers using $cgi->header(), then just dump the \n> remainder of the file to the caller via stdout.\n\nNo need for that. $cgi->header() is to _generate_ the headers, so if\na file is saved with headers, we can just dump it to STDOUT; the possible\nexception is a need to rewrite 'expires' header, if it is used.\n\nPerhaps gitweb should generate it's own ETag instead of messing with\n'expires' header?\n-- \nJakub Narebski\n"},{"id":"297931","messageId":"4579775C.2010608@dawes.za.net","threadId":"43081","inReplyTo":"200612081438.25493.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Rogan Dawes","fromEmail":"discard@dawes.za.net","sentAt":"2006-12-08T14:31:56Z","receivedAt":"2006-12-08T14:31:56Z","isPatch":false,"sender":{"key":"discard@dawes.za.net","avatar":null},"body":"Jakub Narebski wrote:\n> Dnia piątek 8. grudnia 2006 13:57, Rogan Dawes napisał:\n\n>> How about extending gitweb to check to see if there already exists a \n>> cached version of these pages, before recreating them?\n>>\n>> e.g. structure the temp dir in such a way that each project has a place \n>> for cached pages. Then, before performing expensive operations, check to \n>> see if a file corresponding to the requested page already exists. If it \n>> does, simply return the contents of the file, otherwise go ahead and \n>> create the page dynamically, and return it to the user. Do not create \n>> cached pages in gitweb dynamically.\n> \n> This would add the need for directory for temporary files... well,\n> it would be optional now...\n> \nIt would still be optional. If the \"cache\" directory structure exists, \nthen use it, otherwise, continue as usual. All it would cost is a stat() \nor two, I guess.\n\n>> Then, in a post-update hook, for each of the expensive pages, invoke \n>> something like:\n>>\n>> # delete the cached copy of the file, to force gitweb to recreate it\n>> rm -f $git_temp/$project/rss\n>> # get gitweb to recreate the page appropriately\n>> # use a tmp file to prevent gitweb from getting confused\n>> wget -O $git_temp/$project/rss.tmp \\\n>>    http://kernel.org/gitweb.cgi?p=$project;a=rss\n>> # move the tmp file into place\n>> mv $git_temp/$project/rss.tmp $git_temp/$project/rss\n> \n> Good idea... although there are some page views which shouldn't change\n> at all... well, with the possible exception of changes in gitweb output,\n> and even then there are some (blob_plain and snapshot views) which\n> doesn't change at all.\n> \n> It would be good to avoid removing them on push, and only remove\n> them using some tmpwatch-like removal.\n\nWell, my theory was that we would only cache pages that change when new \ndata enters the repo. So, using the push as the trigger is almost \nguaranteed to be the right thing to do. New data indicates new rss \nitems, indicates an updated shortlog page, etc.\n\nNOTE: This caching could be problematic for the \"changed 2 hours ago\" \nnotation for various branches/files, etc. But however we implement the \ncaching, we'd have this problem.\n\n>> This way, we get the exact output returned from the usual gitweb \n>> invocation, but we can now cache the result, and only update it when \n>> there is a new commit that would affect the page output.\n>>\n>> This would also not affect those who do not wish to use this mechanism. \n>> If the file does not exist, gitweb.cgi will simply revert to its usual \n>> behaviour.\n> \n> Good idea. Perhaps I should add it to gitweb TODO file.\n> \n> Hmmm... perhaps it is time for next \"[RFC] gitweb wishlist and TODO list\"\n> thread?\n>  \n>> Possible complications are the content-type headers, etc, but you could \n>> use the -s flag to wget, and store the server headers as well in the \n>> file, and get the necessary headers from the file as you stream it.\n>>\n>> i.e. read the headers looking for ones that are \"interesting\" \n>> (Content-Type, charset, expires) until you get a blank line, print out \n>> the interesting headers using $cgi->header(), then just dump the \n>> remainder of the file to the caller via stdout.\n> \n> No need for that. $cgi->header() is to _generate_ the headers, so if\n> a file is saved with headers, we can just dump it to STDOUT; the possible\n> exception is a need to rewrite 'expires' header, if it is used.\n\nGood point. I guess one thing that will be incorrect in the headers is \nthe server date, but I doubt that anyone cares much. As you say, though, \nthis might relate to the expiry of cached content in upstream caches.\n\n> \n> Perhaps gitweb should generate it's own ETag instead of messing with\n> 'expires' header?\n\nWell, we can possibly eliminate the expires header entirely for dynamic \npages, and check the If-Modified-Since value against the timestamp of \nthe cached file, or the server date in the cached file, and return \"304 \nNot Modified\" responses. That would also help to reduce the load on the \nserver, by only returning the headers, and not the entire response.\n\nThe downside is that it would prevent upstream proxies from caching this \ndata for us.\n\nRegards,\n\n"},{"id":"294745","messageId":"2c6b72b30612080738wbad5938r3dae807729e484c0@mail.gmail.com","threadId":"43081","inReplyTo":"4579775C.2010608@dawes.za.net","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jonas Fonseca","fromEmail":"jonas.fonseca@gmail.com","sentAt":"2006-12-08T15:38:59Z","receivedAt":"2006-12-08T15:38:59Z","isPatch":false,"sender":{"key":"jonas.fonseca@gmail.com","avatar":"https://gravatar.com/avatar/9b7fa23cce50269e5d164312b6ac5ae818a180f837b38f28f3bdf689dd7f96cd?d=mp&s=160"},"body":"On 12/8/06, Rogan Dawes <discard@dawes.za.net> wrote:\n> NOTE: This caching could be problematic for the \"changed 2 hours ago\"\n> notation for various branches/files, etc. But however we implement the\n> caching, we'd have this problem.\n\nIt could be solved using ECMAScript (if that is an option): Include an exact\ntime stamp or something that browsers not supporting ECMAScript can\nshow and others browsers can change the time stamp to make it relative\nand do the coloring/highlighting of recent activity. This could also slightly\nspeed up the script and it might be better to provide an exact time stamp\nby default if aggressive caching is applied.\n\n-- \n"},{"id":"296464","messageId":"45798FE2.9040502@zytor.com","threadId":"43081","inReplyTo":"4579611F.5010303@dawes.za.net","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-08T16:16:34Z","receivedAt":"2006-12-08T16:16:34Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Rogan Dawes wrote:\n> \n> How about extending gitweb to check to see if there already exists a \n> cached version of these pages, before recreating them?\n> \n\nThis goes back to the \"gitweb needs native caching\" again.\n\n"},{"id":"296964","messageId":"Pine.LNX.4.64.0612080830380.3516@woody.osdl.org","threadId":"43081","inReplyTo":"45798FE2.9040502@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-08T16:35:22Z","receivedAt":"2006-12-08T16:35:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 8 Dec 2006, H. Peter Anvin wrote:\n> \n> This goes back to the \"gitweb needs native caching\" again.\n\nIt should be fairly easy to add a caching layer, but I wouldn't do it \ninside gitweb itself - it gets too mixed up. It would be better to have \nit as a separate front-end, that just calls gitweb for anything it doesn't \nfind in the cache.\n\nI could write a simple C caching thing that just hashes the CGI arguments \nand uses a hash to create a cache (and proper lock-files etc to serialize \naccess to a particular cache object while it's being created) fairly \neasily, but I'm pretty sure people would much prefer a mod_perl thing just \nto avoid the fork/exec overhead with Apache (I think mod_perl allows \nApache to run perl scripts without it), and that means I'm not the right \nperson any more.\n\nNot that I'm the right person anyway, since I don't have a web server set \nup on my machine to even test with ;)\n\t\n\t\tLinus\n"},{"id":"293870","messageId":"457995F8.1080405@zytor.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612080830380.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-08T16:42:32Z","receivedAt":"2006-12-08T16:42:32Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 8 Dec 2006, H. Peter Anvin wrote:\n>> This goes back to the \"gitweb needs native caching\" again.\n> \n> It should be fairly easy to add a caching layer, but I wouldn't do it \n> inside gitweb itself - it gets too mixed up. It would be better to have \n> it as a separate front-end, that just calls gitweb for anything it doesn't \n> find in the cache.\n> \n\nIf you want to do side effect generation of cache contents, it might not \nbe possible to do it that way.  At the very least gitweb needs to be \naware of how to explicitly enter things into the cache.\n\nAll of this isn't really all that hard; I have implemented all that \nstuff for diffview, for example (when generating a single diff hunk, you \n  naturally end up producing all of them, so you want to have them \npreemptively cached.)\n\n> I could write a simple C caching thing that just hashes the CGI arguments \n> and uses a hash to create a cache (and proper lock-files etc to serialize \n> access to a particular cache object while it's being created) fairly \n> easily, but I'm pretty sure people would much prefer a mod_perl thing just \n> to avoid the fork/exec overhead with Apache (I think mod_perl allows \n> Apache to run perl scripts without it), and that means I'm not the right \n> person any more.\n\nTrue about mod_perl.  Haven't messed with that myself, either. \nfork/exec really is very cheap on Linux, so it's not a huge deal.\n\n> Not that I'm the right person anyway, since I don't have a web server set \n> up on my machine to even test with ;)\n\nHeh :)\n\n\t-hpa\n"},{"id":"297250","messageId":"457998C8.3050601@garzik.org","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612080830380.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-08T16:54:32Z","receivedAt":"2006-12-08T16:54:32Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Linus Torvalds wrote:\n> I could write a simple C caching thing that just hashes the CGI arguments \n> and uses a hash to create a cache (and proper lock-files etc to serialize \n> access to a particular cache object while it's being created) fairly \n> easily, but I'm pretty sure people would much prefer a mod_perl thing just \n> to avoid the fork/exec overhead with Apache (I think mod_perl allows \n> Apache to run perl scripts without it), and that means I'm not the right \n> person any more.\n> \n> Not that I'm the right person anyway, since I don't have a web server set \n> up on my machine to even test with ;)\n> \t\n> \t\tLinus\n> \n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n\nThis is quite nice and easy, if memory-only caching works for the \nsituation:  http://www.danga.com/memcached/\n\nThere are APIs for C, Perl, and plenty of other languages.\n\n\tJeff\n\n"},{"id":"296048","messageId":"45799B02.3010102@zytor.com","threadId":"43081","inReplyTo":"457998C8.3050601@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-08T17:04:02Z","receivedAt":"2006-12-08T17:04:02Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Jeff Garzik wrote:\n> \n> This is quite nice and easy, if memory-only caching works for the \n> situation:  http://www.danga.com/memcached/\n> \n> There are APIs for C, Perl, and plenty of other languages.\n> \n\nMemory-only caching is kind of nasty.  Memory is a premium resource on \nkernel.org.\n\n"},{"id":"295210","messageId":"4579A390.1080202@garzik.org","threadId":"43081","inReplyTo":"45799B02.3010102@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-08T17:40:32Z","receivedAt":"2006-12-08T17:40:32Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"H. Peter Anvin wrote:\n> Jeff Garzik wrote:\n>>\n>> This is quite nice and easy, if memory-only caching works for the \n>> situation:  http://www.danga.com/memcached/\n>>\n>> There are APIs for C, Perl, and plenty of other languages.\n>>\n> \n> Memory-only caching is kind of nasty.  Memory is a premium resource on \n> kernel.org.\n\nhmmm.  Well, I have been wondering why nobody ever came up with a \nsystem-wide local (==disk) cache for remote and/or calculated objects. \nMaybe its time to do something about that.\n\nI've been in a daemon-writing mood lately.\n\n\tJeff\n\n\n"},{"id":"298502","messageId":"8c5c35580612081149y2042c62j55b1b7b3f23da6ad@mail.gmail.com","threadId":"43081","inReplyTo":"457995F8.1080405@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Lars Hjemli","fromEmail":"lh@elementstorage.no","sentAt":"2006-12-08T19:49:45Z","receivedAt":"2006-12-08T19:49:45Z","isPatch":false,"sender":{"key":"lh@elementstorage.no","avatar":null},"body":"On 12/8/06, H. Peter Anvin <hpa@zytor.com> wrote:\n> Linus Torvalds wrote:\n> > I could write a simple C caching thing that just hashes the CGI arguments\n> > and uses a hash to create a cache (and proper lock-files etc to serialize\n> > access to a particular cache object while it's being created) fairly\n> > easily, but I'm pretty sure people would much prefer a mod_perl thing just\n> > to avoid the fork/exec overhead with Apache (I think mod_perl allows\n> > Apache to run perl scripts without it), and that means I'm not the right\n> > person any more.\n>\n> True about mod_perl.  Haven't messed with that myself, either.\n> fork/exec really is very cheap on Linux, so it's not a huge deal.\n\nI've been playing around with a \"native git\" cgi thingy the last week\n(I call it cgit),  and I've been thinking about adding exactly this\nkind of caching to it. And since it's basically a standard git command\nwritten in C, it should have less overhead than any perl\nimplementation.\n\nIt's far from ready yet, but I'll try to publish some code this\nweekend just in case someone finds it interesting.\n\n-- \n"},{"id":"298016","messageId":"4579C234.5080709@zytor.com","threadId":"43081","inReplyTo":"8c5c35580612081149y2042c62j55b1b7b3f23da6ad@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-08T19:51:16Z","receivedAt":"2006-12-08T19:51:16Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Lars Hjemli wrote:\n> On 12/8/06, H. Peter Anvin <hpa@zytor.com> wrote:\n>> Linus Torvalds wrote:\n>> > I could write a simple C caching thing that just hashes the CGI \n>> arguments\n>> > and uses a hash to create a cache (and proper lock-files etc to \n>> serialize\n>> > access to a particular cache object while it's being created) fairly\n>> > easily, but I'm pretty sure people would much prefer a mod_perl \n>> thing just\n>> > to avoid the fork/exec overhead with Apache (I think mod_perl allows\n>> > Apache to run perl scripts without it), and that means I'm not the \n>> right\n>> > person any more.\n>>\n>> True about mod_perl.  Haven't messed with that myself, either.\n>> fork/exec really is very cheap on Linux, so it's not a huge deal.\n> \n> I've been playing around with a \"native git\" cgi thingy the last week\n> (I call it cgit),  and I've been thinking about adding exactly this\n> kind of caching to it. And since it's basically a standard git command\n> written in C, it should have less overhead than any perl\n> implementation.\n> \n\nTrust me, perl, or CGI, is not the problem.  It's all about I/O traffic \ngenerated by git.\n\n"},{"id":"297442","messageId":"8c5c35580612081159r325bfe1dsa764da050d63606b@mail.gmail.com","threadId":"43081","inReplyTo":"4579C234.5080709@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Lars Hjemli","fromEmail":"lh@elementstorage.no","sentAt":"2006-12-08T19:59:47Z","receivedAt":"2006-12-08T19:59:47Z","isPatch":false,"sender":{"key":"lh@elementstorage.no","avatar":null},"body":"On 12/8/06, H. Peter Anvin <hpa@zytor.com> wrote:\n> Trust me, perl, or CGI, is not the problem.  It's all about I/O traffic\n> generated by git.\n\nYes, I understand. That's why I've been thinking about internal\ncaching of pages.\n\nIt's just a kick doing it in C, playing around with the git internals :-)\n\n-- \n"},{"id":"297832","messageId":"4579C4D5.8070002@zytor.com","threadId":"43081","inReplyTo":"8c5c35580612081159r325bfe1dsa764da050d63606b@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-08T20:02:29Z","receivedAt":"2006-12-08T20:02:29Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Lars Hjemli wrote:\n> On 12/8/06, H. Peter Anvin <hpa@zytor.com> wrote:\n>> Trust me, perl, or CGI, is not the problem.  It's all about I/O traffic\n>> generated by git.\n> \n> Yes, I understand. That's why I've been thinking about internal\n> caching of pages.\n\nCaching, preferrably with smarts, is the key.\n\n> It's just a kick doing it in C, playing around with the git internals :-)\n\nThat's fine, but it does make it harder to maintain.\n\n\t-hpa\n\n"},{"id":"294526","messageId":"Pine.LNX.4.64.0612081453430.3516@woody.osdl.org","threadId":"43081","inReplyTo":"457998C8.3050601@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-08T23:27:42Z","receivedAt":"2006-12-08T23:27:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 8 Dec 2006, Jeff Garzik wrote:\n> \n> This is quite nice and easy, if memory-only caching works for the situation:\n> http://www.danga.com/memcached/\n> \n> There are APIs for C, Perl, and plenty of other languages.\n\nActually, just looking at the examples, it looks like memcached is \nfundamentally flawed, exactly the same way Apache mod_cache is \nfundamentally flawed.\n\nExactly like mod_perl, it appears that if something isn't cached, the \nmemcached server will just return \"not cached\" to everybody, and all the \nclients will, like a stampeding herd, all do the uncached access. Even if \nthey have the exact same query. And you're back to square one: your server \nload went through the roof.\n\nYou can't have a cache architecture where the client just does a \"get\", \nlike memcached does. You need to have a \"read-for-fill\" operation, which \nsays:\n\n - get this cache entry\n\n - if this cache entry does not exist, get an exclusive lock\n\n - if you get that exclusive lock, return NULL, and the client promises \n   that it will fill it (inside the kernel, see for example \n   \"find_get_page()\" vs \"grab_cache_page()\" - the latter will return a \n   locked page whether it exists or not, and if it didn't exist, it will \n   have inserted it into the cache datastructures so that you don't have \n   multiple concurrent readers trying to all create different pages)\n\n - if you block on the exclusive lock, that means that some other client \n   is busy fulfilling it. When you unblock, do a regular \"read\" operation \n   (not a \"repeat\": we only block once, and if that fails, that's it).\n\n - any cachefill operation will release the lock (and allow pending \n   cache queries to succeed)\n\n - the locking client going away will release the lock (and allow pending \n   cache queries to fail, and hopefully cause a \"set cache\" operation)\n\n - a timeout (settable by some method) will also force-release a lock in \n   the case of buggy clients that do \"read-for-modify\" but never do the \n   \"modify\".\n\nThe \"timeout\" thing is to handle the case of buggy clients that crash \nafter trying to get - it will slow down things _enormously_ if that \nhappens, but hey, it's a buggy client. And it will still continue to work.\n\nLooking at the memcached operations, they have the \"read\" op (aka \"get\"), \nbut they seem to have no \"read-for-fill\" op. So memcached fundamentally \ndoesn't fix this problem, at least without explicit serialization by the \nclient.\n\n(The serialization could be done by the client, but that would serialize \n_everything_, and mean that a uncached lookup will hold up all the cached \nones too - which is why you do NOT want to serialize in the caller: you \nreally want to serialize in the layer that does the caching).\n\nIt's fairly easy to do the lock. You could just hash the lookup key using \nsome reasonable hash. It doesn't even have to be a _big_ hash: it's ok to \nhave just a few bits for lock hashing, since it's only going to be for \nmisses.\n\nSo hashing to eight bits and using 256 locks is probably fine, as long as \nthis is done by the cache server. That means that the cache server only \never needs to track that many timeouts, for example (it also indirectly \nsets a limit on the number of possible \"outstanding uncached requests\", \nwhich is _exactly_ what you want - but hash collissions will also \npotentially unlock the _wrong_ bucket, so if you have too many of them, it \ncan make the \"only one outstanding unhashed request per key\" not be as \neffective).\n\nSo assuming you get good cache hit statistics, the locking shouldn't be a \nbig issue. But you definitely want to do it, because the whole point of \ncaching was to not do the same op multiple times. \n\nI still don't understand why apache doesn't do it. I guess it wants to be \nstateless or something.\n\n"},{"id":"296246","messageId":"f2b55d220612081546u1ffa98e5q75be55d31da82a2f@mail.gmail.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081453430.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Michael K. Edwards","fromEmail":"medwards.linux@gmail.com","sentAt":"2006-12-08T23:46:44Z","receivedAt":"2006-12-08T23:46:44Z","isPatch":false,"sender":{"key":"medwards.linux@gmail.com","avatar":null},"body":"On 12/8/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> You can't have a cache architecture where the client just does a \"get\",\n> like memcached does. You need to have a \"read-for-fill\" operation ...\n\nIn Squid 2.6:\n    collapsed_forwarding on\n    refresh_stale_window <seconds>\n(apply the latter only to stanzas where you want \"readahead\" of\nabout-to-expire cache entries)\n\nBrief design description at http://devel.squid-cache.org/collapsed_forwarding/.\n\n(I didn't write this code, everything I know about squid leaked\nthrough the Google-shaped pinhole in my tinfoil hat, etc.  But if you\ngo this way I'd like to be in the loop to understand the scalability\nissues around netfilter-assisted transparent proxying.)\n\nCheers,\n"},{"id":"294544","messageId":"4579F9FF.7050701@zytor.com","threadId":"43081","inReplyTo":"f2b55d220612081546u1ffa98e5q75be55d31da82a2f@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-08T23:49:19Z","receivedAt":"2006-12-08T23:49:19Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Michael K. Edwards wrote:\n> On 12/8/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>> You can't have a cache architecture where the client just does a \"get\",\n>> like memcached does. You need to have a \"read-for-fill\" operation ...\n> \n> In Squid 2.6:\n>    collapsed_forwarding on\n>    refresh_stale_window <seconds>\n> (apply the latter only to stanzas where you want \"readahead\" of\n> about-to-expire cache entries)\n> \n> Brief design description at \n> http://devel.squid-cache.org/collapsed_forwarding/.\n> \n> (I didn't write this code, everything I know about squid leaked\n> through the Google-shaped pinhole in my tinfoil hat, etc.  But if you\n> go this way I'd like to be in the loop to understand the scalability\n> issues around netfilter-assisted transparent proxying.)\n> \n\nThere is another thing that probably will be required, and I'm not sure \nif something in front of Apache (like Squid) rather than behind it can \neasily deal with: on timeout, the process needs to continue in order to \nfeed the cache.  Otherwise, you're still in a failure scenario as soon \nas timeout happens.\n\n"},{"id":"298403","messageId":"f2b55d220612081618v26f7c714l7d4ea5315a964aa@mail.gmail.com","threadId":"43081","inReplyTo":"4579F9FF.7050701@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Michael K. Edwards","fromEmail":"medwards.linux@gmail.com","sentAt":"2006-12-09T00:18:13Z","receivedAt":"2006-12-09T00:18:13Z","isPatch":false,"sender":{"key":"medwards.linux@gmail.com","avatar":null},"body":"On 12/8/06, H. Peter Anvin <hpa@zytor.com> wrote:\n> There is another thing that probably will be required, and I'm not sure\n> if something in front of Apache (like Squid) rather than behind it can\n> easily deal with: on timeout, the process needs to continue in order to\n> feed the cache.  Otherwise, you're still in a failure scenario as soon\n> as timeout happens.\n\nI would think this would be a great deal easier to handle in an\narm's-length \"accelerator\" than in the origin server.  Only restart\nthe hit to the origin server if you think that something has actually\ngone wrong there.  Serve stale data to the client if you have to.\nFrom the page I quoted:\n\n\"In addition an option to shortcut the cache revalidation of\nfrequently accessed objects is added, making further requests\nimmediately return as a cache hit while a cache revalidation is\npending. This may temporarily give slightly stale information to the\nclients, but at the same time allows for optimal response time while a\nfrequently accessed object is being revalidated. This too is an\noptimization only intended for accelerators, and only for accelerators\nwhere minimizing request latency is morer important than freshness.\"\n\nI don't know how sophisticated this logic is currently, but I would\nthink that it wouldn't be that hard to tune up.\n\nCheers,\n"},{"id":"294501","messageId":"457A0214.9010900@zytor.com","threadId":"43081","inReplyTo":"f2b55d220612081618v26f7c714l7d4ea5315a964aa@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-09T00:23:48Z","receivedAt":"2006-12-09T00:23:48Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Michael K. Edwards wrote:\n> On 12/8/06, H. Peter Anvin <hpa@zytor.com> wrote:\n>> There is another thing that probably will be required, and I'm not sure\n>> if something in front of Apache (like Squid) rather than behind it can\n>> easily deal with: on timeout, the process needs to continue in order to\n>> feed the cache.  Otherwise, you're still in a failure scenario as soon\n>> as timeout happens.\n> \n> I would think this would be a great deal easier to handle in an\n> arm's-length \"accelerator\" than in the origin server\n\nTrue, but it needs to run behind Apache rather than in front of it.\n\n\t-hpa\n"},{"id":"295615","messageId":"Pine.LNX.4.64.0612081640400.3516@woody.osdl.org","threadId":"43081","inReplyTo":"4579FABC.5070509@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-09T00:45:16Z","receivedAt":"2006-12-09T00:45:16Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 8 Dec 2006, Jeff Garzik wrote:\n>\n> This is a bit cheesy, and completely untested, but since mod_cache never\n> worked for me either, I bet it works better ;-)\n\nOk, this doesn't do the locking either, so on cache misses or expiry, \nyou're still going to be that thundering herd.\n\nAlso, if you want to be nice to clients, I'd seriously suggest that when \nyou hit in the cache, but it's expired (or it's close to expired), you \nstill serve the cached data back, but you set up a thread in the \nbackground (with some maximum number of active threads, of course!) that \nrefreshes the cached entry and then you extend the expiration time so that \nyou won't end up doing this \"refresh\" _again_.\n\nIt's kind of silly to have people wait for 20 seconds just because a cache \nexpired five seconds ago. Much nicer to say \"ok, we allow a certain \ngrace-period during which we'll do the real lookup, but to make things \n_look_ really responsive, we still use the old cached value\".\n\n"},{"id":"296972","messageId":"457A07B4.2060303@zytor.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081640400.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-09T00:47:48Z","receivedAt":"2006-12-09T00:47:48Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> It's kind of silly to have people wait for 20 seconds just because a cache \n> expired five seconds ago. Much nicer to say \"ok, we allow a certain \n> grace-period during which we'll do the real lookup, but to make things \n> _look_ really responsive, we still use the old cached value\".\n> \n\nYup, DNS does this, and it's a Very Good Thing.\n\n"},{"id":"296920","messageId":"Pine.LNX.4.64.0612081648160.3516@woody.osdl.org","threadId":"43081","inReplyTo":"f2b55d220612081546u1ffa98e5q75be55d31da82a2f@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-09T00:49:31Z","receivedAt":"2006-12-09T00:49:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 8 Dec 2006, Michael K. Edwards wrote:\n> \n> In Squid 2.6:\n>    collapsed_forwarding on\n>    refresh_stale_window <seconds>\n> (apply the latter only to stanzas where you want \"readahead\" of\n> about-to-expire cache entries)\n\nYeah, those look like the Right Thing (tm) to do.\n\nThat said, I'm not personally convinced that there is much point to using \nnetfilter for transparent proxying. Why not just use separate ports for \nsquid and for apache?\n\n"},{"id":"296450","messageId":"457A08AC.8030100@zytor.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081648160.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-09T00:51:56Z","receivedAt":"2006-12-09T00:51:56Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 8 Dec 2006, Michael K. Edwards wrote:\n>> In Squid 2.6:\n>>    collapsed_forwarding on\n>>    refresh_stale_window <seconds>\n>> (apply the latter only to stanzas where you want \"readahead\" of\n>> about-to-expire cache entries)\n> \n> Yeah, those look like the Right Thing (tm) to do.\n> \n> That said, I'm not personally convinced that there is much point to using \n> netfilter for transparent proxying. Why not just use separate ports for \n> squid and for apache?\n> \n\nYeah, this is pretty trivial since one can just do redirects.  However, \nI still think a backend cache is better, since it can detach itself from \nApache when appropriate (e.g. the background refresh scenario, or timeout.)\n\n"},{"id":"297101","messageId":"46a038f90612081728s65d65ccewe64fa1a496de76fa@mail.gmail.com","threadId":"43081","inReplyTo":"200612081438.25493.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-09T01:28:25Z","receivedAt":"2006-12-09T01:28:25Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/9/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> Perhaps gitweb should generate it's own ETag instead of messing with\n> 'expires' header?\n\nThat'll be the winning solution. A combination of\n\n - cache SHA1-based requests forever\n - cache ref-based requests a longish time,  setting an ETag that\ncontains headname+SHA1\n - on 'revalidate', check the ETag vs the ref and only recompute if\nthings have changed\n\nIn the meantime, the code on kernel.org needs to be updated to the\nlatest gitweb. On our server, I'd say the newer gitweb is 3~4 times\nfaster serving the \"expensive\" summary pages. And much smarter in\nterms of caching headers.\n\ncheers\n\n\n"},{"id":"296725","messageId":"46a038f90612081756w1ab4609epcb4a2cbd9f4d8205@mail.gmail.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081453430.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-09T01:56:28Z","receivedAt":"2006-12-09T01:56:28Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/9/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Actually, just looking at the examples, it looks like memcached is\n> fundamentally flawed, exactly the same way Apache mod_cache is\n> fundamentally flawed.\n\nI don't know if fundamentally flawed but (having used memcached) I\ndon't think it's a big win for this at all.\n\nWe can make gitweb to detect mod_perl and a few smarter things if it\nis running inside of it. In fact, we can (ab)use mod_perl and perl\nfacilities a bit to do some serialization which will be a big win for\nsome pages. What we need for that is to set a sensible the ETag and\nuse some IPC to announce/check if other apache/modperl processes are\npreparing content for the same ETag. The first-process-to-announce a\ngiven ETag can then write it to a common temp directory (atomically -\nwrite to a temp-name and move to the expected name) while other\nprocesses wait, polling for the file. Once the file is in place the\nlatecomers can just serve the content of the file and exit.\n\n(I am calling the \"state we are serving\" identifier ETag because I\nthink we should also set it as the ETag in the HTTP headers, so well\nbe able to check the ETag of future requests for staleness - all we\nneed is a ref lookup, and if the SHA1 matches, we are sorted). So\nhaving this 'unique request identifier' doubles up nicely...\n\nThe ETag should probably be:\n - SHA1+displaytype+args for pages that display an object identified by SHA1\n - refname+SHA!+displaytype+args for pages that display something\nidentified by a ref\n - SHA1(names and sha1s of all refs) for the summary page\n\n> You can't have a cache architecture where the client just does a \"get\",\n> like memcached does. You need to have a \"read-for-fill\" operation, which\n> says:\n\nYou _could_ make do with a convention of polling for \"entryname\" and\n\"workingon-entryname\" and if \"workingon-entryname\" is set to 1, you\ncan expect entryname to be filled real soon now. However, memcached is\ncompletely memorybound, so it is only nice for really small stuff or\nfor a large server farm which has gobs of spare ram.\n\n(Note that memcached does have timeouts which means that the\n'workingon' value could have a short timeout in case the request is\ncancelled or the process dies - the nasty bit in the above plan would\nbe the polling.)\n\n> I still don't understand why apache doesn't do it. I guess it wants to be\n> stateless or something.\n\nApache doesn't do it because most web applications don't use the HTTP\nprocol correctly - specially when it comes to the idempotency of GET.\nSo in 99% of the cases, web apps serve truly different pages for the\nsame GET request, depending on your cookie, IP address, time-of-day,\netc.\n\nMost websites deal with very little traffic, so this isn't a problem.\nAnd many large sites that serve a lot of traffic from a dynamic web\napp want to be serving custom ads, let you login and see your\npersonalised toolbar, etc,etc, so this wouldn't work for them either.\n\nSo in practice, serialising speculatively on GET requests for the same\nURL has very little payoff except for static content. And that's quite\nfast anyway.... specially if the underlying OS is smokin' fast ;-)\n\ncheers,\n\n\n\n"},{"id":"298597","messageId":"457A1962.6000401@zytor.com","threadId":"43081","inReplyTo":"46a038f90612081728s65d65ccewe64fa1a496de76fa@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-09T02:03:14Z","receivedAt":"2006-12-09T02:03:14Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Martin Langhoff wrote:\n> On 12/9/06, Jakub Narebski <jnareb@gmail.com> wrote:\n>> Perhaps gitweb should generate it's own ETag instead of messing with\n>> 'expires' header?\n> \n> That'll be the winning solution.\n\nDoesn't solve the thundering herd problem or the timeout problem at all, \nthough.\n\n\t-hpa\n\n"},{"id":"295264","messageId":"46a038f90612081852u63e05da1qe57504636f3578fd@mail.gmail.com","threadId":"43081","inReplyTo":"457A1962.6000401@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-09T02:52:09Z","receivedAt":"2006-12-09T02:52:09Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/9/06, H. Peter Anvin <hpa@zytor.com> wrote:\n> Martin Langhoff wrote:\n> > On 12/9/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> >> Perhaps gitweb should generate it's own ETag instead of messing with\n> >> 'expires' header?\n> >\n> > That'll be the winning solution.\n>\n> Doesn't solve the thundering herd problem or the timeout problem at all,\n> though.\n\nI posted separately about those. And I've been mulling about whether\nthe thundering herd is really such a big problem that we need to\naddress it head-on. If we doHTTP  caching headers right (that is, a\nbit better than now) then the fact that web caches are distributed\nmeans that even a cache restart or cache invalidation won't trigger a\nthundering herd.\n\nAnd gitweb rarely has a \"new\" URL that gets a ton of hits immediately.\nOur real problem is the summary page, and the fact that we aren't\nsetting an effecting ETag there. If we do, a front-end cache plus the\nability to revalidate the ETag cheaply will get us through.\n\nWe get 99% of the benefit from ETags and cheap revalidations,\nspecially if they are coupled with a reverse caching proxy,. The\nremaining 1% of dealing with the highly infrequent thundering herd can\nbe addressed with the scheme I've posted 5 minutes ago.\n\ncheers\n\n\n"},{"id":"296874","messageId":"f2b55d220612082036u1853eb6ei40ad45c2150e1999@mail.gmail.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081648160.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Michael K. Edwards","fromEmail":"medwards.linux@gmail.com","sentAt":"2006-12-09T04:36:18Z","receivedAt":"2006-12-09T04:36:18Z","isPatch":false,"sender":{"key":"medwards.linux@gmail.com","avatar":null},"body":"On 12/8/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> That said, I'm not personally convinced that there is much point to using\n> netfilter for transparent proxying. Why not just use separate ports for\n> squid and for apache?\n\nJust a question of whether you want to be able to yank the squid box\nout if it goes pear-shaped, without touching configs on the apache\nbox.  Some people like to stick the proxy in as a no-op at first, then\ntell netfilter to divert 1% of sessions to squid and see how it holds\nup, retune, ease it in, ease it out, figure out how much operational\nflexibility you will have as demand continues to scale.  If the squid\nand apache are on the same box it's probably less of an issue.\n\nCheers,\n"},{"id":"298761","messageId":"457A44ED.4080606@zytor.com","threadId":"43081","inReplyTo":"46a038f90612081852u63e05da1qe57504636f3578fd@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-09T05:09:01Z","receivedAt":"2006-12-09T05:09:01Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Martin Langhoff wrote:\n> I posted separately about those. And I've been mulling about whether\n> the thundering herd is really such a big problem that we need to\n> address it head-on.\n\nUhm... yes it is.\n\n"},{"id":"294661","messageId":"46a038f90612082134x38be9c8dgca6fe60c087bf100@mail.gmail.com","threadId":"43081","inReplyTo":"457A44ED.4080606@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-09T05:34:18Z","receivedAt":"2006-12-09T05:34:18Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/9/06, H. Peter Anvin <hpa@zytor.com> wrote:\n> Martin Langhoff wrote:\n> > I posted separately about those. And I've been mulling about whether\n> > the thundering herd is really such a big problem that we need to\n> > address it head-on.\n>\n> Uhm... yes it is.\n\nGot some more info, discussion points or links to stuff I should read\nto appreciate why that is? I am trying to articulate why I consider it\nis not a high-payoff task, as well as describing how to tackle it.\n\nTo recap, the reasons it is not high payoff is that:\n\n - the main benefit comes from being cacheable and able to revalidate\nthe cache cheaply (with the ETags-based strategy discussed above)\n - highly distributed caches/proxies means we'll seldom see a true\ncold cache situation\n - we have a huge set of URLs which are seldom hit, and will never see\na thundering anything\n - we have a tiny set of very popular URLs that are the key target for\nthe thundering herd - (projects page, summary page, shortlog, fulllog)\n- but those are in the clear as soon as the caches are populated\n\nWhy do we have to take it head-on? :-)\n\n\n\n"},{"id":"297672","messageId":"457A6C1E.4080302@midwinter.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081453430.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2006-12-09T07:56:14Z","receivedAt":"2006-12-09T07:56:14Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> Looking at the memcached operations, they have the \"read\" op (aka \"get\"), \n> but they seem to have no \"read-for-fill\" op. So memcached fundamentally \n> doesn't fix this problem, at least without explicit serialization by the \n> client.\n>   \n\nActually, memcached does support an operation that would work for this: \nthe \"add\" request, which creates a new cache entry if and only if the \nkey is not already in the cache. If the key is already present, the \nrequest fails. You can use that to implement a simple named mutex, and \nit supports a client-specified timeout. The one thing it doesn't support \nthat you described is a notion of deleting a key when a particular \nclient disconnects, but as you say, that should only happen in the case \nof buggy clients anyway.\n\nMind you, I'm not convinced memcached is necessarily the right answer \nfor this problem, but it does provide a way to implement the required \nlocking semantics.\n\nBTW, I'm one of the main contributors to memcached, so if it does end up \nlooking like a good choice except for some minor issue or another, I may \nbe able to tweak it to cover whatever is missing. For example, the \n\"delete a key on disconnect\" thing would be fairly straightforward, if \nit's actually necessary in practice.\n\n-Steve\n"},{"id":"296659","messageId":"457A7EE8.80207@garzik.org","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081640400.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-09T09:16:24Z","receivedAt":"2006-12-09T09:16:24Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 8 Dec 2006, Jeff Garzik wrote:\n>> This is a bit cheesy, and completely untested, but since mod_cache never\n>> worked for me either, I bet it works better ;-)\n> \n> Ok, this doesn't do the locking either, so on cache misses or expiry, \n> you're still going to be that thundering herd.\n\nWell, gdbm does reader/write locking.\n\nYou still bit a bit of a thundering herd, though.  I suppose I could \nopen the gdbm db for writing before calling the CGI, which would \neffectively get what you're looking for.\n\n\n> Also, if you want to be nice to clients, I'd seriously suggest that when \n> you hit in the cache, but it's expired (or it's close to expired), you \n> still serve the cached data back, but you set up a thread in the \n> background (with some maximum number of active threads, of course!) that \n> refreshes the cached entry and then you extend the expiration time so that \n> you won't end up doing this \"refresh\" _again_.\n> \n> It's kind of silly to have people wait for 20 seconds just because a cache \n> expired five seconds ago. Much nicer to say \"ok, we allow a certain \n> grace-period during which we'll do the real lookup, but to make things \n> _look_ really responsive, we still use the old cached value\".\n\nTrue, should work with gitweb data at least.\n\n\tJeff\n\n"},{"id":"294139","messageId":"457A8197.5090503@garzik.org","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612081648160.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-09T09:27:51Z","receivedAt":"2006-12-09T09:27:51Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Linus Torvalds wrote:\n> That said, I'm not personally convinced that there is much point to using \n> netfilter for transparent proxying. Why not just use separate ports for \n> squid and for apache?\n\n\nThat's what most people using squid in \"http accelerator\" mode do.  They \nput Apache on port 8080 or somesuch.\n\n\tJeff\n\n"},{"id":"294739","messageId":"200612091251.16460.jnareb@gmail.com","threadId":"43081","inReplyTo":"46a038f90612081756w1ab4609epcb4a2cbd9f4d8205@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-09T11:51:15Z","receivedAt":"2006-12-09T11:51:15Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Langhoff wrote:\n\n> We can make gitweb to detect mod_perl and a few smarter things if it\n> is running inside of it. In fact, we can (ab)use mod_perl and perl\n> facilities a bit to do some serialization which will be a big win for\n> some pages. What we need for that is to set a sensible the ETag and\n> use some IPC to announce/check if other apache/modperl processes are\n> preparing content for the same ETag. The first-process-to-announce a\n> given ETag can then write it to a common temp directory (atomically -\n> write to a temp-name and move to the expected name) while other\n> processes wait, polling for the file. Once the file is in place the\n> latecomers can just serve the content of the file and exit.\n\nFirst, it would (and could) work only for serving gitweb over mod_perl.\nI'm not sure if overhead with IPC and complications implementing are\nworth it: this perhaps be better solved by caching engine.\n\nBut let us put aside for a while actual caching (writing HTML version\nof the page to a common temp directory, and serving this static page\nif possible), and talk a bit what gitweb can do with respect to\ncache validation.\n\nIn addition to setting either Expires: header or Cache-Control: max-age\ngitweb should also set Last-Modified: and ETag headers, and also \nprobably respond to If-Modified-Since: and If-None-Match: requests.\n\nWould be worth implementing this?\n \n> (I am calling the \"state we are serving\" identifier ETag because I\n> think we should also set it as the ETag in the HTTP headers, so well\n> be able to check the ETag of future requests for staleness - all we\n> need is a ref lookup, and if the SHA1 matches, we are sorted). So\n> having this 'unique request identifier' doubles up nicely...\n\nFor some pages ETag is natural; for other Last-Modified: would be more\nnatural.\n\n> The ETag should probably be:\n>  - SHA1+displaytype+args for pages that display an object identified\n>    by SHA1\n\nWhat uniquely identifies contents in \"object\" views (\"commit\", \"tag\",\n\"tree\", \"blob\") is either h=SHA1, or hb=SHA1;f=FILENAME (with absence\nof h=SHA1). If both h=SHA1 and hb=SHA1 is present, hb=SHA1 serves as\nbacklink. The \"diff\" views (\"commitdiff\", \"blobdiff\") are uniquely\nidentified by pair of object identifiers (pairs of SHA1, or pairs of\nhb SHA1 + FILENAME).\n\nThree of those views (\"blob\", \"commitdiff\", \"blobdiff\") have their \n\"plain\" version; so ETag should include displaytype (action, 'a' \nparameter).\n\nThe hb=SHA1;f=FILENAME indentifier can be converted at cost of one\ncall to git command (but which is a bit expensive as it recurses\ntrees), namely to git-ls-tree.\n\nETag can be simply args (query), if all h/hb/hbp parameters are SHA1.\nOr ETag can be SHA1 of an object (or pair of SHA1 in the case of diff),\nbut this is little more costly to verify. Although we usually (always?) \nconvert hb=SHA1;f=FILENAME to h=SHA1 anyway when displaying/generating \npage.\n\nUsualy you can compare ETags base on URL alone.\n   \n>  - refname+SHA!+displaytype+args for pages that display something\n>    identified by a ref\n\nFor objects views we can simply convert refname to SHA1. I'm not sure if \nit is worth it. In the cases when for view we have to calculate SHA1 of \nobject anyway, we can return (and validate) ETag with SHA1 as above.\n\n- ETag and/or Last-Modified headers for \"log\" views: \"log\", \n\"shortlog\" (is part of summary view), \"history\", \"rss\"/\"atom\" views.\n\nOn one hand all log views (at least now) are identified by their \nparameters (action/view name, and filename in the case of history view) \nand SHA1 of top commit. On the other hand it might be easier to use \nLast-Modified with date of top commit... Verifying SHA1 based ETag \ncould add some overhead in the case of miss.\n\n>  - SHA1(names and sha1s of all refs) for the summary page\n\nWouldn't it be simplier to just set Last-Modified: header (and check\nit?)\n\n\nP.S. Can anyone post some benchmark comparing gitweb deployed under \nmod_perl as compared to deployed as CGI script? Does kernel.org use \nmod_perl, or CGI version of gitweb?\n\n-- \nJakub Narebski\n"},{"id":"295196","messageId":"457AAF31.2050002@garzik.org","threadId":"43081","inReplyTo":"200612091251.16460.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-09T12:42:25Z","receivedAt":"2006-12-09T12:42:25Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Jakub Narebski wrote:\n> First, it would (and could) work only for serving gitweb over mod_perl.\n> I'm not sure if overhead with IPC and complications implementing are\n> worth it: this perhaps be better solved by caching engine.\n\nIt is.  At least for kernel.org, the issue isn't that CGI is expensive, \nits that I/O is expensive.\n\n\n> In addition to setting either Expires: header or Cache-Control: max-age\n> gitweb should also set Last-Modified: and ETag headers, and also \n> probably respond to If-Modified-Since: and If-None-Match: requests.\n> \n> Would be worth implementing this?\n\nIMO yes, since most major browsers, caches, and spiders support these \nheaders.\n\n\n> For some pages ETag is natural; for other Last-Modified: would be more\n> natural.\n\nYes, a good point to note.\n\n\n> Usualy you can compare ETags base on URL alone.\n\nMostly true:  you must also consider HTTP_ACCEPT\n\n\n> Wouldn't it be simplier to just set Last-Modified: header (and check\n> it?)\n\nThat would be a good start, and suffice for many cases.  If the CGI can \nsimply stat(2) files rather than executing git-* programs, that would \nincrease efficiency quite a bit.\n\nA core problem with cache hints via HTTP headers (last-modified, etc.) \nis that you don't achieve caching across multiple clients, just across \nrepeated queries from the same client (or caching proxy).\n\nAt least for the RSS/Atom feeds and the git main page, it makes no sense \nto regenerate that data repeatedly.\n\nInternally, gitweb would need to do a stat() on key files, and return \npre-generated XML for the feeds if the stat() reveals no changes.  Ditto \nfor the front page.\n\n\n> P.S. Can anyone post some benchmark comparing gitweb deployed under \n> mod_perl as compared to deployed as CGI script? Does kernel.org use \n> mod_perl, or CGI version of gitweb?\n\nCGI version of gitweb.\n\nBut again, mod_perl vs. CGI isn't the issue.\n\n\tJeff\n\n"},{"id":"294925","messageId":"200612091437.01183.jnareb@gmail.com","threadId":"43081","inReplyTo":"457AAF31.2050002@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-09T13:37:00Z","receivedAt":"2006-12-09T13:37:00Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jeff Garzik wrote:\n> Jakub Narebski wrote:\n\n>> In addition to setting either Expires: header or Cache-Control: max-age\n>> gitweb should also set Last-Modified: and ETag headers, and also \n>> probably respond to If-Modified-Since: and If-None-Match: requests.\n>> \n>> Would be worth implementing this?\n> \n> IMO yes, since most major browsers, caches, and spiders support these \n> headers.\n \nSending Last-Modified: should be easy; sending ETag needs some consensus\non the contents: mainly about validation. Responding to If-Modified-Since:\nand If-None-Match: should cut at least _some_ of the page generating time.\nIf ETag can be calculated on URL alone, then we can cut If-None-Match:\njust at beginning of script.\n \n>> For some pages ETag is natural; for other Last-Modified: would be more\n>> natural.\n> \n> Yes, a good point to note.\n> \n>> Usualy you can compare ETags base on URL alone.\n> \n> Mostly true:  you must also consider HTTP_ACCEPT\n\nWell, yes, ETag is HTTP/1.1 header. \n\n>> Wouldn't it be simplier to just set Last-Modified: header (and check\n>> it?)\n> \n> That would be a good start, and suffice for many cases.  If the CGI can \n> simply stat(2) files rather than executing git-* programs, that would \n> increase efficiency quite a bit.\n\nAs I said, I'm not talking (at least now) about saving generated HTML\noutput. This I think is better solved in caching engine like Squid can\nbe. Although even here some git specific can be of help: we can invalidate\ncache on push, and we know that some results doesn't ever change (well,\nwith exception of changing output of gitweb).\n\n> A core problem with cache hints via HTTP headers (last-modified, etc.) \n> is that you don't achieve caching across multiple clients, just across \n> repeated queries from the same client (or caching proxy).\n> \n> At least for the RSS/Atom feeds and the git main page, it makes no sense \n> to regenerate that data repeatedly.\n> \n> Internally, gitweb would need to do a stat() on key files, and return \n> pre-generated XML for the feeds if the stat() reveals no changes.  Ditto \n> for the front page.\n\nI'm not sure if it is worth implementing in gitweb, or is it better left\nto caching engine. With the projects list page and summary page there is\nadditional problem with relative dates, although this can be solved using\nJonas Fonseca idea of using absolute dates in the page and using ECMAScript\n(JavaScript) to convert them to relative: on load, and perhaps on timer ;-)\n\n\nWhat can be _easily_ done:\n * Use post 1.4.4 gitweb, which uses git-for-each-ref to generate summary\n   page; this leads to around 3 times faster summary page.\n * Perhaps using projects list file (which can be now generated by gitweb)\n   instead of scanning directories and stat()-ing for owner would help\n   with time to generate projects lis page\n\nWhat can be quite easy incorporated into gitweb:\n * For immutable pages set Expires: or Cache-Control: max-age (or both)\n   to infinity\n * Calculate hash+action based ETag at least for those actions where it is\n   easy, and respond with 304 Not Modified as soon as it can.\n   This might require some code reorganization to not begin writing output\n   before calculating ETag and ETag comparison (If-Match, If-None-Match).\n * Generate Last-Modified: for those views where it can be calculated,\n   and respond with 304 Not Modified as soon as it can.\n\nWhat can be easily done using caching engine:\n * Select top 10 of common queries, and cache them, invalidating cache on push\n   (depending on query: for example invalidate project list on push to any\n   project, invalidate RSS/Atom feed and summary pages only on push to specific\n   project) - can be done with git hooks.\n-- \nJakub Narebski\n"},{"id":"295680","messageId":"457ACBA1.4090007@garzik.org","threadId":"43081","inReplyTo":"200612091437.01183.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-09T14:43:45Z","receivedAt":"2006-12-09T14:43:45Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Jakub Narebski wrote:\n> Sending Last-Modified: should be easy; sending ETag needs some consensus\n> on the contents: mainly about validation. Responding to If-Modified-Since:\n> and If-None-Match: should cut at least _some_ of the page generating time.\n\nDefinitely.\n\n\n> As I said, I'm not talking (at least now) about saving generated HTML\n> output. This I think is better solved in caching engine like Squid can\n> be. Although even here some git specific can be of help: we can invalidate\n> cache on push, and we know that some results doesn't ever change (well,\n> with exception of changing output of gitweb).\n\nIt depends on how creatively you think ;-)\n\nConsider generating static HTML files on each push, via a hook, for many \nof the toplevel files.  The static HTML would then link to the CGI for \nfurther dynamic querying of the git database.\n\n\n\n> What can be _easily_ done:\n>  * Use post 1.4.4 gitweb, which uses git-for-each-ref to generate summary\n>    page; this leads to around 3 times faster summary page.\n\nThis re-opens the question mentioned earlier, is Kay (or anyone?) still \nactively maintaining gitweb on k.org?\n\n\n>  * Perhaps using projects list file (which can be now generated by gitweb)\n>    instead of scanning directories and stat()-ing for owner would help\n>    with time to generate projects lis page\n\nThis could be statically generated by a robot.  I think everybody would \nshrink in horror if a human needed to maintain such a file.\n\n\n> What can be quite easy incorporated into gitweb:\n>  * For immutable pages set Expires: or Cache-Control: max-age (or both)\n>    to infinity\n\nnice!\n\n\n>  * Generate Last-Modified: for those views where it can be calculated,\n>    and respond with 304 Not Modified as soon as it can.\n\nagreed\n\n\n> What can be easily done using caching engine:\n>  * Select top 10 of common queries, and cache them, invalidating cache on push\n>    (depending on query: for example invalidate project list on push to any\n>    project, invalidate RSS/Atom feed and summary pages only on push to specific\n>    project) - can be done with git hooks.\n\nOr simply generate regular filesystem files into the webspace, as \ntriggered by a hook.  Let the standard filesystem mirroring/caching work \nits magic.\n\n\tJeff\n\n"},{"id":"295991","messageId":"457AE3A7.4080802@zytor.com","threadId":"43081","inReplyTo":"46a038f90612082134x38be9c8dgca6fe60c087bf100@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-09T16:26:15Z","receivedAt":"2006-12-09T16:26:15Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Martin Langhoff wrote:\n> On 12/9/06, H. Peter Anvin <hpa@zytor.com> wrote:\n>> Martin Langhoff wrote:\n>> > I posted separately about those. And I've been mulling about whether\n>> > the thundering herd is really such a big problem that we need to\n>> > address it head-on.\n>>\n>> Uhm... yes it is.\n> \n> Got some more info, discussion points or links to stuff I should read\n> to appreciate why that is? I am trying to articulate why I consider it\n> is not a high-payoff task, as well as describing how to tackle it.\n> \n> To recap, the reasons it is not high payoff is that:\n> \n> - the main benefit comes from being cacheable and able to revalidate\n> the cache cheaply (with the ETags-based strategy discussed above)\n> - highly distributed caches/proxies means we'll seldom see a true\n> cold cache situation\n> - we have a huge set of URLs which are seldom hit, and will never see\n> a thundering anything\n> - we have a tiny set of very popular URLs that are the key target for\n> the thundering herd - (projects page, summary page, shortlog, fulllog)\n> - but those are in the clear as soon as the caches are populated\n> \n> Why do we have to take it head-on? :-)\n> \n\nBecause the primary failure scenario is timeout on the common queries \ndue to excess parallel invocations under high I/O load resulting in \ncatastrophic failure.\n\n"},{"id":"298721","messageId":"200612091802.12810.jnareb@gmail.com","threadId":"43081","inReplyTo":"457ACBA1.4090007@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-09T17:02:12Z","receivedAt":"2006-12-09T17:02:12Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jeff Garzik wrote:\n> Jakub Narebski wrote:\n\n>> As I said, I'm not talking (at least now) about saving generated HTML\n>> output. This I think is better solved in caching engine like Squid can\n>> be. Although even here some git specific can be of help: we can invalidate\n>> cache on push, and we know that some results doesn't ever change (well,\n>> with exception of changing output of gitweb).\n> \n> It depends on how creatively you think ;-)\n> \n> Consider generating static HTML files on each push, via a hook, for many \n> of the toplevel files.  The static HTML would then link to the CGI for \n> further dynamic querying of the git database.\n\nYou mean that the links in this pre-generated HTML would be to CGI\npages?\n \n>> What can be _easily_ done:\n>>  * Use post 1.4.4 gitweb, which uses git-for-each-ref to generate summary\n>>    page; this leads to around 3 times faster summary page.\n> \n> This re-opens the question mentioned earlier, is Kay (or anyone?) still \n> actively maintaining gitweb on k.org?\n\nBy the way, thanks to Martin Waitz it is much easier to install gitweb.\nI for example use the following script to test changes I have made to gitweb:\n\n-- >8 --\n#!/bin/bash\n\nBINDIR=\"/home/local/git\"\n\nfunction make_gitweb()\n{\n\tpushd \"/home/jnareb/git/\"\n\n\tmake GITWEB_PROJECTROOT=\"/home/local/scm\" \\\n\t     GITWEB_CSS=\"/gitweb/gitweb.css\" \\\n\t     GITWEB_LOGO=\"/gitweb/git-logo.png\" \\\n\t     GITWEB_FAVICON=\"/gitweb/git-favicon.png\" \\\n\t     bindir=$BINDIR \\\n\t     gitweb/gitweb.cgi\n\n\tpopd\n}\n\nfunction copy_gitweb()\n{\n\tcp -fv /home/jnareb/git/gitweb/gitweb.{cgi,css} /home/local/gitweb/\n}\n\nmake_gitweb\ncopy_gitweb\n\n# end of gitweb-update.sh\n-- >8 --\n\n>>  * Perhaps using projects list file (which can be now generated by gitweb)\n>>    instead of scanning directories and stat()-ing for owner would help\n>>    with time to generate projects lis page\n> \n> This could be statically generated by a robot.  I think everybody would \n> shrink in horror if a human needed to maintain such a file.\n\nGitweb can generate this file. The problem is that one would have to\ntemporary turn off using index file. This can be done by having the\nfollowing gitweb_list_projects.perl file:\n\n-- >8 --\n#!/usr/bin/perl\n\n$projects_list = \"\";\n-- >8 --\n\nthen use the following invocation to generate project index file:\n\n$ GATEWAY_INTERFACE=\"CGI/1.1\" HTTP_ACCEPT=\"*/*\" REQUEST_METHOD=\"GET\" \\\n  GITWEB_CONFIG=gitweb_list_projects.perl QUERY_STRING=\"a=project_index\" \\\n  gitweb.cgi \n\n-- \nJakub Narebski\n"},{"id":"293822","messageId":"457AF201.2000205@garzik.org","threadId":"43081","inReplyTo":"200612091802.12810.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-09T17:27:29Z","receivedAt":"2006-12-09T17:27:29Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Jakub Narebski wrote:\n> Jeff Garzik wrote:\n>> Jakub Narebski wrote:\n> \n>>> As I said, I'm not talking (at least now) about saving generated HTML\n>>> output. This I think is better solved in caching engine like Squid can\n>>> be. Although even here some git specific can be of help: we can invalidate\n>>> cache on push, and we know that some results doesn't ever change (well,\n>>> with exception of changing output of gitweb).\n>> It depends on how creatively you think ;-)\n>>\n>> Consider generating static HTML files on each push, via a hook, for many \n>> of the toplevel files.  The static HTML would then link to the CGI for \n>> further dynamic querying of the git database.\n> \n> You mean that the links in this pre-generated HTML would be to CGI\n> pages?\n\nYes, they must be.  Otherwise, the gitweb interface changes.\n\nYou don't want to pre-generate HTML for every possible git query, that \nwould cause an explosion of data.\n\nBoth the HTML generator and CGI would need to know which pages were \npre-generated and which are not.\n\n\tJeff\n\n"},{"id":"297374","messageId":"Pine.LNX.4.64.0612090957360.3516@woody.osdl.org","threadId":"43081","inReplyTo":"457AAF31.2050002@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-09T18:04:10Z","receivedAt":"2006-12-09T18:04:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 9 Dec 2006, Jeff Garzik wrote:\n> \n> It is.  At least for kernel.org, the issue isn't that CGI is expensive, its\n> that I/O is expensive.\n\nNote that if we had a new gitweb, we could also used the packed refs. \nThose help CPU usage, but they actually help IO patterns more, exactly \nbecause they avoid all the seeking around in the filesystem.\n\nSo with packed refs, there's no need to go from directory lookup to inode \nlookup to data lookup to object lookup for *each* ref - you can do the \n\"packed-refs\" lookup _once_ (which obviously does the dir->inode->data), \nand you don't need to do the object lookup at all.\n\nOf course, gitweb will then end up doing the object lookup anyway (because \nof getting the dates etc for refs), but if you have packed-refs and a \nreasonably packed repository, that should still really cut down on IO in a \nbig way.\n\nSo there's probably tons of room for making this more efficient: using a \nnewer gitweb, packing refs, using the cgi cache thing.. It sounds like \nwhat it really needs is just somebody with the competence and time to be \nwilling to step up and maintain gitweb on kernel.org...\n\n"},{"id":"293903","messageId":"457B00AB.3030308@zytor.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612090957360.3516@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-09T18:30:03Z","receivedAt":"2006-12-09T18:30:03Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> So there's probably tons of room for making this more efficient: using a \n> newer gitweb, packing refs, using the cgi cache thing.. It sounds like \n> what it really needs is just somebody with the competence and time to be \n> willing to step up and maintain gitweb on kernel.org...\n> \n\nIndeed.  We have a lot of projects on kernel.org which are like this: \nnot at all conceptually hard, but a huge time commitment for Doing It \nRight[TM].  This is why I sometimes think that it would be a Good Thing \nto get paid staff for kernel.org, although I was hoping to defer the \nneed for that until at least we have our 501(c)3 paperwork done, which \nlooks like mid-2007 at this point (assuming no further delays.)\n\n"},{"id":"297613","messageId":"46a038f90612091955i5bdd6e85l749a2f511f27953@mail.gmail.com","threadId":"43081","inReplyTo":"457AAF31.2050002@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-10T03:55:28Z","receivedAt":"2006-12-10T03:55:28Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/10/06, Jeff Garzik <jeff@garzik.org> wrote:\n> > P.S. Can anyone post some benchmark comparing gitweb deployed under\n> > mod_perl as compared to deployed as CGI script? Does kernel.org use\n> > mod_perl, or CGI version of gitweb?\n>\n> CGI version of gitweb.\n>\n> But again, mod_perl vs. CGI isn't the issue.\n\nIO is the issue, and the CGI startup of Perl is quite IO & CPU\nintensive. Even if the caching headers, thundering herds and planet\ncollisions are resolved, I don't think you'll ever be happy with IO\nand CPU load on kernel.org running gitweb as CGI.\n\ncheers,\n\n\n\n"},{"id":"297068","messageId":"46a038f90612092007w4637637aya1a01ec18ff16f6f@mail.gmail.com","threadId":"43081","inReplyTo":"200612091437.01183.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-10T04:07:22Z","receivedAt":"2006-12-10T04:07:22Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/10/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> Jeff Garzik wrote:\n> > Jakub Narebski wrote:\n>\n> >> In addition to setting either Expires: header or Cache-Control: max-age\n> >> gitweb should also set Last-Modified: and ETag headers, and also\n> >> probably respond to If-Modified-Since: and If-None-Match: requests.\n> >>\n> >> Would be worth implementing this?\n> >\n> > IMO yes, since most major browsers, caches, and spiders support these\n> > headers.\n>\n> Sending Last-Modified: should be easy; sending ETag needs some consensus\n> on the contents: mainly about validation. Responding to If-Modified-Since:\n> and If-None-Match: should cut at least _some_ of the page generating time.\n> If ETag can be calculated on URL alone, then we can cut If-None-Match:\n> just at beginning of script.\n\nIndeed. Let me add myself to the pileup agreeing that a combination of\nsetting Last-Modified and checking for If-Modified-Since for\nref-centric pages (log, shortlog, RSS, and summary) is the smartest\nscheme. I got locked into thinking ETags.\n\n> > That would be a good start, and suffice for many cases.  If the CGI can\n> > simply stat(2) files rather than executing git-* programs, that would\n> > increase efficiency quite a bit.\n>\n> As I said, I'm not talking (at least now) about saving generated HTML\n> output. This I think is better solved in caching engine like Squid can\n> be. Although even here some git specific can be of help: we can invalidate\n> cache on push, and we know that some results doesn't ever change (well,\n> with exception of changing output of gitweb).\n\nIndeed - gitweb should not be saving HTML around bit giving the best\npossible hints to squid and friends. And improving our ability to\nshort-cut and send a 304 - Not Modified.\n\n> What can be _easily_ done:\n\nGreat plan. :-)\n\n\ncheers,\n\n\n"},{"id":"296816","messageId":"457BB1A3.2070408@zytor.com","threadId":"43081","inReplyTo":"46a038f90612091955i5bdd6e85l749a2f511f27953@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-10T07:05:07Z","receivedAt":"2006-12-10T07:05:07Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Martin Langhoff wrote:\n> On 12/10/06, Jeff Garzik <jeff@garzik.org> wrote:\n>> > P.S. Can anyone post some benchmark comparing gitweb deployed under\n>> > mod_perl as compared to deployed as CGI script? Does kernel.org use\n>> > mod_perl, or CGI version of gitweb?\n>>\n>> CGI version of gitweb.\n>>\n>> But again, mod_perl vs. CGI isn't the issue.\n> \n> IO is the issue, and the CGI startup of Perl is quite IO & CPU\n> intensive. Even if the caching headers, thundering herds and planet\n> collisions are resolved, I don't think you'll ever be happy with IO\n> and CPU load on kernel.org running gitweb as CGI.\n> \n\nI/O - nonexistent; that stuff will be in memory.\n\nCPU - we have more CPU than you can shake a stick at, and it's 95+% idle.\n\n*NOT AN ISSUE*.\n\n"},{"id":"295753","messageId":"87dcb0bd0612100143t21932358k42fe5044654e1981@mail.gmail.com","threadId":"43081","inReplyTo":"457995F8.1080405@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"rda","fromEmail":"rda@google.com","sentAt":"2006-12-10T09:43:04Z","receivedAt":"2006-12-10T09:43:04Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"On 12/8/06, H. Peter Anvin <hpa@zytor.com> wrote:\n> Linus Torvalds wrote:\n> > I could write a simple C caching thing that just hashes the CGI arguments\n> > and uses a hash to create a cache (and proper lock-files etc to serialize\n> > access to a particular cache object while it's being created) fairly\n> > easily, but I'm pretty sure people would much prefer a mod_perl thing just\n> > to avoid the fork/exec overhead with Apache (I think mod_perl allows\n> > Apache to run perl scripts without it), and that means I'm not the right\n> > person any more.\n>\n> True about mod_perl.  Haven't messed with that myself, either.\n> fork/exec really is very cheap on Linux, so it's not a huge deal.\n\nIn the case of Perl scripts, it's not really the fork/exec overhead,\nbut the Perl startup overhead that you want to try to optimize.  But\ngiven your later statement (lots of spare cpu), this ends up just\nbeing a bit of a latency hit.   In general, I think mod_perl has a\n"},{"id":"294345","messageId":"200612101109.34267.jnareb@gmail.com","threadId":"43081","inReplyTo":"46a038f90612092007w4637637aya1a01ec18ff16f6f@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-10T10:09:33Z","receivedAt":"2006-12-10T10:09:33Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Langhoff wrote:\n> On 12/10/06, Jakub Narebski <jnareb@gmail.com> wrote:\n\n>> Sending Last-Modified: should be easy; sending ETag needs some consensus\n>> on the contents: mainly about validation. Responding to If-Modified-Since:\n>> and If-None-Match: should cut at least _some_ of the page generating time.\n>> If ETag can be calculated on URL alone, then we can cut If-None-Match:\n>> just at beginning of script.\n> \n> Indeed. Let me add myself to the pileup agreeing that a combination of\n> setting Last-Modified and checking for If-Modified-Since for\n> ref-centric pages (log, shortlog, RSS, and summary) is the smartest\n> scheme. I got locked into thinking ETags.\n\nSometimes it is easier to use ETags, sometimes it is easier to use\nLast-Modified:. Usually you can check ETag earlier (after calling\ngit-rev-list) than Last-Modified (after parsing first commit). But\nsome pages doesn't have natural ETag...\n\nBesides, because ETag is HTTP/1.1 we should provide and validate\nboth.\n\nP.S. Any hints to how to do this with CGI Perl module?\n-- \nJakub Narebski\n"},{"id":"297122","messageId":"457C0060.3050605@garzik.org","threadId":"43081","inReplyTo":"200612101109.34267.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-10T12:41:04Z","receivedAt":"2006-12-10T12:41:04Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Jakub Narebski wrote:\n> P.S. Any hints to how to do this with CGI Perl module?\n\nIt's impossible, Apache doesn't supply e-tag info to CGI programs.  (it \ndoes supply HTTP_CACHE_CONTROL though apparently)\n\nYou could probably do it via mod_perl.\n\n\tJeff\n\n"},{"id":"296793","messageId":"200612101402.51363.jnareb@gmail.com","threadId":"43081","inReplyTo":"457C0060.3050605@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-10T13:02:50Z","receivedAt":"2006-12-10T13:02:50Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jeff Garzik wrote:\n> Jakub Narebski wrote:\n>>\n>> P.S. Any hints to how to do this with CGI Perl module?\n> \n> It's impossible, Apache doesn't supply e-tag info to CGI programs.  (it \n> does supply HTTP_CACHE_CONTROL though apparently)\n\nBy ETag info you mean access to HTTP headers sent by browser\nIf-Modified-Since:, If-Match:, If-None-Match: do you?\n \nIt's a pity that CGI interface doesn't cover that...\n\n> You could probably do it via mod_perl.\n\nSo the cache verification should be wrapped in if ($ENV{MOD_PERL}) ?\n-- \nJakub Narebski\n"},{"id":"297378","messageId":"457C0F8F.7030504@garzik.org","threadId":"43081","inReplyTo":"200612101402.51363.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-10T13:45:51Z","receivedAt":"2006-12-10T13:45:51Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Jakub Narebski wrote:\n> Jeff Garzik wrote:\n>> Jakub Narebski wrote:\n>>> P.S. Any hints to how to do this with CGI Perl module?\n>> It's impossible, Apache doesn't supply e-tag info to CGI programs.  (it \n>> does supply HTTP_CACHE_CONTROL though apparently)\n> \n> By ETag info you mean access to HTTP headers sent by browser\n> If-Modified-Since:, If-Match:, If-None-Match: do you?\n\nYou can use this attached shell script as a CGI script, to see precisely \nwhat information Apache gives you.  You can even experiment with passing \nback headers other than Content-type (such as E-tag), to see what sort \nof results are produced.  The script currently passes back both E-Tag \nand Last-Modified of a sample file; modify or delete those lines to suit \nyour experiments.\n\n\n> It's a pity that CGI interface doesn't cover that...\n> \n>> You could probably do it via mod_perl.\n> \n> So the cache verification should be wrapped in if ($ENV{MOD_PERL}) ?\n\nSorry, I was /assuming/ mod_perl would make this available.  The HTTP \nheader info is available to all Apache modules, but I confess I have no \nidea how mod_perl passes that info to scripts.\n\nAlso, an interesting thing while I was testing the attached shell \nscript:  even though repeated hits to the script generate a proper 304 \nresponse to the browse, the CGI script and its output run to completion. \n  So, it didn't save work on the CGI side; the savings was solely in not \ntransmitting the document from server to client.  The server still went \nthrough the work of generating the document (by running the CGI), as one \nwould expect.\n\n\tJeff\n\n\n\n\n#!/bin/sh\n\nFN=/tmp/foo\n\nif [ ! -f \"$FN\" ]\nthen\n\techo \"blah blah blah\" > \"$FN\"\nfi\n\nHASH=`md5sum \"$FN\"`\n\necho \"Content-type: text/plain\"\necho \"E-tag: $HASH\"\necho Last-Modified: `date -r /tmp/foo '+%a, %d %b %Y %T %Z'`\necho \"\"\n\n# don't pollute server environment output with our local additions\nunset FN\nunset HASH\n\nset\n"},{"id":"295998","messageId":"200612102011.52589.jnareb@gmail.com","threadId":"43081","inReplyTo":"457C0F8F.7030504@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-10T19:11:52Z","receivedAt":"2006-12-10T19:11:52Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jeff Garzik wrote:\n> Jakub Narebski wrote:\n>> Jeff Garzik wrote:\n>>> Jakub Narebski wrote:\n>>>>\n>>>> P.S. Any hints to how to do this with CGI Perl module?\n>>>\n>>> It's impossible, Apache doesn't supply e-tag info to CGI programs.  (it \n>>> does supply HTTP_CACHE_CONTROL though apparently)\n>> \n>> By ETag info you mean access to HTTP headers sent by browser\n>> If-Modified-Since:, If-Match:, If-None-Match: do you?\n\nAdn in CGI standard there is a way to access additional HTTP headers\ninfo from CGI script: the envirionmental variables are HTTP_HEADER,\nfor example if browser sent If-Modified-Since: header it's value\ncan be found in HTTP_IF_MODIFIED_SINCE environmental variable.\n\nBut of course gitweb should rather use mod_perl if possible, so\nsomewhere in gitweb there would be the following line:\n\n  $in_date = $ENV{'MOD_PERL'} ?\n    $r->header('If-Modified-Since') :\n    $ENV{'HTTP_IF_MODIFIED_SINCE'};\n\nor something like that...\n \n> You can use this attached shell script as a CGI script, to see precisely \n> what information Apache gives you.  You can even experiment with passing \n> back headers other than Content-type (such as E-tag), to see what sort \n> of results are produced.  The script currently passes back both E-Tag \n> and Last-Modified of a sample file; modify or delete those lines to suit \n> your experiments.\n\nIt is ETag, not E-tag. Besides, I don't see what the attached script is\nmeant to do: it does not output the sample file anyway.\n\n>> It's a pity that CGI interface doesn't cover that...\n>> \n>>> You could probably do it via mod_perl.\n>> \n>> So the cache verification should be wrapped in if ($ENV{MOD_PERL}) ?\n> \n> Sorry, I was /assuming/ mod_perl would make this available.  The HTTP \n> header info is available to all Apache modules, but I confess I have no \n> idea how mod_perl passes that info to scripts.\n> \n> Also, an interesting thing while I was testing the attached shell \n> script:  even though repeated hits to the script generate a proper 304 \n> response to the browse, the CGI script and its output run to completion. \n>   So, it didn't save work on the CGI side; the savings was solely in not \n> transmitting the document from server to client.  The server still went \n> through the work of generating the document (by running the CGI), as one \n> would expect.\n\nThe idea is of course to stop processing in CGI script / mod_perl script\nas soon as possible if cache validates.\n\nI don't know if Apache intercepts and remembers ETag and Last-Modified\nheaders, adds 304 Not Modified HTTP response on finding that cache validates\nand cuts out CGI script output. I.e. if browser provided If-Modified-Since:,\nscript wrote Last-Modified: header, If-Modified-Since: is no earlier than\nLast-Modified: (usually is equal in the case of cache validation), then\nApache provides 304 Not Modified response instead of CGI script output.\n\n-- \nJakub Narebski\n"},{"id":"294039","messageId":"Pine.LNX.4.64.0612101129190.12500@woody.osdl.org","threadId":"43081","inReplyTo":"200612102011.52589.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-10T19:50:15Z","receivedAt":"2006-12-10T19:50:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 10 Dec 2006, Jakub Narebski wrote:\n> >> If-Modified-Since:, If-Match:, If-None-Match: do you?\n> \n> Adn in CGI standard there is a way to access additional HTTP headers\n> info from CGI script: the envirionmental variables are HTTP_HEADER,\n> for example if browser sent If-Modified-Since: header it's value\n> can be found in HTTP_IF_MODIFIED_SINCE environmental variable.\n\nGuys, you're missing something fairly fundamnetal. \n\nIt helps almost _nothing_ to support client-side caching with all these \nfancy \"If-Modified-Since:\" etc crap.\n\nThat's not the _problem_.\n\nIt's usually not one client asking for the gitweb pages: the load comes \nfrom just lots of people independently asking for it. So client-side \ncaching may help a tiny tiny bit, but it's not actually fixing the \nfundamental problem at all.\n\nSo forget about \"If-Modified-Since:\" etc. It may help in benchmarks when \nyou try it yourself, and use \"refresh\" on the client side. But the basic \nproblem is all about lots of clients that do NOT have things cached, \nbecause all teh client caches are all filled up with pr0n, not with gitweb \ndata from yesterday.\n\nSo the thing to help is server-side caching with good access patterns, so \nthat the server won't have to seek all over the disk when clients that \n_don't_ have things in their caches want to see the \"git projects\" summary \noverview (that currently lists something like 200+ projects).\n\nSo to get that list of 200+ projects, right now gitweb will literally walk \nthem all, look at their refs, their descriptions, their ages (which \nrequires looking up the refs, and the objects behing the refs), and if \nthey aren't cached, you're going to have several disk seeks for each \nproject.\n\nAt 200+ projects, the thing that makes it slow is those disk seeks. Even \nwith a fast disk and RAID array, the seeks are all basically going to be \ninterdependent, so there's no room for disk arm movement optimization, and \nin the absense of any other load it's still going to be several seconds \njust for the seeks (say 10ms per seek, four or five seeks per project, \nyou've got 10 seconds _just_ for the seeks to generate the top-level \nsummary page, and quite frankly, five seeks is probably optimistic).\n\nNow, hopefully some of it will be in the disk cache, but when the \nmirroring happens, it will basically blow the disk caches away totally \n(when using the \"--checksum\" option), and then you literally have tens of \nseconds to generate that one top-level page. \n\nAnd when mirroring is blowing out the disk caches, the thing will be doing \nother things _too_ to the disk, of course.\n\nSo what you want is server-side caching, and you basically _never_ want to \nre-generate that data synchronously (because even if the server can take \nthe load, having the clients wait for half a minute or more for the data \nis just NOT FRIENDLY). This is why I suggested the grace-period where we \nfill the cache on he server side in the background _while_at_the_same_time \nactually feeding the clients the old cached contents.\n\nBecause what matters most to _clients_ is not getting the most recent \nup-to-date data within the last few minutes - people who go to the \noverview page want to just get a list of projects, and they want to get \nthem in a second or two, not half a minute later.\n\nAnd btw, all those \"If-Modified-Since:\" things are irrelevant, since quite \noften, the top-level page really technically _has_ been modified in the \nlast few minutes, because with the kernel and git projects, _somebody_ has \nusually pushed out one of the projects within the last hour.\n\nAnd no, people don't just sit there refreshing their browser page all the \ntime. I bet even \"active\" git users do it at most once or twice a day, \nwhich means that their client cache will _never_ be up-to-date.\n\nBut if you do it with server-side caches and grace-periods, you can \ngenerally say \"we have something that is at most five minutes old\", and \nmost importantly, you can hopefully do it without a lot of disk seeks \n(because you just cache the _one_ page as _one_ object), so hopefully you \ncan do it in a few hundred ms even if the thing is on disk and even if \nthere's a lot of other load going on.\n\nI bet the top-level \"all projects\" summary page and the individual \nproject summary pages are the important things to cache. That's what \nprobably most people look at, and they are the ones that have lots of \nserver-side cache locality. Individual commits and diffs probably don't \nget the same kind of \"lots of people looking at them\" and thus don't get \nthe same kind of benefit from caching.\n\n(Individual commits hopefully also need fewer disk seeks, at least with \npacked repositories. So even if you have to re-generate them from scratch, \nthey won't have the seek times themselves taking up tens of seconds, \nunless the project is entirely unpacked and diffing just generates total \ndisk seek hell)\n\n"},{"id":"295151","messageId":"200612102127.05894.jnareb@gmail.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612101129190.12500@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-10T20:27:05Z","receivedAt":"2006-12-10T20:27:05Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n> On Sun, 10 Dec 2006, Jakub Narebski wrote:\n>>>> If-Modified-Since:, If-Match:, If-None-Match: do you?\n>> \n>> And in CGI standard there is a way to access additional HTTP headers\n>> info from CGI script: the envirionmental variables are HTTP_HEADER,\n>> for example if browser sent If-Modified-Since: header it's value\n>> can be found in HTTP_IF_MODIFIED_SINCE environmental variable.\n> \n> Guys, you're missing something fairly fundamnetal. \n> \n> It helps almost _nothing_ to support client-side caching with all these \n> fancy \"If-Modified-Since:\" etc crap.\n> \n> That's not the _problem_.\n> \n> It's usually not one client asking for the gitweb pages: the load comes \n> from just lots of people independently asking for it. So client-side \n> caching may help a tiny tiny bit, but it's not actually fixing the \n> fundamental problem at all.\n\nWell, the idea (perhaps stupid idea: I don't know how caching engines\n/ reverse proxy works) was that there would be caching engine / reverse\nproxy in the front (Squid for example) would cache results and serve it\nto rampaging hordes. But this caching engine has to ask gitweb if the\ncache is valid using \"If-Modified-Since:\" and \"If-None-Match:\" headers.\nIf gitweb returns 304 Not Modified then it serves contents from cache.\n\n> So forget about \"If-Modified-Since:\" etc. It may help in benchmarks when \n> you try it yourself, and use \"refresh\" on the client side. But the basic \n> problem is all about lots of clients that do NOT have things cached, \n> because all teh client caches are all filled up with pr0n, not with gitweb \n> data from yesterday.\n\nWhat about the other idea, the one with raising expires to infinity for\nimmutable pages like \"commit\" view for commit given by SHA-1? Even if\nthe clients won't cache it, the proxies and caches between gitweb and\nclient might cache it...\n\nTalking about most accessed gitweb pages, the project list page changes\non every push, the project summary page and project main RSS feed\n(now in both RSS and Atom formats) changes on every push to given project.\nWith a help of hooks they can be static pages, generated by push...\n...with the exception that projects list and summary pages have _relative_\ndates.\n\n-- \nJakub Narebski\n"},{"id":"296863","messageId":"Pine.LNX.4.64.0612101228590.12500@woody.osdl.org","threadId":"43081","inReplyTo":"200612102127.05894.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-10T20:30:49Z","receivedAt":"2006-12-10T20:30:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 10 Dec 2006, Jakub Narebski wrote:\n>\n> Well, the idea (perhaps stupid idea: I don't know how caching engines\n> / reverse proxy works) was that there would be caching engine / reverse\n> proxy in the front (Squid for example) would cache results and serve it\n> to rampaging hordes.\n\nSure, if the proxies actually do the rigth thing (which they may or may \nnot do)\n\n> What about the other idea, the one with raising expires to infinity for\n> immutable pages like \"commit\" view for commit given by SHA-1? Even if\n> the clients won't cache it, the proxies and caches between gitweb and\n> client might cache it...\n\nI agree, but as mentioned, I think the _real_ problem tends to be the \npages that don't act that way (ie summary pages, both at the individual \nproject level and the top \"all projects\" level).\n\n"},{"id":"296115","messageId":"457C75BE.1010805@zytor.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612101129190.12500@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2006-12-10T21:01:50Z","receivedAt":"2006-12-10T21:01:50Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> Now, hopefully some of it will be in the disk cache, but when the \n> mirroring happens, it will basically blow the disk caches away totally \n> (when using the \"--checksum\" option), and then you literally have tens of \n> seconds to generate that one top-level page. \n> \n\nIf that was the only time that happened, it would be a non-issue, since \nthat only happens once every 96 hours.  However, the problem is that we \nnow have lots of large datasets that blow out the caches on a much more \nfrequent basis.\n\n"},{"id":"294292","messageId":"46a038f90612101401m5f65aefbh78f7adf84725ade4@mail.gmail.com","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612101228590.12500@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-10T22:01:48Z","receivedAt":"2006-12-10T22:01:48Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/11/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Sure, if the proxies actually do the rigth thing (which they may or may\n> not do)\n\nFor a high-traffic setup like kernel.org, you can setup a local\nreverse proxy -- it's a pretty standard practice. That allows you to\ncontrol a well-behaved and locally tuned caching engine just by\nemitting good headers.\n\nIt beats writing and maintaining an internal caching mechanism for\neach CGI script out there by a long mile. It means there'll be no\nfurther tunables or complexity for administrators of other gitweb\ninstalls.\n\ncheers,\n\n\n\n"},{"id":"295695","messageId":"457C84AC.7060105@garzik.org","threadId":"43081","inReplyTo":"200612102011.52589.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-10T22:05:32Z","receivedAt":"2006-12-10T22:05:32Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Jakub Narebski wrote:\n> Adn in CGI standard there is a way to access additional HTTP headers\n> info from CGI script: the envirionmental variables are HTTP_HEADER,\n> for example if browser sent If-Modified-Since: header it's value\n> can be found in HTTP_IF_MODIFIED_SINCE environmental variable.\n\nThe CGI spec does not at all guarantee that the CGI environment will \ncontain all the HTTP headers sent by the client.  That was the point of \nthe environment dump script -- you can see exactly which headers are, \nand are not, passed through to CGI.\n\nCGI only /guarantees/ a bare minimum (things like QUERY_STRING, \nPATH_INFO, etc.)\n\nEven basic server info environment variables are optional.\n\n\n> It is ETag, not E-tag. Besides, I don't see what the attached script is\n> meant to do: it does not output the sample file anyway.\n\nIt's not meant to output the sample file.  It outputs the server \nmetadata sent to the CGI script (the environment variables).  The sample \nfile was simply a way to play around with etag and last-modified metadata.\n\n\n> The idea is of course to stop processing in CGI script / mod_perl script\n> as soon as possible if cache validates.\n\nCertainly.  That should help cut down on I/O.  FWIW though the projects \nlist is particularly painful, with its File::Find call, which you'll \nneed to do in order to return 304-not-modified.\n\n\n> I don't know if Apache intercepts and remembers ETag and Last-Modified\n> headers, adds 304 Not Modified HTTP response on finding that cache validates\n> and cuts out CGI script output. I.e. if browser provided If-Modified-Since:,\n> script wrote Last-Modified: header, If-Modified-Since: is no earlier than\n> Last-Modified: (usually is equal in the case of cache validation), then\n> Apache provides 304 Not Modified response instead of CGI script output.\n\nThis wanders into the realm of mod_cache configuration, I think.  (which \nI have tried to get working as reverse proxy, and failed serveral times) \n  If you are not using mod_*_cache, then Apache must execute the CGI \nscript every time AFAICS, regardless of etag/[if-]last-mod headers.\n\n\tJeff\n\n\n"},{"id":"294446","messageId":"457C8579.5080407@garzik.org","threadId":"43081","inReplyTo":"Pine.LNX.4.64.0612101228590.12500@woody.osdl.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-10T22:08:57Z","receivedAt":"2006-12-10T22:08:57Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Sun, 10 Dec 2006, Jakub Narebski wrote:\n>> Well, the idea (perhaps stupid idea: I don't know how caching engines\n>> / reverse proxy works) was that there would be caching engine / reverse\n>> proxy in the front (Squid for example) would cache results and serve it\n>> to rampaging hordes.\n> \n> Sure, if the proxies actually do the rigth thing (which they may or may \n> not do)\n\nsquid seems to work well as an HTTP accelerator (reverse proxy). \nApache's mem|disk cache stuff fails miserably.\n\nUnfortunately squid development seems to have slowed in recent years.\n\n\tJeff\n\n\n"},{"id":"293842","messageId":"457C86C7.2070000@garzik.org","threadId":"43081","inReplyTo":"46a038f90612101401m5f65aefbh78f7adf84725ade4@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-12-10T22:14:31Z","receivedAt":"2006-12-10T22:14:31Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 12/11/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>> Sure, if the proxies actually do the rigth thing (which they may or may\n>> not do)\n> \n> For a high-traffic setup like kernel.org, you can setup a local\n> reverse proxy -- it's a pretty standard practice. That allows you to\n> control a well-behaved and locally tuned caching engine just by\n> emitting good headers.\n> \n> It beats writing and maintaining an internal caching mechanism for\n> each CGI script out there by a long mile. It means there'll be no\n> further tunables or complexity for administrators of other gitweb\n> installs.\n\nIf gitweb produced cache-friendly headers, squid could definitely serve \nas an HTTP front-end (\"HTTP accelerator\" mode in squid talk).\n\nIn fact, given kernel.org's slave1/slave2<->master setup, that's a \npretty natural fit for caching files and/or cache-aware CGI output.\n\nYou could even replace rsync to the slaves, if squid was serving as the \nfront-end accelerator running on the slaves, communicating to the master.\n\nsquid is smart enough to hold off a thundering herd, and only pulls \nsingle cacheable copies of files as needed.\n\n\tJeff\n\n"},{"id":"297699","messageId":"200612102359.20083.jnareb@gmail.com","threadId":"43081","inReplyTo":"457C84AC.7060105@garzik.org","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-10T22:59:19Z","receivedAt":"2006-12-10T22:59:19Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jeff Garzik wrote:\n> Jakub Narebski wrote:\n>>\n>> And in CGI standard there is a way to access additional HTTP headers\n>> info from CGI script: the envirionmental variables are HTTP_HEADER,\n>> for example if browser sent If-Modified-Since: header it's value\n>> can be found in HTTP_IF_MODIFIED_SINCE environmental variable.\n> \n> The CGI spec does not at all guarantee that the CGI environment will \n> contain all the HTTP headers sent by the client.  That was the point of \n> the environment dump script -- you can see exactly which headers are, \n> and are not, passed through to CGI.\n> \n> CGI only /guarantees/ a bare minimum (things like QUERY_STRING, \n> PATH_INFO, etc.)\n> \n> Even basic server info environment variables are optional.\n\nI have checked that at least Apache 2.0.54 passes HTTP_IF_MODIFIED_SINCE\nwhen getting If-Modified-Since: header (my own script + netcat/nc).\n \n>> It is ETag, not E-tag. Besides, I don't see what the attached script is\n>> meant to do: it does not output the sample file anyway.\n> \n> It's not meant to output the sample file.  It outputs the server \n> metadata sent to the CGI script (the environment variables).  The sample \n> file was simply a way to play around with etag and last-modified metadata.\n\nAh. \n \n>> The idea is of course to stop processing in CGI script / mod_perl script\n>> as soon as possible if cache validates.\n> \n> Certainly.  That should help cut down on I/O.  FWIW though the projects \n> list is particularly painful, with its File::Find call, which you'll \n> need to do in order to return 304-not-modified.\n\nFirst, it is better to use $projects_list which is projects index file\nin the format:\n  <project path> SPC <project owner>\nwhere <project path> is relative to $projectroot and is URI encoded; well\nat least SPC has to be URI (percent) encoded. <project owner> is owner\nof given project, and is also URI encoded (one would usually use '+' in\nthe place of SPC here).\n\nGitweb now can generate projects list in above format, by using\n\"project_index\" action (\"a=project_index\" query string), or by clicking\n'TXT' link at the bottom of the projects list page in new gitweb: see\nhttp://repo.or.cz by Petr Baudis. The problem is that it generates\nprojects list from the list of projects it sees, so to generate it from\nscratch from the filesystem you have for generating \"project_index\"\nto have $projects_list a directory (changing it to something that\nevals to false, e.g. undef or \"\" makes gitweb use $projectroot for\n$projects_list). I have posted how to do this.\n\nThe project list changes rarely, only on addition/removal of project,\nand on changing owner of project; so it can be generated on demand.\n\n\nSecond, even with $projects_list being set to projects index file\nas of now gitweb runs git-for-each-ref (which scans refs and access\npack file for commit date), checks for description file and reads it;\nfor $projects_list being directory it also checks project directory\nowner. I plan to make it configurable to read last activity from\nall heads (all branches) as it is now, from HEAD (current branch)\nas it was before, or given branch (for example 'master').\n\nAssuming that gitweb is configured to read last activity from single\ndefined branch, generating ETag = checksum(sha1 of heads of projects)\nneeds at least read one file from each project.\n \n>> I don't know if Apache intercepts and remembers ETag and Last-Modified\n>> headers, adds 304 Not Modified HTTP response on finding that cache validates\n>> and cuts out CGI script output. I.e. if browser provided If-Modified-Since:,\n>> script wrote Last-Modified: header, If-Modified-Since: is no earlier than\n>> Last-Modified: (usually is equal in the case of cache validation), then\n>> Apache provides 304 Not Modified response instead of CGI script output.\n> \n> This wanders into the realm of mod_cache configuration, I think.  (which \n> I have tried to get working as reverse proxy, and failed serveral times) \n>   If you are not using mod_*_cache, then Apache must execute the CGI \n> script every time AFAICS, regardless of etag/[if-]last-mod headers.\n\nNo, it wanders into realm of header parsing by Apache, and NPH (No Parse\nHeaders) option.\n\nEven if Apache does execute CGI script to completion every time, it might\nnot send the output of the script, but HTTP 304 Not Modified reply. Might.\nI don't know if it does.\n\n-- \nJakub Narebski\n"},{"id":"295343","messageId":"46a038f90612101816j33870bb1j39182358440aaa40@mail.gmail.com","threadId":"43081","inReplyTo":"200612102359.20083.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-11T02:16:23Z","receivedAt":"2006-12-11T02:16:23Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/11/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> Even if Apache does execute CGI script to completion every time, it might\n> not send the output of the script, but HTTP 304 Not Modified reply. Might.\n> I don't know if it does.\n\nIt is up to the script (CGI or via mod_perl) to set the status to 304\nand finish execution. Just setting the status to 304 does not\nforcefully end execution as you may want to cleanup, log, etc.\n\ncheers,\n\n\n"},{"id":"295029","messageId":"200612110959.56492.jnareb@gmail.com","threadId":"43081","inReplyTo":"46a038f90612101816j33870bb1j39182358440aaa40@mail.gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-11T08:59:55Z","receivedAt":"2006-12-11T08:59:55Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Langhoff wrote:\n> On 12/11/06, Jakub Narebski <jnareb@gmail.com> wrote:\n>>\n>> Even if Apache does execute CGI script to completion every time, it might\n>> not send the output of the script, but HTTP 304 Not Modified reply. Might.\n>> I don't know if it does.\n> \n> It is up to the script (CGI or via mod_perl) to set the status to 304\n> and finish execution. Just setting the status to 304 does not\n> forcefully end execution as you may want to cleanup, log, etc.\n\nI was thinking not about ending execution, but about not sending script\noutput but sending HTTP 304 Not Modified reply by Apache.\n\nI meant the following sequence of events:\n 1. Script sends headers, among those Last-Modified and/or ETag\n 2. Apache scans headers (e.g. to add its own), notices that Last-Modified\n    is earlier or equal to If-Modified-Since: sent by browser or reverse\n    proxy, or ETag matches If-None-Match:, and sends 304 instead of script\n    output\n 3. Script finishes execution, it's output sent to /dev/null\n\nAgain, I don't know if Apache (or any other web server) does that. \n-- \nJakub Narebski\n"},{"id":"294113","messageId":"46a038f90612110218u48b7737due56437da57091547@mail.gmail.com","threadId":"43081","inReplyTo":"200612110959.56492.jnareb@gmail.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-12-11T10:18:54Z","receivedAt":"2006-12-11T10:18:54Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 12/11/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> I was thinking not about ending execution, but about not sending script\n> output but sending HTTP 304 Not Modified reply by Apache.\n>\n> I meant the following sequence of events:\n>  1. Script sends headers, among those Last-Modified and/or ETag\n>  2. Apache scans headers (e.g. to add its own), notices that Last-Modified\n>     is earlier or equal to If-Modified-Since: sent by browser or reverse\n>     proxy, or ETag matches If-None-Match:, and sends 304 instead of script\n>     output\n>  3. Script finishes execution, it's output sent to /dev/null\n>\n> Again, I don't know if Apache (or any other web server) does that.\n\nIt doesn't. You want to take the decision to send a 304, cleanup and\nexit _inside_ the CGI. If it was up to apache, then the CGI script\nwould end up creating the (potentially expensive to produce) content\njust to see it sent to /dev/null OR if apache was to terminate\nexecution of the CGI more violently, the CGI wouldn't have a chance to\ncleanup and release resources.\n\nSo it's a matter of setting the header to 304 and exiting.\n\ncheers,\n\n\nmartin\n\n"},{"id":"295166","messageId":"200612122219.30040.jnareb@gmail.com","threadId":"43081","inReplyTo":"457BB1A3.2070408@zytor.com","subject":"Re: kernel.org mirroring (Re: [GIT PULL] MMC update)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-12T21:19:28Z","receivedAt":"2006-12-12T21:19:28Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"By the way, setting Last-Modified: and ETag: and checking for \nIf-Modified-Since: and If-None-Match: is easy only for log-like views: \n\"shortlog\", \"log\", \"history\", \"rss\"/\"atom\". With \"shortlog\" and \n\"history\" we have additional difficulity of using relative dates there.\nAnd even for those views we need reverse proxy / caching engine\n(e.g. Squid in \"HTTP accelerator\" mode) in front.\n\nIt would be easier to pre-generate most common accessed views: \n\"projects_list\", \"summary\" and \"rss\"/\"atom\" main for each project, and \njust serve static pages. I don't know if we need to modify gitweb for \nthat.\n\n\nBTW. for single client (rather stupid benchmark, I know) mod_perl is \nabout twice faster in keepalive mode than CGI version of gitweb for \ngit.git summary page.\n\n-- \nJakub Narebski\n"}]}