{"thread":{"id":"12677","subject":"[PATCH] gitweb: Support caching projects list","startedAt":"2008-03-13T23:14:14Z","lastAt":"2008-03-17T21:09:01Z","messageCount":30,"participants":["Petr Baudis","Jay Soffian","Junio C Hamano","J.H.","Frank Lichtenheld","Jakub Narebski","Miklos Vajna","Theodore Tso"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"72007","messageId":"20080313231413.27966.3383.stgit@rover","threadId":"12677","inReplyTo":null,"subject":"[PATCH] gitweb: Support caching projects list","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-13T23:14:14Z","receivedAt":"2008-03-13T23:14:14Z","isPatch":true,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On repo.or.cz (permanently I/O overloaded and hosting 1050 project +\nforks), the projects list (the default gitweb page) can take more than\na minute to generate. This naive patch adds simple support for caching\nthe projects list data structure so that all the projects do not need\nto get rescanned at every page access.\n\n$projlist_cache_lifetime gitweb configuration variable is introduced,\nby default set to zero. If set to non-zero, it describes the number of\nminutes for which the cache remains valid. Only single project root\nper system can use the cache. Any script running with the same uid\nas gitweb can change the cache trivially - this is for secure installations\nonly.\n\nThe cache itself is stored in /tmp/gitweb.index.cache as a Data::Dumper\ndump of the perl data structure with the list of project details. When\nreusing the cache, the file is simply eval'd back into @projects. For\nclarity, projects scanning and @projects population is separated to\ngit_get_projects_details().\n\nTo prevent contention when multiple accesses coincide with cache\nexpiration, the timeout is postponed to time()+120 when we start\nrefreshing. When showing cached version, a disclaimer is shown at the\ntop of the projects list.\n\nSigned-off-by: Petr Baudis <pasky@suse.cz>\n---\n\n gitweb/gitweb.css  |    6 +++++\n gitweb/gitweb.perl |   59 ++++++++++++++++++++++++++++++++++++++++++++++++----\n 2 files changed, 60 insertions(+), 5 deletions(-)\n\ndiff --git a/gitweb/gitweb.css b/gitweb/gitweb.css\nindex 8e2bf3d..673077a 100644\n--- a/gitweb/gitweb.css\n+++ b/gitweb/gitweb.css\n@@ -85,6 +85,12 @@ div.title, a.title {\n \tcolor: #000000;\n }\n \n+div.stale_info {\n+\tdisplay: block;\n+\ttext-align: right;\n+\tfont-style: italic;\n+}\n+\n div.readme {\n \tpadding: 8px;\n }\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex bcb6193..0eee195 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -122,6 +122,15 @@ our $fallback_encoding = 'latin1';\n # - one might want to include '-B' option, e.g. '-B', '-M'\n our @diff_opts = ('-M'); # taken from git_commit\n \n+# projects list cache for busy sites with many projects;\n+# if you set this to non-zero, it will be used as the cached\n+# index lifetime in minutes\n+# the cached list version is stored in /tmp and can be tweaked\n+# by other scripts running with the same uid as gitweb - use this\n+# only at secure installations; only single gitweb project root per\n+# system is supported!\n+our $projlist_cache_lifetime = 0;\n+\n # information about snapshot formats that gitweb is capable of serving\n our %known_snapshot_formats = (\n \t# name => {\n@@ -3507,10 +3516,8 @@ sub git_patchset_body {\n \n # . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .\n \n-sub git_project_list_body {\n-\tmy ($projlist, $order, $from, $to, $extra, $no_header) = @_;\n-\n-\tmy ($check_forks) = gitweb_check_feature('forks');\n+sub git_get_projects_details {\n+\tmy ($projlist, $check_forks) = @_;\n \n \tmy @projects;\n \tforeach my $pr (@$projlist) {\n@@ -3540,11 +3547,53 @@ sub git_project_list_body {\n \t\t}\n \t\tpush @projects, $pr;\n \t}\n+\treturn @projects;\n+}\n+\n+sub git_project_list_body {\n+\tmy ($projlist, $order, $from, $to, $extra, $no_header, $cache_lifetime) = @_;\n+\n+\tmy ($check_forks) = gitweb_check_feature('forks');\n+\n+\tmy $cache_file = '/tmp/gitweb.index.cache';\n+\tuse File::stat;\n+\n+\tmy @projects;\n+\tmy $stale = 0;\n+\tif ($cache_lifetime and -f $cache_file\n+\t    and stat($cache_file)->mtime + $cache_lifetime * 60 > time()\n+\t    and open (my $fd, $cache_file)) {\n+\t\t$stale = time() - stat($cache_file)->mtime;\n+\t\tmy @dump = <$fd>;\n+\t\tclose $fd;\n+\t\t# Hack zone start\n+\t\tmy $VAR1;\n+\t\teval join(\"\\n\", @dump);\n+\t\t@projects = @$VAR1;\n+\t\t# Hack zone end\n+\t} else {\n+\t\tif ($cache_lifetime and -f $cache_file) {\n+\t\t\t# Postpone timeout by two minutes so that we get\n+\t\t\t# enough time to do our job.\n+\t\t\tmy $time = time() - $cache_lifetime + 120;\n+\t\t\tutime $time, $time, $cache_file;\n+\t\t}\n+\t\t@projects = git_get_projects_details($projlist, $check_forks);\n+\t\tif ($cache_lifetime and open (my $fd, '>'.$cache_file)) {\n+\t\t\tuse Data::Dumper;\n+\t\t\tprint $fd Dumper(\\@projects);\n+\t\t\tclose $fd;\n+\t\t}\n+\t}\n \n \t$order ||= $default_projects_order;\n \t$from = 0 unless defined $from;\n \t$to = $#projects if (!defined $to || $#projects < $to);\n \n+\tif ($cache_lifetime and $stale) {\n+\t\tprint \"<div class=\\\"stale_info\\\">Cached version (${stale}s old)</div>\\n\";\n+\t}\n+\n \tprint \"<table class=\\\"project_list\\\">\\n\";\n \tunless ($no_header) {\n \t\tprint \"<tr>\\n\";\n@@ -3927,7 +3976,7 @@ sub git_project_list {\n \t\tclose $fd;\n \t\tprint \"</div>\\n\";\n \t}\n-\tgit_project_list_body(\\@list, $order);\n+\tgit_project_list_body(\\@list, $order, undef, undef, undef, undef, $projlist_cache_lifetime);\n \tgit_footer_html();\n }\n \n"},{"id":"72008","messageId":"76718490803131707g34fd40d4q21c69391c2597bc@mail.gmail.com","threadId":"12677","inReplyTo":"20080313231413.27966.3383.stgit@rover","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2008-03-14T00:07:09Z","receivedAt":"2008-03-14T00:07:09Z","isPatch":true,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On Thu, Mar 13, 2008 at 7:14 PM, Petr Baudis <pasky@suse.cz> wrote:\n>  diff --git a/gitweb/gitweb.css b/gitweb/gitweb.css\n>  index 8e2bf3d..673077a 100644\n>  --- a/gitweb/gitweb.css\n>  +++ b/gitweb/gitweb.css\n>  @@ -85,6 +85,12 @@ div.title, a.title {\n>         color: #000000;\n>   }\n>\n>  +div.stale_info {\n>  +       display: block;\n>  +       text-align: right;\n>  +       font-style: italic;\n>  +}\n>  +\n>   div.readme {\n>         padding: 8px;\n>   }\n\nWhat does this have to do with it?\n\n>  diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n>  index bcb6193..0eee195 100755\n>  --- a/gitweb/gitweb.perl\n>  +++ b/gitweb/gitweb.perl\n>  @@ -122,6 +122,15 @@ our $fallback_encoding = 'latin1';\n\n...\n\n>  +               if ($cache_lifetime and -f $cache_file) {\n>  +                       # Postpone timeout by two minutes so that we get\n>  +                       # enough time to do our job.\n>  +                       my $time = time() - $cache_lifetime + 120;\n>  +                       utime $time, $time, $cache_file;\n>  +               }\n\nRace condition. I don't see any locking. Nothing keeps multiple instances from\nregenerating the cache concurrently...\n\n>  +               @projects = git_get_projects_details($projlist, $check_forks);\n>  +               if ($cache_lifetime and open (my $fd, '>'.$cache_file)) {\n\n...and then clobbering each other here. You have two choices:\n\n1) Use a lock file for the critical section.\n\n2) Assume the race condition is rare enough, but you still need to account for\nit. In that case, you want to write to a temporary file and then rename to the\ncache file name. The rename is atomic, so though N instances of gitweb may\nregenerate the cache (at some CPU/IO overhead), at least the cache file won't\nget corrupted.\n\nOut of curiosity, repo.or.cz isn't running this as a CGI is it? If so, wouldn't\nrunning it as a FastCGI or modperl be a vast improvement?\n\nj.\n"},{"id":"72012","messageId":"7v63vqw40m.fsf@gitster.siamese.dyndns.org","threadId":"12677","inReplyTo":"20080313231413.27966.3383.stgit@rover","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-03-14T00:19:53Z","receivedAt":"2008-03-14T00:19:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Petr Baudis <pasky@suse.cz> writes:\n\n> To prevent contention when multiple accesses coincide with cache\n> expiration, the timeout is postponed to time()+120 when we start\n> refreshing. When showing cached version, a disclaimer is shown at the\n> top of the projects list.\n\nIsn't this still racy when two requests come at about the same time?\nPerhaps you can avoid it by using a lockfile?\n"},{"id":"72015","messageId":"20080314002205.GL10335@machine.or.cz","threadId":"12677","inReplyTo":"76718490803131707g34fd40d4q21c69391c2597bc@mail.gmail.com","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-14T00:22:05Z","receivedAt":"2008-03-14T00:22:05Z","isPatch":true,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Thu, Mar 13, 2008 at 08:07:09PM -0400, Jay Soffian wrote:\n> On Thu, Mar 13, 2008 at 7:14 PM, Petr Baudis <pasky@suse.cz> wrote:\n> >  diff --git a/gitweb/gitweb.css b/gitweb/gitweb.css\n> >  index 8e2bf3d..673077a 100644\n> >  --- a/gitweb/gitweb.css\n> >  +++ b/gitweb/gitweb.css\n> >  @@ -85,6 +85,12 @@ div.title, a.title {\n> >         color: #000000;\n> >   }\n> >\n> >  +div.stale_info {\n> >  +       display: block;\n> >  +       text-align: right;\n> >  +       font-style: italic;\n> >  +}\n> >  +\n> >   div.readme {\n> >         padding: 8px;\n> >   }\n> \n> What does this have to do with it?\n\nThe box shows that cached information is being shown.\n\n> >  diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> >  index bcb6193..0eee195 100755\n> >  --- a/gitweb/gitweb.perl\n> >  +++ b/gitweb/gitweb.perl\n> >  @@ -122,6 +122,15 @@ our $fallback_encoding = 'latin1';\n> \n> ...\n> \n> >  +               if ($cache_lifetime and -f $cache_file) {\n> >  +                       # Postpone timeout by two minutes so that we get\n> >  +                       # enough time to do our job.\n> >  +                       my $time = time() - $cache_lifetime + 120;\n> >  +                       utime $time, $time, $cache_file;\n> >  +               }\n> \n> Race condition. I don't see any locking. Nothing keeps multiple instances from\n> regenerating the cache concurrently...\n> \n> >  +               @projects = git_get_projects_details($projlist, $check_forks);\n> >  +               if ($cache_lifetime and open (my $fd, '>'.$cache_file)) {\n> \n> ...and then clobbering each other here. You have two choices:\n> \n> 1) Use a lock file for the critical section.\n> \n> 2) Assume the race condition is rare enough, but you still need to account for\n> it. In that case, you want to write to a temporary file and then rename to the\n> cache file name. The rename is atomic, so though N instances of gitweb may\n> regenerate the cache (at some CPU/IO overhead), at least the cache file won't\n> get corrupted.\n\nYou are of course right - I wanted to do the rename, but forgot to write\nit in the actual code. :-)\n\nThere is a more conceptual problem though - in case of such big sites,\nit really makes more sense to explicitly regenerate the cache\nperiodically instead of making random clients to have to wait it out.\nWe could add a 'force_update' parameter to accept from localhost only\nthat will always regenerate the cache, but that feels rather kludgy -\ncan anyone think of a more elegant solution? (I don't think taking the\n@projects generating code out of gitweb and then having to worry during\ngitweb upgrades is any better.)\n\n> Out of curiosity, repo.or.cz isn't running this as a CGI is it? If so, wouldn't\n> running it as a FastCGI or modperl be a vast improvement?\n\nUnlikely. Currently the machine is mostly IO-bound and only small\nportion of CPU usage comes from gitweb itself.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"72013","messageId":"76718490803131727p451967hee96ff26206c97b7@mail.gmail.com","threadId":"12677","inReplyTo":"20080314002205.GL10335@machine.or.cz","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2008-03-14T00:27:34Z","receivedAt":"2008-03-14T00:27:34Z","isPatch":true,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On Thu, Mar 13, 2008 at 8:22 PM, Petr Baudis <pasky@suse.cz> wrote:\n\n>  There is a more conceptual problem though - in case of such big sites,\n>  it really makes more sense to explicitly regenerate the cache\n>  periodically instead of making random clients to have to wait it out.\n\nFork off a child to update the cache?\n\n>  Unlikely. Currently the machine is mostly IO-bound and only small\n>  portion of CPU usage comes from gitweb itself.\n\nExcept that if it were FastCGI or mod_perl you could just keep the cache in\nmemory.\n\nj.\n"},{"id":"72014","messageId":"1205454601.2758.10.camel@localhost.localdomain","threadId":"12677","inReplyTo":"76718490803131727p451967hee96ff26206c97b7@mail.gmail.com","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"J.H.","fromEmail":"warthog19@eaglescrag.net","sentAt":"2008-03-14T00:30:01Z","receivedAt":"2008-03-14T00:30:01Z","isPatch":true,"sender":{"key":"warthog19@eaglescrag.net","avatar":null},"body":"You would be better off using some of the logic I've got in the caching\nversion of gitweb to prevent the race condition.\n\n- John 'Warthog9' Hawley\n\nOn Thu, 2008-03-13 at 20:27 -0400, Jay Soffian wrote:\n> On Thu, Mar 13, 2008 at 8:22 PM, Petr Baudis <pasky@suse.cz> wrote:\n> \n> >  There is a more conceptual problem though - in case of such big sites,\n> >  it really makes more sense to explicitly regenerate the cache\n> >  periodically instead of making random clients to have to wait it out.\n> \n> Fork off a child to update the cache?\n> \n> >  Unlikely. Currently the machine is mostly IO-bound and only small\n> >  portion of CPU usage comes from gitweb itself.\n> \n> Except that if it were FastCGI or mod_perl you could just keep the cache in\n> memory.\n> \n> j.\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"72016","messageId":"1205454999.2758.14.camel@localhost.localdomain","threadId":"12677","inReplyTo":"20080314002205.GL10335@machine.or.cz","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"J.H.","fromEmail":"warthog19@eaglescrag.net","sentAt":"2008-03-14T00:36:39Z","receivedAt":"2008-03-14T00:36:39Z","isPatch":true,"sender":{"key":"warthog19@eaglescrag.net","avatar":null},"body":"\n> You are of course right - I wanted to do the rename, but forgot to write\n> it in the actual code. :-)\n> \n> There is a more conceptual problem though - in case of such big sites,\n> it really makes more sense to explicitly regenerate the cache\n> periodically instead of making random clients to have to wait it out.\n> We could add a 'force_update' parameter to accept from localhost only\n> that will always regenerate the cache, but that feels rather kludgy -\n> can anyone think of a more elegant solution? (I don't think taking the\n> @projects generating code out of gitweb and then having to worry during\n> gitweb upgrades is any better.)\n\nYou could do something similar to the gitweb caching I'm doing,\nbasically if a file isn't generated you make a user wait (no good way\naround this really).  If a cache exists show it to the user unless the\ncache is older than $foo.  If a re-generation needs to happen it happens\nin the background so the user who triggers the regeneration sees\nsomething immediately vs. having to wait (at the cost of showing out of\ndate data)\n\n> > Out of curiosity, repo.or.cz isn't running this as a CGI is it? If so, wouldn't\n> > running it as a FastCGI or modperl be a vast improvement?\n> \n> Unlikely. Currently the machine is mostly IO-bound and only small\n> portion of CPU usage comes from gitweb itself.\n\nThats about the same as what I saw, it's disk bound vs. cpu/memory\nbound.\n\n- John 'Warthog9' Hawley\n"},{"id":"72047","messageId":"20080314083515.GX10103@mail-vs.djpig.de","threadId":"12677","inReplyTo":"20080313231413.27966.3383.stgit@rover","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Frank Lichtenheld","fromEmail":"frank@lichtenheld.de","sentAt":"2008-03-14T08:35:15Z","receivedAt":"2008-03-14T08:35:15Z","isPatch":true,"sender":{"key":"frank@lichtenheld.de","avatar":"https://gravatar.com/avatar/b9f1d4b120e138f157c9e480d0818197c474628923786adb98f30017cdb99c3c?d=mp&s=160"},"body":"On Fri, Mar 14, 2008 at 12:14:14AM +0100, Petr Baudis wrote:\n> +# projects list cache for busy sites with many projects;\n> +# if you set this to non-zero, it will be used as the cached\n> +# index lifetime in minutes\n> +# the cached list version is stored in /tmp and can be tweaked\n> +# by other scripts running with the same uid as gitweb - use this\n> +# only at secure installations; only single gitweb project root per\n> +# system is supported!\n> +our $projlist_cache_lifetime = 0;\n\nI think that would a situation where a uppercase disclaimer would be\nappropriate ;)\n\nIn addition to the race condition problem mentioned in other mails it\nalso has a symlink vulnerability.\n\nI think one should seriously consider reusing an existing caching\nsolution instead of reinventing the wheel here.\nThere are a lot of CPAN modules to do that and at least apache also\nhas modules for that which you could use without any code changes\nat all...\n\nGruesse,\n-- \nFrank Lichtenheld <frank@lichtenheld.de>\nwww: http://www.djpig.de/\n"},{"id":"72067","messageId":"m3hcf9y02p.fsf@localhost.localdomain","threadId":"12677","inReplyTo":"20080313231413.27966.3383.stgit@rover","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-14T12:14:51Z","receivedAt":"2008-03-14T12:14:51Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Petr Baudis <pasky@suse.cz> writes:\n\n> On repo.or.cz (permanently I/O overloaded and hosting 1050 project +\n> forks), the projects list (the default gitweb page) can take more than\n> a minute to generate. This naive patch adds simple support for caching\n> the projects list data structure so that all the projects do not need\n> to get rescanned at every page access.\n\nNice.\n\nBTW adding caching to gitweb is one of proposed ideas (projects) for\nGoogle Summer of Code 2006: http://git.or.cz/gitwiki/SoC2008Ideas\n\n> For clarity, projects scanning and @projects population is separated\n> to git_get_projects_details().\n\nPerhaps this could be submitted as separate patch?\nI could do this if you are otherwise busy...\n\n\n[...]\n> +\tif ($cache_lifetime and -f $cache_file\n> +\t    and stat($cache_file)->mtime + $cache_lifetime * 60 > time()\n> +\t    and open (my $fd, $cache_file)) {\n> +\t\t$stale = time() - stat($cache_file)->mtime;\n> +\t\tmy @dump = <$fd>;\n> +\t\tclose $fd;\n> +\t\t# Hack zone start\n> +\t\tmy $VAR1;\n> +\t\teval join(\"\\n\", @dump);\n> +\t\t@projects = @$VAR1;\n> +\t\t# Hack zone end\n\nWhy do you read line by line, only to join it, i.e.\n  my @dump = <$fd>; ... join(\"\\n\", @dump);\ninstead of slurping all file in one go:\n  local $/ = undef; my $dump = <$fd>; ... $dump;\n\nBesides, why do you use Data::Dumper instead of Storable? Both are\ndistributed with Perl; well, at least both are in perl-5.8.6-24.\n\n[...]\n> -\tgit_project_list_body(\\@list, $order);\n> +\tgit_project_list_body(\\@list, $order, undef, undef, undef, undef, $projlist_cache_lifetime);\n\nThis is ugly. Why not use hash for \"named parameters\", as it is done\nin a few separate places in gitweb (search for '%opts')?\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"72069","messageId":"m3d4pxxzyp.fsf@localhost.localdomain","threadId":"12677","inReplyTo":"1205454601.2758.10.camel@localhost.localdomain","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-14T12:17:17Z","receivedAt":"2008-03-14T12:17:17Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"J.H.\" <warthog19@eaglescrag.net> writes:\n\n> You would be better off using some of the logic I've got in the caching\n> version of gitweb to prevent the race condition.\n\nOr use Cache::FileCache the CGI::Cache uses...\n\n\nBy the way, J.H., would you have time and be interested in becoming\n\"Gitweb caching\" project mentor for Google Summer of Code 2008:\n\n  http://git.or.cz/gitwiki/SoC2008Ideas\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"72101","messageId":"m38x0lxr1k.fsf@localhost.localdomain","threadId":"12677","inReplyTo":"76718490803131707g34fd40d4q21c69391c2597bc@mail.gmail.com","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-14T15:29:47Z","receivedAt":"2008-03-14T15:29:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Jay Soffian\" <jaysoffian@gmail.com> writes:\n\n> On Thu, Mar 13, 2008 at 7:14 PM, Petr Baudis <pasky@suse.cz> wrote:\n> \n> ...\n> \n> >  +               if ($cache_lifetime and -f $cache_file) {\n> >  +                       # Postpone timeout by two minutes so that we get\n> >  +                       # enough time to do our job.\n> >  +                       my $time = time() - $cache_lifetime + 120;\n> >  +                       utime $time, $time, $cache_file;\n> >  +               }\n> \n> Race condition. I don't see any locking. Nothing keeps multiple\n> instances from regenerating the cache concurrently...\n> \n> >  +               @projects = git_get_projects_details($projlist, $check_forks);\n> >  +               if ($cache_lifetime and open (my $fd, '>'.$cache_file)) {\n> \n> ...and then clobbering each other here. You have two choices:\n> \n> 1) Use a lock file for the critical section.\n> \n> 2) Assume the race condition is rare enough, but you still need to\n> account for it. In that case, you want to write to a temporary file\n> and then rename to the cache file name. The rename is atomic, so\n> though N instances of gitweb may regenerate the cache (at some\n> CPU/IO overhead), at least the cache file won't get corrupted.\n\nWhat should the code for this look like? Like below?\n\n        use File::Temp;\n        \n        my ($fh, $temp_file) = tempfile();\n        ...\n        close $fh;\n        rename $temp_file, $cache_file;\n\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"72140","messageId":"76718490803141411v24c31de5x8ba25fcd1654b4e7@mail.gmail.com","threadId":"12677","inReplyTo":"m38x0lxr1k.fsf@localhost.localdomain","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2008-03-14T21:11:53Z","receivedAt":"2008-03-14T21:11:53Z","isPatch":true,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On Fri, Mar 14, 2008 at 11:29 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n>  What should the code for this look like? Like below?\n>\n>         use File::Temp;\n>\n>         my ($fh, $temp_file) = tempfile();\n>         ...\n>         close $fh;\n>         rename $temp_file, $cache_file;\n\nI always use something like:\n\n  my $temp_file = \"$cache_file.tmp$$\";\n  open(my $fh, \">$temp_file\");\n\nto ensure that the temp file is on the same filesystem.\n\nj.\n"},{"id":"72196","messageId":"m3ve3nwtl3.fsf@localhost.localdomain","threadId":"12677","inReplyTo":"20080313231413.27966.3383.stgit@rover","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-15T21:44:42Z","receivedAt":"2008-03-15T21:44:42Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Petr Baudis <pasky@suse.cz> writes:\n\n> On repo.or.cz (permanently I/O overloaded and hosting 1050 project +\n> forks), \n\nIt looks like repo.or.cz is overwhelmed by its success. I hope that\nnow that there are other software hosting sites with git hosting\n(Savannah, GitHub, Gitorious,...) the number of projects wouldn't grow\nas rapidly.\n\n> the projects list (the default gitweb page) can take more than\n> a minute to generate. This naive patch adds simple support for caching\n> the projects list data structure so that all the projects do not need\n> to get rescanned at every page access.\n\nAnother solution would be to divide projects list page into pages,\nperhaps adding search box for searching for a project (by name, by\ndescription and by owner).\n\nNevertheless even with pagination, if we want to have \"sort by last\nupdate\" we do need caching.\n\n[...]\n> +# projects list cache for busy sites with many projects;\n> +# if you set this to non-zero, it will be used as the cached\n> +# index lifetime in minutes\n> +# the cached list version is stored in /tmp and can be tweaked\n> +# by other scripts running with the same uid as gitweb - use this\n> +# only at secure installations; only single gitweb project root per\n> +# system is supported!\n> +our $projlist_cache_lifetime = 0;\n\n[...]\n> +sub git_project_list_body {\n[...]\n> +\tmy $cache_file = '/tmp/gitweb.index.cache';\n> +\tuse File::stat;\n> +\n> +\tmy @projects;\n> +\tmy $stale = 0;\n> +\tif ($cache_lifetime and -f $cache_file\n> +\t    and stat($cache_file)->mtime + $cache_lifetime * 60 > time()\n> +\t    and open (my $fd, $cache_file)) {\n> +\t\t$stale = time() - stat($cache_file)->mtime;\n> +\t\tmy @dump = <$fd>;\n> +\t\tclose $fd;\n> +\t\t# Hack zone start\n> +\t\tmy $VAR1;\n> +\t\teval join(\"\\n\", @dump);\n> +\t\t@projects = @$VAR1;\n> +\t\t# Hack zone end\n> +\t} else {\n> +\t\tif ($cache_lifetime and -f $cache_file) {\n> +\t\t\t# Postpone timeout by two minutes so that we get\n> +\t\t\t# enough time to do our job.\n> +\t\t\tmy $time = time() - $cache_lifetime + 120;\n> +\t\t\tutime $time, $time, $cache_file;\n> +\t\t}\n> +\t\t@projects = git_get_projects_details($projlist, $check_forks);\n> +\t\tif ($cache_lifetime and open (my $fd, '>'.$cache_file)) {\n> +\t\t\tuse Data::Dumper;\n> +\t\t\tprint $fd Dumper(\\@projects);\n> +\t\t\tclose $fd;\n> +\t\t}\n> +\t}\n\nThis could be much simplified with perl-cache (perl-Cache-Cache).\nUnfortunately this is non-standard module, not distributed (yet?)\nwith Perl.\n\nWarning: not tested in gitweb!\n\n+\tuse Cache::FileCache;\n+\n+\tmy $cache;\n+\tmy $projects;\n+\t\n+\tif ($cache_lifetime) {\n+\t\t$cache = new Cache::FileCache(\n+\t\t\t{ namespace => 'gitweb',\n+\t\t\t  default_expires_in => $cache_lifetime\n+\t\t\t});\n+\t\t$projects = $cache->get('projects_list');\n+\t}\n+\tif (!defined $projects) {\n+\t\t$projects = [ git_get_projects_details($projlist, $check_forks); ];\n+\t\t$cache->set('projects_list', $projects)\n+\t\t\tif defined $cache;\n+\t}\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"72199","messageId":"20080316005645.GY2414@genesis.frugalware.org","threadId":"12677","inReplyTo":"m3ve3nwtl3.fsf@localhost.localdomain","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Miklos Vajna","fromEmail":"vmiklos@frugalware.org","sentAt":"2008-03-16T00:56:45Z","receivedAt":"2008-03-16T00:56:45Z","isPatch":true,"sender":{"key":"vmiklos@frugalware.org","avatar":"https://gravatar.com/avatar/401c1cbbb3a5d13e650c691a2c71d6fd0b80df1a01bc74d9f1972675dd58f2bd?d=mp&s=160"},"body":"On Sat, Mar 15, 2008 at 02:44:42PM -0700, Jakub Narebski <jnareb@gmail.com> wrote:\n> It looks like repo.or.cz is overwhelmed by its success. I hope that\n> now that there are other software hosting sites with git hosting\n> (Savannah, GitHub, Gitorious,...) the number of projects wouldn't grow\n> as rapidly.\n\ni think repo.or.cz is still the only one that offers mirroring of git\nrepos while it's quite a handy feature.\n"},{"id":"72217","messageId":"20080316114151.GZ10103@mail-vs.djpig.de","threadId":"12677","inReplyTo":"m3ve3nwtl3.fsf@localhost.localdomain","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Frank Lichtenheld","fromEmail":"frank@lichtenheld.de","sentAt":"2008-03-16T11:41:51Z","receivedAt":"2008-03-16T11:41:51Z","isPatch":true,"sender":{"key":"frank@lichtenheld.de","avatar":"https://gravatar.com/avatar/b9f1d4b120e138f157c9e480d0818197c474628923786adb98f30017cdb99c3c?d=mp&s=160"},"body":"On Sat, Mar 15, 2008 at 02:44:42PM -0700, Jakub Narebski wrote:\n> Petr Baudis <pasky@suse.cz> writes:\n> This could be much simplified with perl-cache (perl-Cache-Cache).\n> Unfortunately this is non-standard module, not distributed (yet?)\n> with Perl.\n\nI think somebody who actually needs this can be bothered to install a\nCPAN perl module. This should probably not enabled by default anyway.\n\n> Warning: not tested in gitweb!\n> \n> +\tuse Cache::FileCache;\n> +\n> +\tmy $cache;\n> +\tmy $projects;\n> +\t\n> +\tif ($cache_lifetime) {\n> +\t\t$cache = new Cache::FileCache(\n> +\t\t\t{ namespace => 'gitweb',\n> +\t\t\t  default_expires_in => $cache_lifetime\n> +\t\t\t});\n> +\t\t$projects = $cache->get('projects_list');\n> +\t}\n> +\tif (!defined $projects) {\n> +\t\t$projects = [ git_get_projects_details($projlist, $check_forks); ];\n> +\t\t$cache->set('projects_list', $projects)\n> +\t\t\tif defined $cache;\n> +\t}\n\nGruesse,\n-- \nFrank Lichtenheld <frank@lichtenheld.de>\nwww: http://www.djpig.de/\n"},{"id":"72220","messageId":"1205686355.2758.31.camel@localhost.localdomain","threadId":"12677","inReplyTo":"20080316114151.GZ10103@mail-vs.djpig.de","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"J.H.","fromEmail":"warthog19@eaglescrag.net","sentAt":"2008-03-16T16:52:34Z","receivedAt":"2008-03-16T16:52:34Z","isPatch":true,"sender":{"key":"warthog19@eaglescrag.net","avatar":null},"body":"On Sun, 2008-03-16 at 12:41 +0100, Frank Lichtenheld wrote:\n> On Sat, Mar 15, 2008 at 02:44:42PM -0700, Jakub Narebski wrote:\n> > Petr Baudis <pasky@suse.cz> writes:\n> > This could be much simplified with perl-cache (perl-Cache-Cache).\n> > Unfortunately this is non-standard module, not distributed (yet?)\n> > with Perl.\n> \n> I think somebody who actually needs this can be bothered to install a\n> CPAN perl module. This should probably not enabled by default anyway.\n\nThe people who need the caching are also likely those who are most\naverse to using things that don't either come with their distribution or\naren't easily and readily available in something like an extras\nrepository or a very well trusted contrib repository.  I can at least\nvouch for one large site that needs this that doesn't install things via\ncpan for a lot of different reasons.\n\n\n- John 'Warthog9' Hawley\n"},{"id":"72225","messageId":"200803161937.07082.jnareb@gmail.com","threadId":"12677","inReplyTo":"1205686355.2758.31.camel@localhost.localdomain","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-16T18:37:05Z","receivedAt":"2008-03-16T18:37:05Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sun, 16 Mar 2008, J.H. wrote:\n> On Sun, 2008-03-16 at 12:41 +0100, Frank Lichtenheld wrote:\n>> On Sat, Mar 15, 2008 at 02:44:42PM -0700, Jakub Narebski wrote:\n>>>  \n>>> This could be much simplified with perl-cache (perl-Cache-Cache).\n>>> Unfortunately this is non-standard module, not distributed (yet?)\n>>> with Perl.\n>> \n>> I think somebody who actually needs this can be bothered to install a\n>> CPAN perl module. This should probably not enabled by default anyway.\n> \n> The people who need the caching are also likely those who are most\n> averse to using things that don't either come with their distribution or\n> aren't easily and readily available in something like an extras\n> repository or a very well trusted contrib repository.  I can at least\n> vouch for one large site that needs this that doesn't install things via\n> cpan for a lot of different reasons.\n\nActually Cache::FileCache, which is part of CacheCache distribution,\nshould be available in contrib or even extras repository. I have\ninstalled it as perl-Cache-Cache RPM (1.05-1.fc4.rf) on my Aurox 11.1\n(which is old Fedora Core 4 based distribution), from Dries RPM\nrepository (part of FreshRPM now, IIRC).\n\nThe problem is that at least according to what documentation of other,\nnever CPAN modules says Cache::FileCache is slow, as it always serialize\nusing Storable (Storable should be part of perl distribution).\n\n\nWe can always install local copy alongside gitweb...\n\n\nP.S. When searching CPAN for existing modules for caching and CGI\ncaching I have found Cache::Adaptive::ByLoad which does what\ncaching-gitweb does, and some solutions in newer caching interfaces,\neither CHI or Cache which try to avoid thundering horde problem.\n\nP.P.S. Does kernel.org use memcached, or some kind of web cache\n(reverse proxy cache) like Varnish or Squid?\n-- \nJakub Narebski\nPoland\n"},{"id":"72252","messageId":"1205707048.2758.35.camel@localhost.localdomain","threadId":"12677","inReplyTo":"200803161937.07082.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"J.H.","fromEmail":"warthog19@eaglescrag.net","sentAt":"2008-03-16T22:37:28Z","receivedAt":"2008-03-16T22:37:28Z","isPatch":true,"sender":{"key":"warthog19@eaglescrag.net","avatar":null},"body":"On Sun, 2008-03-16 at 19:37 +0100, Jakub Narebski wrote:\n> On Sun, 16 Mar 2008, J.H. wrote:\n> > On Sun, 2008-03-16 at 12:41 +0100, Frank Lichtenheld wrote:\n> >> On Sat, Mar 15, 2008 at 02:44:42PM -0700, Jakub Narebski wrote:\n> >>>  \n> >>> This could be much simplified with perl-cache (perl-Cache-Cache).\n> >>> Unfortunately this is non-standard module, not distributed (yet?)\n> >>> with Perl.\n> >> \n> >> I think somebody who actually needs this can be bothered to install a\n> >> CPAN perl module. This should probably not enabled by default anyway.\n> > \n> > The people who need the caching are also likely those who are most\n> > averse to using things that don't either come with their distribution or\n> > aren't easily and readily available in something like an extras\n> > repository or a very well trusted contrib repository.  I can at least\n> > vouch for one large site that needs this that doesn't install things via\n> > cpan for a lot of different reasons.\n> \n> Actually Cache::FileCache, which is part of CacheCache distribution,\n> should be available in contrib or even extras repository. I have\n> installed it as perl-Cache-Cache RPM (1.05-1.fc4.rf) on my Aurox 11.1\n> (which is old Fedora Core 4 based distribution), from Dries RPM\n> repository (part of FreshRPM now, IIRC).\n\nThat would be fine, I don't think the larger sites would have issues\nfinding a copy than.\n\n> \n> The problem is that at least according to what documentation of other,\n> never CPAN modules says Cache::FileCache is slow, as it always serialize\n> using Storable (Storable should be part of perl distribution).\n\nThat makes it much less interesting unfortunately.\n\n> We can always install local copy alongside gitweb...\n> \n> \n> P.S. When searching CPAN for existing modules for caching and CGI\n> caching I have found Cache::Adaptive::ByLoad which does what\n> caching-gitweb does, and some solutions in newer caching interfaces,\n> either CHI or Cache which try to avoid thundering horde problem.\n\nInteresting - my have to take a look at that.\n\n> P.P.S. Does kernel.org use memcached, or some kind of web cache\n> (reverse proxy cache) like Varnish or Squid?\n\nNo - doesn't buy us anything really unfortunately.  And since I'm doing\ncaching inside of gitweb itself having multiple layers of caching just\nmakes things more complicated, adds unnecessary latency to updates, etc.\n\n- John\n"},{"id":"72258","messageId":"200803170039.14634.jnareb@gmail.com","threadId":"12677","inReplyTo":"1205707048.2758.35.camel@localhost.localdomain","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-16T23:39:11Z","receivedAt":"2008-03-16T23:39:11Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"J.H. wrote:\n> Jakub Narebski wrote:\n>>\n>> P.S. When searching CPAN for existing modules for caching and CGI\n>> caching I have found Cache::Adaptive::ByLoad which does what\n>> caching-gitweb does,\n\nI'm not sure about quality of this code, though. It uses Cache::Cache, \nby the way.\n\n>> and some solutions in newer caching interfaces, \n>> either CHI or Cache, which try to avoid thundering horde problem.\n> \n> Interesting - my have to take a look at that.\n\nCHI uses either 'busy_lock [DURATION]' (bump expiration time), or\n'expires_variance [FLOAT]' for fuzzy expiration time matching.\n\nCache has LRU and FIFO removal strategies.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"72303","messageId":"20080317174050.GB10335@machine.or.cz","threadId":"12677","inReplyTo":"m3hcf9y02p.fsf@localhost.localdomain","subject":"Re: [PATCH] gitweb: Support caching projects list","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-17T17:40:50Z","receivedAt":"2008-03-17T17:40:50Z","isPatch":true,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  Hi,\n\nOn Fri, Mar 14, 2008 at 05:14:51AM -0700, Jakub Narebski wrote:\n> Petr Baudis <pasky@suse.cz> writes:\n> [...]\n> > +\tif ($cache_lifetime and -f $cache_file\n> > +\t    and stat($cache_file)->mtime + $cache_lifetime * 60 > time()\n> > +\t    and open (my $fd, $cache_file)) {\n> > +\t\t$stale = time() - stat($cache_file)->mtime;\n> > +\t\tmy @dump = <$fd>;\n> > +\t\tclose $fd;\n> > +\t\t# Hack zone start\n> > +\t\tmy $VAR1;\n> > +\t\teval join(\"\\n\", @dump);\n> > +\t\t@projects = @$VAR1;\n> > +\t\t# Hack zone end\n> \n> Why do you read line by line, only to join it, i.e.\n>   my @dump = <$fd>; ... join(\"\\n\", @dump);\n> instead of slurping all file in one go:\n>   local $/ = undef; my $dump = <$fd>; ... $dump;\n> \n> Besides, why do you use Data::Dumper instead of Storable? Both are\n> distributed with Perl; well, at least both are in perl-5.8.6-24.\n\n  no particular reason - I simply never heard about Storable. I learned\nPerl too long ago it seems. ;-)\n\n> [...]\n> > -\tgit_project_list_body(\\@list, $order);\n> > +\tgit_project_list_body(\\@list, $order, undef, undef, undef, undef, $projlist_cache_lifetime);\n> \n> This is ugly. Why not use hash for \"named parameters\", as it is done\n> in a few separate places in gitweb (search for '%opts')?\n\n  I agree - I was simply too lazy to make another patch. :-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"72304","messageId":"20080317174934.GC6803@machine.or.cz","threadId":"12677","inReplyTo":"1205454999.2758.14.camel@localhost.localdomain","subject":"repo.or.cz renovation","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-17T17:49:34Z","receivedAt":"2008-03-17T17:49:34Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Thu, Mar 13, 2008 at 05:36:39PM -0700, J.H. wrote:\n> \n> > You are of course right - I wanted to do the rename, but forgot to write\n> > it in the actual code. :-)\n> > \n> > There is a more conceptual problem though - in case of such big sites,\n> > it really makes more sense to explicitly regenerate the cache\n> > periodically instead of making random clients to have to wait it out.\n> > We could add a 'force_update' parameter to accept from localhost only\n> > that will always regenerate the cache, but that feels rather kludgy -\n> > can anyone think of a more elegant solution? (I don't think taking the\n> > @projects generating code out of gitweb and then having to worry during\n> > gitweb upgrades is any better.)\n> \n> You could do something similar to the gitweb caching I'm doing,\n> basically if a file isn't generated you make a user wait (no good way\n> around this really).  If a cache exists show it to the user unless the\n> cache is older than $foo.  If a re-generation needs to happen it happens\n> in the background so the user who triggers the regeneration sees\n> something immediately vs. having to wait (at the cost of showing out of\n> date data)\n\nBy the way, the index page is so far really the only bottleneck I'm\nseeing, other than that even project pages for huge repositories are\nshown pretty quickly. Did you ever try to just cache the index page on\nkernel.org? What sort of impact did it have? What evere the hotspots -\nproject pages for the main repositories or some less obvious pages?\n\nJust caching the index would be far less intrusive change than\nintroducing caching everywhere and it might help to bring kernel.org\ngitweb back in sync with mainline. :-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"72305","messageId":"20080317181015.GC10335@machine.or.cz","threadId":"12677","inReplyTo":"m3ve3nwtl3.fsf@localhost.localdomain","subject":"repo.or.cz renovated","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-17T18:10:15Z","receivedAt":"2008-03-17T18:10:15Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Sat, Mar 15, 2008 at 02:44:42PM -0700, Jakub Narebski wrote:\n> Petr Baudis <pasky@suse.cz> writes:\n> \n> > On repo.or.cz (permanently I/O overloaded and hosting 1050 project +\n> > forks), \n> \n> It looks like repo.or.cz is overwhelmed by its success. I hope that\n> now that there are other software hosting sites with git hosting\n> (Savannah, GitHub, Gitorious,...) the number of projects wouldn't grow\n> as rapidly.\n\nActually, it was overwhelmed to so much by its success but by lack of\ngood maintenance. ;-) I gave it some love again for the past week and\nthe improvement was, well, overwhelming. :-)\n\nI finally fixed tons of failures and broken repositories, and most\nimportantly repacked some of the big repositories with object databases\nin pretty horrid shape. The effect has been immense, having everything\nin database of 1/3 the size and single big pack drastically reduced the\nI/O load.\n\nScenario: Site with about 1100 repositories weighting 13GB, running a\nfetch job for about 200 of them hourly. About two git-daemon requests\nper minute and 10 gitweb requests per minute (the last two numbers are\ntaken quite sloppily over a small sample of the last ten minutes ;-).\nSite is running on 2x1GHz P3 with 2G RAM, repository is on hw RAID5.\n(We are currently preparing to migrate it to a more powerful machine.)\n\nBefore, the load on the server would be normally about 6 to 15 _all the\ntime_ and bunch of git-related processes would be permanently eating\nsome CPU and crunch on the disk.\n\nAfter introducing the index caching and repacking the repositories, the\nload seems to be around 1 at most and hardly seems to come above 3; all\nfeels very snappy.\n\nSo for anyone running a hosting site, make sure your repositories are\nnicely packed. It makes huge difference to the I/O load!\n\n> Another solution would be to divide projects list page into pages,\n> perhaps adding search box for searching for a project (by name, by\n> description and by owner).\n> \n> Nevertheless even with pagination, if we want to have \"sort by last\n> update\" we do need caching.\n\nYes, I'm pondering about pagination, but because of web clients, not the\nserver load; it takes firefox on my notebook noticeable time to render\nthis list already, and it's rather big too. Ideas are welcome here.\n\nMy current plan is to have a [Search project] box at the front page,\ntogether with direct link to 'show all'. Other than that, what makes\nsense to display on the front page? I think recently added projects (age\n< 1 week) for sure. I'm not so sure about recently changed projects -\nmaybe it is better to keep the front page cruft-free.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"72306","messageId":"20080317181148.GD6803@machine.or.cz","threadId":"12677","inReplyTo":"20080317174934.GC6803@machine.or.cz","subject":"Re: repo.or.cz renovation","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-17T18:11:48Z","receivedAt":"2008-03-17T18:11:48Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  (Sorry, I messed up the subject here. See the other mail if you are\ninterested about recent speedups of repo.or.cz. And if replying, it\nwould be nice to revert the subject back. :-)\n\n\t\t\t\tPetr \"Pasky\" Baudis\n"},{"id":"72307","messageId":"1205779482.2758.52.camel@localhost.localdomain","threadId":"12677","inReplyTo":"20080317174934.GC6803@machine.or.cz","subject":"Re: repo.or.cz renovation","fromName":"J.H.","fromEmail":"warthog19@eaglescrag.net","sentAt":"2008-03-17T18:44:42Z","receivedAt":"2008-03-17T18:44:42Z","isPatch":false,"sender":{"key":"warthog19@eaglescrag.net","avatar":null},"body":"On Mon, 2008-03-17 at 18:49 +0100, Petr Baudis wrote:\n> On Thu, Mar 13, 2008 at 05:36:39PM -0700, J.H. wrote:\n> > \n> > > You are of course right - I wanted to do the rename, but forgot to write\n> > > it in the actual code. :-)\n> > > \n> > > There is a more conceptual problem though - in case of such big sites,\n> > > it really makes more sense to explicitly regenerate the cache\n> > > periodically instead of making random clients to have to wait it out.\n> > > We could add a 'force_update' parameter to accept from localhost only\n> > > that will always regenerate the cache, but that feels rather kludgy -\n> > > can anyone think of a more elegant solution? (I don't think taking the\n> > > @projects generating code out of gitweb and then having to worry during\n> > > gitweb upgrades is any better.)\n> > \n> > You could do something similar to the gitweb caching I'm doing,\n> > basically if a file isn't generated you make a user wait (no good way\n> > around this really).  If a cache exists show it to the user unless the\n> > cache is older than $foo.  If a re-generation needs to happen it happens\n> > in the background so the user who triggers the regeneration sees\n> > something immediately vs. having to wait (at the cost of showing out of\n> > date data)\n> \n> By the way, the index page is so far really the only bottleneck I'm\n> seeing, other than that even project pages for huge repositories are\n> shown pretty quickly. Did you ever try to just cache the index page on\n> kernel.org? What sort of impact did it have? What evere the hotspots -\n> project pages for the main repositories or some less obvious pages?\n> \n> Just caching the index would be far less intrusive change than\n> introducing caching everywhere and it might help to bring kernel.org\n> gitweb back in sync with mainline. :-)\n\nI think we are likely going to want to keep caching everything vs. just\nthe front page.  There are a few repos that get hit quite a bit and it\nwould be better to have those cache vs. not.  Really I would argue this\nis just a step in the direction of integrating all of my caching changes\nback into gitweb vs. us dropping what we've done so far.\n\nBTW I'm about halfway through refactoring my tree from multiple files\nback to one, which at that point means I can start bringing it back into\nmainline and getting a patch series ready for submission.\n\n- John\n"},{"id":"72309","messageId":"7vhcf5b21n.fsf@gitster.siamese.dyndns.org","threadId":"12677","inReplyTo":"20080317181015.GC10335@machine.or.cz","subject":"Re: repo.or.cz renovated","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-03-17T19:09:24Z","receivedAt":"2008-03-17T19:09:24Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Petr Baudis <pasky@suse.cz> writes:\n\n> My current plan is to have a [Search project] box at the front page,\n> together with direct link to 'show all'. Other than that, what makes\n> sense to display on the front page? I think recently added projects (age\n> < 1 week) for sure. I'm not so sure about recently changed projects -\n> maybe it is better to keep the front page cruft-free.\n\n\"More important projects\"?\n\n;-) Ducks...\n\nHow about asking project owners to categorize (tag) their own projects and\nshow them in different categories?  Maybe many of them will start in\n\"unsorted bin\", but if you organize the top page in such a way that the\nlink to unsorted bin is much less prominent than nicely sorted ones, that\nmay give people incentive to put their project in a real category.\n"},{"id":"72312","messageId":"20080317192528.GE10335@machine.or.cz","threadId":"12677","inReplyTo":"7vhcf5b21n.fsf@gitster.siamese.dyndns.org","subject":"Re: repo.or.cz renovated","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-17T19:25:28Z","receivedAt":"2008-03-17T19:25:28Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Mon, Mar 17, 2008 at 12:09:24PM -0700, Junio C Hamano wrote:\n> Petr Baudis <pasky@suse.cz> writes:\n> \n> > My current plan is to have a [Search project] box at the front page,\n> > together with direct link to 'show all'. Other than that, what makes\n> > sense to display on the front page? I think recently added projects (age\n> > < 1 week) for sure. I'm not so sure about recently changed projects -\n> > maybe it is better to keep the front page cruft-free.\n> \n> \"More important projects\"?\n> \n> ;-) Ducks...\n\nI'm all for that, get access to projects people are the most likely to\nlook for. Would it make sense to count accesses to index pages of each\nproject and then sort by that? Or sort by some activity index?\n\n> How about asking project owners to categorize (tag) their own projects and\n> show them in different categories?  Maybe many of them will start in\n> \"unsorted bin\", but if you organize the top page in such a way that the\n> link to unsorted bin is much less prominent than nicely sorted ones, that\n> may give people incentive to put their project in a real category.\n\nActually, no reason to restrict tagging to project owners. This might be\ninteresting, but again the question is, does anyone really want to\nbrowse the project list based on \"show me all kernel-related projects\"\nor \"show me all xorg-related projects\", especially as we have forks and\nit may not be reliable?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"72314","messageId":"20080317193423.GI8368@mit.edu","threadId":"12677","inReplyTo":"20080317181015.GC10335@machine.or.cz","subject":"Re: repo.or.cz renovated","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-03-17T19:34:23Z","receivedAt":"2008-03-17T19:34:23Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Mar 17, 2008 at 07:10:15PM +0100, Petr Baudis wrote:\n> Actually, it was overwhelmed to so much by its success but by lack of\n> good maintenance. ;-) I gave it some love again for the past week and\n> the improvement was, well, overwhelming. :-)\n> \n> I finally fixed tons of failures and broken repositories, and most\n> importantly repacked some of the big repositories with object databases\n> in pretty horrid shape. The effect has been immense, having everything\n> in database of 1/3 the size and single big pack drastically reduced the\n> I/O load.\n\nAre you making sure that repositories which are forks off of some\nparent repository are using objects/info/alternates to share objects?\n(If so you have to be careful when you prune not to drop objects, but\nit can make a huge difference in disk utilization and I/O bandwidth).\n\nAt least for master.kernel.org, and for those git repositories which I\nown, I make a point of periodically logging in and running git gc,\ncopying over the object packs so I can do a prune operation safely,\netc.  --- and I suspect most of the master.kernel.org git users do\nsomething similar.  On repo.or.cz we don't have shell access, so the\nproject administrators can't do that for you.\n\n> So for anyone running a hosting site, make sure your repositories are\n> nicely packed. It makes huge difference to the I/O load!\n\nIt seems that a Really Good Idea would be do the the packing and\npruning via cron scripts that run during the off hours...\n\n> My current plan is to have a [Search project] box at the front page,\n> together with direct link to 'show all'. Other than that, what makes\n> sense to display on the front page? I think recently added projects (age\n> < 1 week) for sure. I'm not so sure about recently changed projects -\n> maybe it is better to keep the front page cruft-free.\n\nThere are plenty of ways which sites like freshmeat and sourceforge\nhave come up to make it easy to browse a large number of software\nprojects.  One way that might make sense is Sourceforge's Software Map\n(i.e., http://sourceforge.net/softwaremap/).\n\n\n\t\t\t\t\t- Ted\n"},{"id":"72316","messageId":"20080317195422.GF10335@machine.or.cz","threadId":"12677","inReplyTo":"20080317193423.GI8368@mit.edu","subject":"Re: repo.or.cz renovated","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-03-17T19:54:22Z","receivedAt":"2008-03-17T19:54:22Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Mon, Mar 17, 2008 at 03:34:23PM -0400, Theodore Tso wrote:\n> On Mon, Mar 17, 2008 at 07:10:15PM +0100, Petr Baudis wrote:\n> > Actually, it was overwhelmed to so much by its success but by lack of\n> > good maintenance. ;-) I gave it some love again for the past week and\n> > the improvement was, well, overwhelming. :-)\n> > \n> > I finally fixed tons of failures and broken repositories, and most\n> > importantly repacked some of the big repositories with object databases\n> > in pretty horrid shape. The effect has been immense, having everything\n> > in database of 1/3 the size and single big pack drastically reduced the\n> > I/O load.\n> \n> Are you making sure that repositories which are forks off of some\n> parent repository are using objects/info/alternates to share objects?\n> (If so you have to be careful when you prune not to drop objects, but\n> it can make a huge difference in disk utilization and I/O bandwidth).\n\nYes, I reuse objects from parent projects, that has always been so.\n\n> At least for master.kernel.org, and for those git repositories which I\n> own, I make a point of periodically logging in and running git gc,\n> copying over the object packs so I can do a prune operation safely,\n> etc.  --- and I suspect most of the master.kernel.org git users do\n> something similar.  On repo.or.cz we don't have shell access, so the\n> project administrators can't do that for you.\n> \n> > So for anyone running a hosting site, make sure your repositories are\n> > nicely packed. It makes huge difference to the I/O load!\n> \n> It seems that a Really Good Idea would be do the the packing and\n> pruning via cron scripts that run during the off hours...\n\nYes, this was done before too, however repo.or.cz has been around for\nlong time and historically the scripts weren't working very well,\nespecially since I had to be careful about the forks problem.\n\nSince I am repacking on live system, I think the current repacking\nstrategy is still not completely error prone, however I believe that I\nhave encountered no breakage because of pruned objects the last at least\nhalf a year or so it has been running with the current setup (all of the\nbreakages I have encountered seem to be caused by child process of\ngit-repack dying).  Besides, if some fork breaks, it should be possible\nto fix that very easily (I do not backup the object stores at all\nanyway - if the server burns down, you will have to re-push ;-).\n\n> > My current plan is to have a [Search project] box at the front page,\n> > together with direct link to 'show all'. Other than that, what makes\n> > sense to display on the front page? I think recently added projects (age\n> > < 1 week) for sure. I'm not so sure about recently changed projects -\n> > maybe it is better to keep the front page cruft-free.\n> \n> There are plenty of ways which sites like freshmeat and sourceforge\n> have come up to make it easy to browse a large number of software\n> projects.  One way that might make sense is Sourceforge's Software Map\n> (i.e., http://sourceforge.net/softwaremap/).\n\nThis all feels like a real overkill, besides my main doubt is whether\nrepo.or.cz needs something like this *at all* - but I think I will try\nthe tagging system and see how do people like it.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"72319","messageId":"m3myoxw0at.fsf@localhost.localdomain","threadId":"12677","inReplyTo":"1205779482.2758.52.camel@localhost.localdomain","subject":"Re: repo.or.cz renovation","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-17T20:41:39Z","receivedAt":"2008-03-17T20:41:39Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"J.H.\" <warthog19@eaglescrag.net> writes:\n\n> BTW I'm about halfway through refactoring my tree from multiple files\n> back to one, which at that point means I can start bringing it back into\n> mainline and getting a patch series ready for submission.\n\nI plan on sending email with my ideas on gitweb caching somewhere\nbetween now and tomorrow. I'll try to cover my ideas about how to\ncache (support for cache validation for external cache, caching Perl\nstructures / data from git commands, caching final output: HTML, RSS,\netc.) and what solutions can be used.\n\nI wanted first to send my enchancements to Petr Baudis patch adding\ncaching support for projects list, i.e. third patch in the series, and\nget comments (currently none) on *this* patch (idea).\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"72320","messageId":"m3iqzlvz11.fsf@localhost.localdomain","threadId":"12677","inReplyTo":"1205779482.2758.52.camel@localhost.localdomain","subject":"Re: repo.or.cz renovation","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-17T21:09:01Z","receivedAt":"2008-03-17T21:09:01Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"J.H.\" <warthog19@eaglescrag.net> writes:\n\n> BTW I'm about halfway through refactoring my tree from multiple files\n> back to one, which at that point means I can start bringing it back into\n> mainline and getting a patch series ready for submission.\n\nBTW it would be nice to have a merge strategy (blame-based perhaps?)\nwhich would allow to merge changes to split project easily into\noriginal, single file one...  But I guess you are not interested in\nwriting such a merge strategy just for this ;-)\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"}]}