{"thread":{"id":"12001","subject":"Another bench on gitweb","startedAt":"2008-02-10T03:09:19Z","lastAt":"2008-02-15T23:19:08Z","messageCount":11,"participants":["Bruno Cesar Ribas","Jakub Narebski","J.H."],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"68158","messageId":"20080210030919.GA32733@c3sl.ufpr.br","threadId":"12001","inReplyTo":null,"subject":"Another bench on gitweb","fromName":"Bruno Cesar Ribas","fromEmail":"ribas@c3sl.ufpr.br","sentAt":"2008-02-10T03:09:19Z","receivedAt":"2008-02-10T03:09:19Z","isPatch":false,"sender":{"key":"ribas@c3sl.ufpr.br","avatar":null},"body":"Hello,\n\nI made another SIMPLE bench on gitweb. Testing time on git-for-each-ref.\n\nUsing my 1000 projects I ran:\n8<----------------\n#/bin/bash\nPEGAR_ref() { \n    PROJ=projeto$1.git; \n    cd $PROJ; \n    printf \"\\tlastref = $(git-for-each-ref --sort=-committerdate --count=1\\\n            --format='%(committer)')\\n\" >> config; \n    cd -; \n}\ncd $HOME/scm\nfor((i=1;i<=1000;i++)){ PEGAR_ref $i & }\n8<----------------\n\nAnd at the \"git_get_last_activity\" instead of running git-for-each-ref i\nasked to get gitweb.lastref\n\nHere are the results:\n\"dd\" means: dd if=/dev/zero of=$HOME/dd/$i bs=1M count=400000\n\nRunning 2 dd to generate disk IO.  Here comes the results:\nNO projects_list  projects_list\n7m56s55           6m11s95        cached last change, using gitweb.lastref\n16m30s69          15m10s74       default gitweb, using FS's owner\n16m07s40          15m24s34       patched to get gitweb.owner\n\n\nNow results for a 1000projects on an idle machine. (No dd running to\ngenerate IO)\nNO projects_list  projects_list\n0m26s79           0m38s70       cached last change, using gitweb.lastref\n1m19s08           1m09s55       default gitweb, using FS's owner\n1m17s58           1m09s55       patched to get gitweb.owner\n\n\nI found out those VERY interesting, so instead of trying to think a new way\nto store gitweb config, we should think a way to cache those information.\n-- \nBruno Ribas - ribas@c3sl.ufpr.br\nhttp://web.inf.ufpr.br/ribas\nC3SL: http://www.c3sl.ufpr.br \n"},{"id":"68469","messageId":"m363wvdmxr.fsf@localhost.localdomain","threadId":"12001","inReplyTo":"20080210030919.GA32733@c3sl.ufpr.br","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-12T00:44:23Z","receivedAt":"2008-02-12T00:44:23Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Bruno Cesar Ribas <ribas@c3sl.ufpr.br> writes:\n\n> I made another SIMPLE bench on gitweb. Testing time on git-for-each-ref.\n> \n> Using my 1000 projects I ran:\n> 8<----------------\n> #/bin/bash\n> PEGAR_ref() { \n>     PROJ=projeto$1.git; \n>     cd $PROJ; \n>     printf \"\\tlastref = $(git-for-each-ref --sort=-committerdate --count=1\\\n>             --format='%(committer)')\\n\" >> config; \n>     cd -; \n> }\n> cd $HOME/scm\n> for((i=1;i<=1000;i++)){ PEGAR_ref $i & }\n> 8<----------------\n\nCould you please do not mix English and your native language\n(Portuguese?) in shown examples? Mixing two languages in one\nidentifier name (unless it is ref in br too) is especially bad\nform... TIA.\n\nBesides, what I'm more interested in is a script used to generate\nthose 1000 projects...\n \n> And at the \"git_get_last_activity\" instead of running git-for-each-ref i\n> asked to get gitweb.lastref\n> \n> Here are the results:\n> \"dd\" means: dd if=/dev/zero of=$HOME/dd/$i bs=1M count=400000\n> \n> Running 2 dd to generate disk IO.  Here comes the results:\n> NO projects_list  projects_list\n> 7m56s55           6m11s95        cached last change, using gitweb.lastref\n> 16m30s69          15m10s74       default gitweb, using FS's owner\n> 16m07s40          15m24s34       patched to get gitweb.owner\n> \n> Now results for a 1000projects on an idle machine. (No dd running to\n> generate IO)\n> NO projects_list  projects_list\n> 0m26s79           0m38s70       cached last change, using gitweb.lastref\n> 1m19s08           1m09s55       default gitweb, using FS's owner\n> 1m17s58           1m09s55       patched to get gitweb.owner\n\nThose are results of running gitweb as standalone script, or your\nscript runing git-for-each-ref?\n\nBesides, I'd rather see results of running ApacheBench. On Linux it\nusually comes with installed Apache, and it is called by runing\n'ab'. Your tests instead of adding superficial load could try to use\nconcurrent requests, and more than 1 request to get better average.\n \n> I found out those VERY interesting, so instead of trying to think a\n> new way to store gitweb config, we should think a way to cache those\n> information.\n\nBelow there are my thoughts about caching information for gitweb:\n\nFirst, the basis of each otimisation is checking the bottlenecks.\nI think it was posted sometime there that the pages taking most load\nare projects list and feeds. \n\nKernel.org even run modified version of gitweb, with some caching\nsupport; Cgit (git web interface in C) also has caching support.\n\n\nDue to the fact that gitweb produces relative time in output for\nprojects list page and for project summary page, it is unfortunately\nnot easy to just simply cache HTML output: one would have either\nresign from using relative time, or rewrite time from relative to\nabsolute, either on server (in gitweb), or on client (in JavaScript).\nSo perhaps it would be better to cache generating (costly to obtain)\ninformation; like lastchanged time for projects.\n\nOr we can for example assume (i.e. do that if appropriate gitweb\nfeature is set) that projects are bare projects pushed to, and that\ngit-update-server-info is ran on repository update (for example for\nHTTP protocol transport), and stat $GIT_DIR/info/refs and/or\n$GIT_DIR/objects/info/packs instead of running git-for-each-ref.\nOf course then column would be called something like \"Last Update\"\ninstead of \"Last Change\".\n\nThe \"Last Update\" information is especially easy because it can be\ninvalidated / update externally, by the update / post-receive hook,\noutside gitweb. So gitweb doesn't need to implement some caching\ninvalidation mechanism for this.\n\nWe can store lastref / lastchange information in repository config, as\nfor example \"gitweb.lastref\" key. We can store it in gitweb wide\nconfig, for example in $projectroot/gitwebconfig file, as for example\n\"gitweb.<project>.lastref\" key. Or we can store it as hash initializer\nin some sourced Perl file, read from gitweb_config.perl (this I think\ncan be done even now without touching gitweb code at all); we can use\nData::Dumper to save such information.\n\nThe possibilities are many.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"68578","messageId":"20080213004528.GB31455@c3sl.ufpr.br","threadId":"12001","inReplyTo":"m363wvdmxr.fsf@localhost.localdomain","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"Bruno Cesar Ribas","fromEmail":"ribas@c3sl.ufpr.br","sentAt":"2008-02-13T00:45:28Z","receivedAt":"2008-02-13T00:45:28Z","isPatch":false,"sender":{"key":"ribas@c3sl.ufpr.br","avatar":null},"body":"On Mon, Feb 11, 2008 at 04:44:23PM -0800, Jakub Narebski wrote:\n> Bruno Cesar Ribas <ribas@c3sl.ufpr.br> writes:\n> \n> \n> Could you please do not mix English and your native language\n> (Portuguese?) in shown examples? Mixing two languages in one\n> identifier name (unless it is ref in br too) is especially bad\n> form... TIA.\n\nI agree... that's not good =( i'll enforce to send everything in english.\n> \n> Besides, what I'm more interested in is a script used to generate\n> those 1000 projects...\n\nSo.. like I said, i made a simple test so I cloned a very small project[1]\ne replicated it, just generated different owner and descriptions.\n\n\n> > \n> > Running 2 dd to generate disk IO.  Here comes the results:\n> > NO projects_list  projects_list\n> > 7m56s55           6m11s95        cached last change, using gitweb.lastref\n> > 16m30s69          15m10s74       default gitweb, using FS's owner\n> > 16m07s40          15m24s34       patched to get gitweb.owner\n> \n> Those are results of running gitweb as standalone script, or your\n> script runing git-for-each-ref?\n\nRuning gitweb as standalone script.\n\n> \n> Besides, I'd rather see results of running ApacheBench. On Linux it\n> usually comes with installed Apache, and it is called by runing\n> 'ab'. Your tests instead of adding superficial load could try to use\n> concurrent requests, and more than 1 request to get better average.\n\nhmmm I see, but we will bench with it running with filesystem cached.\nThis could be a good idea if the machine runs only git!\nI find interesting running with all those dds to simulate something like my\nenvironment, which is shared with all of our mirrors. I can even run a test\ninside this machine but results may be very different depending on the time\nof the day.\n\nAs soon I get the machine I ran those tests available again i'll run it with\napache. If you have some ideas of which tests to run tell me =) So we don't\nwaste time when i get the machine.\n\n>  \n> > I found out those VERY interesting, so instead of trying to think a\n> > new way to store gitweb config, we should think a way to cache those\n> > information.\n> \n> Below there are my thoughts about caching information for gitweb:\n> \n> First, the basis of each otimisation is checking the bottlenecks.\n> I think it was posted sometime there that the pages taking most load\n> are projects list and feeds. \n> \n> Kernel.org even run modified version of gitweb, with some caching\n> support; Cgit (git web interface in C) also has caching support.\n\nIs this gitweb version for kernel.org available somewhere?\n> \n> \n><snip> \n> The \"Last Update\" information is especially easy because it can be\n> invalidated / update externally, by the update / post-receive hook,\n> outside gitweb. So gitweb doesn't need to implement some caching\n> invalidation mechanism for this.\n\nthat's what i thought.\n> \n> We can store lastref / lastchange information in repository config, as\n> for example \"gitweb.lastref\" key. We can store it in gitweb wide\n> config, for example in $projectroot/gitwebconfig file, as for example\n> \"gitweb.<project>.lastref\" key. Or we can store it as hash initializer\n> in some sourced Perl file, read from gitweb_config.perl (this I think\n> can be done even now without touching gitweb code at all); we can use\n> Data::Dumper to save such information.\n> \n> The possibilities are many.\n\nThat's right. \n\nCaching lastref at $projectroot/gitwebconfig might be a good idea. I think\nthat caching it at $GIT_DIR/config is somehow ugly, because we will have a\nscript modifying this file.\n\nAnd having this $projectroot/gitwebconfig with lastref cached can act as a\nproject_list because we already all the directories we should get gitweb\nconfs, like gitweb.description, gitweb.url  and gitweb.owner (soon?!) and \nothers that will appear.\n\n> \n> -- \n> Jakub Narebski\n> Poland\n> ShadeHawk on #git\n\n-- \nBruno Ribas - ribas@c3sl.ufpr.br\nhttp://web.inf.ufpr.br/ribas\nC3SL: http://www.c3sl.ufpr.br \n"},{"id":"68579","messageId":"20080213005041.GA5808@c3sl.ufpr.br","threadId":"12001","inReplyTo":"20080213004528.GB31455@c3sl.ufpr.br","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"Bruno Cesar Ribas","fromEmail":"ribas@c3sl.ufpr.br","sentAt":"2008-02-13T00:50:41Z","receivedAt":"2008-02-13T00:50:41Z","isPatch":false,"sender":{"key":"ribas@c3sl.ufpr.br","avatar":null},"body":"On Tue, Feb 12, 2008 at 10:45:28PM -0200, Bruno Cesar Ribas wrote:\n> So.. like I said, i made a simple test so I cloned a very small project[1]\n> e replicated it, just generated different owner and descriptions.\n\njust in time: Project is:\nhttp://git.c3sl.ufpr.br/gitweb?p=chessd/bosh.git;a=summary\n-- \nBruno Ribas - ribas@c3sl.ufpr.br\nhttp://web.inf.ufpr.br/ribas\nC3SL: http://www.c3sl.ufpr.br \n"},{"id":"68592","messageId":"1202864250.17207.22.camel@localhost.localdomain","threadId":"12001","inReplyTo":"20080213004528.GB31455@c3sl.ufpr.br","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"J.H.","fromEmail":"warthog19@eaglescrag.net","sentAt":"2008-02-13T00:57:30Z","receivedAt":"2008-02-13T00:57:30Z","isPatch":false,"sender":{"key":"warthog19@eaglescrag.net","avatar":null},"body":" \n> > > I found out those VERY interesting, so instead of trying to think a\n> > > new way to store gitweb config, we should think a way to cache those\n> > > information.\n> > \n> > Below there are my thoughts about caching information for gitweb:\n> > \n> > First, the basis of each otimisation is checking the bottlenecks.\n> > I think it was posted sometime there that the pages taking most load\n> > are projects list and feeds. \n> > \n> > Kernel.org even run modified version of gitweb, with some caching\n> > support; Cgit (git web interface in C) also has caching support.\n> \n> Is this gitweb version for kernel.org available somewhere?\n> > \n> > \n\nIt's available from my git tree on kernel.org\nhttp://git.kernel.org/?p=git/warthog9/gitweb.git;a=summary\n\nor\n\ngit://git.kernel.org/pub/scm/git/warthog9/gitweb.git\n\nMind you my performance on the non-cache state is not going to be any\nbetter than normal gitweb, however the performance on a cache-hit is\norders of magnitude faster - though at a rather expensive cost - disk\nspace.  There is currently something like 20G of disk being used on one\nof kernel.org's machines providing the cache (this does get flushed on\noccasion - I think) but that is providing caching for everything that\nkernel.org has in it's git trees (or 255188 unique urls currently).  My\ncode base is now, horribly, out of date with respect to mainline but it\nworks and it's been solid and reasonably reliable (though I do know of\ntwo bugs in it right now I need to track down - one with respect to a\nfailure of the script - and one that is an array out of bounds error)\n\n- John\n"},{"id":"68593","messageId":"1202864493.17207.24.camel@localhost.localdomain","threadId":"12001","inReplyTo":"20080213004528.GB31455@c3sl.ufpr.br","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"J.H.","fromEmail":"warthog9@kernel.org","sentAt":"2008-02-13T01:01:33Z","receivedAt":"2008-02-13T01:01:33Z","isPatch":false,"sender":{"key":"warthog9@kernel.org","avatar":"https://avatars.githubusercontent.com/u/2334704?v=4"},"body":"> > > I found out those VERY interesting, so instead of trying to think a\n> > > new way to store gitweb config, we should think a way to cache those\n> > > information.\n> > \n> > Below there are my thoughts about caching information for gitweb:\n> > \n> > First, the basis of each otimisation is checking the bottlenecks.\n> > I think it was posted sometime there that the pages taking most load\n> > are projects list and feeds. \n> > \n> > Kernel.org even run modified version of gitweb, with some caching\n> > support; Cgit (git web interface in C) also has caching support.\n> \n> Is this gitweb version for kernel.org available somewhere?\n> > \n> > \n\nIt's available from my git tree on kernel.org\nhttp://git.kernel.org/?p=git/warthog9/gitweb.git;a=summary\n\nor\n\ngit://git.kernel.org/pub/scm/git/warthog9/gitweb.git\n\nMind you my performance on the non-cache state is not going to be any\nbetter than normal gitweb, however the performance on a cache-hit is\norders of magnitude faster - though at a rather expensive cost - disk\nspace.  There is currently something like 20G of disk being used on one\nof kernel.org's machines providing the cache (this does get flushed on\noccasion - I think) but that is providing caching for everything that\nkernel.org has in it's git trees (or 255188 unique urls currently).  My\ncode base is now, horribly, out of date with respect to mainline but it\nworks and it's been solid and reasonably reliable (though I do know of\ntwo bugs in it right now I need to track down - one with respect to a\nfailure of the script - and one that is an array out of bounds error)\n\n- John\n"},{"id":"68637","messageId":"200802131317.48815.jnareb@gmail.com","threadId":"12001","inReplyTo":"1202864493.17207.24.camel@localhost.localdomain","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-13T12:17:46Z","receivedAt":"2008-02-13T12:17:46Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Wed, 13 Feb 2008, J.H. \"Warthog9\" wrote:\n> Bruno Cesar Ribas <ribas@c3sl.ufpr.br> writes:\n>> On Mon, Feb 11, 2008 at 04:44:23PM -0800, Jakub Narebski wrote:\n\n>>> Kernel.org even run modified version of gitweb, with some caching\n>>> support; Cgit (git web interface in C) also has caching support.\n>> \n>> Is this gitweb version for kernel.org available somewhere?\n>> \n> \n> It's available from my git tree on kernel.org\n> http://git.kernel.org/?p=git/warthog9/gitweb.git;a=summary\n> \n> or\n> \n> git://git.kernel.org/pub/scm/git/warthog9/gitweb.git\n> \n> Mind you my performance on the non-cache state is not going to be any\n> better than normal gitweb, however the performance on a cache-hit is\n> orders of magnitude faster - though at a rather expensive cost - disk\n> space.  There is currently something like 20G of disk being used on one\n> of kernel.org's machines providing the cache (this does get flushed on\n> occasion - I think) but that is providing caching for everything that\n> kernel.org has in it's git trees (or 255188 unique urls currently).  My\n> code base is now, horribly, out of date with respect to mainline but it\n> works and it's been solid and reasonably reliable (though I do know of\n> two bugs in it right now I need to track down - one with respect to a\n> failure of the script - and one that is an array out of bounds error)\n\nBTW. did you consider using cgit (C/Caching git web interface) instead\nor in addition to gitweb? Freedesktop.org uses it side by side with\ngitweb. I wonder how it would perform on kernel.org...\n\n(Almost) every optimization should begin with profiling. Could you tell\nus which gitweb pages are most called and perhaps which pages generate\nmost load for kernel.org? How new projects are added (old projects\ndeleted)? Do you control (can add to or can add multiplexing) to update\nor post-receive hooks?\n\nWithout this data we could concentrate on things which are of no\nimportance. BTW. I wonder if slitting projects_list page would help...\n\n-- \nJakub Narebski\nPoland\n"},{"id":"68669","messageId":"1202929923.2687.15.camel@localhost.localdomain","threadId":"12001","inReplyTo":"200802131317.48815.jnareb@gmail.com","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"J.H.","fromEmail":"warthog9@kernel.org","sentAt":"2008-02-13T19:12:03Z","receivedAt":"2008-02-13T19:12:03Z","isPatch":false,"sender":{"key":"warthog9@kernel.org","avatar":"https://avatars.githubusercontent.com/u/2334704?v=4"},"body":"\n> BTW. did you consider using cgit (C/Caching git web interface) instead\n> or in addition to gitweb? Freedesktop.org uses it side by side with\n> gitweb. I wonder how it would perform on kernel.org...\n\nWhen I branched and did the initial work for gitweb-caching CGit had\nonly barely made verion 0.01.  So putting something *that* new into\nproduction on Kernel.org didn't even remotely make sense.  Since than\nthe caching modifications (along with a few other fixes and such) have\nproven to be quite stable and have withstood the onslaught of users\nfairly well.  I have toyed with the idea of giving up on gitweb-caching\n(since I either need to redo it to bring it closer to mainline gitweb,\nand probably give up on breaking it up into multiple files or switch to\nsomething new) but the current question that I, and no one else on the\nkernel.org admin staff has had time to investigate is does cgit use the\nsame url paths.  If so it would be a simple drop-in replacement and that\nwould appeal to us.  If it doesn't we can't use cgit and will have to\nstick with gitweb or a direct derivative there-of.\n\n> \n> (Almost) every optimization should begin with profiling. Could you tell\n> us which gitweb pages are most called and perhaps which pages generate\n> most load for kernel.org?\n\nThat would be correct, though when I did up gitweb-caching the profiling\nwas blatantly obvious, with every single page request git was being\ncalled, git was hammering the disk and it was becoming increasingly\nobvious that going and running git for every page load was completely\nimpractical.  I know git is fast - but it's not *that* fast, and it is a\nbit abusive to the system for certain things requiring a lot of memory,\nchewing cpu or chewing disk.  In order of badness for kernel.org:\nchewing memory, disk, cpu.  Use up too much memory and you force too\nmuch needed content out of ram, chew up disk and you make queries that\nare forced to disk to take longer (and if you've chewed up too much ram\nthis gets *lots* worse).\n\nSo the simplest and obvious thing - take git out of the equation,\ndirectly, for most calls.  If you've ever seen the \"Generating...\" page\non kernel.org that is a stalling mechanism I'm using to let git run in\nthe background and generate the page your going to see.  If you notice\nit can take several seconds for that to complete, and we are on *very*\nfast boxes - now multiply that by hundreds of times a second and you'll\nstart to understand why the caching layer is saving us right now.\n\nAs for the most often hit pages - the front page is by far hit the most,\nwhich should be no surprise to everyone, and it is by far the most\nabusive since it has to query *every* project have.  After that things\ntaper off as people find the project they want and go looking for the\ndata they are interested in.\n\n>  How new projects are added (old projects\n> deleted)?\n\nBy and large - left up to the users - if they don't want their tree\nanymore they delete it (though I don't know of anyone who has) if they\nneed another one - they create it.\n\n>  Do you control (can add to or can add multiplexing) to update\n> or post-receive hooks?\n\nNo.  We do not want to, at all, control in any way the tree's that\npeople put up on Kernel.org.  We just don't have the bandwidth to deal\nwith that for every single tree on kernel.org.  Anything that would\nrequire us to go changing or forcing a user to change something in their\ngit tree means we've already lost.  Taking the caching layer and making\nit 100% transparent to the git tree's owners and generally speaking to\nthe end user makes things very simple for us to deal with.\n\n> \n> Without this data we could concentrate on things which are of no\n> importance. BTW. I wonder if slitting projects_list page would help...\n\nThat would be bad - I know for a fact there are people who will go to\ngit.kernel.org and then search on the page for the things they want - so\nchanging this would probably cause a lot of confusion for minor gain at\nthis point.\n\n- John 'Warthog9' Hawley\n"},{"id":"68691","messageId":"200802140201.55420.jnareb@gmail.com","threadId":"12001","inReplyTo":"1202929923.2687.15.camel@localhost.localdomain","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-14T01:01:53Z","receivedAt":"2008-02-14T01:01:53Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Wed, 13 Feb 2008, J.H. \"Warthog9\" wrote:\n> On Wed, 13 Feb 2008 13:17:46 +0100, Jakub Narebski wrote:\n>> \n>> BTW. did you consider using cgit (C/Caching git web interface) instead\n>> or in addition to gitweb? Freedesktop.org uses it side by side with\n>> gitweb. I wonder how it would perform on kernel.org...\n> \n> When I branched and did the initial work for gitweb-caching CGit had\n> only barely made verion 0.01.  So putting something *that* new into\n> production on Kernel.org didn't even remotely make sense.\n\nIf I remember correctly cgit was _created_ in response (or at least\naround) to discussion on git mailing list that kerne.org needs caching\nfor gitweb.\n\nBTW. I have CC-ed CGit author, Lars Hjemli.\n\n> Since than \n> the caching modifications (along with a few other fixes and such) have\n> proven to be quite stable and have withstood the onslaught of users\n> fairly well.  I have toyed with the idea of giving up on gitweb-caching\n> (since I either need to redo it to bring it closer to mainline gitweb,\n> and probably give up on breaking it up into multiple files or switch to\n> something new)\n\nBy the way, why did you split into so many modules? I would think\nthat separating into generic modules (like for example commit parsing),\nHTML generation modules (specifing to gitweb) and caching module;\nperhaps for easier integration only main gitweb (core version)\nand caching module would be enough.\n\nI was thinking about adding caching, using code from your fork,\nto git.git gitweb, somewhere along the line... I guess I can move\nit earlier in the TODO list, somewhere along CSS cleanup and log-like\nviews generation cleanup, and using feed links depending on page.\n\n> but the current question that I, and no one else on the \n> kernel.org admin staff has had time to investigate is does cgit use the\n> same url paths.  If so it would be a simple drop-in replacement and that\n> would appeal to us.  If it doesn't we can't use cgit and will have to\n> stick with gitweb or a direct derivative there-of.\n\nUnfortunately cgit is not designed to be gitweb compatibile; it is\nsimplier (which might be considered better), doesn't support the\nmultitude of gitweb views, and unfortunately doesn't understand\ngitweb URLs.\n\nOn the other side having gitweb and cgit coexist together on the same\nset of repositories should be quite easy, as shown by FreeDesktop folks:\nhttp://gitweb.freedesktop.org and http://cgit.freedesktop.org\n \n>> (Almost) every optimization should begin with profiling. Could you tell\n>> us which gitweb pages are most called and perhaps which pages generate\n>> most load for kernel.org?\n> \n> That would be correct, though when I did up gitweb-caching the profiling\n> was blatantly obvious, with every single page request git was being\n> called, git was hammering the disk and it was becoming increasingly\n> obvious that going and running git for every page load was completely\n> impractical.\n\n[bringing back old quote]\n>> On Wed, 13 Feb 2008, J.H. \"Warthog9\" wrote:\n>>\n>>>  There is currently something like 20G of disk being used on one\n>>> of kernel.org's machines providing the cache (this does get flushed on\n>>> occasion - I think) but that is providing caching for everything that\n>>> kernel.org has in it's git trees (or 255188 unique urls currently).\n\nWhat I meant here that one should balance between not having cache\nat all (and spending all the CPU), to caching everything under the\nsun (and spending all the HDD space). To that one should know which\npages are most requested, so they would be cached by having static\npage to serve; which generate most load / are less requested, so\nperhaps git commands output would be cached; and which are rare enough\nthat caching is waste of disk space, and hints to caching proxies\nshould be enough.\n\n[...]\n> As for the most often hit pages - the front page is by far hit the most,\n> which should be no surprise to everyone, and it is by far the most\n> abusive since it has to query *every* project have.  After that things\n> taper off as people find the project they want and go looking for the\n> data they are interested in.\n\nBut what of those pages are most requested and generate most load?\n'summary' page? 'rss' or 'atom' feeds? 'tree' view? README 'blob'?\nsnapshot (if enabled)?\n \n>>  How new projects are added (old projects\n>> deleted)?\n> \n> By and large - left up to the users - if they don't want their tree\n> anymore they delete it (though I don't know of anyone who has) if they\n> need another one - they create it.\n\nBummer. If projects were created by some script (like I think is\nthe case for git hosting facilities, like repo.or.cz, GitHub,\nGitorious or TucFamily) we could update projects listing file\n(so gitweb doesn't need to scan directories), and perhaps even\nadd some gitweb-specific hooks (add multiplexer + hooks).\n\n>>  Do you control (can add to or can add multiplexing) to update\n>> or post-receive hooks?\n> \n> No.  We do not want to, at all, control in any way the tree's that\n> people put up on Kernel.org.  We just don't have the bandwidth to deal\n> with that for every single tree on kernel.org.  Anything that would\n> require us to go changing or forcing a user to change something in their\n> git tree means we've already lost.  Taking the caching layer and making\n> it 100% transparent to the git tree's owners and generally speaking to\n> the end user makes things very simple for us to deal with.\n\nThat's bad, because update / post-recive hook could be used for example\nto invalidate 'summary', 'rss' and 'atom' views cache, and perhaps\nregenerate projects list page. First request would generate cache, which\nwould be then used till it was deleted by the hook script.\n \n>> \n>> Without this data we could concentrate on things which are of no\n>> importance. BTW. I wonder if slitting projects_list page would help...\n> \n> That would be bad - I know for a fact there are people who will go to\n> git.kernel.org and then search on the page for the things they want - so\n> changing this would probably cause a lot of confusion for minor gain at\n> this point.\n\nI was thinking about first page being page of categories, perhaps with\n\"search projects\" box. The page with so many projects is a bit unwieldy.\n\nP.S. Do you make use of alternates, or do you left it to users to setup.?\n-- \nJakub Narebski\nPoland\n"},{"id":"68776","messageId":"1203029034.4766.19.camel@localhost.localdomain","threadId":"12001","inReplyTo":"200802140201.55420.jnareb@gmail.com","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"J.H.","fromEmail":"warthog9@kernel.org","sentAt":"2008-02-14T22:43:53Z","receivedAt":"2008-02-14T22:43:53Z","isPatch":false,"sender":{"key":"warthog9@kernel.org","avatar":"https://avatars.githubusercontent.com/u/2334704?v=4"},"body":"On Thu, 2008-02-14 at 02:01 +0100, Jakub Narebski wrote: \n> On Wed, 13 Feb 2008, J.H. \"Warthog9\" wrote:\n> > On Wed, 13 Feb 2008 13:17:46 +0100, Jakub Narebski wrote:\n> >> \n> >> BTW. did you consider using cgit (C/Caching git web interface) instead\n> >> or in addition to gitweb? Freedesktop.org uses it side by side with\n> >> gitweb. I wonder how it would perform on kernel.org...\n> > \n> > When I branched and did the initial work for gitweb-caching CGit had\n> > only barely made verion 0.01.  So putting something *that* new into\n> > production on Kernel.org didn't even remotely make sense.\n> \n> If I remember correctly cgit was _created_ in response (or at least\n> around) to discussion on git mailing list that kerne.org needs caching\n> for gitweb.\n> \n> BTW. I have CC-ed CGit author, Lars Hjemli.\n\nI didn't realize it had come out of those discussions a couple of years\nago, though it's good to hear that it did come out.  From what I have\nused of it, it does seem to be as fast as gitweb-caching and it's\ncaching doesn't seem to be quite the sledge hammer that mine is.\n\n> \n> > Since than \n> > the caching modifications (along with a few other fixes and such) have\n> > proven to be quite stable and have withstood the onslaught of users\n> > fairly well.  I have toyed with the idea of giving up on gitweb-caching\n> > (since I either need to redo it to bring it closer to mainline gitweb,\n> > and probably give up on breaking it up into multiple files or switch to\n> > something new)\n> \n> By the way, why did you split into so many modules? I would think\n> that separating into generic modules (like for example commit parsing),\n> HTML generation modules (specifing to gitweb) and caching module;\n> perhaps for easier integration only main gitweb (core version)\n> and caching module would be enough.\n\nI did it originally for a couple of reasons:\n\n1) a single script that's almost 6000 lines long is a little hard to\nhandle in a single gulp.  I can appreciate why it's being done that way,\nmainly to simplify installation, but...\n\n2) ... I needed something to help me understand the flow of code more -\nripping it apart was a good way to do it and try and group similar\nfunctions.\n\nThere is an obvious downside, in retrospect, to me doing it this way - I\ncan't track upstream *nearly* as easily as I would like.  Which means\nthat the code gitweb-caching is based on is about a year and a half old\n(it has had a major update along the way but that took me two days of\nmanually applying patches to get updated)\n\n> \n> I was thinking about adding caching, using code from your fork,\n> to git.git gitweb, somewhere along the line... I guess I can move\n> it earlier in the TODO list, somewhere along CSS cleanup and log-like\n> views generation cleanup, and using feed links depending on page.\n\nIt's on my todo list as well, basically re-base to current head and\ninstead of breaking stuff up like I did in the original, completely redo\nthe tree so that I can pull from trunk easier.  Considering I'll have\nsome time that I can spend on my OSS projects in the next couple of\nweeks again - I was going to try and get this accomplished and back into\nmy tree.\n\n> \n> > but the current question that I, and no one else on the \n> > kernel.org admin staff has had time to investigate is does cgit use the\n> > same url paths.  If so it would be a simple drop-in replacement and that\n> > would appeal to us.  If it doesn't we can't use cgit and will have to\n> > stick with gitweb or a direct derivative there-of.\n> \n> Unfortunately cgit is not designed to be gitweb compatibile; it is\n> simplier (which might be considered better), doesn't support the\n> multitude of gitweb views, and unfortunately doesn't understand\n> gitweb URLs.\n\nSadly - that makes it more or less unusable to us at this point.  There\nare, I'm sure, a number of links that point back to kernel.org that\nwould be nice to not break - if it's really felt that this isn't the\ncase I would consider switching, but if gitweb + my caching code works\nit might be better for me to try and get the code merged back into trunk\nand not risk changing infrastructure that's been in place for several\nyears now.\n\n> \n> On the other side having gitweb and cgit coexist together on the same\n> set of repositories should be quite easy, as shown by FreeDesktop folks:\n> http://gitweb.freedesktop.org and http://cgit.freedesktop.org\n> \n\nI'm not a big fan of maintaining multiple things that all accomplish\nthe same goal.  I would rather devote the resources to maintaining one\non kernel.org vs. maintaining two.  This prevents users from getting\nconfused and in the event of upgrades forgetting that one needs\nupdating vs. the other one.\n\n>  \n> >> (Almost) every optimization should begin with profiling. Could you tell\n> >> us which gitweb pages are most called and perhaps which pages generate\n> >> most load for kernel.org?\n> > \n> > That would be correct, though when I did up gitweb-caching the profiling\n> > was blatantly obvious, with every single page request git was being\n> > called, git was hammering the disk and it was becoming increasingly\n> > obvious that going and running git for every page load was completely\n> > impractical.\n> \n> [bringing back old quote]\n> >> On Wed, 13 Feb 2008, J.H. \"Warthog9\" wrote:\n> >>\n> >>>  There is currently something like 20G of disk being used on one\n> >>> of kernel.org's machines providing the cache (this does get flushed on\n> >>> occasion - I think) but that is providing caching for everything that\n> >>> kernel.org has in it's git trees (or 255188 unique urls currently).\n> \n> What I meant here that one should balance between not having cache\n> at all (and spending all the CPU), to caching everything under the\n> sun (and spending all the HDD space). To that one should know which\n> pages are most requested, so they would be cached by having static\n> page to serve; which generate most load / are less requested, so\n> perhaps git commands output would be cached; and which are rare enough\n> that caching is waste of disk space, and hints to caching proxies\n> should be enough.\n\nTo a greater extent, disk is cheap, and load is expensive.  While I\nagree it would be worth spending cpu time to not cache things - thats\nnot the way gitweb / git works.  Gitweb forces calls down to disk if\nit's displaying the index page or showing a diff between two\nrevisions.  The problem comes in that git, while fast, is very resource\nintensive, requiring a lot of ram and a lot of disk seeking - on a very\nbusy setup, disk seeking is a killer.  What you want to be able to do\nis take a file and just run it straight out, not have to poke around to\nre-assemble everything and then output.  This is one of the reasons why\nthe index page is so painful - it's reassembling bits from *every*\nrepository.  It's not cpu that's getting hit, it's the disk and ram\nthat's getting hit.\n\nSo yes - while I agree caching can be expensive, in very high volume\nsetups you have to have caching of some sort.  Caching proxies aren't\nsmart enough for something like gitweb either, having a caching layer\ndirectly in gitweb makes a lot more sense - you know what pages need\ncaching, which ones don't, what may be hit harder than others, what you\ncan have a longer cache timeout, etc on.\n\n\n> [...]\n> > As for the most often hit pages - the front page is by far hit the most,\n> > which should be no surprise to everyone, and it is by far the most\n> > abusive since it has to query *every* project have.  After that things\n> > taper off as people find the project they want and go looking for the\n> > data they are interested in.\n> \n> But what of those pages are most requested and generate most load?\n> 'summary' page? 'rss' or 'atom' feeds? 'tree' view? README 'blob'?\n> snapshot (if enabled)?\n\nRight now I don't have explicit statistics on that, though it wouldn't\nbe hard to add in an additional file or small database of some sort\nthat would track generation times.  My gut feeling is the index page is\nthe worst (particularly with the number of trees we have), followed by\nthe summary pages, and from there things will fall off dramatically as\nmost pages after that may not get hit often.\n\n   \n> >>  How new projects are added (old projects\n> >> deleted)?\n> > \n> > By and large - left up to the users - if they don't want their tree\n> > anymore they delete it (though I don't know of anyone who has) if they\n> > need another one - they create it.\n> \n> Bummer. If projects were created by some script (like I think is\n> the case for git hosting facilities, like repo.or.cz, GitHub,\n> Gitorious or TucFamily) we could update projects listing file\n> (so gitweb doesn't need to scan directories), and perhaps even\n> add some gitweb-specific hooks (add multiplexer + hooks).\n\nAt this point the git tree is left up to the user and we have no\nintention of changing it, we don't even force them to turn on the\npost-update hook that will deal with git-update-server-info.\n\n> \n> >>  Do you control (can add to or can add multiplexing) to update\n> >> or post-receive hooks?\n> > \n> > No.  We do not want to, at all, control in any way the tree's that\n> > people put up on Kernel.org.  We just don't have the bandwidth to deal\n> > with that for every single tree on kernel.org.  Anything that would\n> > require us to go changing or forcing a user to change something in their\n> > git tree means we've already lost.  Taking the caching layer and making\n> > it 100% transparent to the git tree's owners and generally speaking to\n> > the end user makes things very simple for us to deal with.\n> \n> That's bad, because update / post-recive hook could be used for example\n> to invalidate 'summary', 'rss' and 'atom' views cache, and perhaps\n> regenerate projects list page. First request would generate cache, which\n> would be then used till it was deleted by the hook script.\n\nThey could be useful, but this is completely left up to the tree's\nowner, we provide a location for them to publish their trees, we don't\nwant to control or limit how or what they do with those trees.\n\n>  \n> >> \n> >> Without this data we could concentrate on things which are of no\n> >> importance. BTW. I wonder if slitting projects_list page would help...\n> > \n> > That would be bad - I know for a fact there are people who will go to\n> > git.kernel.org and then search on the page for the things they want - so\n> > changing this would probably cause a lot of confusion for minor gain at\n> > this point.\n> \n> I was thinking about first page being page of categories, perhaps with\n> \"search projects\" box. The page with so many projects is a bit unwieldy.\n> \n> P.S. Do you make use of alternates, or do you left it to users to setup.?\n\nLeft up to the users, we suggest it's use but I'm sure there are trees\non kernel.org that could use alternates but don't.\n\n- John 'Warthog9' Hawley\n"},{"id":"68852","messageId":"200802160019.09490.jnareb@gmail.com","threadId":"12001","inReplyTo":"1203029034.4766.19.camel@localhost.localdomain","subject":"Re: Another bench on gitweb (also on gitweb caching)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-15T23:19:08Z","receivedAt":"2008-02-15T23:19:08Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Thu, 14 Feb 2008, J.H. wrote:\n> On Thu, 2008-02-14 at 02:01 +0100, Jakub Narebski wrote: \n>> On Wed, 13 Feb 2008, J.H. \"Warthog9\" wrote:\n>>> On Wed, 13 Feb 2008 13:17:46 +0100, Jakub Narebski wrote:\n>> \n>> By the way, why did you split into so many modules? I would think\n>> that separating into generic modules (like for example commit parsing),\n>> HTML generation modules (specifing to gitweb) and caching module;\n>> perhaps for easier integration only main gitweb (core version)\n>> and caching module would be enough.\n> \n> I did it originally for a couple of reasons:\n> \n> 1) a single script that's almost 6000 lines long is a little hard to\n> handle in a single gulp.  I can appreciate why it's being done that way,\n> mainly to simplify installation, but...\n> \n> 2) ... I needed something to help me understand the flow of code more -\n> ripping it apart was a good way to do it and try and group similar\n> functions.\n\nIf I remember correctly the \"great renaming\", renaming subroutines\nand a bit of code restructurization (moving subroutines so similar\nsubroutines are together) came later.\n\n> There is an obvious downside, in retrospect, to me doing it this way - I\n> can't track upstream *nearly* as easily as I would like.  Which means\n> that the code gitweb-caching is based on is about a year and a half old\n> (it has had a major update along the way but that took me two days of\n> manually applying patches to get updated)\n> \n>> I was thinking about adding caching, using code from your fork,\n>> to git.git gitweb, somewhere along the line... I guess I can move\n>> it earlier in the TODO list, somewhere along CSS cleanup and log-like\n>> views generation cleanup, and using feed links depending on page.\n> \n> It's on my todo list as well, basically re-base to current head and\n> instead of breaking stuff up like I did in the original, completely redo\n> the tree so that I can pull from trunk easier.  Considering I'll have\n> some time that I can spend on my OSS projects in the next couple of\n> weeks again - I was going to try and get this accomplished and back into\n> my tree.\n\nBetter you than me. I don't know if I'd have time for it, and obviosly\nI don't know as well the caching code.\n \n>> Unfortunately cgit is not designed to be gitweb compatibile; it is\n>> simplier (which might be considered better), doesn't support the\n>> multitude of gitweb views, and unfortunately doesn't understand\n>> gitweb URLs.\n> \n> Sadly - that makes it more or less unusable to us at this point.  There\n> are, I'm sure, a number of links that point back to kernel.org that\n> would be nice to not break - if it's really felt that this isn't the\n> case I would consider switching, but if gitweb + my caching code works\n> it might be better for me to try and get the code merged back into trunk\n> and not risk changing infrastructure that's been in place for several\n> years now.\n\nTruly, it would be nice if cgit had compatibility mode, accepting gitweb\n(or gitweb-like) URLs, and returning similar page. Or at least\nmod_rewrite rules to rewrite gitweb URLs to equivalent cgit ones.\n\n>> [...] one should balance between not having cache\n>> at all (and spending all the CPU), to caching everything under the\n>> sun (and spending all the HDD space). To that one should know which\n>> pages are most requested, so they would be cached by having static\n>> page to serve; which generate most load / are less requested, so\n>> perhaps git commands output would be cached; and which are rare enough\n>> that caching is waste of disk space, and hints to caching proxies\n>> should be enough.\n> \n> To a greater extent, disk is cheap, and load is expensive.  While I\n> agree it would be worth spending cpu time to not cache things - thats\n> not the way gitweb / git works.  Gitweb forces calls down to disk if\n> it's displaying the index page or showing a diff between two\n> revisions.\n\nWhat I mean here was to have caching in gitweb for pages like projects\nlist, summary for a project, main RSS/Atom feed for a project on one\nhand, but just adding Last-Modified: and (weak?) ETag: headers plus\na year expire (or was it half of a year be infinity according to RFC?)\nfor immutable rarely (I think) accessed pages, like 'blob' view of given\nfile at given revision, or 'commit' view, or even perhaps 'tree' view\n(all for given by immutable sha-1 revision / object-id).\n\nAnd in the middle ground we could habe saving git command output in\na kind if \"git cache\", like storing update / change time for each\nproject; another example would be incremental blame output.\n\n> The problem comes in that git, while fast, is very resource \n> intensive, requiring a lot of ram and a lot of disk seeking - on a very\n> busy setup, disk seeking is a killer.  What you want to be able to do\n> is take a file and just run it straight out, not have to poke around to\n> re-assemble everything and then output.  This is one of the reasons why\n> the index page is so painful - it's reassembling bits from *every*\n> repository.  It's not cpu that's getting hit, it's the disk and ram\n> that's getting hit.\n\nTrue, the projects list especially with so large number of projects\njust have to be cached. It is a pity that due to histerical raisings^W^W\nhistorical reasons called backwards compatibility we cannot just put\nprojects search page, or projects catalogue (projects divided into\ncategories) instead of listing of all projects. Or at least divide\nprojects list in pages... Though the last option is not that good,\nunless you somehow can include most commonly requested projects\non the main page.\n\n> So yes - while I agree caching can be expensive, in very high volume\n> setups you have to have caching of some sort.  Caching proxies aren't\n> smart enough for something like gitweb either, having a caching layer\n> directly in gitweb makes a lot more sense - you know what pages need\n> caching, which ones don't, what may be hit harder than others, what you\n> can have a longer cache timeout, etc on.\n\nI can agree with that.\n\nOn the other hand you are duplicating effort what is already done:\nselecting what to cache, when to invalidate cache, when to prune / purge\ncache etc. But I guess that having gitweb provide hints to caching\nengines in the form of Last-Modified: and ETag:, and responding quickly\nto If-Modified-Since: and If-None-Match: requests / HEAD requests\nmight be not enough; and responding to If-* requests might be not easy;\nwell, not easier than implementing caching inside gitweb.\n \n>> [...]\n>>> As for the most often hit pages - the front page is by far hit the most,\n>>> which should be no surprise to everyone, and it is by far the most\n>>> abusive since it has to query *every* project have.  After that things\n>>> taper off as people find the project they want and go looking for the\n>>> data they are interested in.\n>> \n>> But what of those pages are most requested and generate most load?\n>> 'summary' page? 'rss' or 'atom' feeds? 'tree' view? README 'blob'?\n>> snapshot (if enabled)?\n> \n> Right now I don't have explicit statistics on that, though it wouldn't\n> be hard to add in an additional file or small database of some sort\n> that would track generation times.  \n\nProfiling websites. Debugging websites. I don't think it is easy...\n\n> My gut feeling is the index page is \n> the worst (particularly with the number of trees we have), followed by\n> the summary pages, and from there things will fall off dramatically as\n> most pages after that may not get hit often.\n\nWhat about RSS feeds (as compared to summary page for example)?\n    \n>>>>  How new projects are added (old projects\n>>>> deleted)?\n>>> \n>>> By and large - left up to the users - if they don't want their tree\n>>> anymore they delete it (though I don't know of anyone who has) if they\n>>> need another one - they create it.\n>> \n>> Bummer. If projects were created by some script (like I think is\n>> the case for git hosting facilities, like repo.or.cz, GitHub,\n>> Gitorious or TucFamily) we could update projects listing file\n>> (so gitweb doesn't need to scan directories), and perhaps even\n>> add some gitweb-specific hooks (add multiplexer + hooks).\n> \n> At this point the git tree is left up to the user and we have no\n> intention of changing it, we don't even force them to turn on the\n> post-update hook that will deal with git-update-server-info.\n[...]\n> They could be useful, but this is completely left up to the tree's\n> owner, we provide a location for them to publish their trees, we don't\n> want to control or limit how or what they do with those trees.\n\nDo the project list on kernel.org is then generated by scanning\nfilesystem ($projects_list unset, or set to directory)?\n\n\nI have asked about this because with projects added, renamed and\ndeleted by script you can generate / regenerate projects list file\nwhen chaning a project. I guess that you could in theory watch\nfilesystem for that...\n\nIf you can add gitweb's update / post-receive hook you would be able\nto update file with \"last changed\" information, and delete or regenerate\ncaches for output which depend on the tip of current branch: project\nsummary, RSS feeds etc.\n\nBut if it is truly \"no can do\", then you have to implement storing\ncache and invalidating caches yourself...\n\n-- \nJakub Narebski\nPoland\n"}]}