{"thread":{"id":"13694","subject":"Gitweb caching: Google Summer of Code project","startedAt":"2008-05-27T18:03:43Z","lastAt":"2008-05-31T10:15:34Z","messageCount":18,"participants":["Lea Wiemann","Jakub Narebski","Petr Baudis","Rafael Garcia-Suarez","J.H.","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"77875","messageId":"483C4CFF.2070101@gmail.com","threadId":"13694","inReplyTo":null,"subject":"Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-27T18:03:43Z","receivedAt":"2008-05-27T18:03:43Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Hi everyone,\n\nI just wanted to let everyone know that I'm currently getting started on \na Google Summer of Code project to improve the caching mechanism in gitweb.\n\nSorry for not posting about this earlier...  Anyways, some key data: \nJohn 'warthog9' Hawley (who wrote the current caching system for \nkernel.org) is my mentor, and GSoC is from May 26 to Aug 18, minus a \nvacation from Jul 19 to Aug 9.\n\nWhile I'm planning to keep much of it on the list, if anyone else is \nparticularly interested in helping or providing input, please notify me. \n  (Looking at the logs, Jakub maybe?  Cc'ing him just in case.)\n\nThe current plan is basically to get the gitweb caching fork that's been \nimplemented for kernel.org back to the gitweb mainline, and then \noptimize it (probably move to memcached).  I'm not yet sure how to \napproach this (e.g. whether to merge from the fork to the mainline or \nvice versa), but I'll probably figure this out together with John and \nmight post separately about that later.  In any case, expect patches and \nmessages from me on the list. :)\n\nI'm lea_w (or lea_1) on #git on Freenode, if anyone wants to contact me \nin real time (provided my Pidgin doesn't hiccup).\n\nBest,\n\n     Lea\n"},{"id":"77895","messageId":"200805272353.34319.jnareb@gmail.com","threadId":"13694","inReplyTo":"483C4CFF.2070101@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-05-27T21:53:32Z","receivedAt":"2008-05-27T21:53:32Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 27 May 2008, Lea Wiemann wrote:\n\n> I just wanted to let everyone know that I'm currently getting started on \n> a Google Summer of Code project to improve the caching mechanism in gitweb.\n> \n> Sorry for not posting about this earlier...  Anyways, some key data: \n> John 'warthog9' Hawley (who wrote the current caching system for \n> kernel.org) is my mentor, and GSoC is from May 26 to Aug 18, minus a \n> vacation from Jul 19 to Aug 9.\n\nThanks for the info.\n\n> While I'm planning to keep much of it on the list, if anyone else is \n> particularly interested in helping or providing input, please notify me. \n>   (Looking at the logs, Jakub maybe?  Cc'ing him just in case.)\n\nI'm certainly interested, at least from theoretical point of view, and\nI think I can help (as one of main gitweb contributors).\n\nI guess that Petr Baudis would also be interested, because he maintains\nrepo.or.cz, a public Git hosting site.  Lately he posted a patch\nimplementing projects list caching, in a bit different way from how it\nis done on kernel.org, namely by caching data and not final output:\n  http://thread.gmane.org/gmane.comp.version-control.git/77151\nAFAIK it is implemented in repo.or.cz gitweb: \n  http://repo.or.cz/w/git/repo.git\n\nThis indirectly lead to a bit of research on caching in Perl by yours\ntruly:\n  http://thread.gmane.org/gmane.comp.version-control.git/77529\n(mentioned in http://git.or.cz/gitwiki/SoC2008Projects#gitweb-caching).\n\n\nI think that you can also get some help on caching from Lars Hjemli,\nauthor of cgit, which is caching git web interface written in C.\n\n\n(I have added both Petr Baudis and Lars Hjemli to Cc:)\n\n> The current plan is basically to get the gitweb caching fork that's been \n> implemented for kernel.org back to the gitweb mainline, and then \n> optimize it (probably move to memcached).  I'm not yet sure how to \n> approach this (e.g. whether to merge from the fork to the mainline or \n> vice versa), but I'll probably figure this out together with John and \n> might post separately about that later.  In any case, expect patches and \n> messages from me on the list. :)\n\n>From what I remember correcly from the discussion surrounding\nimplementing caching for kernel.org gitweb, the main culprit of having\nit remain separate from mainline was splitting gitweb into many, many\nfiles.  While it helped John in understanding gitweb, it made it\ndifficult to merge changes back to mainline.\n\nNote also that it is easier to make a site-specific changes, than to\nmake generic, closs-platform and cross-operating system change.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"77900","messageId":"483C912F.6010802@gmail.com","threadId":"13694","inReplyTo":"200805272353.34319.jnareb@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-27T22:54:39Z","receivedAt":"2008-05-27T22:54:39Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Jakub Narebski wrote:\n> Lately he posted a patch\n> implementing projects list caching, in a bit different way from how it\n> is done on kernel.org, namely by caching data and not final output:\n\nThanks for this and all the other pointers.\n\nCaching data and not final output is actually what I'm about to try \nnext.  If I'm not mistaken, the HTML output is significantly larger than \nthe source (repository) data; however, kernel.org still seems to benefit \nfrom caching the HTML, rather than letting Linux' page cache cache the \nsource data.  That leads me to think that the page cache somehow fails \nto cache the source data properly -- I'm not sure why (wild speculation: \nperhaps because of the pack format).  Anyway, I'd hope that I can \nencapsulate the 30-40 git_cmd calls in gitweb.perl and somehow cache \ntheir results (or, to save memory, the parts of their results that are \nactually used) and cache them using memcached.  If that works well, we \ncan stop bothering about frontend (HTML) caching, unless CPU becomes an \nissue, since all HTML pages are generated from cacheable source data.\n\nI'm *kindof* hoping that in the end there will be only few issues with \ncache expiry, since most calls are uniquely identified through hashes. \n(And the ones that are not, like getting the hash of the most recent \ncommit, can perhaps be cached with some fairly low expiry time.)\n\nSo that's what I'll try next.  If you have any comments or warnings off \nthe top of your heads, feel free to send email of course. :)\n\n> the main culprit of [the fork] was splitting gitweb into many, many\n> files.  While it helped John in understanding gitweb, it made it\n> difficult to merge changes back to mainline.\n\nInteresting point, thanks for letting me know.  (I might have gone ahead \nand tried to split the mainline gitweb myself... ^^)  I think it would \nbe nice if gitweb.perl could be split at some point, but I assume there \nare too many patches out there for that to be worth the merge problems, \nright?\n\n-- Lea\n"},{"id":"77930","messageId":"200805281414.36141.jnareb@gmail.com","threadId":"13694","inReplyTo":"483C912F.6010802@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-05-28T12:14:35Z","receivedAt":"2008-05-28T12:14:35Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Wed, 28 May 2008, Lea Wiemann wrote:\n> Jakub Narebski wrote:\n>>\n>> Lately he posted a patch\n>> implementing projects list caching, in a bit different way from how it\n>> is done on kernel.org, namely by caching data and not final output:\n> \n> Thanks for this and all the other pointers.\n> \n> Caching data and not final output is actually what I'm about to try \n> next.\n\nCaching data have its advantages and disadvantages, same as with\ncaching HTML output (or parts of HTML output). I have wrote about\nit in\n  http://thread.gmane.org/gmane.comp.version-control.git/77529\n\nLet me summarize here advantages and disadvantages of caching data\nand of caching HTML output.\n\n1. Caching data\n * advantages:\n   - smaller than caching HTML output\n   - you can use the same data to generate different pages\n     ('summary', 'shortlog', 'log', 'rss'/'atom'; pages or search\n      results of projects list)\n   - you can generate pages with variable data, such as relative dates\n     (\"5 minutes ago\"), staleness info (\"cached data, 5 minutes old\"),\n     or content type: text/html vs application/xhtml+xml\n * disadvantages:\n   - more CPU\n   - need to serialize and deserialize (parse) data\n   - more complicated\n\n2. Caching HTML output\n * advantages:\n   - simple, no need for serialization (pay attention to that in mixed\n     data + output caching solutions)\n   - low CPU (although supposedly[1] gitweb performance is I/O bound,\n     and not CPU bound)\n   - web servers deals very well with static pages\n   - web servers deals with support for HTTP caching (giving ETag and\n     Last-Changed headers, responding to If-Modified-Since, \n     If-None-Match etc. headers from web browsers and caching proxies)\n * disadvantages:\n   - large size of cache data (if most clients support compression, you\n     can store it compressed, at the cost of CPU for non-supporting ones).\n   - difficult to impossible variable output (for example you can still\n     rewrite some HTTP headers for text/html vs application/xhtml+xml\n     or store headers separately, you can use JavaScript to change\n     visible times from absolute to relative dates)\n\nI'm sure John, Lars and Petr can tell you more, and have more experience.\n\n[1] Some evidence both from warthog9 and pasky, but no hard data[2]\n[2] I think it would be good to start with analyse of gitweb statictics,\n    e.g. from Apache logs, from kernel.org and repo.or.cz.\n\n> If I'm not mistaken, the HTML output is significantly larger than  \n> the source (repository) data; however, kernel.org still seems to benefit \n> from caching the HTML, rather than letting Linux' page cache cache the \n> source data.\n\nI don't think kernel.org caches _all_ pages, only the most requested\n(correct me if I'm wrong here, John, please).\n\n> That leads me to think that the page cache somehow fails  \n> to cache the source data properly -- I'm not sure why (wild speculation: \n> perhaps because of the pack format).\n\n>From what I remember one of most costly to generate pages is projects\nlist page (that is why Petr Baudis implemented caching for this page\nin repo.or.cz gitweb, using data caching here).  With 1000+ projects\n(repositories) gitweb has to hit at best 1000+ packfiles, not to\nmention refs, to generate \"Last Changed\" column from git-for-each-ref\noutput (accidentally, also to check if it is truly git repository).\nIn kernel.org case with gitweb working similar to mod_userdir module\nbut for git repositories (as a service, rather than as part of repo\nhosting), gitweb has to hit 1000+ 'summary' files...  That is\ninterspersed with other requests.\n\nHow page cache and filesystem buffers can deal with that?\n\nBTW I'm not sure if kernel.org use CGI or \"legacy\" mod_perl gitweb;\ncurently there is no support for FastCGI in gitweb (although you can\nfind some patches in archive).\n\n(But I'm not an expert in those matters, so please take the above\nwith a pinch of salt, or two).\n\n\nBy the way using pack files besides reducing repository size also\nimproved git performance thanks to better I/O performance and better\nworking with filesystem cache (some say that git is optimized for\nwarm cache).\n\n> Anyway, I'd hope that I can  \n> encapsulate the 30-40 git_cmd calls in gitweb.perl and somehow cache \n> their results (or, to save memory, the parts of their results that are \n> actually used) and cache them using memcached.  If that works well, we \n> can stop bothering about frontend (HTML) caching, unless CPU becomes an \n> issue, since all HTML pages are generated from cacheable source data.\n\nI don't think caching _everything_, including rarely requested pages,\nwould be a good idea.\n\n> I'm *kindof* hoping that in the end there will be only few issues with \n> cache expiry, since most calls are uniquely identified through hashes. \n> (And the ones that are not, like getting the hash of the most recent \n> commit, can perhaps be cached with some fairly low expiry time.)\n\nThe trouble is with those requests which are _not_ uniquely identified\nby hashes requested, such as 'summary', 'log' from given branch (not\nfrom given hash), or web feed for given branch.  For those which are\nnot-changing you can just (as gitweb does even now) give large HTTP\nexpiry (Expires or max-age) and allow web browser or proxies to cache\nit.\n\n> So that's what I'll try next.  If you have any comments or warnings off \n> the top of your heads, feel free to send email of course. :)\n\nI'm afraid that implementing kernel.org caching in mainline in\na generic way would be enough work for a whole GSoC 2008.  I hope\nI am mistaken and you would have time to analyse and implement wider\nreange of caching solutions in gitweb...\n \n>> the main culprit of [the fork] was splitting gitweb into many, many\n>> files.  While it helped John in understanding gitweb, it made it\n>> difficult to merge changes back to mainline.\n> \n> Interesting point, thanks for letting me know.  (I might have gone ahead \n> and tried to split the mainline gitweb myself... ^^)  I think it would \n> be nice if gitweb.perl could be split at some point, but I assume there \n> are too many patches out there for that to be worth the merge problems, \n> right?\n\nOn one hand gitweb.perl in single file makes it easy to install; on the\nother hand if it was split into modules (like git-gui now is) it would\nI think be easier to understand and modify... I think however that it\nwould be better to first make gitweb use Git.pm, adding improving Git.pm\nwhen necessary (for example adding eager config parsing used in gitweb,\ni.e. read whole config into Perl hash at first request, then access hash\ninstead of further calls to git-config).\n\n-- \nJakub Narebski\nPoland\n"},{"id":"77966","messageId":"483DA594.5040803@gmail.com","threadId":"13694","inReplyTo":"200805281414.36141.jnareb@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-28T18:33:56Z","receivedAt":"2008-05-28T18:33:56Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Jakub Narebski wrote:\n> 1. Caching data\n>  * disadvantages:\n>    - more CPU\n>    - need to serialize and deserialize (parse) data\n>    - more complicated\n\nCPU: John told me that so far CPU has *never* been an issue on k.org. \nUnless someone tells me they've had CPU problems, I'll assume that CPU \nis a non-issue until I actually run into it (and then I can optimize the \nparticular pieces where CPU is actually an issue).\n\nSerialization: I was planning to use Storable (memcached's Perl API uses \nit transparently I think).  I'm hoping that this'll just solve it.\n\nIt's true that it's more complicated.  It'll require quite a bit of \nrefactoring, and maybe I'll just back off if I find that it's too hard.\n\n> I'm afraid that implementing kernel.org caching in mainline in\n> a generic way would be enough work for a whole GSoC 2008.\n\nI probably won't reimplement the current caching mechanism.  Do you \nthink that a solution using memcached is generic enough?  I'll still \nneed to add some abstraction layer in the code, but when I'm finished \nthe user will either get the normal uncached gitweb, or activate \nmemcached caching with some configuration setting.\n\nBy the way, I'll be posting about gitweb on this mailing list \noccasionally.  If any of you would like to receive CC's on such \nmessages, please let me know, otherwise I'll assume you get them through \nthe mailing list.\n\n-- Lea\n"},{"id":"78095","messageId":"200805300127.10454.jnareb@gmail.com","threadId":"13694","inReplyTo":"483DA594.5040803@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-05-29T23:27:07Z","receivedAt":"2008-05-29T23:27:07Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Wed, 28 May 2008, Lea Wiemann wrote:\n> Jakub Narebski wrote:\n> >\n> > 1. Caching data\n> >  * disadvantages:\n> >    - more CPU\n> >    - need to serialize and deserialize (parse) data\n> >    - more complicated\n> \n> CPU: John told me that so far CPU has *never* been an issue on k.org. \n> Unless someone tells me they've had CPU problems, I'll assume that CPU \n> is a non-issue until I actually run into it (and then I can optimize the \n> particular pieces where CPU is actually an issue).\n\nTrue.\n\nWhat you have to care about (although I don't think it would be\npartilcularly difficult) is to not repeat bad I/O patterns with\ncache...\n\n> Serialization: I was planning to use Storable (memcached's Perl API uses \n> it transparently I think).  I'm hoping that this'll just solve it.\n\nWhile Storable is part of, I think, any modern Perl installation, there\nmight be problem with memcached API, and memcached API wrappers such as\nCHI one.  Namely you cannot assume that memcached API is installed, so\nyou have to provide some kind of fallback.\n \n> It's true that it's more complicated.  It'll require quite a bit of \n> refactoring, and maybe I'll just back off if I find that it's too hard.\n\nWhat's more, if you want to implement If-Modified-Since and\nIf-None-Match, you would have to implement it by yourself, while\nfor static pages (cahing HTML output) web server would do this\nfor us \"for free\".\n\n> > I'm afraid that implementing kernel.org caching in mainline in\n> > a generic way would be enough work for a whole GSoC 2008.\n> \n> I probably won't reimplement the current caching mechanism.  Do you \n> think that a solution using memcached is generic enough?  I'll still \n> need to add some abstraction layer in the code, but when I'm finished \n> the user will either get the normal uncached gitweb, or activate \n> memcached caching with some configuration setting.\n\nThats good enough, although I think that current caching mechanism in\nkernel.org's gitweb (your implementation follows more what repo.or.cz's\ngitweb does) has some good ideas, like for example adaptive (depending\non load) expiry time.\n\nBy the way what do you think about adding (as an option) information\nabout gitweb performance to the output, in the form of\n  \"Site generated in 0.01 seconds, 2 calls to git commands\"\nor\n  \"Site generated in 0.0023 seconds, cached output, 1m31s old\"\nline somewhere in the page footer?\n\nI hope you have some ideas in gitweb access statistics from kernel.org,\nrepo.or.cz, and perhaps other large git hosting sites (e.g.\nfreedesktop.org), and you plan on benchamrking gitweb caching using\naverage / amortized time to generate page, ApacheBench or equivalent,\nload average on server depending on number of requests, I/O load (using\nfio tool, for example) depending on number of requests etc.\n\n> By the way, I'll be posting about gitweb on this mailing list \n> occasionally.  If any of you would like to receive CC's on such \n> messages, please let me know, otherwise I'll assume you get them through \n> the mailing list.\n\nI read git mailing list via Usenet / news interface (NNTP gateway) from\nGMane. \n\n-- \nJakub Narebski\nPoland\n"},{"id":"78112","messageId":"483FABB4.1010309@gmail.com","threadId":"13694","inReplyTo":"200805300127.10454.jnareb@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-30T07:24:36Z","receivedAt":"2008-05-30T07:24:36Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Jakub Narebski wrote:\n> you cannot assume that memcached API is installed, so\n> you have to provide some kind of fallback.\n\nThat fallback would be to have no caching. :)  I think that's acceptable \n-- I'm not too willing to implement caching for two API's. \n(Incidentally, memcached takes two shell commands to install and get \nrunning on my machine; I think that's acceptably easy.)\n\n> What's more, if you want to implement If-Modified-Since and\n> If-None-Match, you would have to implement it by yourself, while\n> for static pages (cahing HTML output) web server would do this\n> for us \"for free\".\n\nAre web servers doing anything that we can't easily reimplement in a few \nlines (and, on top of that, more easily tailored to different actions, \nprojects, etc.)?\n\n> By the way what do you think about adding (as an option) information\n> about gitweb performance to the [HTML] output,\n\nDefinitely a good idea!\n\n> I hope you have some ideas in gitweb access statistics from kernel.org,\n\nI'm waiting for John to give me SSH access and/or send them my way. :)\n\n> and you plan on benchamrking gitweb caching using [snip]\n\nAbsolutely -- thanks for the suggestions!\n\n-- Lea\n"},{"id":"78126","messageId":"200805301202.25368.jnareb@gmail.com","threadId":"13694","inReplyTo":"483FABB4.1010309@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-05-30T10:02:23Z","receivedAt":"2008-05-30T10:02:23Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Fri, 30 May 2008, Lea Wiemann wrote:\n> Jakub Narebski wrote:\n> >\n> > you cannot assume that memcached API is installed, so\n> > you have to provide some kind of fallback.\n> \n> That fallback would be to have no caching.  I think that's acceptable \n> -- I'm not too willing to implement caching for two API's.\n\nI hope that you would make a wrapper around memcached (caching engine)\nAPI (it's a pity we cannot use CHI unified Perl caching interface),\nso it would be easy for example to change to filesystem based cache,\nor size aware filesystem based cache, or mmap, etc...  I mean here\nthat even if you don't implement two caching API's at least make it\npossible to easy change caching backend.\n\nNote also that memcached may not have sense for single machine\n(single server installation), and does not make sense for memory\nstarved machines... and one can want gitweb caching even in that\nsituation.\n\n> (Incidentally, memcached takes two shell commands to install and get \n> running on my machine; I think that's acceptably easy.)\n\nAs John 'Warthog9' said wrt. using additional Perl modules for gitweb\ncaching, most sites that are used as web servers (and gitweb servers)\nhave strict requirements on stability of installed programs, libraries\nand modules.  IIRC the policy usually is that one can install packages\nfrom main (base) repository for Linux distribution used on server, also\nfrom extras repository; sometimes from trusted contrib package\nrepository.  Modules which are only in CPAN, and programs which require\ncompilation are out of the question, unfortunately.\n\nI think there is no problem wrt. memcached itself, I'm not so sure\nabout Perl APIs: Cache::Memcached and/or Cache::Memcached::Fast (and\noptionally appropriate CHI modules/backends).\n \n> > What's more, if you want to implement If-Modified-Since and\n> > If-None-Match, you would have to implement it by yourself, while\n> > for static pages (cahing HTML output) web server would do this\n> > for us \"for free\".\n> \n> Are web servers doing anything that we can't easily reimplement in a few \n> lines (and, on top of that, more easily tailored to different actions, \n> projects, etc.)?\n\nCan we reimplement it?  I think we can.  Easily?  I'm not sure.  \nHTTP/1.0 If-Modified-Since should be failry easy; it would be harder\nto support fully and correctly ETag (weak vs. strong tags),\nIf-None-Match (from web browsers I think), If-Match (from web caches)\nit would take some work.\n\n> > By the way what do you think about adding (as an option) information\n> > about gitweb performance to the [HTML] output,\n> \n> Definitely a good idea!\n\nI'd try to add it when I'd have a bot more of free time; unless you\nwould do this first.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"78146","messageId":"4840166C.3030903@gmail.com","threadId":"13694","inReplyTo":"200805301202.25368.jnareb@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-30T14:59:56Z","receivedAt":"2008-05-30T14:59:56Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Jakub Narebski wrote:\n> even if you don't implement two caching API's at least make it\n> possible to easy change caching backend.\n\nSure, I'll keep that in mind.\n\n> Note also that memcached may not have sense for single machine [...],\n>  and does not make sense for memory starved machines...\n\nFor single machines, memcached certainly works fine.  For on \nmemory-starved machines with HD caches, you'd have to cache the \naggregate HTML data, not the data in the backend.  So as long as I'm \nworking on the backend (repository) cache, memcached should be fine.\n\n> IIRC the policy usually is that one can install packages\n> from main (base) repository for Linux distribution used on server,\n\nlibcache-memcached-perl is in Debian stable; that's fair enough I think. \n  Cache::Memcached::Fast doesn't seem to be in Debian as of now, but I \nwouldn't worry about performance unless it comes up.\n\n>>> By the way what do you think about adding (as an option) information\n>>> about gitweb performance to the [HTML] output,\n> \n> I'd try to add it when I'd have a bot more of free time\n\nI'd probably wait with this until I've written the Perl Git API.\n\n-- Lea\n"},{"id":"78147","messageId":"20080530150713.GG593@machine.or.cz","threadId":"13694","inReplyTo":"4840166C.3030903@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-05-30T15:07:13Z","receivedAt":"2008-05-30T15:07:13Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Fri, May 30, 2008 at 04:59:56PM +0200, Lea Wiemann wrote:\n> Jakub Narebski wrote:\n>> IIRC the policy usually is that one can install packages\n>> from main (base) repository for Linux distribution used on server,\n>\n> libcache-memcached-perl is in Debian stable; that's fair enough I think.  \n> Cache::Memcached::Fast doesn't seem to be in Debian as of now, but I \n> wouldn't worry about performance unless it comes up.\n\nStill, please make this optional. It is fine for gitweb not to do any\ncaching in the bare setup, but you should be able to get the simple\nversion running without any external dependencies.\n\n>>>> By the way what do you think about adding (as an option) information\n>>>> about gitweb performance to the [HTML] output,\n>> I'd try to add it when I'd have a bot more of free time\n>\n> I'd probably wait with this until I've written the Perl Git API.\n\nHmm, it shouldn't depend on that in any way, should it?\n\nuse Time::HiRes qw(gettimeofday tv_interval);\nmy $t0 = [gettimeofday];\n...\nprint \"<p>This page took \".tv_interval($t0, [gettimeofday]).\"s to generate.</p>\";\n\nI wonder what oldest Perl versions do we aim to support? If <5.8, we\nneed to be more careful about Time::HiRes. It would be useful to\ndocument this with a use perl statement at the top of the script.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"78149","messageId":"48401CFF.4020702@gmail.com","threadId":"13694","inReplyTo":"20080530150713.GG593@machine.or.cz","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-30T15:27:59Z","receivedAt":"2008-05-30T15:27:59Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Petr Baudis wrote:\n> please make [memcached] optional.\n\nOh, absolutely.  I was definitely planning to keep Gitweb runnable \nwithout having Cache::Memcached installed.\n\n> print \"<p>This page took \".tv_interval($t0, [gettimeofday]).\"s to generate.</p>\";\n\nSure -- I'm not sure how useful bare timings are, though.  When I look \nat individual pages, the page cache is usually warm anyway, so the only \nthing I might be interested in is advanced statistics like the number of \ncalls to git or number of cache hits/misses.  To find out how the cache \nperforms timing-wise, you'll have to do larger benchmarks, individual \npage generation times won't help that much.\n\n> I wonder what oldest Perl versions do we aim to support?\n\nI'm thinking about 5.8 or 5.10.  Looking at Debian, Perl 5.10 is not in \nstable (etch), but it's in lenny, which is planned to become stable in \nSept. 08.  So by the time the updated Gitweb/Git.pm has stabilized (and \nshows up as a package in Debian), Perl 5.10 will definitely be available \nwidely enough.\n\n-- Lea\n"},{"id":"78151","messageId":"20080530153822.GH593@machine.or.cz","threadId":"13694","inReplyTo":"48401CFF.4020702@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-05-30T15:38:22Z","receivedAt":"2008-05-30T15:38:22Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Fri, May 30, 2008 at 05:27:59PM +0200, Lea Wiemann wrote:\n> Petr Baudis wrote:\n>> I wonder what oldest Perl versions do we aim to support?\n>\n> I'm thinking about 5.8 or 5.10.  Looking at Debian, Perl 5.10 is not in \n> stable (etch), but it's in lenny, which is planned to become stable in \n> Sept. 08.  So by the time the updated Gitweb/Git.pm has stabilized (and \n> shows up as a package in Debian), Perl 5.10 will definitely be available \n> widely enough.\n\nWow, and here I was wondering if requiring at least 5.6 was not too\nliberal. ;-) I believe 5.8 is the newest possible candidate though, it\nis still too widespread; e.g. Debian-wise, many servers run on Etch and\nare going to stay there even for quite some time after Lenny gets\nreleased. Heck, I still have accounts on plenty of Sarge machines. ;-)\n(Sarge seems to have Perl-5.8.4.)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nWhatever you can do, or dream you can, begin it.\nBoldness has genius, power, and magic in it.\t-- J. W. von Goethe\n"},{"id":"78154","messageId":"b77c1dce0805300904o5b4363efkc4591fc820164bf7@mail.gmail.com","threadId":"13694","inReplyTo":"20080530153822.GH593@machine.or.cz","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Rafael Garcia-Suarez","fromEmail":"rgarciasuarez@gmail.com","sentAt":"2008-05-30T16:04:40Z","receivedAt":"2008-05-30T16:04:40Z","isPatch":false,"sender":{"key":"rgarciasuarez@gmail.com","avatar":null},"body":"2008/5/30 Petr Baudis <pasky@suse.cz>:\n>\n> Wow, and here I was wondering if requiring at least 5.6 was not too\n> liberal. ;-) I believe 5.8 is the newest possible candidate though, it\n> is still too widespread; e.g. Debian-wise, many servers run on Etch and\n> are going to stay there even for quite some time after Lenny gets\n> released. Heck, I still have accounts on plenty of Sarge machines. ;-)\n> (Sarge seems to have Perl-5.8.4.)\n\nI think 5.8.2 is a good _minimum_ perl to support. Before that one,\nUnicode support is next to null (5.6 and below) or too buggy, and\ngitweb needs that.\n"},{"id":"78159","messageId":"48404BD2.6040703@gmail.com","threadId":"13694","inReplyTo":"20080530153822.GH593@machine.or.cz","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-30T18:47:46Z","receivedAt":"2008-05-30T18:47:46Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Petr Baudis wrote:\n> [5.8] is still too widespread;\n\nOkay; I'll keep testing Git.pm and Gitweb with Perl 5.8 then.\n\n-- Lea\n"},{"id":"78160","messageId":"1212173779.26045.77.camel@localhost.localdomain","threadId":"13694","inReplyTo":"b77c1dce0805300904o5b4363efkc4591fc820164bf7@mail.gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"J.H.","fromEmail":"warthog19@eaglescrag.net","sentAt":"2008-05-30T18:56:19Z","receivedAt":"2008-05-30T18:56:19Z","isPatch":false,"sender":{"key":"warthog19@eaglescrag.net","avatar":null},"body":"I would agree - lets try and shoot for 5.8 as a baseline minimum (there\nare lots of people who are slow to upgrade, and it would be nice to be\nable for them to make use of newer gitweb's on things like Centos / RHEL\n4\n\n- John\n\n\nOn Fri, 2008-05-30 at 18:04 +0200, Rafael Garcia-Suarez wrote:\n> 2008/5/30 Petr Baudis <pasky@suse.cz>:\n> >\n> > Wow, and here I was wondering if requiring at least 5.6 was not too\n> > liberal. ;-) I believe 5.8 is the newest possible candidate though, it\n> > is still too widespread; e.g. Debian-wise, many servers run on Etch and\n> > are going to stay there even for quite some time after Lenny gets\n> > released. Heck, I still have accounts on plenty of Sarge machines. ;-)\n> > (Sarge seems to have Perl-5.8.4.)\n> \n> I think 5.8.2 is a good _minimum_ perl to support. Before that one,\n> Unicode support is next to null (5.6 and below) or too buggy, and\n> gitweb needs that.\n"},{"id":"78173","messageId":"7vd4n3k04r.fsf@gitster.siamese.dyndns.org","threadId":"13694","inReplyTo":"1212173779.26045.77.camel@localhost.localdomain","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-05-30T20:28:04Z","receivedAt":"2008-05-30T20:28:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"J.H.\" <warthog19@eaglescrag.net> writes:\n\n> On Fri, 2008-05-30 at 18:04 +0200, Rafael Garcia-Suarez wrote:\n>> 2008/5/30 Petr Baudis <pasky@suse.cz>:\n>> >\n>> > Wow, and here I was wondering if requiring at least 5.6 was not too\n>> > liberal. ;-) I believe 5.8 is the newest possible candidate though, it\n>> > is still too widespread; e.g. Debian-wise, many servers run on Etch and\n>> > are going to stay there even for quite some time after Lenny gets\n>> > released. Heck, I still have accounts on plenty of Sarge machines. ;-)\n>> > (Sarge seems to have Perl-5.8.4.)\n>> \n>> I think 5.8.2 is a good _minimum_ perl to support. Before that one,\n>> Unicode support is next to null (5.6 and below) or too buggy, and\n>> gitweb needs that.\n\n> I would agree - lets try and shoot for 5.8 as a baseline minimum (there\n> are lots of people who are slow to upgrade, and it would be nice to be\n> able for them to make use of newer gitweb's on things like Centos / RHEL\n> 4\n\nI do not think it is unreasonable to require recent Perl for a machine\nthat runs gitweb, as it is not something you would run on your \"customer\nsite that needs to be ultra sta(b)le\" nor on your \"development machine\nthat needs to run the same version as that ultra sta(b)le customer\ninstallation.\"  In other words, gitweb is primarily a developer tool, and\nyou can assume that people can afford to have a dedicated machine they can\nupdate its Perl to recent version.\n\nHowever, introducing dependency on 5.8 to any and all Git.pm users may\nhave a much wider impact.  Right now, these \"use Git\":\n\n    git-add--interactive.perl\n    git-cvsexportcommit.perl\n    git-send-email.perl\n    git-svn.perl\n\nIf you are doing development for some customer application whose end\nproduct needs to land on a machine with a pre-5.8 Perl, it is conceivable\nthat you may pin the Perl running on that development machine to that old\nversion, say 5.6.  Introducing 5.8 dependency to Git.pm in such a way that\n\"use Git\" from these fail might make these people somewhat unhappy.\n"},{"id":"78180","messageId":"48407287.7050900@gmail.com","threadId":"13694","inReplyTo":"7vd4n3k04r.fsf@gitster.siamese.dyndns.org","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Lea Wiemann","fromEmail":"lewiemann@gmail.com","sentAt":"2008-05-30T21:32:55Z","receivedAt":"2008-05-30T21:32:55Z","isPatch":false,"sender":{"key":"lewiemann@gmail.com","avatar":null},"body":"Junio C Hamano wrote:\n> Right now, these \"use Git\": git-add--interactive.perl\n> git-cvsexportcommit.perl git-send-email.perl git-svn.perl\n> \n> Introducing 5.8 dependency to Git.pm in such a way that\n> \"use Git\" from these fail might make these people somewhat unhappy.\n\nGit seems to generally work with Perl 5.6 after installing Scalar::Util \nthrough CPAN.  I'm happy with (sporadically) testing it with 5.6.2, \nthough I don't have any older version installed.  Also, if at some point \nPerl 5.6 compatibility gets in the way (due to lack of Unicode support), \nwe'll have to revisit this issue, but for now that should be fine.\n\nGitweb relies on Unicode support (e.g. \"use Encode\") and will continue \nto be compatible with 5.8 and 5.10 only.\n\nI'm not sure how much changes between Perl's micro versions.  Should we \nboldly claim \"use 5.6.0\" or only \"use 5.6.2\"?  Are people still using \nversions 5.6.0/1 at all?  And for Gitweb, use 5.8.0, or use 5.8.8 (which \nis the version I'm testing with, currently)?  Or should I downgrade?\n\n-- Lea\n"},{"id":"78213","messageId":"200805311215.37233.jnareb@gmail.com","threadId":"13694","inReplyTo":"48401CFF.4020702@gmail.com","subject":"Re: Gitweb caching: Google Summer of Code project","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-05-31T10:15:34Z","receivedAt":"2008-05-31T10:15:34Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Fri, 30 May 2008, Lea Wiemann wrote:\n> Petr Baudis wrote:\n> >\n> > please make [memcached] optional.\n> \n> Oh, absolutely.  I was definitely planning to keep Gitweb runnable \n> without having Cache::Memcached installed.\n\nI think the idea was to have the following options:\n * cache using memcached (Cache::Memcached installed, and memcached on)\n * cache using filesystem, perhaps size aware (with limited cache size)\n * no caching\n\nIt is quite possible that one would want/need gitweb caching, but\neither does not want hassle with memcached, or memcached is not\nfeasible (for example memory starved machine).\n\n> > print \"<p>This page took \".tv_interval($t0, [gettimeofday]).\"s to generate.</p>\";\n> \n> Sure -- I'm not sure how useful bare timings are, though.  When I look \n> at individual pages, the page cache is usually warm anyway, so the only \n> thing I might be interested in is advanced statistics like the number of \n> calls to git\n\nThis should be fairly easy, just modify git_cmd() to count number of\ncalls.\n\n> or number of cache hits/misses.\n\nAnd this I don't think it would be easy.\n\n-- \nJakub Narebski\nPoland\n"}]}