{"thread":{"id":"12753","subject":"Re: [RFD] Gitweb caching, part 1 (long)","startedAt":"2008-03-19T00:54:53Z","lastAt":"2008-03-29T17:13:26Z","messageCount":4,"participants":["Frank Lichtenheld","Jakub Narebski"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"72384","messageId":"200803190154.55532.jnareb@gmail.com","threadId":"12753","inReplyTo":null,"subject":"[RFD] Gitweb caching, part 1 (long)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-19T00:54:53Z","receivedAt":"2008-03-19T00:54:53Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Please Cc: me directly, as I am not subscribed to git mailing list,\n and GMane NNTP (news) interface I use doesn't show currently any new\n posts; I wouldn't want to miss any response.  Thanks in advance.]\n\n\nThis post shows my ideas about how to implement caching in gitweb, my\nthoughts on what are the problems, and what solutions (what code) can\nwe use.\n\n>From what I remember of discussion about gitweb performance and\nbottlenecks on git mailing list, the main culprit is projects list\n(which perhaps should be redesigned), and summary pages for some of\nlarger / more popular projects.  Gitweb performance is I/O bound, not\nCPU bound, so I guess not all existing caching solutions and ideas\nwould work with gitweb.\n\nThere are some troubles with adding generic (as opposed to\nsite-specific) caching to gitweb.  First, gitweb should work both with\nmod_perl and as CGI (perhaps in the future also FastCGI) script.\nSecond, the solution should not depend on additional packages, at\nleast not those that can be found packaged in extras or well trusted\ncontrib repositories; not all admins allow installing packages from\nCPAN.  Third, the solution could be helped but should not depend on\nadding helper hooks to users repositories; while hosting sites like\nrepo.or.cz controls repositories, sites like kernel.org or\nfreedesktop.org, which give shell access, do not.\n\n\nLet talk first about what to cache.\n\n1. Support for caching in HTTP (HTTP accelerators, caching engines)\n\nMy first idea of adding caching support to gitweb was for it to\ngenerate proper \"freshness\" caching headers (Last-Modified: and ETag:)\nand respond to cache validation requests (If-Modified-Since:,\nIf-None-Match: etc.), and for reverse proxy, aka. caching engine,\naka. web accelerator/HTTP accelerator (e.g. Varnish or Squid) take\ncare of caching.  It is better to use existing solution, isn't it?\n\nUnfortunately gitweb to generate date for If-Modified-Since:,\nor a tag for If-None-Match: must do hard work; perhaps not as much as\ngenerating the whole page in the term of CPU, but almost the same in\nthe terms of I/O hit.  So it is not so simple...\n\nNevertheless even if using reverse proxy for gitweb caching is not so\nsimple, gitweb should play well with support for caching in HTTP\nprotocol, so the pages can be cached between gitweb and user, either\nin one of intermediate caches, proxy server with caching support, or\nbrowser cache.  Currently gitweb uses 'Expires:' header with expiry of\n24h / 1d (IIRC cutoff time for caches; also IIRC forever is half\na year) for pages which we know would not change (using full SHA-1\nidentifier and all required information filled).  We should probably\ngenerate Last-Modified: and/or ETag: if it is possible.\n\nHowever if gitweb has some kind of internal caching turned on, it can\nrespond properly to validation requests with low cost.  This way some\nof requests would be handled by intermediate caches, so gitweb\nwouldn't have even to access the cache to return an answer.  But IMHO\nthat is a secondary concern: it could help, but isn't possible to do\nwell without in-gitweb caching (as far as I can see).\n\nBTW. besides optionally marking result as being retrieved from cache\n(\"stale\" or \"cached\"), gitweb I think should also send appropriate\nWarning: header, see sections 13.1.2 and 14.46 of RFC2616, e.g.\n  Warning: 110 git.kernel.org \"Response is stale\"\n\nReferences:\n* \"Caching Tutorial for Web Authors\"\n  http://www.mnot.net/cache_docs/\n* HTTP 1.1 Specification (RFC 2616)\n  http://www.ietf.org/rfc/rfc2616.txt\n\n\n2. Caching Perl structures\n\nOn of solutions (used for example by Petr 'Pasky' Baudis in his last\npost about caching projects list info in gitweb) is be to cache (save)\nPerl structures containing information needed to generate response\n(web page).  Another solution, discussed below, would be to cache\ngenerated output, i.e. web page, optionally with (some) HTTP headers.\n\nThe advantage of storing Perl structures (raw data) in the cache is\nthat the same data can be reused for different pages (e.g. paging\nprojects list if/when it gets implemented), same page with varying\npart (e.g. content type being text/html or application/xhtml+xml\ndepending on what web browser prefers, or transparently compressed\noutput via Transfer-Encoding: depending on web browser capabilities),\nand for replying to cache validation requests.  Additionally we can\ngenerate web pages with correct relative (e.g. \"5 minutes ago\") time\ninfo.  Not that all but first and last are not possible with caching\n[final] output, but it would be, I think, much harder...\n\nThe disadvantage is that we have to decide is what format to use for\nserializing data, i.e. to represent compound complex data as stream of\nbytes in cache... unless of course gitweb would rely on one of already\nexisting caching solutions, which usually take care of this problem\nfor us: see next section (in next installment).\n\nFormats I was considering were:\n - Data::Dumper\n - Storable (binary)\n - YAML Tiny\n - gitconfig tiny\n\n2.1. Data::Dumper\n\nOne of advantages of Data::Dumper format is that it comes with Perl\ninstallation, so there is no problem with installing it (well, at\nleast it comes with perl-5.8.6-24 RPM on Linux).  Another advantage is\nthat it is textual format, thus easy to debug in the case of\nproblems.  \n\nMain problem with Data::Dumper is that it use eval() to thaw (restore)\ndata from serialized form in cache, which is serious security risk in\nless secure environments.\n\n2.2. Storable\n\nAlso comes packaged with Perl distribution.  Offers writing to and\nrestoring from file or to/from opened file handle.  It is fast, from\nmy unscientific tests around 3-4 times as fast as using eval() to read\nData::Dumper data.\n\nOne of disadvantages is the fact that Storable format is binary, so\nyou would have to write separate Perl script to convert it to\nhuman-readable form (e.g. Data::Dumper form).  Format includes format\nand version header, and modern 'file' installations should detect it\ncorrectly, e.g.:\n  filename: perl Storable(v0.7) data (major 2) (minor 6)\n\nAnother nuisance is the fact that while Storable(3pm) manpage states\nthat:\n\n     The [retrieve()] routine returns \"undef\" for I/O problems or\n     other internal error, a true value otherwise. Serious errors are\n     propagated as a \"die\" exception.\n\nBut 'serious errors' include the fact that file is not in correct\nformat, so for safety (because server-side script should return page\nwith error info instead of dying silently) one would have to use \"eval\n{ ... }\" to catch errors.\n\nUsed by Cache::Cache and, I think, also by other caching solutions\n(packages).\n\n\n2.3. YAML Tiny (subset of YAML)\n\nYAML was created as human-readable serialization format, easy to parse\nby machine.  Unfortunately none of YAML parsing modules (YAML,\nYAML::Syck, YAML::Tiny) are packaged with Perl; on the other hand they\ncan be found in Dries RPM repository, so I guess they fill criteria of\nbeing in extras or trusted contrib package repository.\n\nAdditionally at least YAML::Tiny (which implements subset of YAML in\npure Perl code) is slower, around 4-5 times, than even using\nData::Dumper, and more that 10 times slower than using Storable.  This\n_might_ have been caused by the choice of module to implement YAML\nparsing.\n\nWe could write parser (and generator) for even smaller subset of YAML\nto use only those features that are truly needed by gitweb... but then\nwe can go with the next format, which is also text format, and also\ndoesn't have insecurities of using eval() to thaw data (read from\ncache).\n\nYAML was designed from the ground up to be an excellent syntax for\nconfiguration files.  Not necessarily so for cache.\n\nReferences:\n* http://en.wikipedia.org/wiki/YAML\n* YAML::Tiny(3pm)\n  http://search.cpan.org/dist/YAML-Tiny/lib/YAML/Tiny.pm\n* YAML Ain't Markup Language (YAML^TM) Version 1.1, Working Draft\n  http://yaml.org/spec/current.html\n\n2.4. gitconfig tiny (subset of ini-like gitconfig format).\n\nWhat, I think, we would want to cache is usually list of records, or\nin Perl terminology array of hashes; usually ordering of array doesn't\nmatter.\n\nBecause of that I think it would be possible to represent data to be\nsaved (cached) in the ini-like extended git config format.  Then\ngitweb could either (re)use config parser in Perl used by\ngit-cvsserver (which accepts subset of valid config format), or \n\"git config --file=<path> -z -l\" to slurp data in more parseable\nformat... but if we do that, we could choose this format or variation\nof it as our serialization format.\n\nThe cache file could look like this:\n\n   [gitweb \"<primary key value>\"]\n   \tkey1 = value1\n   \tkey2 = \"value with spaces\"\n\nwhere for list of projects info primary key might be path (relative to\nprojectroot) to the repository.\n\n\n3. Caching output: formatted pages\n\nAlternate solution to caching Perl structures is caching final output,\nwith or without (some/all) HTTP headers.  It has the advantage that it\nis simple to implement, and that the same code can be used to cache\nall the pages.  (But we could get similar result by creating something\nsimilar to Tie::Memoize, tying hash or array so it automatically get\ndata either from git command, or from cache... or we can implement\nuniversal API, like Cache::Cache API.)\n\nThis is from what I understand what kernel.org (warthog9) gitweb uses;\nI don't know what cgit (web interface in C) which also has some\ncaching support uses: does it cache data or output?\n\nHow one can simply extend CGI script with support for caching is shown\nby the CGI::Cache (non-standard CPAN Perl module).  On the other hand\ngitweb can afford more extensive surgery.\n\nReferences:\n* CGI::Cache(3pm)\n  http://search.cpan.org/~dcoppit/CGI-Cache-1.4200/lib/CGI/Cache.pm\n\n\nTo be continued...\n\n%% .................................................................. %%\nIn next parts:\n\nCache lifetime and invalidation\n1. static cache, external refreshing e.g. by hooks\n2. stat and/or inotify, to check if repository changed\n3. cache lifetime (trying to avoid \"thundering horde\" problem)\n\nCPAN packages we could use, or take inspiration from\n1. Cache::Cache (standard)\n2. CHI, Unified caching interface\n3. Cache\n4. other (e.g. Cache::Adaptive, using Cache::Cache)\n\n-- \nJakub Narebski\nPoland\n"},{"id":"72382","messageId":"20080319082158.GY18624@mail-vs.djpig.de","threadId":"12753","inReplyTo":"200803190154.55532.jnareb@gmail.com","subject":"Re: [RFD] Gitweb caching, part 1 (long)","fromName":"Frank Lichtenheld","fromEmail":"frank@lichtenheld.de","sentAt":"2008-03-19T08:21:58Z","receivedAt":"2008-03-19T08:21:58Z","isPatch":false,"sender":{"key":"frank@lichtenheld.de","avatar":"https://gravatar.com/avatar/b9f1d4b120e138f157c9e480d0818197c474628923786adb98f30017cdb99c3c?d=mp&s=160"},"body":"On Wed, Mar 19, 2008 at 01:54:53AM +0100, Jakub Narebski wrote:\n> Because of that I think it would be possible to represent data to be\n> saved (cached) in the ini-like extended git config format.  Then\n> gitweb could either (re)use config parser in Perl used by\n> git-cvsserver (which accepts subset of valid config format), or \n> \"git config --file=<path> -z -l\" to slurp data in more parseable\n> format... but if we do that, we could choose this format or variation\n> of it as our serialization format.\n\ngit-cvsserver uses \"git config -l\", too.\n\nGruesse,\n-- \nFrank Lichtenheld <frank@lichtenheld.de>\nwww: http://www.djpig.de/\n"},{"id":"73050","messageId":"200803251806.58290.jnareb@gmail.com","threadId":"12753","inReplyTo":"200803190154.55532.jnareb@gmail.com","subject":"[RFD] Gitweb caching, part 2 (long)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-25T17:06:56Z","receivedAt":"2008-03-25T17:06:56Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"In previous part:\n\nWhat to cache.\n1. Support for caching in HTTP (external caching)\n2. Caching Perl structures (and serialization)\n3. Caching gitweb output: formatted pages\n\n\nTO WHOM IT MAY CONCERN:  John 'Warthog9' Hawley (J.H.) who created\ncaching for gitweb at kernel.org; Petr 'Pasky' Baudis who maintains\nrepo.or.cz fork of gitweb and lately added caching of projects list\ninfo; Lars Hjemli who is the author of cgit, git web interface in C\nwhich includes some caching.  (BTW. I'd like to hear your thoughs on\ngit web interface caching, and about solutions you have implemented).\n\n\nThis is continuation of my thoughts about how to implement caching in\ngitweb, what problems we could encounter, and what existing solutions\n(what code/what packages) can we (re)use.\n\n\nOne of the more important issues to think about when implementing\ncaching is to decide when to regenerate cache, i.e. issues of cache\n(in)validation and lifetime.\n\n\n1. Static cache, external refreshing (invalidation).\n\nThe easiest situation is when cache can be invalidated (removed)\nexternally; caching support in gitweb would then need only to either\nuse cached information or cached output if it exists, and generate\ninformation and/or output and cache appropriate things if cache\ndoesn't exist.\n\nIn closed-up git hosting system like repo.or.cz new contents can\nappear in repository only via push (if repo is manually updated) or\nvia automated fetch (if repo is mirrored automatically).  This means\nthat it is known when infomration about given repository gets stale\n(out of sync).  It would be then enought to make 'update' or\n'post-receive' hook to delete cache, or invalidate parts of cached\ninfo about given repository.  Creation and deletion of repositories\nshould also be handled by scripts; they affect caching too.\n\nThis of course assumes that we can control repository hooks (perhaps\ngit should learn hook multiplexing first, as proposed some time ago on\na mailing list).  This is not the case when developers are given shell\naccess, and gitweb is offered as a part of service rather than as a\npart of git hosting; repositories are not under web administrator\ncontrol.  This is the case (according to J.H. on git mailing list) for\nkernel.org.\n\nSo we have to examine also more generic solutions.\n\n\n2. Checking filesystem (stat and/or inotify).\n\nIf new objects come to repository via commit or via fetch it is enough\n(I guess) to watch for modifications of GIT_DIR of a project (I think\ndue to doing atomic writes via \"create temporary file, then rename it\nto final filename\" of files in GIT_DIR: COMMIT_MSG and FETCH_HEAD).\nSo it should be enough to check and compare stat info for GIT_DIR of a\nproject, or of possible implement some inotify (or equivalent on other\noperating systems than Linux) checking, to see if cache can contain\nstale info.  In practice what we can truly check is that nothing\nchanged with repo.\n\nUnfortunately the above is not the case if objects come to repository\nvia push.  Note that both push resulting in crating a pack (this I\nthink could use the same mechanism, only checking GIT_DIR/objects/pack\ndirectory), and push resulting in creation of loose objects has to be\nsupported; additionally the refs pushed can have deeply hierarchical\nnames.\n\nI would be grateful if somebody could think a way to check if anything\ncould have changed for such situation... but as it is now we have to\ngo to more complicated ways of cache invalidation.\n\n\n3. Cache lifetime.\n\nFinally, for cases such as gitweb where validating cache (checking if\nthe cached information isn't stale, out of sync with reality) is\nalmost as costly as calculating the whole information without using\ncache at all, there is one possible solution to cache validation:\nsimply keep cache for some time.  For longer cache lifetimes gitweb\nperhaps should put some notification that information is from cache,\nperhaps with the time in human readable form how much time ago was\nthis information generated (human readable means no \"1325 seconds ago\"\ninfo ;-).  And if we want to be thorough, put it also in the HTTP\nheader \"Warning:\" (at least for HTTP/1.1, see sections 13.1.2 and\n14.46 of RFC 2616), e.g.:\n\n  Warning: 110 git.kernel.org \"Response is stale\"\n\nThe question is what timeout, or how to choose lifetime of a cache.\nJ.H. kernel.org's gitweb tries to adjust cache lifetime to server\nload, making cache lifetime longer if server load is higher, but\nensuring that cache lifetime stays within specified bounds.  \n\nI have found among CPAN modules Cache::Adaptive where you can also\nspecify bounds for expire time and subroutine to adjust cache\nlifetime, e.g. according to load average, process time for building\nthe cache entry, etc. (it can use specified backend, for example\nCache::FileCache from Cache::Cache distribution).  Its subclass\nCache::Adaptive::ByLoad which tries to adjust cache lifetime for\nbottlenecks under heavy load.  Neither of modules I think is\ndistributed as ready package in extras on trusted contrib packages\nrepositories.  Nevertheless we can \"borrow\" the algorithms used by\nthose modules.\n\nWe should also try to avoid 'thundering herd' problem, namely that\ncache expires, gitweb gets N requests before cache gets re-created,\nand [poorly designed] cache architecture makes all N do the work\nregenerating cache.  There are several ideas of how to deal with this\nproblem:\n\n * If (part of) cache has expired, set its expiration time to the\n   current time plus specified duration (slop) needed to regenerate\n   cache.  It was used by original Pasky solution (and is used by\n   further solutions for caching projects list sent here); in can be\n   used by CHI (caching infrastructure) with busy_lock option... well,\n   kind of.\n\n * Use some kind of locking so only one process does the work and\n   updates the cache.  From what I've briefly checked that is what\n   kernel.org gitweb does (using flock()).\n\n   The patch implementing projects list info caching does protect\n   using O_EXCL on temporary/lock file against more than one process\n   writing the cache, but doesn't protect against more than one\n   process doing the work, unformtunately.\n\n * Allows items to expire a little earlier than the stated expiration\n   time to help prevent cache miss stampedes.  This is what CHI module\n   does with expires_variance option.\n\n   The probability of expiration increases as a function of how far\n   along we are in the potential expiration window, with the\n   probability being near 0 at the beginning of the window and\n   approaching 1 at the end.\n\nIf cache size becomes issue there will be additional complications\nlike which entries (which cached values) to remove first when we go\nover the cache size limit; but lets us leave it for later, if it would\nbe needed at all.\n\n%%\nIn next part:\n\nCPAN packages we could use, or take inspiration from\n1. Cache::Cache (standard)\n2. CHI - Unified cache interface\n3. Cache - the Cache interface \n4. other interesting packages\n  * Cache::Adaptive for adaptive cache lifetime solutions\n  * Cache::Memcached and/or Cache::Swifty\n    for caching using cache daemon \n  * Cache::FastMmap (also example of callbacks),\n    and caching benchmark mentioned there\n\n-- \nJakub Narebski\nPoland\n"},{"id":"73326","messageId":"200803291813.28393.jnareb@gmail.com","threadId":"12753","inReplyTo":"200803251806.58290.jnareb@gmail.com","subject":"[RFD] Gitweb caching, part 3: examining Perl modules for caching (long)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-03-29T17:13:26Z","receivedAt":"2008-03-29T17:13:26Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Hopefully this resend wouldn't be stopped by vger antispam filter]\n\nIn previous parts:\n\nWhat to cache.\n1. Support for caching in HTTP (external caching).\n2. Caching Perl structures (and serialization).\n3. Caching gitweb output: formatted pages.\n\nCache (in)validation and lifetime\n1. Static cache, external refreshing (invalidation).\n2. Checking filesystem (stat and/or inotify).\n3. Cache lifetime (timing-out cached info).\n4. LRU (Least Recently Used) and others <- to be written\n\nIn this part I will write about existing caching solutions, or to be\nmore exact about CPAN packages implementing caching or cache interface\nin Perl.\n\nNote that not on all sites one can install packages from CPAN; only\npackages from main packages repository, from extras repository, or\nsometimes from trusted contrib repository are possible (see J.H. post\nabout this problem).  From mentioned packages only Cache::Cache and\nCache::Mmap are available in contrib repositories for Aurox 11.1\ndistribution (based on Fedora Core 4): Cache::Cache distribution as\nperl-Cache-Cache in Dries RPM Repository, and Cache::Mmap as\nperl-Cache-Mmap in both Dries and Dag Wieers RPM repositories.\n\n\n1. Cache::Cache (standard)\n\nImplements Cache::MemoryCache, Cache::SharedMemoryCache (using\nIPC::ShareLite), Cache::FileCache and Cache::SizeAwareFileCache.  This\nis the standard: various other modules often say that they implement\nCache::Cache interface.  \n\nIt shows a bit its age, so various improvements exists, including CHI\n- unified cache interface (which can use Cache::Cache modules, and\nalso other caching backends like Cache::FastMmap in a unified way),\nand Cache - the cache interface, which tries to improve Cache::Cache\nbut is not yet complete.\n\nHere is some sample code for instantiating and using file system based\ncache (it uses Storable for serialization, IIRC).\n\n  use Cache::FileCache;\n\n  my $cache = new Cache::FileCache({default_expires_in => \"15 minutes\"});\n\n  my $customer = $cache->get($name);\n\n  if (not defined $customer)  {\n    $customer = get_customer_from_db($name);\n    $cache->set($name, $customer, \"10 minutes\");\n  }\n\n  return $customer;\n\nCache::Cache distribution can be found (at least on RPM based Linux\ndistributions) in perl-Cache-Cache package (e.g. in Dries RPM\nrepository).\n\nVarious other modules implements Cache::Cache interface, for example\nCache::BerkeleyDB (compare with Cache::BDB).\n\n\n2. CHI - Unified cache interface\n\nCHI provides a unified caching API.  The CHI interface is implemented\nby driver classes that support fetching, storing and clearing of data.\n\nCHI is intended as an evolution (and successor) of DeWitt Clinton's\nCache::Cache package, adhering to the basic Cache API but adding new\nfeatures and addressing limitations in the Cache::Cache implementation.\n\nMain goals of CHI were performance (minimizing method calls,\nserializing data only when necessary) and making the creation of new\ndrivers as easy as possible.  \n\nThe latter had lead to wrapping most popular caches available on CPAN in\nCHI interface (CHI handles serialization and expiration times) with\nCHI::Driver::CacheCache, CHI::Driver::FastMmap, CHI::Driver::Memcached.\n\"Native\" CHI drivers include 'File' (one file per entry), 'Memory'\n(per-process) and 'Multilevel' (two or more CHI caches, e.g. memcached\nbolstered by a local memory cache).  'DBI' and 'BerkeleyDB' are\nplanned...\n\nCHI provides expire_if [CODEREF] for additional check if cache expired,\nbusy_lock [DURATION] to set expiration time to current time plus\nspecified duration if value has expired to prevent \"cache stampede\", and\nexpires_variance [FLOAT] to allow items to expire little earlier to\nprevent cache miss stampedes (favored over busy_lock).  Even if gitweb\nwouldn't use CHI interface directly, those ideas are worth considering.\n\nIn addition to standard get() and set() methods it implements\ncompute() method which combines get and set operations in a single\ncall.  It also has some methods to process multiple keys and/or values\nat once.\n\nHere is some sample code for instantiating and using file system based\ncache.\n\n    use CHI;\n\n    # Choose a standard driver\n    #\n    my $cache = CHI->new(driver => 'File', \n                         cache_root => '/tmp/cache');\n\n    # Basic cache operations\n    #\n    my $customer = $cache->get($name);\n    if (!defined $customer) {\n        $customer = get_customer_from_db($name);\n        $cache->set($name, $customer, \"10 minutes\");\n    }\n\n    # or simply\n    my $customer = $cache->compute($name, \\&get_customer_from_db,\n                                   \"10 minutes\");\n\n\n3. Cache - the Cache interface \n\nThe Cache modules are a total redesign and reimplementation of Cache::Cache\nand thus not directly compatible.  Contrary to Cache::Cache get() and\nset() methods do not serialize complex data types; you have to freeze()\nand thaw() data explicitely, instead of set/get.  You can get IO::Handle\nby which data can be read from, or written to cache, e.g. when using\nCache::File.  There is no concept of 'namespace' in the basic cache\ninterface.  Purging is done automatically in current implementation.\n\nCurrently only Cache::File (filesystem based implementation, could be\ndone more efficiently, currently supports only LOCK_NFS locking) and\nCache::Memory (per-process memory based implemetation; with namespaces)\ndrivers are implemented.\n\nIn Cache modules one can select removal strategy for the cache.  By\ndefault FIFO (First In First Out: remove oldest) and LRU (Least Recently\nUsed: remove stalest) strategies are available (when cache has size\nlimit?).\n\nCache modules provide callback interface: load_callback to be called\nanytime when a get() is issued for a data that does not exist in cache,\nand validate_callback (for example storing and checking timestamp or\nsimilar).  This means that sample code for instantiating and using file\nsystem based cache can be written as below.\n\n  use Cache::File;\n\n  my $cache = Cache::File->new(cache_root => '/tmp/cacheroot');\n  $cache->set_load_callback(\\&get_customer_from_db);\n\n  # calls get_customer_from_db() if needed\n  my $customer = $cache->get($name);\n\nThe Cache classes can be used via the tie interface, as shown below.\nThis allows the cache to be accessed via a hash.  All the standard\nmethods for accessing the hash are supported, with the exception\nof the 'keys' or 'each' call.\n\n  tie %hash, 'Cache::File', { cache_root => $tempdir };\n\n  $hash{'key'} = 'some data';\n  $data = $hash{'key'};\n\nThe tie interface is especially useful with the load_callback to\nautomatically populate the hash.\n\n\nEven if gitweb wouldn't use Cache modules (perhaps because lack of\nmatority, or/and the fact that they ar not in extras or trusted\ncontrib packages repository) the idea of selectable removal strategy\nand the idea of callback interface are worth considering; perhaps even\ntie interface.  Whether to serialize explicitely or not... that is\nalso to be decided.\n\n\n4. Other interesting caching packages\n\n4.1. Cache::Adaptive for adaptive cache lifetime control\n\nCache::Adaptive is a cache engine with adaptive lifetime control.  Cache\nlifetimes can be increased or decreased by any factor, e.g. load\naverage, process time for building the cache entry, etc.  Can use almost\nany Cache::Cache object as backend (the update algorithm needs reliable\nset() method, so Cache::SizeAwareFileCache cannot be used).\n\nCache::Adaptive::ByLoad is a subclass of Cache::Adaptive, which adjusts\ncache lifetime by two factors; the load average of the platform and the\npercentage of the total time spent by the builder.\n\nCache::Adaptive Introduces additional \n  access({ key => cache_key, builder => sub { ... } })\nmethod, which returns cached entry if possible, or builds the entry by\ncalling the builder function, and optionally stores the build entry to\ncache.  Compare with compute() method from CHI, or callback interfaces.\n\nWorth examining (both interface and implementation) if/when implementing\ncache lifetime control based on load average, like kernel.org gitweb\ntries to do.  J.H. (kernel.org) fork of gitweb uses longer lifetime\nunder heavier load (within specified bounds).\n\n\n4.2. Cache::Memcached and/or Cache::Swifty for caching using cache daemon\n\nFor larger installations, when there is needed caching not only for\ngitweb, it might be worth examining cache daemon solutions, like\nmemcached (and Cache::Memcached, Cache::Memcached::Fast or CHI\nequivalent), a distributed memory cache daemon; or swifty (and\nCache::Swifty), a very fast shared memory cache, in early alpha stages.\n\nThe Cache::Memcached api, besides set/get methods and administrative\nmethods provide add() and replace() methods to set() conditionally, if\nvalue doesn't exists, or does exists in the cache.\n\nMemcached was created to reduce load on high-trafic site with a hight\ndatabase load that contains mostly read threads, so it might be not\napproproate for gitweb, where I/O load is of most concern, not CPU.\nThe main advantage of memcached is its ability to scale out.  Usually\nyou can run memcached together with web serever or database server, as\nmemcached is CPU lean and memory hungry, and web/database server the\nreverse: CPU hungry and usually memory lean.  Note that gitweb (or\nrather git access to repositories in gitweb) is I/O hungry.\n\n\n4.3. Cache::FastMmap (also example of callbacks),\n     and caching benchmark mentioned there\n\nCache::FastMmap uses an mmap'ed file to act as a shared memory\ninterprocess cache.  It uses fcntl locking to ensure multiple processes\ncan safely access the cache at the same time. It uses a basic LRU\nalgorithm to keep the most used entries in the cache, plus (optionally)\ncache timeout.\n\nCache::FastMmap was created to be very fast.\n\nThe class also supports read-through, and write-back or write-through\ncallbacks to access the real data if it's not in the cache.  With those\nthe code to deal with cache can be written simply as\n\n  Cache::FastMmap->new(\n    ...\n    context => $RealDataSourceHandle,\n    read_cb  => sub { $_[0]->get($_[1]) },\n    write_cb => sub { $_[0]->set($_[1], $_[2]) },\n  );\n\n  ...\n\n  my $value = $cache->get($key);\n\n  $cache->set($key, $newvalue);\n\nIt also supports get_and_set() subroutine to atomically retrieve and set\nvalue of given key, and has methods dealing with multiple keys at once.\nThere is also Cache::FastMmap::Tie module to use tie interface to\nCache::FastMap.  Even if gitweb wouldn't use this module, the callback\nbased interface is worth considering to implement.\n\n\n4.4. CGI::Cache to help cache output of time-intensive CGI scripts\n     with minimal changes to CGI script code.\n\nThis module is intended to be used in CGI scripts that may benefit from\ncaching; it is written in such a way that existing CGI code could get\ncaching added with minimal changes to script.  Here's a simple example:\n\n  #!/usr/bin/perl\n\n  use CGI;\n  use CGI::Cache;\n\n  # Set up cache\n  CGI::Cache::setup();\n\n  my $cgi = new CGI;\n\n  # CGI::Vars requires CGI version 2.50 or better\n  CGI::Cache::set_key($cgi->Vars);\n\n  # This should short-circuit the rest of the loop if a cache value is\n  # already there\n  CGI::Cache::start() or exit;\n\n  print $cgi->header, \"\\n\";\n\n  #...\n\n  print <<EOF;\n  This prints to STDOUT, which will be cached.\n  If the next visit is within 24 hours, the cached STDOUT\n  will be served instead of executing this 'print'.\n  EOF\n\nCGI::Cache module ties the output file descriptor (usually STDOUT) to an\ninternal variable to which all output is saved.  This trick (technique)\nis worth considering if we decide on caching final output in gitweb, or\nfinal output without HTTP headers.\n\n\n5. Summary\n\nIf it is decided that gitweb would do caching of Perl structures, we\nwould certainly use Storable, which should be installed as part of Perl\ninstallation on most systems.  Perhaps gitweb could use Cache::Cache\npackages in general (and Cache::FileCache in particular), as it should\nfill \"extras or trusted contrib\" criterion, but I'd rather not add\nanother dependency to gitweb, especially that not all installations need\ncaching.  It could be good solution for gitweb fork, and I guess\nkernel.org gitweb could use it.\n\nIf gitweb is to implement it's own solutions to not introduce extra\ndependencies, and it would cache Perl structures, implementing\nCache::Cache get/set interface, with possible improvements of callback\ninterface would be a good idea.  For very large installations it would\nbe good to check memcached solution (or multilevel cache, see CHI).\n\nIf gitweb is to use caching of output, or output without HTTP headers,\neither using CGI::Cache or using its technique would be a good idea.\n\nGitweb caching is meant to reduce load (mainly I/O load according to\nsome mails send on this mailing list by J.H., kernel.org gitweb admin,\nand Pasky, repo.or.cz gitweb admin).  I think it would be good to try\nand check, benchmarking if possible, different solutions to \"thundering\nhorde\" aka \"cache stampede\" problem, and to adaptive cache lifetime\ncontrol (see Cache::Adaptive).\n\nThoughts? Comments?\n\n%%\nIn the next part I'd like to have thoughts and ideas for gitweb caching\nfrom J.H. and Petr 'Pasky' Baudis...\n\n\nReferences:\n===========\n[1] http://search.cpan.org\n[2] http://code.google.com/p/perl-cache\n[3] http://www.danga.com/memcached/\n[4] http://cpan.robm.fastmail.fm/cache_perf.html\n\n-- \nJakub Narebski\nPoland\n"}]}