{"thread":{"id":"34043","subject":"is there a fast web-interface to git for huge repos?","startedAt":"2013-06-07T01:35:43Z","lastAt":"2013-06-14T10:55:40Z","messageCount":8,"participants":["Constantine A. Murenin","Fredrik Gustafsson","Charles McGarvey","Holger Hellmuth (IKS)"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"219578","messageId":"CAPKkNb4bYfBeqkBKqe-22iJsqjmvrYNSe4oWUnPo7QeghLK59Q@mail.gmail.com","threadId":"34043","inReplyTo":null,"subject":"is there a fast web-interface to git for huge repos?","fromName":"Constantine A. Murenin","fromEmail":"mureninc@gmail.com","sentAt":"2013-06-07T01:35:43Z","receivedAt":"2013-06-07T01:35:43Z","isPatch":false,"sender":{"key":"mureninc@gmail.com","avatar":null},"body":"Hi,\n\nOn a relatively-empty Intel Core i7 975 @ 3.33GHz (quad-core):\n\nCns# cd DragonFly/\n\nCns# time git log sys/sys/sockbuf.h >/dev/null\n0.540u 0.140s 0:04.30 15.8%     0+0k 2754+55io 6484pf+0w\nCns# time git log sys/sys/sockbuf.h > /dev/null\n0.000u 0.030s 0:00.52 5.7%      0+0k 0+0io 0pf+0w\nCns# time git log sys/sys/sockbuf.h > /dev/null\n0.180u 0.020s 0:00.52 38.4%     0+0k 0+2io 0pf+0w\nCns# time git log sys/sys/sockbuf.h > /dev/null\n0.420u 0.020s 0:00.52 84.6%     0+0k 0+0io 0pf+0w\n\nAnd, right away, a semi-cold git-blame:\n\nCns# time git blame sys/sys/sockbuf.h >/dev/null\n0.340u 0.040s 0:01.91 19.8%     0+0k 769+45io 2078pf+0w\nCns# time git blame sys/sys/sockbuf.h > /dev/null\n0.340u 0.010s 0:00.36 97.2%     0+0k 0+2io 0pf+0w\nCns# time git blame sys/sys/sockbuf.h > /dev/null\n0.310u 0.040s 0:00.36 97.2%     0+0k 0+0io 0pf+0w\nCns# time git blame sys/sys/sockbuf.h > /dev/null\n0.310u 0.050s 0:00.36 100.0%    0+0k 0+0io 0pf+0w\n\n\nI'm interested in running a web interface to this and other similar\ngit repositories (FreeBSD and NetBSD git repositories are even much,\nmuch bigger).\n\nSoftware-wise, is there no way to make cold access for git-log and\ngit-blame to be orders of magnitude less than ~5s, and warm access\nless than ~0.5s?\n\nC.\n"},{"id":"219609","messageId":"20130607063353.GB19771@paksenarrion.iveqy.com","threadId":"34043","inReplyTo":"CAPKkNb4bYfBeqkBKqe-22iJsqjmvrYNSe4oWUnPo7QeghLK59Q@mail.gmail.com","subject":"Re: is there a fast web-interface to git for huge repos?","fromName":"Fredrik Gustafsson","fromEmail":"iveqy@iveqy.com","sentAt":"2013-06-07T06:33:53Z","receivedAt":"2013-06-07T06:33:53Z","isPatch":false,"sender":{"key":"iveqy@iveqy.com","avatar":"https://avatars.githubusercontent.com/u/761743?v=4"},"body":"On Thu, Jun 06, 2013 at 06:35:43PM -0700, Constantine A. Murenin wrote:\n> I'm interested in running a web interface to this and other similar\n> git repositories (FreeBSD and NetBSD git repositories are even much,\n> much bigger).\n> \n> Software-wise, is there no way to make cold access for git-log and\n> git-blame to be orders of magnitude less than ~5s, and warm access\n> less than ~0.5s?\n\nThe obvious way would be to cache the results. You can even put an\nupdate cache hook the git repositories to make the cache always be up to\ndate.\n\nThere's some dynamic web frontends like cgit and gitweb out there but\nthere's also static ones like git-arr ( http://blitiri.com.ar/p/git-arr/\n) that might be more of an option to you.\n\n-- \nMed vänliga hälsningar\nFredrik Gustafsson\n\ntel: 0733-608274\ne-post: iveqy@iveqy.com\n"},{"id":"219650","messageId":"CAPKkNb5PyurX1eNsCsckdfiwgM3dqb5KpN9OS0NpLZw1+VsSdg@mail.gmail.com","threadId":"34043","inReplyTo":"20130607063353.GB19771@paksenarrion.iveqy.com","subject":"Re: is there a fast web-interface to git for huge repos?","fromName":"Constantine A. Murenin","fromEmail":"mureninc@gmail.com","sentAt":"2013-06-07T17:05:37Z","receivedAt":"2013-06-07T17:05:37Z","isPatch":false,"sender":{"key":"mureninc@gmail.com","avatar":null},"body":"On 6 June 2013 23:33, Fredrik Gustafsson <iveqy@iveqy.com> wrote:\n> On Thu, Jun 06, 2013 at 06:35:43PM -0700, Constantine A. Murenin wrote:\n>> I'm interested in running a web interface to this and other similar\n>> git repositories (FreeBSD and NetBSD git repositories are even much,\n>> much bigger).\n>>\n>> Software-wise, is there no way to make cold access for git-log and\n>> git-blame to be orders of magnitude less than ~5s, and warm access\n>> less than ~0.5s?\n>\n> The obvious way would be to cache the results. You can even put an\n\nThat would do nothing to prevent slowness of the cold requests, which\nalready run for 5s when completely cold.\n\nIn fact, unless done right, it would actually slow things down, as\nlines would not necessarily show up as they're ready.\n\n> update cache hook the git repositories to make the cache always be up to\n> date.\n\nThat's entirely inefficient.  It'll probably take hours or days to\npre-cache all the html pages with a naive wget and the list of all the\nfiles.  Not a solution at all.\n\n(0.5s x 35k files = 5 hours for log/blame, plus another 5h of cpu time\nfor blame/log)\n\n> There's some dynamic web frontends like cgit and gitweb out there but\n> there's also static ones like git-arr ( http://blitiri.com.ar/p/git-arr/\n> ) that might be more of an option to you.\n\nThe concept for git-arr looks interesting, but it has neither blame\nnor log, so, it's kinda pointless, because the whole thing that's slow\nis exactly blame and log.\n\nThere has to be some way to improve these matters.  Noone wants to\nwait 5 seconds until a page is generated, we're not running enterprise\nsoftware here, latency is important!\n\nC.\n"},{"id":"219656","messageId":"20130607175717.GA25127@paksenarrion.iveqy.com","threadId":"34043","inReplyTo":"CAPKkNb5PyurX1eNsCsckdfiwgM3dqb5KpN9OS0NpLZw1+VsSdg@mail.gmail.com","subject":"Re: is there a fast web-interface to git for huge repos?","fromName":"Fredrik Gustafsson","fromEmail":"iveqy@iveqy.com","sentAt":"2013-06-07T17:57:17Z","receivedAt":"2013-06-07T17:57:17Z","isPatch":false,"sender":{"key":"iveqy@iveqy.com","avatar":"https://avatars.githubusercontent.com/u/761743?v=4"},"body":"On Fri, Jun 07, 2013 at 10:05:37AM -0700, Constantine A. Murenin wrote:\n> On 6 June 2013 23:33, Fredrik Gustafsson <iveqy@iveqy.com> wrote:\n> > On Thu, Jun 06, 2013 at 06:35:43PM -0700, Constantine A. Murenin wrote:\n> >> I'm interested in running a web interface to this and other similar\n> >> git repositories (FreeBSD and NetBSD git repositories are even much,\n> >> much bigger).\n> >>\n> >> Software-wise, is there no way to make cold access for git-log and\n> >> git-blame to be orders of magnitude less than ~5s, and warm access\n> >> less than ~0.5s?\n> >\n> > The obvious way would be to cache the results. You can even put an\n> \n> That would do nothing to prevent slowness of the cold requests, which\n> already run for 5s when completely cold.\n> \n> In fact, unless done right, it would actually slow things down, as\n> lines would not necessarily show up as they're ready.\n\nYou need to cache this _before_ the web-request. Don't let the\nweb-request trigger a cache-update but a git push to the repository.\n\n> \n> > update cache hook the git repositories to make the cache always be up to\n> > date.\n> \n> That's entirely inefficient.  It'll probably take hours or days to\n> pre-cache all the html pages with a naive wget and the list of all the\n> files.  Not a solution at all.\n> \n> (0.5s x 35k files = 5 hours for log/blame, plus another 5h of cpu time\n> for blame/log)\n\nThat's a one-time penalty. Why would that be a problem? And why is wget\neven mentioned? Did we misunderstood eachother?\n\n> \n> > There's some dynamic web frontends like cgit and gitweb out there but\n> > there's also static ones like git-arr ( http://blitiri.com.ar/p/git-arr/\n> > ) that might be more of an option to you.\n> \n> The concept for git-arr looks interesting, but it has neither blame\n> nor log, so, it's kinda pointless, because the whole thing that's slow\n> is exactly blame and log.\n> \n> There has to be some way to improve these matters.  Noone wants to\n> wait 5 seconds until a page is generated, we're not running enterprise\n> software here, latency is important!\n> \n> C.\n\nGit's internal structures make just blame pretty expensive. There's\nnothing you really can do for it algoritm wise (as far as I know, if\nthere was, people would already improved it).\n\nThe solution here is to have a \"hot\" repository to speed up things.\n\nThere's of course little things you can do. I imagine that using git\nrepack in a sane way probably could speed things up, as well as git gc.\n\n-- \nMed vänliga hälsningar\nFredrik Gustafsson\n\ntel: 0733-608274\ne-post: iveqy@iveqy.com\n"},{"id":"219673","messageId":"CAPKkNb4myh9MPNSgLqs5Mku-z1EOsHyWrgK2Qy_3_UOivXvcnw@mail.gmail.com","threadId":"34043","inReplyTo":"20130607175717.GA25127@paksenarrion.iveqy.com","subject":"Re: is there a fast web-interface to git for huge repos?","fromName":"Constantine A. Murenin","fromEmail":"mureninc@gmail.com","sentAt":"2013-06-07T19:02:36Z","receivedAt":"2013-06-07T19:02:36Z","isPatch":false,"sender":{"key":"mureninc@gmail.com","avatar":null},"body":"On 7 June 2013 10:57, Fredrik Gustafsson <iveqy@iveqy.com> wrote:\n> On Fri, Jun 07, 2013 at 10:05:37AM -0700, Constantine A. Murenin wrote:\n>> On 6 June 2013 23:33, Fredrik Gustafsson <iveqy@iveqy.com> wrote:\n>> > On Thu, Jun 06, 2013 at 06:35:43PM -0700, Constantine A. Murenin wrote:\n>> >> I'm interested in running a web interface to this and other similar\n>> >> git repositories (FreeBSD and NetBSD git repositories are even much,\n>> >> much bigger).\n>> >>\n>> >> Software-wise, is there no way to make cold access for git-log and\n>> >> git-blame to be orders of magnitude less than ~5s, and warm access\n>> >> less than ~0.5s?\n>> >\n>> > The obvious way would be to cache the results. You can even put an\n>>\n>> That would do nothing to prevent slowness of the cold requests, which\n>> already run for 5s when completely cold.\n>>\n>> In fact, unless done right, it would actually slow things down, as\n>> lines would not necessarily show up as they're ready.\n>\n> You need to cache this _before_ the web-request. Don't let the\n> web-request trigger a cache-update but a git push to the repository.\n>\n>>\n>> > update cache hook the git repositories to make the cache always be up to\n>> > date.\n>>\n>> That's entirely inefficient.  It'll probably take hours or days to\n>> pre-cache all the html pages with a naive wget and the list of all the\n>> files.  Not a solution at all.\n>>\n>> (0.5s x 35k files = 5 hours for log/blame, plus another 5h of cpu time\n>> for blame/log)\n>\n> That's a one-time penalty. Why would that be a problem? And why is wget\n> even mentioned? Did we misunderstood eachother?\n\n`wget` or `curl --head` would be used to trigger the caching.\n\nI don't understand how it's a one-time penalty.  Noone wants to look\nat an old copy of the repository, so, pretty much, if, say, I want to\nhave a gitweb of all 4 BSDs, updated daily, then, pretty much, even\nwith lots of ram (e.g. to eliminate the cold-case 5s penalty, and\nreduce each page to 0.5s), on a quad-core box, I'd be kinda be lucky\nto complete a generation of all the pages within 12h or so, obviously\nusing the machine at, or above, 50% capacity just for the caching.  Or\nseveral days or even a couple of weeks on an Intel Atom or VIA Nano\nwith 2GB of RAM or so.  Obviously not acceptable, there has to be a\nbetter solution.\n\nOne could, I guess, only regenerate the pages which have changed, but\nit still sounds like an ugly solution, where you'd have to be\ngenerating a list of files that have changed between one gen and the\nnext, and you'd still have to have a very high cpu, cache and storage\nrequirements.\n\nC.\n\n>> > There's some dynamic web frontends like cgit and gitweb out there but\n>> > there's also static ones like git-arr ( http://blitiri.com.ar/p/git-arr/\n>> > ) that might be more of an option to you.\n>>\n>> The concept for git-arr looks interesting, but it has neither blame\n>> nor log, so, it's kinda pointless, because the whole thing that's slow\n>> is exactly blame and log.\n>>\n>> There has to be some way to improve these matters.  Noone wants to\n>> wait 5 seconds until a page is generated, we're not running enterprise\n>> software here, latency is important!\n>>\n>> C.\n>\n> Git's internal structures make just blame pretty expensive. There's\n> nothing you really can do for it algoritm wise (as far as I know, if\n> there was, people would already improved it).\n>\n> The solution here is to have a \"hot\" repository to speed up things.\n>\n> There's of course little things you can do. I imagine that using git\n> repack in a sane way probably could speed things up, as well as git gc.\n>\n> --\n> Med vänliga hälsningar\n> Fredrik Gustafsson\n>\n> tel: 0733-608274\n> e-post: iveqy@iveqy.com\n"},{"id":"219693","messageId":"51B23F01.5020608@brokenzipper.com","threadId":"34043","inReplyTo":"CAPKkNb4myh9MPNSgLqs5Mku-z1EOsHyWrgK2Qy_3_UOivXvcnw@mail.gmail.com","subject":"Re: is there a fast web-interface to git for huge repos?","fromName":"Charles McGarvey","fromEmail":"chazmcgarvey@brokenzipper.com","sentAt":"2013-06-07T20:13:53Z","receivedAt":"2013-06-07T20:13:53Z","isPatch":false,"sender":{"key":"chazmcgarvey@brokenzipper.com","avatar":"https://avatars.githubusercontent.com/u/108998?v=4"},"body":"On 06/07/2013 01:02 PM, Constantine A. Murenin wrote:\n>> That's a one-time penalty. Why would that be a problem? And why is wget\n>> even mentioned? Did we misunderstood eachother?\n> \n> `wget` or `curl --head` would be used to trigger the caching.\n> \n> I don't understand how it's a one-time penalty.  Noone wants to look\n> at an old copy of the repository, so, pretty much, if, say, I want to\n> have a gitweb of all 4 BSDs, updated daily, then, pretty much, even\n> with lots of ram (e.g. to eliminate the cold-case 5s penalty, and\n> reduce each page to 0.5s), on a quad-core box, I'd be kinda be lucky\n> to complete a generation of all the pages within 12h or so, obviously\n> using the machine at, or above, 50% capacity just for the caching.  Or\n> several days or even a couple of weeks on an Intel Atom or VIA Nano\n> with 2GB of RAM or so.  Obviously not acceptable, there has to be a\n> better solution.\n> \n> One could, I guess, only regenerate the pages which have changed, but\n> it still sounds like an ugly solution, where you'd have to be\n> generating a list of files that have changed between one gen and the\n> next, and you'd still have to have a very high cpu, cache and storage\n> requirements.\n\nHave you already ruled out caching on a proxy?  Pages would only be generated\non demand, so the first visitor would still experience the delay but the rest\nwould be fast until the page expires.  Even expiring pages as often as five\nminutes or less would probably provide significant processing savings\n(depending on how many users you have), and that level of staleness and the\noccasional delays may be acceptable to your users.\n\nAs you say, generating the entire cache upfront and continuously is wasteful\nand probably unrealistic, but any type of caching, by definition, is going to\ninvolve users seeing stale content, and I don't see that you have any other\noption but some type of caching.  Well, you could reproduce what git does in a\nbunch of distributed algorithms and run your app on a farm--which, I guess, is\nprobably what GitHub is doing--but throwing up a caching reverse proxy is a\nlot quicker if you can accept the caveats.\n\n-- \nCharles McGarvey\n\n"},{"id":"219694","messageId":"CAPKkNb460fNJcwt6084xkuDa2sWMRnF+FBu+i_G01aJMMiRevA@mail.gmail.com","threadId":"34043","inReplyTo":"51B23F01.5020608@brokenzipper.com","subject":"Re: is there a fast web-interface to git for huge repos?","fromName":"Constantine A. Murenin","fromEmail":"mureninc@gmail.com","sentAt":"2013-06-07T20:21:30Z","receivedAt":"2013-06-07T20:21:30Z","isPatch":false,"sender":{"key":"mureninc@gmail.com","avatar":null},"body":"On 7 June 2013 13:13, Charles McGarvey <chazmcgarvey@brokenzipper.com> wrote:\n> On 06/07/2013 01:02 PM, Constantine A. Murenin wrote:\n>>> That's a one-time penalty. Why would that be a problem? And why is wget\n>>> even mentioned? Did we misunderstood eachother?\n>>\n>> `wget` or `curl --head` would be used to trigger the caching.\n>>\n>> I don't understand how it's a one-time penalty.  Noone wants to look\n>> at an old copy of the repository, so, pretty much, if, say, I want to\n>> have a gitweb of all 4 BSDs, updated daily, then, pretty much, even\n>> with lots of ram (e.g. to eliminate the cold-case 5s penalty, and\n>> reduce each page to 0.5s), on a quad-core box, I'd be kinda be lucky\n>> to complete a generation of all the pages within 12h or so, obviously\n>> using the machine at, or above, 50% capacity just for the caching.  Or\n>> several days or even a couple of weeks on an Intel Atom or VIA Nano\n>> with 2GB of RAM or so.  Obviously not acceptable, there has to be a\n>> better solution.\n>>\n>> One could, I guess, only regenerate the pages which have changed, but\n>> it still sounds like an ugly solution, where you'd have to be\n>> generating a list of files that have changed between one gen and the\n>> next, and you'd still have to have a very high cpu, cache and storage\n>> requirements.\n>\n> Have you already ruled out caching on a proxy?  Pages would only be generated\n> on demand, so the first visitor would still experience the delay but the rest\n> would be fast until the page expires.  Even expiring pages as often as five\n> minutes or less would probably provide significant processing savings\n> (depending on how many users you have), and that level of staleness and the\n> occasional delays may be acceptable to your users.\n>\n> As you say, generating the entire cache upfront and continuously is wasteful\n> and probably unrealistic, but any type of caching, by definition, is going to\n> involve users seeing stale content, and I don't see that you have any other\n> option but some type of caching.  Well, you could reproduce what git does in a\n> bunch of distributed algorithms and run your app on a farm--which, I guess, is\n> probably what GitHub is doing--but throwing up a caching reverse proxy is a\n> lot quicker if you can accept the caveats.\n\nI don't think GitHub / Gitorious / whatever have solved this problem\nat all.  They're terribly slow on big repos, some pages don't even\ngenerate the first time you click on the link.\n\nI'm totally fine with daily updates; but I think there still has to be\nsome better way of doing this than wasting 0.5s of CPU time and 5s of\nHDD time (if completely cold) for each blame / log, at the price of\nmore storage and some pre-caching, and (daily (in my use-case))\nfine-grained incremental updates.\n\nC.\n"},{"id":"220804","messageId":"51BAF6AC.2080504@ira.uka.de","threadId":"34043","inReplyTo":"CAPKkNb460fNJcwt6084xkuDa2sWMRnF+FBu+i_G01aJMMiRevA@mail.gmail.com","subject":"Re: is there a fast web-interface to git for huge repos?","fromName":"Holger Hellmuth (IKS)","fromEmail":"hellmuth@ira.uka.de","sentAt":"2013-06-14T10:55:40Z","receivedAt":"2013-06-14T10:55:40Z","isPatch":false,"sender":{"key":"hellmuth@ira.uka.de","avatar":null},"body":"Am 07.06.2013 22:21, schrieb Constantine A. Murenin:\n> I'm totally fine with daily updates; but I think there still has to be\n> some better way of doing this than wasting 0.5s of CPU time and 5s of\n> HDD time (if completely cold) for each blame / log, at the price of\n> more storage and some pre-caching, and (daily (in my use-case))\n> fine-grained incremental updates.\n\nTo get a feel for the numbers: I would guess 'git blame' is mostly run \nagainst the newest version and the release version of a file, right? I \ncouldn't find the number of files in bsd, so lets take linux instead: \nThat is 25k files for version 2.6.27. Lets say 35k files altogether for \nboth release and newer versions of the files.\n\nA typical page of git blame output on github seems to be in the vicinity \nof 500 kbytes, but that seems to include lots of overhead for comfort \nfunctions. At least that means it is a good upper bound value.\n\n35k files times 500k gives 17.5 Gbytes, a trivial value for a static \n*disk* based cache. It is also a manageable value for affordable SSDs\n"}]}