{"thread":{"id":"21314","subject":"Re: [Foundation-l] Wikipedia meets git","startedAt":"2009-10-21T19:49:27Z","lastAt":"2009-10-22T06:27:14Z","messageCount":6,"participants":["Bernie Innocenti","jamesmikedupont@googlemail.com","Avery Pennarun","Nicolas Pitre","David Gerard","jamesmikedupont-gM/Ye1E23mwN+BqQ9rBEUg@public.gmane.org"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"125635","messageId":"1256154567.1477.87.camel@giskard","threadId":"21314","inReplyTo":"5396c0d10910210543i4c0a3350je5bee4c6389a2292@mail.gmail.com","subject":"Re: [Foundation-l] Wikipedia meets git","fromName":"Bernie Innocenti","fromEmail":"bernie@codewiz.org","sentAt":"2009-10-21T19:49:27Z","receivedAt":"2009-10-21T19:49:27Z","isPatch":false,"sender":{"key":"bernie@codewiz.org","avatar":"https://gravatar.com/avatar/42735a3728f3ff3f989a12e9d6931e6e742865fedc6f5eca63cdb327eb820720?d=mp&s=160"},"body":"[cc+=git@vger.kernel.org]\n\nEl Wed, 21-10-2009 a las 08:43 -0400, Samuel Klein escribió:\n> That sounds like a great idea.  I know a few other people who have\n> worked on git-based wikis and toyed with making them compatible with\n> mediawiki (copying bernie innocenti, one of the most eloquent :).\n\nThen I'll do my best to sound as eloquent as expected :)\n\nWhile I think git's internal structure is wonderfully simple and\nelegant, I'm a little worried about its scalability in the wiki usecase.\n\nThe scenario for which git's repository format was designed is \"patch\noriented\" revision control of a filesystem tree. The central object of a\ngit tree is the \"commit\", which represents a set of changes on multiple\nfiles. I'll disregard all the juicy details on how the changes are\nactually packed together to save disk space, making git's repository\nformat amazingly compact.\n\nCommits are linked to each other in order to represent the history. Git\ncan efficiently represent a highly non-linear history with thousands of\nbranches, each containing hundreds of thousands revisions. Branching and\nmerging huge trees is so fast that one is left wondering if anything has\nhappened at all.\n\nSo far, so good. This commit-oriented design is great if you want to\ntrack the history *the whole tree* at once, applying related changes to\nmultiple files atomically. In Git, as well as most other version control\nsystems, there's no such thing as a *file* revision! Git manages entire\ntrees. Trees are assigned unique revision numbers (in fact, ugly sha-1\nhashes), and can optionally by tagged or branched at will.\n\nAnd here's the the catch: the history of individual files is not\ndirectly represented in a git repository. It is typically scattered\nacross thousands of commit objects, with no direct links to help find\nthem. If you want to retrieve the log of a file that was changed only 6\ntimes in the entire history of the Linux kernel, you'd have to dig\nthrough *all* of the 170K revisions in the \"master\" branch.\n\nAnd it takes some time even if git is blazingly fast:\n\n bernie@giskard:~/src/kernel/linux-2.6$ time git log  --pretty=oneline REPORTING-BUGS  | wc -l \n 6\n\n real\t0m1.668s\n user\t0m1.416s\n sys\t0m0.210s\n\n(my laptop has a low-power CPU. A fast server would be 8-10x faster).\n\n\nNow, the English Wikipedia seems to have slightly more than 3M articles,\nwith--how many? tenths of millions of revisions for sure. Going through\nthem *every time* one needs to consult the history of a file would be\n100x slower. Tens of seconds. Not acceptable, uh?\n\nIt seems to me that the typical usage pattern of an encyclopedia is to\nchange each article individually. Perhaps I'm underestimating the role\nof bots here. Anyway, there's no consistency *requirement* for mass\nchanges to be applied atomically throughout all the encyclopedia, right?\n\nIn conclusion, the \"tree at a time\" design is going to be a performance\nbottleneck for a large wiki, with no useful application. Unless of\ncourse the concept of changesets was exposed in the UI, which would be\nan interesting idea to explore.\n\nMercurial (Hg) seems to have a better repository layout for the \"one\nfile at a time\" access pattern... Unfortunately, it's also much slower\nthan git for almost any other purpose, sometimes by an order of\nmagnitude. I'm not even sure how well Hg would cope with a repository\ncontaining 3M files and some 30M revisions. The largest Hg tree I've\ndealt with is the \"mozilla central\" repo, which is already unbearably\nslow to work with.\n\nIt would be interesting to compare notes with the other DSCM hackers,\ntoo.\n\n-- \n   // Bernie Innocenti - http://codewiz.org/\n \\X/  Sugar Labs       - http://sugarlabs.org/\n"},{"id":"125639","messageId":"ee9cc730910211308u5a0284d6he483b0904ddb6068@mail.gmail.com","threadId":"21314","inReplyTo":"1256154567.1477.87.camel@giskard","subject":"Re: [Foundation-l] Wikipedia meets git","fromName":"jamesmikedupont@googlemail.com","fromEmail":"jamesmikedupont@googlemail.com","sentAt":"2009-10-21T20:08:39Z","receivedAt":"2009-10-21T20:08:39Z","isPatch":false,"sender":{"key":"jamesmikedupont@googlemail.com","avatar":"https://gravatar.com/avatar/cbdff8736805f90186ab15da223e68e3fa91399fd634bfb07ed90a7af1e18140?d=mp&s=160"},"body":"Wow,\nI am impressed.\nLet me remind you of one thing,\nmost people are working on very small subsets of the data. Very few\npeople will want to have all the data, think about getting all the\nversions from all the git repos, it would be the same.\nMy idea is for smaller chapters who want to get started easily, or\ntowns, regions to host their own branches of relevant data.\nGiven a world full of such servers, the sum would be great but the\nindividual branches needed at one time would be small.\n\nmike\n\nOn Wed, Oct 21, 2009 at 9:49 PM, Bernie Innocenti <bernie@codewiz.org> wrote:\n> [cc+=git@vger.kernel.org]\n>\n> El Wed, 21-10-2009 a las 08:43 -0400, Samuel Klein escribió:\n>> That sounds like a great idea.  I know a few other people who have\n>> worked on git-based wikis and toyed with making them compatible with\n>> mediawiki (copying bernie innocenti, one of the most eloquent :).\n>\n> Then I'll do my best to sound as eloquent as expected :)\n>\n> While I think git's internal structure is wonderfully simple and\n> elegant, I'm a little worried about its scalability in the wiki usecase.\n>\n> The scenario for which git's repository format was designed is \"patch\n> oriented\" revision control of a filesystem tree. The central object of a\n> git tree is the \"commit\", which represents a set of changes on multiple\n> files. I'll disregard all the juicy details on how the changes are\n> actually packed together to save disk space, making git's repository\n> format amazingly compact.\n>\n> Commits are linked to each other in order to represent the history. Git\n> can efficiently represent a highly non-linear history with thousands of\n> branches, each containing hundreds of thousands revisions. Branching and\n> merging huge trees is so fast that one is left wondering if anything has\n> happened at all.\n>\n> So far, so good. This commit-oriented design is great if you want to\n> track the history *the whole tree* at once, applying related changes to\n> multiple files atomically. In Git, as well as most other version control\n> systems, there's no such thing as a *file* revision! Git manages entire\n> trees. Trees are assigned unique revision numbers (in fact, ugly sha-1\n> hashes), and can optionally by tagged or branched at will.\n>\n> And here's the the catch: the history of individual files is not\n> directly represented in a git repository. It is typically scattered\n> across thousands of commit objects, with no direct links to help find\n> them. If you want to retrieve the log of a file that was changed only 6\n> times in the entire history of the Linux kernel, you'd have to dig\n> through *all* of the 170K revisions in the \"master\" branch.\n>\n> And it takes some time even if git is blazingly fast:\n>\n>  bernie@giskard:~/src/kernel/linux-2.6$ time git log  --pretty=oneline REPORTING-BUGS  | wc -l\n>  6\n>\n>  real   0m1.668s\n>  user   0m1.416s\n>  sys    0m0.210s\n>\n> (my laptop has a low-power CPU. A fast server would be 8-10x faster).\n>\n>\n> Now, the English Wikipedia seems to have slightly more than 3M articles,\n> with--how many? tenths of millions of revisions for sure. Going through\n> them *every time* one needs to consult the history of a file would be\n> 100x slower. Tens of seconds. Not acceptable, uh?\n>\n> It seems to me that the typical usage pattern of an encyclopedia is to\n> change each article individually. Perhaps I'm underestimating the role\n> of bots here. Anyway, there's no consistency *requirement* for mass\n> changes to be applied atomically throughout all the encyclopedia, right?\n>\n> In conclusion, the \"tree at a time\" design is going to be a performance\n> bottleneck for a large wiki, with no useful application. Unless of\n> course the concept of changesets was exposed in the UI, which would be\n> an interesting idea to explore.\n>\n> Mercurial (Hg) seems to have a better repository layout for the \"one\n> file at a time\" access pattern... Unfortunately, it's also much slower\n> than git for almost any other purpose, sometimes by an order of\n> magnitude. I'm not even sure how well Hg would cope with a repository\n> containing 3M files and some 30M revisions. The largest Hg tree I've\n> dealt with is the \"mozilla central\" repo, which is already unbearably\n> slow to work with.\n>\n> It would be interesting to compare notes with the other DSCM hackers,\n> too.\n>\n> --\n>   // Bernie Innocenti - http://codewiz.org/\n>  \\X/  Sugar Labs       - http://sugarlabs.org/\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"},{"id":"125641","messageId":"32541b130910211331n4f65c2d4ga76ac90816fe45d@mail.gmail.com","threadId":"21314","inReplyTo":"1256154567.1477.87.camel@giskard","subject":"Re: [Foundation-l] Wikipedia meets git","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2009-10-21T20:31:20Z","receivedAt":"2009-10-21T20:31:20Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Oct 21, 2009 at 3:49 PM, Bernie Innocenti <bernie@codewiz.org> wrote:\n> And here's the the catch: the history of individual files is not\n> directly represented in a git repository. It is typically scattered\n> across thousands of commit objects, with no direct links to help find\n> them. If you want to retrieve the log of a file that was changed only 6\n> times in the entire history of the Linux kernel, you'd have to dig\n> through *all* of the 170K revisions in the \"master\" branch.\n>\n> And it takes some time even if git is blazingly fast:\n>\n>  bernie@giskard:~/src/kernel/linux-2.6$ time git log  --pretty=oneline REPORTING-BUGS  | wc -l\n>  6\n>\n>  real   0m1.668s\n>  user   0m1.416s\n>  sys    0m0.210s\n>\n> (my laptop has a low-power CPU. A fast server would be 8-10x faster).\n>\n>\n> Now, the English Wikipedia seems to have slightly more than 3M articles,\n> with--how many? tenths of millions of revisions for sure. Going through\n> them *every time* one needs to consult the history of a file would be\n> 100x slower. Tens of seconds. Not acceptable, uh?\n\nI think this slowness could be overcome using a simple cache of\nfilename -> commitid list, right?\n\nThat is, you run some variant of \"git log --name-only\" and, for each\nfile changed by each commit, add an element to the commit list for\nthat file.  When committing in the future, use a hook that updates the\ncache.  When you want to view the history of a particular file, simply\nretrieve exactly the list of commits in that file's commitlist, not\nother commits.\n\nIt sounds like such a cache could be implemented quite easily outside\nof git itself.\n\nWould that help?\n\nThat said, I'll bet you find other performance glitches when you\nimport millions of files and tens/hundreds of millions of commits.\nBut we probably won't know what those problems are until someone\nimports them :)\n\nHave fun,\n\nAvery\n"},{"id":"125647","messageId":"alpine.LFD.2.00.0910211656350.21460@xanadu.home","threadId":"21314","inReplyTo":"1256154567.1477.87.camel@giskard","subject":"Re: [Foundation-l] Wikipedia meets git","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2009-10-21T21:05:48Z","receivedAt":"2009-10-21T21:05:48Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 21 Oct 2009, Bernie Innocenti wrote:\n\n> And here's the the catch: the history of individual files is not\n> directly represented in a git repository. It is typically scattered\n> across thousands of commit objects, with no direct links to help find\n> them. If you want to retrieve the log of a file that was changed only 6\n> times in the entire history of the Linux kernel, you'd have to dig\n> through *all* of the 170K revisions in the \"master\" branch.\n> \n> And it takes some time even if git is blazingly fast:\n> \n>  bernie@giskard:~/src/kernel/linux-2.6$ time git log  --pretty=oneline REPORTING-BUGS  | wc -l \n>  6\n> \n>  real\t0m1.668s\n>  user\t0m1.416s\n>  sys\t0m0.210s\n> \n> (my laptop has a low-power CPU. A fast server would be 8-10x faster).\n> \n> \n> Now, the English Wikipedia seems to have slightly more than 3M articles,\n> with--how many? tenths of millions of revisions for sure. Going through\n> them *every time* one needs to consult the history of a file would be\n> 100x slower. Tens of seconds. Not acceptable, uh?\n> \n> It seems to me that the typical usage pattern of an encyclopedia is to\n> change each article individually. Perhaps I'm underestimating the role\n> of bots here. Anyway, there's no consistency *requirement* for mass\n> changes to be applied atomically throughout all the encyclopedia, right?\n\nYou certainly don't need to put all files in the same tree then.  \nHaving the whole thing split according to some sections that are \nunlikely to overlap would be the way to go.  Therefore you could arrange \nsubsections to have their own branches with no other files in them, or \neven rely on Git submodules.  The partitioning doesn't necessarily have \nto be one of the two extremes such as one branch per file à la CVS or \nall files in the same branch/tree as Git does by default.\n\n\nNicolas\n"},{"id":"125668","messageId":"fbad4e140910211636hd772962x4535ccbda6faa3c7@mail.gmail.com","threadId":"21314","inReplyTo":"ee9cc730910211308u5a0284d6he483b0904ddb6068@mail.gmail.com","subject":"Re: [Foundation-l] Wikipedia meets git","fromName":"David Gerard","fromEmail":"dgerard@gmail.com","sentAt":"2009-10-21T23:36:59Z","receivedAt":"2009-10-21T23:36:59Z","isPatch":false,"sender":{"key":"dgerard@gmail.com","avatar":null},"body":"2009/10/21 jamesmikedupont@googlemail.com <jamesmikedupont@googlemail.com>:\n\n> most people are working on very small subsets of the data. Very few\n> people will want to have all the data, think about getting all the\n> versions from all the git repos, it would be the same.\n> My idea is for smaller chapters who want to get started easily, or\n> towns, regions to host their own branches of relevant data.\n> Given a world full of such servers, the sum would be great but the\n> individual branches needed at one time would be small.\n\n\nA distributed backend is a nice idea anyway - imagine a meteor hitting\nthe Florida data centres ...\n\nAnd there are third-party users who could benefit from a highly\ndistributed backend, such as Wikileaks.\n\nThis thread should probably move to mediawiki-l ...\n\n\n- d.\n"},{"id":"125675","messageId":"ee9cc730910212327i6ecdbd4fw933e62d08c4c46ec@mail.gmail.com","threadId":"21314","inReplyTo":"fbad4e140910211636hd772962x4535ccbda6faa3c7-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org","subject":"Re: Wikipedia meets git","fromName":"jamesmikedupont-gM/Ye1E23mwN+BqQ9rBEUg@public.gmane.org","fromEmail":"jamesmikedupont-gm/ye1e23mwn+bqq9rbeug@public.gmane.org","sentAt":"2009-10-22T06:27:14Z","receivedAt":"2009-10-22T06:27:14Z","isPatch":false,"sender":{"key":"jamesmikedupont-gm/ye1e23mwn+bqq9rbeug@public.gmane.org","avatar":null},"body":"Ok,\nI have started a google group called mediawiki-vcs\n\n\n\n              http://groups.google.com/group/mediawiki-vcs\n\nWe should just move the discussion there.\nAdditionaly, I did not name it git, but vcs, for the reason that we\nshould support multiple backends via a plugin. I am interested in\nusing git because i think git is great, but others should be free to\nuse cvs if they feel it is needed.\n\nmike\n\n\nOn Thu, Oct 22, 2009 at 1:36 AM, David Gerard <dgerard-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:\n> 2009/10/21 jamesmikedupont-gM/Ye1E23mwN+BqQ9rBEUg@public.gmane.org <jamesmikedupont-gM/Ye1E23mwN+BqQ9rBEUg@public.gmane.org>:\n>\n>> most people are working on very small subsets of the data. Very few\n>> people will want to have all the data, think about getting all the\n>> versions from all the git repos, it would be the same.\n>> My idea is for smaller chapters who want to get started easily, or\n>> towns, regions to host their own branches of relevant data.\n>> Given a world full of such servers, the sum would be great but the\n>> individual branches needed at one time would be small.\n>\n>\n> A distributed backend is a nice idea anyway - imagine a meteor hitting\n> the Florida data centres ...\n>\n> And there are third-party users who could benefit from a highly\n> distributed backend, such as Wikileaks.\n>\n> This thread should probably move to mediawiki-l ...\n>\n>\n> - d.\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n_______________________________________________\nfoundation-l mailing list\nfoundation-l-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org\nUnsubscribe: https://lists.wikimedia.org/mailman/listinfo/foundation-l\n"}]}