threads / discuss / 27470

Git-Mediawiki : Question about Jeff King's import script

Subject: Git-Mediawiki : Question about Jeff King's import script

## tl;dr

4 messages between May 26, 2011 and May 27, 2011.

replies: 3people: 3as markdown or json

Claire Fousse· May 26, 2011, 15:18 UTC · lore

Dear Jeff King, We are the four students in charge of the Git-Mediawiki project proposed by Matthieu Moy. In case you skipped our email, here is a link to our last mail with a few information about the project http://www.spinics.net/lists/git/msg158701.html We based our script on what you called a few months ago the "quick and dirty perl script" for the import part and have a few questions about it. First of all, just in case, here is your original script : http://article.gmane.org/gmane.comp.version-control.git/167560

It seems like you first used a hashmap for it to be transformed later into a flat list / array. What is the reasoning behind this ? Why not create an array right away ?

Thanks for the script and for any information you can give us,

Regards, The Git-Mediawiki team, Arnaud Lacurie, David Amouyal, Claire Fousse & Jeremie Nikaes.

Jeff King· May 26, 2011, 15:42 UTC · re: Claire Fousse · lore

Re: Git-Mediawiki : Question about Jeff King's import script

On Thu, May 26, 2011 at 05:18:11PM +0200, Claire Fousse wrote:
Show 9 quoted lines
> We based our script on what you called a few months ago the "quick and
> dirty perl script" for the import part and have a few questions about
> it.
> First of all, just in case, here is your original script :
> http://article.gmane.org/gmane.comp.version-control.git/167560
> 
> It seems like you first used a hashmap for it to be transformed later
> into a flat list / array. What is the reasoning behind this ? Why not
> create an array right away ?

The hashmap is actually backed by an on-disk key/value database. The purpose of this was to allow resuming an import that had failed in the middle (since even for a moderate-sized wiki like the git wiki, the import was quite slow).

So the hashmap is indexed by page id, and each value contains an array of revisions for that page. If we see a page id that we've already done, we can skip importing it.

If you wanted to do it all at once, yes, you could build a flat array of revisions, with each revision mentioning the page that it came from, and just keep appending to the array as you read more data from the wiki. And then at the end, sort that array based on timestamp to get the chronological ordering of changes.

Hope that helps, -Peff

Alexandre Dulaunoy· May 27, 2011, 12:45 UTC · re: Claire Fousse · lore

Re: Git-Mediawiki : Question about Jeff King's import script

On Thu, May 26, 2011 at 5:18 PM, Claire Fousse <claire.fousse@gmail.com> wrote:
Show 17 quoted lines
> Dear Jeff King,
> We are the four students in charge of the Git-Mediawiki project
> proposed by Matthieu Moy.
> In case you skipped our email, here is a link to our last mail with a
> few information about the project
> http://www.spinics.net/lists/git/msg158701.html
> We based our script on what you called a few months ago the "quick and
> dirty perl script" for the import part and have a few questions about
> it.
> First of all, just in case, here is your original script :
> http://article.gmane.org/gmane.comp.version-control.git/167560
>
> It seems like you first used a hashmap for it to be transformed later
> into a flat list / array. What is the reasoning behind this ? Why not
> create an array right away ?
>
> Thanks for the script and for any information you can give us,
In a similar spirit, there is this Ruby script:
https://github.com/singpolyma/git-mediawiki/blob/master/clone.rb
using the Mediawiki API. Quite nifty.
Hope this helps,
-- 
--                   Alexandre Dulaunoy (adulau) -- http://www.foo.be/
--                             http://www.foo.be/cgi-bin/wiki.pl/Diary
--         "Knowledge can create problems, it is not through ignorance
--                                that we can solve them" Isaac Asimov

← back to recent threads