{"thread":{"id":"30747","subject":"[PATCHv1] git-remote-mediawiki: import \"File:\" attachments","startedAt":"2012-06-08T14:22:56Z","lastAt":"2012-06-08T23:24:13Z","messageCount":5,"participants":["Pavel Volek","Matthieu Moy","Simon.Cathebras","konglu@minatec.inpg.fr"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"193151","messageId":"1339165376-20267-1-git-send-email-Pavel.Volek@ensimag.imag.fr","threadId":"30747","inReplyTo":null,"subject":"[PATCHv1] git-remote-mediawiki: import \"File:\" attachments","fromName":"Pavel Volek","fromEmail":"pavel.volek@ensimag.imag.fr","sentAt":"2012-06-08T14:22:56Z","receivedAt":"2012-06-08T14:22:56Z","isPatch":false,"sender":{"key":"pavel.volek@ensimag.imag.fr","avatar":null},"body":"From: Volek Pavel <me@pavelvolek.cz>\n\nThe current version of the git-remote-mediawiki supports only import and export\nof the pages, doesn't support import and export of file attachements which are\nalso exposed by MediaWiki API. This patch adds the functionality to import the\nlast versions of the files and all versions of description pages for these\nfiles.\n\nSigned-off-by: Pavel Volek <Pavel.Volek@ensimag.imag.fr>\nSigned-off-by: NGUYEN Kim Thuat <Kim-Thuat.Nguyen@ensimag.imag.fr>\nSigned-off-by: ROUCHER IGLESIAS Javier <roucherj@ensimag.imag.fr>\nSigned-off-by: Matthieu Moy <Matthieu.Moy@imag.fr>\n---\n contrib/mw-to-git/git-remote-mediawiki | 290 +++++++++++++++++++++++++++------\n 1 file changed, 244 insertions(+), 46 deletions(-)\n\ndiff --git a/contrib/mw-to-git/git-remote-mediawiki b/contrib/mw-to-git/git-remote-mediawiki\nindex c18bfa1..9f21217 100755\n--- a/contrib/mw-to-git/git-remote-mediawiki\n+++ b/contrib/mw-to-git/git-remote-mediawiki\n@@ -212,59 +212,230 @@ sub get_mw_pages {\n \tmy $user_defined;\n \tif (@tracked_pages) {\n \t\t$user_defined = 1;\n-\t\t# The user provided a list of pages titles, but we\n-\t\t# still need to query the API to get the page IDs.\n-\n-\t\tmy @some_pages = @tracked_pages;\n-\t\twhile (@some_pages) {\n-\t\t\tmy $last = 50;\n-\t\t\tif ($#some_pages < $last) {\n-\t\t\t\t$last = $#some_pages;\n-\t\t\t}\n-\t\t\tmy @slice = @some_pages[0..$last];\n-\t\t\tget_mw_first_pages(\\@slice, \\%pages);\n-\t\t\t@some_pages = @some_pages[51..$#some_pages];\n-\t\t}\n+\t\tget_mw_tracked_pages(\\%pages);\n \t}\n \tif (@tracked_categories) {\n \t\t$user_defined = 1;\n-\t\tforeach my $category (@tracked_categories) {\n-\t\t\tif (index($category, ':') < 0) {\n-\t\t\t\t# Mediawiki requires the Category\n-\t\t\t\t# prefix, but let's not force the user\n-\t\t\t\t# to specify it.\n-\t\t\t\t$category = \"Category:\" . $category;\n-\t\t\t}\n-\t\t\tmy $mw_pages = $mediawiki->list( {\n-\t\t\t\taction => 'query',\n-\t\t\t\tlist => 'categorymembers',\n-\t\t\t\tcmtitle => $category,\n-\t\t\t\tcmlimit => 'max' } )\n-\t\t\t    || die $mediawiki->{error}->{code} . ': ' . $mediawiki->{error}->{details};\n-\t\t\tforeach my $page (@{$mw_pages}) {\n-\t\t\t\t$pages{$page->{title}} = $page;\n-\t\t\t}\n-\t\t}\n+\t\tget_mw_tracked_categories(\\%pages);\n \t}\n \tif (!$user_defined) {\n-\t\t# No user-provided list, get the list of pages from\n-\t\t# the API.\n-\t\tmy $mw_pages = $mediawiki->list({\n-\t\t\taction => 'query',\n-\t\t\tlist => 'allpages',\n-\t\t\taplimit => 500,\n-\t\t});\n-\t\tif (!defined($mw_pages)) {\n-\t\t\tprint STDERR \"fatal: could not get the list of wiki pages.\\n\";\n-\t\t\tprint STDERR \"fatal: '$url' does not appear to be a mediawiki\\n\";\n-\t\t\tprint STDERR \"fatal: make sure '$url/api.php' is a valid page.\\n\";\n-\t\t\texit 1;\n+\t\t get_mw_all_pages(\\%pages);\n+\t}\n+\treturn values(%pages);\n+}\n+\n+sub get_mw_all_pages {\n+\tmy $pages = shift;\n+\t# No user-provided list, get the list of pages from the API.\n+\tmy $mw_pages = $mediawiki->list({\n+\t\taction => 'query',\n+\t\tlist => 'allpages',\n+\t\taplimit => 500,\n+\t});\n+\tif (!defined($mw_pages)) {\n+\t\tprint STDERR \"fatal: could not get the list of wiki pages.\\n\";\n+\t\tprint STDERR \"fatal: '$url' does not appear to be a mediawiki\\n\";\n+\t\tprint STDERR \"fatal: make sure '$url/api.php' is a valid page.\\n\";\n+\t\texit 1;\n+\t}\n+\tforeach my $page (@{$mw_pages}) {\n+\t\t$pages->{$page->{title}} = $page;\n+\t}\n+\n+\t# Attach list of all pages for meadia files from the API,\n+\t# they are in a different namespace, only one namespace\n+\t# can be queried at the same moment\n+\tmy $mw_mediapages = $mediawiki->list({\n+\t\taction => 'query',\n+\t\tlist => 'allpages',\n+\t\tapnamespace => get_mw_namespace_id(\"File\"),\n+\t\taplimit => 500,\n+\t});\n+\tif (!defined($mw_mediapages)) {\n+\t\tprint STDERR \"fatal: could not get the list of media file pages.\\n\";\n+\t\tprint STDERR \"fatal: '$url' does not appear to be a mediawiki\\n\";\n+\t\tprint STDERR \"fatal: make sure '$url/api.php' is a valid page.\\n\";\n+\t\texit 1;\n+\t}\n+\tforeach my $page (@{$mw_mediapages}) {\n+\t\t$pages->{$page->{title}} = $page;\n+\t}\n+}\n+\n+sub get_mw_tracked_pages {\n+\tmy $pages = shift;\n+\t# The user provided a list of pages titles, but we\n+\t# still need to query the API to get the page IDs.\n+\tmy @some_pages = @tracked_pages;\n+\twhile (@some_pages) {\n+\t\tmy $last = 50;\n+\t\tif ($#some_pages < $last) {\n+\t\t\t$last = $#some_pages;\n+\t\t}\n+\t\tmy @slice = @some_pages[0..$last];\n+\t\tget_mw_first_pages(\\@slice, \\%{$pages});\n+\t\t@some_pages = @some_pages[51..$#some_pages];\n+\t}\n+\n+\t# Get pages of related media files.\n+\tget_mw_linked_mediapages(\\@tracked_pages, \\%{$pages});\n+}\n+\n+sub get_mw_tracked_categories {\n+\tmy $pages = shift;\n+\tforeach my $category (@tracked_categories) {\n+\t\tif (index($category, ':') < 0) {\n+\t\t\t# Mediawiki requires the Category\n+\t\t\t# prefix, but let's not force the user\n+\t\t\t# to specify it.\n+\t\t\t$category = \"Category:\" . $category;\n \t\t}\n+\t\tmy $mw_pages = $mediawiki->list( {\n+\t\t\taction => 'query',\n+\t\t\tlist => 'categorymembers',\n+\t\t\tcmtitle => $category,\n+\t\t\tcmlimit => 'max' } )\n+\t\t\t|| die $mediawiki->{error}->{code} . ': '\n+\t\t\t\t. $mediawiki->{error}->{details};\n \t\tforeach my $page (@{$mw_pages}) {\n-\t\t\t$pages{$page->{title}} = $page;\n+\t\t\t$pages->{$page->{title}} = $page;\n+\t\t}\n+\n+\t\tmy @titles = map $_->{title}, @{$mw_pages};\n+\t\t# Get pages of related media files.\n+\t\tget_mw_linked_mediapages(\\@titles, \\%{$pages});\n+\t}\n+}\n+\n+sub get_mw_linked_mediapages {\n+\tmy $titles = shift;\n+\tmy @titles = @{$titles};\n+\tmy $pages = shift;\n+\n+\t# pattern 'page1|page2|...' required by the API\n+\tmy $mw_titles = join('|', @titles);\n+\n+\t# Media files could be included or linked from\n+\t# a page, get all related\n+\tmy $query = {\n+\t\taction => 'query',\n+\t\tprop => 'links|images',\n+\t\ttitles => $mw_titles,\n+\t\tplnamespace => get_mw_namespace_id(\"File\"),\n+\t\tpllimit => 500,\n+\t};\n+\tmy $result = $mediawiki->api($query);\n+\n+\twhile (my ($id, $page) = each(%{$result->{query}->{pages}})) {\n+\t\tmy @titles;\n+\t\tif (defined($page->{links})) {\n+\t\t\tmy @link_titles = map $_->{title}, @{$page->{links}};\n+\t\t\tpush(@titles, @link_titles);\n+\t\t}\n+\t\tif (defined($page->{images})) {\n+\t\t\tmy @image_titles = map $_->{title}, @{$page->{images}};\n+\t\t\tpush(@titles, @image_titles);\n+\t\t}\n+\t\tif (@titles) {\n+\t\t\tget_mw_first_pages(\\@titles, \\%{$pages});\n \t\t}\n \t}\n-\treturn values(%pages);\n+}\n+\n+sub get_mw_medafile_for_mediapage_revision {\n+\t# Name of the file on Wiki, with the prefix.\n+\tmy $mw_filename = shift;\n+\tmy $timestamp = shift;\n+\tmy %mediafile;\n+\n+\t# Search if on MediaWiki exists a media file with given\n+\t# timestamp and in that case download the file.\n+\tmy $query = {\n+\t\taction => 'query',\n+\t\tprop => 'imageinfo',\n+\t\ttitles => $mw_filename,\n+\t\tiistart => $timestamp,\n+\t\tiiend => $timestamp,\n+\t\tiiprop => 'timestamp|archivename',\n+\t\tiilimit => 1,\n+\t};\n+\tmy $result = $mediawiki->api($query);\n+\n+\tmy ($fileid, $file) = each ( %{$result->{query}->{pages}} );\n+\tif (defined($file->{imageinfo})) {\n+\t\tmy $fileinfo = pop(@{$file->{imageinfo}});\n+\t\tif (defined($fileinfo->{archivename})) {\n+\t\t\treturn; # now we are not able to download files from archive\n+\t\t}\n+\n+\t\tmy $filename; # real filename without prefix\n+\t\tif (index($mw_filename, 'File:') == 0) {\n+\t\t\t$filename = substr $mw_filename, 5;\n+\t\t} else {\n+\t\t\t$filename = substr $mw_filename, 6;\n+\t\t}\n+\n+\t\t$mediafile{title} = $filename;\n+\t\t$mediafile{content} = download_mw_mediafile($mw_filename);\n+\t}\n+\treturn %mediafile;\n+}\n+\n+# Returns MediaWiki id for a canonical namespace name.\n+# Ex.: \"File\", \"Project\".\n+# Looks for the namespace id in the local configuration\n+# variables, if it is not found asks MW API.\n+sub get_mw_namespace_id {\n+\tmw_connect_maybe();\n+\n+\tmy $name = shift;\n+\n+\t# Look at configuration file, if the record\n+\t# for that namespace is already stored.\n+\tmy @tracked_namespaces = split(/[ \\n]/, run_git(\"config --get-all remote.\". $remotename .\".namespaces\"));\n+\tchomp(@tracked_namespaces);\n+\tif (@tracked_namespaces) {\n+\t\tforeach my $ns (@tracked_namespaces) {\n+\t\t\tmy @ns_split = split(/=/, $ns);\n+\t\t\tif ($ns_split[0] eq $name) {\n+\t\t\t\treturn $ns_split[1];\n+\t\t\t}\n+\t\t}\n+\t}\n+\n+\t# NS not found => get namespace id from MW and store it in\n+\t# configuration file.\n+\tmy $query = {\n+\t\taction => 'query',\n+\t\tmeta => 'siteinfo',\n+\t\tsiprop => 'namespaces',\n+\t};\n+\tmy $result = $mediawiki->api($query);\n+\n+\twhile (my ($id, $ns) = each(%{$result->{query}->{namespaces}})) {\n+\t\tif (defined($ns->{canonical}) && ($ns->{canonical} eq $name)) {\n+\t\t\trun_git(\"config --add remote.\". $remotename .\".namespaces \". $name .\"=\". $ns->{id});\n+\t\t\treturn $ns->{id};\n+\t\t}\n+\t}\n+\tdie \"Namespace $name was not found on MediaWiki.\";\n+}\n+\n+sub download_mw_mediafile {\n+\tmy $filename = shift;\n+\n+\t$mediawiki->{config}->{files_url} = $url;\n+\n+\tmy $file = $mediawiki->download( { title => $filename } );\n+\tif (!defined($file)) {\n+\t\tprint STDERR \"\\tFile \\'$filename\\' could not be downloaded.\\n\";\n+\t\texit 1;\n+\t} elsif ($file eq \"\") {\n+\t\tprint STDERR \"\\tFile \\'$filename\\' does not exist on the wiki.\\n\";\n+\t\texit 1;\n+\t} else {\n+\t\treturn $file;\n+\t}\n }\n \n sub run_git {\n@@ -466,6 +637,13 @@ sub import_file_revision {\n \tmy %commit = %{$commit};\n \tmy $full_import = shift;\n \tmy $n = shift;\n+\tmy $mediafile_import = shift;\n+\tmy $mediafile;\n+\tmy %mediafile;\n+\tif ($mediafile_import) {\n+\t\t$mediafile = shift;\n+\t\t%mediafile = %{$mediafile};\n+\t}\n \n \tmy $title = $commit{title};\n \tmy $comment = $commit{comment};\n@@ -485,6 +663,10 @@ sub import_file_revision {\n \tif ($content ne DELETED_CONTENT) {\n \t\tprint STDOUT \"M 644 inline $title.mw\\n\";\n \t\tliteral_data($content);\n+\t\tif ($mediafile_import) {\n+\t\t\tprint STDOUT \"M 644 inline $mediafile{title}\\n\";\n+\t\t\tliteral_data($mediafile{content});\n+\t\t}\n \t\tprint STDOUT \"\\n\\n\";\n \t} else {\n \t\tprint STDOUT \"D $title.mw\\n\";\n@@ -580,6 +762,7 @@ sub mw_import_ref {\n \n \t\t$n++;\n \n+\t\tmy $page_title = $result->{query}->{pages}->{$pagerevid->{pageid}}->{title};\n \t\tmy %commit;\n \t\t$commit{author} = $rev->{user} || 'Anonymous';\n \t\t$commit{comment} = $rev->{comment} || '*Empty MediaWiki Message*';\n@@ -596,9 +779,24 @@ sub mw_import_ref {\n \t\t}\n \t\t$commit{date} = DateTime::Format::ISO8601->parse_datetime($last_timestamp);\n \n-\t\tprint STDERR \"$n/\", scalar(@revisions), \": Revision #$pagerevid->{revid} of $commit{title}\\n\";\n+\t\t# differentiates classic pages and media pages\n+\t\tmy @prefix = split (\":\", $page_title);\n \n-\t\timport_file_revision(\\%commit, ($fetch_from == 1), $n);\n+\t\tmy %mediafile;\n+\t\tif ($prefix[0] eq \"File\" || $prefix[0] eq \"Image\") {\n+\t\t\t# The name of the file is the same as the media page.\n+\t\t\tmy $filename = $page_title;\n+\t\t\t%mediafile = get_mw_medafile_for_mediapage_revision($filename, $rev->{timestamp});\n+\t\t}\n+\t\t# If this is a revision of the media page for new version\n+\t\t# of a file do one common commit for both file and media page.\n+\t\t# Else do commit only for that page.\n+\t\tprint STDERR \"$n/\", scalar(@revisions), \": Revision #$pagerevid->{revid} of $commit{title}\\n\";\n+\t\tif (%mediafile) {\n+\t\t\timport_file_revision(\\%commit, ($fetch_from == 1), $n, 1, \\%mediafile);\n+\t\t} else {\n+\t\t\timport_file_revision(\\%commit, ($fetch_from == 1), $n, 0);\n+\t\t}\n \t}\n \n \tif ($fetch_from == 1 && $n == 0) {\n-- \n1.7.10.2.552.gaa3bb87\n"},{"id":"193153","messageId":"vpqmx4drhjm.fsf@bauges.imag.fr","threadId":"30747","inReplyTo":"1339165376-20267-1-git-send-email-Pavel.Volek@ensimag.imag.fr","subject":"Re: [PATCHv1] git-remote-mediawiki: import \"File:\" attachments","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2012-06-08T14:42:53Z","receivedAt":"2012-06-08T14:42:53Z","isPatch":false,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Pavel Volek <Pavel.Volek@ensimag.imag.fr> writes:\n\n> --- a/contrib/mw-to-git/git-remote-mediawiki\n> +++ b/contrib/mw-to-git/git-remote-mediawiki\n\nIf the patch adds support for [[File:...]], then it should remove/adapt\nthe comment at the top of the file :\n\n# Known limitations:\n#\n# - Only wiki pages are managed, no support for [[File:...]]\n#   attachments.\n\n> @@ -212,59 +212,230 @@ sub get_mw_pages {\n>  \tmy $user_defined;\n>  \tif (@tracked_pages) {\n>  \t\t$user_defined = 1;\n> -\t\t# The user provided a list of pages titles, but we\n> -\t\t# still need to query the API to get the page IDs.\n> -\n> -\t\tmy @some_pages = @tracked_pages;\n> -\t\twhile (@some_pages) {\n> -\t\t\tmy $last = 50;\n> -\t\t\tif ($#some_pages < $last) {\n> -\t\t\t\t$last = $#some_pages;\n> -\t\t\t}\n> -\t\t\tmy @slice = @some_pages[0..$last];\n> -\t\t\tget_mw_first_pages(\\@slice, \\%pages);\n> -\t\t\t@some_pages = @some_pages[51..$#some_pages];\n> -\t\t}\n> +\t\tget_mw_tracked_pages(\\%pages);\n>  \t}\n>  \tif (@tracked_categories) {\n>  \t\t$user_defined = 1;\n> -\t\tforeach my $category (@tracked_categories) {\n> -\t\t\tif (index($category, ':') < 0) {\n> -\t\t\t\t# Mediawiki requires the Category\n> -\t\t\t\t# prefix, but let's not force the user\n> -\t\t\t\t# to specify it.\n> -\t\t\t\t$category = \"Category:\" . $category;\n> -\t\t\t}\n> -\t\t\tmy $mw_pages = $mediawiki->list( {\n> -\t\t\t\taction => 'query',\n> -\t\t\t\tlist => 'categorymembers',\n> -\t\t\t\tcmtitle => $category,\n> -\t\t\t\tcmlimit => 'max' } )\n> -\t\t\t    || die $mediawiki->{error}->{code} . ': ' . $mediawiki->{error}->{details};\n> -\t\t\tforeach my $page (@{$mw_pages}) {\n> -\t\t\t\t$pages{$page->{title}} = $page;\n> -\t\t\t}\n> -\t\t}\n> +\t\tget_mw_tracked_categories(\\%pages);\n>  \t}\n>  \tif (!$user_defined) {\n> -\t\t# No user-provided list, get the list of pages from\n> -\t\t# the API.\n> -\t\tmy $mw_pages = $mediawiki->list({\n> -\t\t\taction => 'query',\n> -\t\t\tlist => 'allpages',\n> -\t\t\taplimit => 500,\n> -\t\t});\n> -\t\tif (!defined($mw_pages)) {\n> -\t\t\tprint STDERR \"fatal: could not get the list of wiki pages.\\n\";\n> -\t\t\tprint STDERR \"fatal: '$url' does not appear to be a mediawiki\\n\";\n> -\t\t\tprint STDERR \"fatal: make sure '$url/api.php' is a valid page.\\n\";\n> -\t\t\texit 1;\n> +\t\t get_mw_all_pages(\\%pages);\n> +\t}\n> +\treturn values(%pages);\n> +}\n\nThe refactoring is welcome, but it would have been better to make it in\na separate patch. The patch as you made it is long and hard to review,\nbecause it combines several new features, and refactoring.\n\n> +sub get_mw_tracked_pages {\n> +\tmy $pages = shift;\n> +\t# The user provided a list of pages titles, but we\n> +\t# still need to query the API to get the page IDs.\n> +\tmy @some_pages = @tracked_pages;\n> +\twhile (@some_pages) {\n> +\t\tmy $last = 50;\n> +\t\tif ($#some_pages < $last) {\n> +\t\t\t$last = $#some_pages;\n> +\t\t}\n> +\t\tmy @slice = @some_pages[0..$last];\n> +\t\tget_mw_first_pages(\\@slice, \\%{$pages});\n> +\t\t@some_pages = @some_pages[51..$#some_pages];\n> +\t}\n> +\n> +\t# Get pages of related media files.\n> +\tget_mw_linked_mediapages(\\@tracked_pages, \\%{$pages});\n[...]\n> +sub get_mw_linked_mediapages {\n\nThis is a nice feature, but I think it deserves to be configurable (if\nthe user explicitely specified one page, it actually seems strange to\nimport all the files it links to by default). Also, it should be\nmentionned in the commit message.\n\nShouldn't the function be named get_mw_linked_mediafiles instead? In\ngeneral, the wording \"media page\" is used in many places in the code,\nI prefer \"media file\" which is unambiguous.\n\n> +sub get_mw_medafile_for_mediapage_revision {\n\nmedafile -> mediafile ?\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"193168","messageId":"4FD2266B.3040706@ensimag.imag.fr","threadId":"30747","inReplyTo":"1339165376-20267-1-git-send-email-Pavel.Volek@ensimag.imag.fr","subject":"Re: [PATCHv1] git-remote-mediawiki: import \"File:\" attachments","fromName":"Simon.Cathebras","fromEmail":"simon.cathebras@ensimag.imag.fr","sentAt":"2012-06-08T16:20:59Z","receivedAt":"2012-06-08T16:20:59Z","isPatch":false,"sender":{"key":"simon.cathebras@ensimag.imag.fr","avatar":"https://avatars.githubusercontent.com/u/1764010?v=4"},"body":"\n\nOn 08/06/2012 16:22, Pavel Volek wrote:\n> From: Volek Pavel<me@pavelvolek.cz>\n>\n> The current version of the git-remote-mediawiki supports only import and export\n> of the pages, doesn't support import and export of file attachements which are\n> also exposed by MediaWiki API. This patch adds the functionality to import the\n> last versions of the files and all versions of description pages for these\n> files.\n>\n> Signed-off-by: Pavel Volek<Pavel.Volek@ensimag.imag.fr>\n> Signed-off-by: NGUYEN Kim Thuat<Kim-Thuat.Nguyen@ensimag.imag.fr>\n> Signed-off-by: ROUCHER IGLESIAS Javier<roucherj@ensimag.imag.fr>\n> Signed-off-by: Matthieu Moy<Matthieu.Moy@imag.fr>\n> ---\n\n>   contrib/mw-to-git/git-remote-mediawiki | 290 +++++++++++++++++++++++++++------\n>   1 file changed, 244 insertions(+), 46 deletions(-)\n\nI am wondering why are you showing the removal for a v1 patch ?\n\n>\n> diff --git a/contrib/mw-to-git/git-remote-mediawiki b/contrib/mw-to-git/git-remote-mediawiki\n> index c18bfa1..9f21217 100755\n> --- a/contrib/mw-to-git/git-remote-mediawiki\n> +++ b/contrib/mw-to-git/git-remote-mediawiki\n> @@ -212,59 +212,230 @@ sub get_mw_pages {\n>   \tmy $user_defined;\n>   \tif (@tracked_pages) {\n>   \t\t$user_defined = 1;\n> -\t\t# The user provided a list of pages titles, but we\n> -\t\t# still need to query the API to get the page IDs.\n> -\n> -\t\tmy @some_pages = @tracked_pages;\n> -\t\twhile (@some_pages) {\n> -\t\t\tmy $last = 50;\n> -\t\t\tif ($#some_pages<  $last) {\n> -\t\t\t\t$last = $#some_pages;\n> -\t\t\t}\n> -\t\t\tmy @slice = @some_pages[0..$last];\n> -\t\t\tget_mw_first_pages(\\@slice, \\%pages);\n> -\t\t\t@some_pages = @some_pages[51..$#some_pages];\n> -\t\t}\n> +\t\tget_mw_tracked_pages(\\%pages);\n>   \t}\n>   \tif (@tracked_categories) {\n>   \t\t$user_defined = 1;\n> -\t\tforeach my $category (@tracked_categories) {\n> -\t\t\tif (index($category, ':')<  0) {\n> -\t\t\t\t# Mediawiki requires the Category\n> -\t\t\t\t# prefix, but let's not force the user\n> -\t\t\t\t# to specify it.\n> -\t\t\t\t$category = \"Category:\" . $category;\n> -\t\t\t}\n> -\t\t\tmy $mw_pages = $mediawiki->list( {\n> -\t\t\t\taction =>  'query',\n> -\t\t\t\tlist =>  'categorymembers',\n> -\t\t\t\tcmtitle =>  $category,\n> -\t\t\t\tcmlimit =>  'max' } )\n> -\t\t\t    || die $mediawiki->{error}->{code} . ': ' . $mediawiki->{error}->{details};\n> -\t\t\tforeach my $page (@{$mw_pages}) {\n> -\t\t\t\t$pages{$page->{title}} = $page;\n> -\t\t\t}\n> -\t\t}\n> +\t\tget_mw_tracked_categories(\\%pages);\n>   \t}\n>   \tif (!$user_defined) {\n> -\t\t# No user-provided list, get the list of pages from\n> -\t\t# the API.\n> -\t\tmy $mw_pages = $mediawiki->list({\n> -\t\t\taction =>  'query',\n> -\t\t\tlist =>  'allpages',\n> -\t\t\taplimit =>  500,\n> -\t\t});\n> -\t\tif (!defined($mw_pages)) {\n> -\t\t\tprint STDERR \"fatal: could not get the list of wiki pages.\\n\";\n> -\t\t\tprint STDERR \"fatal: '$url' does not appear to be a mediawiki\\n\";\n> -\t\t\tprint STDERR \"fatal: make sure '$url/api.php' is a valid page.\\n\";\n> -\t\t\texit 1;\n> +\t\t get_mw_all_pages(\\%pages);\n> +\t}\n> +\treturn values(%pages);\n> +}\n> +\n> +sub get_mw_all_pages {\n> +\tmy $pages = shift;\n> +\t# No user-provided list, get the list of pages from the API.\n> +\tmy $mw_pages = $mediawiki->list({\n> +\t\taction =>  'query',\n> +\t\tlist =>  'allpages',\n> +\t\taplimit =>  500,\n> +\t});\n> +\tif (!defined($mw_pages)) {\n> +\t\tprint STDERR \"fatal: could not get the list of wiki pages.\\n\";\n> +\t\tprint STDERR \"fatal: '$url' does not appear to be a mediawiki\\n\";\n> +\t\tprint STDERR \"fatal: make sure '$url/api.php' is a valid page.\\n\";\n> +\t\texit 1;\n> +\t}\n> +\tforeach my $page (@{$mw_pages}) {\n> +\t\t$pages->{$page->{title}} = $page;\n> +\t}\n> +\n> +\t# Attach list of all pages for meadia files from the API,\n> +\t# they are in a different namespace, only one namespace\n> +\t# can be queried at the same moment\n> +\tmy $mw_mediapages = $mediawiki->list({\n> +\t\taction =>  'query',\n> +\t\tlist =>  'allpages',\n> +\t\tapnamespace =>  get_mw_namespace_id(\"File\"),\n> +\t\taplimit =>  500,\n> +\t});\n> +\tif (!defined($mw_mediapages)) {\n> +\t\tprint STDERR \"fatal: could not get the list of media file pages.\\n\";\n> +\t\tprint STDERR \"fatal: '$url' does not appear to be a mediawiki\\n\";\n> +\t\tprint STDERR \"fatal: make sure '$url/api.php' is a valid page.\\n\";\n> +\t\texit 1;\n> +\t}\n> +\tforeach my $page (@{$mw_mediapages}) {\n> +\t\t$pages->{$page->{title}} = $page;\n> +\t}\n> +}\n> +\n> +sub get_mw_tracked_pages {\n> +\tmy $pages = shift;\n> +\t# The user provided a list of pages titles, but we\n> +\t# still need to query the API to get the page IDs.\n> +\tmy @some_pages = @tracked_pages;\n> +\twhile (@some_pages) {\n> +\t\tmy $last = 50;\n> +\t\tif ($#some_pages<  $last) {\n> +\t\t\t$last = $#some_pages;\n> +\t\t}\n> +\t\tmy @slice = @some_pages[0..$last];\n> +\t\tget_mw_first_pages(\\@slice, \\%{$pages});\n> +\t\t@some_pages = @some_pages[51..$#some_pages];\n> +\t}\n> +\n> +\t# Get pages of related media files.\n> +\tget_mw_linked_mediapages(\\@tracked_pages, \\%{$pages});\n> +}\n> +\n> +sub get_mw_tracked_categories {\n> +\tmy $pages = shift;\n> +\tforeach my $category (@tracked_categories) {\n> +\t\tif (index($category, ':')<  0) {\n> +\t\t\t# Mediawiki requires the Category\n> +\t\t\t# prefix, but let's not force the user\n> +\t\t\t# to specify it.\n> +\t\t\t$category = \"Category:\" . $category;\n>   \t\t}\n> +\t\tmy $mw_pages = $mediawiki->list( {\n> +\t\t\taction =>  'query',\n> +\t\t\tlist =>  'categorymembers',\n> +\t\t\tcmtitle =>  $category,\n> +\t\t\tcmlimit =>  'max' } )\n> +\t\t\t|| die $mediawiki->{error}->{code} . ': '\n> +\t\t\t\t. $mediawiki->{error}->{details};\n>   \t\tforeach my $page (@{$mw_pages}) {\n> -\t\t\t$pages{$page->{title}} = $page;\n> +\t\t\t$pages->{$page->{title}} = $page;\n> +\t\t}\n> +\n> +\t\tmy @titles = map $_->{title}, @{$mw_pages};\n> +\t\t# Get pages of related media files.\n> +\t\tget_mw_linked_mediapages(\\@titles, \\%{$pages});\n> +\t}\n> +}\n> +\n> +sub get_mw_linked_mediapages {\n> +\tmy $titles = shift;\n> +\tmy @titles = @{$titles};\n> +\tmy $pages = shift;\n> +\n> +\t# pattern 'page1|page2|...' required by the API\n> +\tmy $mw_titles = join('|', @titles);\n> +\n> +\t# Media files could be included or linked from\n> +\t# a page, get all related\n> +\tmy $query = {\n> +\t\taction =>  'query',\n> +\t\tprop =>  'links|images',\n> +\t\ttitles =>  $mw_titles,\n> +\t\tplnamespace =>  get_mw_namespace_id(\"File\"),\n> +\t\tpllimit =>  500,\n> +\t};\n\nWhy a comma after 500 ?\n\n> +\tmy $result = $mediawiki->api($query);\n\n\nWhat happened if the titles in the query contains special character \nwhich are not allowed by mediawiki for filename like { or [.\nMaybe you should build a test for it and if it doesn't work try out the \nfunctions called:\n     mediawiki_clean/smudge_filename\nin the file git-remote-mediawiki\n\n\n> +\n> +\twhile (my ($id, $page) = each(%{$result->{query}->{pages}})) {\n> +\t\tmy @titles;\n> +\t\tif (defined($page->{links})) {\n> +\t\t\tmy @link_titles = map $_->{title}, @{$page->{links}};\n> +\t\t\tpush(@titles, @link_titles);\n> +\t\t}\n> +\t\tif (defined($page->{images})) {\n> +\t\t\tmy @image_titles = map $_->{title}, @{$page->{images}};\n> +\t\t\tpush(@titles, @image_titles);\n> +\t\t}\n> +\t\tif (@titles) {\n> +\t\t\tget_mw_first_pages(\\@titles, \\%{$pages});\n>   \t\t}\n>   \t}\n> -\treturn values(%pages);\n> +}\n> +\n> +sub get_mw_medafile_for_mediapage_revision {\n> +\t# Name of the file on Wiki, with the prefix.\n> +\tmy $mw_filename = shift;\n> +\tmy $timestamp = shift;\n> +\tmy %mediafile;\n> +\n> +\t# Search if on MediaWiki exists a media file with given\n> +\t# timestamp and in that case download the file.\n> +\tmy $query = {\n> +\t\taction =>  'query',\n> +\t\tprop =>  'imageinfo',\n> +\t\ttitles =>  $mw_filename,\n> +\t\tiistart =>  $timestamp,\n> +\t\tiiend =>  $timestamp,\n> +\t\tiiprop =>  'timestamp|archivename',\n> +\t\tiilimit =>  1,\n> +\t};\n\nWhy a comma after iilimit ? (end of list of parameter here I think...)\n\n> +\tmy $result = $mediawiki->api($query);\n> +\n> +\tmy ($fileid, $file) = each ( %{$result->{query}->{pages}} );\n> +\tif (defined($file->{imageinfo})) {\n> +\t\tmy $fileinfo = pop(@{$file->{imageinfo}});\n> +\t\tif (defined($fileinfo->{archivename})) {\n> +\t\t\treturn; # now we are not able to download files from archive\n> +\t\t}\n> +\n> +\t\tmy $filename; # real filename without prefix\n> +\t\tif (index($mw_filename, 'File:') == 0) {\n> +\t\t\t$filename = substr $mw_filename, 5;\n> +\t\t} else {\n> +\t\t\t$filename = substr $mw_filename, 6;\n> +\t\t}\n> +\n> +\t\t$mediafile{title} = $filename;\n> +\t\t$mediafile{content} = download_mw_mediafile($mw_filename);\n> +\t}\n> +\treturn %mediafile;\n> +}\n> +\n> +# Returns MediaWiki id for a canonical namespace name.\n> +# Ex.: \"File\", \"Project\".\n> +# Looks for the namespace id in the local configuration\n> +# variables, if it is not found asks MW API.\n> +sub get_mw_namespace_id {\n> +\tmw_connect_maybe();\n> +\n> +\tmy $name = shift;\n> +\n> +\t# Look at configuration file, if the record\n> +\t# for that namespace is already stored.\n> +\tmy @tracked_namespaces = split(/[ \\n]/, run_git(\"config --get-all remote.\". $remotename .\".namespaces\"));\n\nBroken indentation/line too long ?\n\n> +\n> +\t# NS not found =>  get namespace id from MW and store it in\n> +\t# configuration file.\n> +\tmy $query = {\n> +\t\taction =>  'query',\n> +\t\tmeta =>  'siteinfo',\n> +\t\tsiprop =>  'namespaces',\n> +\t};\n\nSame here concerning comma.\n\n> +\tmy $result = $mediawiki->api($query);\n> +\n> +\twhile (my ($id, $ns) = each(%{$result->{query}->{namespaces}})) {\n> +\t\tif (defined($ns->{canonical})&&  ($ns->{canonical} eq $name)) {\n> +\t\t\trun_git(\"config --add remote.\". $remotename .\".namespaces \". $name .\"=\". $ns->{id});\n> +\t\t\treturn $ns->{id};\n> +\t\t}\n> +\t}\n> +\tdie \"Namespace $name was not found on MediaWiki.\";\n> +}\n> +\n> +sub download_mw_mediafile {\n> +\tmy $filename = shift;\n> +\n> +\t$mediawiki->{config}->{files_url} = $url;\n> +\n> +\tmy $file = $mediawiki->download( { title =>  $filename } );\n\nJust wondering: What happened if $filename contains some forbidden \ncharacter on wiki's filename such as '{' or '|' ?\nI am worrying about it because i've got some similar issues in my own \nwork on tests for git-remote-mediawiki.\n\nHope I helped :).\n\nSimon\n\n-- \nCATHEBRAS Simon\n\n2A-ENSIMAG\n\nFilière Ingéniérie des Systèmes d'Information\nMembre Bug-Buster\n"},{"id":"193171","messageId":"20120608190305.Horde.szbWGnwdC4BP0jBJMfN12lA@webmail.minatec.grenoble-inp.fr","threadId":"30747","inReplyTo":"4FD2266B.3040706@ensimag.imag.fr","subject":"Re: [PATCHv1] git-remote-mediawiki: import \"File:\" attachments","fromName":"","fromEmail":"konglu@minatec.inpg.fr","sentAt":"2012-06-08T17:03:05Z","receivedAt":"2012-06-08T17:03:05Z","isPatch":false,"sender":{"key":"konglu@minatec.inpg.fr","avatar":null},"body":"\n\"Simon.Cathebras\" <Simon.Cathebras@ensimag.imag.fr> a écrit :\n\n> On 08/06/2012 16:22, Pavel Volek wrote:\n>> From: Volek Pavel<me@pavelvolek.cz>\n>>\n>> The current version of the git-remote-mediawiki supports only  \n>> import and export\n>> of the pages, doesn't support import and export of file  \n>> attachements which are\n>> also exposed by MediaWiki API. This patch adds the functionality to  \n>> import the\n>> last versions of the files and all versions of description pages for these\n>> files.\n>>\n>> Signed-off-by: Pavel Volek<Pavel.Volek@ensimag.imag.fr>\n>> Signed-off-by: NGUYEN Kim Thuat<Kim-Thuat.Nguyen@ensimag.imag.fr>\n>> Signed-off-by: ROUCHER IGLESIAS Javier<roucherj@ensimag.imag.fr>\n>> Signed-off-by: Matthieu Moy<Matthieu.Moy@imag.fr>\n>> ---\n>\n>>  contrib/mw-to-git/git-remote-mediawiki | 290  \n>> +++++++++++++++++++++++++++------\n>>  1 file changed, 244 insertions(+), 46 deletions(-)\n>\n> I am wondering why are you showing the removal for a v1 patch ?\n\nWhy not ? The file already exists on branch master and they are\nworking on it. Anyway, the patch applies correctly on master.\nBTW, are you implying that only v2+ patch could have deletions ?\n(a patch is not meant to be applied on the previous version).\n\nLucien Kong\n"},{"id":"193200","messageId":"4FD2899D.4080407@ensimag.imag.fr","threadId":"30747","inReplyTo":"20120608190305.Horde.szbWGnwdC4BP0jBJMfN12lA@webmail.minatec.grenoble-inp.fr","subject":"Re: [PATCHv1] git-remote-mediawiki: import \"File:\" attachments","fromName":"Simon.Cathebras","fromEmail":"simon.cathebras@ensimag.imag.fr","sentAt":"2012-06-08T23:24:13Z","receivedAt":"2012-06-08T23:24:13Z","isPatch":false,"sender":{"key":"simon.cathebras@ensimag.imag.fr","avatar":"https://avatars.githubusercontent.com/u/1764010?v=4"},"body":"\n\nOn 08/06/2012 19:03, konglu@minatec.inpg.fr wrote:\n>\n> \"Simon.Cathebras\" <Simon.Cathebras@ensimag.imag.fr> a écrit :\n>\n>> On 08/06/2012 16:22, Pavel Volek wrote:\n>>> From: Volek Pavel<me@pavelvolek.cz>\n>>>\n>>> The current version of the git-remote-mediawiki supports only import \n>>> and export\n>>> of the pages, doesn't support import and export of file attachements \n>>> which are\n>>> also exposed by MediaWiki API. This patch adds the functionality to \n>>> import the\n>>> last versions of the files and all versions of description pages for \n>>> these\n>>> files.\n>>>\n>>> Signed-off-by: Pavel Volek<Pavel.Volek@ensimag.imag.fr>\n>>> Signed-off-by: NGUYEN Kim Thuat<Kim-Thuat.Nguyen@ensimag.imag.fr>\n>>> Signed-off-by: ROUCHER IGLESIAS Javier<roucherj@ensimag.imag.fr>\n>>> Signed-off-by: Matthieu Moy<Matthieu.Moy@imag.fr>\n>>> ---\n>>\n>>>  contrib/mw-to-git/git-remote-mediawiki | 290 \n>>> +++++++++++++++++++++++++++------\n>>>  1 file changed, 244 insertions(+), 46 deletions(-)\n>>\n>> I am wondering why are you showing the removal for a v1 patch ?\n>\n> Why not ? The file already exists on branch master and they are\n> working on it.\n\n\nMakes sense... I didn't notice the deletions were on master, my bad.\n\n\n> Anyway, the patch applies correctly on master.\n> BTW, are you implying that only v2+ patch could have deletions ?\n> (a patch is not meant to be applied on the previous version).\n\n\nActually, I was just saying that showing corrections on a patche's code \nduring the development of this one, isn't really necessary.\nBut if it is concerning a modification of a code in a previous version, \nI agree, it is absolutly useful ;).\n\nSimon\n\n-- \nCATHEBRAS Simon\n\n2A-ENSIMAG\n\nFilière Ingéniérie des Systèmes d'Information\nMembre Bug-Buster\n"}]}