{"thread":{"id":"30250","subject":"[PATCHv2] Add details about svn-fe's dumpfile parsing","startedAt":"2012-04-15T16:10:46Z","lastAt":"2012-07-23T01:37:39Z","messageCount":7,"participants":["Andrew Sayers","Junio C Hamano","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"189340","messageId":"4F8AF306.8070804@pileofstuff.org","threadId":"30250","inReplyTo":null,"subject":"[PATCHv2] Add details about svn-fe's dumpfile parsing","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-04-15T16:10:46Z","receivedAt":"2012-04-15T16:10:46Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"The documentation for the SVN dumpfile format says that \"property key/value\npairs may be interpreted as binary data in any encoding by client tools\".\nDocumenting svn-fe's interpretation helps authors of related tools, while\nexplaining limitations helps ordinary users import their SVN repositories.\n\nThe \"INPUT FORMAT\" section is aimed at authors of tools that interact with\nsvn-fe, so it particularly addresses assumptions that authors might make after\ndealing with svn itself.\n\nThe \"BUGS\" section is aimed at ordinary users, so it only explains what readers\nneed to know when importing a repository.  In particular, users don't need to\nknow that other characters in the range 0x01-0x1F are imported correctly, even\nthough they were all disabled in Subversion 1.2.0.  The text in this section is\nbased largely on an example sent by Jonathan Nieder, with minor changes to suit\nthe surrounding style.\n\nSigned-off-by: Andrew Sayers <andrew-git@pileofstuff.org>\n---\n contrib/svn-fe/svn-fe.txt |   13 +++++++++++++\n 1 files changed, 13 insertions(+), 0 deletions(-)\n\ndiff --git a/contrib/svn-fe/svn-fe.txt b/contrib/svn-fe/svn-fe.txt\nindex 1128ab2..3872b9d 100644\n--- a/contrib/svn-fe/svn-fe.txt\n+++ b/contrib/svn-fe/svn-fe.txt\n@@ -32,6 +32,13 @@ Subversion's repository dump format is documented in full in\n Files in this format can be generated using the 'svnadmin dump' or\n 'svk admin dump' command.\n \n+Unlike Subversion, 'svn-fe' interprets property key/value pairs as\n+null-terminated binary strings.  This means it will accept content\n+that Subversion normally wouldn't produce (such as filenames\n+containing tab characters) or would refuse to parse (such as usernames\n+containing Latin-1 characters).  However, like Subversion it will\n+handle newlines incorrectly in filenames (see BUGS below).\n+\n OUTPUT FORMAT\n -------------\n The fast-import format is documented by the git-fast-import(1)\n@@ -65,6 +72,12 @@ Empty directories and unknown properties are silently discarded.\n \n The exit status does not reflect whether an error was detected.\n \n+Due to limitations in the Subversion dumpfile format, 'svn-fe' does\n+not support filenames with newlines.  'svn add' has forbidden such\n+filenames since version 1.2.0, but some historical repositories still\n+contain them.  An import can appear to succeed and produce incorrect\n+results when such pathological filenames are present.\n+\n SEE ALSO\n --------\n git-svn(1), svn2git(1), svk(1), git-filter-branch(1), git-fast-import(1),\n-- \n1.7.1\n"},{"id":"189451","messageId":"7vipgztpaf.fsf@alter.siamese.dyndns.org","threadId":"30250","inReplyTo":"4F8AF306.8070804@pileofstuff.org","subject":"Re: [PATCHv2] Add details about svn-fe's dumpfile parsing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-04-16T20:06:16Z","receivedAt":"2012-04-16T20:06:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andrew Sayers <andrew-git@pileofstuff.org> writes:\n\n> The documentation for the SVN dumpfile format says that \"property key/value\n> pairs may be interpreted as binary data in any encoding by client tools\".\n> Documenting svn-fe's interpretation helps authors of related tools, while\n> explaining limitations helps ordinary users import their SVN repositories.\n>\n> The \"INPUT FORMAT\" section is aimed at authors of tools that interact with\n> svn-fe, so it particularly addresses assumptions that authors might make after\n> dealing with svn itself.\n>\n> The \"BUGS\" section is aimed at ordinary users, so it only explains what readers\n> need to know when importing a repository.  In particular, users don't need to\n> know that other characters in the range 0x01-0x1F are imported correctly, even\n> though they were all disabled in Subversion 1.2.0.  The text in this section is\n> based largely on an example sent by Jonathan Nieder, with minor changes to suit\n> the surrounding style.\n>\n> Signed-off-by: Andrew Sayers <andrew-git@pileofstuff.org>\n> ---\n\nOK, so is this ready for 'master' already?\n\n>  contrib/svn-fe/svn-fe.txt |   13 +++++++++++++\n>  1 files changed, 13 insertions(+), 0 deletions(-)\n>\n> diff --git a/contrib/svn-fe/svn-fe.txt b/contrib/svn-fe/svn-fe.txt\n> index 1128ab2..3872b9d 100644\n> --- a/contrib/svn-fe/svn-fe.txt\n> +++ b/contrib/svn-fe/svn-fe.txt\n> @@ -32,6 +32,13 @@ Subversion's repository dump format is documented in full in\n>  Files in this format can be generated using the 'svnadmin dump' or\n>  'svk admin dump' command.\n>  \n> +Unlike Subversion, 'svn-fe' interprets property key/value pairs as\n> +null-terminated binary strings.  This means it will accept content\n> +that Subversion normally wouldn't produce (such as filenames\n> +containing tab characters) or would refuse to parse (such as usernames\n> +containing Latin-1 characters).  However, like Subversion it will\n> +handle newlines incorrectly in filenames (see BUGS below).\n> +\n\nDo the first two sentences in the above paragraph claim that it a bug that\n'svn-fe' does not mimick what Subversion does?  I am not sure what lessons\nthe authors of tools, whose output is meant to feed svn-fe, are expected\nto learn here.  For example, is the purpose of the above paragraph to make\ntool authors realize that \"NUL terminates key and value, so I have to\nrefrain from using a key or a value that contains a NUL byte?\" [*1*]  Even\nin that case, it is unclear to me what I (as an author of such a tool that\nreads data from somewhere and format it to plesae svn-fe) could do with\nthat knowledge.\n\n[Footnote]\n\n*1* By the way, NULL is a pointer that does not point anywhere.  The name\nof a byte whose value is 0x00 is NUL.\n"},{"id":"189466","messageId":"4F8C909B.7010507@pileofstuff.org","threadId":"30250","inReplyTo":"7vipgztpaf.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCHv2] Add details about svn-fe's dumpfile parsing","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-04-16T21:35:23Z","receivedAt":"2012-04-16T21:35:23Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"On 16/04/12 21:06, Junio C Hamano wrote:\n> Andrew Sayers <andrew-git@pileofstuff.org> writes:\n>>  \n>> +Unlike Subversion, 'svn-fe' interprets property key/value pairs as\n>> +null-terminated binary strings.  This means it will accept content\n>> +that Subversion normally wouldn't produce (such as filenames\n>> +containing tab characters) or would refuse to parse (such as usernames\n>> +containing Latin-1 characters).  However, like Subversion it will\n>> +handle newlines incorrectly in filenames (see BUGS below).\n>> +\n> \n> Do the first two sentences in the above paragraph claim that it a bug that\n> 'svn-fe' does not mimick what Subversion does?  I am not sure what lessons\n> the authors of tools, whose output is meant to feed svn-fe, are expected\n> to learn here.  For example, is the purpose of the above paragraph to make\n> tool authors realize that \"NUL terminates key and value, so I have to\n> refrain from using a key or a value that contains a NUL byte?\" [*1*]  Even\n> in that case, it is unclear to me what I (as an author of such a tool that\n> reads data from somewhere and format it to plesae svn-fe) could do with\n> that knowledge.\n> \n> [Footnote]\n> \n> *1* By the way, NULL is a pointer that does not point anywhere.  The name\n> of a byte whose value is 0x00 is NUL.\n\nThe dumpfile documentation says that \"... property key/value pairs may\nbe interpreted as binary data in any encoding by client tools\"[1], but\nSVN itself interprets the data as UTF-8, so I was surprised to see\nsvn-fe hadn't aped that behaviour.  You could argue this is a bug if you\nwant to call `svnadmin` the reference implementation.  You could even\nargue that treating NUL characters specially is a bug if you want to\ncall the documentation the official standard (albeit a bug shared by\n`svnadmin`).  Personally I don't have a problem with either decision, so\nI've just noted some unobvious behaviour.\n\nLessons to learn will depend on the author, but here are some I took:\n\n1. UTF-8 is the most common encoding, but not the only one.  If your\ntool only allows UTF-8 input and only produces only UTF-8 output then\nyou are the limiting factor in your toolchain.  In my case I think I'll\njust live with that, but I would like to have known before I started.\n\n2. `svn` itself isn't universally considered the reference\nimplementation, only a popular one.  When deciding the correct behaviour\n(or the range of possible behaviours), it's not enough just to look at\nwhat `svn` does.\n\n3. `svn-fe` doesn't slavishly follow either the documentation or `svn`.\n Beyond a certain point you have to actually check assumptions against\nsvn-fe's behaviour (then document what you find so the next guy doesn't\nhave to ;).  Again, I think this is pragmatic but I would like to have\nknown earlier.\n\nI've tried not to labour the above points, because it would be easy to\noverstate them and because other authors will take different lessons.\nIt's certainly possible that some author would come to svn-fe wanting to\ndo something crazy like encode newlines as NULs and come away realising\nC doesn't like that.  A more concrete example is that a GSoC student who\nturns up next year to work on writing commits back to SVN will need to\nknow svn-fe chose not to care about UTF-8 and that there's a nest of\nedge cases waiting for them.\n\n\t- Andrew\n\n[1]http://svn.apache.org/repos/asf/subversion/trunk/notes/dump-load-format.txt\n"},{"id":"189467","messageId":"20120416213910.GP12613@burratino","threadId":"30250","inReplyTo":"4F8C909B.7010507@pileofstuff.org","subject":"Re: [PATCHv2] Add details about svn-fe's dumpfile parsing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-16T21:39:10Z","receivedAt":"2012-04-16T21:39:10Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Andrew Sayers wrote:\n\n> The dumpfile documentation says that \"... property key/value pairs may\n> be interpreted as binary data in any encoding by client tools\"[1], but\n> SVN itself interprets the data as UTF-8\n\nYes, I suspect most of the changes you proposed for the INPUT FORMAT\nsection would actually be better as changes for the\ndump-load-format.txt document.  I imagine that folks on the dev@ list\nmight be able to clarify a few details (e.g., what one is expected to\ndo with historical repositories with non-UTF-8 property data), too.\nWhat do you think?\n\nThe patch for svn-fe(1) already looks pretty good.  I was planning on\napplying it after finding a moment to clarify the patch description.\n\nThanks again,\nJonathan\n"},{"id":"189481","messageId":"4F8C9A13.80906@pileofstuff.org","threadId":"30250","inReplyTo":"20120416213910.GP12613@burratino","subject":"Re: [PATCHv2] Add details about svn-fe's dumpfile parsing","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-04-16T22:15:47Z","receivedAt":"2012-04-16T22:15:47Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"On 16/04/12 22:39, Jonathan Nieder wrote:\n> Andrew Sayers wrote:\n> \n>> The dumpfile documentation says that \"... property key/value pairs may\n>> be interpreted as binary data in any encoding by client tools\"[1], but\n>> SVN itself interprets the data as UTF-8\n> \n> Yes, I suspect most of the changes you proposed for the INPUT FORMAT\n> section would actually be better as changes for the\n> dump-load-format.txt document.  I imagine that folks on the dev@ list\n> might be able to clarify a few details (e.g., what one is expected to\n> do with historical repositories with non-UTF-8 property data), too.\n> What do you think?\n\nHmm, I'd personally be more interested in going to the SVN folks with a\nmore general question.  The SVN Book[1] says \"pathnames can contain only\nlegal XML (1.0) characters, and properties are further limited to ASCII\ncharacters. Subversion also prohibits TAB, CR, and LF characters in path\nnames\".  Code documentation[2] gives a lot of complex rules that don't\nbear much resemblance to the behaviour I've seen so far (albeit only\nlightly tested in SVN 1.6).  The dumpfile docs[3] pretty much declare a\nfree-for-all, and I've yet to see historical documentation properly\nwritten up anywhere.\n\nI guess my question would be something like \"what should a client\nreading or writing SVN dumps do to stay as compatible as possible?\", but\nI feel like I've got a collection of bits that haven't quite coalesced\nwell enough yet to really drive the conversation.\n\nAs a web developer, the SBL work I've been doing is starting to remind\nme of the jump from HTML4 (\"here's what clients should do.  Of course\nit's not what they actually do...\") to HTML5 (\"here's what clients\nactually do.  No we're not allowed to just shoot those people\").  Like\nHTML5, I figure I've got to take the argument to the official body some\nday, but I'd rather have something vaguely mature first.\n\nMy instinct is to put this on the TODO list for after I've finished\nwriting tests, but I'm open to suggestions.\n\n\t- Andrew\n\n[1]http://svnbook.red-bean.com/en/1.7/svn.tour.importing.html#svn.tour.importing.naming\n[2]http://subversion.apache.org/docs/api/latest/group__svn__fs__directories.html#details\n[3]http://svn.apache.org/repos/asf/subversion/trunk/notes/dump-load-format.txt\n"},{"id":"189482","messageId":"20120416222700.GT12613@burratino","threadId":"30250","inReplyTo":"4F8C9A13.80906@pileofstuff.org","subject":"Re: [PATCHv2] Add details about svn-fe's dumpfile parsing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-16T22:27:00Z","receivedAt":"2012-04-16T22:27:00Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Andrew Sayers wrote:\n\n>                                                                   Like\n> HTML5, I figure I've got to take the argument to the official body some\n> day, but I'd rather have something vaguely mature first.\n\nPerhaps unlike the XHTML days at the W3C ;-), in this instance the\npeople responsible for the dump-load-format.txt document are really\nnice and helpful and generally right-minded people.\n\nIf the mailing list is intimidating (or even if not), I'd recommend\nvisiting the IRC channel #svn-dev on freenode to say hello and get\nadvice.\n\nJonathan\n"},{"id":"195470","messageId":"20120723013739.GC3390@burratino","threadId":"30250","inReplyTo":"20120416213910.GP12613@burratino","subject":"Re: [PATCHv2] Add details about svn-fe's dumpfile parsing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-07-23T01:37:39Z","receivedAt":"2012-07-23T01:37:39Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi Andrew,\n\nIn April, you wrote a nice patch[1] for the svn-fe(1) manpage\nclarifying some of its limitations.  I was hoping to offer some\npatches to squash in to polish some of its more confusing edges and\nthen apply it, but time for polishing ended up being scarce.\n\nCan you remind me of the current status of the patch --- e.g., is the\nversion at [1] the latest version?  Do you think it's ready as-is or\nwould you have suggestions for a person wanting to get it ready for\ninclusion?\n\nThanks,\nJonathan\n\n[1] http://thread.gmane.org/gmane.comp.version-control.git/195570\n"}]}