{"thread":{"id":"30240","subject":"[PATCH] Explain how svn-fe parses filenames in SVN dumps","startedAt":"2012-04-14T17:03:09Z","lastAt":"2012-04-15T05:08:51Z","messageCount":7,"participants":["Andrew Sayers","Jonathan Nieder"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"189263","messageId":"4F89ADCD.6000109@pileofstuff.org","threadId":"30240","inReplyTo":null,"subject":"[PATCH] Explain how svn-fe parses filenames in SVN dumps","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-04-14T17:03:09Z","receivedAt":"2012-04-14T17:03:09Z","isPatch":true,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"The documentation for the SVN dumpfile format says that filenames \"may be\ninterpreted as binary data in any encoding by client tools\", but users might be\nsurprised that svn-fe's handling differs from svn's.\n\nBefore version 1.2.0, `svn add` supported files containing characters in the\nrange 0x01-0x1F, and Subversion still supports existing files that contain\nthose characters.  The newline character is explicitly discussed so that users\nwith ancient repositories understand why they can't be supported by tools that\nread the SVN dump format.\n\nThe documentation for the SVN dumpfile format describes records as containing\n\"a group of RFC822-style header lines\", and its full text can be read as\nimplying newline characters are reserved for use by the format.  This reading\nis slightly charitable, but it avoids the need to discuss the format's design\nissues in a context where few readers will be interested.\n\nSigned-off-by: Andrew Sayers <andrew-git@pileofstuff.org>\n---\n contrib/svn-fe/svn-fe.txt |    8 ++++++++\n 1 files changed, 8 insertions(+), 0 deletions(-)\n\ndiff --git a/contrib/svn-fe/svn-fe.txt b/contrib/svn-fe/svn-fe.txt\nindex 1128ab2..c079abe 100644\n--- a/contrib/svn-fe/svn-fe.txt\n+++ b/contrib/svn-fe/svn-fe.txt\n@@ -59,6 +59,14 @@ to put each project in its own repository and to separate the history\n of each branch.  The 'git filter-branch --subdirectory-filter' command\n may be useful for this purpose.\n \n+Filenames are interpreted by svn-fe as binary data, and may contain\n+any character except NUL (0x00) and newline (0x0A).  The NUL\n+character is not valid in git paths, and the newline character is\n+reserved for use by the (line-based) Subversion dumpfile format.\n+This differs from Subversion, which requires filenames to contain\n+only legal XML characters and disallows tabs characters, carriage\n+returns and newlines.\n+\n BUGS\n ----\n Empty directories and unknown properties are silently discarded.\n-- \n1.7.1\n"},{"id":"189264","messageId":"20120414171431.GA4161@burratino","threadId":"30240","inReplyTo":"4F89ADCD.6000109@pileofstuff.org","subject":"Re: [PATCH] Explain how svn-fe parses filenames in SVN dumps","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-14T17:14:31Z","receivedAt":"2012-04-14T17:14:31Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nAndrew Sayers wrote:\n\n> Before version 1.2.0, `svn add` supported files containing characters in the\n> range 0x01-0x1F, and Subversion still supports existing files that contain\n> those characters.\n\nBecause of the above,\n\n[...]\n> +++ b/contrib/svn-fe/svn-fe.txt\n> @@ -59,6 +59,14 @@ to put each project in its own repository and to separate the history\n>  of each branch.  The 'git filter-branch --subdirectory-filter' command\n>  may be useful for this purpose.\n>  \n> +Filenames are interpreted by svn-fe as binary data, and may contain\n> +any character except NUL (0x00) and newline (0x0A).  The NUL\n> +character is not valid in git paths, and the newline character is\n> +reserved for use by the (line-based) Subversion dumpfile format.\n> +This differs from Subversion, which requires filenames to contain\n> +only legal XML characters and disallows tabs characters, carriage\n> +returns and newlines.\n> +\n>  BUGS\n\nthis description and the location of this description seem quite\nmisleading.  Isn't what the reader needs to know something like the\nfollowing?\n\n\tBUGS\n\t----\n\tDue to limitations in the Subversion dumpfile format, svn-fe\n\tdoes not support filenames with newlines.  Since version 1.2.0,\n\t\"svn add\" forbids adding such filenames but some historical\n\trepositories contain them.  An import can appear to succeed and\n\tproduce incorrect results when such pathological filenames are\n\tpresent.\n\nThanks,\nJonathan\n"},{"id":"189267","messageId":"4F89B5C5.3030606@pileofstuff.org","threadId":"30240","inReplyTo":"20120414171431.GA4161@burratino","subject":"Re: [PATCH] Explain how svn-fe parses filenames in SVN dumps","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-04-14T17:37:09Z","receivedAt":"2012-04-14T17:37:09Z","isPatch":true,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"On 14/04/12 18:14, Jonathan Nieder wrote:\n> this description and the location of this description seem quite\n> misleading.  Isn't what the reader needs to know something like the\n> following?\n> \n> \tBUGS\n> \t----\n> \tDue to limitations in the Subversion dumpfile format, svn-fe\n> \tdoes not support filenames with newlines.  Since version 1.2.0,\n> \t\"svn add\" forbids adding such filenames but some historical\n> \trepositories contain them.  An import can appear to succeed and\n> \tproduce incorrect results when such pathological filenames are\n> \tpresent.\n> \n> Thanks,\n> Jonathan\n> \n\nI went back and forth a bit while writing the text.  Newlines are only\none special case, albeit an important one that I hadn't expressed\nclearly enough.  For example, the handling of NUL characters is worse\nthan newlines (a quick test suggests svn-fe terminates parsing\naltogether if it sees one), but SVN has never allowed the creation of\nfiles with NULs so arguably it's even less important.\n\nI'm warming again to the idea of explicitly mentioning that newlines\ncause breakage, but it feels like the wider story is worth telling too.\n If this were a man page I'd consider adding a section or something, but\nI'm not sure what level of verbosity you're looking for in this file.\n\n\t- Andrew\n"},{"id":"189269","messageId":"20120414181357.GA4560@burratino","threadId":"30240","inReplyTo":"4F89B5C5.3030606@pileofstuff.org","subject":"Re: [PATCH] Explain how svn-fe parses filenames in SVN dumps","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-14T18:13:57Z","receivedAt":"2012-04-14T18:13:57Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Andrew Sayers wrote:\n\n> If this were a man page\n\nThis is a man page.  \"git log contrib/svn-fe/svn-fe.txt\" has details.\n"},{"id":"189270","messageId":"20120414181823.GB4560@burratino","threadId":"30240","inReplyTo":"4F89B5C5.3030606@pileofstuff.org","subject":"Re: [PATCH] Explain how svn-fe parses filenames in SVN dumps","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-14T18:18:23Z","receivedAt":"2012-04-14T18:18:23Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Andrew Sayers wrote:\n\n>                  For example, the handling of NUL characters is worse\n> than newlines\n\nFilenames are C strings.  Filenames with NUL bytes simply do not\nexist, at least as long as one is using the ANSI C or POSIX interfaces\nfor file access.\n\nSo as far as I can tell the story really is only about newlines.  (If\nsvn-fe tried to push history back to Subversion, life would presumably\nbe more complicated.)\n\nHoping that clarifies a little,\nJonathan\n"},{"id":"189294","messageId":"4F89F27B.5050706@pileofstuff.org","threadId":"30240","inReplyTo":"20120414181823.GB4560@burratino","subject":"Re: [PATCH] Explain how svn-fe parses filenames in SVN dumps","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-04-14T21:56:11Z","receivedAt":"2012-04-14T21:56:11Z","isPatch":true,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"On 14/04/12 19:18, Jonathan Nieder wrote:\n> So as far as I can tell the story really is only about newlines.  (If\n> svn-fe tried to push history back to Subversion, life would presumably\n> be more complicated.)\n\nI think I've been trying to balance two incompatible use cases -\nordinary users that just need a heads-up about a bug that might bite\nthem, and authors of related tools (i.e. me) that need technical\ninformation about svn-fe.\n\nI think your text serves ordinary users better, so I'll investigate\npushing history back to Subversion tomorrow and think how to write\nsomething more technical.  I'm afraid my patch etiquette fails me here\nthough - is it better for me to roll your text into a new patch, or let\nyou do that and submit my text separately?\n\n\t- Andrew\n"},{"id":"189316","messageId":"20120415050851.GA2544@burratino","threadId":"30240","inReplyTo":"4F89F27B.5050706@pileofstuff.org","subject":"Re: [PATCH] Explain how svn-fe parses filenames in SVN dumps","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-15T05:08:51Z","receivedAt":"2012-04-15T05:08:51Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Andrew Sayers wrote:\n\n>                            I'm afraid my patch etiquette fails me here\n> though - is it better for me to roll your text into a new patch, or let\n> you do that and submit my text separately?\n\nIf the example I sent looks good, please feel free to morph it into a\nnew patch.  Thanks again for your work, and glad I could help in some\nsmall way.\n\nCiao,\nJonathan\n"}]}