{"thread":{"id":"21749","subject":"jgit problems for file paths with non-ASCII characters","startedAt":"2009-11-25T13:47:25Z","lastAt":"2009-11-26T20:03:35Z","messageCount":10,"participants":["Marc Strapetz","Robin Rosenberg","Shawn O. Pearce","Thomas Singer","Johannes Schindelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"128311","messageId":"4B0D356D.1080709@syntevo.com","threadId":"21749","inReplyTo":null,"subject":"jgit problems for file paths with non-ASCII characters","fromName":"Marc Strapetz","fromEmail":"marc.strapetz@syntevo.com","sentAt":"2009-11-25T13:47:25Z","receivedAt":"2009-11-25T13:47:25Z","isPatch":false,"sender":{"key":"marc.strapetz@syntevo.com","avatar":"https://avatars.githubusercontent.com/u/3380730?v=4"},"body":"I have noticed that jgit converts file paths to UTF-8 when querying the\nrepository. Especially,\norg.eclipse.jgit.treewalk.filter.PathFilter#PathFilter performs this\nconversion:\n\n  private PathFilter(final String s) {\n    pathStr = s;\n    pathRaw = Constants.encode(pathStr);\n  }\n\nBecause of this conversion, a TreeWalk fails to identify a file with\nGerman umlauts. When using platform encoding to convert the file path to\nbytes:\n\n  private PathFilter(final String s) {\n    pathStr = s;\n    pathRaw = s.getBytes();\n  }\n\nthe TreeWalk works as expected. Actually, the file path seems to be\nstored with platform encoding in the repository.\n\nIs this a bug or a misconfiguration of my repository? I'm using jgit\n(commit e16af839e8a0cc01c52d3648d2d28e4cb915f80f) on Windows.\n\nThanks!\n\n--\nBest regards,\nMarc Strapetz\n=============\nsyntevo GmbH\nhttp://www.syntevo.com\nhttp://blog.syntevo.com\n"},{"id":"128343","messageId":"200911252211.55137.robin.rosenberg@dewire.com","threadId":"21749","inReplyTo":"4B0D356D.1080709@syntevo.com","subject":"Re: jgit problems for file paths with non-ASCII characters","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg@dewire.com","sentAt":"2009-11-25T21:11:54Z","receivedAt":"2009-11-25T21:11:54Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"onsdag 25 november 2009 14:47:25 skrev  Marc Strapetz:\n> I have noticed that jgit converts file paths to UTF-8 when querying the\n> repository. Especially,\n> org.eclipse.jgit.treewalk.filter.PathFilter#PathFilter performs this\n> conversion:\n>\n>   private PathFilter(final String s) {\n>     pathStr = s;\n>     pathRaw = Constants.encode(pathStr);\n>   }\n>\n> Because of this conversion, a TreeWalk fails to identify a file with\n> German umlauts. When using platform encoding to convert the file path to\n> bytes:\n>\n>   private PathFilter(final String s) {\n>     pathStr = s;\n>     pathRaw = s.getBytes();e pr\n>   }\n>\n> the TreeWalk works as expected. Actually, the file path seems to be\n> stored with platform encoding in the repository.\n>\n> Is this a bug or a misconfiguration of my repository? I'm using jgit\n> (commit e16af839e8a0cc01c52d3648d2d28e4cb915f80f) on Windows.\n\nA bug. \n\nThe problem here is that we need to allow multiple encodings since there\nis no reliable encoding specified anywhere. The approach I advocate is\nthe one we use for handling encoding in general. I.e. if it looks like UTF-8,\ntreat it like that else fallback. This is expensive however and then we have\nall the other issues with case insensitive name and the funny property that\nunicode has when it allows characters to be encoding using multiple sequences\nof code points as empoloyed by Apple.\n\n-- robin\n\n\n-- robin\n"},{"id":"128412","messageId":"20091126005423.GM11919@spearce.org","threadId":"21749","inReplyTo":"200911252211.55137.robin.rosenberg@dewire.com","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-11-26T00:54:23Z","receivedAt":"2009-11-26T00:54:23Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Robin Rosenberg <robin.rosenberg@dewire.com> wrote:\n> onsdag 25 november 2009 14:47:25 skrev  Marc Strapetz:\n> > I have noticed that jgit converts file paths to UTF-8 when querying the\n> > repository.\n...\n> > Is this a bug or a misconfiguration of my repository? I'm using jgit\n> > (commit e16af839e8a0cc01c52d3648d2d28e4cb915f80f) on Windows.\n> \n> A bug. \n> \n> The problem here is that we need to allow multiple encodings since there\n> is no reliable encoding specified anywhere.\n\nThis is a design fault of both Linux and git.  git gets a byte\nsequence from readdir and stores that as-is into the repository.\nWe have no way of knowing what that encoding is.  So now everyone\ntouching a Git repository is screwed.\n\n> The approach I advocate is\n> the one we use for handling encoding in general. I.e. if it looks like UTF-8,\n> treat it like that else fallback. This is expensive however\n\nWe should try to work harder with the git-core folks to get character\nset encoding for file names worked out.  We might be able to use a\nconfiguration setting in the repository to tell us what the proper\nencoding should be, and if not set, assume UTF-8.\n\n> and then we have\n> all the other issues with case insensitive name and the funny property that\n> unicode has when it allows characters to be encoding using multiple sequences\n> of code points as empoloyed by Apple.\n\nBut as you said, this still doesn't make the Apple normal form\nany easier.  Though if we know we are on such a strange filesystem\nwe might be able to assume the paths in the repository are equally\ndamaged.  Or not.\n\n-- \nShawn.\n"},{"id":"128461","messageId":"4B0E7DF5.9040007@syntevo.com","threadId":"21749","inReplyTo":"20091126005423.GM11919@spearce.org","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Thomas Singer","fromEmail":"thomas.singer@syntevo.com","sentAt":"2009-11-26T13:09:09Z","receivedAt":"2009-11-26T13:09:09Z","isPatch":false,"sender":{"key":"thomas.singer@syntevo.com","avatar":null},"body":"> But as you said, this still doesn't make the Apple normal form\n> any easier.  Though if we know we are on such a strange filesystem\n> we might be able to assume the paths in the repository are equally\n> damaged.  Or not.\n\nWell, if the git-core folks could standardize on, e.g., composed UTF-8\n(rather then just UTF-8), for storing file names in the repository, then\neverything should be clear, isn't it?\n\n--\nBest regards,\nThomas Singer\n=============\nsyntevo GmbH\nhttp://www.syntevo.com\nhttp://blog.syntevo.com\n\n\nShawn O. Pearce wrote:\n> Robin Rosenberg <robin.rosenberg@dewire.com> wrote:\n>> onsdag 25 november 2009 14:47:25 skrev  Marc Strapetz:\n>>> I have noticed that jgit converts file paths to UTF-8 when querying the\n>>> repository.\n> ...\n>>> Is this a bug or a misconfiguration of my repository? I'm using jgit\n>>> (commit e16af839e8a0cc01c52d3648d2d28e4cb915f80f) on Windows.\n>> A bug. \n>>\n>> The problem here is that we need to allow multiple encodings since there\n>> is no reliable encoding specified anywhere.\n> \n> This is a design fault of both Linux and git.  git gets a byte\n> sequence from readdir and stores that as-is into the repository.\n> We have no way of knowing what that encoding is.  So now everyone\n> touching a Git repository is screwed.\n> \n>> The approach I advocate is\n>> the one we use for handling encoding in general. I.e. if it looks like UTF-8,\n>> treat it like that else fallback. This is expensive however\n> \n> We should try to work harder with the git-core folks to get character\n> set encoding for file names worked out.  We might be able to use a\n> configuration setting in the repository to tell us what the proper\n> encoding should be, and if not set, assume UTF-8.\n> \n>> and then we have\n>> all the other issues with case insensitive name and the funny property that\n>> unicode has when it allows characters to be encoding using multiple sequences\n>> of code points as empoloyed by Apple.\n> \n> But as you said, this still doesn't make the Apple normal form\n> any easier.  Though if we know we are on such a strange filesystem\n> we might be able to assume the paths in the repository are equally\n> damaged.  Or not.\n> \n"},{"id":"128464","messageId":"4B0E8FF2.8040206@syntevo.com","threadId":"21749","inReplyTo":"20091126005423.GM11919@spearce.org","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Marc Strapetz","fromEmail":"marc.strapetz@syntevo.com","sentAt":"2009-11-26T14:25:54Z","receivedAt":"2009-11-26T14:25:54Z","isPatch":false,"sender":{"key":"marc.strapetz@syntevo.com","avatar":"https://avatars.githubusercontent.com/u/3380730?v=4"},"body":"> We should try to work harder with the git-core folks to get character\n> set encoding for file names worked out.  We might be able to use a\n> configuration setting in the repository to tell us what the proper\n> encoding should be, and if not set, assume UTF-8.\n\nI agree that this should be the ultimate goal, though the default should\nbetter be \"system encoding\" for compatibility with current git\nrepositories and instead have newer git versions always set encoding to\nUTF-8. Thus, for our jgit clone I've introduced a system property to\nconfigure Constants.PATH_ENCODING set to system encoding. It's used by\nPathFilter and this resolves my original problem.\n\nI have tried to switch more usages from Constants.CHARACTER_ENCODING to\nConstants.PATH_ENCODING, but ended up in confusion due to my lack of\nunderstanding: primarily because I couldn't tell anymore whether encoded\nstrings were file names or not. Does it make sense to explicitly\ndistinguish encoding usages in that way? We could try to contribute here\n(and hopefully cause less review effort to jgit developers than the\nchanges itself are worth ;-)\n\n--\nBest regards,\nMarc Strapetz\n=============\nsyntevo GmbH\nhttp://www.syntevo.com\nhttp://blog.syntevo.com\n\n\n\nShawn O. Pearce wrote:\n> Robin Rosenberg <robin.rosenberg@dewire.com> wrote:\n>> onsdag 25 november 2009 14:47:25 skrev  Marc Strapetz:\n>>> I have noticed that jgit converts file paths to UTF-8 when querying the\n>>> repository.\n> ...\n>>> Is this a bug or a misconfiguration of my repository? I'm using jgit\n>>> (commit e16af839e8a0cc01c52d3648d2d28e4cb915f80f) on Windows.\n>> A bug. \n>>\n>> The problem here is that we need to allow multiple encodings since there\n>> is no reliable encoding specified anywhere.\n> \n> This is a design fault of both Linux and git.  git gets a byte\n> sequence from readdir and stores that as-is into the repository.\n> We have no way of knowing what that encoding is.  So now everyone\n> touching a Git repository is screwed.\n> \n>> The approach I advocate is\n>> the one we use for handling encoding in general. I.e. if it looks like UTF-8,\n>> treat it like that else fallback. This is expensive however\n> \n> We should try to work harder with the git-core folks to get character\n> set encoding for file names worked out.  We might be able to use a\n> configuration setting in the repository to tell us what the proper\n> encoding should be, and if not set, assume UTF-8.\n> \n>> and then we have\n>> all the other issues with case insensitive name and the funny property that\n>> unicode has when it allows characters to be encoding using multiple sequences\n>> of code points as empoloyed by Apple.\n> \n> But as you said, this still doesn't make the Apple normal form\n> any easier.  Though if we know we are on such a strange filesystem\n> we might be able to assume the paths in the repository are equally\n> damaged.  Or not.\n> \n"},{"id":"128466","messageId":"alpine.DEB.1.00.0911261546350.7500@intel-tinevez-2-302","threadId":"21749","inReplyTo":"4B0E7DF5.9040007@syntevo.com","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-11-26T14:47:29Z","receivedAt":"2009-11-26T14:47:29Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 26 Nov 2009, Thomas Singer wrote:\n\n> [someone said, Thomas did not say who]\n>\n> > But as you said, this still doesn't make the Apple normal form any \n> > easier.  Though if we know we are on such a strange filesystem we \n> > might be able to assume the paths in the repository are equally \n> > damaged.  Or not.\n> \n> Well, if the git-core folks could standardize on, e.g., composed UTF-8 \n> (rather then just UTF-8), for storing file names in the repository, then \n> everything should be clear, isn't it?\n\nYou mean we should do the same thing as Apple with HFS?  Are you serious?\n\nCiao,\nDscho\n"},{"id":"128470","messageId":"4B0E9F69.9040502@syntevo.com","threadId":"21749","inReplyTo":"alpine.DEB.1.00.0911261546350.7500@intel-tinevez-2-302","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Thomas Singer","fromEmail":"thomas.singer@syntevo.com","sentAt":"2009-11-26T15:31:53Z","receivedAt":"2009-11-26T15:31:53Z","isPatch":false,"sender":{"key":"thomas.singer@syntevo.com","avatar":null},"body":"> You mean we should do the same thing as Apple with HFS?  Are you serious?\n\nYes, I'm serious. IMHO there should be a defined clear encoding used for\nfiles names in the repository. Otherwise you don't know what you can expect\nby reading it - it could mean anything. File names are in fact strings which\nare based on characters. To convert characters to bytes (or visa versa) you\nneed to know the encoding.\n\n--\nBest regards,\nThomas Singer\n=============\nsyntevo GmbH\nhttp://www.syntevo.com\nhttp://blog.syntevo.com\n\n\nJohannes Schindelin wrote:\n> Hi,\n> \n> On Thu, 26 Nov 2009, Thomas Singer wrote:\n> \n>> [someone said, Thomas did not say who]\n>>\n>>> But as you said, this still doesn't make the Apple normal form any \n>>> easier.  Though if we know we are on such a strange filesystem we \n>>> might be able to assume the paths in the repository are equally \n>>> damaged.  Or not.\n>> Well, if the git-core folks could standardize on, e.g., composed UTF-8 \n>> (rather then just UTF-8), for storing file names in the repository, then \n>> everything should be clear, isn't it?\n> \n> You mean we should do the same thing as Apple with HFS?  Are you serious?\n> \n> Ciao,\n> Dscho\n"},{"id":"128477","messageId":"200911261744.43917.robin.rosenberg@dewire.com","threadId":"21749","inReplyTo":"4B0E7DF5.9040007@syntevo.com","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg@dewire.com","sentAt":"2009-11-26T16:44:43Z","receivedAt":"2009-11-26T16:44:43Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"torsdag 26 november 2009 14:09:09 skrev  Thomas Singer:\n> > But as you said, this still doesn't make the Apple normal form\n> > any easier.  Though if we know we are on such a strange filesystem\n> > we might be able to assume the paths in the repository are equally\n> > damaged.  Or not.\n>\n> Well, if the git-core folks could standardize on, e.g., composed UTF-8\n> (rather then just UTF-8), for storing file names in the repository, then\n> everything should be clear, isn't it?\n\nHey, we're trying to enforce composed characters...\n\n-- robin\n"},{"id":"128496","messageId":"20091126195736.GV11919@spearce.org","threadId":"21749","inReplyTo":"4B0E9F69.9040502@syntevo.com","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-11-26T19:57:36Z","receivedAt":"2009-11-26T19:57:36Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Thomas Singer <thomas.singer@syntevo.com> wrote:\n> > You mean we should do the same thing as Apple with HFS?  Are you serious?\n> \n> Yes, I'm serious. IMHO there should be a defined clear encoding used for\n> files names in the repository. Otherwise you don't know what you can expect\n> by reading it - it could mean anything. File names are in fact strings which\n> are based on characters. To convert characters to bytes (or visa versa) you\n> need to know the encoding.\n\nThat's likely not going to fly.  HFS+ has changed their decomposition\nrules at least once, which means the byte sequence for the same\ncharacter sequence would differ, and a tree or commit hash would\ncome out different depending upon which rules you were following.\nSee [1] for details on what HFS+ does.\n\nAlso, Linus has previously stated HFS+ chose the worst possible\nway to encode the names.  Getting Linus to admit he was wrong is\nimpossible, getting Linus to accept the HFS+ encoding rules as the\nstandard format used in a Git repository is not likely to happen.\nFortunately Linus carries a slightly smaller stick in Git than he\nused to, but he is quite vocal and people tend to listen.\n\n[1] http://developer.apple.com/mac/library/technotes/tn/tn1150.html#UnicodeSubtleties\n\n-- \nShawn.\n"},{"id":"128497","messageId":"20091126200335.GW11919@spearce.org","threadId":"21749","inReplyTo":"4B0E8FF2.8040206@syntevo.com","subject":"Re: [egit-dev] Re: jgit problems for file paths with non-ASCII characters","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-11-26T20:03:35Z","receivedAt":"2009-11-26T20:03:35Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Marc Strapetz <marc.strapetz@syntevo.com> wrote:\n> > We should try to work harder with the git-core folks to get character\n> > set encoding for file names worked out.  We might be able to use a\n> > configuration setting in the repository to tell us what the proper\n> > encoding should be, and if not set, assume UTF-8.\n> \n> I agree that this should be the ultimate goal, though the default should\n> better be \"system encoding\" for compatibility with current git\n> repositories and instead have newer git versions always set encoding to\n> UTF-8. Thus, for our jgit clone I've introduced a system property to\n> configure Constants.PATH_ENCODING set to system encoding. It's used by\n> PathFilter and this resolves my original problem.\n\nThat's probably a good point, using the system encoding on a\nrepository may produce the file names in a more compatible way\nwith git-core.  But we probably don't want the encoding to be a\nsingle encoding constant in this JVM, we probably need to support\na per-repository configuration of the encoding for path names so\nthat we can eventually move to a non-platform specific encoding.\n\n> I have tried to switch more usages from Constants.CHARACTER_ENCODING to\n> Constants.PATH_ENCODING, but ended up in confusion due to my lack of\n> understanding: primarily because I couldn't tell anymore whether encoded\n> strings were file names or not.\n\nHeh.  Yea.  There are a number of file name encoding sites.  I think\neverything in the treewalk package, as well as the GitIndex, Tree and\nDirCache* classes.  Also the Patch class and its FileHeader friend.\n\n> Does it make sense to explicitly\n> distinguish encoding usages in that way? We could try to contribute here\n> (and hopefully cause less review effort to jgit developers than the\n> changes itself are worth ;-)\n\nYes, it does.  Because we eventually need to support encodings\nother than the current UTF-8 we assume for file names, especially\nif a repository is using the local filesystem encoding and that\nisn't UTF-8.\n\n-- \nShawn.\n"}]}