Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames
- From
Paul Eggert <eggert@cs.ucla.edu>
- Date
- Oct 14, 2005, 00:57 UTC
- Message-ID
- <87ek6ork3y.fsf@penguin.cs.ucla.edu>
- In-Reply-To
- <Pine.LNX.4.64.0510121355280.15297@g5.osdl.org>
Linus Torvalds <torvalds@osdl.org> writes:
> I find that email is very robust - it's basically 8-bit clean. No > character encoding, no crap. Just a byte stream. It really _is_ the most > reliable format.
I found another amusing bit of info that tends to undercut this claim.
This discussion thread is archived at <http://marc.theaimsgroup.com/?t=112877773400002&r=1&w=2&n=22>. But there's an item missing from the archive: my message with Message-ID <87vf02qy79.fsf@penguin.cs.ucla.edu>. This is the message with the joke "Aach! Those Finns! Always on the trailing edge of technology!".
All my other messages are achived. What was special about this one? Surely there's not a joke filter at theaimsgroup.com!
I nosed around through the archive and here's my guess as to what happened. My message's email header contained this:
Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable
and my guess is that the web archiver can't handle that format.
This is just a guess. I can't confirm it because (among other things) the web archiver won't give me all the bytes of the messages that it archives. Even its "Download message RAW" doesn't do that: it omits the header. But I have a strong suspicion. Let's put it this way: I think mine was the only message in the thread that said "charset=utf-8".
If my guess is right, the archiver dropped my email on the floor simply because it contained UTF-8. This is not a good sign for putting UTF-8 into email, or for relying on email to transmit byte streams.