Re: non-ascii filenames issue
- From
- demerphq <demerphq@gmail.com>
- Date
- Apr 7, 2009, 08:26 UTC
- Message-ID
- <9b18b3110904070126i354fc100l69b6ce3c9cd19d49@mail.gmail.com>
- In-Reply-To
- <alpine.DEB.2.00.0904060823400.21376@ds9.cixit.se>
2009/4/6 Peter Krefting <peter@softwolves.pp.se>:
Show 12 quoted lines
> John Tapsell: > >> Unfortunately not, because for some absolutely crazy reason, there is no >> way at all to tell what encoding the string is in. It never occured to >> anyone that it might actually be useful to be able to read the filename in >> an unambiguous way. > > It comes from the Unix tradition, unfortunately, that file names are just a > stream of bytes, instead of a stream of characters mapped to a byte > sequence. The "stream of bytes" think worked back when everyone used ASCII, > but as soon as other character encodings were used (i.e back in the 1970s or > so), that assumption broke.
Those interested in this subject may find the following document on the creation of utf8 interesting.
http://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt
cheers, Yves
-- perl -Mre=debug -e "/just|another|perl|hacker/"