Re: [PATCH] gitweb: Fix chop_str not to cut in middle of utf8 multibyte chars.
- From
Anders Waldenborg <anders@0x63.nu>
- Date
- May 21, 2008, 07:45 UTC
- Message-ID
- <4833D314.4010904@0x63.nu>
- In-Reply-To
- <7vve185d6s.fsf@gitster.siamese.dyndns.org>
Show 7 quoted lines
> I haven't followed the codepath but what do the callers do to the string > returned from chop_str? Don't they assume the string hasn't been decoded > (because the old implementation of chop_str did not do this to_utf8), and > emit the result directly to the output because it also assumes the > undecoded format is what the outside world wants? In other words, don't > they now need to do different things because returned string has gone > through the to_utf8() processing already?
The to_utf8() (defined in gitweb.perl, not part of perl it self) is kind of sneaky, it checks if the string already is valid utf8. (guess it should be called ensure_utf8())
chop_str needs to work on decoded string, otherwise character count goes all wrong. But maybe it is better to add the to_utf8() to the callsites?
anders