From: Zoltán Füzesi Date: Thu, 06 Aug 2009 08:15:25 GMT Subject: Re: [PATCH] gitweb: parse_commit_text encoding fix Message-ID: <9ab80d150908060115q4b56b2e5xb327e09cda7e2b7a@mail.gmail.com> In-Reply-To: <7viqh43vz3.fsf@alter.siamese.dyndns.org> 2009/8/4 Junio C Hamano : > > Thanks, Zoltán. > > We should be able to set up a script that scrapes the output to test this > kind of thing.  We may not want to have a test pattern that matches too > strictly for the current structure and appearance of the output > (e.g. counting nested
s, presentation styles and such), but if we can > robustly scrape off HTML tags (e.g. "elinks -dump") and check the > remaining payload, it might be enough. > > Jakub what do you think?  I suspect that scraping approach may turn out to > be too fragile for tests to be worth doing, but I am just throwing out a > thought. > This issue comes out when chop_and_escape_str function is called with a non-ascii string (like my name :)) without before calling to_utf8 on it. "author_name" and "committer_name" are two examples, and "author_name" shows up with bad encoding in HTML. Example from one of my repos (little piece from shortlog output): Füzesi Zoltán After applying the patch: Füzesi Zoltán This is an "old" (seen in 1.5.6 version too) and (I think) minor issue. I haven't spent time on thinking how a test script could show this yet. Waiting for Jakub's reaction.