git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: gitweb on kernel.org and UTF-8

From
Junio C Hamano <junkio@cox.net>
Date
Nov 24, 2005, 03:24 UTC
Message-ID
<7vfypm20eh.fsf@assigned-by-dhcp.cox.net>
In-Reply-To
<20051123033526.GA24098@vrfy.org>
Kay Sievers <kay.sievers@vrfy.org> writes:
> Should be fine now. The escapeHTML() garbled the utf8 "ö", and the
> decode() failed that.
Looking better.  Thanks.

This begs for addressing another issue, although I am hesitant to open this can of worms at this moment.

It might be a good idea to have configuration items gitk and gitweb can use to get a hint to decide what the commit log message and the blob data encodings might be. gitweb already adds its own information to the git repository format (.git/description), so this _could_ be stored outside just like that (e.g. .git/commit_log_encoding), but using .git/config is probably better.

How about doing something like this?
	[i18n]
        	commitEncoding = utf8
		blobEncoding = utf8
to mean:
	If you _have_ to make an assumption on an encoding
	commit and blob objects are in, utf8 is your best bet
	(but mistakes can happen, and some blobs can be binary).

Then gitweb and gitk can look at commitEncoding and blobEncoding as a hint to base its display defaults on. Sending everything out in utf8 would be sane and safe choice these days for gitweb, so if commitEncoding is latin-1 it may need to iconv latin-1 to utf8 while reading commits. For blobs, it might be better off asking file(1) or File::MMagic (since you are using Perl in gitweb --- sorry I do not know tcl equivalent of that) what they are; eventually you would want to be able to show repositories full of jpeg pictures anyway ;-).

On the commit-producing side, I could have:
	[i18n]
        	editorEncoding = latin-1

and if editorEncoding is different from commitEncoding, "git-commit -c $commit" would first iconv from utf8 to latin-1 before populating the user's editor, and iconv back from latin-1 to utf8 before feeding what the user edited to commit-tree.

Pathname encoding is the reason why I was hesitant about bring this up. Although it is too late for 1.0 now, we _could_ have declared that the paths recorded in git tree objects and index files are internally utf8, and working tree paths can be in different encoding. As a local repository configuration not project wide configuration, we could have something like:

	[i18n]
        	pathnameEncoding = latin-1

to mean that the filesystem paths returned by readdir(3) and accepted by open(2) and friends are in latin-1. Comparison and movement between working tree files, the index file, and tree objects have to involve iconv and do the right thing. So if you fetch from such a repository into a filesystem that stores pathnames in utf8, the right thing should happen.

I personally feel any sane project should restrict its pathname to ASCII only, so this issue might be moot (or is the right word "mute"?), but something like this _might_ be useful in later versions of git.

But not in 1.0.
Previous: H. Peter AnvinNext: Ryan Anderson
Message 5 of 7 in “gitweb on kernel.org and UTF-8”
  1. Junio C HamanoNov 23, 2005
  2. H. Peter AnvinNov 23, 2005
  3. Kay SieversNov 23, 2005
  4. H. Peter AnvinNov 23, 2005
  5. Junio C HamanoNov 24, 2005
  6. Ryan AndersonNov 24, 2005
  7. Junio C HamanoNov 24, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.