git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Re: Moving a directory into another fails

From
Shawn Pearce <spearce@spearce.org>
Date
Dec 4, 2006, 20:54 UTC
Message-ID
<20061204205407.GB6764@spearce.org>
In-Reply-To
<Pine.LNX.4.64.0612041114240.3476@woody.osdl.org>
Linus Torvalds <torvalds@osdl.org> wrote:
Show 5 quoted lines
> You guys are ignoring the _real_ problem. 
> 
> It has nothing at all to do with dependencies on external packages. The 
> REAL problem is that if you do locale-dependent trees and other git 
> objects, git will STOP WORKING.
Yes!

In jgit I assumed all tree entry names were encoded in UTF8. Then I later learned they aren't. Foolish me.

As Linus points out its a HUGE problem that the caller of git-write-tree gets to decide what encoding should be used for that tree. Especially if someone else wants to use a different encoding for the same filename (think ISO-8859-1 vs. UTF-8)!

I'd rather just force the tree entry names to be encoded in UTF-8 always, as its compact for most western texts (which many filenames are), and at least degrades to supporting the non western texts.

A per-project setting is essentially impossible as we have
no such concept today, and a per-repository setting (like
i18n.commitEncoding) lets two different users encode the same
filename differently, which means two different tree SHA1s with
the exact same content... not correct!
 
Show 6 quoted lines
> This is true for all levels of the git archive. It's true for blob 
> content, it's true for filenames in trees, and it is true for commits. The 
> commit message is actually somewhat easier (because we have nothing to 
> "compare" it to afterwards in the checked-out tree), so the commit message 
> is the _one_ thing we can kind of play games with, but even there, once 
> it's done, it's done, and it's just a stream of bytes.

Commit encoding is a problem. Clearly the "header parts" (tree, parent) are US-ASCII but the author and committer lines can be anything. So can the body. And we have no way of knowing what encoding was used years later, we can only guess and display it wrong.

We really should either normalize all commit messages to a single encoding (again, UTF-8) or embed the encoding as part of the headers somehow (e.g. look at how XML embeds the document encoding in the start of the document).

Previous: Linus TorvaldsNext: Jakub Narebski
Message 14 of 24 in “Moving a directory into another fails”
  1. Jon SmirlJul 26, 2006
  2. Nicolas VilzJul 26, 2006
  3. Jon SmirlJul 26, 2006
  4. Nicolas VilzJul 26, 2006
  5. Petr BaudisJul 28, 2006
  6. Stefan PfetzingDec 4, 2006
  7. Jakub NarebskiDec 4, 2006
  8. Johannes SchindelinDec 4, 2006
  9. Jakub NarebskiDec 4, 2006
  10. Johannes SchindelinDec 4, 2006
  11. Jakub NarebskiDec 4, 2006
  12. Linus TorvaldsDec 4, 2006
  13. Linus TorvaldsDec 4, 2006
  14. Shawn PearceDec 4, 2006
  15. Jakub NarebskiDec 4, 2006
  16. Johannes SchindelinDec 4, 2006
  17. Linus TorvaldsDec 4, 2006
  18. Johannes SchindelinDec 5, 2006
  19. Jakub NarebskiDec 5, 2006
  20. filesystem encodings and gitweb tests, was Re: Moving a directory into another failsJohannes Schindelin, Dec 5, 2006
  21. Jakub NarebskiDec 5, 2006
  22. Johannes SchindelinDec 5, 2006
  23. Linus TorvaldsDec 5, 2006
  24. Johannes SchindelinDec 4, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.