git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v4 1/2] diff.c: When appropriate, use utf8_strwidth(), part1

From
Junio C Hamano <gitster@pobox.com>
Date
Sep 7, 2022, 18:31 UTC
Message-ID
<xmqqedwnujce.fsf@gitster.g>
In-Reply-To
<20220907043040.idqqivi3jt35jyst@tb-raspi4>
Torsten Bögershausen <tboegi@web.de> writes:
Show 9 quoted lines
>
> OK - the comment can be removed.
>
> I didn't know how to read this comment:
>>...but the former may chomp a single multi-byte letter in the middle,
>> which would need to be corrected as a part of this change.
>
> After diffing into the code some more times, I think that we don't
> chomp a single byte out of an UTF-8 sequence.

When turning a/b/c vs a/B/c into a/{b->B}/c, two steps are involved. Take common prefix and suffix (in this case 'a' and 'c') and turn 'b' vs 'B' into {b->B} is one step. The other is what to do when prefix and suffix are long. After turning aaaaa/b/c vs aaaaa/B/c into aaaaa/{b->B}/c, if the result is overly long, how we shorten the prefix (i.e. aaaaa) and the suffix?

I knew the code that produces {b->B} honored '/' boundary, but I just did not remember offhand what diff.c::pprint_rename() did in its latter half, specifically, if it just chomped pfx and sfx as a sequence of bytes (which would have been wrong) or insisted that the common sequence search honors '/' boundary (which would be OK, as byte '/' will not appear in the middle of a single multi-byte UTF-8 "letter"). I think iti s doing the latter, so it should be fine.

Thanks.
Previous: Torsten BögershausenNext: tboegi@web.de
Message 30 of 42 in “[BUG] Unicode filenames handling in `git log --stat`”
  1. Alexander MeshcheryakovAug 9, 2022
  2. Calvin WanAug 9, 2022
  3. Alexander MeshcheryakovAug 9, 2022
  4. Calvin WanAug 9, 2022
  5. Junio C HamanoAug 10, 2022
  6. Torsten BögershausenAug 10, 2022
  7. Alexander MeshcheryakovAug 10, 2022
  8. Torsten BögershausenAug 10, 2022
  9. Torsten BögershausenAug 10, 2022
  10. Junio C HamanoAug 10, 2022
  11. Torsten BögershausenAug 10, 2022
  12. 1/1 diff.c: When appropriate, use utf8_strwidth()tboegi@web.de, Aug 14, 2022
  13. Junio C HamanoAug 14, 2022
  14. Torsten BögershausenAug 15, 2022
  15. Junio C HamanoAug 18, 2022
  16. 1/1 diff.c: When appropriate, use utf8_strwidth()tboegi@web.de, Aug 27, 2022
  17. Torsten BögershausenAug 27, 2022
  18. Eric SunshineAug 27, 2022
  19. Johannes SchindelinAug 29, 2022
  20. Torsten BögershausenAug 29, 2022
  21. Junio C HamanoAug 29, 2022
  22. Johannes SchindelinSep 2, 2022
  23. 2/2 diff.c: More changes and tests around utf8_strwidth()tboegi@web.de, Sep 2, 2022
  24. Johannes SchindelinSep 2, 2022
  25. 1/2 diff.c: When appropriate, use utf8_strwidth(), part1tboegi@web.de, Sep 2, 2022
  26. Johannes SchindelinSep 2, 2022
  27. 1/2 diff.c: When appropriate, use utf8_strwidth(), part1tboegi@web.de, Sep 3, 2022
  28. Junio C HamanoSep 5, 2022
  29. Torsten BögershausenSep 7, 2022
  30. Junio C HamanoSep 7, 2022
  31. 2/2 diff.c: More changes and tests around utf8_strwidth()tboegi@web.de, Sep 3, 2022
  32. Johannes SchindelinSep 5, 2022
  33. 1/1 diff.c: When appropriate, use utf8_strwidth()tboegi@web.de, Sep 14, 2022
  34. Junio C HamanoSep 14, 2022
  35. Torsten BögershausenSep 26, 2022
  36. Junio C HamanoOct 10, 2022
  37. Torsten BögershausenOct 20, 2022
  38. Junio C HamanoOct 20, 2022
  39. Torsten BögershausenOct 21, 2022
  40. Junio C HamanoOct 21, 2022
  41. Torsten BögershausenOct 23, 2022
  42. Junio C HamanoSep 15, 2022

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.