git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v4] pretty-formats: add hard truncation, without ellipsis, options

From
Philip Oakley <philipoakley@iee.email>
Date
Nov 28, 2022, 13:39 UTC
Message-ID
<b7b84dde-a723-0773-279f-c04c7f35cb7f@iee.email>
In-Reply-To
<xmqq35a5cnhq.fsf@gitster.g>
On 26/11/2022 23:19, Junio C Hamano wrote:
Show 39 quoted lines
> Philip Oakley <philipoakley@iee.email> writes:
>
>>>  in that they may do "[][].." or "[][][]" when told to
>>> "trunc" fill a string with four or more double-width letters into a
>>> 5 display space.  But the point is at least for these with ellipsis
>>> it is fairly clear what the desired behaviour is.
>> That "is fairly clear" is probably the problem. In retrospect it's not
>> clear in the docs that the "%<(N" format is (would appear to be) about
>> defining the display width, in terminal character columns, that the
>> selected parameter is to be displayed within.
>>
>> The code already pads the displayed parameter with spaces as required if
>> the parameter is shorter than the display width - the else condition in
>> pretty.c L1750
>>
>>>   For "trunc" in
>>> the above example, I think the right thing for it to do would be to
>>> do "[][].", i.e. consume exactly 5 display columns, and avoid
>>> exceeding the given space by not giving two dots but just one.
>> The existing choice is padding "[][]" with a single space to reach 5
>> display chars.
>> For the 6-char "[][][]" truncation it is "[][..", i.e. 3 chars from
>> "[][][]", then the two ".." dots of the ellipsis.
> Here, I realize that I did not explain the scenario well.  The
> message you are responding to was meant to be a clarification of my
> earlier message and it should have done a better job but apparently
> I failed.  Sorry, and let me try again.
>
> The single example I meant to use to illustrate the scenario I worry
> about is this.  There is a string, in which there are four (or more)
> letters, each of which occupies two display columns.  And '[]' in my
> earlier messages stood for a SINGLE such letter (I just wanted to
> stick to ASCII, instead of using East Asian script, for
> illustration).  So "[][.." is not possible (you are chomping the
> second such letter in half).
>
> I could use East Asian 一二三四 (there are four letters, denoting
> one, two, three, and four, each occupying two display spaces when
> typeset in a fixed width font),

Thanks for that clarification, I'd been thinking it was about c char (bytes) such as ASCII and multi-byte characters (code points), e.g. European umlaut style distinctions.

I hadn't really picked up on the distinction between wide and narrow
'glyphs' (if that's the right term to use).
 
I see that the code does properly count the widths of narrow and wide
code points as 1 and 2 columns of the display, but then doesn't
explicitly try any adjustment for the wide code point problem you noted.
Show 18 quoted lines
>  but to make it easier to see in
> ASCII only text, let's pretend "[1]", "[2]", "[3]", "[4]" are such
> letters.  You cannot chomp them in the middle (and please pretend
> each of them occupy two, not three, display spaces).
>
> When the given display space is 6 columns, we can fit 2 such letters
> plus ".." in the space.  If the original string were [1][2][3][4],
> it is clear trunk and ltrunk can do "[1][2].." (remember [n] stands
> for a single letter whose width is 2 columns, so that takes 6
> columns) and "..[3][4]", respectively.  It also is clear that Trunk
> and Ltrunk can do "[1][2][3]" and "[2][3][4]", respectively.  We
> truncate the given string so that we fill the alloted display
> columns fully.
>
> If the given display space is 5 columns, the desirable behaviour for
> trunk and ltrunk is still clear.  Instead of consuming two dots, we
> could use a single dot as the filler.  As I said, I suspect that the
> implementation of trunk and ltrunc does this correctly, though.

I believe there is a possible solution that, if we detect a column over-run, then we can go back and replace the current two column double dot with a narrow U+2026 Horizontal ellipsis, to regain the needed column.

>
> My worry is it is not clear what Trunk and Ltrunk should do in that
> case.  There is no way to fit a substring of [1][2][3][4] into 5
> columns without any filler.

For this case where the final code point overruns, my solution could/would be to use the Vertical ellipsis U+22EE "⋮" to re-write that final character (though the Unicode Replacement Character "�" could be used, but that's ugly)

I suspect the code would need some close reading to ensure that the column counting and replacement would correctly cope with the 'off by one' wide width case inside the strbuf_utf8_replace().

I.e. given the same off-by-one position and replacement length, get back to the same point to replace either the double dot or the final code point in an idempotent manner.

The logic feels sound, as long as there are no three wide crocodile code-points. Either we counted the right number of columns, or we over-ran by one, so we go back and substitute with a one-for-two replacement.

Philip

For watchers, https://github.com/microsoft/terminal/issues/4345 shows some of the issues in the general case.

Previous: Junio C HamanoNext: Junio C Hamano
Message 21 of 30 in “extend the truncating pretty formats”
  1. 0/1 extend the truncating pretty formatsPhilip Oakley, Oct 30, 2022
  2. 1/1 pretty-formats: add hard truncation, without ellipsis, optionsPhilip Oakley, Oct 30, 2022
  3. Taylor BlauOct 30, 2022
  4. Philip OakleyOct 30, 2022
  5. Taylor BlauOct 30, 2022
  6. Philip OakleyOct 30, 2022
  7. 0/1 extend the truncating pretty formatsPhilip Oakley, Nov 1, 2022
  8. 1/1 pretty-formats: add hard truncation, without ellipsis, optionsPhilip Oakley, Nov 1, 2022
  9. Philip OakleyNov 1, 2022
  10. Taylor BlauNov 2, 2022
  11. pretty-formats: add hard truncation, without ellipsis, optionsPhilip Oakley, Nov 2, 2022
  12. pretty-formats: add hard truncation, without ellipsis, optionsPhilip Oakley, Nov 12, 2022
  13. Junio C HamanoNov 21, 2022
  14. Philip OakleyNov 21, 2022
  15. Junio C HamanoNov 22, 2022
  16. Philip OakleyNov 23, 2022
  17. Junio C HamanoNov 25, 2022
  18. Philip OakleyNov 26, 2022
  19. Philip OakleyNov 26, 2022
  20. Junio C HamanoNov 26, 2022
  21. Philip OakleyNov 28, 2022
  22. Junio C HamanoNov 29, 2022
  23. Philip OakleyDec 7, 2022
  24. Junio C HamanoDec 7, 2022
  25. 0/5 Pretty formats: Clarify column alignmentPhilip Oakley, Jan 19, 2023
  26. 3/5 doc: pretty-formats document negative column alignmentsPhilip Oakley, Jan 19, 2023
  27. 1/5 doc: pretty-formats: separate parameters from placeholdersPhilip Oakley, Jan 19, 2023
  28. 2/5 doc: pretty-formats: delineate `%<|(` parameter valuesPhilip Oakley, Jan 19, 2023
  29. 4/5 doc: pretty-formats describe use of ellipsis in truncationPhilip Oakley, Jan 19, 2023
  30. 5/5 doc: pretty-formats note wide char limitations, and add testsPhilip Oakley, Jan 19, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.