git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language

From
Carlo Arenas <carenas@gmail.com>
Date
Jul 11, 2025, 22:12 UTC
Message-ID
<CAPUEsphkzaibm2FMBoj-9nbFch7UgRvyvmzErmno0z+2k5X+OA@mail.gmail.com>
In-Reply-To
<aHGCRLGHEB0m_cXZ@fruit.crustytoothpaste.net>

On Fri, Jul 11, 2025 at 2:29 PM brian m. carlson <sandals@crustytoothpaste.net> wrote:

Show 30 quoted lines
>
> On 2025-07-11 at 20:57:03, Carlo Marcelo Arenas Belón wrote:
> > except that it would be incorrect, as language tags are defined in RFC5646
> > and are larger than that.
> >
> > most importantly, deriving language tags from locales provides some very
> > useful tags when including the characters after the _, because zh_CN and
> > zh_HK use completely different scripts, for example.
>
> Yes, that's true.  You have some private use and some irregular tags and
> you also have some tags that include scripts or country codes.
>
> For instance, Swahili can be written in Latin or Arabic script.  As I
> understand it, the Arabic script form is older and less common these
> days, so if I learned Swahili (which I would like to), then I might only
> learn the Latin script variant in a course.  I would need to specify
> that script in the language code to be sure that I was presented with
> content in a form that I could read and understand.  Similar concerns
> exist with the variants of Serbo-Croatian: some are written in Latin
> scripts, some in Cyrillic, and some in both, and it's not guaranteed
> that all speakers understand all forms.
>
> And then there's pt-PT and pt-BR, which are not always mutually
> intelligible.  Most free software I've seen ships these as separate
> translations.
>
> I don't want to implement language tag parsing here since we don't need
> to do that.  I would like to do the simple thing to prevent commonly
> used locales that don't represent actual language tags from being
> included and not overengineer this design

I think that your design of filtering C and POSIX accomplishes that, even if it might seem like hardcoding those two values is a little dirty.

Moving the logic (including the filtering, which is already happening for the `!NO_GETTEXT `code path adds several chances to modernize and cleanup the code though which will be beneficial (ex: using and strvec or even a hashtable to process the candidates, improve validation and tests)

Carlo

CC Yi EungJun at a hopefully working email address with link to thread https://lore.kernel.org/git/20250710221641.857081-1-sandals@crustytoothpaste.net/

.
> --
> brian m. carlson (they/them)
> Toronto, Ontario, CA
Previous: brian m. carlsonNext: Collin Funk
Message 9 of 18 in “Filter C and POSIX out of Accept-Language”
  1. 0/1 Filter C and POSIX out of Accept-Languagebrian m. carlson, Jul 10, 2025
  2. 1/1 http: don't send C or POSIX in Accept-Languagebrian m. carlson, Jul 10, 2025
  3. Junio C HamanoJul 10, 2025
  4. brian m. carlsonJul 10, 2025
  5. Justin ToblerJul 11, 2025
  6. Collin FunkJul 11, 2025
  7. Carlo Marcelo Arenas BelónJul 11, 2025
  8. brian m. carlsonJul 11, 2025
  9. Carlo ArenasJul 11, 2025
  10. Collin FunkJul 11, 2025
  11. Junio C HamanoJul 11, 2025
  12. Carlo Marcelo Arenas BelónJul 11, 2025
  13. Eli SchwartzJul 15, 2025
  14. Junio C HamanoJul 10, 2025
  15. brian m. carlsonJul 10, 2025
  16. Collin FunkJul 10, 2025
  17. Han YoungJul 11, 2025
  18. Junio C HamanoJul 11, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.