# [PATCH 0/1] Filter C and POSIX out of Accept-Language

18 messages from 2025-07-10 to 2025-07-15. Participants: brian m. carlson, Junio C Hamano, Collin Funk, Han Young, Justin Tobler, Carlo Marcelo Arenas Belón, Carlo Arenas, Eli Schwartz.
Thread: https://gitlist.dev/t/63780

## brian m. carlson, 2025-07-10 22:16

Subject: [PATCH 0/1] Filter C and POSIX out of Accept-Language
Message-ID: <20250710221641.857081-1-sandals@crustytoothpaste.net>
URL: https://gitlist.dev/e/20250710221641.857081-1-sandals%40crustytoothpaste.net

```
At work, I've seen some cases where people provide "C" in the
Accept-Language header of their Git requests, such as when they provide
us with debugging traces, but "C" and "POSIX", while valid locales, are
not valid languages and do not belong in the Accept-Language header.

It turns out this is actually very easy to reproduce and fix, so there's
a patch to filter these out.  I have not actually myself seen "POSIX" in
the header, but it's equivalent to "C" and I've seen it in non-Git
requests in various places online, so we reject that as well.

This can be seen in GitLab's issues as well at
https://gitlab.com/gitlab-org/gitlab/-/issues/412077.

brian m. carlson (1):
  http: don't send C or POSIX in Accept-Language

 http.c                     |  8 ++++++++
 t/t5541-http-push-smart.sh | 18 ++++++++++++++++++
 2 files changed, 26 insertions(+)


```

## brian m. carlson, 2025-07-10 22:16

Subject: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <20250710221641.857081-2-sandals@crustytoothpaste.net>
URL: https://gitlist.dev/e/20250710221641.857081-2-sandals%40crustytoothpaste.net
In-Reply-To: <20250710221641.857081-1-sandals@crustytoothpaste.net>

```
The LANGUAGE environment variable is not specified by POSIX, but a
variety of programs using GNU gettext accept it.  The Linux manpages
state that it can contain a colon-separated list of locales.

However, not all locales are valid as languages.  The C and POSIX
locales, for instance, are not languages and are not registered with
IANA, nor are they a part of ISO 639.  In fact, "C" is too short to
match the ABNF production for a language, which must be at least two
characters in length.

Nonetheless, many users provide these values in the LANGUAGE environment
variable for unknown reasons and if they do, we do not want to send a
malformed Accept-Language header to the server.  If there are no other
valid language tags, then send no header; otherwise, send only the valid
tags, ignoring "C" and "POSIX" wherever they may appear, as well as any
variants (such as the "C.UTF-8" locale found on some Linux systems).

We do not reject all possible invalid language tags since doing so
would require bundling a copy of the IANA database and would risk poor
behavior in the face of uncommon languages or values that are not
registered but meet the production for private use or other restricted
interchange.  However, these two values are widely used in the LANGUAGE
header, are well-known and widely used non-language locales, and have
been seen in the wild on the server side.

Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>
---
 http.c                     |  8 ++++++++
 t/t5541-http-push-smart.sh | 18 ++++++++++++++++++
 2 files changed, 26 insertions(+)

diff --git a/http.c b/http.c
index d88e79fbde..a96df4fcdb 100644
--- a/http.c
+++ b/http.c
@@ -2022,6 +2022,14 @@ static void write_accept_language(struct strbuf *buf)
 			s++;
 
 		if (tag.len) {
+			/*
+			 * These are not valid languages: do not send them to
+			 * the server.
+			 */
+			if (!strcmp(tag.buf, "C") || !strcmp(tag.buf, "POSIX")) {
+				strbuf_reset(&tag);
+				continue;
+			}
 			num_langs++;
 			REALLOC_ARRAY(language_tags, num_langs);
 			language_tags[num_langs - 1] = strbuf_detach(&tag, NULL);
diff --git a/t/t5541-http-push-smart.sh b/t/t5541-http-push-smart.sh
index 538b603f03..96a6833e67 100755
--- a/t/t5541-http-push-smart.sh
+++ b/t/t5541-http-push-smart.sh
@@ -86,6 +86,24 @@ test_expect_success 'push to remote repository (standard) with sending Accept-La
 	GIT_TRACE_CURL=true LANGUAGE="ko_KR.UTF-8" git push -v -v 2>err &&
 	! grep "Expect: 100-continue" err &&
 
+	grep "=> Send header: Accept-Language:" err >err.language &&
+	test_cmp exp err.language &&
+
+	test_commit C-is-not-a-language &&
+	GIT_TRACE_CURL=true LANGUAGE="C" git push -v -v 2>err &&
+
+	! grep "=> Send header: Accept-Language:" err >err.language &&
+	test_must_be_empty err.language &&
+
+	test_commit POSIX-is-not-a-language-either &&
+	GIT_TRACE_CURL=true LANGUAGE="POSIX" git push -v -v 2>err &&
+
+	! grep "=> Send header: Accept-Language:" err >err.language &&
+	test_must_be_empty err.language &&
+
+	test_commit ignore-C-and-POSIX-as-languages-wherever-provided &&
+	GIT_TRACE_CURL=true LANGUAGE="C.UTF-8:ko_KR.UTF-8:POSIX" git push -v -v 2>err &&
+
 	grep "=> Send header: Accept-Language:" err >err.language &&
 	test_cmp exp err.language
 '

```

## Junio C Hamano, 2025-07-10 22:45

Subject: Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language
Message-ID: <xmqqfrf34qdb.fsf@gitster.g>
URL: https://gitlist.dev/e/xmqqfrf34qdb.fsf%40gitster.g
In-Reply-To: <20250710221641.857081-1-sandals@crustytoothpaste.net>

```
"brian m. carlson" <sandals@crustytoothpaste.net> writes:

> At work, I've seen some cases where people provide "C" in the
> Accept-Language header of their Git requests, such as when they provide
> us with debugging traces, but "C" and "POSIX", while valid locales, are
> not valid languages and do not belong in the Accept-Language header.
>
> It turns out this is actually very easy to reproduce and fix, so there's
> a patch to filter these out.  I have not actually myself seen "POSIX" in
> the header, but it's equivalent to "C" and I've seen it in non-Git
> requests in various places online, so we reject that as well.
>
> This can be seen in GitLab's issues as well at
> https://gitlab.com/gitlab-org/gitlab/-/issues/412077.

Sorry, I am confused.  Is that Authentication failure in the cited
issue "caused by" the client sending "Accept-Language: C"?

"reproduce and fix" makes it sound like a correct exchange between
such a client and a server is somehow broken (i.e. unable to clone,
unable to authenticate, etc.) if the client sends C (or POSIX) as if
it were a langauge, but is there a breakage there?

I understand and agree with the change in patch 1/1 that it is the
right thing to do (to more strictly adhere to the standard in what
we send out) for hygiene.  I just want to understand if this caused
real problems, or if it is primarily a preemptive clean-up to avoid
non-standard behaviour causing problems in the future.

Thanks.

> brian m. carlson (1):
>   http: don't send C or POSIX in Accept-Language
>
>  http.c                     |  8 ++++++++
>  t/t5541-http-push-smart.sh | 18 ++++++++++++++++++
>  2 files changed, 26 insertions(+)

```

## Junio C Hamano, 2025-07-10 22:47

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <xmqqbjpr4q8z.fsf@gitster.g>
URL: https://gitlist.dev/e/xmqqbjpr4q8z.fsf%40gitster.g
In-Reply-To: <20250710221641.857081-2-sandals@crustytoothpaste.net>

```
"brian m. carlson" <sandals@crustytoothpaste.net> writes:

> ...  However, these two values are widely used in the LANGUAGE
> header, are well-known and widely used non-language locales, and have
> been seen in the wild on the server side.

"header" -> "environment variable" I presume?  


```

## brian m. carlson, 2025-07-10 23:08

Subject: Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language
Message-ID: <aHBH0nRLPxBg2HAj@fruit.crustytoothpaste.net>
URL: https://gitlist.dev/e/aHBH0nRLPxBg2HAj%40fruit.crustytoothpaste.net
In-Reply-To: <xmqqfrf34qdb.fsf@gitster.g>

```
[Dropping Yi EungJun from CC because their email bounced.]

On 2025-07-10 at 22:45:20, Junio C Hamano wrote:
> "brian m. carlson" <sandals@crustytoothpaste.net> writes:
> 
> > At work, I've seen some cases where people provide "C" in the
> > Accept-Language header of their Git requests, such as when they provide
> > us with debugging traces, but "C" and "POSIX", while valid locales, are
> > not valid languages and do not belong in the Accept-Language header.
> >
> > It turns out this is actually very easy to reproduce and fix, so there's
> > a patch to filter these out.  I have not actually myself seen "POSIX" in
> > the header, but it's equivalent to "C" and I've seen it in non-Git
> > requests in various places online, so we reject that as well.
> >
> > This can be seen in GitLab's issues as well at
> > https://gitlab.com/gitlab-org/gitlab/-/issues/412077.
> 
> Sorry, I am confused.  Is that Authentication failure in the cited
> issue "caused by" the client sending "Accept-Language: C"?
> 
> "reproduce and fix" makes it sound like a correct exchange between
> such a client and a server is somehow broken (i.e. unable to clone,
> unable to authenticate, etc.) if the client sends C (or POSIX) as if
> it were a langauge, but is there a breakage there?

No, sorry.  I just meant that the trace in that issue demonstrates the
incorrect Accept-Language header; it's unrelated to the authentication
problem that the issue is about (which I think is a GitLab issue).

> I understand and agree with the change in patch 1/1 that it is the
> right thing to do (to more strictly adhere to the standard in what
> we send out) for hygiene.  I just want to understand if this caused
> real problems, or if it is primarily a preemptive clean-up to avoid
> non-standard behaviour causing problems in the future.

I'm not aware of it causing any practical problems for people, although
I could imagine some cases where it could, in theory, break things.  I
merely noticed this in trace output and thought we should tidy it up.
If users are using the header and expecting a localized response, this
will make it more likely that they get the one they were expecting.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

```

## brian m. carlson, 2025-07-10 23:08

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <aHBH-3io3rqw4pK4@fruit.crustytoothpaste.net>
URL: https://gitlist.dev/e/aHBH-3io3rqw4pK4%40fruit.crustytoothpaste.net
In-Reply-To: <xmqqbjpr4q8z.fsf@gitster.g>

```
On 2025-07-10 at 22:47:56, Junio C Hamano wrote:
> "brian m. carlson" <sandals@crustytoothpaste.net> writes:
> 
> > ...  However, these two values are widely used in the LANGUAGE
> > header, are well-known and widely used non-language locales, and have
> > been seen in the wild on the server side.
> 
> "header" -> "environment variable" I presume?  

Ah, yes.  I'll fix that for a v2 in a few days while I wait for other
comments.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

```

## Collin Funk, 2025-07-10 23:26

Subject: Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language
Message-ID: <m1h5zjk4pv.fsf@gmail.com>
URL: https://gitlist.dev/e/m1h5zjk4pv.fsf%40gmail.com
In-Reply-To: <aHBH0nRLPxBg2HAj@fruit.crustytoothpaste.net>

```
"brian m. carlson" <sandals@crustytoothpaste.net> writes:

> I'm not aware of it causing any practical problems for people, although
> I could imagine some cases where it could, in theory, break things.  I
> merely noticed this in trace output and thought we should tidy it up.
> If users are using the header and expecting a localized response, this
> will make it more likely that they get the one they were expecting.

I feel like it is a bit strange to only exclude "C" or "POSIX".

I think the correct behavior would be to accept any values, or convert
the current locale to the closest BCP 47 language tag.

But as you mentioned converting them would require a database of all
tags...

Collin

```

## Han Young, 2025-07-11 02:49

Subject: Re: [External] Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language
Message-ID: <CAG1j3zGn5fS=_Oftu7bBmWsoMc-aCa84AtDXdfxgL8QFEkp+yA@mail.gmail.com>
URL: https://gitlist.dev/e/CAG1j3zGn5fS%3D_Oftu7bBmWsoMc-aCa84AtDXdfxgL8QFEkp%2ByA%40mail.gmail.com
In-Reply-To: <m1h5zjk4pv.fsf@gmail.com>

```
On Fri, Jul 11, 2025 at 7:26 AM Collin Funk <collin.funk1@gmail.com> wrote:

> I think the correct behavior would be to accept any values, or convert
> the current locale to the closest BCP 47 language tag.

On some Linux systems, not all BCP languages are supported. Not all
Linux distributions generate all the locales, and musl doesn't even support
locales. Converting to the closest BCP 47 language alone does not
ensure the locale is valid. Not to mention the tricky heuristics of language
matching (pt_PT or pt_BR if LANGUAGE is pt?).

> But as you mentioned converting them would require a database of all
> tags...

Hardcoding all the locale names in the code should be fine, I guess?
Though the problem of filtering out locales unsupported by glibc is more
troublesome.

```

## Justin Tobler, 2025-07-11 15:23

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <r34i7fhxwbxhppc4ia7lpyr3xqj4tgusaeikaaonpwtywlywxw@ygfmv3f3q67u>
URL: https://gitlist.dev/e/r34i7fhxwbxhppc4ia7lpyr3xqj4tgusaeikaaonpwtywlywxw%40ygfmv3f3q67u
In-Reply-To: <20250710221641.857081-2-sandals@crustytoothpaste.net>

```
On 25/07/10 10:16PM, brian m. carlson wrote:
> The LANGUAGE environment variable is not specified by POSIX, but a
> variety of programs using GNU gettext accept it.  The Linux manpages
> state that it can contain a colon-separated list of locales.
> 
> However, not all locales are valid as languages.  The C and POSIX
> locales, for instance, are not languages and are not registered with
> IANA, nor are they a part of ISO 639.  In fact, "C" is too short to
> match the ABNF production for a language, which must be at least two
> characters in length.
> 
> Nonetheless, many users provide these values in the LANGUAGE environment
> variable for unknown reasons and if they do, we do not want to send a
> malformed Accept-Language header to the server.  If there are no other
> valid language tags, then send no header; otherwise, send only the valid
> tags, ignoring "C" and "POSIX" wherever they may appear, as well as any
> variants (such as the "C.UTF-8" locale found on some Linux systems).

Ok so the languages returned by `get_preferred_languages()` are used to
write the Accept-Language header when making requests.

Looking at `get_preferred_languages()` when NO_GETTEXT is defined, we
already filter out "C" and "POSIX". So doing this for the LANGUAGE
environment variable when writing the header also makes sense.

> We do not reject all possible invalid language tags since doing so
> would require bundling a copy of the IANA database and would risk poor
> behavior in the face of uncommon languages or values that are not
> registered but meet the production for private use or other restricted
> interchange.  However, these two values are widely used in the LANGUAGE
> header, are well-known and widely used non-language locales, and have
> been seen in the wild on the server side.
> 
> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>
> ---
>  http.c                     |  8 ++++++++
>  t/t5541-http-push-smart.sh | 18 ++++++++++++++++++
>  2 files changed, 26 insertions(+)
> 
> diff --git a/http.c b/http.c
> index d88e79fbde..a96df4fcdb 100644
> --- a/http.c
> +++ b/http.c
> @@ -2022,6 +2022,14 @@ static void write_accept_language(struct strbuf *buf)
>  			s++;
>  
>  		if (tag.len) {
> +			/*
> +			 * These are not valid languages: do not send them to
> +			 * the server.
> +			 */
> +			if (!strcmp(tag.buf, "C") || !strcmp(tag.buf, "POSIX")) {
> +				strbuf_reset(&tag);
> +				continue;
> +			}

From my understanding, each language is expected to be defined in the
following form:

  language[_territory][.codeset][@modifier]

When we parse the list of languages we only care about the
`language[_territory]` part though.

From looking at ISO 639 language codes, only codes with two or three
characters are valid. If we wanted to be a bit more strict, we could
check the length of the language code (everything before the first '_')
and filter out anything outside of those limits. This would naturally
filter out "C" and "POSIX" without having to mention them explicitly.

Not sure if being more strict adds much more value here in practice
though. So it may be fine to keep it as-is. :)

>  			num_langs++;
>  			REALLOC_ARRAY(language_tags, num_langs);
>  			language_tags[num_langs - 1] = strbuf_detach(&tag, NULL);
> diff --git a/t/t5541-http-push-smart.sh b/t/t5541-http-push-smart.sh
> index 538b603f03..96a6833e67 100755
> --- a/t/t5541-http-push-smart.sh
> +++ b/t/t5541-http-push-smart.sh
> @@ -86,6 +86,24 @@ test_expect_success 'push to remote repository (standard) with sending Accept-La
>  	GIT_TRACE_CURL=true LANGUAGE="ko_KR.UTF-8" git push -v -v 2>err &&
>  	! grep "Expect: 100-continue" err &&
>  
> +	grep "=> Send header: Accept-Language:" err >err.language &&
> +	test_cmp exp err.language &&
> +
> +	test_commit C-is-not-a-language &&
> +	GIT_TRACE_CURL=true LANGUAGE="C" git push -v -v 2>err &&
> +
> +	! grep "=> Send header: Accept-Language:" err >err.language &&
> +	test_must_be_empty err.language &&
> +
> +	test_commit POSIX-is-not-a-language-either &&
> +	GIT_TRACE_CURL=true LANGUAGE="POSIX" git push -v -v 2>err &&
> +
> +	! grep "=> Send header: Accept-Language:" err >err.language &&
> +	test_must_be_empty err.language &&

The above two tests demonstrate that the Accept-Language header is not
sent if no valid languages are found.

> +
> +	test_commit ignore-C-and-POSIX-as-languages-wherever-provided &&
> +	GIT_TRACE_CURL=true LANGUAGE="C.UTF-8:ko_KR.UTF-8:POSIX" git push -v -v 2>err &&
> +
>  	grep "=> Send header: Accept-Language:" err >err.language &&
>  	test_cmp exp err.language
>  '

And here we see only the valid languages sent in the header. Looks good!

-Justin

```

## Junio C Hamano, 2025-07-11 16:46

Subject: Re: [External] Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language
Message-ID: <xmqq5xfy3cbb.fsf@gitster.g>
URL: https://gitlist.dev/e/xmqq5xfy3cbb.fsf%40gitster.g
In-Reply-To: <CAG1j3zGn5fS=_Oftu7bBmWsoMc-aCa84AtDXdfxgL8QFEkp+yA@mail.gmail.com>

```
Han Young <hanyang.tony@bytedance.com> writes:

>> But as you mentioned converting them would require a database of all
>> tags...
>
> Hardcoding all the locale names in the code should be fine, I guess?

That is exactly the "database" we do not want to have to maintain,
so not fine.

> Though the problem of filtering out locales unsupported by glibc is more
> troublesome.

```

## Collin Funk, 2025-07-11 17:02

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <875xfypsom.fsf@gmail.com>
URL: https://gitlist.dev/e/875xfypsom.fsf%40gmail.com
In-Reply-To: <r34i7fhxwbxhppc4ia7lpyr3xqj4tgusaeikaaonpwtywlywxw@ygfmv3f3q67u>

```
Justin Tobler <jltobler@gmail.com> writes:

> From my understanding, each language is expected to be defined in the
> following form:
>
>   language[_territory][.codeset][@modifier]
>
> When we parse the list of languages we only care about the
> `language[_territory]` part though.
>
> From looking at ISO 639 language codes, only codes with two or three
> characters are valid. If we wanted to be a bit more strict, we could
> check the length of the language code (everything before the first '_')
> and filter out anything outside of those limits. This would naturally
> filter out "C" and "POSIX" without having to mention them explicitly.
>
> Not sure if being more strict adds much more value here in practice
> though. So it may be fine to keep it as-is. :)

Filtering out anything that isn't 2-3 letters seems like a good
heuristic to me.

It seems better than only filtering out "C" and "POSIX" and allowing
anything else. And it keeps us from having to keep a list of updated BCP
47 language tags.

Collin

```

## Junio C Hamano, 2025-07-11 18:32

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <xmqqldou1suk.fsf@gitster.g>
URL: https://gitlist.dev/e/xmqqldou1suk.fsf%40gitster.g
In-Reply-To: <r34i7fhxwbxhppc4ia7lpyr3xqj4tgusaeikaaonpwtywlywxw@ygfmv3f3q67u>

```
Justin Tobler <jltobler@gmail.com> writes:

> Looking at `get_preferred_languages()` when NO_GETTEXT is defined, we
> already filter out "C" and "POSIX". So doing this for the LANGUAGE
> environment variable when writing the header also makes sense.

True.  I wonder if it makes sense to do the check in that helper
function, though.  I.e. something like

diff --git c/gettext.c w/gettext.c
index 8d08a61f84..e2e0fe339d 100644
--- c/gettext.c
+++ w/gettext.c
@@ -41,6 +41,16 @@ static const char *locale_charset(void)
 
 static const char *charset;
 
+static const char *filter_out_non_languages(const char *candidate)
+{
+	if (candidate && *candidate &&
+	    strcmp(candidate, "C") &&
+	    strcmp(candidate, "POSIX"))
+		return candidate;
+	else
+		return NULL;
+}
+
 /*
  * Guess the user's preferred languages from the value in LANGUAGE environment
  * variable and LC_MESSAGES locale category if NO_GETTEXT is not defined.
@@ -51,15 +61,13 @@ const char *get_preferred_languages(void)
 {
 	const char *retval;
 
-	retval = getenv("LANGUAGE");
-	if (retval && *retval)
+	retval = filter_out_non_languages(getenv("LANGUAGE"));
+	if (retval)
 		return retval;
 
 #ifndef NO_GETTEXT
-	retval = setlocale(LC_MESSAGES, NULL);
-	if (retval && *retval &&
-		strcmp(retval, "C") &&
-		strcmp(retval, "POSIX"))
+	retval = filter_out_non_languages(setlocale(LC_MESSAGES, NULL));
+	if (retval)
 		return retval;
 #endif
 

In the production code, we should have a comment before that new
helper function that explains why we exclude C and POSIX, if we were
to go that route.

> Not sure if being more strict adds much more value here in practice
> though. So it may be fine to keep it as-is. :)

Yup.  I care more about having a single place that checks using the
same logic, than what that logic exactly is ;-).

Thanks.

```

## Carlo Marcelo Arenas Belón, 2025-07-11 20:22

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <owlvoi7beap4mx2tejuny26xo4jpzzpxkz7243erlhpgu7oa2u@q7fuycw32ds6>
URL: https://gitlist.dev/e/owlvoi7beap4mx2tejuny26xo4jpzzpxkz7243erlhpgu7oa2u%40q7fuycw32ds6
In-Reply-To: <xmqqldou1suk.fsf@gitster.g>

```
On Fri, Jul 11, 2025 at 11:32:19AM -0800, Junio C Hamano wrote:
> Justin Tobler <jltobler@gmail.com> writes:
> 
> > Looking at `get_preferred_languages()` when NO_GETTEXT is defined, we
> > already filter out "C" and "POSIX". So doing this for the LANGUAGE
> > environment variable when writing the header also makes sense.
> 
> True.  I wonder if it makes sense to do the check in that helper
> function, though.  I.e. something like

Definitely, and might also fix another bug, as IMHO the current logic have
a couple of issues:

* LANGUAGE is not meant to be relevant unless LANG is set to a valid locale
  as per the SPEC[1], allthough for our use case it might be better to still
  do, specially if there are users in the wild setting C and POSIX there.
* it might make more sense to use the union of LANGUAGE and LC_MESSAGES
  instead.

Carlo

```

## Carlo Marcelo Arenas Belón, 2025-07-11 20:57

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <idgdx2au3zgpowozspvu6ttvehybtwwuqf5kwqga4yok7uo2uj@wno7evyjg6pq>
URL: https://gitlist.dev/e/idgdx2au3zgpowozspvu6ttvehybtwwuqf5kwqga4yok7uo2uj%40wno7evyjg6pq
In-Reply-To: <875xfypsom.fsf@gmail.com>

```
On Fri, Jul 11, 2025 at 10:02:01AM -0800, Collin Funk wrote:
> Justin Tobler <jltobler@gmail.com> writes:
> 
> > From my understanding, each language is expected to be defined in the
> > following form:
> >
> >   language[_territory][.codeset][@modifier]
> >
> > When we parse the list of languages we only care about the
> > `language[_territory]` part though.
> >
> > From looking at ISO 639 language codes, only codes with two or three
> > characters are valid. If we wanted to be a bit more strict, we could
> > check the length of the language code (everything before the first '_')
> > and filter out anything outside of those limits. This would naturally
> > filter out "C" and "POSIX" without having to mention them explicitly.
> 
> Filtering out anything that isn't 2-3 letters seems like a good
> heuristic to me.

except that it would be incorrect, as language tags are defined in RFC5646
and are larger than that.

most importantly, deriving language tags from locales provides some very
useful tags when including the characters after the _, because zh_CN and
zh_HK use completely different scripts, for example.

Carlo

```

## brian m. carlson, 2025-07-11 21:29

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <aHGCRLGHEB0m_cXZ@fruit.crustytoothpaste.net>
URL: https://gitlist.dev/e/aHGCRLGHEB0m_cXZ%40fruit.crustytoothpaste.net
In-Reply-To: <idgdx2au3zgpowozspvu6ttvehybtwwuqf5kwqga4yok7uo2uj@wno7evyjg6pq>

```
On 2025-07-11 at 20:57:03, Carlo Marcelo Arenas Belón wrote:
> except that it would be incorrect, as language tags are defined in RFC5646
> and are larger than that.
> 
> most importantly, deriving language tags from locales provides some very
> useful tags when including the characters after the _, because zh_CN and
> zh_HK use completely different scripts, for example.

Yes, that's true.  You have some private use and some irregular tags and
you also have some tags that include scripts or country codes.

For instance, Swahili can be written in Latin or Arabic script.  As I
understand it, the Arabic script form is older and less common these
days, so if I learned Swahili (which I would like to), then I might only
learn the Latin script variant in a course.  I would need to specify
that script in the language code to be sure that I was presented with
content in a form that I could read and understand.  Similar concerns
exist with the variants of Serbo-Croatian: some are written in Latin
scripts, some in Cyrillic, and some in both, and it's not guaranteed
that all speakers understand all forms.

And then there's pt-PT and pt-BR, which are not always mutually
intelligible.  Most free software I've seen ships these as separate
translations.

I don't want to implement language tag parsing here since we don't need
to do that.  I would like to do the simple thing to prevent commonly
used locales that don't represent actual language tags from being
included and not overengineer this design.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

```

## Carlo Arenas, 2025-07-11 22:12

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <CAPUEsphkzaibm2FMBoj-9nbFch7UgRvyvmzErmno0z+2k5X+OA@mail.gmail.com>
URL: https://gitlist.dev/e/CAPUEsphkzaibm2FMBoj-9nbFch7UgRvyvmzErmno0z%2B2k5X%2BOA%40mail.gmail.com
In-Reply-To: <aHGCRLGHEB0m_cXZ@fruit.crustytoothpaste.net>

```
On Fri, Jul 11, 2025 at 2:29 PM brian m. carlson
<sandals@crustytoothpaste.net> wrote:
>
> On 2025-07-11 at 20:57:03, Carlo Marcelo Arenas Belón wrote:
> > except that it would be incorrect, as language tags are defined in RFC5646
> > and are larger than that.
> >
> > most importantly, deriving language tags from locales provides some very
> > useful tags when including the characters after the _, because zh_CN and
> > zh_HK use completely different scripts, for example.
>
> Yes, that's true.  You have some private use and some irregular tags and
> you also have some tags that include scripts or country codes.
>
> For instance, Swahili can be written in Latin or Arabic script.  As I
> understand it, the Arabic script form is older and less common these
> days, so if I learned Swahili (which I would like to), then I might only
> learn the Latin script variant in a course.  I would need to specify
> that script in the language code to be sure that I was presented with
> content in a form that I could read and understand.  Similar concerns
> exist with the variants of Serbo-Croatian: some are written in Latin
> scripts, some in Cyrillic, and some in both, and it's not guaranteed
> that all speakers understand all forms.
>
> And then there's pt-PT and pt-BR, which are not always mutually
> intelligible.  Most free software I've seen ships these as separate
> translations.
>
> I don't want to implement language tag parsing here since we don't need
> to do that.  I would like to do the simple thing to prevent commonly
> used locales that don't represent actual language tags from being
> included and not overengineer this design

I think that your design of filtering C and POSIX accomplishes that,
even if it might seem like hardcoding those two values is a little dirty.

Moving the logic (including the filtering, which is already happening
for the `!NO_GETTEXT `code path adds several chances to modernize
and cleanup the code though which will be beneficial (ex: using and
strvec or even a hashtable to process the candidates, improve
validation and tests)

Carlo

CC Yi EungJun at a hopefully working email address with link to thread
https://lore.kernel.org/git/20250710221641.857081-1-sandals@crustytoothpaste.net/

.
> --
> brian m. carlson (they/them)
> Toronto, Ontario, CA

```

## Collin Funk, 2025-07-11 22:17

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <87cya6xthl.fsf@gmail.com>
URL: https://gitlist.dev/e/87cya6xthl.fsf%40gmail.com
In-Reply-To: <CAPUEsphkzaibm2FMBoj-9nbFch7UgRvyvmzErmno0z+2k5X+OA@mail.gmail.com>

```
Carlo Arenas <carenas@gmail.com> writes:

>> I don't want to implement language tag parsing here since we don't need
>> to do that.  I would like to do the simple thing to prevent commonly
>> used locales that don't represent actual language tags from being
>> included and not overengineer this design
>
> I think that your design of filtering C and POSIX accomplishes that,
> even if it might seem like hardcoding those two values is a little dirty.

Thanks for correcting me regarding language tag length Carlo.

I guess I am fine with filtering out "C" and "POSIX" now. Not perfect,
but I think everyone agrees that we don't want to maintain a database of
language tags just for this.

Collin

```

## Eli Schwartz, 2025-07-15 04:38

Subject: Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language
Message-ID: <77879bc4-9b8b-4f8f-a9d5-ea0114937e9b@gentoo.org>
URL: https://gitlist.dev/e/77879bc4-9b8b-4f8f-a9d5-ea0114937e9b%40gentoo.org
In-Reply-To: <20250710221641.857081-2-sandals@crustytoothpaste.net>

```
On 7/10/25 6:16 PM, brian m. carlson wrote:
> The LANGUAGE environment variable is not specified by POSIX, but a
> variety of programs using GNU gettext accept it.  The Linux manpages
> state that it can contain a colon-separated list of locales.
> 
> However, not all locales are valid as languages.  The C and POSIX
> locales, for instance, are not languages and are not registered with
> IANA, nor are they a part of ISO 639.  In fact, "C" is too short to
> match the ABNF production for a language, which must be at least two
> characters in length.
> 
> Nonetheless, many users provide these values in the LANGUAGE environment
> variable for unknown reasons and if they do, we do not want to send a
> malformed Accept-Language header to the server.  If there are no other
> valid language tags, then send no header; otherwise, send only the valid
> tags, ignoring "C" and "POSIX" wherever they may appear, as well as any
> variants (such as the "C.UTF-8" locale found on some Linux systems).


Better docs -- the gettext manpages suck:
https://www.gnu.org/software/gettext/manual/html_node/Locale-Names.html
https://www.gnu.org/software/gettext/manual/html_node/The-LANGUAGE-variable.html


At minimum this commit message needs revising. Gettext was adopted into
POSIX 2024 (Issue 8).


Respected by tools of course:
https://pubs.opengroup.org/onlinepubs/9799919799/utilities/gettext.html#tag_20_54_08
https://pubs.opengroup.org/onlinepubs/9799919799/functions/gettext.html

$LANGUAGE docs can be found at

https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/V1_chap08.html#tag_08_02


"""
The value of LANGUAGE shall be a list of locale names separated by a
<colon> (':') character. If LANGUAGE is set to a non-empty string, each
locale name shall be tried in the specified order and if a messages
object is found, it shall be used for translation. If a locale name has
the format language[_territory][.codeset][@modifier], additional
searches of locale names without .codeset (if present), without
_territory (if present), and without @modifier (if present) may be performed
"""


And, for locale name values,

"""
If the locale value is "C" or "POSIX", the POSIX locale shall be used
and the standard utilities behave in accordance with the rules in 7.2
POSIX Locale for the associated category.

If the locale value begins with a <slash>, it shall be interpreted as
the pathname of a file that was created in the output format used by the
localedef utility; see OUTPUT FILES under localedef. Referencing such a
pathname shall result in that locale being used for the indicated category.

[XSI] [Option Start] If the locale value has the form:

language[_territory][.codeset]

it refers to an implementation-provided locale, where settings of
language, territory, and codeset are implementation-defined.

LC_COLLATE , LC_CTYPE , LC_MESSAGES , LC_MONETARY , LC_NUMERIC , and
LC_TIME are defined to accept an additional field @modifier, which
allows the user to select a specific instance of localization data
within a single category (for example, for selecting the dictionary as
opposed to the character ordering of data). The syntax for these
environment variables is thus defined as:

[language[_territory][.codeset][@modifier]]
"""


Your tests and code are probably broken -- they appear to normalize
nearly none of the standard grammar into valid Accept-Language entries.
Of course, "surely nobody actually does that" (except when they do!) --
but it's a relatively simple grammar structure, simply getting the
"shape" correct seems like a good idea.


-- 
Eli Schwartz

```
