{"thread":{"id":"63780","subject":"[PATCH 0/1] Filter C and POSIX out of Accept-Language","startedAt":"2025-07-10T22:16:58Z","lastAt":"2025-07-15T04:38:24Z","messageCount":18,"participants":["brian m. carlson","Junio C Hamano","Collin Funk","Han Young","Justin Tobler","Carlo Marcelo Arenas Belón","Carlo Arenas","Eli Schwartz"],"isPatch":true,"patchVersion":1,"patchTotal":1},"messages":[{"id":"521784","messageId":"20250710221641.857081-1-sandals@crustytoothpaste.net","threadId":"63780","inReplyTo":null,"subject":"[PATCH 0/1] Filter C and POSIX out of Accept-Language","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2025-07-10T22:16:40Z","receivedAt":"2025-07-10T22:16:58Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"At work, I've seen some cases where people provide \"C\" in the\nAccept-Language header of their Git requests, such as when they provide\nus with debugging traces, but \"C\" and \"POSIX\", while valid locales, are\nnot valid languages and do not belong in the Accept-Language header.\n\nIt turns out this is actually very easy to reproduce and fix, so there's\na patch to filter these out.  I have not actually myself seen \"POSIX\" in\nthe header, but it's equivalent to \"C\" and I've seen it in non-Git\nrequests in various places online, so we reject that as well.\n\nThis can be seen in GitLab's issues as well at\nhttps://gitlab.com/gitlab-org/gitlab/-/issues/412077.\n\nbrian m. carlson (1):\n  http: don't send C or POSIX in Accept-Language\n\n http.c                     |  8 ++++++++\n t/t5541-http-push-smart.sh | 18 ++++++++++++++++++\n 2 files changed, 26 insertions(+)\n\n"},{"id":"521785","messageId":"20250710221641.857081-2-sandals@crustytoothpaste.net","threadId":"63780","inReplyTo":"20250710221641.857081-1-sandals@crustytoothpaste.net","subject":"[PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2025-07-10T22:16:41Z","receivedAt":"2025-07-10T22:16:58Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"The LANGUAGE environment variable is not specified by POSIX, but a\nvariety of programs using GNU gettext accept it.  The Linux manpages\nstate that it can contain a colon-separated list of locales.\n\nHowever, not all locales are valid as languages.  The C and POSIX\nlocales, for instance, are not languages and are not registered with\nIANA, nor are they a part of ISO 639.  In fact, \"C\" is too short to\nmatch the ABNF production for a language, which must be at least two\ncharacters in length.\n\nNonetheless, many users provide these values in the LANGUAGE environment\nvariable for unknown reasons and if they do, we do not want to send a\nmalformed Accept-Language header to the server.  If there are no other\nvalid language tags, then send no header; otherwise, send only the valid\ntags, ignoring \"C\" and \"POSIX\" wherever they may appear, as well as any\nvariants (such as the \"C.UTF-8\" locale found on some Linux systems).\n\nWe do not reject all possible invalid language tags since doing so\nwould require bundling a copy of the IANA database and would risk poor\nbehavior in the face of uncommon languages or values that are not\nregistered but meet the production for private use or other restricted\ninterchange.  However, these two values are widely used in the LANGUAGE\nheader, are well-known and widely used non-language locales, and have\nbeen seen in the wild on the server side.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n http.c                     |  8 ++++++++\n t/t5541-http-push-smart.sh | 18 ++++++++++++++++++\n 2 files changed, 26 insertions(+)\n\ndiff --git a/http.c b/http.c\nindex d88e79fbde..a96df4fcdb 100644\n--- a/http.c\n+++ b/http.c\n@@ -2022,6 +2022,14 @@ static void write_accept_language(struct strbuf *buf)\n \t\t\ts++;\n \n \t\tif (tag.len) {\n+\t\t\t/*\n+\t\t\t * These are not valid languages: do not send them to\n+\t\t\t * the server.\n+\t\t\t */\n+\t\t\tif (!strcmp(tag.buf, \"C\") || !strcmp(tag.buf, \"POSIX\")) {\n+\t\t\t\tstrbuf_reset(&tag);\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tnum_langs++;\n \t\t\tREALLOC_ARRAY(language_tags, num_langs);\n \t\t\tlanguage_tags[num_langs - 1] = strbuf_detach(&tag, NULL);\ndiff --git a/t/t5541-http-push-smart.sh b/t/t5541-http-push-smart.sh\nindex 538b603f03..96a6833e67 100755\n--- a/t/t5541-http-push-smart.sh\n+++ b/t/t5541-http-push-smart.sh\n@@ -86,6 +86,24 @@ test_expect_success 'push to remote repository (standard) with sending Accept-La\n \tGIT_TRACE_CURL=true LANGUAGE=\"ko_KR.UTF-8\" git push -v -v 2>err &&\n \t! grep \"Expect: 100-continue\" err &&\n \n+\tgrep \"=> Send header: Accept-Language:\" err >err.language &&\n+\ttest_cmp exp err.language &&\n+\n+\ttest_commit C-is-not-a-language &&\n+\tGIT_TRACE_CURL=true LANGUAGE=\"C\" git push -v -v 2>err &&\n+\n+\t! grep \"=> Send header: Accept-Language:\" err >err.language &&\n+\ttest_must_be_empty err.language &&\n+\n+\ttest_commit POSIX-is-not-a-language-either &&\n+\tGIT_TRACE_CURL=true LANGUAGE=\"POSIX\" git push -v -v 2>err &&\n+\n+\t! grep \"=> Send header: Accept-Language:\" err >err.language &&\n+\ttest_must_be_empty err.language &&\n+\n+\ttest_commit ignore-C-and-POSIX-as-languages-wherever-provided &&\n+\tGIT_TRACE_CURL=true LANGUAGE=\"C.UTF-8:ko_KR.UTF-8:POSIX\" git push -v -v 2>err &&\n+\n \tgrep \"=> Send header: Accept-Language:\" err >err.language &&\n \ttest_cmp exp err.language\n '\n"},{"id":"521789","messageId":"xmqqfrf34qdb.fsf@gitster.g","threadId":"63780","inReplyTo":"20250710221641.857081-1-sandals@crustytoothpaste.net","subject":"Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-10T22:45:20Z","receivedAt":"2025-07-10T22:45:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> At work, I've seen some cases where people provide \"C\" in the\n> Accept-Language header of their Git requests, such as when they provide\n> us with debugging traces, but \"C\" and \"POSIX\", while valid locales, are\n> not valid languages and do not belong in the Accept-Language header.\n>\n> It turns out this is actually very easy to reproduce and fix, so there's\n> a patch to filter these out.  I have not actually myself seen \"POSIX\" in\n> the header, but it's equivalent to \"C\" and I've seen it in non-Git\n> requests in various places online, so we reject that as well.\n>\n> This can be seen in GitLab's issues as well at\n> https://gitlab.com/gitlab-org/gitlab/-/issues/412077.\n\nSorry, I am confused.  Is that Authentication failure in the cited\nissue \"caused by\" the client sending \"Accept-Language: C\"?\n\n\"reproduce and fix\" makes it sound like a correct exchange between\nsuch a client and a server is somehow broken (i.e. unable to clone,\nunable to authenticate, etc.) if the client sends C (or POSIX) as if\nit were a langauge, but is there a breakage there?\n\nI understand and agree with the change in patch 1/1 that it is the\nright thing to do (to more strictly adhere to the standard in what\nwe send out) for hygiene.  I just want to understand if this caused\nreal problems, or if it is primarily a preemptive clean-up to avoid\nnon-standard behaviour causing problems in the future.\n\nThanks.\n\n> brian m. carlson (1):\n>   http: don't send C or POSIX in Accept-Language\n>\n>  http.c                     |  8 ++++++++\n>  t/t5541-http-push-smart.sh | 18 ++++++++++++++++++\n>  2 files changed, 26 insertions(+)\n"},{"id":"521790","messageId":"xmqqbjpr4q8z.fsf@gitster.g","threadId":"63780","inReplyTo":"20250710221641.857081-2-sandals@crustytoothpaste.net","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-10T22:47:56Z","receivedAt":"2025-07-10T22:47:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> ...  However, these two values are widely used in the LANGUAGE\n> header, are well-known and widely used non-language locales, and have\n> been seen in the wild on the server side.\n\n\"header\" -> \"environment variable\" I presume?  \n\n"},{"id":"521793","messageId":"aHBH0nRLPxBg2HAj@fruit.crustytoothpaste.net","threadId":"63780","inReplyTo":"xmqqfrf34qdb.fsf@gitster.g","subject":"Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2025-07-10T23:08:02Z","receivedAt":"2025-07-10T23:08:10Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"[Dropping Yi EungJun from CC because their email bounced.]\n\nOn 2025-07-10 at 22:45:20, Junio C Hamano wrote:\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > At work, I've seen some cases where people provide \"C\" in the\n> > Accept-Language header of their Git requests, such as when they provide\n> > us with debugging traces, but \"C\" and \"POSIX\", while valid locales, are\n> > not valid languages and do not belong in the Accept-Language header.\n> >\n> > It turns out this is actually very easy to reproduce and fix, so there's\n> > a patch to filter these out.  I have not actually myself seen \"POSIX\" in\n> > the header, but it's equivalent to \"C\" and I've seen it in non-Git\n> > requests in various places online, so we reject that as well.\n> >\n> > This can be seen in GitLab's issues as well at\n> > https://gitlab.com/gitlab-org/gitlab/-/issues/412077.\n> \n> Sorry, I am confused.  Is that Authentication failure in the cited\n> issue \"caused by\" the client sending \"Accept-Language: C\"?\n> \n> \"reproduce and fix\" makes it sound like a correct exchange between\n> such a client and a server is somehow broken (i.e. unable to clone,\n> unable to authenticate, etc.) if the client sends C (or POSIX) as if\n> it were a langauge, but is there a breakage there?\n\nNo, sorry.  I just meant that the trace in that issue demonstrates the\nincorrect Accept-Language header; it's unrelated to the authentication\nproblem that the issue is about (which I think is a GitLab issue).\n\n> I understand and agree with the change in patch 1/1 that it is the\n> right thing to do (to more strictly adhere to the standard in what\n> we send out) for hygiene.  I just want to understand if this caused\n> real problems, or if it is primarily a preemptive clean-up to avoid\n> non-standard behaviour causing problems in the future.\n\nI'm not aware of it causing any practical problems for people, although\nI could imagine some cases where it could, in theory, break things.  I\nmerely noticed this in trace output and thought we should tidy it up.\nIf users are using the header and expecting a localized response, this\nwill make it more likely that they get the one they were expecting.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"521794","messageId":"aHBH-3io3rqw4pK4@fruit.crustytoothpaste.net","threadId":"63780","inReplyTo":"xmqqbjpr4q8z.fsf@gitster.g","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2025-07-10T23:08:43Z","receivedAt":"2025-07-10T23:08:45Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2025-07-10 at 22:47:56, Junio C Hamano wrote:\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > ...  However, these two values are widely used in the LANGUAGE\n> > header, are well-known and widely used non-language locales, and have\n> > been seen in the wild on the server side.\n> \n> \"header\" -> \"environment variable\" I presume?  \n\nAh, yes.  I'll fix that for a v2 in a few days while I wait for other\ncomments.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"521801","messageId":"m1h5zjk4pv.fsf@gmail.com","threadId":"63780","inReplyTo":"aHBH0nRLPxBg2HAj@fruit.crustytoothpaste.net","subject":"Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language","fromName":"Collin Funk","fromEmail":"collin.funk1@gmail.com","sentAt":"2025-07-10T23:26:20Z","receivedAt":"2025-07-10T23:26:22Z","isPatch":true,"sender":{"key":"collin.funk1@gmail.com","avatar":"https://avatars.githubusercontent.com/u/65689063?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> I'm not aware of it causing any practical problems for people, although\n> I could imagine some cases where it could, in theory, break things.  I\n> merely noticed this in trace output and thought we should tidy it up.\n> If users are using the header and expecting a localized response, this\n> will make it more likely that they get the one they were expecting.\n\nI feel like it is a bit strange to only exclude \"C\" or \"POSIX\".\n\nI think the correct behavior would be to accept any values, or convert\nthe current locale to the closest BCP 47 language tag.\n\nBut as you mentioned converting them would require a database of all\ntags...\n\nCollin\n"},{"id":"521807","messageId":"CAG1j3zGn5fS=_Oftu7bBmWsoMc-aCa84AtDXdfxgL8QFEkp+yA@mail.gmail.com","threadId":"63780","inReplyTo":"m1h5zjk4pv.fsf@gmail.com","subject":"Re: [External] Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language","fromName":"Han Young","fromEmail":"hanyang.tony@bytedance.com","sentAt":"2025-07-11T02:49:03Z","receivedAt":"2025-07-11T02:49:15Z","isPatch":true,"sender":{"key":"hanyang.tony@bytedance.com","avatar":"https://avatars.githubusercontent.com/u/108711387?v=4"},"body":"On Fri, Jul 11, 2025 at 7:26 AM Collin Funk <collin.funk1@gmail.com> wrote:\n\n> I think the correct behavior would be to accept any values, or convert\n> the current locale to the closest BCP 47 language tag.\n\nOn some Linux systems, not all BCP languages are supported. Not all\nLinux distributions generate all the locales, and musl doesn't even support\nlocales. Converting to the closest BCP 47 language alone does not\nensure the locale is valid. Not to mention the tricky heuristics of language\nmatching (pt_PT or pt_BR if LANGUAGE is pt?).\n\n> But as you mentioned converting them would require a database of all\n> tags...\n\nHardcoding all the locale names in the code should be fine, I guess?\nThough the problem of filtering out locales unsupported by glibc is more\ntroublesome.\n"},{"id":"521823","messageId":"r34i7fhxwbxhppc4ia7lpyr3xqj4tgusaeikaaonpwtywlywxw@ygfmv3f3q67u","threadId":"63780","inReplyTo":"20250710221641.857081-2-sandals@crustytoothpaste.net","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-07-11T15:23:38Z","receivedAt":"2025-07-11T15:29:17Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/07/10 10:16PM, brian m. carlson wrote:\n> The LANGUAGE environment variable is not specified by POSIX, but a\n> variety of programs using GNU gettext accept it.  The Linux manpages\n> state that it can contain a colon-separated list of locales.\n> \n> However, not all locales are valid as languages.  The C and POSIX\n> locales, for instance, are not languages and are not registered with\n> IANA, nor are they a part of ISO 639.  In fact, \"C\" is too short to\n> match the ABNF production for a language, which must be at least two\n> characters in length.\n> \n> Nonetheless, many users provide these values in the LANGUAGE environment\n> variable for unknown reasons and if they do, we do not want to send a\n> malformed Accept-Language header to the server.  If there are no other\n> valid language tags, then send no header; otherwise, send only the valid\n> tags, ignoring \"C\" and \"POSIX\" wherever they may appear, as well as any\n> variants (such as the \"C.UTF-8\" locale found on some Linux systems).\n\nOk so the languages returned by `get_preferred_languages()` are used to\nwrite the Accept-Language header when making requests.\n\nLooking at `get_preferred_languages()` when NO_GETTEXT is defined, we\nalready filter out \"C\" and \"POSIX\". So doing this for the LANGUAGE\nenvironment variable when writing the header also makes sense.\n\n> We do not reject all possible invalid language tags since doing so\n> would require bundling a copy of the IANA database and would risk poor\n> behavior in the face of uncommon languages or values that are not\n> registered but meet the production for private use or other restricted\n> interchange.  However, these two values are widely used in the LANGUAGE\n> header, are well-known and widely used non-language locales, and have\n> been seen in the wild on the server side.\n> \n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> ---\n>  http.c                     |  8 ++++++++\n>  t/t5541-http-push-smart.sh | 18 ++++++++++++++++++\n>  2 files changed, 26 insertions(+)\n> \n> diff --git a/http.c b/http.c\n> index d88e79fbde..a96df4fcdb 100644\n> --- a/http.c\n> +++ b/http.c\n> @@ -2022,6 +2022,14 @@ static void write_accept_language(struct strbuf *buf)\n>  \t\t\ts++;\n>  \n>  \t\tif (tag.len) {\n> +\t\t\t/*\n> +\t\t\t * These are not valid languages: do not send them to\n> +\t\t\t * the server.\n> +\t\t\t */\n> +\t\t\tif (!strcmp(tag.buf, \"C\") || !strcmp(tag.buf, \"POSIX\")) {\n> +\t\t\t\tstrbuf_reset(&tag);\n> +\t\t\t\tcontinue;\n> +\t\t\t}\n\nFrom my understanding, each language is expected to be defined in the\nfollowing form:\n\n  language[_territory][.codeset][@modifier]\n\nWhen we parse the list of languages we only care about the\n`language[_territory]` part though.\n\nFrom looking at ISO 639 language codes, only codes with two or three\ncharacters are valid. If we wanted to be a bit more strict, we could\ncheck the length of the language code (everything before the first '_')\nand filter out anything outside of those limits. This would naturally\nfilter out \"C\" and \"POSIX\" without having to mention them explicitly.\n\nNot sure if being more strict adds much more value here in practice\nthough. So it may be fine to keep it as-is. :)\n\n>  \t\t\tnum_langs++;\n>  \t\t\tREALLOC_ARRAY(language_tags, num_langs);\n>  \t\t\tlanguage_tags[num_langs - 1] = strbuf_detach(&tag, NULL);\n> diff --git a/t/t5541-http-push-smart.sh b/t/t5541-http-push-smart.sh\n> index 538b603f03..96a6833e67 100755\n> --- a/t/t5541-http-push-smart.sh\n> +++ b/t/t5541-http-push-smart.sh\n> @@ -86,6 +86,24 @@ test_expect_success 'push to remote repository (standard) with sending Accept-La\n>  \tGIT_TRACE_CURL=true LANGUAGE=\"ko_KR.UTF-8\" git push -v -v 2>err &&\n>  \t! grep \"Expect: 100-continue\" err &&\n>  \n> +\tgrep \"=> Send header: Accept-Language:\" err >err.language &&\n> +\ttest_cmp exp err.language &&\n> +\n> +\ttest_commit C-is-not-a-language &&\n> +\tGIT_TRACE_CURL=true LANGUAGE=\"C\" git push -v -v 2>err &&\n> +\n> +\t! grep \"=> Send header: Accept-Language:\" err >err.language &&\n> +\ttest_must_be_empty err.language &&\n> +\n> +\ttest_commit POSIX-is-not-a-language-either &&\n> +\tGIT_TRACE_CURL=true LANGUAGE=\"POSIX\" git push -v -v 2>err &&\n> +\n> +\t! grep \"=> Send header: Accept-Language:\" err >err.language &&\n> +\ttest_must_be_empty err.language &&\n\nThe above two tests demonstrate that the Accept-Language header is not\nsent if no valid languages are found.\n\n> +\n> +\ttest_commit ignore-C-and-POSIX-as-languages-wherever-provided &&\n> +\tGIT_TRACE_CURL=true LANGUAGE=\"C.UTF-8:ko_KR.UTF-8:POSIX\" git push -v -v 2>err &&\n> +\n>  \tgrep \"=> Send header: Accept-Language:\" err >err.language &&\n>  \ttest_cmp exp err.language\n>  '\n\nAnd here we see only the valid languages sent in the header. Looks good!\n\n-Justin\n"},{"id":"521832","messageId":"xmqq5xfy3cbb.fsf@gitster.g","threadId":"63780","inReplyTo":"CAG1j3zGn5fS=_Oftu7bBmWsoMc-aCa84AtDXdfxgL8QFEkp+yA@mail.gmail.com","subject":"Re: [External] Re: [PATCH 0/1] Filter C and POSIX out of Accept-Language","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-11T16:46:32Z","receivedAt":"2025-07-11T16:46:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Han Young <hanyang.tony@bytedance.com> writes:\n\n>> But as you mentioned converting them would require a database of all\n>> tags...\n>\n> Hardcoding all the locale names in the code should be fine, I guess?\n\nThat is exactly the \"database\" we do not want to have to maintain,\nso not fine.\n\n> Though the problem of filtering out locales unsupported by glibc is more\n> troublesome.\n"},{"id":"521833","messageId":"875xfypsom.fsf@gmail.com","threadId":"63780","inReplyTo":"r34i7fhxwbxhppc4ia7lpyr3xqj4tgusaeikaaonpwtywlywxw@ygfmv3f3q67u","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Collin Funk","fromEmail":"collin.funk1@gmail.com","sentAt":"2025-07-11T17:02:01Z","receivedAt":"2025-07-11T17:02:05Z","isPatch":true,"sender":{"key":"collin.funk1@gmail.com","avatar":"https://avatars.githubusercontent.com/u/65689063?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> From my understanding, each language is expected to be defined in the\n> following form:\n>\n>   language[_territory][.codeset][@modifier]\n>\n> When we parse the list of languages we only care about the\n> `language[_territory]` part though.\n>\n> From looking at ISO 639 language codes, only codes with two or three\n> characters are valid. If we wanted to be a bit more strict, we could\n> check the length of the language code (everything before the first '_')\n> and filter out anything outside of those limits. This would naturally\n> filter out \"C\" and \"POSIX\" without having to mention them explicitly.\n>\n> Not sure if being more strict adds much more value here in practice\n> though. So it may be fine to keep it as-is. :)\n\nFiltering out anything that isn't 2-3 letters seems like a good\nheuristic to me.\n\nIt seems better than only filtering out \"C\" and \"POSIX\" and allowing\nanything else. And it keeps us from having to keep a list of updated BCP\n47 language tags.\n\nCollin\n"},{"id":"521838","messageId":"xmqqldou1suk.fsf@gitster.g","threadId":"63780","inReplyTo":"r34i7fhxwbxhppc4ia7lpyr3xqj4tgusaeikaaonpwtywlywxw@ygfmv3f3q67u","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-11T18:32:19Z","receivedAt":"2025-07-11T18:32:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> Looking at `get_preferred_languages()` when NO_GETTEXT is defined, we\n> already filter out \"C\" and \"POSIX\". So doing this for the LANGUAGE\n> environment variable when writing the header also makes sense.\n\nTrue.  I wonder if it makes sense to do the check in that helper\nfunction, though.  I.e. something like\n\ndiff --git c/gettext.c w/gettext.c\nindex 8d08a61f84..e2e0fe339d 100644\n--- c/gettext.c\n+++ w/gettext.c\n@@ -41,6 +41,16 @@ static const char *locale_charset(void)\n \n static const char *charset;\n \n+static const char *filter_out_non_languages(const char *candidate)\n+{\n+\tif (candidate && *candidate &&\n+\t    strcmp(candidate, \"C\") &&\n+\t    strcmp(candidate, \"POSIX\"))\n+\t\treturn candidate;\n+\telse\n+\t\treturn NULL;\n+}\n+\n /*\n  * Guess the user's preferred languages from the value in LANGUAGE environment\n  * variable and LC_MESSAGES locale category if NO_GETTEXT is not defined.\n@@ -51,15 +61,13 @@ const char *get_preferred_languages(void)\n {\n \tconst char *retval;\n \n-\tretval = getenv(\"LANGUAGE\");\n-\tif (retval && *retval)\n+\tretval = filter_out_non_languages(getenv(\"LANGUAGE\"));\n+\tif (retval)\n \t\treturn retval;\n \n #ifndef NO_GETTEXT\n-\tretval = setlocale(LC_MESSAGES, NULL);\n-\tif (retval && *retval &&\n-\t\tstrcmp(retval, \"C\") &&\n-\t\tstrcmp(retval, \"POSIX\"))\n+\tretval = filter_out_non_languages(setlocale(LC_MESSAGES, NULL));\n+\tif (retval)\n \t\treturn retval;\n #endif\n \n\nIn the production code, we should have a comment before that new\nhelper function that explains why we exclude C and POSIX, if we were\nto go that route.\n\n> Not sure if being more strict adds much more value here in practice\n> though. So it may be fine to keep it as-is. :)\n\nYup.  I care more about having a single place that checks using the\nsame logic, than what that logic exactly is ;-).\n\nThanks.\n"},{"id":"521845","messageId":"owlvoi7beap4mx2tejuny26xo4jpzzpxkz7243erlhpgu7oa2u@q7fuycw32ds6","threadId":"63780","inReplyTo":"xmqqldou1suk.fsf@gitster.g","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Carlo Marcelo Arenas Belón","fromEmail":"carenas@gmail.com","sentAt":"2025-07-11T20:22:42Z","receivedAt":"2025-07-11T20:22:45Z","isPatch":true,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Fri, Jul 11, 2025 at 11:32:19AM -0800, Junio C Hamano wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > Looking at `get_preferred_languages()` when NO_GETTEXT is defined, we\n> > already filter out \"C\" and \"POSIX\". So doing this for the LANGUAGE\n> > environment variable when writing the header also makes sense.\n> \n> True.  I wonder if it makes sense to do the check in that helper\n> function, though.  I.e. something like\n\nDefinitely, and might also fix another bug, as IMHO the current logic have\na couple of issues:\n\n* LANGUAGE is not meant to be relevant unless LANG is set to a valid locale\n  as per the SPEC[1], allthough for our use case it might be better to still\n  do, specially if there are users in the wild setting C and POSIX there.\n* it might make more sense to use the union of LANGUAGE and LC_MESSAGES\n  instead.\n\nCarlo\n"},{"id":"521847","messageId":"idgdx2au3zgpowozspvu6ttvehybtwwuqf5kwqga4yok7uo2uj@wno7evyjg6pq","threadId":"63780","inReplyTo":"875xfypsom.fsf@gmail.com","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Carlo Marcelo Arenas Belón","fromEmail":"carenas@gmail.com","sentAt":"2025-07-11T20:57:03Z","receivedAt":"2025-07-11T20:57:05Z","isPatch":true,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Fri, Jul 11, 2025 at 10:02:01AM -0800, Collin Funk wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > From my understanding, each language is expected to be defined in the\n> > following form:\n> >\n> >   language[_territory][.codeset][@modifier]\n> >\n> > When we parse the list of languages we only care about the\n> > `language[_territory]` part though.\n> >\n> > From looking at ISO 639 language codes, only codes with two or three\n> > characters are valid. If we wanted to be a bit more strict, we could\n> > check the length of the language code (everything before the first '_')\n> > and filter out anything outside of those limits. This would naturally\n> > filter out \"C\" and \"POSIX\" without having to mention them explicitly.\n> \n> Filtering out anything that isn't 2-3 letters seems like a good\n> heuristic to me.\n\nexcept that it would be incorrect, as language tags are defined in RFC5646\nand are larger than that.\n\nmost importantly, deriving language tags from locales provides some very\nuseful tags when including the characters after the _, because zh_CN and\nzh_HK use completely different scripts, for example.\n\nCarlo\n"},{"id":"521850","messageId":"aHGCRLGHEB0m_cXZ@fruit.crustytoothpaste.net","threadId":"63780","inReplyTo":"idgdx2au3zgpowozspvu6ttvehybtwwuqf5kwqga4yok7uo2uj@wno7evyjg6pq","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2025-07-11T21:29:40Z","receivedAt":"2025-07-11T21:29:42Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2025-07-11 at 20:57:03, Carlo Marcelo Arenas Belón wrote:\n> except that it would be incorrect, as language tags are defined in RFC5646\n> and are larger than that.\n> \n> most importantly, deriving language tags from locales provides some very\n> useful tags when including the characters after the _, because zh_CN and\n> zh_HK use completely different scripts, for example.\n\nYes, that's true.  You have some private use and some irregular tags and\nyou also have some tags that include scripts or country codes.\n\nFor instance, Swahili can be written in Latin or Arabic script.  As I\nunderstand it, the Arabic script form is older and less common these\ndays, so if I learned Swahili (which I would like to), then I might only\nlearn the Latin script variant in a course.  I would need to specify\nthat script in the language code to be sure that I was presented with\ncontent in a form that I could read and understand.  Similar concerns\nexist with the variants of Serbo-Croatian: some are written in Latin\nscripts, some in Cyrillic, and some in both, and it's not guaranteed\nthat all speakers understand all forms.\n\nAnd then there's pt-PT and pt-BR, which are not always mutually\nintelligible.  Most free software I've seen ships these as separate\ntranslations.\n\nI don't want to implement language tag parsing here since we don't need\nto do that.  I would like to do the simple thing to prevent commonly\nused locales that don't represent actual language tags from being\nincluded and not overengineer this design.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"521852","messageId":"CAPUEsphkzaibm2FMBoj-9nbFch7UgRvyvmzErmno0z+2k5X+OA@mail.gmail.com","threadId":"63780","inReplyTo":"aHGCRLGHEB0m_cXZ@fruit.crustytoothpaste.net","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Carlo Arenas","fromEmail":"carenas@gmail.com","sentAt":"2025-07-11T22:12:01Z","receivedAt":"2025-07-11T22:12:14Z","isPatch":true,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Fri, Jul 11, 2025 at 2:29 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> On 2025-07-11 at 20:57:03, Carlo Marcelo Arenas Belón wrote:\n> > except that it would be incorrect, as language tags are defined in RFC5646\n> > and are larger than that.\n> >\n> > most importantly, deriving language tags from locales provides some very\n> > useful tags when including the characters after the _, because zh_CN and\n> > zh_HK use completely different scripts, for example.\n>\n> Yes, that's true.  You have some private use and some irregular tags and\n> you also have some tags that include scripts or country codes.\n>\n> For instance, Swahili can be written in Latin or Arabic script.  As I\n> understand it, the Arabic script form is older and less common these\n> days, so if I learned Swahili (which I would like to), then I might only\n> learn the Latin script variant in a course.  I would need to specify\n> that script in the language code to be sure that I was presented with\n> content in a form that I could read and understand.  Similar concerns\n> exist with the variants of Serbo-Croatian: some are written in Latin\n> scripts, some in Cyrillic, and some in both, and it's not guaranteed\n> that all speakers understand all forms.\n>\n> And then there's pt-PT and pt-BR, which are not always mutually\n> intelligible.  Most free software I've seen ships these as separate\n> translations.\n>\n> I don't want to implement language tag parsing here since we don't need\n> to do that.  I would like to do the simple thing to prevent commonly\n> used locales that don't represent actual language tags from being\n> included and not overengineer this design\n\nI think that your design of filtering C and POSIX accomplishes that,\neven if it might seem like hardcoding those two values is a little dirty.\n\nMoving the logic (including the filtering, which is already happening\nfor the `!NO_GETTEXT `code path adds several chances to modernize\nand cleanup the code though which will be beneficial (ex: using and\nstrvec or even a hashtable to process the candidates, improve\nvalidation and tests)\n\nCarlo\n\nCC Yi EungJun at a hopefully working email address with link to thread\nhttps://lore.kernel.org/git/20250710221641.857081-1-sandals@crustytoothpaste.net/\n\n.\n> --\n> brian m. carlson (they/them)\n> Toronto, Ontario, CA\n"},{"id":"521853","messageId":"87cya6xthl.fsf@gmail.com","threadId":"63780","inReplyTo":"CAPUEsphkzaibm2FMBoj-9nbFch7UgRvyvmzErmno0z+2k5X+OA@mail.gmail.com","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Collin Funk","fromEmail":"collin.funk1@gmail.com","sentAt":"2025-07-11T22:17:26Z","receivedAt":"2025-07-11T22:17:28Z","isPatch":true,"sender":{"key":"collin.funk1@gmail.com","avatar":"https://avatars.githubusercontent.com/u/65689063?v=4"},"body":"Carlo Arenas <carenas@gmail.com> writes:\n\n>> I don't want to implement language tag parsing here since we don't need\n>> to do that.  I would like to do the simple thing to prevent commonly\n>> used locales that don't represent actual language tags from being\n>> included and not overengineer this design\n>\n> I think that your design of filtering C and POSIX accomplishes that,\n> even if it might seem like hardcoding those two values is a little dirty.\n\nThanks for correcting me regarding language tag length Carlo.\n\nI guess I am fine with filtering out \"C\" and \"POSIX\" now. Not perfect,\nbut I think everyone agrees that we don't want to maintain a database of\nlanguage tags just for this.\n\nCollin\n"},{"id":"521938","messageId":"77879bc4-9b8b-4f8f-a9d5-ea0114937e9b@gentoo.org","threadId":"63780","inReplyTo":"20250710221641.857081-2-sandals@crustytoothpaste.net","subject":"Re: [PATCH 1/1] http: don't send C or POSIX in Accept-Language","fromName":"Eli Schwartz","fromEmail":"eschwartz@gentoo.org","sentAt":"2025-07-15T04:38:19Z","receivedAt":"2025-07-15T04:38:24Z","isPatch":true,"sender":{"key":"eschwartz@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/6551424?v=4"},"body":"On 7/10/25 6:16 PM, brian m. carlson wrote:\n> The LANGUAGE environment variable is not specified by POSIX, but a\n> variety of programs using GNU gettext accept it.  The Linux manpages\n> state that it can contain a colon-separated list of locales.\n> \n> However, not all locales are valid as languages.  The C and POSIX\n> locales, for instance, are not languages and are not registered with\n> IANA, nor are they a part of ISO 639.  In fact, \"C\" is too short to\n> match the ABNF production for a language, which must be at least two\n> characters in length.\n> \n> Nonetheless, many users provide these values in the LANGUAGE environment\n> variable for unknown reasons and if they do, we do not want to send a\n> malformed Accept-Language header to the server.  If there are no other\n> valid language tags, then send no header; otherwise, send only the valid\n> tags, ignoring \"C\" and \"POSIX\" wherever they may appear, as well as any\n> variants (such as the \"C.UTF-8\" locale found on some Linux systems).\n\n\nBetter docs -- the gettext manpages suck:\nhttps://www.gnu.org/software/gettext/manual/html_node/Locale-Names.html\nhttps://www.gnu.org/software/gettext/manual/html_node/The-LANGUAGE-variable.html\n\n\nAt minimum this commit message needs revising. Gettext was adopted into\nPOSIX 2024 (Issue 8).\n\n\nRespected by tools of course:\nhttps://pubs.opengroup.org/onlinepubs/9799919799/utilities/gettext.html#tag_20_54_08\nhttps://pubs.opengroup.org/onlinepubs/9799919799/functions/gettext.html\n\n$LANGUAGE docs can be found at\n\nhttps://pubs.opengroup.org/onlinepubs/9799919799/basedefs/V1_chap08.html#tag_08_02\n\n\n\"\"\"\nThe value of LANGUAGE shall be a list of locale names separated by a\n<colon> (':') character. If LANGUAGE is set to a non-empty string, each\nlocale name shall be tried in the specified order and if a messages\nobject is found, it shall be used for translation. If a locale name has\nthe format language[_territory][.codeset][@modifier], additional\nsearches of locale names without .codeset (if present), without\n_territory (if present), and without @modifier (if present) may be performed\n\"\"\"\n\n\nAnd, for locale name values,\n\n\"\"\"\nIf the locale value is \"C\" or \"POSIX\", the POSIX locale shall be used\nand the standard utilities behave in accordance with the rules in 7.2\nPOSIX Locale for the associated category.\n\nIf the locale value begins with a <slash>, it shall be interpreted as\nthe pathname of a file that was created in the output format used by the\nlocaledef utility; see OUTPUT FILES under localedef. Referencing such a\npathname shall result in that locale being used for the indicated category.\n\n[XSI] [Option Start] If the locale value has the form:\n\nlanguage[_territory][.codeset]\n\nit refers to an implementation-provided locale, where settings of\nlanguage, territory, and codeset are implementation-defined.\n\nLC_COLLATE , LC_CTYPE , LC_MESSAGES , LC_MONETARY , LC_NUMERIC , and\nLC_TIME are defined to accept an additional field @modifier, which\nallows the user to select a specific instance of localization data\nwithin a single category (for example, for selecting the dictionary as\nopposed to the character ordering of data). The syntax for these\nenvironment variables is thus defined as:\n\n[language[_territory][.codeset][@modifier]]\n\"\"\"\n\n\nYour tests and code are probably broken -- they appear to normalize\nnearly none of the standard grammar into valid Accept-Language entries.\nOf course, \"surely nobody actually does that\" (except when they do!) --\nbut it's a relatively simple grammar structure, simply getting the\n\"shape\" correct seems like a good idea.\n\n\n-- \nEli Schwartz\n"}]}