{"thread":{"id":"65030","subject":"[RFC] send-email: UTF-8 encoding in subject line","startedAt":"2026-02-20T14:51:40Z","lastAt":"2026-03-03T19:07:25Z","messageCount":27,"participants":["Shreyansh Paliwal","Ben Knoble","Junio C Hamano","D. Ben Knoble","Philip Oakley"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"536523","messageId":"20260220145126.131651-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":null,"subject":"[RFC] send-email: UTF-8 encoding in subject line","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-20T14:50:46Z","receivedAt":"2026-02-20T14:51:40Z","isPatch":false,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"Hi,\n\nWhile using git send-email I ran into some confusion around the prompt that\nappears when any 8-bit (non-ASCII) content is detected.\n\nWhen prompted with,\n\n  Which 8bit encoding should I declare [UTF-8]? y\n  Are you sure you want to use <y> [y/N]? y\n\nI initially assumed this was a yes/no style confirmation and answered \"y\",\nand ignored the 'which' part (this was due to my oversight). This resulted\nin the charset being set to \"y\", which later produced a subject line like,\n\n  =?y?q?...?=\n\nMail clients like Gmail still displayed the message correctly, but the\nmailing list archive showed the raw encoded form[1].\n\nAfterwards, I realized the prompt expects a charset name (e.g., \"UTF-8\")\nrather than a yes/no answer, and pressing enter would have selected the\ndefault (which is UTF-8).\n\nI had also encountered this earlier when the non-ASCII character was in the\nmessage body rather than the subject, in that case the result appeared to\nwork fine even with the mistaken input, which made the issue less obvious\nto me at first.\n\nThis made me wonder whether the current UX around the prompts or input\nvalidation could be improved in any way to reduce the chance of accidental\ninput being interpreted as a charset name.\n\nBest,\nShreyansh\n\n[1]- https://lore.kernel.org/git/20260219181154.66814-1-shreyanshpaliwalcmsmn@gmail.com/\n"},{"id":"536570","messageId":"5EDD26EE-51B6-4BE2-A7C7-E1E0991537E4@gmail.com","threadId":"65030","inReplyTo":"20260220145126.131651-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-02-21T02:28:32Z","receivedAt":"2026-02-21T02:28:44Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 20 févr. 2026 à 09:51, Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> a écrit :\n> \n> ﻿Hi,\n> \n> While using git send-email I ran into some confusion around the prompt that\n> appears when any 8-bit (non-ASCII) content is detected.\n> \n> When prompted with,\n> \n>  Which 8bit encoding should I declare [UTF-8]? y\n>  Are you sure you want to use <y> [y/N]? y\n\nYeah, that was a bit confusing for me until I got used to it. Maybe saying “[default: UTF-8]” would be a small and definite improvement?\n\n> I initially assumed this was a yes/no style confirmation and answered \"y\",\n> and ignored the 'which' part (this was due to my oversight). This resulted\n> in the charset being set to \"y\", which later produced a subject line like,\n> \n>  =?y?q?...?=\n> \n> Mail clients like Gmail still displayed the message correctly, but the\n> mailing list archive showed the raw encoded form[1].\n> \n> Afterwards, I realized the prompt expects a charset name (e.g., \"UTF-8\")\n> rather than a yes/no answer, and pressing enter would have selected the\n> default (which is UTF-8).\n> \n> I had also encountered this earlier when the non-ASCII character was in the\n> message body rather than the subject, in that case the result appeared to\n> work fine even with the mistaken input, which made the issue less obvious\n> to me at first.\n> \n> This made me wonder whether the current UX around the prompts or input\n> validation could be improved in any way to reduce the chance of accidental\n> input being interpreted as a charset name.\n> \n> Best,\n> Shreyansh\n> \n> [1]- https://lore.kernel.org/git/20260219181154.66814-1-shreyanshpaliwalcmsmn@gmail.com/\n\nThanks for thinking on this; better that I never needed to get used to the oddity ;)"},{"id":"536586","messageId":"20260221140049.579922-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"5EDD26EE-51B6-4BE2-A7C7-E1E0991537E4@gmail.com","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-21T13:38:39Z","receivedAt":"2026-02-21T14:01:13Z","isPatch":false,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> > Hi,\n> >\n> > While using git send-email I ran into some confusion around the prompt that\n> > appears when any 8-bit (non-ASCII) content is detected.\n> >\n> > When prompted with,\n> >\n> >  Which 8bit encoding should I declare [UTF-8]? y\n> >  Are you sure you want to use <y> [y/N]? y\n>\n> Yeah, that was a bit confusing for me until I got used to it. Maybe\n> saying “[default: UTF-8]” would be a small and definite improvement?\n\nThat makes sense, I tried it below.\nI also wondered whether, in addition to this, it might be helpful to warn on\nan invalid charset, and/or possibly fall back to UTF-8.\n\nLet me know what you think.\n\nSigned-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n---\n git-send-email.perl | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/git-send-email.perl b/git-send-email.perl\nindex cd4b316ddc..12d0e7e6c9 100755\n--- a/git-send-email.perl\n+++ b/git-send-email.perl\n@@ -1044,7 +1044,7 @@ sub file_declares_8bit_cte {\n \tforeach my $f (sort keys %broken_encoding) {\n \t\tprint \"    $f\\n\";\n \t}\n-\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n+\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [default: UTF-8]? \"),\n \t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n \t\t\t\t  default => \"UTF-8\");\n }\n--\n2.53.0\n\n"},{"id":"536600","messageId":"xmqqldgmrom9.fsf@gitster.g","threadId":"65030","inReplyTo":"20260221140049.579922-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-21T17:30:06Z","receivedAt":"2026-02-21T17:30:08Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n\n>> Yeah, that was a bit confusing for me until I got used to it. Maybe\n>> saying “[default: UTF-8]” would be a small and definite improvement?\n\nThe current message can be mistaken, if the reader does not READ, if\nit is asking a yes/no question, but with the \"default\" label, you\ncannot imagine answering \"yes\", which is clearly not one of the\nthings in the same class as \"UTF-8\" that is given as the default,\nwhich also serves as an example.\n\nThis is indeed a clever hack (not hack on computer code but hack on\nthe mind of human who is reading the message).  \n\n> That makes sense, I tried it below.\n> I also wondered whether, in addition to this, it might be helpful to warn on\n> an invalid charset, and/or possibly fall back to UTF-8.\n\nAgreed on the first half of the statement, if we have an easy and\nportable way to tell if a given random string names a valid charset.\nI do not recommend to \"fall back\" to anything, if we are asking an\ninput from the user.\n\nThanks.\n"},{"id":"536642","messageId":"20260222140737.1760413-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"xmqqldgmrom9.fsf@gitster.g","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-22T14:03:52Z","receivedAt":"2026-02-22T14:07:55Z","isPatch":false,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> > That makes sense, I tried it below.\n> > I also wondered whether, in addition to this, it might be helpful to warn on\n> > an invalid charset, and/or possibly fall back to UTF-8.\n>\n> Agreed on the first half of the statement, if we have an easy and\n> portable way to tell if a given random string names a valid charset.\n> I do not recommend to \"fall back\" to anything, if we are asking an\n> input from the user.\n\nFollowing up on this, I tried adding a warning when the provided charset\ndoes not appear to be valid. Current flow is,\n\n  Which 8bit encoding should I declare [UTF-8]? y\n  Are you sure you want to use <y> [y/N]? y\n\nWith the additional check, it becomes,\n\n  Which 8bit encoding should I declare [default: UTF-8]? y\n  warning: 'y' does not appear to be a valid charset name.\n  Are you sure you want to use <y> [y/N]?\n\nThis uses find_encoding() from Perl’s Encode module to detect any\nunrecognized charset names.\n\nLet me know what you think.\nAlso, is there any new test that should be added for this change?\n\nSigned-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n---\n git-send-email.perl | 23 ++++++++++++++++++++---\n 1 file changed, 20 insertions(+), 3 deletions(-)\n\ndiff --git a/git-send-email.perl b/git-send-email.perl\nindex cd4b316ddc..e62fa259ba 100755\n--- a/git-send-email.perl\n+++ b/git-send-email.perl\n@@ -23,6 +23,7 @@\n use Git::LoadCPAN::Error qw(:try);\n use Git;\n use Git::I18N;\n+use Encode qw(find_encoding);\n \n Getopt::Long::Configure qw/ pass_through /;\n \n@@ -1044,9 +1045,25 @@ sub file_declares_8bit_cte {\n \tforeach my $f (sort keys %broken_encoding) {\n \t\tprint \"    $f\\n\";\n \t}\n-\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n-\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n-\t\t\t\t  default => \"UTF-8\");\n+\twhile (1) {\n+\t\tmy $encoding = ask(__(\"Which 8bit encoding should I declare [default: UTF-8]? \"),\n+\t\t\tvalid_re => qr/^\\S+$/,\n+\t\t\tdefault  => \"UTF-8\");\n+\t\tnext unless defined $encoding;\n+\t\tif (find_encoding($encoding)) {\n+\t\t\t$auto_8bit_encoding = $encoding;\n+\t\t\tlast;\n+\t\t}\n+\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n+\t\tmy $yesno = ask(\n+\t\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $encoding),\n+\t\t\tvalid_re => qr/^(?:y|n)/i,\n+\t\t\tdefault  => 'n');\n+\t\tif (defined $yesno && $yesno =~ /^y/i) {\n+\t\t\t$auto_8bit_encoding = $encoding;\n+\t\t\tlast;\n+\t\t}\n+\t}\n }\n \n if (!$force) {\n-- \n2.53.0\n"},{"id":"536644","messageId":"CALnO6CALy48cmpqSp7TMjkg0ZuuMwhw_hsvLU09yxp5MSEi=Wg@mail.gmail.com","threadId":"65030","inReplyTo":"xmqqldgmrom9.fsf@gitster.g","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-02-22T14:53:10Z","receivedAt":"2026-02-22T14:53:22Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Sat, Feb 21, 2026 at 12:30 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n>\n> >> Yeah, that was a bit confusing for me until I got used to it. Maybe\n> >> saying “[default: UTF-8]” would be a small and definite improvement?\n>\n> The current message can be mistaken, if the reader does not READ, if\n> it is asking a yes/no question, but with the \"default\" label, you\n> cannot imagine answering \"yes\", which is clearly not one of the\n> things in the same class as \"UTF-8\" that is given as the default,\n> which also serves as an example.\n>\n> This is indeed a clever hack (not hack on computer code but hack on\n> the mind of human who is reading the message).\n\nYep. I've seen it somewhere before; I thought it was from Portage's\nemerge in \"ask\" mode, but checking that doesn't seem to be the case.\n\n-- \nD. Ben Knoble\n"},{"id":"536645","messageId":"CALnO6CBhB+O-CBCw3f+2n5yaHO7Wk7-Adaa9_4shXZvciGpUPA@mail.gmail.com","threadId":"65030","inReplyTo":"20260222140737.1760413-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-02-22T15:00:28Z","receivedAt":"2026-02-22T15:00:40Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Sun, Feb 22, 2026 at 9:07 AM Shreyansh Paliwal\n<shreyanshpaliwalcmsmn@gmail.com> wrote:\n>\n> > > That makes sense, I tried it below.\n> > > I also wondered whether, in addition to this, it might be helpful to warn on\n> > > an invalid charset, and/or possibly fall back to UTF-8.\n> >\n> > Agreed on the first half of the statement, if we have an easy and\n> > portable way to tell if a given random string names a valid charset.\n> > I do not recommend to \"fall back\" to anything, if we are asking an\n> > input from the user.\n>\n> Following up on this, I tried adding a warning when the provided charset\n> does not appear to be valid. Current flow is,\n>\n>   Which 8bit encoding should I declare [UTF-8]? y\n>   Are you sure you want to use <y> [y/N]? y\n>\n> With the additional check, it becomes,\n>\n>   Which 8bit encoding should I declare [default: UTF-8]? y\n>   warning: 'y' does not appear to be a valid charset name.\n>   Are you sure you want to use <y> [y/N]?\n>\n> This uses find_encoding() from Perl’s Encode module to detect any\n> unrecognized charset names.\n>\n> Let me know what you think.\n> Also, is there any new test that should be added for this change?\n>\n> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> ---\n>  git-send-email.perl | 23 ++++++++++++++++++++---\n>  1 file changed, 20 insertions(+), 3 deletions(-)\n>\n> diff --git a/git-send-email.perl b/git-send-email.perl\n> index cd4b316ddc..e62fa259ba 100755\n> --- a/git-send-email.perl\n> +++ b/git-send-email.perl\n> @@ -23,6 +23,7 @@\n>  use Git::LoadCPAN::Error qw(:try);\n>  use Git;\n>  use Git::I18N;\n> +use Encode qw(find_encoding);\n>\n>  Getopt::Long::Configure qw/ pass_through /;\n>\n> @@ -1044,9 +1045,25 @@ sub file_declares_8bit_cte {\n>         foreach my $f (sort keys %broken_encoding) {\n>                 print \"    $f\\n\";\n>         }\n> -       $auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n> -                                 valid_re => qr/.{4}/, confirm_only => 1,\n> -                                 default => \"UTF-8\");\n> +       while (1) {\n> +               my $encoding = ask(__(\"Which 8bit encoding should I declare [default: UTF-8]? \"),\n> +                       valid_re => qr/^\\S+$/,\n> +                       default  => \"UTF-8\");\n\nHere we change things, right?\n\n- The original validation is \"at least 4 characters\", the new\nvalidation is \"at least one non-blank.\" I'm not sure why we'd prefer\none or the other, frankly. The original goes to 852a15d748\n(send-email: ask confirmation if given encoding name is very short,\n2015-02-13), which is motivated by the same problem we're discussing\nhere!\n- We get rid of confirm_only, since we're about to roll our own\nconfirmation below:\n\n> +               next unless defined $encoding;\n> +               if (find_encoding($encoding)) {\n> +                       $auto_8bit_encoding = $encoding;\n> +                       last;\n> +               }\n> +               printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n> +               my $yesno = ask(\n> +                       sprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $encoding),\n> +                       valid_re => qr/^(?:y|n)/i,\n> +                       default  => 'n');\n\n…which might want refactored a bit so it can stay close to the original? idk.\n\n> +               if (defined $yesno && $yesno =~ /^y/i) {\n> +                       $auto_8bit_encoding = $encoding;\n> +                       last;\n> +               }\n> +       }\n>  }\n>\n>  if (!$force) {\n> --\n> 2.53.0\n\n\n\n-- \nD. Ben Knoble\n"},{"id":"536646","messageId":"335b1189-f5c3-4e7c-ad3a-266810a0ca90@iee.email","threadId":"65030","inReplyTo":"20260222140737.1760413-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2026-02-22T14:53:32Z","receivedAt":"2026-02-22T15:38:04Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 22/02/2026 14:03, Shreyansh Paliwal wrote:\n>>> That makes sense, I tried it below.\n>>> I also wondered whether, in addition to this, it might be helpful to warn on\n>>> an invalid charset, and/or possibly fall back to UTF-8.\n>>\n>> Agreed on the first half of the statement, if we have an easy and\n>> portable way to tell if a given random string names a valid charset.\n>> I do not recommend to \"fall back\" to anything, if we are asking an\n>> input from the user.\n> \n> Following up on this, I tried adding a warning when the provided charset\n> does not appear to be valid. Current flow is,\n> \n>   Which 8bit encoding should I declare [UTF-8]? y\n\nPerhaps swap around the 'Which-declare' to \"Declare which' to to get\naway from the obviousness of 'Which' being the classic y/n binary\nquestion. Action first?\n\n\tDeclare which 8bit encoding to use [default:UTF-8]?\n\nChecking validity of the encoding is a reasonable follow on.\n\nPhilip\n\n>   Are you sure you want to use <y> [y/N]? y\n> \n> With the additional check, it becomes,\n> \n>   Which 8bit encoding should I declare [default: UTF-8]? y\n>   warning: 'y' does not appear to be a valid charset name.\n>   Are you sure you want to use <y> [y/N]?\n> \n> This uses find_encoding() from Perl’s Encode module to detect any\n> unrecognized charset names.\n> \n> Let me know what you think.\n> Also, is there any new test that should be added for this change?\n> \n> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> ---\n>  git-send-email.perl | 23 ++++++++++++++++++++---\n>  1 file changed, 20 insertions(+), 3 deletions(-)\n> \n> diff --git a/git-send-email.perl b/git-send-email.perl\n> index cd4b316ddc..e62fa259ba 100755\n> --- a/git-send-email.perl\n> +++ b/git-send-email.perl\n> @@ -23,6 +23,7 @@\n>  use Git::LoadCPAN::Error qw(:try);\n>  use Git;\n>  use Git::I18N;\n> +use Encode qw(find_encoding);\n>  \n>  Getopt::Long::Configure qw/ pass_through /;\n>  \n> @@ -1044,9 +1045,25 @@ sub file_declares_8bit_cte {\n>  \tforeach my $f (sort keys %broken_encoding) {\n>  \t\tprint \"    $f\\n\";\n>  \t}\n> -\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n> -\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n> -\t\t\t\t  default => \"UTF-8\");\n> +\twhile (1) {\n> +\t\tmy $encoding = ask(__(\"Which 8bit encoding should I declare [default: UTF-8]? \"),\n> +\t\t\tvalid_re => qr/^\\S+$/,\n> +\t\t\tdefault  => \"UTF-8\");\n> +\t\tnext unless defined $encoding;\n> +\t\tif (find_encoding($encoding)) {\n> +\t\t\t$auto_8bit_encoding = $encoding;\n> +\t\t\tlast;\n> +\t\t}\n> +\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n> +\t\tmy $yesno = ask(\n> +\t\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $encoding),\n> +\t\t\tvalid_re => qr/^(?:y|n)/i,\n> +\t\t\tdefault  => 'n');\n> +\t\tif (defined $yesno && $yesno =~ /^y/i) {\n> +\t\t\t$auto_8bit_encoding = $encoding;\n> +\t\t\tlast;\n> +\t\t}\n> +\t}\n>  }\n>  \n>  if (!$force) {\n\n"},{"id":"536648","messageId":"20260222155559.1777883-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"CALnO6CBhB+O-CBCw3f+2n5yaHO7Wk7-Adaa9_4shXZvciGpUPA@mail.gmail.com","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-22T15:52:01Z","receivedAt":"2026-02-22T15:56:11Z","isPatch":false,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> On Sun, Feb 22, 2026 at 9:07 AM Shreyansh Paliwal\n> <shreyanshpaliwalcmsmn@gmail.com> wrote:\n> >\n> > > > That makes sense, I tried it below.\n> > > > I also wondered whether, in addition to this, it might be helpful to warn on\n> > > > an invalid charset, and/or possibly fall back to UTF-8.\n> > >\n> > > Agreed on the first half of the statement, if we have an easy and\n> > > portable way to tell if a given random string names a valid charset.\n> > > I do not recommend to \"fall back\" to anything, if we are asking an\n> > > input from the user.\n> >\n> > Following up on this, I tried adding a warning when the provided charset\n> > does not appear to be valid. Current flow is,\n> >\n> >   Which 8bit encoding should I declare [UTF-8]? y\n> >   Are you sure you want to use <y> [y/N]? y\n> >\n> > With the additional check, it becomes,\n> >\n> >   Which 8bit encoding should I declare [default: UTF-8]? y\n> >   warning: 'y' does not appear to be a valid charset name.\n> >   Are you sure you want to use <y> [y/N]?\n> >\n> > This uses find_encoding() from Perl’s Encode module to detect any\n> > unrecognized charset names.\n> >\n> > Let me know what you think.\n> > Also, is there any new test that should be added for this change?\n> >\n> > Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> > ---\n> >  git-send-email.perl | 23 ++++++++++++++++++++---\n> >  1 file changed, 20 insertions(+), 3 deletions(-)\n> >\n> > diff --git a/git-send-email.perl b/git-send-email.perl\n> > index cd4b316ddc..e62fa259ba 100755\n> > --- a/git-send-email.perl\n> > +++ b/git-send-email.perl\n> > @@ -23,6 +23,7 @@\n> >  use Git::LoadCPAN::Error qw(:try);\n> >  use Git;\n> >  use Git::I18N;\n> > +use Encode qw(find_encoding);\n> >\n> >  Getopt::Long::Configure qw/ pass_through /;\n> >\n> > @@ -1044,9 +1045,25 @@ sub file_declares_8bit_cte {\n> >         foreach my $f (sort keys %broken_encoding) {\n> >                 print \"    $f\\n\";\n> >         }\n> > -       $auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n> > -                                 valid_re => qr/.{4}/, confirm_only => 1,\n> > -                                 default => \"UTF-8\");\n> > +       while (1) {\n> > +               my $encoding = ask(__(\"Which 8bit encoding should I declare [default: UTF-8]? \"),\n> > +                       valid_re => qr/^\\S+$/,\n> > +                       default  => \"UTF-8\");\n>\n> Here we change things, right?\n>\n> - The original validation is \"at least 4 characters\", the new\n> validation is \"at least one non-blank.\" I'm not sure why we'd prefer\n> one or the other, frankly. The original goes to 852a15d748\n> (send-email: ask confirmation if given encoding name is very short,\n> 2015-02-13), which is motivated by the same problem we're discussing\n> here!\n\nI see.\nMy understanding of the earlier change (852a15d748) is that the\nlength check was intended as a heuristic check to catch obviously invalid\ninputs like \"y\" and trigger an extra confirmation based on the fact that\ncharset names would be at least 4 letters.\n\nWith the additional find_encoding() check, the validation becomes semantic\nrather than length-based, recognized charset names are accepted directly,\nwhile unrecognized ones trigger a warning and still require explicit\nconfirmation. The relaxed regex (at least one non-blank) is only meant to\nensure we receive some non-empty input before passing it to find_encoding().\n\n> - We get rid of confirm_only, since we're about to roll our own\n> confirmation below:\n>\n> > +               next unless defined $encoding;\n> > +               if (find_encoding($encoding)) {\n> > +                       $auto_8bit_encoding = $encoding;\n> > +                       last;\n> > +               }\n> > +               printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n> > +               my $yesno = ask(\n> > +                       sprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $encoding),\n> > +                       valid_re => qr/^(?:y|n)/i,\n> > +                       default  => 'n');\n>\n> …which might want refactored a bit so it can stay close to the original? idk.\n>\n\nActually the flow needed to change slightly to insert the validity warning\nbefore the final confirmation step. Since ask() handles confirmation internally\nusing confrim_only and is used in multiple places, it seemed simpler to keep the\nadditional confirmation local here rather than modifying ask() itself.\n\nLet me know what you think.\n\nBest,\nShreyansh\n"},{"id":"536885","messageId":"43DCEEB9-33C4-4EE2-9FF3-49DCB9B837E0@gmail.com","threadId":"65030","inReplyTo":"20260222155559.1777883-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [RFC] send-email: UTF-8 encoding in subject line","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-02-23T21:38:31Z","receivedAt":"2026-02-23T21:38:43Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 22 févr. 2026 à 10:56, Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> a écrit :\n> \n> ﻿\n>> \n>>> On Sun, Feb 22, 2026 at 9:07 AM Shreyansh Paliwal\n>>> <shreyanshpaliwalcmsmn@gmail.com> wrote:\n>>> \n>>>>> That makes sense, I tried it below.\n>>>>> I also wondered whether, in addition to this, it might be helpful to warn on\n>>>>> an invalid charset, and/or possibly fall back to UTF-8.\n>>>> \n>>>> Agreed on the first half of the statement, if we have an easy and\n>>>> portable way to tell if a given random string names a valid charset.\n>>>> I do not recommend to \"fall back\" to anything, if we are asking an\n>>>> input from the user.\n>>> \n>>> Following up on this, I tried adding a warning when the provided charset\n>>> does not appear to be valid. Current flow is,\n>>> \n>>>  Which 8bit encoding should I declare [UTF-8]? y\n>>>  Are you sure you want to use <y> [y/N]? y\n>>> \n>>> With the additional check, it becomes,\n>>> \n>>>  Which 8bit encoding should I declare [default: UTF-8]? y\n>>>  warning: 'y' does not appear to be a valid charset name.\n>>>  Are you sure you want to use <y> [y/N]?\n>>> \n>>> This uses find_encoding() from Perl’s Encode module to detect any\n>>> unrecognized charset names.\n>>> \n>>> Let me know what you think.\n>>> Also, is there any new test that should be added for this change?\n>>> \n>>> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n>>> ---\n>>> git-send-email.perl | 23 ++++++++++++++++++++---\n>>> 1 file changed, 20 insertions(+), 3 deletions(-)\n>>> \n>>> diff --git a/git-send-email.perl b/git-send-email.perl\n>>> index cd4b316ddc..e62fa259ba 100755\n>>> --- a/git-send-email.perl\n>>> +++ b/git-send-email.perl\n>>> @@ -23,6 +23,7 @@\n>>> use Git::LoadCPAN::Error qw(:try);\n>>> use Git;\n>>> use Git::I18N;\n>>> +use Encode qw(find_encoding);\n>>> \n>>> Getopt::Long::Configure qw/ pass_through /;\n>>> \n>>> @@ -1044,9 +1045,25 @@ sub file_declares_8bit_cte {\n>>>        foreach my $f (sort keys %broken_encoding) {\n>>>                print \"    $f\\n\";\n>>>        }\n>>> -       $auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n>>> -                                 valid_re => qr/.{4}/, confirm_only => 1,\n>>> -                                 default => \"UTF-8\");\n>>> +       while (1) {\n>>> +               my $encoding = ask(__(\"Which 8bit encoding should I declare [default: UTF-8]? \"),\n>>> +                       valid_re => qr/^\\S+$/,\n>>> +                       default  => \"UTF-8\");\n>> \n>> Here we change things, right?\n>> \n>> - The original validation is \"at least 4 characters\", the new\n>> validation is \"at least one non-blank.\" I'm not sure why we'd prefer\n>> one or the other, frankly. The original goes to 852a15d748\n>> (send-email: ask confirmation if given encoding name is very short,\n>> 2015-02-13), which is motivated by the same problem we're discussing\n>> here!\n> \n> I see.\n> My understanding of the earlier change (852a15d748) is that the\n> length check was intended as a heuristic check to catch obviously invalid\n> inputs like \"y\" and trigger an extra confirmation based on the fact that\n> charset names would be at least 4 letters.\n> \n> With the additional find_encoding() check, the validation becomes semantic\n> rather than length-based, recognized charset names are accepted directly,\n> while unrecognized ones trigger a warning and still require explicit\n> confirmation. The relaxed regex (at least one non-blank) is only meant to\n> ensure we receive some non-empty input before passing it to find_encoding().\n> \n>> - We get rid of confirm_only, since we're about to roll our own\n>> confirmation below:\n>> \n>>> +               next unless defined $encoding;\n>>> +               if (find_encoding($encoding)) {\n>>> +                       $auto_8bit_encoding = $encoding;\n>>> +                       last;\n>>> +               }\n>>> +               printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n>>> +               my $yesno = ask(\n>>> +                       sprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $encoding),\n>>> +                       valid_re => qr/^(?:y|n)/i,\n>>> +                       default  => 'n');\n>> \n>> …which might want refactored a bit so it can stay close to the original? idk.\n>> \n> \n> Actually the flow needed to change slightly to insert the validity warning\n> before the final confirmation step. Since ask() handles confirmation internally\n> using confrim_only and is used in multiple places, it seemed simpler to keep the\n> additional confirmation local here rather than modifying ask() itself.\n> \n> Let me know what you think.\n> \n> Best,\n> Shreyansh\n\nAh, my mistake for being ambiguous. I meant:\n\nThe code is similar enough to the original that perhaps a helper can be introduced, or at least we should keep the equivalent strings together to help those who change one. "},{"id":"536930","messageId":"20260224075650.1885050-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"43DCEEB9-33C4-4EE2-9FF3-49DCB9B837E0@gmail.com","subject":"Re: [GSOC] Discuss: Refactoring in order to reduce global state","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-24T07:55:44Z","receivedAt":"2026-02-24T07:57:13Z","isPatch":false,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> > Le 22 févr. 2026 à 10:56, Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> a écrit :\n> >\n> > ﻿\n> >>\n> >>> On Sun, Feb 22, 2026 at 9:07 AM Shreyansh Paliwal\n> >>> <shreyanshpaliwalcmsmn@gmail.com> wrote:\n> >>>\n> >>>>> That makes sense, I tried it below.\n> >>>>> I also wondered whether, in addition to this, it might be helpful to warn on\n> >>>>> an invalid charset, and/or possibly fall back to UTF-8.\n> >>>>\n> >>>> Agreed on the first half of the statement, if we have an easy and\n> >>>> portable way to tell if a given random string names a valid charset.\n> >>>> I do not recommend to \"fall back\" to anything, if we are asking an\n> >>>> input from the user.\n> >>>\n> >>> Following up on this, I tried adding a warning when the provided charset\n> >>> does not appear to be valid. Current flow is,\n> >>>\n> >>>  Which 8bit encoding should I declare [UTF-8]? y\n> >>>  Are you sure you want to use <y> [y/N]? y\n> >>>\n> >>> With the additional check, it becomes,\n> >>>\n> >>>  Which 8bit encoding should I declare [default: UTF-8]? y\n> >>>  warning: 'y' does not appear to be a valid charset name.\n> >>>  Are you sure you want to use <y> [y/N]?\n> >>>\n> >>> This uses find_encoding() from Perl’s Encode module to detect any\n> >>> unrecognized charset names.\n> >>>\n> >>> Let me know what you think.\n> >>> Also, is there any new test that should be added for this change?\n> >>>\n> >>> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> >>> ---\n> >>> git-send-email.perl | 23 ++++++++++++++++++++---\n> >>> 1 file changed, 20 insertions(+), 3 deletions(-)\n> >>>\n> >>> diff --git a/git-send-email.perl b/git-send-email.perl\n> >>> index cd4b316ddc..e62fa259ba 100755\n> >>> --- a/git-send-email.perl\n> >>> +++ b/git-send-email.perl\n> >>> @@ -23,6 +23,7 @@\n> >>> use Git::LoadCPAN::Error qw(:try);\n> >>> use Git;\n> >>> use Git::I18N;\n> >>> +use Encode qw(find_encoding);\n> >>>\n> >>> Getopt::Long::Configure qw/ pass_through /;\n> >>>\n> >>> @@ -1044,9 +1045,25 @@ sub file_declares_8bit_cte {\n> >>>        foreach my $f (sort keys %broken_encoding) {\n> >>>                print \"    $f\\n\";\n> >>>        }\n> >>> -       $auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n> >>> -                                 valid_re => qr/.{4}/, confirm_only => 1,\n> >>> -                                 default => \"UTF-8\");\n> >>> +       while (1) {\n> >>> +               my $encoding = ask(__(\"Which 8bit encoding should I declare [default: UTF-8]? \"),\n> >>> +                       valid_re => qr/^\\S+$/,\n> >>> +                       default  => \"UTF-8\");\n> >>\n> >> Here we change things, right?\n> >>\n> >> - The original validation is \"at least 4 characters\", the new\n> >> validation is \"at least one non-blank.\" I'm not sure why we'd prefer\n> >> one or the other, frankly. The original goes to 852a15d748\n> >> (send-email: ask confirmation if given encoding name is very short,\n> >> 2015-02-13), which is motivated by the same problem we're discussing\n> >> here!\n> >\n> > I see.\n> > My understanding of the earlier change (852a15d748) is that the\n> > length check was intended as a heuristic check to catch obviously invalid\n> > inputs like \"y\" and trigger an extra confirmation based on the fact that\n> > charset names would be at least 4 letters.\n> >\n> > With the additional find_encoding() check, the validation becomes semantic\n> > rather than length-based, recognized charset names are accepted directly,\n> > while unrecognized ones trigger a warning and still require explicit\n> > confirmation. The relaxed regex (at least one non-blank) is only meant to\n> > ensure we receive some non-empty input before passing it to find_encoding().\n> >\n> >> - We get rid of confirm_only, since we're about to roll our own\n> >> confirmation below:\n> >>\n> >>> +               next unless defined $encoding;\n> >>> +               if (find_encoding($encoding)) {\n> >>> +                       $auto_8bit_encoding = $encoding;\n> >>> +                       last;\n> >>> +               }\n> >>> +               printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n> >>> +               my $yesno = ask(\n> >>> +                       sprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $encoding),\n> >>> +                       valid_re => qr/^(?:y|n)/i,\n> >>> +                       default  => 'n');\n> >>\n> >> …which might want refactored a bit so it can stay close to the original? idk.\n> >>\n> >\n> > Actually the flow needed to change slightly to insert the validity warning\n> > before the final confirmation step. Since ask() handles confirmation internally\n> > using confrim_only and is used in multiple places, it seemed simpler to keep the\n> > additional confirmation local here rather than modifying ask() itself.\n> >\n> > Let me know what you think.\n> >\n> > Best,\n> > Shreyansh\n>\n> Ah, my mistake for being ambiguous. I meant:\n>\n> The code is similar enough to the original that perhaps a helper can be\n> introduced, or at least we should keep the equivalent strings together to\n> help those who change one.\n\nThanks for clarifying, that makes sense.\nI'll refactor and send a revised patch on this.\n\nBest,\nShreyansh\n"},{"id":"536975","messageId":"20260224143624.23678-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"20260220145126.131651-1-shreyanshpaliwalcmsmn@gmail.com","subject":"[PATCH] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-24T14:33:52Z","receivedAt":"2026-02-24T14:37:01Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"When a non-ASCII character is detected in the body or subject of the email\nthe user is prompted with,\n\n  Which 8bit encoding should I declare [UTF-8]? foo\n\nAfter this the input string is validated by the regex, based on the fact\nthat the charset string will be minimum 4 characters [1]. If the string is\nmore than 4 letters the email is sent, if not then a second prompt to\nconfirm is asked to the user,\n\n  Are you sure you want to use <foo> [y/N]? y\n\nThis relies on a length based regex heuristic check to validate the user\ninput, and can allow invalid charset names to pass if the input is greater\nthan 4 characters.\n\nAdd a semantic validation of the charset name using the\nEncode::find_encoding() module of perl. If the encoding is not recognized,\nwarn the user and ask for confirmation before proceeding. After this\nvalidation the lenght based validation becomes redundant and also breaks\nflow, so change the regex of valid input to any non blank string.\n\nAdditionally, the wording of the first prompt can confuse the user if not\nread properly or under any default assumptions for a yes/no prompt. Change\nthe wording to make it explicitly clear to the user that the prompt needs a\nstring input, UTF-8 being the default.\n\nThe intended flow is,\n\n  Declare which 8bit encoding to use [default: UTF-8]? foobar\n  warning: 'foobar' does not appear to be a valid charset name.\n  Are you sure you want to use <foobar> [y/N]?\n\n[1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n\nSigned-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n---\n git-send-email.perl   | 15 ++++++++++++---\n t/t9001-send-email.sh |  2 +-\n 2 files changed, 13 insertions(+), 4 deletions(-)\n\ndiff --git a/git-send-email.perl b/git-send-email.perl\nindex cd4b316ddc..dc4e5418d3 100755\n--- a/git-send-email.perl\n+++ b/git-send-email.perl\n@@ -23,6 +23,7 @@\n use Git::LoadCPAN::Error qw(:try);\n use Git;\n use Git::I18N;\n+use Encode qw(find_encoding);\n \n Getopt::Long::Configure qw/ pass_through /;\n \n@@ -987,6 +988,7 @@ sub get_patch_subject {\n sub ask {\n \tmy ($prompt, %arg) = @_;\n \tmy $valid_re = $arg{valid_re};\n+\tmy $warn_invalid = $arg{warn_invalid};\n \tmy $default = $arg{default};\n \tmy $confirm_only = $arg{confirm_only};\n \tmy $resp;\n@@ -1005,7 +1007,13 @@ sub ask {\n \t\t\treturn $default;\n \t\t}\n \t\tif (!defined $valid_re or $resp =~ /$valid_re/) {\n-\t\t\treturn $resp;\n+\t\t\tif ($warn_invalid) {\n+\t\t\t\tif (find_encoding($resp))\n+\t\t\t\t\treturn $resp;\n+\t\t\t\telse\n+\t\t\t\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n+\t\t\t} else\n+\t\t\t\treturn $resp;\n \t\t}\n \t\tif ($confirm_only) {\n \t\t\tmy $yesno = $term->readline(\n@@ -1044,8 +1052,9 @@ sub file_declares_8bit_cte {\n \tforeach my $f (sort keys %broken_encoding) {\n \t\tprint \"    $f\\n\";\n \t}\n-\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n-\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n+\t$auto_8bit_encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n+\t\t\t\t  valid_re => qr/^\\S+$/, confirm_only => 1,\n+\t\t\t\t  warn_invalid => 1,\n \t\t\t\t  default => \"UTF-8\");\n }\n \ndiff --git a/t/t9001-send-email.sh b/t/t9001-send-email.sh\nindex e56e0c8d77..24f6c76aee 100755\n--- a/t/t9001-send-email.sh\n+++ b/t/t9001-send-email.sh\n@@ -1691,7 +1691,7 @@ test_expect_success $PREREQ 'asks about and fixes 8bit encodings' '\n \t\t\temail-using-8bit >stdout &&\n \tgrep \"do not declare a Content-Transfer-Encoding\" stdout &&\n \tgrep email-using-8bit stdout &&\n-\tgrep \"Which 8bit encoding\" stdout &&\n+\tgrep \"Declare which 8bit encoding to use\" stdout &&\n \tgrep -E \"Content|MIME\" msgtxt1 >actual &&\n \ttest_cmp content-type-decl actual\n '\n-- \n2.53.0.155.g35e93594f7.dirty\n"},{"id":"537018","messageId":"xmqqbjhdg847.fsf@gitster.g","threadId":"65030","inReplyTo":"20260224143624.23678-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [PATCH] send-email: validate charset name in 8bit encoding prompt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-24T21:11:04Z","receivedAt":"2026-02-24T21:11:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n\n> +\t\t\tif ($warn_invalid) {\n> +\t\t\t\tif (find_encoding($resp))\n> +\t\t\t\t\treturn $resp;\n> +\t\t\t\telse\n> +\t\t\t\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n> +\t\t\t} else\n> +\t\t\t\treturn $resp;\n\nThis is not C but Perl.\n\n git-send-email.perl | 8 +++++---\n 1 file changed, 5 insertions(+), 3 deletions(-)\n\ndiff --git i/git-send-email.perl w/git-send-email.perl\nindex dc4e5418d3..15387ac377 100755\n--- i/git-send-email.perl\n+++ w/git-send-email.perl\n@@ -1008,12 +1008,14 @@ sub ask {\n \t\t}\n \t\tif (!defined $valid_re or $resp =~ /$valid_re/) {\n \t\t\tif ($warn_invalid) {\n-\t\t\t\tif (find_encoding($resp))\n+\t\t\t\tif (find_encoding($resp)) {\n \t\t\t\t\treturn $resp;\n-\t\t\t\telse\n+\t\t\t\t} else {\n \t\t\t\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n-\t\t\t} else\n+\t\t\t\t}\n+\t\t\t} else {\n \t\t\t\treturn $resp;\n+\t\t\t}\n \t\t}\n \t\tif ($confirm_only) {\n \t\t\tmy $yesno = $term->readline(\n"},{"id":"537023","messageId":"20260224213932.92364-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"20260224143624.23678-1-shreyanshpaliwalcmsmn@gmail.com","subject":"[PATCH v2] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-24T21:37:39Z","receivedAt":"2026-02-24T21:39:48Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"When a non-ASCII character is detected in the body or subject of the email\nthe user is prompted with,\n\n  Which 8bit encoding should I declare [UTF-8]? foo\n\nAfter this the input string is validated by the regex, based on the fact\nthat the charset string will be minimum 4 characters [1]. If the string is\nmore than 4 letters the email is sent, if not then a second prompt to\nconfirm is asked to the user,\n\n  Are you sure you want to use <foo> [y/N]? y\n\nThis relies on a length based regex heuristic check to validate the user\ninput, and can allow clearly invalid charset names to pass if the input is\ngreater than 4 characters.\n\nAdd a semantic validation of the charset name using the\nEncode::find_encoding() module of perl. If the encoding is not recognized,\nwarn the user and ask for confirmation before proceeding. After this\nvalidation the lenght based validation becomes redundant and also breaks\nflow, so change the regex of valid input to any non blank string.\n\nAdditionally, the wording of the first prompt can confuse the user if not\nread properly or under any default assumptions for a yes/no prompt. Change\nthe wording to make it explicitly clear to the user that the prompt needs a\nstring input, UTF-8 being the default.\n\nThe intended flow is,\n\n  Declare which 8bit encoding to use [default: UTF-8]? foobar\n  warning: 'foobar' does not appear to be a valid charset name.\n  Are you sure you want to use <foobar> [y/N]?\n\n[1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n\nSigned-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n---\nChanges in v2:\n - Added braces in if-else block.\n\n git-send-email.perl   | 17 ++++++++++++++---\n t/t9001-send-email.sh |  2 +-\n 2 files changed, 15 insertions(+), 4 deletions(-)\n\ndiff --git a/git-send-email.perl b/git-send-email.perl\nindex cd4b316ddc..15387ac377 100755\n--- a/git-send-email.perl\n+++ b/git-send-email.perl\n@@ -23,6 +23,7 @@\n use Git::LoadCPAN::Error qw(:try);\n use Git;\n use Git::I18N;\n+use Encode qw(find_encoding);\n\n Getopt::Long::Configure qw/ pass_through /;\n\n@@ -987,6 +988,7 @@ sub get_patch_subject {\n sub ask {\n \tmy ($prompt, %arg) = @_;\n \tmy $valid_re = $arg{valid_re};\n+\tmy $warn_invalid = $arg{warn_invalid};\n \tmy $default = $arg{default};\n \tmy $confirm_only = $arg{confirm_only};\n \tmy $resp;\n@@ -1005,7 +1007,15 @@ sub ask {\n \t\t\treturn $default;\n \t\t}\n \t\tif (!defined $valid_re or $resp =~ /$valid_re/) {\n-\t\t\treturn $resp;\n+\t\t\tif ($warn_invalid) {\n+\t\t\t\tif (find_encoding($resp)) {\n+\t\t\t\t\treturn $resp;\n+\t\t\t\t} else {\n+\t\t\t\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n+\t\t\t\t}\n+\t\t\t} else {\n+\t\t\t\treturn $resp;\n+\t\t\t}\n \t\t}\n \t\tif ($confirm_only) {\n \t\t\tmy $yesno = $term->readline(\n@@ -1044,8 +1054,9 @@ sub file_declares_8bit_cte {\n \tforeach my $f (sort keys %broken_encoding) {\n \t\tprint \"    $f\\n\";\n \t}\n-\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n-\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n+\t$auto_8bit_encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n+\t\t\t\t  valid_re => qr/^\\S+$/, confirm_only => 1,\n+\t\t\t\t  warn_invalid => 1,\n \t\t\t\t  default => \"UTF-8\");\n }\n\ndiff --git a/t/t9001-send-email.sh b/t/t9001-send-email.sh\nindex e56e0c8d77..24f6c76aee 100755\n--- a/t/t9001-send-email.sh\n+++ b/t/t9001-send-email.sh\n@@ -1691,7 +1691,7 @@ test_expect_success $PREREQ 'asks about and fixes 8bit encodings' '\n \t\t\temail-using-8bit >stdout &&\n \tgrep \"do not declare a Content-Transfer-Encoding\" stdout &&\n \tgrep email-using-8bit stdout &&\n-\tgrep \"Which 8bit encoding\" stdout &&\n+\tgrep \"Declare which 8bit encoding to use\" stdout &&\n \tgrep -E \"Content|MIME\" msgtxt1 >actual &&\n \ttest_cmp content-type-decl actual\n '\n\nRange-diff against v1:\n1:  70fa4d2899 ! 1:  954c1dae9f send-email: validate charset name in 8bit encoding prompt\n    @@ git-send-email.perl: sub ask {\n      \t\tif (!defined $valid_re or $resp =~ /$valid_re/) {\n     -\t\t\treturn $resp;\n     +\t\t\tif ($warn_invalid) {\n    -+\t\t\t\tif (find_encoding($resp))\n    ++\t\t\t\tif (find_encoding($resp)) {\n     +\t\t\t\t\treturn $resp;\n    -+\t\t\t\telse\n    ++\t\t\t\t} else {\n     +\t\t\t\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n    -+\t\t\t} else\n    ++\t\t\t\t}\n    ++\t\t\t} else {\n     +\t\t\t\treturn $resp;\n    ++\t\t\t}\n      \t\t}\n      \t\tif ($confirm_only) {\n      \t\t\tmy $yesno = $term->readline(\n--\n2.53.0\n"},{"id":"537026","messageId":"xmqqqzq9er01.fsf@gitster.g","threadId":"65030","inReplyTo":"20260224213932.92364-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [PATCH v2] send-email: validate charset name in 8bit encoding prompt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-24T22:06:06Z","receivedAt":"2026-02-24T22:06:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n\n> When a non-ASCII character is detected in the body or subject of the email\n> the user is prompted with,\n>\n>   Which 8bit encoding should I declare [UTF-8]? foo\n>\n> After this the input string is validated by the regex, based on the fact\n> that the charset string will be minimum 4 characters [1]. If the string is\n> more than 4 letters the email is sent, if not then a second prompt to\n> confirm is asked to the user,\n>\n>   Are you sure you want to use <foo> [y/N]? y\n>\n> This relies on a length based regex heuristic check to validate the user\n> input, and can allow clearly invalid charset names to pass if the input is\n> greater than 4 characters.\n>\n> Add a semantic validation of the charset name using the\n> Encode::find_encoding() module of perl. If the encoding is not recognized,\n> warn the user and ask for confirmation before proceeding. After this\n> validation the lenght based validation becomes redundant and also breaks\n> flow, so change the regex of valid input to any non blank string.\n>\n> Additionally, the wording of the first prompt can confuse the user if not\n> read properly or under any default assumptions for a yes/no prompt. Change\n> the wording to make it explicitly clear to the user that the prompt needs a\n> string input, UTF-8 being the default.\n>\n> The intended flow is,\n>\n>   Declare which 8bit encoding to use [default: UTF-8]? foobar\n>   warning: 'foobar' does not appear to be a valid charset name.\n>   Are you sure you want to use <foobar> [y/N]?\n>\n> [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n>\n> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> ---\n> Changes in v2:\n>  - Added braces in if-else block.\n>\n>  git-send-email.perl   | 17 ++++++++++++++---\n>  t/t9001-send-email.sh |  2 +-\n>  2 files changed, 15 insertions(+), 4 deletions(-)\n\nCurious.  This change to t9001 was there even in the previous\niteration that did not even work.  How did you test it?\n\nWill replace.\n"},{"id":"537031","messageId":"20260224222156.13712-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"xmqqqzq9er01.fsf@gitster.g","subject":"Re: [PATCH v2] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-24T22:20:36Z","receivedAt":"2026-02-24T22:22:27Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n>\n> > When a non-ASCII character is detected in the body or subject of the email\n> > the user is prompted with,\n> >\n> >   Which 8bit encoding should I declare [UTF-8]? foo\n> >\n> > After this the input string is validated by the regex, based on the fact\n> > that the charset string will be minimum 4 characters [1]. If the string is\n> > more than 4 letters the email is sent, if not then a second prompt to\n> > confirm is asked to the user,\n> >\n> >   Are you sure you want to use <foo> [y/N]? y\n> >\n> > This relies on a length based regex heuristic check to validate the user\n> > input, and can allow clearly invalid charset names to pass if the input is\n> > greater than 4 characters.\n> >\n> > Add a semantic validation of the charset name using the\n> > Encode::find_encoding() module of perl. If the encoding is not recognized,\n> > warn the user and ask for confirmation before proceeding. After this\n> > validation the lenght based validation becomes redundant and also breaks\n> > flow, so change the regex of valid input to any non blank string.\n> >\n> > Additionally, the wording of the first prompt can confuse the user if not\n> > read properly or under any default assumptions for a yes/no prompt. Change\n> > the wording to make it explicitly clear to the user that the prompt needs a\n> > string input, UTF-8 being the default.\n> >\n> > The intended flow is,\n> >\n> >   Declare which 8bit encoding to use [default: UTF-8]? foobar\n> >   warning: 'foobar' does not appear to be a valid charset name.\n> >   Are you sure you want to use <foobar> [y/N]?\n> >\n> > [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n> >\n> > Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> > ---\n> > Changes in v2:\n> >  - Added braces in if-else block.\n> >\n> >  git-send-email.perl   | 17 ++++++++++++++---\n> >  t/t9001-send-email.sh |  2 +-\n> >  2 files changed, 15 insertions(+), 4 deletions(-)\n>\n> Curious.  This change to t9001 was there even in the previous\n> iteration that did not even work.  How did you test it?\n\nInitially I had the braces in place, when I made the wording change to the failing\ntest, so all the tests passed. But just before sending the patch, I saw that the\nindentation looked a bit off in the annotated git send-email, so I fixed that and\nalso removed what I thought were unnecessary braces, but I neglected to rerun the\ntests after that. My bad.\n"},{"id":"537099","messageId":"CALnO6CDSJPnVi-1RUsr7tFMwa0_xTJkiQmzTL_b-BGq=6PSz0A@mail.gmail.com","threadId":"65030","inReplyTo":"20260224213932.92364-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [PATCH v2] send-email: validate charset name in 8bit encoding prompt","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-02-25T16:37:57Z","receivedAt":"2026-02-25T16:38:09Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Tue, Feb 24, 2026 at 4:39 PM Shreyansh Paliwal\n<shreyanshpaliwalcmsmn@gmail.com> wrote:\n>\n> When a non-ASCII character is detected in the body or subject of the email\n> the user is prompted with,\n>\n>   Which 8bit encoding should I declare [UTF-8]? foo\n>\n> After this the input string is validated by the regex, based on the fact\n> that the charset string will be minimum 4 characters [1]. If the string is\n> more than 4 letters the email is sent, if not then a second prompt to\n> confirm is asked to the user,\n>\n>   Are you sure you want to use <foo> [y/N]? y\n>\n> This relies on a length based regex heuristic check to validate the user\n> input, and can allow clearly invalid charset names to pass if the input is\n> greater than 4 characters.\n>\n> Add a semantic validation of the charset name using the\n> Encode::find_encoding() module of perl. If the encoding is not recognized,\n> warn the user and ask for confirmation before proceeding. After this\n> validation the lenght based validation becomes redundant and also breaks\n> flow, so change the regex of valid input to any non blank string.\n>\n> Additionally, the wording of the first prompt can confuse the user if not\n> read properly or under any default assumptions for a yes/no prompt. Change\n> the wording to make it explicitly clear to the user that the prompt needs a\n> string input, UTF-8 being the default.\n>\n> The intended flow is,\n>\n>   Declare which 8bit encoding to use [default: UTF-8]? foobar\n>   warning: 'foobar' does not appear to be a valid charset name.\n>   Are you sure you want to use <foobar> [y/N]?\n>\n> [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n>\n> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> ---\n> Changes in v2:\n>  - Added braces in if-else block.\n>\n>  git-send-email.perl   | 17 ++++++++++++++---\n>  t/t9001-send-email.sh |  2 +-\n>  2 files changed, 15 insertions(+), 4 deletions(-)\n>\n> diff --git a/git-send-email.perl b/git-send-email.perl\n> index cd4b316ddc..15387ac377 100755\n> --- a/git-send-email.perl\n> +++ b/git-send-email.perl\n> @@ -23,6 +23,7 @@\n>  use Git::LoadCPAN::Error qw(:try);\n>  use Git;\n>  use Git::I18N;\n> +use Encode qw(find_encoding);\n>\n>  Getopt::Long::Configure qw/ pass_through /;\n>\n> @@ -987,6 +988,7 @@ sub get_patch_subject {\n>  sub ask {\n>         my ($prompt, %arg) = @_;\n>         my $valid_re = $arg{valid_re};\n> +       my $warn_invalid = $arg{warn_invalid};\n>         my $default = $arg{default};\n>         my $confirm_only = $arg{confirm_only};\n>         my $resp;\n> @@ -1005,7 +1007,15 @@ sub ask {\n>                         return $default;\n>                 }\n>                 if (!defined $valid_re or $resp =~ /$valid_re/) {\n> -                       return $resp;\n> +                       if ($warn_invalid) {\n> +                               if (find_encoding($resp)) {\n> +                                       return $resp;\n> +                               } else {\n> +                                       printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n> +                               }\n> +                       } else {\n> +                               return $resp;\n> +                       }\n\nI think this is asking \"ask\" to do too much, since only encoding\naskers can use warn_invalid.\n\nWhat I rather meant was to extract relevant helper procedures so that\nopen-coding ask around the encoding question would be easier to\nmaintain.\n\n>                 }\n>                 if ($confirm_only) {\n>                         my $yesno = $term->readline(\n> @@ -1044,8 +1054,9 @@ sub file_declares_8bit_cte {\n>         foreach my $f (sort keys %broken_encoding) {\n>                 print \"    $f\\n\";\n>         }\n> -       $auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n> -                                 valid_re => qr/.{4}/, confirm_only => 1,\n> +       $auto_8bit_encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n> +                                 valid_re => qr/^\\S+$/, confirm_only => 1,\n> +                                 warn_invalid => 1,\n>                                   default => \"UTF-8\");\n>  }\n>\n> diff --git a/t/t9001-send-email.sh b/t/t9001-send-email.sh\n> index e56e0c8d77..24f6c76aee 100755\n> --- a/t/t9001-send-email.sh\n> +++ b/t/t9001-send-email.sh\n> @@ -1691,7 +1691,7 @@ test_expect_success $PREREQ 'asks about and fixes 8bit encodings' '\n>                         email-using-8bit >stdout &&\n>         grep \"do not declare a Content-Transfer-Encoding\" stdout &&\n>         grep email-using-8bit stdout &&\n> -       grep \"Which 8bit encoding\" stdout &&\n> +       grep \"Declare which 8bit encoding to use\" stdout &&\n>         grep -E \"Content|MIME\" msgtxt1 >actual &&\n>         test_cmp content-type-decl actual\n>  '\n>\n> Range-diff against v1:\n> 1:  70fa4d2899 ! 1:  954c1dae9f send-email: validate charset name in 8bit encoding prompt\n>     @@ git-send-email.perl: sub ask {\n>                 if (!defined $valid_re or $resp =~ /$valid_re/) {\n>      -                  return $resp;\n>      +                  if ($warn_invalid) {\n>     -+                          if (find_encoding($resp))\n>     ++                          if (find_encoding($resp)) {\n>      +                                  return $resp;\n>     -+                          else\n>     ++                          } else {\n>      +                                  printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n>     -+                  } else\n>     ++                          }\n>     ++                  } else {\n>      +                          return $resp;\n>     ++                  }\n>                 }\n>                 if ($confirm_only) {\n>                         my $yesno = $term->readline(\n> --\n> 2.53.0\n\n\n\n-- \nD. Ben Knoble\n"},{"id":"537208","messageId":"20260226165559.187261-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"20260224143624.23678-1-shreyanshpaliwalcmsmn@gmail.com","subject":"[PATCH v3] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-26T16:16:34Z","receivedAt":"2026-02-26T16:56:18Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"When a non-ASCII character is detected in the body or subject of the email\nthe user is prompted with,\n\n        Which 8bit encoding should I declare [UTF-8]? foo\n\nAfter this the input string is validated by the regex, based on the fact\nthat the charset string will be minimum 4 characters [1]. If the string is\nmore than 4 letters the email is sent, if not then a second prompt to\nconfirm is asked to the user,\n\n        Are you sure you want to use <foo> [y/N]? y\n\nThis relies on a length based regex heuristic check to validate the user\ninput, and can allow clearly invalid charset names to pass if the input is\ngreater than 4 characters.\n\nAdd a semantic validation of the charset name using the\nEncode::find_encoding() module of perl. If the encoding is not recognized,\nwarn the user and ask for confirmation before proceeding. After this\nvalidation the lenght based validation becomes redundant and also breaks\nflow, so change the regex of valid input to any non blank string.\n\nIntroduce a dedicated helper for confirmation handling that can be reused\nboth by ask() and the custom 8bit prompt flow. Make the encoding warning\nlogic specific to the 8bit prompt, this reduces the load on ask(), and\nimproves maintainability.\n\nAdditionally, the wording of the first prompt can confuse the user if not\nread properly or under any default assumptions for a yes/no prompt. Change\nthe wording to make it explicitly clear to the user that the prompt needs a\nstring input, UTF-8 being the default.\n\nThe intended flow is,\n\n        Declare which 8bit encoding to use [default: UTF-8]? foobar\n        warning: 'foobar' does not appear to be a valid charset name.\n        Are you sure you want to use <foobar> [y/N]?\n\n[1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n\nSigned-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n---\nChanges in v2:\n - Added a helper function confirm_ask() to handle yes/no confirmation prompts.\n - Added the validation and warning logic in the 8bit prompt instead of ask().\n\n git-send-email.perl   | 36 +++++++++++++++++++++++++++++-------\n t/t9001-send-email.sh |  2 +-\n 2 files changed, 30 insertions(+), 8 deletions(-)\n\ndiff --git a/git-send-email.perl b/git-send-email.perl\nindex cd4b316ddc..3230b80701 100755\n--- a/git-send-email.perl\n+++ b/git-send-email.perl\n@@ -23,6 +23,7 @@\n use Git::LoadCPAN::Error qw(:try);\n use Git;\n use Git::I18N;\n+use Encode qw(find_encoding);\n\n Getopt::Long::Configure qw/ pass_through /;\n\n@@ -984,6 +985,18 @@ sub get_patch_subject {\n \t}\n }\n\n+sub confirm_ask {\n+\tmy ($resp) = @_;\n+\tmy $term = term();\n+\treturn 0\n+\t\tunless defined $term->IN and defined fileno($term->IN) and\n+\t\t       defined $term->OUT and defined fileno($term->OUT);\n+\tmy $yesno = $term->readline(\n+\t\t# TRANSLATORS: please keep [y/N] as is.\n+\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n+\treturn defined $yesno && $yesno =~ /y/i;\n+}\n+\n sub ask {\n \tmy ($prompt, %arg) = @_;\n \tmy $valid_re = $arg{valid_re};\n@@ -1008,10 +1021,7 @@ sub ask {\n \t\t\treturn $resp;\n \t\t}\n \t\tif ($confirm_only) {\n-\t\t\tmy $yesno = $term->readline(\n-\t\t\t\t# TRANSLATORS: please keep [y/N] as is.\n-\t\t\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n-\t\t\tif (defined $yesno && $yesno =~ /y/i) {\n+\t\t\tif (confirm_ask($resp)) {\n \t\t\t\treturn $resp;\n \t\t\t}\n \t\t}\n@@ -1044,9 +1054,21 @@ sub file_declares_8bit_cte {\n \tforeach my $f (sort keys %broken_encoding) {\n \t\tprint \"    $f\\n\";\n \t}\n-\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n-\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n-\t\t\t\t  default => \"UTF-8\");\n+\twhile(1) {\n+\t\tmy $encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n+\t\tvalid_re => qr/^\\S+$/,\n+\t\tdefault  => \"UTF-8\");\n+\t\tnext unless defined $encoding;\n+\t\tif (find_encoding($encoding)) {\n+\t\t\t$auto_8bit_encoding = $encoding;\n+\t\t\tlast;\n+\t\t}\n+\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n+\t\tif (confirm_ask($encoding)) {\n+\t\t\t$auto_8bit_encoding = $encoding;\n+\t\t\tlast;\n+\t\t}\n+\t}\n }\n\n if (!$force) {\ndiff --git a/t/t9001-send-email.sh b/t/t9001-send-email.sh\nindex e56e0c8d77..24f6c76aee 100755\n--- a/t/t9001-send-email.sh\n+++ b/t/t9001-send-email.sh\n@@ -1691,7 +1691,7 @@ test_expect_success $PREREQ 'asks about and fixes 8bit encodings' '\n \t\t\temail-using-8bit >stdout &&\n \tgrep \"do not declare a Content-Transfer-Encoding\" stdout &&\n \tgrep email-using-8bit stdout &&\n-\tgrep \"Which 8bit encoding\" stdout &&\n+\tgrep \"Declare which 8bit encoding to use\" stdout &&\n \tgrep -E \"Content|MIME\" msgtxt1 >actual &&\n \ttest_cmp content-type-decl actual\n '\n\nRange-diff against v2:\n1:  954c1dae9f ! 1:  748bb03a00 send-email: validate charset name in 8bit encoding prompt\n    @@ Commit message\n         When a non-ASCII character is detected in the body or subject of the email\n         the user is prompted with,\n\n    -      Which 8bit encoding should I declare [UTF-8]? foo\n    +            Which 8bit encoding should I declare [UTF-8]? foo\n\n         After this the input string is validated by the regex, based on the fact\n         that the charset string will be minimum 4 characters [1]. If the string is\n         more than 4 letters the email is sent, if not then a second prompt to\n         confirm is asked to the user,\n\n    -      Are you sure you want to use <foo> [y/N]? y\n    +            Are you sure you want to use <foo> [y/N]? y\n\n         This relies on a length based regex heuristic check to validate the user\n         input, and can allow clearly invalid charset names to pass if the input is\n    @@ Commit message\n         validation the lenght based validation becomes redundant and also breaks\n         flow, so change the regex of valid input to any non blank string.\n\n    +    Introduce a dedicated helper for confirmation handling that can be reused\n    +    both by ask() and the custom 8bit prompt flow. This makes the encoding\n    +    warning logic specific to the 8bit prompt, reduces the load on ask(), and\n    +    improves maintainability.\n    +\n         Additionally, the wording of the first prompt can confuse the user if not\n         read properly or under any default assumptions for a yes/no prompt. Change\n         the wording to make it explicitly clear to the user that the prompt needs a\n         string input, UTF-8 being the default.\n\n    +    The intended flow is,\n    +\n    +            Declare which 8bit encoding to use [default: UTF-8]? foobar\n    +            warning: 'foobar' does not appear to be a valid charset name.\n    +            Are you sure you want to use <foobar> [y/N]?\n    +\n    +    [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n    +\n         Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n\n      ## git-send-email.perl ##\n    @@ git-send-email.perl\n      Getopt::Long::Configure qw/ pass_through /;\n\n     @@ git-send-email.perl: sub get_patch_subject {\n    + \t}\n    + }\n    +\n    ++sub confirm_ask {\n    ++\tmy ($resp) = @_;\n    ++\tmy $term = term();\n    ++\treturn 0\n    ++\t\tunless defined $term->IN and defined fileno($term->IN) and\n    ++\t\t       defined $term->OUT and defined fileno($term->OUT);\n    ++\tmy $yesno = $term->readline(\n    ++\t\t# TRANSLATORS: please keep [y/N] as is.\n    ++\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n    ++\treturn defined $yesno && $yesno =~ /y/i;\n    ++}\n    ++\n      sub ask {\n      \tmy ($prompt, %arg) = @_;\n      \tmy $valid_re = $arg{valid_re};\n    -+\tmy $warn_invalid = $arg{warn_invalid};\n    - \tmy $default = $arg{default};\n    - \tmy $confirm_only = $arg{confirm_only};\n    - \tmy $resp;\n     @@ git-send-email.perl: sub ask {\n    - \t\t\treturn $default;\n    - \t\t}\n    - \t\tif (!defined $valid_re or $resp =~ /$valid_re/) {\n    --\t\t\treturn $resp;\n    -+\t\t\tif ($warn_invalid) {\n    -+\t\t\t\tif (find_encoding($resp)) {\n    -+\t\t\t\t\treturn $resp;\n    -+\t\t\t\t} else {\n    -+\t\t\t\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n    -+\t\t\t\t}\n    -+\t\t\t} else {\n    -+\t\t\t\treturn $resp;\n    -+\t\t\t}\n    + \t\t\treturn $resp;\n      \t\t}\n      \t\tif ($confirm_only) {\n    - \t\t\tmy $yesno = $term->readline(\n    +-\t\t\tmy $yesno = $term->readline(\n    +-\t\t\t\t# TRANSLATORS: please keep [y/N] as is.\n    +-\t\t\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n    +-\t\t\tif (defined $yesno && $yesno =~ /y/i) {\n    ++\t\t\tif (confirm_ask($resp)) {\n    + \t\t\t\treturn $resp;\n    + \t\t\t}\n    + \t\t}\n     @@ git-send-email.perl: sub file_declares_8bit_cte {\n      \tforeach my $f (sort keys %broken_encoding) {\n      \t\tprint \"    $f\\n\";\n      \t}\n     -\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n     -\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n    -+\t$auto_8bit_encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n    -+\t\t\t\t  valid_re => qr/^\\S+$/, confirm_only => 1,\n    -+\t\t\t\t  warn_invalid => 1,\n    - \t\t\t\t  default => \"UTF-8\");\n    +-\t\t\t\t  default => \"UTF-8\");\n    ++\twhile(1) {\n    ++\t\tmy $encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n    ++\t\tvalid_re => qr/^\\S+$/,\n    ++\t\tdefault  => \"UTF-8\");\n    ++\t\tnext unless defined $encoding;\n    ++\t\tif (find_encoding($encoding)) {\n    ++\t\t\t$auto_8bit_encoding = $encoding;\n    ++\t\t\tlast;\n    ++\t\t}\n    ++\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n    ++\t\tif (confirm_ask($encoding)) {\n    ++\t\t\t$auto_8bit_encoding = $encoding;\n    ++\t\t\tlast;\n    ++\t\t}\n    ++\t}\n      }\n\n    + if (!$force) {\n\n      ## t/t9001-send-email.sh ##\n     @@ t/t9001-send-email.sh: test_expect_success $PREREQ 'asks about and fixes 8bit encodings' '\n--\n2.53.0.154.g7c02d39fc2.dirty\n"},{"id":"537211","messageId":"20260226173336.194601-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"CALnO6CDSJPnVi-1RUsr7tFMwa0_xTJkiQmzTL_b-BGq=6PSz0A@mail.gmail.com","subject":"Re: [PATCH v2] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-26T17:32:47Z","receivedAt":"2026-02-26T17:33:54Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> On Tue, Feb 24, 2026 at 4:39 PM Shreyansh Paliwal\n> <shreyanshpaliwalcmsmn@gmail.com> wrote:\n> >\n> > When a non-ASCII character is detected in the body or subject of the email\n> > the user is prompted with,\n> >\n> >   Which 8bit encoding should I declare [UTF-8]? foo\n> >\n> > After this the input string is validated by the regex, based on the fact\n> > that the charset string will be minimum 4 characters [1]. If the string is\n> > more than 4 letters the email is sent, if not then a second prompt to\n> > confirm is asked to the user,\n> >\n> >   Are you sure you want to use <foo> [y/N]? y\n> >\n> > This relies on a length based regex heuristic check to validate the user\n> > input, and can allow clearly invalid charset names to pass if the input is\n> > greater than 4 characters.\n> >\n> > Add a semantic validation of the charset name using the\n> > Encode::find_encoding() module of perl. If the encoding is not recognized,\n> > warn the user and ask for confirmation before proceeding. After this\n> > validation the lenght based validation becomes redundant and also breaks\n> > flow, so change the regex of valid input to any non blank string.\n> >\n> > Additionally, the wording of the first prompt can confuse the user if not\n> > read properly or under any default assumptions for a yes/no prompt. Change\n> > the wording to make it explicitly clear to the user that the prompt needs a\n> > string input, UTF-8 being the default.\n> >\n> > The intended flow is,\n> >\n> >   Declare which 8bit encoding to use [default: UTF-8]? foobar\n> >   warning: 'foobar' does not appear to be a valid charset name.\n> >   Are you sure you want to use <foobar> [y/N]?\n> >\n> > [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n> >\n> > Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> > ---\n> > Changes in v2:\n> >  - Added braces in if-else block.\n> >\n> >  git-send-email.perl   | 17 ++++++++++++++---\n> >  t/t9001-send-email.sh |  2 +-\n> >  2 files changed, 15 insertions(+), 4 deletions(-)\n> >\n> > diff --git a/git-send-email.perl b/git-send-email.perl\n> > index cd4b316ddc..15387ac377 100755\n> > --- a/git-send-email.perl\n> > +++ b/git-send-email.perl\n> > @@ -23,6 +23,7 @@\n> >  use Git::LoadCPAN::Error qw(:try);\n> >  use Git;\n> >  use Git::I18N;\n> > +use Encode qw(find_encoding);\n> >\n> >  Getopt::Long::Configure qw/ pass_through /;\n> >\n> > @@ -987,6 +988,7 @@ sub get_patch_subject {\n> >  sub ask {\n> >         my ($prompt, %arg) = @_;\n> >         my $valid_re = $arg{valid_re};\n> > +       my $warn_invalid = $arg{warn_invalid};\n> >         my $default = $arg{default};\n> >         my $confirm_only = $arg{confirm_only};\n> >         my $resp;\n> > @@ -1005,7 +1007,15 @@ sub ask {\n> >                         return $default;\n> >                 }\n> >                 if (!defined $valid_re or $resp =~ /$valid_re/) {\n> > -                       return $resp;\n> > +                       if ($warn_invalid) {\n> > +                               if (find_encoding($resp)) {\n> > +                                       return $resp;\n> > +                               } else {\n> > +                                       printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $resp;\n> > +                               }\n> > +                       } else {\n> > +                               return $resp;\n> > +                       }\n>\n> I think this is asking \"ask\" to do too much, since only encoding\n> askers can use warn_invalid.\n>\n> What I rather meant was to extract relevant helper procedures so that\n> open-coding ask around the encoding question would be easier to\n> maintain.\n\nHi,\n\nI have sent a v3 on this, in which I introduced a helper for the confirmation\nprompt and also made validation logic specific to the 8bit prompt.\nDo you think it would also be better to move the encoding validation\ninto a separate helper, or does the current split look reasonable?\n\nBest,\nShreyansh\n"},{"id":"537217","messageId":"xmqq8qcf2vk8.fsf@gitster.g","threadId":"65030","inReplyTo":"20260226165559.187261-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [PATCH v3] send-email: validate charset name in 8bit encoding prompt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-26T18:45:11Z","receivedAt":"2026-02-26T18:45:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n\n> diff --git a/git-send-email.perl b/git-send-email.perl\n> index cd4b316ddc..3230b80701 100755\n> --- a/git-send-email.perl\n> +++ b/git-send-email.perl\n> @@ -23,6 +23,7 @@\n>  use Git::LoadCPAN::Error qw(:try);\n>  use Git;\n>  use Git::I18N;\n> +use Encode qw(find_encoding);\n\nI wonder how common is this module already installed on users'\nsystems (not asking \"how widely available\"---which is \"can users\neasily make it work?\", but asking \"would this work out of box with\nwhat users already have?\").\n\n> +sub confirm_ask {\n> +\tmy ($resp) = @_;\n> +\tmy $term = term();\n> +\treturn 0\n> +\t\tunless defined $term->IN and defined fileno($term->IN) and\n> +\t\t       defined $term->OUT and defined fileno($term->OUT);\n> +\tmy $yesno = $term->readline(\n> +\t\t# TRANSLATORS: please keep [y/N] as is.\n> +\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n> +\treturn defined $yesno && $yesno =~ /y/i;\n> +}\n\nThis is a bit incosistent with what \"sub ask\" (the only caller of\nthis sub) does, isn't it?  Before entering the loop that makes a\ncall into this, it does this:\n\n        sub ask {\n                my ($prompt, %arg) = @_;\n                my $valid_re = $arg{valid_re};\n                my $default = $arg{default};\n                my $confirm_only = $arg{confirm_only};\n                my $resp;\n                my $i = 0;\n                my $term = term();\n                return defined $default ? $default : undef\n                        unless defined $term->IN and defined fileno($term->IN) and\n                               defined $term->OUT and defined fileno($term->OUT);\n\nIf $term is not usable for interactive prompt, it uses the default\nsetting.  But the new confirm_ask always says \"no\".\n\nconfirm_ask does its own \"check term() to see it is usable\" because\nit is called from another code path which does not have its own\nlogic, but it may be a wrong abstraction to give uneven interface.\nIt would make it more clear what is going on if you just do the\ninteractive $term->readline() thing in \"sub ask\", instead of calling\n\"sub confirm_ask\" that does tghe $term thing redundantly.\n\nCan't the other confirm_ask() caller call a normal \"sub ask\"?  \n\nI am not sure why we want to add a dedicated sub, just to ask \"are\nyou sure you want to use X [y/N]? \".\n\n> The intended flow is,\n>\n>         Declare which 8bit encoding to use [default: UTF-8]? foobar\n>         warning: 'foobar' does not appear to be a valid charset name.\n>         Are you sure you want to use <foobar> [y/N]?\n\nIt somehow looks uneven to have three lines, two of them\ncapitalizing their first word while the other one is all lowercase.\nI wonder if this would be simpler?\n\n    Declare which 8bit encoding to use [default: UTF-8]?  foobar<RET>\n    Do you really mean 'foobar', not a valid charset name [y/N]?\n\n\n\nSo, taking all of the above together, perhaps:\n\n * Discard changes to \"sub ask\" and addition of \"sub confirm_ask\".\n\n * Tweak this part a bit to call ask().\n\n> +\twhile(1) {\n\nStyle.  missing SP before \"(\".\n\n> +\t\tmy $encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n\nOverly long line.\n\n> +\t\tvalid_re => qr/^\\S+$/,\n> +\t\tdefault  => \"UTF-8\");\n> +\t\tnext unless defined $encoding;\n> +\t\tif (find_encoding($encoding)) {\n> +\t\t\t$auto_8bit_encoding = $encoding;\n> +\t\t\tlast;\n> +\t\t}\n\n> +\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n> +\t\tif (confirm_ask($encoding)) {\n\nUse ask() to ask \n\n    Do you really mean 'foobar', not a valid charset name [y/N]?\n\nhere, perhaps?\n\n> +\t\t\t$auto_8bit_encoding = $encoding;\n> +\t\t\tlast;\n> +\t\t}\n> +\t}\n>  }\n"},{"id":"537218","messageId":"xmqq4in32ulj.fsf@gitster.g","threadId":"65030","inReplyTo":"xmqq8qcf2vk8.fsf@gitster.g","subject":"Re: [PATCH v3] send-email: validate charset name in 8bit encoding prompt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-02-26T19:06:00Z","receivedAt":"2026-02-26T19:06:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n>\n>> diff --git a/git-send-email.perl b/git-send-email.perl\n>> index cd4b316ddc..3230b80701 100755\n>> --- a/git-send-email.perl\n>> +++ b/git-send-email.perl\n>> @@ -23,6 +23,7 @@\n>>  use Git::LoadCPAN::Error qw(:try);\n>>  use Git;\n>>  use Git::I18N;\n>> +use Encode qw(find_encoding);\n>\n> I wonder how common is this module already installed on users'\n> systems (not asking \"how widely available\"---which is \"can users\n> easily make it work?\", but asking \"would this work out of box with\n> what users already have?\").\n\nAnswering my own question: \"yes\".\n\nWe use Encode::find_encoding as well as Encode::{de,en}code in\ngitweb and git-svn, so it is very likely that anybody who has a full\ninstallation of Git would already have it on their system.  Also\nEncode.pm is distributed as part of Perl itself, if I am not\nmistaken.\n\n"},{"id":"537381","messageId":"20260228083803.238503-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"xmqq8qcf2vk8.fsf@gitster.g","subject":"Re: [PATCH v3] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-28T08:36:15Z","receivedAt":"2026-02-28T08:38:16Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"[...]\n> > +sub confirm_ask {\n> > +\tmy ($resp) = @_;\n> > +\tmy $term = term();\n> > +\treturn 0\n> > +\t\tunless defined $term->IN and defined fileno($term->IN) and\n> > +\t\t       defined $term->OUT and defined fileno($term->OUT);\n> > +\tmy $yesno = $term->readline(\n> > +\t\t# TRANSLATORS: please keep [y/N] as is.\n> > +\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n> > +\treturn defined $yesno && $yesno =~ /y/i;\n> > +}\n>\n> This is a bit incosistent with what \"sub ask\" (the only caller of\n> this sub) does, isn't it?  Before entering the loop that makes a\n> call into this, it does this:\n>\n>         sub ask {\n>                 my ($prompt, %arg) = @_;\n>                 my $valid_re = $arg{valid_re};\n>                 my $default = $arg{default};\n>                 my $confirm_only = $arg{confirm_only};\n>                 my $resp;\n>                 my $i = 0;\n>                 my $term = term();\n>                 return defined $default ? $default : undef\n>                         unless defined $term->IN and defined fileno($term->IN) and\n>                                defined $term->OUT and defined fileno($term->OUT);\n>\n> If $term is not usable for interactive prompt, it uses the default\n> setting.  But the new confirm_ask always says \"no\".\n>\n> confirm_ask does its own \"check term() to see it is usable\" because\n> it is called from another code path which does not have its own\n> logic, but it may be a wrong abstraction to give uneven interface.\n> It would make it more clear what is going on if you just do the\n> interactive $term->readline() thing in \"sub ask\", instead of calling\n> \"sub confirm_ask\" that does tghe $term thing redundantly.\n>\n> Can't the other confirm_ask() caller call a normal \"sub ask\"?\n>\n> I am not sure why we want to add a dedicated sub, just to ask \"are\n> you sure you want to use X [y/N]? \".\n>\n> > The intended flow is,\n> >\n> >         Declare which 8bit encoding to use [default: UTF-8]? foobar\n> >         warning: 'foobar' does not appear to be a valid charset name.\n> >         Are you sure you want to use <foobar> [y/N]?\n>\n> It somehow looks uneven to have three lines, two of them\n> capitalizing their first word while the other one is all lowercase.\n> I wonder if this would be simpler?\n>\n>     Declare which 8bit encoding to use [default: UTF-8]?  foobar<RET>\n>     Do you really mean 'foobar', not a valid charset name [y/N]?\n>\n\nActually that makes sense, because if we need to add a special warning\nin between the two prompts (what I was aiming for), either we need to\nmodify ask() to add the warning into the flow, or we had to seperate the\nconfirm_ask because we have to change the flow in any case, but if we\ndrop the additional warning, and instead warn/confirm together in the\nsecond prompt we dont need this abstraction.\n\n>\n>\n> So, taking all of the above together, perhaps:\n>\n>  * Discard changes to \"sub ask\" and addition of \"sub confirm_ask\".\n>\n>  * Tweak this part a bit to call ask().\n>\n> > +\twhile(1) {\n>\n> Style.  missing SP before \"(\".\n>\n> > +\t\tmy $encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n>\n> Overly long line.\n>\n\nmy bad. will fix.\n\n> > +\t\tvalid_re => qr/^\\S+$/,\n> > +\t\tdefault  => \"UTF-8\");\n> > +\t\tnext unless defined $encoding;\n> > +\t\tif (find_encoding($encoding)) {\n> > +\t\t\t$auto_8bit_encoding = $encoding;\n> > +\t\t\tlast;\n> > +\t\t}\n>\n> > +\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n> > +\t\tif (confirm_ask($encoding)) {\n>\n> Use ask() to ask\n>\n>     Do you really mean 'foobar', not a valid charset name [y/N]?\n>\n> here, perhaps?\n>\n\nUnderstood. I am hoping now this doesn't need any additional\nabstraction as Ben suggested.\nSorry for the delay in response to the review.\nI will send a reroll.\n\nThanks,\nShreyansh\n"},{"id":"537382","messageId":"20260228084217.239120-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"xmqq4in32ulj.fsf@gitster.g","subject":"Re: [PATCH v3] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-28T08:41:34Z","receivedAt":"2026-02-28T08:42:28Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n> >\n> >> diff --git a/git-send-email.perl b/git-send-email.perl\n> >> index cd4b316ddc..3230b80701 100755\n> >> --- a/git-send-email.perl\n> >> +++ b/git-send-email.perl\n> >> @@ -23,6 +23,7 @@\n> >>  use Git::LoadCPAN::Error qw(:try);\n> >>  use Git;\n> >>  use Git::I18N;\n> >> +use Encode qw(find_encoding);\n> >\n> > I wonder how common is this module already installed on users'\n> > systems (not asking \"how widely available\"---which is \"can users\n> > easily make it work?\", but asking \"would this work out of box with\n> > what users already have?\").\n>\n> Answering my own question: \"yes\".\n>\n> We use Encode::find_encoding as well as Encode::{de,en}code in\n> gitweb and git-svn, so it is very likely that anybody who has a full\n> installation of Git would already have it on their system.  Also\n> Encode.pm is distributed as part of Perl itself, if I am not\n> mistaken.\n\nThat's right, Encode is bundled with Perl, so users do not need to\ninstall anything extra, other than what is already required for\nbuilding Git.\n"},{"id":"537390","messageId":"20260228112210.270273-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"20260224143624.23678-1-shreyanshpaliwalcmsmn@gmail.com","subject":"[PATCH v4] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-02-28T11:20:45Z","receivedAt":"2026-02-28T11:22:29Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"When a non-ASCII character is detected in the body or subject of the email\nthe user is prompted with,\n\n        Which 8bit encoding should I declare [UTF-8]? foo\n\nAfter this the input string is validated by the regex, based on the fact\nthat the charset string will be minimum 4 characters [1]. If the string is\nmore than 4 letters the email is sent, if not then a second prompt to\nconfirm is asked to the user,\n\n        Are you sure you want to use <foo> [y/N]? y\n\nThis relies on a length based regex heuristic check to validate the user\ninput, and can allow clearly invalid charset names to pass if the input is\ngreater than 4 characters.\n\nAdd a semantic validation of the charset name using the\nEncode::find_encoding() which is a bundled module of perl. If the encoding\nis not recognized, warn the user and ask for confirmation before proceeding.\nAfter this validation the lenght based validation becomes redundant and also\nbreaks flow, so change the regex of valid input to any non blank string.\n\nMake the encoding warning logic specific to the 8bit prompt, also add a\nunique confirmation prompt which  reduces the load on ask(), and improves\nmaintainability.\n\nAdditionally, the wording of the first prompt can confuse the user if not\nread properly or under any default assumptions for a yes/no prompt. Change\nthe wording to make it explicitly clear to the user that the prompt needs a\nstring input, UTF-8 being the default.\n\nThe intended flow is,\n\n        Declare which 8bit encoding to use [default: UTF-8]? foobar\n        <foobar> does not appear to be a valid charset name. Use it anyway [y/N]?\n\n[1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n\nSigned-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n---\nChanges in v4:\n - removed the confirm_ask() helper and changes to ask().\n - make a new warning/confirmation prompt specific to the 8bit encoding flow.\n\n git-send-email.perl   | 25 ++++++++++++++++++++++---\n t/t9001-send-email.sh |  2 +-\n 2 files changed, 23 insertions(+), 4 deletions(-)\n\ndiff --git a/git-send-email.perl b/git-send-email.perl\nindex cd4b316ddc..3186104709 100755\n--- a/git-send-email.perl\n+++ b/git-send-email.perl\n@@ -23,6 +23,7 @@\n use Git::LoadCPAN::Error qw(:try);\n use Git;\n use Git::I18N;\n+use Encode qw(find_encoding);\n\n Getopt::Long::Configure qw/ pass_through /;\n\n@@ -1044,9 +1045,27 @@ sub file_declares_8bit_cte {\n \tforeach my $f (sort keys %broken_encoding) {\n \t\tprint \"    $f\\n\";\n \t}\n-\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n-\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n-\t\t\t\t  default => \"UTF-8\");\n+\twhile (1) {\n+\t\tmy $encoding = ask(\n+\t\t\t__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n+\t\t\tvalid_re => qr/^\\S+$/,\n+\t\t\tdefault  => \"UTF-8\");\n+\t\tnext unless defined $encoding;\n+\t\tif (find_encoding($encoding)) {\n+\t\t\t$auto_8bit_encoding = $encoding;\n+\t\t\tlast;\n+\t\t}\n+\t\tmy $yesno = ask(\n+\t\t\tsprintf(\n+\t\t\t__(\"'%s' does not appear to be a valid charset name. Use it anyway [y/N]? \"),\n+\t\t\t$encoding),\n+\t\t\tvalid_re => qr/^(?:y|n)/i,\n+\t\t\tdefault => \"n\");\n+\t\tif (defined $yesno && $yesno =~ /^y/i) {\n+\t\t\t$auto_8bit_encoding = $encoding;\n+\t\t\tlast;\n+\t\t}\n+\t}\n }\n\n if (!$force) {\ndiff --git a/t/t9001-send-email.sh b/t/t9001-send-email.sh\nindex e56e0c8d77..24f6c76aee 100755\n--- a/t/t9001-send-email.sh\n+++ b/t/t9001-send-email.sh\n@@ -1691,7 +1691,7 @@ test_expect_success $PREREQ 'asks about and fixes 8bit encodings' '\n \t\t\temail-using-8bit >stdout &&\n \tgrep \"do not declare a Content-Transfer-Encoding\" stdout &&\n \tgrep email-using-8bit stdout &&\n-\tgrep \"Which 8bit encoding\" stdout &&\n+\tgrep \"Declare which 8bit encoding to use\" stdout &&\n \tgrep -E \"Content|MIME\" msgtxt1 >actual &&\n \ttest_cmp content-type-decl actual\n '\n\nRange-diff against v3:\n1:  748bb03a00 ! 1:  37e17eac68 send-email: validate charset name in 8bit encoding prompt\n    @@ Commit message\n         validation the lenght based validation becomes redundant and also breaks\n         flow, so change the regex of valid input to any non blank string.\n\n    -    Introduce a dedicated helper for confirmation handling that can be reused\n    -    both by ask() and the custom 8bit prompt flow. This makes the encoding\n    -    warning logic specific to the 8bit prompt, reduces the load on ask(), and\n    -    improves maintainability.\n    +    Make the encoding warning logic specific to the 8bit prompt, also add a\n    +    unique confirmation prompt which  reduces the load on ask(), and improves\n    +    maintainability.\n\n         Additionally, the wording of the first prompt can confuse the user if not\n         read properly or under any default assumptions for a yes/no prompt. Change\n    @@ Commit message\n         The intended flow is,\n\n                 Declare which 8bit encoding to use [default: UTF-8]? foobar\n    -            warning: 'foobar' does not appear to be a valid charset name.\n    -            Are you sure you want to use <foobar> [y/N]?\n    +            <foobar> does not appear to be a valid charset name. Use it anyway [y/N]?\n\n         [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n\n    @@ git-send-email.perl\n\n      Getopt::Long::Configure qw/ pass_through /;\n\n    -@@ git-send-email.perl: sub get_patch_subject {\n    - \t}\n    - }\n    -\n    -+sub confirm_ask {\n    -+\tmy ($resp) = @_;\n    -+\tmy $term = term();\n    -+\treturn 0\n    -+\t\tunless defined $term->IN and defined fileno($term->IN) and\n    -+\t\t       defined $term->OUT and defined fileno($term->OUT);\n    -+\tmy $yesno = $term->readline(\n    -+\t\t# TRANSLATORS: please keep [y/N] as is.\n    -+\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n    -+\treturn defined $yesno && $yesno =~ /y/i;\n    -+}\n    -+\n    - sub ask {\n    - \tmy ($prompt, %arg) = @_;\n    - \tmy $valid_re = $arg{valid_re};\n    -@@ git-send-email.perl: sub ask {\n    - \t\t\treturn $resp;\n    - \t\t}\n    - \t\tif ($confirm_only) {\n    --\t\t\tmy $yesno = $term->readline(\n    --\t\t\t\t# TRANSLATORS: please keep [y/N] as is.\n    --\t\t\t\tsprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n    --\t\t\tif (defined $yesno && $yesno =~ /y/i) {\n    -+\t\t\tif (confirm_ask($resp)) {\n    - \t\t\t\treturn $resp;\n    - \t\t\t}\n    - \t\t}\n     @@ git-send-email.perl: sub file_declares_8bit_cte {\n      \tforeach my $f (sort keys %broken_encoding) {\n      \t\tprint \"    $f\\n\";\n    @@ git-send-email.perl: sub file_declares_8bit_cte {\n     -\t$auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n     -\t\t\t\t  valid_re => qr/.{4}/, confirm_only => 1,\n     -\t\t\t\t  default => \"UTF-8\");\n    -+\twhile(1) {\n    -+\t\tmy $encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n    -+\t\tvalid_re => qr/^\\S+$/,\n    -+\t\tdefault  => \"UTF-8\");\n    ++\twhile (1) {\n    ++\t\tmy $encoding = ask(\n    ++\t\t\t__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n    ++\t\t\tvalid_re => qr/^\\S+$/,\n    ++\t\t\tdefault  => \"UTF-8\");\n     +\t\tnext unless defined $encoding;\n     +\t\tif (find_encoding($encoding)) {\n     +\t\t\t$auto_8bit_encoding = $encoding;\n     +\t\t\tlast;\n     +\t\t}\n    -+\t\tprintf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n    -+\t\tif (confirm_ask($encoding)) {\n    ++\t\tmy $yesno = ask(\n    ++\t\t\tsprintf(\n    ++\t\t\t__(\"'%s' does not appear to be a valid charset name. Use it anyway [y/N]? \"),\n    ++\t\t\t$encoding),\n    ++\t\t\tvalid_re => qr/^(?:y|n)/i,\n    ++\t\t\tdefault => \"n\");\n    ++\t\tif (defined $yesno && $yesno =~ /^y/i) {\n     +\t\t\t$auto_8bit_encoding = $encoding;\n     +\t\t\tlast;\n     +\t\t}\n--\n2.53.0.155.g748bb03a00.dirty\n"},{"id":"537409","messageId":"CALnO6CD0jvtaTpvNHQvvpDUVXZmzp9cq9oiuDMyX2BPt0ibFYw@mail.gmail.com","threadId":"65030","inReplyTo":"20260228112210.270273-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [PATCH v4] send-email: validate charset name in 8bit encoding prompt","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-02-28T21:16:16Z","receivedAt":"2026-02-28T21:16:27Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Sat, Feb 28, 2026 at 6:22 AM Shreyansh Paliwal\n<shreyanshpaliwalcmsmn@gmail.com> wrote:\n>\n> When a non-ASCII character is detected in the body or subject of the email\n> the user is prompted with,\n>\n>         Which 8bit encoding should I declare [UTF-8]? foo\n>\n> After this the input string is validated by the regex, based on the fact\n> that the charset string will be minimum 4 characters [1]. If the string is\n> more than 4 letters the email is sent, if not then a second prompt to\n> confirm is asked to the user,\n>\n>         Are you sure you want to use <foo> [y/N]? y\n>\n> This relies on a length based regex heuristic check to validate the user\n> input, and can allow clearly invalid charset names to pass if the input is\n> greater than 4 characters.\n>\n> Add a semantic validation of the charset name using the\n> Encode::find_encoding() which is a bundled module of perl. If the encoding\n> is not recognized, warn the user and ask for confirmation before proceeding.\n> After this validation the lenght based validation becomes redundant and also\n> breaks flow, so change the regex of valid input to any non blank string.\n>\n> Make the encoding warning logic specific to the 8bit prompt, also add a\n> unique confirmation prompt which  reduces the load on ask(), and improves\n> maintainability.\n>\n> Additionally, the wording of the first prompt can confuse the user if not\n> read properly or under any default assumptions for a yes/no prompt. Change\n> the wording to make it explicitly clear to the user that the prompt needs a\n> string input, UTF-8 being the default.\n>\n> The intended flow is,\n>\n>         Declare which 8bit encoding to use [default: UTF-8]? foobar\n>         <foobar> does not appear to be a valid charset name. Use it anyway [y/N]?\n>\n> [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n>\n> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> ---\n> Changes in v4:\n>  - removed the confirm_ask() helper and changes to ask().\n>  - make a new warning/confirmation prompt specific to the 8bit encoding flow.\n>\n>  git-send-email.perl   | 25 ++++++++++++++++++++++---\n>  t/t9001-send-email.sh |  2 +-\n>  2 files changed, 23 insertions(+), 4 deletions(-)\n>\n> diff --git a/git-send-email.perl b/git-send-email.perl\n> index cd4b316ddc..3186104709 100755\n> --- a/git-send-email.perl\n> +++ b/git-send-email.perl\n> @@ -23,6 +23,7 @@\n>  use Git::LoadCPAN::Error qw(:try);\n>  use Git;\n>  use Git::I18N;\n> +use Encode qw(find_encoding);\n>\n>  Getopt::Long::Configure qw/ pass_through /;\n>\n> @@ -1044,9 +1045,27 @@ sub file_declares_8bit_cte {\n>         foreach my $f (sort keys %broken_encoding) {\n>                 print \"    $f\\n\";\n>         }\n> -       $auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n> -                                 valid_re => qr/.{4}/, confirm_only => 1,\n> -                                 default => \"UTF-8\");\n> +       while (1) {\n> +               my $encoding = ask(\n> +                       __(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n> +                       valid_re => qr/^\\S+$/,\n> +                       default  => \"UTF-8\");\n> +               next unless defined $encoding;\n> +               if (find_encoding($encoding)) {\n> +                       $auto_8bit_encoding = $encoding;\n> +                       last;\n> +               }\n> +               my $yesno = ask(\n> +                       sprintf(\n> +                       __(\"'%s' does not appear to be a valid charset name. Use it anyway [y/N]? \"),\n> +                       $encoding),\n> +                       valid_re => qr/^(?:y|n)/i,\n> +                       default => \"n\");\n> +               if (defined $yesno && $yesno =~ /^y/i) {\n> +                       $auto_8bit_encoding = $encoding;\n> +                       last;\n> +               }\n> +       }\n>  }\n>\n>  if (!$force) {\n> diff --git a/t/t9001-send-email.sh b/t/t9001-send-email.sh\n> index e56e0c8d77..24f6c76aee 100755\n> --- a/t/t9001-send-email.sh\n> +++ b/t/t9001-send-email.sh\n> @@ -1691,7 +1691,7 @@ test_expect_success $PREREQ 'asks about and fixes 8bit encodings' '\n>                         email-using-8bit >stdout &&\n>         grep \"do not declare a Content-Transfer-Encoding\" stdout &&\n>         grep email-using-8bit stdout &&\n> -       grep \"Which 8bit encoding\" stdout &&\n> +       grep \"Declare which 8bit encoding to use\" stdout &&\n>         grep -E \"Content|MIME\" msgtxt1 >actual &&\n>         test_cmp content-type-decl actual\n>  '\n>\n> Range-diff against v3:\n> 1:  748bb03a00 ! 1:  37e17eac68 send-email: validate charset name in 8bit encoding prompt\n>     @@ Commit message\n>          validation the lenght based validation becomes redundant and also breaks\n>          flow, so change the regex of valid input to any non blank string.\n>\n>     -    Introduce a dedicated helper for confirmation handling that can be reused\n>     -    both by ask() and the custom 8bit prompt flow. This makes the encoding\n>     -    warning logic specific to the 8bit prompt, reduces the load on ask(), and\n>     -    improves maintainability.\n>     +    Make the encoding warning logic specific to the 8bit prompt, also add a\n>     +    unique confirmation prompt which  reduces the load on ask(), and improves\n>     +    maintainability.\n>\n>          Additionally, the wording of the first prompt can confuse the user if not\n>          read properly or under any default assumptions for a yes/no prompt. Change\n>     @@ Commit message\n>          The intended flow is,\n>\n>                  Declare which 8bit encoding to use [default: UTF-8]? foobar\n>     -            warning: 'foobar' does not appear to be a valid charset name.\n>     -            Are you sure you want to use <foobar> [y/N]?\n>     +            <foobar> does not appear to be a valid charset name. Use it anyway [y/N]?\n>\n>          [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n>\n>     @@ git-send-email.perl\n>\n>       Getopt::Long::Configure qw/ pass_through /;\n>\n>     -@@ git-send-email.perl: sub get_patch_subject {\n>     -   }\n>     - }\n>     -\n>     -+sub confirm_ask {\n>     -+  my ($resp) = @_;\n>     -+  my $term = term();\n>     -+  return 0\n>     -+          unless defined $term->IN and defined fileno($term->IN) and\n>     -+                 defined $term->OUT and defined fileno($term->OUT);\n>     -+  my $yesno = $term->readline(\n>     -+          # TRANSLATORS: please keep [y/N] as is.\n>     -+          sprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n>     -+  return defined $yesno && $yesno =~ /y/i;\n>     -+}\n>     -+\n>     - sub ask {\n>     -   my ($prompt, %arg) = @_;\n>     -   my $valid_re = $arg{valid_re};\n>     -@@ git-send-email.perl: sub ask {\n>     -                   return $resp;\n>     -           }\n>     -           if ($confirm_only) {\n>     --                  my $yesno = $term->readline(\n>     --                          # TRANSLATORS: please keep [y/N] as is.\n>     --                          sprintf(__(\"Are you sure you want to use <%s> [y/N]? \"), $resp));\n>     --                  if (defined $yesno && $yesno =~ /y/i) {\n>     -+                  if (confirm_ask($resp)) {\n>     -                           return $resp;\n>     -                   }\n>     -           }\n>      @@ git-send-email.perl: sub file_declares_8bit_cte {\n>         foreach my $f (sort keys %broken_encoding) {\n>                 print \"    $f\\n\";\n>     @@ git-send-email.perl: sub file_declares_8bit_cte {\n>      -  $auto_8bit_encoding = ask(__(\"Which 8bit encoding should I declare [UTF-8]? \"),\n>      -                            valid_re => qr/.{4}/, confirm_only => 1,\n>      -                            default => \"UTF-8\");\n>     -+  while(1) {\n>     -+          my $encoding = ask(__(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n>     -+          valid_re => qr/^\\S+$/,\n>     -+          default  => \"UTF-8\");\n>     ++  while (1) {\n>     ++          my $encoding = ask(\n>     ++                  __(\"Declare which 8bit encoding to use [default: UTF-8]? \"),\n>     ++                  valid_re => qr/^\\S+$/,\n>     ++                  default  => \"UTF-8\");\n>      +          next unless defined $encoding;\n>      +          if (find_encoding($encoding)) {\n>      +                  $auto_8bit_encoding = $encoding;\n>      +                  last;\n>      +          }\n>     -+          printf STDERR __(\"warning: '%s' does not appear to be a valid charset name.\\n\"), $encoding;\n>     -+          if (confirm_ask($encoding)) {\n>     ++          my $yesno = ask(\n>     ++                  sprintf(\n>     ++                  __(\"'%s' does not appear to be a valid charset name. Use it anyway [y/N]? \"),\n>     ++                  $encoding),\n>     ++                  valid_re => qr/^(?:y|n)/i,\n>     ++                  default => \"n\");\n>     ++          if (defined $yesno && $yesno =~ /^y/i) {\n>      +                  $auto_8bit_encoding = $encoding;\n>      +                  last;\n>      +          }\n> --\n> 2.53.0.155.g748bb03a00.dirty\n\nThis version looks nice to my eyes. Thanks!\n\n-- \nD. Ben Knoble\n"},{"id":"537544","messageId":"xmqqo6l643ga.fsf@gitster.g","threadId":"65030","inReplyTo":"20260228112210.270273-1-shreyanshpaliwalcmsmn@gmail.com","subject":"Re: [PATCH v4] send-email: validate charset name in 8bit encoding prompt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-02T16:10:45Z","receivedAt":"2026-03-02T16:10:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com> writes:\n\n> Additionally, the wording of the first prompt can confuse the user if not\n> read properly or under any default assumptions for a yes/no prompt. Change\n> the wording to make it explicitly clear to the user that the prompt needs a\n> string input, UTF-8 being the default.\n>\n> The intended flow is,\n>\n>         Declare which 8bit encoding to use [default: UTF-8]? foobar\n>         <foobar> does not appear to be a valid charset name. Use it anyway [y/N]?\n>\n> [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n>\n> Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> ---\n> Changes in v4:\n>  - removed the confirm_ask() helper and changes to ask().\n>  - make a new warning/confirmation prompt specific to the 8bit encoding flow.\n\nLooking quite straight-forward.  Will replace.\n\nShall we declare victory and mark the topic for 'next'?\n\nThanks.\n"},{"id":"537721","messageId":"20260303190713.153825-1-shreyanshpaliwalcmsmn@gmail.com","threadId":"65030","inReplyTo":"xmqqo6l643ga.fsf@gitster.g","subject":"Re: [PATCH v4] send-email: validate charset name in 8bit encoding prompt","fromName":"Shreyansh Paliwal","fromEmail":"shreyanshpaliwalcmsmn@gmail.com","sentAt":"2026-03-03T19:06:58Z","receivedAt":"2026-03-03T19:07:25Z","isPatch":true,"sender":{"key":"shreyanshpaliwalcmsmn@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152720574?v=4"},"body":"> > Additionally, the wording of the first prompt can confuse the user if not\n> > read properly or under any default assumptions for a yes/no prompt. Change\n> > the wording to make it explicitly clear to the user that the prompt needs a\n> > string input, UTF-8 being the default.\n> >\n> > The intended flow is,\n> >\n> >         Declare which 8bit encoding to use [default: UTF-8]? foobar\n> >         <foobar> does not appear to be a valid charset name. Use it anyway [y/N]?\n> >\n> > [1]- https://github.com/git/git/commit/852a15d748034eec87adbee73a72689c8936fb8b\n> >\n> > Signed-off-by: Shreyansh Paliwal <shreyanshpaliwalcmsmn@gmail.com>\n> > ---\n> > Changes in v4:\n> >  - removed the confirm_ask() helper and changes to ask().\n> >  - make a new warning/confirmation prompt specific to the 8bit encoding flow.\n>\n> Looking quite straight-forward.  Will replace.\n>\n> Shall we declare victory and mark the topic for 'next'?\n>\n> Thanks.\n\nYup, it is good to go from my side, presumably from Ben too.\n\nBest,\nShreyansh\n"}]}