{"thread":{"id":"64880","subject":"[RFC] config --get-regexp: avoid rewriting regex patterns; consider REG_ICASE","startedAt":"2026-01-28T14:12:48Z","lastAt":"2026-01-29T12:22:50Z","messageCount":3,"participants":["Pushkar Singh","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"534761","messageId":"CALE2CrQD11Qa+wGVhsF8JwkuwkLWkDf9kGvs1NM2dsYFuPgUKA@mail.gmail.com","threadId":"64880","inReplyTo":null,"subject":"[RFC] config --get-regexp: avoid rewriting regex patterns; consider REG_ICASE","fromName":"Pushkar Singh","fromEmail":"pushkarkumarsingh1970@gmail.com","sentAt":"2026-01-28T14:12:34Z","receivedAt":"2026-01-28T14:12:48Z","isPatch":false,"sender":{"key":"pushkarkumarsingh1970@gmail.com","avatar":"https://avatars.githubusercontent.com/u/173247767?v=4"},"body":"Hi,\n\nWhile looking at builtin/config.c I noticed the following NEEDSWORK comment\nin get_value():\n\n  /*\n   * NEEDSWORK: this naive pattern lowercasing obviously does not\n   * work for more complex patterns like \"^[^.]*Foo.*\".\n   */\n\nCurrently, git config --get-regexp emulates case-insensitive matching by\nlowercasing parts of the user-provided regex before compiling it. This\nbreaks valid regular expressions and makes it impossible to express more\ncomplex patterns.\n\nFor example:\n\n  git config --add Foo.Bar baz\n  git config --add foo.Baz qux\n  git config --get-regexp '^[^.]*Foo.*'\n\ndoes not behave as expected because the pattern is rewritten before\nregcomp().\n\nPOSIX regex also does not support inline modifiers like (?i), so users\ncurrently have no way to explicitly request case-insensitive matching.\n\nThe documentation says matching is performed against a canonicalized\nlowercase key, but the current implementation achieves this by modifying\nthe regex itself.\n\nWould it make sense to stop rewriting the pattern and instead use REG_ICASE\nwhen compiling the regex? This would preserve user-provided regexes, support\nmore complex expressions, simplify the code, and eliminate the NEEDSWORK.\n\nIf this direction sounds reasonable, I’d be happy to follow up with a patch.\n\nThanks,\nPushkar\n"},{"id":"534808","messageId":"20260129112446.GB1285720@coredump.intra.peff.net","threadId":"64880","inReplyTo":"CALE2CrQD11Qa+wGVhsF8JwkuwkLWkDf9kGvs1NM2dsYFuPgUKA@mail.gmail.com","subject":"Re: [RFC] config --get-regexp: avoid rewriting regex patterns; consider REG_ICASE","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-01-29T11:24:46Z","receivedAt":"2026-01-29T11:24:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jan 28, 2026 at 07:42:34PM +0530, Pushkar Singh wrote:\n\n> The documentation says matching is performed against a canonicalized\n> lowercase key, but the current implementation achieves this by modifying\n> the regex itself.\n> \n> Would it make sense to stop rewriting the pattern and instead use REG_ICASE\n> when compiling the regex? This would preserve user-provided regexes, support\n> more complex expressions, simplify the code, and eliminate the NEEDSWORK.\n> \n> If this direction sounds reasonable, I’d be happy to follow up with a patch.\n\nNo, I don't think that would yield correct results, because the whole\nconfig key is not case-insensitive. The \"subsection\" (the middle part of\na key with two dots, like \"section.SubSection.key\") is case sensitive.\nThat's why we only lowercase the regex up to the first dot (and after\nthe last dot).\n\nSo as a concrete example:\n\n  git config foo.Bar.baz value\n  git config --get-regexp foo.bar.baz\n\nshould not match (and does not currently). Whereas:\n\n  git config --get-regexp foo.Bar.baz\n\nwould (and does).\n\nIf anything, I think we should consider deprecating the auto-lowercasing\nof the regex.\n\nThe \"right\" thing for callers to do is to downcase their regexes\nthemselves, in order to match the canonicalized name (which we do\ndocument; I think it was added around the same time as the comment you\nfound). So any lowercasing we do is a favor to callers to make their\nlives easier. The fact that we can't do it as thoroughly as possible is\nperhaps OK, but I also think we could actually screw up their regex in\nsome extreme cases (say, by thinking we found a dot as a section\nseparator that isn't really one). Which is gross, but nobody seems to\nhave cared too much.\n\nOTOH, it does help the regex queries match the regular ones. We will\ncanonicalize \"git config Foo.Bar.Baz\" into \"foo.Bar.baz\", which we can\ndo unambiguously. And especially since we document names as camelCase,\npeople tend to write things like \"fetch.unpackLimit\" in their queries. I\ndon't think we'd ever want to stop making that work (though\ninterestingly, I do not think we document that anywhere).\n\nSo probably I'd do nothing. ;)\n\n-Peff\n"},{"id":"534812","messageId":"CALE2CrQVnv7wcvD+ewDD2js7=mUE4soLmsud1U2AoorBRwAoNw@mail.gmail.com","threadId":"64880","inReplyTo":"20260129112446.GB1285720@coredump.intra.peff.net","subject":"Re: [RFC] config --get-regexp: avoid rewriting regex patterns; consider REG_ICASE","fromName":"Pushkar Singh","fromEmail":"pushkarkumarsingh1970@gmail.com","sentAt":"2026-01-29T12:22:39Z","receivedAt":"2026-01-29T12:22:50Z","isPatch":false,"sender":{"key":"pushkarkumarsingh1970@gmail.com","avatar":"https://avatars.githubusercontent.com/u/173247767?v=4"},"body":"Thanks for the detailed explanation, that makes sense.\n\nI had missed the subsection case-sensitivity, so REG_ICASE would\nindeed be incorrect.\n\nAppreciate the context on the historical behavior and tradeoffs here.\nI'll drop this for now.\n\nThanks,\nPushkar\n\nOn Thu, Jan 29, 2026 at 4:54 PM Jeff King <peff@peff.net> wrote:\n>\n> On Wed, Jan 28, 2026 at 07:42:34PM +0530, Pushkar Singh wrote:\n>\n> > The documentation says matching is performed against a canonicalized\n> > lowercase key, but the current implementation achieves this by modifying\n> > the regex itself.\n> >\n> > Would it make sense to stop rewriting the pattern and instead use REG_ICASE\n> > when compiling the regex? This would preserve user-provided regexes, support\n> > more complex expressions, simplify the code, and eliminate the NEEDSWORK.\n> >\n> > If this direction sounds reasonable, I’d be happy to follow up with a patch.\n>\n> No, I don't think that would yield correct results, because the whole\n> config key is not case-insensitive. The \"subsection\" (the middle part of\n> a key with two dots, like \"section.SubSection.key\") is case sensitive.\n> That's why we only lowercase the regex up to the first dot (and after\n> the last dot).\n>\n> So as a concrete example:\n>\n>   git config foo.Bar.baz value\n>   git config --get-regexp foo.bar.baz\n>\n> should not match (and does not currently). Whereas:\n>\n>   git config --get-regexp foo.Bar.baz\n>\n> would (and does).\n>\n> If anything, I think we should consider deprecating the auto-lowercasing\n> of the regex.\n>\n> The \"right\" thing for callers to do is to downcase their regexes\n> themselves, in order to match the canonicalized name (which we do\n> document; I think it was added around the same time as the comment you\n> found). So any lowercasing we do is a favor to callers to make their\n> lives easier. The fact that we can't do it as thoroughly as possible is\n> perhaps OK, but I also think we could actually screw up their regex in\n> some extreme cases (say, by thinking we found a dot as a section\n> separator that isn't really one). Which is gross, but nobody seems to\n> have cared too much.\n>\n> OTOH, it does help the regex queries match the regular ones. We will\n> canonicalize \"git config Foo.Bar.Baz\" into \"foo.Bar.baz\", which we can\n> do unambiguously. And especially since we document names as camelCase,\n> people tend to write things like \"fetch.unpackLimit\" in their queries. I\n> don't think we'd ever want to stop making that work (though\n> interestingly, I do not think we document that anywhere).\n>\n> So probably I'd do nothing. ;)\n>\n> -Peff\n"}]}