{"thread":{"id":"66087","subject":"[RFC PATCH 0/6] Git 3.0: restrict hex object IDs to lowercase only","startedAt":"2026-07-29T23:32:19Z","lastAt":"2026-09-07T20:00:01Z","messageCount":41,"participants":["brian m. carlson","Junio C Hamano","Jeff King","Michael Montalbo","Phillip Wood","Elijah Newren"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"549246","messageId":"20260729233215.398654-1-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":null,"subject":"[RFC PATCH 0/6] Git 3.0: restrict hex object IDs to lowercase only","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T23:32:09Z","receivedAt":"2026-07-29T23:32:19Z","isPatch":true,"body":"As far as I can tell, Git has always emitted hex object IDs in\nlowercase, but our object ID parser accepts both uppercase and\nlowercase.  This leads to much software relying on hex object IDs being\nbroken because it doesn't handle uppercase object IDs and this can even\nlead to security problems when people assume that an object ID has a\nunique hex form.\n\nThis series proposes to remove the ability to use uppercase hex in\nobject IDs in Git 3.0.  It is RFC simply because it's not clear if\nthere's the desire to do this, although the series should be fully\nfunctional.\n\nAs further evidence of why we should do this, I'll note that there is\nexactly one testcase in our testsuite that fails due to this change\n(fixed in the last patch) and it's not clear that it fails\nintentionally.  If we decide not to adopt this series, it would probably\nbe prudent to add some additional tests for the uppercase variant of hex\nobject IDs.\n\nbrian m. carlson (6):\n  hex: add functionality for lowercase-only hex\n  hex: allow specifying hex type with hex2chr\n  hex: make hex_to_bytes accept kind of hex to use\n  hex: label usages of hex parsing for object IDs\n  object-name: use hexval\n  hex: allow only lowercase object IDs in breaking changes mode\n\n Documentation/BreakingChanges.adoc |  5 ++++\n builtin/index-pack.c               |  2 +-\n color.c                            |  2 +-\n diagnose.c                         |  2 +-\n hex-ll.c                           | 39 ++++++++++++++++++++++++++++--\n hex-ll.h                           | 24 +++++++++++++-----\n hex.c                              |  2 +-\n http-push.c                        |  5 ++--\n mailinfo.c                         |  2 +-\n notes.c                            |  5 ++--\n object-file.c                      |  2 +-\n object-name.c                      | 13 +++-------\n pkt-line.c                         |  8 +++---\n ref-filter.c                       |  2 +-\n strbuf.c                           |  2 +-\n t/t1503-rev-parse-verify.sh        |  5 ++++\n t/t5324-split-commit-graph.sh      |  4 +--\n url.c                              |  2 +-\n urlmatch.c                         |  2 +-\n 19 files changed, 90 insertions(+), 38 deletions(-)\n\n"},{"id":"549247","messageId":"20260729233215.398654-3-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"[RFC PATCH 2/6] hex: allow specifying hex type with hex2chr","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T23:32:11Z","receivedAt":"2026-07-29T23:32:19Z","isPatch":true,"body":"We have several places where we use hex2chr.  One of those is parsing\nobject IDs, but others decode quoted-printable or percent encoding.  All\nof them accept both uppercase and lowercase hex.\n\nIn a future commit, we'll change some of these cases, so make hex2chr\naccept the kind of encoding to use: lowercase only hex or any kind of\nhex.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n hex-ll.h     | 6 +++---\n hex.c        | 2 +-\n mailinfo.c   | 2 +-\n ref-filter.c | 2 +-\n strbuf.c     | 2 +-\n url.c        | 2 +-\n urlmatch.c   | 2 +-\n 7 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/hex-ll.h b/hex-ll.h\nindex da1b5239b2..26847c7b2f 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -17,10 +17,10 @@ static inline unsigned int hexval(unsigned char c, enum hexkind kind)\n  * Convert two consecutive hexadecimal digits into a char.  Return a\n  * negative value on error.  Don't run over the end of short strings.\n  */\n-static inline int hex2chr(const char *s)\n+static inline int hex2chr(const char *s, enum hexkind kind)\n {\n-\tunsigned int val = hexval(s[0], HEX_KIND_MIXED);\n-\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1], HEX_KIND_MIXED);\n+\tunsigned int val = hexval(s[0], kind);\n+\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1], kind);\n }\n \n /*\ndiff --git a/hex.c b/hex.c\nindex f02832140d..6150bdcbf8 100644\n--- a/hex.c\n+++ b/hex.c\n@@ -9,7 +9,7 @@ static int get_hash_hex_algop(const char *hex, unsigned char *hash,\n \t\t\t      const struct git_hash_algo *algop)\n {\n \tfor (size_t i = 0; i < algop->rawsz; i++) {\n-\t\tint val = hex2chr(hex);\n+\t\tint val = hex2chr(hex, HEX_KIND_MIXED);\n \t\tif (val < 0)\n \t\t\treturn -1;\n \t\t*hash++ = val;\ndiff --git a/mailinfo.c b/mailinfo.c\nindex 13949ff31e..85c3119048 100644\n--- a/mailinfo.c\n+++ b/mailinfo.c\n@@ -396,7 +396,7 @@ static int decode_q_segment(struct strbuf *out, const struct strbuf *q_seg,\n \t\t\tint ch, d = *in;\n \t\t\tif (d == '\\n' || !d)\n \t\t\t\tbreak; /* drop trailing newline */\n-\t\t\tch = hex2chr(in);\n+\t\t\tch = hex2chr(in, HEX_KIND_MIXED);\n \t\t\tif (ch >= 0) {\n \t\t\t\tstrbuf_addch(out, ch);\n \t\t\t\tin += 2;\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 29aca08ce7..884bcd8fc5 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -3567,7 +3567,7 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n \t\t\tif (cp[1] == '%')\n \t\t\t\tcp++;\n \t\t\telse {\n-\t\t\t\tint ch = hex2chr(cp + 1);\n+\t\t\t\tint ch = hex2chr(cp + 1, HEX_KIND_MIXED);\n \t\t\t\tif (0 <= ch) {\n \t\t\t\t\tstrbuf_addch(s, ch);\n \t\t\t\t\tcp += 3;\ndiff --git a/strbuf.c b/strbuf.c\nindex 44955669e8..88d23f8ac5 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -457,7 +457,7 @@ size_t strbuf_expand_literal(struct strbuf *sb, const char *placeholder)\n \t\treturn 1;\n \tcase 'x':\n \t\t/* %x00 == NUL, %x0a == LF, etc. */\n-\t\tch = hex2chr(placeholder + 1);\n+\t\tch = hex2chr(placeholder + 1, HEX_KIND_MIXED);\n \t\tif (ch < 0)\n \t\t\treturn 0;\n \t\tstrbuf_addch(sb, ch);\ndiff --git a/url.c b/url.c\nindex a59818278f..b4d72f784a 100644\n--- a/url.c\n+++ b/url.c\n@@ -62,7 +62,7 @@ static char *url_decode_internal(const char **query, int len,\n \t\t}\n \n \t\tif (c == '%' && (len < 0 || len >= 3)) {\n-\t\t\tint val = hex2chr(q + 1);\n+\t\t\tint val = hex2chr(q + 1, HEX_KIND_MIXED);\n \t\t\tif (0 < val) {\n \t\t\t\tstrbuf_addch(out, val);\n \t\t\t\tq += 3;\ndiff --git a/urlmatch.c b/urlmatch.c\nindex 20bc2d009c..989f1d794b 100644\n--- a/urlmatch.c\n+++ b/urlmatch.c\n@@ -50,7 +50,7 @@ static int append_normalized_escapes(struct strbuf *buf,\n \t\tif (ch == '%') {\n \t\t\tif (from_len < 2)\n \t\t\t\treturn 0;\n-\t\t\tch = hex2chr(from);\n+\t\t\tch = hex2chr(from, HEX_KIND_MIXED);\n \t\t\tif (ch < 0)\n \t\t\t\treturn 0;\n \t\t\tfrom += 2;\n"},{"id":"549248","messageId":"20260729233215.398654-5-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"[RFC PATCH 4/6] hex: label usages of hex parsing for object IDs","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T23:32:13Z","receivedAt":"2026-07-29T23:32:19Z","isPatch":true,"body":"In preparation for a future change, label the hex parsing we're doing\nfor object IDs by defining a constant called HEX_KIND_OID.  This is\ncurrently the same as HEX_KIND_MIXED, so there is no functional change\nhere.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n diagnose.c    | 2 +-\n hex-ll.h      | 2 ++\n hex.c         | 2 +-\n http-push.c   | 4 ++--\n notes.c       | 2 +-\n object-file.c | 2 +-\n 6 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/diagnose.c b/diagnose.c\nindex fc11cea229..9c652d36a6 100644\n--- a/diagnose.c\n+++ b/diagnose.c\n@@ -112,7 +112,7 @@ static void loose_objs_stats(struct strbuf *buf, const char *path)\n \twhile ((e = readdir_skip_dot_and_dotdot(dir)) != NULL)\n \t\tif (get_dtype(e, &count_path, 0) == DT_DIR &&\n \t\t    strlen(e->d_name) == 2 &&\n-\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_MIXED)) {\n+\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_OID)) {\n \t\t\tstrbuf_setlen(&count_path, base_path_len);\n \t\t\tstrbuf_addf(&count_path, \"%s/\", e->d_name);\n \t\t\ttotal += (count = count_files(&count_path));\ndiff --git a/hex-ll.h b/hex-ll.h\nindex fe698f0c76..9da76f17e8 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -6,6 +6,8 @@ enum hexkind {\n \tHEX_KIND_LOWER = 1,\n };\n \n+#define HEX_KIND_OID HEX_KIND_MIXED\n+\n extern const signed char hexval_table[256];\n extern const signed char hexval_lc_table[256];\n static inline unsigned int hexval(unsigned char c, enum hexkind kind)\ndiff --git a/hex.c b/hex.c\nindex 6150bdcbf8..4e1e81af3f 100644\n--- a/hex.c\n+++ b/hex.c\n@@ -9,7 +9,7 @@ static int get_hash_hex_algop(const char *hex, unsigned char *hash,\n \t\t\t      const struct git_hash_algo *algop)\n {\n \tfor (size_t i = 0; i < algop->rawsz; i++) {\n-\t\tint val = hex2chr(hex, HEX_KIND_MIXED);\n+\t\tint val = hex2chr(hex, HEX_KIND_OID);\n \t\tif (val < 0)\n \t\t\treturn -1;\n \t\t*hash++ = val;\ndiff --git a/http-push.c b/http-push.c\nindex 0cc990d395..132d26d6a1 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1030,13 +1030,13 @@ static int get_oid_hex_from_objpath(const char *path, struct object_id *oid)\n \tif (strlen(path) != the_hash_algo->hexsz + 1)\n \t\treturn -1;\n \n-\tif (hex_to_bytes(oid->hash, path, 1, HEX_KIND_MIXED))\n+\tif (hex_to_bytes(oid->hash, path, 1, HEX_KIND_OID))\n \t\treturn -1;\n \tpath += 2;\n \tpath++; /* skip '/' */\n \n \treturn hex_to_bytes(oid->hash + 1, path, the_hash_algo->rawsz - 1,\n-\t\t\t    HEX_KIND_MIXED);\n+\t\t\t    HEX_KIND_OID);\n }\n \n static void process_ls_object(struct remote_ls_ctx *ls)\ndiff --git a/notes.c b/notes.c\nindex 99b8b15d81..7e9e3eb2d2 100644\n--- a/notes.c\n+++ b/notes.c\n@@ -443,7 +443,7 @@ static void load_subtree(struct notes_tree *t, struct leaf_node *subtree,\n \t\t\t\tgoto handle_non_note;\n \n \t\t\tif (hex_to_bytes(object_oid.hash + len++, entry.path, 1,\n-\t\t\t\t\t HEX_KIND_MIXED))\n+\t\t\t\t\t HEX_KIND_OID))\n \t\t\t\tgoto handle_non_note; /* entry.path is not a SHA1 */\n \n \t\t\t/*\ndiff --git a/object-file.c b/object-file.c\nindex 8427b2802a..4bcff66442 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1473,7 +1473,7 @@ int for_each_file_in_obj_subdir(unsigned int subdir_nr,\n \t\tstrbuf_add(path, de->d_name, namelen);\n \t\tif (namelen == algop->hexsz - 2 &&\n \t\t    !hex_to_bytes(oid.hash + 1, de->d_name,\n-\t\t\t\t  algop->rawsz - 1, HEX_KIND_MIXED)) {\n+\t\t\t\t  algop->rawsz - 1, HEX_KIND_OID)) {\n \t\t\toid_set_algo(&oid, algop);\n \t\t\tmemset(oid.hash + algop->rawsz, 0,\n \t\t\t       GIT_MAX_RAWSZ - algop->rawsz);\n"},{"id":"549249","messageId":"20260729233215.398654-2-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"[RFC PATCH 1/6] hex: add functionality for lowercase-only hex","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T23:32:10Z","receivedAt":"2026-07-29T23:32:19Z","isPatch":true,"body":"We currently allow both upper and lower case for all hex values in Git.\nHowever, in a future commit, we'll want to change that to allow only\nlowercase values in some cases.  To prepare for that case, provide a\ntable to convert hex values using lowercase only and an enum to let us\nchoose which we want, wiring it up to the hexval function.\n\nFor now, keep things completely the same by specifying only the\nvariant that accepts both lowercase and uppercase to avoid changing\nbehavior.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n color.c    |  2 +-\n hex-ll.c   | 37 ++++++++++++++++++++++++++++++++++++-\n hex-ll.h   | 14 ++++++++++----\n pkt-line.c |  8 ++++----\n 4 files changed, 51 insertions(+), 10 deletions(-)\n\ndiff --git a/color.c b/color.c\nindex 00b53f97ac..9015d0faf1 100644\n--- a/color.c\n+++ b/color.c\n@@ -72,7 +72,7 @@ static int get_hex_color(const char **inp, int width, unsigned char *out)\n \tunsigned int val;\n \n \tassert(width == 1 || width == 2);\n-\tval = (hexval(in[0]) << 4) | hexval(in[width - 1]);\n+\tval = (hexval(in[0], HEX_KIND_MIXED) << 4) | hexval(in[width - 1], HEX_KIND_MIXED);\n \tif (val & ~0xff)\n \t\treturn -1;\n \t*inp += width;\ndiff --git a/hex-ll.c b/hex-ll.c\nindex 4d7ece1de5..fa85e91827 100644\n--- a/hex-ll.c\n+++ b/hex-ll.c\n@@ -36,10 +36,45 @@ const signed char hexval_table[256] = {\n \t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n };\n \n+const signed char hexval_lc_table[256] = {\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 00-07 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 08-0f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 10-17 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 18-1f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 20-27 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 28-2f */\n+\t  0,  1,  2,  3,  4,  5,  6,  7,\t\t/* 30-37 */\n+\t  8,  9, -1, -1, -1, -1, -1, -1,\t\t/* 38-3f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 40-47 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 48-4f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 50-57 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 58-5f */\n+\t -1, 10, 11, 12, 13, 14, 15, -1,\t\t/* 60-67 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 68-67 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 70-77 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 78-7f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 80-87 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 88-8f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 90-97 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 98-9f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* a0-a7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* a8-af */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* b0-b7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* b8-bf */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* c0-c7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* c8-cf */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* d0-d7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* d8-df */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* e0-e7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* e8-ef */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f0-f7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n+};\n+\n int hex_to_bytes(unsigned char *binary, const char *hex, size_t len)\n {\n \tfor (; len; len--, hex += 2) {\n-\t\tunsigned int val = (hexval(hex[0]) << 4) | hexval(hex[1]);\n+\t\tunsigned int val = (hexval(hex[0], HEX_KIND_MIXED) << 4) | hexval(hex[1], HEX_KIND_MIXED);\n \n \t\tif (val & ~0xff)\n \t\t\treturn -1;\ndiff --git a/hex-ll.h b/hex-ll.h\nindex a381fa8556..da1b5239b2 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -1,10 +1,16 @@\n #ifndef HEX_LL_H\n #define HEX_LL_H\n \n+enum hexkind {\n+\tHEX_KIND_MIXED = 0,\n+\tHEX_KIND_LOWER = 1,\n+};\n+\n extern const signed char hexval_table[256];\n-static inline unsigned int hexval(unsigned char c)\n+extern const signed char hexval_lc_table[256];\n+static inline unsigned int hexval(unsigned char c, enum hexkind kind)\n {\n-\treturn hexval_table[c];\n+\treturn kind == HEX_KIND_MIXED ? hexval_table[c] : hexval_lc_table[c];\n }\n \n /*\n@@ -13,8 +19,8 @@ static inline unsigned int hexval(unsigned char c)\n  */\n static inline int hex2chr(const char *s)\n {\n-\tunsigned int val = hexval(s[0]);\n-\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1]);\n+\tunsigned int val = hexval(s[0], HEX_KIND_MIXED);\n+\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1], HEX_KIND_MIXED);\n }\n \n /*\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 3fc3e9ea70..338075558c 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -378,10 +378,10 @@ int packet_length(const char lenbuf_hex[4], size_t size)\n {\n \tif (size < 4)\n \t\tBUG(\"buffer too small\");\n-\treturn\thexval(lenbuf_hex[0]) << 12 |\n-\t\thexval(lenbuf_hex[1]) <<  8 |\n-\t\thexval(lenbuf_hex[2]) <<  4 |\n-\t\thexval(lenbuf_hex[3]);\n+\treturn\thexval(lenbuf_hex[0], HEX_KIND_MIXED) << 12 |\n+\t\thexval(lenbuf_hex[1], HEX_KIND_MIXED) <<  8 |\n+\t\thexval(lenbuf_hex[2], HEX_KIND_MIXED) <<  4 |\n+\t\thexval(lenbuf_hex[3], HEX_KIND_MIXED);\n }\n \n static const char *find_packfile_uri_path(const char *buffer)\n"},{"id":"549250","messageId":"20260729233215.398654-4-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"[RFC PATCH 3/6] hex: make hex_to_bytes accept kind of hex to use","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T23:32:12Z","receivedAt":"2026-07-29T23:32:19Z","isPatch":true,"body":"Similarly to the previous commit, introduce an option for hex_to_bytes\nto allow us to specify the kind of hex to use: lowercase only or not.\nFor now, everything remains the same as before, but we will change\nthings in a future commit.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n builtin/index-pack.c | 2 +-\n diagnose.c           | 2 +-\n hex-ll.c             | 4 ++--\n hex-ll.h             | 2 +-\n http-push.c          | 5 +++--\n notes.c              | 5 +++--\n object-file.c        | 2 +-\n 7 files changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex bc86925ad0..9660e5967b 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1866,7 +1866,7 @@ static void repack_local_links(void)\n \twhile (strbuf_getline_lf(&line, out) != EOF) {\n \t\tunsigned char binary[GIT_MAX_RAWSZ];\n \t\tif (line.len != the_hash_algo->hexsz ||\n-\t\t    !hex_to_bytes(binary, line.buf, line.len))\n+\t\t    !hex_to_bytes(binary, line.buf, line.len, HEX_KIND_MIXED))\n \t\t\tdie(_(\"index-pack: Expecting full hex object ID lines only from pack-objects.\"));\n \n \t\t/*\ndiff --git a/diagnose.c b/diagnose.c\nindex 5092bf80d3..fc11cea229 100644\n--- a/diagnose.c\n+++ b/diagnose.c\n@@ -112,7 +112,7 @@ static void loose_objs_stats(struct strbuf *buf, const char *path)\n \twhile ((e = readdir_skip_dot_and_dotdot(dir)) != NULL)\n \t\tif (get_dtype(e, &count_path, 0) == DT_DIR &&\n \t\t    strlen(e->d_name) == 2 &&\n-\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_MIXED)) {\n \t\t\tstrbuf_setlen(&count_path, base_path_len);\n \t\t\tstrbuf_addf(&count_path, \"%s/\", e->d_name);\n \t\t\ttotal += (count = count_files(&count_path));\ndiff --git a/hex-ll.c b/hex-ll.c\nindex fa85e91827..b2e9684693 100644\n--- a/hex-ll.c\n+++ b/hex-ll.c\n@@ -71,10 +71,10 @@ const signed char hexval_lc_table[256] = {\n \t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n };\n \n-int hex_to_bytes(unsigned char *binary, const char *hex, size_t len)\n+int hex_to_bytes(unsigned char *binary, const char *hex, size_t len, enum hexkind kind)\n {\n \tfor (; len; len--, hex += 2) {\n-\t\tunsigned int val = (hexval(hex[0], HEX_KIND_MIXED) << 4) | hexval(hex[1], HEX_KIND_MIXED);\n+\t\tunsigned int val = (hexval(hex[0], kind) << 4) | hexval(hex[1], kind);\n \n \t\tif (val & ~0xff)\n \t\t\treturn -1;\ndiff --git a/hex-ll.h b/hex-ll.h\nindex 26847c7b2f..fe698f0c76 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -28,6 +28,6 @@ static inline int hex2chr(const char *s, enum hexkind kind)\n  * values to `binary` as `len` bytes. Return 0 on success, or -1 if\n  * the input does not consist of hex digits).\n  */\n-int hex_to_bytes(unsigned char *binary, const char *hex, size_t len);\n+int hex_to_bytes(unsigned char *binary, const char *hex, size_t len, enum hexkind kind);\n \n #endif\ndiff --git a/http-push.c b/http-push.c\nindex 94a1fac9ab..0cc990d395 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1030,12 +1030,13 @@ static int get_oid_hex_from_objpath(const char *path, struct object_id *oid)\n \tif (strlen(path) != the_hash_algo->hexsz + 1)\n \t\treturn -1;\n \n-\tif (hex_to_bytes(oid->hash, path, 1))\n+\tif (hex_to_bytes(oid->hash, path, 1, HEX_KIND_MIXED))\n \t\treturn -1;\n \tpath += 2;\n \tpath++; /* skip '/' */\n \n-\treturn hex_to_bytes(oid->hash + 1, path, the_hash_algo->rawsz - 1);\n+\treturn hex_to_bytes(oid->hash + 1, path, the_hash_algo->rawsz - 1,\n+\t\t\t    HEX_KIND_MIXED);\n }\n \n static void process_ls_object(struct remote_ls_ctx *ls)\ndiff --git a/notes.c b/notes.c\nindex ec9c2cb150..99b8b15d81 100644\n--- a/notes.c\n+++ b/notes.c\n@@ -428,7 +428,7 @@ static void load_subtree(struct notes_tree *t, struct leaf_node *subtree,\n \t\t\t\tgoto handle_non_note;\n \n \t\t\tif (hex_to_bytes(object_oid.hash + prefix_len, entry.path,\n-\t\t\t\t\t hashsz - prefix_len))\n+\t\t\t\t\t hashsz - prefix_len, HEX_KIND_MIXED))\n \t\t\t\tgoto handle_non_note; /* entry.path is not a SHA1 */\n \n \t\t\tmemset(object_oid.hash + hashsz, 0, GIT_MAX_RAWSZ - hashsz);\n@@ -442,7 +442,8 @@ static void load_subtree(struct notes_tree *t, struct leaf_node *subtree,\n \t\t\t\t/* internal nodes must be trees */\n \t\t\t\tgoto handle_non_note;\n \n-\t\t\tif (hex_to_bytes(object_oid.hash + len++, entry.path, 1))\n+\t\t\tif (hex_to_bytes(object_oid.hash + len++, entry.path, 1,\n+\t\t\t\t\t HEX_KIND_MIXED))\n \t\t\t\tgoto handle_non_note; /* entry.path is not a SHA1 */\n \n \t\t\t/*\ndiff --git a/object-file.c b/object-file.c\nindex 7ff2b730ac..8427b2802a 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1473,7 +1473,7 @@ int for_each_file_in_obj_subdir(unsigned int subdir_nr,\n \t\tstrbuf_add(path, de->d_name, namelen);\n \t\tif (namelen == algop->hexsz - 2 &&\n \t\t    !hex_to_bytes(oid.hash + 1, de->d_name,\n-\t\t\t\t  algop->rawsz - 1)) {\n+\t\t\t\t  algop->rawsz - 1, HEX_KIND_MIXED)) {\n \t\t\toid_set_algo(&oid, algop);\n \t\t\tmemset(oid.hash + algop->rawsz, 0,\n \t\t\t       GIT_MAX_RAWSZ - algop->rawsz);\n"},{"id":"549251","messageId":"20260729233215.398654-6-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"[RFC PATCH 5/6] object-name: use hexval","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T23:32:14Z","receivedAt":"2026-07-29T23:32:21Z","isPatch":true,"body":"We've open-coded a different implementation of parsing hex values here\nwhen we already have a perfectly good one in hexval.  This\nimplementation will almost certainly be slower because it isn't\ntable-driven, unlike the other one, and since it's not constant time it\nhas no other advantages either.  To tidy things up and prepare for\nfuture work, switch to hexval in this case.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n object-name.c | 13 +++----------\n 1 file changed, 3 insertions(+), 10 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 83efba0ba6..d2d81b3511 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -236,17 +236,10 @@ static int parse_oid_prefix(const char *name, int len,\n {\n \tfor (int i = 0; i < len; i++) {\n \t\tunsigned char c = name[i];\n-\t\tunsigned char val;\n-\t\tif (c >= '0' && c <= '9') {\n-\t\t\tval = c - '0';\n-\t\t} else if (c >= 'a' && c <= 'f') {\n-\t\t\tval = c - 'a' + 10;\n-\t\t} else if (c >= 'A' && c <='F') {\n-\t\t\tval = c - 'A' + 10;\n-\t\t\tc -= 'A' - 'a';\n-\t\t} else {\n+\t\tint val = hexval(c, HEX_KIND_OID);\n+\n+\t\tif (val < 0)\n \t\t\treturn -1;\n-\t\t}\n \n \t\tif (hex_out)\n \t\t\thex_out[i] = c;\n"},{"id":"549252","messageId":"20260729233215.398654-7-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"[RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T23:32:15Z","receivedAt":"2026-07-29T23:32:21Z","isPatch":true,"body":"Git has historically allowed either lowercase or uppercase hex for\nobject IDs, but it has always emitted only lowercase.  This has caused\npeople to expect only lowercase and not handle uppercase.\n\nAs an example, Git's own example hooks look for \"[0-9a-f]\" in several\nplaces, but there are many other Git-adjacent pieces of software,\nincluding Gitolite, which make the assumption that object IDs are always\nlowercase.  This is not to criticize the authors of these projects, but\nrather to point out how common this assumption is.  In fact, it's so\ncommon that we have only one test in our codebase that fails when we\nreject uppercase object IDs.\n\nMore critically, it leads people to make security-based assumptions that\nan object ID either does not contain uppercase characters or that an\nobject ID can be expressed uniquely in hex form, neither of which are\ncurrently true.  Git itself normally uses binary object IDs, which\navoids many of these problems, but most other projects deal primarily in\nhex object IDs, so they are more affected.\n\nIn preparation for Git 3.0, only allow lowercase hex object IDs in\nbreaking changes mode and document this as well.  Update the single\nfailing test and add a new one to verify we reject new uppercase object\nIDs.  Note that in t5324, we change the hex character from \"A\" to \"b\"\nbecause in SHA-256 mode, \"a\" is the correct value, so our test_must_fail\nassertion will unexpectedly succeed in that case.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n Documentation/BreakingChanges.adoc | 5 +++++\n hex-ll.h                           | 4 ++++\n t/t1503-rev-parse-verify.sh        | 5 +++++\n t/t5324-split-commit-graph.sh      | 4 ++--\n 4 files changed, 16 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/BreakingChanges.adoc b/Documentation/BreakingChanges.adoc\nindex 73bb939359..dbc46d14e3 100644\n--- a/Documentation/BreakingChanges.adoc\n+++ b/Documentation/BreakingChanges.adoc\n@@ -171,6 +171,11 @@ JGit, libgit2 and Gitoxide need to support it.\n   matches the default branch name used in new repositories by many of the\n   big Git forges.\n \n+* Git will accept hex object IDs only in lowercase. The fact that Git has\n+\thistorically allowed uppercase characters in hex object IDs has been the\n+\tsource of a variety of bugs and security problems in software using Git. We\n+\tdon't expect most users to notice any change.\n+\n * Git will require Rust as a mandatory part of the build process. While Git\n   already started to adopt Rust in Git 2.49, all parts written in Rust are\n   optional for the time being. This includes:\ndiff --git a/hex-ll.h b/hex-ll.h\nindex 9da76f17e8..2f9c8d7c25 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -6,7 +6,11 @@ enum hexkind {\n \tHEX_KIND_LOWER = 1,\n };\n \n+#ifdef WITH_BREAKING_CHANGES\n+#define HEX_KIND_OID HEX_KIND_LOWER\n+#else\n #define HEX_KIND_OID HEX_KIND_MIXED\n+#endif\n \n extern const signed char hexval_table[256];\n extern const signed char hexval_lc_table[256];\ndiff --git a/t/t1503-rev-parse-verify.sh b/t/t1503-rev-parse-verify.sh\nindex 87638a4a2c..f07b45de5a 100755\n--- a/t/t1503-rev-parse-verify.sh\n+++ b/t/t1503-rev-parse-verify.sh\n@@ -60,6 +60,11 @@ test_expect_success 'works with one good rev' '\n \ttest \"$rev_head\" = \"$HASH4\"\n '\n \n+test_expect_success WITH_BREAKING_CHANGES 'rejects uppercase revs' '\n+\tUC_HASH=$(echo \"$HASH1\" | tr a-f A-F) &&\n+\ttest_must_fail git rev-parse --verify \"$UC_HASH\"\n+'\n+\n test_expect_success 'fails with any bad rev or many good revs' '\n \ttest_must_fail git rev-parse --verify 2>error &&\n \ttest_grep \"single revision\" error &&\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex bf7ba0e558..29db815c77 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -349,7 +349,7 @@ test_expect_success 'verify after commit-graph-chain corruption (base)' '\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"invalid commit-graph chain\" err &&\n-\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"A\" &&\n+\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"a\" &&\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"unable to find all commit-graph files\" err\n@@ -364,7 +364,7 @@ test_expect_success 'verify after commit-graph-chain corruption (tip)' '\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"invalid commit-graph chain\" err &&\n-\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"A\" &&\n+\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"b\" &&\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"unable to find all commit-graph files\" err\n"},{"id":"549264","messageId":"xmqqjyqclwf9.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 0/6] Git 3.0: restrict hex object IDs to lowercase only","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-30T08:21:46Z","receivedAt":"2026-07-30T08:21:49Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> As far as I can tell, Git has always emitted hex object IDs in\n> lowercase, but our object ID parser accepts both uppercase and\n> lowercase.  This leads to much software relying on hex object IDs being\n> broken because it doesn't handle uppercase object IDs and this can even\n> lead to security problems when people assume that an object ID has a\n> unique hex form.\n>\n> This series proposes to remove the ability to use uppercase hex in\n> object IDs in Git 3.0.  It is RFC simply because it's not clear if\n> there's the desire to do this, although the series should be fully\n> functional.\n>\n> As further evidence of why we should do this, I'll note that there is\n> exactly one testcase in our testsuite that fails due to this change\n> (fixed in the last patch) and it's not clear that it fails\n> intentionally.  If we decide not to adopt this series, it would probably\n> be prudent to add some additional tests for the uppercase variant of hex\n> object IDs.\n\nBefore going there, we should hear a solid argument why doing this\nmight be beneficial longer term.  \"Just because we might be able to\nwithout harming too many users\" is probably not good enough, when it\nis not accompanied by \"... the (low) risk may be worth taking because\nwe will gain such and such benefit\".\n"},{"id":"549326","messageId":"amu_rzanuYc_2lww@fruit.crustytoothpaste.net","threadId":"66087","inReplyTo":"xmqqjyqclwf9.fsf@gitster.g","subject":"Re: [RFC PATCH 0/6] Git 3.0: restrict hex object IDs to lowercase only","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-30T21:18:40Z","receivedAt":"2026-07-30T21:18:42Z","isPatch":true,"body":"On 2026-07-30 at 08:21:46, Junio C Hamano wrote:\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > As far as I can tell, Git has always emitted hex object IDs in\n> > lowercase, but our object ID parser accepts both uppercase and\n> > lowercase.  This leads to much software relying on hex object IDs being\n> > broken because it doesn't handle uppercase object IDs and this can even\n> > lead to security problems when people assume that an object ID has a\n> > unique hex form.\n> >\n> > This series proposes to remove the ability to use uppercase hex in\n> > object IDs in Git 3.0.  It is RFC simply because it's not clear if\n> > there's the desire to do this, although the series should be fully\n> > functional.\n> >\n> > As further evidence of why we should do this, I'll note that there is\n> > exactly one testcase in our testsuite that fails due to this change\n> > (fixed in the last patch) and it's not clear that it fails\n> > intentionally.  If we decide not to adopt this series, it would probably\n> > be prudent to add some additional tests for the uppercase variant of hex\n> > object IDs.\n> \n> Before going there, we should hear a solid argument why doing this\n> might be beneficial longer term.  \"Just because we might be able to\n> without harming too many users\" is probably not good enough, when it\n> is not accompanied by \"... the (low) risk may be worth taking because\n> we will gain such and such benefit\".\n\nWhile I can't speak about the details, I've actually seen multiple\nsecurity vulnerabilities show up because people didn't realize that\nuppercase hex was a thing in Git object IDs and so filtering or other\nsanitizing was ineffective.  That's the real motivation behind this\nchange.\n\nAlso, as patch 6 says, a large amount of Git-adjacent software,\nincluding common implementations such as Gitolite, don't accept them or\ndon't handle them correctly.  A quick code search for `[0-9a-f]{40}` on\nGitHub shows a lot of these tools.  Our own hook examples even use a\nsimilar pattern, and although in that case they are accepting only Git's\noutput, users see those as examples of how to parse object IDs.\n\nThe situation is presently that Git will accept them and this leads to\nsurprising behaviour, but almost all adjacent software rejects or\nmishandles them.  I'm arguing that we should stop accepting hex object\nID formats that cannot be effectively used in the Git ecosystem but\nwhose presence is effectively only ever the source of misbehaviour and\nsecurity vulnerabilities.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"549328","messageId":"xmqq1pcjkfi1.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-5-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 4/6] hex: label usages of hex parsing for object IDs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-31T03:24:54Z","receivedAt":"2026-07-31T03:24:56Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> In preparation for a future change, label the hex parsing we're doing\n> for object IDs by defining a constant called HEX_KIND_OID.  This is\n> currently the same as HEX_KIND_MIXED, so there is no functional change\n> here.\n>\n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> ---\n>  diagnose.c    | 2 +-\n>  hex-ll.h      | 2 ++\n>  hex.c         | 2 +-\n>  http-push.c   | 4 ++--\n>  notes.c       | 2 +-\n>  object-file.c | 2 +-\n>  6 files changed, 8 insertions(+), 6 deletions(-)\n\nOK.  It makes sense to say \"we are reading object names\", than \"we\nare reading hex spelled in both cases\".  Are we throwing the \"not\nobject names but derived from the same hash function\" things like\npackname and rerere database key into the same category?\n\n> diff --git a/diagnose.c b/diagnose.c\n> index fc11cea229..9c652d36a6 100644\n> --- a/diagnose.c\n> +++ b/diagnose.c\n> @@ -112,7 +112,7 @@ static void loose_objs_stats(struct strbuf *buf, const char *path)\n>  \twhile ((e = readdir_skip_dot_and_dotdot(dir)) != NULL)\n>  \t\tif (get_dtype(e, &count_path, 0) == DT_DIR &&\n>  \t\t    strlen(e->d_name) == 2 &&\n> -\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_MIXED)) {\n> +\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_OID)) {\n>  \t\t\tstrbuf_setlen(&count_path, base_path_len);\n>  \t\t\tstrbuf_addf(&count_path, \"%s/\", e->d_name);\n>  \t\t\ttotal += (count = count_files(&count_path));\n"},{"id":"549332","messageId":"xmqq4ihfip7d.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-2-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 1/6] hex: add functionality for lowercase-only hex","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-31T07:38:14Z","receivedAt":"2026-07-31T07:38:17Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> We currently allow both upper and lower case for all hex values in Git.\n> However, in a future commit, we'll want to change that to allow only\n> lowercase values in some cases.  To prepare for that case, provide a\n> table to convert hex values using lowercase only and an enum to let us\n> choose which we want, wiring it up to the hexval function.\n>\n> For now, keep things completely the same by specifying only the\n> variant that accepts both lowercase and uppercase to avoid changing\n> behavior.\n>\n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> ---\n\nIn \"some\" cases?  I wonder what other cases there are that we MUST\naccept uppercase variants.  Obviously the network protocol where we\nare willing to talk to reimplementation of Git by others is one.  I\ndo not think we historically produced anything in uppercase.\n\n> diff --git a/color.c b/color.c\n> index 00b53f97ac..9015d0faf1 100644\n> --- a/color.c\n> +++ b/color.c\n> @@ -72,7 +72,7 @@ static int get_hex_color(const char **inp, int width, unsigned char *out)\n>  \tunsigned int val;\n>  \n>  \tassert(width == 1 || width == 2);\n> -\tval = (hexval(in[0]) << 4) | hexval(in[width - 1]);\n> +\tval = (hexval(in[0], HEX_KIND_MIXED) << 4) | hexval(in[width - 1], HEX_KIND_MIXED);\n>  \tif (val & ~0xff)\n>  \t\treturn -1;\n>  \t*inp += width;\n> diff --git a/hex-ll.c b/hex-ll.c\n> index 4d7ece1de5..fa85e91827 100644\n> --- a/hex-ll.c\n> +++ b/hex-ll.c\n> @@ -36,10 +36,45 @@ const signed char hexval_table[256] = {\n>  \t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n>  };\n>  \n> +const signed char hexval_lc_table[256] = {\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 00-07 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 08-0f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 10-17 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 18-1f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 20-27 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 28-2f */\n> +\t  0,  1,  2,  3,  4,  5,  6,  7,\t\t/* 30-37 */\n> +\t  8,  9, -1, -1, -1, -1, -1, -1,\t\t/* 38-3f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 40-47 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 48-4f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 50-57 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 58-5f */\n> +\t -1, 10, 11, 12, 13, 14, 15, -1,\t\t/* 60-67 */\n> ...\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n> +};\n> +\n>  int hex_to_bytes(unsigned char *binary, const char *hex, size_t len)\n>  {\n>  \tfor (; len; len--, hex += 2) {\n> -\t\tunsigned int val = (hexval(hex[0]) << 4) | hexval(hex[1]);\n> +\t\tunsigned int val = (hexval(hex[0], HEX_KIND_MIXED) << 4) | hexval(hex[1], HEX_KIND_MIXED);\n>  \n>  \t\tif (val & ~0xff)\n>  \t\t\treturn -1;\n>\n> diff --git a/hex-ll.h b/hex-ll.h\n> index a381fa8556..da1b5239b2 100644\n> --- a/hex-ll.h\n> +++ b/hex-ll.h\n> @@ -1,10 +1,16 @@\n>  #ifndef HEX_LL_H\n>  #define HEX_LL_H\n>  \n> +enum hexkind {\n> +\tHEX_KIND_MIXED = 0,\n> +\tHEX_KIND_LOWER = 1,\n> +};\n> +\n>  extern const signed char hexval_table[256];\n> -static inline unsigned int hexval(unsigned char c)\n> +extern const signed char hexval_lc_table[256];\n> +static inline unsigned int hexval(unsigned char c, enum hexkind kind)\n>  {\n> -\treturn hexval_table[c];\n> +\treturn kind == HEX_KIND_MIXED ? hexval_table[c] : hexval_lc_table[c];\n>  }\n\nIt is very welcome to make sure we are conservative in what we\nproduce, but be liberal in what we accept.  In that sense, use of\nHEX_KIND_LOWER goes directly against the Robustness Principle.\n\n>  \n>  /*\n> @@ -13,8 +19,8 @@ static inline unsigned int hexval(unsigned char c)\n>   */\n>  static inline int hex2chr(const char *s)\n>  {\n> -\tunsigned int val = hexval(s[0]);\n> -\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1]);\n> +\tunsigned int val = hexval(s[0], HEX_KIND_MIXED);\n> +\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1], HEX_KIND_MIXED);\n>  }\n>  \n>  /*\n> diff --git a/pkt-line.c b/pkt-line.c\n> index 3fc3e9ea70..338075558c 100644\n> --- a/pkt-line.c\n> +++ b/pkt-line.c\n> @@ -378,10 +378,10 @@ int packet_length(const char lenbuf_hex[4], size_t size)\n>  {\n>  \tif (size < 4)\n>  \t\tBUG(\"buffer too small\");\n> -\treturn\thexval(lenbuf_hex[0]) << 12 |\n> -\t\thexval(lenbuf_hex[1]) <<  8 |\n> -\t\thexval(lenbuf_hex[2]) <<  4 |\n> -\t\thexval(lenbuf_hex[3]);\n> +\treturn\thexval(lenbuf_hex[0], HEX_KIND_MIXED) << 12 |\n> +\t\thexval(lenbuf_hex[1], HEX_KIND_MIXED) <<  8 |\n> +\t\thexval(lenbuf_hex[2], HEX_KIND_MIXED) <<  4 |\n> +\t\thexval(lenbuf_hex[3], HEX_KIND_MIXED);\n>  }\n>  \n>  static const char *find_packfile_uri_path(const char *buffer)\n"},{"id":"549333","messageId":"xmqqzez7hamu.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-4-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 3/6] hex: make hex_to_bytes accept kind of hex to use","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-31T07:38:17Z","receivedAt":"2026-07-31T07:38:19Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> -int hex_to_bytes(unsigned char *binary, const char *hex, size_t len)\n> +int hex_to_bytes(unsigned char *binary, const char *hex, size_t len, enum hexkind kind)\n>  {\n>  \tfor (; len; len--, hex += 2) {\n> -\t\tunsigned int val = (hexval(hex[0], HEX_KIND_MIXED) << 4) | hexval(hex[1], HEX_KIND_MIXED);\n> +\t\tunsigned int val = (hexval(hex[0], kind) << 4) | hexval(hex[1], kind);\n>  \n>  \t\tif (val & ~0xff)\n>  \t\t\treturn -1;\n\nIt depends on how big 'len' would be to matter, but if we are\nlooping for a long stretch, choosing which one of the two hexval\ntables to use outside the loop and using that inside may of course\nbe more performant.\n\nI wondered how ugly such a restructure of the API would look like,\nand it does not look _too_ bad.\n\n\tvoid *hextable = hex_table(HEX_KIND_MIXED);\n\n\tfor (; len; len--, hex += 2) {\n\t\tunsigned int val =\n\t\t\t(hexval(hex[0], hextable) << 4) | hexval(hex[1], hextable);\n\t\t...\n\t}\n\nThe true type of hextable would be \"signed char [256]\", but the\ncallers of the hexval() function do not need to know it, hence I\nchose \"void *\" here.\n"},{"id":"549334","messageId":"xmqqv79vha69.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-7-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-31T07:48:14Z","receivedAt":"2026-07-31T07:48:17Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> Git has historically allowed either lowercase or uppercase hex for\n> object IDs, but it has always emitted only lowercase.  This has caused\n> people to expect only lowercase and not handle uppercase.\n\nIt is violation of Postel's Law by other people.  We do not\nnecessarily have to follow suit.\n\nEven though I said throwing object names in a single category makes\nsense, it may make sense to treat the object names that we locally\nuse to access our own object database and those that we use when\ntalking with _other_ people on the net separately for the Robustness\nprinciple, we keep being strict in what we produce and stick to\nlowercase, while accepting uppercase produced by those third-party\nreimplementations of Git.\n\n"},{"id":"549345","messageId":"xmqqtspffidw.fsf@gitster.g","threadId":"66087","inReplyTo":"xmqqv79vha69.fsf@gitster.g","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-31T12:33:47Z","receivedAt":"2026-07-31T12:33:50Z","isPatch":true,"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n>\n>> Git has historically allowed either lowercase or uppercase hex for\n>> object IDs, but it has always emitted only lowercase.  This has caused\n>> people to expect only lowercase and not handle uppercase.\n>\n> It is violation of Postel's Law by other people.  We do not\n> necessarily have to follow suit.\n\nImagine we somehow misbehave badly when we have two loose object\nfiles storing the same object's contents.  Let us further imagine\nthat we can download these individual loose object files from\nothers, perhaps via the dumb HTTP transport.\n\nIf we tried to be robust, we would be liberal in what we accept,\neven though we try to be strict in what we produce.  In this\nhypothetical scenario, if we talk to someone else over the dumb\nHTTP transport and find that they have an objects/AB/ directory, we\nmay try to be liberal and say, \"Ah, that is a fan-out directory\nhousing all their loose objects whose names begin with 'ab'.\"\n\nThis is the right thing to do for those who liberally accept\nothers' data.\n\nBut we might further say, \"Let us enumerate and download what we do\nnot have locally.  They have a file 012345...EF (38 hex characters)\nin that directory, which stores the object AB012345...EF (40 hex\ncharacters) in loose object form,\" and then conclude, \"and we do not\nhave it,\" even when we actually have the file ab/012345...ef in\nall-lowercase form locally!\n\nHowever, liberally accepting AB/012345...EF and storing it verbatim\nin our own store will break things because, in this hypothetical\nscenario, we will misbehave when we have both ab/012345...ef (which\nwe had from the start) and AB/012345...EF (which we just downloaded)\nat the same time.\n\nThe approach taken by this RFC series is to stop recognizing their\nobjects/AB/ as a valid fan-out directory and their AB/012345...EF as\na valid loose object file.  While I agree that this is certainly\none way to avoid entering such a state and triggering bad behavior,\nI think the real solution that honors the robustness principle is to\nstill recognize objects/AB/012345...EF as valid, recognize it as a\nloose object file for ab012345...ef, and notice that it represents\nthe same object ab012345...ef we already have.  Then we can avoid\nmisbehaving without being less liberal than we used to be.\n\nIf the system had been case-sensitive from day one, and ignoring\nuppercase hex had been the norm from the beginning, I would not have\nfound it so disturbing that we reject case-insensitive object names\nand being stricter than folks with those other systems may feel is\nnecessary.\n\nTightening the rule after twenty years is the part I am most\nhesitant to accept.  So, I dunno.\n"},{"id":"549394","messageId":"20260801143513.GE2041176@coredump.intra.peff.net","threadId":"66087","inReplyTo":"xmqqzez7hamu.fsf@gitster.g","subject":"Re: [RFC PATCH 3/6] hex: make hex_to_bytes accept kind of hex to use","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-08-01T14:35:13Z","receivedAt":"2026-08-01T14:35:17Z","isPatch":true,"body":"On Fri, Jul 31, 2026 at 12:38:17AM -0700, Junio C Hamano wrote:\n\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > -int hex_to_bytes(unsigned char *binary, const char *hex, size_t len)\n> > +int hex_to_bytes(unsigned char *binary, const char *hex, size_t len, enum hexkind kind)\n> >  {\n> >  \tfor (; len; len--, hex += 2) {\n> > -\t\tunsigned int val = (hexval(hex[0], HEX_KIND_MIXED) << 4) | hexval(hex[1], HEX_KIND_MIXED);\n> > +\t\tunsigned int val = (hexval(hex[0], kind) << 4) | hexval(hex[1], kind);\n> >  \n> >  \t\tif (val & ~0xff)\n> >  \t\t\treturn -1;\n> \n> It depends on how big 'len' would be to matter, but if we are\n> looping for a long stretch, choosing which one of the two hexval\n> tables to use outside the loop and using that inside may of course\n> be more performant.\n> \n> I wondered how ugly such a restructure of the API would look like,\n> and it does not look _too_ bad.\n> \n> \tvoid *hextable = hex_table(HEX_KIND_MIXED);\n> \n> \tfor (; len; len--, hex += 2) {\n> \t\tunsigned int val =\n> \t\t\t(hexval(hex[0], hextable) << 4) | hexval(hex[1], hextable);\n> \t\t...\n> \t}\n> \n> The true type of hextable would be \"signed char [256]\", but the\n> callers of the hexval() function do not need to know it, hence I\n> chose \"void *\" here.\n\nI had the same thought when reading this, but I wondered if the compiler\nmight be able to hoist the comparison out of the loop itself (because\nhexval() it inlined anyway). It doesn't seem to do so, though (at least\nwith gcc-15). It loads both table addresses into registers, but there's\nstill a branch in the loop to decide which table to use.\n\nSo in theory this kind of manual hoisting could help.  Might not be that\nbig a deal with branch prediction, though.\n\n-Peff\n"},{"id":"549395","messageId":"20260801144527.GF2041176@coredump.intra.peff.net","threadId":"66087","inReplyTo":"amu_rzanuYc_2lww@fruit.crustytoothpaste.net","subject":"Re: [RFC PATCH 0/6] Git 3.0: restrict hex object IDs to lowercase only","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-08-01T14:45:27Z","receivedAt":"2026-08-01T14:45:29Z","isPatch":true,"body":"On Thu, Jul 30, 2026 at 09:18:40PM +0000, brian m. carlson wrote:\n\n> The situation is presently that Git will accept them and this leads to\n> surprising behaviour, but almost all adjacent software rejects or\n> mishandles them.  I'm arguing that we should stop accepting hex object\n> ID formats that cannot be effectively used in the Git ecosystem but\n> whose presence is effectively only ever the source of misbehaviour and\n> security vulnerabilities.\n\nAnother interesting case is upper-case hex within objects:\n\n  $ git rev-parse HEAD\n  b85b9595a8136c79551340c3d73443a62eddd893\n\n  $ git cat-file commit HEAD |\n    perl -lpe '\n        if (/^parent (.*)/) {\n\t\t$_ = \"parent \" . uc($1);\n\t}\n    ' |\n    git hash-object -w -t commit --stdin\n  5a08c6b3f06d91c4a09c8d7ea6e9c8ce200b7698\n\nNow there's a parallel history of otherwise identical commits. I think\nthis is mostly \"if it hurts don't do it\", but we generally try to avoid\nmultiple representations of the same data within the object model.\n\nI think only commits and tags are subject to this (because the tree\nhashes are binary). I don't know if you'd be able to stumble into this\naccidentally with most Git commands. We don't intentionally normalize\ncase anywhere, but I think most code will round-trip through a binary\nhash at some point (so \"git commit-tree 1234ABCD\" would incidentally\nnormalize the case).\n\n-Peff\n"},{"id":"549410","messageId":"xmqqfr0x8zuu.fsf@gitster.g","threadId":"66087","inReplyTo":"20260801144527.GF2041176@coredump.intra.peff.net","subject":"Re: [RFC PATCH 0/6] Git 3.0: restrict hex object IDs to lowercase only","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-01T18:22:49Z","receivedAt":"2026-08-01T18:22:51Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> Now there's a parallel history of otherwise identical commits. I think\n> this is mostly \"if it hurts don't do it\", but we generally try to avoid\n> multiple representations of the same data within the object model.\n\nTrue.  Already almost an empty rebase that delays the clock by 1\nseconds is a perfectly normal thing, and the way we treat such a\nparallel history with the original would be the same as such an\nuppercase parallel history, so I do not think it is a huge issue.\n\nA tree object hierarchy consists fully of binary links, so we are\nOK.  Only the top-level object name within a commit/tag may have\nmultiple representation of the same tree objects.\n\n'git diff COMMIT-A COMMIT-B' would notice that they record the same\ntree without opening two tree objects, even if these two commits\nrecord a normal tree and uppercase equivalent tree, as diff_tree()\nlayer will be called with the binary representation of the tree\nobject names.\n\nAs Brian alluded to in his discussion starter message, we normalize\nthe case by going binary in many places (we cannot unfortunately say\n\"in strategic places\"), and these are such cases.\n\n> I think only commits and tags are subject to this (because the tree\n> hashes are binary). I don't know if you'd be able to stumble into this\n> accidentally with most Git commands. We don't intentionally normalize\n> case anywhere, but I think most code will round-trip through a binary\n> hash at some point (so \"git commit-tree 1234ABCD\" would incidentally\n> normalize the case).\n\n"},{"id":"549451","messageId":"am-8vm5QwLQhiXaO@fruit.crustytoothpaste.net","threadId":"66087","inReplyTo":"20260801144527.GF2041176@coredump.intra.peff.net","subject":"Re: [RFC PATCH 0/6] Git 3.0: restrict hex object IDs to lowercase only","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-08-02T21:55:10Z","receivedAt":"2026-08-02T21:55:12Z","isPatch":true,"body":"On 2026-08-01 at 14:45:27, Jeff King wrote:\n> Another interesting case is upper-case hex within objects:\n> \n>   $ git rev-parse HEAD\n>   b85b9595a8136c79551340c3d73443a62eddd893\n> \n>   $ git cat-file commit HEAD |\n>     perl -lpe '\n>         if (/^parent (.*)/) {\n> \t\t$_ = \"parent \" . uc($1);\n> \t}\n>     ' |\n>     git hash-object -w -t commit --stdin\n>   5a08c6b3f06d91c4a09c8d7ea6e9c8ce200b7698\n> \n> Now there's a parallel history of otherwise identical commits. I think\n> this is mostly \"if it hurts don't do it\", but we generally try to avoid\n> multiple representations of the same data within the object model.\n> \n> I think only commits and tags are subject to this (because the tree\n> hashes are binary). I don't know if you'd be able to stumble into this\n> accidentally with most Git commands. We don't intentionally normalize\n> case anywhere, but I think most code will round-trip through a binary\n> hash at some point (so \"git commit-tree 1234ABCD\" would incidentally\n> normalize the case).\n\nYes, this is true.  I agree that multiple representations is a problem,\nand although that can be an issue with signatures, we shouldn't make it\nworse.\n\nIn addition, those objects cannot be round-tripped through the\ninteroperability code (which only writes lowercase object IDs), so\nthey're effectively locked to SHA-1 only.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"549452","messageId":"am_AL9dymrkidizF@fruit.crustytoothpaste.net","threadId":"66087","inReplyTo":"xmqqv79vha69.fsf@gitster.g","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-08-02T22:09:52Z","receivedAt":"2026-08-02T22:09:54Z","isPatch":true,"body":"On 2026-07-31 at 07:48:14, Junio C Hamano wrote:\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > Git has historically allowed either lowercase or uppercase hex for\n> > object IDs, but it has always emitted only lowercase.  This has caused\n> > people to expect only lowercase and not handle uppercase.\n> \n> It is violation of Postel's Law by other people.  We do not\n> necessarily have to follow suit.\n\nPostel's Law was a great idea on the early Internet, but it is\nunfortunately no longer a good idea.  The problem is that being liberal\nin what you accept these days usually has security implications.\n\nTLS cannot be liberal in what it accepts because that means potentially\nallowing attacker-controlled data.  Even HTTP cannot do that because\nwe've seen where refusing to reject requests with both Content-Length\nand Transfer-Encoding: chunked means that two parts of a backend can\ndisagree on the content, allowing request smuggling.\n\nWe've seen these problems in our code where not caring about CR comes\nback to bite us on Windows in a security-sensitive way.\n\nModern development effectively requires being clear and definitive about\nwhat data is accepted and what is not, as well as what meaning is given\nto the data that is accepted.\n\n> Even though I said throwing object names in a single category makes\n> sense, it may make sense to treat the object names that we locally\n> use to access our own object database and those that we use when\n> talking with _other_ people on the net separately for the Robustness\n> principle, we keep being strict in what we produce and stick to\n> lowercase, while accepting uppercase produced by those third-party\n> reimplementations of Git.\n\nUnfortunately, that also doesn't fix most of the security problems I've\nseen, which involve object IDs that get passed on the command line when\ntools invoke Git.  It does fix the problem with round-tripping objects\nbetween hash algorithms, though, but I don't really want to audit every\nuse of oid_to_hex in our codebase to half-fix this situation.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"549607","messageId":"xmqqldalvfzx.fsf@gitster.g","threadId":"66087","inReplyTo":"am_AL9dymrkidizF@fruit.crustytoothpaste.net","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-04T19:32:18Z","receivedAt":"2026-08-04T19:32:21Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> Postel's Law was a great idea on the early Internet, but it is\n> unfortunately no longer a good idea.  The problem is that being liberal\n> in what you accept these days usually has security implications.\n\nI am afraid that is debatable, though.\n\nI would grant you that you can increase the attack surface by being\ncarelessly liberal.  Recall my example of allowing mixed-case names\nfor loose object files and storing them verbatim on a case-sensitive\nfilesystem without normalizing the names; that is an example of\nbeing carelessly liberal.\n\nBut is it a good excuse to give up being careful, declare it is\nimpossible to be careful enough, and punt?\n\nWill queue, but I invite others to chime in.  My practical side says\nwe should just take the series as it is much less work for us to\ndeclare that any incompatibility fallout is the problem of other\npeople who have reimplementations of Git, but my more principled\nside feels dirty, just for saying this ;-).\n\nThanks.\n\n"},{"id":"549627","messageId":"anJdvLV7-raJK67B@fruit.crustytoothpaste.net","threadId":"66087","inReplyTo":"xmqqldalvfzx.fsf@gitster.g","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-08-04T21:46:37Z","receivedAt":"2026-08-04T21:46:45Z","isPatch":true,"body":"On 2026-08-04 at 19:32:18, Junio C Hamano wrote:\n> Will queue, but I invite others to chime in.  My practical side says\n> we should just take the series as it is much less work for us to\n> declare that any incompatibility fallout is the problem of other\n> people who have reimplementations of Git, but my more principled\n> side feels dirty, just for saying this ;-).\n\nI appreciate that, and I do welcome other viewpoints here.  While I\nthink this series a good idea (or I wouldn't have sent it, obviously),\nit's really up to the project what the right thing is and if the\nconsensus is that this should be dropped, then we can do that.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"549633","messageId":"CAC2QwmKS+ojHd31oagdHv1G3h=Sa-BttWpQLv49=kC=PC-1BTQ@mail.gmail.com","threadId":"66087","inReplyTo":"am_AL9dymrkidizF@fruit.crustytoothpaste.net","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Michael Montalbo","fromEmail":"mmontalbo@gmail.com","sentAt":"2026-08-05T03:09:53Z","receivedAt":"2026-08-05T03:10:06Z","isPatch":true,"body":"On Sun, Aug 2, 2026 at 3:10 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> Modern development effectively requires being clear and definitive about\n> what data is accepted and what is not, as well as what meaning is given\n> to the data that is accepted.\n>\n\nI agree with this idea, and the topic inspired me to explore how mixing\nupper and lowercase hex oids might be \"abused\" today. Interestingly, I\nfound out it is possible to mix upper and lowercase formats within one\noid. Depending on output path, Git either normalizes the casing or\npreserves the raw form stored. I didn't come up with a specific way to\ntake advantage of this behavior yet, but one could imagine a scenario\nwhere downstream consumers of Git output each have their own way\nof parsing such an ambiguously formatted oid and behave in different\nways that at best may cause confusion and at worst could be used\nmaliciously.\n\nTo reproduce a mixed case oid scenario:\n\n    #!/bin/sh\n    set -eu\n\n    ( repo=$(mktemp -d); cd \"$repo\"\n      export GIT_PAGER=cat\n      git init -q\n      git config user.name Tester && git config user.email tester@example.com\n\n      echo one >f && git add f && git commit -qm base\n      echo two >f && git commit -qam child\n\n      # Re-spell the child's parent OID with a mixed-case tail, re-store the\n      # object.  Nothing else about the commit changes.\n      parent=$(git rev-parse HEAD^)\n      upper=$(printf %s \"$parent\" | tr '[:lower:]' '[:upper:]')\n      mixed=$(printf %s \"$parent\" | cut -c1-10)$(printf %s \"$parent\" |\ncut -c11- | tr '[:lower:]' '[:upper:]')\n      git cat-file commit HEAD | sed \"s/^parent .*/parent $mixed/\" >crafted-obj\n      crafted=$(git hash-object -w -t commit --stdin <crafted-obj)\n      git update-ref refs/heads/mixed \"$crafted\"\n\n      echo \"== the two commit objects differ by one field, case only ==\"\n      git cat-file commit HEAD >canon-obj\n      diff canon-obj crafted-obj || true\n\n      echo\n      echo \"== they are distinct commits ==\"\n      printf 'canonical: %s\\nmixed    : %s\\n' \"$(git rev-parse HEAD)\" \"$crafted\"\n\n      echo\n      echo \"== both parent spellings resolve to the same object ==\"\n      git rev-parse \"$parent\" \"$upper\"\n\n      echo\n      echo \"== the one parent is spelled two ways, by output path ==\"\n      printf 'raw (as stored) : '; git log -1 --pretty=raw mixed | sed\n-n 's/^parent //p'\n      printf '%%P  (normalized): '; git log -1 --format='%P' mixed\n    )\n\nproduces:\n\n    == the two commit objects differ by one field, case only ==\n    2c2\n    < parent f3b4ff525d82ddcec0cc8597842c73b72e4c5aba\n    ---\n    > parent f3b4ff525d82DDCEC0CC8597842C73B72E4C5ABA\n\n    == they are distinct commits ==\n    canonical: 52d41f6a261b49a18d38933b59a65f2dc919f9ac\n    mixed    : ed20e85251a37eb01009faaaa8fbc23baf8bdc72\n\n    == both parent spellings resolve to the same object ==\n    f3b4ff525d82ddcec0cc8597842c73b72e4c5aba\n    f3b4ff525d82ddcec0cc8597842c73b72e4c5aba\n\n    == the one parent is spelled two ways, by output path ==\n    raw (as stored) : f3b4ff525d82DDCEC0CC8597842C73B72E4C5ABA\n    %P  (normalized): f3b4ff525d82ddcec0cc8597842c73b72e4c5aba\n"},{"id":"551170","messageId":"d6940aa6-9336-481b-8ee5-5e3d9f3d3a50@gmail.com","threadId":"66087","inReplyTo":"20260729233215.398654-7-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-08-25T09:04:36Z","receivedAt":"2026-08-25T09:04:47Z","isPatch":true,"body":"Hi brian\n\nOn 30/07/2026 00:32, brian m. carlson wrote:\n> Git has historically allowed either lowercase or uppercase hex for\n> object IDs, but it has always emitted only lowercase.  This has caused\n> people to expect only lowercase and not handle uppercase.\n> \n> As an example, Git's own example hooks look for \"[0-9a-f]\" in several\n> places, but there are many other Git-adjacent pieces of software,\n> including Gitolite, which make the assumption that object IDs are always\n> lowercase.  This is not to criticize the authors of these projects, but\n> rather to point out how common this assumption is.  In fact, it's so\n> common that we have only one test in our codebase that fails when we\n> reject uppercase object IDs.\n> \n> More critically, it leads people to make security-based assumptions that\n> an object ID either does not contain uppercase characters or that an\n> object ID can be expressed uniquely in hex form, neither of which are\n> currently true.  Git itself normally uses binary object IDs, which\n> avoids many of these problems, but most other projects deal primarily in\n> hex object IDs, so they are more affected.\n\nCan you say a bit more about the security problems please - I'm trying \nto understand why ABCDEF is a security risk when abcdef^0 isn't.\n\nThanks\n\nPhillip\n\n> In preparation for Git 3.0, only allow lowercase hex object IDs in\n> breaking changes mode and document this as well.  Update the single\n> failing test and add a new one to verify we reject new uppercase object\n> IDs.  Note that in t5324, we change the hex character from \"A\" to \"b\"\n> because in SHA-256 mode, \"a\" is the correct value, so our test_must_fail\n> assertion will unexpectedly succeed in that case.\n> \n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> ---\n>   Documentation/BreakingChanges.adoc | 5 +++++\n>   hex-ll.h                           | 4 ++++\n>   t/t1503-rev-parse-verify.sh        | 5 +++++\n>   t/t5324-split-commit-graph.sh      | 4 ++--\n>   4 files changed, 16 insertions(+), 2 deletions(-)\n> \n> diff --git a/Documentation/BreakingChanges.adoc b/Documentation/BreakingChanges.adoc\n> index 73bb939359..dbc46d14e3 100644\n> --- a/Documentation/BreakingChanges.adoc\n> +++ b/Documentation/BreakingChanges.adoc\n> @@ -171,6 +171,11 @@ JGit, libgit2 and Gitoxide need to support it.\n>     matches the default branch name used in new repositories by many of the\n>     big Git forges.\n>   \n> +* Git will accept hex object IDs only in lowercase. The fact that Git has\n> +\thistorically allowed uppercase characters in hex object IDs has been the\n> +\tsource of a variety of bugs and security problems in software using Git. We\n> +\tdon't expect most users to notice any change.\n> +\n>   * Git will require Rust as a mandatory part of the build process. While Git\n>     already started to adopt Rust in Git 2.49, all parts written in Rust are\n>     optional for the time being. This includes:\n> diff --git a/hex-ll.h b/hex-ll.h\n> index 9da76f17e8..2f9c8d7c25 100644\n> --- a/hex-ll.h\n> +++ b/hex-ll.h\n> @@ -6,7 +6,11 @@ enum hexkind {\n>   \tHEX_KIND_LOWER = 1,\n>   };\n>   \n> +#ifdef WITH_BREAKING_CHANGES\n> +#define HEX_KIND_OID HEX_KIND_LOWER\n> +#else\n>   #define HEX_KIND_OID HEX_KIND_MIXED\n> +#endif\n>   \n>   extern const signed char hexval_table[256];\n>   extern const signed char hexval_lc_table[256];\n> diff --git a/t/t1503-rev-parse-verify.sh b/t/t1503-rev-parse-verify.sh\n> index 87638a4a2c..f07b45de5a 100755\n> --- a/t/t1503-rev-parse-verify.sh\n> +++ b/t/t1503-rev-parse-verify.sh\n> @@ -60,6 +60,11 @@ test_expect_success 'works with one good rev' '\n>   \ttest \"$rev_head\" = \"$HASH4\"\n>   '\n>   \n> +test_expect_success WITH_BREAKING_CHANGES 'rejects uppercase revs' '\n> +\tUC_HASH=$(echo \"$HASH1\" | tr a-f A-F) &&\n> +\ttest_must_fail git rev-parse --verify \"$UC_HASH\"\n> +'\n> +\n>   test_expect_success 'fails with any bad rev or many good revs' '\n>   \ttest_must_fail git rev-parse --verify 2>error &&\n>   \ttest_grep \"single revision\" error &&\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index bf7ba0e558..29db815c77 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -349,7 +349,7 @@ test_expect_success 'verify after commit-graph-chain corruption (base)' '\n>   \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>   \t\tgrep -v \"^+\" test_err >err &&\n>   \t\ttest_grep \"invalid commit-graph chain\" err &&\n> -\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"A\" &&\n> +\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"a\" &&\n>   \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>   \t\tgrep -v \"^+\" test_err >err &&\n>   \t\ttest_grep \"unable to find all commit-graph files\" err\n> @@ -364,7 +364,7 @@ test_expect_success 'verify after commit-graph-chain corruption (tip)' '\n>   \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>   \t\tgrep -v \"^+\" test_err >err &&\n>   \t\ttest_grep \"invalid commit-graph chain\" err &&\n> -\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"A\" &&\n> +\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"b\" &&\n>   \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>   \t\tgrep -v \"^+\" test_err >err &&\n>   \t\ttest_grep \"unable to find all commit-graph files\" err\n> \n\n"},{"id":"551194","messageId":"xmqq5x0yp5ts.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-2-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 1/6] hex: add functionality for lowercase-only hex","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-25T15:39:43Z","receivedAt":"2026-08-25T15:39:46Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> We currently allow both upper and lower case for all hex values in Git.\n> However, in a future commit, we'll want to change that to allow only\n> lowercase values in some cases.  To prepare for that case, provide a\n> table to convert hex values using lowercase only and an enum to let us\n> choose which we want, wiring it up to the hexval function.\n>\n> For now, keep things completely the same by specifying only the\n> variant that accepts both lowercase and uppercase to avoid changing\n> behavior.\n>\n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> ---\n>  color.c    |  2 +-\n>  hex-ll.c   | 37 ++++++++++++++++++++++++++++++++++++-\n>  hex-ll.h   | 14 ++++++++++----\n>  pkt-line.c |  8 ++++----\n>  4 files changed, 51 insertions(+), 10 deletions(-)\n\n\nNow this is an embarrassingly late review.  I hope this is not a\nsign that nobody is paying attention on the list these days X-<.\n\n> +const signed char hexval_lc_table[256] = {\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 00-07 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 08-0f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 10-17 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 18-1f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 20-27 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 28-2f */\n> +\t  0,  1,  2,  3,  4,  5,  6,  7,\t\t/* 30-37 */\n> +\t  8,  9, -1, -1, -1, -1, -1, -1,\t\t/* 38-3f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 40-47 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 48-4f */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 50-57 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 58-5f */\n> +\t -1, 10, 11, 12, 13, 14, 15, -1,\t\t/* 60-67 */\n> +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 68-67 */\n\nThat's 68-6f if I am not mistaken ;-).\n\n"},{"id":"551196","messageId":"xmqqh5kinps1.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-5-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 4/6] hex: label usages of hex parsing for object IDs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-25T16:11:42Z","receivedAt":"2026-08-25T16:11:44Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> In preparation for a future change, label the hex parsing we're doing\n> for object IDs by defining a constant called HEX_KIND_OID.  This is\n> currently the same as HEX_KIND_MIXED, so there is no functional change\n> here.\n>\n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> ---\n>  diagnose.c    | 2 +-\n>  hex-ll.h      | 2 ++\n>  hex.c         | 2 +-\n>  http-push.c   | 4 ++--\n>  notes.c       | 2 +-\n>  object-file.c | 2 +-\n>  6 files changed, 8 insertions(+), 6 deletions(-)\n\nThis is a hard-to-review patch in the sense that what we see in the\npatch may be perfectly good, but we cannot see what is left out,\neither by mistake or by misdesign.  So I checked out the state with\nthis patch (and no later ones) applied, and eyeballed the output of\n\n    $ git grep -n -e HEX_KIND_MIXED\n\nAt this step, a few explicit uses of HEX_KIND_MIXED remain that I\nthink should have been converted to HEX_KIND_OID.\n\n * builtin/index-pack.c:repack_local_links() spawns a pack-objects\n   process and reads its output.  As we are reading from a known\n   version of Git (i.e., pack-objects that came with the index-pack\n   that runs this code), we do not need to be lenient and can use\n   HEX_KIND_OID here.\n\n * notes.c:load_subtree() has two calls to hex_to_bytes() to read\n   paths in a notes tree, and this patch updates only one to use\n   HEX_KIND_OID, leaving the other one HEX_KIND_MIXED, which we\n   probably should change at the same time (if there is a valid\n   reason, it deserves an in-code comment to explain it).\n\n\nThe remaining uses of HEX_KIND_MIXED look mostly OK.\n\n - color.c uses MIXED to decode things like #AAFF00, which will be\n   correct forever.\n\n - mailinfo.c uses MIXED to decode Q encoding, and we have no power\n   or business to forbid uppercase hex there.\n\n - pkt-line.c:packet_length() uses MIXED to decode the packet length\n   expressed in the four hex digits at the beginning.  We could\n   forbid uppercase hex there (our length bytes have always been\n   lowercase) if we wanted to, but HEX_KIND_OID is not the enum to\n   use to do so.\n\n - ref-filter.c:append_literal() is similar to the next one.\n\n - strbuf.c:strbuf_expand_literal() uses MIXED to decode %0A into line\n   feed, etc.  We could forbid uppercase hex there if we wanted to,\n   but HEX_KIND_OID is not the enum to use to do so.\n\n - url.c:url_decode_internal() uses MIXED to decode %2F into '/',\n   etc., and we have no power or business to forbid uppercase hex\n   there.\n\n - urlmatch.c:append_normalized_escapes() uses MIXED to decode %2F\n   into '/' before escaping it back with %02X.\n"},{"id":"551197","messageId":"xmqqcxv6npf0.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-6-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 5/6] object-name: use hexval","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-25T16:19:31Z","receivedAt":"2026-08-25T16:19:34Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> We've open-coded a different implementation of parsing hex values here\n> when we already have a perfectly good one in hexval.  This\n> implementation will almost certainly be slower because it isn't\n> table-driven, unlike the other one, and since it's not constant time it\n> has no other advantages either.  To tidy things up and prepare for\n> future work, switch to hexval in this case.\n>\n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> ---\n>  object-name.c | 13 +++----------\n>  1 file changed, 3 insertions(+), 10 deletions(-)\n>\n> diff --git a/object-name.c b/object-name.c\n> index 83efba0ba6..d2d81b3511 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n> @@ -236,17 +236,10 @@ static int parse_oid_prefix(const char *name, int len,\n>  {\n>  \tfor (int i = 0; i < len; i++) {\n>  \t\tunsigned char c = name[i];\n> -\t\tunsigned char val;\n> -\t\tif (c >= '0' && c <= '9') {\n> -\t\t\tval = c - '0';\n> -\t\t} else if (c >= 'a' && c <= 'f') {\n> -\t\t\tval = c - 'a' + 10;\n> -\t\t} else if (c >= 'A' && c <='F') {\n> -\t\t\tval = c - 'A' + 10;\n> -\t\t\tc -= 'A' - 'a';\n> -\t\t} else {\n> +\t\tint val = hexval(c, HEX_KIND_OID);\n> +\n> +\t\tif (val < 0)\n>  \t\t\treturn -1;\n> -\t\t}\n>  \n>  \t\tif (hex_out)\n>  \t\t\thex_out[i] = c;\n\nWhen hex_out[] is given by the caller, they used to get a downcased\nversion of object name.  After your planned transition to forbid\nuppercase hex, they will get an error, which is exactly as you\nintend.\n\nHowever, during transition, they will *not* get an error (as\nKIND_OID is still KIND_MIXED before the transition), and they will\nsee the hex_out[] filled with object names in the original case,\nwithout canonicalization that the original code gave them.\n\nWhile seemingly harmless, because repo_for_each_abbrev() doesn't\nseem to malfunction on uppercase string metadata for\ndisambiguation), we may want to mention that this changes API\ncontract (until we forbid uppercase input altogether).\n"},{"id":"551198","messageId":"xmqq8q5unomp.fsf@gitster.g","threadId":"66087","inReplyTo":"20260729233215.398654-7-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-25T16:36:30Z","receivedAt":"2026-08-25T16:36:33Z","isPatch":true,"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> IDs.  Note that in t5324, we change the hex character from \"A\" to \"b\"\n> because in SHA-256 mode, \"a\" is the correct value, so our test_must_fail\n> assertion will unexpectedly succeed in that case.\n\nThis was a bit hard to read and puzzled me, as you have two \"A\" and\nchange only one of them to \"b\".\n\nIs the idea that we wanted to make sure we use lowercase letters,\nbecause we do not want to see the tested \"verify\" command fail for\nnow-forbidden uppercase hex but we want the command to read the data\nas valid hex and fail because it notices the corruption?  So the\nfirst hunk is a no-op change (i.e., the first hash identifier on the\nfirst line is corrupt with the 30-th char in the file replaced with\neither 'a' or 'A'), while the second hunk is not (i.e., the second\nhash identifier on the second line in the file is corrupt with the\n70-th char in the file replaced with 'A' but it is OK with 'a'\nbecause in the SHA-256 mode, the correct character for the place\nhappens to be 'a')?  It is puzzling if that is the case, because\nwhat this series wanted to tighten was that we used to treat hex\nchars case insensitively.  So, if 'a' happened to be the right\nuncorrupted value for position 70, how did the original that\nreplaced it to 'A' tested a \"corrupted\" state?\n\n\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index bf7ba0e558..29db815c77 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -349,7 +349,7 @@ test_expect_success 'verify after commit-graph-chain corruption (base)' '\n>  \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>  \t\tgrep -v \"^+\" test_err >err &&\n>  \t\ttest_grep \"invalid commit-graph chain\" err &&\n> -\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"A\" &&\n> +\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"a\" &&\n>  \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>  \t\tgrep -v \"^+\" test_err >err &&\n>  \t\ttest_grep \"unable to find all commit-graph files\" err\n> @@ -364,7 +364,7 @@ test_expect_success 'verify after commit-graph-chain corruption (tip)' '\n>  \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>  \t\tgrep -v \"^+\" test_err >err &&\n>  \t\ttest_grep \"invalid commit-graph chain\" err &&\n> -\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"A\" &&\n> +\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"b\" &&\n>  \t\ttest_must_fail git commit-graph verify 2>test_err &&\n>  \t\tgrep -v \"^+\" test_err >err &&\n>  \t\ttest_grep \"unable to find all commit-graph files\" err\n\n"},{"id":"551232","messageId":"CABPp-BFDaWdahoOnNRGQjshzQXin1YLuROv94W_PrajnLWDAuQ@mail.gmail.com","threadId":"66087","inReplyTo":"20260729233215.398654-6-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 5/6] object-name: use hexval","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2026-08-25T19:44:34Z","receivedAt":"2026-08-25T19:44:47Z","isPatch":true,"body":"On Wed, Jul 29, 2026 at 4:33 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> We've open-coded a different implementation of parsing hex values here\n> when we already have a perfectly good one in hexval.  This\n> implementation will almost certainly be slower because it isn't\n> table-driven, unlike the other one, and since it's not constant time it\n> has no other advantages either.  To tidy things up and prepare for\n> future work, switch to hexval in this case.\n\nAs Junio noted, you may want to call out that your replacement drops\nthe case-normalization that the former parse_oid_prefix() provided.\n\n[...]\n> -               unsigned char val;\n[...]\n> +               int val = hexval(c, HEX_KIND_OID);\n> +\n> +               if (val < 0)\n>                         return -1;\n[...]\n>                if (oid_out) {\n>                        if (!(i & 1))\n>                                val <<= 4;\n>                        oid_out->hash[i >> 1] |= val;\n\nhexval returns unsigned int.  Is there a risk that someone \"tries to\nfix\" that discrepancy by changing val to unsigned int here,\ninadvertently causing the `if` immediately below to become dead code?\n\nIn patch 1, in hex2chr, you used a (val & ~0xf) check together with an\nunsigned int val; would that make sense here, or is that overkill?\n"},{"id":"551233","messageId":"CABPp-BEAx+YZ547ig52EQaB65Yg6aEXb0qdLsWsChekhacqCSw@mail.gmail.com","threadId":"66087","inReplyTo":"20260729233215.398654-7-sandals@crustytoothpaste.net","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2026-08-25T19:44:43Z","receivedAt":"2026-08-25T19:44:56Z","isPatch":true,"body":"On Wed, Jul 29, 2026 at 4:33 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> In preparation for Git 3.0, only allow lowercase hex object IDs in\n> breaking changes mode and document this as well.  Update the single\n> failing test and add a new one to verify we reject new uppercase object\n> IDs.  Note that in t5324, we change the hex character from \"A\" to \"b\"\n> because in SHA-256 mode, \"a\" is the correct value, so our test_must_fail\n> assertion will unexpectedly succeed in that case.\n[...snip...]\n> -               corrupt_file \"$graphdir/commit-graph-chain\" 30 \"A\" &&\n> +               corrupt_file \"$graphdir/commit-graph-chain\" 30 \"a\" &&\n[...]\n> -               corrupt_file \"$graphdir/commit-graph-chain\" 70 \"A\" &&\n> +               corrupt_file \"$graphdir/commit-graph-chain\" 70 \"b\" &&\n\nwhich \"A\" is the commit message referring to?\n\n> +* Git will accept hex object IDs only in lowercase. The fact that Git has\n> +       historically allowed uppercase characters in hex object IDs has been the\n> +       source of a variety of bugs and security problems in software using Git. We\n> +       don't expect most users to notice any change.\n\nYou've indented with tabs here while the surrounding paragraphs use\nspaces; is that going to mess up rendering?\n"},{"id":"551249","messageId":"ao4K44RP66mjnpd7@fruit.crustytoothpaste.net","threadId":"66087","inReplyTo":"d6940aa6-9336-481b-8ee5-5e3d9f3d3a50@gmail.com","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-08-25T21:36:36Z","receivedAt":"2026-08-25T21:36:39Z","isPatch":true,"body":"On 2026-08-25 at 09:04:36, Phillip Wood wrote:\n> Hi brian\n> \n> On 30/07/2026 00:32, brian m. carlson wrote:\n> > Git has historically allowed either lowercase or uppercase hex for\n> > object IDs, but it has always emitted only lowercase.  This has caused\n> > people to expect only lowercase and not handle uppercase.\n> > \n> > As an example, Git's own example hooks look for \"[0-9a-f]\" in several\n> > places, but there are many other Git-adjacent pieces of software,\n> > including Gitolite, which make the assumption that object IDs are always\n> > lowercase.  This is not to criticize the authors of these projects, but\n> > rather to point out how common this assumption is.  In fact, it's so\n> > common that we have only one test in our codebase that fails when we\n> > reject uppercase object IDs.\n> > \n> > More critically, it leads people to make security-based assumptions that\n> > an object ID either does not contain uppercase characters or that an\n> > object ID can be expressed uniquely in hex form, neither of which are\n> > currently true.  Git itself normally uses binary object IDs, which\n> > avoids many of these problems, but most other projects deal primarily in\n> > hex object IDs, so they are more affected.\n> \n> Can you say a bit more about the security problems please - I'm trying to\n> understand why ABCDEF is a security risk when abcdef^0 isn't.\n\nThere's two cases I've seen.  The first is that people assume an object\nID is unique in hex form.  So if we have some policy to enforce, say,\nthat we can't allow certain objects, people will check against the\nlowercase version when they may get the uppercase version somewhere\n(say, user input or a specially crafted protocol message), which\nbypasses the check.\n\nThe other case is where we try to distinguish between an object ID and a\nref, branch, or tag.  If our regexp has `[0-9a-f]{40}` or `[0-9a-f]{64}`\nand we assume that if it matches it's an object ID and if it's not it's\na ref, that's not correct here.  We'd need to match the uppercase\nversion as well, but experience shows that people overwhelmingly do not\ndo that.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"551251","messageId":"ao4MJQgp6Ai4tJxi@fruit.crustytoothpaste.net","threadId":"66087","inReplyTo":"CABPp-BFDaWdahoOnNRGQjshzQXin1YLuROv94W_PrajnLWDAuQ@mail.gmail.com","subject":"Re: [RFC PATCH 5/6] object-name: use hexval","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-08-25T21:41:57Z","receivedAt":"2026-08-25T21:41:58Z","isPatch":true,"body":"On 2026-08-25 at 19:44:34, Elijah Newren wrote:\n> On Wed, Jul 29, 2026 at 4:33 PM brian m. carlson\n> <sandals@crustytoothpaste.net> wrote:\n> >\n> > We've open-coded a different implementation of parsing hex values here\n> > when we already have a perfectly good one in hexval.  This\n> > implementation will almost certainly be slower because it isn't\n> > table-driven, unlike the other one, and since it's not constant time it\n> > has no other advantages either.  To tidy things up and prepare for\n> > future work, switch to hexval in this case.\n> \n> As Junio noted, you may want to call out that your replacement drops\n> the case-normalization that the former parse_oid_prefix() provided.\n\nWill fix in v2.\n\n> [...]\n> > -               unsigned char val;\n> [...]\n> > +               int val = hexval(c, HEX_KIND_OID);\n> > +\n> > +               if (val < 0)\n> >                         return -1;\n> [...]\n> >                if (oid_out) {\n> >                        if (!(i & 1))\n> >                                val <<= 4;\n> >                        oid_out->hash[i >> 1] |= val;\n> \n> hexval returns unsigned int.  Is there a risk that someone \"tries to\n> fix\" that discrepancy by changing val to unsigned int here,\n> inadvertently causing the `if` immediately below to become dead code?\n> \n> In patch 1, in hex2chr, you used a (val & ~0xf) check together with an\n> unsigned int val; would that make sense here, or is that overkill?\n\nI can re-roll with an appropriate change, sure.  I think that we'd need\nto have a slightly different check, but I'll tidy it up accordingly.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"551252","messageId":"ao4MqtDxZJaMEBBI@fruit.crustytoothpaste.net","threadId":"66087","inReplyTo":"xmqq5x0yp5ts.fsf@gitster.g","subject":"Re: [RFC PATCH 1/6] hex: add functionality for lowercase-only hex","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-08-25T21:44:11Z","receivedAt":"2026-08-25T21:44:12Z","isPatch":true,"body":"On 2026-08-25 at 15:39:43, Junio C Hamano wrote:\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > We currently allow both upper and lower case for all hex values in Git.\n> > However, in a future commit, we'll want to change that to allow only\n> > lowercase values in some cases.  To prepare for that case, provide a\n> > table to convert hex values using lowercase only and an enum to let us\n> > choose which we want, wiring it up to the hexval function.\n> >\n> > For now, keep things completely the same by specifying only the\n> > variant that accepts both lowercase and uppercase to avoid changing\n> > behavior.\n> >\n> > Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> > ---\n> >  color.c    |  2 +-\n> >  hex-ll.c   | 37 ++++++++++++++++++++++++++++++++++++-\n> >  hex-ll.h   | 14 ++++++++++----\n> >  pkt-line.c |  8 ++++----\n> >  4 files changed, 51 insertions(+), 10 deletions(-)\n> \n> \n> Now this is an embarrassingly late review.  I hope this is not a\n> sign that nobody is paying attention on the list these days X-<.\n> \n> > +const signed char hexval_lc_table[256] = {\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 00-07 */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 08-0f */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 10-17 */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 18-1f */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 20-27 */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 28-2f */\n> > +\t  0,  1,  2,  3,  4,  5,  6,  7,\t\t/* 30-37 */\n> > +\t  8,  9, -1, -1, -1, -1, -1, -1,\t\t/* 38-3f */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 40-47 */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 48-4f */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 50-57 */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 58-5f */\n> > +\t -1, 10, 11, 12, 13, 14, 15, -1,\t\t/* 60-67 */\n> > +\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 68-67 */\n> \n> That's 68-6f if I am not mistaken ;-).\n\nSo it is.  Will fix in v2.\n\nI think I accidentally included the uppercase but not lowercase variants\nwhen creating the original array and then copied and pasted the line,\nbut messed up the comment.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"552142","messageId":"9fb6317d-ce64-4f53-bce2-81cb2dc12056@gmail.com","threadId":"66087","inReplyTo":"ao4K44RP66mjnpd7@fruit.crustytoothpaste.net","subject":"Re: [RFC PATCH 6/6] hex: allow only lowercase object IDs in breaking changes mode","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-09-07T13:37:04Z","receivedAt":"2026-09-07T13:37:07Z","isPatch":true,"body":"Hi brian\n\nThanks for the examples, sorry it has taken so long for me to respond\n\nOn 25/08/2026 22:36, brian m. carlson wrote:\n> On 2026-08-25 at 09:04:36, Phillip Wood wrote:\n>> Hi brian\n>>\n>> On 30/07/2026 00:32, brian m. carlson wrote:\n>>> Git has historically allowed either lowercase or uppercase hex for\n>>> object IDs, but it has always emitted only lowercase.  This has caused\n>>> people to expect only lowercase and not handle uppercase.\n>>>\n>>> As an example, Git's own example hooks look for \"[0-9a-f]\" in several\n>>> places, but there are many other Git-adjacent pieces of software,\n>>> including Gitolite, which make the assumption that object IDs are always\n>>> lowercase.  This is not to criticize the authors of these projects, but\n>>> rather to point out how common this assumption is.  In fact, it's so\n>>> common that we have only one test in our codebase that fails when we\n>>> reject uppercase object IDs.\n>>>\n>>> More critically, it leads people to make security-based assumptions that\n>>> an object ID either does not contain uppercase characters or that an\n>>> object ID can be expressed uniquely in hex form, neither of which are\n>>> currently true.  Git itself normally uses binary object IDs, which\n>>> avoids many of these problems, but most other projects deal primarily in\n>>> hex object IDs, so they are more affected.\n>>\n>> Can you say a bit more about the security problems please - I'm trying to\n>> understand why ABCDEF is a security risk when abcdef^0 isn't.\n> \n> There's two cases I've seen.  The first is that people assume an object\n> ID is unique in hex form.  So if we have some policy to enforce, say,\n> that we can't allow certain objects, people will check against the\n> lowercase version when they may get the uppercase version somewhere\n> (say, user input or a specially crafted protocol message), which\n> bypasses the check.\n\nI'm a bit unclear how upper case hex can defeat that policy but ref \nnames dont. If the input is not being checked to ensure it is a hex \nobject id wont a ref pointing to a commit we're trying to restrict \naccess to also defeat the check?\n\n> The other case is where we try to distinguish between an object ID and a\n> ref, branch, or tag.  If our regexp has `[0-9a-f]{40}` or `[0-9a-f]{64}`\n> and we assume that if it matches it's an object ID and if it's not it's\n> a ref, that's not correct here.  We'd need to match the uppercase\n> version as well, but experience shows that people overwhelmingly do not\n> do that.\n\nThat makes more sense to me. It also makes me wonder if we should forbid \nrefnames where the last component looks like an object id.\n\nThanks\n\nPhillip\n\n"},{"id":"552157","messageId":"20260907195941.1024289-1-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260729233215.398654-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 0/7] Git 3.0: restrict hex object IDs to lowercase only","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:33Z","receivedAt":"2026-09-07T19:59:58Z","isPatch":true,"body":"As far as I can tell, Git has always emitted hex object IDs in\nlowercase, but our object ID parser accepts both uppercase and\nlowercase.  This leads to much software relying on hex object IDs being\nbroken because it doesn't handle uppercase object IDs and this can even\nlead to security problems when people assume that an object ID has a\nunique hex form.\n\nThis series removes the ability to use uppercase hex in\nobject IDs in Git 3.0.\n\nChanges from v1:\n\n* Fix incorrect hex range in comment.\n* Update commit messages.\n* Move t5324 fixes to a separate commit and expand commit message.\n* Restore lowercasing in parse_oid_prefix.\n* Possibly other miscellaneous changes which I have forgotten.\n\nbrian m. carlson (7):\n  hex: add functionality for lowercase-only hex\n  hex: allow specifying hex type with hex2chr\n  hex: make hex_to_bytes accept kind of hex to use\n  hex: label usages of hex parsing for object IDs\n  object-name: use hexval\n  t5324: adjust tests for corrupt commit-graph\n  hex: allow only lowercase object IDs in breaking changes mode\n\n Documentation/BreakingChanges.adoc |  5 ++++\n builtin/index-pack.c               |  2 +-\n color.c                            |  2 +-\n diagnose.c                         |  2 +-\n hex-ll.c                           | 39 ++++++++++++++++++++++++++++--\n hex-ll.h                           | 24 +++++++++++++-----\n hex.c                              |  2 +-\n http-push.c                        |  5 ++--\n mailinfo.c                         |  2 +-\n notes.c                            |  5 ++--\n object-file.c                      |  2 +-\n object-name.c                      | 15 +++---------\n pkt-line.c                         |  8 +++---\n ref-filter.c                       |  2 +-\n strbuf.c                           |  2 +-\n t/t1503-rev-parse-verify.sh        |  5 ++++\n t/t5324-split-commit-graph.sh      |  4 +--\n url.c                              |  2 +-\n urlmatch.c                         |  2 +-\n 19 files changed, 91 insertions(+), 39 deletions(-)\n\n"},{"id":"552158","messageId":"20260907195941.1024289-5-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260907195941.1024289-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 4/7] hex: label usages of hex parsing for object IDs","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:37Z","receivedAt":"2026-09-07T19:59:58Z","isPatch":true,"body":"In preparation for a future change, label the hex parsing we're doing\nfor object IDs by defining a constant called HEX_KIND_OID.  This is\ncurrently the same as HEX_KIND_MIXED, so there is no functional change\nhere.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n diagnose.c    | 2 +-\n hex-ll.h      | 2 ++\n hex.c         | 2 +-\n http-push.c   | 4 ++--\n notes.c       | 2 +-\n object-file.c | 2 +-\n 6 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/diagnose.c b/diagnose.c\nindex fc11cea229..9c652d36a6 100644\n--- a/diagnose.c\n+++ b/diagnose.c\n@@ -112,7 +112,7 @@ static void loose_objs_stats(struct strbuf *buf, const char *path)\n \twhile ((e = readdir_skip_dot_and_dotdot(dir)) != NULL)\n \t\tif (get_dtype(e, &count_path, 0) == DT_DIR &&\n \t\t    strlen(e->d_name) == 2 &&\n-\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_MIXED)) {\n+\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_OID)) {\n \t\t\tstrbuf_setlen(&count_path, base_path_len);\n \t\t\tstrbuf_addf(&count_path, \"%s/\", e->d_name);\n \t\t\ttotal += (count = count_files(&count_path));\ndiff --git a/hex-ll.h b/hex-ll.h\nindex fe698f0c76..9da76f17e8 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -6,6 +6,8 @@ enum hexkind {\n \tHEX_KIND_LOWER = 1,\n };\n \n+#define HEX_KIND_OID HEX_KIND_MIXED\n+\n extern const signed char hexval_table[256];\n extern const signed char hexval_lc_table[256];\n static inline unsigned int hexval(unsigned char c, enum hexkind kind)\ndiff --git a/hex.c b/hex.c\nindex 6150bdcbf8..4e1e81af3f 100644\n--- a/hex.c\n+++ b/hex.c\n@@ -9,7 +9,7 @@ static int get_hash_hex_algop(const char *hex, unsigned char *hash,\n \t\t\t      const struct git_hash_algo *algop)\n {\n \tfor (size_t i = 0; i < algop->rawsz; i++) {\n-\t\tint val = hex2chr(hex, HEX_KIND_MIXED);\n+\t\tint val = hex2chr(hex, HEX_KIND_OID);\n \t\tif (val < 0)\n \t\t\treturn -1;\n \t\t*hash++ = val;\ndiff --git a/http-push.c b/http-push.c\nindex b5c5bad3db..43b4b61c70 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1031,13 +1031,13 @@ static int get_oid_hex_from_objpath(const char *path, struct object_id *oid)\n \tif (strlen(path) != the_hash_algo->hexsz + 1)\n \t\treturn -1;\n \n-\tif (hex_to_bytes(oid->hash, path, 1, HEX_KIND_MIXED))\n+\tif (hex_to_bytes(oid->hash, path, 1, HEX_KIND_OID))\n \t\treturn -1;\n \tpath += 2;\n \tpath++; /* skip '/' */\n \n \treturn hex_to_bytes(oid->hash + 1, path, the_hash_algo->rawsz - 1,\n-\t\t\t    HEX_KIND_MIXED);\n+\t\t\t    HEX_KIND_OID);\n }\n \n static void process_ls_object(struct remote_ls_ctx *ls)\ndiff --git a/notes.c b/notes.c\nindex 99b8b15d81..7e9e3eb2d2 100644\n--- a/notes.c\n+++ b/notes.c\n@@ -443,7 +443,7 @@ static void load_subtree(struct notes_tree *t, struct leaf_node *subtree,\n \t\t\t\tgoto handle_non_note;\n \n \t\t\tif (hex_to_bytes(object_oid.hash + len++, entry.path, 1,\n-\t\t\t\t\t HEX_KIND_MIXED))\n+\t\t\t\t\t HEX_KIND_OID))\n \t\t\t\tgoto handle_non_note; /* entry.path is not a SHA1 */\n \n \t\t\t/*\ndiff --git a/object-file.c b/object-file.c\nindex 892be4bbb2..c4bb3263cc 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1079,7 +1079,7 @@ int for_each_file_in_obj_subdir(unsigned int subdir_nr,\n \t\tstrbuf_add(path, de->d_name, namelen);\n \t\tif (namelen == algop->hexsz - 2 &&\n \t\t    !hex_to_bytes(oid.hash + 1, de->d_name,\n-\t\t\t\t  algop->rawsz - 1, HEX_KIND_MIXED)) {\n+\t\t\t\t  algop->rawsz - 1, HEX_KIND_OID)) {\n \t\t\toid_set_algo(&oid, algop);\n \t\t\tmemset(oid.hash + algop->rawsz, 0,\n \t\t\t       GIT_MAX_RAWSZ - algop->rawsz);\n"},{"id":"552159","messageId":"20260907195941.1024289-3-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260907195941.1024289-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 2/7] hex: allow specifying hex type with hex2chr","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:35Z","receivedAt":"2026-09-07T19:59:58Z","isPatch":true,"body":"We have several places where we use hex2chr.  One of those is parsing\nobject IDs, but others decode quoted-printable or percent encoding.  All\nof them accept both uppercase and lowercase hex.\n\nIn a future commit, we'll change some of these cases, so make hex2chr\naccept the kind of encoding to use: lowercase only hex or any kind of\nhex.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n hex-ll.h     | 6 +++---\n hex.c        | 2 +-\n mailinfo.c   | 2 +-\n ref-filter.c | 2 +-\n strbuf.c     | 2 +-\n url.c        | 2 +-\n urlmatch.c   | 2 +-\n 7 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/hex-ll.h b/hex-ll.h\nindex da1b5239b2..26847c7b2f 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -17,10 +17,10 @@ static inline unsigned int hexval(unsigned char c, enum hexkind kind)\n  * Convert two consecutive hexadecimal digits into a char.  Return a\n  * negative value on error.  Don't run over the end of short strings.\n  */\n-static inline int hex2chr(const char *s)\n+static inline int hex2chr(const char *s, enum hexkind kind)\n {\n-\tunsigned int val = hexval(s[0], HEX_KIND_MIXED);\n-\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1], HEX_KIND_MIXED);\n+\tunsigned int val = hexval(s[0], kind);\n+\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1], kind);\n }\n \n /*\ndiff --git a/hex.c b/hex.c\nindex f02832140d..6150bdcbf8 100644\n--- a/hex.c\n+++ b/hex.c\n@@ -9,7 +9,7 @@ static int get_hash_hex_algop(const char *hex, unsigned char *hash,\n \t\t\t      const struct git_hash_algo *algop)\n {\n \tfor (size_t i = 0; i < algop->rawsz; i++) {\n-\t\tint val = hex2chr(hex);\n+\t\tint val = hex2chr(hex, HEX_KIND_MIXED);\n \t\tif (val < 0)\n \t\t\treturn -1;\n \t\t*hash++ = val;\ndiff --git a/mailinfo.c b/mailinfo.c\nindex 13949ff31e..85c3119048 100644\n--- a/mailinfo.c\n+++ b/mailinfo.c\n@@ -396,7 +396,7 @@ static int decode_q_segment(struct strbuf *out, const struct strbuf *q_seg,\n \t\t\tint ch, d = *in;\n \t\t\tif (d == '\\n' || !d)\n \t\t\t\tbreak; /* drop trailing newline */\n-\t\t\tch = hex2chr(in);\n+\t\t\tch = hex2chr(in, HEX_KIND_MIXED);\n \t\t\tif (ch >= 0) {\n \t\t\t\tstrbuf_addch(out, ch);\n \t\t\t\tin += 2;\ndiff --git a/ref-filter.c b/ref-filter.c\nindex bdf54f6f59..cd02677e42 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -3636,7 +3636,7 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n \t\t\tif (cp[1] == '%')\n \t\t\t\tcp++;\n \t\t\telse {\n-\t\t\t\tint ch = hex2chr(cp + 1);\n+\t\t\t\tint ch = hex2chr(cp + 1, HEX_KIND_MIXED);\n \t\t\t\tif (0 <= ch) {\n \t\t\t\t\tstrbuf_addch(s, ch);\n \t\t\t\t\tcp += 3;\ndiff --git a/strbuf.c b/strbuf.c\nindex 44955669e8..88d23f8ac5 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -457,7 +457,7 @@ size_t strbuf_expand_literal(struct strbuf *sb, const char *placeholder)\n \t\treturn 1;\n \tcase 'x':\n \t\t/* %x00 == NUL, %x0a == LF, etc. */\n-\t\tch = hex2chr(placeholder + 1);\n+\t\tch = hex2chr(placeholder + 1, HEX_KIND_MIXED);\n \t\tif (ch < 0)\n \t\t\treturn 0;\n \t\tstrbuf_addch(sb, ch);\ndiff --git a/url.c b/url.c\nindex a59818278f..b4d72f784a 100644\n--- a/url.c\n+++ b/url.c\n@@ -62,7 +62,7 @@ static char *url_decode_internal(const char **query, int len,\n \t\t}\n \n \t\tif (c == '%' && (len < 0 || len >= 3)) {\n-\t\t\tint val = hex2chr(q + 1);\n+\t\t\tint val = hex2chr(q + 1, HEX_KIND_MIXED);\n \t\t\tif (0 < val) {\n \t\t\t\tstrbuf_addch(out, val);\n \t\t\t\tq += 3;\ndiff --git a/urlmatch.c b/urlmatch.c\nindex 20bc2d009c..989f1d794b 100644\n--- a/urlmatch.c\n+++ b/urlmatch.c\n@@ -50,7 +50,7 @@ static int append_normalized_escapes(struct strbuf *buf,\n \t\tif (ch == '%') {\n \t\t\tif (from_len < 2)\n \t\t\t\treturn 0;\n-\t\t\tch = hex2chr(from);\n+\t\t\tch = hex2chr(from, HEX_KIND_MIXED);\n \t\t\tif (ch < 0)\n \t\t\t\treturn 0;\n \t\t\tfrom += 2;\n"},{"id":"552160","messageId":"20260907195941.1024289-4-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260907195941.1024289-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 3/7] hex: make hex_to_bytes accept kind of hex to use","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:36Z","receivedAt":"2026-09-07T19:59:58Z","isPatch":true,"body":"Similarly to the previous commit, introduce an option for hex_to_bytes\nto allow us to specify the kind of hex to use: lowercase only or not.\nFor now, everything remains the same as before, but we will change\nthings in a future commit.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n builtin/index-pack.c | 2 +-\n diagnose.c           | 2 +-\n hex-ll.c             | 4 ++--\n hex-ll.h             | 2 +-\n http-push.c          | 5 +++--\n notes.c              | 5 +++--\n object-file.c        | 2 +-\n 7 files changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 6b2a87e2d3..44ac72ef1e 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1866,7 +1866,7 @@ static void repack_local_links(void)\n \twhile (strbuf_getline_lf(&line, out) != EOF) {\n \t\tunsigned char binary[GIT_MAX_RAWSZ];\n \t\tif (line.len != the_hash_algo->hexsz ||\n-\t\t    !hex_to_bytes(binary, line.buf, line.len))\n+\t\t    !hex_to_bytes(binary, line.buf, line.len, HEX_KIND_MIXED))\n \t\t\tdie(_(\"index-pack: Expecting full hex object ID lines only from pack-objects.\"));\n \n \t\t/*\ndiff --git a/diagnose.c b/diagnose.c\nindex 5092bf80d3..fc11cea229 100644\n--- a/diagnose.c\n+++ b/diagnose.c\n@@ -112,7 +112,7 @@ static void loose_objs_stats(struct strbuf *buf, const char *path)\n \twhile ((e = readdir_skip_dot_and_dotdot(dir)) != NULL)\n \t\tif (get_dtype(e, &count_path, 0) == DT_DIR &&\n \t\t    strlen(e->d_name) == 2 &&\n-\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t    !hex_to_bytes(&c, e->d_name, 1, HEX_KIND_MIXED)) {\n \t\t\tstrbuf_setlen(&count_path, base_path_len);\n \t\t\tstrbuf_addf(&count_path, \"%s/\", e->d_name);\n \t\t\ttotal += (count = count_files(&count_path));\ndiff --git a/hex-ll.c b/hex-ll.c\nindex 8f5a4e4644..3b7d22a824 100644\n--- a/hex-ll.c\n+++ b/hex-ll.c\n@@ -71,10 +71,10 @@ const signed char hexval_lc_table[256] = {\n \t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n };\n \n-int hex_to_bytes(unsigned char *binary, const char *hex, size_t len)\n+int hex_to_bytes(unsigned char *binary, const char *hex, size_t len, enum hexkind kind)\n {\n \tfor (; len; len--, hex += 2) {\n-\t\tunsigned int val = (hexval(hex[0], HEX_KIND_MIXED) << 4) | hexval(hex[1], HEX_KIND_MIXED);\n+\t\tunsigned int val = (hexval(hex[0], kind) << 4) | hexval(hex[1], kind);\n \n \t\tif (val & ~0xff)\n \t\t\treturn -1;\ndiff --git a/hex-ll.h b/hex-ll.h\nindex 26847c7b2f..fe698f0c76 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -28,6 +28,6 @@ static inline int hex2chr(const char *s, enum hexkind kind)\n  * values to `binary` as `len` bytes. Return 0 on success, or -1 if\n  * the input does not consist of hex digits).\n  */\n-int hex_to_bytes(unsigned char *binary, const char *hex, size_t len);\n+int hex_to_bytes(unsigned char *binary, const char *hex, size_t len, enum hexkind kind);\n \n #endif\ndiff --git a/http-push.c b/http-push.c\nindex b8f3faaed9..b5c5bad3db 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1031,12 +1031,13 @@ static int get_oid_hex_from_objpath(const char *path, struct object_id *oid)\n \tif (strlen(path) != the_hash_algo->hexsz + 1)\n \t\treturn -1;\n \n-\tif (hex_to_bytes(oid->hash, path, 1))\n+\tif (hex_to_bytes(oid->hash, path, 1, HEX_KIND_MIXED))\n \t\treturn -1;\n \tpath += 2;\n \tpath++; /* skip '/' */\n \n-\treturn hex_to_bytes(oid->hash + 1, path, the_hash_algo->rawsz - 1);\n+\treturn hex_to_bytes(oid->hash + 1, path, the_hash_algo->rawsz - 1,\n+\t\t\t    HEX_KIND_MIXED);\n }\n \n static void process_ls_object(struct remote_ls_ctx *ls)\ndiff --git a/notes.c b/notes.c\nindex ec9c2cb150..99b8b15d81 100644\n--- a/notes.c\n+++ b/notes.c\n@@ -428,7 +428,7 @@ static void load_subtree(struct notes_tree *t, struct leaf_node *subtree,\n \t\t\t\tgoto handle_non_note;\n \n \t\t\tif (hex_to_bytes(object_oid.hash + prefix_len, entry.path,\n-\t\t\t\t\t hashsz - prefix_len))\n+\t\t\t\t\t hashsz - prefix_len, HEX_KIND_MIXED))\n \t\t\t\tgoto handle_non_note; /* entry.path is not a SHA1 */\n \n \t\t\tmemset(object_oid.hash + hashsz, 0, GIT_MAX_RAWSZ - hashsz);\n@@ -442,7 +442,8 @@ static void load_subtree(struct notes_tree *t, struct leaf_node *subtree,\n \t\t\t\t/* internal nodes must be trees */\n \t\t\t\tgoto handle_non_note;\n \n-\t\t\tif (hex_to_bytes(object_oid.hash + len++, entry.path, 1))\n+\t\t\tif (hex_to_bytes(object_oid.hash + len++, entry.path, 1,\n+\t\t\t\t\t HEX_KIND_MIXED))\n \t\t\t\tgoto handle_non_note; /* entry.path is not a SHA1 */\n \n \t\t\t/*\ndiff --git a/object-file.c b/object-file.c\nindex a4cbf8b081..892be4bbb2 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1079,7 +1079,7 @@ int for_each_file_in_obj_subdir(unsigned int subdir_nr,\n \t\tstrbuf_add(path, de->d_name, namelen);\n \t\tif (namelen == algop->hexsz - 2 &&\n \t\t    !hex_to_bytes(oid.hash + 1, de->d_name,\n-\t\t\t\t  algop->rawsz - 1)) {\n+\t\t\t\t  algop->rawsz - 1, HEX_KIND_MIXED)) {\n \t\t\toid_set_algo(&oid, algop);\n \t\t\tmemset(oid.hash + algop->rawsz, 0,\n \t\t\t       GIT_MAX_RAWSZ - algop->rawsz);\n"},{"id":"552161","messageId":"20260907195941.1024289-2-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260907195941.1024289-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 1/7] hex: add functionality for lowercase-only hex","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:34Z","receivedAt":"2026-09-07T19:59:58Z","isPatch":true,"body":"We currently allow both upper and lower case for all hex values in Git.\nHowever, in a future commit, we'll want to change that to allow only\nlowercase values in some cases.  To prepare for that case, provide a\ntable to convert hex values using lowercase only and an enum to let us\nchoose which we want, wiring it up to the hexval function.\n\nFor now, keep things completely the same by specifying only the\nvariant that accepts both lowercase and uppercase to avoid changing\nbehavior.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n color.c    |  2 +-\n hex-ll.c   | 37 ++++++++++++++++++++++++++++++++++++-\n hex-ll.h   | 14 ++++++++++----\n pkt-line.c |  8 ++++----\n 4 files changed, 51 insertions(+), 10 deletions(-)\n\ndiff --git a/color.c b/color.c\nindex 00b53f97ac..9015d0faf1 100644\n--- a/color.c\n+++ b/color.c\n@@ -72,7 +72,7 @@ static int get_hex_color(const char **inp, int width, unsigned char *out)\n \tunsigned int val;\n \n \tassert(width == 1 || width == 2);\n-\tval = (hexval(in[0]) << 4) | hexval(in[width - 1]);\n+\tval = (hexval(in[0], HEX_KIND_MIXED) << 4) | hexval(in[width - 1], HEX_KIND_MIXED);\n \tif (val & ~0xff)\n \t\treturn -1;\n \t*inp += width;\ndiff --git a/hex-ll.c b/hex-ll.c\nindex 4d7ece1de5..8f5a4e4644 100644\n--- a/hex-ll.c\n+++ b/hex-ll.c\n@@ -36,10 +36,45 @@ const signed char hexval_table[256] = {\n \t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n };\n \n+const signed char hexval_lc_table[256] = {\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 00-07 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 08-0f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 10-17 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 18-1f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 20-27 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 28-2f */\n+\t  0,  1,  2,  3,  4,  5,  6,  7,\t\t/* 30-37 */\n+\t  8,  9, -1, -1, -1, -1, -1, -1,\t\t/* 38-3f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 40-47 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 48-4f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 50-57 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 58-5f */\n+\t -1, 10, 11, 12, 13, 14, 15, -1,\t\t/* 60-67 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 68-6f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 70-77 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 78-7f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 80-87 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 88-8f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 90-97 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* 98-9f */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* a0-a7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* a8-af */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* b0-b7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* b8-bf */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* c0-c7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* c8-cf */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* d0-d7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* d8-df */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* e0-e7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* e8-ef */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f0-f7 */\n+\t -1, -1, -1, -1, -1, -1, -1, -1,\t\t/* f8-ff */\n+};\n+\n int hex_to_bytes(unsigned char *binary, const char *hex, size_t len)\n {\n \tfor (; len; len--, hex += 2) {\n-\t\tunsigned int val = (hexval(hex[0]) << 4) | hexval(hex[1]);\n+\t\tunsigned int val = (hexval(hex[0], HEX_KIND_MIXED) << 4) | hexval(hex[1], HEX_KIND_MIXED);\n \n \t\tif (val & ~0xff)\n \t\t\treturn -1;\ndiff --git a/hex-ll.h b/hex-ll.h\nindex a381fa8556..da1b5239b2 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -1,10 +1,16 @@\n #ifndef HEX_LL_H\n #define HEX_LL_H\n \n+enum hexkind {\n+\tHEX_KIND_MIXED = 0,\n+\tHEX_KIND_LOWER = 1,\n+};\n+\n extern const signed char hexval_table[256];\n-static inline unsigned int hexval(unsigned char c)\n+extern const signed char hexval_lc_table[256];\n+static inline unsigned int hexval(unsigned char c, enum hexkind kind)\n {\n-\treturn hexval_table[c];\n+\treturn kind == HEX_KIND_MIXED ? hexval_table[c] : hexval_lc_table[c];\n }\n \n /*\n@@ -13,8 +19,8 @@ static inline unsigned int hexval(unsigned char c)\n  */\n static inline int hex2chr(const char *s)\n {\n-\tunsigned int val = hexval(s[0]);\n-\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1]);\n+\tunsigned int val = hexval(s[0], HEX_KIND_MIXED);\n+\treturn (val & ~0xf) ? val : (val << 4) | hexval(s[1], HEX_KIND_MIXED);\n }\n \n /*\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 3fc3e9ea70..338075558c 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -378,10 +378,10 @@ int packet_length(const char lenbuf_hex[4], size_t size)\n {\n \tif (size < 4)\n \t\tBUG(\"buffer too small\");\n-\treturn\thexval(lenbuf_hex[0]) << 12 |\n-\t\thexval(lenbuf_hex[1]) <<  8 |\n-\t\thexval(lenbuf_hex[2]) <<  4 |\n-\t\thexval(lenbuf_hex[3]);\n+\treturn\thexval(lenbuf_hex[0], HEX_KIND_MIXED) << 12 |\n+\t\thexval(lenbuf_hex[1], HEX_KIND_MIXED) <<  8 |\n+\t\thexval(lenbuf_hex[2], HEX_KIND_MIXED) <<  4 |\n+\t\thexval(lenbuf_hex[3], HEX_KIND_MIXED);\n }\n \n static const char *find_packfile_uri_path(const char *buffer)\n"},{"id":"552162","messageId":"20260907195941.1024289-6-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260907195941.1024289-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 5/7] object-name: use hexval","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:38Z","receivedAt":"2026-09-07T20:00:01Z","isPatch":true,"body":"We've open-coded a different implementation of parsing hex values here\nwhen we already have a perfectly good one in hexval.  This\nimplementation will almost certainly be slower because it isn't\ntable-driven, unlike the other one, and since it's not constant time it\nhas no other advantages either.  To tidy things up and prepare for\nfuture work, switch to hexval in this case.\n\nBecause hexval returns an unsigned int, check to see if the value is\ninvalid by looking for any bits beyond a single unsigned character.  In\naddition, be sure to continue to force the hexadecimal value to\nlowercase.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n object-name.c | 15 ++++-----------\n 1 file changed, 4 insertions(+), 11 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 4eda8c8eac..8f2da51547 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -236,20 +236,13 @@ static int parse_oid_prefix(const char *name, int len,\n {\n \tfor (int i = 0; i < len; i++) {\n \t\tunsigned char c = name[i];\n-\t\tunsigned char val;\n-\t\tif (c >= '0' && c <= '9') {\n-\t\t\tval = c - '0';\n-\t\t} else if (c >= 'a' && c <= 'f') {\n-\t\t\tval = c - 'a' + 10;\n-\t\t} else if (c >= 'A' && c <='F') {\n-\t\t\tval = c - 'A' + 10;\n-\t\t\tc -= 'A' - 'a';\n-\t\t} else {\n+\t\tint val = hexval(c, HEX_KIND_OID);\n+\n+\t\tif (val & ~0xff)\n \t\t\treturn -1;\n-\t\t}\n \n \t\tif (hex_out)\n-\t\t\thex_out[i] = c;\n+\t\t\thex_out[i] = tolower(c);\n \t\tif (oid_out) {\n \t\t\tif (!(i & 1))\n \t\t\t\tval <<= 4;\n"},{"id":"552163","messageId":"20260907195941.1024289-7-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260907195941.1024289-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 6/7] t5324: adjust tests for corrupt commit-graph","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:39Z","receivedAt":"2026-09-07T20:00:01Z","isPatch":true,"body":"In a future commit, we'll no longer allow uppercase object IDs.  When\nthat happens, t5324 will fail because the error message from corrupting\ncommit-graph files will be different.  The current test looks for an\nerror message that occurs when the commit-graph file has a name that is\nparseable as a valid object ID but is not a valid commit-graph file\n(which, on case-sensitive systems, is of course checked case\nsensitively).\n\nHowever, in the future, we need the object ID to continue to be valid,\nand that means that the hex character we use must be lowercase.  In\nSHA-256, though, one of the characters we want to corrupt is\nlegitimately already an \"a\", so switch to using the character \"b\" in\nboth of these cases, which is the correct character for neither SHA-1\nnor SHA-256.  If we adopt another hash algorithm in the future, we may\nneed to adjust these values again.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n t/t5324-split-commit-graph.sh | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex bf7ba0e558..53cddfe437 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -349,7 +349,7 @@ test_expect_success 'verify after commit-graph-chain corruption (base)' '\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"invalid commit-graph chain\" err &&\n-\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"A\" &&\n+\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 30 \"b\" &&\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"unable to find all commit-graph files\" err\n@@ -364,7 +364,7 @@ test_expect_success 'verify after commit-graph-chain corruption (tip)' '\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"invalid commit-graph chain\" err &&\n-\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"A\" &&\n+\t\tcorrupt_file \"$graphdir/commit-graph-chain\" 70 \"b\" &&\n \t\ttest_must_fail git commit-graph verify 2>test_err &&\n \t\tgrep -v \"^+\" test_err >err &&\n \t\ttest_grep \"unable to find all commit-graph files\" err\n"},{"id":"552164","messageId":"20260907195941.1024289-8-sandals@crustytoothpaste.net","threadId":"66087","inReplyTo":"20260907195941.1024289-1-sandals@crustytoothpaste.net","subject":"[PATCH v2 7/7] hex: allow only lowercase object IDs in breaking changes mode","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:59:40Z","receivedAt":"2026-09-07T20:00:01Z","isPatch":true,"body":"Git has historically allowed either lowercase or uppercase hex for\nobject IDs, but it has always emitted only lowercase.  This has caused\npeople to expect only lowercase and not handle uppercase.\n\nAs an example, Git's own example hooks look for \"[0-9a-f]\" in several\nplaces, but there are many other Git-adjacent pieces of software,\nincluding Gitolite, which make the assumption that object IDs are always\nlowercase.  This is not to criticize the authors of these projects, but\nrather to point out how common this assumption is.  In fact, it's so\ncommon that we had only one test in our codebase that failed when we\nreject uppercase object IDs.\n\nMore critically, it leads people to make security-based assumptions that\nan object ID either does not contain uppercase characters or that an\nobject ID can be expressed uniquely in hex form, neither of which are\ncurrently true.  Git itself normally uses binary object IDs, which\navoids many of these problems, but most other projects deal primarily in\nhex object IDs, so they are more affected.\n\nIn preparation for Git 3.0, only allow lowercase hex object IDs in\nbreaking changes mode and document this as well.  Add a new test to\nverify we reject new uppercase object IDs.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n Documentation/BreakingChanges.adoc | 5 +++++\n hex-ll.h                           | 4 ++++\n t/t1503-rev-parse-verify.sh        | 5 +++++\n 3 files changed, 14 insertions(+)\n\ndiff --git a/Documentation/BreakingChanges.adoc b/Documentation/BreakingChanges.adoc\nindex 73bb939359..dbc46d14e3 100644\n--- a/Documentation/BreakingChanges.adoc\n+++ b/Documentation/BreakingChanges.adoc\n@@ -171,6 +171,11 @@ JGit, libgit2 and Gitoxide need to support it.\n   matches the default branch name used in new repositories by many of the\n   big Git forges.\n \n+* Git will accept hex object IDs only in lowercase. The fact that Git has\n+\thistorically allowed uppercase characters in hex object IDs has been the\n+\tsource of a variety of bugs and security problems in software using Git. We\n+\tdon't expect most users to notice any change.\n+\n * Git will require Rust as a mandatory part of the build process. While Git\n   already started to adopt Rust in Git 2.49, all parts written in Rust are\n   optional for the time being. This includes:\ndiff --git a/hex-ll.h b/hex-ll.h\nindex 9da76f17e8..2f9c8d7c25 100644\n--- a/hex-ll.h\n+++ b/hex-ll.h\n@@ -6,7 +6,11 @@ enum hexkind {\n \tHEX_KIND_LOWER = 1,\n };\n \n+#ifdef WITH_BREAKING_CHANGES\n+#define HEX_KIND_OID HEX_KIND_LOWER\n+#else\n #define HEX_KIND_OID HEX_KIND_MIXED\n+#endif\n \n extern const signed char hexval_table[256];\n extern const signed char hexval_lc_table[256];\ndiff --git a/t/t1503-rev-parse-verify.sh b/t/t1503-rev-parse-verify.sh\nindex 87638a4a2c..f07b45de5a 100755\n--- a/t/t1503-rev-parse-verify.sh\n+++ b/t/t1503-rev-parse-verify.sh\n@@ -60,6 +60,11 @@ test_expect_success 'works with one good rev' '\n \ttest \"$rev_head\" = \"$HASH4\"\n '\n \n+test_expect_success WITH_BREAKING_CHANGES 'rejects uppercase revs' '\n+\tUC_HASH=$(echo \"$HASH1\" | tr a-f A-F) &&\n+\ttest_must_fail git rev-parse --verify \"$UC_HASH\"\n+'\n+\n test_expect_success 'fails with any bad rev or many good revs' '\n \ttest_must_fail git rev-parse --verify 2>error &&\n \ttest_grep \"single revision\" error &&\n"}]}