{"thread":{"id":"64326","subject":"[PATCH 0/9] Xdiff cleanup part2","startedAt":"2025-10-15T21:18:24Z","lastAt":"2025-11-19T04:15:00Z","messageCount":118,"participants":["Ezekiel Newren via GitGitGadget","Junio C Hamano","Kristoffer Haugsbakk","Ezekiel Newren","Patrick Steinhardt","Phillip Wood","Chris Torek","Ramsay Jones","Ben Knoble","D. Ben Knoble"],"isPatch":true,"patchVersion":1,"patchTotal":9},"messages":[{"id":"528842","messageId":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":null,"subject":"[PATCH 0/9] Xdiff cleanup part2","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:12Z","receivedAt":"2025-10-15T21:18:24Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"Maintainer note: This patch series builds on top of en/xdiff-cleanup and\nam/xdiff-hash-tweak (both of which are now in master).\n\nThe primary goal of this patch series is to convert every field's type in\nxrecord_t and xdfile_t to be unambiguous, in preparation to make it more\nRust FFI friendly. Additionally the ha field in xrecord_t is split into\nline_hash and minimal_perfect hash.\n\nThe order of some of the fields has changed as called out by the commit\nmessages.\n\nBefore:\n\ntypedef struct s_xrecord {\n\tchar const *ptr;\n\tlong size;\n\tunsigned long ha;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tlong nrec;\n\tlong dstart, dend;\n\tbool *changed;\n\tlong *rindex;\n\tlong nreff;\n} xdfile_t;\n\n\nAfter part 2\n\ntypedef struct s_xrecord {\n\tuint8_t const *ptr;\n\tsize_t size;\n\tuint64_t line_hash;\n\tsize_t minimal_perfect_hash;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tsize_t nrec;\n\tbool *changed;\n\tsize_t *reference_index;\n\tsize_t nreff;\n\tssize_t dstart, dend;\n} xdfile_t;\n\n\nEzekiel Newren (9):\n  xdiff: use ssize_t for dstart/dend, make them last in xdfile_t\n  xdiff: make xrecord_t.ptr a uint8_t instead of char\n  xdiff: use size_t for xrecord_t.size\n  xdiff: use unambiguous types in xdl_hash_record()\n  xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n  xdiff: make xdfile_t.nrec a size_t instead of long\n  xdiff: make xdfile_t.nreff a size_t instead of long\n  xdiff: change rindex from long to size_t in xdfile_t\n  xdiff: rename rindex -> reference_index\n\n xdiff-interface.c  |  2 +-\n xdiff/xdiffi.c     | 29 +++++++++++------------\n xdiff/xemit.c      | 28 +++++++++++-----------\n xdiff/xhistogram.c |  4 ++--\n xdiff/xmerge.c     | 30 ++++++++++++------------\n xdiff/xpatience.c  | 14 +++++------\n xdiff/xprepare.c   | 58 +++++++++++++++++++++++-----------------------\n xdiff/xtypes.h     | 15 ++++++------\n xdiff/xutils.c     | 32 ++++++++++++-------------\n xdiff/xutils.h     |  6 ++---\n 10 files changed, 109 insertions(+), 109 deletions(-)\n\n\nbase-commit: 143f58ef7535f8f8a80d810768a18bdf3807de26\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2070%2Fezekielnewren%2Fxdiff_cleanup_part2-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2070/ezekielnewren/xdiff_cleanup_part2-v1\nPull-Request: https://github.com/git/git/pull/2070\n-- \ngitgitgadget\n"},{"id":"528843","messageId":"1fa9a7d7d1c309f2f651da351ba7bc0b36272d91.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 1/9] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:13Z","receivedAt":"2025-10-15T21:18:25Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nssize_t is appropriate for dstart and dend because they both describe\npositive or negative offsets relative to a pointer.\n\nA future patch will move these fields to a different struct. Moving\nthem to the end of xdfile_t now, means the field order of xdfile_t will\nbe disturbed less.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex f145abba3e..3514bb1684 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,10 +47,10 @@ typedef struct s_xrecord {\n typedef struct s_xdfile {\n \txrecord_t *recs;\n \tlong nrec;\n-\tlong dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n+\tssize_t dstart, dend;\n } xdfile_t;\n \n typedef struct s_xdfenv {\n-- \ngitgitgadget\n\n"},{"id":"528844","messageId":"7b9e8961d42e0f367ba0782e7d932607aa7e0b0a.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:14Z","receivedAt":"2025-10-15T21:18:26Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nRust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\nreferring to bytes in memory, rather than unicode code points, use\nuint8_t instead of char.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     |  6 +++---\n xdiff/xmerge.c    | 14 +++++++-------\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  |  8 ++++----\n xdiff/xtypes.h    |  2 +-\n xdiff/xutils.c    |  4 ++--\n 7 files changed, 22 insertions(+), 22 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 6f3998ee54..411a8aa69f 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n \tint ret = 0;\n \n \tfor (i = 0; i < rec->size; i++) {\n-\t\tchar c = rec->ptr[i];\n+\t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n \t\t\treturn ret;\n@@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\n@@ -1008,7 +1008,7 @@ static int record_matches_regex(xrecord_t *rec, xpparam_t const *xpp) {\n \tsize_t i;\n \n \tfor (i = 0; i < xpp->ignore_regex_nr; i++)\n-\t\tif (!regexec_buf(xpp->ignore_regex[i], rec->ptr, rec->size, 1,\n+\t\tif (!regexec_buf(xpp->ignore_regex[i], (const char *)rec->ptr, rec->size, 1,\n \t\t\t\t &regmatch, 0))\n \t\t\treturn 1;\n \ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex b2f1f30cd3..ead930088a 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec(rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff(rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func(rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex fd600cbb5d..75cb3e76a2 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch(rec1[i].ptr, rec1[i].size,\n-\t\t\trec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch(rec1->ptr, rec1->size,\n-\t\t\t    rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n+\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n }\n \n /*\n@@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n \t\t * we have a very simple mmfile structure.\n \t\t */\n \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n-\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n+\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n \t\tt2.ptr = (char *)xe2->xdf2.recs[m->i2].ptr;\n-\t\tt2.size = xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n+\t\tt2.size = (char *)xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n \t\t\t+ xe2->xdf2.recs[m->i2 + m->chg2 - 1].size - t2.ptr;\n \t\tif (xdl_do_diff(&t1, &t2, xpp, &xe) < 0)\n \t\t\treturn -1;\n@@ -440,7 +440,7 @@ static int line_contains_alnum(const char *ptr, long size)\n static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n-\t\tif (line_contains_alnum(xe->xdf2.recs[i].ptr,\n+\t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n \t\t\t\txe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 669b653580..bb61354f22 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -121,7 +121,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t\treturn;\n \tmap->entries[index].line1 = line;\n \tmap->entries[index].hash = record->ha;\n-\tmap->entries[index].anchor = is_anchor(xpp, map->env->xdf1.recs[line - 1].ptr);\n+\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n \tif (map->last) {\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 192334f1b7..4cb18b2b88 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch(rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\trec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -156,8 +156,8 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = prev;\n-\t\t\tcrec->size = (long) (cur - prev);\n+\t\t\tcrec->ptr = (uint8_t const *)prev;\n+\t\t\tcrec->size =(long) ( cur - prev);\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 3514bb1684..57983627f5 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -39,7 +39,7 @@ typedef struct s_chastore {\n } chastore_t;\n \n typedef struct s_xrecord {\n-\tchar const *ptr;\n+\tuint8_t const *ptr;\n \tlong size;\n \tunsigned long ha;\n } xrecord_t;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 447e66c719..7be063bfb6 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -465,10 +465,10 @@ int xdl_fall_back_diff(xdfenv_t *diff_env, xpparam_t const *xpp,\n \txdfenv_t env;\n \n \tsubfile1.ptr = (char *)diff_env->xdf1.recs[line1 - 1].ptr;\n-\tsubfile1.size = diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n+\tsubfile1.size = (char *)diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n \t\tdiff_env->xdf1.recs[line1 + count1 - 2].size - subfile1.ptr;\n \tsubfile2.ptr = (char *)diff_env->xdf2.recs[line2 - 1].ptr;\n-\tsubfile2.size = diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n+\tsubfile2.size = (char *)diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n \t\tdiff_env->xdf2.recs[line2 + count2 - 2].size - subfile2.ptr;\n \tif (xdl_do_diff(&subfile1, &subfile2, xpp, &env) < 0)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"528845","messageId":"ae15ed712123c151b7856b56a2da9393fa5943fa.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 3/9] xdiff: use size_t for xrecord_t.size","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:15Z","receivedAt":"2025-10-15T21:18:27Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is the appropriate type because size is describing the number of\nelements, bytes in this case, in memory.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  7 +++----\n xdiff/xemit.c    |  8 ++++----\n xdiff/xmerge.c   | 16 ++++++++--------\n xdiff/xprepare.c |  6 +++---\n xdiff/xtypes.h   |  2 +-\n 5 files changed, 19 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 411a8aa69f..edd05466df 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -403,10 +403,9 @@ static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n  */\n static int get_indent(xrecord_t *rec)\n {\n-\tlong i;\n \tint ret = 0;\n \n-\tfor (i = 0; i < rec->size; i++) {\n+\tfor (size_t i = 0; i < rec->size; i++) {\n \t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n@@ -993,11 +992,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex ead930088a..2f8007753c 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, (long)rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, (long)rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, (long)rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\n@@ -151,7 +151,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n static int is_empty_rec(xdfile_t *xdf, long ri)\n {\n \txrecord_t *rec = &xdf->recs[ri];\n-\tlong i = 0;\n+\tsize_t i = 0;\n \n \tfor (; i < rec->size && XDL_ISSPACE(rec->ptr[i]); i++);\n \ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 75cb3e76a2..0dd4558a32 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n-\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, (long)rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, (long)rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -119,11 +119,11 @@ static int xdl_recs_copy_0(int use_orig, xdfenv_t *xe, int i, int count, int nee\n \tif (count < 1)\n \t\treturn 0;\n \n-\tfor (i = 0; i < count; size += recs[i++].size)\n+\tfor (i = 0; i < count; size += (int)recs[i++].size)\n \t\tif (dest)\n \t\t\tmemcpy(dest + size, recs[i].ptr, recs[i].size);\n \tif (add_nl) {\n-\t\ti = recs[count - 1].size;\n+\t\ti = (int)recs[count - 1].size;\n \t\tif (i == 0 || recs[count - 1].ptr[i - 1] != '\\n') {\n \t\t\tif (needs_cr) {\n \t\t\t\tif (dest)\n@@ -156,7 +156,7 @@ static int xdl_orig_copy(xdfenv_t *xe, int i, int count, int needs_cr, int add_n\n  */\n static int is_eol_crlf(xdfile_t *file, int i)\n {\n-\tlong size;\n+\tsize_t size;\n \n \tif (i < file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n-\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, (long)rec1->size,\n+\t\t\t    (const char *)rec2->ptr, (long)rec2->size, flags);\n }\n \n /*\n@@ -441,7 +441,7 @@ static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n \t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n-\t\t\t\txe->xdf2.recs[i].size))\n+\t\t\t\t(long)xe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 4cb18b2b88..b3219aed3e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -157,7 +157,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = (uint8_t const *)prev;\n-\t\t\tcrec->size =(long) ( cur - prev);\n+\t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 57983627f5..00d2d8c8cd 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -40,7 +40,7 @@ typedef struct s_chastore {\n \n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n-\tlong size;\n+\tsize_t size;\n \tunsigned long ha;\n } xrecord_t;\n \n-- \ngitgitgadget\n\n"},{"id":"528846","messageId":"7fcd83c99076404960302b64a4f0c8fa1c13feba.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 4/9] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:16Z","receivedAt":"2025-10-15T21:18:28Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff-interface.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xutils.c    | 28 ++++++++++++++--------------\n xdiff/xutils.h    |  6 +++---\n 4 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff-interface.c b/xdiff-interface.c\nindex 4971f722b3..1a35556380 100644\n--- a/xdiff-interface.c\n+++ b/xdiff-interface.c\n@@ -300,7 +300,7 @@ void xdiff_clear_find_func(xdemitconf_t *xecfg)\n \n unsigned long xdiff_hash_string(const char *s, size_t len, long flags)\n {\n-\treturn xdl_hash_record(&s, s + len, flags);\n+\treturn xdl_hash_record((uint8_t const**)&s, (uint8_t const*)s + len, flags);\n }\n \n int xdiff_compare_lines(const char *l1, long s1,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex b3219aed3e..85e56021da 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -137,8 +137,8 @@ static void xdl_free_ctx(xdfile_t *xdf)\n static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_t const *xpp,\n \t\t\t   xdlclassifier_t *cf, xdfile_t *xdf) {\n \tlong bsize;\n-\tunsigned long hav;\n-\tchar const *blk, *cur, *top, *prev;\n+\tuint64_t hav;\n+\tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n \txdf->rindex = NULL;\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = (uint8_t const *)prev;\n+\t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 7be063bfb6..77ee1ad9c8 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -249,11 +249,11 @@ int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags)\n \treturn 1;\n }\n \n-unsigned long xdl_hash_record_with_whitespace(char const **data,\n-\t\tchar const *top, long flags) {\n-\tunsigned long ha = 5381;\n-\tchar const *ptr = *data;\n-\tint cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data,\n+\t\tuint8_t const *top, uint64_t flags) {\n+\tuint64_t ha = 5381;\n+\tuint8_t const *ptr = *data;\n+\tbool cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n \n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tif (cr_at_eol_only) {\n@@ -263,8 +263,8 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\t\tcontinue;\n \t\t}\n \t\telse if (XDL_ISSPACE(*ptr)) {\n-\t\t\tconst char *ptr2 = ptr;\n-\t\t\tint at_eol;\n+\t\t\tconst uint8_t *ptr2 = ptr;\n+\t\t\tbool at_eol;\n \t\t\twhile (ptr + 1 < top && XDL_ISSPACE(ptr[1])\n \t\t\t\t\t&& ptr[1] != '\\n')\n \t\t\t\tptr++;\n@@ -274,20 +274,20 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_CHANGE\n \t\t\t\t && !at_eol) {\n \t\t\t\tha += (ha << 5);\n-\t\t\t\tha ^= (unsigned long) ' ';\n+\t\t\t\tha ^= (uint64_t) ' ';\n \t\t\t}\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_AT_EOL\n \t\t\t\t && !at_eol) {\n \t\t\t\twhile (ptr2 != ptr + 1) {\n \t\t\t\t\tha += (ha << 5);\n-\t\t\t\t\tha ^= (unsigned long) *ptr2;\n+\t\t\t\t\tha ^= (uint64_t) *ptr2;\n \t\t\t\t\tptr2++;\n \t\t\t\t}\n \t\t\t}\n \t\t\tcontinue;\n \t\t}\n \t\tha += (ha << 5);\n-\t\tha ^= (unsigned long) *ptr;\n+\t\tha ^= (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n \n@@ -304,9 +304,9 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n #define REASSOC_FENCE(x, y)\n #endif\n \n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n-\tunsigned long ha = 5381, c0, c1;\n-\tchar const *ptr = *data;\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top) {\n+\tuint64_t ha = 5381, c0, c1;\n+\tuint8_t const *ptr = *data;\n #if 0\n \t/*\n \t * The baseline form of the optimized loop below. This is the djb2\n@@ -314,7 +314,7 @@ unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n \t */\n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tha += (ha << 5);\n-\t\tha += (unsigned long) *ptr;\n+\t\tha += (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n #else\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 13f6831047..615b4a9d35 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -34,9 +34,9 @@ void *xdl_cha_alloc(chastore_t *cha);\n long xdl_guess_lines(mmfile_t *mf, long sample);\n int xdl_blankline(const char *line, long size, long flags);\n int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags);\n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top);\n-unsigned long xdl_hash_record_with_whitespace(char const **data, char const *top, long flags);\n-static inline unsigned long xdl_hash_record(char const **data, char const *top, long flags)\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top);\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data, uint8_t const *top, uint64_t flags);\n+static inline uint64_t xdl_hash_record(uint8_t const **data, uint8_t const *top, uint64_t flags)\n {\n \tif (flags & XDF_WHITESPACE_FLAGS)\n \t\treturn xdl_hash_record_with_whitespace(data, top, flags);\n-- \ngitgitgadget\n\n"},{"id":"528847","messageId":"a3e706ecdae51434fd5ee112c13f8cf374faf6ed.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 5/9] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:17Z","receivedAt":"2025-10-15T21:18:30Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe ha field is serving two different purposes, which makes the code\nharder to read. At first glance it looks like many places assume\nthere could never be hash collisions between lines of the two input\nfiles. In reality, line_hash is used together with xdl_recmatch() to\nensure correct comparisons of lines, even when collisions occur.\n\nTo make this clearer, the old ha field has been split:\n  * line_hash: The straightforward hash of a line, requiring no\n    additional context.\n  * minimal_perfect_hash: Not a new concept, but now a separate\n    field. It comes from the classifier's general-purpose hash table,\n    which assigns each line a unique and minimal hash across the two\n    files.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c     |  6 +++---\n xdiff/xhistogram.c |  4 ++--\n xdiff/xpatience.c  | 10 +++++-----\n xdiff/xprepare.c   | 16 ++++++++--------\n xdiff/xtypes.h     |  3 ++-\n 5 files changed, 20 insertions(+), 19 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex edd05466df..436c34697d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -22,9 +22,9 @@\n \n #include \"xinclude.h\"\n \n-static unsigned long get_hash(xdfile_t *xdf, long index)\n+static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].ha;\n+\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -385,7 +385,7 @@ static xdchange_t *xdl_add_change(xdchange_t *xscr, long i1, long i2, long chg1,\n \n static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n {\n-\treturn (rec1->ha == rec2->ha);\n+\treturn rec1->minimal_perfect_hash == rec2->minimal_perfect_hash;\n }\n \n /*\ndiff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\nindex 6dc450b1fe..5ae1282c27 100644\n--- a/xdiff/xhistogram.c\n+++ b/xdiff/xhistogram.c\n@@ -90,7 +90,7 @@ struct region {\n \n static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n {\n-\treturn r1->ha == r2->ha;\n+\treturn r1->minimal_perfect_hash == r2->minimal_perfect_hash;\n \n }\n \n@@ -98,7 +98,7 @@ static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n \t(cmp_recs(REC(i->env, s1, l1), REC(i->env, s2, l2)))\n \n #define TABLE_HASH(index, side, line) \\\n-\tXDL_HASHLONG((REC(index->env, side, line))->ha, index->table_bits)\n+\tXDL_HASHLONG((REC(index->env, side, line))->minimal_perfect_hash, index->table_bits)\n \n static int scanA(struct histindex *index, int line1, int count1)\n {\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex bb61354f22..cc53266f3b 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -48,7 +48,7 @@\n struct hashmap {\n \tint nr, alloc;\n \tstruct entry {\n-\t\tunsigned long hash;\n+\t\tsize_t minimal_perfect_hash;\n \t\t/*\n \t\t * 0 = unused entry, 1 = first line, 2 = second, etc.\n \t\t * line2 is NON_UNIQUE if the line is not unique\n@@ -101,10 +101,10 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t * So we multiply ha by 2 in the hope that the hashing was\n \t * \"unique enough\".\n \t */\n-\tint index = (int)((record->ha << 1) % map->alloc);\n+\tint index = (int)((record->minimal_perfect_hash << 1) % map->alloc);\n \n \twhile (map->entries[index].line1) {\n-\t\tif (map->entries[index].hash != record->ha) {\n+\t\tif (map->entries[index].minimal_perfect_hash != record->minimal_perfect_hash) {\n \t\t\tif (++index >= map->alloc)\n \t\t\t\tindex = 0;\n \t\t\tcontinue;\n@@ -120,7 +120,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \tif (pass == 2)\n \t\treturn;\n \tmap->entries[index].line1 = line;\n-\tmap->entries[index].hash = record->ha;\n+\tmap->entries[index].minimal_perfect_hash = record->minimal_perfect_hash;\n \tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n@@ -248,7 +248,7 @@ static int match(struct hashmap *map, int line1, int line2)\n {\n \txrecord_t *record1 = &map->env->xdf1.recs[line1 - 1];\n \txrecord_t *record2 = &map->env->xdf2.recs[line2 - 1];\n-\treturn record1->ha == record2->ha;\n+\treturn record1->minimal_perfect_hash == record2->minimal_perfect_hash;\n }\n \n static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 85e56021da..16236bd045 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -96,9 +96,9 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \tlong hi;\n \txdlclass_t *rcrec;\n \n-\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n+\thi = (long) XDL_HASHLONG(rec->line_hash, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n-\t\tif (rcrec->rec.ha == rec->ha &&\n+\t\tif (rcrec->rec.line_hash == rec->line_hash &&\n \t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n \t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n@@ -120,7 +120,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n \t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n \n-\trec->ha = (unsigned long) rcrec->idx;\n+\trec->minimal_perfect_hash = (size_t)rcrec->idx;\n \n \treturn 0;\n }\n@@ -158,7 +158,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n-\t\t\tcrec->ha = hav;\n+\t\t\tcrec->line_hash = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\n \t\t}\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -298,7 +298,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -350,7 +350,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs2 = xdf2->recs;\n \tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dstart = xdf2->dstart = i;\n@@ -358,7 +358,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs1 = xdf1->recs + xdf1->nrec - 1;\n \trecs2 = xdf2->recs + xdf2->nrec - 1;\n \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dend = xdf1->nrec - i - 1;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 00d2d8c8cd..a57a8c2c12 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -41,7 +41,8 @@ typedef struct s_chastore {\n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n \tsize_t size;\n-\tunsigned long ha;\n+\tuint64_t line_hash;\n+\tsize_t minimal_perfect_hash;\n } xrecord_t;\n \n typedef struct s_xdfile {\n-- \ngitgitgadget\n\n"},{"id":"528848","messageId":"5767ba4ee87e7a01283b12745a71c14f1b0088a7.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 6/9] xdiff: make xdfile_t.nrec a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:18Z","receivedAt":"2025-10-15T21:18:32Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nrec describes the number of elements in memory\nfor recs, and the number of elements in memory for 'changed' + 2.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     | 20 ++++++++++----------\n xdiff/xmerge.c    |  8 ++++----\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  | 12 ++++++------\n xdiff/xtypes.h    |  2 +-\n 6 files changed, 26 insertions(+), 26 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 436c34697d..759193fe5d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -483,7 +483,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n {\n \tlong i;\n \n-\tif (split >= xdf->nrec) {\n+\tif (split >= (long)xdf->nrec) {\n \t\tm->end_of_file = 1;\n \t\tm->indent = -1;\n \t} else {\n@@ -506,7 +506,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n \n \tm->post_blank = 0;\n \tm->post_indent = -1;\n-\tfor (i = split + 1; i < xdf->nrec; i++) {\n+\tfor (i = split + 1; i < (long)xdf->nrec; i++) {\n \t\tm->post_indent = get_indent(&xdf->recs[i]);\n \t\tif (m->post_indent != -1)\n \t\t\tbreak;\n@@ -717,7 +717,7 @@ static void group_init(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static inline int group_next(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end == xdf->nrec)\n+\tif (g->end == (long)xdf->nrec)\n \t\treturn -1;\n \n \tg->start = g->end + 1;\n@@ -750,7 +750,7 @@ static inline int group_previous(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static int group_slide_down(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end < xdf->nrec &&\n+\tif (g->end < (long)xdf->nrec &&\n \t    recs_match(&xdf->recs[g->start], &xdf->recs[g->end])) {\n \t\txdf->changed[g->start++] = false;\n \t\txdf->changed[g->end++] = true;\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex 2f8007753c..04f7e9193b 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -137,7 +137,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n \tbuf = func_line ? func_line->buf : dummy;\n \tsize = func_line ? sizeof(func_line->buf) : sizeof(dummy);\n \n-\tfor (l = start; l != limit && 0 <= l && l < xe->xdf1.nrec; l += step) {\n+\tfor (l = start; l != limit && 0 <= l && l < (long)xe->xdf1.nrec; l += step) {\n \t\tlong len = match_func_rec(&xe->xdf1, xecfg, l, buf, size);\n \t\tif (len >= 0) {\n \t\t\tif (func_line)\n@@ -179,14 +179,14 @@ pre_context_calculation:\n \t\t\tlong fs1, i1 = xch->i1;\n \n \t\t\t/* Appended chunk? */\n-\t\t\tif (i1 >= xe->xdf1.nrec) {\n+\t\t\tif (i1 >= (long)xe->xdf1.nrec) {\n \t\t\t\tlong i2 = xch->i2;\n \n \t\t\t\t/*\n \t\t\t\t * We don't need additional context if\n \t\t\t\t * a whole function was added.\n \t\t\t\t */\n-\t\t\t\twhile (i2 < xe->xdf2.nrec) {\n+\t\t\t\twhile (i2 < (long)xe->xdf2.nrec) {\n \t\t\t\t\tif (is_func_rec(&xe->xdf2, xecfg, i2))\n \t\t\t\t\t\tgoto post_context_calculation;\n \t\t\t\t\ti2++;\n@@ -196,7 +196,7 @@ pre_context_calculation:\n \t\t\t\t * Otherwise get more context from the\n \t\t\t\t * pre-image.\n \t\t\t\t */\n-\t\t\t\ti1 = xe->xdf1.nrec - 1;\n+\t\t\t\ti1 = (long)xe->xdf1.nrec - 1;\n \t\t\t}\n \n \t\t\tfs1 = get_func_line(xe, xecfg, NULL, i1, -1);\n@@ -228,8 +228,8 @@ pre_context_calculation:\n \n  post_context_calculation:\n \t\tlctx = xecfg->ctxlen;\n-\t\tlctx = XDL_MIN(lctx, xe->xdf1.nrec - (xche->i1 + xche->chg1));\n-\t\tlctx = XDL_MIN(lctx, xe->xdf2.nrec - (xche->i2 + xche->chg2));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf1.nrec - (xche->i1 + xche->chg1));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf2.nrec - (xche->i2 + xche->chg2));\n \n \t\te1 = xche->i1 + xche->chg1 + lctx;\n \t\te2 = xche->i2 + xche->chg2 + lctx;\n@@ -237,13 +237,13 @@ pre_context_calculation:\n \t\tif (xecfg->flags & XDL_EMIT_FUNCCONTEXT) {\n \t\t\tlong fe1 = get_func_line(xe, xecfg, NULL,\n \t\t\t\t\t\t xche->i1 + xche->chg1,\n-\t\t\t\t\t\t xe->xdf1.nrec);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec);\n \t\t\twhile (fe1 > 0 && is_empty_rec(&xe->xdf1, fe1 - 1))\n \t\t\t\tfe1--;\n \t\t\tif (fe1 < 0)\n-\t\t\t\tfe1 = xe->xdf1.nrec;\n+\t\t\t\tfe1 = (long)xe->xdf1.nrec;\n \t\t\tif (fe1 > e1) {\n-\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), xe->xdf2.nrec);\n+\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), (long)xe->xdf2.nrec);\n \t\t\t\te1 = fe1;\n \t\t\t}\n \n@@ -254,7 +254,7 @@ pre_context_calculation:\n \t\t\t */\n \t\t\tif (xche->next) {\n \t\t\t\tlong l = XDL_MIN(xche->next->i1,\n-\t\t\t\t\t\t xe->xdf1.nrec - 1);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec - 1);\n \t\t\t\tif (l - xecfg->ctxlen <= e1 ||\n \t\t\t\t    get_func_line(xe, xecfg, NULL, l, e1) < 0) {\n \t\t\t\t\txche = xche->next;\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 0dd4558a32..29dad98c49 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -158,7 +158,7 @@ static int is_eol_crlf(xdfile_t *file, int i)\n {\n \tsize_t size;\n \n-\tif (i < file->nrec - 1)\n+\tif (i < (long)file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n \t\treturn (size = file->recs[i].size) > 1 &&\n \t\t\tfile->recs[i].ptr[size - 2] == '\\r';\n@@ -317,7 +317,7 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \t\t\tcontinue;\n \t\ti = m->i1 + m->chg1;\n \t}\n-\tsize += xdl_recs_copy(xe1, i, xe1->xdf2.nrec - i, 0, 0,\n+\tsize += xdl_recs_copy(xe1, i, (int)xe1->xdf2.nrec - i, 0, 0,\n \t\t\t      dest ? dest + size : NULL);\n \treturn size;\n }\n@@ -622,7 +622,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\t\tchanges = c;\n \t\ti0 = xscr1->i1;\n \t\ti1 = xscr1->i2;\n-\t\ti2 = xscr1->i1 + xe2->xdf2.nrec - xe2->xdf1.nrec;\n+\t\ti2 = xscr1->i1 + (long)xe2->xdf2.nrec - (long)xe2->xdf1.nrec;\n \t\tchg0 = xscr1->chg1;\n \t\tchg1 = xscr1->chg2;\n \t\tchg2 = xscr1->chg1;\n@@ -637,7 +637,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\tif (!changes)\n \t\t\tchanges = c;\n \t\ti0 = xscr2->i1;\n-\t\ti1 = xscr2->i1 + xe1->xdf2.nrec - xe1->xdf1.nrec;\n+\t\ti1 = xscr2->i1 + (long)xe1->xdf2.nrec - (long)xe1->xdf1.nrec;\n \t\ti2 = xscr2->i2;\n \t\tchg0 = xscr2->chg1;\n \t\tchg1 = xscr2->chg1;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex cc53266f3b..a0b31eb5d8 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -370,5 +370,5 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n \n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n-\treturn patience_diff(xpp, env, 1, env->xdf1.nrec, 1, env->xdf2.nrec);\n+\treturn patience_diff(xpp, env, 1, (int)env->xdf1.nrec, 1, (int)env->xdf2.nrec);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 16236bd045..4ee9fb60cd 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -153,7 +153,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\tfor (top = blk + bsize; cur < top; ) {\n \t\t\tprev = cur;\n \t\t\thav = xdl_hash_record(&cur, top, xpp->flags);\n-\t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n+\t\t\tif (XDL_ALLOC_GROW(xdf->recs, (long)xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n@@ -287,7 +287,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -295,7 +295,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -348,7 +348,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \n \trecs1 = xdf1->recs;\n \trecs2 = xdf2->recs;\n-\tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n+\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n@@ -361,8 +361,8 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n-\txdf1->dend = xdf1->nrec - i - 1;\n-\txdf2->dend = xdf2->nrec - i - 1;\n+\txdf1->dend = (long)xdf1->nrec - i - 1;\n+\txdf2->dend = (long)xdf2->nrec - i - 1;\n \n \treturn 0;\n }\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex a57a8c2c12..179ae2ae89 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,7 +47,7 @@ typedef struct s_xrecord {\n \n typedef struct s_xdfile {\n \txrecord_t *recs;\n-\tlong nrec;\n+\tsize_t nrec;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n-- \ngitgitgadget\n\n"},{"id":"528849","messageId":"4caa6a466977483c42f4e37bd0067dc1ca3b28aa.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 7/9] xdiff: make xdfile_t.nreff a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:19Z","receivedAt":"2025-10-15T21:18:33Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nreff describes the number of elements in memory\nfor rindex.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 +++++++-------\n xdiff/xtypes.h   |  2 +-\n 2 files changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 4ee9fb60cd..c690bafeb1 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -264,7 +264,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, nreff, mlim;\n+\tlong i, nm, mlim;\n \txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n@@ -307,29 +307,29 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * Use temporary arrays to decide if changed[i] should remain\n \t * false, or become true.\n \t */\n-\tfor (nreff = 0, i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n+\txdf1->nreff = 0;\n+\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[nreff++] = i;\n+\t\t\txdf1->rindex[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf1->nreff = nreff;\n \n-\tfor (nreff = 0, i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n+\txdf2->nreff = 0;\n+\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[nreff++] = i;\n+\t\t\txdf2->rindex[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf2->nreff = nreff;\n \n cleanup:\n \txdl_free(action1);\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 179ae2ae89..e9473bfd45 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -50,7 +50,7 @@ typedef struct s_xdfile {\n \tsize_t nrec;\n \tbool *changed;\n \tlong *rindex;\n-\tlong nreff;\n+\tsize_t nreff;\n \tssize_t dstart, dend;\n } xdfile_t;\n \n-- \ngitgitgadget\n\n"},{"id":"528850","messageId":"6dca5e6222e1d02092d4ba8296b757b123b85afa.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 8/9] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:20Z","receivedAt":"2025-10-15T21:18:34Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nrindex describes a index offset which means it's an index into memory\nwhich should use size_t. dstart and dend will be deleted in a future\npatch series. Move them to the end to help avoid refactor conflicts.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex e9473bfd45..8016222de9 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -49,7 +49,7 @@ typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n \tbool *changed;\n-\tlong *rindex;\n+\tsize_t *rindex;\n \tsize_t nreff;\n \tssize_t dstart, dend;\n } xdfile_t;\n-- \ngitgitgadget\n\n"},{"id":"528851","messageId":"518e5f5557e9bb30727d0d26433d64117269d159.1760563101.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH 9/9] xdiff: rename rindex -> reference_index","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-15T21:18:21Z","receivedAt":"2025-10-15T21:18:35Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe classic diff adds only the lines that it's going to consider,\nduring the diff, to an array. A mapping between the compacted\narray, and the lines of the file that they reference, are\nfacilitated by this array.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  6 +++---\n xdiff/xprepare.c | 10 +++++-----\n xdiff/xtypes.h   |  2 +-\n 3 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 759193fe5d..8eb664be3e 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -24,7 +24,7 @@\n \n static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n+\treturn xdf->recs[xdf->reference_index[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -278,10 +278,10 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n \t */\n \tif (off1 == lim1) {\n \t\tfor (; off2 < lim2; off2++)\n-\t\t\txdf2->changed[xdf2->rindex[off2]] = true;\n+\t\t\txdf2->changed[xdf2->reference_index[off2]] = true;\n \t} else if (off2 == lim2) {\n \t\tfor (; off1 < lim1; off1++)\n-\t\t\txdf1->changed[xdf1->rindex[off1]] = true;\n+\t\t\txdf1->changed[xdf1->reference_index[off1]] = true;\n \t} else {\n \t\txdpsplit_t spl;\n \t\tspl.i1 = spl.i2 = 0;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex c690bafeb1..1dd420a2ff 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -128,7 +128,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n static void xdl_free_ctx(xdfile_t *xdf)\n {\n-\txdl_free(xdf->rindex);\n+\txdl_free(xdf->reference_index);\n \txdl_free(xdf->changed - 1);\n \txdl_free(xdf->recs);\n }\n@@ -141,7 +141,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n-\txdf->rindex = NULL;\n+\txdf->reference_index = NULL;\n \txdf->changed = NULL;\n \txdf->recs = NULL;\n \n@@ -169,7 +169,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF)) {\n-\t\tif (!XDL_ALLOC_ARRAY(xdf->rindex, xdf->nrec + 1))\n+\t\tif (!XDL_ALLOC_ARRAY(xdf->reference_index, xdf->nrec + 1))\n \t\t\tgoto abort;\n \t}\n \n@@ -312,7 +312,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[xdf1->nreff++] = i;\n+\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n@@ -324,7 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[xdf2->nreff++] = i;\n+\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 8016222de9..373ccefa28 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -49,7 +49,7 @@ typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n \tbool *changed;\n-\tsize_t *rindex;\n+\tsize_t *reference_index;\n \tsize_t nreff;\n \tssize_t dstart, dend;\n } xdfile_t;\n-- \ngitgitgadget\n"},{"id":"528860","messageId":"xmqqa51rua5u.fsf@gitster.g","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] Xdiff cleanup part2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-15T21:28:45Z","receivedAt":"2025-10-15T21:28:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> The primary goal of this patch series is to convert every field's type in\n> xrecord_t and xdfile_t to be unambiguous, in preparation to make it more\n> Rust FFI friendly. Additionally the ha field in xrecord_t is split into\n> line_hash and minimal_perfect hash.\n>\n> The order of some of the fields has changed as called out by the commit\n> messages.\n>\n> Before:\n>\n> typedef struct s_xrecord {\n> \tchar const *ptr;\n> \tlong size;\n> \tunsigned long ha;\n> } xrecord_t;\n>\n> typedef struct s_xdfile {\n> \txrecord_t *recs;\n> \tlong nrec;\n> \tlong dstart, dend;\n> \tbool *changed;\n> \tlong *rindex;\n> \tlong nreff;\n> } xdfile_t;\n>\n>\n> After part 2\n>\n> typedef struct s_xrecord {\n> \tuint8_t const *ptr;\n> \tsize_t size;\n> \tuint64_t line_hash;\n> \tsize_t minimal_perfect_hash;\n> } xrecord_t;\n>\n> typedef struct s_xdfile {\n> \txrecord_t *recs;\n> \tsize_t nrec;\n> \tbool *changed;\n> \tsize_t *reference_index;\n> \tsize_t nreff;\n> \tssize_t dstart, dend;\n> } xdfile_t;\n\nExcellent summary.\n\n>\n>\n> Ezekiel Newren (9):\n>   xdiff: use ssize_t for dstart/dend, make them last in xdfile_t\n>   xdiff: make xrecord_t.ptr a uint8_t instead of char\n>   xdiff: use size_t for xrecord_t.size\n>   xdiff: use unambiguous types in xdl_hash_record()\n>   xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n>   xdiff: make xdfile_t.nrec a size_t instead of long\n>   xdiff: make xdfile_t.nreff a size_t instead of long\n>   xdiff: change rindex from long to size_t in xdfile_t\n>   xdiff: rename rindex -> reference_index\n>\n>  xdiff-interface.c  |  2 +-\n>  xdiff/xdiffi.c     | 29 +++++++++++------------\n>  xdiff/xemit.c      | 28 +++++++++++-----------\n>  xdiff/xhistogram.c |  4 ++--\n>  xdiff/xmerge.c     | 30 ++++++++++++------------\n>  xdiff/xpatience.c  | 14 +++++------\n>  xdiff/xprepare.c   | 58 +++++++++++++++++++++++-----------------------\n>  xdiff/xtypes.h     | 15 ++++++------\n>  xdiff/xutils.c     | 32 ++++++++++++-------------\n>  xdiff/xutils.h     |  6 ++---\n>  10 files changed, 109 insertions(+), 109 deletions(-)\n>\n>\n> base-commit: 143f58ef7535f8f8a80d810768a18bdf3807de26\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2070%2Fezekielnewren%2Fxdiff_cleanup_part2-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2070/ezekielnewren/xdiff_cleanup_part2-v1\n> Pull-Request: https://github.com/git/git/pull/2070\n"},{"id":"529028","messageId":"5af53ed1-2f84-46e4-9da3-b44871c3cbce@app.fastmail.com","threadId":"64326","inReplyTo":"7b9e8961d42e0f367ba0782e7d932607aa7e0b0a.1760563101.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-10-16T21:51:51Z","receivedAt":"2025-10-16T21:52:12Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"On Wed, Oct 15, 2025, at 23:18, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> Rust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\n> referring to bytes in memory, rather than unicode code points, use\n\ns/unicode/Unicode/\n\n> uint8_t instead of char.\n>\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>[snip]\n"},{"id":"529213","messageId":"CAH=ZcbAjX=V_VvJsRzvQEA+CMM7dWQx6E5=d4FL5CD3s+ozjBg@mail.gmail.com","threadId":"64326","inReplyTo":"a3e706ecdae51434fd5ee112c13f8cf374faf6ed.1760563101.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 5/9] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-20T23:29:25Z","receivedAt":"2025-10-20T23:29:39Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Wed, Oct 15, 2025 at 3:18 PM Ezekiel Newren via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> The ha field is serving two different purposes, which makes the code\n> harder to read. At first glance it looks like many places assume\n> there could never be hash collisions between lines of the two input\n> files. In reality, line_hash is used together with xdl_recmatch() to\n> ensure correct comparisons of lines, even when collisions occur.\n>\n> To make this clearer, the old ha field has been split:\n>   * line_hash: The straightforward hash of a line, requiring no\n>     additional context.\n>   * minimal_perfect_hash: Not a new concept, but now a separate\n>     field. It comes from the classifier's general-purpose hash table,\n>     which assigns each line a unique and minimal hash across the two\n>     files.\n>\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n\nI'm a bit surprised that nobody has commented on this patch. I thought\nthat someone would have criticized the length of the name\n\"minimal_perfect_hash\" or asked me why I was splitting one field into\ntwo.\n\nI don't see any reason why this patch series shouldn't move forward.\n"},{"id":"529214","messageId":"xmqqv7k8yh45.fsf@gitster.g","threadId":"64326","inReplyTo":"CAH=ZcbAjX=V_VvJsRzvQEA+CMM7dWQx6E5=d4FL5CD3s+ozjBg@mail.gmail.com","subject":"Re: [PATCH 5/9] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-21T05:10:50Z","receivedAt":"2025-10-21T05:10:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ezekiel Newren <ezekielnewren@gmail.com> writes:\n\n> I'm a bit surprised that nobody has commented on this patch. I thought\n> that someone would have criticized the length of the name\n> \"minimal_perfect_hash\" or asked me why I was splitting one field into\n> two.\n\nSometimes there aren't enough round tuits to go around, and when\npeople have been too busy to review it, we see no comment, either\npositive ones or negative ones.\n\n> I don't see any reason why this patch series shouldn't move forward.\n\nA patch series needs a positive reason to move forward;\nunfortunately we cannot tell much from lack of negative comments.\n\n"},{"id":"529241","messageId":"aPdFZp8GokGoshol@pks.im","threadId":"64326","inReplyTo":"7b9e8961d42e0f367ba0782e7d932607aa7e0b0a.1760563101.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-21T08:33:42Z","receivedAt":"2025-10-21T08:33:48Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Oct 15, 2025 at 09:18:14PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> diff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\n> index 6f3998ee54..411a8aa69f 100644\n> --- a/xdiff/xdiffi.c\n> +++ b/xdiff/xdiffi.c\n> @@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n>  \n>  \t\trec = &xe->xdf1.recs[xch->i1];\n>  \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n> -\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> +\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n>  \n>  \t\trec = &xe->xdf2.recs[xch->i2];\n>  \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n> -\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> +\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n>  \n>  \t\txch->ignore = ignore;\n>  \t}\n\nOkay. Seemingly, we convert the structure itself, but we don't convert\nany of the functions to accept an `uint8_t`. I guess you drew the line\nhere so that we don't have to also touch up dozens of function\nsignatures?\n\nAnd how did you end up verifying that you added all casts? Does the\ncompiler flag those as warnings?\n\nIn any case, it might be nice to explain both of these details in the\ncommit message.\n\nPatrick\n"},{"id":"529242","messageId":"aPdFbPN-60MVo3cv@pks.im","threadId":"64326","inReplyTo":"7fcd83c99076404960302b64a4f0c8fa1c13feba.1760563101.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 4/9] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-21T08:33:48Z","receivedAt":"2025-10-21T08:33:52Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Oct 15, 2025 at 09:18:16PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThis should have a commit message explaining what exactly you're doing\nhere.\n\nPatrick\n"},{"id":"529244","messageId":"aPdFczDXrlu6CJLZ@pks.im","threadId":"64326","inReplyTo":"CAH=ZcbAjX=V_VvJsRzvQEA+CMM7dWQx6E5=d4FL5CD3s+ozjBg@mail.gmail.com","subject":"Re: [PATCH 5/9] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-21T08:33:55Z","receivedAt":"2025-10-21T08:33:59Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Oct 20, 2025 at 05:29:25PM -0600, Ezekiel Newren wrote:\n> On Wed, Oct 15, 2025 at 3:18 PM Ezekiel Newren via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > The ha field is serving two different purposes, which makes the code\n> > harder to read. At first glance it looks like many places assume\n> > there could never be hash collisions between lines of the two input\n> > files. In reality, line_hash is used together with xdl_recmatch() to\n> > ensure correct comparisons of lines, even when collisions occur.\n> >\n> > To make this clearer, the old ha field has been split:\n> >   * line_hash: The straightforward hash of a line, requiring no\n> >     additional context.\n> >   * minimal_perfect_hash: Not a new concept, but now a separate\n> >     field. It comes from the classifier's general-purpose hash table,\n> >     which assigns each line a unique and minimal hash across the two\n> >     files.\n> >\n> > Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> I'm a bit surprised that nobody has commented on this patch. I thought\n> that someone would have criticized the length of the name\n> \"minimal_perfect_hash\" or asked me why I was splitting one field into\n> two.\n\nI actually appreciate the longer name. I'm not a fan of abbreviations\nthat are hard to understand myself. Sure, they are easier to type, but\nin many cases they end up making the code way harder to understand if\nyou are not deeply familiar with it. There's of course exceptions to\nthis, but I don't really think that your patch falls into them.\n\nPatrick\n"},{"id":"529243","messageId":"aPdFeHZKEsRw1cTX@pks.im","threadId":"64326","inReplyTo":"6dca5e6222e1d02092d4ba8296b757b123b85afa.1760563101.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 8/9] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-21T08:34:00Z","receivedAt":"2025-10-21T08:34:04Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Oct 15, 2025 at 09:18:20PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> rindex describes a index offset which means it's an index into memory\n> which should use size_t. dstart and dend will be deleted in a future\n> patch series. Move them to the end to help avoid refactor conflicts.\n\nIn a patch like this I would appreciate some explanation why we can\nchange the type without adapting any of its users. So basically explain\nwhy this refactoring is safe to do and won't cause any issues.\n\nPatrick\n"},{"id":"529249","messageId":"a0711cfe-6e44-44d6-b66b-84a296e113d2@gmail.com","threadId":"64326","inReplyTo":"CAH=ZcbAjX=V_VvJsRzvQEA+CMM7dWQx6E5=d4FL5CD3s+ozjBg@mail.gmail.com","subject":"Re: [PATCH 5/9] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-10-21T10:03:37Z","receivedAt":"2025-10-21T10:03:42Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ezekiel\n\nOn 21/10/2025 00:29, Ezekiel Newren wrote:\n> On Wed, Oct 15, 2025 at 3:18 PM Ezekiel Newren via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>>\n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>\n>> The ha field is serving two different purposes, which makes the code\n>> harder to read. At first glance it looks like many places assume\n>> there could never be hash collisions between lines of the two input\n>> files. In reality, line_hash is used together with xdl_recmatch() to\n>> ensure correct comparisons of lines, even when collisions occur.\n>>\n>> To make this clearer, the old ha field has been split:\n>>    * line_hash: The straightforward hash of a line, requiring no\n>>      additional context.\n>>    * minimal_perfect_hash: Not a new concept, but now a separate\n>>      field. It comes from the classifier's general-purpose hash table,\n>>      which assigns each line a unique and minimal hash across the two\n>>      files.\n>>\n>> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> I'm a bit surprised that nobody has commented on this patch.\n\nI've been off the list and I haven't caught up with this series yet.\n\n> I thought\n> that someone would have criticized the length of the name\n> \"minimal_perfect_hash\" or asked me why I was splitting one field into\n> two.\n\nI think \"perfect_hash\" would be fine if we want a shorter name. More \nimportantly it would be helpful to explain why the two fields have \ndifferent types. I assume it is because the perfect_hash is used as an \narray index and therefore size_t is a better match for rust's usize than \nuint64_t. How much more memory do we end up using by adding second hash \nmember to the struct? If the aim is to show that only one of them is \nused at a time then a union might be more appropriate but I doubt that \nplays well with rust.\n\nI'll try and have a look at the other patches later this week. I think \nthe type changes are going to need careful review.\n\nThanks\n\nPhillip\n"},{"id":"529251","messageId":"CAPx1Gvdd4KW=P=0te6ZeBXJPSp8NgyXnrEnJLb5g1uLcjNYnXQ@mail.gmail.com","threadId":"64326","inReplyTo":"a0711cfe-6e44-44d6-b66b-84a296e113d2@gmail.com","subject":"Re: [PATCH 5/9] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Chris Torek","fromEmail":"chris.torek@gmail.com","sentAt":"2025-10-21T11:16:53Z","receivedAt":"2025-10-21T11:17:08Z","isPatch":true,"sender":{"key":"chris.torek@gmail.com","avatar":"https://avatars.githubusercontent.com/u/16826774?v=4"},"body":"On Tue, Oct 21, 2025 at 3:04 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n...\n> uint64_t. How much more memory do we end up using by adding second hash\n> member to the struct?\n\nAs in any string-to-string algorithm of this sort, there's one per \"symbol\",\nbut in this case a \"symbol\" is a line in a file. So if files are M and N lines\nlong, there are M+N symbols. Take the difference of the size of the two\nrecords and multiply by this.\n\nAssuming \"sane\" input file sizes (under a million lines each) it's a few\nmegabytes maximum...\n\nChris\n"},{"id":"529253","messageId":"9eafee4d-ea94-4382-ada0-58000d229d2e@gmail.com","threadId":"64326","inReplyTo":"1fa9a7d7d1c309f2f651da351ba7bc0b36272d91.1760563101.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/9] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-10-21T11:32:55Z","receivedAt":"2025-10-21T11:32:59Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 15/10/2025 22:18, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> ssize_t is appropriate for dstart and dend because they both describe\n> positive or negative offsets relative to a pointer.\n\nIsn't ptrdiff_t the appropriate type for an offset to a pointer? ssize_t \nis not guaranteed to be the same width as size_t (this has caused \nproblems in the past[1]) and is only defined by POSIX, not the C standard.\n\nThanks\n\nPhillip\n\n[1] https://lore.kernel.org/git/loom.20150207T174514-727@post.gmane.org/\n\n> A future patch will move these fields to a different struct. Moving\n> them to the end of xdfile_t now, means the field order of xdfile_t will\n> be disturbed less.\n> \n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xtypes.h | 2 +-\n>   1 file changed, 1 insertion(+), 1 deletion(-)\n> \n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index f145abba3e..3514bb1684 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -47,10 +47,10 @@ typedef struct s_xrecord {\n>   typedef struct s_xdfile {\n>   \txrecord_t *recs;\n>   \tlong nrec;\n> -\tlong dstart, dend;\n>   \tbool *changed;\n>   \tlong *rindex;\n>   \tlong nreff;\n> +\tssize_t dstart, dend;\n>   } xdfile_t;\n>   \n>   typedef struct s_xdfenv {\n\n"},{"id":"529266","messageId":"786d6c19-0a13-4e55-8f4b-39b57dd6ea28@gmail.com","threadId":"64326","inReplyTo":"7b9e8961d42e0f367ba0782e7d932607aa7e0b0a.1760563101.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-10-21T13:13:52Z","receivedAt":"2025-10-21T13:13:58Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 15/10/2025 22:18, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Rust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\n> referring to bytes in memory, rather than unicode code points, use\n> uint8_t instead of char.\n\nIt C \"char\" never refers to a unicode code point so I don't follow the \nreasoning here. Isn't the reason you want to change from \"char\" to \n\"uint8_t\" to match rust? Given \"char\" and \"uint8_t\" are the same width \nwhy can't we use \"char\" in the C struct and \"u8\" in the rust struct as \nthe two structs would still have the same layout?\n\nI agree with Patrick's comments on this patch - it would be nice to know \nhow you decided where to add casts. Given that rust is going to be \noptional for at least a year we should take care to leave the C code in \ngood shape with a minimum number of casts.\n\nThanks\n\nPhillip\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xdiffi.c    |  8 ++++----\n>   xdiff/xemit.c     |  6 +++---\n>   xdiff/xmerge.c    | 14 +++++++-------\n>   xdiff/xpatience.c |  2 +-\n>   xdiff/xprepare.c  |  8 ++++----\n>   xdiff/xtypes.h    |  2 +-\n>   xdiff/xutils.c    |  4 ++--\n>   7 files changed, 22 insertions(+), 22 deletions(-)\n> \n> diff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\n> index 6f3998ee54..411a8aa69f 100644\n> --- a/xdiff/xdiffi.c\n> +++ b/xdiff/xdiffi.c\n> @@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n>   \tint ret = 0;\n>   \n>   \tfor (i = 0; i < rec->size; i++) {\n> -\t\tchar c = rec->ptr[i];\n> +\t\tuint8_t c = rec->ptr[i];\n>   \n>   \t\tif (!XDL_ISSPACE(c))\n>   \t\t\treturn ret;\n> @@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n>   \n>   \t\trec = &xe->xdf1.recs[xch->i1];\n>   \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n> -\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> +\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n>   \n>   \t\trec = &xe->xdf2.recs[xch->i2];\n>   \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n> -\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> +\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n>   \n>   \t\txch->ignore = ignore;\n>   \t}\n> @@ -1008,7 +1008,7 @@ static int record_matches_regex(xrecord_t *rec, xpparam_t const *xpp) {\n>   \tsize_t i;\n>   \n>   \tfor (i = 0; i < xpp->ignore_regex_nr; i++)\n> -\t\tif (!regexec_buf(xpp->ignore_regex[i], rec->ptr, rec->size, 1,\n> +\t\tif (!regexec_buf(xpp->ignore_regex[i], (const char *)rec->ptr, rec->size, 1,\n>   \t\t\t\t &regmatch, 0))\n>   \t\t\treturn 1;\n>   \n> diff --git a/xdiff/xemit.c b/xdiff/xemit.c\n> index b2f1f30cd3..ead930088a 100644\n> --- a/xdiff/xemit.c\n> +++ b/xdiff/xemit.c\n> @@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n>   {\n>   \txrecord_t *rec = &xdf->recs[ri];\n>   \n> -\tif (xdl_emit_diffrec(rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n> +\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n>   \t\treturn -1;\n>   \n>   \treturn 0;\n> @@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n>   \txrecord_t *rec = &xdf->recs[ri];\n>   \n>   \tif (!xecfg->find_func)\n> -\t\treturn def_ff(rec->ptr, rec->size, buf, sz);\n> -\treturn xecfg->find_func(rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n> +\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n> +\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n>   }\n>   \n>   static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\n> diff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\n> index fd600cbb5d..75cb3e76a2 100644\n> --- a/xdiff/xmerge.c\n> +++ b/xdiff/xmerge.c\n> @@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n>   \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n>   \n>   \tfor (i = 0; i < line_count; i++) {\n> -\t\tint result = xdl_recmatch(rec1[i].ptr, rec1[i].size,\n> -\t\t\trec2[i].ptr, rec2[i].size, flags);\n> +\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n> +\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n>   \t\tif (!result)\n>   \t\t\treturn -1;\n>   \t}\n> @@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n>   \n>   static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n>   {\n> -\treturn xdl_recmatch(rec1->ptr, rec1->size,\n> -\t\t\t    rec2->ptr, rec2->size, flags);\n> +\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n> +\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n>   }\n>   \n>   /*\n> @@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n>   \t\t * we have a very simple mmfile structure.\n>   \t\t */\n>   \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n> -\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n> +\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n>   \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n>   \t\tt2.ptr = (char *)xe2->xdf2.recs[m->i2].ptr;\n> -\t\tt2.size = xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n> +\t\tt2.size = (char *)xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n>   \t\t\t+ xe2->xdf2.recs[m->i2 + m->chg2 - 1].size - t2.ptr;\n>   \t\tif (xdl_do_diff(&t1, &t2, xpp, &xe) < 0)\n>   \t\t\treturn -1;\n> @@ -440,7 +440,7 @@ static int line_contains_alnum(const char *ptr, long size)\n>   static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n>   {\n>   \tfor (; chg; chg--, i++)\n> -\t\tif (line_contains_alnum(xe->xdf2.recs[i].ptr,\n> +\t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n>   \t\t\t\txe->xdf2.recs[i].size))\n>   \t\t\treturn 1;\n>   \treturn 0;\n> diff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\n> index 669b653580..bb61354f22 100644\n> --- a/xdiff/xpatience.c\n> +++ b/xdiff/xpatience.c\n> @@ -121,7 +121,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n>   \t\treturn;\n>   \tmap->entries[index].line1 = line;\n>   \tmap->entries[index].hash = record->ha;\n> -\tmap->entries[index].anchor = is_anchor(xpp, map->env->xdf1.recs[line - 1].ptr);\n> +\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n>   \tif (!map->first)\n>   \t\tmap->first = map->entries + index;\n>   \tif (map->last) {\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 192334f1b7..4cb18b2b88 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n>   \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n>   \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n>   \t\tif (rcrec->rec.ha == rec->ha &&\n> -\t\t\t\txdl_recmatch(rcrec->rec.ptr, rcrec->rec.size,\n> -\t\t\t\t\trec->ptr, rec->size, cf->flags))\n> +\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n> +\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n>   \t\t\tbreak;\n>   \n>   \tif (!rcrec) {\n> @@ -156,8 +156,8 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n>   \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n>   \t\t\t\tgoto abort;\n>   \t\t\tcrec = &xdf->recs[xdf->nrec++];\n> -\t\t\tcrec->ptr = prev;\n> -\t\t\tcrec->size = (long) (cur - prev);\n> +\t\t\tcrec->ptr = (uint8_t const *)prev;\n> +\t\t\tcrec->size =(long) ( cur - prev);\n>   \t\t\tcrec->ha = hav;\n>   \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n>   \t\t\t\tgoto abort;\n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index 3514bb1684..57983627f5 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -39,7 +39,7 @@ typedef struct s_chastore {\n>   } chastore_t;\n>   \n>   typedef struct s_xrecord {\n> -\tchar const *ptr;\n> +\tuint8_t const *ptr;\n>   \tlong size;\n>   \tunsigned long ha;\n>   } xrecord_t;\n> diff --git a/xdiff/xutils.c b/xdiff/xutils.c\n> index 447e66c719..7be063bfb6 100644\n> --- a/xdiff/xutils.c\n> +++ b/xdiff/xutils.c\n> @@ -465,10 +465,10 @@ int xdl_fall_back_diff(xdfenv_t *diff_env, xpparam_t const *xpp,\n>   \txdfenv_t env;\n>   \n>   \tsubfile1.ptr = (char *)diff_env->xdf1.recs[line1 - 1].ptr;\n> -\tsubfile1.size = diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n> +\tsubfile1.size = (char *)diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n>   \t\tdiff_env->xdf1.recs[line1 + count1 - 2].size - subfile1.ptr;\n>   \tsubfile2.ptr = (char *)diff_env->xdf2.recs[line2 - 1].ptr;\n> -\tsubfile2.size = diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n> +\tsubfile2.size = (char *)diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n>   \t\tdiff_env->xdf2.recs[line2 + count2 - 2].size - subfile2.ptr;\n>   \tif (xdl_do_diff(&subfile1, &subfile2, xpp, &env) < 0)\n>   \t\treturn -1;\n\n"},{"id":"529268","messageId":"93ec3dbf-ad98-4038-84e9-9ca12b7481a0@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] Xdiff cleanup part2","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-10-21T13:28:32Z","receivedAt":"2025-10-21T13:28:39Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ezekiel\n\nOn 15/10/2025 22:18, Ezekiel Newren via GitGitGadget wrote:\n> Maintainer note: This patch series builds on top of en/xdiff-cleanup and\n> am/xdiff-hash-tweak (both of which are now in master).\n> \n> The primary goal of this patch series is to convert every field's type in\n> xrecord_t and xdfile_t to be unambiguous, in preparation to make it more\n> Rust FFI friendly. Additionally the ha field in xrecord_t is split into\n> line_hash and minimal_perfect hash.\n\nGiven that this series changes the types of all the \"long\" struct \nmembers to \"size_t\" I was surprised to see that it adds so many \"(long)\" \ncasts. At the end of this series there are 38 lines in xdiff/ that \ncontain \"(long)\" compared to just 4 in master. I had expected that as \nwe'd converted all the members to \"size_t\" there would be no need to \nkeep using \"long\" in the code. As rust is going to be optional for quite \na while I think we should clean up the C code to avoid casting between \n\"long\" and \"size_t\"\n\nThanks\n\nPhillip\n\n> The order of some of the fields has changed as called out by the commit\n> messages.\n> \n> Before:\n> \n> typedef struct s_xrecord {\n> \tchar const *ptr;\n> \tlong size;\n> \tunsigned long ha;\n> } xrecord_t;\n> \n> typedef struct s_xdfile {\n> \txrecord_t *recs;\n> \tlong nrec;\n> \tlong dstart, dend;\n> \tbool *changed;\n> \tlong *rindex;\n> \tlong nreff;\n> } xdfile_t;\n> \n> \n> After part 2\n> \n> typedef struct s_xrecord {\n> \tuint8_t const *ptr;\n> \tsize_t size;\n> \tuint64_t line_hash;\n> \tsize_t minimal_perfect_hash;\n> } xrecord_t;\n> \n> typedef struct s_xdfile {\n> \txrecord_t *recs;\n> \tsize_t nrec;\n> \tbool *changed;\n> \tsize_t *reference_index;\n> \tsize_t nreff;\n> \tssize_t dstart, dend;\n> } xdfile_t;\n> \n> \n> Ezekiel Newren (9):\n>    xdiff: use ssize_t for dstart/dend, make them last in xdfile_t\n>    xdiff: make xrecord_t.ptr a uint8_t instead of char\n>    xdiff: use size_t for xrecord_t.size\n>    xdiff: use unambiguous types in xdl_hash_record()\n>    xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n>    xdiff: make xdfile_t.nrec a size_t instead of long\n>    xdiff: make xdfile_t.nreff a size_t instead of long\n>    xdiff: change rindex from long to size_t in xdfile_t\n>    xdiff: rename rindex -> reference_index\n> \n>   xdiff-interface.c  |  2 +-\n>   xdiff/xdiffi.c     | 29 +++++++++++------------\n>   xdiff/xemit.c      | 28 +++++++++++-----------\n>   xdiff/xhistogram.c |  4 ++--\n>   xdiff/xmerge.c     | 30 ++++++++++++------------\n>   xdiff/xpatience.c  | 14 +++++------\n>   xdiff/xprepare.c   | 58 +++++++++++++++++++++++-----------------------\n>   xdiff/xtypes.h     | 15 ++++++------\n>   xdiff/xutils.c     | 32 ++++++++++++-------------\n>   xdiff/xutils.h     |  6 ++---\n>   10 files changed, 109 insertions(+), 109 deletions(-)\n> \n> \n> base-commit: 143f58ef7535f8f8a80d810768a18bdf3807de26\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2070%2Fezekielnewren%2Fxdiff_cleanup_part2-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2070/ezekielnewren/xdiff_cleanup_part2-v1\n> Pull-Request: https://github.com/git/git/pull/2070\n\n"},{"id":"529269","messageId":"xmqqqzuwxthp.fsf@gitster.g","threadId":"64326","inReplyTo":"93ec3dbf-ad98-4038-84e9-9ca12b7481a0@gmail.com","subject":"Re: [PATCH 0/9] Xdiff cleanup part2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-21T13:41:06Z","receivedAt":"2025-10-21T13:41:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> Given that this series changes the types of all the \"long\" struct \n> members to \"size_t\" I was surprised to see that it adds so many \"(long)\" \n> casts. At the end of this series there are 38 lines in xdiff/ that \n> contain \"(long)\" compared to just 4 in master. I had expected that as \n> we'd converted all the members to \"size_t\" there would be no need to \n> keep using \"long\" in the code. As rust is going to be optional for quite \n> a while I think we should clean up the C code to avoid casting between \n> \"long\" and \"size_t\"\n\nEither we cast here or have existing code that used to use long to\nuse another type, that needs to be done carefully as we would be\nmoving code that used signed type to now use unsigned.  While I\nagree with you in principle that we shouldn't try to interface\nbetween code pieces with impedance mismatch (for which the need to\ncast is an indication), we'd need to draw a line somewhere.\n\nThanks.\n\n"},{"id":"529306","messageId":"xmqqecqww4u7.fsf@gitster.g","threadId":"64326","inReplyTo":"9eafee4d-ea94-4382-ada0-58000d229d2e@gmail.com","subject":"Re: [PATCH 1/9] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-21T17:18:56Z","receivedAt":"2025-10-21T17:18:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> On 15/10/2025 22:18, Ezekiel Newren via GitGitGadget wrote:\n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>> \n>> ssize_t is appropriate for dstart and dend because they both describe\n>> positive or negative offsets relative to a pointer.\n>\n> Isn't ptrdiff_t the appropriate type for an offset to a pointer? ssize_t \n> is not guaranteed to be the same width as size_t (this has caused \n> problems in the past[1]) and is only defined by POSIX, not the C standard.\n>\n> Thanks\n>\n> Phillip\n>\n> [1] https://lore.kernel.org/git/loom.20150207T174514-727@post.gmane.org/\n\nThanks for bringing up a very good point.\n\nWe often consider that a function that yields what we would normally\nput in a size_t variable, when we _know_ that the return value would\nnot be so big to exceed half the range of size_t, can instead return\nssize_t and use the negative half of the range to signal error\nconditions, but as the cited incident shows that it is an easy\nmistake to make.\n"},{"id":"529309","messageId":"xmqqplagunnm.fsf@gitster.g","threadId":"64326","inReplyTo":"786d6c19-0a13-4e55-8f4b-39b57dd6ea28@gmail.com","subject":"Re: [PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-21T18:15:25Z","receivedAt":"2025-10-21T18:15:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> It C \"char\" never refers to a unicode code point so I don't follow the \n> reasoning here. Isn't the reason you want to change from \"char\" to \n> \"uint8_t\" to match rust? Given \"char\" and \"uint8_t\" are the same width \n> why can't we use \"char\" in the C struct and \"u8\" in the rust struct as \n> the two structs would still have the same layout?\n\nAnd forcing u8 makes sure both sides of the ffi agrees on the\nsignedness (C \"char\"'s signedness is implementation defined),\nwhich is a good thing.\n\nI 100% agree that being honest about the motivation to sell this\nchange would be a good thing to do here.  I do not think \"in this\nseries, I want to match the types used at the interface to be of\nRust's\" is a position to be ashamed of ;-)\n\n> I agree with Patrick's comments on this patch - it would be nice to know \n> how you decided where to add casts. Given that rust is going to be \n> optional for at least a year we should take care to leave the C code in \n> good shape with a minimum number of casts.\n\nThanks.\n"},{"id":"529414","messageId":"d863c518-3246-4752-83f3-469592b1de69@gmail.com","threadId":"64326","inReplyTo":"xmqqplagunnm.fsf@gitster.g","subject":"Re: [PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-10-22T13:27:09Z","receivedAt":"2025-10-22T13:27:12Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 21/10/2025 19:15, Junio C Hamano wrote:\n> Phillip Wood <phillip.wood123@gmail.com> writes:\n> \n>> It C \"char\" never refers to a unicode code point so I don't follow the\n>> reasoning here. Isn't the reason you want to change from \"char\" to\n>> \"uint8_t\" to match rust? Given \"char\" and \"uint8_t\" are the same width\n>> why can't we use \"char\" in the C struct and \"u8\" in the rust struct as\n>> the two structs would still have the same layout?\n> \n> And forcing u8 makes sure both sides of the ffi agrees on the\n> signedness (C \"char\"'s signedness is implementation defined),\n> which is a good thing.\n\nThat's true and ignoring the signedness would be hacky but I'm not sure \nit matters in practice. Both C and rust would use the same bit patterns \nfor \"abc\" and b\"abc\\0\" and in general C plays fast and loose with the \nsignedness of variables all over the place. The trade off for respecting \nthe signedness is that we either have casts all over the place or \nmassive churn converting the rest of the code to use uint8_t. This \nproblem isn't limited to xdiff, it will be true wherever we share \nbytestrings such as the contents of objects between C and rust as we \ntend to use char rather than uint8_t in our code.\n\nThanks\n\nPhillip\n\n> I 100% agree that being honest about the motivation to sell this\n> change would be a good thing to do here.  I do not think \"in this\n> series, I want to match the types used at the interface to be of\n> Rust's\" is a position to be ashamed of ;-)\n> \n>> I agree with Patrick's comments on this patch - it would be nice to know\n>> how you decided where to add casts. Given that rust is going to be\n>> optional for at least a year we should take care to leave the C code in\n>> good shape with a minimum number of casts.\n> \n> Thanks.\n\n"},{"id":"529448","messageId":"CAH=ZcbALmH1LRKpLXygUOPiNJeoG2Uqvkb0fuy_i412W=z2oeQ@mail.gmail.com","threadId":"64326","inReplyTo":"d863c518-3246-4752-83f3-469592b1de69@gmail.com","subject":"Re: [PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-22T20:55:10Z","receivedAt":"2025-10-22T20:55:23Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Wed, Oct 22, 2025 at 7:27 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n> > I 100% agree that being honest about the motivation to sell this\n> > change would be a good thing to do here.  I do not think \"in this\n> > series, I want to match the types used at the interface to be of\n> > Rust's\" is a position to be ashamed of ;-)\n> >\n> >> I agree with Patrick's comments on this patch - it would be nice to know\n> >> how you decided where to add casts. Given that rust is going to be\n> >> optional for at least a year we should take care to leave the C code in\n> >> good shape with a minimum number of casts.\n> >\n> > Thanks.\n\nI'm not arguing that uint8_t should be used everywhere in Git, only\nthat it is used everywhere in xdiff. xrecord_t and xdfile_t are\nfundamental to how xdiff passes data around and they need to be\ntransparent to both sides. I'm trying to leave the rest of the data\nstructures alone in order to avoid refactor churn. Refactoring C to\nuse unambiguous types, outside of xdiff, is outside the scope of this\npatch series.\n\nAnother problem with using char instead of uint8_t is that tools like\ncbindgen and bindgen don't translate char to u8. Bindgen will see char\nand will produce std::ffi::c_char on the Rust side, see [1] for why\nthat's a problem. The other way around is a problem too. When cbindgen\nsees u8 it will generate uint8_t on the C side and then `make\nDEVELOPER=1` won't compile because uint8_t and char differer in\nsignedness.\n\n[1] Problems with C types\nhttps://lore.kernel.org/git/CAH=ZcbA_8JM1hdUAfFe3ho0ShuniguEpV1308S0nCkCHOCsmmg@mail.gmail.com/\n"},{"id":"529449","messageId":"CAH=ZcbBmdWCBh9zH1Y1JxcnNS-E9AU6Q4rRXPhMOtDBmkxLd8g@mail.gmail.com","threadId":"64326","inReplyTo":"xmqqecqww4u7.fsf@gitster.g","subject":"Re: [PATCH 1/9] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-22T21:07:40Z","receivedAt":"2025-10-22T21:07:54Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Oct 21, 2025 at 11:18 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Phillip Wood <phillip.wood123@gmail.com> writes:\n>\n> > On 15/10/2025 22:18, Ezekiel Newren via GitGitGadget wrote:\n> >> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >>\n> >> ssize_t is appropriate for dstart and dend because they both describe\n> >> positive or negative offsets relative to a pointer.\n> >\n> > Isn't ptrdiff_t the appropriate type for an offset to a pointer? ssize_t\n> > is not guaranteed to be the same width as size_t (this has caused\n> > problems in the past[1]) and is only defined by POSIX, not the C standard.\n> >\n> > Thanks\n> >\n> > Phillip\n> >\n> > [1] https://lore.kernel.org/git/loom.20150207T174514-727@post.gmane.org/\n>\n> Thanks for bringing up a very good point.\n>\n> We often consider that a function that yields what we would normally\n> put in a size_t variable, when we _know_ that the return value would\n> not be so big to exceed half the range of size_t, can instead return\n> ssize_t and use the negative half of the range to signal error\n> conditions, but as the cited incident shows that it is an easy\n> mistake to make.\n\nIn my compat/rust_types.h file (which was dropped) I defined isize\nusing ptrdiff_t rather than ssize_t. Maybe that file should be revived\nso that we don't have confusion in code reviews when structs are being\nexpressly converted for the purpose of Rust FFI? I'd really like to\nbring that file back so that everyone has a clear reference for how C\ntypes map to Rust, but no one seemed to like it except me. Maybe it\nshould be an adoc file rather than a header?\n\n[1] compat/rust_types.h\nhttps://lore.kernel.org/git/2a7d5b05c18d4a96f1905b7043d47c62d367cd2a.1757274320.git.gitgitgadget@gmail.com/\n"},{"id":"529450","messageId":"CAH=ZcbDOY2yDQbBJeoKHesZzZCBvscqf7SoqbX4j3oHCBY5p8g@mail.gmail.com","threadId":"64326","inReplyTo":"aPdFZp8GokGoshol@pks.im","subject":"Re: [PATCH 2/9] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-22T21:12:50Z","receivedAt":"2025-10-22T21:13:04Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Oct 21, 2025 at 2:33 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Wed, Oct 15, 2025 at 09:18:14PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> > diff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\n> > index 6f3998ee54..411a8aa69f 100644\n> > --- a/xdiff/xdiffi.c\n> > +++ b/xdiff/xdiffi.c\n> > @@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n> >\n> >               rec = &xe->xdf1.recs[xch->i1];\n> >               for (i = 0; i < xch->chg1 && ignore; i++)\n> > -                     ignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> > +                     ignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n> >\n> >               rec = &xe->xdf2.recs[xch->i2];\n> >               for (i = 0; i < xch->chg2 && ignore; i++)\n> > -                     ignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> > +                     ignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n> >\n> >               xch->ignore = ignore;\n> >       }\n>\n> Okay. Seemingly, we convert the structure itself, but we don't convert\n> any of the functions to accept an `uint8_t`. I guess you drew the line\n> here so that we don't have to also touch up dozens of function\n> signatures?\n\nThat is correct. I wanted to avoid _boiling the ocean_ just to change\nthe type of ptr.\n\n> And how did you end up verifying that you added all casts? Does the\n> compiler flag those as warnings?\n\nI used CLion to search for all uses of that field and then added casts\nwhere the types differ. Another way to do that is to run `make\nDEVELOPER=1` and address all of the `uint8_t differs in signedness\nfrom char` errors that are spat out.\n\n> In any case, it might be nice to explain both of these details in the\n> commit message.\n\nI will update it.\n\nThanks.\n"},{"id":"529452","messageId":"CAH=ZcbBeDNqW6PqhhzU75wttND86RfMRuNS2ga6KP1fN7AhFnw@mail.gmail.com","threadId":"64326","inReplyTo":"aPdFbPN-60MVo3cv@pks.im","subject":"Re: [PATCH 4/9] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-22T21:20:32Z","receivedAt":"2025-10-22T21:20:46Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Oct 21, 2025 at 2:33 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Wed, Oct 15, 2025 at 09:18:16PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> This should have a commit message explaining what exactly you're doing\n> here.\n\nI thought I did have a commit message justifying my changes. Maybe it\ngot deleted through a rebase. How about a message like:\n\nConvert the function signature and body to use unambiguous types. char\nis changed to uint8_t because this function processes bytes in memory.\nunsigned long to uint64_t so that the hash output is consistent across\nplatforms. `flags` was changed from long to uint64_t to ensure the\nhigh order bits are not dropped on platforms that treat long as 32\nbits.\n"},{"id":"529453","messageId":"CAH=ZcbD7FeRHtYvN_4=qHApB-AwK18=KRU2SGWNg8ADkrFM-Fw@mail.gmail.com","threadId":"64326","inReplyTo":"a0711cfe-6e44-44d6-b66b-84a296e113d2@gmail.com","subject":"Re: [PATCH 5/9] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-22T21:31:57Z","receivedAt":"2025-10-22T21:32:10Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Oct 21, 2025 at 4:03 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 21/10/2025 00:29, Ezekiel Newren wrote:\n> > On Wed, Oct 15, 2025 at 3:18 PM Ezekiel Newren via GitGitGadget\n> > <gitgitgadget@gmail.com> wrote:\n> >>\n> >> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >>\n> >> The ha field is serving two different purposes, which makes the code\n> >> harder to read. At first glance it looks like many places assume\n> >> there could never be hash collisions between lines of the two input\n> >> files. In reality, line_hash is used together with xdl_recmatch() to\n> >> ensure correct comparisons of lines, even when collisions occur.\n> >>\n> >> To make this clearer, the old ha field has been split:\n> >>    * line_hash: The straightforward hash of a line, requiring no\n> >>      additional context.\n> >>    * minimal_perfect_hash: Not a new concept, but now a separate\n> >>      field. It comes from the classifier's general-purpose hash table,\n> >>      which assigns each line a unique and minimal hash across the two\n> >>      files.\n> >>\n> >> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > I'm a bit surprised that nobody has commented on this patch.\n>\n> I've been off the list and I haven't caught up with this series yet.\n>\n> > I thought\n> > that someone would have criticized the length of the name\n> > \"minimal_perfect_hash\" or asked me why I was splitting one field into\n> > two.\n>\n> I think \"perfect_hash\" would be fine if we want a shorter name. More\n> importantly it would be helpful to explain why the two fields have\n> different types. I assume it is because the perfect_hash is used as an\n> array index and therefore size_t is a better match for rust's usize than\n> uint64_t.\n\nYour understanding is correct. line_hash is fixed width while\nminimal_perfect_hash is meant to be used as an array index into\nmemory. I'll update my commit message to make this more clear.\n\n> How much more memory do we end up using by adding second hash\n> member to the struct? If the aim is to show that only one of them is\n> used at a time then a union might be more appropriate but I doubt that\n> plays well with rust.\n\nxrecord_t used to be defined with a pointer, so we're at the same\nsize. But more importantly I plan on splitting minimal_perfect_hash\nout of xrecord_t into its own array. I think the diff algorithms end\nup being a little bit faster with a separate array because each\nelement is only 8 bytes instead of 32.\n\nIn v2.51.0:\ntypedef struct s_xrecord {\n       struct s_xrecord *next;\n       char const *ptr;\n       long size;\n       unsigned long ha;\n} xrecord_t;\n\nThis patch series:\ntypedef struct s_xrecord {\n       uint8_t const *ptr;\n       size_t size;\n       uint64_t line_hash;\n       size_t minimal_perfect_hash;\n} xrecord_t;\n\n> I'll try and have a look at the other patches later this week. I think\n> the type changes are going to need careful review.\n\nI appreciate the careful review. I figured it would be best to limit\nthe scope of this patch series to type changes, so that it wasn't\nbogged down by other stuff.\n"},{"id":"529455","messageId":"xmqqqzuuwra4.fsf@gitster.g","threadId":"64326","inReplyTo":"CAH=ZcbBmdWCBh9zH1Y1JxcnNS-E9AU6Q4rRXPhMOtDBmkxLd8g@mail.gmail.com","subject":"Re: [PATCH 1/9] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-22T21:38:43Z","receivedAt":"2025-10-22T21:38:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ezekiel Newren <ezekielnewren@gmail.com> writes:\n\n> In my compat/rust_types.h file (which was dropped) I defined isize\n> using ptrdiff_t rather than ssize_t. Maybe that file should be revived\n> so that we don't have confusion in code reviews when structs are being\n> expressly converted for the purpose of Rust FFI? I'd really like to\n> bring that file back so that everyone has a clear reference for how C\n> types map to Rust, but no one seemed to like it except me. Maybe it\n> should be an adoc file rather than a header?\n\nI may be mistaken, but I thought that the latest agreement was to\nuse conceptually the \"same\" type in each language, have each\nlanguage call that type in its native way, and if needed convert at\nthe FFI boundary.  So if we agree to use, for example, 64-bit signed\ninteger type for counting things plus returning error conditions via\nnegative values, maybe C-side can agree to use i64 for it, without\nhaving to worry about how that thing is called in Rust side.\n\nI am not sure in what way <compat/rust_types.h> should be used, and\nperhaps a documentation file may be sufficient as you suggest, but\nin any case, I agree that it should be made clear to everybody what\nC-types are to be mapped to what Rust types and vice versa, and if\nsome C-types have no corresponding Rust type in that mapping, or if\nsome Rust types have no corresponding C-type, that type needs to be\nconverted before they reach the FFI boundary.\n\n"},{"id":"529456","messageId":"CAH=ZcbDVBWcRzOmJM7OWvtap2F-84qJ0zcU+Z8u8yX4p7CWb=Q@mail.gmail.com","threadId":"64326","inReplyTo":"xmqqqzuuwra4.fsf@gitster.g","subject":"Re: [PATCH 1/9] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-22T21:51:14Z","receivedAt":"2025-10-22T21:51:27Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Wed, Oct 22, 2025 at 3:38 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Ezekiel Newren <ezekielnewren@gmail.com> writes:\n>\n> > In my compat/rust_types.h file (which was dropped) I defined isize\n> > using ptrdiff_t rather than ssize_t. Maybe that file should be revived\n> > so that we don't have confusion in code reviews when structs are being\n> > expressly converted for the purpose of Rust FFI? I'd really like to\n> > bring that file back so that everyone has a clear reference for how C\n> > types map to Rust, but no one seemed to like it except me. Maybe it\n> > should be an adoc file rather than a header?\n>\n> I may be mistaken, but I thought that the latest agreement was to\n> use conceptually the \"same\" type in each language, have each\n> language call that type in its native way, and if needed convert at\n> the FFI boundary.  So if we agree to use, for example, 64-bit signed\n> integer type for counting things plus returning error conditions via\n> negative values, maybe C-side can agree to use i64 for it, without\n> having to worry about how that thing is called in Rust side.\n\nYour understanding is correct. Would\nDocumentation/unambiguous_types.adoc be an appropriate place for this\ndocumentation?\n\n> I am not sure in what way <compat/rust_types.h> should be used, and\n> perhaps a documentation file may be sufficient as you suggest, but\n> in any case, I agree that it should be made clear to everybody what\n> C-types are to be mapped to what Rust types and vice versa, and if\n> some C-types have no corresponding Rust type in that mapping, or if\n> some Rust types have no corresponding C-type, that type needs to be\n> converted before they reach the FFI boundary.\n\nAlright. I guess I'll drop the idea of compat/rust_types.h permanently.\n"},{"id":"529457","messageId":"CAH=ZcbBbnoiBndEYryMpDzav+-iHFA7_3BPNw8hgOBiaFjCq0A@mail.gmail.com","threadId":"64326","inReplyTo":"aPdFeHZKEsRw1cTX@pks.im","subject":"Re: [PATCH 8/9] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-10-22T22:14:42Z","receivedAt":"2025-10-22T22:14:55Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Oct 21, 2025 at 2:34 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Wed, Oct 15, 2025 at 09:18:20PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > rindex describes a index offset which means it's an index into memory\n> > which should use size_t. dstart and dend will be deleted in a future\n> > patch series. Move them to the end to help avoid refactor conflicts.\n>\n> In a patch like this I would appreciate some explanation why we can\n> change the type without adapting any of its users. So basically explain\n> why this refactoring is safe to do and won't cause any issues.\n\nThe values of rindex are only used in 3 places. get_hash() which was\ncreated in [1]. and 2 places in xdl_recs_cmp(). All of them use rindex\nas an index into another array directly so there's no cascading\nrefactor impact. get_hash() was created precisely to reduce refactor\nchurn. How about a commit message like:\n\nChanging the type of rindex from long to size_t has no cascading\nrefactor impact because it is only ever used to directly index other\narrays.\n\n[1] create get_hash()\nhttps://lore.kernel.org/git/637d1032abbd33b7673d3c101267816fbf1a343c.1758926520.git.gitgitgadget@gmail.com/\n"},{"id":"529463","messageId":"aPnB_jEjn-nnRg82@pks.im","threadId":"64326","inReplyTo":"CAH=ZcbBeDNqW6PqhhzU75wttND86RfMRuNS2ga6KP1fN7AhFnw@mail.gmail.com","subject":"Re: [PATCH 4/9] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-23T05:49:50Z","receivedAt":"2025-10-23T05:49:56Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Oct 22, 2025 at 03:20:32PM -0600, Ezekiel Newren wrote:\n> On Tue, Oct 21, 2025 at 2:33 AM Patrick Steinhardt <ps@pks.im> wrote:\n> >\n> > On Wed, Oct 15, 2025 at 09:18:16PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> > > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > This should have a commit message explaining what exactly you're doing\n> > here.\n> \n> I thought I did have a commit message justifying my changes. Maybe it\n> got deleted through a rebase. How about a message like:\n> \n> Convert the function signature and body to use unambiguous types. char\n> is changed to uint8_t because this function processes bytes in memory.\n> unsigned long to uint64_t so that the hash output is consistent across\n> platforms. `flags` was changed from long to uint64_t to ensure the\n> high order bits are not dropped on platforms that treat long as 32\n> bits.\n\nWorks for me, I guess. Thanks!\n\nPatrick\n"},{"id":"529464","messageId":"aPnCA7lzREhUETKc@pks.im","threadId":"64326","inReplyTo":"CAH=ZcbBbnoiBndEYryMpDzav+-iHFA7_3BPNw8hgOBiaFjCq0A@mail.gmail.com","subject":"Re: [PATCH 8/9] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-10-23T05:49:55Z","receivedAt":"2025-10-23T05:50:00Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Oct 22, 2025 at 04:14:42PM -0600, Ezekiel Newren wrote:\n> On Tue, Oct 21, 2025 at 2:34 AM Patrick Steinhardt <ps@pks.im> wrote:\n> >\n> > On Wed, Oct 15, 2025 at 09:18:20PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> > > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> > >\n> > > rindex describes a index offset which means it's an index into memory\n> > > which should use size_t. dstart and dend will be deleted in a future\n> > > patch series. Move them to the end to help avoid refactor conflicts.\n> >\n> > In a patch like this I would appreciate some explanation why we can\n> > change the type without adapting any of its users. So basically explain\n> > why this refactoring is safe to do and won't cause any issues.\n> \n> The values of rindex are only used in 3 places. get_hash() which was\n> created in [1]. and 2 places in xdl_recs_cmp(). All of them use rindex\n> as an index into another array directly so there's no cascading\n> refactor impact. get_hash() was created precisely to reduce refactor\n> churn. How about a commit message like:\n> \n> Changing the type of rindex from long to size_t has no cascading\n> refactor impact because it is only ever used to directly index other\n> arrays.\n\nSounds good to me, thanks!\n\nPatrick\n"},{"id":"529894","messageId":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.git.git.1760563101.gitgitgadget@gmail.com","subject":"[PATCH v2 00/10] Xdiff cleanup part2","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:38Z","receivedAt":"2025-10-29T22:19:51Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"Changes in v2:\n\n * Added documentation about unambiguous types and FFI\n * Addressed comments on the mailing list\n\n\nOriginal cover letter below:\n============================\n\nMaintainer note: This patch series builds on top of en/xdiff-cleanup and\nam/xdiff-hash-tweak (both of which are now in master).\n\nThe primary goal of this patch series is to convert every field's type in\nxrecord_t and xdfile_t to be unambiguous, in preparation to make it more\nRust FFI friendly. Additionally the ha field in xrecord_t is split into\nline_hash and minimal_perfect hash.\n\nThe order of some of the fields has changed as called out by the commit\nmessages.\n\nBefore:\n\ntypedef struct s_xrecord {\n\tchar const *ptr;\n\tlong size;\n\tunsigned long ha;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tlong nrec;\n\tlong dstart, dend;\n\tbool *changed;\n\tlong *rindex;\n\tlong nreff;\n} xdfile_t;\n\n\nAfter part 2\n\ntypedef struct s_xrecord {\n\tuint8_t const *ptr;\n\tsize_t size;\n\tuint64_t line_hash;\n\tsize_t minimal_perfect_hash;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tsize_t nrec;\n\tbool *changed;\n\tsize_t *reference_index;\n\tsize_t nreff;\n\tssize_t dstart, dend;\n} xdfile_t;\n\n\nEzekiel Newren (10):\n  doc: define unambiguous type mappings across C and Rust\n  xdiff: use ssize_t for dstart/dend, make them last in xdfile_t\n  xdiff: make xrecord_t.ptr a uint8_t instead of char\n  xdiff: use size_t for xrecord_t.size\n  xdiff: use unambiguous types in xdl_hash_record()\n  xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n  xdiff: make xdfile_t.nrec a size_t instead of long\n  xdiff: make xdfile_t.nreff a size_t instead of long\n  xdiff: change rindex from long to size_t in xdfile_t\n  xdiff: rename rindex -> reference_index\n\n .../technical/unambiguous-types.adoc          | 229 ++++++++++++++++++\n xdiff-interface.c                             |   2 +-\n xdiff/xdiffi.c                                |  29 ++-\n xdiff/xemit.c                                 |  28 +--\n xdiff/xhistogram.c                            |   4 +-\n xdiff/xmerge.c                                |  30 +--\n xdiff/xpatience.c                             |  14 +-\n xdiff/xprepare.c                              |  58 ++---\n xdiff/xtypes.h                                |  15 +-\n xdiff/xutils.c                                |  32 +--\n xdiff/xutils.h                                |   6 +-\n 11 files changed, 338 insertions(+), 109 deletions(-)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\n\nbase-commit: 143f58ef7535f8f8a80d810768a18bdf3807de26\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2070%2Fezekielnewren%2Fxdiff_cleanup_part2-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2070/ezekielnewren/xdiff_cleanup_part2-v2\nPull-Request: https://github.com/git/git/pull/2070\n\nRange-diff vs v1:\n\n  -:  ---------- >  1:  88133848d1 doc: define unambiguous type mappings across C and Rust\n  1:  1fa9a7d7d1 !  2:  9197903add xdiff: use ssize_t for dstart/dend, make them last in xdfile_t\n     @@ xdiff/xtypes.h: typedef struct s_xrecord {\n       \tbool *changed;\n       \tlong *rindex;\n       \tlong nreff;\n     -+\tssize_t dstart, dend;\n     ++\tptrdiff_t dstart, dend;\n       } xdfile_t;\n       \n       typedef struct s_xdfenv {\n  2:  7b9e8961d4 !  3:  46bc1b3e25 xdiff: make xrecord_t.ptr a uint8_t instead of char\n     @@ Commit message\n          xdiff: make xrecord_t.ptr a uint8_t instead of char\n      \n          Rust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\n     -    referring to bytes in memory, rather than unicode code points, use\n     +    referring to bytes in memory, rather than Unicode code points, use\n          uint8_t instead of char.\n      \n     +    Every usage of this field was inspected and cast to char*, or similar,\n     +    to avoid signedness warnings/errors from the compiler. Casting was used\n     +    so that the whole of xdiff doesn't need to be refactored in order to\n     +    change the type of this field.\n     +\n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xdiffi.c ##\n  3:  ae15ed7121 =  4:  07e28aad3b xdiff: use size_t for xrecord_t.size\n  4:  7fcd83c990 !  5:  1ade7d8165 xdiff: use unambiguous types in xdl_hash_record()\n     @@ Metadata\n       ## Commit message ##\n          xdiff: use unambiguous types in xdl_hash_record()\n      \n     +    Convert the function signature and body to use unambiguous types. char\n     +    is changed to uint8_t because this function processes bytes in memory.\n     +    unsigned long to uint64_t so that the hash output is consistent across\n     +    platforms. `flags` was changed from long to uint64_t to ensure the\n     +    high order bits are not dropped on platforms that treat long as 32\n     +    bits.\n     +\n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff-interface.c ##\n  5:  a3e706ecda =  6:  59054ea0cb xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n  6:  5767ba4ee8 =  7:  f91be17858 xdiff: make xdfile_t.nrec a size_t instead of long\n  7:  4caa6a4669 !  8:  e2a6a23cc4 xdiff: make xdfile_t.nreff a size_t instead of long\n     @@ xdiff/xtypes.h: typedef struct s_xdfile {\n       \tlong *rindex;\n      -\tlong nreff;\n      +\tsize_t nreff;\n     - \tssize_t dstart, dend;\n     + \tptrdiff_t dstart, dend;\n       } xdfile_t;\n       \n  8:  6dca5e6222 !  9:  3b6054945f xdiff: change rindex from long to size_t in xdfile_t\n     @@ Commit message\n          xdiff: change rindex from long to size_t in xdfile_t\n      \n          rindex describes a index offset which means it's an index into memory\n     -    which should use size_t. dstart and dend will be deleted in a future\n     -    patch series. Move them to the end to help avoid refactor conflicts.\n     +    which should use size_t.\n     +\n     +    Changing the type of rindex from long to size_t has no cascading\n     +    refactor impact because it is only ever used to directly index other\n     +    arrays.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n     @@ xdiff/xtypes.h: typedef struct s_xdfile {\n      -\tlong *rindex;\n      +\tsize_t *rindex;\n       \tsize_t nreff;\n     - \tssize_t dstart, dend;\n     + \tptrdiff_t dstart, dend;\n       } xdfile_t;\n  9:  518e5f5557 ! 10:  1856a29026 xdiff: rename rindex -> reference_index\n     @@ xdiff/xtypes.h: typedef struct s_xdfile {\n      -\tsize_t *rindex;\n      +\tsize_t *reference_index;\n       \tsize_t nreff;\n     - \tssize_t dstart, dend;\n     + \tptrdiff_t dstart, dend;\n       } xdfile_t;\n\n-- \ngitgitgadget\n"},{"id":"529895","messageId":"88133848d1a317f8a95c19ee5482b828a3f8705f.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:39Z","receivedAt":"2025-10-29T22:19:52Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nDocument other nuances with crossing the FFI boundary. Other language\nmappings may be added in the future.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n .../technical/unambiguous-types.adoc          | 229 ++++++++++++++++++\n 1 file changed, 229 insertions(+)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\ndiff --git a/Documentation/technical/unambiguous-types.adoc b/Documentation/technical/unambiguous-types.adoc\nnew file mode 100644\nindex 0000000000..658a5b578e\n--- /dev/null\n+++ b/Documentation/technical/unambiguous-types.adoc\n@@ -0,0 +1,229 @@\n+= Unambiguous types\n+\n+Most of these mappings are obvious, but there are some nuances and gotchas with\n+Rust FFI (Foreign Function Interface).\n+\n+This document defines clear, one-to-one mappings between primitive types in C,\n+Rust (and possible other languages in the future). Its purpose is to eliminate\n+ambiguity in type widths, signedness, and binary representation across\n+platforms and languages.\n+\n+For Git, the only header required to use these unambiguous types in C is\n+`git-compat-util.h`.\n+\n+== Boolean types\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| bool^1^       | bool\n+|===\n+\n+== Integer types\n+\n+In C, `<stdint.h>` (or an equivalent) must be included.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| uint8_t    | u8\n+| uint16_t   | u16\n+| uint32_t   | u32\n+| uint64_t   | u64\n+\n+| int8_t     | i8\n+| int16_t    | i16\n+| int32_t    | i32\n+| int64_t    | i64\n+|===\n+\n+== Floating-point types\n+\n+Rust requires IEEE-754 semantics.\n+In C, that is typically true, but not guaranteed by the standard.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| float^2^      | f32\n+| double^2^     | f64\n+|===\n+\n+== Size types\n+\n+These types represent pointer-sized integers and are typically defined in\n+`<stddef.h>` or an equivalent header.\n+\n+Size types should be used any time pointer arithmetic is performed e.g.\n+indexing an array, describing the number of elements in memory, etc...\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| size_t^3^     | usize\n+| ptrdiff_t^4^  | isize\n+|===\n+\n+== Character types\n+\n+This is where C and Rust don't have a clean one-to-one mapping. A C `char` is\n+an 8-bit type that is signless (neither signed nor unsigned) which causes\n+problems with e.g. `make DEVELOPER=1`. Rust's `char` type is an unsigned 32-bit\n+integer that is used to describe Unicode code points. Even though a C `char`\n+is the same width as `u8`, `char` should be converted to u8 where it is\n+describing bytes in memory. If a C `char` is not describing bytes, then it\n+should be converted to a more accurate unambiguous type.\n+\n+While you could specify `char` in the C code and `u8` in Rust code, it's not as\n+clear what the appropriate type is, but it would work across the FFI boundary.\n+However the bigger problem comes from code generation tools like cbindgen and\n+bindgen. When cbindgen see u8 in Rust it will generate uint8_t on the C side\n+which will cause differ in signedness warnings/errors. Similaraly if bindgen\n+see `char` on the C side it will generate `std::ffi::c_char` which has its own\n+problems.\n+\n+=== Notes\n+^1^ This is only true if stdbool.h (or equivalent) is used. +\n+^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the\n+platform/arch for C does not follow IEEE-754 then this equivalence does not\n+hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n+there may be a strange platform/arch where even this isn't true. +\n+^3^ C also defines uintptr_t, but this should not be used in Git. +\n+^4^ C also defines ssize_t and intptr_t, but these should not be used in Git. +\n+\n+== Problems with std::ffi::c_* types in Rust\n+TL;DR: They're not guaranteed to match C types for all possible C\n+compilers/platforms/architectures.\n+\n+Only a few of Rust's C FFI types are considered safe and semantically clear to\n+use: +\n+\n+* `c_void`\n+* `CStr`\n+* `CString`\n+\n+Even then, they should be used sparingly, and only where the semantics match\n+exactly.\n+\n+The std::os::raw::c_* (which is deprecated) directly inherits the problems of\n+core::ffi, which changes over time and seems to make a best guess at the\n+correct definition for a given platform/target. This probably isn't a problem\n+for all platforms that Rust supports currently, but can anyone say that Rust\n+got it right for all C compilers of all platforms/targets?\n+\n+On top of all of that we're targeting an older version of Rust which doesn't\n+have the latest mappings.\n+\n+To give an example: c_long is defined in\n+footnote:[https://doc.rust-lang.org/1.63.0/src/core/ffi/mod.rs.html#175-189[c_long in 1.63.0]]\n+footnote:[https://doc.rust-lang.org/1.89.0/src/core/ffi/primitives.rs.html#135-151[c_long in 1.89.0]]\n+\n+=== Rust version 1.63.0\n+\n+[source]\n+----\n+mod c_long_definition {\n+    cfg_if! {\n+        if #[cfg(all(target_pointer_width = \"64\", not(windows)))] {\n+            pub type c_long = i64;\n+            pub type NonZero_c_long = crate::num::NonZeroI64;\n+            pub type c_ulong = u64;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU64;\n+        } else {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub type c_long = i32;\n+            pub type NonZero_c_long = crate::num::NonZeroI32;\n+            pub type c_ulong = u32;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU32;\n+        }\n+    }\n+}\n+----\n+\n+=== Rust version 1.89.0\n+\n+[source]\n+----\n+mod c_long_definition {\n+    crate::cfg_select! {\n+        any(\n+            all(target_pointer_width = \"64\", not(windows)),\n+            // wasm32 Linux ABI uses 64-bit long\n+            all(target_arch = \"wasm32\", target_os = \"linux\")\n+        ) => {\n+            pub(super) type c_long = i64;\n+            pub(super) type c_ulong = u64;\n+        }\n+        _ => {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub(super) type c_long = i32;\n+            pub(super) type c_ulong = u32;\n+        }\n+    }\n+}\n+----\n+\n+Even for the cases where C types are correctly mapped to Rust types via\n+std::ffi::c_* there are still problems. Let's take c_char for example. On some\n+platforms it's u8 on others it's i8.\n+\n+=== Subtraction underflow in debug mode\n+\n+The following code will panic in debug on platforms that define c_char as u8,\n+but won't if it's an i8.\n+\n+[source]\n+----\n+let mut x: std::ffi::c_char = 0;\n+x -= 1;\n+----\n+\n+=== Inconsistent shift behavior\n+\n+`x` will be 0xC0 for platforms that use i8, but will be 0x40 where it's u8.\n+\n+[source]\n+----\n+let mut x: std::ffi::c_char = 0x80;\n+x >>= 1;\n+----\n+\n+=== Equality fails to compile on some platforms\n+\n+The following will not compile on platforms that define c_char as i8, but will\n+if it's u8. You can cast x e.g. `assert_eq!(x as u8, b'a');`, but then you get\n+a warning on platforms that use u8 and a clean compilation where i8 is used.\n+\n+[source]\n+----\n+let mut x: std::ffi::c_char = 0x61;\n+assert_eq!(x, b'a');\n+----\n+\n+== Enum types\n+Rust enum types should not be used as FFI types. Rust enum types are more like\n+C union types than C enum's. For something like:\n+\n+[source]\n+----\n+#[repr(C, u8)]\n+enum Fruit {\n+    Apple,\n+    Banana,\n+    Cherry,\n+}\n+----\n+\n+It's easy enough to make sure the Rust enum matches what C would expect, but a\n+more complex type like.\n+\n+[source]\n+----\n+enum HashResult {\n+    SHA1([u8; 20]),\n+    SHA256([u8; 32]),\n+}\n+----\n+\n+The Rust compiler has to add a discriminant to the enum to distinguish between\n+the variants. The width, location, and values for that discriminant is up to\n+the Rust compiler and is not ABI stable.\n-- \ngitgitgadget\n\n"},{"id":"529896","messageId":"9197903add26e5b8af0bb2dd25bf115670e18e8c.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 02/10] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:40Z","receivedAt":"2025-10-29T22:19:53Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nssize_t is appropriate for dstart and dend because they both describe\npositive or negative offsets relative to a pointer.\n\nA future patch will move these fields to a different struct. Moving\nthem to the end of xdfile_t now, means the field order of xdfile_t will\nbe disturbed less.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex f145abba3e..7c8c057bca 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,10 +47,10 @@ typedef struct s_xrecord {\n typedef struct s_xdfile {\n \txrecord_t *recs;\n \tlong nrec;\n-\tlong dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n+\tptrdiff_t dstart, dend;\n } xdfile_t;\n \n typedef struct s_xdfenv {\n-- \ngitgitgadget\n\n"},{"id":"529897","messageId":"46bc1b3e25885fbd324a6428ee7ac3b5d272c4ce.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:41Z","receivedAt":"2025-10-29T22:19:55Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nRust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\nreferring to bytes in memory, rather than Unicode code points, use\nuint8_t instead of char.\n\nEvery usage of this field was inspected and cast to char*, or similar,\nto avoid signedness warnings/errors from the compiler. Casting was used\nso that the whole of xdiff doesn't need to be refactored in order to\nchange the type of this field.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     |  6 +++---\n xdiff/xmerge.c    | 14 +++++++-------\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  |  8 ++++----\n xdiff/xtypes.h    |  2 +-\n xdiff/xutils.c    |  4 ++--\n 7 files changed, 22 insertions(+), 22 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 6f3998ee54..411a8aa69f 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n \tint ret = 0;\n \n \tfor (i = 0; i < rec->size; i++) {\n-\t\tchar c = rec->ptr[i];\n+\t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n \t\t\treturn ret;\n@@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\n@@ -1008,7 +1008,7 @@ static int record_matches_regex(xrecord_t *rec, xpparam_t const *xpp) {\n \tsize_t i;\n \n \tfor (i = 0; i < xpp->ignore_regex_nr; i++)\n-\t\tif (!regexec_buf(xpp->ignore_regex[i], rec->ptr, rec->size, 1,\n+\t\tif (!regexec_buf(xpp->ignore_regex[i], (const char *)rec->ptr, rec->size, 1,\n \t\t\t\t &regmatch, 0))\n \t\t\treturn 1;\n \ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex b2f1f30cd3..ead930088a 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec(rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff(rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func(rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex fd600cbb5d..75cb3e76a2 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch(rec1[i].ptr, rec1[i].size,\n-\t\t\trec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch(rec1->ptr, rec1->size,\n-\t\t\t    rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n+\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n }\n \n /*\n@@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n \t\t * we have a very simple mmfile structure.\n \t\t */\n \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n-\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n+\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n \t\tt2.ptr = (char *)xe2->xdf2.recs[m->i2].ptr;\n-\t\tt2.size = xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n+\t\tt2.size = (char *)xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n \t\t\t+ xe2->xdf2.recs[m->i2 + m->chg2 - 1].size - t2.ptr;\n \t\tif (xdl_do_diff(&t1, &t2, xpp, &xe) < 0)\n \t\t\treturn -1;\n@@ -440,7 +440,7 @@ static int line_contains_alnum(const char *ptr, long size)\n static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n-\t\tif (line_contains_alnum(xe->xdf2.recs[i].ptr,\n+\t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n \t\t\t\txe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 669b653580..bb61354f22 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -121,7 +121,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t\treturn;\n \tmap->entries[index].line1 = line;\n \tmap->entries[index].hash = record->ha;\n-\tmap->entries[index].anchor = is_anchor(xpp, map->env->xdf1.recs[line - 1].ptr);\n+\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n \tif (map->last) {\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 192334f1b7..4cb18b2b88 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch(rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\trec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -156,8 +156,8 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = prev;\n-\t\t\tcrec->size = (long) (cur - prev);\n+\t\t\tcrec->ptr = (uint8_t const *)prev;\n+\t\t\tcrec->size =(long) ( cur - prev);\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 7c8c057bca..b1c520a378 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -39,7 +39,7 @@ typedef struct s_chastore {\n } chastore_t;\n \n typedef struct s_xrecord {\n-\tchar const *ptr;\n+\tuint8_t const *ptr;\n \tlong size;\n \tunsigned long ha;\n } xrecord_t;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 447e66c719..7be063bfb6 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -465,10 +465,10 @@ int xdl_fall_back_diff(xdfenv_t *diff_env, xpparam_t const *xpp,\n \txdfenv_t env;\n \n \tsubfile1.ptr = (char *)diff_env->xdf1.recs[line1 - 1].ptr;\n-\tsubfile1.size = diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n+\tsubfile1.size = (char *)diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n \t\tdiff_env->xdf1.recs[line1 + count1 - 2].size - subfile1.ptr;\n \tsubfile2.ptr = (char *)diff_env->xdf2.recs[line2 - 1].ptr;\n-\tsubfile2.size = diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n+\tsubfile2.size = (char *)diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n \t\tdiff_env->xdf2.recs[line2 + count2 - 2].size - subfile2.ptr;\n \tif (xdl_do_diff(&subfile1, &subfile2, xpp, &env) < 0)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"529898","messageId":"07e28aad3b5dc453967456b0017ae4751c9275bc.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 04/10] xdiff: use size_t for xrecord_t.size","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:42Z","receivedAt":"2025-10-29T22:19:56Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is the appropriate type because size is describing the number of\nelements, bytes in this case, in memory.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  7 +++----\n xdiff/xemit.c    |  8 ++++----\n xdiff/xmerge.c   | 16 ++++++++--------\n xdiff/xprepare.c |  6 +++---\n xdiff/xtypes.h   |  2 +-\n 5 files changed, 19 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 411a8aa69f..edd05466df 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -403,10 +403,9 @@ static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n  */\n static int get_indent(xrecord_t *rec)\n {\n-\tlong i;\n \tint ret = 0;\n \n-\tfor (i = 0; i < rec->size; i++) {\n+\tfor (size_t i = 0; i < rec->size; i++) {\n \t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n@@ -993,11 +992,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex ead930088a..2f8007753c 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, (long)rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, (long)rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, (long)rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\n@@ -151,7 +151,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n static int is_empty_rec(xdfile_t *xdf, long ri)\n {\n \txrecord_t *rec = &xdf->recs[ri];\n-\tlong i = 0;\n+\tsize_t i = 0;\n \n \tfor (; i < rec->size && XDL_ISSPACE(rec->ptr[i]); i++);\n \ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 75cb3e76a2..0dd4558a32 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n-\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, (long)rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, (long)rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -119,11 +119,11 @@ static int xdl_recs_copy_0(int use_orig, xdfenv_t *xe, int i, int count, int nee\n \tif (count < 1)\n \t\treturn 0;\n \n-\tfor (i = 0; i < count; size += recs[i++].size)\n+\tfor (i = 0; i < count; size += (int)recs[i++].size)\n \t\tif (dest)\n \t\t\tmemcpy(dest + size, recs[i].ptr, recs[i].size);\n \tif (add_nl) {\n-\t\ti = recs[count - 1].size;\n+\t\ti = (int)recs[count - 1].size;\n \t\tif (i == 0 || recs[count - 1].ptr[i - 1] != '\\n') {\n \t\t\tif (needs_cr) {\n \t\t\t\tif (dest)\n@@ -156,7 +156,7 @@ static int xdl_orig_copy(xdfenv_t *xe, int i, int count, int needs_cr, int add_n\n  */\n static int is_eol_crlf(xdfile_t *file, int i)\n {\n-\tlong size;\n+\tsize_t size;\n \n \tif (i < file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n-\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, (long)rec1->size,\n+\t\t\t    (const char *)rec2->ptr, (long)rec2->size, flags);\n }\n \n /*\n@@ -441,7 +441,7 @@ static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n \t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n-\t\t\t\txe->xdf2.recs[i].size))\n+\t\t\t\t(long)xe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 4cb18b2b88..b3219aed3e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -157,7 +157,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = (uint8_t const *)prev;\n-\t\t\tcrec->size =(long) ( cur - prev);\n+\t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex b1c520a378..88b1fe4649 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -40,7 +40,7 @@ typedef struct s_chastore {\n \n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n-\tlong size;\n+\tsize_t size;\n \tunsigned long ha;\n } xrecord_t;\n \n-- \ngitgitgadget\n\n"},{"id":"529899","messageId":"1ade7d8165406bab37007c73627e72f9b9143773.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 05/10] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:43Z","receivedAt":"2025-10-29T22:19:58Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nConvert the function signature and body to use unambiguous types. char\nis changed to uint8_t because this function processes bytes in memory.\nunsigned long to uint64_t so that the hash output is consistent across\nplatforms. `flags` was changed from long to uint64_t to ensure the\nhigh order bits are not dropped on platforms that treat long as 32\nbits.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff-interface.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xutils.c    | 28 ++++++++++++++--------------\n xdiff/xutils.h    |  6 +++---\n 4 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff-interface.c b/xdiff-interface.c\nindex 4971f722b3..1a35556380 100644\n--- a/xdiff-interface.c\n+++ b/xdiff-interface.c\n@@ -300,7 +300,7 @@ void xdiff_clear_find_func(xdemitconf_t *xecfg)\n \n unsigned long xdiff_hash_string(const char *s, size_t len, long flags)\n {\n-\treturn xdl_hash_record(&s, s + len, flags);\n+\treturn xdl_hash_record((uint8_t const**)&s, (uint8_t const*)s + len, flags);\n }\n \n int xdiff_compare_lines(const char *l1, long s1,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex b3219aed3e..85e56021da 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -137,8 +137,8 @@ static void xdl_free_ctx(xdfile_t *xdf)\n static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_t const *xpp,\n \t\t\t   xdlclassifier_t *cf, xdfile_t *xdf) {\n \tlong bsize;\n-\tunsigned long hav;\n-\tchar const *blk, *cur, *top, *prev;\n+\tuint64_t hav;\n+\tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n \txdf->rindex = NULL;\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = (uint8_t const *)prev;\n+\t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 7be063bfb6..77ee1ad9c8 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -249,11 +249,11 @@ int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags)\n \treturn 1;\n }\n \n-unsigned long xdl_hash_record_with_whitespace(char const **data,\n-\t\tchar const *top, long flags) {\n-\tunsigned long ha = 5381;\n-\tchar const *ptr = *data;\n-\tint cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data,\n+\t\tuint8_t const *top, uint64_t flags) {\n+\tuint64_t ha = 5381;\n+\tuint8_t const *ptr = *data;\n+\tbool cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n \n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tif (cr_at_eol_only) {\n@@ -263,8 +263,8 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\t\tcontinue;\n \t\t}\n \t\telse if (XDL_ISSPACE(*ptr)) {\n-\t\t\tconst char *ptr2 = ptr;\n-\t\t\tint at_eol;\n+\t\t\tconst uint8_t *ptr2 = ptr;\n+\t\t\tbool at_eol;\n \t\t\twhile (ptr + 1 < top && XDL_ISSPACE(ptr[1])\n \t\t\t\t\t&& ptr[1] != '\\n')\n \t\t\t\tptr++;\n@@ -274,20 +274,20 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_CHANGE\n \t\t\t\t && !at_eol) {\n \t\t\t\tha += (ha << 5);\n-\t\t\t\tha ^= (unsigned long) ' ';\n+\t\t\t\tha ^= (uint64_t) ' ';\n \t\t\t}\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_AT_EOL\n \t\t\t\t && !at_eol) {\n \t\t\t\twhile (ptr2 != ptr + 1) {\n \t\t\t\t\tha += (ha << 5);\n-\t\t\t\t\tha ^= (unsigned long) *ptr2;\n+\t\t\t\t\tha ^= (uint64_t) *ptr2;\n \t\t\t\t\tptr2++;\n \t\t\t\t}\n \t\t\t}\n \t\t\tcontinue;\n \t\t}\n \t\tha += (ha << 5);\n-\t\tha ^= (unsigned long) *ptr;\n+\t\tha ^= (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n \n@@ -304,9 +304,9 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n #define REASSOC_FENCE(x, y)\n #endif\n \n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n-\tunsigned long ha = 5381, c0, c1;\n-\tchar const *ptr = *data;\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top) {\n+\tuint64_t ha = 5381, c0, c1;\n+\tuint8_t const *ptr = *data;\n #if 0\n \t/*\n \t * The baseline form of the optimized loop below. This is the djb2\n@@ -314,7 +314,7 @@ unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n \t */\n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tha += (ha << 5);\n-\t\tha += (unsigned long) *ptr;\n+\t\tha += (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n #else\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 13f6831047..615b4a9d35 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -34,9 +34,9 @@ void *xdl_cha_alloc(chastore_t *cha);\n long xdl_guess_lines(mmfile_t *mf, long sample);\n int xdl_blankline(const char *line, long size, long flags);\n int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags);\n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top);\n-unsigned long xdl_hash_record_with_whitespace(char const **data, char const *top, long flags);\n-static inline unsigned long xdl_hash_record(char const **data, char const *top, long flags)\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top);\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data, uint8_t const *top, uint64_t flags);\n+static inline uint64_t xdl_hash_record(uint8_t const **data, uint8_t const *top, uint64_t flags)\n {\n \tif (flags & XDF_WHITESPACE_FLAGS)\n \t\treturn xdl_hash_record_with_whitespace(data, top, flags);\n-- \ngitgitgadget\n\n"},{"id":"529900","messageId":"59054ea0cb65718dbac500d342bc960bdb5066c1.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:44Z","receivedAt":"2025-10-29T22:20:00Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe ha field is serving two different purposes, which makes the code\nharder to read. At first glance it looks like many places assume\nthere could never be hash collisions between lines of the two input\nfiles. In reality, line_hash is used together with xdl_recmatch() to\nensure correct comparisons of lines, even when collisions occur.\n\nTo make this clearer, the old ha field has been split:\n  * line_hash: The straightforward hash of a line, requiring no\n    additional context.\n  * minimal_perfect_hash: Not a new concept, but now a separate\n    field. It comes from the classifier's general-purpose hash table,\n    which assigns each line a unique and minimal hash across the two\n    files.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c     |  6 +++---\n xdiff/xhistogram.c |  4 ++--\n xdiff/xpatience.c  | 10 +++++-----\n xdiff/xprepare.c   | 16 ++++++++--------\n xdiff/xtypes.h     |  3 ++-\n 5 files changed, 20 insertions(+), 19 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex edd05466df..436c34697d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -22,9 +22,9 @@\n \n #include \"xinclude.h\"\n \n-static unsigned long get_hash(xdfile_t *xdf, long index)\n+static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].ha;\n+\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -385,7 +385,7 @@ static xdchange_t *xdl_add_change(xdchange_t *xscr, long i1, long i2, long chg1,\n \n static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n {\n-\treturn (rec1->ha == rec2->ha);\n+\treturn rec1->minimal_perfect_hash == rec2->minimal_perfect_hash;\n }\n \n /*\ndiff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\nindex 6dc450b1fe..5ae1282c27 100644\n--- a/xdiff/xhistogram.c\n+++ b/xdiff/xhistogram.c\n@@ -90,7 +90,7 @@ struct region {\n \n static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n {\n-\treturn r1->ha == r2->ha;\n+\treturn r1->minimal_perfect_hash == r2->minimal_perfect_hash;\n \n }\n \n@@ -98,7 +98,7 @@ static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n \t(cmp_recs(REC(i->env, s1, l1), REC(i->env, s2, l2)))\n \n #define TABLE_HASH(index, side, line) \\\n-\tXDL_HASHLONG((REC(index->env, side, line))->ha, index->table_bits)\n+\tXDL_HASHLONG((REC(index->env, side, line))->minimal_perfect_hash, index->table_bits)\n \n static int scanA(struct histindex *index, int line1, int count1)\n {\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex bb61354f22..cc53266f3b 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -48,7 +48,7 @@\n struct hashmap {\n \tint nr, alloc;\n \tstruct entry {\n-\t\tunsigned long hash;\n+\t\tsize_t minimal_perfect_hash;\n \t\t/*\n \t\t * 0 = unused entry, 1 = first line, 2 = second, etc.\n \t\t * line2 is NON_UNIQUE if the line is not unique\n@@ -101,10 +101,10 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t * So we multiply ha by 2 in the hope that the hashing was\n \t * \"unique enough\".\n \t */\n-\tint index = (int)((record->ha << 1) % map->alloc);\n+\tint index = (int)((record->minimal_perfect_hash << 1) % map->alloc);\n \n \twhile (map->entries[index].line1) {\n-\t\tif (map->entries[index].hash != record->ha) {\n+\t\tif (map->entries[index].minimal_perfect_hash != record->minimal_perfect_hash) {\n \t\t\tif (++index >= map->alloc)\n \t\t\t\tindex = 0;\n \t\t\tcontinue;\n@@ -120,7 +120,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \tif (pass == 2)\n \t\treturn;\n \tmap->entries[index].line1 = line;\n-\tmap->entries[index].hash = record->ha;\n+\tmap->entries[index].minimal_perfect_hash = record->minimal_perfect_hash;\n \tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n@@ -248,7 +248,7 @@ static int match(struct hashmap *map, int line1, int line2)\n {\n \txrecord_t *record1 = &map->env->xdf1.recs[line1 - 1];\n \txrecord_t *record2 = &map->env->xdf2.recs[line2 - 1];\n-\treturn record1->ha == record2->ha;\n+\treturn record1->minimal_perfect_hash == record2->minimal_perfect_hash;\n }\n \n static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 85e56021da..16236bd045 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -96,9 +96,9 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \tlong hi;\n \txdlclass_t *rcrec;\n \n-\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n+\thi = (long) XDL_HASHLONG(rec->line_hash, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n-\t\tif (rcrec->rec.ha == rec->ha &&\n+\t\tif (rcrec->rec.line_hash == rec->line_hash &&\n \t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n \t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n@@ -120,7 +120,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n \t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n \n-\trec->ha = (unsigned long) rcrec->idx;\n+\trec->minimal_perfect_hash = (size_t)rcrec->idx;\n \n \treturn 0;\n }\n@@ -158,7 +158,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n-\t\t\tcrec->ha = hav;\n+\t\t\tcrec->line_hash = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\n \t\t}\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -298,7 +298,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -350,7 +350,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs2 = xdf2->recs;\n \tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dstart = xdf2->dstart = i;\n@@ -358,7 +358,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs1 = xdf1->recs + xdf1->nrec - 1;\n \trecs2 = xdf2->recs + xdf2->nrec - 1;\n \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dend = xdf1->nrec - i - 1;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 88b1fe4649..742b81bf3b 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -41,7 +41,8 @@ typedef struct s_chastore {\n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n \tsize_t size;\n-\tunsigned long ha;\n+\tuint64_t line_hash;\n+\tsize_t minimal_perfect_hash;\n } xrecord_t;\n \n typedef struct s_xdfile {\n-- \ngitgitgadget\n\n"},{"id":"529901","messageId":"f91be17858ab39292ab6667636d9573da937d248.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 07/10] xdiff: make xdfile_t.nrec a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:45Z","receivedAt":"2025-10-29T22:20:01Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nrec describes the number of elements in memory\nfor recs, and the number of elements in memory for 'changed' + 2.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     | 20 ++++++++++----------\n xdiff/xmerge.c    |  8 ++++----\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  | 12 ++++++------\n xdiff/xtypes.h    |  2 +-\n 6 files changed, 26 insertions(+), 26 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 436c34697d..759193fe5d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -483,7 +483,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n {\n \tlong i;\n \n-\tif (split >= xdf->nrec) {\n+\tif (split >= (long)xdf->nrec) {\n \t\tm->end_of_file = 1;\n \t\tm->indent = -1;\n \t} else {\n@@ -506,7 +506,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n \n \tm->post_blank = 0;\n \tm->post_indent = -1;\n-\tfor (i = split + 1; i < xdf->nrec; i++) {\n+\tfor (i = split + 1; i < (long)xdf->nrec; i++) {\n \t\tm->post_indent = get_indent(&xdf->recs[i]);\n \t\tif (m->post_indent != -1)\n \t\t\tbreak;\n@@ -717,7 +717,7 @@ static void group_init(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static inline int group_next(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end == xdf->nrec)\n+\tif (g->end == (long)xdf->nrec)\n \t\treturn -1;\n \n \tg->start = g->end + 1;\n@@ -750,7 +750,7 @@ static inline int group_previous(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static int group_slide_down(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end < xdf->nrec &&\n+\tif (g->end < (long)xdf->nrec &&\n \t    recs_match(&xdf->recs[g->start], &xdf->recs[g->end])) {\n \t\txdf->changed[g->start++] = false;\n \t\txdf->changed[g->end++] = true;\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex 2f8007753c..04f7e9193b 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -137,7 +137,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n \tbuf = func_line ? func_line->buf : dummy;\n \tsize = func_line ? sizeof(func_line->buf) : sizeof(dummy);\n \n-\tfor (l = start; l != limit && 0 <= l && l < xe->xdf1.nrec; l += step) {\n+\tfor (l = start; l != limit && 0 <= l && l < (long)xe->xdf1.nrec; l += step) {\n \t\tlong len = match_func_rec(&xe->xdf1, xecfg, l, buf, size);\n \t\tif (len >= 0) {\n \t\t\tif (func_line)\n@@ -179,14 +179,14 @@ pre_context_calculation:\n \t\t\tlong fs1, i1 = xch->i1;\n \n \t\t\t/* Appended chunk? */\n-\t\t\tif (i1 >= xe->xdf1.nrec) {\n+\t\t\tif (i1 >= (long)xe->xdf1.nrec) {\n \t\t\t\tlong i2 = xch->i2;\n \n \t\t\t\t/*\n \t\t\t\t * We don't need additional context if\n \t\t\t\t * a whole function was added.\n \t\t\t\t */\n-\t\t\t\twhile (i2 < xe->xdf2.nrec) {\n+\t\t\t\twhile (i2 < (long)xe->xdf2.nrec) {\n \t\t\t\t\tif (is_func_rec(&xe->xdf2, xecfg, i2))\n \t\t\t\t\t\tgoto post_context_calculation;\n \t\t\t\t\ti2++;\n@@ -196,7 +196,7 @@ pre_context_calculation:\n \t\t\t\t * Otherwise get more context from the\n \t\t\t\t * pre-image.\n \t\t\t\t */\n-\t\t\t\ti1 = xe->xdf1.nrec - 1;\n+\t\t\t\ti1 = (long)xe->xdf1.nrec - 1;\n \t\t\t}\n \n \t\t\tfs1 = get_func_line(xe, xecfg, NULL, i1, -1);\n@@ -228,8 +228,8 @@ pre_context_calculation:\n \n  post_context_calculation:\n \t\tlctx = xecfg->ctxlen;\n-\t\tlctx = XDL_MIN(lctx, xe->xdf1.nrec - (xche->i1 + xche->chg1));\n-\t\tlctx = XDL_MIN(lctx, xe->xdf2.nrec - (xche->i2 + xche->chg2));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf1.nrec - (xche->i1 + xche->chg1));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf2.nrec - (xche->i2 + xche->chg2));\n \n \t\te1 = xche->i1 + xche->chg1 + lctx;\n \t\te2 = xche->i2 + xche->chg2 + lctx;\n@@ -237,13 +237,13 @@ pre_context_calculation:\n \t\tif (xecfg->flags & XDL_EMIT_FUNCCONTEXT) {\n \t\t\tlong fe1 = get_func_line(xe, xecfg, NULL,\n \t\t\t\t\t\t xche->i1 + xche->chg1,\n-\t\t\t\t\t\t xe->xdf1.nrec);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec);\n \t\t\twhile (fe1 > 0 && is_empty_rec(&xe->xdf1, fe1 - 1))\n \t\t\t\tfe1--;\n \t\t\tif (fe1 < 0)\n-\t\t\t\tfe1 = xe->xdf1.nrec;\n+\t\t\t\tfe1 = (long)xe->xdf1.nrec;\n \t\t\tif (fe1 > e1) {\n-\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), xe->xdf2.nrec);\n+\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), (long)xe->xdf2.nrec);\n \t\t\t\te1 = fe1;\n \t\t\t}\n \n@@ -254,7 +254,7 @@ pre_context_calculation:\n \t\t\t */\n \t\t\tif (xche->next) {\n \t\t\t\tlong l = XDL_MIN(xche->next->i1,\n-\t\t\t\t\t\t xe->xdf1.nrec - 1);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec - 1);\n \t\t\t\tif (l - xecfg->ctxlen <= e1 ||\n \t\t\t\t    get_func_line(xe, xecfg, NULL, l, e1) < 0) {\n \t\t\t\t\txche = xche->next;\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 0dd4558a32..29dad98c49 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -158,7 +158,7 @@ static int is_eol_crlf(xdfile_t *file, int i)\n {\n \tsize_t size;\n \n-\tif (i < file->nrec - 1)\n+\tif (i < (long)file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n \t\treturn (size = file->recs[i].size) > 1 &&\n \t\t\tfile->recs[i].ptr[size - 2] == '\\r';\n@@ -317,7 +317,7 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \t\t\tcontinue;\n \t\ti = m->i1 + m->chg1;\n \t}\n-\tsize += xdl_recs_copy(xe1, i, xe1->xdf2.nrec - i, 0, 0,\n+\tsize += xdl_recs_copy(xe1, i, (int)xe1->xdf2.nrec - i, 0, 0,\n \t\t\t      dest ? dest + size : NULL);\n \treturn size;\n }\n@@ -622,7 +622,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\t\tchanges = c;\n \t\ti0 = xscr1->i1;\n \t\ti1 = xscr1->i2;\n-\t\ti2 = xscr1->i1 + xe2->xdf2.nrec - xe2->xdf1.nrec;\n+\t\ti2 = xscr1->i1 + (long)xe2->xdf2.nrec - (long)xe2->xdf1.nrec;\n \t\tchg0 = xscr1->chg1;\n \t\tchg1 = xscr1->chg2;\n \t\tchg2 = xscr1->chg1;\n@@ -637,7 +637,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\tif (!changes)\n \t\t\tchanges = c;\n \t\ti0 = xscr2->i1;\n-\t\ti1 = xscr2->i1 + xe1->xdf2.nrec - xe1->xdf1.nrec;\n+\t\ti1 = xscr2->i1 + (long)xe1->xdf2.nrec - (long)xe1->xdf1.nrec;\n \t\ti2 = xscr2->i2;\n \t\tchg0 = xscr2->chg1;\n \t\tchg1 = xscr2->chg1;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex cc53266f3b..a0b31eb5d8 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -370,5 +370,5 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n \n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n-\treturn patience_diff(xpp, env, 1, env->xdf1.nrec, 1, env->xdf2.nrec);\n+\treturn patience_diff(xpp, env, 1, (int)env->xdf1.nrec, 1, (int)env->xdf2.nrec);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 16236bd045..4ee9fb60cd 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -153,7 +153,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\tfor (top = blk + bsize; cur < top; ) {\n \t\t\tprev = cur;\n \t\t\thav = xdl_hash_record(&cur, top, xpp->flags);\n-\t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n+\t\t\tif (XDL_ALLOC_GROW(xdf->recs, (long)xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n@@ -287,7 +287,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -295,7 +295,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -348,7 +348,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \n \trecs1 = xdf1->recs;\n \trecs2 = xdf2->recs;\n-\tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n+\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n@@ -361,8 +361,8 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n-\txdf1->dend = xdf1->nrec - i - 1;\n-\txdf2->dend = xdf2->nrec - i - 1;\n+\txdf1->dend = (long)xdf1->nrec - i - 1;\n+\txdf2->dend = (long)xdf2->nrec - i - 1;\n \n \treturn 0;\n }\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 742b81bf3b..17cafd8b6e 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,7 +47,7 @@ typedef struct s_xrecord {\n \n typedef struct s_xdfile {\n \txrecord_t *recs;\n-\tlong nrec;\n+\tsize_t nrec;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n-- \ngitgitgadget\n\n"},{"id":"529902","messageId":"e2a6a23cc473108f5a79aed88eb2b4df6661c612.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 08/10] xdiff: make xdfile_t.nreff a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:46Z","receivedAt":"2025-10-29T22:20:02Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nreff describes the number of elements in memory\nfor rindex.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 +++++++-------\n xdiff/xtypes.h   |  2 +-\n 2 files changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 4ee9fb60cd..c690bafeb1 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -264,7 +264,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, nreff, mlim;\n+\tlong i, nm, mlim;\n \txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n@@ -307,29 +307,29 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * Use temporary arrays to decide if changed[i] should remain\n \t * false, or become true.\n \t */\n-\tfor (nreff = 0, i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n+\txdf1->nreff = 0;\n+\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[nreff++] = i;\n+\t\t\txdf1->rindex[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf1->nreff = nreff;\n \n-\tfor (nreff = 0, i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n+\txdf2->nreff = 0;\n+\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[nreff++] = i;\n+\t\t\txdf2->rindex[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf2->nreff = nreff;\n \n cleanup:\n \txdl_free(action1);\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 17cafd8b6e..df4c5cab1a 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -50,7 +50,7 @@ typedef struct s_xdfile {\n \tsize_t nrec;\n \tbool *changed;\n \tlong *rindex;\n-\tlong nreff;\n+\tsize_t nreff;\n \tptrdiff_t dstart, dend;\n } xdfile_t;\n \n-- \ngitgitgadget\n\n"},{"id":"529903","messageId":"3b6054945f200def2b8ce77867f34b096bdfd0ab.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 09/10] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:47Z","receivedAt":"2025-10-29T22:20:03Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nrindex describes a index offset which means it's an index into memory\nwhich should use size_t.\n\nChanging the type of rindex from long to size_t has no cascading\nrefactor impact because it is only ever used to directly index other\narrays.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex df4c5cab1a..3bcc0920e0 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -49,7 +49,7 @@ typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n \tbool *changed;\n-\tlong *rindex;\n+\tsize_t *rindex;\n \tsize_t nreff;\n \tptrdiff_t dstart, dend;\n } xdfile_t;\n-- \ngitgitgadget\n\n"},{"id":"529904","messageId":"1856a29026d8c3d824723e253dc68b052a5d8b9a.1761776388.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v2 10/10] xdiff: rename rindex -> reference_index","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-10-29T22:19:48Z","receivedAt":"2025-10-29T22:20:04Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe classic diff adds only the lines that it's going to consider,\nduring the diff, to an array. A mapping between the compacted\narray, and the lines of the file that they reference, are\nfacilitated by this array.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  6 +++---\n xdiff/xprepare.c | 10 +++++-----\n xdiff/xtypes.h   |  2 +-\n 3 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 759193fe5d..8eb664be3e 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -24,7 +24,7 @@\n \n static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n+\treturn xdf->recs[xdf->reference_index[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -278,10 +278,10 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n \t */\n \tif (off1 == lim1) {\n \t\tfor (; off2 < lim2; off2++)\n-\t\t\txdf2->changed[xdf2->rindex[off2]] = true;\n+\t\t\txdf2->changed[xdf2->reference_index[off2]] = true;\n \t} else if (off2 == lim2) {\n \t\tfor (; off1 < lim1; off1++)\n-\t\t\txdf1->changed[xdf1->rindex[off1]] = true;\n+\t\t\txdf1->changed[xdf1->reference_index[off1]] = true;\n \t} else {\n \t\txdpsplit_t spl;\n \t\tspl.i1 = spl.i2 = 0;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex c690bafeb1..1dd420a2ff 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -128,7 +128,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n static void xdl_free_ctx(xdfile_t *xdf)\n {\n-\txdl_free(xdf->rindex);\n+\txdl_free(xdf->reference_index);\n \txdl_free(xdf->changed - 1);\n \txdl_free(xdf->recs);\n }\n@@ -141,7 +141,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n-\txdf->rindex = NULL;\n+\txdf->reference_index = NULL;\n \txdf->changed = NULL;\n \txdf->recs = NULL;\n \n@@ -169,7 +169,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF)) {\n-\t\tif (!XDL_ALLOC_ARRAY(xdf->rindex, xdf->nrec + 1))\n+\t\tif (!XDL_ALLOC_ARRAY(xdf->reference_index, xdf->nrec + 1))\n \t\t\tgoto abort;\n \t}\n \n@@ -312,7 +312,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[xdf1->nreff++] = i;\n+\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n@@ -324,7 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[xdf2->nreff++] = i;\n+\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 3bcc0920e0..5accbec284 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -49,7 +49,7 @@ typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n \tbool *changed;\n-\tsize_t *rindex;\n+\tsize_t *reference_index;\n \tsize_t nreff;\n \tptrdiff_t dstart, dend;\n } xdfile_t;\n-- \ngitgitgadget\n"},{"id":"529967","messageId":"xmqqcy6479e5.fsf@gitster.g","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 00/10] Xdiff cleanup part2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-10-30T14:26:58Z","receivedAt":"2025-10-30T14:27:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n>  * Added documentation about unambiguous types and FFI\n\nNicely written; a few footnote entries may be a bit too strict,\nmisleading, and may need rephrasing, though.  For example, we may\nwant to be suspicious when we see code that uses ssize_t as if it is\nhalf the size_t plus error indication, it does not immediately mean\nthat the type \"should not be used in Git\". It is perfectly sensible\nto assign to or compare with returned value from write(2), for\nexample.\n\nWill queue.  Thanks.\n\n"},{"id":"530303","messageId":"995f77a3-b94c-46df-87d3-22c7b2a3c762@gmail.com","threadId":"64326","inReplyTo":"88133848d1a317f8a95c19ee5482b828a3f8705f.1761776388.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-11-06T09:55:31Z","receivedAt":"2025-11-06T09:55:37Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ezekiel\n\nOn 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Document other nuances with crossing the FFI boundary. Other language\n> mappings may be added in the future.\n\nThanks for adding this, I've left a few comments below. Overall I \nthought it was very well written. I tried building an html version of \nthis but even after adding it to the list of TECH_DOCS in \nDocumentation/Makefile with\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 47208269a2e..2699f0b24af 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -143,6 +143,7 @@ TECH_DOCS += technical/shallow\n  TECH_DOCS += technical/sparse-checkout\n  TECH_DOCS += technical/sparse-index\n  TECH_DOCS += technical/trivial-merge\n+TECH_DOCS += technical/unambiguous-types\n  TECH_DOCS += technical/unit-tests\n  SP_ARTICLES += $(TECH_DOCS)\n  SP_ARTICLES += technical/api-index\n\nit fails with\n\n$ make -C Documentation/ technical/unambiguous-types.html \n                                       Merge branch \n'ps/object-source-loose' into seen\nmake: Entering directory '/home/phil/src/git/Documentation'\n     GEN asciidoc.conf\n     * new asciidoc flags\n     ASCIIDOC technical/unambiguous-types.html\nasciidoc: ERROR: unambiguous-types.adoc: line 139: undefined filter \nattribute in command: source-highlight --gen-version -f xhtml -s \n{language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} \n{args=}\nasciidoc: ERROR: unambiguous-types.adoc: line 162: undefined filter \nattribute in command: source-highlight --gen-version -f xhtml -s \n{language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} \n{args=}\nasciidoc: ERROR: unambiguous-types.adoc: line 177: undefined filter \nattribute in command: source-highlight --gen-version -f xhtml -s \n{language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} \n{args=}\nasciidoc: ERROR: unambiguous-types.adoc: line 187: undefined filter \nattribute in command: source-highlight --gen-version -f xhtml -s \n{language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} \n{args=}\nasciidoc: ERROR: unambiguous-types.adoc: line 199: undefined filter \nattribute in command: source-highlight --gen-version -f xhtml -s \n{language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} \n{args=}\nasciidoc: ERROR: unambiguous-types.adoc: line 213: undefined filter \nattribute in command: source-highlight --gen-version -f xhtml -s \n{language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} \n{args=}\nasciidoc: ERROR: unambiguous-types.adoc: line 224: undefined filter \nattribute in command: source-highlight --gen-version -f xhtml -s \n{language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} \n{args=}\nmake: *** [Makefile:396: technical/unambiguous-types.html] Error 1\nmake: *** Deleting file 'technical/unambiguous-types.html'\nmake: Leaving directory '/home/phil/src/git/Documentation'\n\n> +== Character types\n> +\n> +This is where C and Rust don't have a clean one-to-one mapping. A C `char` is\n> +an 8-bit type that is signless (neither signed nor unsigned) \n\nI found this a bit confusing. Isn't the signedness of \"char\" \nimplementation defined rather than it being \"signless\"\n\n> which causes\n> +problems with e.g. `make DEVELOPER=1`.\n\nI'm not sure what this is referring to - maybe -Wsign-compare?\n\n> Rust's `char` type is an unsigned 32-bit\n> +integer that is used to describe Unicode code points. Even though a C `char`\n> +is the same width as `u8`, `char` should be converted to u8 where it is\n> +describing bytes in memory. \n\nI'm dreading the point where we start sharing \"struct strbuf\" with rust \nand have to change the \"buf\" member from \"char*\" to \"uint8_t*\". While it \nis not used in the xdiff code it is ubiquitous everywhere else and there \nare lots of places where be pass the \"buf\" member to functions expecting \na \"char*\".\n\n\tgit grep -E '(\\.|->)buf\\W'\n\nhas over 4000 matches\n\n> If a C `char` is not describing bytes, then it\n> +should be converted to a more accurate unambiguous type.\n\nThat's a good point.\n\n> +While you could specify `char` in the C code and `u8` in Rust code, it's not as\n> +clear what the appropriate type is, but it would work across the FFI boundary.\n> +However the bigger problem comes from code generation tools like cbindgen and\n> +bindgen. When cbindgen see u8 in Rust it will generate uint8_t on the C side\n> +which will cause differ in signedness warnings/errors. Similaraly if bindgen\n> +see `char` on the C side it will generate `std::ffi::c_char` which has its own\n> +problems.\n\nYeah, we definitely don't want to be using \"std::ffi::c_char\" in our \nrust implementations. I do wonder if we might want to use it (or CStr) \njudiciously in function parameters and immediately convert it to u8 in \nthe function body where the function is called from C though.\n\n> +=== Notes\n> +^1^ This is only true if stdbool.h (or equivalent) is used. +\n> +^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the\n> +platform/arch for C does not follow IEEE-754 then this equivalence does not\n> +hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n> +there may be a strange platform/arch where even this isn't true. +\n> +^3^ C also defines uintptr_t, but this should not be used in Git. +\n> +^4^ C also defines ssize_t and intptr_t, but these should not be used in Git. +\n\n[u]intptr_t and ssize_t are used in git already. As Junio has pointed \nout there are sane uses for these types but we don't want to use them in \nstructs or function parameters where the struct or function is shared \nwith rust.\n\n> +\n> +== Problems with std::ffi::c_* types in Rust\n> +TL;DR: They're not guaranteed to match C types for all possible C\n> +compilers/platforms/architectures.\n\nIs this official policy of the rust project?\n\nThanks\n\nPhillip\n\n> +Only a few of Rust's C FFI types are considered safe and semantically clear to\n> +use: +\n> +\n> +* `c_void`\n> +* `CStr`\n> +* `CString`\n> +\n> +Even then, they should be used sparingly, and only where the semantics match\n> +exactly.\n> +\n> +The std::os::raw::c_* (which is deprecated) directly inherits the problems of\n> +core::ffi, which changes over time and seems to make a best guess at the\n> +correct definition for a given platform/target. This probably isn't a problem\n> +for all platforms that Rust supports currently, but can anyone say that Rust\n> +got it right for all C compilers of all platforms/targets?\n> +\n> +On top of all of that we're targeting an older version of Rust which doesn't\n> +have the latest mappings.\n> +\n> +To give an example: c_long is defined in\n> +footnote:[https://doc.rust-lang.org/1.63.0/src/core/ffi/mod.rs.html#175-189[c_long in 1.63.0]]\n> +footnote:[https://doc.rust-lang.org/1.89.0/src/core/ffi/primitives.rs.html#135-151[c_long in 1.89.0]]\n> +\n> +=== Rust version 1.63.0\n> +\n> +[source]\n> +----\n> +mod c_long_definition {\n> +    cfg_if! {\n> +        if #[cfg(all(target_pointer_width = \"64\", not(windows)))] {\n> +            pub type c_long = i64;\n> +            pub type NonZero_c_long = crate::num::NonZeroI64;\n> +            pub type c_ulong = u64;\n> +            pub type NonZero_c_ulong = crate::num::NonZeroU64;\n> +        } else {\n> +            // The minimal size of `long` in the C standard is 32 bits\n> +            pub type c_long = i32;\n> +            pub type NonZero_c_long = crate::num::NonZeroI32;\n> +            pub type c_ulong = u32;\n> +            pub type NonZero_c_ulong = crate::num::NonZeroU32;\n> +        }\n> +    }\n> +}\n> +----\n> +\n> +=== Rust version 1.89.0\n> +\n> +[source]\n> +----\n> +mod c_long_definition {\n> +    crate::cfg_select! {\n> +        any(\n> +            all(target_pointer_width = \"64\", not(windows)),\n> +            // wasm32 Linux ABI uses 64-bit long\n> +            all(target_arch = \"wasm32\", target_os = \"linux\")\n> +        ) => {\n> +            pub(super) type c_long = i64;\n> +            pub(super) type c_ulong = u64;\n> +        }\n> +        _ => {\n> +            // The minimal size of `long` in the C standard is 32 bits\n> +            pub(super) type c_long = i32;\n> +            pub(super) type c_ulong = u32;\n> +        }\n> +    }\n> +}\n> +----\n> +\n> +Even for the cases where C types are correctly mapped to Rust types via\n> +std::ffi::c_* there are still problems. Let's take c_char for example. On some\n> +platforms it's u8 on others it's i8.\n> +\n> +=== Subtraction underflow in debug mode\n> +\n> +The following code will panic in debug on platforms that define c_char as u8,\n> +but won't if it's an i8.\n> +\n> +[source]\n> +----\n> +let mut x: std::ffi::c_char = 0;\n> +x -= 1;\n> +----\n> +\n> +=== Inconsistent shift behavior\n> +\n> +`x` will be 0xC0 for platforms that use i8, but will be 0x40 where it's u8.\n> +\n> +[source]\n> +----\n> +let mut x: std::ffi::c_char = 0x80;\n> +x >>= 1;\n> +----\n> +\n> +=== Equality fails to compile on some platforms\n> +\n> +The following will not compile on platforms that define c_char as i8, but will\n> +if it's u8. You can cast x e.g. `assert_eq!(x as u8, b'a');`, but then you get\n> +a warning on platforms that use u8 and a clean compilation where i8 is used.\n> +\n> +[source]\n> +----\n> +let mut x: std::ffi::c_char = 0x61;\n> +assert_eq!(x, b'a');\n> +----\n> +\n> +== Enum types\n> +Rust enum types should not be used as FFI types. Rust enum types are more like\n> +C union types than C enum's. For something like:\n> +\n> +[source]\n> +----\n> +#[repr(C, u8)]\n> +enum Fruit {\n> +    Apple,\n> +    Banana,\n> +    Cherry,\n> +}\n> +----\n> +\n> +It's easy enough to make sure the Rust enum matches what C would expect, but a\n> +more complex type like.\n> +\n> +[source]\n> +----\n> +enum HashResult {\n> +    SHA1([u8; 20]),\n> +    SHA256([u8; 32]),\n> +}\n> +----\n> +\n> +The Rust compiler has to add a discriminant to the enum to distinguish between\n> +the variants. The width, location, and values for that discriminant is up to\n> +the Rust compiler and is not ABI stable.\n\n"},{"id":"530304","messageId":"14496da7-3d9e-4e07-8893-0a5414fbbe70@gmail.com","threadId":"64326","inReplyTo":"9197903add26e5b8af0bb2dd25bf115670e18e8c.1761776388.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 02/10] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-11-06T09:55:42Z","receivedAt":"2025-11-06T09:55:47Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ezekiel\n\nOn 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> ssize_t is appropriate for dstart and dend because they both describe\n> positive or negative offsets relative to a pointer.\n\nThis paragraph and the subject need updating to match the change from \nssize_t to ptrdiff_t.\n\n> A future patch will move these fields to a different struct. Moving\n> them to the end of xdfile_t now, means the field order of xdfile_t will\n> be disturbed less.\n\nI'm not sure why that matters but I also don't object\n\nThanks\n\nPhillip\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xtypes.h | 2 +-\n>   1 file changed, 1 insertion(+), 1 deletion(-)\n> \n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index f145abba3e..7c8c057bca 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -47,10 +47,10 @@ typedef struct s_xrecord {\n>   typedef struct s_xdfile {\n>   \txrecord_t *recs;\n>   \tlong nrec;\n> -\tlong dstart, dend;\n>   \tbool *changed;\n>   \tlong *rindex;\n>   \tlong nreff;\n> +\tptrdiff_t dstart, dend;\n>   } xdfile_t;\n>   \n>   typedef struct s_xdfenv {\n\n"},{"id":"530305","messageId":"299e25d6-caaf-4672-8160-53fdafe96134@gmail.com","threadId":"64326","inReplyTo":"46bc1b3e25885fbd324a6428ee7ac3b5d272c4ce.1761776388.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-11-06T10:49:05Z","receivedAt":"2025-11-06T10:49:09Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ezekiel\n\nOn 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Rust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\n> referring to bytes in memory, rather than Unicode code points, use\n> uint8_t instead of char.\n\nThe reference to unicode code points here still makes no sense to me. I \nthought the reason for the conversion was to match rust's u8.\n\n> Every usage of this field was inspected and cast to char*, or similar,\n> to avoid signedness warnings/errors from the compiler. Casting was used\n> so that the whole of xdiff doesn't need to be refactored in order to\n> change the type of this field.\n\nThanks for adding this. Having played a little with changing some \nfunction parameters to avoid adding these casts I agree this patch is a \ngood place to stop as the number of changes required quickly spiraled \nout of control.\n\nThanks\n\nPhillip\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xdiffi.c    |  8 ++++----\n>   xdiff/xemit.c     |  6 +++---\n>   xdiff/xmerge.c    | 14 +++++++-------\n>   xdiff/xpatience.c |  2 +-\n>   xdiff/xprepare.c  |  8 ++++----\n>   xdiff/xtypes.h    |  2 +-\n>   xdiff/xutils.c    |  4 ++--\n>   7 files changed, 22 insertions(+), 22 deletions(-)\n> \n> diff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\n> index 6f3998ee54..411a8aa69f 100644\n> --- a/xdiff/xdiffi.c\n> +++ b/xdiff/xdiffi.c\n> @@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n>   \tint ret = 0;\n>   \n>   \tfor (i = 0; i < rec->size; i++) {\n> -\t\tchar c = rec->ptr[i];\n> +\t\tuint8_t c = rec->ptr[i];\n>   \n>   \t\tif (!XDL_ISSPACE(c))\n>   \t\t\treturn ret;\n> @@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n>   \n>   \t\trec = &xe->xdf1.recs[xch->i1];\n>   \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n> -\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> +\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n>   \n>   \t\trec = &xe->xdf2.recs[xch->i2];\n>   \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n> -\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n> +\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n>   \n>   \t\txch->ignore = ignore;\n>   \t}\n> @@ -1008,7 +1008,7 @@ static int record_matches_regex(xrecord_t *rec, xpparam_t const *xpp) {\n>   \tsize_t i;\n>   \n>   \tfor (i = 0; i < xpp->ignore_regex_nr; i++)\n> -\t\tif (!regexec_buf(xpp->ignore_regex[i], rec->ptr, rec->size, 1,\n> +\t\tif (!regexec_buf(xpp->ignore_regex[i], (const char *)rec->ptr, rec->size, 1,\n>   \t\t\t\t &regmatch, 0))\n>   \t\t\treturn 1;\n>   \n> diff --git a/xdiff/xemit.c b/xdiff/xemit.c\n> index b2f1f30cd3..ead930088a 100644\n> --- a/xdiff/xemit.c\n> +++ b/xdiff/xemit.c\n> @@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n>   {\n>   \txrecord_t *rec = &xdf->recs[ri];\n>   \n> -\tif (xdl_emit_diffrec(rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n> +\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n>   \t\treturn -1;\n>   \n>   \treturn 0;\n> @@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n>   \txrecord_t *rec = &xdf->recs[ri];\n>   \n>   \tif (!xecfg->find_func)\n> -\t\treturn def_ff(rec->ptr, rec->size, buf, sz);\n> -\treturn xecfg->find_func(rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n> +\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n> +\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n>   }\n>   \n>   static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\n> diff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\n> index fd600cbb5d..75cb3e76a2 100644\n> --- a/xdiff/xmerge.c\n> +++ b/xdiff/xmerge.c\n> @@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n>   \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n>   \n>   \tfor (i = 0; i < line_count; i++) {\n> -\t\tint result = xdl_recmatch(rec1[i].ptr, rec1[i].size,\n> -\t\t\trec2[i].ptr, rec2[i].size, flags);\n> +\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n> +\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n>   \t\tif (!result)\n>   \t\t\treturn -1;\n>   \t}\n> @@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n>   \n>   static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n>   {\n> -\treturn xdl_recmatch(rec1->ptr, rec1->size,\n> -\t\t\t    rec2->ptr, rec2->size, flags);\n> +\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n> +\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n>   }\n>   \n>   /*\n> @@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n>   \t\t * we have a very simple mmfile structure.\n>   \t\t */\n>   \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n> -\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n> +\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n>   \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n>   \t\tt2.ptr = (char *)xe2->xdf2.recs[m->i2].ptr;\n> -\t\tt2.size = xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n> +\t\tt2.size = (char *)xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n>   \t\t\t+ xe2->xdf2.recs[m->i2 + m->chg2 - 1].size - t2.ptr;\n>   \t\tif (xdl_do_diff(&t1, &t2, xpp, &xe) < 0)\n>   \t\t\treturn -1;\n> @@ -440,7 +440,7 @@ static int line_contains_alnum(const char *ptr, long size)\n>   static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n>   {\n>   \tfor (; chg; chg--, i++)\n> -\t\tif (line_contains_alnum(xe->xdf2.recs[i].ptr,\n> +\t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n>   \t\t\t\txe->xdf2.recs[i].size))\n>   \t\t\treturn 1;\n>   \treturn 0;\n> diff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\n> index 669b653580..bb61354f22 100644\n> --- a/xdiff/xpatience.c\n> +++ b/xdiff/xpatience.c\n> @@ -121,7 +121,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n>   \t\treturn;\n>   \tmap->entries[index].line1 = line;\n>   \tmap->entries[index].hash = record->ha;\n> -\tmap->entries[index].anchor = is_anchor(xpp, map->env->xdf1.recs[line - 1].ptr);\n> +\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n>   \tif (!map->first)\n>   \t\tmap->first = map->entries + index;\n>   \tif (map->last) {\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 192334f1b7..4cb18b2b88 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n>   \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n>   \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n>   \t\tif (rcrec->rec.ha == rec->ha &&\n> -\t\t\t\txdl_recmatch(rcrec->rec.ptr, rcrec->rec.size,\n> -\t\t\t\t\trec->ptr, rec->size, cf->flags))\n> +\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n> +\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n>   \t\t\tbreak;\n>   \n>   \tif (!rcrec) {\n> @@ -156,8 +156,8 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n>   \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n>   \t\t\t\tgoto abort;\n>   \t\t\tcrec = &xdf->recs[xdf->nrec++];\n> -\t\t\tcrec->ptr = prev;\n> -\t\t\tcrec->size = (long) (cur - prev);\n> +\t\t\tcrec->ptr = (uint8_t const *)prev;\n> +\t\t\tcrec->size =(long) ( cur - prev);\n>   \t\t\tcrec->ha = hav;\n>   \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n>   \t\t\t\tgoto abort;\n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index 7c8c057bca..b1c520a378 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -39,7 +39,7 @@ typedef struct s_chastore {\n>   } chastore_t;\n>   \n>   typedef struct s_xrecord {\n> -\tchar const *ptr;\n> +\tuint8_t const *ptr;\n>   \tlong size;\n>   \tunsigned long ha;\n>   } xrecord_t;\n> diff --git a/xdiff/xutils.c b/xdiff/xutils.c\n> index 447e66c719..7be063bfb6 100644\n> --- a/xdiff/xutils.c\n> +++ b/xdiff/xutils.c\n> @@ -465,10 +465,10 @@ int xdl_fall_back_diff(xdfenv_t *diff_env, xpparam_t const *xpp,\n>   \txdfenv_t env;\n>   \n>   \tsubfile1.ptr = (char *)diff_env->xdf1.recs[line1 - 1].ptr;\n> -\tsubfile1.size = diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n> +\tsubfile1.size = (char *)diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n>   \t\tdiff_env->xdf1.recs[line1 + count1 - 2].size - subfile1.ptr;\n>   \tsubfile2.ptr = (char *)diff_env->xdf2.recs[line2 - 1].ptr;\n> -\tsubfile2.size = diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n> +\tsubfile2.size = (char *)diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n>   \t\tdiff_env->xdf2.recs[line2 + count2 - 2].size - subfile2.ptr;\n>   \tif (xdl_do_diff(&subfile1, &subfile2, xpp, &env) < 0)\n>   \t\treturn -1;\n\n"},{"id":"530307","messageId":"3f7bbb5e-0d67-4ef2-82fb-e0b00683c178@gmail.com","threadId":"64326","inReplyTo":"46bc1b3e25885fbd324a6428ee7ac3b5d272c4ce.1761776388.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-11-06T10:55:39Z","receivedAt":"2025-11-06T10:55:42Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> @@ -156,8 +156,8 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n>   \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n>   \t\t\t\tgoto abort;\n>   \t\t\tcrec = &xdf->recs[xdf->nrec++];\n> -\t\t\tcrec->ptr = prev;\n> -\t\t\tcrec->size = (long) (cur - prev);\n> +\t\t\tcrec->ptr = (uint8_t const *)prev;\n> +\t\t\tcrec->size =(long) ( cur - prev);\n\nThe changes to crec->size here look unintentional\n\nThanks\n\nPhillip\n\n"},{"id":"530308","messageId":"a66fb440-058e-4cd8-8971-9c320c0387e8@gmail.com","threadId":"64326","inReplyTo":"59054ea0cb65718dbac500d342bc960bdb5066c1.1761776388.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-11-06T11:00:34Z","receivedAt":"2025-11-06T11:00:37Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ezekiel\n\nOn 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> The ha field is serving two different purposes, which makes the code\n> harder to read. At first glance it looks like many places assume\n> there could never be hash collisions between lines of the two input\n> files. In reality, line_hash is used together with xdl_recmatch() to\n> ensure correct comparisons of lines, even when collisions occur.\n> \n> To make this clearer, the old ha field has been split:\n>    * line_hash: The straightforward hash of a line, requiring no\n>      additional context.\n>    * minimal_perfect_hash: Not a new concept, but now a separate\n>      field. It comes from the classifier's general-purpose hash table,\n>      which assigns each line a unique and minimal hash across the two\n>      files.\n\nIt would be nice to explain the differing types for the two fields in \nthe commit message.\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 85e56021da..16236bd045 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -96,9 +96,9 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n>   \tlong hi;\n>   \txdlclass_t *rcrec;\n>   \n> -\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n> +\thi = (long) XDL_HASHLONG(rec->line_hash, cf->hbits);\n\n\"hi\" is only used as an array index so it might be nicer to change it to \nsize_t and avoid this cast instead.\n\nThanks\n\nPhillip\n\n"},{"id":"530348","messageId":"CAH=ZcbA25eyMhQpvK7eh=ydZkg5RdzbdRFEdj-22T+d1VuTazA@mail.gmail.com","threadId":"64326","inReplyTo":"995f77a3-b94c-46df-87d3-22c7b2a3c762@gmail.com","subject":"Re: [PATCH v2 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-06T22:52:39Z","receivedAt":"2025-11-06T22:52:52Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Thu, Nov 6, 2025 at 2:55 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > Document other nuances with crossing the FFI boundary. Other language\n> > mappings may be added in the future.\n>\n> Thanks for adding this, I've left a few comments below. Overall I\n> thought it was very well written.\n\nThanks.\n\nI felt it was necessary since C vs Rust types keep coming up over and\nover again. I'm flexible with the wording of this document. I was just\ntrying to convey a firm and clear stance on what is and isn't proper\nin Git.\n\n> I tried building an html version of\n> this but even after adding it to the list of TECH_DOCS in\n> Documentation/Makefile with\n>\n> diff --git a/Documentation/Makefile b/Documentation/Makefile\n> index 47208269a2e..2699f0b24af 100644\n> --- a/Documentation/Makefile\n> +++ b/Documentation/Makefile\n> @@ -143,6 +143,7 @@ TECH_DOCS += technical/shallow\n>   TECH_DOCS += technical/sparse-checkout\n>   TECH_DOCS += technical/sparse-index\n>   TECH_DOCS += technical/trivial-merge\n> +TECH_DOCS += technical/unambiguous-types\n>   TECH_DOCS += technical/unit-tests\n>   SP_ARTICLES += $(TECH_DOCS)\n>   SP_ARTICLES += technical/api-index\n>\n> it fails with\n>\n> $ make -C Documentation/ technical/unambiguous-types.html\n>                                        Merge branch\n> 'ps/object-source-loose' into seen\n> make: Entering directory '/home/phil/src/git/Documentation'\n>      GEN asciidoc.conf\n>      * new asciidoc flags\n>      ASCIIDOC technical/unambiguous-types.html\n> asciidoc: ERROR: unambiguous-types.adoc: line 139: undefined filter\n> attribute in command: source-highlight --gen-version -f xhtml -s\n> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n> {args=}\n> asciidoc: ERROR: unambiguous-types.adoc: line 162: undefined filter\n> attribute in command: source-highlight --gen-version -f xhtml -s\n> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n> {args=}\n> asciidoc: ERROR: unambiguous-types.adoc: line 177: undefined filter\n> attribute in command: source-highlight --gen-version -f xhtml -s\n> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n> {args=}\n> asciidoc: ERROR: unambiguous-types.adoc: line 187: undefined filter\n> attribute in command: source-highlight --gen-version -f xhtml -s\n> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n> {args=}\n> asciidoc: ERROR: unambiguous-types.adoc: line 199: undefined filter\n> attribute in command: source-highlight --gen-version -f xhtml -s\n> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n> {args=}\n> asciidoc: ERROR: unambiguous-types.adoc: line 213: undefined filter\n> attribute in command: source-highlight --gen-version -f xhtml -s\n> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n> {args=}\n> asciidoc: ERROR: unambiguous-types.adoc: line 224: undefined filter\n> attribute in command: source-highlight --gen-version -f xhtml -s\n> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n> {args=}\n> make: *** [Makefile:396: technical/unambiguous-types.html] Error 1\n> make: *** Deleting file 'technical/unambiguous-types.html'\n> make: Leaving directory '/home/phil/src/git/Documentation'\n\nI've never created documentation for Git before, so this helps. I'll\nincorporate your suggestions.\n\n> > +== Character types\n> > +\n> > +This is where C and Rust don't have a clean one-to-one mapping. A C `char` is\n> > +an 8-bit type that is signless (neither signed nor unsigned)\n>\n> I found this a bit confusing. Isn't the signedness of \"char\"\n> implementation defined rather than it being \"signless\"\n>\n> > which causes\n> > +problems with e.g. `make DEVELOPER=1`.\n>\n> I'm not sure what this is referring to - maybe -Wsign-compare?\n\nWhen I build Git with `make DEVELOPER=1` and I compare uint8_t with\nchar it complains about a difference in signedness. When I compare\nint8_t with char it also complains about a difference in signedness.\nSo it is implementation defined, but it's also neither signed nor\nunsigned according to DEVELOPER=1 since it complains either way.\n\n> > Rust's `char` type is an unsigned 32-bit\n> > +integer that is used to describe Unicode code points. Even though a C `char`\n> > +is the same width as `u8`, `char` should be converted to u8 where it is\n> > +describing bytes in memory.\n>\n> I'm dreading the point where we start sharing \"struct strbuf\" with rust\n> and have to change the \"buf\" member from \"char*\" to \"uint8_t*\". While it\n> is not used in the xdiff code it is ubiquitous everywhere else and there\n> are lots of places where be pass the \"buf\" member to functions expecting\n> a \"char*\".\n>\n>         git grep -E '(\\.|->)buf\\W'\n>\n> has over 4000 matches\n\nThis is why I started in Xdiff since its code is mostly isolated. I\nthink that we might have to bite the bullet and deal with the ugly\nmapping of char on the C side and u8 on the Rust side when dealing\nwith strbuf. Maybe as we translate more of C into Rust someone will\nhave a better suggestion. I think my ivec type would be better since\nstrbuf is almost a special case of my ivec type, but dealing with\nstrbuf is outside the scope of this patch series.\n\n> > If a C `char` is not describing bytes, then it\n> > +should be converted to a more accurate unambiguous type.\n>\n> That's a good point.\n>\n> > +While you could specify `char` in the C code and `u8` in Rust code, it's not as\n> > +clear what the appropriate type is, but it would work across the FFI boundary.\n> > +However the bigger problem comes from code generation tools like cbindgen and\n> > +bindgen. When cbindgen see u8 in Rust it will generate uint8_t on the C side\n> > +which will cause differ in signedness warnings/errors. Similarly if bindgen\n> > +see `char` on the C side it will generate `std::ffi::c_char` which has its own\n> > +problems.\n>\n> Yeah, we definitely don't want to be using \"std::ffi::c_char\" in our\n> rust implementations. I do wonder if we might want to use it (or CStr)\n> judiciously in function parameters and immediately convert it to u8 in\n> the function body where the function is called from C though.\n\nThat's basically the design pattern I've been using.\n\nIn many of my translations from C to Rust I create a Rust stub\nfunction that takes pointer types and wraps them into safe types which\nthen get handed off to a safe Rust function. I think that in the cases\nwhere CString/CStr is required the Rust stub function would create a\n&[u8] slice for the safe function to operate on.\n\n> > +=== Notes\n> > +^1^ This is only true if stdbool.h (or equivalent) is used. +\n> > +^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the\n> > +platform/arch for C does not follow IEEE-754 then this equivalence does not\n> > +hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n> > +there may be a strange platform/arch where even this isn't true. +\n> > +^3^ C also defines uintptr_t, but this should not be used in Git. +\n> > +^4^ C also defines ssize_t and intptr_t, but these should not be used in Git. +\n>\n> [u]intptr_t and ssize_t are used in git already. As Junio has pointed\n> out there are sane uses for these types but we don't want to use them in\n> structs or function parameters where the struct or function is shared\n> with rust.\n\nYou're right, I should update the phrasing. Something like: \"These\ntypes shouldn't be used if their explicit purpose is for FFI. Whether\nas a field in a struct or part of a function signature.\" I'll update\nthe wording.\n\n> > +\n> > +== Problems with std::ffi::c_* types in Rust\n> > +TL;DR: They're not guaranteed to match C types for all possible C\n> > +compilers/platforms/architectures.\n>\n> Is this official policy of the rust project?\n\nNo, this is a personal inference based on logical deduction. The c_*\ndefinitions have changed over time with new Rust version releases, and\nGit targets more platforms/architectures than what Rust officially\nsupports. While it's not guaranteed that it won't work everywhere.\nIt's also not guaranteed to work everywhere either. On top of that\nwe're targeting 1.63.0 who's c_* definitions are different in 1.89.0\nwhich I show an example of with c_long_definition. Can anyone say with\ncertainty that Rust got these mappings right or wrong for all possible\nC compilers/architectures/platforms? If so (which I highly doubt)\ncould someone provide a link?\n"},{"id":"530349","messageId":"CAH=ZcbDe+3Bdxz4OYZw7VMRSXaR1PsWx0GipogD5gfD0n=+XYA@mail.gmail.com","threadId":"64326","inReplyTo":"14496da7-3d9e-4e07-8893-0a5414fbbe70@gmail.com","subject":"Re: [PATCH v2 02/10] xdiff: use ssize_t for dstart/dend, make them last in xdfile_t","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-06T22:56:43Z","receivedAt":"2025-11-06T22:56:56Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Thu, Nov 6, 2025 at 2:55 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > ssize_t is appropriate for dstart and dend because they both describe\n> > positive or negative offsets relative to a pointer.\n>\n> This paragraph and the subject need updating to match the change from\n> ssize_t to ptrdiff_t.\n\nYou're right. I thought I updated that. I'll make that change for the\nnext version.\n"},{"id":"530350","messageId":"CAH=ZcbA7d5Z7d=VT2_o=+M8pYrGzO7TgAaLisk2k0p7CuQuSPQ@mail.gmail.com","threadId":"64326","inReplyTo":"299e25d6-caaf-4672-8160-53fdafe96134@gmail.com","subject":"Re: [PATCH v2 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-06T23:13:51Z","receivedAt":"2025-11-06T23:14:04Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Thu, Nov 6, 2025 at 3:49 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > Rust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\n> > referring to bytes in memory, rather than Unicode code points, use\n> > uint8_t instead of char.\n>\n> The reference to unicode code points here still makes no sense to me. I\n> thought the reason for the conversion was to match rust's u8.\n\nIt is to match Rust's u8 type, but I was also trying to convey that\nptr is referring to bytes and not characters _because_ xdiff performs\ntextual differences. It's not spelled out anywhere in Xdiff that it\ndoes or doesn't take Unicode into consideration. Would comparing\nUnicode code points change how Xdiff behaves? Should it behave\ndifferently? I don't know. My understanding is that whether the bytes\nare utf-8, utf-16le, utf-16be, or some other encoding of Unicode.\nXdiff doesn't care and treats the lines in a file as raw byte strings.\n\nThere's also the question of \"Should the Rust side of Xdiff treat\nlines in a file as &[u8] or &str?\" The reason why this matters is\nbecause in order to get a &str from &[u8] in Rust you need to call a\nfunction like:\n\n```\nlet raw_bytes = b\"abc\\n\";\nlet result = std::str::from_utf8(raw_bytes);\nif let Ok(line) = result {\n    // do something\n}\n```\n\nWhat happens if it's not utf8 encoded? What if it's malformed utf8? To\navoid these problems I only use &[u8] in xdiff and perform differences\non raw byte strings rather than considering Unicode at all like how\nXdiff already does.\n\nDoes that explain my comment about Unicode or does it still seem out\nof place to you? I can remove the mention of Unicode from the commit\nmessage if this still doesn't make any sense to you.\n\n> > Every usage of this field was inspected and cast to char*, or similar,\n> > to avoid signedness warnings/errors from the compiler. Casting was used\n> > so that the whole of xdiff doesn't need to be refactored in order to\n> > change the type of this field.\n>\n> Thanks for adding this. Having played a little with changing some\n> function parameters to avoid adding these casts I agree this patch is a\n> good place to stop as the number of changes required quickly spiraled\n> out of control.\n\nI'm not excited about the casts either, but these 2 structs are\nfundamental to how Xdiff passes data around, and so they need to be\nFFI friendly. I don't plan on converting other structs or function\nsignatures in Xdiff unless I really have to.\n"},{"id":"530351","messageId":"CAH=ZcbDTnvgrkfKYe_uyPHh5Xd2Pbw4532UvNcz+nz+rpbhUiQ@mail.gmail.com","threadId":"64326","inReplyTo":"3f7bbb5e-0d67-4ef2-82fb-e0b00683c178@gmail.com","subject":"Re: [PATCH v2 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-06T23:14:30Z","receivedAt":"2025-11-06T23:14:43Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Thu, Nov 6, 2025 at 3:55 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> > @@ -156,8 +156,8 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n> >                       if (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n> >                               goto abort;\n> >                       crec = &xdf->recs[xdf->nrec++];\n> > -                     crec->ptr = prev;\n> > -                     crec->size = (long) (cur - prev);\n> > +                     crec->ptr = (uint8_t const *)prev;\n> > +                     crec->size =(long) ( cur - prev);\n>\n> The changes to crec->size here look unintentional\n\nI agree. I'll change that.\n"},{"id":"530352","messageId":"CAH=ZcbCnH_3C9D4fppVAL3yZpVKGvJ6SM1FrPwpj5fufEbFNtQ@mail.gmail.com","threadId":"64326","inReplyTo":"a66fb440-058e-4cd8-8971-9c320c0387e8@gmail.com","subject":"Re: [PATCH v2 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-06T23:20:21Z","receivedAt":"2025-11-06T23:20:34Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Thu, Nov 6, 2025 at 4:00 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > The ha field is serving two different purposes, which makes the code\n> > harder to read. At first glance it looks like many places assume\n> > there could never be hash collisions between lines of the two input\n> > files. In reality, line_hash is used together with xdl_recmatch() to\n> > ensure correct comparisons of lines, even when collisions occur.\n> >\n> > To make this clearer, the old ha field has been split:\n> >    * line_hash: The straightforward hash of a line, requiring no\n> >      additional context.\n> >    * minimal_perfect_hash: Not a new concept, but now a separate\n> >      field. It comes from the classifier's general-purpose hash table,\n> >      which assigns each line a unique and minimal hash across the two\n> >      files.\n>\n> It would be nice to explain the differing types for the two fields in\n> the commit message.\n\nI'll add something like:\nline_hash is a uint64_t because it is the output of a fixed width hash\nfunction. minimal_perfect_hash is size_t because its purpose is to\nindex into an array. This also avoids the problem of having to cast to\nusize on the Rust side every time minimal_perfect_hash is used to\nindex a slice.\n\n> > diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> > index 85e56021da..16236bd045 100644\n> > --- a/xdiff/xprepare.c\n> > +++ b/xdiff/xprepare.c\n> > @@ -96,9 +96,9 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n> >       long hi;\n> >       xdlclass_t *rcrec;\n> >\n> > -     hi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n> > +     hi = (long) XDL_HASHLONG(rec->line_hash, cf->hbits);\n>\n> \"hi\" is only used as an array index so it might be nicer to change it to\n> size_t and avoid this cast instead.\n\nI agree. I'll make that change.\n"},{"id":"530425","messageId":"fa95b29a-077c-4df5-9c59-34e0c1447e70@gmail.com","threadId":"64326","inReplyTo":"CAH=ZcbA25eyMhQpvK7eh=ydZkg5RdzbdRFEdj-22T+d1VuTazA@mail.gmail.com","subject":"Re: [PATCH v2 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-11-09T14:14:11Z","receivedAt":"2025-11-09T14:14:19Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 06/11/2025 22:52, Ezekiel Newren wrote:\n> On Thu, Nov 6, 2025 at 2:55 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>> On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote:\n>>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>>\n>>> Document other nuances with crossing the FFI boundary. Other language\n>>> mappings may be added in the future.\n>>\n>> Thanks for adding this, I've left a few comments below. Overall I\n>> thought it was very well written.\n> \n> Thanks.\n> \n> I felt it was necessary since C vs Rust types keep coming up over and\n> over again. I'm flexible with the wording of this document. I was just\n> trying to convey a firm and clear stance on what is and isn't proper\n> in Git.\n\nThat will definitely be useful as we add more rust code. In the future \nwe may want to add a summary of which types to use to \nDocumentation/CodingGuidelines but that doesn't need to be done in this \nseries.\n\n>> I tried building an html version of\n>> this but even after adding it to the list of TECH_DOCS in\n>> Documentation/Makefile with\n>>\n>> diff --git a/Documentation/Makefile b/Documentation/Makefile\n>> index 47208269a2e..2699f0b24af 100644\n>> --- a/Documentation/Makefile\n>> +++ b/Documentation/Makefile\n>> @@ -143,6 +143,7 @@ TECH_DOCS += technical/shallow\n>>    TECH_DOCS += technical/sparse-checkout\n>>    TECH_DOCS += technical/sparse-index\n>>    TECH_DOCS += technical/trivial-merge\n>> +TECH_DOCS += technical/unambiguous-types\n>>    TECH_DOCS += technical/unit-tests\n>>    SP_ARTICLES += $(TECH_DOCS)\n>>    SP_ARTICLES += technical/api-index\n>>\n>> it fails with\n>>\n>> $ make -C Documentation/ technical/unambiguous-types.html\n>>                                         Merge branch\n>> 'ps/object-source-loose' into seen\n>> make: Entering directory '/home/phil/src/git/Documentation'\n>>       GEN asciidoc.conf\n>>       * new asciidoc flags\n>>       ASCIIDOC technical/unambiguous-types.html\n>> asciidoc: ERROR: unambiguous-types.adoc: line 139: undefined filter\n>> attribute in command: source-highlight --gen-version -f xhtml -s\n>> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n>> {args=}\n>> asciidoc: ERROR: unambiguous-types.adoc: line 162: undefined filter\n>> attribute in command: source-highlight --gen-version -f xhtml -s\n>> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n>> {args=}\n>> asciidoc: ERROR: unambiguous-types.adoc: line 177: undefined filter\n>> attribute in command: source-highlight --gen-version -f xhtml -s\n>> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n>> {args=}\n>> asciidoc: ERROR: unambiguous-types.adoc: line 187: undefined filter\n>> attribute in command: source-highlight --gen-version -f xhtml -s\n>> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n>> {args=}\n>> asciidoc: ERROR: unambiguous-types.adoc: line 199: undefined filter\n>> attribute in command: source-highlight --gen-version -f xhtml -s\n>> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n>> {args=}\n>> asciidoc: ERROR: unambiguous-types.adoc: line 213: undefined filter\n>> attribute in command: source-highlight --gen-version -f xhtml -s\n>> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n>> {args=}\n>> asciidoc: ERROR: unambiguous-types.adoc: line 224: undefined filter\n>> attribute in command: source-highlight --gen-version -f xhtml -s\n>> {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}}\n>> {args=}\n>> make: *** [Makefile:396: technical/unambiguous-types.html] Error 1\n>> make: *** Deleting file 'technical/unambiguous-types.html'\n>> make: Leaving directory '/home/phil/src/git/Documentation'\n> \n> I've never created documentation for Git before, so this helps. I'll\n> incorporate your suggestions.\n\nWe should also add this file to Documentation/technical/meson.build. It \nseems those errors above are due to some incompatibility between \nasciidoc and asciidoctor as I just tried running\n\n     make -C Documentation/ USE_ASCIIDOCTOR=1 \ntechnical/unambiguous-types.html\n\nand it worked just fine. I'm afraid I don't know enough asciidoc to make \nany helpful suggestions on how to fix it.\n\n>>> +== Character types\n>>> +\n>>> +This is where C and Rust don't have a clean one-to-one mapping. A C `char` is\n>>> +an 8-bit type that is signless (neither signed nor unsigned)\n>>\n>> I found this a bit confusing. Isn't the signedness of \"char\"\n>> implementation defined rather than it being \"signless\"\n>>\n>>> which causes\n>>> +problems with e.g. `make DEVELOPER=1`.\n>>\n>> I'm not sure what this is referring to - maybe -Wsign-compare?\n> \n> When I build Git with `make DEVELOPER=1` and I compare uint8_t with\n> char it complains about a difference in signedness. When I compare\n> int8_t with char it also complains about a difference in signedness.\n> So it is implementation defined, but it's also neither signed nor\n> unsigned according to DEVELOPER=1 since it complains either way.\n\nOh, I see - this is saying mixing \"char\" and \"uint8_t\" causes problems. \nI agree, perhaps we could expand this slightly to mention comparison \nwith uint8_t to make it clearer.\n\n>>> Rust's `char` type is an unsigned 32-bit\n>>> +integer that is used to describe Unicode code points. Even though a C `char`\n>>> +is the same width as `u8`, `char` should be converted to u8 where it is\n>>> +describing bytes in memory.\n>>\n>> I'm dreading the point where we start sharing \"struct strbuf\" with rust\n>> and have to change the \"buf\" member from \"char*\" to \"uint8_t*\". While it\n>> is not used in the xdiff code it is ubiquitous everywhere else and there\n>> are lots of places where be pass the \"buf\" member to functions expecting\n>> a \"char*\".\n>>\n>>          git grep -E '(\\.|->)buf\\W'\n>>\n>> has over 4000 matches\n> \n> This is why I started in Xdiff since its code is mostly isolated.\n\nGood plan!\n\n> I\n> think that we might have to bite the bullet and deal with the ugly\n> mapping of char on the C side and u8 on the Rust side when dealing\n> with strbuf. Maybe as we translate more of C into Rust someone will\n> have a better suggestion. I think my ivec type would be better since\n> strbuf is almost a special case of my ivec type, but dealing with\n> strbuf is outside the scope of this patch series.\n\nYes, hopefully it will become clearer what the least painful route \nforward is as we get more experience with rust <=> C iterop.\n\n>>> +While you could specify `char` in the C code and `u8` in Rust code, it's not as\n>>> +clear what the appropriate type is, but it would work across the FFI boundary.\n>>> +However the bigger problem comes from code generation tools like cbindgen and\n>>> +bindgen. When cbindgen see u8 in Rust it will generate uint8_t on the C side\n>>> +which will cause differ in signedness warnings/errors. Similarly if bindgen\n>>> +see `char` on the C side it will generate `std::ffi::c_char` which has its own\n>>> +problems.\n>>\n>> Yeah, we definitely don't want to be using \"std::ffi::c_char\" in our\n>> rust implementations. I do wonder if we might want to use it (or CStr)\n>> judiciously in function parameters and immediately convert it to u8 in\n>> the function body where the function is called from C though.\n> \n> That's basically the design pattern I've been using.\n> \n> In many of my translations from C to Rust I create a Rust stub\n> function that takes pointer types and wraps them into safe types which\n> then get handed off to a safe Rust function. I think that in the cases\n> where CString/CStr is required the Rust stub function would create a\n> &[u8] slice for the safe function to operate on.\n\nThat sounds like a good pattern - we get a nice interface for the C code \nand the rust implementation uses the idiomatic rust types.\n\nThanks\n\nPhillip\n\n>>> +=== Notes\n>>> +^1^ This is only true if stdbool.h (or equivalent) is used. +\n>>> +^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the\n>>> +platform/arch for C does not follow IEEE-754 then this equivalence does not\n>>> +hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n>>> +there may be a strange platform/arch where even this isn't true. +\n>>> +^3^ C also defines uintptr_t, but this should not be used in Git. +\n>>> +^4^ C also defines ssize_t and intptr_t, but these should not be used in Git. +\n>>\n>> [u]intptr_t and ssize_t are used in git already. As Junio has pointed\n>> out there are sane uses for these types but we don't want to use them in\n>> structs or function parameters where the struct or function is shared\n>> with rust.\n> \n> You're right, I should update the phrasing. Something like: \"These\n> types shouldn't be used if their explicit purpose is for FFI. Whether\n> as a field in a struct or part of a function signature.\" I'll update\n> the wording.\n> \n>>> +\n>>> +== Problems with std::ffi::c_* types in Rust\n>>> +TL;DR: They're not guaranteed to match C types for all possible C\n>>> +compilers/platforms/architectures.\n>>\n>> Is this official policy of the rust project?\n> \n> No, this is a personal inference based on logical deduction. The c_*\n> definitions have changed over time with new Rust version releases, and\n> Git targets more platforms/architectures than what Rust officially\n> supports. While it's not guaranteed that it won't work everywhere.\n> It's also not guaranteed to work everywhere either. On top of that\n> we're targeting 1.63.0 who's c_* definitions are different in 1.89.0\n> which I show an example of with c_long_definition. Can anyone say with\n> certainty that Rust got these mappings right or wrong for all possible\n> C compilers/architectures/platforms? If so (which I highly doubt)\n> could someone provide a link?\n> \n\n"},{"id":"530526","messageId":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v2.git.git.1761776388.gitgitgadget@gmail.com","subject":"[PATCH v3 00/10] Xdiff cleanup part2","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:22Z","receivedAt":"2025-11-11T19:42:35Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"Changes in v3:\n\n * Address comments about commit messages and documentation\n * Add unambiguous-types.adoc to Makefile and Meson\n * Use markdown style to avoid asciidoc issues\n\nChanges in v2:\n\n * Added documentation about unambiguous types and FFI\n * Addressed comments on the mailing list\n\n\nOriginal cover letter below:\n============================\n\nMaintainer note: This patch series builds on top of en/xdiff-cleanup and\nam/xdiff-hash-tweak (both of which are now in master).\n\nThe primary goal of this patch series is to convert every field's type in\nxrecord_t and xdfile_t to be unambiguous, in preparation to make it more\nRust FFI friendly. Additionally the ha field in xrecord_t is split into\nline_hash and minimal_perfect hash.\n\nThe order of some of the fields has changed as called out by the commit\nmessages.\n\nBefore:\n\ntypedef struct s_xrecord {\n\tchar const *ptr;\n\tlong size;\n\tunsigned long ha;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tlong nrec;\n\tlong dstart, dend;\n\tbool *changed;\n\tlong *rindex;\n\tlong nreff;\n} xdfile_t;\n\n\nAfter part 2\n\ntypedef struct s_xrecord {\n\tuint8_t const *ptr;\n\tsize_t size;\n\tuint64_t line_hash;\n\tsize_t minimal_perfect_hash;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tsize_t nrec;\n\tbool *changed;\n\tsize_t *reference_index;\n\tsize_t nreff;\n\tssize_t dstart, dend;\n} xdfile_t;\n\n\nEzekiel Newren (10):\n  doc: define unambiguous type mappings across C and Rust\n  xdiff: use ptrdiff_t for dstart/dend\n  xdiff: make xrecord_t.ptr a uint8_t instead of char\n  xdiff: use size_t for xrecord_t.size\n  xdiff: use unambiguous types in xdl_hash_record()\n  xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n  xdiff: make xdfile_t.nrec a size_t instead of long\n  xdiff: make xdfile_t.nreff a size_t instead of long\n  xdiff: change rindex from long to size_t in xdfile_t\n  xdiff: rename rindex -> reference_index\n\n Documentation/Makefile                        |   1 +\n Documentation/technical/meson.build           |   1 +\n .../technical/unambiguous-types.adoc          | 239 ++++++++++++++++++\n xdiff-interface.c                             |   2 +-\n xdiff/xdiffi.c                                |  29 +--\n xdiff/xemit.c                                 |  28 +-\n xdiff/xhistogram.c                            |   4 +-\n xdiff/xmerge.c                                |  30 +--\n xdiff/xpatience.c                             |  14 +-\n xdiff/xprepare.c                              |  60 ++---\n xdiff/xtypes.h                                |  15 +-\n xdiff/xutils.c                                |  32 +--\n xdiff/xutils.h                                |   6 +-\n 13 files changed, 351 insertions(+), 110 deletions(-)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\n\nbase-commit: a99f379adf116d53eb11957af5bab5214915f91d\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2070%2Fezekielnewren%2Fxdiff_cleanup_part2-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2070/ezekielnewren/xdiff_cleanup_part2-v3\nPull-Request: https://github.com/git/git/pull/2070\n\nRange-diff vs v2:\n\n  1:  88133848d1 !  1:  e5d084d340 doc: define unambiguous type mappings across C and Rust\n     @@ Commit message\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n     + ## Documentation/Makefile ##\n     +@@ Documentation/Makefile: TECH_DOCS += technical/shallow\n     + TECH_DOCS += technical/sparse-checkout\n     + TECH_DOCS += technical/sparse-index\n     + TECH_DOCS += technical/trivial-merge\n     ++TECH_DOCS += technical/unambiguous-types\n     + TECH_DOCS += technical/unit-tests\n     + SP_ARTICLES += $(TECH_DOCS)\n     + SP_ARTICLES += technical/api-index\n     +\n     + ## Documentation/technical/meson.build ##\n     +@@ Documentation/technical/meson.build: articles = [\n     +   'sparse-checkout.adoc',\n     +   'sparse-index.adoc',\n     +   'trivial-merge.adoc',\n     ++  'unambiguous-types.adoc',\n     +   'unit-tests.adoc',\n     + ]\n     + \n     +\n       ## Documentation/technical/unambiguous-types.adoc (new) ##\n      @@\n      += Unambiguous types\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +|===\n      +| C Type | Rust Type\n      +| size_t^3^     | usize\n     -+| ptrdiff_t^4^  | isize\n     ++| ptrdiff_t^3^  | isize\n      +|===\n      +\n      +== Character types\n      +\n     -+This is where C and Rust don't have a clean one-to-one mapping. A C `char` is\n     -+an 8-bit type that is signless (neither signed nor unsigned) which causes\n     -+problems with e.g. `make DEVELOPER=1`. Rust's `char` type is an unsigned 32-bit\n     -+integer that is used to describe Unicode code points. Even though a C `char`\n     -+is the same width as `u8`, `char` should be converted to u8 where it is\n     -+describing bytes in memory. If a C `char` is not describing bytes, then it\n     -+should be converted to a more accurate unambiguous type.\n     ++This is where C and Rust don't have a clean one-to-one mapping.\n     ++\n     ++C comparison problem: While the sign of `char` is implementation defined, it's\n     ++also signless (neither signed nor unsigned). When building with\n     ++`make DEVELOPER=1` it will complain about a \"differ in signedness\" when `char`\n     ++is compared with `uint8_t` or `int8_t`.\n     ++\n     ++Rust's `char` type is an unsigned 32-bit integer that is used to describe\n     ++Unicode code points. Even though a C `char` is the same width as `u8`, `char`\n     ++should be converted to u8 where it is describing bytes in memory. If a C\n     ++`char` is not describing bytes, then it should be converted to a more accurate\n     ++unambiguous type. The reason for mentioning Unicode here is because of how &str\n     ++is defined in Rust and how to create a &str from &[u8]. Rust assumes that &str\n     ++is a correctly encoded utf-8 string, i.e. text in memory. Where as a C `char`\n     ++makes no assumption about the bytes that it is representing.\n     ++\n     ++```\n     ++let raw_bytes = b\"abc\\n\";\n     ++let result = std::str::from_utf8(raw_bytes);\n     ++if let Ok(line) = result {\n     ++    // do something with text\n     ++}\n     ++```\n      +\n      +While you could specify `char` in the C code and `u8` in Rust code, it's not as\n      +clear what the appropriate type is, but it would work across the FFI boundary.\n     -+However the bigger problem comes from code generation tools like cbindgen and\n     -+bindgen. When cbindgen see u8 in Rust it will generate uint8_t on the C side\n     -+which will cause differ in signedness warnings/errors. Similaraly if bindgen\n     -+see `char` on the C side it will generate `std::ffi::c_char` which has its own\n     ++However, the bigger problem comes from code generation tools like cbindgen and\n     ++bindgen. When cbindgen sees u8 in Rust it will generate uint8_t on the C side\n     ++which will cause differ in signedness warnings/errors. Similarly if bindgen\n     ++sees `char` on the C side it will generate `std::ffi::c_char` which has its own\n      +problems.\n      +\n      +=== Notes\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +platform/arch for C does not follow IEEE-754 then this equivalence does not\n      +hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n      +there may be a strange platform/arch where even this isn't true. +\n     -+^3^ C also defines uintptr_t, but this should not be used in Git. +\n     -+^4^ C also defines ssize_t and intptr_t, but these should not be used in Git. +\n     ++^3^ C also defines uintptr_t, ssize_t and intptr_t, but these types are\n     ++discouraged for FFI purposes. For functions like `read()` and `write()` ssize_t\n     ++should be cast to a different, and unambiguous, type before being passed over\n     ++the FFI boundary. +\n      +\n      +== Problems with std::ffi::c_* types in Rust\n     -+TL;DR: They're not guaranteed to match C types for all possible C\n     -+compilers/platforms/architectures.\n     ++TL;DR: In practice, Rust's `c_*` types aren't guaranteed to match C types for\n     ++all possible C compilers, platforms, or architectures, because Rust only\n     ++ensures correctness of C types on officially supported targets. These\n     ++definitions have changed over time to match more targets which means that the\n     ++c_* definitions will differ based on which Rust version Git chooses to use.\n      +\n     -+Only a few of Rust's C FFI types are considered safe and semantically clear to\n     -+use: +\n     ++Current list of safe, Rust side, FFI types in Git: +\n      +\n      +* `c_void`\n      +* `CStr`\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +Even then, they should be used sparingly, and only where the semantics match\n      +exactly.\n      +\n     -+The std::os::raw::c_* (which is deprecated) directly inherits the problems of\n     -+core::ffi, which changes over time and seems to make a best guess at the\n     -+correct definition for a given platform/target. This probably isn't a problem\n     -+for all platforms that Rust supports currently, but can anyone say that Rust\n     -+got it right for all C compilers of all platforms/targets?\n     -+\n     -+On top of all of that we're targeting an older version of Rust which doesn't\n     -+have the latest mappings.\n     ++The std::os::raw::c_* directly inherits the problems of core::ffi, which\n     ++changes over time and seems to make a best guess at the correct definition for\n     ++a given platform/target. This probably isn't a problem for all other platforms\n     ++that Rust supports currently, but can anyone say that Rust got it right for all\n     ++C compilers of all platforms/targets?\n      +\n      +To give an example: c_long is defined in\n      +footnote:[https://doc.rust-lang.org/1.63.0/src/core/ffi/mod.rs.html#175-189[c_long in 1.63.0]]\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +\n      +=== Rust version 1.63.0\n      +\n     -+[source]\n     -+----\n     ++```\n      +mod c_long_definition {\n      +    cfg_if! {\n      +        if #[cfg(all(target_pointer_width = \"64\", not(windows)))] {\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +        }\n      +    }\n      +}\n     -+----\n     ++```\n      +\n      +=== Rust version 1.89.0\n      +\n     -+[source]\n     -+----\n     ++```\n      +mod c_long_definition {\n      +    crate::cfg_select! {\n      +        any(\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +        }\n      +    }\n      +}\n     -+----\n     ++```\n      +\n      +Even for the cases where C types are correctly mapped to Rust types via\n      +std::ffi::c_* there are still problems. Let's take c_char for example. On some\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +The following code will panic in debug on platforms that define c_char as u8,\n      +but won't if it's an i8.\n      +\n     -+[source]\n     -+----\n     ++```\n      +let mut x: std::ffi::c_char = 0;\n      +x -= 1;\n     -+----\n     ++```\n      +\n      +=== Inconsistent shift behavior\n      +\n      +`x` will be 0xC0 for platforms that use i8, but will be 0x40 where it's u8.\n      +\n     -+[source]\n     -+----\n     ++```\n      +let mut x: std::ffi::c_char = 0x80;\n      +x >>= 1;\n     -+----\n     ++```\n      +\n      +=== Equality fails to compile on some platforms\n      +\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +if it's u8. You can cast x e.g. `assert_eq!(x as u8, b'a');`, but then you get\n      +a warning on platforms that use u8 and a clean compilation where i8 is used.\n      +\n     -+[source]\n     -+----\n     ++```\n      +let mut x: std::ffi::c_char = 0x61;\n      +assert_eq!(x, b'a');\n     -+----\n     ++```\n      +\n      +== Enum types\n      +Rust enum types should not be used as FFI types. Rust enum types are more like\n      +C union types than C enum's. For something like:\n      +\n     -+[source]\n     -+----\n     ++```\n      +#[repr(C, u8)]\n      +enum Fruit {\n      +    Apple,\n      +    Banana,\n      +    Cherry,\n      +}\n     -+----\n     ++```\n      +\n      +It's easy enough to make sure the Rust enum matches what C would expect, but a\n      +more complex type like.\n      +\n     -+[source]\n     -+----\n     ++```\n      +enum HashResult {\n      +    SHA1([u8; 20]),\n      +    SHA256([u8; 32]),\n      +}\n     -+----\n     ++```\n      +\n      +The Rust compiler has to add a discriminant to the enum to distinguish between\n      +the variants. The width, location, and values for that discriminant is up to\n  2:  9197903add !  2:  52e3f589b1 xdiff: use ssize_t for dstart/dend, make them last in xdfile_t\n     @@ Metadata\n      Author: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## Commit message ##\n     -    xdiff: use ssize_t for dstart/dend, make them last in xdfile_t\n     +    xdiff: use ptrdiff_t for dstart/dend\n      \n     -    ssize_t is appropriate for dstart and dend because they both describe\n     +    ptrdiff_t is appropriate for dstart and dend because they both describe\n          positive or negative offsets relative to a pointer.\n      \n          A future patch will move these fields to a different struct. Moving\n  3:  46bc1b3e25 !  3:  83e7bf180a xdiff: make xrecord_t.ptr a uint8_t instead of char\n     @@ Metadata\n       ## Commit message ##\n          xdiff: make xrecord_t.ptr a uint8_t instead of char\n      \n     -    Rust uses u8 to refer to bytes in memory. Since xrecord_t.ptr is also\n     -    referring to bytes in memory, rather than Unicode code points, use\n     -    uint8_t instead of char.\n     +    Make xrecord_t.ptr uint8_t because it's referring to bytes in memory.\n      \n          Every usage of this field was inspected and cast to char*, or similar,\n          to avoid signedness warnings/errors from the compiler. Casting was used\n     @@ xdiff/xprepare.c: static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, lo\n       \t\t\t\tgoto abort;\n       \t\t\tcrec = &xdf->recs[xdf->nrec++];\n      -\t\t\tcrec->ptr = prev;\n     --\t\t\tcrec->size = (long) (cur - prev);\n      +\t\t\tcrec->ptr = (uint8_t const *)prev;\n     -+\t\t\tcrec->size =(long) ( cur - prev);\n     + \t\t\tcrec->size = (long) (cur - prev);\n       \t\t\tcrec->ha = hav;\n       \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n     - \t\t\t\tgoto abort;\n      \n       ## xdiff/xtypes.h ##\n      @@ xdiff/xtypes.h: typedef struct s_chastore {\n  4:  07e28aad3b !  4:  da2b80ea0b xdiff: use size_t for xrecord_t.size\n     @@ xdiff/xprepare.c: static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, lo\n       \t\t\t\tgoto abort;\n       \t\t\tcrec = &xdf->recs[xdf->nrec++];\n       \t\t\tcrec->ptr = (uint8_t const *)prev;\n     --\t\t\tcrec->size =(long) ( cur - prev);\n     +-\t\t\tcrec->size = (long) (cur - prev);\n      +\t\t\tcrec->size = cur - prev;\n       \t\t\tcrec->ha = hav;\n       \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n  5:  1ade7d8165 =  5:  c6ba630ac5 xdiff: use unambiguous types in xdl_hash_record()\n  6:  59054ea0cb !  6:  3834ea8f9b xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n     @@ Commit message\n          xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n      \n          The ha field is serving two different purposes, which makes the code\n     -    harder to read. At first glance it looks like many places assume\n     +    harder to read. At first glance, it looks like many places assume\n          there could never be hash collisions between lines of the two input\n          files. In reality, line_hash is used together with xdl_recmatch() to\n          ensure correct comparisons of lines, even when collisions occur.\n      \n          To make this clearer, the old ha field has been split:\n     -      * line_hash: The straightforward hash of a line, requiring no\n     -        additional context.\n     +      * line_hash: a straightforward hash of a line, independent of any\n     +        external context. Its type is uint64_t, as it comes from a fixed\n     +        width hash function.\n            * minimal_perfect_hash: Not a new concept, but now a separate\n              field. It comes from the classifier's general-purpose hash table,\n              which assigns each line a unique and minimal hash across the two\n     -        files.\n     +        files. A size_t is used here because it's meant to be used to\n     +        index an array. This also this avoids ` as usize` casts on the Rust\n     +        side when using it to index a slice.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n     @@ xdiff/xpatience.c: static int match(struct hashmap *map, int line1, int line2)\n       static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n      \n       ## xdiff/xprepare.c ##\n     -@@ xdiff/xprepare.c: static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n     - \tlong hi;\n     +@@ xdiff/xprepare.c: static void xdl_free_classifier(xdlclassifier_t *cf) {\n     + \n     + \n     + static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n     +-\tlong hi;\n     ++\tsize_t hi;\n       \txdlclass_t *rcrec;\n       \n      -\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n     -+\thi = (long) XDL_HASHLONG(rec->line_hash, cf->hbits);\n     ++\thi = XDL_HASHLONG(rec->line_hash, cf->hbits);\n       \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n      -\t\tif (rcrec->rec.ha == rec->ha &&\n      +\t\tif (rcrec->rec.line_hash == rec->line_hash &&\n  7:  f91be17858 !  7:  e2a2c7530c xdiff: make xdfile_t.nrec a size_t instead of long\n     @@ Metadata\n       ## Commit message ##\n          xdiff: make xdfile_t.nrec a size_t instead of long\n      \n     -    size_t is used because nrec describes the number of elements in memory\n     -    for recs, and the number of elements in memory for 'changed' + 2.\n     +    size_t is used because nrec describes the number of elements for both\n     +    recs, and for 'changed' + 2.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n  8:  e2a6a23cc4 =  8:  31cd2a1aa4 xdiff: make xdfile_t.nreff a size_t instead of long\n  9:  3b6054945f !  9:  aee0d3958b xdiff: change rindex from long to size_t in xdfile_t\n     @@ Metadata\n       ## Commit message ##\n          xdiff: change rindex from long to size_t in xdfile_t\n      \n     -    rindex describes a index offset which means it's an index into memory\n     -    which should use size_t.\n     +    The field rindex describes an index offset for other arrays. Change it\n     +    to size_t.\n      \n          Changing the type of rindex from long to size_t has no cascading\n          refactor impact because it is only ever used to directly index other\n 10:  1856a29026 ! 10:  75c26fe160 xdiff: rename rindex -> reference_index\n     @@ Commit message\n      \n          The classic diff adds only the lines that it's going to consider,\n          during the diff, to an array. A mapping between the compacted\n     -    array, and the lines of the file that they reference, are\n     +    array, and the lines of the file that they reference, is\n          facilitated by this array.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n\n-- \ngitgitgadget\n"},{"id":"530527","messageId":"e5d084d340e874be52e7c3b056ada15ab5557877.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:23Z","receivedAt":"2025-11-11T19:42:36Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nDocument other nuances with crossing the FFI boundary. Other language\nmappings may be added in the future.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n Documentation/Makefile                        |   1 +\n Documentation/technical/meson.build           |   1 +\n .../technical/unambiguous-types.adoc          | 239 ++++++++++++++++++\n 3 files changed, 241 insertions(+)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 04e9e10b27..bc1adb2d9d 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -142,6 +142,7 @@ TECH_DOCS += technical/shallow\n TECH_DOCS += technical/sparse-checkout\n TECH_DOCS += technical/sparse-index\n TECH_DOCS += technical/trivial-merge\n+TECH_DOCS += technical/unambiguous-types\n TECH_DOCS += technical/unit-tests\n SP_ARTICLES += $(TECH_DOCS)\n SP_ARTICLES += technical/api-index\ndiff --git a/Documentation/technical/meson.build b/Documentation/technical/meson.build\nindex be698ef22a..89a6e26821 100644\n--- a/Documentation/technical/meson.build\n+++ b/Documentation/technical/meson.build\n@@ -32,6 +32,7 @@ articles = [\n   'sparse-checkout.adoc',\n   'sparse-index.adoc',\n   'trivial-merge.adoc',\n+  'unambiguous-types.adoc',\n   'unit-tests.adoc',\n ]\n \ndiff --git a/Documentation/technical/unambiguous-types.adoc b/Documentation/technical/unambiguous-types.adoc\nnew file mode 100644\nindex 0000000000..6bca39209b\n--- /dev/null\n+++ b/Documentation/technical/unambiguous-types.adoc\n@@ -0,0 +1,239 @@\n+= Unambiguous types\n+\n+Most of these mappings are obvious, but there are some nuances and gotchas with\n+Rust FFI (Foreign Function Interface).\n+\n+This document defines clear, one-to-one mappings between primitive types in C,\n+Rust (and possible other languages in the future). Its purpose is to eliminate\n+ambiguity in type widths, signedness, and binary representation across\n+platforms and languages.\n+\n+For Git, the only header required to use these unambiguous types in C is\n+`git-compat-util.h`.\n+\n+== Boolean types\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| bool^1^       | bool\n+|===\n+\n+== Integer types\n+\n+In C, `<stdint.h>` (or an equivalent) must be included.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| uint8_t    | u8\n+| uint16_t   | u16\n+| uint32_t   | u32\n+| uint64_t   | u64\n+\n+| int8_t     | i8\n+| int16_t    | i16\n+| int32_t    | i32\n+| int64_t    | i64\n+|===\n+\n+== Floating-point types\n+\n+Rust requires IEEE-754 semantics.\n+In C, that is typically true, but not guaranteed by the standard.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| float^2^      | f32\n+| double^2^     | f64\n+|===\n+\n+== Size types\n+\n+These types represent pointer-sized integers and are typically defined in\n+`<stddef.h>` or an equivalent header.\n+\n+Size types should be used any time pointer arithmetic is performed e.g.\n+indexing an array, describing the number of elements in memory, etc...\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| size_t^3^     | usize\n+| ptrdiff_t^3^  | isize\n+|===\n+\n+== Character types\n+\n+This is where C and Rust don't have a clean one-to-one mapping.\n+\n+C comparison problem: While the sign of `char` is implementation defined, it's\n+also signless (neither signed nor unsigned). When building with\n+`make DEVELOPER=1` it will complain about a \"differ in signedness\" when `char`\n+is compared with `uint8_t` or `int8_t`.\n+\n+Rust's `char` type is an unsigned 32-bit integer that is used to describe\n+Unicode code points. Even though a C `char` is the same width as `u8`, `char`\n+should be converted to u8 where it is describing bytes in memory. If a C\n+`char` is not describing bytes, then it should be converted to a more accurate\n+unambiguous type. The reason for mentioning Unicode here is because of how &str\n+is defined in Rust and how to create a &str from &[u8]. Rust assumes that &str\n+is a correctly encoded utf-8 string, i.e. text in memory. Where as a C `char`\n+makes no assumption about the bytes that it is representing.\n+\n+```\n+let raw_bytes = b\"abc\\n\";\n+let result = std::str::from_utf8(raw_bytes);\n+if let Ok(line) = result {\n+    // do something with text\n+}\n+```\n+\n+While you could specify `char` in the C code and `u8` in Rust code, it's not as\n+clear what the appropriate type is, but it would work across the FFI boundary.\n+However, the bigger problem comes from code generation tools like cbindgen and\n+bindgen. When cbindgen sees u8 in Rust it will generate uint8_t on the C side\n+which will cause differ in signedness warnings/errors. Similarly if bindgen\n+sees `char` on the C side it will generate `std::ffi::c_char` which has its own\n+problems.\n+\n+=== Notes\n+^1^ This is only true if stdbool.h (or equivalent) is used. +\n+^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the\n+platform/arch for C does not follow IEEE-754 then this equivalence does not\n+hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n+there may be a strange platform/arch where even this isn't true. +\n+^3^ C also defines uintptr_t, ssize_t and intptr_t, but these types are\n+discouraged for FFI purposes. For functions like `read()` and `write()` ssize_t\n+should be cast to a different, and unambiguous, type before being passed over\n+the FFI boundary. +\n+\n+== Problems with std::ffi::c_* types in Rust\n+TL;DR: In practice, Rust's `c_*` types aren't guaranteed to match C types for\n+all possible C compilers, platforms, or architectures, because Rust only\n+ensures correctness of C types on officially supported targets. These\n+definitions have changed over time to match more targets which means that the\n+c_* definitions will differ based on which Rust version Git chooses to use.\n+\n+Current list of safe, Rust side, FFI types in Git: +\n+\n+* `c_void`\n+* `CStr`\n+* `CString`\n+\n+Even then, they should be used sparingly, and only where the semantics match\n+exactly.\n+\n+The std::os::raw::c_* directly inherits the problems of core::ffi, which\n+changes over time and seems to make a best guess at the correct definition for\n+a given platform/target. This probably isn't a problem for all other platforms\n+that Rust supports currently, but can anyone say that Rust got it right for all\n+C compilers of all platforms/targets?\n+\n+To give an example: c_long is defined in\n+footnote:[https://doc.rust-lang.org/1.63.0/src/core/ffi/mod.rs.html#175-189[c_long in 1.63.0]]\n+footnote:[https://doc.rust-lang.org/1.89.0/src/core/ffi/primitives.rs.html#135-151[c_long in 1.89.0]]\n+\n+=== Rust version 1.63.0\n+\n+```\n+mod c_long_definition {\n+    cfg_if! {\n+        if #[cfg(all(target_pointer_width = \"64\", not(windows)))] {\n+            pub type c_long = i64;\n+            pub type NonZero_c_long = crate::num::NonZeroI64;\n+            pub type c_ulong = u64;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU64;\n+        } else {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub type c_long = i32;\n+            pub type NonZero_c_long = crate::num::NonZeroI32;\n+            pub type c_ulong = u32;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU32;\n+        }\n+    }\n+}\n+```\n+\n+=== Rust version 1.89.0\n+\n+```\n+mod c_long_definition {\n+    crate::cfg_select! {\n+        any(\n+            all(target_pointer_width = \"64\", not(windows)),\n+            // wasm32 Linux ABI uses 64-bit long\n+            all(target_arch = \"wasm32\", target_os = \"linux\")\n+        ) => {\n+            pub(super) type c_long = i64;\n+            pub(super) type c_ulong = u64;\n+        }\n+        _ => {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub(super) type c_long = i32;\n+            pub(super) type c_ulong = u32;\n+        }\n+    }\n+}\n+```\n+\n+Even for the cases where C types are correctly mapped to Rust types via\n+std::ffi::c_* there are still problems. Let's take c_char for example. On some\n+platforms it's u8 on others it's i8.\n+\n+=== Subtraction underflow in debug mode\n+\n+The following code will panic in debug on platforms that define c_char as u8,\n+but won't if it's an i8.\n+\n+```\n+let mut x: std::ffi::c_char = 0;\n+x -= 1;\n+```\n+\n+=== Inconsistent shift behavior\n+\n+`x` will be 0xC0 for platforms that use i8, but will be 0x40 where it's u8.\n+\n+```\n+let mut x: std::ffi::c_char = 0x80;\n+x >>= 1;\n+```\n+\n+=== Equality fails to compile on some platforms\n+\n+The following will not compile on platforms that define c_char as i8, but will\n+if it's u8. You can cast x e.g. `assert_eq!(x as u8, b'a');`, but then you get\n+a warning on platforms that use u8 and a clean compilation where i8 is used.\n+\n+```\n+let mut x: std::ffi::c_char = 0x61;\n+assert_eq!(x, b'a');\n+```\n+\n+== Enum types\n+Rust enum types should not be used as FFI types. Rust enum types are more like\n+C union types than C enum's. For something like:\n+\n+```\n+#[repr(C, u8)]\n+enum Fruit {\n+    Apple,\n+    Banana,\n+    Cherry,\n+}\n+```\n+\n+It's easy enough to make sure the Rust enum matches what C would expect, but a\n+more complex type like.\n+\n+```\n+enum HashResult {\n+    SHA1([u8; 20]),\n+    SHA256([u8; 32]),\n+}\n+```\n+\n+The Rust compiler has to add a discriminant to the enum to distinguish between\n+the variants. The width, location, and values for that discriminant is up to\n+the Rust compiler and is not ABI stable.\n-- \ngitgitgadget\n\n"},{"id":"530528","messageId":"52e3f589b1ce25085921453eea14b9c9d7c8f362.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 02/10] xdiff: use ptrdiff_t for dstart/dend","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:24Z","receivedAt":"2025-11-11T19:42:37Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nptrdiff_t is appropriate for dstart and dend because they both describe\npositive or negative offsets relative to a pointer.\n\nA future patch will move these fields to a different struct. Moving\nthem to the end of xdfile_t now, means the field order of xdfile_t will\nbe disturbed less.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex f145abba3e..7c8c057bca 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,10 +47,10 @@ typedef struct s_xrecord {\n typedef struct s_xdfile {\n \txrecord_t *recs;\n \tlong nrec;\n-\tlong dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n+\tptrdiff_t dstart, dend;\n } xdfile_t;\n \n typedef struct s_xdfenv {\n-- \ngitgitgadget\n\n"},{"id":"530529","messageId":"83e7bf180a380a625c9dc324333ba4a46a4c17c1.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:25Z","receivedAt":"2025-11-11T19:42:38Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake xrecord_t.ptr uint8_t because it's referring to bytes in memory.\n\nEvery usage of this field was inspected and cast to char*, or similar,\nto avoid signedness warnings/errors from the compiler. Casting was used\nso that the whole of xdiff doesn't need to be refactored in order to\nchange the type of this field.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     |  6 +++---\n xdiff/xmerge.c    | 14 +++++++-------\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xtypes.h    |  2 +-\n xdiff/xutils.c    |  4 ++--\n 7 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 6f3998ee54..411a8aa69f 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n \tint ret = 0;\n \n \tfor (i = 0; i < rec->size; i++) {\n-\t\tchar c = rec->ptr[i];\n+\t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n \t\t\treturn ret;\n@@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\n@@ -1008,7 +1008,7 @@ static int record_matches_regex(xrecord_t *rec, xpparam_t const *xpp) {\n \tsize_t i;\n \n \tfor (i = 0; i < xpp->ignore_regex_nr; i++)\n-\t\tif (!regexec_buf(xpp->ignore_regex[i], rec->ptr, rec->size, 1,\n+\t\tif (!regexec_buf(xpp->ignore_regex[i], (const char *)rec->ptr, rec->size, 1,\n \t\t\t\t &regmatch, 0))\n \t\t\treturn 1;\n \ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex b2f1f30cd3..ead930088a 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec(rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff(rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func(rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex fd600cbb5d..75cb3e76a2 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch(rec1[i].ptr, rec1[i].size,\n-\t\t\trec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch(rec1->ptr, rec1->size,\n-\t\t\t    rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n+\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n }\n \n /*\n@@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n \t\t * we have a very simple mmfile structure.\n \t\t */\n \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n-\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n+\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n \t\tt2.ptr = (char *)xe2->xdf2.recs[m->i2].ptr;\n-\t\tt2.size = xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n+\t\tt2.size = (char *)xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n \t\t\t+ xe2->xdf2.recs[m->i2 + m->chg2 - 1].size - t2.ptr;\n \t\tif (xdl_do_diff(&t1, &t2, xpp, &xe) < 0)\n \t\t\treturn -1;\n@@ -440,7 +440,7 @@ static int line_contains_alnum(const char *ptr, long size)\n static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n-\t\tif (line_contains_alnum(xe->xdf2.recs[i].ptr,\n+\t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n \t\t\t\txe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 669b653580..bb61354f22 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -121,7 +121,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t\treturn;\n \tmap->entries[index].line1 = line;\n \tmap->entries[index].hash = record->ha;\n-\tmap->entries[index].anchor = is_anchor(xpp, map->env->xdf1.recs[line - 1].ptr);\n+\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n \tif (map->last) {\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 192334f1b7..4c56467076 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch(rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\trec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = prev;\n+\t\t\tcrec->ptr = (uint8_t const *)prev;\n \t\t\tcrec->size = (long) (cur - prev);\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 7c8c057bca..b1c520a378 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -39,7 +39,7 @@ typedef struct s_chastore {\n } chastore_t;\n \n typedef struct s_xrecord {\n-\tchar const *ptr;\n+\tuint8_t const *ptr;\n \tlong size;\n \tunsigned long ha;\n } xrecord_t;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 447e66c719..7be063bfb6 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -465,10 +465,10 @@ int xdl_fall_back_diff(xdfenv_t *diff_env, xpparam_t const *xpp,\n \txdfenv_t env;\n \n \tsubfile1.ptr = (char *)diff_env->xdf1.recs[line1 - 1].ptr;\n-\tsubfile1.size = diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n+\tsubfile1.size = (char *)diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n \t\tdiff_env->xdf1.recs[line1 + count1 - 2].size - subfile1.ptr;\n \tsubfile2.ptr = (char *)diff_env->xdf2.recs[line2 - 1].ptr;\n-\tsubfile2.size = diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n+\tsubfile2.size = (char *)diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n \t\tdiff_env->xdf2.recs[line2 + count2 - 2].size - subfile2.ptr;\n \tif (xdl_do_diff(&subfile1, &subfile2, xpp, &env) < 0)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"530530","messageId":"da2b80ea0be3470cbfe04ff4d39727e6d5921a9a.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 04/10] xdiff: use size_t for xrecord_t.size","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:26Z","receivedAt":"2025-11-11T19:42:40Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is the appropriate type because size is describing the number of\nelements, bytes in this case, in memory.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  7 +++----\n xdiff/xemit.c    |  8 ++++----\n xdiff/xmerge.c   | 16 ++++++++--------\n xdiff/xprepare.c |  6 +++---\n xdiff/xtypes.h   |  2 +-\n 5 files changed, 19 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 411a8aa69f..edd05466df 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -403,10 +403,9 @@ static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n  */\n static int get_indent(xrecord_t *rec)\n {\n-\tlong i;\n \tint ret = 0;\n \n-\tfor (i = 0; i < rec->size; i++) {\n+\tfor (size_t i = 0; i < rec->size; i++) {\n \t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n@@ -993,11 +992,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex ead930088a..2f8007753c 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, (long)rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, (long)rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, (long)rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\n@@ -151,7 +151,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n static int is_empty_rec(xdfile_t *xdf, long ri)\n {\n \txrecord_t *rec = &xdf->recs[ri];\n-\tlong i = 0;\n+\tsize_t i = 0;\n \n \tfor (; i < rec->size && XDL_ISSPACE(rec->ptr[i]); i++);\n \ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 75cb3e76a2..0dd4558a32 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n-\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, (long)rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, (long)rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -119,11 +119,11 @@ static int xdl_recs_copy_0(int use_orig, xdfenv_t *xe, int i, int count, int nee\n \tif (count < 1)\n \t\treturn 0;\n \n-\tfor (i = 0; i < count; size += recs[i++].size)\n+\tfor (i = 0; i < count; size += (int)recs[i++].size)\n \t\tif (dest)\n \t\t\tmemcpy(dest + size, recs[i].ptr, recs[i].size);\n \tif (add_nl) {\n-\t\ti = recs[count - 1].size;\n+\t\ti = (int)recs[count - 1].size;\n \t\tif (i == 0 || recs[count - 1].ptr[i - 1] != '\\n') {\n \t\t\tif (needs_cr) {\n \t\t\t\tif (dest)\n@@ -156,7 +156,7 @@ static int xdl_orig_copy(xdfenv_t *xe, int i, int count, int needs_cr, int add_n\n  */\n static int is_eol_crlf(xdfile_t *file, int i)\n {\n-\tlong size;\n+\tsize_t size;\n \n \tif (i < file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n-\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, (long)rec1->size,\n+\t\t\t    (const char *)rec2->ptr, (long)rec2->size, flags);\n }\n \n /*\n@@ -441,7 +441,7 @@ static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n \t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n-\t\t\t\txe->xdf2.recs[i].size))\n+\t\t\t\t(long)xe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 4c56467076..b3219aed3e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -157,7 +157,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = (uint8_t const *)prev;\n-\t\t\tcrec->size = (long) (cur - prev);\n+\t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex b1c520a378..88b1fe4649 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -40,7 +40,7 @@ typedef struct s_chastore {\n \n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n-\tlong size;\n+\tsize_t size;\n \tunsigned long ha;\n } xrecord_t;\n \n-- \ngitgitgadget\n\n"},{"id":"530531","messageId":"c6ba630ac53ea56d567ccfabe263d2d8c31719b5.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 05/10] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:27Z","receivedAt":"2025-11-11T19:42:41Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nConvert the function signature and body to use unambiguous types. char\nis changed to uint8_t because this function processes bytes in memory.\nunsigned long to uint64_t so that the hash output is consistent across\nplatforms. `flags` was changed from long to uint64_t to ensure the\nhigh order bits are not dropped on platforms that treat long as 32\nbits.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff-interface.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xutils.c    | 28 ++++++++++++++--------------\n xdiff/xutils.h    |  6 +++---\n 4 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff-interface.c b/xdiff-interface.c\nindex 4971f722b3..1a35556380 100644\n--- a/xdiff-interface.c\n+++ b/xdiff-interface.c\n@@ -300,7 +300,7 @@ void xdiff_clear_find_func(xdemitconf_t *xecfg)\n \n unsigned long xdiff_hash_string(const char *s, size_t len, long flags)\n {\n-\treturn xdl_hash_record(&s, s + len, flags);\n+\treturn xdl_hash_record((uint8_t const**)&s, (uint8_t const*)s + len, flags);\n }\n \n int xdiff_compare_lines(const char *l1, long s1,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex b3219aed3e..85e56021da 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -137,8 +137,8 @@ static void xdl_free_ctx(xdfile_t *xdf)\n static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_t const *xpp,\n \t\t\t   xdlclassifier_t *cf, xdfile_t *xdf) {\n \tlong bsize;\n-\tunsigned long hav;\n-\tchar const *blk, *cur, *top, *prev;\n+\tuint64_t hav;\n+\tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n \txdf->rindex = NULL;\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = (uint8_t const *)prev;\n+\t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 7be063bfb6..77ee1ad9c8 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -249,11 +249,11 @@ int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags)\n \treturn 1;\n }\n \n-unsigned long xdl_hash_record_with_whitespace(char const **data,\n-\t\tchar const *top, long flags) {\n-\tunsigned long ha = 5381;\n-\tchar const *ptr = *data;\n-\tint cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data,\n+\t\tuint8_t const *top, uint64_t flags) {\n+\tuint64_t ha = 5381;\n+\tuint8_t const *ptr = *data;\n+\tbool cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n \n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tif (cr_at_eol_only) {\n@@ -263,8 +263,8 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\t\tcontinue;\n \t\t}\n \t\telse if (XDL_ISSPACE(*ptr)) {\n-\t\t\tconst char *ptr2 = ptr;\n-\t\t\tint at_eol;\n+\t\t\tconst uint8_t *ptr2 = ptr;\n+\t\t\tbool at_eol;\n \t\t\twhile (ptr + 1 < top && XDL_ISSPACE(ptr[1])\n \t\t\t\t\t&& ptr[1] != '\\n')\n \t\t\t\tptr++;\n@@ -274,20 +274,20 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_CHANGE\n \t\t\t\t && !at_eol) {\n \t\t\t\tha += (ha << 5);\n-\t\t\t\tha ^= (unsigned long) ' ';\n+\t\t\t\tha ^= (uint64_t) ' ';\n \t\t\t}\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_AT_EOL\n \t\t\t\t && !at_eol) {\n \t\t\t\twhile (ptr2 != ptr + 1) {\n \t\t\t\t\tha += (ha << 5);\n-\t\t\t\t\tha ^= (unsigned long) *ptr2;\n+\t\t\t\t\tha ^= (uint64_t) *ptr2;\n \t\t\t\t\tptr2++;\n \t\t\t\t}\n \t\t\t}\n \t\t\tcontinue;\n \t\t}\n \t\tha += (ha << 5);\n-\t\tha ^= (unsigned long) *ptr;\n+\t\tha ^= (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n \n@@ -304,9 +304,9 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n #define REASSOC_FENCE(x, y)\n #endif\n \n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n-\tunsigned long ha = 5381, c0, c1;\n-\tchar const *ptr = *data;\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top) {\n+\tuint64_t ha = 5381, c0, c1;\n+\tuint8_t const *ptr = *data;\n #if 0\n \t/*\n \t * The baseline form of the optimized loop below. This is the djb2\n@@ -314,7 +314,7 @@ unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n \t */\n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tha += (ha << 5);\n-\t\tha += (unsigned long) *ptr;\n+\t\tha += (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n #else\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 13f6831047..615b4a9d35 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -34,9 +34,9 @@ void *xdl_cha_alloc(chastore_t *cha);\n long xdl_guess_lines(mmfile_t *mf, long sample);\n int xdl_blankline(const char *line, long size, long flags);\n int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags);\n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top);\n-unsigned long xdl_hash_record_with_whitespace(char const **data, char const *top, long flags);\n-static inline unsigned long xdl_hash_record(char const **data, char const *top, long flags)\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top);\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data, uint8_t const *top, uint64_t flags);\n+static inline uint64_t xdl_hash_record(uint8_t const **data, uint8_t const *top, uint64_t flags)\n {\n \tif (flags & XDF_WHITESPACE_FLAGS)\n \t\treturn xdl_hash_record_with_whitespace(data, top, flags);\n-- \ngitgitgadget\n\n"},{"id":"530532","messageId":"3834ea8f9becc9d6e1b407679e8a95dc6c9d56de.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:28Z","receivedAt":"2025-11-11T19:42:42Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe ha field is serving two different purposes, which makes the code\nharder to read. At first glance, it looks like many places assume\nthere could never be hash collisions between lines of the two input\nfiles. In reality, line_hash is used together with xdl_recmatch() to\nensure correct comparisons of lines, even when collisions occur.\n\nTo make this clearer, the old ha field has been split:\n  * line_hash: a straightforward hash of a line, independent of any\n    external context. Its type is uint64_t, as it comes from a fixed\n    width hash function.\n  * minimal_perfect_hash: Not a new concept, but now a separate\n    field. It comes from the classifier's general-purpose hash table,\n    which assigns each line a unique and minimal hash across the two\n    files. A size_t is used here because it's meant to be used to\n    index an array. This also this avoids ` as usize` casts on the Rust\n    side when using it to index a slice.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c     |  6 +++---\n xdiff/xhistogram.c |  4 ++--\n xdiff/xpatience.c  | 10 +++++-----\n xdiff/xprepare.c   | 18 +++++++++---------\n xdiff/xtypes.h     |  3 ++-\n 5 files changed, 21 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex edd05466df..436c34697d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -22,9 +22,9 @@\n \n #include \"xinclude.h\"\n \n-static unsigned long get_hash(xdfile_t *xdf, long index)\n+static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].ha;\n+\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -385,7 +385,7 @@ static xdchange_t *xdl_add_change(xdchange_t *xscr, long i1, long i2, long chg1,\n \n static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n {\n-\treturn (rec1->ha == rec2->ha);\n+\treturn rec1->minimal_perfect_hash == rec2->minimal_perfect_hash;\n }\n \n /*\ndiff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\nindex 6dc450b1fe..5ae1282c27 100644\n--- a/xdiff/xhistogram.c\n+++ b/xdiff/xhistogram.c\n@@ -90,7 +90,7 @@ struct region {\n \n static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n {\n-\treturn r1->ha == r2->ha;\n+\treturn r1->minimal_perfect_hash == r2->minimal_perfect_hash;\n \n }\n \n@@ -98,7 +98,7 @@ static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n \t(cmp_recs(REC(i->env, s1, l1), REC(i->env, s2, l2)))\n \n #define TABLE_HASH(index, side, line) \\\n-\tXDL_HASHLONG((REC(index->env, side, line))->ha, index->table_bits)\n+\tXDL_HASHLONG((REC(index->env, side, line))->minimal_perfect_hash, index->table_bits)\n \n static int scanA(struct histindex *index, int line1, int count1)\n {\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex bb61354f22..cc53266f3b 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -48,7 +48,7 @@\n struct hashmap {\n \tint nr, alloc;\n \tstruct entry {\n-\t\tunsigned long hash;\n+\t\tsize_t minimal_perfect_hash;\n \t\t/*\n \t\t * 0 = unused entry, 1 = first line, 2 = second, etc.\n \t\t * line2 is NON_UNIQUE if the line is not unique\n@@ -101,10 +101,10 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t * So we multiply ha by 2 in the hope that the hashing was\n \t * \"unique enough\".\n \t */\n-\tint index = (int)((record->ha << 1) % map->alloc);\n+\tint index = (int)((record->minimal_perfect_hash << 1) % map->alloc);\n \n \twhile (map->entries[index].line1) {\n-\t\tif (map->entries[index].hash != record->ha) {\n+\t\tif (map->entries[index].minimal_perfect_hash != record->minimal_perfect_hash) {\n \t\t\tif (++index >= map->alloc)\n \t\t\t\tindex = 0;\n \t\t\tcontinue;\n@@ -120,7 +120,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \tif (pass == 2)\n \t\treturn;\n \tmap->entries[index].line1 = line;\n-\tmap->entries[index].hash = record->ha;\n+\tmap->entries[index].minimal_perfect_hash = record->minimal_perfect_hash;\n \tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n@@ -248,7 +248,7 @@ static int match(struct hashmap *map, int line1, int line2)\n {\n \txrecord_t *record1 = &map->env->xdf1.recs[line1 - 1];\n \txrecord_t *record2 = &map->env->xdf2.recs[line2 - 1];\n-\treturn record1->ha == record2->ha;\n+\treturn record1->minimal_perfect_hash == record2->minimal_perfect_hash;\n }\n \n static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 85e56021da..bea0992b5e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -93,12 +93,12 @@ static void xdl_free_classifier(xdlclassifier_t *cf) {\n \n \n static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n-\tlong hi;\n+\tsize_t hi;\n \txdlclass_t *rcrec;\n \n-\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n+\thi = XDL_HASHLONG(rec->line_hash, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n-\t\tif (rcrec->rec.ha == rec->ha &&\n+\t\tif (rcrec->rec.line_hash == rec->line_hash &&\n \t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n \t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n@@ -120,7 +120,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n \t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n \n-\trec->ha = (unsigned long) rcrec->idx;\n+\trec->minimal_perfect_hash = (size_t)rcrec->idx;\n \n \treturn 0;\n }\n@@ -158,7 +158,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n-\t\t\tcrec->ha = hav;\n+\t\t\tcrec->line_hash = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\n \t\t}\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -298,7 +298,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -350,7 +350,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs2 = xdf2->recs;\n \tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dstart = xdf2->dstart = i;\n@@ -358,7 +358,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs1 = xdf1->recs + xdf1->nrec - 1;\n \trecs2 = xdf2->recs + xdf2->nrec - 1;\n \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dend = xdf1->nrec - i - 1;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 88b1fe4649..742b81bf3b 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -41,7 +41,8 @@ typedef struct s_chastore {\n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n \tsize_t size;\n-\tunsigned long ha;\n+\tuint64_t line_hash;\n+\tsize_t minimal_perfect_hash;\n } xrecord_t;\n \n typedef struct s_xdfile {\n-- \ngitgitgadget\n\n"},{"id":"530533","messageId":"e2a2c7530cd4c1ea16f39cc4f1ad9a6ec293d7c5.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 07/10] xdiff: make xdfile_t.nrec a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:29Z","receivedAt":"2025-11-11T19:42:43Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nrec describes the number of elements for both\nrecs, and for 'changed' + 2.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     | 20 ++++++++++----------\n xdiff/xmerge.c    |  8 ++++----\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  | 12 ++++++------\n xdiff/xtypes.h    |  2 +-\n 6 files changed, 26 insertions(+), 26 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 436c34697d..759193fe5d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -483,7 +483,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n {\n \tlong i;\n \n-\tif (split >= xdf->nrec) {\n+\tif (split >= (long)xdf->nrec) {\n \t\tm->end_of_file = 1;\n \t\tm->indent = -1;\n \t} else {\n@@ -506,7 +506,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n \n \tm->post_blank = 0;\n \tm->post_indent = -1;\n-\tfor (i = split + 1; i < xdf->nrec; i++) {\n+\tfor (i = split + 1; i < (long)xdf->nrec; i++) {\n \t\tm->post_indent = get_indent(&xdf->recs[i]);\n \t\tif (m->post_indent != -1)\n \t\t\tbreak;\n@@ -717,7 +717,7 @@ static void group_init(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static inline int group_next(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end == xdf->nrec)\n+\tif (g->end == (long)xdf->nrec)\n \t\treturn -1;\n \n \tg->start = g->end + 1;\n@@ -750,7 +750,7 @@ static inline int group_previous(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static int group_slide_down(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end < xdf->nrec &&\n+\tif (g->end < (long)xdf->nrec &&\n \t    recs_match(&xdf->recs[g->start], &xdf->recs[g->end])) {\n \t\txdf->changed[g->start++] = false;\n \t\txdf->changed[g->end++] = true;\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex 2f8007753c..04f7e9193b 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -137,7 +137,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n \tbuf = func_line ? func_line->buf : dummy;\n \tsize = func_line ? sizeof(func_line->buf) : sizeof(dummy);\n \n-\tfor (l = start; l != limit && 0 <= l && l < xe->xdf1.nrec; l += step) {\n+\tfor (l = start; l != limit && 0 <= l && l < (long)xe->xdf1.nrec; l += step) {\n \t\tlong len = match_func_rec(&xe->xdf1, xecfg, l, buf, size);\n \t\tif (len >= 0) {\n \t\t\tif (func_line)\n@@ -179,14 +179,14 @@ pre_context_calculation:\n \t\t\tlong fs1, i1 = xch->i1;\n \n \t\t\t/* Appended chunk? */\n-\t\t\tif (i1 >= xe->xdf1.nrec) {\n+\t\t\tif (i1 >= (long)xe->xdf1.nrec) {\n \t\t\t\tlong i2 = xch->i2;\n \n \t\t\t\t/*\n \t\t\t\t * We don't need additional context if\n \t\t\t\t * a whole function was added.\n \t\t\t\t */\n-\t\t\t\twhile (i2 < xe->xdf2.nrec) {\n+\t\t\t\twhile (i2 < (long)xe->xdf2.nrec) {\n \t\t\t\t\tif (is_func_rec(&xe->xdf2, xecfg, i2))\n \t\t\t\t\t\tgoto post_context_calculation;\n \t\t\t\t\ti2++;\n@@ -196,7 +196,7 @@ pre_context_calculation:\n \t\t\t\t * Otherwise get more context from the\n \t\t\t\t * pre-image.\n \t\t\t\t */\n-\t\t\t\ti1 = xe->xdf1.nrec - 1;\n+\t\t\t\ti1 = (long)xe->xdf1.nrec - 1;\n \t\t\t}\n \n \t\t\tfs1 = get_func_line(xe, xecfg, NULL, i1, -1);\n@@ -228,8 +228,8 @@ pre_context_calculation:\n \n  post_context_calculation:\n \t\tlctx = xecfg->ctxlen;\n-\t\tlctx = XDL_MIN(lctx, xe->xdf1.nrec - (xche->i1 + xche->chg1));\n-\t\tlctx = XDL_MIN(lctx, xe->xdf2.nrec - (xche->i2 + xche->chg2));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf1.nrec - (xche->i1 + xche->chg1));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf2.nrec - (xche->i2 + xche->chg2));\n \n \t\te1 = xche->i1 + xche->chg1 + lctx;\n \t\te2 = xche->i2 + xche->chg2 + lctx;\n@@ -237,13 +237,13 @@ pre_context_calculation:\n \t\tif (xecfg->flags & XDL_EMIT_FUNCCONTEXT) {\n \t\t\tlong fe1 = get_func_line(xe, xecfg, NULL,\n \t\t\t\t\t\t xche->i1 + xche->chg1,\n-\t\t\t\t\t\t xe->xdf1.nrec);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec);\n \t\t\twhile (fe1 > 0 && is_empty_rec(&xe->xdf1, fe1 - 1))\n \t\t\t\tfe1--;\n \t\t\tif (fe1 < 0)\n-\t\t\t\tfe1 = xe->xdf1.nrec;\n+\t\t\t\tfe1 = (long)xe->xdf1.nrec;\n \t\t\tif (fe1 > e1) {\n-\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), xe->xdf2.nrec);\n+\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), (long)xe->xdf2.nrec);\n \t\t\t\te1 = fe1;\n \t\t\t}\n \n@@ -254,7 +254,7 @@ pre_context_calculation:\n \t\t\t */\n \t\t\tif (xche->next) {\n \t\t\t\tlong l = XDL_MIN(xche->next->i1,\n-\t\t\t\t\t\t xe->xdf1.nrec - 1);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec - 1);\n \t\t\t\tif (l - xecfg->ctxlen <= e1 ||\n \t\t\t\t    get_func_line(xe, xecfg, NULL, l, e1) < 0) {\n \t\t\t\t\txche = xche->next;\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 0dd4558a32..29dad98c49 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -158,7 +158,7 @@ static int is_eol_crlf(xdfile_t *file, int i)\n {\n \tsize_t size;\n \n-\tif (i < file->nrec - 1)\n+\tif (i < (long)file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n \t\treturn (size = file->recs[i].size) > 1 &&\n \t\t\tfile->recs[i].ptr[size - 2] == '\\r';\n@@ -317,7 +317,7 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \t\t\tcontinue;\n \t\ti = m->i1 + m->chg1;\n \t}\n-\tsize += xdl_recs_copy(xe1, i, xe1->xdf2.nrec - i, 0, 0,\n+\tsize += xdl_recs_copy(xe1, i, (int)xe1->xdf2.nrec - i, 0, 0,\n \t\t\t      dest ? dest + size : NULL);\n \treturn size;\n }\n@@ -622,7 +622,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\t\tchanges = c;\n \t\ti0 = xscr1->i1;\n \t\ti1 = xscr1->i2;\n-\t\ti2 = xscr1->i1 + xe2->xdf2.nrec - xe2->xdf1.nrec;\n+\t\ti2 = xscr1->i1 + (long)xe2->xdf2.nrec - (long)xe2->xdf1.nrec;\n \t\tchg0 = xscr1->chg1;\n \t\tchg1 = xscr1->chg2;\n \t\tchg2 = xscr1->chg1;\n@@ -637,7 +637,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\tif (!changes)\n \t\t\tchanges = c;\n \t\ti0 = xscr2->i1;\n-\t\ti1 = xscr2->i1 + xe1->xdf2.nrec - xe1->xdf1.nrec;\n+\t\ti1 = xscr2->i1 + (long)xe1->xdf2.nrec - (long)xe1->xdf1.nrec;\n \t\ti2 = xscr2->i2;\n \t\tchg0 = xscr2->chg1;\n \t\tchg1 = xscr2->chg1;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex cc53266f3b..a0b31eb5d8 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -370,5 +370,5 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n \n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n-\treturn patience_diff(xpp, env, 1, env->xdf1.nrec, 1, env->xdf2.nrec);\n+\treturn patience_diff(xpp, env, 1, (int)env->xdf1.nrec, 1, (int)env->xdf2.nrec);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex bea0992b5e..705ddd1ae0 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -153,7 +153,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\tfor (top = blk + bsize; cur < top; ) {\n \t\t\tprev = cur;\n \t\t\thav = xdl_hash_record(&cur, top, xpp->flags);\n-\t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n+\t\t\tif (XDL_ALLOC_GROW(xdf->recs, (long)xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n@@ -287,7 +287,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -295,7 +295,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -348,7 +348,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \n \trecs1 = xdf1->recs;\n \trecs2 = xdf2->recs;\n-\tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n+\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n@@ -361,8 +361,8 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n-\txdf1->dend = xdf1->nrec - i - 1;\n-\txdf2->dend = xdf2->nrec - i - 1;\n+\txdf1->dend = (long)xdf1->nrec - i - 1;\n+\txdf2->dend = (long)xdf2->nrec - i - 1;\n \n \treturn 0;\n }\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 742b81bf3b..17cafd8b6e 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,7 +47,7 @@ typedef struct s_xrecord {\n \n typedef struct s_xdfile {\n \txrecord_t *recs;\n-\tlong nrec;\n+\tsize_t nrec;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n-- \ngitgitgadget\n\n"},{"id":"530534","messageId":"31cd2a1aa4c4205f0875f5608013b3ef9adf7984.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 08/10] xdiff: make xdfile_t.nreff a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:30Z","receivedAt":"2025-11-11T19:42:44Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nreff describes the number of elements in memory\nfor rindex.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 +++++++-------\n xdiff/xtypes.h   |  2 +-\n 2 files changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 705ddd1ae0..39fd79d9d4 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -264,7 +264,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, nreff, mlim;\n+\tlong i, nm, mlim;\n \txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n@@ -307,29 +307,29 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * Use temporary arrays to decide if changed[i] should remain\n \t * false, or become true.\n \t */\n-\tfor (nreff = 0, i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n+\txdf1->nreff = 0;\n+\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[nreff++] = i;\n+\t\t\txdf1->rindex[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf1->nreff = nreff;\n \n-\tfor (nreff = 0, i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n+\txdf2->nreff = 0;\n+\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[nreff++] = i;\n+\t\t\txdf2->rindex[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf2->nreff = nreff;\n \n cleanup:\n \txdl_free(action1);\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 17cafd8b6e..df4c5cab1a 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -50,7 +50,7 @@ typedef struct s_xdfile {\n \tsize_t nrec;\n \tbool *changed;\n \tlong *rindex;\n-\tlong nreff;\n+\tsize_t nreff;\n \tptrdiff_t dstart, dend;\n } xdfile_t;\n \n-- \ngitgitgadget\n\n"},{"id":"530535","messageId":"aee0d3958b8b72a93773466d1c1ff386a3ac4872.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 09/10] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:31Z","receivedAt":"2025-11-11T19:42:45Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe field rindex describes an index offset for other arrays. Change it\nto size_t.\n\nChanging the type of rindex from long to size_t has no cascading\nrefactor impact because it is only ever used to directly index other\narrays.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex df4c5cab1a..3bcc0920e0 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -49,7 +49,7 @@ typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n \tbool *changed;\n-\tlong *rindex;\n+\tsize_t *rindex;\n \tsize_t nreff;\n \tptrdiff_t dstart, dend;\n } xdfile_t;\n-- \ngitgitgadget\n\n"},{"id":"530536","messageId":"75c26fe16049122f35b4fffb15f15429ae55f8e7.1762890152.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v3 10/10] xdiff: rename rindex -> reference_index","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-11T19:42:32Z","receivedAt":"2025-11-11T19:42:46Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe classic diff adds only the lines that it's going to consider,\nduring the diff, to an array. A mapping between the compacted\narray, and the lines of the file that they reference, is\nfacilitated by this array.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  6 +++---\n xdiff/xprepare.c | 10 +++++-----\n xdiff/xtypes.h   |  2 +-\n 3 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 759193fe5d..8eb664be3e 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -24,7 +24,7 @@\n \n static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n+\treturn xdf->recs[xdf->reference_index[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -278,10 +278,10 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n \t */\n \tif (off1 == lim1) {\n \t\tfor (; off2 < lim2; off2++)\n-\t\t\txdf2->changed[xdf2->rindex[off2]] = true;\n+\t\t\txdf2->changed[xdf2->reference_index[off2]] = true;\n \t} else if (off2 == lim2) {\n \t\tfor (; off1 < lim1; off1++)\n-\t\t\txdf1->changed[xdf1->rindex[off1]] = true;\n+\t\t\txdf1->changed[xdf1->reference_index[off1]] = true;\n \t} else {\n \t\txdpsplit_t spl;\n \t\tspl.i1 = spl.i2 = 0;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 39fd79d9d4..34c82e4f8e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -128,7 +128,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n static void xdl_free_ctx(xdfile_t *xdf)\n {\n-\txdl_free(xdf->rindex);\n+\txdl_free(xdf->reference_index);\n \txdl_free(xdf->changed - 1);\n \txdl_free(xdf->recs);\n }\n@@ -141,7 +141,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n-\txdf->rindex = NULL;\n+\txdf->reference_index = NULL;\n \txdf->changed = NULL;\n \txdf->recs = NULL;\n \n@@ -169,7 +169,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF)) {\n-\t\tif (!XDL_ALLOC_ARRAY(xdf->rindex, xdf->nrec + 1))\n+\t\tif (!XDL_ALLOC_ARRAY(xdf->reference_index, xdf->nrec + 1))\n \t\t\tgoto abort;\n \t}\n \n@@ -312,7 +312,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[xdf1->nreff++] = i;\n+\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n@@ -324,7 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[xdf2->nreff++] = i;\n+\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 3bcc0920e0..5accbec284 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -49,7 +49,7 @@ typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n \tbool *changed;\n-\tsize_t *rindex;\n+\tsize_t *reference_index;\n \tsize_t nreff;\n \tptrdiff_t dstart, dend;\n } xdfile_t;\n-- \ngitgitgadget\n"},{"id":"530543","messageId":"xmqq7bvwwauh.fsf@gitster.g","threadId":"64326","inReplyTo":"e5d084d340e874be52e7c3b056ada15ab5557877.1762890152.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T20:52:38Z","receivedAt":"2025-11-11T20:52:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +== Character types\n> +\n> +This is where C and Rust don't have a clean one-to-one mapping.\n> +\n> +C comparison problem: While the sign of `char` is implementation defined, it's\n> +also signless (neither signed nor unsigned). When building with\n> +`make DEVELOPER=1` it will complain about a \"differ in signedness\" when `char`\n> +is compared with `uint8_t` or `int8_t`.\n> +\n> +Rust's `char` type is an unsigned 32-bit integer that is used to describe\n> +Unicode code points. Even though a C `char` is the same width as `u8`, `char`\n> +should be converted to u8 where it is describing bytes in memory. If a C\n> +`char` is not describing bytes, then it should be converted to a more accurate\n> +unambiguous type. The reason for mentioning Unicode here is because of how &str\n> +is defined in Rust and how to create a &str from &[u8]. Rust assumes that &str\n> +is a correctly encoded utf-8 string, i.e. text in memory. Where as a C `char`\n> +makes no assumption about the bytes that it is representing.\n\nEven though you write excuses for bringing up Unicode here, I am\nafraid that most of the above is irrelevant tangent that makes the\npoint of this documentation muddier.  Anybody who is involved in\nthis effort would at least know that C's char is not about\nrepresenting Unicode codepoints (it is way too narrow for that),\nwhile Rust's char type exactly is, and I do not see much point in\nmaking such an apples-and-oranges comparison to spend extra words\nhere.\n\nAnother thing I found confusing is your mention of &[u8] vs &str.\nSurely, Rust will have trouble if an array of u8 we FFI an array of\nbytes we have on the C side, if the byte sequence were a broken\nUTF-8.  But that would not be fixed if you only rewrote C code to\nuse `uint8_t[]` where it originally used `char[]`, would it?  If we\nhave on C-side char[] that has iso8859-1 in it, we still would want\nto use uint8_t[] when we smuggle the result of passing it to iconv()\nto translate that into UTF-8 into Rust.  Or we may pass such an\niso8859-1 encoded string directly as an uint8_t[] byte array to Rust\nand let Rust side run an equivalent of iconv() to obtain char array.\n\nThe point is that \"your byte sequence has to be valid UTF-8\" does\nnot fit well in the narrative here.  If we want to move/interface\nthe handling of \"encoding\" header in commit objects with code\nwritten in Rust, this starts to matter.\n\nSo even if it is technically correct, it is another irrelevant\ntangent when we discuss why we want to use uint8_t on the C side to\nhelp cbindgen/bindgen to map it to u8 on Rust side.\n\nWouldn't just directly going into\n\n    If a piece of C code uses `char` to represent a byte, it makes\n    it easier to interface with Rust to rewrite it to use uint8_t\n    and let cbindgen/bindgen map it to u8 on the Rust side.\n\nbe clearer, would it?  We never deal with a single Unicode codepoint\nor an array of them (we do deal with utf8 encoded array of bytes,\nthough) on the C side, and I do not think it is likely to change, so\nthere is nothing lost if we did not talk about how `char` in Rust\nbehaves at all.\n\nAnd of course, not talking about `char` in Rust does not mean that\nwe need a rule like \"if you want to interface with C, never use\n`char` on the Rust side\".  `char` may have its uses on Rust side,\njust like `char` may have its uses on C side.\n\nAlso I do not quite get your precondition \"If a C `char` is not\ndescribing bytes\".  What `char` in C on modern platforms would\ndescribe something _other_ _than_ bytes?  Even the way things like\nvarint use `char` is exactly for accessing individual bytes.  Even\nwhen it is used as a space-saver in a structure member whose value\nwould never exceed 100, i.e., a small integer, we would know and be\nimplicitly relying on the fact that the member is a byte-wide.\n\nThanks.\n"},{"id":"530544","messageId":"xmqq1pm4wa8o.fsf@gitster.g","threadId":"64326","inReplyTo":"e5d084d340e874be52e7c3b056ada15ab5557877.1762890152.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T21:05:43Z","receivedAt":"2025-11-11T21:05:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/Documentation/technical/unambiguous-types.adoc b/Documentation/technical/unambiguous-types.adoc\n> new file mode 100644\n> index 0000000000..6bca39209b\n> --- /dev/null\n> +++ b/Documentation/technical/unambiguous-types.adoc\n> @@ -0,0 +1,239 @@\n> += Unambiguous types\n> +\n> +Most of these mappings are obvious, but there are some nuances and gotchas with\n> +Rust FFI (Foreign Function Interface).\n> +\n> +This document defines clear, one-to-one mappings between primitive types in C,\n> +Rust (and possible other languages in the future). Its purpose is to eliminate\n> +ambiguity in type widths, signedness, and binary representation across\n> +platforms and languages.\n\nThis is a laudable goal.  It does a lot more than \"to eliminate\nambiguity\" at least in some sections.  The section on character\ntypes I already commented on, for example, is full of good points to\nlist the concerns that developers need to be aware of and careful\nabout.\n\n> +== Enum types\n> +Rust enum types should not be used as FFI types. Rust enum types are more like\n> +C union types than C enum's. For something like:\n> +\n> +```\n> +#[repr(C, u8)]\n> +enum Fruit {\n> +    Apple,\n> +    Banana,\n> +    Cherry,\n> +}\n> +```\n> +\n> +It's easy enough to make sure the Rust enum matches what C would expect, but a\n> +more complex type like.\n> +\n> +```\n> +enum HashResult {\n> +    SHA1([u8; 20]),\n> +    SHA256([u8; 32]),\n> +}\n> +```\n> +\n> +The Rust compiler has to add a discriminant to the enum to distinguish between\n> +the variants. The width, location, and values for that discriminant is up to\n> +the Rust compiler and is not ABI stable.\n\nGood example, as we already do use this one, if I am not mistaken\n;-)\n"},{"id":"530549","messageId":"xmqqms4sus2w.fsf@gitster.g","threadId":"64326","inReplyTo":"52e3f589b1ce25085921453eea14b9c9d7c8f362.1762890152.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 02/10] xdiff: use ptrdiff_t for dstart/dend","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T22:23:19Z","receivedAt":"2025-11-11T22:23:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> ptrdiff_t is appropriate for dstart and dend because they both describe\n> positive or negative offsets relative to a pointer.\n\nMakes sense.\n\n> A future patch will move these fields to a different struct. Moving\n> them to the end of xdfile_t now, means the field order of xdfile_t will\n> be disturbed less.\n\nIf these members will be gone from this struct, it wouldn't make any\ndifference in the end.  I am not sure what you mean by \"disturbed\nless\".  Right now there is a gap between changed and nrec members,\nand at some later point, these two members may be adjacent with each\nother.  I do not think it would make that much difference if they\nbecome adjacent after this step [02/10], after step [10/10], or in a\nseparate series (xdiff-cleanup-3?).\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  xdiff/xtypes.h | 2 +-\n>  1 file changed, 1 insertion(+), 1 deletion(-)\n>\n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index f145abba3e..7c8c057bca 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -47,10 +47,10 @@ typedef struct s_xrecord {\n>  typedef struct s_xdfile {\n>  \txrecord_t *recs;\n>  \tlong nrec;\n> -\tlong dstart, dend;\n>  \tbool *changed;\n>  \tlong *rindex;\n>  \tlong nreff;\n> +\tptrdiff_t dstart, dend;\n>  } xdfile_t;\n>  \n>  typedef struct s_xdfenv {\n"},{"id":"530552","messageId":"xmqqbjl8uqp6.fsf@gitster.g","threadId":"64326","inReplyTo":"83e7bf180a380a625c9dc324333ba4a46a4c17c1.1762890152.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T22:53:09Z","receivedAt":"2025-11-11T22:53:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> Make xrecord_t.ptr uint8_t because it's referring to bytes in memory.\n>\n> Every usage of this field was inspected and cast to char*, or similar,\n\n\"inspected and changed to cast to\"?\n\n> to avoid signedness warnings/errors from the compiler. Casting was used\n> so that the whole of xdiff doesn't need to be refactored in order to\n> change the type of this field.\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  xdiff/xdiffi.c    |  8 ++++----\n>  xdiff/xemit.c     |  6 +++---\n>  xdiff/xmerge.c    | 14 +++++++-------\n>  xdiff/xpatience.c |  2 +-\n>  xdiff/xprepare.c  |  6 +++---\n>  xdiff/xtypes.h    |  2 +-\n>  xdiff/xutils.c    |  4 ++--\n>  7 files changed, 21 insertions(+), 21 deletions(-)\n>\n> diff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\n> index 6f3998ee54..411a8aa69f 100644\n> --- a/xdiff/xdiffi.c\n> +++ b/xdiff/xdiffi.c\n> @@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n>  \tint ret = 0;\n>  \n>  \tfor (i = 0; i < rec->size; i++) {\n> -\t\tchar c = rec->ptr[i];\n> +\t\tuint8_t c = rec->ptr[i];\n\nrec->ptr[] is now an array of uint8_t, so this is not \"inspected and\ncast to\".  It is unclear from limited context lines how 'c' is used\nby the existing code, but one example here ...\n\n>  \t\tif (!XDL_ISSPACE(c))\n>  \t\t\treturn ret;\n\n... in the post context assumes that XDL_ISSPACE(), which was\ndesigned to work with `char` (of implementation-defined signedness),\nwould safely accept an `unsigned char` (let's admit it; for all\npractical purposes, uint8_t is equivalent to unsigned char while we\nare looking at C code) so the updated code should work fine.\n\nThe definition of XDL_ISSPACE(c) indeed casts `c` to \"unsigned char\"\nas the first thing it does, and other tests in this if/else if\ncascade (hidden in the post context of this hunk) are equality\ncomparisons with ' ' and '\\t', so this conversion is safe.\n\nEither way would work so it is a minor point, but instead of\nchanging type of `c` to u8 than casting it to `char`, as the\nproposed log message explained, i.e.,\n\n\t\tchar c = (char)rec->ptr[i];\n\nwould have been much easier to reason about why this code after the\npatch is still correct.\n\n> @@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n>  \t\t * we have a very simple mmfile structure.\n>  \t\t */\n>  \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n> -\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n> +\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n>  \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n\nThe ptr member in the t1 and t2 struct is still of type (char *), so\nthe size computation is performed as ptrdiff between two (char *),\nwhich makes sense.\n\n> @@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n>  \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n>  \t\t\t\tgoto abort;\n>  \t\t\tcrec = &xdf->recs[xdf->nrec++];\n> -\t\t\tcrec->ptr = prev;\n> +\t\t\tcrec->ptr = (uint8_t const *)prev;\n>  \t\t\tcrec->size = (long) (cur - prev);\n\nHmm, it is tempting to fix this while at it, but I guess the \".size\"\nmember being \"long\" will be updated to use ptrdiff_t or something\nmore appropriate in a later step.\n\nLooking sensible.\n"},{"id":"530553","messageId":"xmqq346kupzm.fsf@gitster.g","threadId":"64326","inReplyTo":"da2b80ea0be3470cbfe04ff4d39727e6d5921a9a.1762890152.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 04/10] xdiff: use size_t for xrecord_t.size","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T23:08:29Z","receivedAt":"2025-11-11T23:08:31Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> size_t is the appropriate type because size is describing the number of\n> elements, bytes in this case, in memory.\n>\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  xdiff/xdiffi.c   |  7 +++----\n>  xdiff/xemit.c    |  8 ++++----\n>  xdiff/xmerge.c   | 16 ++++++++--------\n>  xdiff/xprepare.c |  6 +++---\n>  xdiff/xtypes.h   |  2 +-\n>  5 files changed, 19 insertions(+), 20 deletions(-)\n\nThis step looks mostly OK but it is messy in some places.\n\n> diff --git a/xdiff/xemit.c b/xdiff/xemit.c\n> index ead930088a..2f8007753c 100644\n> --- a/xdiff/xemit.c\n> +++ b/xdiff/xemit.c\n> @@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n>  {\n>  \txrecord_t *rec = &xdf->recs[ri];\n>  \n> -\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n> +\tif (xdl_emit_diffrec((char const *)rec->ptr, (long)rec->size, pre, strlen(pre), ecb) < 0)\n\nOn platforms where long is narrower than size_t, we'd tentatively\nleave things broken until we update xdl_emit_diffrec() to take\nsize_t, as it would become too noisy to change it in the same patch,\nI guess?\n\n> @@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n>  \txrecord_t *rec = &xdf->recs[ri];\n>  \n>  \tif (!xecfg->find_func)\n> -\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n> -\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n> +\t\treturn def_ff((const char *)rec->ptr, (long)rec->size, buf, sz);\n> +\treturn xecfg->find_func((const char *)rec->ptr, (long)rec->size, buf, sz, xecfg->find_func_priv);\n\nDitto.\n\n> diff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\n> index 75cb3e76a2..0dd4558a32 100644\n> --- a/xdiff/xmerge.c\n> +++ b/xdiff/xmerge.c\n> @@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n>  \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n>  \n>  \tfor (i = 0; i < line_count; i++) {\n> -\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n> -\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n> +\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, (long)rec1[i].size,\n> +\t\t\t(const char *)rec2[i].ptr, (long)rec2[i].size, flags);\n\nDitto.\n\n> @@ -119,11 +119,11 @@ static int xdl_recs_copy_0(int use_orig, xdfenv_t *xe, int i, int count, int nee\n>  \tif (count < 1)\n>  \t\treturn 0;\n>  \n> -\tfor (i = 0; i < count; size += recs[i++].size)\n> +\tfor (i = 0; i < count; size += (int)recs[i++].size)\n>  \t\tif (dest)\n>  \t\t\tmemcpy(dest + size, recs[i].ptr, recs[i].size);\n>  \tif (add_nl) {\n> -\t\ti = recs[count - 1].size;\n> +\t\ti = (int)recs[count - 1].size;\n>  \t\tif (i == 0 || recs[count - 1].ptr[i - 1] != '\\n') {\n>  \t\t\tif (needs_cr) {\n>  \t\t\t\tif (dest)\n\nThis is messier than I expected.  Before the precontext of this\nhunk, \"i\" and \"count\" are both incoming parameters of type \"int\", so\nthe same \"what if size_t is wider?\" puzzlement applies here.  At\nleast, the reason why \"i\" and \"count\" is \"int\" is not because they\nwant to be able to express negative values, so it shouldn't involve\ntoo much hassle if we later want to change them to size_t to lose\nthese casts.\n\n> @@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n>  \n>  static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n>  {\n> -\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n> -\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n> +\treturn xdl_recmatch((const char *)rec1->ptr, (long)rec1->size,\n> +\t\t\t    (const char *)rec2->ptr, (long)rec2->size, flags);\n>  }\n\nSame \"long may not be wide enough, in which case we'd need further\nfixes\" applies here.\n\n> @@ -441,7 +441,7 @@ static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n>  {\n>  \tfor (; chg; chg--, i++)\n>  \t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n> -\t\t\t\txe->xdf2.recs[i].size))\n> +\t\t\t\t(long)xe->xdf2.recs[i].size))\n>  \t\t\treturn 1;\n>  \treturn 0;\n>  }\n\nDitto.\n\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 4c56467076..b3219aed3e 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n>  \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n>  \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n>  \t\tif (rcrec->rec.ha == rec->ha &&\n> -\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n> -\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n> +\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n> +\t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n\nDitto.\n\n> @@ -157,7 +157,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n>  \t\t\t\tgoto abort;\n>  \t\t\tcrec = &xdf->recs[xdf->nrec++];\n>  \t\t\tcrec->ptr = (uint8_t const *)prev;\n> -\t\t\tcrec->size = (long) (cur - prev);\n> +\t\t\tcrec->size = cur - prev;\n\nYay!\n\n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index b1c520a378..88b1fe4649 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -40,7 +40,7 @@ typedef struct s_chastore {\n>  \n>  typedef struct s_xrecord {\n>  \tuint8_t const *ptr;\n> -\tlong size;\n> +\tsize_t size;\n\nYay, too!\n\n>  \tunsigned long ha;\n>  } xrecord_t;\n\n"},{"id":"530554","messageId":"xmqqwm3wtat8.fsf@gitster.g","threadId":"64326","inReplyTo":"3834ea8f9becc9d6e1b407679e8a95dc6c9d56de.1762890152.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T23:21:39Z","receivedAt":"2025-11-11T23:21:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> To make this clearer, the old ha field has been split:\n>   * line_hash: a straightforward hash of a line, independent of any\n>     external context. Its type is uint64_t, as it comes from a fixed\n>     width hash function.\n>   * minimal_perfect_hash: Not a new concept, but now a separate\n>     field. It comes from the classifier's general-purpose hash table,\n>     which assigns each line a unique and minimal hash across the two\n>     files. A size_t is used here because it's meant to be used to\n>     index an array. This also this avoids ` as usize` casts on the Rust\n>     side when using it to index a slice.\n\nHow much extra memory pressure does this change cause?  In a single\ninstance of xrecord_t, we used to have a single ulong plus a pointer\nand a size_t; now we replaced the single ulong with two 8-byte words,\nso 33% more memory per record, which is not so huge a deal?\n\n>  static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n> -\tlong hi;\n> +\tsize_t hi;\n>  \txdlclass_t *rcrec;\n>  \n> -\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n> +\thi = XDL_HASHLONG(rec->line_hash, cf->hbits);\n\nVery nice that we can lose these random-looking casts.\n\n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index 88b1fe4649..742b81bf3b 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -41,7 +41,8 @@ typedef struct s_chastore {\n>  typedef struct s_xrecord {\n>  \tuint8_t const *ptr;\n>  \tsize_t size;\n> -\tunsigned long ha;\n> +\tuint64_t line_hash;\n> +\tsize_t minimal_perfect_hash;\n>  } xrecord_t;\n>  \n>  typedef struct s_xdfile {\n"},{"id":"530555","messageId":"xmqqqzu4t9yc.fsf@gitster.g","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 00/10] Xdiff cleanup part2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-11T23:40:11Z","receivedAt":"2025-11-11T23:40:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> The primary goal of this patch series is to convert every field's type in\n> xrecord_t and xdfile_t to be unambiguous, in preparation to make it more\n> Rust FFI friendly. Additionally the ha field in xrecord_t is split into\n> line_hash and minimal_perfect hash.\n\nAfter having read the series to its end, I am left with this feeling\nthat it does only half the things that it needs to do.  It does all\nwhat the above paragraph claims it does, sure, in that the relevant\ndata structures now use not \"long\" but \"size_t\", not \"char\" but\n\"uint8_t\", etc., and I do find the resulting data structures sensibly\ndescribed.\n\nBut for the code to be truly consistent between the data structures\nand the operations that work on them, types of on-stack variables\nand function parameters would need to be updated to match these\nstruct members.  As we convert one structure member at a time, casts\nmay need to be sprinkled for assignments to these variables and\npassing these struct members as parameters to functions (which I\ncommented on one of these patche) to keep the blast radius of the\nchanges in each step manageable, but I would have expected that\nfunctions that used to take, say, an \"int\", would be updated to take\n\"size_t\" if the value coming to the parameter is from these struct\nmembers.\n\nPerhaps that would be the theme for \"Xdiff cleanup part 3\" series\nthat we will eventually see after the dust settles from this round?\n\nThanks.\n\n"},{"id":"530676","messageId":"CAH=ZcbCJ4MXnHpspuT+KkeR6LRTQrzh-7v5ep9S8WPRjdteR8g@mail.gmail.com","threadId":"64326","inReplyTo":"xmqqwm3wtat8.fsf@gitster.g","subject":"Re: [PATCH v3 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-14T05:41:15Z","receivedAt":"2025-11-14T05:41:28Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Nov 11, 2025 at 4:21 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > To make this clearer, the old ha field has been split:\n> >   * line_hash: a straightforward hash of a line, independent of any\n> >     external context. Its type is uint64_t, as it comes from a fixed\n> >     width hash function.\n> >   * minimal_perfect_hash: Not a new concept, but now a separate\n> >     field. It comes from the classifier's general-purpose hash table,\n> >     which assigns each line a unique and minimal hash across the two\n> >     files. A size_t is used here because it's meant to be used to\n> >     index an array. This also this avoids ` as usize` casts on the Rust\n> >     side when using it to index a slice.\n>\n> How much extra memory pressure does this change cause?  In a single\n> instance of xrecord_t, we used to have a single ulong plus a pointer\n> and a size_t; now we replaced the single ulong with two 8-byte words,\n> so 33% more memory per record, which is not so huge a deal?\n\nThis was asked and answered earlier in this patch series [1].\n\n> >  static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n> > -     long hi;\n> > +     size_t hi;\n> >       xdlclass_t *rcrec;\n> >\n> > -     hi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n> > +     hi = XDL_HASHLONG(rec->line_hash, cf->hbits);\n>\n> Very nice that we can lose these random-looking casts.\n\nThis was Phillip's suggestion [2]. Thanks Phillip.\n\n[1] https://lore.kernel.org/git/CAH=ZcbD7FeRHtYvN_4=qHApB-AwK18=KRU2SGWNg8ADkrFM-Fw@mail.gmail.com/\n[2] https://lore.kernel.org/git/a66fb440-058e-4cd8-8971-9c320c0387e8@gmail.com/\n"},{"id":"530677","messageId":"CAH=ZcbAQ5fCUuL3cpETQmGNXsPE_5UMf4CqVgjj0vvmXmU7-Vg@mail.gmail.com","threadId":"64326","inReplyTo":"xmqqqzu4t9yc.fsf@gitster.g","subject":"Re: [PATCH v3 00/10] Xdiff cleanup part2","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-14T05:52:06Z","receivedAt":"2025-11-14T05:52:19Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Nov 11, 2025 at 4:40 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > The primary goal of this patch series is to convert every field's type in\n> > xrecord_t and xdfile_t to be unambiguous, in preparation to make it more\n> > Rust FFI friendly. Additionally the ha field in xrecord_t is split into\n> > line_hash and minimal_perfect hash.\n>\n> After having read the series to its end, I am left with this feeling\n> that it does only half the things that it needs to do.  It does all\n> what the above paragraph claims it does, sure, in that the relevant\n> data structures now use not \"long\" but \"size_t\", not \"char\" but\n> \"uint8_t\", etc., and I do find the resulting data structures sensibly\n> described.\n\nThis patch series is already 10 commits long, and it's been a\nchallenge to chunk cleanups of Xdiff because its code is so tangled.\nI'm hoping that future maintenance of Xdiff (after my xdiff cleanup\nseries is complete) will be much easier.\n\n> But for the code to be truly consistent between the data structures\n> and the operations that work on them, types of on-stack variables\n> and function parameters would need to be updated to match these\n> struct members.  As we convert one structure member at a time, casts\n> may need to be sprinkled for assignments to these variables and\n> passing these struct members as parameters to functions (which I\n> commented on one of these patches) to keep the blast radius of the\n> changes in each step manageable, but I would have expected that\n> functions that used to take, say, an \"int\", would be updated to take\n> \"size_t\" if the value coming to the parameter is from these struct\n> members.\n\nI had to draw the line somewhere, and I plan on making more changes to\ndelete more idiosyncrasies in Xdiff.\n\n> Perhaps that would be the theme for \"Xdiff cleanup part 3\" series\n> that we will eventually see after the dust settles from this round?\n\nNot just part 3, but the entire xdiff cleanup series will be about\ncorrecting types among many other code cleanups. This patch series\nalone is unsatisfactory, but it is only 1 of many patch series to\ncome.\n"},{"id":"530681","messageId":"CAH=ZcbBNSNqU3i4DSruVixvYzCEs_MxLCvX6D5W7FsXRqpvALw@mail.gmail.com","threadId":"64326","inReplyTo":"xmqq346kupzm.fsf@gitster.g","subject":"Re: [PATCH v3 04/10] xdiff: use size_t for xrecord_t.size","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-14T06:02:16Z","receivedAt":"2025-11-14T06:02:30Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Tue, Nov 11, 2025 at 4:08 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > size_t is the appropriate type because size is describing the number of\n> > elements, bytes in this case, in memory.\n> >\n> > Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> > ---\n> >  xdiff/xdiffi.c   |  7 +++----\n> >  xdiff/xemit.c    |  8 ++++----\n> >  xdiff/xmerge.c   | 16 ++++++++--------\n> >  xdiff/xprepare.c |  6 +++---\n> >  xdiff/xtypes.h   |  2 +-\n> >  5 files changed, 19 insertions(+), 20 deletions(-)\n>\n> This step looks mostly OK but it is messy in some places.\n>\n> > diff --git a/xdiff/xemit.c b/xdiff/xemit.c\n> > index ead930088a..2f8007753c 100644\n> > --- a/xdiff/xemit.c\n> > +++ b/xdiff/xemit.c\n> > @@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n> >  {\n> >       xrecord_t *rec = &xdf->recs[ri];\n> >\n> > -     if (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n> > +     if (xdl_emit_diffrec((char const *)rec->ptr, (long)rec->size, pre, strlen(pre), ecb) < 0)\n>\n> On platforms where long is narrower than size_t, we'd tentatively\n> leave things broken until we update xdl_emit_diffrec() to take\n> size_t, as it would become too noisy to change it in the same patch,\n> I guess?\n>\n> > @@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n> >       xrecord_t *rec = &xdf->recs[ri];\n> >\n> >       if (!xecfg->find_func)\n> > -             return def_ff((const char *)rec->ptr, rec->size, buf, sz);\n> > -     return xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n> > +             return def_ff((const char *)rec->ptr, (long)rec->size, buf, sz);\n> > +     return xecfg->find_func((const char *)rec->ptr, (long)rec->size, buf, sz, xecfg->find_func_priv);\n>\n> Ditto.\n>\n> > diff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\n> > index 75cb3e76a2..0dd4558a32 100644\n> > --- a/xdiff/xmerge.c\n> > +++ b/xdiff/xmerge.c\n> > @@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n> >       xrecord_t *rec2 = xe2->xdf2.recs + i2;\n> >\n> >       for (i = 0; i < line_count; i++) {\n> > -             int result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n> > -                     (const char *)rec2[i].ptr, rec2[i].size, flags);\n> > +             int result = xdl_recmatch((const char *)rec1[i].ptr, (long)rec1[i].size,\n> > +                     (const char *)rec2[i].ptr, (long)rec2[i].size, flags);\n>\n> Ditto.\n>\n> > @@ -119,11 +119,11 @@ static int xdl_recs_copy_0(int use_orig, xdfenv_t *xe, int i, int count, int nee\n> >       if (count < 1)\n> >               return 0;\n> >\n> > -     for (i = 0; i < count; size += recs[i++].size)\n> > +     for (i = 0; i < count; size += (int)recs[i++].size)\n> >               if (dest)\n> >                       memcpy(dest + size, recs[i].ptr, recs[i].size);\n> >       if (add_nl) {\n> > -             i = recs[count - 1].size;\n> > +             i = (int)recs[count - 1].size;\n> >               if (i == 0 || recs[count - 1].ptr[i - 1] != '\\n') {\n> >                       if (needs_cr) {\n> >                               if (dest)\n>\n> This is messier than I expected.  Before the precontext of this\n> hunk, \"i\" and \"count\" are both incoming parameters of type \"int\", so\n> the same \"what if size_t is wider?\" puzzlement applies here.  At\n> least, the reason why \"i\" and \"count\" is \"int\" is not because they\n> want to be able to express negative values, so it shouldn't involve\n> too much hassle if we later want to change them to size_t to lose\n> these casts.\n>\n> > @@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n> >\n> >  static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n> >  {\n> > -     return xdl_recmatch((const char *)rec1->ptr, rec1->size,\n> > -                         (const char *)rec2->ptr, rec2->size, flags);\n> > +     return xdl_recmatch((const char *)rec1->ptr, (long)rec1->size,\n> > +                         (const char *)rec2->ptr, (long)rec2->size, flags);\n> >  }\n>\n> Same \"long may not be wide enough, in which case we'd need further\n> fixes\" applies here.\n>\n> > @@ -441,7 +441,7 @@ static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n> >  {\n> >       for (; chg; chg--, i++)\n> >               if (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n> > -                             xe->xdf2.recs[i].size))\n> > +                             (long)xe->xdf2.recs[i].size))\n> >                       return 1;\n> >       return 0;\n> >  }\n>\n> Ditto.\n>\n> > diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> > index 4c56467076..b3219aed3e 100644\n> > --- a/xdiff/xprepare.c\n> > +++ b/xdiff/xprepare.c\n> > @@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n> >       hi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n> >       for (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n> >               if (rcrec->rec.ha == rec->ha &&\n> > -                             xdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n> > -                                     (const char *)rec->ptr, rec->size, cf->flags))\n> > +                             xdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n> > +                                     (const char *)rec->ptr, (long)rec->size, cf->flags))\n>\n> Ditto.\n\nmmbuffer_t holds all of the bytes of the file in memory, so the number\nof lines referenced in mmbuffer_t has to be less than or equal to\nthat, which makes the point about long vs size_t moot for this patch\nseries. Maybe int vs size_t is a different story, but there are many\nother places that use `int` that limit the number of lines in a file\nthat aren't touched at all in this patch series. I will update these\ntypes, but in a future patch series because they cause a refactor\navalanche in many places.\n\nI don't like the current state that Xdiff is in either. That's why I\nintend to keep going with my xdiff cleanup series.\n\n> > @@ -157,7 +157,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n> >                               goto abort;\n> >                       crec = &xdf->recs[xdf->nrec++];\n> >                       crec->ptr = (uint8_t const *)prev;\n> > -                     crec->size = (long) (cur - prev);\n> > +                     crec->size = cur - prev;\n>\n> Yay!\n>\n> > diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> > index b1c520a378..88b1fe4649 100644\n> > --- a/xdiff/xtypes.h\n> > +++ b/xdiff/xtypes.h\n> > @@ -40,7 +40,7 @@ typedef struct s_chastore {\n> >\n> >  typedef struct s_xrecord {\n> >       uint8_t const *ptr;\n> > -     long size;\n> > +     size_t size;\n>\n> Yay, too!\n>\n> >       unsigned long ha;\n> >  } xrecord_t;\n\nI agree. It's nice to see some clean code in this patch series.\n"},{"id":"530693","messageId":"xmqqecq0k23r.fsf@gitster.g","threadId":"64326","inReplyTo":"CAH=ZcbBNSNqU3i4DSruVixvYzCEs_MxLCvX6D5W7FsXRqpvALw@mail.gmail.com","subject":"Re: [PATCH v3 04/10] xdiff: use size_t for xrecord_t.size","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-14T16:31:20Z","receivedAt":"2025-11-14T16:31:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ezekiel Newren <ezekielnewren@gmail.com> writes:\n\n> On Tue, Nov 11, 2025 at 4:08 PM Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> ...\n> mmbuffer_t holds all of the bytes of the file in memory, so the number\n> of lines referenced in mmbuffer_t has to be less than or equal to\n> that, which makes the point about long vs size_t moot for this patch\n> series.\n\n... because size there is still \"long\"?\n\n> I don't like the current state that Xdiff is in either. That's why I\n> intend to keep going with my xdiff cleanup series.\n\nGreat, and we already have seen improvements; an intermediate state,\nas we already discussed in this thread, may be noisier with casts\nbut that cannot be avoided.\n\n> I agree. It's nice to see some clean code in this patch series.\n\nThanks.\n"},{"id":"530699","messageId":"xmqq346gidlc.fsf@gitster.g","threadId":"64326","inReplyTo":"CAH=ZcbCJ4MXnHpspuT+KkeR6LRTQrzh-7v5ep9S8WPRjdteR8g@mail.gmail.com","subject":"Re: [PATCH v3 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-14T20:06:07Z","receivedAt":"2025-11-14T20:06:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ezekiel Newren <ezekielnewren@gmail.com> writes:\n\n>> How much extra memory pressure does this change cause?  In a single\n>> instance of xrecord_t, we used to have a single ulong plus a pointer\n>> and a size_t; now we replaced the single ulong with two 8-byte words,\n>> so 33% more memory per record, which is not so huge a deal?\n>\n> This was asked and answered earlier in this patch series [1].\n\nIn short, this step does bloat, but the memory usage will shrink\nwhen the members are moved elsewhere in future patches?\n\n> [1] https://lore.kernel.org/git/CAH=ZcbD7FeRHtYvN_4=qHApB-AwK18=KRU2SGWNg8ADkrFM-Fw@mail.gmail.com/\n"},{"id":"530708","messageId":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v3.git.git.1762890152.gitgitgadget@gmail.com","subject":"[PATCH v4 00/10] Xdiff cleanup part2","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:46Z","receivedAt":"2025-11-14T22:36:59Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"Changes in v4:\n\n * Update documentation to not mention Unicode except once\n * Don't move dstart/dend with in the xdfile_t struct\n * Rephrase justification on changing xrecord_t.ptr's type\n\nChanges in v3:\n\n * Address comments about commit messages and documentation\n * Add unambiguous-types.adoc to Makefile and Meson\n * Use markdown style to avoid asciidoc issues\n\nChanges in v2:\n\n * Added documentation about unambiguous types and FFI\n * Addressed comments on the mailing list\n\n\nOriginal cover letter below:\n============================\n\nMaintainer note: This patch series builds on top of en/xdiff-cleanup and\nam/xdiff-hash-tweak (both of which are now in master).\n\nThe primary goal of this patch series is to convert every field's type in\nxrecord_t and xdfile_t to be unambiguous, in preparation to make it more\nRust FFI friendly. Additionally the ha field in xrecord_t is split into\nline_hash and minimal_perfect hash.\n\nThe order of some of the fields has changed as called out by the commit\nmessages.\n\nBefore:\n\ntypedef struct s_xrecord {\n\tchar const *ptr;\n\tlong size;\n\tunsigned long ha;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tlong nrec;\n\tlong dstart, dend;\n\tbool *changed;\n\tlong *rindex;\n\tlong nreff;\n} xdfile_t;\n\n\nAfter part 2\n\ntypedef struct s_xrecord {\n\tuint8_t const *ptr;\n\tsize_t size;\n\tuint64_t line_hash;\n\tsize_t minimal_perfect_hash;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tsize_t nrec;\n\tptrdiff_t dstart, dend;\n\tbool *changed;\n\tsize_t *reference_index;\n\tsize_t nreff;\n} xdfile_t;\n\n\nEzekiel Newren (10):\n  doc: define unambiguous type mappings across C and Rust\n  xdiff: use ptrdiff_t for dstart/dend\n  xdiff: make xrecord_t.ptr a uint8_t instead of char\n  xdiff: use size_t for xrecord_t.size\n  xdiff: use unambiguous types in xdl_hash_record()\n  xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n  xdiff: make xdfile_t.nrec a size_t instead of long\n  xdiff: make xdfile_t.nreff a size_t instead of long\n  xdiff: change rindex from long to size_t in xdfile_t\n  xdiff: rename rindex -> reference_index\n\n Documentation/Makefile                        |   1 +\n Documentation/technical/meson.build           |   1 +\n .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n xdiff-interface.c                             |   2 +-\n xdiff/xdiffi.c                                |  29 ++-\n xdiff/xemit.c                                 |  28 +--\n xdiff/xhistogram.c                            |   4 +-\n xdiff/xmerge.c                                |  30 +--\n xdiff/xpatience.c                             |  14 +-\n xdiff/xprepare.c                              |  60 ++---\n xdiff/xtypes.h                                |  15 +-\n xdiff/xutils.c                                |  32 +--\n xdiff/xutils.h                                |   6 +-\n 13 files changed, 336 insertions(+), 110 deletions(-)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\n\nbase-commit: a99f379adf116d53eb11957af5bab5214915f91d\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2070%2Fezekielnewren%2Fxdiff_cleanup_part2-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2070/ezekielnewren/xdiff_cleanup_part2-v4\nPull-Request: https://github.com/git/git/pull/2070\n\nRange-diff vs v3:\n\n  1:  e5d084d340 !  1:  af732beb69 doc: define unambiguous type mappings across C and Rust\n     @@ Metadata\n       ## Commit message ##\n          doc: define unambiguous type mappings across C and Rust\n      \n     -    Document other nuances with crossing the FFI boundary. Other language\n     +    Document other nuances when crossing the FFI boundary. Other language\n          mappings may be added in the future.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +\n      +This is where C and Rust don't have a clean one-to-one mapping.\n      +\n     ++A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n     ++a `char` will have the same size as the corresponding Rust struct using `u8`.\n     ++In that sense, such structs are safe to pass over the FFI boundary, because\n     ++their fields will be laid out identically. However, beyond bit width, C `char`\n     ++has additional semantics and platform-dependent behavior that can cause\n     ++problems, as discussed below.\n     ++\n      +C comparison problem: While the sign of `char` is implementation defined, it's\n      +also signless (neither signed nor unsigned). When building with\n      +`make DEVELOPER=1` it will complain about a \"differ in signedness\" when `char`\n      +is compared with `uint8_t` or `int8_t`.\n      +\n     -+Rust's `char` type is an unsigned 32-bit integer that is used to describe\n     -+Unicode code points. Even though a C `char` is the same width as `u8`, `char`\n     -+should be converted to u8 where it is describing bytes in memory. If a C\n     -+`char` is not describing bytes, then it should be converted to a more accurate\n     -+unambiguous type. The reason for mentioning Unicode here is because of how &str\n     -+is defined in Rust and how to create a &str from &[u8]. Rust assumes that &str\n     -+is a correctly encoded utf-8 string, i.e. text in memory. Where as a C `char`\n     -+makes no assumption about the bytes that it is representing.\n     -+\n     -+```\n     -+let raw_bytes = b\"abc\\n\";\n     -+let result = std::str::from_utf8(raw_bytes);\n     -+if let Ok(line) = result {\n     -+    // do something with text\n     -+}\n     -+```\n     -+\n     -+While you could specify `char` in the C code and `u8` in Rust code, it's not as\n     -+clear what the appropriate type is, but it would work across the FFI boundary.\n     -+However, the bigger problem comes from code generation tools like cbindgen and\n     -+bindgen. When cbindgen sees u8 in Rust it will generate uint8_t on the C side\n     -+which will cause differ in signedness warnings/errors. Similarly if bindgen\n     -+sees `char` on the C side it will generate `std::ffi::c_char` which has its own\n     -+problems.\n     ++Note: Rust's `char` type is an unsigned 32-bit integer that is used to describe\n     ++Unicode code points.\n      +\n      +=== Notes\n      +^1^ This is only true if stdbool.h (or equivalent) is used. +\n  2:  52e3f589b1 !  2:  b60a03eb31 xdiff: use ptrdiff_t for dstart/dend\n     @@ Commit message\n          ptrdiff_t is appropriate for dstart and dend because they both describe\n          positive or negative offsets relative to a pointer.\n      \n     -    A future patch will move these fields to a different struct. Moving\n     -    them to the end of xdfile_t now, means the field order of xdfile_t will\n     -    be disturbed less.\n     -\n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xtypes.h ##\n     @@ xdiff/xtypes.h: typedef struct s_xrecord {\n       \txrecord_t *recs;\n       \tlong nrec;\n      -\tlong dstart, dend;\n     ++\tptrdiff_t dstart, dend;\n       \tbool *changed;\n       \tlong *rindex;\n       \tlong nreff;\n     -+\tptrdiff_t dstart, dend;\n     - } xdfile_t;\n     - \n     - typedef struct s_xdfenv {\n  3:  83e7bf180a !  3:  042fbb11d0 xdiff: make xrecord_t.ptr a uint8_t instead of char\n     @@ Commit message\n      \n          Make xrecord_t.ptr uint8_t because it's referring to bytes in memory.\n      \n     -    Every usage of this field was inspected and cast to char*, or similar,\n     -    to avoid signedness warnings/errors from the compiler. Casting was used\n     -    so that the whole of xdiff doesn't need to be refactored in order to\n     -    change the type of this field.\n     +    In order to avoid a refactor avalanche, many uses of this field were\n     +    cast to char* or similar. One exception is in get_indent() where the\n     +    local variable `char c` was changed to `uint8_t c`.\n     +\n     +    Places where casting was unnecessary:\n     +    xemit.c:156\n     +    xmerge.c:124\n     +    xmerge.c:127\n     +    xmerge.c:164\n     +    xmerge.c:169\n     +    xmerge.c:172\n     +    xmerge.c:178\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n  4:  da2b80ea0b =  4:  c103fa6bea xdiff: use size_t for xrecord_t.size\n  5:  c6ba630ac5 =  5:  2ee9a74653 xdiff: use unambiguous types in xdl_hash_record()\n  6:  3834ea8f9b !  6:  f044274bd5 xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n     @@ Commit message\n              field. It comes from the classifier's general-purpose hash table,\n              which assigns each line a unique and minimal hash across the two\n              files. A size_t is used here because it's meant to be used to\n     -        index an array. This also this avoids ` as usize` casts on the Rust\n     +        index an array. This also avoids ` as usize` casts on the Rust\n              side when using it to index a slice.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n  7:  e2a2c7530c !  7:  f7a3731d94 xdiff: make xdfile_t.nrec a size_t instead of long\n     @@ xdiff/xtypes.h: typedef struct s_xrecord {\n       \txrecord_t *recs;\n      -\tlong nrec;\n      +\tsize_t nrec;\n     + \tptrdiff_t dstart, dend;\n       \tbool *changed;\n       \tlong *rindex;\n     - \tlong nreff;\n  8:  31cd2a1aa4 !  8:  93f84ae72e xdiff: make xdfile_t.nreff a size_t instead of long\n     @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n      \n       ## xdiff/xtypes.h ##\n      @@ xdiff/xtypes.h: typedef struct s_xdfile {\n     - \tsize_t nrec;\n     + \tptrdiff_t dstart, dend;\n       \tbool *changed;\n       \tlong *rindex;\n      -\tlong nreff;\n      +\tsize_t nreff;\n     - \tptrdiff_t dstart, dend;\n       } xdfile_t;\n       \n     + typedef struct s_xdfenv {\n  9:  aee0d3958b !  9:  39369becc8 xdiff: change rindex from long to size_t in xdfile_t\n     @@ Commit message\n      \n       ## xdiff/xtypes.h ##\n      @@ xdiff/xtypes.h: typedef struct s_xdfile {\n     - \txrecord_t *recs;\n       \tsize_t nrec;\n     + \tptrdiff_t dstart, dend;\n       \tbool *changed;\n      -\tlong *rindex;\n      +\tsize_t *rindex;\n       \tsize_t nreff;\n     - \tptrdiff_t dstart, dend;\n       } xdfile_t;\n     + \n 10:  75c26fe160 ! 10:  950d1e6193 xdiff: rename rindex -> reference_index\n     @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n      \n       ## xdiff/xtypes.h ##\n      @@ xdiff/xtypes.h: typedef struct s_xdfile {\n     - \txrecord_t *recs;\n       \tsize_t nrec;\n     + \tptrdiff_t dstart, dend;\n       \tbool *changed;\n      -\tsize_t *rindex;\n      +\tsize_t *reference_index;\n       \tsize_t nreff;\n     - \tptrdiff_t dstart, dend;\n       } xdfile_t;\n     + \n\n-- \ngitgitgadget\n"},{"id":"530707","messageId":"af732beb6904d8b9b7801ecc2487dabdea05a571.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:47Z","receivedAt":"2025-11-14T22:37:00Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nDocument other nuances when crossing the FFI boundary. Other language\nmappings may be added in the future.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n Documentation/Makefile                        |   1 +\n Documentation/technical/meson.build           |   1 +\n .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n 3 files changed, 226 insertions(+)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 04e9e10b27..bc1adb2d9d 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -142,6 +142,7 @@ TECH_DOCS += technical/shallow\n TECH_DOCS += technical/sparse-checkout\n TECH_DOCS += technical/sparse-index\n TECH_DOCS += technical/trivial-merge\n+TECH_DOCS += technical/unambiguous-types\n TECH_DOCS += technical/unit-tests\n SP_ARTICLES += $(TECH_DOCS)\n SP_ARTICLES += technical/api-index\ndiff --git a/Documentation/technical/meson.build b/Documentation/technical/meson.build\nindex be698ef22a..89a6e26821 100644\n--- a/Documentation/technical/meson.build\n+++ b/Documentation/technical/meson.build\n@@ -32,6 +32,7 @@ articles = [\n   'sparse-checkout.adoc',\n   'sparse-index.adoc',\n   'trivial-merge.adoc',\n+  'unambiguous-types.adoc',\n   'unit-tests.adoc',\n ]\n \ndiff --git a/Documentation/technical/unambiguous-types.adoc b/Documentation/technical/unambiguous-types.adoc\nnew file mode 100644\nindex 0000000000..9a42c72890\n--- /dev/null\n+++ b/Documentation/technical/unambiguous-types.adoc\n@@ -0,0 +1,224 @@\n+= Unambiguous types\n+\n+Most of these mappings are obvious, but there are some nuances and gotchas with\n+Rust FFI (Foreign Function Interface).\n+\n+This document defines clear, one-to-one mappings between primitive types in C,\n+Rust (and possible other languages in the future). Its purpose is to eliminate\n+ambiguity in type widths, signedness, and binary representation across\n+platforms and languages.\n+\n+For Git, the only header required to use these unambiguous types in C is\n+`git-compat-util.h`.\n+\n+== Boolean types\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| bool^1^       | bool\n+|===\n+\n+== Integer types\n+\n+In C, `<stdint.h>` (or an equivalent) must be included.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| uint8_t    | u8\n+| uint16_t   | u16\n+| uint32_t   | u32\n+| uint64_t   | u64\n+\n+| int8_t     | i8\n+| int16_t    | i16\n+| int32_t    | i32\n+| int64_t    | i64\n+|===\n+\n+== Floating-point types\n+\n+Rust requires IEEE-754 semantics.\n+In C, that is typically true, but not guaranteed by the standard.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| float^2^      | f32\n+| double^2^     | f64\n+|===\n+\n+== Size types\n+\n+These types represent pointer-sized integers and are typically defined in\n+`<stddef.h>` or an equivalent header.\n+\n+Size types should be used any time pointer arithmetic is performed e.g.\n+indexing an array, describing the number of elements in memory, etc...\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| size_t^3^     | usize\n+| ptrdiff_t^3^  | isize\n+|===\n+\n+== Character types\n+\n+This is where C and Rust don't have a clean one-to-one mapping.\n+\n+A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n+a `char` will have the same size as the corresponding Rust struct using `u8`.\n+In that sense, such structs are safe to pass over the FFI boundary, because\n+their fields will be laid out identically. However, beyond bit width, C `char`\n+has additional semantics and platform-dependent behavior that can cause\n+problems, as discussed below.\n+\n+C comparison problem: While the sign of `char` is implementation defined, it's\n+also signless (neither signed nor unsigned). When building with\n+`make DEVELOPER=1` it will complain about a \"differ in signedness\" when `char`\n+is compared with `uint8_t` or `int8_t`.\n+\n+Note: Rust's `char` type is an unsigned 32-bit integer that is used to describe\n+Unicode code points.\n+\n+=== Notes\n+^1^ This is only true if stdbool.h (or equivalent) is used. +\n+^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the\n+platform/arch for C does not follow IEEE-754 then this equivalence does not\n+hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n+there may be a strange platform/arch where even this isn't true. +\n+^3^ C also defines uintptr_t, ssize_t and intptr_t, but these types are\n+discouraged for FFI purposes. For functions like `read()` and `write()` ssize_t\n+should be cast to a different, and unambiguous, type before being passed over\n+the FFI boundary. +\n+\n+== Problems with std::ffi::c_* types in Rust\n+TL;DR: In practice, Rust's `c_*` types aren't guaranteed to match C types for\n+all possible C compilers, platforms, or architectures, because Rust only\n+ensures correctness of C types on officially supported targets. These\n+definitions have changed over time to match more targets which means that the\n+c_* definitions will differ based on which Rust version Git chooses to use.\n+\n+Current list of safe, Rust side, FFI types in Git: +\n+\n+* `c_void`\n+* `CStr`\n+* `CString`\n+\n+Even then, they should be used sparingly, and only where the semantics match\n+exactly.\n+\n+The std::os::raw::c_* directly inherits the problems of core::ffi, which\n+changes over time and seems to make a best guess at the correct definition for\n+a given platform/target. This probably isn't a problem for all other platforms\n+that Rust supports currently, but can anyone say that Rust got it right for all\n+C compilers of all platforms/targets?\n+\n+To give an example: c_long is defined in\n+footnote:[https://doc.rust-lang.org/1.63.0/src/core/ffi/mod.rs.html#175-189[c_long in 1.63.0]]\n+footnote:[https://doc.rust-lang.org/1.89.0/src/core/ffi/primitives.rs.html#135-151[c_long in 1.89.0]]\n+\n+=== Rust version 1.63.0\n+\n+```\n+mod c_long_definition {\n+    cfg_if! {\n+        if #[cfg(all(target_pointer_width = \"64\", not(windows)))] {\n+            pub type c_long = i64;\n+            pub type NonZero_c_long = crate::num::NonZeroI64;\n+            pub type c_ulong = u64;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU64;\n+        } else {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub type c_long = i32;\n+            pub type NonZero_c_long = crate::num::NonZeroI32;\n+            pub type c_ulong = u32;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU32;\n+        }\n+    }\n+}\n+```\n+\n+=== Rust version 1.89.0\n+\n+```\n+mod c_long_definition {\n+    crate::cfg_select! {\n+        any(\n+            all(target_pointer_width = \"64\", not(windows)),\n+            // wasm32 Linux ABI uses 64-bit long\n+            all(target_arch = \"wasm32\", target_os = \"linux\")\n+        ) => {\n+            pub(super) type c_long = i64;\n+            pub(super) type c_ulong = u64;\n+        }\n+        _ => {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub(super) type c_long = i32;\n+            pub(super) type c_ulong = u32;\n+        }\n+    }\n+}\n+```\n+\n+Even for the cases where C types are correctly mapped to Rust types via\n+std::ffi::c_* there are still problems. Let's take c_char for example. On some\n+platforms it's u8 on others it's i8.\n+\n+=== Subtraction underflow in debug mode\n+\n+The following code will panic in debug on platforms that define c_char as u8,\n+but won't if it's an i8.\n+\n+```\n+let mut x: std::ffi::c_char = 0;\n+x -= 1;\n+```\n+\n+=== Inconsistent shift behavior\n+\n+`x` will be 0xC0 for platforms that use i8, but will be 0x40 where it's u8.\n+\n+```\n+let mut x: std::ffi::c_char = 0x80;\n+x >>= 1;\n+```\n+\n+=== Equality fails to compile on some platforms\n+\n+The following will not compile on platforms that define c_char as i8, but will\n+if it's u8. You can cast x e.g. `assert_eq!(x as u8, b'a');`, but then you get\n+a warning on platforms that use u8 and a clean compilation where i8 is used.\n+\n+```\n+let mut x: std::ffi::c_char = 0x61;\n+assert_eq!(x, b'a');\n+```\n+\n+== Enum types\n+Rust enum types should not be used as FFI types. Rust enum types are more like\n+C union types than C enum's. For something like:\n+\n+```\n+#[repr(C, u8)]\n+enum Fruit {\n+    Apple,\n+    Banana,\n+    Cherry,\n+}\n+```\n+\n+It's easy enough to make sure the Rust enum matches what C would expect, but a\n+more complex type like.\n+\n+```\n+enum HashResult {\n+    SHA1([u8; 20]),\n+    SHA256([u8; 32]),\n+}\n+```\n+\n+The Rust compiler has to add a discriminant to the enum to distinguish between\n+the variants. The width, location, and values for that discriminant is up to\n+the Rust compiler and is not ABI stable.\n-- \ngitgitgadget\n\n"},{"id":"530709","messageId":"b60a03eb31bba747db546258e843ac94a4b70950.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 02/10] xdiff: use ptrdiff_t for dstart/dend","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:48Z","receivedAt":"2025-11-14T22:37:01Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nptrdiff_t is appropriate for dstart and dend because they both describe\npositive or negative offsets relative to a pointer.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex f145abba3e..7a2d429ec5 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,7 +47,7 @@ typedef struct s_xrecord {\n typedef struct s_xdfile {\n \txrecord_t *recs;\n \tlong nrec;\n-\tlong dstart, dend;\n+\tptrdiff_t dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n-- \ngitgitgadget\n\n"},{"id":"530711","messageId":"042fbb11d03606879503846e86fac65e6e74d02a.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:49Z","receivedAt":"2025-11-14T22:37:03Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake xrecord_t.ptr uint8_t because it's referring to bytes in memory.\n\nIn order to avoid a refactor avalanche, many uses of this field were\ncast to char* or similar. One exception is in get_indent() where the\nlocal variable `char c` was changed to `uint8_t c`.\n\nPlaces where casting was unnecessary:\nxemit.c:156\nxmerge.c:124\nxmerge.c:127\nxmerge.c:164\nxmerge.c:169\nxmerge.c:172\nxmerge.c:178\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     |  6 +++---\n xdiff/xmerge.c    | 14 +++++++-------\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xtypes.h    |  2 +-\n xdiff/xutils.c    |  4 ++--\n 7 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 6f3998ee54..411a8aa69f 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n \tint ret = 0;\n \n \tfor (i = 0; i < rec->size; i++) {\n-\t\tchar c = rec->ptr[i];\n+\t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n \t\t\treturn ret;\n@@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\n@@ -1008,7 +1008,7 @@ static int record_matches_regex(xrecord_t *rec, xpparam_t const *xpp) {\n \tsize_t i;\n \n \tfor (i = 0; i < xpp->ignore_regex_nr; i++)\n-\t\tif (!regexec_buf(xpp->ignore_regex[i], rec->ptr, rec->size, 1,\n+\t\tif (!regexec_buf(xpp->ignore_regex[i], (const char *)rec->ptr, rec->size, 1,\n \t\t\t\t &regmatch, 0))\n \t\t\treturn 1;\n \ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex b2f1f30cd3..ead930088a 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec(rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff(rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func(rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex fd600cbb5d..75cb3e76a2 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch(rec1[i].ptr, rec1[i].size,\n-\t\t\trec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch(rec1->ptr, rec1->size,\n-\t\t\t    rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n+\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n }\n \n /*\n@@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n \t\t * we have a very simple mmfile structure.\n \t\t */\n \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n-\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n+\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n \t\tt2.ptr = (char *)xe2->xdf2.recs[m->i2].ptr;\n-\t\tt2.size = xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n+\t\tt2.size = (char *)xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n \t\t\t+ xe2->xdf2.recs[m->i2 + m->chg2 - 1].size - t2.ptr;\n \t\tif (xdl_do_diff(&t1, &t2, xpp, &xe) < 0)\n \t\t\treturn -1;\n@@ -440,7 +440,7 @@ static int line_contains_alnum(const char *ptr, long size)\n static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n-\t\tif (line_contains_alnum(xe->xdf2.recs[i].ptr,\n+\t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n \t\t\t\txe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 669b653580..bb61354f22 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -121,7 +121,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t\treturn;\n \tmap->entries[index].line1 = line;\n \tmap->entries[index].hash = record->ha;\n-\tmap->entries[index].anchor = is_anchor(xpp, map->env->xdf1.recs[line - 1].ptr);\n+\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n \tif (map->last) {\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 192334f1b7..4c56467076 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch(rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\trec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = prev;\n+\t\t\tcrec->ptr = (uint8_t const *)prev;\n \t\t\tcrec->size = (long) (cur - prev);\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 7a2d429ec5..69727fb299 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -39,7 +39,7 @@ typedef struct s_chastore {\n } chastore_t;\n \n typedef struct s_xrecord {\n-\tchar const *ptr;\n+\tuint8_t const *ptr;\n \tlong size;\n \tunsigned long ha;\n } xrecord_t;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 447e66c719..7be063bfb6 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -465,10 +465,10 @@ int xdl_fall_back_diff(xdfenv_t *diff_env, xpparam_t const *xpp,\n \txdfenv_t env;\n \n \tsubfile1.ptr = (char *)diff_env->xdf1.recs[line1 - 1].ptr;\n-\tsubfile1.size = diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n+\tsubfile1.size = (char *)diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n \t\tdiff_env->xdf1.recs[line1 + count1 - 2].size - subfile1.ptr;\n \tsubfile2.ptr = (char *)diff_env->xdf2.recs[line2 - 1].ptr;\n-\tsubfile2.size = diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n+\tsubfile2.size = (char *)diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n \t\tdiff_env->xdf2.recs[line2 + count2 - 2].size - subfile2.ptr;\n \tif (xdl_do_diff(&subfile1, &subfile2, xpp, &env) < 0)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"530710","messageId":"c103fa6bea97b119c01ac139cd18ca5e272a4b29.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 04/10] xdiff: use size_t for xrecord_t.size","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:50Z","receivedAt":"2025-11-14T22:37:04Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is the appropriate type because size is describing the number of\nelements, bytes in this case, in memory.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  7 +++----\n xdiff/xemit.c    |  8 ++++----\n xdiff/xmerge.c   | 16 ++++++++--------\n xdiff/xprepare.c |  6 +++---\n xdiff/xtypes.h   |  2 +-\n 5 files changed, 19 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 411a8aa69f..edd05466df 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -403,10 +403,9 @@ static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n  */\n static int get_indent(xrecord_t *rec)\n {\n-\tlong i;\n \tint ret = 0;\n \n-\tfor (i = 0; i < rec->size; i++) {\n+\tfor (size_t i = 0; i < rec->size; i++) {\n \t\tuint8_t c = rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n@@ -993,11 +992,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex ead930088a..2f8007753c 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, (long)rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, (long)rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, (long)rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\n@@ -151,7 +151,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n static int is_empty_rec(xdfile_t *xdf, long ri)\n {\n \txrecord_t *rec = &xdf->recs[ri];\n-\tlong i = 0;\n+\tsize_t i = 0;\n \n \tfor (; i < rec->size && XDL_ISSPACE(rec->ptr[i]); i++);\n \ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 75cb3e76a2..0dd4558a32 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n-\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, (long)rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, (long)rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -119,11 +119,11 @@ static int xdl_recs_copy_0(int use_orig, xdfenv_t *xe, int i, int count, int nee\n \tif (count < 1)\n \t\treturn 0;\n \n-\tfor (i = 0; i < count; size += recs[i++].size)\n+\tfor (i = 0; i < count; size += (int)recs[i++].size)\n \t\tif (dest)\n \t\t\tmemcpy(dest + size, recs[i].ptr, recs[i].size);\n \tif (add_nl) {\n-\t\ti = recs[count - 1].size;\n+\t\ti = (int)recs[count - 1].size;\n \t\tif (i == 0 || recs[count - 1].ptr[i - 1] != '\\n') {\n \t\t\tif (needs_cr) {\n \t\t\t\tif (dest)\n@@ -156,7 +156,7 @@ static int xdl_orig_copy(xdfenv_t *xe, int i, int count, int needs_cr, int add_n\n  */\n static int is_eol_crlf(xdfile_t *file, int i)\n {\n-\tlong size;\n+\tsize_t size;\n \n \tif (i < file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n-\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, (long)rec1->size,\n+\t\t\t    (const char *)rec2->ptr, (long)rec2->size, flags);\n }\n \n /*\n@@ -441,7 +441,7 @@ static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n \t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n-\t\t\t\txe->xdf2.recs[i].size))\n+\t\t\t\t(long)xe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 4c56467076..b3219aed3e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -157,7 +157,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = (uint8_t const *)prev;\n-\t\t\tcrec->size = (long) (cur - prev);\n+\t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 69727fb299..354349b523 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -40,7 +40,7 @@ typedef struct s_chastore {\n \n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n-\tlong size;\n+\tsize_t size;\n \tunsigned long ha;\n } xrecord_t;\n \n-- \ngitgitgadget\n\n"},{"id":"530715","messageId":"2ee9a74653e77c395659c8540d9139179478e3fd.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 05/10] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:51Z","receivedAt":"2025-11-14T22:37:05Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nConvert the function signature and body to use unambiguous types. char\nis changed to uint8_t because this function processes bytes in memory.\nunsigned long to uint64_t so that the hash output is consistent across\nplatforms. `flags` was changed from long to uint64_t to ensure the\nhigh order bits are not dropped on platforms that treat long as 32\nbits.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff-interface.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xutils.c    | 28 ++++++++++++++--------------\n xdiff/xutils.h    |  6 +++---\n 4 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff-interface.c b/xdiff-interface.c\nindex 4971f722b3..1a35556380 100644\n--- a/xdiff-interface.c\n+++ b/xdiff-interface.c\n@@ -300,7 +300,7 @@ void xdiff_clear_find_func(xdemitconf_t *xecfg)\n \n unsigned long xdiff_hash_string(const char *s, size_t len, long flags)\n {\n-\treturn xdl_hash_record(&s, s + len, flags);\n+\treturn xdl_hash_record((uint8_t const**)&s, (uint8_t const*)s + len, flags);\n }\n \n int xdiff_compare_lines(const char *l1, long s1,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex b3219aed3e..85e56021da 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -137,8 +137,8 @@ static void xdl_free_ctx(xdfile_t *xdf)\n static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_t const *xpp,\n \t\t\t   xdlclassifier_t *cf, xdfile_t *xdf) {\n \tlong bsize;\n-\tunsigned long hav;\n-\tchar const *blk, *cur, *top, *prev;\n+\tuint64_t hav;\n+\tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n \txdf->rindex = NULL;\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = (uint8_t const *)prev;\n+\t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 7be063bfb6..77ee1ad9c8 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -249,11 +249,11 @@ int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags)\n \treturn 1;\n }\n \n-unsigned long xdl_hash_record_with_whitespace(char const **data,\n-\t\tchar const *top, long flags) {\n-\tunsigned long ha = 5381;\n-\tchar const *ptr = *data;\n-\tint cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data,\n+\t\tuint8_t const *top, uint64_t flags) {\n+\tuint64_t ha = 5381;\n+\tuint8_t const *ptr = *data;\n+\tbool cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n \n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tif (cr_at_eol_only) {\n@@ -263,8 +263,8 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\t\tcontinue;\n \t\t}\n \t\telse if (XDL_ISSPACE(*ptr)) {\n-\t\t\tconst char *ptr2 = ptr;\n-\t\t\tint at_eol;\n+\t\t\tconst uint8_t *ptr2 = ptr;\n+\t\t\tbool at_eol;\n \t\t\twhile (ptr + 1 < top && XDL_ISSPACE(ptr[1])\n \t\t\t\t\t&& ptr[1] != '\\n')\n \t\t\t\tptr++;\n@@ -274,20 +274,20 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_CHANGE\n \t\t\t\t && !at_eol) {\n \t\t\t\tha += (ha << 5);\n-\t\t\t\tha ^= (unsigned long) ' ';\n+\t\t\t\tha ^= (uint64_t) ' ';\n \t\t\t}\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_AT_EOL\n \t\t\t\t && !at_eol) {\n \t\t\t\twhile (ptr2 != ptr + 1) {\n \t\t\t\t\tha += (ha << 5);\n-\t\t\t\t\tha ^= (unsigned long) *ptr2;\n+\t\t\t\t\tha ^= (uint64_t) *ptr2;\n \t\t\t\t\tptr2++;\n \t\t\t\t}\n \t\t\t}\n \t\t\tcontinue;\n \t\t}\n \t\tha += (ha << 5);\n-\t\tha ^= (unsigned long) *ptr;\n+\t\tha ^= (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n \n@@ -304,9 +304,9 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n #define REASSOC_FENCE(x, y)\n #endif\n \n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n-\tunsigned long ha = 5381, c0, c1;\n-\tchar const *ptr = *data;\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top) {\n+\tuint64_t ha = 5381, c0, c1;\n+\tuint8_t const *ptr = *data;\n #if 0\n \t/*\n \t * The baseline form of the optimized loop below. This is the djb2\n@@ -314,7 +314,7 @@ unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n \t */\n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tha += (ha << 5);\n-\t\tha += (unsigned long) *ptr;\n+\t\tha += (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n #else\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 13f6831047..615b4a9d35 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -34,9 +34,9 @@ void *xdl_cha_alloc(chastore_t *cha);\n long xdl_guess_lines(mmfile_t *mf, long sample);\n int xdl_blankline(const char *line, long size, long flags);\n int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags);\n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top);\n-unsigned long xdl_hash_record_with_whitespace(char const **data, char const *top, long flags);\n-static inline unsigned long xdl_hash_record(char const **data, char const *top, long flags)\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top);\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data, uint8_t const *top, uint64_t flags);\n+static inline uint64_t xdl_hash_record(uint8_t const **data, uint8_t const *top, uint64_t flags)\n {\n \tif (flags & XDF_WHITESPACE_FLAGS)\n \t\treturn xdl_hash_record_with_whitespace(data, top, flags);\n-- \ngitgitgadget\n\n"},{"id":"530713","messageId":"f044274bd586aa389dc4142e9d3cfa0544800a34.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:52Z","receivedAt":"2025-11-14T22:37:06Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe ha field is serving two different purposes, which makes the code\nharder to read. At first glance, it looks like many places assume\nthere could never be hash collisions between lines of the two input\nfiles. In reality, line_hash is used together with xdl_recmatch() to\nensure correct comparisons of lines, even when collisions occur.\n\nTo make this clearer, the old ha field has been split:\n  * line_hash: a straightforward hash of a line, independent of any\n    external context. Its type is uint64_t, as it comes from a fixed\n    width hash function.\n  * minimal_perfect_hash: Not a new concept, but now a separate\n    field. It comes from the classifier's general-purpose hash table,\n    which assigns each line a unique and minimal hash across the two\n    files. A size_t is used here because it's meant to be used to\n    index an array. This also avoids ` as usize` casts on the Rust\n    side when using it to index a slice.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c     |  6 +++---\n xdiff/xhistogram.c |  4 ++--\n xdiff/xpatience.c  | 10 +++++-----\n xdiff/xprepare.c   | 18 +++++++++---------\n xdiff/xtypes.h     |  3 ++-\n 5 files changed, 21 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex edd05466df..436c34697d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -22,9 +22,9 @@\n \n #include \"xinclude.h\"\n \n-static unsigned long get_hash(xdfile_t *xdf, long index)\n+static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].ha;\n+\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -385,7 +385,7 @@ static xdchange_t *xdl_add_change(xdchange_t *xscr, long i1, long i2, long chg1,\n \n static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n {\n-\treturn (rec1->ha == rec2->ha);\n+\treturn rec1->minimal_perfect_hash == rec2->minimal_perfect_hash;\n }\n \n /*\ndiff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\nindex 6dc450b1fe..5ae1282c27 100644\n--- a/xdiff/xhistogram.c\n+++ b/xdiff/xhistogram.c\n@@ -90,7 +90,7 @@ struct region {\n \n static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n {\n-\treturn r1->ha == r2->ha;\n+\treturn r1->minimal_perfect_hash == r2->minimal_perfect_hash;\n \n }\n \n@@ -98,7 +98,7 @@ static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n \t(cmp_recs(REC(i->env, s1, l1), REC(i->env, s2, l2)))\n \n #define TABLE_HASH(index, side, line) \\\n-\tXDL_HASHLONG((REC(index->env, side, line))->ha, index->table_bits)\n+\tXDL_HASHLONG((REC(index->env, side, line))->minimal_perfect_hash, index->table_bits)\n \n static int scanA(struct histindex *index, int line1, int count1)\n {\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex bb61354f22..cc53266f3b 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -48,7 +48,7 @@\n struct hashmap {\n \tint nr, alloc;\n \tstruct entry {\n-\t\tunsigned long hash;\n+\t\tsize_t minimal_perfect_hash;\n \t\t/*\n \t\t * 0 = unused entry, 1 = first line, 2 = second, etc.\n \t\t * line2 is NON_UNIQUE if the line is not unique\n@@ -101,10 +101,10 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t * So we multiply ha by 2 in the hope that the hashing was\n \t * \"unique enough\".\n \t */\n-\tint index = (int)((record->ha << 1) % map->alloc);\n+\tint index = (int)((record->minimal_perfect_hash << 1) % map->alloc);\n \n \twhile (map->entries[index].line1) {\n-\t\tif (map->entries[index].hash != record->ha) {\n+\t\tif (map->entries[index].minimal_perfect_hash != record->minimal_perfect_hash) {\n \t\t\tif (++index >= map->alloc)\n \t\t\t\tindex = 0;\n \t\t\tcontinue;\n@@ -120,7 +120,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \tif (pass == 2)\n \t\treturn;\n \tmap->entries[index].line1 = line;\n-\tmap->entries[index].hash = record->ha;\n+\tmap->entries[index].minimal_perfect_hash = record->minimal_perfect_hash;\n \tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n@@ -248,7 +248,7 @@ static int match(struct hashmap *map, int line1, int line2)\n {\n \txrecord_t *record1 = &map->env->xdf1.recs[line1 - 1];\n \txrecord_t *record2 = &map->env->xdf2.recs[line2 - 1];\n-\treturn record1->ha == record2->ha;\n+\treturn record1->minimal_perfect_hash == record2->minimal_perfect_hash;\n }\n \n static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 85e56021da..bea0992b5e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -93,12 +93,12 @@ static void xdl_free_classifier(xdlclassifier_t *cf) {\n \n \n static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n-\tlong hi;\n+\tsize_t hi;\n \txdlclass_t *rcrec;\n \n-\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n+\thi = XDL_HASHLONG(rec->line_hash, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n-\t\tif (rcrec->rec.ha == rec->ha &&\n+\t\tif (rcrec->rec.line_hash == rec->line_hash &&\n \t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n \t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n@@ -120,7 +120,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n \t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n \n-\trec->ha = (unsigned long) rcrec->idx;\n+\trec->minimal_perfect_hash = (size_t)rcrec->idx;\n \n \treturn 0;\n }\n@@ -158,7 +158,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n-\t\t\tcrec->ha = hav;\n+\t\t\tcrec->line_hash = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\n \t\t}\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -298,7 +298,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -350,7 +350,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs2 = xdf2->recs;\n \tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dstart = xdf2->dstart = i;\n@@ -358,7 +358,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs1 = xdf1->recs + xdf1->nrec - 1;\n \trecs2 = xdf2->recs + xdf2->nrec - 1;\n \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dend = xdf1->nrec - i - 1;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 354349b523..d4e9cd2e76 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -41,7 +41,8 @@ typedef struct s_chastore {\n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n \tsize_t size;\n-\tunsigned long ha;\n+\tuint64_t line_hash;\n+\tsize_t minimal_perfect_hash;\n } xrecord_t;\n \n typedef struct s_xdfile {\n-- \ngitgitgadget\n\n"},{"id":"530712","messageId":"f7a3731d941481fee8cb953c9dd792be500ee96a.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 07/10] xdiff: make xdfile_t.nrec a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:53Z","receivedAt":"2025-11-14T22:37:08Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nrec describes the number of elements for both\nrecs, and for 'changed' + 2.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     | 20 ++++++++++----------\n xdiff/xmerge.c    |  8 ++++----\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  | 12 ++++++------\n xdiff/xtypes.h    |  2 +-\n 6 files changed, 26 insertions(+), 26 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 436c34697d..759193fe5d 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -483,7 +483,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n {\n \tlong i;\n \n-\tif (split >= xdf->nrec) {\n+\tif (split >= (long)xdf->nrec) {\n \t\tm->end_of_file = 1;\n \t\tm->indent = -1;\n \t} else {\n@@ -506,7 +506,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n \n \tm->post_blank = 0;\n \tm->post_indent = -1;\n-\tfor (i = split + 1; i < xdf->nrec; i++) {\n+\tfor (i = split + 1; i < (long)xdf->nrec; i++) {\n \t\tm->post_indent = get_indent(&xdf->recs[i]);\n \t\tif (m->post_indent != -1)\n \t\t\tbreak;\n@@ -717,7 +717,7 @@ static void group_init(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static inline int group_next(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end == xdf->nrec)\n+\tif (g->end == (long)xdf->nrec)\n \t\treturn -1;\n \n \tg->start = g->end + 1;\n@@ -750,7 +750,7 @@ static inline int group_previous(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static int group_slide_down(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end < xdf->nrec &&\n+\tif (g->end < (long)xdf->nrec &&\n \t    recs_match(&xdf->recs[g->start], &xdf->recs[g->end])) {\n \t\txdf->changed[g->start++] = false;\n \t\txdf->changed[g->end++] = true;\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex 2f8007753c..04f7e9193b 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -137,7 +137,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n \tbuf = func_line ? func_line->buf : dummy;\n \tsize = func_line ? sizeof(func_line->buf) : sizeof(dummy);\n \n-\tfor (l = start; l != limit && 0 <= l && l < xe->xdf1.nrec; l += step) {\n+\tfor (l = start; l != limit && 0 <= l && l < (long)xe->xdf1.nrec; l += step) {\n \t\tlong len = match_func_rec(&xe->xdf1, xecfg, l, buf, size);\n \t\tif (len >= 0) {\n \t\t\tif (func_line)\n@@ -179,14 +179,14 @@ pre_context_calculation:\n \t\t\tlong fs1, i1 = xch->i1;\n \n \t\t\t/* Appended chunk? */\n-\t\t\tif (i1 >= xe->xdf1.nrec) {\n+\t\t\tif (i1 >= (long)xe->xdf1.nrec) {\n \t\t\t\tlong i2 = xch->i2;\n \n \t\t\t\t/*\n \t\t\t\t * We don't need additional context if\n \t\t\t\t * a whole function was added.\n \t\t\t\t */\n-\t\t\t\twhile (i2 < xe->xdf2.nrec) {\n+\t\t\t\twhile (i2 < (long)xe->xdf2.nrec) {\n \t\t\t\t\tif (is_func_rec(&xe->xdf2, xecfg, i2))\n \t\t\t\t\t\tgoto post_context_calculation;\n \t\t\t\t\ti2++;\n@@ -196,7 +196,7 @@ pre_context_calculation:\n \t\t\t\t * Otherwise get more context from the\n \t\t\t\t * pre-image.\n \t\t\t\t */\n-\t\t\t\ti1 = xe->xdf1.nrec - 1;\n+\t\t\t\ti1 = (long)xe->xdf1.nrec - 1;\n \t\t\t}\n \n \t\t\tfs1 = get_func_line(xe, xecfg, NULL, i1, -1);\n@@ -228,8 +228,8 @@ pre_context_calculation:\n \n  post_context_calculation:\n \t\tlctx = xecfg->ctxlen;\n-\t\tlctx = XDL_MIN(lctx, xe->xdf1.nrec - (xche->i1 + xche->chg1));\n-\t\tlctx = XDL_MIN(lctx, xe->xdf2.nrec - (xche->i2 + xche->chg2));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf1.nrec - (xche->i1 + xche->chg1));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf2.nrec - (xche->i2 + xche->chg2));\n \n \t\te1 = xche->i1 + xche->chg1 + lctx;\n \t\te2 = xche->i2 + xche->chg2 + lctx;\n@@ -237,13 +237,13 @@ pre_context_calculation:\n \t\tif (xecfg->flags & XDL_EMIT_FUNCCONTEXT) {\n \t\t\tlong fe1 = get_func_line(xe, xecfg, NULL,\n \t\t\t\t\t\t xche->i1 + xche->chg1,\n-\t\t\t\t\t\t xe->xdf1.nrec);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec);\n \t\t\twhile (fe1 > 0 && is_empty_rec(&xe->xdf1, fe1 - 1))\n \t\t\t\tfe1--;\n \t\t\tif (fe1 < 0)\n-\t\t\t\tfe1 = xe->xdf1.nrec;\n+\t\t\t\tfe1 = (long)xe->xdf1.nrec;\n \t\t\tif (fe1 > e1) {\n-\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), xe->xdf2.nrec);\n+\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), (long)xe->xdf2.nrec);\n \t\t\t\te1 = fe1;\n \t\t\t}\n \n@@ -254,7 +254,7 @@ pre_context_calculation:\n \t\t\t */\n \t\t\tif (xche->next) {\n \t\t\t\tlong l = XDL_MIN(xche->next->i1,\n-\t\t\t\t\t\t xe->xdf1.nrec - 1);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec - 1);\n \t\t\t\tif (l - xecfg->ctxlen <= e1 ||\n \t\t\t\t    get_func_line(xe, xecfg, NULL, l, e1) < 0) {\n \t\t\t\t\txche = xche->next;\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 0dd4558a32..29dad98c49 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -158,7 +158,7 @@ static int is_eol_crlf(xdfile_t *file, int i)\n {\n \tsize_t size;\n \n-\tif (i < file->nrec - 1)\n+\tif (i < (long)file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n \t\treturn (size = file->recs[i].size) > 1 &&\n \t\t\tfile->recs[i].ptr[size - 2] == '\\r';\n@@ -317,7 +317,7 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \t\t\tcontinue;\n \t\ti = m->i1 + m->chg1;\n \t}\n-\tsize += xdl_recs_copy(xe1, i, xe1->xdf2.nrec - i, 0, 0,\n+\tsize += xdl_recs_copy(xe1, i, (int)xe1->xdf2.nrec - i, 0, 0,\n \t\t\t      dest ? dest + size : NULL);\n \treturn size;\n }\n@@ -622,7 +622,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\t\tchanges = c;\n \t\ti0 = xscr1->i1;\n \t\ti1 = xscr1->i2;\n-\t\ti2 = xscr1->i1 + xe2->xdf2.nrec - xe2->xdf1.nrec;\n+\t\ti2 = xscr1->i1 + (long)xe2->xdf2.nrec - (long)xe2->xdf1.nrec;\n \t\tchg0 = xscr1->chg1;\n \t\tchg1 = xscr1->chg2;\n \t\tchg2 = xscr1->chg1;\n@@ -637,7 +637,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\tif (!changes)\n \t\t\tchanges = c;\n \t\ti0 = xscr2->i1;\n-\t\ti1 = xscr2->i1 + xe1->xdf2.nrec - xe1->xdf1.nrec;\n+\t\ti1 = xscr2->i1 + (long)xe1->xdf2.nrec - (long)xe1->xdf1.nrec;\n \t\ti2 = xscr2->i2;\n \t\tchg0 = xscr2->chg1;\n \t\tchg1 = xscr2->chg1;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex cc53266f3b..a0b31eb5d8 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -370,5 +370,5 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n \n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n-\treturn patience_diff(xpp, env, 1, env->xdf1.nrec, 1, env->xdf2.nrec);\n+\treturn patience_diff(xpp, env, 1, (int)env->xdf1.nrec, 1, (int)env->xdf2.nrec);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex bea0992b5e..705ddd1ae0 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -153,7 +153,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\tfor (top = blk + bsize; cur < top; ) {\n \t\t\tprev = cur;\n \t\t\thav = xdl_hash_record(&cur, top, xpp->flags);\n-\t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n+\t\t\tif (XDL_ALLOC_GROW(xdf->recs, (long)xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n@@ -287,7 +287,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -295,7 +295,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -348,7 +348,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \n \trecs1 = xdf1->recs;\n \trecs2 = xdf2->recs;\n-\tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n+\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n@@ -361,8 +361,8 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n-\txdf1->dend = xdf1->nrec - i - 1;\n-\txdf2->dend = xdf2->nrec - i - 1;\n+\txdf1->dend = (long)xdf1->nrec - i - 1;\n+\txdf2->dend = (long)xdf2->nrec - i - 1;\n \n \treturn 0;\n }\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex d4e9cd2e76..4c4d9bd147 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,7 +47,7 @@ typedef struct s_xrecord {\n \n typedef struct s_xdfile {\n \txrecord_t *recs;\n-\tlong nrec;\n+\tsize_t nrec;\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n-- \ngitgitgadget\n\n"},{"id":"530714","messageId":"93f84ae72e42c7321e8ad028e86b8a2e8c7f8f6d.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 08/10] xdiff: make xdfile_t.nreff a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:54Z","receivedAt":"2025-11-14T22:37:09Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nreff describes the number of elements in memory\nfor rindex.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 +++++++-------\n xdiff/xtypes.h   |  2 +-\n 2 files changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 705ddd1ae0..39fd79d9d4 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -264,7 +264,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, nreff, mlim;\n+\tlong i, nm, mlim;\n \txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n@@ -307,29 +307,29 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * Use temporary arrays to decide if changed[i] should remain\n \t * false, or become true.\n \t */\n-\tfor (nreff = 0, i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n+\txdf1->nreff = 0;\n+\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[nreff++] = i;\n+\t\t\txdf1->rindex[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf1->nreff = nreff;\n \n-\tfor (nreff = 0, i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n+\txdf2->nreff = 0;\n+\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[nreff++] = i;\n+\t\t\txdf2->rindex[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf2->nreff = nreff;\n \n cleanup:\n \txdl_free(action1);\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 4c4d9bd147..1f495f987f 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -51,7 +51,7 @@ typedef struct s_xdfile {\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n-\tlong nreff;\n+\tsize_t nreff;\n } xdfile_t;\n \n typedef struct s_xdfenv {\n-- \ngitgitgadget\n\n"},{"id":"530716","messageId":"39369becc87842315480b6c351ff8c6a3352be68.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 09/10] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:55Z","receivedAt":"2025-11-14T22:37:10Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe field rindex describes an index offset for other arrays. Change it\nto size_t.\n\nChanging the type of rindex from long to size_t has no cascading\nrefactor impact because it is only ever used to directly index other\narrays.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 1f495f987f..9074cdadd1 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -50,7 +50,7 @@ typedef struct s_xdfile {\n \tsize_t nrec;\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n-\tlong *rindex;\n+\tsize_t *rindex;\n \tsize_t nreff;\n } xdfile_t;\n \n-- \ngitgitgadget\n\n"},{"id":"530717","messageId":"950d1e6193877faba44899ca6b9c274a889a22f9.1763159816.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v4 10/10] xdiff: rename rindex -> reference_index","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-14T22:36:56Z","receivedAt":"2025-11-14T22:37:11Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe classic diff adds only the lines that it's going to consider,\nduring the diff, to an array. A mapping between the compacted\narray, and the lines of the file that they reference, is\nfacilitated by this array.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  6 +++---\n xdiff/xprepare.c | 10 +++++-----\n xdiff/xtypes.h   |  2 +-\n 3 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 759193fe5d..8eb664be3e 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -24,7 +24,7 @@\n \n static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n+\treturn xdf->recs[xdf->reference_index[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -278,10 +278,10 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n \t */\n \tif (off1 == lim1) {\n \t\tfor (; off2 < lim2; off2++)\n-\t\t\txdf2->changed[xdf2->rindex[off2]] = true;\n+\t\t\txdf2->changed[xdf2->reference_index[off2]] = true;\n \t} else if (off2 == lim2) {\n \t\tfor (; off1 < lim1; off1++)\n-\t\t\txdf1->changed[xdf1->rindex[off1]] = true;\n+\t\t\txdf1->changed[xdf1->reference_index[off1]] = true;\n \t} else {\n \t\txdpsplit_t spl;\n \t\tspl.i1 = spl.i2 = 0;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 39fd79d9d4..34c82e4f8e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -128,7 +128,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n static void xdl_free_ctx(xdfile_t *xdf)\n {\n-\txdl_free(xdf->rindex);\n+\txdl_free(xdf->reference_index);\n \txdl_free(xdf->changed - 1);\n \txdl_free(xdf->recs);\n }\n@@ -141,7 +141,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n-\txdf->rindex = NULL;\n+\txdf->reference_index = NULL;\n \txdf->changed = NULL;\n \txdf->recs = NULL;\n \n@@ -169,7 +169,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF)) {\n-\t\tif (!XDL_ALLOC_ARRAY(xdf->rindex, xdf->nrec + 1))\n+\t\tif (!XDL_ALLOC_ARRAY(xdf->reference_index, xdf->nrec + 1))\n \t\t\tgoto abort;\n \t}\n \n@@ -312,7 +312,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[xdf1->nreff++] = i;\n+\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n@@ -324,7 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[xdf2->nreff++] = i;\n+\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 9074cdadd1..979586f20a 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -50,7 +50,7 @@ typedef struct s_xdfile {\n \tsize_t nrec;\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n-\tsize_t *rindex;\n+\tsize_t *reference_index;\n \tsize_t nreff;\n } xdfile_t;\n \n-- \ngitgitgadget\n"},{"id":"530726","messageId":"23b7fd8a-2b50-4da3-bc8a-3727ee99654f@ramsayjones.plus.com","threadId":"64326","inReplyTo":"af732beb6904d8b9b7801ecc2487dabdea05a571.1763159816.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsayjones.plus.com","sentAt":"2025-11-15T03:06:06Z","receivedAt":"2025-11-15T03:09:16Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"\n\nOn 14/11/2025 10:36 pm, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Document other nuances when crossing the FFI boundary. Other language\n> mappings may be added in the future.\n> \n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  Documentation/Makefile                        |   1 +\n>  Documentation/technical/meson.build           |   1 +\n>  .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n>  3 files changed, 226 insertions(+)\n>  create mode 100644 Documentation/technical/unambiguous-types.adoc\n> \n[snip]\n\n> +== Character types\n> +\n> +This is where C and Rust don't have a clean one-to-one mapping.\n> +\n> +A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n> +a `char` will have the same size as the corresponding Rust struct using `u8`.\n> +In that sense, such structs are safe to pass over the FFI boundary, because\n> +their fields will be laid out identically. However, beyond bit width, C `char`\n> +has additional semantics and platform-dependent behavior that can cause\n> +problems, as discussed below.\n> +\n> +C comparison problem: While the sign of `char` is implementation defined, it's\n> +also signless (neither signed nor unsigned). When building with\n\nHmm, this sets my teeth on edge. The C char type is not 'signless' (whatever that is\nsupposed to mean), it's 'sign-ness' is implementation-defined behaviour. This means\nthat it is 'unspecified behavior where each implementation documents how the choice\nis made'. In particular, it has to document:\n\n  \"Which of signed char or unsigned char has the same range, representation, and\n   behavior as \"plain\" char (6.2.5, 6.3.1.1).\"\n\n(it is still a distinct type, however). Note that some compilers even allow you to\nspecify which you want for a given compilation! (see gcc options -f[un]signed-char\nand their inverse 'no' options!)\n\n\nATB,\nRamsay Jones\n\n\n"},{"id":"530728","messageId":"5A740EE4-D545-4828-8D38-E0E5E9F87A3E@gmail.com","threadId":"64326","inReplyTo":"23b7fd8a-2b50-4da3-bc8a-3727ee99654f@ramsayjones.plus.com","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-11-15T03:41:36Z","receivedAt":"2025-11-15T03:41:48Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 14 nov. 2025 à 22:09, Ramsay Jones <ramsay@ramsayjones.plus.com> a écrit :\n> \n> ﻿\n> \n>> On 14/11/2025 10:36 pm, Ezekiel Newren via GitGitGadget wrote:\n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>> \n>> Document other nuances when crossing the FFI boundary. Other language\n>> mappings may be added in the future.\n>> \n>> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>> ---\n>> Documentation/Makefile                        |   1 +\n>> Documentation/technical/meson.build           |   1 +\n>> .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n>> 3 files changed, 226 insertions(+)\n>> create mode 100644 Documentation/technical/unambiguous-types.adoc\n>> \n> [snip]\n> \n>> +== Character types\n>> +\n>> +This is where C and Rust don't have a clean one-to-one mapping.\n>> +\n>> +A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n>> +a `char` will have the same size as the corresponding Rust struct using `u8`.\n>> +In that sense, such structs are safe to pass over the FFI boundary, because\n>> +their fields will be laid out identically. However, beyond bit width, C `char`\n>> +has additional semantics and platform-dependent behavior that can cause\n>> +problems, as discussed below.\n>> +\n>> +C comparison problem: While the sign of `char` is implementation defined, it's\n>> +also signless (neither signed nor unsigned). When building with\n> \n> Hmm, this sets my teeth on edge. The C char type is not 'signless' (whatever that is\n> supposed to mean), it's 'sign-ness' is implementation-defined behaviour. This means\n> that it is 'unspecified behavior where each implementation documents how the choice\n> is made'. In particular, it has to document:\n> \n>  \"Which of signed char or unsigned char has the same range, representation, and\n>   behavior as \"plain\" char (6.2.5, 6.3.1.1).\"\n> \n> (it is still a distinct type, however). Note that some compilers even allow you to\n> specify which you want for a given compilation! (see gcc options -f[un]signed-char\n> and their inverse 'no' options!)\n> \n> \n> ATB,\n> Ramsay Jones\n\nThis was discussed briefly in replies to v2’s 2/10, where Ezekiel said that DEVELOPER=1 warned about sign issues whether char was compared to int or unsigned. [From mobile I cannot reliably paste the message ID or link and preserve a plain-text email, apologies for the oblique reference.]"},{"id":"530731","messageId":"xmqqy0o7g0rk.fsf@gitster.g","threadId":"64326","inReplyTo":"042fbb11d03606879503846e86fac65e6e74d02a.1763159816.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-15T08:26:07Z","receivedAt":"2025-11-15T08:26:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> In order to avoid a refactor avalanche, many uses of this field were\n> cast to char* or similar. One exception is in get_indent() where the\n> local variable `char c` was changed to `uint8_t c`.\n\nI actually think keeping \"char c\" as in the original is a lot more\nlogical for that particular case, as the existing use of that local\nvariable are _all_ about C's 'char', and not about a very short\nunsigned integer.  The variable is compared with C's character\nconstants like ' ' (whitespace) and '\\t' (horizontal tab), or is\ngiven to XDL_ISSPACE() macro, which is also about C's character.\n\nBut because it is so minor a thing, I do not think that it deserves\na reroll on its own.  Just in case if there are other things that\nneed to change and the series needs a reroll, here is the only\nchange required for this.\n\n\n xdiff/xdiffi.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git c/xdiff/xdiffi.c w/xdiff/xdiffi.c\nindex 8eb664be3e..4376f943db 100644\n--- c/xdiff/xdiffi.c\n+++ w/xdiff/xdiffi.c\n@@ -406,7 +406,7 @@ static int get_indent(xrecord_t *rec)\n \tint ret = 0;\n \n \tfor (size_t i = 0; i < rec->size; i++) {\n-\t\tuint8_t c = rec->ptr[i];\n+\t\tchar c = (char) rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n \t\t\treturn ret;\n"},{"id":"530742","messageId":"a30ad114-61c2-4eed-a24e-033b3b9d6d0c@ramsayjones.plus.com","threadId":"64326","inReplyTo":"5A740EE4-D545-4828-8D38-E0E5E9F87A3E@gmail.com","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsayjones.plus.com","sentAt":"2025-11-15T14:55:07Z","receivedAt":"2025-11-15T14:58:17Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"\n\nOn 15/11/2025 3:41 am, Ben Knoble wrote:\n> \n>> Le 14 nov. 2025 à 22:09, Ramsay Jones <ramsay@ramsayjones.plus.com> a écrit :\n>>\n>> ﻿\n>>\n>>> On 14/11/2025 10:36 pm, Ezekiel Newren via GitGitGadget wrote:\n>>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>>\n>>> Document other nuances when crossing the FFI boundary. Other language\n>>> mappings may be added in the future.\n>>>\n>>> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>>> ---\n>>> Documentation/Makefile                        |   1 +\n>>> Documentation/technical/meson.build           |   1 +\n>>> .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n>>> 3 files changed, 226 insertions(+)\n>>> create mode 100644 Documentation/technical/unambiguous-types.adoc\n>>>\n>> [snip]\n>>\n>>> +== Character types\n>>> +\n>>> +This is where C and Rust don't have a clean one-to-one mapping.\n>>> +\n>>> +A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n>>> +a `char` will have the same size as the corresponding Rust struct using `u8`.\n>>> +In that sense, such structs are safe to pass over the FFI boundary, because\n>>> +their fields will be laid out identically. However, beyond bit width, C `char`\n>>> +has additional semantics and platform-dependent behavior that can cause\n>>> +problems, as discussed below.\n>>> +\n>>> +C comparison problem: While the sign of `char` is implementation defined, it's\n>>> +also signless (neither signed nor unsigned). When building with\n>>\n>> Hmm, this sets my teeth on edge. The C char type is not 'signless' (whatever that is\n>> supposed to mean), it's 'sign-ness' is implementation-defined behaviour. This means\n>> that it is 'unspecified behavior where each implementation documents how the choice\n>> is made'. In particular, it has to document:\n>>\n>>  \"Which of signed char or unsigned char has the same range, representation, and\n>>   behavior as \"plain\" char (6.2.5, 6.3.1.1).\"\n>>\n>> (it is still a distinct type, however). Note that some compilers even allow you to\n>> specify which you want for a given compilation! (see gcc options -f[un]signed-char\n>> and their inverse 'no' options!)\n>>\n>>\n>> ATB,\n>> Ramsay Jones\n> \n> This was discussed briefly in replies to v2’s 2/10, where Ezekiel said that DEVELOPER=1 warned about sign issues whether char was compared to int or unsigned. [From mobile I cannot reliably paste the message ID or link and preserve a plain-text email, apologies for the oblique reference.]\n\nErr... sorry, but I don't see how this comment relates to my email. puzzled! ;)\n\nATB,\nRamsay Jones\n\n\n\n"},{"id":"530746","messageId":"xmqqpl9jfdso.fsf@gitster.g","threadId":"64326","inReplyTo":"a30ad114-61c2-4eed-a24e-033b3b9d6d0c@ramsayjones.plus.com","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-15T16:42:15Z","receivedAt":"2025-11-15T16:42:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ramsay Jones <ramsay@ramsayjones.plus.com> writes:\n\n>> This was discussed briefly in replies to v2’s 2/10, where\n>> Ezekiel said that DEVELOPER=1 warned about sign issues whether\n>> char was compared to int or unsigned. [From mobile I cannot\n>> reliably paste the message ID or link and preserve a plain-text\n>> email, apologies for the oblique reference.]\n>\n> Err... sorry, but I don't see how this comment relates to my\n> email. puzzled! ;)\n\nMe neither, but I suspect it may mostly use of non-word \"signless\"\nthat is the issue.  It is understandable for the -Wsign-compare\nwarning (especially given that it very often complains about\nperfectly good pieces of code) to complain when you compare a \"char\"\nwith a signed integer, saying \"on a platform where 'char' is\nunsigned, you would be comparing signed and unsigned values with\nthis expression\", and at the same time complain when you compare a\n\"char\" with an unsigned integer, saying \"on a platform where 'char'\nis signed...\".\n\nI'd say it shows more about how garbage -Wsign-compare is than about\nhow 'char' is ambiguous and should be avoided, but others may have\ndifferent opinions.\n\n\n"},{"id":"530749","messageId":"CALnO6CA-6waRpkqzLxR+f2yzwfhmf_jvbtEZC7FAFN9NLkqkXg@mail.gmail.com","threadId":"64326","inReplyTo":"xmqqpl9jfdso.fsf@gitster.g","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-11-15T16:59:08Z","receivedAt":"2025-11-15T16:59:21Z","isPatch":true,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Sat, Nov 15, 2025 at 11:42 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Ramsay Jones <ramsay@ramsayjones.plus.com> writes:\n>\n> >> This was discussed briefly in replies to v2’s 2/10, where\n> >> Ezekiel said that DEVELOPER=1 warned about sign issues whether\n> >> char was compared to int or unsigned. [From mobile I cannot\n> >> reliably paste the message ID or link and preserve a plain-text\n> >> email, apologies for the oblique reference.]\n> >\n> > Err... sorry, but I don't see how this comment relates to my\n> > email. puzzled! ;)\n>\n> Me neither, but I suspect it may mostly use of non-word \"signless\"\n> that is the issue.  It is understandable for the -Wsign-compare\n> warning (especially given that it very often complains about\n> perfectly good pieces of code) to complain when you compare a \"char\"\n> with a signed integer, saying \"on a platform where 'char' is\n> unsigned, you would be comparing signed and unsigned values with\n> this expression\", and at the same time complain when you compare a\n> \"char\" with an unsigned integer, saying \"on a platform where 'char'\n> is signed...\".\n\nAgreed, and I suspect this is roughly the implementation.\n\nMy point was that Ezekiel seemed to justify (?) the use of \"signless\"\nby pointing to those warnings (I personally am on the fence for how to\ntreat the combination of facts, but it seems useful to consider that\nchar is not easily comparable with integers of various signedness).\n\n-- \nD. Ben Knoble\n"},{"id":"530758","messageId":"xmqqo6p3dpw4.fsf@gitster.g","threadId":"64326","inReplyTo":"CALnO6CA-6waRpkqzLxR+f2yzwfhmf_jvbtEZC7FAFN9NLkqkXg@mail.gmail.com","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-15T20:03:55Z","receivedAt":"2025-11-15T20:03:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"D. Ben Knoble\" <ben.knoble@gmail.com> writes:\n\n>> Me neither, but I suspect it may mostly use of non-word \"signless\"\n>> that is the issue.\n> ...\n> Agreed, and I suspect this is roughly the implementation.\n>\n> My point was that Ezekiel seemed to justify (?) the use of \"signless\"\n> by pointing to those warnings (I personally am on the fence for how to\n> treat the combination of facts, but it seems useful to consider that\n> char is not easily comparable with integers of various signedness).\n\nAgreed.  I think your point matches my suspicion that the use of the\nnon-word \"signless\" was what Ramsay reacted.  The 'char' with the\nimplementation defined signedness is making -Wsign-compare even more\nquirky than it already is.\n"},{"id":"530785","messageId":"xmqqzf8la20o.fsf@gitster.g","threadId":"64326","inReplyTo":"xmqqpl9jfdso.fsf@gitster.g","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-17T01:20:07Z","receivedAt":"2025-11-17T01:20:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Me neither, but I suspect it may mostly use of non-word \"signless\"\n> that is the issue.\n\nSo, the patch text that claims C's \"char\" is \"signless\" still needs\nto be updated, I think.  The problematic paragraph (with a bit of\nrewrapping) reads like this:\n\n    C comparison problem: While the sign of `char` is implementation\n    defined, it's also signless (neither signed nor unsigned). When\n    building with `make DEVELOPER=1` it will complain about a\n    \"differ in signedness\" when `char` is compared with `uint8_t` or\n    `int8_t`.\n\nPerhaps\n\n    The C language leaves the signedness of `char` implementation\n    defined.  Because our developer build enables -Wsign-compare,\n    comparison of a value of `char` type with either signed or\n    unsigned integers will trigger warnings from the compiler.\n    Avoiding `char` of implementation defined signedness helps us\n    being a bit more explicit.\n\nor something is sufficient?\n\n"},{"id":"530786","messageId":"dd012ab5-d239-49b8-8635-8d22e16c9f1c@ramsayjones.plus.com","threadId":"64326","inReplyTo":"xmqqzf8la20o.fsf@gitster.g","subject":"Re: [PATCH v4 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsayjones.plus.com","sentAt":"2025-11-17T02:08:42Z","receivedAt":"2025-11-17T02:08:51Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"\n\nOn 17/11/2025 1:20 am, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n>> Me neither, but I suspect it may mostly use of non-word \"signless\"\n>> that is the issue.\n> \n> So, the patch text that claims C's \"char\" is \"signless\" still needs\n> to be updated, I think.  The problematic paragraph (with a bit of\n> rewrapping) reads like this:\n\nSorry for being AFK for a over a day! :) I didn't think this would\ngenerate so much traffic.\n\n>     C comparison problem: While the sign of `char` is implementation\n>     defined, it's also signless (neither signed nor unsigned). When\n>     building with `make DEVELOPER=1` it will complain about a\n>     \"differ in signedness\" when `char` is compared with `uint8_t` or\n>     `int8_t`.\n\nYes, the 'signless' nonsense is what 'triggered' me. ;)\n\n> \n> Perhaps\n> \n>     The C language leaves the signedness of `char` implementation\n>     defined.  Because our developer build enables -Wsign-compare,\n>     comparison of a value of `char` type with either signed or\n>     unsigned integers will trigger warnings from the compiler.\n\ns/will/may/ - it depends!\n\n>     Avoiding `char` of implementation defined signedness helps us\n>     being a bit more explicit.\n> \n> or something is sufficient?\n\nYes, this looks good to me (but then I am not particularly good at\nword-smithing).\n\nThanks.\n\nATB,\nRamsay Jones\n\n\n\n"},{"id":"530924","messageId":"CAH=ZcbDzERvz7ZZ+yFOgEhtoBw3Ym1_2YPL3mbj2p0k7AK0v8w@mail.gmail.com","threadId":"64326","inReplyTo":"xmqqy0o7g0rk.fsf@gitster.g","subject":"Re: [PATCH v4 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2025-11-18T20:55:45Z","receivedAt":"2025-11-18T20:55:58Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"On Sat, Nov 15, 2025 at 1:26 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > In order to avoid a refactor avalanche, many uses of this field were\n> > cast to char* or similar. One exception is in get_indent() where the\n> > local variable `char c` was changed to `uint8_t c`.\n>\n> I actually think keeping \"char c\" as in the original is a lot more\n> logical for that particular case, as the existing use of that local\n> variable are _all_ about C's 'char', and not about a very short\n> unsigned integer.  The variable is compared with C's character\n> constants like ' ' (whitespace) and '\\t' (horizontal tab), or is\n> given to XDL_ISSPACE() macro, which is also about C's character.\n>\n> But because it is so minor a thing, I do not think that it deserves\n> a reroll on its own.  Just in case if there are other things that\n> need to change and the series needs a reroll, here is the only\n> change required for this.\n>\n>\n>  xdiff/xdiffi.c | 2 +-\n>  1 file changed, 1 insertion(+), 1 deletion(-)\n>\n> diff --git c/xdiff/xdiffi.c w/xdiff/xdiffi.c\n> index 8eb664be3e..4376f943db 100644\n> --- c/xdiff/xdiffi.c\n> +++ w/xdiff/xdiffi.c\n> @@ -406,7 +406,7 @@ static int get_indent(xrecord_t *rec)\n>         int ret = 0;\n>\n>         for (size_t i = 0; i < rec->size; i++) {\n> -               uint8_t c = rec->ptr[i];\n> +               char c = (char) rec->ptr[i];\n>\n>                 if (!XDL_ISSPACE(c))\n>                         return ret;\n\nI have v5 ready to go, but there seems to be a problem with\ngitgitgadget. Once that's resolved I'll post the new version.\n"},{"id":"530928","messageId":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v4.git.git.1763159816.gitgitgadget@gmail.com","subject":"[PATCH v5 00/10] Xdiff cleanup part2","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:12Z","receivedAt":"2025-11-18T22:34:24Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"Changes in v5:\n\n * Remove the non-word 'signless', and rephrase that paragraph in\n   unambiguous-types.adoc\n * Cast to char in xdiffi.c:get_indent() rather than changing the local\n   variable to uint8_t\n\nChanges in v4:\n\n * Update documentation to not mention Unicode except once\n * Don't move dstart/dend with in the xdfile_t struct\n * Rephrase justification on changing xrecord_t.ptr's type\n\nChanges in v3:\n\n * Address comments about commit messages and documentation\n * Add unambiguous-types.adoc to Makefile and Meson\n * Use markdown style to avoid asciidoc issues\n\nChanges in v2:\n\n * Added documentation about unambiguous types and FFI\n * Addressed comments on the mailing list\n\n\nOriginal cover letter below:\n============================\n\nMaintainer note: This patch series builds on top of en/xdiff-cleanup and\nam/xdiff-hash-tweak (both of which are now in master).\n\nThe primary goal of this patch series is to convert every field's type in\nxrecord_t and xdfile_t to be unambiguous, in preparation to make it more\nRust FFI friendly. Additionally the ha field in xrecord_t is split into\nline_hash and minimal_perfect hash.\n\nThe order of some of the fields has changed as called out by the commit\nmessages.\n\nBefore:\n\ntypedef struct s_xrecord {\n\tchar const *ptr;\n\tlong size;\n\tunsigned long ha;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tlong nrec;\n\tlong dstart, dend;\n\tbool *changed;\n\tlong *rindex;\n\tlong nreff;\n} xdfile_t;\n\n\nAfter part 2\n\ntypedef struct s_xrecord {\n\tuint8_t const *ptr;\n\tsize_t size;\n\tuint64_t line_hash;\n\tsize_t minimal_perfect_hash;\n} xrecord_t;\n\ntypedef struct s_xdfile {\n\txrecord_t *recs;\n\tsize_t nrec;\n\tptrdiff_t dstart, dend;\n\tbool *changed;\n\tsize_t *reference_index;\n\tsize_t nreff;\n} xdfile_t;\n\n\nEzekiel Newren (10):\n  doc: define unambiguous type mappings across C and Rust\n  xdiff: use ptrdiff_t for dstart/dend\n  xdiff: make xrecord_t.ptr a uint8_t instead of char\n  xdiff: use size_t for xrecord_t.size\n  xdiff: use unambiguous types in xdl_hash_record()\n  xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n  xdiff: make xdfile_t.nrec a size_t instead of long\n  xdiff: make xdfile_t.nreff a size_t instead of long\n  xdiff: change rindex from long to size_t in xdfile_t\n  xdiff: rename rindex -> reference_index\n\n Documentation/Makefile                        |   1 +\n Documentation/technical/meson.build           |   1 +\n .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n xdiff-interface.c                             |   2 +-\n xdiff/xdiffi.c                                |  29 ++-\n xdiff/xemit.c                                 |  28 +--\n xdiff/xhistogram.c                            |   4 +-\n xdiff/xmerge.c                                |  30 +--\n xdiff/xpatience.c                             |  14 +-\n xdiff/xprepare.c                              |  60 ++---\n xdiff/xtypes.h                                |  15 +-\n xdiff/xutils.c                                |  32 +--\n xdiff/xutils.h                                |   6 +-\n 13 files changed, 336 insertions(+), 110 deletions(-)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\n\nbase-commit: a99f379adf116d53eb11957af5bab5214915f91d\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2070%2Fezekielnewren%2Fxdiff_cleanup_part2-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2070/ezekielnewren/xdiff_cleanup_part2-v5\nPull-Request: https://github.com/git/git/pull/2070\n\nRange-diff vs v4:\n\n  1:  af732beb69 !  1:  8b56bf1172 doc: define unambiguous type mappings across C and Rust\n     @@ Documentation/technical/unambiguous-types.adoc (new)\n      +has additional semantics and platform-dependent behavior that can cause\n      +problems, as discussed below.\n      +\n     -+C comparison problem: While the sign of `char` is implementation defined, it's\n     -+also signless (neither signed nor unsigned). When building with\n     -+`make DEVELOPER=1` it will complain about a \"differ in signedness\" when `char`\n     -+is compared with `uint8_t` or `int8_t`.\n     ++The C language leaves the signedness of `char` implementation defined. Because\n     ++our developer build enables -Wsign-compare, comparison of a value of `char`\n     ++type with either signed or unsigned integers may trigger warnings from the\n     ++compiler.\n      +\n      +Note: Rust's `char` type is an unsigned 32-bit integer that is used to describe\n      +Unicode code points.\n  2:  b60a03eb31 =  2:  c4193d11f5 xdiff: use ptrdiff_t for dstart/dend\n  3:  042fbb11d0 !  3:  dd76d4f586 xdiff: make xrecord_t.ptr a uint8_t instead of char\n     @@ Commit message\n          Make xrecord_t.ptr uint8_t because it's referring to bytes in memory.\n      \n          In order to avoid a refactor avalanche, many uses of this field were\n     -    cast to char* or similar. One exception is in get_indent() where the\n     -    local variable `char c` was changed to `uint8_t c`.\n     +    cast to char* or similar.\n      \n          Places where casting was unnecessary:\n          xemit.c:156\n     @@ xdiff/xdiffi.c: static int get_indent(xrecord_t *rec)\n       \n       \tfor (i = 0; i < rec->size; i++) {\n      -\t\tchar c = rec->ptr[i];\n     -+\t\tuint8_t c = rec->ptr[i];\n     ++\t\tchar c = (char) rec->ptr[i];\n       \n       \t\tif (!XDL_ISSPACE(c))\n       \t\t\treturn ret;\n  4:  c103fa6bea !  4:  11cec1d2ec xdiff: use size_t for xrecord_t.size\n     @@ xdiff/xdiffi.c: static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n       \n      -\tfor (i = 0; i < rec->size; i++) {\n      +\tfor (size_t i = 0; i < rec->size; i++) {\n     - \t\tuint8_t c = rec->ptr[i];\n     + \t\tchar c = (char) rec->ptr[i];\n       \n       \t\tif (!XDL_ISSPACE(c))\n      @@ xdiff/xdiffi.c: static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n  5:  2ee9a74653 =  5:  6f267360b7 xdiff: use unambiguous types in xdl_hash_record()\n  6:  f044274bd5 =  6:  78af0f16f4 xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n  7:  f7a3731d94 =  7:  5c19f9ded3 xdiff: make xdfile_t.nrec a size_t instead of long\n  8:  93f84ae72e =  8:  d1f498edb1 xdiff: make xdfile_t.nreff a size_t instead of long\n  9:  39369becc8 =  9:  bc4941c146 xdiff: change rindex from long to size_t in xdfile_t\n 10:  950d1e6193 = 10:  dcc9d6bfaf xdiff: rename rindex -> reference_index\n\n-- \ngitgitgadget\n"},{"id":"530929","messageId":"8b56bf117289ca3be25533a36da1ea0c178ccfca.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:13Z","receivedAt":"2025-11-18T22:34:25Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nDocument other nuances when crossing the FFI boundary. Other language\nmappings may be added in the future.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n Documentation/Makefile                        |   1 +\n Documentation/technical/meson.build           |   1 +\n .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n 3 files changed, 226 insertions(+)\n create mode 100644 Documentation/technical/unambiguous-types.adoc\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 04e9e10b27..bc1adb2d9d 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -142,6 +142,7 @@ TECH_DOCS += technical/shallow\n TECH_DOCS += technical/sparse-checkout\n TECH_DOCS += technical/sparse-index\n TECH_DOCS += technical/trivial-merge\n+TECH_DOCS += technical/unambiguous-types\n TECH_DOCS += technical/unit-tests\n SP_ARTICLES += $(TECH_DOCS)\n SP_ARTICLES += technical/api-index\ndiff --git a/Documentation/technical/meson.build b/Documentation/technical/meson.build\nindex be698ef22a..89a6e26821 100644\n--- a/Documentation/technical/meson.build\n+++ b/Documentation/technical/meson.build\n@@ -32,6 +32,7 @@ articles = [\n   'sparse-checkout.adoc',\n   'sparse-index.adoc',\n   'trivial-merge.adoc',\n+  'unambiguous-types.adoc',\n   'unit-tests.adoc',\n ]\n \ndiff --git a/Documentation/technical/unambiguous-types.adoc b/Documentation/technical/unambiguous-types.adoc\nnew file mode 100644\nindex 0000000000..9a4990847c\n--- /dev/null\n+++ b/Documentation/technical/unambiguous-types.adoc\n@@ -0,0 +1,224 @@\n+= Unambiguous types\n+\n+Most of these mappings are obvious, but there are some nuances and gotchas with\n+Rust FFI (Foreign Function Interface).\n+\n+This document defines clear, one-to-one mappings between primitive types in C,\n+Rust (and possible other languages in the future). Its purpose is to eliminate\n+ambiguity in type widths, signedness, and binary representation across\n+platforms and languages.\n+\n+For Git, the only header required to use these unambiguous types in C is\n+`git-compat-util.h`.\n+\n+== Boolean types\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| bool^1^       | bool\n+|===\n+\n+== Integer types\n+\n+In C, `<stdint.h>` (or an equivalent) must be included.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| uint8_t    | u8\n+| uint16_t   | u16\n+| uint32_t   | u32\n+| uint64_t   | u64\n+\n+| int8_t     | i8\n+| int16_t    | i16\n+| int32_t    | i32\n+| int64_t    | i64\n+|===\n+\n+== Floating-point types\n+\n+Rust requires IEEE-754 semantics.\n+In C, that is typically true, but not guaranteed by the standard.\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| float^2^      | f32\n+| double^2^     | f64\n+|===\n+\n+== Size types\n+\n+These types represent pointer-sized integers and are typically defined in\n+`<stddef.h>` or an equivalent header.\n+\n+Size types should be used any time pointer arithmetic is performed e.g.\n+indexing an array, describing the number of elements in memory, etc...\n+\n+[cols=\"1,1\", options=\"header\"]\n+|===\n+| C Type | Rust Type\n+| size_t^3^     | usize\n+| ptrdiff_t^3^  | isize\n+|===\n+\n+== Character types\n+\n+This is where C and Rust don't have a clean one-to-one mapping.\n+\n+A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n+a `char` will have the same size as the corresponding Rust struct using `u8`.\n+In that sense, such structs are safe to pass over the FFI boundary, because\n+their fields will be laid out identically. However, beyond bit width, C `char`\n+has additional semantics and platform-dependent behavior that can cause\n+problems, as discussed below.\n+\n+The C language leaves the signedness of `char` implementation defined. Because\n+our developer build enables -Wsign-compare, comparison of a value of `char`\n+type with either signed or unsigned integers may trigger warnings from the\n+compiler.\n+\n+Note: Rust's `char` type is an unsigned 32-bit integer that is used to describe\n+Unicode code points.\n+\n+=== Notes\n+^1^ This is only true if stdbool.h (or equivalent) is used. +\n+^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the\n+platform/arch for C does not follow IEEE-754 then this equivalence does not\n+hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but\n+there may be a strange platform/arch where even this isn't true. +\n+^3^ C also defines uintptr_t, ssize_t and intptr_t, but these types are\n+discouraged for FFI purposes. For functions like `read()` and `write()` ssize_t\n+should be cast to a different, and unambiguous, type before being passed over\n+the FFI boundary. +\n+\n+== Problems with std::ffi::c_* types in Rust\n+TL;DR: In practice, Rust's `c_*` types aren't guaranteed to match C types for\n+all possible C compilers, platforms, or architectures, because Rust only\n+ensures correctness of C types on officially supported targets. These\n+definitions have changed over time to match more targets which means that the\n+c_* definitions will differ based on which Rust version Git chooses to use.\n+\n+Current list of safe, Rust side, FFI types in Git: +\n+\n+* `c_void`\n+* `CStr`\n+* `CString`\n+\n+Even then, they should be used sparingly, and only where the semantics match\n+exactly.\n+\n+The std::os::raw::c_* directly inherits the problems of core::ffi, which\n+changes over time and seems to make a best guess at the correct definition for\n+a given platform/target. This probably isn't a problem for all other platforms\n+that Rust supports currently, but can anyone say that Rust got it right for all\n+C compilers of all platforms/targets?\n+\n+To give an example: c_long is defined in\n+footnote:[https://doc.rust-lang.org/1.63.0/src/core/ffi/mod.rs.html#175-189[c_long in 1.63.0]]\n+footnote:[https://doc.rust-lang.org/1.89.0/src/core/ffi/primitives.rs.html#135-151[c_long in 1.89.0]]\n+\n+=== Rust version 1.63.0\n+\n+```\n+mod c_long_definition {\n+    cfg_if! {\n+        if #[cfg(all(target_pointer_width = \"64\", not(windows)))] {\n+            pub type c_long = i64;\n+            pub type NonZero_c_long = crate::num::NonZeroI64;\n+            pub type c_ulong = u64;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU64;\n+        } else {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub type c_long = i32;\n+            pub type NonZero_c_long = crate::num::NonZeroI32;\n+            pub type c_ulong = u32;\n+            pub type NonZero_c_ulong = crate::num::NonZeroU32;\n+        }\n+    }\n+}\n+```\n+\n+=== Rust version 1.89.0\n+\n+```\n+mod c_long_definition {\n+    crate::cfg_select! {\n+        any(\n+            all(target_pointer_width = \"64\", not(windows)),\n+            // wasm32 Linux ABI uses 64-bit long\n+            all(target_arch = \"wasm32\", target_os = \"linux\")\n+        ) => {\n+            pub(super) type c_long = i64;\n+            pub(super) type c_ulong = u64;\n+        }\n+        _ => {\n+            // The minimal size of `long` in the C standard is 32 bits\n+            pub(super) type c_long = i32;\n+            pub(super) type c_ulong = u32;\n+        }\n+    }\n+}\n+```\n+\n+Even for the cases where C types are correctly mapped to Rust types via\n+std::ffi::c_* there are still problems. Let's take c_char for example. On some\n+platforms it's u8 on others it's i8.\n+\n+=== Subtraction underflow in debug mode\n+\n+The following code will panic in debug on platforms that define c_char as u8,\n+but won't if it's an i8.\n+\n+```\n+let mut x: std::ffi::c_char = 0;\n+x -= 1;\n+```\n+\n+=== Inconsistent shift behavior\n+\n+`x` will be 0xC0 for platforms that use i8, but will be 0x40 where it's u8.\n+\n+```\n+let mut x: std::ffi::c_char = 0x80;\n+x >>= 1;\n+```\n+\n+=== Equality fails to compile on some platforms\n+\n+The following will not compile on platforms that define c_char as i8, but will\n+if it's u8. You can cast x e.g. `assert_eq!(x as u8, b'a');`, but then you get\n+a warning on platforms that use u8 and a clean compilation where i8 is used.\n+\n+```\n+let mut x: std::ffi::c_char = 0x61;\n+assert_eq!(x, b'a');\n+```\n+\n+== Enum types\n+Rust enum types should not be used as FFI types. Rust enum types are more like\n+C union types than C enum's. For something like:\n+\n+```\n+#[repr(C, u8)]\n+enum Fruit {\n+    Apple,\n+    Banana,\n+    Cherry,\n+}\n+```\n+\n+It's easy enough to make sure the Rust enum matches what C would expect, but a\n+more complex type like.\n+\n+```\n+enum HashResult {\n+    SHA1([u8; 20]),\n+    SHA256([u8; 32]),\n+}\n+```\n+\n+The Rust compiler has to add a discriminant to the enum to distinguish between\n+the variants. The width, location, and values for that discriminant is up to\n+the Rust compiler and is not ABI stable.\n-- \ngitgitgadget\n\n"},{"id":"530930","messageId":"c4193d11f552547195e04c1318f0c2745d6a59ff.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 02/10] xdiff: use ptrdiff_t for dstart/dend","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:14Z","receivedAt":"2025-11-18T22:34:26Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nptrdiff_t is appropriate for dstart and dend because they both describe\npositive or negative offsets relative to a pointer.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex f145abba3e..7a2d429ec5 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,7 +47,7 @@ typedef struct s_xrecord {\n typedef struct s_xdfile {\n \txrecord_t *recs;\n \tlong nrec;\n-\tlong dstart, dend;\n+\tptrdiff_t dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n \tlong nreff;\n-- \ngitgitgadget\n\n"},{"id":"530931","messageId":"dd76d4f58630fe165998f26f010fb3a4c69fd1fb.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 03/10] xdiff: make xrecord_t.ptr a uint8_t instead of char","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:15Z","receivedAt":"2025-11-18T22:34:28Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake xrecord_t.ptr uint8_t because it's referring to bytes in memory.\n\nIn order to avoid a refactor avalanche, many uses of this field were\ncast to char* or similar.\n\nPlaces where casting was unnecessary:\nxemit.c:156\nxmerge.c:124\nxmerge.c:127\nxmerge.c:164\nxmerge.c:169\nxmerge.c:172\nxmerge.c:178\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     |  6 +++---\n xdiff/xmerge.c    | 14 +++++++-------\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xtypes.h    |  2 +-\n xdiff/xutils.c    |  4 ++--\n 7 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 6f3998ee54..95989b6af1 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -407,7 +407,7 @@ static int get_indent(xrecord_t *rec)\n \tint ret = 0;\n \n \tfor (i = 0; i < rec->size; i++) {\n-\t\tchar c = rec->ptr[i];\n+\t\tchar c = (char) rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n \t\t\treturn ret;\n@@ -993,11 +993,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline(rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\n@@ -1008,7 +1008,7 @@ static int record_matches_regex(xrecord_t *rec, xpparam_t const *xpp) {\n \tsize_t i;\n \n \tfor (i = 0; i < xpp->ignore_regex_nr; i++)\n-\t\tif (!regexec_buf(xpp->ignore_regex[i], rec->ptr, rec->size, 1,\n+\t\tif (!regexec_buf(xpp->ignore_regex[i], (const char *)rec->ptr, rec->size, 1,\n \t\t\t\t &regmatch, 0))\n \t\t\treturn 1;\n \ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex b2f1f30cd3..ead930088a 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec(rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff(rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func(rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex fd600cbb5d..75cb3e76a2 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch(rec1[i].ptr, rec1[i].size,\n-\t\t\trec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch(rec1->ptr, rec1->size,\n-\t\t\t    rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n+\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n }\n \n /*\n@@ -382,10 +382,10 @@ static int xdl_refine_conflicts(xdfenv_t *xe1, xdfenv_t *xe2, xdmerge_t *m,\n \t\t * we have a very simple mmfile structure.\n \t\t */\n \t\tt1.ptr = (char *)xe1->xdf2.recs[m->i1].ptr;\n-\t\tt1.size = xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n+\t\tt1.size = (char *)xe1->xdf2.recs[m->i1 + m->chg1 - 1].ptr\n \t\t\t+ xe1->xdf2.recs[m->i1 + m->chg1 - 1].size - t1.ptr;\n \t\tt2.ptr = (char *)xe2->xdf2.recs[m->i2].ptr;\n-\t\tt2.size = xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n+\t\tt2.size = (char *)xe2->xdf2.recs[m->i2 + m->chg2 - 1].ptr\n \t\t\t+ xe2->xdf2.recs[m->i2 + m->chg2 - 1].size - t2.ptr;\n \t\tif (xdl_do_diff(&t1, &t2, xpp, &xe) < 0)\n \t\t\treturn -1;\n@@ -440,7 +440,7 @@ static int line_contains_alnum(const char *ptr, long size)\n static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n-\t\tif (line_contains_alnum(xe->xdf2.recs[i].ptr,\n+\t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n \t\t\t\txe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 669b653580..bb61354f22 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -121,7 +121,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t\treturn;\n \tmap->entries[index].line1 = line;\n \tmap->entries[index].hash = record->ha;\n-\tmap->entries[index].anchor = is_anchor(xpp, map->env->xdf1.recs[line - 1].ptr);\n+\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n \tif (map->last) {\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 192334f1b7..4c56467076 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch(rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\trec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = prev;\n+\t\t\tcrec->ptr = (uint8_t const *)prev;\n \t\t\tcrec->size = (long) (cur - prev);\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 7a2d429ec5..69727fb299 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -39,7 +39,7 @@ typedef struct s_chastore {\n } chastore_t;\n \n typedef struct s_xrecord {\n-\tchar const *ptr;\n+\tuint8_t const *ptr;\n \tlong size;\n \tunsigned long ha;\n } xrecord_t;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 447e66c719..7be063bfb6 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -465,10 +465,10 @@ int xdl_fall_back_diff(xdfenv_t *diff_env, xpparam_t const *xpp,\n \txdfenv_t env;\n \n \tsubfile1.ptr = (char *)diff_env->xdf1.recs[line1 - 1].ptr;\n-\tsubfile1.size = diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n+\tsubfile1.size = (char *)diff_env->xdf1.recs[line1 + count1 - 2].ptr +\n \t\tdiff_env->xdf1.recs[line1 + count1 - 2].size - subfile1.ptr;\n \tsubfile2.ptr = (char *)diff_env->xdf2.recs[line2 - 1].ptr;\n-\tsubfile2.size = diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n+\tsubfile2.size = (char *)diff_env->xdf2.recs[line2 + count2 - 2].ptr +\n \t\tdiff_env->xdf2.recs[line2 + count2 - 2].size - subfile2.ptr;\n \tif (xdl_do_diff(&subfile1, &subfile2, xpp, &env) < 0)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"530932","messageId":"11cec1d2ec7defbaf130fbe544c806ffb5a485ed.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 04/10] xdiff: use size_t for xrecord_t.size","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:16Z","receivedAt":"2025-11-18T22:34:29Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is the appropriate type because size is describing the number of\nelements, bytes in this case, in memory.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  7 +++----\n xdiff/xemit.c    |  8 ++++----\n xdiff/xmerge.c   | 16 ++++++++--------\n xdiff/xprepare.c |  6 +++---\n xdiff/xtypes.h   |  2 +-\n 5 files changed, 19 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 95989b6af1..cb8e412c7b 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -403,10 +403,9 @@ static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n  */\n static int get_indent(xrecord_t *rec)\n {\n-\tlong i;\n \tint ret = 0;\n \n-\tfor (i = 0; i < rec->size; i++) {\n+\tfor (size_t i = 0; i < rec->size; i++) {\n \t\tchar c = (char) rec->ptr[i];\n \n \t\tif (!XDL_ISSPACE(c))\n@@ -993,11 +992,11 @@ static void xdl_mark_ignorable_lines(xdchange_t *xscr, xdfenv_t *xe, long flags)\n \n \t\trec = &xe->xdf1.recs[xch->i1];\n \t\tfor (i = 0; i < xch->chg1 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\trec = &xe->xdf2.recs[xch->i2];\n \t\tfor (i = 0; i < xch->chg2 && ignore; i++)\n-\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, rec[i].size, flags);\n+\t\t\tignore = xdl_blankline((const char *)rec[i].ptr, (long)rec[i].size, flags);\n \n \t\txch->ignore = ignore;\n \t}\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex ead930088a..2f8007753c 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -27,7 +27,7 @@ static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *\n {\n \txrecord_t *rec = &xdf->recs[ri];\n \n-\tif (xdl_emit_diffrec((char const *)rec->ptr, rec->size, pre, strlen(pre), ecb) < 0)\n+\tif (xdl_emit_diffrec((char const *)rec->ptr, (long)rec->size, pre, strlen(pre), ecb) < 0)\n \t\treturn -1;\n \n \treturn 0;\n@@ -113,8 +113,8 @@ static long match_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri,\n \txrecord_t *rec = &xdf->recs[ri];\n \n \tif (!xecfg->find_func)\n-\t\treturn def_ff((const char *)rec->ptr, rec->size, buf, sz);\n-\treturn xecfg->find_func((const char *)rec->ptr, rec->size, buf, sz, xecfg->find_func_priv);\n+\t\treturn def_ff((const char *)rec->ptr, (long)rec->size, buf, sz);\n+\treturn xecfg->find_func((const char *)rec->ptr, (long)rec->size, buf, sz, xecfg->find_func_priv);\n }\n \n static int is_func_rec(xdfile_t *xdf, xdemitconf_t const *xecfg, long ri)\n@@ -151,7 +151,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n static int is_empty_rec(xdfile_t *xdf, long ri)\n {\n \txrecord_t *rec = &xdf->recs[ri];\n-\tlong i = 0;\n+\tsize_t i = 0;\n \n \tfor (; i < rec->size && XDL_ISSPACE(rec->ptr[i]); i++);\n \ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 75cb3e76a2..0dd4558a32 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -101,8 +101,8 @@ static int xdl_merge_cmp_lines(xdfenv_t *xe1, int i1, xdfenv_t *xe2, int i2,\n \txrecord_t *rec2 = xe2->xdf2.recs + i2;\n \n \tfor (i = 0; i < line_count; i++) {\n-\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, rec1[i].size,\n-\t\t\t(const char *)rec2[i].ptr, rec2[i].size, flags);\n+\t\tint result = xdl_recmatch((const char *)rec1[i].ptr, (long)rec1[i].size,\n+\t\t\t(const char *)rec2[i].ptr, (long)rec2[i].size, flags);\n \t\tif (!result)\n \t\t\treturn -1;\n \t}\n@@ -119,11 +119,11 @@ static int xdl_recs_copy_0(int use_orig, xdfenv_t *xe, int i, int count, int nee\n \tif (count < 1)\n \t\treturn 0;\n \n-\tfor (i = 0; i < count; size += recs[i++].size)\n+\tfor (i = 0; i < count; size += (int)recs[i++].size)\n \t\tif (dest)\n \t\t\tmemcpy(dest + size, recs[i].ptr, recs[i].size);\n \tif (add_nl) {\n-\t\ti = recs[count - 1].size;\n+\t\ti = (int)recs[count - 1].size;\n \t\tif (i == 0 || recs[count - 1].ptr[i - 1] != '\\n') {\n \t\t\tif (needs_cr) {\n \t\t\t\tif (dest)\n@@ -156,7 +156,7 @@ static int xdl_orig_copy(xdfenv_t *xe, int i, int count, int needs_cr, int add_n\n  */\n static int is_eol_crlf(xdfile_t *file, int i)\n {\n-\tlong size;\n+\tsize_t size;\n \n \tif (i < file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n@@ -324,8 +324,8 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \n static int recmatch(xrecord_t *rec1, xrecord_t *rec2, unsigned long flags)\n {\n-\treturn xdl_recmatch((const char *)rec1->ptr, rec1->size,\n-\t\t\t    (const char *)rec2->ptr, rec2->size, flags);\n+\treturn xdl_recmatch((const char *)rec1->ptr, (long)rec1->size,\n+\t\t\t    (const char *)rec2->ptr, (long)rec2->size, flags);\n }\n \n /*\n@@ -441,7 +441,7 @@ static int lines_contain_alnum(xdfenv_t *xe, int i, int chg)\n {\n \tfor (; chg; chg--, i++)\n \t\tif (line_contains_alnum((const char *)xe->xdf2.recs[i].ptr,\n-\t\t\t\txe->xdf2.recs[i].size))\n+\t\t\t\t(long)xe->xdf2.recs[i].size))\n \t\t\treturn 1;\n \treturn 0;\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 4c56467076..b3219aed3e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -99,8 +99,8 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n \t\tif (rcrec->rec.ha == rec->ha &&\n-\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, rcrec->rec.size,\n-\t\t\t\t\t(const char *)rec->ptr, rec->size, cf->flags))\n+\t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n+\t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n \n \tif (!rcrec) {\n@@ -157,7 +157,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = (uint8_t const *)prev;\n-\t\t\tcrec->size = (long) (cur - prev);\n+\t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 69727fb299..354349b523 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -40,7 +40,7 @@ typedef struct s_chastore {\n \n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n-\tlong size;\n+\tsize_t size;\n \tunsigned long ha;\n } xrecord_t;\n \n-- \ngitgitgadget\n\n"},{"id":"530933","messageId":"6f267360b705e6d5ee62a67c22b1de2d3ff38196.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 05/10] xdiff: use unambiguous types in xdl_hash_record()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:17Z","receivedAt":"2025-11-18T22:34:31Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nConvert the function signature and body to use unambiguous types. char\nis changed to uint8_t because this function processes bytes in memory.\nunsigned long to uint64_t so that the hash output is consistent across\nplatforms. `flags` was changed from long to uint64_t to ensure the\nhigh order bits are not dropped on platforms that treat long as 32\nbits.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff-interface.c |  2 +-\n xdiff/xprepare.c  |  6 +++---\n xdiff/xutils.c    | 28 ++++++++++++++--------------\n xdiff/xutils.h    |  6 +++---\n 4 files changed, 21 insertions(+), 21 deletions(-)\n\ndiff --git a/xdiff-interface.c b/xdiff-interface.c\nindex 4971f722b3..1a35556380 100644\n--- a/xdiff-interface.c\n+++ b/xdiff-interface.c\n@@ -300,7 +300,7 @@ void xdiff_clear_find_func(xdemitconf_t *xecfg)\n \n unsigned long xdiff_hash_string(const char *s, size_t len, long flags)\n {\n-\treturn xdl_hash_record(&s, s + len, flags);\n+\treturn xdl_hash_record((uint8_t const**)&s, (uint8_t const*)s + len, flags);\n }\n \n int xdiff_compare_lines(const char *l1, long s1,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex b3219aed3e..85e56021da 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -137,8 +137,8 @@ static void xdl_free_ctx(xdfile_t *xdf)\n static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_t const *xpp,\n \t\t\t   xdlclassifier_t *cf, xdfile_t *xdf) {\n \tlong bsize;\n-\tunsigned long hav;\n-\tchar const *blk, *cur, *top, *prev;\n+\tuint64_t hav;\n+\tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n \txdf->rindex = NULL;\n@@ -156,7 +156,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n-\t\t\tcrec->ptr = (uint8_t const *)prev;\n+\t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n \t\t\tcrec->ha = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 7be063bfb6..77ee1ad9c8 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -249,11 +249,11 @@ int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags)\n \treturn 1;\n }\n \n-unsigned long xdl_hash_record_with_whitespace(char const **data,\n-\t\tchar const *top, long flags) {\n-\tunsigned long ha = 5381;\n-\tchar const *ptr = *data;\n-\tint cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data,\n+\t\tuint8_t const *top, uint64_t flags) {\n+\tuint64_t ha = 5381;\n+\tuint8_t const *ptr = *data;\n+\tbool cr_at_eol_only = (flags & XDF_WHITESPACE_FLAGS) == XDF_IGNORE_CR_AT_EOL;\n \n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tif (cr_at_eol_only) {\n@@ -263,8 +263,8 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\t\tcontinue;\n \t\t}\n \t\telse if (XDL_ISSPACE(*ptr)) {\n-\t\t\tconst char *ptr2 = ptr;\n-\t\t\tint at_eol;\n+\t\t\tconst uint8_t *ptr2 = ptr;\n+\t\t\tbool at_eol;\n \t\t\twhile (ptr + 1 < top && XDL_ISSPACE(ptr[1])\n \t\t\t\t\t&& ptr[1] != '\\n')\n \t\t\t\tptr++;\n@@ -274,20 +274,20 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_CHANGE\n \t\t\t\t && !at_eol) {\n \t\t\t\tha += (ha << 5);\n-\t\t\t\tha ^= (unsigned long) ' ';\n+\t\t\t\tha ^= (uint64_t) ' ';\n \t\t\t}\n \t\t\telse if (flags & XDF_IGNORE_WHITESPACE_AT_EOL\n \t\t\t\t && !at_eol) {\n \t\t\t\twhile (ptr2 != ptr + 1) {\n \t\t\t\t\tha += (ha << 5);\n-\t\t\t\t\tha ^= (unsigned long) *ptr2;\n+\t\t\t\t\tha ^= (uint64_t) *ptr2;\n \t\t\t\t\tptr2++;\n \t\t\t\t}\n \t\t\t}\n \t\t\tcontinue;\n \t\t}\n \t\tha += (ha << 5);\n-\t\tha ^= (unsigned long) *ptr;\n+\t\tha ^= (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n \n@@ -304,9 +304,9 @@ unsigned long xdl_hash_record_with_whitespace(char const **data,\n #define REASSOC_FENCE(x, y)\n #endif\n \n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n-\tunsigned long ha = 5381, c0, c1;\n-\tchar const *ptr = *data;\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top) {\n+\tuint64_t ha = 5381, c0, c1;\n+\tuint8_t const *ptr = *data;\n #if 0\n \t/*\n \t * The baseline form of the optimized loop below. This is the djb2\n@@ -314,7 +314,7 @@ unsigned long xdl_hash_record_verbatim(char const **data, char const *top) {\n \t */\n \tfor (; ptr < top && *ptr != '\\n'; ptr++) {\n \t\tha += (ha << 5);\n-\t\tha += (unsigned long) *ptr;\n+\t\tha += (uint64_t) *ptr;\n \t}\n \t*data = ptr < top ? ptr + 1: ptr;\n #else\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 13f6831047..615b4a9d35 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -34,9 +34,9 @@ void *xdl_cha_alloc(chastore_t *cha);\n long xdl_guess_lines(mmfile_t *mf, long sample);\n int xdl_blankline(const char *line, long size, long flags);\n int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags);\n-unsigned long xdl_hash_record_verbatim(char const **data, char const *top);\n-unsigned long xdl_hash_record_with_whitespace(char const **data, char const *top, long flags);\n-static inline unsigned long xdl_hash_record(char const **data, char const *top, long flags)\n+uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top);\n+uint64_t xdl_hash_record_with_whitespace(uint8_t const **data, uint8_t const *top, uint64_t flags);\n+static inline uint64_t xdl_hash_record(uint8_t const **data, uint8_t const *top, uint64_t flags)\n {\n \tif (flags & XDF_WHITESPACE_FLAGS)\n \t\treturn xdl_hash_record_with_whitespace(data, top, flags);\n-- \ngitgitgadget\n\n"},{"id":"530934","messageId":"78af0f16f4a1c3af0b8474169ceb75bda2b94b49.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:18Z","receivedAt":"2025-11-18T22:34:33Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe ha field is serving two different purposes, which makes the code\nharder to read. At first glance, it looks like many places assume\nthere could never be hash collisions between lines of the two input\nfiles. In reality, line_hash is used together with xdl_recmatch() to\nensure correct comparisons of lines, even when collisions occur.\n\nTo make this clearer, the old ha field has been split:\n  * line_hash: a straightforward hash of a line, independent of any\n    external context. Its type is uint64_t, as it comes from a fixed\n    width hash function.\n  * minimal_perfect_hash: Not a new concept, but now a separate\n    field. It comes from the classifier's general-purpose hash table,\n    which assigns each line a unique and minimal hash across the two\n    files. A size_t is used here because it's meant to be used to\n    index an array. This also avoids ` as usize` casts on the Rust\n    side when using it to index a slice.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c     |  6 +++---\n xdiff/xhistogram.c |  4 ++--\n xdiff/xpatience.c  | 10 +++++-----\n xdiff/xprepare.c   | 18 +++++++++---------\n xdiff/xtypes.h     |  3 ++-\n 5 files changed, 21 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex cb8e412c7b..8d96074414 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -22,9 +22,9 @@\n \n #include \"xinclude.h\"\n \n-static unsigned long get_hash(xdfile_t *xdf, long index)\n+static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].ha;\n+\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -385,7 +385,7 @@ static xdchange_t *xdl_add_change(xdchange_t *xscr, long i1, long i2, long chg1,\n \n static int recs_match(xrecord_t *rec1, xrecord_t *rec2)\n {\n-\treturn (rec1->ha == rec2->ha);\n+\treturn rec1->minimal_perfect_hash == rec2->minimal_perfect_hash;\n }\n \n /*\ndiff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\nindex 6dc450b1fe..5ae1282c27 100644\n--- a/xdiff/xhistogram.c\n+++ b/xdiff/xhistogram.c\n@@ -90,7 +90,7 @@ struct region {\n \n static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n {\n-\treturn r1->ha == r2->ha;\n+\treturn r1->minimal_perfect_hash == r2->minimal_perfect_hash;\n \n }\n \n@@ -98,7 +98,7 @@ static int cmp_recs(xrecord_t *r1, xrecord_t *r2)\n \t(cmp_recs(REC(i->env, s1, l1), REC(i->env, s2, l2)))\n \n #define TABLE_HASH(index, side, line) \\\n-\tXDL_HASHLONG((REC(index->env, side, line))->ha, index->table_bits)\n+\tXDL_HASHLONG((REC(index->env, side, line))->minimal_perfect_hash, index->table_bits)\n \n static int scanA(struct histindex *index, int line1, int count1)\n {\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex bb61354f22..cc53266f3b 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -48,7 +48,7 @@\n struct hashmap {\n \tint nr, alloc;\n \tstruct entry {\n-\t\tunsigned long hash;\n+\t\tsize_t minimal_perfect_hash;\n \t\t/*\n \t\t * 0 = unused entry, 1 = first line, 2 = second, etc.\n \t\t * line2 is NON_UNIQUE if the line is not unique\n@@ -101,10 +101,10 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t * So we multiply ha by 2 in the hope that the hashing was\n \t * \"unique enough\".\n \t */\n-\tint index = (int)((record->ha << 1) % map->alloc);\n+\tint index = (int)((record->minimal_perfect_hash << 1) % map->alloc);\n \n \twhile (map->entries[index].line1) {\n-\t\tif (map->entries[index].hash != record->ha) {\n+\t\tif (map->entries[index].minimal_perfect_hash != record->minimal_perfect_hash) {\n \t\t\tif (++index >= map->alloc)\n \t\t\t\tindex = 0;\n \t\t\tcontinue;\n@@ -120,7 +120,7 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \tif (pass == 2)\n \t\treturn;\n \tmap->entries[index].line1 = line;\n-\tmap->entries[index].hash = record->ha;\n+\tmap->entries[index].minimal_perfect_hash = record->minimal_perfect_hash;\n \tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n@@ -248,7 +248,7 @@ static int match(struct hashmap *map, int line1, int line2)\n {\n \txrecord_t *record1 = &map->env->xdf1.recs[line1 - 1];\n \txrecord_t *record2 = &map->env->xdf2.recs[line2 - 1];\n-\treturn record1->ha == record2->ha;\n+\treturn record1->minimal_perfect_hash == record2->minimal_perfect_hash;\n }\n \n static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 85e56021da..bea0992b5e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -93,12 +93,12 @@ static void xdl_free_classifier(xdlclassifier_t *cf) {\n \n \n static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n-\tlong hi;\n+\tsize_t hi;\n \txdlclass_t *rcrec;\n \n-\thi = (long) XDL_HASHLONG(rec->ha, cf->hbits);\n+\thi = XDL_HASHLONG(rec->line_hash, cf->hbits);\n \tfor (rcrec = cf->rchash[hi]; rcrec; rcrec = rcrec->next)\n-\t\tif (rcrec->rec.ha == rec->ha &&\n+\t\tif (rcrec->rec.line_hash == rec->line_hash &&\n \t\t\t\txdl_recmatch((const char *)rcrec->rec.ptr, (long)rcrec->rec.size,\n \t\t\t\t\t(const char *)rec->ptr, (long)rec->size, cf->flags))\n \t\t\tbreak;\n@@ -120,7 +120,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n \t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n \n-\trec->ha = (unsigned long) rcrec->idx;\n+\trec->minimal_perfect_hash = (size_t)rcrec->idx;\n \n \treturn 0;\n }\n@@ -158,7 +158,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n-\t\t\tcrec->ha = hav;\n+\t\t\tcrec->line_hash = hav;\n \t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n \t\t\t\tgoto abort;\n \t\t}\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -298,7 +298,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->ha];\n+\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -350,7 +350,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs2 = xdf2->recs;\n \tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dstart = xdf2->dstart = i;\n@@ -358,7 +358,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \trecs1 = xdf1->recs + xdf1->nrec - 1;\n \trecs2 = xdf2->recs + xdf2->nrec - 1;\n \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n-\t\tif (recs1->ha != recs2->ha)\n+\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n \txdf1->dend = xdf1->nrec - i - 1;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 354349b523..d4e9cd2e76 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -41,7 +41,8 @@ typedef struct s_chastore {\n typedef struct s_xrecord {\n \tuint8_t const *ptr;\n \tsize_t size;\n-\tunsigned long ha;\n+\tuint64_t line_hash;\n+\tsize_t minimal_perfect_hash;\n } xrecord_t;\n \n typedef struct s_xdfile {\n-- \ngitgitgadget\n\n"},{"id":"530935","messageId":"5c19f9ded397272da5a42e00b7ccd4f9cd484ee8.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 07/10] xdiff: make xdfile_t.nrec a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:19Z","receivedAt":"2025-11-18T22:34:34Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nrec describes the number of elements for both\nrecs, and for 'changed' + 2.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c    |  8 ++++----\n xdiff/xemit.c     | 20 ++++++++++----------\n xdiff/xmerge.c    |  8 ++++----\n xdiff/xpatience.c |  2 +-\n xdiff/xprepare.c  | 12 ++++++------\n xdiff/xtypes.h    |  2 +-\n 6 files changed, 26 insertions(+), 26 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 8d96074414..21d06bce96 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -483,7 +483,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n {\n \tlong i;\n \n-\tif (split >= xdf->nrec) {\n+\tif (split >= (long)xdf->nrec) {\n \t\tm->end_of_file = 1;\n \t\tm->indent = -1;\n \t} else {\n@@ -506,7 +506,7 @@ static void measure_split(const xdfile_t *xdf, long split,\n \n \tm->post_blank = 0;\n \tm->post_indent = -1;\n-\tfor (i = split + 1; i < xdf->nrec; i++) {\n+\tfor (i = split + 1; i < (long)xdf->nrec; i++) {\n \t\tm->post_indent = get_indent(&xdf->recs[i]);\n \t\tif (m->post_indent != -1)\n \t\t\tbreak;\n@@ -717,7 +717,7 @@ static void group_init(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static inline int group_next(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end == xdf->nrec)\n+\tif (g->end == (long)xdf->nrec)\n \t\treturn -1;\n \n \tg->start = g->end + 1;\n@@ -750,7 +750,7 @@ static inline int group_previous(xdfile_t *xdf, struct xdlgroup *g)\n  */\n static int group_slide_down(xdfile_t *xdf, struct xdlgroup *g)\n {\n-\tif (g->end < xdf->nrec &&\n+\tif (g->end < (long)xdf->nrec &&\n \t    recs_match(&xdf->recs[g->start], &xdf->recs[g->end])) {\n \t\txdf->changed[g->start++] = false;\n \t\txdf->changed[g->end++] = true;\ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex 2f8007753c..04f7e9193b 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -137,7 +137,7 @@ static long get_func_line(xdfenv_t *xe, xdemitconf_t const *xecfg,\n \tbuf = func_line ? func_line->buf : dummy;\n \tsize = func_line ? sizeof(func_line->buf) : sizeof(dummy);\n \n-\tfor (l = start; l != limit && 0 <= l && l < xe->xdf1.nrec; l += step) {\n+\tfor (l = start; l != limit && 0 <= l && l < (long)xe->xdf1.nrec; l += step) {\n \t\tlong len = match_func_rec(&xe->xdf1, xecfg, l, buf, size);\n \t\tif (len >= 0) {\n \t\t\tif (func_line)\n@@ -179,14 +179,14 @@ pre_context_calculation:\n \t\t\tlong fs1, i1 = xch->i1;\n \n \t\t\t/* Appended chunk? */\n-\t\t\tif (i1 >= xe->xdf1.nrec) {\n+\t\t\tif (i1 >= (long)xe->xdf1.nrec) {\n \t\t\t\tlong i2 = xch->i2;\n \n \t\t\t\t/*\n \t\t\t\t * We don't need additional context if\n \t\t\t\t * a whole function was added.\n \t\t\t\t */\n-\t\t\t\twhile (i2 < xe->xdf2.nrec) {\n+\t\t\t\twhile (i2 < (long)xe->xdf2.nrec) {\n \t\t\t\t\tif (is_func_rec(&xe->xdf2, xecfg, i2))\n \t\t\t\t\t\tgoto post_context_calculation;\n \t\t\t\t\ti2++;\n@@ -196,7 +196,7 @@ pre_context_calculation:\n \t\t\t\t * Otherwise get more context from the\n \t\t\t\t * pre-image.\n \t\t\t\t */\n-\t\t\t\ti1 = xe->xdf1.nrec - 1;\n+\t\t\t\ti1 = (long)xe->xdf1.nrec - 1;\n \t\t\t}\n \n \t\t\tfs1 = get_func_line(xe, xecfg, NULL, i1, -1);\n@@ -228,8 +228,8 @@ pre_context_calculation:\n \n  post_context_calculation:\n \t\tlctx = xecfg->ctxlen;\n-\t\tlctx = XDL_MIN(lctx, xe->xdf1.nrec - (xche->i1 + xche->chg1));\n-\t\tlctx = XDL_MIN(lctx, xe->xdf2.nrec - (xche->i2 + xche->chg2));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf1.nrec - (xche->i1 + xche->chg1));\n+\t\tlctx = XDL_MIN(lctx, (long)xe->xdf2.nrec - (xche->i2 + xche->chg2));\n \n \t\te1 = xche->i1 + xche->chg1 + lctx;\n \t\te2 = xche->i2 + xche->chg2 + lctx;\n@@ -237,13 +237,13 @@ pre_context_calculation:\n \t\tif (xecfg->flags & XDL_EMIT_FUNCCONTEXT) {\n \t\t\tlong fe1 = get_func_line(xe, xecfg, NULL,\n \t\t\t\t\t\t xche->i1 + xche->chg1,\n-\t\t\t\t\t\t xe->xdf1.nrec);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec);\n \t\t\twhile (fe1 > 0 && is_empty_rec(&xe->xdf1, fe1 - 1))\n \t\t\t\tfe1--;\n \t\t\tif (fe1 < 0)\n-\t\t\t\tfe1 = xe->xdf1.nrec;\n+\t\t\t\tfe1 = (long)xe->xdf1.nrec;\n \t\t\tif (fe1 > e1) {\n-\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), xe->xdf2.nrec);\n+\t\t\t\te2 = XDL_MIN(e2 + (fe1 - e1), (long)xe->xdf2.nrec);\n \t\t\t\te1 = fe1;\n \t\t\t}\n \n@@ -254,7 +254,7 @@ pre_context_calculation:\n \t\t\t */\n \t\t\tif (xche->next) {\n \t\t\t\tlong l = XDL_MIN(xche->next->i1,\n-\t\t\t\t\t\t xe->xdf1.nrec - 1);\n+\t\t\t\t\t\t (long)xe->xdf1.nrec - 1);\n \t\t\t\tif (l - xecfg->ctxlen <= e1 ||\n \t\t\t\t    get_func_line(xe, xecfg, NULL, l, e1) < 0) {\n \t\t\t\t\txche = xche->next;\ndiff --git a/xdiff/xmerge.c b/xdiff/xmerge.c\nindex 0dd4558a32..29dad98c49 100644\n--- a/xdiff/xmerge.c\n+++ b/xdiff/xmerge.c\n@@ -158,7 +158,7 @@ static int is_eol_crlf(xdfile_t *file, int i)\n {\n \tsize_t size;\n \n-\tif (i < file->nrec - 1)\n+\tif (i < (long)file->nrec - 1)\n \t\t/* All lines before the last *must* end in LF */\n \t\treturn (size = file->recs[i].size) > 1 &&\n \t\t\tfile->recs[i].ptr[size - 2] == '\\r';\n@@ -317,7 +317,7 @@ static int xdl_fill_merge_buffer(xdfenv_t *xe1, const char *name1,\n \t\t\tcontinue;\n \t\ti = m->i1 + m->chg1;\n \t}\n-\tsize += xdl_recs_copy(xe1, i, xe1->xdf2.nrec - i, 0, 0,\n+\tsize += xdl_recs_copy(xe1, i, (int)xe1->xdf2.nrec - i, 0, 0,\n \t\t\t      dest ? dest + size : NULL);\n \treturn size;\n }\n@@ -622,7 +622,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\t\tchanges = c;\n \t\ti0 = xscr1->i1;\n \t\ti1 = xscr1->i2;\n-\t\ti2 = xscr1->i1 + xe2->xdf2.nrec - xe2->xdf1.nrec;\n+\t\ti2 = xscr1->i1 + (long)xe2->xdf2.nrec - (long)xe2->xdf1.nrec;\n \t\tchg0 = xscr1->chg1;\n \t\tchg1 = xscr1->chg2;\n \t\tchg2 = xscr1->chg1;\n@@ -637,7 +637,7 @@ static int xdl_do_merge(xdfenv_t *xe1, xdchange_t *xscr1,\n \t\tif (!changes)\n \t\t\tchanges = c;\n \t\ti0 = xscr2->i1;\n-\t\ti1 = xscr2->i1 + xe1->xdf2.nrec - xe1->xdf1.nrec;\n+\t\ti1 = xscr2->i1 + (long)xe1->xdf2.nrec - (long)xe1->xdf1.nrec;\n \t\ti2 = xscr2->i2;\n \t\tchg0 = xscr2->chg1;\n \t\tchg1 = xscr2->chg1;\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex cc53266f3b..a0b31eb5d8 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -370,5 +370,5 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n \n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n-\treturn patience_diff(xpp, env, 1, env->xdf1.nrec, 1, env->xdf2.nrec);\n+\treturn patience_diff(xpp, env, 1, (int)env->xdf1.nrec, 1, (int)env->xdf2.nrec);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex bea0992b5e..705ddd1ae0 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -153,7 +153,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \t\tfor (top = blk + bsize; cur < top; ) {\n \t\t\tprev = cur;\n \t\t\thav = xdl_hash_record(&cur, top, xpp->flags);\n-\t\t\tif (XDL_ALLOC_GROW(xdf->recs, xdf->nrec + 1, narec))\n+\t\t\tif (XDL_ALLOC_GROW(xdf->recs, (long)xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n@@ -287,7 +287,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -295,7 +295,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n@@ -348,7 +348,7 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \n \trecs1 = xdf1->recs;\n \trecs2 = xdf2->recs;\n-\tfor (i = 0, lim = XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n+\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n \t     i++, recs1++, recs2++)\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n@@ -361,8 +361,8 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n-\txdf1->dend = xdf1->nrec - i - 1;\n-\txdf2->dend = xdf2->nrec - i - 1;\n+\txdf1->dend = (long)xdf1->nrec - i - 1;\n+\txdf2->dend = (long)xdf2->nrec - i - 1;\n \n \treturn 0;\n }\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex d4e9cd2e76..4c4d9bd147 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -47,7 +47,7 @@ typedef struct s_xrecord {\n \n typedef struct s_xdfile {\n \txrecord_t *recs;\n-\tlong nrec;\n+\tsize_t nrec;\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n-- \ngitgitgadget\n\n"},{"id":"530936","messageId":"d1f498edb113e925326039a3e6a7fd31a24ba43d.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 08/10] xdiff: make xdfile_t.nreff a size_t instead of long","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:20Z","receivedAt":"2025-11-18T22:34:35Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nsize_t is used because nreff describes the number of elements in memory\nfor rindex.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 +++++++-------\n xdiff/xtypes.h   |  2 +-\n 2 files changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 705ddd1ae0..39fd79d9d4 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -264,7 +264,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, nreff, mlim;\n+\tlong i, nm, mlim;\n \txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n@@ -307,29 +307,29 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * Use temporary arrays to decide if changed[i] should remain\n \t * false, or become true.\n \t */\n-\tfor (nreff = 0, i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n+\txdf1->nreff = 0;\n+\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[nreff++] = i;\n+\t\t\txdf1->rindex[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf1->nreff = nreff;\n \n-\tfor (nreff = 0, i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n+\txdf2->nreff = 0;\n+\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[nreff++] = i;\n+\t\t\txdf2->rindex[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n-\txdf2->nreff = nreff;\n \n cleanup:\n \txdl_free(action1);\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 4c4d9bd147..1f495f987f 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -51,7 +51,7 @@ typedef struct s_xdfile {\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n \tlong *rindex;\n-\tlong nreff;\n+\tsize_t nreff;\n } xdfile_t;\n \n typedef struct s_xdfenv {\n-- \ngitgitgadget\n\n"},{"id":"530937","messageId":"bc4941c14668984e882d43baaeddd08312efc217.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 09/10] xdiff: change rindex from long to size_t in xdfile_t","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:21Z","receivedAt":"2025-11-18T22:34:36Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe field rindex describes an index offset for other arrays. Change it\nto size_t.\n\nChanging the type of rindex from long to size_t has no cascading\nrefactor impact because it is only ever used to directly index other\narrays.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xtypes.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 1f495f987f..9074cdadd1 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -50,7 +50,7 @@ typedef struct s_xdfile {\n \tsize_t nrec;\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n-\tlong *rindex;\n+\tsize_t *rindex;\n \tsize_t nreff;\n } xdfile_t;\n \n-- \ngitgitgadget\n\n"},{"id":"530938","messageId":"dcc9d6bfafb69993b83b13e824fc52a9f4aaa256.1763505262.git.gitgitgadget@gmail.com","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"[PATCH v5 10/10] xdiff: rename rindex -> reference_index","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-18T22:34:22Z","receivedAt":"2025-11-18T22:34:37Z","isPatch":true,"sender":{"key":"ezekielnewren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/18313324?v=4"},"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe classic diff adds only the lines that it's going to consider,\nduring the diff, to an array. A mapping between the compacted\narray, and the lines of the file that they reference, is\nfacilitated by this array.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   |  6 +++---\n xdiff/xprepare.c | 10 +++++-----\n xdiff/xtypes.h   |  2 +-\n 3 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 21d06bce96..4376f943db 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -24,7 +24,7 @@\n \n static size_t get_hash(xdfile_t *xdf, long index)\n {\n-\treturn xdf->recs[xdf->rindex[index]].minimal_perfect_hash;\n+\treturn xdf->recs[xdf->reference_index[index]].minimal_perfect_hash;\n }\n \n #define XDL_MAX_COST_MIN 256\n@@ -278,10 +278,10 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n \t */\n \tif (off1 == lim1) {\n \t\tfor (; off2 < lim2; off2++)\n-\t\t\txdf2->changed[xdf2->rindex[off2]] = true;\n+\t\t\txdf2->changed[xdf2->reference_index[off2]] = true;\n \t} else if (off2 == lim2) {\n \t\tfor (; off1 < lim1; off1++)\n-\t\t\txdf1->changed[xdf1->rindex[off1]] = true;\n+\t\t\txdf1->changed[xdf1->reference_index[off1]] = true;\n \t} else {\n \t\txdpsplit_t spl;\n \t\tspl.i1 = spl.i2 = 0;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 39fd79d9d4..34c82e4f8e 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -128,7 +128,7 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \n static void xdl_free_ctx(xdfile_t *xdf)\n {\n-\txdl_free(xdf->rindex);\n+\txdl_free(xdf->reference_index);\n \txdl_free(xdf->changed - 1);\n \txdl_free(xdf->recs);\n }\n@@ -141,7 +141,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n \n-\txdf->rindex = NULL;\n+\txdf->reference_index = NULL;\n \txdf->changed = NULL;\n \txdf->recs = NULL;\n \n@@ -169,7 +169,7 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF)) {\n-\t\tif (!XDL_ALLOC_ARRAY(xdf->rindex, xdf->nrec + 1))\n+\t\tif (!XDL_ALLOC_ARRAY(xdf->reference_index, xdf->nrec + 1))\n \t\t\tgoto abort;\n \t}\n \n@@ -312,7 +312,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf1->dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->rindex[xdf1->nreff++] = i;\n+\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf1->changed[i] = true;\n@@ -324,7 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t     i <= xdf2->dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->rindex[xdf2->nreff++] = i;\n+\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n \t\t\txdf2->changed[i] = true;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 9074cdadd1..979586f20a 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -50,7 +50,7 @@ typedef struct s_xdfile {\n \tsize_t nrec;\n \tptrdiff_t dstart, dend;\n \tbool *changed;\n-\tsize_t *rindex;\n+\tsize_t *reference_index;\n \tsize_t nreff;\n } xdfile_t;\n \n-- \ngitgitgadget\n"},{"id":"530941","messageId":"xmqqms4jc4wj.fsf@gitster.g","threadId":"64326","inReplyTo":"pull.2070.v5.git.git.1763505262.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 00/10] Xdiff cleanup part2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-18T23:11:40Z","receivedAt":"2025-11-18T23:11:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Changes in v5:\n>\n>  * Remove the non-word 'signless', and rephrase that paragraph in\n>    unambiguous-types.adoc\n>  * Cast to char in xdiffi.c:get_indent() rather than changing the local\n>    variable to uint8_t\n> ...\n>\n> Ezekiel Newren (10):\n>   doc: define unambiguous type mappings across C and Rust\n>   xdiff: use ptrdiff_t for dstart/dend\n>   xdiff: make xrecord_t.ptr a uint8_t instead of char\n>   xdiff: use size_t for xrecord_t.size\n>   xdiff: use unambiguous types in xdl_hash_record()\n>   xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash\n>   xdiff: make xdfile_t.nrec a size_t instead of long\n>   xdiff: make xdfile_t.nreff a size_t instead of long\n>   xdiff: change rindex from long to size_t in xdfile_t\n>   xdiff: rename rindex -> reference_index\n>\n>  Documentation/Makefile                        |   1 +\n>  Documentation/technical/meson.build           |   1 +\n>  .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n>  xdiff-interface.c                             |   2 +-\n>  xdiff/xdiffi.c                                |  29 ++-\n>  xdiff/xemit.c                                 |  28 +--\n>  xdiff/xhistogram.c                            |   4 +-\n>  xdiff/xmerge.c                                |  30 +--\n>  xdiff/xpatience.c                             |  14 +-\n>  xdiff/xprepare.c                              |  60 ++---\n>  xdiff/xtypes.h                                |  15 +-\n>  xdiff/xutils.c                                |  32 +--\n>  xdiff/xutils.h                                |   6 +-\n>  13 files changed, 336 insertions(+), 110 deletions(-)\n>  create mode 100644 Documentation/technical/unambiguous-types.adoc\n\nThis round looks good to me.  Shall we mark it for 'next'?\n\nThanks.\n"},{"id":"530942","messageId":"9c7a7d09-2cc0-40f7-b37a-befef5339d76@ramsayjones.plus.com","threadId":"64326","inReplyTo":"8b56bf117289ca3be25533a36da1ea0c178ccfca.1763505262.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsayjones.plus.com","sentAt":"2025-11-18T23:46:47Z","receivedAt":"2025-11-18T23:49:58Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"\n\nOn 18/11/2025 10:34 pm, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Document other nuances when crossing the FFI boundary. Other language\n> mappings may be added in the future.\n> \n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  Documentation/Makefile                        |   1 +\n>  Documentation/technical/meson.build           |   1 +\n>  .../technical/unambiguous-types.adoc          | 224 ++++++++++++++++++\n>  3 files changed, 226 insertions(+)\n>  create mode 100644 Documentation/technical/unambiguous-types.adoc\n> \n[snip]\n\n> diff --git a/Documentation/technical/unambiguous-types.adoc b/Documentation/technical/unambiguous-types.adoc\n> new file mode 100644\n> index 0000000000..9a4990847c\n> --- /dev/null\n> +++ b/Documentation/technical/unambiguous-types.adoc\n> @@ -0,0 +1,224 @@\n> += Unambiguous types\n> +\n> +Most of these mappings are obvious, but there are some nuances and gotchas with\n> +Rust FFI (Foreign Function Interface).\n> +\n> +This document defines clear, one-to-one mappings between primitive types in C,\n> +Rust (and possible other languages in the future). Its purpose is to eliminate\n> +ambiguity in type widths, signedness, and binary representation across\n> +platforms and languages.\n> +\n> +For Git, the only header required to use these unambiguous types in C is\n> +`git-compat-util.h`.\n> +\n> +== Boolean types\n> +[cols=\"1,1\", options=\"header\"]\n> +|===\n> +| C Type | Rust Type\n> +| bool^1^       | bool\n> +|===\n> +\n> +== Integer types\n> +\n> +In C, `<stdint.h>` (or an equivalent) must be included.\n> +\n> +[cols=\"1,1\", options=\"header\"]\n> +|===\n> +| C Type | Rust Type\n> +| uint8_t    | u8\n> +| uint16_t   | u16\n> +| uint32_t   | u32\n> +| uint64_t   | u64\n> +\n> +| int8_t     | i8\n> +| int16_t    | i16\n> +| int32_t    | i32\n> +| int64_t    | i64\n> +|===\n> +\n> +== Floating-point types\n> +\n> +Rust requires IEEE-754 semantics.\n> +In C, that is typically true, but not guaranteed by the standard.\n> +\n> +[cols=\"1,1\", options=\"header\"]\n> +|===\n> +| C Type | Rust Type\n> +| float^2^      | f32\n> +| double^2^     | f64\n> +|===\n> +\n> +== Size types\n> +\n> +These types represent pointer-sized integers and are typically defined in\n> +`<stddef.h>` or an equivalent header.\n> +\n> +Size types should be used any time pointer arithmetic is performed e.g.\n> +indexing an array, describing the number of elements in memory, etc...\n> +\n> +[cols=\"1,1\", options=\"header\"]\n> +|===\n> +| C Type | Rust Type\n> +| size_t^3^     | usize\n> +| ptrdiff_t^3^  | isize\n> +|===\n> +\n> +== Character types\n> +\n> +This is where C and Rust don't have a clean one-to-one mapping.\n> +\n> +A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n> +a `char` will have the same size as the corresponding Rust struct using `u8`.\n> +In that sense, such structs are safe to pass over the FFI boundary, because\n> +their fields will be laid out identically. However, beyond bit width, C `char`\n> +has additional semantics and platform-dependent behavior that can cause\n> +problems, as discussed below.\n> +\n> +The C language leaves the signedness of `char` implementation defined. Because\n> +our developer build enables -Wsign-compare, comparison of a value of `char`\n> +type with either signed or unsigned integers may trigger warnings from the\n> +compiler.\n\nYep, much better. Thanks!\n\nATB,\nRamsay Jones\n\n\n"},{"id":"530943","messageId":"87h5uqk69w.fsf@gitster.g","threadId":"64326","inReplyTo":"9c7a7d09-2cc0-40f7-b37a-befef5339d76@ramsayjones.plus.com","subject":"Re: [PATCH v5 01/10] doc: define unambiguous type mappings across C and Rust","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-19T04:14:51Z","receivedAt":"2025-11-19T04:15:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ramsay Jones <ramsay@ramsayjones.plus.com> writes:\n\n>> +== Character types\n>> +\n>> +This is where C and Rust don't have a clean one-to-one mapping.\n>> +\n>> +A C `char` and a Rust `u8` share the same bit width, so any C struct containing\n>> +a `char` will have the same size as the corresponding Rust struct using `u8`.\n>> +In that sense, such structs are safe to pass over the FFI boundary, because\n>> +their fields will be laid out identically. However, beyond bit width, C `char`\n>> +has additional semantics and platform-dependent behavior that can cause\n>> +problems, as discussed below.\n>> +\n>> +The C language leaves the signedness of `char` implementation defined. Because\n>> +our developer build enables -Wsign-compare, comparison of a value of `char`\n>> +type with either signed or unsigned integers may trigger warnings from the\n>> +compiler.\n>\n> Yep, much better. Thanks!\n>\n> ATB,\n> Ramsay Jones\n\nIndeed.\nThanks, both.\n"}]}