{"thread":{"id":"64712","subject":"[PATCH 00/10] Xdiff cleanup part 3","startedAt":"2026-01-02T18:52:26Z","lastAt":"2026-05-04T00:59:49Z","messageCount":124,"participants":["Ezekiel Newren via GitGitGadget","Junio C Hamano","Yee Cheng Chin","Phillip Wood","Ezekiel Newren","René Scharfe","Jeff King","D. Ben Knoble","SZEDER Gábor"],"isPatch":true,"patchVersion":1,"patchTotal":10},"messages":[{"id":"532925","messageId":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":null,"subject":"[PATCH 00/10] Xdiff cleanup part 3","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:14Z","receivedAt":"2026-01-02T18:52:26Z","isPatch":true,"body":"Patch series summary:\n\n * patch 1: Introduce the ivec type\n * patch 2: Create the function xdl_do_classic_diff()\n * patches 3-4: generic cleanup\n * patches 5-8: convert from dstart/dend (in xdfile_t) to\n   delta_start/delta_end (in xdfenv_t)\n * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n   xdiffi.c\n\nThings that will be addressed in future patch series:\n\n * Make xdl_cleanup_records() easier to read\n * convert recs/nrec into an ivec\n * convert changed to an ivec\n * remove reference_index/nreff from xdfile_t and turn it into an ivec\n * splitting minimal_perfect_hash out as its own ivec\n * improve the performance of the classifier and parsing/hashing lines\n\n=== before this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\nsize_t nreff; } xdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n\n=== after this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\nxdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\ndelta_end; size_t mph_size; } xdfenv_t;\n\nEzekiel Newren (10):\n  ivec: introduce the C side of ivec\n  xdiff: make classic diff explicit by creating xdl_do_classic_diff()\n  xdiff: don't waste time guessing the number of lines\n  xdiff: let patience and histogram benefit from xdl_trim_ends()\n  xdiff: use xdfenv_t in xdl_trim_ends() and xdl_cleanup_records()\n  xdiff: cleanup xdl_trim_ends()\n  xdiff: replace xdfile_t.dstart with xdfenv_t.delta_start\n  xdiff: replace xdfile_t.dend with xdfenv_t.delta_end\n  xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()\n  xdiff: move xdl_cleanup_records() from xprepare.c to xdiffi.c\n\n Makefile           |   1 +\n compat/ivec.c      | 113 ++++++++++++++++++\n compat/ivec.h      |  52 +++++++++\n meson.build        |   1 +\n xdiff/xdiffi.c     | 221 +++++++++++++++++++++++++++++++++---\n xdiff/xdiffi.h     |   1 +\n xdiff/xhistogram.c |   7 +-\n xdiff/xpatience.c  |   7 +-\n xdiff/xprepare.c   | 277 ++++++++-------------------------------------\n xdiff/xtypes.h     |   3 +-\n xdiff/xutils.c     |  20 ----\n xdiff/xutils.h     |   1 -\n 12 files changed, 432 insertions(+), 272 deletions(-)\n create mode 100644 compat/ivec.c\n create mode 100644 compat/ivec.h\n\n\nbase-commit: 66ce5f8e8872f0183bb137911c52b07f1f242d13\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v1\nPull-Request: https://github.com/git/git/pull/2156\n-- \ngitgitgadget\n"},{"id":"532926","messageId":"adf1395d201e916f23accc7644d21aff4f58368b.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:15Z","receivedAt":"2026-01-02T18:52:28Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nTrying to use Rust's Vec in C, or git's ALLOC_GROW() macros (via\nwrapper functions) in Rust is painful because:\n\n  * C doesn't define its own vector type, and even though Rust does\n    have Vec its painful to use on the C side (more on that below).\n    However its still not viable to use Rust's Vec type because Git\n    needs to be able to compile without Rust. So ivec was created\n    expressley to be interoperable between C and Rust without needing\n    Rust.\n  * C doing vector things the Rust way would require wrapper functions,\n    and Rust doing vector things the C way would require wrapper\n    functions, so ivec was created to ensure a consistent contract\n    between the 2 languages for how to manipulate a vector.\n  * Currently, Rust defines its own 'Vec' type that is generic, but its\n    memory allocator and struct layout weren't designed for\n    interoperability with C (or any language for that matter), meaning\n    that the C side cannot push to or expand a 'Vec' without defining\n    wrapper functions in Rust that C can call. Without special care,\n    the two languages might use different allocators (malloc/free on\n    the C side, and possibly something else in Rust), which would make\n    it difficult for a function in one language to free elements\n    allocated by a call from a function in the other language.\n  * Similarly, git defines ALLOC_GROW() and related macros in\n    git-compat-util.h. While we could add functions allowing Rust to\n    invoke something similar to those macros, passing three variables\n    (pointer, length, allocated_size) instead of a single variable\n    (vector) across the language boundary requires more cognitive\n    overhead for readers to keep track of and makes it easier to make\n    mistakes. Further, for low-level components that we want to\n    eventually convert to pure Rust, such triplets would feel very out\n    of place.\n\nTo address these issue, introduce a new type, ivec -- short for\ninteroperable vector. (We refer to it as 'ivec' generally, though on\nthe Rust side the struct is called IVec to match Rust style.)  This new\ntype is specifically designed for FFI purposes, so that both languages\nhandle the vector in the same way, though it could be used on either\nside independently. This type is designed such that it can easily be\nreplaced by a Rust 'Vec' once interoperability is no longer a concern.\n\nOne particular item to note is that Git's macros to handle vec\noperations infer the amount that a vec needs to grow from the size of\na pointer, but that makes it somewhat specific to the macros used in C.\nTo avoid defining every ivec function as a macro I opted to also\ninclude an element_size field that allows concrete functions like\npush() to know how much to grow the memory. This element_size also\nhelps in verifying that the ivec is correct when passing from C to\nRust.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n Makefile      |   1 +\n compat/ivec.c | 113 ++++++++++++++++++++++++++++++++++++++++++++++++++\n compat/ivec.h |  52 +++++++++++++++++++++++\n meson.build   |   1 +\n 4 files changed, 167 insertions(+)\n create mode 100644 compat/ivec.c\n create mode 100644 compat/ivec.h\n\ndiff --git a/Makefile b/Makefile\nindex 89d8d73ec0..f923b307d6 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1107,6 +1107,7 @@ LIB_OBJS += commit-reach.o\n LIB_OBJS += commit.o\n LIB_OBJS += common-exit.o\n LIB_OBJS += common-init.o\n+LIB_OBJS += compat/ivec.o\n LIB_OBJS += compat/nonblock.o\n LIB_OBJS += compat/obstack.o\n LIB_OBJS += compat/open.o\ndiff --git a/compat/ivec.c b/compat/ivec.c\nnew file mode 100644\nindex 0000000000..0a777e78dc\n--- /dev/null\n+++ b/compat/ivec.c\n@@ -0,0 +1,113 @@\n+#include \"ivec.h\"\n+\n+struct IVec_c_void {\n+\tvoid *ptr;\n+\tsize_t length;\n+\tsize_t capacity;\n+\tsize_t element_size;\n+};\n+\n+static void _set_capacity(void *self_, size_t new_capacity)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\n+\tif (new_capacity == self->capacity) {\n+\t\treturn;\n+\t}\n+\tif (new_capacity == 0) {\n+\t\tfree(self->ptr);\n+\t\tself->ptr = NULL;\n+\t} else {\n+\t\tself->ptr = realloc(self->ptr, new_capacity * self->element_size);\n+\t}\n+\tself->capacity = new_capacity;\n+}\n+\n+\n+void ivec_init(void *self_, size_t element_size)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\n+\tself->ptr = NULL;\n+\tself->length = 0;\n+\tself->capacity = 0;\n+\tself->element_size = element_size;\n+}\n+\n+void ivec_zero(void *self_, size_t capacity)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\n+\tself->ptr = calloc(capacity, self->element_size);\n+\tself->length = capacity;\n+\tself->capacity = capacity;\n+\t// DO NOT MODIFY element_size!!!\n+}\n+\n+void ivec_reserve_exact(void *self_, size_t additional)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\n+\t_set_capacity(self, self->capacity + additional);\n+}\n+\n+void ivec_reserve(void *self_, size_t additional)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\n+\tsize_t growby = 128;\n+\tif (self->capacity > growby)\n+\t\tgrowby = self->capacity;\n+\tif (additional > growby)\n+\t\tgrowby = additional;\n+\n+\t_set_capacity(self, self->capacity + growby);\n+}\n+\n+void ivec_shrink_to_fit(void *self_)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\n+\t_set_capacity(self, self->length);\n+}\n+\n+void ivec_push(void *self_, const void *value)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\tvoid *dst = NULL;\n+\n+\tif (self->length == self->capacity)\n+\t\tivec_reserve(self, 1);\n+\n+\tdst = (uint8_t*)self->ptr + self->length * self->element_size;\n+\tmemcpy(dst, value, self->element_size);\n+\tself->length++;\n+}\n+\n+void ivec_free(void *self_)\n+{\n+\tstruct IVec_c_void *self = self_;\n+\n+\tfree(self->ptr);\n+\tself->ptr = NULL;\n+\tself->length = 0;\n+\tself->capacity = 0;\n+\t// DO NOT MODIFY element_size!!!\n+}\n+\n+void ivec_move(void *src_, void *dst_)\n+{\n+\tstruct IVec_c_void *src = src_;\n+\tstruct IVec_c_void *dst = dst_;\n+\n+\tivec_free(dst);\n+\tdst->ptr = src->ptr;\n+\tdst->length = src->length;\n+\tdst->capacity = src->capacity;\n+\t// DO NOT MODIFY element_size!!!\n+\n+\tsrc->ptr = NULL;\n+\tsrc->length = 0;\n+\tsrc->capacity = 0;\n+\t// DO NOT MODIFY element_size!!!\n+}\ndiff --git a/compat/ivec.h b/compat/ivec.h\nnew file mode 100644\nindex 0000000000..654a05c506\n--- /dev/null\n+++ b/compat/ivec.h\n@@ -0,0 +1,52 @@\n+#ifndef IVEC_H\n+#define IVEC_H\n+\n+#include <git-compat-util.h>\n+\n+#define IVEC_INIT(variable) ivec_init(&(variable), sizeof(*(variable).ptr))\n+\n+#ifndef CBINDGEN\n+#define DEFINE_IVEC_TYPE(type, suffix) \\\n+struct IVec_##suffix { \\\n+\ttype* ptr; \\\n+\tsize_t length; \\\n+\tsize_t capacity; \\\n+\tsize_t element_size; \\\n+}\n+\n+DEFINE_IVEC_TYPE(bool, bool);\n+\n+DEFINE_IVEC_TYPE(uint8_t, u8);\n+DEFINE_IVEC_TYPE(uint16_t, u16);\n+DEFINE_IVEC_TYPE(uint32_t, u32);\n+DEFINE_IVEC_TYPE(uint64_t, u64);\n+\n+DEFINE_IVEC_TYPE(int8_t, i8);\n+DEFINE_IVEC_TYPE(int16_t, i16);\n+DEFINE_IVEC_TYPE(int32_t, i32);\n+DEFINE_IVEC_TYPE(int64_t, i64);\n+\n+DEFINE_IVEC_TYPE(float, f32);\n+DEFINE_IVEC_TYPE(double, f64);\n+\n+DEFINE_IVEC_TYPE(size_t, usize);\n+DEFINE_IVEC_TYPE(ssize_t, isize);\n+#endif\n+\n+void ivec_init(void *self_, size_t element_size);\n+\n+void ivec_zero(void *self_, size_t capacity);\n+\n+void ivec_reserve_exact(void *self_, size_t additional);\n+\n+void ivec_reserve(void *self_, size_t additional);\n+\n+void ivec_shrink_to_fit(void *self_);\n+\n+void ivec_push(void *self_, const void *value);\n+\n+void ivec_free(void *self_);\n+\n+void ivec_move(void *src, void *dst);\n+\n+#endif /* IVEC_H */\ndiff --git a/meson.build b/meson.build\nindex dd52efd1c8..42ac0c8c42 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -302,6 +302,7 @@ libgit_sources = [\n   'commit.c',\n   'common-exit.c',\n   'common-init.c',\n+  'compat/ivec.c',\n   'compat/nonblock.c',\n   'compat/obstack.c',\n   'compat/open.c',\n-- \ngitgitgadget\n\n"},{"id":"532927","messageId":"9bd01bce9f0763d9dcc962ff94fcda36346bafc4.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 02/10] xdiff: make classic diff explicit by creating xdl_do_classic_diff()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:16Z","receivedAt":"2026-01-02T18:52:29Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nLater patches will prepare xdl_cleanup_records() to be moved into xdiffi.c\nsince only the classic diff uses that function.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c | 43 +++++++++++++++++++++++++++----------------\n xdiff/xdiffi.h |  1 +\n 2 files changed, 28 insertions(+), 16 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 4376f943db..e3196c7245 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -311,26 +311,13 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n }\n \n \n-int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n-\t\txdfenv_t *xe) {\n+int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags)\n+{\n \tlong ndiags;\n \tlong *kvd, *kvdf, *kvdb;\n \txdalgoenv_t xenv;\n \tint res;\n \n-\tif (xdl_prepare_env(mf1, mf2, xpp, xe) < 0)\n-\t\treturn -1;\n-\n-\tif (XDF_DIFF_ALG(xpp->flags) == XDF_PATIENCE_DIFF) {\n-\t\tres = xdl_do_patience_diff(xpp, xe);\n-\t\tgoto out;\n-\t}\n-\n-\tif (XDF_DIFF_ALG(xpp->flags) == XDF_HISTOGRAM_DIFF) {\n-\t\tres = xdl_do_histogram_diff(xpp, xe);\n-\t\tgoto out;\n-\t}\n-\n \t/*\n \t * Allocate and setup K vectors to be used by the differential\n \t * algorithm.\n@@ -355,9 +342,33 @@ int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \txenv.heur_min = XDL_HEUR_MIN_COST;\n \n \tres = xdl_recs_cmp(&xe->xdf1, 0, xe->xdf1.nreff, &xe->xdf2, 0, xe->xdf2.nreff,\n-\t\t\t   kvdf, kvdb, (xpp->flags & XDF_NEED_MINIMAL) != 0,\n+\t\t\t   kvdf, kvdb, (flags & XDF_NEED_MINIMAL) != 0,\n \t\t\t   &xenv);\n+\n \txdl_free(kvd);\n+\n+\treturn res;\n+}\n+\n+\n+int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n+\t\txdfenv_t *xe) {\n+\tint res;\n+\n+\tif (xdl_prepare_env(mf1, mf2, xpp, xe) < 0)\n+\t\treturn -1;\n+\n+\tif (XDF_DIFF_ALG(xpp->flags) == XDF_PATIENCE_DIFF) {\n+\t\tres = xdl_do_patience_diff(xpp, xe);\n+\t\tgoto out;\n+\t}\n+\n+\tif (XDF_DIFF_ALG(xpp->flags) == XDF_HISTOGRAM_DIFF) {\n+\t\tres = xdl_do_histogram_diff(xpp, xe);\n+\t\tgoto out;\n+\t}\n+\n+\tres = xdl_do_classic_diff(xe, xpp->flags);\n  out:\n \tif (res < 0)\n \t\txdl_free_env(xe);\ndiff --git a/xdiff/xdiffi.h b/xdiff/xdiffi.h\nindex 49e52c67f9..8bf4c20373 100644\n--- a/xdiff/xdiffi.h\n+++ b/xdiff/xdiffi.h\n@@ -42,6 +42,7 @@ typedef struct s_xdchange {\n int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n \t\t xdfile_t *xdf2, long off2, long lim2,\n \t\t long *kvdf, long *kvdb, int need_min, xdalgoenv_t *xenv);\n+int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags);\n int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \t\txdfenv_t *xe);\n int xdl_change_compact(xdfile_t *xdf, xdfile_t *xdfo, long flags);\n-- \ngitgitgadget\n\n"},{"id":"532928","messageId":"53e4840c1653772379dc8d5c883b34717b81ac43.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 03/10] xdiff: don't waste time guessing the number of lines","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:17Z","receivedAt":"2026-01-02T18:52:30Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nAll lines must be read anyway, so classify them after they're read in.\nAlso move the memset() into xdl_init_classifier().\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 52 +++++++++++++++++++-----------------------------\n xdiff/xutils.c   | 20 -------------------\n xdiff/xutils.h   |  1 -\n 3 files changed, 21 insertions(+), 52 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 34c82e4f8e..96a32cc5e9 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -26,8 +26,6 @@\n #define XDL_KPDIS_RUN 4\n #define XDL_MAX_EQLIMIT 1024\n #define XDL_SIMSCAN_WINDOW 100\n-#define XDL_GUESS_NLINES1 256\n-#define XDL_GUESS_NLINES2 20\n \n #define DISCARD 0\n #define KEEP 1\n@@ -55,6 +53,8 @@ typedef struct s_xdlclassifier {\n \n \n static int xdl_init_classifier(xdlclassifier_t *cf, long size, long flags) {\n+\tmemset(cf, 0, sizeof(xdlclassifier_t));\n+\n \tcf->flags = flags;\n \n \tcf->hbits = xdl_hashbits((unsigned int) size);\n@@ -134,12 +134,12 @@ static void xdl_free_ctx(xdfile_t *xdf)\n }\n \n \n-static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_t const *xpp,\n-\t\t\t   xdlclassifier_t *cf, xdfile_t *xdf) {\n+static int xdl_prepare_ctx(mmfile_t *mf, xdfile_t *xdf, uint64_t flags) {\n \tlong bsize;\n \tuint64_t hav;\n \tuint8_t const *blk, *cur, *top, *prev;\n \txrecord_t *crec;\n+\tlong narec = 8;\n \n \txdf->reference_index = NULL;\n \txdf->changed = NULL;\n@@ -152,23 +152,21 @@ static int xdl_prepare_ctx(unsigned int pass, mmfile_t *mf, long narec, xpparam_\n \tif ((cur = blk = xdl_mmfile_first(mf, &bsize))) {\n \t\tfor (top = blk + bsize; cur < top; ) {\n \t\t\tprev = cur;\n-\t\t\thav = xdl_hash_record(&cur, top, xpp->flags);\n+\t\t\thav = xdl_hash_record(&cur, top, flags);\n \t\t\tif (XDL_ALLOC_GROW(xdf->recs, (long)xdf->nrec + 1, narec))\n \t\t\t\tgoto abort;\n \t\t\tcrec = &xdf->recs[xdf->nrec++];\n \t\t\tcrec->ptr = prev;\n \t\t\tcrec->size = cur - prev;\n \t\t\tcrec->line_hash = hav;\n-\t\t\tif (xdl_classify_record(pass, cf, crec) < 0)\n-\t\t\t\tgoto abort;\n \t\t}\n \t}\n \n \tif (!XDL_CALLOC_ARRAY(xdf->changed, xdf->nrec + 2))\n \t\tgoto abort;\n \n-\tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n-\t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF)) {\n+\tif ((XDF_DIFF_ALG(flags) != XDF_PATIENCE_DIFF) &&\n+\t    (XDF_DIFF_ALG(flags) != XDF_HISTOGRAM_DIFF)) {\n \t\tif (!XDL_ALLOC_ARRAY(xdf->reference_index, xdf->nrec + 1))\n \t\t\tgoto abort;\n \t}\n@@ -381,37 +379,29 @@ static int xdl_optimize_ctxs(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2\n \n int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \t\t    xdfenv_t *xe) {\n-\tlong enl1, enl2, sample;\n \txdlclassifier_t cf;\n \n-\tmemset(&cf, 0, sizeof(cf));\n-\n-\t/*\n-\t * For histogram diff, we can afford a smaller sample size and\n-\t * thus a poorer estimate of the number of lines, as the hash\n-\t * table (rhash) won't be filled up/grown. The number of lines\n-\t * (nrecs) will be updated correctly anyway by\n-\t * xdl_prepare_ctx().\n-\t */\n-\tsample = (XDF_DIFF_ALG(xpp->flags) == XDF_HISTOGRAM_DIFF\n-\t\t  ? XDL_GUESS_NLINES2 : XDL_GUESS_NLINES1);\n+\tif (xdl_prepare_ctx(mf1, &xe->xdf1, xpp->flags) < 0) {\n \n-\tenl1 = xdl_guess_lines(mf1, sample) + 1;\n-\tenl2 = xdl_guess_lines(mf2, sample) + 1;\n-\n-\tif (xdl_init_classifier(&cf, enl1 + enl2 + 1, xpp->flags) < 0)\n \t\treturn -1;\n+\t}\n+\tif (xdl_prepare_ctx(mf2, &xe->xdf2, xpp->flags) < 0) {\n \n-\tif (xdl_prepare_ctx(1, mf1, enl1, xpp, &cf, &xe->xdf1) < 0) {\n-\n-\t\txdl_free_classifier(&cf);\n+\t\txdl_free_ctx(&xe->xdf1);\n \t\treturn -1;\n \t}\n-\tif (xdl_prepare_ctx(2, mf2, enl2, xpp, &cf, &xe->xdf2) < 0) {\n \n-\t\txdl_free_ctx(&xe->xdf1);\n-\t\txdl_free_classifier(&cf);\n+\tif (xdl_init_classifier(&cf, xe->xdf1.nrec + xe->xdf2.nrec + 1, xpp->flags) < 0)\n \t\treturn -1;\n+\n+\tfor (size_t i = 0; i < xe->xdf1.nrec; i++) {\n+\t\txrecord_t *rec = &xe->xdf1.recs[i];\n+\t\txdl_classify_record(1, &cf, rec);\n+\t}\n+\n+\tfor (size_t i = 0; i < xe->xdf2.nrec; i++) {\n+\t\txrecord_t *rec = &xe->xdf2.recs[i];\n+\t\txdl_classify_record(2, &cf, rec);\n \t}\n \n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 77ee1ad9c8..b3d51197c1 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -118,26 +118,6 @@ void *xdl_cha_alloc(chastore_t *cha) {\n \treturn data;\n }\n \n-long xdl_guess_lines(mmfile_t *mf, long sample) {\n-\tlong nl = 0, size, tsize = 0;\n-\tchar const *data, *cur, *top;\n-\n-\tif ((cur = data = xdl_mmfile_first(mf, &size))) {\n-\t\tfor (top = data + size; nl < sample && cur < top; ) {\n-\t\t\tnl++;\n-\t\t\tif (!(cur = memchr(cur, '\\n', top - cur)))\n-\t\t\t\tcur = top;\n-\t\t\telse\n-\t\t\t\tcur++;\n-\t\t}\n-\t\ttsize += (long) (cur - data);\n-\t}\n-\n-\tif (nl && tsize)\n-\t\tnl = xdl_mmfile_size(mf) / (tsize / nl);\n-\n-\treturn nl + 1;\n-}\n \n int xdl_blankline(const char *line, long size, long flags)\n {\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 615b4a9d35..d800840dd0 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -31,7 +31,6 @@ int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n int xdl_cha_init(chastore_t *cha, long isize, long icount);\n void xdl_cha_free(chastore_t *cha);\n void *xdl_cha_alloc(chastore_t *cha);\n-long xdl_guess_lines(mmfile_t *mf, long sample);\n int xdl_blankline(const char *line, long size, long flags);\n int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags);\n uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top);\n-- \ngitgitgadget\n\n"},{"id":"532929","messageId":"70040ea1351451243be90d59d26cf1a403f3000a.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 04/10] xdiff: let patience and histogram benefit from xdl_trim_ends()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:18Z","receivedAt":"2026-01-02T18:52:32Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThe patience diff is set up the exact same way as histogram, see\nxdl_do_historgram_diff() in xhistogram.c. xdl_optimize_ctxs() is\nredundant now, delete it.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xpatience.c |  4 +++-\n xdiff/xprepare.c  | 14 ++------------\n 2 files changed, 5 insertions(+), 13 deletions(-)\n\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 9580d18032..2bce07cf48 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -373,5 +373,7 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n \n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n-\treturn patience_diff(xpp, env, 1, (int)env->xdf1.nrec, 1, (int)env->xdf2.nrec);\n+\treturn patience_diff(xpp, env,\n+\t\tenv->xdf1.dstart + 1, env->xdf1.dend - env->xdf1.dstart + 1,\n+\t\tenv->xdf2.dstart + 1, env->xdf2.dend - env->xdf2.dstart + 1);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 96a32cc5e9..0d7d9f6146 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -366,17 +366,6 @@ static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n }\n \n \n-static int xdl_optimize_ctxs(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\n-\tif (xdl_trim_ends(xdf1, xdf2) < 0 ||\n-\t    xdl_cleanup_records(cf, xdf1, xdf2) < 0) {\n-\n-\t\treturn -1;\n-\t}\n-\n-\treturn 0;\n-}\n-\n int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \t\t    xdfenv_t *xe) {\n \txdlclassifier_t cf;\n@@ -404,9 +393,10 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \t\txdl_classify_record(2, &cf, rec);\n \t}\n \n+\txdl_trim_ends(&xe->xdf1, &xe->xdf2);\n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n-\t    xdl_optimize_ctxs(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n+\t    xdl_cleanup_records(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n \n \t\txdl_free_ctx(&xe->xdf2);\n \t\txdl_free_ctx(&xe->xdf1);\n-- \ngitgitgadget\n\n"},{"id":"532930","messageId":"742f2d381af52cd8314a6a643e84b6b9daf99c7c.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 05/10] xdiff: use xdfenv_t in xdl_trim_ends() and xdl_cleanup_records()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:19Z","receivedAt":"2026-01-02T18:52:33Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nView with --color-words. Prepare these functions to use the fields:\ndelta_start, delta_end. A future patch will add these fields to\nxdfenv_t.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 60 ++++++++++++++++++++++++------------------------\n 1 file changed, 30 insertions(+), 30 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 0d7d9f6146..0acb3437d4 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -261,7 +261,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * matches on the other file. Also, lines that have multiple matches\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n-static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n+static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \tlong i, nm, mlim;\n \txrecord_t *recs;\n \txdlclass_t *rcrec;\n@@ -273,11 +273,11 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * Create temporary arrays that will help us decide if\n \t * changed[i] should remain false, or become true.\n \t */\n-\tif (!XDL_CALLOC_ARRAY(action1, xdf1->nrec + 1)) {\n+\tif (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n \t\tret = -1;\n \t\tgoto cleanup;\n \t}\n-\tif (!XDL_CALLOC_ARRAY(action2, xdf2->nrec + 1)) {\n+\tif (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n \t\tret = -1;\n \t\tgoto cleanup;\n \t}\n@@ -285,17 +285,17 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n+\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart]; i <= xe->xdf1.dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n+\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart]; i <= xe->xdf2.dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n@@ -305,27 +305,27 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * Use temporary arrays to decide if changed[i] should remain\n \t * false, or become true.\n \t */\n-\txdf1->nreff = 0;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n-\t     i <= xdf1->dend; i++, recs++) {\n+\txe->xdf1.nreff = 0;\n+\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart];\n+\t     i <= xe->xdf1.dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n-\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n+\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->xdf1.dstart, xe->xdf1.dend))) {\n+\t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n-\t\t\txdf1->changed[i] = true;\n+\t\t\txe->xdf1.changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n \n-\txdf2->nreff = 0;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n-\t     i <= xdf2->dend; i++, recs++) {\n+\txe->xdf2.nreff = 0;\n+\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart];\n+\t     i <= xe->xdf2.dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n-\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n+\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->xdf2.dstart, xe->xdf2.dend))) {\n+\t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n-\t\t\txdf2->changed[i] = true;\n+\t\t\txe->xdf2.changed[i] = true;\n \t\t\t/* i.e. discard */\n \t}\n \n@@ -340,27 +340,27 @@ cleanup:\n /*\n  * Early trim initial and terminal matching records.\n  */\n-static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n+static int xdl_trim_ends(xdfenv_t *xe) {\n \tlong i, lim;\n \txrecord_t *recs1, *recs2;\n \n-\trecs1 = xdf1->recs;\n-\trecs2 = xdf2->recs;\n-\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n+\trecs1 = xe->xdf1.recs;\n+\trecs2 = xe->xdf2.recs;\n+\tfor (i = 0, lim = (long)XDL_MIN(xe->xdf1.nrec, xe->xdf2.nrec); i < lim;\n \t     i++, recs1++, recs2++)\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n-\txdf1->dstart = xdf2->dstart = i;\n+\txe->xdf1.dstart = xe->xdf2.dstart = i;\n \n-\trecs1 = xdf1->recs + xdf1->nrec - 1;\n-\trecs2 = xdf2->recs + xdf2->nrec - 1;\n+\trecs1 = xe->xdf1.recs + xe->xdf1.nrec - 1;\n+\trecs2 = xe->xdf2.recs + xe->xdf2.nrec - 1;\n \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n \t\t\tbreak;\n \n-\txdf1->dend = (long)xdf1->nrec - i - 1;\n-\txdf2->dend = (long)xdf2->nrec - i - 1;\n+\txe->xdf1.dend = (long)xe->xdf1.nrec - i - 1;\n+\txe->xdf2.dend = (long)xe->xdf2.nrec - i - 1;\n \n \treturn 0;\n }\n@@ -393,10 +393,10 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \t\txdl_classify_record(2, &cf, rec);\n \t}\n \n-\txdl_trim_ends(&xe->xdf1, &xe->xdf2);\n+\txdl_trim_ends(xe);\n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n-\t    xdl_cleanup_records(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n+\t    xdl_cleanup_records(&cf, xe) < 0) {\n \n \t\txdl_free_ctx(&xe->xdf2);\n \t\txdl_free_ctx(&xe->xdf1);\n-- \ngitgitgadget\n\n"},{"id":"532931","messageId":"65da408da9589420ec341368d0853e6183aee922.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 06/10] xdiff: cleanup xdl_trim_ends()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:20Z","receivedAt":"2026-01-02T18:52:34Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThis patch is best viewed with a before and after of the whole\nfunction.\n\nRather than using 2 pointers and walking them. Use direct indexing with\nlocal variables of what is being compared to make it easier to follow\nalong.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 40 ++++++++++++++++++++--------------------\n 1 file changed, 20 insertions(+), 20 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 0acb3437d4..06b6a6f804 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -340,29 +340,29 @@ cleanup:\n /*\n  * Early trim initial and terminal matching records.\n  */\n-static int xdl_trim_ends(xdfenv_t *xe) {\n-\tlong i, lim;\n-\txrecord_t *recs1, *recs2;\n-\n-\trecs1 = xe->xdf1.recs;\n-\trecs2 = xe->xdf2.recs;\n-\tfor (i = 0, lim = (long)XDL_MIN(xe->xdf1.nrec, xe->xdf2.nrec); i < lim;\n-\t     i++, recs1++, recs2++)\n-\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n+static void xdl_trim_ends(xdfenv_t *xe)\n+{\n+\tsize_t lim = XDL_MIN(xe->xdf1.nrec, xe->xdf2.nrec);\n+\n+\tfor (size_t i = 0; i < lim; i++) {\n+\t\tsize_t mph1 = xe->xdf1.recs[i].minimal_perfect_hash;\n+\t\tsize_t mph2 = xe->xdf2.recs[i].minimal_perfect_hash;\n+\t\tif (mph1 != mph2) {\n+\t\t\txe->xdf1.dstart = xe->xdf2.dstart = (ssize_t)i;\n+\t\t\tlim -= i;\n \t\t\tbreak;\n+\t\t}\n+\t}\n \n-\txe->xdf1.dstart = xe->xdf2.dstart = i;\n-\n-\trecs1 = xe->xdf1.recs + xe->xdf1.nrec - 1;\n-\trecs2 = xe->xdf2.recs + xe->xdf2.nrec - 1;\n-\tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n-\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n+\tfor (size_t i = 0; i < lim; i++) {\n+\t\tsize_t mph1 = xe->xdf1.recs[xe->xdf1.nrec - 1 - i].minimal_perfect_hash;\n+\t\tsize_t mph2 = xe->xdf2.recs[xe->xdf2.nrec - 1 - i].minimal_perfect_hash;\n+\t\tif (mph1 != mph2) {\n+\t\t\txe->xdf1.dend = xe->xdf1.nrec - 1 - i;\n+\t\t\txe->xdf2.dend = xe->xdf2.nrec - 1 - i;\n \t\t\tbreak;\n-\n-\txe->xdf1.dend = (long)xe->xdf1.nrec - i - 1;\n-\txe->xdf2.dend = (long)xe->xdf2.nrec - i - 1;\n-\n-\treturn 0;\n+\t\t}\n+\t}\n }\n \n \n-- \ngitgitgadget\n\n"},{"id":"532932","messageId":"d74722538b693fb26e8684f9dd3fbc319a2a575e.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 07/10] xdiff: replace xdfile_t.dstart with xdfenv_t.delta_start","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:21Z","receivedAt":"2026-01-02T18:52:36Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nPlacing delta_start in xdfenv_t instead of xdfile_t provides a more\nappropriate context since this variable only makes sense with a pair\nof files. View with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xhistogram.c |  4 ++--\n xdiff/xpatience.c  |  4 ++--\n xdiff/xprepare.c   | 17 +++++++++--------\n xdiff/xtypes.h     |  3 ++-\n 4 files changed, 15 insertions(+), 13 deletions(-)\n\ndiff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\nindex 5ae1282c27..eb6a52d9ba 100644\n--- a/xdiff/xhistogram.c\n+++ b/xdiff/xhistogram.c\n@@ -365,6 +365,6 @@ out:\n int xdl_do_histogram_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n \treturn histogram_diff(xpp, env,\n-\t\tenv->xdf1.dstart + 1, env->xdf1.dend - env->xdf1.dstart + 1,\n-\t\tenv->xdf2.dstart + 1, env->xdf2.dend - env->xdf2.dstart + 1);\n+\t\tenv->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n+\t\tenv->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n }\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 2bce07cf48..bd0ffbb417 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -374,6 +374,6 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n \treturn patience_diff(xpp, env,\n-\t\tenv->xdf1.dstart + 1, env->xdf1.dend - env->xdf1.dstart + 1,\n-\t\tenv->xdf2.dstart + 1, env->xdf2.dend - env->xdf2.dstart + 1);\n+\t\tenv->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n+\t\tenv->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 06b6a6f804..e88468e74c 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -173,7 +173,6 @@ static int xdl_prepare_ctx(mmfile_t *mf, xdfile_t *xdf, uint64_t flags) {\n \n \txdf->changed += 1;\n \txdf->nreff = 0;\n-\txdf->dstart = 0;\n \txdf->dend = xdf->nrec - 1;\n \n \treturn 0;\n@@ -287,7 +286,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \t */\n \tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart]; i <= xe->xdf1.dend; i++, recs++) {\n+\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= xe->xdf1.dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n@@ -295,7 +294,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \n \tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart]; i <= xe->xdf2.dend; i++, recs++) {\n+\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= xe->xdf2.dend; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n@@ -306,10 +305,10 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \t * false, or become true.\n \t */\n \txe->xdf1.nreff = 0;\n-\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart];\n+\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n \t     i <= xe->xdf1.dend; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->xdf1.dstart, xe->xdf1.dend))) {\n+\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, xe->xdf1.dend))) {\n \t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n@@ -318,10 +317,10 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \t}\n \n \txe->xdf2.nreff = 0;\n-\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart];\n+\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n \t     i <= xe->xdf2.dend; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->xdf2.dstart, xe->xdf2.dend))) {\n+\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, xe->xdf2.dend))) {\n \t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n@@ -348,7 +347,7 @@ static void xdl_trim_ends(xdfenv_t *xe)\n \t\tsize_t mph1 = xe->xdf1.recs[i].minimal_perfect_hash;\n \t\tsize_t mph2 = xe->xdf2.recs[i].minimal_perfect_hash;\n \t\tif (mph1 != mph2) {\n-\t\t\txe->xdf1.dstart = xe->xdf2.dstart = (ssize_t)i;\n+\t\t\txe->delta_start = (ssize_t)i;\n \t\t\tlim -= i;\n \t\t\tbreak;\n \t\t}\n@@ -370,6 +369,8 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \t\t    xdfenv_t *xe) {\n \txdlclassifier_t cf;\n \n+\txe->delta_start = 0;\n+\n \tif (xdl_prepare_ctx(mf1, &xe->xdf1, xpp->flags) < 0) {\n \n \t\treturn -1;\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex 979586f20a..bda1f85eb0 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -48,7 +48,7 @@ typedef struct s_xrecord {\n typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n-\tptrdiff_t dstart, dend;\n+\tptrdiff_t dend;\n \tbool *changed;\n \tsize_t *reference_index;\n \tsize_t nreff;\n@@ -56,6 +56,7 @@ typedef struct s_xdfile {\n \n typedef struct s_xdfenv {\n \txdfile_t xdf1, xdf2;\n+\tsize_t delta_start;\n } xdfenv_t;\n \n \n-- \ngitgitgadget\n\n"},{"id":"532933","messageId":"d0ef5b23c4483c32069594374489234d48050384.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 08/10] xdiff: replace xdfile_t.dend with xdfenv_t.delta_end","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:22Z","receivedAt":"2026-01-02T18:52:37Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nView with --color-words. Same argument as delta_start.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xhistogram.c |  7 +++++--\n xdiff/xpatience.c  |  7 +++++--\n xdiff/xprepare.c   | 19 ++++++++++---------\n xdiff/xtypes.h     |  3 +--\n 4 files changed, 21 insertions(+), 15 deletions(-)\n\ndiff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\nindex eb6a52d9ba..b4d6f88748 100644\n--- a/xdiff/xhistogram.c\n+++ b/xdiff/xhistogram.c\n@@ -364,7 +364,10 @@ out:\n \n int xdl_do_histogram_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n+\tptrdiff_t dend1 = env->xdf1.nrec - 1 - env->delta_end;\n+\tptrdiff_t dend2 = env->xdf2.nrec - 1 - env->delta_end;\n+\n \treturn histogram_diff(xpp, env,\n-\t\tenv->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n-\t\tenv->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n+\t\tenv->delta_start + 1, dend1 - env->delta_start + 1,\n+\t\tenv->delta_start + 1, dend2 - env->delta_start + 1);\n }\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex bd0ffbb417..5b8bb34d2b 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -373,7 +373,10 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n \n int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n {\n+\tptrdiff_t dend1 = env->xdf1.nrec - 1 - env->delta_end;\n+\tptrdiff_t dend2 = env->xdf2.nrec - 1 - env->delta_end;\n+\n \treturn patience_diff(xpp, env,\n-\t\tenv->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n-\t\tenv->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n+\t\tenv->delta_start + 1, dend1 - env->delta_start + 1,\n+\t\tenv->delta_start + 1, dend2 - env->delta_start + 1);\n }\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex e88468e74c..d3cdb6ac02 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -173,7 +173,6 @@ static int xdl_prepare_ctx(mmfile_t *mf, xdfile_t *xdf, uint64_t flags) {\n \n \txdf->changed += 1;\n \txdf->nreff = 0;\n-\txdf->dend = xdf->nrec - 1;\n \n \treturn 0;\n \n@@ -267,6 +266,8 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n \tint ret = 0;\n+\tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n+\tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n \n \t/*\n \t * Create temporary arrays that will help us decide if\n@@ -286,7 +287,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \t */\n \tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= xe->xdf1.dend; i++, recs++) {\n+\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n@@ -294,7 +295,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \n \tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= xe->xdf2.dend; i++, recs++) {\n+\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n@@ -306,9 +307,9 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \t */\n \txe->xdf1.nreff = 0;\n \tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n-\t     i <= xe->xdf1.dend; i++, recs++) {\n+\t     i <= dend1; i++, recs++) {\n \t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, xe->xdf1.dend))) {\n+\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, dend1))) {\n \t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n@@ -318,9 +319,9 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \n \txe->xdf2.nreff = 0;\n \tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n-\t     i <= xe->xdf2.dend; i++, recs++) {\n+\t     i <= dend2; i++, recs++) {\n \t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, xe->xdf2.dend))) {\n+\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, dend2))) {\n \t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n \t\t} else\n@@ -357,8 +358,7 @@ static void xdl_trim_ends(xdfenv_t *xe)\n \t\tsize_t mph1 = xe->xdf1.recs[xe->xdf1.nrec - 1 - i].minimal_perfect_hash;\n \t\tsize_t mph2 = xe->xdf2.recs[xe->xdf2.nrec - 1 - i].minimal_perfect_hash;\n \t\tif (mph1 != mph2) {\n-\t\t\txe->xdf1.dend = xe->xdf1.nrec - 1 - i;\n-\t\t\txe->xdf2.dend = xe->xdf2.nrec - 1 - i;\n+\t\t\txe->delta_end = i;\n \t\t\tbreak;\n \t\t}\n \t}\n@@ -370,6 +370,7 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \txdlclassifier_t cf;\n \n \txe->delta_start = 0;\n+\txe->delta_end = 0;\n \n \tif (xdl_prepare_ctx(mf1, &xe->xdf1, xpp->flags) < 0) {\n \ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex bda1f85eb0..a939396064 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -48,7 +48,6 @@ typedef struct s_xrecord {\n typedef struct s_xdfile {\n \txrecord_t *recs;\n \tsize_t nrec;\n-\tptrdiff_t dend;\n \tbool *changed;\n \tsize_t *reference_index;\n \tsize_t nreff;\n@@ -56,7 +55,7 @@ typedef struct s_xdfile {\n \n typedef struct s_xdfenv {\n \txdfile_t xdf1, xdf2;\n-\tsize_t delta_start;\n+\tsize_t delta_start, delta_end;\n } xdfenv_t;\n \n \n-- \ngitgitgadget\n\n"},{"id":"532934","messageId":"f9b10e71d23f8b4fa34dcffb371cf5a173760409.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 09/10] xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:23Z","receivedAt":"2026-01-02T18:52:39Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nDisentangle xdl_cleanup_records() from the classifier so that it can be\nmoved from xprepare.c into xdiffi.c.\n\nThe classic diff is the only algorithm that needs to count the number\nof times each line occurs in each file. Make xdl_cleanup_records()\ncount the number of lines instead of the classifier so it won't slow\ndown patience or histogram.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 52 +++++++++++++++++++++++++++++++++---------------\n xdiff/xtypes.h   |  1 +\n 2 files changed, 37 insertions(+), 16 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex d3cdb6ac02..b53a3b80c4 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -21,6 +21,7 @@\n  */\n \n #include \"xinclude.h\"\n+#include \"compat/ivec.h\"\n \n \n #define XDL_KPDIS_RUN 4\n@@ -35,7 +36,6 @@ typedef struct s_xdlclass {\n \tstruct s_xdlclass *next;\n \txrecord_t rec;\n \tlong idx;\n-\tlong len1, len2;\n } xdlclass_t;\n \n typedef struct s_xdlclassifier {\n@@ -92,7 +92,7 @@ static void xdl_free_classifier(xdlclassifier_t *cf) {\n }\n \n \n-static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n+static int xdl_classify_record(xdlclassifier_t *cf, xrecord_t *rec) {\n \tsize_t hi;\n \txdlclass_t *rcrec;\n \n@@ -113,13 +113,10 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n \t\t\t\treturn -1;\n \t\tcf->rcrecs[rcrec->idx] = rcrec;\n \t\trcrec->rec = *rec;\n-\t\trcrec->len1 = rcrec->len2 = 0;\n \t\trcrec->next = cf->rchash[hi];\n \t\tcf->rchash[hi] = rcrec;\n \t}\n \n-\t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n-\n \trec->minimal_perfect_hash = (size_t)rcrec->idx;\n \n \treturn 0;\n@@ -253,22 +250,44 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n \treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n }\n \n+struct xoccurrence\n+{\n+\tsize_t file1, file2;\n+};\n+\n+\n+DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n+\n \n /*\n  * Try to reduce the problem complexity, discard records that have no\n  * matches on the other file. Also, lines that have multiple matches\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n-static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n-\tlong i, nm, mlim;\n+static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n+\tlong i;\n+\tsize_t nm, mlim;\n \txrecord_t *recs;\n-\txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n-\tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n+\tstruct IVec_xoccurrence occ;\n+\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n \tint ret = 0;\n \tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n \tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n \n+\tIVEC_INIT(occ);\n+\tivec_zero(&occ, xe->mph_size);\n+\n+\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n+\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n+\t\tocc.ptr[mph1].file1 += 1;\n+\t}\n+\n+\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n+\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n+\t\tocc.ptr[mph2].file2 += 1;\n+\t}\n+\n \t/*\n \t * Create temporary arrays that will help us decide if\n \t * changed[i] should remain false, or become true.\n@@ -288,16 +307,14 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n \tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n-\t\tnm = rcrec ? rcrec->len2 : 0;\n+\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n \tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n-\t\tnm = rcrec ? rcrec->len1 : 0;\n+\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n@@ -332,6 +349,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n cleanup:\n \txdl_free(action1);\n \txdl_free(action2);\n+\tivec_free(&occ);\n \n \treturn ret;\n }\n@@ -387,18 +405,20 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \n \tfor (size_t i = 0; i < xe->xdf1.nrec; i++) {\n \t\txrecord_t *rec = &xe->xdf1.recs[i];\n-\t\txdl_classify_record(1, &cf, rec);\n+\t\txdl_classify_record(&cf, rec);\n \t}\n \n \tfor (size_t i = 0; i < xe->xdf2.nrec; i++) {\n \t\txrecord_t *rec = &xe->xdf2.recs[i];\n-\t\txdl_classify_record(2, &cf, rec);\n+\t\txdl_classify_record(&cf, rec);\n \t}\n \n+\txe->mph_size = cf.count;\n+\n \txdl_trim_ends(xe);\n \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n-\t    xdl_cleanup_records(&cf, xe) < 0) {\n+\t    xdl_cleanup_records(xe, xpp->flags) < 0) {\n \n \t\txdl_free_ctx(&xe->xdf2);\n \t\txdl_free_ctx(&xe->xdf1);\ndiff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\nindex a939396064..2528bd37e8 100644\n--- a/xdiff/xtypes.h\n+++ b/xdiff/xtypes.h\n@@ -56,6 +56,7 @@ typedef struct s_xdfile {\n typedef struct s_xdfenv {\n \txdfile_t xdf1, xdf2;\n \tsize_t delta_start, delta_end;\n+\tsize_t mph_size;\n } xdfenv_t;\n \n \n-- \ngitgitgadget\n\n"},{"id":"532935","messageId":"1dba6b34aa5c3eec06ae50a74d133c37b1d2404e.1767379944.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH 10/10] xdiff: move xdl_cleanup_records() from xprepare.c to xdiffi.c","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-02T18:52:24Z","receivedAt":"2026-01-02T18:52:40Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nOnly the classic diff uses xdl_cleanup_records(). Move it,\nxdl_clean_mmatch(), and the macros to xdiffi.c and call\nxdl_cleanup_records() inside of xdl_do_classic_diff(). This better\norganizes the code related to the classic diff.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   | 180 ++++++++++++++++++++++++++++++++++++++++++++\n xdiff/xprepare.c | 191 +----------------------------------------------\n 2 files changed, 181 insertions(+), 190 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex e3196c7245..0f1fd7cf80 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -21,6 +21,7 @@\n  */\n \n #include \"xinclude.h\"\n+#include \"compat/ivec.h\"\n \n static size_t get_hash(xdfile_t *xdf, long index)\n {\n@@ -33,6 +34,14 @@ static size_t get_hash(xdfile_t *xdf, long index)\n #define XDL_SNAKE_CNT 20\n #define XDL_K_HEUR 4\n \n+#define XDL_KPDIS_RUN 4\n+#define XDL_MAX_EQLIMIT 1024\n+#define XDL_SIMSCAN_WINDOW 100\n+\n+#define DISCARD 0\n+#define KEEP 1\n+#define INVESTIGATE 2\n+\n typedef struct s_xdpsplit {\n \tlong i1, i2;\n \tint min_lo, min_hi;\n@@ -311,6 +320,175 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n }\n \n \n+static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n+\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n+\n+\t/*\n+\t * Limits the window that is examined during the similar-lines\n+\t * scan. The loops below stops when action[i - r] == KEEP\n+\t * (line that has no match), but there are corner cases where\n+\t * the loop proceed all the way to the extremities by causing\n+\t * huge performance penalties in case of big files.\n+\t */\n+\tif (i - s > XDL_SIMSCAN_WINDOW)\n+\t\ts = i - XDL_SIMSCAN_WINDOW;\n+\tif (e - i > XDL_SIMSCAN_WINDOW)\n+\t\te = i + XDL_SIMSCAN_WINDOW;\n+\n+\t/*\n+\t * Scans the lines before 'i' to find a run of lines that either\n+\t * have no match (action[j] == DISCARD) or have multiple matches\n+\t * (action[j] == INVESTIGATE). Note that we always call this\n+\t * function with action[i] == INVESTIGATE, so the current line\n+\t * (i) is already a multimatch line.\n+\t */\n+\tfor (r = 1, rdis0 = 0, rpdis0 = 1; (i - r) >= s; r++) {\n+\t\tif (action[i - r] == DISCARD)\n+\t\t\trdis0++;\n+\t\telse if (action[i - r] == INVESTIGATE)\n+\t\t\trpdis0++;\n+\t\telse if (action[i - r] == KEEP)\n+\t\t\tbreak;\n+\t\telse\n+\t\t\tBUG(\"Illegal value for action[i - r]\");\n+\t}\n+\t/*\n+\t * If the run before the line 'i' found only multimatch lines,\n+\t * we return false and hence we don't make the current line (i)\n+\t * discarded. We want to discard multimatch lines only when\n+\t * they appear in the middle of runs with nomatch lines\n+\t * (action[j] == DISCARD).\n+\t */\n+\tif (rdis0 == 0)\n+\t\treturn 0;\n+\tfor (r = 1, rdis1 = 0, rpdis1 = 1; (i + r) <= e; r++) {\n+\t\tif (action[i + r] == DISCARD)\n+\t\t\trdis1++;\n+\t\telse if (action[i + r] == INVESTIGATE)\n+\t\t\trpdis1++;\n+\t\telse if (action[i + r] == KEEP)\n+\t\t\tbreak;\n+\t\telse\n+\t\t\tBUG(\"Illegal value for action[i + r]\");\n+\t}\n+\t/*\n+\t * If the run after the line 'i' found only multimatch lines,\n+\t * we return false and hence we don't make the current line (i)\n+\t * discarded.\n+\t */\n+\tif (rdis1 == 0)\n+\t\treturn false;\n+\trdis1 += rdis0;\n+\trpdis1 += rpdis0;\n+\n+\treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n+}\n+\n+struct xoccurrence\n+{\n+\tsize_t file1, file2;\n+};\n+\n+\n+DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n+\n+\n+/*\n+ * Try to reduce the problem complexity, discard records that have no\n+ * matches on the other file. Also, lines that have multiple matches\n+ * might be potentially discarded if they appear in a run of discardable.\n+ */\n+static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n+\tlong i;\n+\tsize_t nm, mlim;\n+\txrecord_t *recs;\n+\tuint8_t *action1 = NULL, *action2 = NULL;\n+\tstruct IVec_xoccurrence occ;\n+\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n+\tint ret = 0;\n+\tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n+\tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n+\n+\tIVEC_INIT(occ);\n+\tivec_zero(&occ, xe->mph_size);\n+\n+\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n+\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n+\t\tocc.ptr[mph1].file1 += 1;\n+\t}\n+\n+\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n+\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n+\t\tocc.ptr[mph2].file2 += 1;\n+\t}\n+\n+\t/*\n+\t * Create temporary arrays that will help us decide if\n+\t * changed[i] should remain false, or become true.\n+\t */\n+\tif (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\tif (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/*\n+\t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n+\t */\n+\tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n+\t\tmlim = XDL_MAX_EQLIMIT;\n+\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n+\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n+\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t}\n+\n+\tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n+\t\tmlim = XDL_MAX_EQLIMIT;\n+\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n+\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n+\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t}\n+\n+\t/*\n+\t * Use temporary arrays to decide if changed[i] should remain\n+\t * false, or become true.\n+\t */\n+\txe->xdf1.nreff = 0;\n+\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n+\t     i <= dend1; i++, recs++) {\n+\t\tif (action1[i] == KEEP ||\n+\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, dend1))) {\n+\t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n+\t\t\t/* changed[i] remains false, i.e. keep */\n+\t\t} else\n+\t\t\txe->xdf1.changed[i] = true;\n+\t\t\t/* i.e. discard */\n+\t}\n+\n+\txe->xdf2.nreff = 0;\n+\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n+\t     i <= dend2; i++, recs++) {\n+\t\tif (action2[i] == KEEP ||\n+\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, dend2))) {\n+\t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n+\t\t\t/* changed[i] remains false, i.e. keep */\n+\t\t} else\n+\t\t\txe->xdf2.changed[i] = true;\n+\t\t\t/* i.e. discard */\n+\t}\n+\n+cleanup:\n+\txdl_free(action1);\n+\txdl_free(action2);\n+\tivec_free(&occ);\n+\n+\treturn ret;\n+}\n+\n+\n int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags)\n {\n \tlong ndiags;\n@@ -318,6 +496,8 @@ int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags)\n \txdalgoenv_t xenv;\n \tint res;\n \n+\txdl_cleanup_records(xe, flags);\n+\n \t/*\n \t * Allocate and setup K vectors to be used by the differential\n \t * algorithm.\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex b53a3b80c4..3f555e29f4 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -24,14 +24,6 @@\n #include \"compat/ivec.h\"\n \n \n-#define XDL_KPDIS_RUN 4\n-#define XDL_MAX_EQLIMIT 1024\n-#define XDL_SIMSCAN_WINDOW 100\n-\n-#define DISCARD 0\n-#define KEEP 1\n-#define INVESTIGATE 2\n-\n typedef struct s_xdlclass {\n \tstruct s_xdlclass *next;\n \txrecord_t rec;\n@@ -50,8 +42,6 @@ typedef struct s_xdlclassifier {\n } xdlclassifier_t;\n \n \n-\n-\n static int xdl_init_classifier(xdlclassifier_t *cf, long size, long flags) {\n \tmemset(cf, 0, sizeof(xdlclassifier_t));\n \n@@ -186,175 +176,6 @@ void xdl_free_env(xdfenv_t *xe) {\n }\n \n \n-static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n-\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n-\n-\t/*\n-\t * Limits the window that is examined during the similar-lines\n-\t * scan. The loops below stops when action[i - r] == KEEP\n-\t * (line that has no match), but there are corner cases where\n-\t * the loop proceed all the way to the extremities by causing\n-\t * huge performance penalties in case of big files.\n-\t */\n-\tif (i - s > XDL_SIMSCAN_WINDOW)\n-\t\ts = i - XDL_SIMSCAN_WINDOW;\n-\tif (e - i > XDL_SIMSCAN_WINDOW)\n-\t\te = i + XDL_SIMSCAN_WINDOW;\n-\n-\t/*\n-\t * Scans the lines before 'i' to find a run of lines that either\n-\t * have no match (action[j] == DISCARD) or have multiple matches\n-\t * (action[j] == INVESTIGATE). Note that we always call this\n-\t * function with action[i] == INVESTIGATE, so the current line\n-\t * (i) is already a multimatch line.\n-\t */\n-\tfor (r = 1, rdis0 = 0, rpdis0 = 1; (i - r) >= s; r++) {\n-\t\tif (action[i - r] == DISCARD)\n-\t\t\trdis0++;\n-\t\telse if (action[i - r] == INVESTIGATE)\n-\t\t\trpdis0++;\n-\t\telse if (action[i - r] == KEEP)\n-\t\t\tbreak;\n-\t\telse\n-\t\t\tBUG(\"Illegal value for action[i - r]\");\n-\t}\n-\t/*\n-\t * If the run before the line 'i' found only multimatch lines,\n-\t * we return false and hence we don't make the current line (i)\n-\t * discarded. We want to discard multimatch lines only when\n-\t * they appear in the middle of runs with nomatch lines\n-\t * (action[j] == DISCARD).\n-\t */\n-\tif (rdis0 == 0)\n-\t\treturn 0;\n-\tfor (r = 1, rdis1 = 0, rpdis1 = 1; (i + r) <= e; r++) {\n-\t\tif (action[i + r] == DISCARD)\n-\t\t\trdis1++;\n-\t\telse if (action[i + r] == INVESTIGATE)\n-\t\t\trpdis1++;\n-\t\telse if (action[i + r] == KEEP)\n-\t\t\tbreak;\n-\t\telse\n-\t\t\tBUG(\"Illegal value for action[i + r]\");\n-\t}\n-\t/*\n-\t * If the run after the line 'i' found only multimatch lines,\n-\t * we return false and hence we don't make the current line (i)\n-\t * discarded.\n-\t */\n-\tif (rdis1 == 0)\n-\t\treturn false;\n-\trdis1 += rdis0;\n-\trpdis1 += rpdis0;\n-\n-\treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n-}\n-\n-struct xoccurrence\n-{\n-\tsize_t file1, file2;\n-};\n-\n-\n-DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n-\n-\n-/*\n- * Try to reduce the problem complexity, discard records that have no\n- * matches on the other file. Also, lines that have multiple matches\n- * might be potentially discarded if they appear in a run of discardable.\n- */\n-static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n-\tlong i;\n-\tsize_t nm, mlim;\n-\txrecord_t *recs;\n-\tuint8_t *action1 = NULL, *action2 = NULL;\n-\tstruct IVec_xoccurrence occ;\n-\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n-\tint ret = 0;\n-\tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n-\tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n-\n-\tIVEC_INIT(occ);\n-\tivec_zero(&occ, xe->mph_size);\n-\n-\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n-\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n-\t\tocc.ptr[mph1].file1 += 1;\n-\t}\n-\n-\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n-\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n-\t\tocc.ptr[mph2].file2 += 1;\n-\t}\n-\n-\t/*\n-\t * Create temporary arrays that will help us decide if\n-\t * changed[i] should remain false, or become true.\n-\t */\n-\tif (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n-\t\tret = -1;\n-\t\tgoto cleanup;\n-\t}\n-\tif (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n-\t\tret = -1;\n-\t\tgoto cleanup;\n-\t}\n-\n-\t/*\n-\t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n-\t */\n-\tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n-\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n-\t}\n-\n-\tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n-\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n-\t}\n-\n-\t/*\n-\t * Use temporary arrays to decide if changed[i] should remain\n-\t * false, or become true.\n-\t */\n-\txe->xdf1.nreff = 0;\n-\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n-\t     i <= dend1; i++, recs++) {\n-\t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, dend1))) {\n-\t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n-\t\t\txe->xdf1.changed[i] = true;\n-\t\t\t/* i.e. discard */\n-\t}\n-\n-\txe->xdf2.nreff = 0;\n-\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n-\t     i <= dend2; i++, recs++) {\n-\t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, dend2))) {\n-\t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n-\t\t\txe->xdf2.changed[i] = true;\n-\t\t\t/* i.e. discard */\n-\t}\n-\n-cleanup:\n-\txdl_free(action1);\n-\txdl_free(action2);\n-\tivec_free(&occ);\n-\n-\treturn ret;\n-}\n-\n-\n /*\n  * Early trim initial and terminal matching records.\n  */\n@@ -414,19 +235,9 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \t}\n \n \txe->mph_size = cf.count;\n+\txdl_free_classifier(&cf);\n \n \txdl_trim_ends(xe);\n-\tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n-\t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n-\t    xdl_cleanup_records(xe, xpp->flags) < 0) {\n-\n-\t\txdl_free_ctx(&xe->xdf2);\n-\t\txdl_free_ctx(&xe->xdf1);\n-\t\txdl_free_classifier(&cf);\n-\t\treturn -1;\n-\t}\n-\n-\txdl_free_classifier(&cf);\n \n \treturn 0;\n }\n-- \ngitgitgadget\n"},{"id":"532967","messageId":"xmqq5x9ip05a.fsf@gitster.g","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"Re: [PATCH 00/10] Xdiff cleanup part 3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-01-04T02:44:17Z","receivedAt":"2026-01-04T02:44:19Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n>  compat/ivec.c      | 113 ++++++++++++++++++\n>  compat/ivec.h      |  52 +++++++++\n\nI very much like the general direction, but I wonder if we expect\nmany more \"rust-to-C interface layer\" files to come, which I suspect\nis generally true, and in which case I think it is a good idea to\nrethink the use of \"compat/\" for this purpose from early days, as\n\"compat/\" is not about \"compat between C and something else\", but is\nabout \"compat between platform peculiarity and (idealized) POSIX\nenvironment our code assumes\".\n\n"},{"id":"532975","messageId":"xmqq4ip2ndse.fsf@gitster.g","threadId":"64712","inReplyTo":"adf1395d201e916f23accc7644d21aff4f58368b.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-01-04T05:32:33Z","receivedAt":"2026-01-04T05:32:35Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +\tif (new_capacity == 0) {\n> +\t\tfree(self->ptr);\n> +\t\tself->ptr = NULL;\n\n\tif (!new_capacity)\n\t\tFREE_AND_NULL(self->ptr);\n\telse\n\t\t...;\n\n> +void ivec_free(void *self_)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tfree(self->ptr);\n> +\tself->ptr = NULL;\n\nLikewise.  Otherwise the code will fail coccicheck.\n\n> +\tself->length = 0;\n> +\tself->capacity = 0;\n> +\t// DO NOT MODIFY element_size!!!\n\n\t/* A single-liner comment in our codebase looks like this */\n\n"},{"id":"532976","messageId":"CAHTeOx_saiv_ftwS9fo8jLJS6VZyWufNzX4Rzbgaa8NmRJS8EQ@mail.gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"Re: [PATCH 00/10] Xdiff cleanup part 3","fromName":"Yee Cheng Chin","fromEmail":"ychin.git@gmail.com","sentAt":"2026-01-04T06:01:45Z","receivedAt":"2026-01-04T06:02:23Z","isPatch":true,"body":"Hi Ezekiel, I wonder if you saw my proposed patch \"xdiff: fix outdated\nxpatience comments referring to \"ha\" member var\"?\n(https://lore.kernel.org/pull.2139.git.git.1766464905719.gitgitgadget@gmail.com)\nfrom 2 weeks ago? It simply cleans up a stale comment after a previous\nxdiff cleanup when the \"ha\" member variable was split. I don't think\nit conflicts with this part 3 (it's a small comments clean up) but I\nwonder if you could take a look? Just to avoid future conflicts.\n\nOn Fri, Jan 2, 2026 at 10:52 AM Ezekiel Newren via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> Patch series summary:\n>\n>  * patch 1: Introduce the ivec type\n>  * patch 2: Create the function xdl_do_classic_diff()\n>  * patches 3-4: generic cleanup\n>  * patches 5-8: convert from dstart/dend (in xdfile_t) to\n>    delta_start/delta_end (in xdfenv_t)\n>  * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n>    xdiffi.c\n>\n> Things that will be addressed in future patch series:\n>\n>  * Make xdl_cleanup_records() easier to read\n>  * convert recs/nrec into an ivec\n>  * convert changed to an ivec\n>  * remove reference_index/nreff from xdfile_t and turn it into an ivec\n>  * splitting minimal_perfect_hash out as its own ivec\n>  * improve the performance of the classifier and parsing/hashing lines\n>\n> === before this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\n> size_t nreff; } xdfile_t;\n>\n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n>\n> === after this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\n> xdfile_t;\n>\n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\n> delta_end; size_t mph_size; } xdfenv_t;\n>\n> Ezekiel Newren (10):\n>   ivec: introduce the C side of ivec\n>   xdiff: make classic diff explicit by creating xdl_do_classic_diff()\n>   xdiff: don't waste time guessing the number of lines\n>   xdiff: let patience and histogram benefit from xdl_trim_ends()\n>   xdiff: use xdfenv_t in xdl_trim_ends() and xdl_cleanup_records()\n>   xdiff: cleanup xdl_trim_ends()\n>   xdiff: replace xdfile_t.dstart with xdfenv_t.delta_start\n>   xdiff: replace xdfile_t.dend with xdfenv_t.delta_end\n>   xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()\n>   xdiff: move xdl_cleanup_records() from xprepare.c to xdiffi.c\n>\n>  Makefile           |   1 +\n>  compat/ivec.c      | 113 ++++++++++++++++++\n>  compat/ivec.h      |  52 +++++++++\n>  meson.build        |   1 +\n>  xdiff/xdiffi.c     | 221 +++++++++++++++++++++++++++++++++---\n>  xdiff/xdiffi.h     |   1 +\n>  xdiff/xhistogram.c |   7 +-\n>  xdiff/xpatience.c  |   7 +-\n>  xdiff/xprepare.c   | 277 ++++++++-------------------------------------\n>  xdiff/xtypes.h     |   3 +-\n>  xdiff/xutils.c     |  20 ----\n>  xdiff/xutils.h     |   1 -\n>  12 files changed, 432 insertions(+), 272 deletions(-)\n>  create mode 100644 compat/ivec.c\n>  create mode 100644 compat/ivec.h\n>\n>\n> base-commit: 66ce5f8e8872f0183bb137911c52b07f1f242d13\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v1\n> Pull-Request: https://github.com/git/git/pull/2156\n> --\n> gitgitgadget\n>\n"},{"id":"533284","messageId":"0437b899-5a36-4499-a30a-c2a074a80f7e@gmail.com","threadId":"64712","inReplyTo":"adf1395d201e916f23accc7644d21aff4f58368b.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-08T14:34:48Z","receivedAt":"2026-01-08T14:34:53Z","isPatch":true,"body":"Hi Ezekiel\n\nOn 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Trying to use Rust's Vec in C, or git's ALLOC_GROW() macros (via\n> wrapper functions) in Rust is painful because:\n> \n>    * C doesn't define its own vector type, and even though Rust does\n>      have Vec its painful to use on the C side (more on that below).\n>      However its still not viable to use Rust's Vec type because Git\n>      needs to be able to compile without Rust. So ivec was created\n>      expressley to be interoperable between C and Rust without needing\n>      Rust.\n>    * C doing vector things the Rust way would require wrapper functions,\n>      and Rust doing vector things the C way would require wrapper\n>      functions, so ivec was created to ensure a consistent contract\n>      between the 2 languages for how to manipulate a vector.\n>    * Currently, Rust defines its own 'Vec' type that is generic, but its\n>      memory allocator and struct layout weren't designed for\n>      interoperability with C (or any language for that matter), meaning\n>      that the C side cannot push to or expand a 'Vec' without defining\n>      wrapper functions in Rust that C can call. Without special care,\n>      the two languages might use different allocators (malloc/free on\n>      the C side, and possibly something else in Rust), which would make\n>      it difficult for a function in one language to free elements\n>      allocated by a call from a function in the other language.\n>    * Similarly, git defines ALLOC_GROW() and related macros in\n>      git-compat-util.h. While we could add functions allowing Rust to\n>      invoke something similar to those macros, passing three variables\n>      (pointer, length, allocated_size) instead of a single variable\n>      (vector) across the language boundary requires more cognitive\n>      overhead for readers to keep track of and makes it easier to make\n>      mistakes. Further, for low-level components that we want to\n>      eventually convert to pure Rust, such triplets would feel very out\n>      of place.\n> \n> To address these issue, introduce a new type, ivec -- short for\n> interoperable vector. (We refer to it as 'ivec' generally, though on\n> the Rust side the struct is called IVec to match Rust style.)  This new\n> type is specifically designed for FFI purposes, so that both languages\n> handle the vector in the same way, though it could be used on either\n> side independently. This type is designed such that it can easily be\n> replaced by a Rust 'Vec' once interoperability is no longer a concern.\n> \n> One particular item to note is that Git's macros to handle vec\n> operations infer the amount that a vec needs to grow from the size of\n> a pointer, but that makes it somewhat specific to the macros used in C.\n> To avoid defining every ivec function as a macro I opted to also\n> include an element_size field that allows concrete functions like\n> push() to know how much to grow the memory. This element_size also\n> helps in verifying that the ivec is correct when passing from C to\n> Rust.\n\nI've left some comments below but I think this is a sensible direction.\n\n> diff --git a/compat/ivec.c b/compat/ivec.c\n> new file mode 100644\n> index 0000000000..0a777e78dc\n> --- /dev/null\n> +++ b/compat/ivec.c\n> @@ -0,0 +1,113 @@\n> +#include \"ivec.h\"\n> +\n> +struct IVec_c_void {\n\nWe normally use all lower case names for structs but as this is shared \nwith rust it maybe makes sense to use CamelCase so the names are the \nsame in both languages.\n\n> +\tvoid *ptr;\n> +\tsize_t length;\n> +\tsize_t capacity;\n> +\tsize_t element_size;\n> +};\n> +\n> +static void _set_capacity(void *self_, size_t new_capacity)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n\nPassing any of the ivec variants defined below to this function invokes \nundefined behavior because we're not casting the pointer back to the \norginal type. However I think on the platforms we care about \nsizeof(void*) == sizeof(T*) for all T so maybe we can look the other way.\n\n> +\n> +\tif (new_capacity == self->capacity) {\n> +\t\treturn;\n> +\t}\n> +\tif (new_capacity == 0) {\n> +\t\tfree(self->ptr);\n> +\t\tself->ptr = NULL;\n> +\t} else {\n> +\t\tself->ptr = realloc(self->ptr, new_capacity * self->element_size);\n> +\t}\n> +\tself->capacity = new_capacity;\n\nNot if realloc() returns NULL. We should check for that, probably by \nusing xrealloc().\n\n> +void ivec_zero(void *self_, size_t capacity)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tself->ptr = calloc(capacity, self->element_size);\n\nWe should be handling allocation failures here probably by using xcalloc().\n\n> +void ivec_reserve(void *self_, size_t additional)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tsize_t growby = 128;\n> +\tif (self->capacity > growby)\n> +\t\tgrowby = self->capacity;\n> +\tif (additional > growby)\n> +\t\tgrowby = additional;\n\nThis growth strategy differs from both ALLOC_GROW() and \nXDL_ALLOC_GROW(), if there isn't a good reason for that we should \nperhaps just use ALLOC_GROW() here.\n\n> +void ivec_push(void *self_, const void *value)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\tvoid *dst = NULL;\n> +\n> +\tif (self->length == self->capacity)\n> +\t\tivec_reserve(self, 1);\n> +\n> +\tdst = (uint8_t*)self->ptr + self->length * self->element_size;\n> +\tmemcpy(dst, value, self->element_size);\n\nIf self->element_size was a compile time constant the compiler could \neasily optimize this call away. I'm not sure that is easy to achieve though.\n\n> +\tself->length++;\n> +}\n> +\n> +void ivec_free(void *self_)\n\nNormally we'd call a like this that free the allocations and \nre-initializes the members ivec_clear()\n\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tfree(self->ptr);\n> +\tself->ptr = NULL;\n> +\tself->length = 0;\n> +\tself->capacity = 0;\n> +\t// DO NOT MODIFY element_size!!!\n> +}\n> +\n> +void ivec_move(void *src_, void *dst_)\n> +{\n> +\tstruct IVec_c_void *src = src_;\n> +\tstruct IVec_c_void *dst = dst_;\n\nMaybe we should add\n\n\tif (src->element_size != dst->element_size)\n\t\tBUG(\"moving incompatible arrays\");\n> +\n> +\tivec_free(dst);\n> +\tdst->ptr = src->ptr;\n> +\tdst->length = src->length;\n> +\tdst->capacity = src->capacity;\n> +\t// DO NOT MODIFY element_size!!!\n\nAs the element sizes must match maybe *dst = *src would be clearer?\n\n> +\n> +\tsrc->ptr = NULL;\n> +\tsrc->length = 0;\n> +\tsrc->capacity = 0;\n> +\t// DO NOT MODIFY element_size!!!\n> +}\n> diff --git a/compat/ivec.h b/compat/ivec.h\n> new file mode 100644\n> index 0000000000..654a05c506\n> --- /dev/null\n> +++ b/compat/ivec.h\n> @@ -0,0 +1,52 @@\n> +#ifndef IVEC_H\n> +#define IVEC_H\n> +\n> +#include <git-compat-util.h>\n\nIt would be nice to have some documentation in this header, see the \nexamples in strvec.h and hashmap.h\n\n> +#define IVEC_INIT(variable) ivec_init(&(variable), sizeof(*(variable).ptr))\n\nThis is a bit cumbersome to use compared to our usual *_INIT macros. I'm \nstruggling to see how we can make it nicer though as DEFINE_IVEC_TYPE \ncannot define a per-type initializer macro and I we cannot initialize \nthe element size without knowing the type.\n\n> +\n> +#ifndef CBINDGEN\n> +#define DEFINE_IVEC_TYPE(type, suffix) \\\n> +struct IVec_##suffix { \\\n> +\ttype* ptr; \\\n> +\tsize_t length; \\\n> +\tsize_t capacity; \\\n> +\tsize_t element_size; \\\n> +}\n\nI wonder if we want to define type safe inline safe wrappers for the \nivec_* functions here. I think the only functions where the element type \nmatters are ivec_move() and ivec_push(), for the others like \nivec_zero(), ivec_reserve() and ivec_free() the element type does not \nmatter. ivec_push() would certainly be easier to use with a wrapper as \nmeans we can avoid forcing the caller to take the address of the value.\n\nstatic inline ivec_##suffix##_push(struct IVec_##suffix *self, type \nvalue) { \\\n\tconst void *ptr = &value; \\\n\tivec_push(self, ptr); \\\n}\n\nI'll try and take a look at the rest of this series next week\n\nThanks\n\nPhillip\n\n> +\n> +DEFINE_IVEC_TYPE(bool, bool);\n> +\n> +DEFINE_IVEC_TYPE(uint8_t, u8);\n> +DEFINE_IVEC_TYPE(uint16_t, u16);\n> +DEFINE_IVEC_TYPE(uint32_t, u32);\n> +DEFINE_IVEC_TYPE(uint64_t, u64);\n> +\n> +DEFINE_IVEC_TYPE(int8_t, i8);\n> +DEFINE_IVEC_TYPE(int16_t, i16);\n> +DEFINE_IVEC_TYPE(int32_t, i32);\n> +DEFINE_IVEC_TYPE(int64_t, i64);\n> +\n> +DEFINE_IVEC_TYPE(float, f32);\n> +DEFINE_IVEC_TYPE(double, f64);\n> +\n> +DEFINE_IVEC_TYPE(size_t, usize);\n> +DEFINE_IVEC_TYPE(ssize_t, isize);\n> +#endif\n> +\n> +void ivec_init(void *self_, size_t element_size);\n> +\n> +void ivec_zero(void *self_, size_t capacity);\n> +\n> +void ivec_reserve_exact(void *self_, size_t additional);\n> +\n> +void ivec_reserve(void *self_, size_t additional);\n> +\n> +void ivec_shrink_to_fit(void *self_);\n> +\n> +void ivec_push(void *self_, const void *value);\n> +\n> +void ivec_free(void *self_);\n> +\n> +void ivec_move(void *src, void *dst);\n> +\n> +#endif /* IVEC_H */\n> diff --git a/meson.build b/meson.build\n> index dd52efd1c8..42ac0c8c42 100644\n> --- a/meson.build\n> +++ b/meson.build\n> @@ -302,6 +302,7 @@ libgit_sources = [\n>     'commit.c',\n>     'common-exit.c',\n>     'common-init.c',\n> +  'compat/ivec.c',\n>     'compat/nonblock.c',\n>     'compat/obstack.c',\n>     'compat/open.c',\n\n"},{"id":"533966","messageId":"CAH=ZcbA_HgEO2T2smn4Yg6gf4sm4jrR8A0ek1v9nqsa1MXbRJw@mail.gmail.com","threadId":"64712","inReplyTo":"0437b899-5a36-4499-a30a-c2a074a80f7e@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-15T15:55:58Z","receivedAt":"2026-01-15T15:56:11Z","isPatch":true,"body":"On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n> > diff --git a/compat/ivec.c b/compat/ivec.c\n> > new file mode 100644\n> > index 0000000000..0a777e78dc\n> > --- /dev/null\n> > +++ b/compat/ivec.c\n> > @@ -0,0 +1,113 @@\n> > +#include \"ivec.h\"\n> > +\n> > +struct IVec_c_void {\n>\n> We normally use all lower case names for structs but as this is shared\n> with rust it maybe makes sense to use CamelCase so the names are the\n> same in both languages.\n\nMy preference would be all lowercase, but cbindgen insists on using\nthe same casing as was used in Rust. I don't think there's a way to\nmake cbindgen use all lowercase for structs.\n\n> > +     void *ptr;\n> > +     size_t length;\n> > +     size_t capacity;\n> > +     size_t element_size;\n> > +};\n> > +\n> > +static void _set_capacity(void *self_, size_t new_capacity)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n>\n> Passing any of the ivec variants defined below to this function invokes\n> undefined behavior because we're not casting the pointer back to the\n> orginal type. However I think on the platforms we care about\n> sizeof(void*) == sizeof(T*) for all T so maybe we can look the other way.\n\nIf someone finds that this code does not work because of this\nassumption I'd like to know. But I can't fathom a case where it\nwouldn't work.\n\n> > +\n> > +     if (new_capacity == self->capacity) {\n> > +             return;\n> > +     }\n> > +     if (new_capacity == 0) {\n> > +             free(self->ptr);\n> > +             self->ptr = NULL;\n> > +     } else {\n> > +             self->ptr = realloc(self->ptr, new_capacity * self->element_size);\n> > +     }\n> > +     self->capacity = new_capacity;\n>\n> Not if realloc() returns NULL. We should check for that, probably by\n> using xrealloc().\n>\n> > +void ivec_zero(void *self_, size_t capacity)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     self->ptr = calloc(capacity, self->element_size);\n>\n> We should be handling allocation failures here probably by using xcalloc().\n\nI've changed it to xrealloc() similar for the calloc() call.\n\n\n> > +void ivec_reserve(void *self_, size_t additional)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     size_t growby = 128;\n> > +     if (self->capacity > growby)\n> > +             growby = self->capacity;\n> > +     if (additional > growby)\n> > +             growby = additional;\n>\n> This growth strategy differs from both ALLOC_GROW() and\n> XDL_ALLOC_GROW(), if there isn't a good reason for that we should\n> perhaps just use ALLOC_GROW() here.\n\nXDL_ALLOW_GROW() can't be used because the pointer is always a void*\nin this function.\n\n> > +void ivec_push(void *self_, const void *value)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +     void *dst = NULL;\n> > +\n> > +     if (self->length == self->capacity)\n> > +             ivec_reserve(self, 1);\n> > +\n> > +     dst = (uint8_t*)self->ptr + self->length * self->element_size;\n> > +     memcpy(dst, value, self->element_size);\n>\n> If self->element_size was a compile time constant the compiler could\n> easily optimize this call away. I'm not sure that is easy to achieve though.\n\nThe problem is that I didn't want all of ivec to be macros that looked\nlike function calls. I wanted to minimize use of macros so that it was\neasier to port and verify that the Rust implementation matches the\nbehavior of the C implementation.\n\n> > +void ivec_free(void *self_)\n>\n> Normally we'd call a like this that free the allocations and\n> re-initializes the members ivec_clear()\n\nIn Rust Vec.clear() means to set length to zero, but leaves the\nallocation alone. The reason why I'm zeroing the struct is to help\navoid FFI issues. If not zero then what should the members be set to,\nto indicate that using the struct is not valid anymore? In Rust an\nobject is freed when it goes out of scope and _cannot_ be accessed\nafterward.\n\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     free(self->ptr);\n> > +     self->ptr = NULL;\n> > +     self->length = 0;\n> > +     self->capacity = 0;\n> > +     // DO NOT MODIFY element_size!!!\n> > +}\n> > +\n> > +void ivec_move(void *src_, void *dst_)\n> > +{\n> > +     struct IVec_c_void *src = src_;\n> > +     struct IVec_c_void *dst = dst_;\n>\n> Maybe we should add\n>\n>         if (src->element_size != dst->element_size)\n>                 BUG(\"moving incompatible arrays\");\n\nI'll do that.\n\n> > +\n> > +     ivec_free(dst);\n> > +     dst->ptr = src->ptr;\n> > +     dst->length = src->length;\n> > +     dst->capacity = src->capacity;\n> > +     // DO NOT MODIFY element_size!!!\n>\n> As the element sizes must match maybe *dst = *src would be clearer?\n\nThat seems fine.\n\n> > +\n> > +     src->ptr = NULL;\n> > +     src->length = 0;\n> > +     src->capacity = 0;\n> > +     // DO NOT MODIFY element_size!!!\n> > +}\n> > diff --git a/compat/ivec.h b/compat/ivec.h\n> > new file mode 100644\n> > index 0000000000..654a05c506\n> > --- /dev/null\n> > +++ b/compat/ivec.h\n> > @@ -0,0 +1,52 @@\n> > +#ifndef IVEC_H\n> > +#define IVEC_H\n> > +\n> > +#include <git-compat-util.h>\n>\n> It would be nice to have some documentation in this header, see the\n> examples in strvec.h and hashmap.h\n>\n> > +#define IVEC_INIT(variable) ivec_init(&(variable), sizeof(*(variable).ptr))\n>\n> This is a bit cumbersome to use compared to our usual *_INIT macros. I'm\n> struggling to see how we can make it nicer though as DEFINE_IVEC_TYPE\n> cannot define a per-type initializer macro and I we cannot initialize\n> the element size without knowing the type.\n\nI don't see what's cumbersome about it. Maybe an example use case\nwould clarify things.\n\n```\nDEFINE_IVEC_TYPE(xrecord_t, xrecord);\n\nvoid some_function() {\n    struct IVec_xrecord rec;\n    IVEC_INIT(rec);  // i.e. ivec_init(&rec, sizeof(*rec.ptr);\n\n    // use concrete functions to manipulate vector or access the array\ndirectly via ptr\n}\n```\n\nIVEC_INIT() should be used on the concrete type.\n\n> > +\n> > +#ifndef CBINDGEN\n> > +#define DEFINE_IVEC_TYPE(type, suffix) \\\n> > +struct IVec_##suffix { \\\n> > +     type* ptr; \\\n> > +     size_t length; \\\n> > +     size_t capacity; \\\n> > +     size_t element_size; \\\n> > +}\n>\n> I wonder if we want to define type safe inline safe wrappers for the\n> ivec_* functions here. I think the only functions where the element type\n> matters are ivec_move() and ivec_push(), for the others like\n> ivec_zero(), ivec_reserve() and ivec_free() the element type does not\n> matter. ivec_push() would certainly be easier to use with a wrapper as\n> means we can avoid forcing the caller to take the address of the value.\n>\n> static inline ivec_##suffix##_push(struct IVec_##suffix *self, type\n> value) { \\\n>         const void *ptr = &value; \\\n>         ivec_push(self, ptr); \\\n> }\n\nI turned ivec_push() into a macro, but the rest will remain as\nconcrete functions.\n"},{"id":"534025","messageId":"c2d9a432-0753-4786-8de9-c3dcfe69ac36@gmail.com","threadId":"64712","inReplyTo":"CAH=ZcbA_HgEO2T2smn4Yg6gf4sm4jrR8A0ek1v9nqsa1MXbRJw@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-16T10:39:28Z","receivedAt":"2026-01-16T10:39:31Z","isPatch":true,"body":"I've Cc'd Peff and René for a second opinion if you have time please.\n\nOn 15/01/2026 15:55, Ezekiel Newren wrote:\n> On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n >\n>>> +static void _set_capacity(void *self_, size_t new_capacity)\n>>> +{\n>>> +     struct IVec_c_void *self = self_;\n>>\n>> Passing any of the ivec variants defined below to this function invokes\n>> undefined behavior because we're not casting the pointer back to the\n>> orginal type. However I think on the platforms we care about\n>> sizeof(void*) == sizeof(T*) for all T so maybe we can look the other way.\n> \n> If someone finds that this code does not work because of this\n> assumption I'd like to know. But I can't fathom a case where it\n> wouldn't work.\n\nSo we have two different structs\n\nstruct IVec_c_void {\n\tvoid *ptr;\n\tsize_t length;\n\tsize_t capacity;\n\tsize_t element_size;\n}\n\nand\n\nstruct Ivec_u8 {\n\tuint8_t *ptr;\n\tsize_t length;\n\tsize_t capacity;\n\tsize_t element_size;\n}\n\nOne the platforms we care about they will have the same memory layout as \nall pointers have the same representation. However I don't think they \nare \"compatible types\" in the language of the C standard because the \ntype of the \"ptr\" member differs. That means casting IVec_u8* to \nIVec_c_void* either directly or via void* is undefined and so\n\n\tstruct IVec_u8 vec;\n\tivec_init(&vec, sizeof(*vec.ptr));\n\nis undefined. For the compiler to see the undefined cast it needs to \nlook across translation units because the implementation of ivec_init() \nwill be in a separate file to where it is called. Maybe that and the \nfact they have the same memory layout saves us from having to worry too \nmuch though I'm always nervous of undefined behavior.\n\nAn alternative would be to pass the individual struct members as \nfunction parameters\n\n\tvoid ivec_init(void **vec, size_t &length, size_t &capacity,\n\t\t       size_t &element_size_, size_t element_size)\n\t{\n\t\t*vec = NULL;\n\t\t*length = 0;\n\t\t*capacity = 0;\n\t\t*element_size_ = element_size;\n\t}\n\nand have DEFINE_IVEC_TYPE create typesafe wrappers\n\n\tstatic inline void ivec_u8_init(struct IVec_u8 *vec)\n\t{\n\t\tvoid *ptr = vec->ptr;\n\t\tivec_init(&ptr, &v->length, &v->capacity,\n\t\t\t  &v->element_size, sizeof(*(v->ptr));\n\t\tvec->ptr = ptr;\n\t}\n\nThat's safe because we cast the \"ptr\" member to \"void*\" and then back to \nthe original type. On the rust side the implementation of IVec<T> would \nalso need to split out the individual struct members when it calls \nivec_init() etc. It's all a bit more effort but the benefit is that we \ndon't have any undefined behavior and we have a nice typesafe C \ninterface to 'struct IVec_*'.\n\nThanks\n\nPhillip\n\n"},{"id":"534082","messageId":"0a306227-5db8-4d12-865c-fa0efe5c6beb@web.de","threadId":"64712","inReplyTo":"adf1395d201e916f23accc7644d21aff4f58368b.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2026-01-16T20:19:02Z","receivedAt":"2026-01-16T20:19:04Z","isPatch":true,"body":"On 1/2/26 7:52 PM, Ezekiel Newren via GitGitGadget wrote:\n> diff --git a/compat/ivec.c b/compat/ivec.c\n> new file mode 100644\n> index 0000000000..0a777e78dc\n> --- /dev/null\n> +++ b/compat/ivec.c\n> @@ -0,0 +1,113 @@\n> +#include \"ivec.h\"\n> +\n> +struct IVec_c_void {\n> +\tvoid *ptr;\n> +\tsize_t length;\n> +\tsize_t capacity;\n> +\tsize_t element_size;\n> +};\n> +\n> +static void _set_capacity(void *self_, size_t new_capacity)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tif (new_capacity == self->capacity) {\n> +\t\treturn;\n> +\t}\n> +\tif (new_capacity == 0) {\n> +\t\tfree(self->ptr);\n> +\t\tself->ptr = NULL;\n> +\t} else {\n> +\t\tself->ptr = realloc(self->ptr, new_capacity * self->element_size);\n> +\t}\n> +\tself->capacity = new_capacity;\n> +}\n> +\n> +\n> +void ivec_init(void *self_, size_t element_size)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tself->ptr = NULL;\n> +\tself->length = 0;\n> +\tself->capacity = 0;\n> +\tself->element_size = element_size;\n> +}\n> +\n> +void ivec_zero(void *self_, size_t capacity)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tself->ptr = calloc(capacity, self->element_size);\n> +\tself->length = capacity;\n> +\tself->capacity = capacity;\n> +\t// DO NOT MODIFY element_size!!!\n> +}\n> +\n> +void ivec_reserve_exact(void *self_, size_t additional)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\t_set_capacity(self, self->capacity + additional);\n> +}\n> +\n> +void ivec_reserve(void *self_, size_t additional)\n> +{\n> +\tstruct IVec_c_void *self = self_;\n> +\n> +\tsize_t growby = 128;\n> +\tif (self->capacity > growby)\n> +\t\tgrowby = self->capacity;\n> +\tif (additional > growby)\n> +\t\tgrowby = additional;\n> +\n> +\t_set_capacity(self, self->capacity + growby);\n> +}\n\nConstant growth steps like these cause linear growth and quadratic\ncomplexity.  ALLOC_GROW does exponential growth with factor 1.5 to\nget linear complexity.  Here's an old plea to do the same:\nhttps://blog.mozilla.org/nnethercote/2014/11/04/please-grow-your-buffers-exponentially/\n\nRené\n\n"},{"id":"534083","messageId":"fc291b3a-5ee5-4488-9b01-d3de32f7c257@web.de","threadId":"64712","inReplyTo":"c2d9a432-0753-4786-8de9-c3dcfe69ac36@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2026-01-16T20:19:01Z","receivedAt":"2026-01-16T20:19:05Z","isPatch":true,"body":"On 1/16/26 11:39 AM, Phillip Wood wrote:\n> I've Cc'd Peff and René for a second opinion if you have time please.\n> \n> On 15/01/2026 15:55, Ezekiel Newren wrote:\n>> On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>\n>>>> +static void _set_capacity(void *self_, size_t new_capacity)\n>>>> +{\n>>>> +     struct IVec_c_void *self = self_;\n>>>\n>>> Passing any of the ivec variants defined below to this function invokes\n>>> undefined behavior because we're not casting the pointer back to the\n>>> orginal type. However I think on the platforms we care about\n>>> sizeof(void*) == sizeof(T*) for all T so maybe we can look the other way.\n>>\n>> If someone finds that this code does not work because of this\n>> assumption I'd like to know. But I can't fathom a case where it\n>> wouldn't work.\n> \n> So we have two different structs\n> \n> struct IVec_c_void {\n>     void *ptr;\n>     size_t length;\n>     size_t capacity;\n>     size_t element_size;\n> }\n> \n> and\n> \n> struct Ivec_u8 {\n>     uint8_t *ptr;\n>     size_t length;\n>     size_t capacity;\n>     size_t element_size;\n> }\n> \n> One the platforms we care about they will have the same memory\n> layout as all pointers have the same representation. However I don't\n> think they are \"compatible types\" in the language of the C standard\n> because the type of the \"ptr\" member differs. That means casting\n> IVec_u8* to IVec_c_void* either directly or via void* is undefined\n> and so\n> \n>     struct IVec_u8 vec;\n>     ivec_init(&vec, sizeof(*vec.ptr));\n> \n> is undefined. For the compiler to see the undefined cast it needs to\n> look across translation units because the implementation of\n> ivec_init() will be in a separate file to where it is called. Maybe\n> that and the fact they have the same memory layout saves us from\n> having to worry too much though I'm always nervous of undefined\n> behavior.\n\nTrue.  The GCC docs give a fun example of what a compiler might do\nwhen using different struct types to access the same memory:\n\nhttps://www.gnu.org/software/c-intro-and-ref/manual/html_node/Aliasing-Type-Rules.html\n\nNot sure it applies to this case, but the point is that compilers\ncan and will do terrifying things when they smell UB, with little\nconcern for safety or original intent.\n\n> An alternative would be to pass the individual struct members as function parameters\n> \n>     void ivec_init(void **vec, size_t &length, size_t &capacity,\n>                size_t &element_size_, size_t element_size)\n>     {\n>         *vec = NULL;\n>         *length = 0;\n>         *capacity = 0;\n>         *element_size_ = element_size;\n>     }\n\nThe ampersands (&) should be asterisks (*), right?\n\n> and have DEFINE_IVEC_TYPE create typesafe wrappers\n> \n>     static inline void ivec_u8_init(struct IVec_u8 *vec)\n>     {\n>         void *ptr = vec->ptr;\n>         ivec_init(&ptr, &v->length, &v->capacity,\n>               &v->element_size, sizeof(*(v->ptr));\n>         vec->ptr = ptr;\n>     }\n\nMixes \"v\" and \"vec\", misses a closing parenthesis.  Looks viable,\nthough, and this method should be applicable to the rest of the\nfunctions as well (on the C side).\n\nI guess this doesn't require an element_size member anymore as\neach wrapper can pass in the sizeof value.\n\n> That's safe because we cast the \"ptr\" member to \"void*\" and then\n> back to the original type. On the rust side the implementation of\n> IVec<T> would also need to split out the individual struct members\n> when it calls ivec_init() etc. It's all a bit more effort but the\n> benefit is that we don't have any undefined behavior and we have a\n> nice typesafe C interface to 'struct IVec_*'.\nRight.  No idea how ugly this would be on the Rust side, though.\n\nRené\n\n"},{"id":"534084","messageId":"07ca298a-ad32-4998-88ff-d69c04418fdd@web.de","threadId":"64712","inReplyTo":"f9b10e71d23f8b4fa34dcffb371cf5a173760409.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 09/10] xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2026-01-16T20:19:06Z","receivedAt":"2026-01-16T20:19:09Z","isPatch":true,"body":"On 1/2/26 7:52 PM, Ezekiel Newren via GitGitGadget wrote:\n> @@ -253,22 +250,44 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n>  \treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n>  }\n>  \n> +struct xoccurrence\n> +{\n> +\tsize_t file1, file2;\n> +};\n> +\n> +\n> +DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n> +\n>  \n>  /*\n>   * Try to reduce the problem complexity, discard records that have no\n>   * matches on the other file. Also, lines that have multiple matches\n>   * might be potentially discarded if they appear in a run of discardable.\n>   */\n> -static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n> -\tlong i, nm, mlim;\n> +static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n> +\tlong i;\n> +\tsize_t nm, mlim;\n>  \txrecord_t *recs;\n> -\txdlclass_t *rcrec;\n>  \tuint8_t *action1 = NULL, *action2 = NULL;\n> -\tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> +\tstruct IVec_xoccurrence occ;\n> +\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n>  \tint ret = 0;\n>  \tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n>  \tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n>  \n> +\tIVEC_INIT(occ);\n> +\tivec_zero(&occ, xe->mph_size);\n\nThis array is presized here.  It is neither grown nor shrunken.\nCALLOC_ARRAY would work just as well, at least at this point, no?\n\n> +\n> +\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n> +\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n> +\t\tocc.ptr[mph1].file1 += 1;\n> +\t}\n> +\n> +\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n> +\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n> +\t\tocc.ptr[mph2].file2 += 1;\n> +\t}\n> +\n>  \t/*\n>  \t * Create temporary arrays that will help us decide if\n>  \t * changed[i] should remain false, or become true.\n> @@ -288,16 +307,14 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>  \tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>  \t\tmlim = XDL_MAX_EQLIMIT;\n>  \tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n> -\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> -\t\tnm = rcrec ? rcrec->len2 : 0;\n> +\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n>  \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>  \t}\n>  \n>  \tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>  \t\tmlim = XDL_MAX_EQLIMIT;\n>  \tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n> -\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> -\t\tnm = rcrec ? rcrec->len1 : 0;\n> +\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n>  \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>  \t}\n>  \n> @@ -332,6 +349,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>  cleanup:\n>  \txdl_free(action1);\n>  \txdl_free(action2);\n> +\tivec_free(&occ);\n>  \n>  \treturn ret;\n>  }\n"},{"id":"534110","messageId":"ed06232c-4c48-4d5a-a269-8663b32787ea@gmail.com","threadId":"64712","inReplyTo":"fc291b3a-5ee5-4488-9b01-d3de32f7c257@web.de","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-17T13:55:46Z","receivedAt":"2026-01-17T13:55:49Z","isPatch":true,"body":"On 16/01/2026 20:19, René Scharfe wrote:\n> On 1/16/26 11:39 AM, Phillip Wood wrote:\n>> I've Cc'd Peff and René for a second opinion if you have time please.\n>>\n>> On 15/01/2026 15:55, Ezekiel Newren wrote:\n>>> On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>>\n>>>>> +static void _set_capacity(void *self_, size_t new_capacity)\n>>>>> +{\n>>>>> +     struct IVec_c_void *self = self_;\n>>>>\n>>>> Passing any of the ivec variants defined below to this function invokes\n>>>> undefined behavior because we're not casting the pointer back to the\n>>>> orginal type. However I think on the platforms we care about\n>>>> sizeof(void*) == sizeof(T*) for all T so maybe we can look the other way.\n>>>\n>>> If someone finds that this code does not work because of this\n>>> assumption I'd like to know. But I can't fathom a case where it\n>>> wouldn't work.\n>>\n>> So we have two different structs\n>>\n>> struct IVec_c_void {\n>>      void *ptr;\n>>      size_t length;\n>>      size_t capacity;\n>>      size_t element_size;\n>> }\n>>\n>> and\n>>\n>> struct Ivec_u8 {\n>>      uint8_t *ptr;\n>>      size_t length;\n>>      size_t capacity;\n>>      size_t element_size;\n>> }\n>>\n>> One the platforms we care about they will have the same memory\n>> layout as all pointers have the same representation. However I don't\n>> think they are \"compatible types\" in the language of the C standard\n>> because the type of the \"ptr\" member differs. That means casting\n>> IVec_u8* to IVec_c_void* either directly or via void* is undefined\n>> and so\n>>\n>>      struct IVec_u8 vec;\n>>      ivec_init(&vec, sizeof(*vec.ptr));\n>>\n>> is undefined. For the compiler to see the undefined cast it needs to\n>> look across translation units because the implementation of\n>> ivec_init() will be in a separate file to where it is called. Maybe\n>> that and the fact they have the same memory layout saves us from\n>> having to worry too much though I'm always nervous of undefined\n>> behavior.\n> \n> True.  The GCC docs give a fun example of what a compiler might do\n> when using different struct types to access the same memory:\n> \n> https://www.gnu.org/software/c-intro-and-ref/manual/html_node/Aliasing-Type-Rules.html\n\nThanks for the link\n\n> Not sure it applies to this case, but the point is that compilers\n> can and will do terrifying things when they smell UB, with little\n> concern for safety or original intent.\n> \n>> An alternative would be to pass the individual struct members as function parameters\n>>\n>>      void ivec_init(void **vec, size_t &length, size_t &capacity,\n>>                 size_t &element_size_, size_t element_size)\n>>      {\n>>          *vec = NULL;\n>>          *length = 0;\n>>          *capacity = 0;\n>>          *element_size_ = element_size;\n>>      }\n> \n> The ampersands (&) should be asterisks (*), right?\n\nIndeed, that's embarrassing - I must have been thinking of the caller.\n\n>> and have DEFINE_IVEC_TYPE create typesafe wrappers\n>>\n>>      static inline void ivec_u8_init(struct IVec_u8 *vec)\n>>      {\n>>          void *ptr = vec->ptr;\n>>          ivec_init(&ptr, &v->length, &v->capacity,\n>>                &v->element_size, sizeof(*(v->ptr));\n>>          vec->ptr = ptr;\n>>      }\n> \n> Mixes \"v\" and \"vec\", misses a closing parenthesis.  Looks viable,\n> though, and this method should be applicable to the rest of the\n> functions as well (on the C side).\n> \n> I guess this doesn't require an element_size member anymore as\n> each wrapper can pass in the sizeof value.\n\nGood point\n\n>> That's safe because we cast the \"ptr\" member to \"void*\" and then\n>> back to the original type. On the rust side the implementation of\n>> IVec<T> would also need to split out the individual struct members\n>> when it calls ivec_init() etc. It's all a bit more effort but the\n>> benefit is that we don't have any undefined behavior and we have a\n>> nice typesafe C interface to 'struct IVec_*'.\n> Right.  No idea how ugly this would be on the Rust side, though.\n\nI'm hoping it's not too bad and `impl IVec<T>` just contains the \nequivalent of the wrappers generated by DEFINE_IVEC_TYPE()\n\nThanks\n\nPhillip\n> \n> René\n> \n\n"},{"id":"534114","messageId":"CAH=ZcbB=Yf=wn2O273adrvpUpE0bJGKwrAjOAjmB8AgJrjz5Bg@mail.gmail.com","threadId":"64712","inReplyTo":"0a306227-5db8-4d12-865c-fa0efe5c6beb@web.de","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-17T15:58:13Z","receivedAt":"2026-01-17T15:58:26Z","isPatch":true,"body":"On Fri, Jan 16, 2026 at 1:19 PM René Scharfe <l.s.r@web.de> wrote:\n>\n> On 1/2/26 7:52 PM, Ezekiel Newren via GitGitGadget wrote:\n> > diff --git a/compat/ivec.c b/compat/ivec.c\n> > new file mode 100644\n> > index 0000000000..0a777e78dc\n> > --- /dev/null\n> > +++ b/compat/ivec.c\n> > @@ -0,0 +1,113 @@\n> > +#include \"ivec.h\"\n> > +\n> > +struct IVec_c_void {\n> > +     void *ptr;\n> > +     size_t length;\n> > +     size_t capacity;\n> > +     size_t element_size;\n> > +};\n> > +\n> > +static void _set_capacity(void *self_, size_t new_capacity)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     if (new_capacity == self->capacity) {\n> > +             return;\n> > +     }\n> > +     if (new_capacity == 0) {\n> > +             free(self->ptr);\n> > +             self->ptr = NULL;\n> > +     } else {\n> > +             self->ptr = realloc(self->ptr, new_capacity * self->element_size);\n> > +     }\n> > +     self->capacity = new_capacity;\n> > +}\n> > +\n> > +\n> > +void ivec_init(void *self_, size_t element_size)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     self->ptr = NULL;\n> > +     self->length = 0;\n> > +     self->capacity = 0;\n> > +     self->element_size = element_size;\n> > +}\n> > +\n> > +void ivec_zero(void *self_, size_t capacity)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     self->ptr = calloc(capacity, self->element_size);\n> > +     self->length = capacity;\n> > +     self->capacity = capacity;\n> > +     // DO NOT MODIFY element_size!!!\n> > +}\n> > +\n> > +void ivec_reserve_exact(void *self_, size_t additional)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     _set_capacity(self, self->capacity + additional);\n> > +}\n> > +\n> > +void ivec_reserve(void *self_, size_t additional)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     size_t growby = 128;\n> > +     if (self->capacity > growby)\n> > +             growby = self->capacity;\n> > +     if (additional > growby)\n> > +             growby = additional;\n> > +\n> > +     _set_capacity(self, self->capacity + growby);\n> > +}\n>\n> Constant growth steps like these cause linear growth and quadratic\n> complexity.  ALLOC_GROW does exponential growth with factor 1.5 to\n> get linear complexity.  Here's an old plea to do the same:\n> https://blog.mozilla.org/nnethercote/2014/11/04/please-grow-your-buffers-exponentially/\n>\n> René\n\nIt _is_ exponential. ivec_reserve(&vec, 1) means grow by _at least_ 1.\nI'm not using typical memory management as defined in\ngit-compat-util.h because I'm trying to get ivec to behave very\nsimilarly to Rust's Vec so that when Rust is introduced into the code,\nC programmers will already be familiar with how Vec operates _and_ so\nthat converting from IVec to Vec is as simple as refactoring IVec\ndeclarations to Vec.\n\nSince C does not support generics there is no _proper_ solution. What\nI have come up with on the C side for ivec is my best effort\ncompromise.\n"},{"id":"534115","messageId":"CAH=ZcbCuY22WCqzyK-=Adw924a6ZJqnMYjWK9fxwoFn5xK9q-w@mail.gmail.com","threadId":"64712","inReplyTo":"ed06232c-4c48-4d5a-a269-8663b32787ea@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-17T16:04:22Z","receivedAt":"2026-01-17T16:04:35Z","isPatch":true,"body":"On Sat, Jan 17, 2026 at 6:55 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> On 16/01/2026 20:19, René Scharfe wrote:\n> > On 1/16/26 11:39 AM, Phillip Wood wrote:\n> >> I've Cc'd Peff and René for a second opinion if you have time please.\n> >>\n> >> On 15/01/2026 15:55, Ezekiel Newren wrote:\n> >>> On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n> >>>\n> >>>>> +static void _set_capacity(void *self_, size_t new_capacity)\n> >>>>> +{\n> >>>>> +     struct IVec_c_void *self = self_;\n> >>>>\n> >>>> Passing any of the ivec variants defined below to this function invokes\n> >>>> undefined behavior because we're not casting the pointer back to the\n> >>>> orginal type. However I think on the platforms we care about\n> >>>> sizeof(void*) == sizeof(T*) for all T so maybe we can look the other way.\n> >>>\n> >>> If someone finds that this code does not work because of this\n> >>> assumption I'd like to know. But I can't fathom a case where it\n> >>> wouldn't work.\n> >>\n> >> So we have two different structs\n> >>\n> >> struct IVec_c_void {\n> >>      void *ptr;\n> >>      size_t length;\n> >>      size_t capacity;\n> >>      size_t element_size;\n> >> }\n> >>\n> >> and\n> >>\n> >> struct Ivec_u8 {\n> >>      uint8_t *ptr;\n> >>      size_t length;\n> >>      size_t capacity;\n> >>      size_t element_size;\n> >> }\n> >>\n> >> One the platforms we care about they will have the same memory\n> >> layout as all pointers have the same representation. However I don't\n> >> think they are \"compatible types\" in the language of the C standard\n> >> because the type of the \"ptr\" member differs. That means casting\n> >> IVec_u8* to IVec_c_void* either directly or via void* is undefined\n> >> and so\n> >>\n> >>      struct IVec_u8 vec;\n> >>      ivec_init(&vec, sizeof(*vec.ptr));\n> >>\n> >> is undefined. For the compiler to see the undefined cast it needs to\n> >> look across translation units because the implementation of\n> >> ivec_init() will be in a separate file to where it is called. Maybe\n> >> that and the fact they have the same memory layout saves us from\n> >> having to worry too much though I'm always nervous of undefined\n> >> behavior.\n> >\n> > True.  The GCC docs give a fun example of what a compiler might do\n> > when using different struct types to access the same memory:\n> >\n> > https://www.gnu.org/software/c-intro-and-ref/manual/html_node/Aliasing-Type-Rules.html\n>\n> Thanks for the link\n>\n> > Not sure it applies to this case, but the point is that compilers\n> > can and will do terrifying things when they smell UB, with little\n> > concern for safety or original intent.\n> >\n> >> An alternative would be to pass the individual struct members as function parameters\n> >>\n> >>      void ivec_init(void **vec, size_t &length, size_t &capacity,\n> >>                 size_t &element_size_, size_t element_size)\n> >>      {\n> >>          *vec = NULL;\n> >>          *length = 0;\n> >>          *capacity = 0;\n> >>          *element_size_ = element_size;\n> >>      }\n> >\n> > The ampersands (&) should be asterisks (*), right?\n>\n> Indeed, that's embarrassing - I must have been thinking of the caller.\n>\n> >> and have DEFINE_IVEC_TYPE create typesafe wrappers\n> >>\n> >>      static inline void ivec_u8_init(struct IVec_u8 *vec)\n> >>      {\n> >>          void *ptr = vec->ptr;\n> >>          ivec_init(&ptr, &v->length, &v->capacity,\n> >>                &v->element_size, sizeof(*(v->ptr));\n> >>          vec->ptr = ptr;\n> >>      }\n> >\n> > Mixes \"v\" and \"vec\", misses a closing parenthesis.  Looks viable,\n> > though, and this method should be applicable to the rest of the\n> > functions as well (on the C side).\n> >\n> > I guess this doesn't require an element_size member anymore as\n> > each wrapper can pass in the sizeof value.\n>\n> Good point\n>\n> >> That's safe because we cast the \"ptr\" member to \"void*\" and then\n> >> back to the original type. On the rust side the implementation of\n> >> IVec<T> would also need to split out the individual struct members\n> >> when it calls ivec_init() etc. It's all a bit more effort but the\n> >> benefit is that we don't have any undefined behavior and we have a\n> >> nice typesafe C interface to 'struct IVec_*'.\n> > Right.  No idea how ugly this would be on the Rust side, though.\n>\n> I'm hoping it's not too bad and `impl IVec<T>` just contains the\n> equivalent of the wrappers generated by DEFINE_IVEC_TYPE()\n>\n> Thanks\n>\n> Phillip\n> >\n> > René\n> >\n>\n\nI don't like this solution. ivec_push() is the only function that\ndeals with actual values. The rest are just generic memory management\nfunctions. What if we used:\n\n#define ivec_init(vec) { \\\n    (vec)->ptr = NULL; \\\n    (vec)->length = 0; \\\n    (vec)->capacity = 0; \\\n    (vec)->element_size = sizeof(*(vec)->ptr); \\\n}\n\n#define ivec_push_unsafe(vec, value) (vec)->ptr[(vec)->length++] = (value)\n\n/*\n * grow by at least 1\n */\n#define ivec_push(vec, value) { \\\n    if ((vec)->length == (vec)->capacity) \\\n       ivec_reserve(vec, 1); \\\n    ivec_push_unsafe(vec, value); \\\n}\n\nInstead of concrete functions?\n"},{"id":"534116","messageId":"CAH=ZcbCfRziL7Aimq_9Z0k_8MqLRRy_NnD=HRhySYvu3HH3Y6w@mail.gmail.com","threadId":"64712","inReplyTo":"xmqq4ip2ndse.fsf@gitster.g","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-17T16:06:36Z","receivedAt":"2026-01-17T16:06:49Z","isPatch":true,"body":"On Sat, Jan 3, 2026 at 10:32 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > +     if (new_capacity == 0) {\n> > +             free(self->ptr);\n> > +             self->ptr = NULL;\n>\n>         if (!new_capacity)\n>                 FREE_AND_NULL(self->ptr);\n>         else\n>                 ...;\n>\n> > +void ivec_free(void *self_)\n> > +{\n> > +     struct IVec_c_void *self = self_;\n> > +\n> > +     free(self->ptr);\n> > +     self->ptr = NULL;\n>\n> Likewise.  Otherwise the code will fail coccicheck.\n>\n> > +     self->length = 0;\n> > +     self->capacity = 0;\n> > +     // DO NOT MODIFY element_size!!!\n>\n>         /* A single-liner comment in our codebase looks like this */\n>\n\nI will make these changes.\n"},{"id":"534117","messageId":"CAH=ZcbAogCpqg0RkKg1WjuAcuKyArDs4aP+k=McCs_byDT2Weg@mail.gmail.com","threadId":"64712","inReplyTo":"c2d9a432-0753-4786-8de9-c3dcfe69ac36@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-17T16:14:54Z","receivedAt":"2026-01-17T16:15:08Z","isPatch":true,"body":"On Fri, Jan 16, 2026 at 3:39 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> I've Cc'd Peff and René for a second opinion if you have time please.\n>\n> On 15/01/2026 15:55, Ezekiel Newren wrote:\n> > On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>  >\n> >>> +static void _set_capacity(void *self_, size_t new_capacity)\n> >>> +{\n> >>> +     struct IVec_c_void *self = self_;\n> >>\n> >> Passing any of the ivec variants defined below to this function invokes\n> >> undefined behavior because we're not casting the pointer back to the\n> >> orginal type. However I think on the platforms we care about\n> >> sizeof(void*) == sizeof(T*) for all T so maybe we can look the other way.\n> >\n> > If someone finds that this code does not work because of this\n> > assumption I'd like to know. But I can't fathom a case where it\n> > wouldn't work.\n>\n> So we have two different structs\n>\n> struct IVec_c_void {\n>         void *ptr;\n>         size_t length;\n>         size_t capacity;\n>         size_t element_size;\n> }\n>\n> and\n>\n> struct Ivec_u8 {\n>         uint8_t *ptr;\n>         size_t length;\n>         size_t capacity;\n>         size_t element_size;\n> }\n>\n> One the platforms we care about they will have the same memory layout as\n> all pointers have the same representation. However I don't think they\n> are \"compatible types\" in the language of the C standard because the\n> type of the \"ptr\" member differs. That means casting IVec_u8* to\n> IVec_c_void* either directly or via void* is undefined and so\n>\n>         struct IVec_u8 vec;\n>         ivec_init(&vec, sizeof(*vec.ptr));\n>\n> is undefined. For the compiler to see the undefined cast it needs to\n> look across translation units because the implementation of ivec_init()\n> will be in a separate file to where it is called. Maybe that and the\n> fact they have the same memory layout saves us from having to worry too\n> much though I'm always nervous of undefined behavior.\n>\n> An alternative would be to pass the individual struct members as\n> function parameters\n>\n>         void ivec_init(void **vec, size_t &length, size_t &capacity,\n>                        size_t &element_size_, size_t element_size)\n>         {\n>                 *vec = NULL;\n>                 *length = 0;\n>                 *capacity = 0;\n>                 *element_size_ = element_size;\n>         }\n>\n> and have DEFINE_IVEC_TYPE create typesafe wrappers\n>\n>         static inline void ivec_u8_init(struct IVec_u8 *vec)\n>         {\n>                 void *ptr = vec->ptr;\n>                 ivec_init(&ptr, &v->length, &v->capacity,\n>                           &v->element_size, sizeof(*(v->ptr));\n>                 vec->ptr = ptr;\n>         }\n>\n> That's safe because we cast the \"ptr\" member to \"void*\" and then back to\n> the original type. On the rust side the implementation of IVec<T> would\n> also need to split out the individual struct members when it calls\n> ivec_init() etc. It's all a bit more effort but the benefit is that we\n> don't have any undefined behavior and we have a nice typesafe C\n> interface to 'struct IVec_*'.\n>\n> Thanks\n>\n> Phillip\n>\n\nIf the size of different kinds of pointers ever differed from the size\nof void* then wouldn't that make all calls to malloc undefined? I\ndon't see this as a problem since I'm not casting between structs with\ndifferent members that are not pointers. I could use void* for\neverything, but then we'd need an accessor like *(T*)ivec_at(&vec, i),\nbut this is much more painful and error prone than simply vec.ptr[i].\n\nI agree that the example referenced by Rene is problematic, but\nirrelevant to ivec in my opinion.\n"},{"id":"534118","messageId":"CAH=ZcbCmMCYd7m-nrjSM4i3Tyr76C50ekJGQgDtRveMC7UxvwA@mail.gmail.com","threadId":"64712","inReplyTo":"CAH=ZcbAogCpqg0RkKg1WjuAcuKyArDs4aP+k=McCs_byDT2Weg@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-17T16:16:56Z","receivedAt":"2026-01-17T16:17:09Z","isPatch":true,"body":"> If the size of different kinds of pointers ever differed from the size\n> of void* then wouldn't that make all calls to malloc undefined? I\n\nI meant to say undefined behavior, not simply undefined.\n"},{"id":"534119","messageId":"CAH=ZcbDw0_Od3+zuGLsy3Z=bLR-4ByH8Fguiuw_MyLTi=U7gcQ@mail.gmail.com","threadId":"64712","inReplyTo":"07ca298a-ad32-4998-88ff-d69c04418fdd@web.de","subject":"Re: [PATCH 09/10] xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-17T16:34:06Z","receivedAt":"2026-01-17T16:34:19Z","isPatch":true,"body":"On Fri, Jan 16, 2026 at 1:19 PM René Scharfe <l.s.r@web.de> wrote:\n>\n> On 1/2/26 7:52 PM, Ezekiel Newren via GitGitGadget wrote:\n> > @@ -253,22 +250,44 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n> >       return rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n> >  }\n> >\n> > +struct xoccurrence\n> > +{\n> > +     size_t file1, file2;\n> > +};\n> > +\n> > +\n> > +DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n> > +\n> >\n> >  /*\n> >   * Try to reduce the problem complexity, discard records that have no\n> >   * matches on the other file. Also, lines that have multiple matches\n> >   * might be potentially discarded if they appear in a run of discardable.\n> >   */\n> > -static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n> > -     long i, nm, mlim;\n> > +static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n> > +     long i;\n> > +     size_t nm, mlim;\n> >       xrecord_t *recs;\n> > -     xdlclass_t *rcrec;\n> >       uint8_t *action1 = NULL, *action2 = NULL;\n> > -     bool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> > +     struct IVec_xoccurrence occ;\n> > +     bool need_min = !!(flags & XDF_NEED_MINIMAL);\n> >       int ret = 0;\n> >       ptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n> >       ptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n> >\n> > +     IVEC_INIT(occ);\n> > +     ivec_zero(&occ, xe->mph_size);\n>\n> This array is presized here.  It is neither grown nor shrunken.\n> CALLOC_ARRAY would work just as well, at least at this point, no?\n>\n> > +\n> > +     for (size_t j = 0; j < xe->xdf1.nrec; j++) {\n> > +             size_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n> > +             occ.ptr[mph1].file1 += 1;\n> > +     }\n> > +\n> > +     for (size_t j = 0; j < xe->xdf2.nrec; j++) {\n> > +             size_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n> > +             occ.ptr[mph2].file2 += 1;\n> > +     }\n> > +\n> >       /*\n> >        * Create temporary arrays that will help us decide if\n> >        * changed[i] should remain false, or become true.\n> > @@ -288,16 +307,14 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n> >       if ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n> >               mlim = XDL_MAX_EQLIMIT;\n> >       for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n> > -             rcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> > -             nm = rcrec ? rcrec->len2 : 0;\n> > +             nm = occ.ptr[recs->minimal_perfect_hash].file2;\n> >               action1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> >       }\n> >\n> >       if ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n> >               mlim = XDL_MAX_EQLIMIT;\n> >       for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n> > -             rcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> > -             nm = rcrec ? rcrec->len1 : 0;\n> > +             nm = occ.ptr[recs->minimal_perfect_hash].file1;\n> >               action2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> >       }\n> >\n> > @@ -332,6 +349,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n> >  cleanup:\n> >       xdl_free(action1);\n> >       xdl_free(action2);\n> > +     ivec_free(&occ);\n> >\n> >       return ret;\n> >  }\n\nIn Rust the memory management macros defined in git-compat-util.h will\nnot be available. ivec was built expressly to bridge the gap between C\nand Rust. I'm avoiding using those macros because I'm trying to get C\nprogrammers familiar with how Rust's Vec operates without forcing them\nto read and write in Rust. Also, it makes converting from IVec to Vec\nsuper easy.\n\nivec_zero() also sets length and capacity. Also CALLOC_ARRAY needs to\nknow the type of the pointer which ivec_zero() does not have access\nto. This is one of the few ivec functions that does not have a direct\nequivalent in Rust's Vec, but is faster than what is logically\nequivalent in Rust.\n\nIn Rust the closest safe equivalent would look like:\n\nlet size = 35;\nlet mut vec = Vec::<u64>::new();\nvec.reserve_exact(size);\nvec.fill(0);  // requires that T implements the `Copy` trait\n\nThe unsafe version would look like:\nlet size = 35;\nlet mut vec = Vec::<u64>::new();\nvec.reserve_exact(size);\nunsafe {\n    std::ptr::write_bytes(vec.as_mut_ptr(), 0, size * size_of::<u64>());\n}\n"},{"id":"534120","messageId":"6ae80903-3cc5-4017-9eac-0b3100b93b04@gmail.com","threadId":"64712","inReplyTo":"CAH=ZcbAogCpqg0RkKg1WjuAcuKyArDs4aP+k=McCs_byDT2Weg@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-17T17:40:08Z","receivedAt":"2026-01-17T17:40:10Z","isPatch":true,"body":"On 17/01/2026 16:14, Ezekiel Newren wrote:\n> \n> If the size of different kinds of pointers ever differed from the size\n> of void* then wouldn't that make all calls to malloc undefined?\n\nI believe there are (Havard architecture?) platforms where function \npointers are a different width to data pointers, and that's why you \ncannot store a function pointer in void*. I agree it would be weird for \nchar* to have a different width to int*, I suspect the restrictions on \ncasting from one type to another are about alignment.\n\n> I\n> don't see this as a problem since I'm not casting between structs with\n> different members that are not pointers.\n\nBut isn't that is still undefined behavior as far as the C standard is \nconcerned? It might make sense for it to work, but common sense has \nlittle to do with undefined behavior.\n\n> I could use void* for\n> everything, but then we'd need an accessor like *(T*)ivec_at(&vec, i),\n> but this is much more painful and error prone than simply vec.ptr[i].\n\nYeah that's horrible\n\nThanks\n\nPhillip\n\n"},{"id":"534145","messageId":"5c7e853d-f368-4d8a-a5f0-f4d485c1c3d6@web.de","threadId":"64712","inReplyTo":"CAH=ZcbB=Yf=wn2O273adrvpUpE0bJGKwrAjOAjmB8AgJrjz5Bg@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2026-01-18T14:55:23Z","receivedAt":"2026-01-18T14:55:25Z","isPatch":true,"body":"On 1/17/26 4:58 PM, Ezekiel Newren wrote:\n> On Fri, Jan 16, 2026 at 1:19 PM René Scharfe <l.s.r@web.de> wrote:\n>>\n>>> +void ivec_reserve(void *self_, size_t additional)\n>>> +{\n>>> +     struct IVec_c_void *self = self_;\n>>> +\n>>> +     size_t growby = 128;\n>>> +     if (self->capacity > growby)\n>>> +             growby = self->capacity;\n>>> +     if (additional > growby)\n>>> +             growby = additional;\n>>> +\n>>> +     _set_capacity(self, self->capacity + growby);\n>>> +}\n>>\n>> Constant growth steps like these cause linear growth and quadratic\n>> complexity.  ALLOC_GROW does exponential growth with factor 1.5 to\n>> get linear complexity.  Here's an old plea to do the same:\n>> https://blog.mozilla.org/nnethercote/2014/11/04/please-grow-your-buffers-exponentially/\n>>\n>> René\n> \n> It _is_ exponential. ivec_reserve(&vec, 1) means grow by _at least_ 1.\nD'oh!  Right, it grows with factor 2, as growby is at least as big as\n->capacity.  I can't read.\n\nRené\n\n"},{"id":"534146","messageId":"f5b36fd0-1942-499c-bf4a-1107a3afd951@web.de","threadId":"64712","inReplyTo":"CAH=ZcbCuY22WCqzyK-=Adw924a6ZJqnMYjWK9fxwoFn5xK9q-w@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2026-01-18T14:58:27Z","receivedAt":"2026-01-18T14:58:29Z","isPatch":true,"body":"On 1/17/26 5:04 PM, Ezekiel Newren wrote:\n> \n> I don't like this solution. ivec_push() is the only function that\n> deals with actual values. The rest are just generic memory management\n> functions. What if we used:\n> \n> #define ivec_init(vec) { \\\n>     (vec)->ptr = NULL; \\\n>     (vec)->length = 0; \\\n>     (vec)->capacity = 0; \\\n>     (vec)->element_size = sizeof(*(vec)->ptr); \\\n> }\n> \n> #define ivec_push_unsafe(vec, value) (vec)->ptr[(vec)->length++] = (value)\n> \n> /*\n>  * grow by at least 1\n>  */\n> #define ivec_push(vec, value) { \\\n>     if ((vec)->length == (vec)->capacity) \\\n>        ivec_reserve(vec, 1); \\\n>     ivec_push_unsafe(vec, value); \\\n> }\n> \n> Instead of concrete functions?\n\nThese macros are OK on the C side in respect to type-safety.\n\nI guess they would have to be duplicated somehow in Rust?\n\nHow would ivec_reserve() look like?\n\nThe macros use their parameter \"vec\" multiple times, though, so callers\nmust not pass in an expression with a side-effect, as it would be\nevaluated more than once.  We have a few of those already.  You have to\nbe careful not to do stuff like this (example of calling a _push-like\nfunction with an argument with a side-effect from\nstrbuf.c::strbuf_join_argv()):\n\n\twhile (--argc)\n\t\tstrbuf_addstr(buf, *(++argv));\n\nAlso they can't be used like a function -- you'd have to call them\nwithout a trailing semicolon.  That's a small issue and easily\novercome by wrapping their body in \"do { } while (0)\".\n\nRené\n\n"},{"id":"534149","messageId":"914e4157-557e-4ea4-9b17-b6b1cb078283@web.de","threadId":"64712","inReplyTo":"CAH=ZcbDw0_Od3+zuGLsy3Z=bLR-4ByH8Fguiuw_MyLTi=U7gcQ@mail.gmail.com","subject":"Re: [PATCH 09/10] xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2026-01-18T18:23:48Z","receivedAt":"2026-01-18T18:23:52Z","isPatch":true,"body":"On 1/17/26 5:34 PM, Ezekiel Newren wrote:\n> On Fri, Jan 16, 2026 at 1:19 PM René Scharfe <l.s.r@web.de> wrote:\n>>\n>> On 1/2/26 7:52 PM, Ezekiel Newren via GitGitGadget wrote:\n>>> @@ -253,22 +250,44 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n>>>       return rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n>>>  }\n>>>\n>>> +struct xoccurrence\n>>> +{\n>>> +     size_t file1, file2;\n>>> +};\n>>> +\n>>> +\n>>> +DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n>>> +\n>>>\n>>>  /*\n>>>   * Try to reduce the problem complexity, discard records that have no\n>>>   * matches on the other file. Also, lines that have multiple matches\n>>>   * might be potentially discarded if they appear in a run of discardable.\n>>>   */\n>>> -static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>>> -     long i, nm, mlim;\n>>> +static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n>>> +     long i;\n>>> +     size_t nm, mlim;\n>>>       xrecord_t *recs;\n>>> -     xdlclass_t *rcrec;\n>>>       uint8_t *action1 = NULL, *action2 = NULL;\n>>> -     bool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n>>> +     struct IVec_xoccurrence occ;\n>>> +     bool need_min = !!(flags & XDF_NEED_MINIMAL);\n>>>       int ret = 0;\n>>>       ptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n>>>       ptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n>>>\n>>> +     IVEC_INIT(occ);\n>>> +     ivec_zero(&occ, xe->mph_size);\n>>\n>> This array is presized here.  It is neither grown nor shrunken.\n>> CALLOC_ARRAY would work just as well, at least at this point, no?\n>>\n>>> +\n>>> +     for (size_t j = 0; j < xe->xdf1.nrec; j++) {\n>>> +             size_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n>>> +             occ.ptr[mph1].file1 += 1;\n>>> +     }\n>>> +\n>>> +     for (size_t j = 0; j < xe->xdf2.nrec; j++) {\n>>> +             size_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n>>> +             occ.ptr[mph2].file2 += 1;\n>>> +     }\n>>> +\n>>>       /*\n>>>        * Create temporary arrays that will help us decide if\n>>>        * changed[i] should remain false, or become true.\n>>> @@ -288,16 +307,14 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>>>       if ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>>>               mlim = XDL_MAX_EQLIMIT;\n>>>       for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n>>> -             rcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>>> -             nm = rcrec ? rcrec->len2 : 0;\n>>> +             nm = occ.ptr[recs->minimal_perfect_hash].file2;\n>>>               action1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>>>       }\n>>>\n>>>       if ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>>>               mlim = XDL_MAX_EQLIMIT;\n>>>       for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n>>> -             rcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>>> -             nm = rcrec ? rcrec->len1 : 0;\n>>> +             nm = occ.ptr[recs->minimal_perfect_hash].file1;\n>>>               action2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>>>       }\n>>>\n>>> @@ -332,6 +349,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>>>  cleanup:\n>>>       xdl_free(action1);\n>>>       xdl_free(action2);\n>>> +     ivec_free(&occ);\n>>>\n>>>       return ret;\n>>>  }\n> \n> In Rust the memory management macros defined in git-compat-util.h will\n> not be available. ivec was built expressly to bridge the gap between C\n> and Rust. I'm avoiding using those macros because I'm trying to get C\n> programmers familiar with how Rust's Vec operates without forcing them\n> to read and write in Rust. Also, it makes converting from IVec to Vec\n> super easy.\n> \n> ivec_zero() also sets length and capacity. Also CALLOC_ARRAY needs to\n> know the type of the pointer which ivec_zero() does not have access\n> to. This is one of the few ivec functions that does not have a direct\n> equivalent in Rust's Vec, but is faster than what is logically\n> equivalent in Rust.\n> \n> In Rust the closest safe equivalent would look like:\n> \n> let size = 35;\n> let mut vec = Vec::<u64>::new();\n> vec.reserve_exact(size);\n> vec.fill(0);  // requires that T implements the `Copy` trait\n> \n> The unsafe version would look like:\n> let size = 35;\n> let mut vec = Vec::<u64>::new();\n> vec.reserve_exact(size);\n> unsafe {\n>     std::ptr::write_bytes(vec.as_mut_ptr(), 0, size * size_of::<u64>());\n> }\n\nI was being unclear and made a few assumptions here.  My point was just\nthat this is a fixed-size array and doesn't need to be stored in a\nvariable-sized container.  This is the first Ivec user, and I would have\nexpected it to exercise the push function.  I assume accessing a\nfixed-size array via FFI would be a lot easier since allocation and\ngrowth are out of the picture.\n\nRené\n\n"},{"id":"534165","messageId":"20260119055947.GA3100271@coredump.intra.peff.net","threadId":"64712","inReplyTo":"6ae80903-3cc5-4017-9eac-0b3100b93b04@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-01-19T05:59:47Z","receivedAt":"2026-01-19T05:59:48Z","isPatch":true,"body":"On Sat, Jan 17, 2026 at 05:40:08PM +0000, Phillip Wood wrote:\n\n> On 17/01/2026 16:14, Ezekiel Newren wrote:\n> > \n> > If the size of different kinds of pointers ever differed from the size\n> > of void* then wouldn't that make all calls to malloc undefined?\n> \n> I believe there are (Havard architecture?) platforms where function pointers\n> are a different width to data pointers, and that's why you cannot store a\n> function pointer in void*. I agree it would be weird for char* to have a\n> different width to int*, I suspect the restrictions on casting from one type\n> to another are about alignment.\n\nThe standard does allow for different pointer sizes for char and int.\nThe key thing is that a void pointer has to be able to represent any. So\nyou can cast a smaller pointer to void and vice versa (and the latter\nwould presumably throw away some of the bits, which is OK as long as the\nvoid was made from one of those smaller pointers originally).\n\nMore discussion at:\n\n  https://c-faq.com/null/machexamp.html\n\nI don't know how malloc worked on those platforms, though. The caller\nknows that malloc returns a void pointer, so it could cast to the\nsmaller format in the usual way at the call-site. But I don't know how\nyou would tell malloc() in a standard way what type of pointer you\nwanted to get out of it. I suspect they may have had specialized\nallocation functions. Or maybe it was enough to just throw away the low\nbits if you only cared about a word-addressable pointer.\n\nAt any rate, yeah, I agree with your original concern that the two\nstructs are not compatible. The layouts could be totally different. And\nnot just due to pointer size, but IIRC pointers to different types could\nhave different alignment requirements. So:\n\n  struct foo_void {\n\tsize_t len;\n\tvoid *ptr;\n  };\n\n  struct foo_u8 {\n\tsize_t len;\n\tuint8_t *ptr;\n  };\n\nmight need different padding to properly align the pointers. In the case\nunder discussion the pointers are always at the start, though, so I\nthink it wouldn't matter.\n\n-Peff\n"},{"id":"534192","messageId":"CAH=ZcbCXAB3vzRbyHkunQh09njyLk4WXvfLVxynXaswEkBv+DA@mail.gmail.com","threadId":"64712","inReplyTo":"20260119055947.GA3100271@coredump.intra.peff.net","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-19T20:21:04Z","receivedAt":"2026-01-19T20:21:18Z","isPatch":true,"body":"On Sun, Jan 18, 2026 at 10:59 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Jan 17, 2026 at 05:40:08PM +0000, Phillip Wood wrote:\n>\n> > On 17/01/2026 16:14, Ezekiel Newren wrote:\n> > >\n> > > If the size of different kinds of pointers ever differed from the size\n> > > of void* then wouldn't that make all calls to malloc undefined?\n> >\n> > I believe there are (Havard architecture?) platforms where function pointers\n> > are a different width to data pointers, and that's why you cannot store a\n> > function pointer in void*. I agree it would be weird for char* to have a\n> > different width to int*, I suspect the restrictions on casting from one type\n> > to another are about alignment.\n>\n> The standard does allow for different pointer sizes for char and int.\n> The key thing is that a void pointer has to be able to represent any. So\n> you can cast a smaller pointer to void and vice versa (and the latter\n> would presumably throw away some of the bits, which is OK as long as the\n> void was made from one of those smaller pointers originally).\n>\n> More discussion at:\n>\n>   https://c-faq.com/null/machexamp.html\n>\n> I don't know how malloc worked on those platforms, though. The caller\n> knows that malloc returns a void pointer, so it could cast to the\n> smaller format in the usual way at the call-site. But I don't know how\n> you would tell malloc() in a standard way what type of pointer you\n> wanted to get out of it. I suspect they may have had specialized\n> allocation functions. Or maybe it was enough to just throw away the low\n> bits if you only cared about a word-addressable pointer.\n>\n> At any rate, yeah, I agree with your original concern that the two\n> structs are not compatible. The layouts could be totally different. And\n> not just due to pointer size, but IIRC pointers to different types could\n> have different alignment requirements. So:\n>\n>   struct foo_void {\n>         size_t len;\n>         void *ptr;\n>   };\n>\n>   struct foo_u8 {\n>         size_t len;\n>         uint8_t *ptr;\n>   };\n>\n> might need different padding to properly align the pointers. In the case\n> under discussion the pointers are always at the start, though, so I\n> think it wouldn't matter.\n>\n> -Peff\n\nOk..., is there a way to pad a field to the largest size needed so\nthat this also works on the harvard architecture? If C isn't even self\nconsistent then how are these structs going to be passed between C and\nRust (which is THE point of ivec)?\n\nOr do we just tell the arcane Harvard architecture \"too bad\" Git won't\nrun on it anymore?\n"},{"id":"534193","messageId":"20260119204010.GA3148606@coredump.intra.peff.net","threadId":"64712","inReplyTo":"CAH=ZcbCXAB3vzRbyHkunQh09njyLk4WXvfLVxynXaswEkBv+DA@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-01-19T20:40:10Z","receivedAt":"2026-01-19T20:40:13Z","isPatch":true,"body":"On Mon, Jan 19, 2026 at 01:21:04PM -0700, Ezekiel Newren wrote:\n\n> Ok..., is there a way to pad a field to the largest size needed so\n> that this also works on the harvard architecture? If C isn't even self\n> consistent then how are these structs going to be passed between C and\n> Rust (which is THE point of ivec)?\n\nIf you make a union of the pointers, it will require the largest size\nand the strictest alignment requirement. So:\n\n  struct foo {\n\tunion {\n\t\tvoid *v;\n\t\tuint8_t *u8;\n\t} ptr;\n\tsize_t len;\n  };\n\nwould be a single struct you could use to store a void pointer _or_ a u8\npointer. The one thing you shouldn't do there, though, is assign via one\nunion member and read from the other. So I don't know if that helps you\nor not (I confess I have not followed this rust discussion at all, and\nknow nothing about rust/c ABI compatibility, and just got roped in on C\nesoterica).\n\n> Or do we just tell the arcane Harvard architecture \"too bad\" Git won't\n> run on it anymore?\n\nMinor nit: the Harvard architecture is one where function pointers are\nnot the same as data pointers. An int/char distinction can happen even\non more common (von Neumann) machines.\n\nBut I think we can rephrase your question as: are there real-world\nmachines we care about that will have different pointer sizes, or can we\nignore this issue for practical purposes?\n\nI don't know the answer. I suspect it probably is OK for Git not to run\non the machines mentioned in that C faq. But:\n\n  1. Sometimes there are subtle implications of undefined behavior that\n     may cause a compiler (even for a sensible machine) to do unexpected\n     things. I don't know offhand if that is the case here.\n\n  2. There are some modern platforms in which pointers are a bit more\n     opaque than just numeric addresses. For example, we've had a few\n     patches dealing with questionable pointer usage to make things work\n     on CHERI Arm systems. I'm not sure if any of that would matter\n     here, though (IIRC, it was mostly that pointers were unexpectedly\n     large and had matching alignment requirements, but all of them\n     equally so).\n\n-Peff\n"},{"id":"534206","messageId":"CALnO6CCf9zEVgWHjK_-kHzALa_JrOUDT_CSA1a_xa5gfPB3LtQ@mail.gmail.com","threadId":"64712","inReplyTo":"20260119204010.GA3148606@coredump.intra.peff.net","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-01-20T02:36:44Z","receivedAt":"2026-01-20T02:36:56Z","isPatch":true,"body":"On Mon, Jan 19, 2026 at 3:41 PM Jeff King <peff@peff.net> wrote:\n>\n>   2. There are some modern platforms in which pointers are a bit more\n>      opaque than just numeric addresses. For example, we've had a few\n>      patches dealing with questionable pointer usage to make things work\n>      on CHERI Arm systems. I'm not sure if any of that would matter\n>      here, though (IIRC, it was mostly that pointers were unexpectedly\n>      large and had matching alignment requirements, but all of them\n>      equally so).\n\nArguably on all modern platforms, pointers are more than just numeric\naddresses, due to provenance ;)\n\n- https://www.ralfj.de/blog/2018/07/24/pointers-and-bytes.html\n- https://www.ralfj.de/blog/2020/12/14/provenance.html\n- https://www.ralfj.de/blog/2022/04/11/provenance-exposed.html\n\n-- \nD. Ben Knoble\n"},{"id":"534232","messageId":"c1846365-10d8-4252-bb38-59bd652fcfd0@gmail.com","threadId":"64712","inReplyTo":"20260119055947.GA3100271@coredump.intra.peff.net","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T13:46:24Z","receivedAt":"2026-01-20T13:46:27Z","isPatch":true,"body":"Hi Peff\n\nOn 19/01/2026 05:59, Jeff King wrote:\n> On Sat, Jan 17, 2026 at 05:40:08PM +0000, Phillip Wood wrote:\n> \n>> On 17/01/2026 16:14, Ezekiel Newren wrote:\n>>>\n>>> If the size of different kinds of pointers ever differed from the size\n>>> of void* then wouldn't that make all calls to malloc undefined?\n>>\n>> I believe there are (Havard architecture?) platforms where function pointers\n>> are a different width to data pointers, and that's why you cannot store a\n>> function pointer in void*. I agree it would be weird for char* to have a\n>> different width to int*, I suspect the restrictions on casting from one type\n>> to another are about alignment.\n> \n> The standard does allow for different pointer sizes for char and int.\n> The key thing is that a void pointer has to be able to represent any. So\n> you can cast a smaller pointer to void and vice versa (and the latter\n> would presumably throw away some of the bits, which is OK as long as the\n> void was made from one of those smaller pointers originally).\n> \n> More discussion at:\n> \n>    https://c-faq.com/null/machexamp.html\n\nThanks for the clarification and the link - the C FAQ is always an \ninteresting read.\n\nPhillip\n\n> I don't know how malloc worked on those platforms, though. The caller\n> knows that malloc returns a void pointer, so it could cast to the\n> smaller format in the usual way at the call-site. But I don't know how\n> you would tell malloc() in a standard way what type of pointer you\n> wanted to get out of it. I suspect they may have had specialized\n> allocation functions. Or maybe it was enough to just throw away the low\n> bits if you only cared about a word-addressable pointer.\n> \n> At any rate, yeah, I agree with your original concern that the two\n> structs are not compatible. The layouts could be totally different. And\n> not just due to pointer size, but IIRC pointers to different types could\n> have different alignment requirements. So:\n> \n>    struct foo_void {\n> \tsize_t len;\n> \tvoid *ptr;\n>    };\n> \n>    struct foo_u8 {\n> \tsize_t len;\n> \tuint8_t *ptr;\n>    };\n> \n> might need different padding to properly align the pointers. In the case\n> under discussion the pointers are always at the start, though, so I\n> think it wouldn't matter.\n> \n> -Peff\n\n"},{"id":"534236","messageId":"08318339-03c3-4068-92fa-7a711bd13da0@gmail.com","threadId":"64712","inReplyTo":"CAH=ZcbA_HgEO2T2smn4Yg6gf4sm4jrR8A0ek1v9nqsa1MXbRJw@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T14:06:48Z","receivedAt":"2026-01-20T14:06:51Z","isPatch":true,"body":"Hi Ezekiel\n\nOn 15/01/2026 15:55, Ezekiel Newren wrote:\n> On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>> +void ivec_reserve(void *self_, size_t additional)\n>>> +{\n>>> +     struct IVec_c_void *self = self_;\n>>> +\n>>> +     size_t growby = 128;\n>>> +     if (self->capacity > growby)\n>>> +             growby = self->capacity;\n>>> +     if (additional > growby)\n>>> +             growby = additional;\n>>\n>> This growth strategy differs from both ALLOC_GROW() and\n>> XDL_ALLOC_GROW(), if there isn't a good reason for that we should\n>> perhaps just use ALLOC_GROW() here.\n> \n> XDL_ALLOW_GROW() can't be used because the pointer is always a void*\n> in this function.\n\nOh right. I'm not sure that's not a reason to use a different growth \nstrategy though. The minimum size of 128 elements is probably good for \nthe xdiff code that creates arrays with one element per line but if this \nis supposed to be for general use it is going to waste space when we're \nallocating a lot of small arrays. ALLOC_GROW() uses alloc_nr() to \ncalculate the new side so perhaps we could use that here?\n\n>>> +void ivec_push(void *self_, const void *value)\n>>> +{\n>>> +     struct IVec_c_void *self = self_;\n>>> +     void *dst = NULL;\n>>> +\n>>> +     if (self->length == self->capacity)\n>>> +             ivec_reserve(self, 1);\n>>> +\n>>> +     dst = (uint8_t*)self->ptr + self->length * self->element_size;\n>>> +     memcpy(dst, value, self->element_size);\n>>\n>> If self->element_size was a compile time constant the compiler could\n>> easily optimize this call away. I'm not sure that is easy to achieve though.\n> \n> The problem is that I didn't want all of ivec to be macros that looked\n> like function calls. I wanted to minimize use of macros so that it was\n> easier to port and verify that the Rust implementation matches the\n> behavior of the C implementation.\n\nI think that's a reasonable concern. So is the plan to have a parallel \nrust implementation of these functions rather than call the C \nimplementation from rust?\n\n>>> +void ivec_free(void *self_)\n>>\n>> Normally we'd call a like this that free the allocations and\n>> re-initializes the members ivec_clear()\n> \n> In Rust Vec.clear() means to set length to zero, but leaves the\n> allocation alone. The reason why I'm zeroing the struct is to help\n> avoid FFI issues. If not zero then what should the members be set to,\n> to indicate that using the struct is not valid anymore? In Rust an\n> object is freed when it goes out of scope and _cannot_ be accessed\n> afterward.\n\nI'm aware that Vec::clear() has different semantics (it does what \nstrbuf_reset() does). That's unfortunate but this function has different \nsemantics to all the other *_free() functions in git. Our coding \nguidelines say\n\n  - There are several common idiomatic names for functions performing\n    specific tasks on a structure `S`:\n\n     - `S_init()` initializes a structure without allocating the\n       structure itself.\n\n     - `S_release()` releases a structure's contents without freeing the\n       structure.\n\n     - `S_clear()` is equivalent to `S_release()` followed by `S_init()`\n       such that the structure is directly usable after clearing it. When\n       `S_clear()` is provided, `S_init()` shall not allocate resources\n       that need to be released again.\n\n     - `S_free()` releases a structure's contents and frees the\n       structure.\n\nAs we write more rust code and so wrap more of our existing structs \nwe're going to be wrapping C code that uses the definitions above so I \nthink we should do the same with struct IVec_*.\n\n>>> diff --git a/compat/ivec.h b/compat/ivec.h\n>>> new file mode 100644\n>>> index 0000000000..654a05c506\n>>> --- /dev/null\n>>> +++ b/compat/ivec.h\n>>> @@ -0,0 +1,52 @@\n>>> +#ifndef IVEC_H\n>>> +#define IVEC_H\n>>> +\n>>> +#include <git-compat-util.h>\n>>\n>> It would be nice to have some documentation in this header, see the\n>> examples in strvec.h and hashmap.h\n>>\n>>> +#define IVEC_INIT(variable) ivec_init(&(variable), sizeof(*(variable).ptr))\n>>\n>> This is a bit cumbersome to use compared to our usual *_INIT macros. I'm\n>> struggling to see how we can make it nicer though as DEFINE_IVEC_TYPE\n>> cannot define a per-type initializer macro and I we cannot initialize\n>> the element size without knowing the type.\n> \n> I don't see what's cumbersome about it. Maybe an example use case\n> would clarify things.\n\nIt is cumbersome because it separates the initialization from the \ndeclaration. Normally our *_INIT macros are initializer lists so we can \nwrite\n\n\tstruct strbuf = STRBUF_INIT;\n\nwhich keeps the declaration and initialization together. Although \nthey're on adjacent lines in your example in real code the \ninitialization likely to be separated from the declaration by other \nvariable declarations.\n\n> ```\n> DEFINE_IVEC_TYPE(xrecord_t, xrecord);\n> \n> void some_function() {\n>      struct IVec_xrecord rec;\n>      IVEC_INIT(rec);  // i.e. ivec_init(&rec, sizeof(*rec.ptr);\n\nThanks\n\nPhillip\n"},{"id":"534242","messageId":"1c46f551-0040-481e-9476-bc1b85f92636@gmail.com","threadId":"64712","inReplyTo":"9bd01bce9f0763d9dcc962ff94fcda36346bafc4.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 02/10] xdiff: make classic diff explicit by creating xdl_do_classic_diff()","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T15:01:39Z","receivedAt":"2026-01-20T15:01:44Z","isPatch":true,"body":"On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Later patches will prepare xdl_cleanup_records() to be moved into xdiffi.c\n> since only the classic diff uses that function.\n\nI assume that's to make it easier to covert the myers implementation to \nrust without affecting the rest of the code? If so it would be nice to \nsay that.\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n\n> +int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n> +\t\txdfenv_t *xe) {\n> +\tint res;\n> +\n> +\tif (xdl_prepare_env(mf1, mf2, xpp, xe) < 0)\n> +\t\treturn -1;\n> +\n> +\tif (XDF_DIFF_ALG(xpp->flags) == XDF_PATIENCE_DIFF) {\n> +\t\tres = xdl_do_patience_diff(xpp, xe);\n> +\t\tgoto out;\n> +\t}\n> +\n> +\tif (XDF_DIFF_ALG(xpp->flags) == XDF_HISTOGRAM_DIFF) {\n> +\t\tres = xdl_do_histogram_diff(xpp, xe);\n> +\t\tgoto out;\n> +\t}\n> +\n> +\tres = xdl_do_classic_diff(xe, xpp->flags);\n\nThis might be clearer that we're calling only one of the three functions \nif we wrote this as\n\n\tif (XDF_DIFF_ALG(xpp->flags) == XDIF_PATIENCE_DIFF)\n\t\tres = xdl_do_patience_diff(xpp, xe);\n\telse if (XDF_DIFF_ALG(xpp->flags) == XDF_HISTOGRAM_DIFF)\n\t\tres = xdl_do_histogram_diff(xpp, xe);\n\telse\n\t\tres = xdl_do_classic_diff(xe, xpp->flags);\n\nand then we can drop the out: label\n\nThanks\n\nPhillip\n\n>    out:\n>   \tif (res < 0)\n>   \t\txdl_free_env(xe);\n> diff --git a/xdiff/xdiffi.h b/xdiff/xdiffi.h\n> index 49e52c67f9..8bf4c20373 100644\n> --- a/xdiff/xdiffi.h\n> +++ b/xdiff/xdiffi.h\n> @@ -42,6 +42,7 @@ typedef struct s_xdchange {\n>   int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n>   \t\t xdfile_t *xdf2, long off2, long lim2,\n>   \t\t long *kvdf, long *kvdb, int need_min, xdalgoenv_t *xenv);\n> +int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags);\n>   int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>   \t\txdfenv_t *xe);\n>   int xdl_change_compact(xdfile_t *xdf, xdfile_t *xdfo, long flags);\n\n"},{"id":"534243","messageId":"208da094-8a5d-4f16-b42b-5d5204576b5f@gmail.com","threadId":"64712","inReplyTo":"53e4840c1653772379dc8d5c883b34717b81ac43.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 03/10] xdiff: don't waste time guessing the number of lines","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T15:02:28Z","receivedAt":"2026-01-20T15:02:33Z","isPatch":true,"body":"On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> All lines must be read anyway, so classify them after they're read in.\n> Also move the memset() into xdl_init_classifier().\n\nSo instead of looping over the input lines one and a bit times (the bit \nbeing from xdl_guess_lines) we now loop over them twice as we split them \nfirst and then classify them in a separate loop. It does save some work \nnot to call xdl_guess_lines but it is unclear if that offsets \nclassifying them in a separate loop.\n\n> +\tfor (size_t i = 0; i < xe->xdf1.nrec; i++) {\n> +\t\txrecord_t *rec = &xe->xdf1.recs[i];\n> +\t\txdl_classify_record(1, &cf, rec);\n\nWe seem to have lost the error handling if xdl_classify_record() fails.\n\nThanks\n\nPhillip\n\n> +\t}\n> +\n> +\tfor (size_t i = 0; i < xe->xdf2.nrec; i++) {\n> +\t\txrecord_t *rec = &xe->xdf2.recs[i];\n> +\t\txdl_classify_record(2, &cf, rec);\n>   \t}\n>   \n>   \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n> diff --git a/xdiff/xutils.c b/xdiff/xutils.c\n> index 77ee1ad9c8..b3d51197c1 100644\n> --- a/xdiff/xutils.c\n> +++ b/xdiff/xutils.c\n> @@ -118,26 +118,6 @@ void *xdl_cha_alloc(chastore_t *cha) {\n>   \treturn data;\n>   }\n>   \n> -long xdl_guess_lines(mmfile_t *mf, long sample) {\n> -\tlong nl = 0, size, tsize = 0;\n> -\tchar const *data, *cur, *top;\n> -\n> -\tif ((cur = data = xdl_mmfile_first(mf, &size))) {\n> -\t\tfor (top = data + size; nl < sample && cur < top; ) {\n> -\t\t\tnl++;\n> -\t\t\tif (!(cur = memchr(cur, '\\n', top - cur)))\n> -\t\t\t\tcur = top;\n> -\t\t\telse\n> -\t\t\t\tcur++;\n> -\t\t}\n> -\t\ttsize += (long) (cur - data);\n> -\t}\n> -\n> -\tif (nl && tsize)\n> -\t\tnl = xdl_mmfile_size(mf) / (tsize / nl);\n> -\n> -\treturn nl + 1;\n> -}\n>   \n>   int xdl_blankline(const char *line, long size, long flags)\n>   {\n> diff --git a/xdiff/xutils.h b/xdiff/xutils.h\n> index 615b4a9d35..d800840dd0 100644\n> --- a/xdiff/xutils.h\n> +++ b/xdiff/xutils.h\n> @@ -31,7 +31,6 @@ int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n>   int xdl_cha_init(chastore_t *cha, long isize, long icount);\n>   void xdl_cha_free(chastore_t *cha);\n>   void *xdl_cha_alloc(chastore_t *cha);\n> -long xdl_guess_lines(mmfile_t *mf, long sample);\n>   int xdl_blankline(const char *line, long size, long flags);\n>   int xdl_recmatch(const char *l1, long s1, const char *l2, long s2, long flags);\n>   uint64_t xdl_hash_record_verbatim(uint8_t const **data, uint8_t const *top);\n\n"},{"id":"534244","messageId":"f17adb7a-8776-42b3-b753-f6306145250a@gmail.com","threadId":"64712","inReplyTo":"70040ea1351451243be90d59d26cf1a403f3000a.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 04/10] xdiff: let patience and histogram benefit from xdl_trim_ends()","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T15:02:40Z","receivedAt":"2026-01-20T15:02:44Z","isPatch":true,"body":"On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> The patience diff is set up the exact same way as histogram, see\n> xdl_do_historgram_diff() in xhistogram.c. xdl_optimize_ctxs() is\n> redundant now, delete it.\n\nDoes this change the output? The patience diff looks for unique context \nlines and builds the context out from those. For files that look like\n\nOld\tNew\nA\tA\nB\tB\nC\tA\nB\tB\nA\tC\n\tB\n\tA\n\nThat will give a hunk\n\n@@ -1,3 +0,5 @@\n+A\n+B\n  A\n  B\n  C\n\nbut trimming the common prefix first would give\n\n@@ -1,5 +1,7\n  A\n  B\n+A\n+B\n  C\n  B\n  A\n\nThough it seems like the diff silder causes us to output the same diff \nin both cases for that simple test so maybe it is not an issue. It would \ncertainly be helpful to comment on any possible changes in the commit \nmessage as it could have been a deliberate choice not to trim the ends \nfor those algorithms.\n\n> -static int xdl_optimize_ctxs(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n> -\n> -\tif (xdl_trim_ends(xdf1, xdf2) < 0 ||\n> -\t    xdl_cleanup_records(cf, xdf1, xdf2) < 0) {\n> -\n> -\t\treturn -1;\n> -\t}\n> -\n> -\treturn 0;\n> -}\n> -\n>   int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>   \t\t    xdfenv_t *xe) {\n>   \txdlclassifier_t cf;\n> @@ -404,9 +393,10 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>   \t\txdl_classify_record(2, &cf, rec);\n>   \t}\n>   \n> +\txdl_trim_ends(&xe->xdf1, &xe->xdf2);\n\nIt would be clear that this was safe if you changed the function \nsignature to return void as the way it is called in xdl_optimize_ctxs() \nmakes it look like it can return an error.\n\nThanks\n\nPhillip\n\n>   \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n>   \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n> -\t    xdl_optimize_ctxs(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n> +\t    xdl_cleanup_records(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n>   \n>   \t\txdl_free_ctx(&xe->xdf2);\n>   \t\txdl_free_ctx(&xe->xdf1);\n\n"},{"id":"534270","messageId":"e2ccd068-3ea7-4be1-9ed1-71b39382c064@gmail.com","threadId":"64712","inReplyTo":"742f2d381af52cd8314a6a643e84b6b9daf99c7c.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 05/10] xdiff: use xdfenv_t in xdl_trim_ends() and xdl_cleanup_records()","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T16:32:01Z","receivedAt":"2026-01-20T16:32:04Z","isPatch":true,"body":"On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> View with --color-words. Prepare these functions to use the fields:\n> delta_start, delta_end. A future patch will add these fields to\n> xdfenv_t.\n\nI'm afraid this message doesn't make much sense to me. What are these \nnew fields? I think it would help to explain what this up comming change \nis going to do and why.\n\nOh, having read patch 7 we're removing dstart and dend from xdfile_t and \nreplacing them with delta_start and delta_end in xdfenv_t. It would be \nuseful to say that here.\n> -static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n> +static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n\nMaybe we could add xdf1 and xdf2 as local variables to avoid having to \nchange the code that accesses the members of xdfile_t that are not going \nto be moved to xdfenv_t.\n\nThanks\n\nPhillip\n\n>   \tlong i, nm, mlim;\n>   \txrecord_t *recs;\n>   \txdlclass_t *rcrec;\n> @@ -273,11 +273,11 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t * Create temporary arrays that will help us decide if\n>   \t * changed[i] should remain false, or become true.\n>   \t */\n> -\tif (!XDL_CALLOC_ARRAY(action1, xdf1->nrec + 1)) {\n> +\tif (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n>   \t\tret = -1;\n>   \t\tgoto cleanup;\n>   \t}\n> -\tif (!XDL_CALLOC_ARRAY(action2, xdf2->nrec + 1)) {\n> +\tif (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n>   \t\tret = -1;\n>   \t\tgoto cleanup;\n>   \t}\n> @@ -285,17 +285,17 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t/*\n>   \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>   \t */\n> -\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n> +\tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>   \t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n> +\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart]; i <= xe->xdf1.dend; i++, recs++) {\n>   \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>   \t\tnm = rcrec ? rcrec->len2 : 0;\n>   \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>   \t}\n>   \n> -\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n> +\tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>   \t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n> +\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart]; i <= xe->xdf2.dend; i++, recs++) {\n>   \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>   \t\tnm = rcrec ? rcrec->len1 : 0;\n>   \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> @@ -305,27 +305,27 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t * Use temporary arrays to decide if changed[i] should remain\n>   \t * false, or become true.\n>   \t */\n> -\txdf1->nreff = 0;\n> -\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n> -\t     i <= xdf1->dend; i++, recs++) {\n> +\txe->xdf1.nreff = 0;\n> +\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart];\n> +\t     i <= xe->xdf1.dend; i++, recs++) {\n>   \t\tif (action1[i] == KEEP ||\n> -\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n> -\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n> +\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->xdf1.dstart, xe->xdf1.dend))) {\n> +\t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n>   \t\t\t/* changed[i] remains false, i.e. keep */\n>   \t\t} else\n> -\t\t\txdf1->changed[i] = true;\n> +\t\t\txe->xdf1.changed[i] = true;\n>   \t\t\t/* i.e. discard */\n>   \t}\n>   \n> -\txdf2->nreff = 0;\n> -\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n> -\t     i <= xdf2->dend; i++, recs++) {\n> +\txe->xdf2.nreff = 0;\n> +\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart];\n> +\t     i <= xe->xdf2.dend; i++, recs++) {\n>   \t\tif (action2[i] == KEEP ||\n> -\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n> -\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n> +\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->xdf2.dstart, xe->xdf2.dend))) {\n> +\t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n>   \t\t\t/* changed[i] remains false, i.e. keep */\n>   \t\t} else\n> -\t\t\txdf2->changed[i] = true;\n> +\t\t\txe->xdf2.changed[i] = true;\n>   \t\t\t/* i.e. discard */\n>   \t}\n>   \n> @@ -340,27 +340,27 @@ cleanup:\n>   /*\n>    * Early trim initial and terminal matching records.\n>    */\n> -static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n> +static int xdl_trim_ends(xdfenv_t *xe) {\n>   \tlong i, lim;\n>   \txrecord_t *recs1, *recs2;\n>   \n> -\trecs1 = xdf1->recs;\n> -\trecs2 = xdf2->recs;\n> -\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n> +\trecs1 = xe->xdf1.recs;\n> +\trecs2 = xe->xdf2.recs;\n> +\tfor (i = 0, lim = (long)XDL_MIN(xe->xdf1.nrec, xe->xdf2.nrec); i < lim;\n>   \t     i++, recs1++, recs2++)\n>   \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n>   \t\t\tbreak;\n>   \n> -\txdf1->dstart = xdf2->dstart = i;\n> +\txe->xdf1.dstart = xe->xdf2.dstart = i;\n>   \n> -\trecs1 = xdf1->recs + xdf1->nrec - 1;\n> -\trecs2 = xdf2->recs + xdf2->nrec - 1;\n> +\trecs1 = xe->xdf1.recs + xe->xdf1.nrec - 1;\n> +\trecs2 = xe->xdf2.recs + xe->xdf2.nrec - 1;\n>   \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n>   \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n>   \t\t\tbreak;\n>   \n> -\txdf1->dend = (long)xdf1->nrec - i - 1;\n> -\txdf2->dend = (long)xdf2->nrec - i - 1;\n> +\txe->xdf1.dend = (long)xe->xdf1.nrec - i - 1;\n> +\txe->xdf2.dend = (long)xe->xdf2.nrec - i - 1;\n>   \n>   \treturn 0;\n>   }\n> @@ -393,10 +393,10 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>   \t\txdl_classify_record(2, &cf, rec);\n>   \t}\n>   \n> -\txdl_trim_ends(&xe->xdf1, &xe->xdf2);\n> +\txdl_trim_ends(xe);\n>   \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n>   \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n> -\t    xdl_cleanup_records(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n> +\t    xdl_cleanup_records(&cf, xe) < 0) {\n>   \n>   \t\txdl_free_ctx(&xe->xdf2);\n>   \t\txdl_free_ctx(&xe->xdf1);\n\n"},{"id":"534271","messageId":"2b6592f6-1d20-4cdd-8af7-f524920ad1a7@gmail.com","threadId":"64712","inReplyTo":"65da408da9589420ec341368d0853e6183aee922.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 06/10] xdiff: cleanup xdl_trim_ends()","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T16:32:41Z","receivedAt":"2026-01-20T16:32:44Z","isPatch":true,"body":"\n\nOn 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> This patch is best viewed with a before and after of the whole\n> function.\n> \n> Rather than using 2 pointers and walking them. Use direct indexing with\n> local variables of what is being compared to make it easier to follow\n> along.\n\nI think using direct indexing makes things clearer, but I'm not sure \nthis is a faithful conversion (see below).\n\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 0acb3437d4..06b6a6f804 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -340,29 +340,29 @@ cleanup:\n>   /*\n>    * Early trim initial and terminal matching records.\n>    */\n> -static int xdl_trim_ends(xdfenv_t *xe) {\n> -\tlong i, lim;\n> -\txrecord_t *recs1, *recs2;\n> -\n> -\trecs1 = xe->xdf1.recs;\n> -\trecs2 = xe->xdf2.recs;\n> -\tfor (i = 0, lim = (long)XDL_MIN(xe->xdf1.nrec, xe->xdf2.nrec); i < lim;\n> -\t     i++, recs1++, recs2++)\n> -\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n> +static void xdl_trim_ends(xdfenv_t *xe)\n> +{\n> +\tsize_t lim = XDL_MIN(xe->xdf1.nrec, xe->xdf2.nrec);\n> +\n> +\tfor (size_t i = 0; i < lim; i++) {\n> +\t\tsize_t mph1 = xe->xdf1.recs[i].minimal_perfect_hash;\n> +\t\tsize_t mph2 = xe->xdf2.recs[i].minimal_perfect_hash;\n> +\t\tif (mph1 != mph2) {\n> +\t\t\txe->xdf1.dstart = xe->xdf2.dstart = (ssize_t)i;\n\nThe type of dstart is ptrdiff_t, not ssize_t.\n\nThe original set dstart and dend unconditionally but here they are not \nset if all the lines match.\n\nThanks\n\nPhillip\n\n> +\t\t\tlim -= i;\n>   \t\t\tbreak;\n> +\t\t}\n> +\t}\n>   \n> -\txe->xdf1.dstart = xe->xdf2.dstart = i;\n> -\n> -\trecs1 = xe->xdf1.recs + xe->xdf1.nrec - 1;\n> -\trecs2 = xe->xdf2.recs + xe->xdf2.nrec - 1;\n> -\tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n> -\t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n> +\tfor (size_t i = 0; i < lim; i++) {\n> +\t\tsize_t mph1 = xe->xdf1.recs[xe->xdf1.nrec - 1 - i].minimal_perfect_hash;\n> +\t\tsize_t mph2 = xe->xdf2.recs[xe->xdf2.nrec - 1 - i].minimal_perfect_hash;\n> +\t\tif (mph1 != mph2) {\n> +\t\t\txe->xdf1.dend = xe->xdf1.nrec - 1 - i;\n> +\t\t\txe->xdf2.dend = xe->xdf2.nrec - 1 - i;\n>   \t\t\tbreak;\n> -\n> -\txe->xdf1.dend = (long)xe->xdf1.nrec - i - 1;\n> -\txe->xdf2.dend = (long)xe->xdf2.nrec - i - 1;\n> -\n> -\treturn 0;\n> +\t\t}\n> +\t}\n>   }\n>   \n>   \n\n"},{"id":"534272","messageId":"e9a031fd-072d-4810-b7e0-0d64ffedce10@gmail.com","threadId":"64712","inReplyTo":"d74722538b693fb26e8684f9dd3fbc319a2a575e.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 07/10] xdiff: replace xdfile_t.dstart with xdfenv_t.delta_start","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-20T16:32:54Z","receivedAt":"2026-01-20T16:32:56Z","isPatch":true,"body":"On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Placing delta_start in xdfenv_t instead of xdfile_t provides a more\n> appropriate context since this variable only makes sense with a pair\n> of files. View with --color-words.\n\nSo as dstart and dend must be the same for both files we now store the \nvalues once in xdfenv_t. That explains why we start passing xdfenv_t \naround rather than xdfile_t in patch 5.\n\nThanks\n\nPhillip\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xhistogram.c |  4 ++--\n>   xdiff/xpatience.c  |  4 ++--\n>   xdiff/xprepare.c   | 17 +++++++++--------\n>   xdiff/xtypes.h     |  3 ++-\n>   4 files changed, 15 insertions(+), 13 deletions(-)\n> \n> diff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\n> index 5ae1282c27..eb6a52d9ba 100644\n> --- a/xdiff/xhistogram.c\n> +++ b/xdiff/xhistogram.c\n> @@ -365,6 +365,6 @@ out:\n>   int xdl_do_histogram_diff(xpparam_t const *xpp, xdfenv_t *env)\n>   {\n>   \treturn histogram_diff(xpp, env,\n> -\t\tenv->xdf1.dstart + 1, env->xdf1.dend - env->xdf1.dstart + 1,\n> -\t\tenv->xdf2.dstart + 1, env->xdf2.dend - env->xdf2.dstart + 1);\n> +\t\tenv->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n> +\t\tenv->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n>   }\n> diff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\n> index 2bce07cf48..bd0ffbb417 100644\n> --- a/xdiff/xpatience.c\n> +++ b/xdiff/xpatience.c\n> @@ -374,6 +374,6 @@ static int patience_diff(xpparam_t const *xpp, xdfenv_t *env,\n>   int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n>   {\n>   \treturn patience_diff(xpp, env,\n> -\t\tenv->xdf1.dstart + 1, env->xdf1.dend - env->xdf1.dstart + 1,\n> -\t\tenv->xdf2.dstart + 1, env->xdf2.dend - env->xdf2.dstart + 1);\n> +\t\tenv->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n> +\t\tenv->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n>   }\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 06b6a6f804..e88468e74c 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -173,7 +173,6 @@ static int xdl_prepare_ctx(mmfile_t *mf, xdfile_t *xdf, uint64_t flags) {\n>   \n>   \txdf->changed += 1;\n>   \txdf->nreff = 0;\n> -\txdf->dstart = 0;\n>   \txdf->dend = xdf->nrec - 1;\n>   \n>   \treturn 0;\n> @@ -287,7 +286,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>   \t */\n>   \tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>   \t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart]; i <= xe->xdf1.dend; i++, recs++) {\n> +\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= xe->xdf1.dend; i++, recs++) {\n>   \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>   \t\tnm = rcrec ? rcrec->len2 : 0;\n>   \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> @@ -295,7 +294,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>   \n>   \tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>   \t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart]; i <= xe->xdf2.dend; i++, recs++) {\n> +\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= xe->xdf2.dend; i++, recs++) {\n>   \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>   \t\tnm = rcrec ? rcrec->len1 : 0;\n>   \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> @@ -306,10 +305,10 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>   \t * false, or become true.\n>   \t */\n>   \txe->xdf1.nreff = 0;\n> -\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart];\n> +\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n>   \t     i <= xe->xdf1.dend; i++, recs++) {\n>   \t\tif (action1[i] == KEEP ||\n> -\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->xdf1.dstart, xe->xdf1.dend))) {\n> +\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, xe->xdf1.dend))) {\n>   \t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n>   \t\t\t/* changed[i] remains false, i.e. keep */\n>   \t\t} else\n> @@ -318,10 +317,10 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>   \t}\n>   \n>   \txe->xdf2.nreff = 0;\n> -\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart];\n> +\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n>   \t     i <= xe->xdf2.dend; i++, recs++) {\n>   \t\tif (action2[i] == KEEP ||\n> -\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->xdf2.dstart, xe->xdf2.dend))) {\n> +\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, xe->xdf2.dend))) {\n>   \t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n>   \t\t\t/* changed[i] remains false, i.e. keep */\n>   \t\t} else\n> @@ -348,7 +347,7 @@ static void xdl_trim_ends(xdfenv_t *xe)\n>   \t\tsize_t mph1 = xe->xdf1.recs[i].minimal_perfect_hash;\n>   \t\tsize_t mph2 = xe->xdf2.recs[i].minimal_perfect_hash;\n>   \t\tif (mph1 != mph2) {\n> -\t\t\txe->xdf1.dstart = xe->xdf2.dstart = (ssize_t)i;\n> +\t\t\txe->delta_start = (ssize_t)i;\n>   \t\t\tlim -= i;\n>   \t\t\tbreak;\n>   \t\t}\n> @@ -370,6 +369,8 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>   \t\t    xdfenv_t *xe) {\n>   \txdlclassifier_t cf;\n>   \n> +\txe->delta_start = 0;\n> +\n>   \tif (xdl_prepare_ctx(mf1, &xe->xdf1, xpp->flags) < 0) {\n>   \n>   \t\treturn -1;\n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index 979586f20a..bda1f85eb0 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -48,7 +48,7 @@ typedef struct s_xrecord {\n>   typedef struct s_xdfile {\n>   \txrecord_t *recs;\n>   \tsize_t nrec;\n> -\tptrdiff_t dstart, dend;\n> +\tptrdiff_t dend;\n>   \tbool *changed;\n>   \tsize_t *reference_index;\n>   \tsize_t nreff;\n> @@ -56,6 +56,7 @@ typedef struct s_xdfile {\n>   \n>   typedef struct s_xdfenv {\n>   \txdfile_t xdf1, xdf2;\n> +\tsize_t delta_start;\n>   } xdfenv_t;\n>   \n>   \n\n"},{"id":"534353","messageId":"6d533cfd-d308-4004-8d8e-4ae730c76086@gmail.com","threadId":"64712","inReplyTo":"f17adb7a-8776-42b3-b753-f6306145250a@gmail.com","subject":"Re: [PATCH 04/10] xdiff: let patience and histogram benefit from xdl_trim_ends()","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-21T14:49:58Z","receivedAt":"2026-01-21T14:50:01Z","isPatch":true,"body":"On 20/01/2026 15:02, Phillip Wood wrote:\n> On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>\n>> The patience diff is set up the exact same way as histogram, see\n>> xdl_do_historgram_diff() in xhistogram.c. xdl_optimize_ctxs() is\n>> redundant now, delete it.\n> \n> Does this change the output? The patience diff looks for unique context \n> lines and builds the context out from those. For files that look like\n> \n> Old    New\n> A    A\n> B    B\n> C    A\n> B    B\n> A    C\n>      B\n>      A\n> \n> That will give a hunk\n> \n> @@ -1,3 +0,5 @@\n> +A\n> +B\n>   A\n>   B\n>   C\n> \n> but trimming the common prefix first would give\n> \n> @@ -1,5 +1,7\n>   A\n>   B\n> +A\n> +B\n>   C\n>   B\n>   A\n> \n> Though it seems like the diff silder causes us to output the same diff \n> in both cases for that simple test so maybe it is not an issue.\n\nIt does change larger diffs. If you run\n\ngit show --diff-algorithm=patience --diff-merges=first-parent f406b89552\n\nYou get a different diff with this series applied.\n\nThanks\n\nPhillip\n\n> It would \n> certainly be helpful to comment on any possible changes in the commit \n> message as it could have been a deliberate choice not to trim the ends \n> for those algorithms.\n> \n>> -static int xdl_optimize_ctxs(xdlclassifier_t *cf, xdfile_t *xdf1, \n>> xdfile_t *xdf2) {\n>> -\n>> -    if (xdl_trim_ends(xdf1, xdf2) < 0 ||\n>> -        xdl_cleanup_records(cf, xdf1, xdf2) < 0) {\n>> -\n>> -        return -1;\n>> -    }\n>> -\n>> -    return 0;\n>> -}\n>> -\n>>   int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>>               xdfenv_t *xe) {\n>>       xdlclassifier_t cf;\n>> @@ -404,9 +393,10 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, \n>> xpparam_t const *xpp,\n>>           xdl_classify_record(2, &cf, rec);\n>>       }\n>> +    xdl_trim_ends(&xe->xdf1, &xe->xdf2);\n> \n> It would be clear that this was safe if you changed the function \n> signature to return void as the way it is called in xdl_optimize_ctxs() \n> makes it look like it can return an error.\n> \n> Thanks\n> \n> Phillip\n> \n>>       if ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n>>           (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n>> -        xdl_optimize_ctxs(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n>> +        xdl_cleanup_records(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n>>           xdl_free_ctx(&xe->xdf2);\n>>           xdl_free_ctx(&xe->xdf1);\n> \n> \n\n"},{"id":"534354","messageId":"99b28295-fb14-4f4d-98d9-2caa9be88e33@gmail.com","threadId":"64712","inReplyTo":"f9b10e71d23f8b4fa34dcffb371cf5a173760409.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 09/10] xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-21T15:01:12Z","receivedAt":"2026-01-21T15:01:16Z","isPatch":true,"body":"Hi Ezekiel\n\nOn 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Disentangle xdl_cleanup_records() from the classifier so that it can be\n> moved from xprepare.c into xdiffi.c.\n> \n> The classic diff is the only algorithm that needs to count the number\n> of times each line occurs in each file. Make xdl_cleanup_records()\n> count the number of lines instead of the classifier so it won't slow\n> down patience or histogram.\n\nHave you measured the speed up that this gives? It looks like it saves \nvery little work for the patience or histogram algorithms and means we \nnow make a second pass over the data in the myers case. If there is a \nreason to do this related to the rust conversion then that might be a \nmore convincing argument. As Rene has said already this isn't a \nparticularly interesting demonstration of struct IVec - it would be nice \nto see more of the API exercised.\n\nThanks\n\nPhillip\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xprepare.c | 52 +++++++++++++++++++++++++++++++++---------------\n>   xdiff/xtypes.h   |  1 +\n>   2 files changed, 37 insertions(+), 16 deletions(-)\n> \n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index d3cdb6ac02..b53a3b80c4 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -21,6 +21,7 @@\n>    */\n>   \n>   #include \"xinclude.h\"\n> +#include \"compat/ivec.h\"\n>   \n>   \n>   #define XDL_KPDIS_RUN 4\n> @@ -35,7 +36,6 @@ typedef struct s_xdlclass {\n>   \tstruct s_xdlclass *next;\n>   \txrecord_t rec;\n>   \tlong idx;\n> -\tlong len1, len2;\n>   } xdlclass_t;\n>   \n>   typedef struct s_xdlclassifier {\n> @@ -92,7 +92,7 @@ static void xdl_free_classifier(xdlclassifier_t *cf) {\n>   }\n>   \n>   \n> -static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n> +static int xdl_classify_record(xdlclassifier_t *cf, xrecord_t *rec) {\n>   \tsize_t hi;\n>   \txdlclass_t *rcrec;\n>   \n> @@ -113,13 +113,10 @@ static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n>   \t\t\t\treturn -1;\n>   \t\tcf->rcrecs[rcrec->idx] = rcrec;\n>   \t\trcrec->rec = *rec;\n> -\t\trcrec->len1 = rcrec->len2 = 0;\n>   \t\trcrec->next = cf->rchash[hi];\n>   \t\tcf->rchash[hi] = rcrec;\n>   \t}\n>   \n> -\t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n> -\n>   \trec->minimal_perfect_hash = (size_t)rcrec->idx;\n>   \n>   \treturn 0;\n> @@ -253,22 +250,44 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n>   \treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n>   }\n>   \n> +struct xoccurrence\n> +{\n> +\tsize_t file1, file2;\n> +};\n> +\n> +\n> +DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n> +\n>   \n>   /*\n>    * Try to reduce the problem complexity, discard records that have no\n>    * matches on the other file. Also, lines that have multiple matches\n>    * might be potentially discarded if they appear in a run of discardable.\n>    */\n> -static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n> -\tlong i, nm, mlim;\n> +static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n> +\tlong i;\n> +\tsize_t nm, mlim;\n>   \txrecord_t *recs;\n> -\txdlclass_t *rcrec;\n>   \tuint8_t *action1 = NULL, *action2 = NULL;\n> -\tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> +\tstruct IVec_xoccurrence occ;\n> +\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n>   \tint ret = 0;\n>   \tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n>   \tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n>   \n> +\tIVEC_INIT(occ);\n> +\tivec_zero(&occ, xe->mph_size);\n> +\n> +\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n> +\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n> +\t\tocc.ptr[mph1].file1 += 1;\n> +\t}\n> +\n> +\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n> +\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n> +\t\tocc.ptr[mph2].file2 += 1;\n> +\t}\n> +\n>   \t/*\n>   \t * Create temporary arrays that will help us decide if\n>   \t * changed[i] should remain false, or become true.\n> @@ -288,16 +307,14 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>   \tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>   \t\tmlim = XDL_MAX_EQLIMIT;\n>   \tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n> -\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> -\t\tnm = rcrec ? rcrec->len2 : 0;\n> +\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n>   \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>   \t}\n>   \n>   \tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>   \t\tmlim = XDL_MAX_EQLIMIT;\n>   \tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n> -\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> -\t\tnm = rcrec ? rcrec->len1 : 0;\n> +\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n>   \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>   \t}\n>   \n> @@ -332,6 +349,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n>   cleanup:\n>   \txdl_free(action1);\n>   \txdl_free(action2);\n> +\tivec_free(&occ);\n>   \n>   \treturn ret;\n>   }\n> @@ -387,18 +405,20 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>   \n>   \tfor (size_t i = 0; i < xe->xdf1.nrec; i++) {\n>   \t\txrecord_t *rec = &xe->xdf1.recs[i];\n> -\t\txdl_classify_record(1, &cf, rec);\n> +\t\txdl_classify_record(&cf, rec);\n>   \t}\n>   \n>   \tfor (size_t i = 0; i < xe->xdf2.nrec; i++) {\n>   \t\txrecord_t *rec = &xe->xdf2.recs[i];\n> -\t\txdl_classify_record(2, &cf, rec);\n> +\t\txdl_classify_record(&cf, rec);\n>   \t}\n>   \n> +\txe->mph_size = cf.count;\n> +\n>   \txdl_trim_ends(xe);\n>   \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n>   \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n> -\t    xdl_cleanup_records(&cf, xe) < 0) {\n> +\t    xdl_cleanup_records(xe, xpp->flags) < 0) {\n>   \n>   \t\txdl_free_ctx(&xe->xdf2);\n>   \t\txdl_free_ctx(&xe->xdf1);\n> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n> index a939396064..2528bd37e8 100644\n> --- a/xdiff/xtypes.h\n> +++ b/xdiff/xtypes.h\n> @@ -56,6 +56,7 @@ typedef struct s_xdfile {\n>   typedef struct s_xdfenv {\n>   \txdfile_t xdf1, xdf2;\n>   \tsize_t delta_start, delta_end;\n> +\tsize_t mph_size;\n>   } xdfenv_t;\n>   \n>   \n\n"},{"id":"534355","messageId":"2a31e36a-8e36-4544-a54b-d877a85af8a3@gmail.com","threadId":"64712","inReplyTo":"1dba6b34aa5c3eec06ae50a74d133c37b1d2404e.1767379944.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 10/10] xdiff: move xdl_cleanup_records() from xprepare.c to xdiffi.c","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-21T15:01:20Z","receivedAt":"2026-01-21T15:01:24Z","isPatch":true,"body":"Hi Ezekiel\n\nOn 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Only the classic diff uses xdl_cleanup_records(). Move it,\n> xdl_clean_mmatch(), and the macros to xdiffi.c and call\n> xdl_cleanup_records() inside of xdl_do_classic_diff(). This better\n> organizes the code related to the classic diff.\n\nI think calling xdl_cleanup_records() from inside xdl_do_classic_diff() \nmakes sense. I don't have a strong opinion either way on the code \nmovement. You should remove '#include \"compat/ivec.h\"' from xprepare.c \nif you're moving the only code that uses it out of that file.\n\nThanks\n\nPhillip\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xdiffi.c   | 180 ++++++++++++++++++++++++++++++++++++++++++++\n>   xdiff/xprepare.c | 191 +----------------------------------------------\n>   2 files changed, 181 insertions(+), 190 deletions(-)\n> \n> diff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\n> index e3196c7245..0f1fd7cf80 100644\n> --- a/xdiff/xdiffi.c\n> +++ b/xdiff/xdiffi.c\n> @@ -21,6 +21,7 @@\n>    */\n>   \n>   #include \"xinclude.h\"\n> +#include \"compat/ivec.h\"\n>   \n>   static size_t get_hash(xdfile_t *xdf, long index)\n>   {\n> @@ -33,6 +34,14 @@ static size_t get_hash(xdfile_t *xdf, long index)\n>   #define XDL_SNAKE_CNT 20\n>   #define XDL_K_HEUR 4\n>   \n> +#define XDL_KPDIS_RUN 4\n> +#define XDL_MAX_EQLIMIT 1024\n> +#define XDL_SIMSCAN_WINDOW 100\n> +\n> +#define DISCARD 0\n> +#define KEEP 1\n> +#define INVESTIGATE 2\n> +\n>   typedef struct s_xdpsplit {\n>   \tlong i1, i2;\n>   \tint min_lo, min_hi;\n> @@ -311,6 +320,175 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long lim1,\n>   }\n>   \n>   \n> +static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n> +\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n> +\n> +\t/*\n> +\t * Limits the window that is examined during the similar-lines\n> +\t * scan. The loops below stops when action[i - r] == KEEP\n> +\t * (line that has no match), but there are corner cases where\n> +\t * the loop proceed all the way to the extremities by causing\n> +\t * huge performance penalties in case of big files.\n> +\t */\n> +\tif (i - s > XDL_SIMSCAN_WINDOW)\n> +\t\ts = i - XDL_SIMSCAN_WINDOW;\n> +\tif (e - i > XDL_SIMSCAN_WINDOW)\n> +\t\te = i + XDL_SIMSCAN_WINDOW;\n> +\n> +\t/*\n> +\t * Scans the lines before 'i' to find a run of lines that either\n> +\t * have no match (action[j] == DISCARD) or have multiple matches\n> +\t * (action[j] == INVESTIGATE). Note that we always call this\n> +\t * function with action[i] == INVESTIGATE, so the current line\n> +\t * (i) is already a multimatch line.\n> +\t */\n> +\tfor (r = 1, rdis0 = 0, rpdis0 = 1; (i - r) >= s; r++) {\n> +\t\tif (action[i - r] == DISCARD)\n> +\t\t\trdis0++;\n> +\t\telse if (action[i - r] == INVESTIGATE)\n> +\t\t\trpdis0++;\n> +\t\telse if (action[i - r] == KEEP)\n> +\t\t\tbreak;\n> +\t\telse\n> +\t\t\tBUG(\"Illegal value for action[i - r]\");\n> +\t}\n> +\t/*\n> +\t * If the run before the line 'i' found only multimatch lines,\n> +\t * we return false and hence we don't make the current line (i)\n> +\t * discarded. We want to discard multimatch lines only when\n> +\t * they appear in the middle of runs with nomatch lines\n> +\t * (action[j] == DISCARD).\n> +\t */\n> +\tif (rdis0 == 0)\n> +\t\treturn 0;\n> +\tfor (r = 1, rdis1 = 0, rpdis1 = 1; (i + r) <= e; r++) {\n> +\t\tif (action[i + r] == DISCARD)\n> +\t\t\trdis1++;\n> +\t\telse if (action[i + r] == INVESTIGATE)\n> +\t\t\trpdis1++;\n> +\t\telse if (action[i + r] == KEEP)\n> +\t\t\tbreak;\n> +\t\telse\n> +\t\t\tBUG(\"Illegal value for action[i + r]\");\n> +\t}\n> +\t/*\n> +\t * If the run after the line 'i' found only multimatch lines,\n> +\t * we return false and hence we don't make the current line (i)\n> +\t * discarded.\n> +\t */\n> +\tif (rdis1 == 0)\n> +\t\treturn false;\n> +\trdis1 += rdis0;\n> +\trpdis1 += rpdis0;\n> +\n> +\treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n> +}\n> +\n> +struct xoccurrence\n> +{\n> +\tsize_t file1, file2;\n> +};\n> +\n> +\n> +DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n> +\n> +\n> +/*\n> + * Try to reduce the problem complexity, discard records that have no\n> + * matches on the other file. Also, lines that have multiple matches\n> + * might be potentially discarded if they appear in a run of discardable.\n> + */\n> +static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n> +\tlong i;\n> +\tsize_t nm, mlim;\n> +\txrecord_t *recs;\n> +\tuint8_t *action1 = NULL, *action2 = NULL;\n> +\tstruct IVec_xoccurrence occ;\n> +\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n> +\tint ret = 0;\n> +\tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n> +\tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n> +\n> +\tIVEC_INIT(occ);\n> +\tivec_zero(&occ, xe->mph_size);\n> +\n> +\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n> +\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n> +\t\tocc.ptr[mph1].file1 += 1;\n> +\t}\n> +\n> +\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n> +\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n> +\t\tocc.ptr[mph2].file2 += 1;\n> +\t}\n> +\n> +\t/*\n> +\t * Create temporary arrays that will help us decide if\n> +\t * changed[i] should remain false, or become true.\n> +\t */\n> +\tif (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n> +\t\tret = -1;\n> +\t\tgoto cleanup;\n> +\t}\n> +\tif (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n> +\t\tret = -1;\n> +\t\tgoto cleanup;\n> +\t}\n> +\n> +\t/*\n> +\t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n> +\t */\n> +\tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n> +\t\tmlim = XDL_MAX_EQLIMIT;\n> +\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n> +\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n> +\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t}\n> +\n> +\tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n> +\t\tmlim = XDL_MAX_EQLIMIT;\n> +\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n> +\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n> +\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t}\n> +\n> +\t/*\n> +\t * Use temporary arrays to decide if changed[i] should remain\n> +\t * false, or become true.\n> +\t */\n> +\txe->xdf1.nreff = 0;\n> +\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n> +\t     i <= dend1; i++, recs++) {\n> +\t\tif (action1[i] == KEEP ||\n> +\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, dend1))) {\n> +\t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n> +\t\t\t/* changed[i] remains false, i.e. keep */\n> +\t\t} else\n> +\t\t\txe->xdf1.changed[i] = true;\n> +\t\t\t/* i.e. discard */\n> +\t}\n> +\n> +\txe->xdf2.nreff = 0;\n> +\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n> +\t     i <= dend2; i++, recs++) {\n> +\t\tif (action2[i] == KEEP ||\n> +\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, dend2))) {\n> +\t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n> +\t\t\t/* changed[i] remains false, i.e. keep */\n> +\t\t} else\n> +\t\t\txe->xdf2.changed[i] = true;\n> +\t\t\t/* i.e. discard */\n> +\t}\n> +\n> +cleanup:\n> +\txdl_free(action1);\n> +\txdl_free(action2);\n> +\tivec_free(&occ);\n> +\n> +\treturn ret;\n> +}\n> +\n> +\n>   int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags)\n>   {\n>   \tlong ndiags;\n> @@ -318,6 +496,8 @@ int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags)\n>   \txdalgoenv_t xenv;\n>   \tint res;\n>   \n> +\txdl_cleanup_records(xe, flags);\n> +\n>   \t/*\n>   \t * Allocate and setup K vectors to be used by the differential\n>   \t * algorithm.\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index b53a3b80c4..3f555e29f4 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -24,14 +24,6 @@\n>   #include \"compat/ivec.h\"\n>   \n>   \n> -#define XDL_KPDIS_RUN 4\n> -#define XDL_MAX_EQLIMIT 1024\n> -#define XDL_SIMSCAN_WINDOW 100\n> -\n> -#define DISCARD 0\n> -#define KEEP 1\n> -#define INVESTIGATE 2\n> -\n>   typedef struct s_xdlclass {\n>   \tstruct s_xdlclass *next;\n>   \txrecord_t rec;\n> @@ -50,8 +42,6 @@ typedef struct s_xdlclassifier {\n>   } xdlclassifier_t;\n>   \n>   \n> -\n> -\n>   static int xdl_init_classifier(xdlclassifier_t *cf, long size, long flags) {\n>   \tmemset(cf, 0, sizeof(xdlclassifier_t));\n>   \n> @@ -186,175 +176,6 @@ void xdl_free_env(xdfenv_t *xe) {\n>   }\n>   \n>   \n> -static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n> -\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n> -\n> -\t/*\n> -\t * Limits the window that is examined during the similar-lines\n> -\t * scan. The loops below stops when action[i - r] == KEEP\n> -\t * (line that has no match), but there are corner cases where\n> -\t * the loop proceed all the way to the extremities by causing\n> -\t * huge performance penalties in case of big files.\n> -\t */\n> -\tif (i - s > XDL_SIMSCAN_WINDOW)\n> -\t\ts = i - XDL_SIMSCAN_WINDOW;\n> -\tif (e - i > XDL_SIMSCAN_WINDOW)\n> -\t\te = i + XDL_SIMSCAN_WINDOW;\n> -\n> -\t/*\n> -\t * Scans the lines before 'i' to find a run of lines that either\n> -\t * have no match (action[j] == DISCARD) or have multiple matches\n> -\t * (action[j] == INVESTIGATE). Note that we always call this\n> -\t * function with action[i] == INVESTIGATE, so the current line\n> -\t * (i) is already a multimatch line.\n> -\t */\n> -\tfor (r = 1, rdis0 = 0, rpdis0 = 1; (i - r) >= s; r++) {\n> -\t\tif (action[i - r] == DISCARD)\n> -\t\t\trdis0++;\n> -\t\telse if (action[i - r] == INVESTIGATE)\n> -\t\t\trpdis0++;\n> -\t\telse if (action[i - r] == KEEP)\n> -\t\t\tbreak;\n> -\t\telse\n> -\t\t\tBUG(\"Illegal value for action[i - r]\");\n> -\t}\n> -\t/*\n> -\t * If the run before the line 'i' found only multimatch lines,\n> -\t * we return false and hence we don't make the current line (i)\n> -\t * discarded. We want to discard multimatch lines only when\n> -\t * they appear in the middle of runs with nomatch lines\n> -\t * (action[j] == DISCARD).\n> -\t */\n> -\tif (rdis0 == 0)\n> -\t\treturn 0;\n> -\tfor (r = 1, rdis1 = 0, rpdis1 = 1; (i + r) <= e; r++) {\n> -\t\tif (action[i + r] == DISCARD)\n> -\t\t\trdis1++;\n> -\t\telse if (action[i + r] == INVESTIGATE)\n> -\t\t\trpdis1++;\n> -\t\telse if (action[i + r] == KEEP)\n> -\t\t\tbreak;\n> -\t\telse\n> -\t\t\tBUG(\"Illegal value for action[i + r]\");\n> -\t}\n> -\t/*\n> -\t * If the run after the line 'i' found only multimatch lines,\n> -\t * we return false and hence we don't make the current line (i)\n> -\t * discarded.\n> -\t */\n> -\tif (rdis1 == 0)\n> -\t\treturn false;\n> -\trdis1 += rdis0;\n> -\trpdis1 += rpdis0;\n> -\n> -\treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n> -}\n> -\n> -struct xoccurrence\n> -{\n> -\tsize_t file1, file2;\n> -};\n> -\n> -\n> -DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n> -\n> -\n> -/*\n> - * Try to reduce the problem complexity, discard records that have no\n> - * matches on the other file. Also, lines that have multiple matches\n> - * might be potentially discarded if they appear in a run of discardable.\n> - */\n> -static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n> -\tlong i;\n> -\tsize_t nm, mlim;\n> -\txrecord_t *recs;\n> -\tuint8_t *action1 = NULL, *action2 = NULL;\n> -\tstruct IVec_xoccurrence occ;\n> -\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n> -\tint ret = 0;\n> -\tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n> -\tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n> -\n> -\tIVEC_INIT(occ);\n> -\tivec_zero(&occ, xe->mph_size);\n> -\n> -\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n> -\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n> -\t\tocc.ptr[mph1].file1 += 1;\n> -\t}\n> -\n> -\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n> -\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n> -\t\tocc.ptr[mph2].file2 += 1;\n> -\t}\n> -\n> -\t/*\n> -\t * Create temporary arrays that will help us decide if\n> -\t * changed[i] should remain false, or become true.\n> -\t */\n> -\tif (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n> -\t\tret = -1;\n> -\t\tgoto cleanup;\n> -\t}\n> -\tif (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n> -\t\tret = -1;\n> -\t\tgoto cleanup;\n> -\t}\n> -\n> -\t/*\n> -\t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n> -\t */\n> -\tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n> -\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n> -\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> -\t}\n> -\n> -\tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n> -\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n> -\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> -\t}\n> -\n> -\t/*\n> -\t * Use temporary arrays to decide if changed[i] should remain\n> -\t * false, or become true.\n> -\t */\n> -\txe->xdf1.nreff = 0;\n> -\tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n> -\t     i <= dend1; i++, recs++) {\n> -\t\tif (action1[i] == KEEP ||\n> -\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->delta_start, dend1))) {\n> -\t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> -\t\t\txe->xdf1.changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> -\t}\n> -\n> -\txe->xdf2.nreff = 0;\n> -\tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n> -\t     i <= dend2; i++, recs++) {\n> -\t\tif (action2[i] == KEEP ||\n> -\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->delta_start, dend2))) {\n> -\t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> -\t\t\txe->xdf2.changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> -\t}\n> -\n> -cleanup:\n> -\txdl_free(action1);\n> -\txdl_free(action2);\n> -\tivec_free(&occ);\n> -\n> -\treturn ret;\n> -}\n> -\n> -\n>   /*\n>    * Early trim initial and terminal matching records.\n>    */\n> @@ -414,19 +235,9 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n>   \t}\n>   \n>   \txe->mph_size = cf.count;\n> +\txdl_free_classifier(&cf);\n>   \n>   \txdl_trim_ends(xe);\n> -\tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n> -\t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n> -\t    xdl_cleanup_records(xe, xpp->flags) < 0) {\n> -\n> -\t\txdl_free_ctx(&xe->xdf2);\n> -\t\txdl_free_ctx(&xe->xdf1);\n> -\t\txdl_free_classifier(&cf);\n> -\t\treturn -1;\n> -\t}\n> -\n> -\txdl_free_classifier(&cf);\n>   \n>   \treturn 0;\n>   }\n\n"},{"id":"534384","messageId":"CAH=ZcbCNeYATxqAeXcGd9kkHzJq2y5BpMrChSzb215EHAjHsbg@mail.gmail.com","threadId":"64712","inReplyTo":"20260119204010.GA3148606@coredump.intra.peff.net","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-21T21:00:15Z","receivedAt":"2026-01-21T21:00:30Z","isPatch":true,"body":"On Mon, Jan 19, 2026 at 1:40 PM Jeff King <peff@peff.net> wrote:\n>\n> On Mon, Jan 19, 2026 at 01:21:04PM -0700, Ezekiel Newren wrote:\n>\n> > Ok..., is there a way to pad a field to the largest size needed so\n> > that this also works on the harvard architecture? If C isn't even self\n> > consistent then how are these structs going to be passed between C and\n> > Rust (which is THE point of ivec)?\n>\n> If you make a union of the pointers, it will require the largest size\n> and the strictest alignment requirement. So:\n>\n>   struct foo {\n>         union {\n>                 void *v;\n>                 uint8_t *u8;\n>         } ptr;\n>         size_t len;\n>   };\n>\n> would be a single struct you could use to store a void pointer _or_ a u8\n> pointer. The one thing you shouldn't do there, though, is assign via one\n> union member and read from the other. So I don't know if that helps you\n> or not (I confess I have not followed this rust discussion at all, and\n> know nothing about rust/c ABI compatibility, and just got roped in on C\n> esoterica).\n>\n> > Or do we just tell the arcane Harvard architecture \"too bad\" Git won't\n> > run on it anymore?\n>\n> Minor nit: the Harvard architecture is one where function pointers are\n> not the same as data pointers. An int/char distinction can happen even\n> on more common (von Neumann) machines.\n>\n> But I think we can rephrase your question as: are there real-world\n> machines we care about that will have different pointer sizes, or can we\n> ignore this issue for practical purposes?\n>\n> I don't know the answer. I suspect it probably is OK for Git not to run\n> on the machines mentioned in that C faq. But:\n>\n>   1. Sometimes there are subtle implications of undefined behavior that\n>      may cause a compiler (even for a sensible machine) to do unexpected\n>      things. I don't know offhand if that is the case here.\n>\n>   2. There are some modern platforms in which pointers are a bit more\n>      opaque than just numeric addresses. For example, we've had a few\n>      patches dealing with questionable pointer usage to make things work\n>      on CHERI Arm systems. I'm not sure if any of that would matter\n>      here, though (IIRC, it was mostly that pointers were unexpectedly\n>      large and had matching alignment requirements, but all of them\n>      equally so).\n>\n> -Peff\n\nWhat about adding clar unit tests to make sure that different ivec\ntypes have the same size and layout? e.g. sizeof(IVec_c_void) ==\nsizeof(IVec_u8);\nsizeof(IVec_c_void) == sizeof(IVec_u16);\nsizeof(IVec_c_void) == sizeof(IVec_u32);\nsizeof(IVec_c_void) == sizeof(IVec_u64);\n...\n\nAs well as other tests for ivec.\n"},{"id":"534387","messageId":"CAH=ZcbAb-iQM81Sd79KtFV0nf1gv4gfnBBvJ2AvcxCTN9xOr7Q@mail.gmail.com","threadId":"64712","inReplyTo":"1c46f551-0040-481e-9476-bc1b85f92636@gmail.com","subject":"Re: [PATCH 02/10] xdiff: make classic diff explicit by creating xdl_do_classic_diff()","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-21T21:05:42Z","receivedAt":"2026-01-21T21:05:56Z","isPatch":true,"body":"On Tue, Jan 20, 2026 at 8:01 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > Later patches will prepare xdl_cleanup_records() to be moved into xdiffi.c\n> > since only the classic diff uses that function.\n>\n> I assume that's to make it easier to covert the myers implementation to\n> rust without affecting the rest of the code? If so it would be nice to\n> say that.\n\nMaking it easier to port to Rust is a side effect. The primary goal is\nto simplify the job of xprepare to only parsing and hashing lines in a\nfile. xdl_cleanup_records() is only used by classic diff\n(myers/minimal) which means it doesn't belong in xprepare because it's\npart of a diff algorithm and isn't relevant to preparing the file for\na diff algorithm. Perhaps xdl_trim_ends() should be moved into\nxdl_do_diff() too...\n\n> > Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> > +int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n> > +             xdfenv_t *xe) {\n> > +     int res;\n> > +\n> > +     if (xdl_prepare_env(mf1, mf2, xpp, xe) < 0)\n> > +             return -1;\n> > +\n> > +     if (XDF_DIFF_ALG(xpp->flags) == XDF_PATIENCE_DIFF) {\n> > +             res = xdl_do_patience_diff(xpp, xe);\n> > +             goto out;\n> > +     }\n> > +\n> > +     if (XDF_DIFF_ALG(xpp->flags) == XDF_HISTOGRAM_DIFF) {\n> > +             res = xdl_do_histogram_diff(xpp, xe);\n> > +             goto out;\n> > +     }\n> > +\n> > +     res = xdl_do_classic_diff(xe, xpp->flags);\n>\n> This might be clearer that we're calling only one of the three functions\n> if we wrote this as\n>\n>         if (XDF_DIFF_ALG(xpp->flags) == XDIF_PATIENCE_DIFF)\n>                 res = xdl_do_patience_diff(xpp, xe);\n>         else if (XDF_DIFF_ALG(xpp->flags) == XDF_HISTOGRAM_DIFF)\n>                 res = xdl_do_histogram_diff(xpp, xe);\n>         else\n>                 res = xdl_do_classic_diff(xe, xpp->flags);\n>\n> and then we can drop the out: label\n\nIn a later cleanup, I make this exact change :)\n"},{"id":"534391","messageId":"CAH=ZcbCbz6MB9-9Ehskk2+27GMXXewmAzRcGyN_bBi8s5Ksxjg@mail.gmail.com","threadId":"64712","inReplyTo":"208da094-8a5d-4f16-b42b-5d5204576b5f@gmail.com","subject":"Re: [PATCH 03/10] xdiff: don't waste time guessing the number of lines","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-21T21:12:40Z","receivedAt":"2026-01-21T21:12:54Z","isPatch":true,"body":"On Tue, Jan 20, 2026 at 8:02 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > All lines must be read anyway, so classify them after they're read in.\n> > Also move the memset() into xdl_init_classifier().\n>\n> So instead of looping over the input lines one and a bit times (the bit\n> being from xdl_guess_lines) we now loop over them twice as we split them\n> first and then classify them in a separate loop. It does save some work\n> not to call xdl_guess_lines but it is unclear if that offsets\n> classifying them in a separate loop.\n>\n> > +     for (size_t i = 0; i < xe->xdf1.nrec; i++) {\n> > +             xrecord_t *rec = &xe->xdf1.recs[i];\n> > +             xdl_classify_record(1, &cf, rec);\n>\n> We seem to have lost the error handling if xdl_classify_record() fails.\n\nThe error handling was not \"lost\" it was deliberately removed. The\nonly way in which xdl_classify_record() could fail is by a failed\nmemory allocation. On the Rust side this would result in a panic\n(panic means something different in Rust vs C) in which case C could\nnot possibly recover. Also for operations like Vec.push() in Rust it's\nassumed that memory management functions will never fail and if they\ndo they crash the program with no chance of recovery (unless you\naccount for panic unwinding which is really ugly). It seems a lot of\narguments about ivec and my xdiff cleanups are \"We don't do things\nthis way in Git/C\" I'm aware of many of these arguments and I'm trying\nto address them with a more specific answer of \"Yes, but that's not\nhow things are done in Rust and all of this is to prepare the code for\nconversion to Rust and some things shouldn't, or even, cannot be done\nthe C way in Rust.\"\n"},{"id":"534392","messageId":"20260121212024.GC723458@coredump.intra.peff.net","threadId":"64712","inReplyTo":"CAH=ZcbCNeYATxqAeXcGd9kkHzJq2y5BpMrChSzb215EHAjHsbg@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-01-21T21:20:24Z","receivedAt":"2026-01-21T21:20:26Z","isPatch":true,"body":"On Wed, Jan 21, 2026 at 02:00:15PM -0700, Ezekiel Newren wrote:\n\n> What about adding clar unit tests to make sure that different ivec\n> types have the same size and layout? e.g. sizeof(IVec_c_void) ==\n> sizeof(IVec_u8);\n> sizeof(IVec_c_void) == sizeof(IVec_u16);\n> sizeof(IVec_c_void) == sizeof(IVec_u32);\n> sizeof(IVec_c_void) == sizeof(IVec_u64);\n> ...\n> \n> As well as other tests for ivec.\n\nI'm a little hesitant in general to have run-time tests for properties\naround undefined behavior, just because the compiler is allowed to do a\nlot of tricky things when we get into that territory. Plus it is not\nreally _solving_ the problem, but perhaps just alerting us slightly\nsooner than the production code itself crashing and burning.\n\nYou'd also need to check the pointer field sizes directly due to\npadding. I don't think it's sufficient, due to padding. If one pointer\nis 4 bytes and another is 8 (for example), but the element afterwards\nrequires 8-byte alignment, then the compiler will have to insert 4 bytes\nof padding. And the resulting struct size will be the same. You'd have\nto more directly check that sizeof(uint_t*) == sizeof(void *), I think.\n\nSo I dunno. I am not a compiler expert, nor a rust expert, nor really\nknow anything about rust/C ABI boundaries. There might be no problem at\nall here, and I'm only commenting on what I know is possible (albeit\nunlikely) from the C side.\n\n-Peff\n"},{"id":"534393","messageId":"xmqqv7gur6t4.fsf@gitster.g","threadId":"64712","inReplyTo":"20260121212024.GC723458@coredump.intra.peff.net","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-01-21T21:31:51Z","receivedAt":"2026-01-21T21:31:57Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> On Wed, Jan 21, 2026 at 02:00:15PM -0700, Ezekiel Newren wrote:\n>\n>> What about adding clar unit tests to make sure that different ivec\n>> types have the same size and layout? e.g. sizeof(IVec_c_void) ==\n>> sizeof(IVec_u8);\n>> sizeof(IVec_c_void) == sizeof(IVec_u16);\n>> sizeof(IVec_c_void) == sizeof(IVec_u32);\n>> sizeof(IVec_c_void) == sizeof(IVec_u64);\n>> ...\n>> \n>> As well as other tests for ivec.\n>\n> I'm a little hesitant in general to have run-time tests for properties\n> around undefined behavior, just because the compiler is allowed to do a\n> lot of tricky things when we get into that territory. Plus it is not\n> really _solving_ the problem, but perhaps just alerting us slightly\n> sooner than the production code itself crashing and burning.\n\nYup, by definition, testing undefined behaviour with code is more or\nless pointless.  Implementation defined behaviour, maybe, but not\nundefined ones, please.\n\nI thought you already gave them that having different possibilities\nin a union would work correctly, but perhaps I was reading a\ndifferent thread?  I dunno...\n\n\n\n"},{"id":"534394","messageId":"CAH=ZcbAiGONrOyma7YjNKKLqNFoisU5LG=nGWjtOJ1wLfqX4cQ@mail.gmail.com","threadId":"64712","inReplyTo":"08318339-03c3-4068-92fa-7a711bd13da0@gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-21T21:39:24Z","receivedAt":"2026-01-21T21:39:39Z","isPatch":true,"body":"On Tue, Jan 20, 2026 at 7:06 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 15/01/2026 15:55, Ezekiel Newren wrote:\n> > On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n> >>> +void ivec_reserve(void *self_, size_t additional)\n> >>> +{\n> >>> +     struct IVec_c_void *self = self_;\n> >>> +\n> >>> +     size_t growby = 128;\n> >>> +     if (self->capacity > growby)\n> >>> +             growby = self->capacity;\n> >>> +     if (additional > growby)\n> >>> +             growby = additional;\n> >>\n> >> This growth strategy differs from both ALLOC_GROW() and\n> >> XDL_ALLOC_GROW(), if there isn't a good reason for that we should\n> >> perhaps just use ALLOC_GROW() here.\n> >\n> > XDL_ALLOW_GROW() can't be used because the pointer is always a void*\n> > in this function.\n>\n> Oh right. I'm not sure that's not a reason to use a different growth\n> strategy though. The minimum size of 128 elements is probably good for\n> the xdiff code that creates arrays with one element per line but if this\n> is supposed to be for general use it is going to waste space when we're\n> allocating a lot of small arrays. ALLOC_GROW() uses alloc_nr() to\n> calculate the new side so perhaps we could use that here?\n\nIf ivec_reserve() isn't suitable then ivec_reserve_exact() should be\nused instead.\n\n> >>> +void ivec_push(void *self_, const void *value)\n> >>> +{\n> >>> +     struct IVec_c_void *self = self_;\n> >>> +     void *dst = NULL;\n> >>> +\n> >>> +     if (self->length == self->capacity)\n> >>> +             ivec_reserve(self, 1);\n> >>> +\n> >>> +     dst = (uint8_t*)self->ptr + self->length * self->element_size;\n> >>> +     memcpy(dst, value, self->element_size);\n> >>\n> >> If self->element_size was a compile time constant the compiler could\n> >> easily optimize this call away. I'm not sure that is easy to achieve though.\n> >\n> > The problem is that I didn't want all of ivec to be macros that looked\n> > like function calls. I wanted to minimize use of macros so that it was\n> > easier to port and verify that the Rust implementation matches the\n> > behavior of the C implementation.\n>\n> I think that's a reasonable concern. So is the plan to have a parallel\n> rust implementation of these functions rather than call the C\n> implementation from rust?\n\nYes, the Rust implementation will be independent of the C\nimplementation, but will behave the same way. That's why I'm calling\nit an interoperable vec as opposed to a compatible vec. Rust can't\ncall the C ivec functions and C can't call the Rust ivec functions,\nbut they'll behave the same way.\n\n> >>> +void ivec_free(void *self_)\n> >>\n> >> Normally we'd call a like this that free the allocations and\n> >> re-initializes the members ivec_clear()\n> >\n> > In Rust Vec.clear() means to set length to zero, but leaves the\n> > allocation alone. The reason why I'm zeroing the struct is to help\n> > avoid FFI issues. If not zero then what should the members be set to,\n> > to indicate that using the struct is not valid anymore? In Rust an\n> > object is freed when it goes out of scope and _cannot_ be accessed\n> > afterward.\n\nMaybe I should call this ivec_drop(). Though the notion of explicitly\nfreeing an object in Rust is _almost_ nonsense. The way you free\nsomething in Rust is to let it go out of scope.\n\n> I'm aware that Vec::clear() has different semantics (it does what\n> strbuf_reset() does). That's unfortunate but this function has different\n> semantics to all the other *_free() functions in git. Our coding\n> guidelines say\n>\n>   - There are several common idiomatic names for functions performing\n>     specific tasks on a structure `S`:\n>\n>      - `S_init()` initializes a structure without allocating the\n>        structure itself.\n>\n>      - `S_release()` releases a structure's contents without freeing the\n>        structure.\n>\n>      - `S_clear()` is equivalent to `S_release()` followed by `S_init()`\n>        such that the structure is directly usable after clearing it. When\n>        `S_clear()` is provided, `S_init()` shall not allocate resources\n>        that need to be released again.\n>\n>      - `S_free()` releases a structure's contents and frees the\n>        structure.\n>\n> As we write more rust code and so wrap more of our existing structs\n> we're going to be wrapping C code that uses the definitions above so I\n> think we should do the same with struct IVec_*.\n\nI disagree. IVec isn't a wrapper around an existing struct. ivec is\nmeant to very closely mimic Rust's Vec while guaranteeing\ninteroperability. For things like strbuf I haven't conceived of a\nsolution for that yet. Making ivec diverge from Rust's Vec will result\nin POLA violations due to different behavior when refactoring an\nIVec<your_type_here> to Vec<your_type_here>.\n\n> >>> diff --git a/compat/ivec.h b/compat/ivec.h\n> >>> new file mode 100644\n> >>> index 0000000000..654a05c506\n> >>> --- /dev/null\n> >>> +++ b/compat/ivec.h\n> >>> @@ -0,0 +1,52 @@\n> >>> +#ifndef IVEC_H\n> >>> +#define IVEC_H\n> >>> +\n> >>> +#include <git-compat-util.h>\n> >>\n> >> It would be nice to have some documentation in this header, see the\n> >> examples in strvec.h and hashmap.h\n> >>\n> >>> +#define IVEC_INIT(variable) ivec_init(&(variable), sizeof(*(variable).ptr))\n> >>\n> >> This is a bit cumbersome to use compared to our usual *_INIT macros. I'm\n> >> struggling to see how we can make it nicer though as DEFINE_IVEC_TYPE\n> >> cannot define a per-type initializer macro and I we cannot initialize\n> >> the element size without knowing the type.\n> >\n> > I don't see what's cumbersome about it. Maybe an example use case\n> > would clarify things.\n>\n> It is cumbersome because it separates the initialization from the\n> declaration. Normally our *_INIT macros are initializer lists so we can\n> write\n>\n>         struct strbuf = STRBUF_INIT;\n>\n> which keeps the declaration and initialization together. Although\n> they're on adjacent lines in your example in real code the\n> initialization likely to be separated from the declaration by other\n> variable declarations.\n\nAh I see what you mean now. I'll experiment with making IVEC_INIT()\nwork like that. One wrinkle is that STRBUF_INIT is a single concrete\ntype whereas IVEC_INIT() is meant for generic types.\n"},{"id":"534395","messageId":"CAH=ZcbDYjW5=8jNOA1=Cw8eaAuKBMshAn2nAFgUCGHmBC=zGEA@mail.gmail.com","threadId":"64712","inReplyTo":"xmqqv7gur6t4.fsf@gitster.g","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-01-21T21:45:58Z","receivedAt":"2026-01-21T21:46:13Z","isPatch":true,"body":"On Wed, Jan 21, 2026 at 2:31 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Jeff King <peff@peff.net> writes:\n>\n> > On Wed, Jan 21, 2026 at 02:00:15PM -0700, Ezekiel Newren wrote:\n> >\n> >> What about adding clar unit tests to make sure that different ivec\n> >> types have the same size and layout? e.g. sizeof(IVec_c_void) ==\n> >> sizeof(IVec_u8);\n> >> sizeof(IVec_c_void) == sizeof(IVec_u16);\n> >> sizeof(IVec_c_void) == sizeof(IVec_u32);\n> >> sizeof(IVec_c_void) == sizeof(IVec_u64);\n> >> ...\n> >>\n> >> As well as other tests for ivec.\n> >\n> > I'm a little hesitant in general to have run-time tests for properties\n> > around undefined behavior, just because the compiler is allowed to do a\n> > lot of tricky things when we get into that territory. Plus it is not\n> > really _solving_ the problem, but perhaps just alerting us slightly\n> > sooner than the production code itself crashing and burning.\n>\n> Yup, by definition, testing undefined behaviour with code is more or\n> less pointless.  Implementation defined behaviour, maybe, but not\n> undefined ones, please.\n>\n> I thought you already gave them that having different possibilities\n> in a union would work correctly, but perhaps I was reading a\n> different thread?  I dunno...\n>\n\nIn my opinion the proper solution to this is to document that any\nplatform with different size pointers for different types is not\nsupported by Git. Which would make using Git on those platforms \"use\nat your own risk\".\n"},{"id":"534437","messageId":"8f7ec565-f91c-4950-91d7-781a31d6fb6e@gmail.com","threadId":"64712","inReplyTo":"CAH=ZcbCbz6MB9-9Ehskk2+27GMXXewmAzRcGyN_bBi8s5Ksxjg@mail.gmail.com","subject":"Re: [PATCH 03/10] xdiff: don't waste time guessing the number of lines","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-22T10:16:02Z","receivedAt":"2026-01-22T10:16:08Z","isPatch":true,"body":"On 21/01/2026 21:12, Ezekiel Newren wrote:\n> On Tue, Jan 20, 2026 at 8:02 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>\n>> On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n>>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>>\n>>> All lines must be read anyway, so classify them after they're read in.\n>>> Also move the memset() into xdl_init_classifier().\n>>\n>> So instead of looping over the input lines one and a bit times (the bit\n>> being from xdl_guess_lines) we now loop over them twice as we split them\n>> first and then classify them in a separate loop. It does save some work\n>> not to call xdl_guess_lines but it is unclear if that offsets\n>> classifying them in a separate loop.\n>>\n>>> +     for (size_t i = 0; i < xe->xdf1.nrec; i++) {\n>>> +             xrecord_t *rec = &xe->xdf1.recs[i];\n>>> +             xdl_classify_record(1, &cf, rec);\n>>\n>> We seem to have lost the error handling if xdl_classify_record() fails.\n> \n> The error handling was not \"lost\" it was deliberately removed. \n\nThat's the sort of thing that needs to be explained in the commit message.\n\n> The\n> only way in which xdl_classify_record() could fail is by a failed\n> memory allocation. On the Rust side this would result in a panic\n> (panic means something different in Rust vs C) in which case C could\n> not possibly recover.\n\nThere is no rust code in xdiff at the moment so we don't panic on \nfailure. In git we'll die() because xdl_malloc() and friends are defined \nas xmalloc() etc. which die on allocation failure. However anyone else \npicking up this code and using a different allocator that does not die \non allocation failure will expect the error to be propagated.\n\nIf you want to stop supporting other allocators then you should propose \na patch to do so, not silently slip the change into this patch.\n\nThanks\n\nPhillip\n\n> Also for operations like Vec.push() in Rust it's\n> assumed that memory management functions will never fail and if they\n> do they crash the program with no chance of recovery (unless you\n> account for panic unwinding which is really ugly). It seems a lot of\n> arguments about ivec and my xdiff cleanups are \"We don't do things\n> this way in Git/C\" I'm aware of many of these arguments and I'm trying\n> to address them with a more specific answer of \"Yes, but that's not\n> how things are done in Rust and all of this is to prepare the code for\n> conversion to Rust and some things shouldn't, or even, cannot be done\n> the C way in Rust.\"\n\n"},{"id":"534753","messageId":"79ea1b9c-47cd-4702-bcb2-05417adf9eae@gmail.com","threadId":"64712","inReplyTo":"e9a031fd-072d-4810-b7e0-0d64ffedce10@gmail.com","subject":"Re: [PATCH 07/10] xdiff: replace xdfile_t.dstart with xdfenv_t.delta_start","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-28T10:51:49Z","receivedAt":"2026-01-28T10:51:53Z","isPatch":true,"body":"On 20/01/2026 16:32, Phillip Wood wrote:\n> On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>\n>> Placing delta_start in xdfenv_t instead of xdfile_t provides a more\n>> appropriate context since this variable only makes sense with a pair\n>> of files. View with --color-words.\n> \n> So as dstart and dend must be the same for both files we now store the \n> values once in xdfenv_t. \n\nExcept it's only dstart that's the same, dend is different because it \nconvinently stores an index, not an offset from the end. Having realized \nthat, moving them to xdfenv_t makes less sense as having to calculate \nthe dend index from an offset from the end of the array each time is a \npain and sooner or later we'll make a mistake.\n\nThanks\n\nPhillip\n\n> That explains why we start passing xdfenv_t \n> around rather than xdfile_t in patch 5.\n> \n> Thanks\n> \n> Phillip\n> \n>> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>> ---\n>>   xdiff/xhistogram.c |  4 ++--\n>>   xdiff/xpatience.c  |  4 ++--\n>>   xdiff/xprepare.c   | 17 +++++++++--------\n>>   xdiff/xtypes.h     |  3 ++-\n>>   4 files changed, 15 insertions(+), 13 deletions(-)\n>>\n>> diff --git a/xdiff/xhistogram.c b/xdiff/xhistogram.c\n>> index 5ae1282c27..eb6a52d9ba 100644\n>> --- a/xdiff/xhistogram.c\n>> +++ b/xdiff/xhistogram.c\n>> @@ -365,6 +365,6 @@ out:\n>>   int xdl_do_histogram_diff(xpparam_t const *xpp, xdfenv_t *env)\n>>   {\n>>       return histogram_diff(xpp, env,\n>> -        env->xdf1.dstart + 1, env->xdf1.dend - env->xdf1.dstart + 1,\n>> -        env->xdf2.dstart + 1, env->xdf2.dend - env->xdf2.dstart + 1);\n>> +        env->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n>> +        env->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n>>   }\n>> diff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\n>> index 2bce07cf48..bd0ffbb417 100644\n>> --- a/xdiff/xpatience.c\n>> +++ b/xdiff/xpatience.c\n>> @@ -374,6 +374,6 @@ static int patience_diff(xpparam_t const *xpp, \n>> xdfenv_t *env,\n>>   int xdl_do_patience_diff(xpparam_t const *xpp, xdfenv_t *env)\n>>   {\n>>       return patience_diff(xpp, env,\n>> -        env->xdf1.dstart + 1, env->xdf1.dend - env->xdf1.dstart + 1,\n>> -        env->xdf2.dstart + 1, env->xdf2.dend - env->xdf2.dstart + 1);\n>> +        env->delta_start + 1, env->xdf1.dend - env->delta_start + 1,\n>> +        env->delta_start + 1, env->xdf2.dend - env->delta_start + 1);\n>>   }\n>> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n>> index 06b6a6f804..e88468e74c 100644\n>> --- a/xdiff/xprepare.c\n>> +++ b/xdiff/xprepare.c\n>> @@ -173,7 +173,6 @@ static int xdl_prepare_ctx(mmfile_t *mf, xdfile_t \n>> *xdf, uint64_t flags) {\n>>       xdf->changed += 1;\n>>       xdf->nreff = 0;\n>> -    xdf->dstart = 0;\n>>       xdf->dend = xdf->nrec - 1;\n>>       return 0;\n>> @@ -287,7 +286,7 @@ static int xdl_cleanup_records(xdlclassifier_t \n>> *cf, xdfenv_t *xe) {\n>>        */\n>>       if ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>>           mlim = XDL_MAX_EQLIMIT;\n>> -    for (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart]; \n>> i <= xe->xdf1.dend; i++, recs++) {\n>> +    for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; \n>> i <= xe->xdf1.dend; i++, recs++) {\n>>           rcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>>           nm = rcrec ? rcrec->len2 : 0;\n>>           action1[i] = (nm == 0) ? DISCARD: (nm >= mlim && ! \n>> need_min) ? INVESTIGATE: KEEP;\n>> @@ -295,7 +294,7 @@ static int xdl_cleanup_records(xdlclassifier_t \n>> *cf, xdfenv_t *xe) {\n>>       if ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>>           mlim = XDL_MAX_EQLIMIT;\n>> -    for (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart]; \n>> i <= xe->xdf2.dend; i++, recs++) {\n>> +    for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; \n>> i <= xe->xdf2.dend; i++, recs++) {\n>>           rcrec = cf->rcrecs[recs->minimal_perfect_hash];\n>>           nm = rcrec ? rcrec->len1 : 0;\n>>           action2[i] = (nm == 0) ? DISCARD: (nm >= mlim && ! \n>> need_min) ? INVESTIGATE: KEEP;\n>> @@ -306,10 +305,10 @@ static int xdl_cleanup_records(xdlclassifier_t \n>> *cf, xdfenv_t *xe) {\n>>        * false, or become true.\n>>        */\n>>       xe->xdf1.nreff = 0;\n>> -    for (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart];\n>> +    for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n>>            i <= xe->xdf1.dend; i++, recs++) {\n>>           if (action1[i] == KEEP ||\n>> -            (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, \n>> i, xe->xdf1.dstart, xe->xdf1.dend))) {\n>> +            (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, \n>> i, xe->delta_start, xe->xdf1.dend))) {\n>>               xe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n>>               /* changed[i] remains false, i.e. keep */\n>>           } else\n>> @@ -318,10 +317,10 @@ static int xdl_cleanup_records(xdlclassifier_t \n>> *cf, xdfenv_t *xe) {\n>>       }\n>>       xe->xdf2.nreff = 0;\n>> -    for (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart];\n>> +    for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n>>            i <= xe->xdf2.dend; i++, recs++) {\n>>           if (action2[i] == KEEP ||\n>> -            (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, \n>> i, xe->xdf2.dstart, xe->xdf2.dend))) {\n>> +            (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, \n>> i, xe->delta_start, xe->xdf2.dend))) {\n>>               xe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n>>               /* changed[i] remains false, i.e. keep */\n>>           } else\n>> @@ -348,7 +347,7 @@ static void xdl_trim_ends(xdfenv_t *xe)\n>>           size_t mph1 = xe->xdf1.recs[i].minimal_perfect_hash;\n>>           size_t mph2 = xe->xdf2.recs[i].minimal_perfect_hash;\n>>           if (mph1 != mph2) {\n>> -            xe->xdf1.dstart = xe->xdf2.dstart = (ssize_t)i;\n>> +            xe->delta_start = (ssize_t)i;\n>>               lim -= i;\n>>               break;\n>>           }\n>> @@ -370,6 +369,8 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, \n>> xpparam_t const *xpp,\n>>               xdfenv_t *xe) {\n>>       xdlclassifier_t cf;\n>> +    xe->delta_start = 0;\n>> +\n>>       if (xdl_prepare_ctx(mf1, &xe->xdf1, xpp->flags) < 0) {\n>>           return -1;\n>> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h\n>> index 979586f20a..bda1f85eb0 100644\n>> --- a/xdiff/xtypes.h\n>> +++ b/xdiff/xtypes.h\n>> @@ -48,7 +48,7 @@ typedef struct s_xrecord {\n>>   typedef struct s_xdfile {\n>>       xrecord_t *recs;\n>>       size_t nrec;\n>> -    ptrdiff_t dstart, dend;\n>> +    ptrdiff_t dend;\n>>       bool *changed;\n>>       size_t *reference_index;\n>>       size_t nreff;\n>> @@ -56,6 +56,7 @@ typedef struct s_xdfile {\n>>   typedef struct s_xdfenv {\n>>       xdfile_t xdf1, xdf2;\n>> +    size_t delta_start;\n>>   } xdfenv_t;\n> \n> \n\n"},{"id":"534754","messageId":"7513600a-24bb-4ea2-847f-8e9a1dbe7ef3@gmail.com","threadId":"64712","inReplyTo":"2a31e36a-8e36-4544-a54b-d877a85af8a3@gmail.com","subject":"Re: [PATCH 10/10] xdiff: move xdl_cleanup_records() from xprepare.c to xdiffi.c","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-28T10:56:16Z","receivedAt":"2026-01-28T10:56:19Z","isPatch":true,"body":"On 21/01/2026 15:01, Phillip Wood wrote:\n> Hi Ezekiel\n> \n> On 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>\n>> Only the classic diff uses xdl_cleanup_records(). Move it,\n>> xdl_clean_mmatch(), and the macros to xdiffi.c and call\n>> xdl_cleanup_records() inside of xdl_do_classic_diff(). This better\n>> organizes the code related to the classic diff.\n> \n> I think calling xdl_cleanup_records() from inside xdl_do_classic_diff() \n> makes sense. I don't have a strong opinion either way on the code \n> movement.\n\nHaving thought about it I'm not so sure the code movement here makes \nsense. Having utility functions in a separate file is perfectly \nreasonable (afterall xprepare.c existed before the histogram and \npatientce algorithms were added). It's not like the code xdiffi.c is \nonly about the myers diff there is generic code for diff sliders in \nthere as well.\n\nThanks\n\nPhillip\n\n  You should remove '#include \"compat/ivec.h\"' from xprepare.c\n> if you're moving the only code that uses it out of that file.\n> \n> Thanks\n> \n> Phillip\n> \n>> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>> ---\n>>   xdiff/xdiffi.c   | 180 ++++++++++++++++++++++++++++++++++++++++++++\n>>   xdiff/xprepare.c | 191 +----------------------------------------------\n>>   2 files changed, 181 insertions(+), 190 deletions(-)\n>>\n>> diff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\n>> index e3196c7245..0f1fd7cf80 100644\n>> --- a/xdiff/xdiffi.c\n>> +++ b/xdiff/xdiffi.c\n>> @@ -21,6 +21,7 @@\n>>    */\n>>   #include \"xinclude.h\"\n>> +#include \"compat/ivec.h\"\n>>   static size_t get_hash(xdfile_t *xdf, long index)\n>>   {\n>> @@ -33,6 +34,14 @@ static size_t get_hash(xdfile_t *xdf, long index)\n>>   #define XDL_SNAKE_CNT 20\n>>   #define XDL_K_HEUR 4\n>> +#define XDL_KPDIS_RUN 4\n>> +#define XDL_MAX_EQLIMIT 1024\n>> +#define XDL_SIMSCAN_WINDOW 100\n>> +\n>> +#define DISCARD 0\n>> +#define KEEP 1\n>> +#define INVESTIGATE 2\n>> +\n>>   typedef struct s_xdpsplit {\n>>       long i1, i2;\n>>       int min_lo, min_hi;\n>> @@ -311,6 +320,175 @@ int xdl_recs_cmp(xdfile_t *xdf1, long off1, long \n>> lim1,\n>>   }\n>> +static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, \n>> long e) {\n>> +    long r, rdis0, rpdis0, rdis1, rpdis1;\n>> +\n>> +    /*\n>> +     * Limits the window that is examined during the similar-lines\n>> +     * scan. The loops below stops when action[i - r] == KEEP\n>> +     * (line that has no match), but there are corner cases where\n>> +     * the loop proceed all the way to the extremities by causing\n>> +     * huge performance penalties in case of big files.\n>> +     */\n>> +    if (i - s > XDL_SIMSCAN_WINDOW)\n>> +        s = i - XDL_SIMSCAN_WINDOW;\n>> +    if (e - i > XDL_SIMSCAN_WINDOW)\n>> +        e = i + XDL_SIMSCAN_WINDOW;\n>> +\n>> +    /*\n>> +     * Scans the lines before 'i' to find a run of lines that either\n>> +     * have no match (action[j] == DISCARD) or have multiple matches\n>> +     * (action[j] == INVESTIGATE). Note that we always call this\n>> +     * function with action[i] == INVESTIGATE, so the current line\n>> +     * (i) is already a multimatch line.\n>> +     */\n>> +    for (r = 1, rdis0 = 0, rpdis0 = 1; (i - r) >= s; r++) {\n>> +        if (action[i - r] == DISCARD)\n>> +            rdis0++;\n>> +        else if (action[i - r] == INVESTIGATE)\n>> +            rpdis0++;\n>> +        else if (action[i - r] == KEEP)\n>> +            break;\n>> +        else\n>> +            BUG(\"Illegal value for action[i - r]\");\n>> +    }\n>> +    /*\n>> +     * If the run before the line 'i' found only multimatch lines,\n>> +     * we return false and hence we don't make the current line (i)\n>> +     * discarded. We want to discard multimatch lines only when\n>> +     * they appear in the middle of runs with nomatch lines\n>> +     * (action[j] == DISCARD).\n>> +     */\n>> +    if (rdis0 == 0)\n>> +        return 0;\n>> +    for (r = 1, rdis1 = 0, rpdis1 = 1; (i + r) <= e; r++) {\n>> +        if (action[i + r] == DISCARD)\n>> +            rdis1++;\n>> +        else if (action[i + r] == INVESTIGATE)\n>> +            rpdis1++;\n>> +        else if (action[i + r] == KEEP)\n>> +            break;\n>> +        else\n>> +            BUG(\"Illegal value for action[i + r]\");\n>> +    }\n>> +    /*\n>> +     * If the run after the line 'i' found only multimatch lines,\n>> +     * we return false and hence we don't make the current line (i)\n>> +     * discarded.\n>> +     */\n>> +    if (rdis1 == 0)\n>> +        return false;\n>> +    rdis1 += rdis0;\n>> +    rpdis1 += rpdis0;\n>> +\n>> +    return rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n>> +}\n>> +\n>> +struct xoccurrence\n>> +{\n>> +    size_t file1, file2;\n>> +};\n>> +\n>> +\n>> +DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n>> +\n>> +\n>> +/*\n>> + * Try to reduce the problem complexity, discard records that have no\n>> + * matches on the other file. Also, lines that have multiple matches\n>> + * might be potentially discarded if they appear in a run of \n>> discardable.\n>> + */\n>> +static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n>> +    long i;\n>> +    size_t nm, mlim;\n>> +    xrecord_t *recs;\n>> +    uint8_t *action1 = NULL, *action2 = NULL;\n>> +    struct IVec_xoccurrence occ;\n>> +    bool need_min = !!(flags & XDF_NEED_MINIMAL);\n>> +    int ret = 0;\n>> +    ptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n>> +    ptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n>> +\n>> +    IVEC_INIT(occ);\n>> +    ivec_zero(&occ, xe->mph_size);\n>> +\n>> +    for (size_t j = 0; j < xe->xdf1.nrec; j++) {\n>> +        size_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n>> +        occ.ptr[mph1].file1 += 1;\n>> +    }\n>> +\n>> +    for (size_t j = 0; j < xe->xdf2.nrec; j++) {\n>> +        size_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n>> +        occ.ptr[mph2].file2 += 1;\n>> +    }\n>> +\n>> +    /*\n>> +     * Create temporary arrays that will help us decide if\n>> +     * changed[i] should remain false, or become true.\n>> +     */\n>> +    if (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n>> +        ret = -1;\n>> +        goto cleanup;\n>> +    }\n>> +    if (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n>> +        ret = -1;\n>> +        goto cleanup;\n>> +    }\n>> +\n>> +    /*\n>> +     * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>> +     */\n>> +    if ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>> +        mlim = XDL_MAX_EQLIMIT;\n>> +    for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; \n>> i <= dend1; i++, recs++) {\n>> +        nm = occ.ptr[recs->minimal_perfect_hash].file2;\n>> +        action1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? \n>> INVESTIGATE: KEEP;\n>> +    }\n>> +\n>> +    if ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>> +        mlim = XDL_MAX_EQLIMIT;\n>> +    for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; \n>> i <= dend2; i++, recs++) {\n>> +        nm = occ.ptr[recs->minimal_perfect_hash].file1;\n>> +        action2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? \n>> INVESTIGATE: KEEP;\n>> +    }\n>> +\n>> +    /*\n>> +     * Use temporary arrays to decide if changed[i] should remain\n>> +     * false, or become true.\n>> +     */\n>> +    xe->xdf1.nreff = 0;\n>> +    for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n>> +         i <= dend1; i++, recs++) {\n>> +        if (action1[i] == KEEP ||\n>> +            (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, \n>> i, xe->delta_start, dend1))) {\n>> +            xe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n>> +            /* changed[i] remains false, i.e. keep */\n>> +        } else\n>> +            xe->xdf1.changed[i] = true;\n>> +            /* i.e. discard */\n>> +    }\n>> +\n>> +    xe->xdf2.nreff = 0;\n>> +    for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n>> +         i <= dend2; i++, recs++) {\n>> +        if (action2[i] == KEEP ||\n>> +            (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, \n>> i, xe->delta_start, dend2))) {\n>> +            xe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n>> +            /* changed[i] remains false, i.e. keep */\n>> +        } else\n>> +            xe->xdf2.changed[i] = true;\n>> +            /* i.e. discard */\n>> +    }\n>> +\n>> +cleanup:\n>> +    xdl_free(action1);\n>> +    xdl_free(action2);\n>> +    ivec_free(&occ);\n>> +\n>> +    return ret;\n>> +}\n>> +\n>> +\n>>   int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags)\n>>   {\n>>       long ndiags;\n>> @@ -318,6 +496,8 @@ int xdl_do_classic_diff(xdfenv_t *xe, uint64_t flags)\n>>       xdalgoenv_t xenv;\n>>       int res;\n>> +    xdl_cleanup_records(xe, flags);\n>> +\n>>       /*\n>>        * Allocate and setup K vectors to be used by the differential\n>>        * algorithm.\n>> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n>> index b53a3b80c4..3f555e29f4 100644\n>> --- a/xdiff/xprepare.c\n>> +++ b/xdiff/xprepare.c\n>> @@ -24,14 +24,6 @@\n>>   #include \"compat/ivec.h\"\n>> -#define XDL_KPDIS_RUN 4\n>> -#define XDL_MAX_EQLIMIT 1024\n>> -#define XDL_SIMSCAN_WINDOW 100\n>> -\n>> -#define DISCARD 0\n>> -#define KEEP 1\n>> -#define INVESTIGATE 2\n>> -\n>>   typedef struct s_xdlclass {\n>>       struct s_xdlclass *next;\n>>       xrecord_t rec;\n>> @@ -50,8 +42,6 @@ typedef struct s_xdlclassifier {\n>>   } xdlclassifier_t;\n>> -\n>> -\n>>   static int xdl_init_classifier(xdlclassifier_t *cf, long size, long \n>> flags) {\n>>       memset(cf, 0, sizeof(xdlclassifier_t));\n>> @@ -186,175 +176,6 @@ void xdl_free_env(xdfenv_t *xe) {\n>>   }\n>> -static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, \n>> long e) {\n>> -    long r, rdis0, rpdis0, rdis1, rpdis1;\n>> -\n>> -    /*\n>> -     * Limits the window that is examined during the similar-lines\n>> -     * scan. The loops below stops when action[i - r] == KEEP\n>> -     * (line that has no match), but there are corner cases where\n>> -     * the loop proceed all the way to the extremities by causing\n>> -     * huge performance penalties in case of big files.\n>> -     */\n>> -    if (i - s > XDL_SIMSCAN_WINDOW)\n>> -        s = i - XDL_SIMSCAN_WINDOW;\n>> -    if (e - i > XDL_SIMSCAN_WINDOW)\n>> -        e = i + XDL_SIMSCAN_WINDOW;\n>> -\n>> -    /*\n>> -     * Scans the lines before 'i' to find a run of lines that either\n>> -     * have no match (action[j] == DISCARD) or have multiple matches\n>> -     * (action[j] == INVESTIGATE). Note that we always call this\n>> -     * function with action[i] == INVESTIGATE, so the current line\n>> -     * (i) is already a multimatch line.\n>> -     */\n>> -    for (r = 1, rdis0 = 0, rpdis0 = 1; (i - r) >= s; r++) {\n>> -        if (action[i - r] == DISCARD)\n>> -            rdis0++;\n>> -        else if (action[i - r] == INVESTIGATE)\n>> -            rpdis0++;\n>> -        else if (action[i - r] == KEEP)\n>> -            break;\n>> -        else\n>> -            BUG(\"Illegal value for action[i - r]\");\n>> -    }\n>> -    /*\n>> -     * If the run before the line 'i' found only multimatch lines,\n>> -     * we return false and hence we don't make the current line (i)\n>> -     * discarded. We want to discard multimatch lines only when\n>> -     * they appear in the middle of runs with nomatch lines\n>> -     * (action[j] == DISCARD).\n>> -     */\n>> -    if (rdis0 == 0)\n>> -        return 0;\n>> -    for (r = 1, rdis1 = 0, rpdis1 = 1; (i + r) <= e; r++) {\n>> -        if (action[i + r] == DISCARD)\n>> -            rdis1++;\n>> -        else if (action[i + r] == INVESTIGATE)\n>> -            rpdis1++;\n>> -        else if (action[i + r] == KEEP)\n>> -            break;\n>> -        else\n>> -            BUG(\"Illegal value for action[i + r]\");\n>> -    }\n>> -    /*\n>> -     * If the run after the line 'i' found only multimatch lines,\n>> -     * we return false and hence we don't make the current line (i)\n>> -     * discarded.\n>> -     */\n>> -    if (rdis1 == 0)\n>> -        return false;\n>> -    rdis1 += rdis0;\n>> -    rpdis1 += rpdis0;\n>> -\n>> -    return rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n>> -}\n>> -\n>> -struct xoccurrence\n>> -{\n>> -    size_t file1, file2;\n>> -};\n>> -\n>> -\n>> -DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n>> -\n>> -\n>> -/*\n>> - * Try to reduce the problem complexity, discard records that have no\n>> - * matches on the other file. Also, lines that have multiple matches\n>> - * might be potentially discarded if they appear in a run of \n>> discardable.\n>> - */\n>> -static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n>> -    long i;\n>> -    size_t nm, mlim;\n>> -    xrecord_t *recs;\n>> -    uint8_t *action1 = NULL, *action2 = NULL;\n>> -    struct IVec_xoccurrence occ;\n>> -    bool need_min = !!(flags & XDF_NEED_MINIMAL);\n>> -    int ret = 0;\n>> -    ptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n>> -    ptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n>> -\n>> -    IVEC_INIT(occ);\n>> -    ivec_zero(&occ, xe->mph_size);\n>> -\n>> -    for (size_t j = 0; j < xe->xdf1.nrec; j++) {\n>> -        size_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n>> -        occ.ptr[mph1].file1 += 1;\n>> -    }\n>> -\n>> -    for (size_t j = 0; j < xe->xdf2.nrec; j++) {\n>> -        size_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n>> -        occ.ptr[mph2].file2 += 1;\n>> -    }\n>> -\n>> -    /*\n>> -     * Create temporary arrays that will help us decide if\n>> -     * changed[i] should remain false, or become true.\n>> -     */\n>> -    if (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n>> -        ret = -1;\n>> -        goto cleanup;\n>> -    }\n>> -    if (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n>> -        ret = -1;\n>> -        goto cleanup;\n>> -    }\n>> -\n>> -    /*\n>> -     * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>> -     */\n>> -    if ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n>> -        mlim = XDL_MAX_EQLIMIT;\n>> -    for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; \n>> i <= dend1; i++, recs++) {\n>> -        nm = occ.ptr[recs->minimal_perfect_hash].file2;\n>> -        action1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? \n>> INVESTIGATE: KEEP;\n>> -    }\n>> -\n>> -    if ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n>> -        mlim = XDL_MAX_EQLIMIT;\n>> -    for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; \n>> i <= dend2; i++, recs++) {\n>> -        nm = occ.ptr[recs->minimal_perfect_hash].file1;\n>> -        action2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? \n>> INVESTIGATE: KEEP;\n>> -    }\n>> -\n>> -    /*\n>> -     * Use temporary arrays to decide if changed[i] should remain\n>> -     * false, or become true.\n>> -     */\n>> -    xe->xdf1.nreff = 0;\n>> -    for (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start];\n>> -         i <= dend1; i++, recs++) {\n>> -        if (action1[i] == KEEP ||\n>> -            (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, \n>> i, xe->delta_start, dend1))) {\n>> -            xe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n>> -            /* changed[i] remains false, i.e. keep */\n>> -        } else\n>> -            xe->xdf1.changed[i] = true;\n>> -            /* i.e. discard */\n>> -    }\n>> -\n>> -    xe->xdf2.nreff = 0;\n>> -    for (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start];\n>> -         i <= dend2; i++, recs++) {\n>> -        if (action2[i] == KEEP ||\n>> -            (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, \n>> i, xe->delta_start, dend2))) {\n>> -            xe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n>> -            /* changed[i] remains false, i.e. keep */\n>> -        } else\n>> -            xe->xdf2.changed[i] = true;\n>> -            /* i.e. discard */\n>> -    }\n>> -\n>> -cleanup:\n>> -    xdl_free(action1);\n>> -    xdl_free(action2);\n>> -    ivec_free(&occ);\n>> -\n>> -    return ret;\n>> -}\n>> -\n>> -\n>>   /*\n>>    * Early trim initial and terminal matching records.\n>>    */\n>> @@ -414,19 +235,9 @@ int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, \n>> xpparam_t const *xpp,\n>>       }\n>>       xe->mph_size = cf.count;\n>> +    xdl_free_classifier(&cf);\n>>       xdl_trim_ends(xe);\n>> -    if ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n>> -        (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n>> -        xdl_cleanup_records(xe, xpp->flags) < 0) {\n>> -\n>> -        xdl_free_ctx(&xe->xdf2);\n>> -        xdl_free_ctx(&xe->xdf1);\n>> -        xdl_free_classifier(&cf);\n>> -        return -1;\n>> -    }\n>> -\n>> -    xdl_free_classifier(&cf);\n>>       return 0;\n>>   }\n> \n> \n\n"},{"id":"534757","messageId":"7d6dd22d-286e-4f7b-a211-44e00e711401@gmail.com","threadId":"64712","inReplyTo":"CAH=ZcbAiGONrOyma7YjNKKLqNFoisU5LG=nGWjtOJ1wLfqX4cQ@mail.gmail.com","subject":"Re: [PATCH 01/10] ivec: introduce the C side of ivec","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-28T11:15:22Z","receivedAt":"2026-01-28T11:15:26Z","isPatch":true,"body":"On 21/01/2026 21:39, Ezekiel Newren wrote:\n> On Tue, Jan 20, 2026 at 7:06 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>\n>> Hi Ezekiel\n>>\n>> On 15/01/2026 15:55, Ezekiel Newren wrote:\n>>> On Thu, Jan 8, 2026 at 7:34 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>>>> +void ivec_reserve(void *self_, size_t additional)\n>>>>> +{\n>>>>> +     struct IVec_c_void *self = self_;\n>>>>> +\n>>>>> +     size_t growby = 128;\n>>>>> +     if (self->capacity > growby)\n>>>>> +             growby = self->capacity;\n>>>>> +     if (additional > growby)\n>>>>> +             growby = additional;\n>>>>\n>>>> This growth strategy differs from both ALLOC_GROW() and\n>>>> XDL_ALLOC_GROW(), if there isn't a good reason for that we should\n>>>> perhaps just use ALLOC_GROW() here.\n>>>\n>>> XDL_ALLOW_GROW() can't be used because the pointer is always a void*\n>>> in this function.\n>>\n>> Oh right. I'm not sure that's not a reason to use a different growth\n>> strategy though. The minimum size of 128 elements is probably good for\n>> the xdiff code that creates arrays with one element per line but if this\n>> is supposed to be for general use it is going to waste space when we're\n>> allocating a lot of small arrays. ALLOC_GROW() uses alloc_nr() to\n>> calculate the new side so perhaps we could use that here?\n> \n> If ivec_reserve() isn't suitable then ivec_reserve_exact() should be\n> used instead.\n\nIf some C code that pushes one element at a time to an array using \nALLOC_GROW() is converted to use an ivec then we don't want to change \nthe code behaves - that means it should grow the array in the same way. \nI don't see how the suggestion to use ivec_reserve_exact() helps in that \nsituation. What is the advantage in having a different growth \ncharacteristic?\n\n>>>>> +void ivec_push(void *self_, const void *value)\n>>>>> +{\n>>>>> +     struct IVec_c_void *self = self_;\n>>>>> +     void *dst = NULL;\n>>>>> +\n>>>>> +     if (self->length == self->capacity)\n>>>>> +             ivec_reserve(self, 1);\n>>>>> +\n>>>>> +     dst = (uint8_t*)self->ptr + self->length * self->element_size;\n>>>>> +     memcpy(dst, value, self->element_size);\n>>>>\n>>>> If self->element_size was a compile time constant the compiler could\n>>>> easily optimize this call away. I'm not sure that is easy to achieve though.\n>>>\n>>> The problem is that I didn't want all of ivec to be macros that looked\n>>> like function calls. I wanted to minimize use of macros so that it was\n>>> easier to port and verify that the Rust implementation matches the\n>>> behavior of the C implementation.\n>>\n>> I think that's a reasonable concern. So is the plan to have a parallel\n>> rust implementation of these functions rather than call the C\n>> implementation from rust?\n> \n> Yes, the Rust implementation will be independent of the C\n> implementation, but will behave the same way. That's why I'm calling\n> it an interoperable vec as opposed to a compatible vec. Rust can't\n> call the C ivec functions and C can't call the Rust ivec functions,\n> but they'll behave the same way.\n\nInteresting - I'm curious what the advantage of that is over having rust \ncall the C implementation? I can see you wouldn't want to be calling \ninto C for each ivec.push() call, but checking if there is room to push \nthe new element in rust and calling into C to extend the vector if not \nshould be reasonable and then you don't have to re-implement everything \nin rust.\n\n>>>>> +void ivec_free(void *self_)\n>>>>\n>>>> Normally we'd call a like this that free the allocations and\n>>>> re-initializes the members ivec_clear()\n>>>\n>>> In Rust Vec.clear() means to set length to zero, but leaves the\n>>> allocation alone. The reason why I'm zeroing the struct is to help\n>>> avoid FFI issues. If not zero then what should the members be set to,\n>>> to indicate that using the struct is not valid anymore? In Rust an\n>>> object is freed when it goes out of scope and _cannot_ be accessed\n>>> afterward.\n> \n> Maybe I should call this ivec_drop(). Though the notion of explicitly\n> freeing an object in Rust is _almost_ nonsense. The way you free\n> something in Rust is to let it go out of scope.\n\nIndeed - which means this wont be a public function in rust and so why \ndo we worry about naming it ivec_clear()? At least ivec_drop() does not \nconflict with any of the standard function suffixes that we're already \nusing in git.\n\n>> I'm aware that Vec::clear() has different semantics (it does what\n>> strbuf_reset() does). That's unfortunate but this function has different\n>> semantics to all the other *_free() functions in git. Our coding\n>> guidelines say\n>>\n>>    - There are several common idiomatic names for functions performing\n>>      specific tasks on a structure `S`:\n>>\n>>       - `S_init()` initializes a structure without allocating the\n>>         structure itself.\n>>\n>>       - `S_release()` releases a structure's contents without freeing the\n>>         structure.\n>>\n>>       - `S_clear()` is equivalent to `S_release()` followed by `S_init()`\n>>         such that the structure is directly usable after clearing it. When\n>>         `S_clear()` is provided, `S_init()` shall not allocate resources\n>>         that need to be released again.\n>>\n>>       - `S_free()` releases a structure's contents and frees the\n>>         structure.\n>>\n>> As we write more rust code and so wrap more of our existing structs\n>> we're going to be wrapping C code that uses the definitions above so I\n>> think we should do the same with struct IVec_*.\n> \n> I disagree. IVec isn't a wrapper around an existing struct.\n\nSo just because it is a new stuct it shouldn't have to follow the \nexisting naming conventions?\n\n> ivec is\n> meant to very closely mimic Rust's Vec while guaranteeing\n> interoperability. For things like strbuf I haven't conceived of a\n> solution for that yet. Making ivec diverge from Rust's Vec will result\n> in POLA violations due to different behavior when refactoring an\n> IVec<your_type_here> to Vec<your_type_here>.\n\nOn the other hand, vec.reset() does not exist so you'd get a compiler \nerror if you forgot to rename those calls when changing from IVec to Vec \nand the rust code wouldn't be calling ivec.clear(). I'm not sure citing \nPOLA concerns is very convincing as ivec_free() in C is a POLA violation \nfor anyone familiar with git's code base so it's not like there's a \nchoice that avoids that concern.\n\nThanks\n\nPhillip\n\n>>>>> diff --git a/compat/ivec.h b/compat/ivec.h\n>>>>> new file mode 100644\n>>>>> index 0000000000..654a05c506\n>>>>> --- /dev/null\n>>>>> +++ b/compat/ivec.h\n>>>>> @@ -0,0 +1,52 @@\n>>>>> +#ifndef IVEC_H\n>>>>> +#define IVEC_H\n>>>>> +\n>>>>> +#include <git-compat-util.h>\n>>>>\n>>>> It would be nice to have some documentation in this header, see the\n>>>> examples in strvec.h and hashmap.h\n>>>>\n>>>>> +#define IVEC_INIT(variable) ivec_init(&(variable), sizeof(*(variable).ptr))\n>>>>\n>>>> This is a bit cumbersome to use compared to our usual *_INIT macros. I'm\n>>>> struggling to see how we can make it nicer though as DEFINE_IVEC_TYPE\n>>>> cannot define a per-type initializer macro and I we cannot initialize\n>>>> the element size without knowing the type.\n>>>\n>>> I don't see what's cumbersome about it. Maybe an example use case\n>>> would clarify things.\n>>\n>> It is cumbersome because it separates the initialization from the\n>> declaration. Normally our *_INIT macros are initializer lists so we can\n>> write\n>>\n>>          struct strbuf = STRBUF_INIT;\n>>\n>> which keeps the declaration and initialization together. Although\n>> they're on adjacent lines in your example in real code the\n>> initialization likely to be separated from the declaration by other\n>> variable declarations.\n> \n> Ah I see what you mean now. I'll experiment with making IVEC_INIT()\n> work like that. One wrinkle is that STRBUF_INIT is a single concrete\n> type whereas IVEC_INIT() is meant for generic types.\n\nIf you can get it to work that would be great, but I can't think of a \nway of getting it to work for a generic type.\n\nThanks\n\nPhillip\n\n"},{"id":"534766","messageId":"93942e80-207d-4fbd-9965-7c072693b6e1@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"Re: [PATCH 00/10] Xdiff cleanup part 3","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-28T14:40:17Z","receivedAt":"2026-01-28T14:40:20Z","isPatch":true,"body":"The discussion of this series has got rather spread out so I thought it \nmight be helpful to write a summary of my thoughts here.\n\nOn 02/01/2026 18:52, Ezekiel Newren via GitGitGadget wrote:\n> Patch series summary:\n> \n>   * patch 1: Introduce the ivec type\n\nI agree this is a good idea to allow rust and C code to operate on the \nsome data structure. The implementation needs a bit of work to avoid \nundefined behavior.\n\n>   * patch 2: Create the function xdl_do_classic_diff()\n\nThis is sensible\n\n>   * patches 3-4: generic cleanup\n\nPatch 3 claims to \"stop wasting time\" but it introduces an extra pass \nover the input records without any explanation of why that is more \nefficient.\n\nPatch 4 removes the common lines from the beginning and end of the input \nfiles before passing them on to the patience or histogram algorithms. \nThat should speed things up (though we should measure by how much). It \nchanges the output because excluding the common lines at beginning and \nend of the file changes the longest sequence of unique context lines in \nthe lines that remain. If the different output is easier to read then \nthat's clearly a good thing but you would need to do some analysis to \nshow that.\n\n>   * patches 5-8: convert from dstart/dend (in xdfile_t) to\n>     delta_start/delta_end (in xdfenv_t)\n\ndstart a dend in xdfile_t are the index of the first and last line after \nremoving any common lines from the beginning and end. The proposal is to \nstore the offset from the beginning and end instead in xdfenv_t. Looking \nat where dstart and dend are used I think storing the indices is more \nconvenient - if we store offsets we end up calculating the indices from \nthem which is a pain and introduces an opportunity to make an error.\n\n>   * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n>     xdiffi.c\n\nHere we finally get to use the ivec data structre introduced in patch 1. \nHowever it is just replacing a fixed size array and so does not \ndemonstrate the more interesting parts of the API which concern growing \nthe array as we push more elements to it. I'm also not convinced by the \nclaim that this change saves time as it introduces an extra pass over \nthe input records.\n\nOverall I struggled to see how the cleanups proposed here linked to the \nintroduction of the ivec data structure.\n\nThanks\n\nPhillip\n\n> Things that will be addressed in future patch series:\n> \n>   * Make xdl_cleanup_records() easier to read\n>   * convert recs/nrec into an ivec\n>   * convert changed to an ivec\n>   * remove reference_index/nreff from xdfile_t and turn it into an ivec\n>   * splitting minimal_perfect_hash out as its own ivec\n>   * improve the performance of the classifier and parsing/hashing lines\n> \n> === before this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\n> size_t nreff; } xdfile_t;\n> \n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n> \n> === after this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\n> xdfile_t;\n> \n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\n> delta_end; size_t mph_size; } xdfenv_t;\n> \n> Ezekiel Newren (10):\n>    ivec: introduce the C side of ivec\n>    xdiff: make classic diff explicit by creating xdl_do_classic_diff()\n>    xdiff: don't waste time guessing the number of lines\n>    xdiff: let patience and histogram benefit from xdl_trim_ends()\n>    xdiff: use xdfenv_t in xdl_trim_ends() and xdl_cleanup_records()\n>    xdiff: cleanup xdl_trim_ends()\n>    xdiff: replace xdfile_t.dstart with xdfenv_t.delta_start\n>    xdiff: replace xdfile_t.dend with xdfenv_t.delta_end\n>    xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()\n>    xdiff: move xdl_cleanup_records() from xprepare.c to xdiffi.c\n> \n>   Makefile           |   1 +\n>   compat/ivec.c      | 113 ++++++++++++++++++\n>   compat/ivec.h      |  52 +++++++++\n>   meson.build        |   1 +\n>   xdiff/xdiffi.c     | 221 +++++++++++++++++++++++++++++++++---\n>   xdiff/xdiffi.h     |   1 +\n>   xdiff/xhistogram.c |   7 +-\n>   xdiff/xpatience.c  |   7 +-\n>   xdiff/xprepare.c   | 277 ++++++++-------------------------------------\n>   xdiff/xtypes.h     |   3 +-\n>   xdiff/xutils.c     |  20 ----\n>   xdiff/xutils.h     |   1 -\n>   12 files changed, 432 insertions(+), 272 deletions(-)\n>   create mode 100644 compat/ivec.c\n>   create mode 100644 compat/ivec.h\n> \n> \n> base-commit: 66ce5f8e8872f0183bb137911c52b07f1f242d13\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v1\n> Pull-Request: https://github.com/git/git/pull/2156\n\n"},{"id":"538134","messageId":"xmqqy0k4wogg.fsf@gitster.g","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"Re: [PATCH 00/10] Xdiff cleanup part 3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-06T23:03:11Z","receivedAt":"2026-03-06T23:03:13Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Patch series summary:\n>\n>  * patch 1: Introduce the ivec type\n>  * patch 2: Create the function xdl_do_classic_diff()\n>  * patches 3-4: generic cleanup\n>  * patches 5-8: convert from dstart/dend (in xdfile_t) to\n>    delta_start/delta_end (in xdfenv_t)\n>  * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n>    xdiffi.c\n\nIs this topic still viable?\n\nWe had to stop merging this series to the integration branches as\nanother topic <cover.1769424529.git.phillip.wood@dunelm.org.uk> with\nsmaller footprint was making conflicting clean-up.  Since the other\ntopic was merged at 5465d368 (Merge branch 'pw/xdiff-cleanups',\n2026-02-20) a few weeks ago, we may want to resurrect this topic by\nrebasing on top of a more recent 'master' branch.\n\nThanks.\n"},{"id":"538310","messageId":"CAH=ZcbBVSqUG89n65MBpN+HMCmmjzmADGaVnuEJn_cYN0SYknw@mail.gmail.com","threadId":"64712","inReplyTo":"xmqqy0k4wogg.fsf@gitster.g","subject":"Re: [PATCH 00/10] Xdiff cleanup part 3","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-03-09T19:06:26Z","receivedAt":"2026-03-09T19:06:40Z","isPatch":true,"body":"On Fri, Mar 6, 2026 at 4:03 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > Patch series summary:\n> >\n> >  * patch 1: Introduce the ivec type\n> >  * patch 2: Create the function xdl_do_classic_diff()\n> >  * patches 3-4: generic cleanup\n> >  * patches 5-8: convert from dstart/dend (in xdfile_t) to\n> >    delta_start/delta_end (in xdfenv_t)\n> >  * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n> >    xdiffi.c\n>\n> Is this topic still viable?\n>\n> We had to stop merging this series to the integration branches as\n> another topic <cover.1769424529.git.phillip.wood@dunelm.org.uk> with\n> smaller footprint was making conflicting clean-up.  Since the other\n> topic was merged at 5465d368 (Merge branch 'pw/xdiff-cleanups',\n> 2026-02-20) a few weeks ago, we may want to resurrect this topic by\n> rebasing on top of a more recent 'master' branch.\n>\n> Thanks.\n\nI plan on rebasing on top of master with lots of changes. v2 will be\nquite different from v1.\n"},{"id":"538338","messageId":"xmqq8qc01swv.fsf@gitster.g","threadId":"64712","inReplyTo":"CAH=ZcbBVSqUG89n65MBpN+HMCmmjzmADGaVnuEJn_cYN0SYknw@mail.gmail.com","subject":"Re: [PATCH 00/10] Xdiff cleanup part 3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-09T23:31:44Z","receivedAt":"2026-03-09T23:31:46Z","isPatch":true,"body":"Ezekiel Newren <ezekielnewren@gmail.com> writes:\n\n> I plan on rebasing on top of master with lots of changes. v2 will be\n> quite different from v1.\n\nThanks.\n"},{"id":"540009","messageId":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.git.git.1767379944.gitgitgadget@gmail.com","subject":"[PATCH v2 0/5] Xdiff cleanup part 3","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-25T21:11:00Z","receivedAt":"2026-03-25T21:11:08Z","isPatch":true,"body":"v2 is a radical departure from v1 Changes in v2:\n\n * make the flow of xdl_cleanup_records() easier to follow\n\nThere is no performance or behavioral change introduced in this patch\nseries.\n\n=== original cover letter bellow ===\n\nPatch series summary:\n\n * patch 1: Introduce the ivec type\n * patch 2: Create the function xdl_do_classic_diff()\n * patches 3-4: generic cleanup\n * patches 5-8: convert from dstart/dend (in xdfile_t) to\n   delta_start/delta_end (in xdfenv_t)\n * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n   xdiffi.c\n\nThings that will be addressed in future patch series:\n\n * Make xdl_cleanup_records() easier to read\n * convert recs/nrec into an ivec\n * convert changed to an ivec\n * remove reference_index/nreff from xdfile_t and turn it into an ivec\n * splitting minimal_perfect_hash out as its own ivec\n * improve the performance of the classifier and parsing/hashing lines\n\n=== before this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\nsize_t nreff; } xdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n\n=== after this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\nxdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\ndelta_end; size_t mph_size; } xdfenv_t;\n\nEzekiel Newren (5):\n  xdiff/xdl_cleanup_records: delete local recs pointer\n  xdiff/xdl_cleanup_records: make limits more clear\n  xdiff/xdl_cleanup_records: make setting action easier to follow\n  xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n  xdiff/xdl_cleanup_records: use unambiguous types\n\n xdiff/xprepare.c | 89 ++++++++++++++++++++++++++++++++----------------\n 1 file changed, 59 insertions(+), 30 deletions(-)\n\n\nbase-commit: ca1db8a0f7dc0dbea892e99f5b37c5fe5861be71\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v2\nPull-Request: https://github.com/git/git/pull/2156\n\nRange-diff vs v1:\n\n  1:  adf1395d20 <  -:  ---------- ivec: introduce the C side of ivec\n  2:  9bd01bce9f <  -:  ---------- xdiff: make classic diff explicit by creating xdl_do_classic_diff()\n  3:  53e4840c16 <  -:  ---------- xdiff: don't waste time guessing the number of lines\n  4:  70040ea135 <  -:  ---------- xdiff: let patience and histogram benefit from xdl_trim_ends()\n  5:  742f2d381a !  1:  8f9165d477 xdiff: use xdfenv_t in xdl_trim_ends() and xdl_cleanup_records()\n     @@ Metadata\n      Author: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## Commit message ##\n     -    xdiff: use xdfenv_t in xdl_trim_ends() and xdl_cleanup_records()\n     +    xdiff/xdl_cleanup_records: delete local recs pointer\n      \n     -    View with --color-words. Prepare these functions to use the fields:\n     -    delta_start, delta_end. A future patch will add these fields to\n     -    xdfenv_t.\n     +    Simplify the first 2 for loops by directly indexing the xdfile.recs.\n     +    recs is unused in the last 2 for loops, remove it. Best viewed with\n     +    --color-words.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xprepare.c ##\n      @@ xdiff/xprepare.c: static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n     -  * matches on the other file. Also, lines that have multiple matches\n     -  * might be potentially discarded if they appear in a run of discardable.\n        */\n     --static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n     -+static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n     + static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n       \tlong i, nm, mlim;\n     - \txrecord_t *recs;\n     +-\txrecord_t *recs;\n       \txdlclass_t *rcrec;\n     + \tuint8_t *action1 = NULL, *action2 = NULL;\n     + \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n      @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \t * Create temporary arrays that will help us decide if\n     - \t * changed[i] should remain false, or become true.\n       \t */\n     --\tif (!XDL_CALLOC_ARRAY(action1, xdf1->nrec + 1)) {\n     -+\tif (!XDL_CALLOC_ARRAY(action1, xe->xdf1.nrec + 1)) {\n     - \t\tret = -1;\n     - \t\tgoto cleanup;\n     - \t}\n     --\tif (!XDL_CALLOC_ARRAY(action2, xdf2->nrec + 1)) {\n     -+\tif (!XDL_CALLOC_ARRAY(action2, xe->xdf2.nrec + 1)) {\n     - \t\tret = -1;\n     - \t\tgoto cleanup;\n     - \t}\n     -@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \t/*\n     - \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n     - \t */\n     --\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n     -+\tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n     + \tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n       \t\tmlim = XDL_MAX_EQLIMIT;\n      -\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n     -+\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart]; i <= xe->xdf1.dend; i++, recs++) {\n     - \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n     +-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n     ++\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n     ++\t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n     ++\t\trcrec = cf->rcrecs[mph1];\n       \t\tnm = rcrec ? rcrec->len2 : 0;\n       \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n       \t}\n       \n     --\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n     -+\tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n     + \tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n       \t\tmlim = XDL_MAX_EQLIMIT;\n      -\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n     -+\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart]; i <= xe->xdf2.dend; i++, recs++) {\n     - \t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n     +-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n     ++\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n     ++\t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n     ++\t\trcrec = cf->rcrecs[mph2];\n       \t\tnm = rcrec ? rcrec->len1 : 0;\n       \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n     + \t}\n      @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \t * Use temporary arrays to decide if changed[i] should remain\n       \t * false, or become true.\n       \t */\n     --\txdf1->nreff = 0;\n     + \txdf1->nreff = 0;\n      -\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n      -\t     i <= xdf1->dend; i++, recs++) {\n     -+\txe->xdf1.nreff = 0;\n     -+\tfor (i = xe->xdf1.dstart, recs = &xe->xdf1.recs[xe->xdf1.dstart];\n     -+\t     i <= xe->xdf1.dend; i++, recs++) {\n     ++\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n       \t\tif (action1[i] == KEEP ||\n     --\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n     --\t\t\txdf1->reference_index[xdf1->nreff++] = i;\n     -+\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xe->xdf1.dstart, xe->xdf1.dend))) {\n     -+\t\t\txe->xdf1.reference_index[xe->xdf1.nreff++] = i;\n     - \t\t\t/* changed[i] remains false, i.e. keep */\n     - \t\t} else\n     --\t\t\txdf1->changed[i] = true;\n     -+\t\t\txe->xdf1.changed[i] = true;\n     - \t\t\t/* i.e. discard */\n     + \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n     + \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n     +@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n       \t}\n       \n     --\txdf2->nreff = 0;\n     + \txdf2->nreff = 0;\n      -\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n      -\t     i <= xdf2->dend; i++, recs++) {\n     -+\txe->xdf2.nreff = 0;\n     -+\tfor (i = xe->xdf2.dstart, recs = &xe->xdf2.recs[xe->xdf2.dstart];\n     -+\t     i <= xe->xdf2.dend; i++, recs++) {\n     ++\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n       \t\tif (action2[i] == KEEP ||\n     --\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n     --\t\t\txdf2->reference_index[xdf2->nreff++] = i;\n     -+\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xe->xdf2.dstart, xe->xdf2.dend))) {\n     -+\t\t\txe->xdf2.reference_index[xe->xdf2.nreff++] = i;\n     - \t\t\t/* changed[i] remains false, i.e. keep */\n     - \t\t} else\n     --\t\t\txdf2->changed[i] = true;\n     -+\t\t\txe->xdf2.changed[i] = true;\n     - \t\t\t/* i.e. discard */\n     - \t}\n     - \n     -@@ xdiff/xprepare.c: cleanup:\n     - /*\n     -  * Early trim initial and terminal matching records.\n     -  */\n     --static int xdl_trim_ends(xdfile_t *xdf1, xdfile_t *xdf2) {\n     -+static int xdl_trim_ends(xdfenv_t *xe) {\n     - \tlong i, lim;\n     - \txrecord_t *recs1, *recs2;\n     - \n     --\trecs1 = xdf1->recs;\n     --\trecs2 = xdf2->recs;\n     --\tfor (i = 0, lim = (long)XDL_MIN(xdf1->nrec, xdf2->nrec); i < lim;\n     -+\trecs1 = xe->xdf1.recs;\n     -+\trecs2 = xe->xdf2.recs;\n     -+\tfor (i = 0, lim = (long)XDL_MIN(xe->xdf1.nrec, xe->xdf2.nrec); i < lim;\n     - \t     i++, recs1++, recs2++)\n     - \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n     - \t\t\tbreak;\n     - \n     --\txdf1->dstart = xdf2->dstart = i;\n     -+\txe->xdf1.dstart = xe->xdf2.dstart = i;\n     - \n     --\trecs1 = xdf1->recs + xdf1->nrec - 1;\n     --\trecs2 = xdf2->recs + xdf2->nrec - 1;\n     -+\trecs1 = xe->xdf1.recs + xe->xdf1.nrec - 1;\n     -+\trecs2 = xe->xdf2.recs + xe->xdf2.nrec - 1;\n     - \tfor (lim -= i, i = 0; i < lim; i++, recs1--, recs2--)\n     - \t\tif (recs1->minimal_perfect_hash != recs2->minimal_perfect_hash)\n     - \t\t\tbreak;\n     - \n     --\txdf1->dend = (long)xdf1->nrec - i - 1;\n     --\txdf2->dend = (long)xdf2->nrec - i - 1;\n     -+\txe->xdf1.dend = (long)xe->xdf1.nrec - i - 1;\n     -+\txe->xdf2.dend = (long)xe->xdf2.nrec - i - 1;\n     - \n     - \treturn 0;\n     - }\n     -@@ xdiff/xprepare.c: int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n     - \t\txdl_classify_record(2, &cf, rec);\n     - \t}\n     - \n     --\txdl_trim_ends(&xe->xdf1, &xe->xdf2);\n     -+\txdl_trim_ends(xe);\n     - \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n     - \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n     --\t    xdl_cleanup_records(&cf, &xe->xdf1, &xe->xdf2) < 0) {\n     -+\t    xdl_cleanup_records(&cf, xe) < 0) {\n     - \n     - \t\txdl_free_ctx(&xe->xdf2);\n     - \t\txdl_free_ctx(&xe->xdf1);\n     + \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n     + \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n  6:  65da408da9 <  -:  ---------- xdiff: cleanup xdl_trim_ends()\n  7:  d74722538b <  -:  ---------- xdiff: replace xdfile_t.dstart with xdfenv_t.delta_start\n  8:  d0ef5b23c4 <  -:  ---------- xdiff: replace xdfile_t.dend with xdfenv_t.delta_end\n  9:  f9b10e71d2 !  2:  62adaa8e5a xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()\n     @@ Metadata\n      Author: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## Commit message ##\n     -    xdiff: remove dependence on xdlclassifier from xdl_cleanup_records()\n     +    xdiff/xdl_cleanup_records: make limits more clear\n      \n     -    Disentangle xdl_cleanup_records() from the classifier so that it can be\n     -    moved from xprepare.c into xdiffi.c.\n     -\n     -    The classic diff is the only algorithm that needs to count the number\n     -    of times each line occurs in each file. Make xdl_cleanup_records()\n     -    count the number of lines instead of the classifier so it won't slow\n     -    down patience or histogram.\n     +    Make the handling of per-file limits and the minimal-case clearer.\n     +      * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n     +        them.\n     +      * The additional condition `!need_min` is redudant now, remove it.\n     +    Best viewed with --color-words.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xprepare.c ##\n     -@@\n     -  */\n     - \n     - #include \"xinclude.h\"\n     -+#include \"compat/ivec.h\"\n     - \n     - \n     - #define XDL_KPDIS_RUN 4\n     -@@ xdiff/xprepare.c: typedef struct s_xdlclass {\n     - \tstruct s_xdlclass *next;\n     - \txrecord_t rec;\n     - \tlong idx;\n     --\tlong len1, len2;\n     - } xdlclass_t;\n     - \n     - typedef struct s_xdlclassifier {\n     -@@ xdiff/xprepare.c: static void xdl_free_classifier(xdlclassifier_t *cf) {\n     - }\n     - \n     - \n     --static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {\n     -+static int xdl_classify_record(xdlclassifier_t *cf, xrecord_t *rec) {\n     - \tsize_t hi;\n     - \txdlclass_t *rcrec;\n     - \n     -@@ xdiff/xprepare.c: static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t\n     - \t\t\t\treturn -1;\n     - \t\tcf->rcrecs[rcrec->idx] = rcrec;\n     - \t\trcrec->rec = *rec;\n     --\t\trcrec->len1 = rcrec->len2 = 0;\n     - \t\trcrec->next = cf->rchash[hi];\n     - \t\tcf->rchash[hi] = rcrec;\n     - \t}\n     - \n     --\t(pass == 1) ? rcrec->len1++ : rcrec->len2++;\n     --\n     - \trec->minimal_perfect_hash = (size_t)rcrec->idx;\n     - \n     - \treturn 0;\n      @@ xdiff/xprepare.c: static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n     - \treturn rpdis1 * XDL_KPDIS_RUN < (rpdis1 + rdis1);\n     - }\n     - \n     -+struct xoccurrence\n     -+{\n     -+\tsize_t file1, file2;\n     -+};\n     -+\n     -+\n     -+DEFINE_IVEC_TYPE(struct xoccurrence, xoccurrence);\n     -+\n     - \n     - /*\n     -  * Try to reduce the problem complexity, discard records that have no\n     -  * matches on the other file. Also, lines that have multiple matches\n        * might be potentially discarded if they appear in a run of discardable.\n        */\n     --static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n     + static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n      -\tlong i, nm, mlim;\n     -+static int xdl_cleanup_records(xdfenv_t *xe, uint64_t flags) {\n     -+\tlong i;\n     -+\tsize_t nm, mlim;\n     - \txrecord_t *recs;\n     --\txdlclass_t *rcrec;\n     ++\tlong i, nm;\n     ++\tsize_t mlim1, mlim2;\n     + \txdlclass_t *rcrec;\n       \tuint8_t *action1 = NULL, *action2 = NULL;\n     --\tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n     -+\tstruct IVec_xoccurrence occ;\n     -+\tbool need_min = !!(flags & XDF_NEED_MINIMAL);\n     - \tint ret = 0;\n     - \tptrdiff_t dend1 = xe->xdf1.nrec - 1 - xe->delta_end;\n     - \tptrdiff_t dend2 = xe->xdf2.nrec - 1 - xe->delta_end;\n     + \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n     +@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     + \t\tgoto cleanup;\n     + \t}\n       \n     -+\tIVEC_INIT(occ);\n     -+\tivec_zero(&occ, xe->mph_size);\n     -+\n     -+\tfor (size_t j = 0; j < xe->xdf1.nrec; j++) {\n     -+\t\tsize_t mph1 = xe->xdf1.recs[j].minimal_perfect_hash;\n     -+\t\tocc.ptr[mph1].file1 += 1;\n     -+\t}\n     -+\n     -+\tfor (size_t j = 0; j < xe->xdf2.nrec; j++) {\n     -+\t\tsize_t mph2 = xe->xdf2.recs[j].minimal_perfect_hash;\n     -+\t\tocc.ptr[mph2].file2 += 1;\n     ++\tif (need_min) {\n     ++\t\t/* i.e. infinity */\n     ++\t\tmlim1 = SIZE_MAX;\n     ++\t\tmlim2 = SIZE_MAX;\n     ++\t} else {\n     ++\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n     ++\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n      +\t}\n      +\n       \t/*\n     - \t * Create temporary arrays that will help us decide if\n     - \t * changed[i] should remain false, or become true.\n     -@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n     - \tif ((mlim = xdl_bogosqrt((long)xe->xdf1.nrec)) > XDL_MAX_EQLIMIT)\n     - \t\tmlim = XDL_MAX_EQLIMIT;\n     - \tfor (i = xe->delta_start, recs = &xe->xdf1.recs[xe->delta_start]; i <= dend1; i++, recs++) {\n     --\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n     --\t\tnm = rcrec ? rcrec->len2 : 0;\n     -+\t\tnm = occ.ptr[recs->minimal_perfect_hash].file2;\n     - \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n     + \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n     + \t */\n     +-\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n     +-\t\tmlim = XDL_MAX_EQLIMIT;\n     + \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n     + \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n     + \t\trcrec = cf->rcrecs[mph1];\n     + \t\tnm = rcrec ? rcrec->len2 : 0;\n     +-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n     ++\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n       \t}\n       \n     - \tif ((mlim = xdl_bogosqrt((long)xe->xdf2.nrec)) > XDL_MAX_EQLIMIT)\n     - \t\tmlim = XDL_MAX_EQLIMIT;\n     - \tfor (i = xe->delta_start, recs = &xe->xdf2.recs[xe->delta_start]; i <= dend2; i++, recs++) {\n     --\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n     --\t\tnm = rcrec ? rcrec->len1 : 0;\n     -+\t\tnm = occ.ptr[recs->minimal_perfect_hash].file1;\n     - \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n     - \t}\n     - \n     -@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfenv_t *xe) {\n     - cleanup:\n     - \txdl_free(action1);\n     - \txdl_free(action2);\n     -+\tivec_free(&occ);\n     - \n     - \treturn ret;\n     - }\n     -@@ xdiff/xprepare.c: int xdl_prepare_env(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n     - \n     - \tfor (size_t i = 0; i < xe->xdf1.nrec; i++) {\n     - \t\txrecord_t *rec = &xe->xdf1.recs[i];\n     --\t\txdl_classify_record(1, &cf, rec);\n     -+\t\txdl_classify_record(&cf, rec);\n     +-\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n     +-\t\tmlim = XDL_MAX_EQLIMIT;\n     + \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n     + \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n     + \t\trcrec = cf->rcrecs[mph2];\n     + \t\tnm = rcrec ? rcrec->len1 : 0;\n     +-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n     ++\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n       \t}\n       \n     - \tfor (size_t i = 0; i < xe->xdf2.nrec; i++) {\n     - \t\txrecord_t *rec = &xe->xdf2.recs[i];\n     --\t\txdl_classify_record(2, &cf, rec);\n     -+\t\txdl_classify_record(&cf, rec);\n     - \t}\n     - \n     -+\txe->mph_size = cf.count;\n     -+\n     - \txdl_trim_ends(xe);\n     - \tif ((XDF_DIFF_ALG(xpp->flags) != XDF_PATIENCE_DIFF) &&\n     - \t    (XDF_DIFF_ALG(xpp->flags) != XDF_HISTOGRAM_DIFF) &&\n     --\t    xdl_cleanup_records(&cf, xe) < 0) {\n     -+\t    xdl_cleanup_records(xe, xpp->flags) < 0) {\n     - \n     - \t\txdl_free_ctx(&xe->xdf2);\n     - \t\txdl_free_ctx(&xe->xdf1);\n     -\n     - ## xdiff/xtypes.h ##\n     -@@ xdiff/xtypes.h: typedef struct s_xdfile {\n     - typedef struct s_xdfenv {\n     - \txdfile_t xdf1, xdf2;\n     - \tsize_t delta_start, delta_end;\n     -+\tsize_t mph_size;\n     - } xdfenv_t;\n     - \n     - \n     + \t/*\n 10:  1dba6b34aa <  -:  ---------- xdiff: move xdl_cleanup_records() from xprepare.c to xdiffi.c\n  -:  ---------- >  3:  8be7e4781a xdiff/xdl_cleanup_records: make setting action easier to follow\n  -:  ---------- >  4:  6abd052c34 xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n  -:  ---------- >  5:  a52787f019 xdiff/xdl_cleanup_records: use unambiguous types\n\n-- \ngitgitgadget\n"},{"id":"540010","messageId":"8f9165d477ca1dcd2c1915623740a53093b1c258.1774473065.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","subject":"[PATCH v2 1/5] xdiff/xdl_cleanup_records: delete local recs pointer","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-25T21:11:01Z","receivedAt":"2026-03-25T21:11:10Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nSimplify the first 2 for loops by directly indexing the xdfile.recs.\nrecs is unused in the last 2 for loops, remove it. Best viewed with\n--color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 17 ++++++++---------\n 1 file changed, 8 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex cd4fc405eb..d6e1901d2d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -269,7 +269,6 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n \tlong i, nm, mlim;\n-\txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -293,16 +292,18 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n+\t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n \tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n+\t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -312,8 +313,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * false, or become true.\n \t */\n \txdf1->nreff = 0;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n-\t     i <= xdf1->dend; i++, recs++) {\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n@@ -324,8 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t}\n \n \txdf2->nreff = 0;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n-\t     i <= xdf2->dend; i++, recs++) {\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-- \ngitgitgadget\n\n"},{"id":"540011","messageId":"62adaa8e5a5aed585a4b4214c34ece74757d54c7.1774473065.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","subject":"[PATCH v2 2/5] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-25T21:11:02Z","receivedAt":"2026-03-25T21:11:11Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake the handling of per-file limits and the minimal-case clearer.\n  * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n    them.\n  * The additional condition `!need_min` is redudant now, remove it.\nBest viewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 20 +++++++++++++-------\n 1 file changed, 13 insertions(+), 7 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex d6e1901d2d..756a5b8dcc 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -268,7 +268,8 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, mlim;\n+\tlong i, nm;\n+\tsize_t mlim1, mlim2;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -287,25 +288,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tgoto cleanup;\n \t}\n \n+\tif (need_min) {\n+\t\t/* i.e. infinity */\n+\t\tmlim1 = SIZE_MAX;\n+\t\tmlim2 = SIZE_MAX;\n+\t} else {\n+\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n+\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n+\t}\n+\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"540012","messageId":"8be7e4781a9914e7f051a0fc94cb5bb79e258304.1774473065.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","subject":"[PATCH v2 3/5] xdiff/xdl_cleanup_records: make setting action easier to follow","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-25T21:11:03Z","receivedAt":"2026-03-25T21:11:13Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nRewrite nested ternaries with a clear if/else ladder for\naction1/action2 to improve readability while preserving\nbehavior.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 ++++++++++++--\n 1 file changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 756a5b8dcc..127848b764 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -304,14 +304,24 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction1[i] = DISCARD;\n+\t\telse if (nm < mlim1)\n+\t\t\taction1[i] = KEEP;\n+\t\telse /* nm >= mlim1 */\n+\t\t\taction1[i] = INVESTIGATE;\n \t}\n \n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction2[i] = DISCARD;\n+\t\telse if (nm < mlim2)\n+\t\t\taction2[i] = KEEP;\n+\t\telse /* nm >= mlim2 */\n+\t\t\taction2[i] = INVESTIGATE;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"540013","messageId":"6abd052c347025610d26197ba5de8cd11fc1b618.1774473065.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","subject":"[PATCH v2 4/5] xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-25T21:11:04Z","receivedAt":"2026-03-25T21:11:14Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake it clear that INVESTIGATE is turned into KEEP or DISCARD based on\nthe result of xdl_clean_mmatch() which reduces actionX[i] into a\nboolean value.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 34 ++++++++++++++++++++++++----------\n 1 file changed, 24 insertions(+), 10 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 127848b764..dd595cf8a1 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -330,24 +330,38 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \txdf1->nreff = 0;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n-\t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n+\t\tif (action1[i] == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n+\t\t\t\taction1[i] = KEEP;\n+\t\t\telse\n+\t\t\t\taction1[i] = DISCARD;\n+\t\t}\n+\n+\t\tif (action1[i] == KEEP) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action1[i] == DISCARD)\n \t\t\txdf1->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action1[i]\");\n \t}\n \n \txdf2->nreff = 0;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n-\t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n+\t\tif (action2[i] == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n+\t\t\t\taction2[i] = KEEP;\n+\t\t\telse\n+\t\t\t\taction2[i] = DISCARD;\n+\t\t}\n+\n+\t\tif (action2[i] == KEEP) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action2[i] == DISCARD)\n \t\t\txdf2->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action2[i]\");\n \t}\n \n cleanup:\n-- \ngitgitgadget\n\n"},{"id":"540014","messageId":"a52787f0194bf9f7d1e0abe024c423b8d93754fc.1774473065.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","subject":"[PATCH v2 5/5] xdiff/xdl_cleanup_records: use unambiguous types","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-25T21:11:05Z","receivedAt":"2026-03-25T21:11:16Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nChange the parameters of xdl_clean_mmatch() and the local variables\ni, nm in xdl_cleanup_records() to use unambiguous types. Best viewed\nwith --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 12 ++++++------\n 1 file changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex dd595cf8a1..39e48ad33a 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -197,8 +197,8 @@ void xdl_free_env(xdfenv_t *xe) {\n }\n \n \n-static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n-\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n+static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, ptrdiff_t e) {\n+\tptrdiff_t r, rdis0, rpdis0, rdis1, rpdis1;\n \n \t/*\n \t * Limits the window that is examined during the similar-lines\n@@ -268,8 +268,8 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm;\n-\tsize_t mlim1, mlim2;\n+\tptrdiff_t i;\n+\tsize_t nm, mlim1, mlim2;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -303,7 +303,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n-\t\tnm = rcrec ? rcrec->len2 : 0;\n+\t\tnm = rcrec ? (size_t)rcrec->len2 : 0;\n \t\tif (nm == 0)\n \t\t\taction1[i] = DISCARD;\n \t\telse if (nm < mlim1)\n@@ -315,7 +315,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n-\t\tnm = rcrec ? rcrec->len1 : 0;\n+\t\tnm = rcrec ? (size_t)rcrec->len1 : 0;\n \t\tif (nm == 0)\n \t\t\taction2[i] = DISCARD;\n \t\telse if (nm < mlim2)\n-- \ngitgitgadget\n"},{"id":"540018","messageId":"xmqqldffsh9s.fsf@gitster.g","threadId":"64712","inReplyTo":"a52787f0194bf9f7d1e0abe024c423b8d93754fc.1774473065.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 5/5] xdiff/xdl_cleanup_records: use unambiguous types","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-25T21:58:39Z","receivedAt":"2026-03-25T21:58:41Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> Change the parameters of xdl_clean_mmatch() and the local variables\n> i, nm in xdl_cleanup_records() to use unambiguous types. Best viewed\n> with --color-words.\n>\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  xdiff/xprepare.c | 12 ++++++------\n>  1 file changed, 6 insertions(+), 6 deletions(-)\n>\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index dd595cf8a1..39e48ad33a 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -197,8 +197,8 @@ void xdl_free_env(xdfenv_t *xe) {\n>  }\n>  \n>  \n> -static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n> -\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n> +static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, ptrdiff_t e) {\n> +\tptrdiff_t r, rdis0, rpdis0, rdis1, rpdis1;\n>  \n>  \t/*\n>  \t * Limits the window that is examined during the similar-lines\n> @@ -268,8 +268,8 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n>   * might be potentially discarded if they appear in a run of discardable.\n>   */\n>  static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n> -\tlong i, nm;\n> -\tsize_t mlim1, mlim2;\n> +\tptrdiff_t i;\n> +\tsize_t nm, mlim1, mlim2;\n\nLooking good.  Moving away from platform native \"long\" and to types\nthat have more specific meaning makes sense.\n\n>  \txdlclass_t *rcrec;\n>  \tuint8_t *action1 = NULL, *action2 = NULL;\n>  \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> @@ -303,7 +303,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>  \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>  \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>  \t\trcrec = cf->rcrecs[mph1];\n> -\t\tnm = rcrec ? rcrec->len2 : 0;\n> +\t\tnm = rcrec ? (size_t)rcrec->len2 : 0;\n>  \t\tif (nm == 0)\n>  \t\t\taction1[i] = DISCARD;\n>  \t\telse if (nm < mlim1)\n> @@ -315,7 +315,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>  \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>  \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>  \t\trcrec = cf->rcrecs[mph2];\n> -\t\tnm = rcrec ? rcrec->len1 : 0;\n> +\t\tnm = rcrec ? (size_t)rcrec->len1 : 0;\n>  \t\tif (nm == 0)\n>  \t\t\taction2[i] = DISCARD;\n>  \t\telse if (nm < mlim2)\n"},{"id":"540036","messageId":"acTRg4+8/c/BfE7d@szeder.dev","threadId":"64712","inReplyTo":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 0/5] Xdiff cleanup part 3","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2026-03-26T06:26:11Z","receivedAt":"2026-03-26T06:26:43Z","isPatch":true,"body":"On Wed, Mar 25, 2026 at 09:11:00PM +0000, Ezekiel Newren via GitGitGadget wrote:\n> v2 is a radical departure from v1 Changes in v2:\n> \n>  * make the flow of xdl_cleanup_records() easier to follow\n> \n> There is no performance or behavioral change introduced in this patch\n> series.\n> \n> === original cover letter bellow ===\n> \n> Patch series summary:\n> \n>  * patch 1: Introduce the ivec type\n>  * patch 2: Create the function xdl_do_classic_diff()\n>  * patches 3-4: generic cleanup\n>  * patches 5-8: convert from dstart/dend (in xdfile_t) to\n>    delta_start/delta_end (in xdfenv_t)\n>  * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n>    xdiffi.c\n> \n> Things that will be addressed in future patch series:\n> \n>  * Make xdl_cleanup_records() easier to read\n>  * convert recs/nrec into an ivec\n>  * convert changed to an ivec\n>  * remove reference_index/nreff from xdfile_t and turn it into an ivec\n>  * splitting minimal_perfect_hash out as its own ivec\n>  * improve the performance of the classifier and parsing/hashing lines\n> \n> === before this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\n> size_t nreff; } xdfile_t;\n> \n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n> \n> === after this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\n> xdfile_t;\n> \n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\n> delta_end; size_t mph_size; } xdfenv_t;\n\nPlease make sure that each commit in this series can be built with\nDEVELOPER=1, which enables a bunch of additional compiler warnings.\nWhile the last commit can be built with all those warnings, the three\nin the middle fail with sign comparison errors.\n\n> Ezekiel Newren (5):\n>   xdiff/xdl_cleanup_records: delete local recs pointer\n>   xdiff/xdl_cleanup_records: make limits more clear\n\n        CC xdiff/xprepare.o\n    xdiff/xprepare.c: In function ‘xdl_cleanup_records’:\n    xdiff/xprepare.c:307:54: error: comparison of integer expressions of different signedness: ‘long int’ and ‘size_t’ {aka ‘long unsigned int’} [-Werror=sign-compare]\n      307 |                 action1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n          |                                                      ^~\n    xdiff/xprepare.c:314:54: error: comparison of integer expressions of different signedness: ‘long int’ and ‘size_t’ {aka ‘long unsigned int’} [-Werror=sign-compare]\n      314 |                 action2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n          |                                                      ^~\n    cc1: all warnings being treated as errors\n    make: *** [Makefile:2923: xdiff/xprepare.o] Error 1\n\n>   xdiff/xdl_cleanup_records: make setting action easier to follow\n\n      CC xdiff/xprepare.o\n  xdiff/xprepare.c: In function ‘xdl_cleanup_records’:\n  xdiff/xprepare.c:309:29: error: comparison of integer expressions of different signedness: ‘long int’ and ‘size_t’ {aka ‘long unsigned int’} [-Werror=sign-compare]\n    309 |                 else if (nm < mlim1)\n        |                             ^\n  xdiff/xprepare.c:321:29: error: comparison of integer expressions of different signedness: ‘long int’ and ‘size_t’ {aka ‘long unsigned int’} [-Werror=sign-compare]\n    321 |                 else if (nm < mlim2)\n        |                             ^\n  cc1: all warnings being treated as errors\n  make: *** [Makefile:2923: xdiff/xprepare.o] Error 1\n\n>   xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n\nSame error as the last one.\n\n>   xdiff/xdl_cleanup_records: use unambiguous types\n\nGood.\n\n> \n>  xdiff/xprepare.c | 89 ++++++++++++++++++++++++++++++++----------------\n>  1 file changed, 59 insertions(+), 30 deletions(-)\n"},{"id":"540217","messageId":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v2.git.git.1774473065.gitgitgadget@gmail.com","subject":"[PATCH v3 0/6] Xdiff cleanup part 3","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-27T19:23:47Z","receivedAt":"2026-03-27T19:23:57Z","isPatch":true,"body":"Changes in v3:\n\n * run make DEVELOPER=1 on each commit and fix all compiler issues\n\nv2 is a radical departure from v1 Changes in v2:\n\n * make the flow of xdl_cleanup_records() easier to follow\n\nThere is no performance or behavioral change introduced in this patch\nseries.\n\n=== original cover letter bellow ===\n\nPatch series summary:\n\n * patch 1: Introduce the ivec type\n * patch 2: Create the function xdl_do_classic_diff()\n * patches 3-4: generic cleanup\n * patches 5-8: convert from dstart/dend (in xdfile_t) to\n   delta_start/delta_end (in xdfenv_t)\n * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n   xdiffi.c\n\nThings that will be addressed in future patch series:\n\n * Make xdl_cleanup_records() easier to read\n * convert recs/nrec into an ivec\n * convert changed to an ivec\n * remove reference_index/nreff from xdfile_t and turn it into an ivec\n * splitting minimal_perfect_hash out as its own ivec\n * improve the performance of the classifier and parsing/hashing lines\n\n=== before this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\nsize_t nreff; } xdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n\n=== after this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\nxdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\ndelta_end; size_t mph_size; } xdfenv_t;\n\nEzekiel Newren (6):\n  xdiff/xdl_cleanup_records: delete local recs pointer\n  xdiff: use unambiguous types in xdl_bogo_sqrt()\n  xdiff/xdl_cleanup_records: use unambiguous types\n  xdiff/xdl_cleanup_records: make limits more clear\n  xdiff/xdl_cleanup_records: make setting action easier to follow\n  xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n\n xdiff/xdiffi.c   |  2 +-\n xdiff/xprepare.c | 84 ++++++++++++++++++++++++++++++++----------------\n xdiff/xutils.c   |  4 +--\n xdiff/xutils.h   |  2 +-\n 4 files changed, 60 insertions(+), 32 deletions(-)\n\n\nbase-commit: ca1db8a0f7dc0dbea892e99f5b37c5fe5861be71\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v3\nPull-Request: https://github.com/git/git/pull/2156\n\nRange-diff vs v2:\n\n 1:  8f9165d477 = 1:  da32a9747c xdiff/xdl_cleanup_records: delete local recs pointer\n -:  ---------- > 2:  86b0ad100c xdiff: use unambiguous types in xdl_bogo_sqrt()\n 5:  a52787f019 ! 3:  39a35365ae xdiff/xdl_cleanup_records: use unambiguous types\n     @@ Commit message\n          xdiff/xdl_cleanup_records: use unambiguous types\n      \n          Change the parameters of xdl_clean_mmatch() and the local variables\n     -    i, nm in xdl_cleanup_records() to use unambiguous types. Best viewed\n     -    with --color-words.\n     +    i, nm, mlim in xdl_cleanup_records() to use unambiguous types. Best\n     +    viewed with --color-words.\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n     @@ xdiff/xprepare.c: static bool xdl_clean_mmatch(uint8_t const *action, long i, lo\n        * might be potentially discarded if they appear in a run of discardable.\n        */\n       static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n     --\tlong i, nm;\n     --\tsize_t mlim1, mlim2;\n     -+\tptrdiff_t i;\n     -+\tsize_t nm, mlim1, mlim2;\n     +-\tlong i, nm, mlim;\n     ++\tptrdiff_t i, nm, mlim;\n       \txdlclass_t *rcrec;\n       \tuint8_t *action1 = NULL, *action2 = NULL;\n       \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n     -@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n     - \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n     - \t\trcrec = cf->rcrecs[mph1];\n     --\t\tnm = rcrec ? rcrec->len2 : 0;\n     -+\t\tnm = rcrec ? (size_t)rcrec->len2 : 0;\n     - \t\tif (nm == 0)\n     - \t\t\taction1[i] = DISCARD;\n     - \t\telse if (nm < mlim1)\n     -@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n     - \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n     - \t\trcrec = cf->rcrecs[mph2];\n     --\t\tnm = rcrec ? rcrec->len1 : 0;\n     -+\t\tnm = rcrec ? (size_t)rcrec->len1 : 0;\n     - \t\tif (nm == 0)\n     - \t\t\taction2[i] = DISCARD;\n     - \t\telse if (nm < mlim2)\n 2:  62adaa8e5a ! 4:  86dd98db9b xdiff/xdl_cleanup_records: make limits more clear\n     @@ Commit message\n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xprepare.c ##\n     -@@ xdiff/xprepare.c: static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n     +@@ xdiff/xprepare.c: static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n        * might be potentially discarded if they appear in a run of discardable.\n        */\n       static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n     --\tlong i, nm, mlim;\n     -+\tlong i, nm;\n     -+\tsize_t mlim1, mlim2;\n     +-\tptrdiff_t i, nm, mlim;\n     ++\tptrdiff_t i, nm, mlim1, mlim2;\n       \txdlclass_t *rcrec;\n       \tuint8_t *action1 = NULL, *action2 = NULL;\n       \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n     @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n       \t/*\n       \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n       \t */\n     --\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n     +-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n      -\t\tmlim = XDL_MAX_EQLIMIT;\n       \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n       \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n     @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n      +\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n       \t}\n       \n     --\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n     +-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n      -\t\tmlim = XDL_MAX_EQLIMIT;\n       \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n       \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n 3:  8be7e4781a = 5:  ecc25be32f xdiff/xdl_cleanup_records: make setting action easier to follow\n 4:  6abd052c34 = 6:  8f4def8814 xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n\n-- \ngitgitgadget\n"},{"id":"540218","messageId":"da32a9747c7bde88b4fe33e43ae48c7092d57d9d.1774639433.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","subject":"[PATCH v3 1/6] xdiff/xdl_cleanup_records: delete local recs pointer","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-27T19:23:48Z","receivedAt":"2026-03-27T19:23:59Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nSimplify the first 2 for loops by directly indexing the xdfile.recs.\nrecs is unused in the last 2 for loops, remove it. Best viewed with\n--color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 17 ++++++++---------\n 1 file changed, 8 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex cd4fc405eb..d6e1901d2d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -269,7 +269,6 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n \tlong i, nm, mlim;\n-\txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -293,16 +292,18 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n+\t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n \tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n+\t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -312,8 +313,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * false, or become true.\n \t */\n \txdf1->nreff = 0;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n-\t     i <= xdf1->dend; i++, recs++) {\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n@@ -324,8 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t}\n \n \txdf2->nreff = 0;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n-\t     i <= xdf2->dend; i++, recs++) {\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-- \ngitgitgadget\n\n"},{"id":"540219","messageId":"86b0ad100ccbcd1812b24eabd0abe1987592daa0.1774639433.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","subject":"[PATCH v3 2/6] xdiff: use unambiguous types in xdl_bogo_sqrt()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-27T19:23:49Z","receivedAt":"2026-03-27T19:24:00Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThere is no real square root for a negative number and size_t may not\nbe large enough for certain applications, replace long with uint64_t.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   | 2 +-\n xdiff/xprepare.c | 4 ++--\n xdiff/xutils.c   | 4 ++--\n xdiff/xutils.h   | 2 +-\n 4 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 4376f943db..88708c12a3 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -348,7 +348,7 @@ int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \tkvdf += xe->xdf2.nreff + 1;\n \tkvdb += xe->xdf2.nreff + 1;\n \n-\txenv.mxcost = xdl_bogosqrt(ndiags);\n+\txenv.mxcost = (long)xdl_bogosqrt((uint64_t)ndiags);\n \tif (xenv.mxcost < XDL_MAX_COST_MIN)\n \t\txenv.mxcost = XDL_MAX_COST_MIN;\n \txenv.snake_cnt = XDL_SNAKE_CNT;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex d6e1901d2d..48fb5ce6fe 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n@@ -299,7 +299,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 77ee1ad9c8..9a999acdc0 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -23,8 +23,8 @@\n #include \"xinclude.h\"\n \n \n-long xdl_bogosqrt(long n) {\n-\tlong i;\n+uint64_t xdl_bogosqrt(uint64_t n) {\n+\tuint64_t i;\n \n \t/*\n \t * Classical integer square root approximation using shifts.\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 615b4a9d35..58f9d74cda 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -25,7 +25,7 @@\n \n \n \n-long xdl_bogosqrt(long n);\n+uint64_t xdl_bogosqrt(uint64_t n);\n int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n \t\t     xdemitcb_t *ecb);\n int xdl_cha_init(chastore_t *cha, long isize, long icount);\n-- \ngitgitgadget\n\n"},{"id":"540220","messageId":"39a35365ae85f630f6c69c2ab0393ef087becdbd.1774639433.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","subject":"[PATCH v3 3/6] xdiff/xdl_cleanup_records: use unambiguous types","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-27T19:23:50Z","receivedAt":"2026-03-27T19:24:02Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nChange the parameters of xdl_clean_mmatch() and the local variables\ni, nm, mlim in xdl_cleanup_records() to use unambiguous types. Best\nviewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 48fb5ce6fe..386668a92d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -197,8 +197,8 @@ void xdl_free_env(xdfenv_t *xe) {\n }\n \n \n-static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n-\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n+static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, ptrdiff_t e) {\n+\tptrdiff_t r, rdis0, rpdis0, rdis1, rpdis1;\n \n \t/*\n \t * Limits the window that is examined during the similar-lines\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, mlim;\n+\tptrdiff_t i, nm, mlim;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n-- \ngitgitgadget\n\n"},{"id":"540221","messageId":"86dd98db9b93651b21adaa41ccd44917910fedcc.1774639433.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","subject":"[PATCH v3 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-27T19:23:51Z","receivedAt":"2026-03-27T19:24:03Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake the handling of per-file limits and the minimal-case clearer.\n  * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n    them.\n  * The additional condition `!need_min` is redudant now, remove it.\nBest viewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 19 ++++++++++++-------\n 1 file changed, 12 insertions(+), 7 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 386668a92d..2cf1f8d1a8 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tptrdiff_t i, nm, mlim;\n+\tptrdiff_t i, nm, mlim1, mlim2;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -287,25 +287,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tgoto cleanup;\n \t}\n \n+\tif (need_min) {\n+\t\t/* i.e. infinity */\n+\t\tmlim1 = SIZE_MAX;\n+\t\tmlim2 = SIZE_MAX;\n+\t} else {\n+\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n+\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n+\t}\n+\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"540222","messageId":"ecc25be32f394280f3a1a59418140254d7b0811e.1774639433.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","subject":"[PATCH v3 5/6] xdiff/xdl_cleanup_records: make setting action easier to follow","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-27T19:23:52Z","receivedAt":"2026-03-27T19:24:04Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nRewrite nested ternaries with a clear if/else ladder for\naction1/action2 to improve readability while preserving\nbehavior.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 ++++++++++++--\n 1 file changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 2cf1f8d1a8..3d5c61249f 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -303,14 +303,24 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction1[i] = DISCARD;\n+\t\telse if (nm < mlim1)\n+\t\t\taction1[i] = KEEP;\n+\t\telse /* nm >= mlim1 */\n+\t\t\taction1[i] = INVESTIGATE;\n \t}\n \n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction2[i] = DISCARD;\n+\t\telse if (nm < mlim2)\n+\t\t\taction2[i] = KEEP;\n+\t\telse /* nm >= mlim2 */\n+\t\t\taction2[i] = INVESTIGATE;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"540223","messageId":"8f4def8814e21f0ca2772ad6a4426b6cca9d1d14.1774639433.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","subject":"[PATCH v3 6/6] xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-27T19:23:53Z","receivedAt":"2026-03-27T19:24:06Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake it clear that INVESTIGATE is turned into KEEP or DISCARD based on\nthe result of xdl_clean_mmatch() which reduces actionX[i] into a\nboolean value.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 34 ++++++++++++++++++++++++----------\n 1 file changed, 24 insertions(+), 10 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 3d5c61249f..195148442b 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -329,24 +329,38 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \txdf1->nreff = 0;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n-\t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n+\t\tif (action1[i] == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n+\t\t\t\taction1[i] = KEEP;\n+\t\t\telse\n+\t\t\t\taction1[i] = DISCARD;\n+\t\t}\n+\n+\t\tif (action1[i] == KEEP) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action1[i] == DISCARD)\n \t\t\txdf1->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action1[i]\");\n \t}\n \n \txdf2->nreff = 0;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n-\t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n+\t\tif (action2[i] == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n+\t\t\t\taction2[i] = KEEP;\n+\t\t\telse\n+\t\t\t\taction2[i] = DISCARD;\n+\t\t}\n+\n+\t\tif (action2[i] == KEEP) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action2[i] == DISCARD)\n \t\t\txdf2->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action2[i]\");\n \t}\n \n cleanup:\n-- \ngitgitgadget\n"},{"id":"540236","messageId":"xmqqy0jdhtd0.fsf@gitster.g","threadId":"64712","inReplyTo":"86dd98db9b93651b21adaa41ccd44917910fedcc.1774639433.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-27T21:09:47Z","receivedAt":"2026-03-27T21:09:50Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> Make the handling of per-file limits and the minimal-case clearer.\n>   * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n>     them.\n>   * The additional condition `!need_min` is redudant now, remove it.\n> Best viewed with --color-words.\n>\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  xdiff/xprepare.c | 19 ++++++++++++-------\n>  1 file changed, 12 insertions(+), 7 deletions(-)\n\nt4071 and t8015 do not like this step, even though they are happy\nwith 1-3/6 applied.\n\n\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 386668a92d..2cf1f8d1a8 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n>   * might be potentially discarded if they appear in a run of discardable.\n>   */\n>  static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n> -\tptrdiff_t i, nm, mlim;\n> +\tptrdiff_t i, nm, mlim1, mlim2;\n>  \txdlclass_t *rcrec;\n>  \tuint8_t *action1 = NULL, *action2 = NULL;\n>  \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> @@ -287,25 +287,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>  \t\tgoto cleanup;\n>  \t}\n>  \n> +\tif (need_min) {\n> +\t\t/* i.e. infinity */\n> +\t\tmlim1 = SIZE_MAX;\n> +\t\tmlim2 = SIZE_MAX;\n> +\t} else {\n> +\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n> +\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n> +\t}\n> +\n>  \t/*\n>  \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>  \t */\n> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n>  \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>  \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>  \t\trcrec = cf->rcrecs[mph1];\n>  \t\tnm = rcrec ? rcrec->len2 : 0;\n> -\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n>  \t}\n>  \n> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n>  \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>  \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>  \t\trcrec = cf->rcrecs[mph2];\n>  \t\tnm = rcrec ? rcrec->len1 : 0;\n> -\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n>  \t}\n>  \n>  \t/*\n"},{"id":"540244","messageId":"xmqqcy0oj2s1.fsf@gitster.g","threadId":"64712","inReplyTo":"xmqqy0jdhtd0.fsf@gitster.g","subject":"Re: [PATCH v3 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-27T23:01:02Z","receivedAt":"2026-03-27T23:01:05Z","isPatch":true,"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>\n>> Make the handling of per-file limits and the minimal-case clearer.\n>>   * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n>>     them.\n>>   * The additional condition `!need_min` is redudant now, remove it.\n>> Best viewed with --color-words.\n>>\n>> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>> ---\n>>  xdiff/xprepare.c | 19 ++++++++++++-------\n>>  1 file changed, 12 insertions(+), 7 deletions(-)\n>\n> t4071 and t8015 do not like this step, even though they are happy\n> with 1-3/6 applied.\n>\n>\n>> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n>> index 386668a92d..2cf1f8d1a8 100644\n>> --- a/xdiff/xprepare.c\n>> +++ b/xdiff/xprepare.c\n>> @@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n>>   * might be potentially discarded if they appear in a run of discardable.\n>>   */\n>>  static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n>> -\tptrdiff_t i, nm, mlim;\n>> +\tptrdiff_t i, nm, mlim1, mlim2;\n\nAh, the problem may manifest itself in this step in the series, but\nthe root cause might be before this step.  ptrdiff_t is signed and\nthat is the type used for mlim/mlim1/mlim2 here, and before this\nseries these counters count in \"long\" that is signed.\n\n>> +\tif (need_min) {\n>> +\t\t/* i.e. infinity */\n>> +\t\tmlim1 = SIZE_MAX;\n>> +\t\tmlim2 = SIZE_MAX;\n\nBut SIZE_MAX is the maximum that a size_t (unsigned) can take.  No\nwonder assigning it to ptrdiff_t and assuming that any other\nsensible ptrdiff_t value can ever reach it.  Instead, this\nessentially assigns -1 to mlim1 and mlim2 when need_min is true.\n\n>> +\t} else {\n>> +\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n>> +\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n\nThis side I do not think has much to do with the breakage, but the\nway XDL_MIN() is implemented, it must be noted that xdl_bogosqrt()\nis called twice on the same value with this rewrite ...\n\n>> +\t}\n>> +\n>>  \t/*\n>>  \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>>  \t */\n>> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n>> -\t\tmlim = XDL_MAX_EQLIMIT;\n\n... as opposed to computing the value only once, in the original.\n\n>>  \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>>  \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>>  \t\trcrec = cf->rcrecs[mph1];\n>>  \t\tnm = rcrec ? rcrec->len2 : 0;\n>> -\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n\nSo the original said, \"if nm is not zero and need_min is true, do\nnot bother comparing nm with anything, and always use KEEP.  If\nneed_min is false, we use INVESTIGAGE only when nm is large enough,\notherwise KEEP.\n\n>> +\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n\nUpdated code, when nm is not zero, does something different.  if\nneed_min is true, mlim1 is set to -1 and presumably nm is a count or\nlength that is bounded on its lower end with 0, so it is larger than\nmlim1 (== -1), and we always take INVESTIGATE and never kEEP.\n\nSo the rewritten code is broken when need_min is true?\n\nI suspect the remainder of the patch is broken exactly the same way,\nso the remedy would be similar?\n\n>>  \t}\n>>  \n>> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n>> -\t\tmlim = XDL_MAX_EQLIMIT;\n>>  \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>>  \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>>  \t\trcrec = cf->rcrecs[mph2];\n>>  \t\tnm = rcrec ? rcrec->len1 : 0;\n>> -\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>> +\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n>>  \t}\n>>  \n>>  \t/*\n"},{"id":"540395","messageId":"CAH=ZcbAKwtq9jiv=XWi_P0ZD1hz7XEpEtMPONB9n=_EcOPPSRg@mail.gmail.com","threadId":"64712","inReplyTo":"xmqqcy0oj2s1.fsf@gitster.g","subject":"Re: [PATCH v3 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-03-30T16:00:18Z","receivedAt":"2026-03-30T16:00:30Z","isPatch":true,"body":"On Fri, Mar 27, 2026 at 5:01 PM Junio C Hamano <gitster@pobox.com> wrote:\n> Updated code, when nm is not zero, does something different.  if\n> need_min is true, mlim1 is set to -1 and presumably nm is a count or\n> length that is bounded on its lower end with 0, so it is larger than\n> mlim1 (== -1), and we always take INVESTIGATE and never KEEP.\n>\n> So the rewritten code is broken when need_min is true?\n>\n> I suspect the remainder of the patch is broken exactly the same way,\n> so the remedy would be similar?\n\nYour assessment is correct, PTRDIFF_MAX should be used instead of\nSIZE_MAX. I realized my mistake a few hours after I pushed. This will\nbe fixed in the next version.\n"},{"id":"540396","messageId":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v3.git.git.1774639433.gitgitgadget@gmail.com","subject":"[PATCH v4 0/6] Xdiff cleanup part 3","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-30T16:59:57Z","receivedAt":"2026-03-30T17:00:07Z","isPatch":true,"body":"Changes in v3:\n\n * run make DEVELOPER=1 on each commit and fix all compiler issues\n\nv2 is a radical departure from v1 Changes in v2:\n\n * make the flow of xdl_cleanup_records() easier to follow\n\nThere is no performance or behavioral change introduced in this patch\nseries.\n\n=== original cover letter bellow ===\n\nPatch series summary:\n\n * patch 1: Introduce the ivec type\n * patch 2: Create the function xdl_do_classic_diff()\n * patches 3-4: generic cleanup\n * patches 5-8: convert from dstart/dend (in xdfile_t) to\n   delta_start/delta_end (in xdfenv_t)\n * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n   xdiffi.c\n\nThings that will be addressed in future patch series:\n\n * Make xdl_cleanup_records() easier to read\n * convert recs/nrec into an ivec\n * convert changed to an ivec\n * remove reference_index/nreff from xdfile_t and turn it into an ivec\n * splitting minimal_perfect_hash out as its own ivec\n * improve the performance of the classifier and parsing/hashing lines\n\n=== before this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\nsize_t nreff; } xdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n\n=== after this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\nxdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\ndelta_end; size_t mph_size; } xdfenv_t;\n\nEzekiel Newren (6):\n  xdiff/xdl_cleanup_records: delete local recs pointer\n  xdiff: use unambiguous types in xdl_bogo_sqrt()\n  xdiff/xdl_cleanup_records: use unambiguous types\n  xdiff/xdl_cleanup_records: make limits more clear\n  xdiff/xdl_cleanup_records: make setting action easier to follow\n  xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n\n xdiff/xdiffi.c   |  2 +-\n xdiff/xprepare.c | 84 ++++++++++++++++++++++++++++++++----------------\n xdiff/xutils.c   |  4 +--\n xdiff/xutils.h   |  2 +-\n 4 files changed, 60 insertions(+), 32 deletions(-)\n\n\nbase-commit: ca1db8a0f7dc0dbea892e99f5b37c5fe5861be71\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v4\nPull-Request: https://github.com/git/git/pull/2156\n\nRange-diff vs v3:\n\n 1:  da32a9747c = 1:  da32a9747c xdiff/xdl_cleanup_records: delete local recs pointer\n 2:  86b0ad100c = 2:  86b0ad100c xdiff: use unambiguous types in xdl_bogo_sqrt()\n 3:  39a35365ae = 3:  39a35365ae xdiff/xdl_cleanup_records: use unambiguous types\n 4:  86dd98db9b ! 4:  75fe3ea125 xdiff/xdl_cleanup_records: make limits more clear\n     @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n       \n      +\tif (need_min) {\n      +\t\t/* i.e. infinity */\n     -+\t\tmlim1 = SIZE_MAX;\n     -+\t\tmlim2 = SIZE_MAX;\n     ++\t\tmlim1 = PTRDIFF_MAX;\n     ++\t\tmlim2 = PTRDIFF_MAX;\n      +\t} else {\n      +\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n      +\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n 5:  ecc25be32f = 5:  0cf1412d01 xdiff/xdl_cleanup_records: make setting action easier to follow\n 6:  8f4def8814 = 6:  fd14ccafc4 xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n\n-- \ngitgitgadget\n"},{"id":"540397","messageId":"da32a9747c7bde88b4fe33e43ae48c7092d57d9d.1774890003.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"[PATCH v4 1/6] xdiff/xdl_cleanup_records: delete local recs pointer","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-30T16:59:58Z","receivedAt":"2026-03-30T17:00:08Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nSimplify the first 2 for loops by directly indexing the xdfile.recs.\nrecs is unused in the last 2 for loops, remove it. Best viewed with\n--color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 17 ++++++++---------\n 1 file changed, 8 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex cd4fc405eb..d6e1901d2d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -269,7 +269,6 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n \tlong i, nm, mlim;\n-\txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -293,16 +292,18 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n+\t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n \tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n+\t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -312,8 +313,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * false, or become true.\n \t */\n \txdf1->nreff = 0;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n-\t     i <= xdf1->dend; i++, recs++) {\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n@@ -324,8 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t}\n \n \txdf2->nreff = 0;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n-\t     i <= xdf2->dend; i++, recs++) {\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-- \ngitgitgadget\n\n"},{"id":"540398","messageId":"86b0ad100ccbcd1812b24eabd0abe1987592daa0.1774890003.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"[PATCH v4 2/6] xdiff: use unambiguous types in xdl_bogo_sqrt()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-30T16:59:59Z","receivedAt":"2026-03-30T17:00:09Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThere is no real square root for a negative number and size_t may not\nbe large enough for certain applications, replace long with uint64_t.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   | 2 +-\n xdiff/xprepare.c | 4 ++--\n xdiff/xutils.c   | 4 ++--\n xdiff/xutils.h   | 2 +-\n 4 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 4376f943db..88708c12a3 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -348,7 +348,7 @@ int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \tkvdf += xe->xdf2.nreff + 1;\n \tkvdb += xe->xdf2.nreff + 1;\n \n-\txenv.mxcost = xdl_bogosqrt(ndiags);\n+\txenv.mxcost = (long)xdl_bogosqrt((uint64_t)ndiags);\n \tif (xenv.mxcost < XDL_MAX_COST_MIN)\n \t\txenv.mxcost = XDL_MAX_COST_MIN;\n \txenv.snake_cnt = XDL_SNAKE_CNT;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex d6e1901d2d..48fb5ce6fe 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n@@ -299,7 +299,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 77ee1ad9c8..9a999acdc0 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -23,8 +23,8 @@\n #include \"xinclude.h\"\n \n \n-long xdl_bogosqrt(long n) {\n-\tlong i;\n+uint64_t xdl_bogosqrt(uint64_t n) {\n+\tuint64_t i;\n \n \t/*\n \t * Classical integer square root approximation using shifts.\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 615b4a9d35..58f9d74cda 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -25,7 +25,7 @@\n \n \n \n-long xdl_bogosqrt(long n);\n+uint64_t xdl_bogosqrt(uint64_t n);\n int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n \t\t     xdemitcb_t *ecb);\n int xdl_cha_init(chastore_t *cha, long isize, long icount);\n-- \ngitgitgadget\n\n"},{"id":"540399","messageId":"39a35365ae85f630f6c69c2ab0393ef087becdbd.1774890003.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"[PATCH v4 3/6] xdiff/xdl_cleanup_records: use unambiguous types","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-30T17:00:00Z","receivedAt":"2026-03-30T17:00:10Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nChange the parameters of xdl_clean_mmatch() and the local variables\ni, nm, mlim in xdl_cleanup_records() to use unambiguous types. Best\nviewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 48fb5ce6fe..386668a92d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -197,8 +197,8 @@ void xdl_free_env(xdfenv_t *xe) {\n }\n \n \n-static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n-\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n+static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, ptrdiff_t e) {\n+\tptrdiff_t r, rdis0, rpdis0, rdis1, rpdis1;\n \n \t/*\n \t * Limits the window that is examined during the similar-lines\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, mlim;\n+\tptrdiff_t i, nm, mlim;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n-- \ngitgitgadget\n\n"},{"id":"540400","messageId":"75fe3ea1250ab7dfa4e029f49f2ad353185afded.1774890003.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"[PATCH v4 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-30T17:00:01Z","receivedAt":"2026-03-30T17:00:12Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake the handling of per-file limits and the minimal-case clearer.\n  * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n    them.\n  * The additional condition `!need_min` is redudant now, remove it.\nBest viewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 19 ++++++++++++-------\n 1 file changed, 12 insertions(+), 7 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 386668a92d..bd8baf214d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tptrdiff_t i, nm, mlim;\n+\tptrdiff_t i, nm, mlim1, mlim2;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -287,25 +287,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tgoto cleanup;\n \t}\n \n+\tif (need_min) {\n+\t\t/* i.e. infinity */\n+\t\tmlim1 = PTRDIFF_MAX;\n+\t\tmlim2 = PTRDIFF_MAX;\n+\t} else {\n+\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n+\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n+\t}\n+\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"540401","messageId":"0cf1412d01cc4895aa945b6f3ead3b2d79716523.1774890003.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"[PATCH v4 5/6] xdiff/xdl_cleanup_records: make setting action easier to follow","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-30T17:00:02Z","receivedAt":"2026-03-30T17:00:13Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nRewrite nested ternaries with a clear if/else ladder for\naction1/action2 to improve readability while preserving\nbehavior.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 ++++++++++++--\n 1 file changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex bd8baf214d..471d9567c9 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -303,14 +303,24 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction1[i] = DISCARD;\n+\t\telse if (nm < mlim1)\n+\t\t\taction1[i] = KEEP;\n+\t\telse /* nm >= mlim1 */\n+\t\t\taction1[i] = INVESTIGATE;\n \t}\n \n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction2[i] = DISCARD;\n+\t\telse if (nm < mlim2)\n+\t\t\taction2[i] = KEEP;\n+\t\telse /* nm >= mlim2 */\n+\t\t\taction2[i] = INVESTIGATE;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"540402","messageId":"fd14ccafc494aeda4bb9d05b83ac09f35bec8b52.1774890003.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"[PATCH v4 6/6] xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-30T17:00:03Z","receivedAt":"2026-03-30T17:00:15Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake it clear that INVESTIGATE is turned into KEEP or DISCARD based on\nthe result of xdl_clean_mmatch() which reduces actionX[i] into a\nboolean value.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 34 ++++++++++++++++++++++++----------\n 1 file changed, 24 insertions(+), 10 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 471d9567c9..1f2e8c6b4b 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -329,24 +329,38 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \txdf1->nreff = 0;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n-\t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n+\t\tif (action1[i] == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n+\t\t\t\taction1[i] = KEEP;\n+\t\t\telse\n+\t\t\t\taction1[i] = DISCARD;\n+\t\t}\n+\n+\t\tif (action1[i] == KEEP) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action1[i] == DISCARD)\n \t\t\txdf1->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action1[i]\");\n \t}\n \n \txdf2->nreff = 0;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n-\t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n+\t\tif (action2[i] == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n+\t\t\t\taction2[i] = KEEP;\n+\t\t\telse\n+\t\t\t\taction2[i] = DISCARD;\n+\t\t}\n+\n+\t\tif (action2[i] == KEEP) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action2[i] == DISCARD)\n \t\t\txdf2->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action2[i]\");\n \t}\n \n cleanup:\n-- \ngitgitgadget\n"},{"id":"540406","messageId":"CAH=ZcbA661Ho2ttq0VjFf4R4k8ZKg4yf=8rJPa+Nu+PWA_+wkA@mail.gmail.com","threadId":"64712","inReplyTo":"da32a9747c7bde88b4fe33e43ae48c7092d57d9d.1774890003.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 1/6] xdiff/xdl_cleanup_records: delete local recs pointer","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-03-30T17:23:36Z","receivedAt":"2026-03-30T17:23:49Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 11:00 AM Ezekiel Newren via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\nI forgot to update the cover letter. It would have said:\n\nChanges in v4:\n  * Change SIZE_MAX to PTRDIFF_MAX.\n"},{"id":"540422","messageId":"xmqqtstxdr6v.fsf@gitster.g","threadId":"64712","inReplyTo":"CAH=ZcbAKwtq9jiv=XWi_P0ZD1hz7XEpEtMPONB9n=_EcOPPSRg@mail.gmail.com","subject":"Re: [PATCH v3 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-30T19:59:20Z","receivedAt":"2026-03-30T19:59:22Z","isPatch":true,"body":"Ezekiel Newren <ezekielnewren@gmail.com> writes:\n\n> On Fri, Mar 27, 2026 at 5:01 PM Junio C Hamano <gitster@pobox.com> wrote:\n>> Updated code, when nm is not zero, does something different.  if\n>> need_min is true, mlim1 is set to -1 and presumably nm is a count or\n>> length that is bounded on its lower end with 0, so it is larger than\n>> mlim1 (== -1), and we always take INVESTIGATE and never KEEP.\n>>\n>> So the rewritten code is broken when need_min is true?\n>>\n>> I suspect the remainder of the patch is broken exactly the same way,\n>> so the remedy would be similar?\n>\n> Your assessment is correct, PTRDIFF_MAX should be used instead of\n> SIZE_MAX. I realized my mistake a few hours after I pushed. This will\n> be fixed in the next version.\n\nYeah, using PTRDIFF_MAX is fine.  When I reported the breakage I was\nhinting that everything may want to become unsigned, but since the\noriginal does use signed quantities and variables, it is far safer\nto stick to signed arithmetic---until a full audit says it is safe\nto switch to size_t of course.\n\nThanks.\n"},{"id":"540438","messageId":"xmqq7bqt6i9y.fsf@gitster.g","threadId":"64712","inReplyTo":"da32a9747c7bde88b4fe33e43ae48c7092d57d9d.1774890003.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 1/6] xdiff/xdl_cleanup_records: delete local recs pointer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-30T22:53:45Z","receivedAt":"2026-03-30T22:53:49Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> Simplify the first 2 for loops by directly indexing the xdfile.recs.\n> recs is unused in the last 2 for loops, remove it. Best viewed with\n> --color-words.\n>\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  xdiff/xprepare.c | 17 ++++++++---------\n>  1 file changed, 8 insertions(+), 9 deletions(-)\n\nInteresting that the latter loops did not even have to have the\nextra pointer variable.  Nice clean-up.\n\n>\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index cd4fc405eb..d6e1901d2d 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -269,7 +269,6 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n>   */\n>  static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n>  \tlong i, nm, mlim;\n> -\txrecord_t *recs;\n>  \txdlclass_t *rcrec;\n>  \tuint8_t *action1 = NULL, *action2 = NULL;\n>  \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> @@ -293,16 +292,18 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>  \t */\n>  \tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n>  \t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n> -\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> +\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n> +\t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n> +\t\trcrec = cf->rcrecs[mph1];\n>  \t\tnm = rcrec ? rcrec->len2 : 0;\n>  \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>  \t}\n>  \n>  \tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n>  \t\tmlim = XDL_MAX_EQLIMIT;\n> -\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n> -\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n> +\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n> +\t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n> +\t\trcrec = cf->rcrecs[mph2];\n>  \t\tnm = rcrec ? rcrec->len1 : 0;\n>  \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n>  \t}\n> @@ -312,8 +313,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>  \t * false, or become true.\n>  \t */\n>  \txdf1->nreff = 0;\n> -\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n> -\t     i <= xdf1->dend; i++, recs++) {\n> +\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>  \t\tif (action1[i] == KEEP ||\n>  \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n>  \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n> @@ -324,8 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>  \t}\n>  \n>  \txdf2->nreff = 0;\n> -\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n> -\t     i <= xdf2->dend; i++, recs++) {\n> +\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>  \t\tif (action2[i] == KEEP ||\n>  \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n>  \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n"},{"id":"540440","messageId":"xmqq341g7wka.fsf@gitster.g","threadId":"64712","inReplyTo":"86b0ad100ccbcd1812b24eabd0abe1987592daa0.1774890003.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 2/6] xdiff: use unambiguous types in xdl_bogo_sqrt()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-30T22:59:49Z","receivedAt":"2026-03-30T22:59:51Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> -\txenv.mxcost = xdl_bogosqrt(ndiags);\n> +\txenv.mxcost = (long)xdl_bogosqrt((uint64_t)ndiags);\n\nThere is nothing actionable, but this makes me wonder if we want to\nupdate the type of .mxcost member (which seems to never go negative)\nsomehow.  I also wonder if uint32_t should be sufficiently wide for\nxdl_bogosqrt() that takes uint64_t.\n\n"},{"id":"540442","messageId":"xmqqy0j86hva.fsf@gitster.g","threadId":"64712","inReplyTo":"0cf1412d01cc4895aa945b6f3ead3b2d79716523.1774890003.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 5/6] xdiff/xdl_cleanup_records: make setting action easier to follow","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-30T23:02:33Z","receivedAt":"2026-03-30T23:02:35Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>\n> Rewrite nested ternaries with a clear if/else ladder for\n> action1/action2 to improve readability while preserving\n> behavior.\n>\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>  xdiff/xprepare.c | 14 ++++++++++++--\n>  1 file changed, 12 insertions(+), 2 deletions(-)\n\nOh, I love this kind of rewrite that makes it more trivial to follwo\nwhat the code is doing.  Looking good.\n\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index bd8baf214d..471d9567c9 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -303,14 +303,24 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>  \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>  \t\trcrec = cf->rcrecs[mph1];\n>  \t\tnm = rcrec ? rcrec->len2 : 0;\n> -\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n> +\t\tif (nm == 0)\n> +\t\t\taction1[i] = DISCARD;\n> +\t\telse if (nm < mlim1)\n> +\t\t\taction1[i] = KEEP;\n> +\t\telse /* nm >= mlim1 */\n> +\t\t\taction1[i] = INVESTIGATE;\n>  \t}\n>  \n>  \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>  \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>  \t\trcrec = cf->rcrecs[mph2];\n>  \t\tnm = rcrec ? rcrec->len1 : 0;\n> -\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n> +\t\tif (nm == 0)\n> +\t\t\taction2[i] = DISCARD;\n> +\t\telse if (nm < mlim2)\n> +\t\t\taction2[i] = KEEP;\n> +\t\telse /* nm >= mlim2 */\n> +\t\t\taction2[i] = INVESTIGATE;\n>  \t}\n>  \n>  \t/*\n"},{"id":"540443","messageId":"xmqqtstw6hs2.fsf@gitster.g","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 0/6] Xdiff cleanup part 3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-30T23:04:29Z","receivedAt":"2026-03-30T23:04:31Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Changes in v3:\n>\n>  * run make DEVELOPER=1 on each commit and fix all compiler issues\n\nThis round looks very good to me.  Let me mark it for 'next' unless\nothers bring up problems I failed to see in a few days.\n\nThanks.\n"},{"id":"540449","messageId":"CAH=ZcbA_1pZYDjg0Q7bEB11vY8-T76o-r-v9g--NUSwbfZigsQ@mail.gmail.com","threadId":"64712","inReplyTo":"xmqqtstxdr6v.fsf@gitster.g","subject":"Re: [PATCH v3 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-03-31T01:29:57Z","receivedAt":"2026-03-31T01:30:11Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 1:59 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Ezekiel Newren <ezekielnewren@gmail.com> writes:\n>\n> > On Fri, Mar 27, 2026 at 5:01 PM Junio C Hamano <gitster@pobox.com> wrote:\n> >> Updated code, when nm is not zero, does something different.  if\n> >> need_min is true, mlim1 is set to -1 and presumably nm is a count or\n> >> length that is bounded on its lower end with 0, so it is larger than\n> >> mlim1 (== -1), and we always take INVESTIGATE and never KEEP.\n> >>\n> >> So the rewritten code is broken when need_min is true?\n> >>\n> >> I suspect the remainder of the patch is broken exactly the same way,\n> >> so the remedy would be similar?\n> >\n> > Your assessment is correct, PTRDIFF_MAX should be used instead of\n> > SIZE_MAX. I realized my mistake a few hours after I pushed. This will\n> > be fixed in the next version.\n>\n> Yeah, using PTRDIFF_MAX is fine.  When I reported the breakage I was\n> hinting that everything may want to become unsigned, but since the\n> original does use signed quantities and variables, it is far safer\n> to stick to signed arithmetic---until a full audit says it is safe\n> to switch to size_t of course.\n\nI would prefer to make everything size_t, but dend can be negative if\nthe number of lines in a file is 0 and that breaks the current code if\nunsigned is forced. I can cleanup the code to use unsigned, but I\ndidn't want to distract from the readability, of this patch series, of\nxdl_cleanup_records() with other refactorings.\n\nIn fact dstart is never negative, but I thought that it would be more\nconfusing to change dstart to unsigned and keep dend signed and\nexplain why there is a discrepancy in types between the 2.\n"},{"id":"540494","messageId":"40589b6f-6694-4d9c-8367-3f6352e45e7b@gmail.com","threadId":"64712","inReplyTo":"fd14ccafc494aeda4bb9d05b83ac09f35bec8b52.1774890003.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 6/6] xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-03-31T09:43:49Z","receivedAt":"2026-03-31T09:43:52Z","isPatch":true,"body":"Hi Ezekiel\n\nOn 30/03/2026 18:00, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Make it clear that INVESTIGATE is turned into KEEP or DISCARD based on\n> the result of xdl_clean_mmatch() which reduces actionX[i] into a\n> boolean value.\n> \n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xprepare.c | 34 ++++++++++++++++++++++++----------\n>   1 file changed, 24 insertions(+), 10 deletions(-)\n> \n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 471d9567c9..1f2e8c6b4b 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -329,24 +329,38 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t */\n>   \txdf1->nreff = 0;\n>   \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n> -\t\tif (action1[i] == KEEP ||\n> -\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n> +\t\tif (action1[i] == INVESTIGATE) {\n> +\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n> +\t\t\t\taction1[i] = KEEP;\n> +\t\t\telse\n> +\t\t\t\taction1[i] = DISCARD;\n> +\t\t}\n> +\n> +\t\tif (action1[i] == KEEP) {\n>   \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> +\t\t\t/* changed[i] remains false */\n> +\t\t} else if (action1[i] == DISCARD)\n\nAs one clause uses braces, they all should. Apart from that this looks \nlike another nice improvement in readability.\n\nThanks\n\nPhillip\n\n>   \t\t\txdf1->changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> +\t\telse\n> +\t\t\tBUG(\"Illegal state for action1[i]\");\n>   \t}\n>   \n>   \txdf2->nreff = 0;\n>   \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n> -\t\tif (action2[i] == KEEP ||\n> -\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n> +\t\tif (action2[i] == INVESTIGATE) {\n> +\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n> +\t\t\t\taction2[i] = KEEP;\n> +\t\t\telse\n> +\t\t\t\taction2[i] = DISCARD;\n> +\t\t}\n> +\n> +\t\tif (action2[i] == KEEP) {\n>   \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> +\t\t\t/* changed[i] remains false */\n> +\t\t} else if (action2[i] == DISCARD)\n>   \t\t\txdf2->changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> +\t\telse\n> +\t\t\tBUG(\"Illegal state for action2[i]\");\n>   \t}\n>   \n>   cleanup:\n\n"},{"id":"540495","messageId":"6d099729-d28c-4c1b-b61b-26aaa6b48ec8@gmail.com","threadId":"64712","inReplyTo":"xmqqy0j86hva.fsf@gitster.g","subject":"Re: [PATCH v4 5/6] xdiff/xdl_cleanup_records: make setting action easier to follow","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-03-31T09:44:05Z","receivedAt":"2026-03-31T09:44:08Z","isPatch":true,"body":"On 31/03/2026 00:02, Junio C Hamano wrote:\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Ezekiel Newren <ezekielnewren@gmail.com>\n>>\n>> Rewrite nested ternaries with a clear if/else ladder for\n>> action1/action2 to improve readability while preserving\n>> behavior.\n>>\n>> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>> ---\n>>   xdiff/xprepare.c | 14 ++++++++++++--\n>>   1 file changed, 12 insertions(+), 2 deletions(-)\n> \n> Oh, I love this kind of rewrite that makes it more trivial to follwo\n> what the code is doing.  Looking good.\n\nYes, this is a nice improvement in readability\n\nThanks\n\nPhillip\n\n> \n>> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n>> index bd8baf214d..471d9567c9 100644\n>> --- a/xdiff/xprepare.c\n>> +++ b/xdiff/xprepare.c\n>> @@ -303,14 +303,24 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>>   \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>>   \t\trcrec = cf->rcrecs[mph1];\n>>   \t\tnm = rcrec ? rcrec->len2 : 0;\n>> -\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n>> +\t\tif (nm == 0)\n>> +\t\t\taction1[i] = DISCARD;\n>> +\t\telse if (nm < mlim1)\n>> +\t\t\taction1[i] = KEEP;\n>> +\t\telse /* nm >= mlim1 */\n>> +\t\t\taction1[i] = INVESTIGATE;\n>>   \t}\n>>   \n>>   \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>>   \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>>   \t\trcrec = cf->rcrecs[mph2];\n>>   \t\tnm = rcrec ? rcrec->len1 : 0;\n>> -\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n>> +\t\tif (nm == 0)\n>> +\t\t\taction2[i] = DISCARD;\n>> +\t\telse if (nm < mlim2)\n>> +\t\t\taction2[i] = KEEP;\n>> +\t\telse /* nm >= mlim2 */\n>> +\t\t\taction2[i] = INVESTIGATE;\n>>   \t}\n>>   \n>>   \t/*\n\n"},{"id":"540496","messageId":"32c34d0d-9358-43e3-9d58-5999b3ffd6c2@gmail.com","threadId":"64712","inReplyTo":"75fe3ea1250ab7dfa4e029f49f2ad353185afded.1774890003.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-03-31T09:44:15Z","receivedAt":"2026-03-31T09:44:18Z","isPatch":true,"body":"Hi Ezekiel\n\nOn 30/03/2026 18:00, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Make the handling of per-file limits and the minimal-case clearer.\n>    * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n>      them.\n>    * The additional condition `!need_min` is redudant now, remove it.\n> Best viewed with --color-words.\n> \n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xprepare.c | 19 ++++++++++++-------\n>   1 file changed, 12 insertions(+), 7 deletions(-)\n> \n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 386668a92d..bd8baf214d 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n>    * might be potentially discarded if they appear in a run of discardable.\n>    */\n>   static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n> -\tptrdiff_t i, nm, mlim;\n> +\tptrdiff_t i, nm, mlim1, mlim2;\n>   \txdlclass_t *rcrec;\n>   \tuint8_t *action1 = NULL, *action2 = NULL;\n>   \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> @@ -287,25 +287,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t\tgoto cleanup;\n>   \t}\n>   \n> +\tif (need_min) {\n> +\t\t/* i.e. infinity */\n> +\t\tmlim1 = PTRDIFF_MAX;\n> +\t\tmlim2 = PTRDIFF_MAX;\n\nThis is a nice improvement as it simplifies the checks below\n\n> +\t} else {\n> +\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n> +\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n\nAs Junio has pointed out we now evaluate xdl_bogosqrt() twice which is \nunfortunate. It would have been nice to mention that in the commit \nmessage and explain why it does not matter. Personally I find the old \ncode that set the limit just before each loop quite readable, the new \nversion sets mlim2 a long way before it is used.\n\nThanks\n\nPhillip\n\n> +\t}\n> +\n>   \t/*\n>   \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>   \t */\n> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n>   \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>   \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>   \t\trcrec = cf->rcrecs[mph1];\n>   \t\tnm = rcrec ? rcrec->len2 : 0;\n> -\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n>   \t}\n>   \n> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n>   \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>   \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>   \t\trcrec = cf->rcrecs[mph2];\n>   \t\tnm = rcrec ? rcrec->len1 : 0;\n> -\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n>   \t}\n>   \n>   \t/*\n\n"},{"id":"540497","messageId":"d4b72cd3-9429-458f-970c-8aa83dbf7286@gmail.com","threadId":"64712","inReplyTo":"xmqqtstw6hs2.fsf@gitster.g","subject":"Re: [PATCH v4 0/6] Xdiff cleanup part 3","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-03-31T09:45:10Z","receivedAt":"2026-03-31T09:45:13Z","isPatch":true,"body":"On 31/03/2026 00:04, Junio C Hamano wrote:\n> \"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> Changes in v3:\n>>\n>>   * run make DEVELOPER=1 on each commit and fix all compiler issues\n> \n> This round looks very good to me.  Let me mark it for 'next' unless\n> others bring up problems I failed to see in a few days.\n\nI've left a couple of comments, they're pretty minor though\n\nThanks\n\nPhillip\n\n"},{"id":"540533","messageId":"xmqq8qb82czd.fsf@gitster.g","threadId":"64712","inReplyTo":"32c34d0d-9358-43e3-9d58-5999b3ffd6c2@gmail.com","subject":"Re: [PATCH v4 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-31T16:13:58Z","receivedAt":"2026-03-31T16:14:00Z","isPatch":true,"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n>> +\t} else {\n>> +\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n>> +\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n>\n> As Junio has pointed out we now evaluate xdl_bogosqrt() twice which is \n> unfortunate. It would have been nice to mention that in the commit \n> message and explain why it does not matter.\n\nYup, that completely slipped my mind.  Personally I too find the\noriginal perfectly readable, but the updated one is not too bad,\neither.\n\nThanks.\n"},{"id":"540650","messageId":"87a54698-396d-4de8-bd9d-cd72f8d1e8df@gmail.com","threadId":"64712","inReplyTo":"fd14ccafc494aeda4bb9d05b83ac09f35bec8b52.1774890003.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 6/6] xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-04-01T16:00:34Z","receivedAt":"2026-04-01T16:00:36Z","isPatch":true,"body":"On 30/03/2026 18:00, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Make it clear that INVESTIGATE is turned into KEEP or DISCARD based on\n> the result of xdl_clean_mmatch() which reduces actionX[i] into a\n> boolean value.\n\nThis patch changes the diff output. If I compare the output of\n\n     git log --diff-merges=1 --diff-algorithm=myers -n1000 origin/master\n\nwith git built from 702083c4820 (xdiff/xdl_cleanup_records: make setting\naction easier to follow, 2026-03-30) and from 7ff1460b62f (xdiff/\nxdl_cleanup_records: simplify INVESTIGATE handling for clarity,\n2026-03-30) I see the diff below. I think the problem is that\nxdl_clean_mmatch() depends on the action arrays and this patch modifies\nthem because it converts INVESTIGATE to KEEP or DISCARD. We can probably\nwork around that by using a local variable rather than modifing\nactionX[i].\n\nThanks\n\nPhillip\n\ndiff --git a/tmp/p-sub-KADADDFP b/tmp/p-sub-PDCNCAOG\n--- a/tmp/p-sub-KADADDFP\n+++ b/tmp/p-sub-PDCNCAOG\n@@ -4258,14 +4258,15 @@ index 6485cb67068..0ff2e45aa7a 100644\n   \t\tctx.progress = NULL;\n\n  -\tctx.to_include = packs_to_include;\n+-\n+-\tfor_each_file_in_pack_dir(source->path, add_pack_to_midx, &ctx);\n  +\tif (ctx.compact) {\n  +\t\tint bitmap_order = 0;\n  +\t\tif (opts->preferred_pack_name)\n  +\t\t\tbitmap_order |= 1;\n  +\t\telse if (opts->flags & (MIDX_WRITE_REV_INDEX | MIDX_WRITE_BITMAP))\n  +\t\t\tbitmap_order |= 1;\n-\n--\tfor_each_file_in_pack_dir(source->path, add_pack_to_midx, &ctx);\n++\n  +\t\tfill_packs_from_midx_range(&ctx, bitmap_order);\n  +\t} else {\n  +\t\tctx.to_include = opts->packs_to_include;\n@@ -50839,7 +50840,8 @@ index 41b0750e5af..acaf42b2d93 100644\n  -\n  -\tif (opts->in_place)\n  -\t\toutfile = create_in_place_tempfile(file);\n--\n++\tread_input_file(&input, file);\n+\n  -\ttrailer_block = parse_trailers(opts, sb.buf, &head);\n  -\n  -\t/* Print the lines before the trailer block */\n@@ -50848,8 +50850,7 @@ index 41b0750e5af..acaf42b2d93 100644\n  -\n  -\tif (!opts->only_trailers && !blank_line_before_trailer_block(trailer_block))\n  -\t\tfprintf(outfile, \"\\n\");\n-+\tread_input_file(&input, file);\n-\n+-\n  -\n  -\tif (!opts->only_input) {\n  -\t\tLIST_HEAD(config_head);\n@@ -58880,12 +58881,12 @@ index 776de5356c9..84a31084d38 100644\n   \todb_prepare_alternates(odb);\n  -\tfor (source = odb->sources; source; source = source->next) {\n  -\t\tif (packfile_store_freshen_object(source->packfiles, oid))\n+-\t\t\treturn 1;\n+-\n+-\t\tif (odb_source_loose_freshen_object(source, oid))\n  +\tfor (source = odb->sources; source; source = source->next)\n  +\t\tif (odb_source_freshen_object(source, oid))\n\nThanks\n\nPhillip\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xprepare.c | 34 ++++++++++++++++++++++++----------\n>   1 file changed, 24 insertions(+), 10 deletions(-)\n> \n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 471d9567c9..1f2e8c6b4b 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -329,24 +329,38 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t */\n>   \txdf1->nreff = 0;\n>   \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n> -\t\tif (action1[i] == KEEP ||\n> -\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n> +\t\tif (action1[i] == INVESTIGATE) {\n> +\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n> +\t\t\t\taction1[i] = KEEP;\n> +\t\t\telse\n> +\t\t\t\taction1[i] = DISCARD;\n> +\t\t}\n> +\n> +\t\tif (action1[i] == KEEP) {\n>   \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> +\t\t\t/* changed[i] remains false */\n> +\t\t} else if (action1[i] == DISCARD)\n>   \t\t\txdf1->changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> +\t\telse\n> +\t\t\tBUG(\"Illegal state for action1[i]\");\n>   \t}\n>   \n>   \txdf2->nreff = 0;\n>   \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n> -\t\tif (action2[i] == KEEP ||\n> -\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n> +\t\tif (action2[i] == INVESTIGATE) {\n> +\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n> +\t\t\t\taction2[i] = KEEP;\n> +\t\t\telse\n> +\t\t\t\taction2[i] = DISCARD;\n> +\t\t}\n> +\n> +\t\tif (action2[i] == KEEP) {\n>   \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> +\t\t\t/* changed[i] remains false */\n> +\t\t} else if (action2[i] == DISCARD)\n>   \t\t\txdf2->changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> +\t\telse\n> +\t\t\tBUG(\"Illegal state for action2[i]\");\n>   \t}\n>   \n>   cleanup:\n\n"},{"id":"541172","messageId":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v4.git.git.1774890003.gitgitgadget@gmail.com","subject":"[PATCH v5 0/6] Xdiff cleanup part 3","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-08T20:26:22Z","receivedAt":"2026-04-08T20:26:31Z","isPatch":true,"body":"Changes in v5:\n\n * drop commit \"xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for\n   clarity\".\n * add braces around the else clause\n\nI didn't see a better way to rewrite how action is used so I reverted to\nwhat it used to be.\n\nChanges in v4:\n\n * Change SIZE_MAX to PTRDIFF_MAX.\n\nChanges in v3:\n\n * run make DEVELOPER=1 on each commit and fix all compiler issues\n\nv2 is a radical departure from v1 Changes in v2:\n\n * make the flow of xdl_cleanup_records() easier to follow\n\nThere is no performance or behavioral change introduced in this patch\nseries.\n\n=== original cover letter bellow ===\n\nPatch series summary:\n\n * patch 1: Introduce the ivec type\n * patch 2: Create the function xdl_do_classic_diff()\n * patches 3-4: generic cleanup\n * patches 5-8: convert from dstart/dend (in xdfile_t) to\n   delta_start/delta_end (in xdfenv_t)\n * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n   xdiffi.c\n\nThings that will be addressed in future patch series:\n\n * Make xdl_cleanup_records() easier to read\n * convert recs/nrec into an ivec\n * convert changed to an ivec\n * remove reference_index/nreff from xdfile_t and turn it into an ivec\n * splitting minimal_perfect_hash out as its own ivec\n * improve the performance of the classifier and parsing/hashing lines\n\n=== before this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\nsize_t nreff; } xdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n\n=== after this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\nxdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\ndelta_end; size_t mph_size; } xdfenv_t;\n\nEzekiel Newren (6):\n  xdiff/xdl_cleanup_records: delete local recs pointer\n  xdiff: use unambiguous types in xdl_bogo_sqrt()\n  xdiff/xdl_cleanup_records: use unambiguous types\n  xdiff/xdl_cleanup_records: make limits more clear\n  xdiff/xdl_cleanup_records: make setting action easier to follow\n  xdiff/xdl_cleanup_records: put braces around the else clause\n\n xdiff/xdiffi.c   |  2 +-\n xdiff/xprepare.c | 56 +++++++++++++++++++++++++++++++-----------------\n xdiff/xutils.c   |  4 ++--\n xdiff/xutils.h   |  2 +-\n 4 files changed, 40 insertions(+), 24 deletions(-)\n\n\nbase-commit: ca1db8a0f7dc0dbea892e99f5b37c5fe5861be71\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v5\nPull-Request: https://github.com/git/git/pull/2156\n\nRange-diff vs v4:\n\n 1:  da32a9747c = 1:  b31924a949 xdiff/xdl_cleanup_records: delete local recs pointer\n 2:  86b0ad100c = 2:  1822166fef xdiff: use unambiguous types in xdl_bogo_sqrt()\n 3:  39a35365ae = 3:  85aa0da90c xdiff/xdl_cleanup_records: use unambiguous types\n 4:  75fe3ea125 = 4:  fec2b0f38a xdiff/xdl_cleanup_records: make limits more clear\n 5:  0cf1412d01 = 5:  88c68fa89a xdiff/xdl_cleanup_records: make setting action easier to follow\n 6:  fd14ccafc4 ! 6:  699e198fa9 xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n     @@ Metadata\n      Author: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## Commit message ##\n     -    xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for clarity\n     -\n     -    Make it clear that INVESTIGATE is turned into KEEP or DISCARD based on\n     -    the result of xdl_clean_mmatch() which reduces actionX[i] into a\n     -    boolean value.\n     +    xdiff/xdl_cleanup_records: put braces around the else clause\n      \n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xprepare.c ##\n      @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \t */\n     - \txdf1->nreff = 0;\n     - \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n     --\t\tif (action1[i] == KEEP ||\n     --\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n     -+\t\tif (action1[i] == INVESTIGATE) {\n     -+\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n     -+\t\t\t\taction1[i] = KEEP;\n     -+\t\t\telse\n     -+\t\t\t\taction1[i] = DISCARD;\n     -+\t\t}\n     -+\n     -+\t\tif (action1[i] == KEEP) {\n     + \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n       \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n     --\t\t\t/* changed[i] remains false, i.e. keep */\n     + \t\t\t/* changed[i] remains false, i.e. keep */\n      -\t\t} else\n     -+\t\t\t/* changed[i] remains false */\n     -+\t\t} else if (action1[i] == DISCARD)\n     ++\t\t} else {\n       \t\t\txdf1->changed[i] = true;\n     --\t\t\t/* i.e. discard */\n     -+\t\telse\n     -+\t\t\tBUG(\"Illegal state for action1[i]\");\n     + \t\t\t/* i.e. discard */\n     ++\t\t}\n       \t}\n       \n       \txdf2->nreff = 0;\n     - \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n     --\t\tif (action2[i] == KEEP ||\n     --\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n     -+\t\tif (action2[i] == INVESTIGATE) {\n     -+\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n     -+\t\t\t\taction2[i] = KEEP;\n     -+\t\t\telse\n     -+\t\t\t\taction2[i] = DISCARD;\n     -+\t\t}\n     -+\n     -+\t\tif (action2[i] == KEEP) {\n     +@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     + \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n       \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n     --\t\t\t/* changed[i] remains false, i.e. keep */\n     + \t\t\t/* changed[i] remains false, i.e. keep */\n      -\t\t} else\n     -+\t\t\t/* changed[i] remains false */\n     -+\t\t} else if (action2[i] == DISCARD)\n     ++\t\t} else {\n       \t\t\txdf2->changed[i] = true;\n     --\t\t\t/* i.e. discard */\n     -+\t\telse\n     -+\t\t\tBUG(\"Illegal state for action2[i]\");\n     + \t\t\t/* i.e. discard */\n     ++\t\t}\n       \t}\n       \n       cleanup:\n\n-- \ngitgitgadget\n"},{"id":"541173","messageId":"b31924a94966686883079feff5dcbff071bc57e1.1775679988.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"[PATCH v5 1/6] xdiff/xdl_cleanup_records: delete local recs pointer","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-08T20:26:23Z","receivedAt":"2026-04-08T20:26:33Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nSimplify the first 2 for loops by directly indexing the xdfile.recs.\nrecs is unused in the last 2 for loops, remove it. Best viewed with\n--color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 17 ++++++++---------\n 1 file changed, 8 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex cd4fc405eb..d6e1901d2d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -269,7 +269,6 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n \tlong i, nm, mlim;\n-\txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -293,16 +292,18 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n+\t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n \tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n+\t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -312,8 +313,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * false, or become true.\n \t */\n \txdf1->nreff = 0;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n-\t     i <= xdf1->dend; i++, recs++) {\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n@@ -324,8 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t}\n \n \txdf2->nreff = 0;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n-\t     i <= xdf2->dend; i++, recs++) {\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-- \ngitgitgadget\n\n"},{"id":"541174","messageId":"1822166fef0c5dcafb4f3c717eff235db6404342.1775679988.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"[PATCH v5 2/6] xdiff: use unambiguous types in xdl_bogo_sqrt()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-08T20:26:24Z","receivedAt":"2026-04-08T20:26:34Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThere is no real square root for a negative number and size_t may not\nbe large enough for certain applications, replace long with uint64_t.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   | 2 +-\n xdiff/xprepare.c | 4 ++--\n xdiff/xutils.c   | 4 ++--\n xdiff/xutils.h   | 2 +-\n 4 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 4376f943db..88708c12a3 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -348,7 +348,7 @@ int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \tkvdf += xe->xdf2.nreff + 1;\n \tkvdb += xe->xdf2.nreff + 1;\n \n-\txenv.mxcost = xdl_bogosqrt(ndiags);\n+\txenv.mxcost = (long)xdl_bogosqrt((uint64_t)ndiags);\n \tif (xenv.mxcost < XDL_MAX_COST_MIN)\n \t\txenv.mxcost = XDL_MAX_COST_MIN;\n \txenv.snake_cnt = XDL_SNAKE_CNT;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex d6e1901d2d..48fb5ce6fe 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n@@ -299,7 +299,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 77ee1ad9c8..9a999acdc0 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -23,8 +23,8 @@\n #include \"xinclude.h\"\n \n \n-long xdl_bogosqrt(long n) {\n-\tlong i;\n+uint64_t xdl_bogosqrt(uint64_t n) {\n+\tuint64_t i;\n \n \t/*\n \t * Classical integer square root approximation using shifts.\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 615b4a9d35..58f9d74cda 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -25,7 +25,7 @@\n \n \n \n-long xdl_bogosqrt(long n);\n+uint64_t xdl_bogosqrt(uint64_t n);\n int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n \t\t     xdemitcb_t *ecb);\n int xdl_cha_init(chastore_t *cha, long isize, long icount);\n-- \ngitgitgadget\n\n"},{"id":"541175","messageId":"85aa0da90c62d9217ac2a3f907c37855a95298b9.1775679988.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"[PATCH v5 3/6] xdiff/xdl_cleanup_records: use unambiguous types","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-08T20:26:25Z","receivedAt":"2026-04-08T20:26:36Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nChange the parameters of xdl_clean_mmatch() and the local variables\ni, nm, mlim in xdl_cleanup_records() to use unambiguous types. Best\nviewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 48fb5ce6fe..386668a92d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -197,8 +197,8 @@ void xdl_free_env(xdfenv_t *xe) {\n }\n \n \n-static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n-\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n+static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, ptrdiff_t e) {\n+\tptrdiff_t r, rdis0, rpdis0, rdis1, rpdis1;\n \n \t/*\n \t * Limits the window that is examined during the similar-lines\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, mlim;\n+\tptrdiff_t i, nm, mlim;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n-- \ngitgitgadget\n\n"},{"id":"541176","messageId":"fec2b0f38ae30deceda104ec140fa4d404324103.1775679988.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"[PATCH v5 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-08T20:26:26Z","receivedAt":"2026-04-08T20:26:38Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake the handling of per-file limits and the minimal-case clearer.\n  * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n    them.\n  * The additional condition `!need_min` is redudant now, remove it.\nBest viewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 19 ++++++++++++-------\n 1 file changed, 12 insertions(+), 7 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 386668a92d..bd8baf214d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tptrdiff_t i, nm, mlim;\n+\tptrdiff_t i, nm, mlim1, mlim2;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -287,25 +287,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tgoto cleanup;\n \t}\n \n+\tif (need_min) {\n+\t\t/* i.e. infinity */\n+\t\tmlim1 = PTRDIFF_MAX;\n+\t\tmlim2 = PTRDIFF_MAX;\n+\t} else {\n+\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n+\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n+\t}\n+\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"541177","messageId":"88c68fa89a263a6031fbfc201e930e0c3d9ec6db.1775679988.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"[PATCH v5 5/6] xdiff/xdl_cleanup_records: make setting action easier to follow","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-08T20:26:27Z","receivedAt":"2026-04-08T20:26:39Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nRewrite nested ternaries with a clear if/else ladder for\naction1/action2 to improve readability while preserving\nbehavior.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 ++++++++++++--\n 1 file changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex bd8baf214d..471d9567c9 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -303,14 +303,24 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction1[i] = DISCARD;\n+\t\telse if (nm < mlim1)\n+\t\t\taction1[i] = KEEP;\n+\t\telse /* nm >= mlim1 */\n+\t\t\taction1[i] = INVESTIGATE;\n \t}\n \n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction2[i] = DISCARD;\n+\t\telse if (nm < mlim2)\n+\t\t\taction2[i] = KEEP;\n+\t\telse /* nm >= mlim2 */\n+\t\t\taction2[i] = INVESTIGATE;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"541178","messageId":"699e198fa9bdd4b6829d7fbd550b7d387bb884d0.1775679988.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"[PATCH v5 6/6] xdiff/xdl_cleanup_records: put braces around the else clause","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-08T20:26:28Z","receivedAt":"2026-04-08T20:26:41Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 6 ++++--\n 1 file changed, 4 insertions(+), 2 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 471d9567c9..18ee7e815c 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -333,9 +333,10 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t} else {\n \t\t\txdf1->changed[i] = true;\n \t\t\t/* i.e. discard */\n+\t\t}\n \t}\n \n \txdf2->nreff = 0;\n@@ -344,9 +345,10 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n \t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t} else {\n \t\t\txdf2->changed[i] = true;\n \t\t\t/* i.e. discard */\n+\t\t}\n \t}\n \n cleanup:\n-- \ngitgitgadget\n"},{"id":"541185","messageId":"xmqqh5plxhta.fsf@gitster.g","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 0/6] Xdiff cleanup part 3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-08T21:28:49Z","receivedAt":"2026-04-08T21:28:52Z","isPatch":true,"body":"\"Ezekiel Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Changes in v5:\n>\n>  * drop commit \"xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for\n>    clarity\".\n>  * add braces around the else clause\n>\n> I didn't see a better way to rewrite how action is used so I reverted to\n> what it used to be.\n\nThanks, will replace.\n\nPhillip's pw/xdiff-shrink-memory-consumption topic was built on the\nprevious iteration of this topic, so I took the liberty of rebasing\nit on top.\n\nPhillip, can you double check for mistakes when I push the result\nout later today?  These two topics should appear near the tip of\n'seen' next to each other.  Thanks.\n\n\n\n"},{"id":"541272","messageId":"1ef8dc54-871d-4a3e-80c1-a689fcb883a8@gmail.com","threadId":"64712","inReplyTo":"xmqqh5plxhta.fsf@gitster.g","subject":"Re: [PATCH v5 0/6] Xdiff cleanup part 3","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-04-09T14:01:05Z","receivedAt":"2026-04-09T14:01:15Z","isPatch":true,"body":"On 08/04/2026 22:28, Junio C Hamano wrote:\n> \n> Thanks, will replace.\n> \n> Phillip's pw/xdiff-shrink-memory-consumption topic was built on the\n> previous iteration of this topic, so I took the liberty of rebasing\n> it on top.\n> \n> Phillip, can you double check for mistakes when I push the result\n> out later today?  These two topics should appear near the tip of\n> 'seen' next to each other.  Thanks.\n\nThanks for rebasing that topic. To check I rebased my branch locally and \nit matches what you have in seen.\n\nThanks\n\nPhillip\n\n"},{"id":"541539","messageId":"df244360-e9a9-44c0-946d-29288e6dd269@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 0/6] Xdiff cleanup part 3","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-04-14T10:08:11Z","receivedAt":"2026-04-14T10:08:31Z","isPatch":true,"body":"On 08/04/2026 21:26, Ezekiel Newren via GitGitGadget wrote:\n> Changes in v5:\n> \n>   * drop commit \"xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for\n>     clarity\".\n>   * add braces around the else clause\n> \n> I didn't see a better way to rewrite how action is used so I reverted to\n> what it used to be.\n\nThat's a shame, the diff below uses a local variable to avoid altering the\narrays as suggested in [1]. The comments about the double evaluation of\nxdl_bogosort() in patch 4 [2,3] also seem to have been overlooked.\n\nThanks\n\nPhillip\n\n[1] https://lore.kernel.org/git/87a54698-396d-4de8-bd9d-cd72f8d1e8df@gmail.com\n[2] https://lore.kernel.org/git/32c34d0d-9358-43e3-9d58-5999b3ffd6c2@gmail.com\n[3] https://lore.kernel.org/git/xmqqcy0oj2s1.fsf@gitster.g\n\n---- 8< ----\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 471d9567c9..9966b4715d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -329,24 +329,42 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n  \t */\n  \txdf1->nreff = 0;\n  \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n-\t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n+\t\tuint8_t action = action1[i];\n+\n+\t\tif (action == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n+\t\t\t\taction = KEEP;\n+\t\t\telse\n+\t\t\t\taction = DISCARD;\n+\t\t}\n+\n+\t\tif (action == KEEP) {\n  \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action == DISCARD)\n  \t\t\txdf1->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action1[i]\");\n  \t}\n\n  \txdf2->nreff = 0;\n  \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n-\t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n+\t\tuint8_t action = action2[i];\n+\n+\t\tif (action == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n+\t\t\t\taction = KEEP;\n+\t\t\telse\n+\t\t\t\taction = DISCARD;\n+\t\t}\n+\n+\t\tif (action == KEEP) {\n  \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action == DISCARD)\n  \t\t\txdf2->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\telse\n+\t\t\tBUG(\"Illegal state for action2[i]\");\n  \t}\n\n  cleanup:\n\n"},{"id":"541540","messageId":"d88af7e1-e8dd-4423-9c6c-977e1f1dc074@gmail.com","threadId":"64712","inReplyTo":"fec2b0f38ae30deceda104ec140fa4d404324103.1775679988.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-04-14T10:09:00Z","receivedAt":"2026-04-14T10:09:12Z","isPatch":true,"body":"On 08/04/2026 21:26, Ezekiel Newren via GitGitGadget wrote:\n> From: Ezekiel Newren <ezekielnewren@gmail.com>\n> \n> Make the handling of per-file limits and the minimal-case clearer.\n>    * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n>      them.\n>    * The additional condition `!need_min` is redudant now, remove it.\n> Best viewed with --color-words.\n\nThis still suffers from double evaluation in XDL_MIN(), perhaps we\ncould do something like the diff below instead?\n\nThanks\n\nPhillip\n\n---- 8< ----\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 386668a92d7..bde3cc7f3e5 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -290,22 +290,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n  \t/*\n  \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n  \t */\n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n+\tif (need_min) {\n+\t\tmlim = PTRDIFF_MAX;\n+\t} else {\n+\t\tmlim = xdl_bogosqrt((uint64_t)xdf1->nrec);\n+\t\tif (mlim > XDL_MAX_EQLIMIT)\n+\t\t\tmlim = XDL_MAX_EQLIMIT;\n+\t}\n  \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n  \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n  \t\trcrec = cf->rcrecs[mph1];\n  \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim ? INVESTIGATE: KEEP;\n  \t}\n\n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n+\tif (!need_min) {\n+\t\tmlim = xdl_bogosqrt((uint64_t)xdf2->nrec);\n+\t\tif (mlim > XDL_MAX_EQLIMIT)\n+\t\t\tmlim = XDL_MAX_EQLIMIT;\n+\t}\n  \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n  \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n  \t\trcrec = cf->rcrecs[mph2];\n  \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim ? INVESTIGATE: KEEP;\n  \t}\n\n  \t/*\n---- >8 ----\n\n> Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> ---\n>   xdiff/xprepare.c | 19 ++++++++++++-------\n>   1 file changed, 12 insertions(+), 7 deletions(-)\n> \n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 386668a92d..bd8baf214d 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n>    * might be potentially discarded if they appear in a run of discardable.\n>    */\n>   static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n> -\tptrdiff_t i, nm, mlim;\n> +\tptrdiff_t i, nm, mlim1, mlim2;\n>   \txdlclass_t *rcrec;\n>   \tuint8_t *action1 = NULL, *action2 = NULL;\n>   \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> @@ -287,25 +287,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t\tgoto cleanup;\n>   \t}\n>   \n> +\tif (need_min) {\n> +\t\t/* i.e. infinity */\n> +\t\tmlim1 = PTRDIFF_MAX;\n> +\t\tmlim2 = PTRDIFF_MAX;\n> +\t} else {\n> +\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n> +\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n> +\t}\n> +\n>   \t/*\n>   \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>   \t */\n> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n>   \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>   \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>   \t\trcrec = cf->rcrecs[mph1];\n>   \t\tnm = rcrec ? rcrec->len2 : 0;\n> -\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n>   \t}\n>   \n> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n>   \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>   \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>   \t\trcrec = cf->rcrecs[mph2];\n>   \t\tnm = rcrec ? rcrec->len1 : 0;\n> -\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n> +\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n>   \t}\n>   \n>   \t/*\n\n"},{"id":"541571","messageId":"xmqqeckhbhep.fsf@gitster.g","threadId":"64712","inReplyTo":"df244360-e9a9-44c0-946d-29288e6dd269@gmail.com","subject":"Re: [PATCH v5 0/6] Xdiff cleanup part 3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-14T17:06:38Z","receivedAt":"2026-04-14T17:06:40Z","isPatch":true,"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> On 08/04/2026 21:26, Ezekiel Newren via GitGitGadget wrote:\n>> Changes in v5:\n>> \n>>   * drop commit \"xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for\n>>     clarity\".\n>>   * add braces around the else clause\n>> \n>> I didn't see a better way to rewrite how action is used so I reverted to\n>> what it used to be.\n>\n> That's a shame, the diff below uses a local variable to avoid altering the\n> arrays as suggested in [1]. The comments about the double evaluation of\n> xdl_bogosort() in patch 4 [2,3] also seem to have been overlooked.\n\nThanks for keeping an eye on this topic.  Very much appreciated.\n\n>\n> Thanks\n>\n> Phillip\n>\n> [1] https://lore.kernel.org/git/87a54698-396d-4de8-bd9d-cd72f8d1e8df@gmail.com\n> [2] https://lore.kernel.org/git/32c34d0d-9358-43e3-9d58-5999b3ffd6c2@gmail.com\n> [3] https://lore.kernel.org/git/xmqqcy0oj2s1.fsf@gitster.g\n>\n> ---- 8< ----\n> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> index 471d9567c9..9966b4715d 100644\n> --- a/xdiff/xprepare.c\n> +++ b/xdiff/xprepare.c\n> @@ -329,24 +329,42 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>   \t */\n>   \txdf1->nreff = 0;\n>   \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n> -\t\tif (action1[i] == KEEP ||\n> -\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n> +\t\tuint8_t action = action1[i];\n> +\n> +\t\tif (action == INVESTIGATE) {\n> +\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n> +\t\t\t\taction = KEEP;\n> +\t\t\telse\n> +\t\t\t\taction = DISCARD;\n> +\t\t}\n> +\n> +\t\tif (action == KEEP) {\n>   \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> +\t\t\t/* changed[i] remains false */\n> +\t\t} else if (action == DISCARD)\n>   \t\t\txdf1->changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> +\t\telse\n> +\t\t\tBUG(\"Illegal state for action1[i]\");\n>   \t}\n>\n>   \txdf2->nreff = 0;\n>   \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n> -\t\tif (action2[i] == KEEP ||\n> -\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n> +\t\tuint8_t action = action2[i];\n> +\n> +\t\tif (action == INVESTIGATE) {\n> +\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n> +\t\t\t\taction = KEEP;\n> +\t\t\telse\n> +\t\t\t\taction = DISCARD;\n> +\t\t}\n> +\n> +\t\tif (action == KEEP) {\n>   \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n> -\t\t\t/* changed[i] remains false, i.e. keep */\n> -\t\t} else\n> +\t\t\t/* changed[i] remains false */\n> +\t\t} else if (action == DISCARD)\n>   \t\t\txdf2->changed[i] = true;\n> -\t\t\t/* i.e. discard */\n> +\t\telse\n> +\t\t\tBUG(\"Illegal state for action2[i]\");\n>   \t}\n>\n>   cleanup:\n"},{"id":"541596","messageId":"CAH=ZcbCX8FEs4ueU7+groQp8XhiaP0QPHMeGqT+Ap1FjeW9foQ@mail.gmail.com","threadId":"64712","inReplyTo":"32c34d0d-9358-43e3-9d58-5999b3ffd6c2@gmail.com","subject":"Re: [PATCH v4 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-04-14T21:58:03Z","receivedAt":"2026-04-14T21:58:15Z","isPatch":true,"body":"On Tue, Mar 31, 2026 at 3:44 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 30/03/2026 18:00, Ezekiel Newren via GitGitGadget wrote:\n> > From: Ezekiel Newren <ezekielnewren@gmail.com>\n> >\n> > Make the handling of per-file limits and the minimal-case clearer.\n> >    * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n> >      them.\n> >    * The additional condition `!need_min` is redudant now, remove it.\n> > Best viewed with --color-words.\n> >\n> > Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n> > ---\n> >   xdiff/xprepare.c | 19 ++++++++++++-------\n> >   1 file changed, 12 insertions(+), 7 deletions(-)\n> >\n> > diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\n> > index 386668a92d..bd8baf214d 100644\n> > --- a/xdiff/xprepare.c\n> > +++ b/xdiff/xprepare.c\n> > @@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n> >    * might be potentially discarded if they appear in a run of discardable.\n> >    */\n> >   static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n> > -     ptrdiff_t i, nm, mlim;\n> > +     ptrdiff_t i, nm, mlim1, mlim2;\n> >       xdlclass_t *rcrec;\n> >       uint8_t *action1 = NULL, *action2 = NULL;\n> >       bool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n> > @@ -287,25 +287,30 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n> >               goto cleanup;\n> >       }\n> >\n> > +     if (need_min) {\n> > +             /* i.e. infinity */\n> > +             mlim1 = PTRDIFF_MAX;\n> > +             mlim2 = PTRDIFF_MAX;\n>\n> This is a nice improvement as it simplifies the checks below\n>\n> > +     } else {\n> > +             mlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n> > +             mlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n>\n> As Junio has pointed out we now evaluate xdl_bogosqrt() twice which is\n> unfortunate. It would have been nice to mention that in the commit\n> message and explain why it does not matter.\n\nIt doesn't matter because xdl_bogosqrt() was being called twice before\nand is being called twice now. There is no change in that regard.\nThat's why I split mlim into 2 variables to make it more clear.\n\nIt looks like you and Junio have both missed that xdl_bogo_sqrt() is\nbeing called on different values.\n"},{"id":"541601","messageId":"xmqqik9t89yt.fsf@gitster.g","threadId":"64712","inReplyTo":"CAH=ZcbCX8FEs4ueU7+groQp8XhiaP0QPHMeGqT+Ap1FjeW9foQ@mail.gmail.com","subject":"Re: [PATCH v4 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-14T22:15:38Z","receivedAt":"2026-04-14T22:15:41Z","isPatch":true,"body":"Ezekiel Newren <ezekielnewren@gmail.com> writes:\n\n>> > +             mlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n>> > +             mlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n>>\n>> As Junio has pointed out we now evaluate xdl_bogosqrt() twice which is\n>> unfortunate. It would have been nice to mention that in the commit\n>> message and explain why it does not matter.\n>\n> It doesn't matter because xdl_bogosqrt() was being called twice before\n> and is being called twice now. There is no change in that regard.\n> That's why I split mlim into 2 variables to make it more clear.\n>\n> It looks like you and Junio have both missed that xdl_bogo_sqrt() is\n> being called on different values.\n\nI think Phillip's point is that XDL_MIN(a, b) would evaluate (a)\ntwice.\n\n\t#define XDL_MIN(a, b) ((a) < (b) ? (a): (b))\n\nSo the code you have above\n\n\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n\nactually is\n\n\tmlim1 = ((xdl_bogosqrt(xdf1->nrec) < XDL_MAX_EQLIMIT) \n\t\t? xdl_bogosqrt(xdf1->nrec)\n\t\t: XDL_MAX_EQLIMIT);\n\nIf you are lucky and xdf1->nrec is so large, there is only one call\nto xdl_bogosqrt() before mlim1 gets assigned XDL_MAX_EQLIMIT, but\nusually you'll call it on the same xdf1->nrec twice before you\nassign the result to mlim1, no?\n\nThe original lost by the patch looked like this:\n\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n\nwhich computed it once, assigned it to mlim, and then clamped.\n"},{"id":"541656","messageId":"50ea6e41-6b29-46ad-aa97-0eaa289db7cf@gmail.com","threadId":"64712","inReplyTo":"xmqqik9t89yt.fsf@gitster.g","subject":"Re: [PATCH v4 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-04-15T13:54:40Z","receivedAt":"2026-04-15T13:54:46Z","isPatch":true,"body":"On 14/04/2026 23:15, Junio C Hamano wrote:\n> Ezekiel Newren <ezekielnewren@gmail.com> writes:\n> \n>>>> +             mlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n>>>> +             mlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n>>>\n>>> As Junio has pointed out we now evaluate xdl_bogosqrt() twice which is\n>>> unfortunate. It would have been nice to mention that in the commit\n>>> message and explain why it does not matter.\n>>\n>> It doesn't matter because xdl_bogosqrt() was being called twice before\n>> and is being called twice now. There is no change in that regard.\n>> That's why I split mlim into 2 variables to make it more clear.\n>>\n>> It looks like you and Junio have both missed that xdl_bogo_sqrt() is\n>> being called on different values.\n> \n> I think Phillip's point is that XDL_MIN(a, b) would evaluate (a)\n> twice.\n\nExactly, thanks for clarifying\n\nPhillip\n\n> \n> \t#define XDL_MIN(a, b) ((a) < (b) ? (a): (b))\n> \n> So the code you have above\n> \n> \tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n> \n> actually is\n> \n> \tmlim1 = ((xdl_bogosqrt(xdf1->nrec) < XDL_MAX_EQLIMIT)\n> \t\t? xdl_bogosqrt(xdf1->nrec)\n> \t\t: XDL_MAX_EQLIMIT);\n> \n> If you are lucky and xdf1->nrec is so large, there is only one call\n> to xdl_bogosqrt() before mlim1 gets assigned XDL_MAX_EQLIMIT, but\n> usually you'll call it on the same xdf1->nrec twice before you\n> assign the result to mlim1, no?\n> \n> The original lost by the patch looked like this:\n> \n>   \t/*\n>   \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>   \t */\n> -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n> -\t\tmlim = XDL_MAX_EQLIMIT;\n> \n> which computed it once, assigned it to mlim, and then clamped.\n> \n\n"},{"id":"542482","messageId":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v5.git.git.1775679988.gitgitgadget@gmail.com","subject":"[PATCH v6 0/6] Xdiff cleanup part 3","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-29T22:08:09Z","receivedAt":"2026-04-29T22:08:18Z","isPatch":true,"body":"Changes in v6:\n\n * implement suggestions by Phillip Wood [1,2]\n\nPhillip's second \"if\" in [1] differs from his first one. In my changes I\nmade both of them structurally the same.\n\nSomething I'm confused by is the range-diff of patch 5. I'm confused why\nrange-diff states that this is different at all. I don't think this is a\nproblem, I just don't like not being able to explain a difference pointed\nout by range-diff.\n\n5: 88c68fa89a ! 5: 099b08c33f xdiff/xdl_cleanup_records: make setting action\neasier to follow @@ xdiff/xprepare.c: static int\nxdl_cleanup_records(xdlclassifier_t *cf, xdfile_t * + action1[i] =\nINVESTIGATE; }\n\n-   for (i = xdf2->dstart; i <= xdf2->dend; i++) {\n+   if (need_min) {\n+@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n            size_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n            rcrec = cf->rcrecs[mph2];\n            nm = rcrec ? rcrec->len1 : 0;\n\n\n[1] limits\nhttps://lore.kernel.org/git/d88af7e1-e8dd-4423-9c6c-977e1f1dc074@gmail.com/\n[2] action execution\nhttps://lore.kernel.org/git/df244360-e9a9-44c0-946d-29288e6dd269@gmail.com/\n\nChanges in v5:\n\n * drop commit \"xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for\n   clarity\".\n * add braces around the else clause\n\nI didn't see a better way to rewrite how action is used so I reverted to\nwhat it used to be.\n\nChanges in v4:\n\n * Change SIZE_MAX to PTRDIFF_MAX.\n\nChanges in v3:\n\n * run make DEVELOPER=1 on each commit and fix all compiler issues\n\nv2 is a radical departure from v1 Changes in v2:\n\n * make the flow of xdl_cleanup_records() easier to follow\n\nThere is no performance or behavioral change introduced in this patch\nseries.\n\n=== original cover letter bellow ===\n\nPatch series summary:\n\n * patch 1: Introduce the ivec type\n * patch 2: Create the function xdl_do_classic_diff()\n * patches 3-4: generic cleanup\n * patches 5-8: convert from dstart/dend (in xdfile_t) to\n   delta_start/delta_end (in xdfenv_t)\n * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n   xdiffi.c\n\nThings that will be addressed in future patch series:\n\n * Make xdl_cleanup_records() easier to read\n * convert recs/nrec into an ivec\n * convert changed to an ivec\n * remove reference_index/nreff from xdfile_t and turn it into an ivec\n * splitting minimal_perfect_hash out as its own ivec\n * improve the performance of the classifier and parsing/hashing lines\n\n=== before this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\nsize_t nreff; } xdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n\n=== after this patch series typedef struct s_xdfile { xrecord_t *recs;\nsize_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\nxdfile_t;\n\ntypedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\ndelta_end; size_t mph_size; } xdfenv_t;\n\nEzekiel Newren (6):\n  xdiff/xdl_cleanup_records: delete local recs pointer\n  xdiff: use unambiguous types in xdl_bogo_sqrt()\n  xdiff/xdl_cleanup_records: use unambiguous types\n  xdiff/xdl_cleanup_records: make limits more clear\n  xdiff/xdl_cleanup_records: make setting action easier to follow\n  xdiff/xdl_cleanup_records: make execution of action easier to follow\n\n xdiff/xdiffi.c   |  2 +-\n xdiff/xprepare.c | 97 ++++++++++++++++++++++++++++++++++--------------\n xdiff/xutils.c   |  4 +-\n xdiff/xutils.h   |  2 +-\n 4 files changed, 73 insertions(+), 32 deletions(-)\n\n\nbase-commit: ca1db8a0f7dc0dbea892e99f5b37c5fe5861be71\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v6\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v6\nPull-Request: https://github.com/git/git/pull/2156\n\nRange-diff vs v5:\n\n 1:  b31924a949 = 1:  b31924a949 xdiff/xdl_cleanup_records: delete local recs pointer\n 2:  1822166fef = 2:  1822166fef xdiff: use unambiguous types in xdl_bogo_sqrt()\n 3:  85aa0da90c = 3:  85aa0da90c xdiff/xdl_cleanup_records: use unambiguous types\n 4:  fec2b0f38a ! 4:  51c62ed454 xdiff/xdl_cleanup_records: make limits more clear\n     @@ Commit message\n            * The additional condition `!need_min` is redudant now, remove it.\n          Best viewed with --color-words.\n      \n     +    Helped-by: Phillip Wood\n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xprepare.c ##\n     @@ xdiff/xprepare.c: static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t\n       \tuint8_t *action1 = NULL, *action2 = NULL;\n       \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n      @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \t\tgoto cleanup;\n     - \t}\n     - \n     -+\tif (need_min) {\n     -+\t\t/* i.e. infinity */\n     -+\t\tmlim1 = PTRDIFF_MAX;\n     -+\t\tmlim2 = PTRDIFF_MAX;\n     -+\t} else {\n     -+\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n     -+\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n     -+\t}\n     -+\n       \t/*\n       \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n       \t */\n      -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n      -\t\tmlim = XDL_MAX_EQLIMIT;\n     ++\tif (need_min) {\n     ++\t\t/* i.e. infinity */\n     ++\t\tmlim1 = PTRDIFF_MAX;\n     ++\t} else {\n     ++\t\tmlim1 = xdl_bogosqrt((uint64_t)xdf1->nrec);\n     ++\t\tif (mlim1 > XDL_MAX_EQLIMIT)\n     ++\t\t\tmlim1 = XDL_MAX_EQLIMIT;\n     ++\t}\n       \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n       \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n       \t\trcrec = cf->rcrecs[mph1];\n     @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n       \n      -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n      -\t\tmlim = XDL_MAX_EQLIMIT;\n     ++\tif (need_min) {\n     ++\t\t/* i.e. infinity */\n     ++\t\tmlim2 = PTRDIFF_MAX;\n     ++\t} else {\n     ++\t\tmlim2 = xdl_bogosqrt((uint64_t)xdf2->nrec);\n     ++\t\tif (mlim2 > XDL_MAX_EQLIMIT)\n     ++\t\t\tmlim2 = XDL_MAX_EQLIMIT;\n     ++\t}\n       \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n       \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n       \t\trcrec = cf->rcrecs[mph2];\n 5:  88c68fa89a ! 5:  45ad2ae62d xdiff/xdl_cleanup_records: make setting action easier to follow\n     @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n      +\t\t\taction1[i] = INVESTIGATE;\n       \t}\n       \n     - \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n     + \tif (need_min) {\n     +@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n       \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n       \t\trcrec = cf->rcrecs[mph2];\n       \t\tnm = rcrec ? rcrec->len1 : 0;\n 6:  699e198fa9 ! 6:  a5174802f4 xdiff/xdl_cleanup_records: put braces around the else clause\n     @@ Metadata\n      Author: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## Commit message ##\n     -    xdiff/xdl_cleanup_records: put braces around the else clause\n     +    xdiff/xdl_cleanup_records: make execution of action easier to follow\n      \n     +    Helped-by: Phillip Wood\n          Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n      \n       ## xdiff/xprepare.c ##\n      @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n     + \t */\n     + \txdf1->nreff = 0;\n     + \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n     +-\t\tif (action1[i] == KEEP ||\n     +-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n     ++\t\tuint8_t action = action1[i];\n     ++\n     ++\t\tif (action == INVESTIGATE) {\n     ++\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n     ++\t\t\t\taction = KEEP;\n     ++\t\t\telse\n     ++\t\t\t\taction = DISCARD;\n     ++\t\t}\n     ++\n     ++\t\tif (action == KEEP) {\n       \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n     - \t\t\t/* changed[i] remains false, i.e. keep */\n     +-\t\t\t/* changed[i] remains false, i.e. keep */\n      -\t\t} else\n     -+\t\t} else {\n     ++\t\t\t/* changed[i] remains false */\n     ++\t\t} else if (action == DISCARD) {\n       \t\t\txdf1->changed[i] = true;\n     - \t\t\t/* i.e. discard */\n     +-\t\t\t/* i.e. discard */\n     ++\t\t} else {\n     ++\t\t\tBUG(\"Illegal state for action\");\n      +\t\t}\n       \t}\n       \n       \txdf2->nreff = 0;\n     -@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n     - \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n     + \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n     +-\t\tif (action2[i] == KEEP ||\n     +-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n     ++\t\tuint8_t action = action2[i];\n     ++\n     ++\t\tif (action == INVESTIGATE) {\n     ++\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n     ++\t\t\t\taction = KEEP;\n     ++\t\t\telse\n     ++\t\t\t\taction = DISCARD;\n     ++\t\t}\n     ++\n     ++\t\tif (action == KEEP) {\n       \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n     - \t\t\t/* changed[i] remains false, i.e. keep */\n     +-\t\t\t/* changed[i] remains false, i.e. keep */\n      -\t\t} else\n     -+\t\t} else {\n     ++\t\t\t/* changed[i] remains false */\n     ++\t\t} else if (action == DISCARD) {\n       \t\t\txdf2->changed[i] = true;\n     - \t\t\t/* i.e. discard */\n     +-\t\t\t/* i.e. discard */\n     ++\t\t} else {\n     ++\t\t\tBUG(\"Illegal state for action\");\n      +\t\t}\n       \t}\n       \n\n-- \ngitgitgadget\n"},{"id":"542483","messageId":"b31924a94966686883079feff5dcbff071bc57e1.1777500495.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","subject":"[PATCH v6 1/6] xdiff/xdl_cleanup_records: delete local recs pointer","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-29T22:08:10Z","receivedAt":"2026-04-29T22:08:19Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nSimplify the first 2 for loops by directly indexing the xdfile.recs.\nrecs is unused in the last 2 for loops, remove it. Best viewed with\n--color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 17 ++++++++---------\n 1 file changed, 8 insertions(+), 9 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex cd4fc405eb..d6e1901d2d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -269,7 +269,6 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n \tlong i, nm, mlim;\n-\txrecord_t *recs;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -293,16 +292,18 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n+\t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n \tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {\n-\t\trcrec = cf->rcrecs[recs->minimal_perfect_hash];\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n+\t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n+\t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n \t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n@@ -312,8 +313,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t * false, or become true.\n \t */\n \txdf1->nreff = 0;\n-\tfor (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];\n-\t     i <= xdf1->dend; i++, recs++) {\n+\tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tif (action1[i] == KEEP ||\n \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n@@ -324,8 +324,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t}\n \n \txdf2->nreff = 0;\n-\tfor (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart];\n-\t     i <= xdf2->dend; i++, recs++) {\n+\tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tif (action2[i] == KEEP ||\n \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-- \ngitgitgadget\n\n"},{"id":"542484","messageId":"1822166fef0c5dcafb4f3c717eff235db6404342.1777500495.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","subject":"[PATCH v6 2/6] xdiff: use unambiguous types in xdl_bogo_sqrt()","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-29T22:08:11Z","receivedAt":"2026-04-29T22:08:20Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nThere is no real square root for a negative number and size_t may not\nbe large enough for certain applications, replace long with uint64_t.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xdiffi.c   | 2 +-\n xdiff/xprepare.c | 4 ++--\n xdiff/xutils.c   | 4 ++--\n xdiff/xutils.h   | 2 +-\n 4 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/xdiff/xdiffi.c b/xdiff/xdiffi.c\nindex 4376f943db..88708c12a3 100644\n--- a/xdiff/xdiffi.c\n+++ b/xdiff/xdiffi.c\n@@ -348,7 +348,7 @@ int xdl_do_diff(mmfile_t *mf1, mmfile_t *mf2, xpparam_t const *xpp,\n \tkvdf += xe->xdf2.nreff + 1;\n \tkvdb += xe->xdf2.nreff + 1;\n \n-\txenv.mxcost = xdl_bogosqrt(ndiags);\n+\txenv.mxcost = (long)xdl_bogosqrt((uint64_t)ndiags);\n \tif (xenv.mxcost < XDL_MAX_COST_MIN)\n \t\txenv.mxcost = XDL_MAX_COST_MIN;\n \txenv.snake_cnt = XDL_SNAKE_CNT;\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex d6e1901d2d..48fb5ce6fe 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -290,7 +290,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = xdl_bogosqrt((long)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n@@ -299,7 +299,7 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = xdl_bogosqrt((long)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n+\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n \t\tmlim = XDL_MAX_EQLIMIT;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 77ee1ad9c8..9a999acdc0 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -23,8 +23,8 @@\n #include \"xinclude.h\"\n \n \n-long xdl_bogosqrt(long n) {\n-\tlong i;\n+uint64_t xdl_bogosqrt(uint64_t n) {\n+\tuint64_t i;\n \n \t/*\n \t * Classical integer square root approximation using shifts.\ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex 615b4a9d35..58f9d74cda 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -25,7 +25,7 @@\n \n \n \n-long xdl_bogosqrt(long n);\n+uint64_t xdl_bogosqrt(uint64_t n);\n int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n \t\t     xdemitcb_t *ecb);\n int xdl_cha_init(chastore_t *cha, long isize, long icount);\n-- \ngitgitgadget\n\n"},{"id":"542485","messageId":"85aa0da90c62d9217ac2a3f907c37855a95298b9.1777500495.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","subject":"[PATCH v6 3/6] xdiff/xdl_cleanup_records: use unambiguous types","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-29T22:08:12Z","receivedAt":"2026-04-29T22:08:21Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nChange the parameters of xdl_clean_mmatch() and the local variables\ni, nm, mlim in xdl_cleanup_records() to use unambiguous types. Best\nviewed with --color-words.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 48fb5ce6fe..386668a92d 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -197,8 +197,8 @@ void xdl_free_env(xdfenv_t *xe) {\n }\n \n \n-static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n-\tlong r, rdis0, rpdis0, rdis1, rpdis1;\n+static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, ptrdiff_t e) {\n+\tptrdiff_t r, rdis0, rpdis0, rdis1, rpdis1;\n \n \t/*\n \t * Limits the window that is examined during the similar-lines\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, long i, long s, long e) {\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tlong i, nm, mlim;\n+\tptrdiff_t i, nm, mlim;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n-- \ngitgitgadget\n\n"},{"id":"542487","messageId":"51c62ed454cd66e884fdbbf3635603ca66966bc8.1777500495.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","subject":"[PATCH v6 4/6] xdiff/xdl_cleanup_records: make limits more clear","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-29T22:08:13Z","receivedAt":"2026-04-29T22:08:22Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nMake the handling of per-file limits and the minimal-case clearer.\n  * Use explicit per-file limit variables (mlim1, mlim2) and initialize\n    them.\n  * The additional condition `!need_min` is redudant now, remove it.\nBest viewed with --color-words.\n\nHelped-by: Phillip Wood\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 26 +++++++++++++++++++-------\n 1 file changed, 19 insertions(+), 7 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 386668a92d..7141dbc058 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -268,7 +268,7 @@ static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t i, ptrdiff_t s, pt\n  * might be potentially discarded if they appear in a run of discardable.\n  */\n static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {\n-\tptrdiff_t i, nm, mlim;\n+\tptrdiff_t i, nm, mlim1, mlim2;\n \txdlclass_t *rcrec;\n \tuint8_t *action1 = NULL, *action2 = NULL;\n \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n@@ -290,22 +290,34 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t/*\n \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n \t */\n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n+\tif (need_min) {\n+\t\t/* i.e. infinity */\n+\t\tmlim1 = PTRDIFF_MAX;\n+\t} else {\n+\t\tmlim1 = xdl_bogosqrt((uint64_t)xdf1->nrec);\n+\t\tif (mlim1 > XDL_MAX_EQLIMIT)\n+\t\t\tmlim1 = XDL_MAX_EQLIMIT;\n+\t}\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n \t}\n \n-\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n-\t\tmlim = XDL_MAX_EQLIMIT;\n+\tif (need_min) {\n+\t\t/* i.e. infinity */\n+\t\tmlim2 = PTRDIFF_MAX;\n+\t} else {\n+\t\tmlim2 = xdl_bogosqrt((uint64_t)xdf2->nrec);\n+\t\tif (mlim2 > XDL_MAX_EQLIMIT)\n+\t\t\tmlim2 = XDL_MAX_EQLIMIT;\n+\t}\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: (nm >= mlim && !need_min) ? INVESTIGATE: KEEP;\n+\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"542486","messageId":"45ad2ae62de99de598088fd041559ff3a23ef82c.1777500495.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","subject":"[PATCH v6 5/6] xdiff/xdl_cleanup_records: make setting action easier to follow","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-29T22:08:14Z","receivedAt":"2026-04-29T22:08:24Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nRewrite nested ternaries with a clear if/else ladder for\naction1/action2 to improve readability while preserving\nbehavior.\n\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 14 ++++++++++++--\n 1 file changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex 7141dbc058..ddd0577676 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -302,7 +302,12 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph1];\n \t\tnm = rcrec ? rcrec->len2 : 0;\n-\t\taction1[i] = (nm == 0) ? DISCARD: nm >= mlim1 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction1[i] = DISCARD;\n+\t\telse if (nm < mlim1)\n+\t\t\taction1[i] = KEEP;\n+\t\telse /* nm >= mlim1 */\n+\t\t\taction1[i] = INVESTIGATE;\n \t}\n \n \tif (need_min) {\n@@ -317,7 +322,12 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n \t\trcrec = cf->rcrecs[mph2];\n \t\tnm = rcrec ? rcrec->len1 : 0;\n-\t\taction2[i] = (nm == 0) ? DISCARD: nm >= mlim2 ? INVESTIGATE: KEEP;\n+\t\tif (nm == 0)\n+\t\t\taction2[i] = DISCARD;\n+\t\telse if (nm < mlim2)\n+\t\t\taction2[i] = KEEP;\n+\t\telse /* nm >= mlim2 */\n+\t\t\taction2[i] = INVESTIGATE;\n \t}\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"542488","messageId":"a5174802f453c3f26f950efcc5416ff961a6e4ab.1777500495.git.gitgitgadget@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","subject":"[PATCH v6 6/6] xdiff/xdl_cleanup_records: make execution of action easier to follow","fromName":"Ezekiel Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-29T22:08:15Z","receivedAt":"2026-04-29T22:08:25Z","isPatch":true,"body":"From: Ezekiel Newren <ezekielnewren@gmail.com>\n\nHelped-by: Phillip Wood\nSigned-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n---\n xdiff/xprepare.c | 40 ++++++++++++++++++++++++++++++----------\n 1 file changed, 30 insertions(+), 10 deletions(-)\n\ndiff --git a/xdiff/xprepare.c b/xdiff/xprepare.c\nindex ddd0577676..beef711067 100644\n--- a/xdiff/xprepare.c\n+++ b/xdiff/xprepare.c\n@@ -336,24 +336,44 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n \t */\n \txdf1->nreff = 0;\n \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n-\t\tif (action1[i] == KEEP ||\n-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n+\t\tuint8_t action = action1[i];\n+\n+\t\tif (action == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n+\t\t\t\taction = KEEP;\n+\t\t\telse\n+\t\t\t\taction = DISCARD;\n+\t\t}\n+\n+\t\tif (action == KEEP) {\n \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action == DISCARD) {\n \t\t\txdf1->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\t} else {\n+\t\t\tBUG(\"Illegal state for action\");\n+\t\t}\n \t}\n \n \txdf2->nreff = 0;\n \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n-\t\tif (action2[i] == KEEP ||\n-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n+\t\tuint8_t action = action2[i];\n+\n+\t\tif (action == INVESTIGATE) {\n+\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n+\t\t\t\taction = KEEP;\n+\t\t\telse\n+\t\t\t\taction = DISCARD;\n+\t\t}\n+\n+\t\tif (action == KEEP) {\n \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n-\t\t\t/* changed[i] remains false, i.e. keep */\n-\t\t} else\n+\t\t\t/* changed[i] remains false */\n+\t\t} else if (action == DISCARD) {\n \t\t\txdf2->changed[i] = true;\n-\t\t\t/* i.e. discard */\n+\t\t} else {\n+\t\t\tBUG(\"Illegal state for action\");\n+\t\t}\n \t}\n \n cleanup:\n-- \ngitgitgadget\n"},{"id":"542530","messageId":"c8b48c6a-5a20-4981-9cd4-999b40c618fc@gmail.com","threadId":"64712","inReplyTo":"pull.2156.v6.git.git.1777500495.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 0/6] Xdiff cleanup part 3","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-04-30T13:35:49Z","receivedAt":"2026-04-30T13:35:54Z","isPatch":true,"body":"Hi Ezekiel\n\nOn 29/04/2026 23:08, Ezekiel Newren via GitGitGadget wrote:\n> Changes in v6:\n> \n>   * implement suggestions by Phillip Wood [1,2]\n> \n> Phillip's second \"if\" in [1] differs from his first one. In my changes I\n> made both of them structurally the same.\n\nI was in two minds about whether to do that or not, all the changes here \nlook good to me.\n\nJuino - are you happy to rebase pw/xdiff-shrink-memory-consumption, or \ndo you want be to send a re-roll?\n\n> Something I'm confused by is the range-diff of patch 5. I'm confused why\n> range-diff states that this is different at all. I don't think this is a\n> problem, I just don't like not being able to explain a difference pointed\n> out by range-diff.\n\nThe context line below the insertion of \"action1[i] = INVESTIGATE;\" has \nchanged do to the changes to patch 4. It would be nice if there was a \nway to tell range-diff to ignore hunks where only the context lines have \nchanged but nobody has implemented that yet.\n\nThanks\n\nPhillip\n\n> 5: 88c68fa89a ! 5: 099b08c33f xdiff/xdl_cleanup_records: make setting action easier to follow\n > @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t \ncf, xdfile_t *\n >  +\t\t\taction1[i] = > INVESTIGATE;\n >   \t}>\n> - \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n> + \tif (need_min) {\n> +@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>              size_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>              rcrec = cf->rcrecs[mph2];\n>              nm = rcrec ? rcrec->len1 : 0;\n> \n> \n> [1] limits\n> https://lore.kernel.org/git/d88af7e1-e8dd-4423-9c6c-977e1f1dc074@gmail.com/\n> [2] action execution\n> https://lore.kernel.org/git/df244360-e9a9-44c0-946d-29288e6dd269@gmail.com/\n> \n> Changes in v5:\n> \n>   * drop commit \"xdiff/xdl_cleanup_records: simplify INVESTIGATE handling for\n>     clarity\".\n>   * add braces around the else clause\n> \n> I didn't see a better way to rewrite how action is used so I reverted to\n> what it used to be.\n> \n> Changes in v4:\n> \n>   * Change SIZE_MAX to PTRDIFF_MAX.\n> \n> Changes in v3:\n> \n>   * run make DEVELOPER=1 on each commit and fix all compiler issues\n> \n> v2 is a radical departure from v1 Changes in v2:\n> \n>   * make the flow of xdl_cleanup_records() easier to follow\n> \n> There is no performance or behavioral change introduced in this patch\n> series.\n> \n> === original cover letter bellow ===\n> \n> Patch series summary:\n> \n>   * patch 1: Introduce the ivec type\n>   * patch 2: Create the function xdl_do_classic_diff()\n>   * patches 3-4: generic cleanup\n>   * patches 5-8: convert from dstart/dend (in xdfile_t) to\n>     delta_start/delta_end (in xdfenv_t)\n>   * patches 9-10: move xdl_cleanup_records(), and related, from xprepare.c to\n>     xdiffi.c\n> \n> Things that will be addressed in future patch series:\n> \n>   * Make xdl_cleanup_records() easier to read\n>   * convert recs/nrec into an ivec\n>   * convert changed to an ivec\n>   * remove reference_index/nreff from xdfile_t and turn it into an ivec\n>   * splitting minimal_perfect_hash out as its own ivec\n>   * improve the performance of the classifier and parsing/hashing lines\n> \n> === before this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; ptrdiff_t dstart, dend; bool *changed; size_t *reference_index;\n> size_t nreff; } xdfile_t;\n> \n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; } xdfenv_t;\n> \n> === after this patch series typedef struct s_xdfile { xrecord_t *recs;\n> size_t nrec; bool *changed; size_t *reference_index; size_t nreff; }\n> xdfile_t;\n> \n> typedef struct s_xdfenv { xdfile_t xdf1, xdf2; size_t delta_start,\n> delta_end; size_t mph_size; } xdfenv_t;\n> \n> Ezekiel Newren (6):\n>    xdiff/xdl_cleanup_records: delete local recs pointer\n>    xdiff: use unambiguous types in xdl_bogo_sqrt()\n>    xdiff/xdl_cleanup_records: use unambiguous types\n>    xdiff/xdl_cleanup_records: make limits more clear\n>    xdiff/xdl_cleanup_records: make setting action easier to follow\n>    xdiff/xdl_cleanup_records: make execution of action easier to follow\n> \n>   xdiff/xdiffi.c   |  2 +-\n>   xdiff/xprepare.c | 97 ++++++++++++++++++++++++++++++++++--------------\n>   xdiff/xutils.c   |  4 +-\n>   xdiff/xutils.h   |  2 +-\n>   4 files changed, 73 insertions(+), 32 deletions(-)\n> \n> \n> base-commit: ca1db8a0f7dc0dbea892e99f5b37c5fe5861be71\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2156%2Fezekielnewren%2Fxdiff-cleanup-3-v6\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2156/ezekielnewren/xdiff-cleanup-3-v6\n> Pull-Request: https://github.com/git/git/pull/2156\n> \n> Range-diff vs v5:\n> \n>   1:  b31924a949 = 1:  b31924a949 xdiff/xdl_cleanup_records: delete local recs pointer\n>   2:  1822166fef = 2:  1822166fef xdiff: use unambiguous types in xdl_bogo_sqrt()\n>   3:  85aa0da90c = 3:  85aa0da90c xdiff/xdl_cleanup_records: use unambiguous types\n>   4:  fec2b0f38a ! 4:  51c62ed454 xdiff/xdl_cleanup_records: make limits more clear\n>       @@ Commit message\n>              * The additional condition `!need_min` is redudant now, remove it.\n>            Best viewed with --color-words.\n>        \n>       +    Helped-by: Phillip Wood\n>            Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>        \n>         ## xdiff/xprepare.c ##\n>       @@ xdiff/xprepare.c: static bool xdl_clean_mmatch(uint8_t const *action, ptrdiff_t\n>         \tuint8_t *action1 = NULL, *action2 = NULL;\n>         \tbool need_min = !!(cf->flags & XDF_NEED_MINIMAL);\n>        @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>       - \t\tgoto cleanup;\n>       - \t}\n>       -\n>       -+\tif (need_min) {\n>       -+\t\t/* i.e. infinity */\n>       -+\t\tmlim1 = PTRDIFF_MAX;\n>       -+\t\tmlim2 = PTRDIFF_MAX;\n>       -+\t} else {\n>       -+\t\tmlim1 = XDL_MIN(xdl_bogosqrt(xdf1->nrec), XDL_MAX_EQLIMIT);\n>       -+\t\tmlim2 = XDL_MIN(xdl_bogosqrt(xdf2->nrec), XDL_MAX_EQLIMIT);\n>       -+\t}\n>       -+\n>         \t/*\n>         \t * Initialize temporary arrays with DISCARD, KEEP, or INVESTIGATE.\n>         \t */\n>        -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf1->nrec)) > XDL_MAX_EQLIMIT)\n>        -\t\tmlim = XDL_MAX_EQLIMIT;\n>       ++\tif (need_min) {\n>       ++\t\t/* i.e. infinity */\n>       ++\t\tmlim1 = PTRDIFF_MAX;\n>       ++\t} else {\n>       ++\t\tmlim1 = xdl_bogosqrt((uint64_t)xdf1->nrec);\n>       ++\t\tif (mlim1 > XDL_MAX_EQLIMIT)\n>       ++\t\t\tmlim1 = XDL_MAX_EQLIMIT;\n>       ++\t}\n>         \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>         \t\tsize_t mph1 = xdf1->recs[i].minimal_perfect_hash;\n>         \t\trcrec = cf->rcrecs[mph1];\n>       @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n>         \n>        -\tif ((mlim = (long)xdl_bogosqrt((uint64_t)xdf2->nrec)) > XDL_MAX_EQLIMIT)\n>        -\t\tmlim = XDL_MAX_EQLIMIT;\n>       ++\tif (need_min) {\n>       ++\t\t/* i.e. infinity */\n>       ++\t\tmlim2 = PTRDIFF_MAX;\n>       ++\t} else {\n>       ++\t\tmlim2 = xdl_bogosqrt((uint64_t)xdf2->nrec);\n>       ++\t\tif (mlim2 > XDL_MAX_EQLIMIT)\n>       ++\t\t\tmlim2 = XDL_MAX_EQLIMIT;\n>       ++\t}\n>         \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>         \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>         \t\trcrec = cf->rcrecs[mph2];\n>   5:  88c68fa89a ! 5:  45ad2ae62d xdiff/xdl_cleanup_records: make setting action easier to follow\n>       @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *\n>        +\t\t\taction1[i] = INVESTIGATE;\n>         \t}\n>         \n>       - \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>       + \tif (need_min) {\n>       +@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>         \t\tsize_t mph2 = xdf2->recs[i].minimal_perfect_hash;\n>         \t\trcrec = cf->rcrecs[mph2];\n>         \t\tnm = rcrec ? rcrec->len1 : 0;\n>   6:  699e198fa9 ! 6:  a5174802f4 xdiff/xdl_cleanup_records: put braces around the else clause\n>       @@ Metadata\n>        Author: Ezekiel Newren <ezekielnewren@gmail.com>\n>        \n>         ## Commit message ##\n>       -    xdiff/xdl_cleanup_records: put braces around the else clause\n>       +    xdiff/xdl_cleanup_records: make execution of action easier to follow\n>        \n>       +    Helped-by: Phillip Wood\n>            Signed-off-by: Ezekiel Newren <ezekielnewren@gmail.com>\n>        \n>         ## xdiff/xprepare.c ##\n>        @@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>       - \t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n>       + \t */\n>       + \txdf1->nreff = 0;\n>       + \tfor (i = xdf1->dstart; i <= xdf1->dend; i++) {\n>       +-\t\tif (action1[i] == KEEP ||\n>       +-\t\t    (action1[i] == INVESTIGATE && !xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))) {\n>       ++\t\tuint8_t action = action1[i];\n>       ++\n>       ++\t\tif (action == INVESTIGATE) {\n>       ++\t\t\tif (!xdl_clean_mmatch(action1, i, xdf1->dstart, xdf1->dend))\n>       ++\t\t\t\taction = KEEP;\n>       ++\t\t\telse\n>       ++\t\t\t\taction = DISCARD;\n>       ++\t\t}\n>       ++\n>       ++\t\tif (action == KEEP) {\n>         \t\t\txdf1->reference_index[xdf1->nreff++] = i;\n>       - \t\t\t/* changed[i] remains false, i.e. keep */\n>       +-\t\t\t/* changed[i] remains false, i.e. keep */\n>        -\t\t} else\n>       -+\t\t} else {\n>       ++\t\t\t/* changed[i] remains false */\n>       ++\t\t} else if (action == DISCARD) {\n>         \t\t\txdf1->changed[i] = true;\n>       - \t\t\t/* i.e. discard */\n>       +-\t\t\t/* i.e. discard */\n>       ++\t\t} else {\n>       ++\t\t\tBUG(\"Illegal state for action\");\n>        +\t\t}\n>         \t}\n>         \n>         \txdf2->nreff = 0;\n>       -@@ xdiff/xprepare.c: static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd\n>       - \t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n>       + \tfor (i = xdf2->dstart; i <= xdf2->dend; i++) {\n>       +-\t\tif (action2[i] == KEEP ||\n>       +-\t\t    (action2[i] == INVESTIGATE && !xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))) {\n>       ++\t\tuint8_t action = action2[i];\n>       ++\n>       ++\t\tif (action == INVESTIGATE) {\n>       ++\t\t\tif (!xdl_clean_mmatch(action2, i, xdf2->dstart, xdf2->dend))\n>       ++\t\t\t\taction = KEEP;\n>       ++\t\t\telse\n>       ++\t\t\t\taction = DISCARD;\n>       ++\t\t}\n>       ++\n>       ++\t\tif (action == KEEP) {\n>         \t\t\txdf2->reference_index[xdf2->nreff++] = i;\n>       - \t\t\t/* changed[i] remains false, i.e. keep */\n>       +-\t\t\t/* changed[i] remains false, i.e. keep */\n>        -\t\t} else\n>       -+\t\t} else {\n>       ++\t\t\t/* changed[i] remains false */\n>       ++\t\t} else if (action == DISCARD) {\n>         \t\t\txdf2->changed[i] = true;\n>       - \t\t\t/* i.e. discard */\n>       +-\t\t\t/* i.e. discard */\n>       ++\t\t} else {\n>       ++\t\t\tBUG(\"Illegal state for action\");\n>        +\t\t}\n>         \t}\n>         \n> \n\n"},{"id":"542538","messageId":"CAH=ZcbBqtE4AYhPbrVPoBEVk3-f+pdpzqxooYawrGJrcJuSarQ@mail.gmail.com","threadId":"64712","inReplyTo":"c8b48c6a-5a20-4981-9cd4-999b40c618fc@gmail.com","subject":"Re: [PATCH v6 0/6] Xdiff cleanup part 3","fromName":"Ezekiel Newren","fromEmail":"ezekielnewren@gmail.com","sentAt":"2026-04-30T21:08:35Z","receivedAt":"2026-04-30T21:08:48Z","isPatch":true,"body":"On Thu, Apr 30, 2026 at 7:35 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ezekiel\n>\n> On 29/04/2026 23:08, Ezekiel Newren via GitGitGadget wrote:\n> > Changes in v6:\n> >\n> >   * implement suggestions by Phillip Wood [1,2]\n> >\n> > Phillip's second \"if\" in [1] differs from his first one. In my changes I\n> > made both of them structurally the same.\n>\n> I was in two minds about whether to do that or not, all the changes here\n> look good to me.\n\nHopefully this is the last revision. These changes took way longer\nthan I expected to get through the review process. What do you think\nJunio?\n"},{"id":"542639","messageId":"xmqq8qa0q9a4.fsf@gitster.g","threadId":"64712","inReplyTo":"c8b48c6a-5a20-4981-9cd4-999b40c618fc@gmail.com","subject":"Re: [PATCH v6 0/6] Xdiff cleanup part 3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-04T00:59:47Z","receivedAt":"2026-05-04T00:59:49Z","isPatch":true,"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> Juino - are you happy to rebase pw/xdiff-shrink-memory-consumption, or \n> do you want be to send a re-roll?\n\nI'd rather not risk botched rebase by a third-party and prefer to\ntake a fresh submission by the original author.  Thanks.\n"}]}