Re: [PATCH v3 06/10] xdiff: split xrecord_t.ha into line_hash and minimal_perfect_hash
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Nov 11, 2025, 23:21 UTC
- Message-ID
- <xmqqwm3wtat8.fsf@gitster.g>
- In-Reply-To
- <3834ea8f9becc9d6e1b407679e8a95dc6c9d56de.1762890152.git.gitgitgadget@gmail.com>
"Ezekiel Newren via GitGitGadget" <gitgitgadget@gmail.com> writes:
Show 10 quoted lines
> To make this clearer, the old ha field has been split: > * line_hash: a straightforward hash of a line, independent of any > external context. Its type is uint64_t, as it comes from a fixed > width hash function. > * minimal_perfect_hash: Not a new concept, but now a separate > field. It comes from the classifier's general-purpose hash table, > which assigns each line a unique and minimal hash across the two > files. A size_t is used here because it's meant to be used to > index an array. This also this avoids ` as usize` casts on the Rust > side when using it to index a slice.
How much extra memory pressure does this change cause? In a single instance of xrecord_t, we used to have a single ulong plus a pointer and a size_t; now we replaced the single ulong with two 8-byte words, so 33% more memory per record, which is not so huge a deal?
Show 7 quoted lines
> static int xdl_classify_record(unsigned int pass, xdlclassifier_t *cf, xrecord_t *rec) {
> - long hi;
> + size_t hi;
> xdlclass_t *rcrec;
>
> - hi = (long) XDL_HASHLONG(rec->ha, cf->hbits);
> + hi = XDL_HASHLONG(rec->line_hash, cf->hbits);Very nice that we can lose these random-looking casts.
Show 14 quoted lines
> diff --git a/xdiff/xtypes.h b/xdiff/xtypes.h
> index 88b1fe4649..742b81bf3b 100644
> --- a/xdiff/xtypes.h
> +++ b/xdiff/xtypes.h
> @@ -41,7 +41,8 @@ typedef struct s_chastore {
> typedef struct s_xrecord {
> uint8_t const *ptr;
> size_t size;
> - unsigned long ha;
> + uint64_t line_hash;
> + size_t minimal_perfect_hash;
> } xrecord_t;
>
> typedef struct s_xdfile {