git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Bug: git log --numstat counts wrong

From
Tay Ray Chuan <rctay89@gmail.com>
Date
Sep 23, 2011, 06:30 UTC
Message-ID
<CALUzUxprUFGMR-WVEMOOvYiwkev1cfxHOyBmZq9bKJcHq5E2VA@mail.gmail.com>
In-Reply-To
<4E7B5F28.2020204@lsrfire.ath.cx>

On Fri, Sep 23, 2011 at 12:15 AM, René Scharfe <rene.scharfe@lsrfire.ath.cx> wrote:

Show 46 quoted lines
> The patch below reverts a part of 27af01d5523 that's not explained in its
> commit message and doesn't seem to contribute to the intended speedup.  It
> seems to restore the original diff output.  I don't know how it's actually
> doing that, though, as I haven't dug into the code at all.
>
> [snip]
>
> diff --git a/xdiff/xprepare.c b/xdiff/xprepare.c
> index 5a33d1a..e419f4f 100644
> --- a/xdiff/xprepare.c
> +++ b/xdiff/xprepare.c
> @@ -383,7 +383,7 @@ static int xdl_clean_mmatch(char const *dis, long i, long s, long e) {
>  * might be potentially discarded if they happear in a run of discardable.
>  */
>  static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xdf2) {
> -       long i, nm, nreff;
> +       long i, nm, nreff, mlim;
>        xrecord_t **recs;
>        xdlclass_t *rcrec;
>        char *dis, *dis1, *dis2;
> @@ -396,16 +396,20 @@ static int xdl_cleanup_records(xdlclassifier_t *cf, xdfile_t *xdf1, xdfile_t *xd
>        dis1 = dis;
>        dis2 = dis1 + xdf1->nrec + 1;
>
> +       if ((mlim = xdl_bogosqrt(xdf1->nrec)) > XDL_MAX_EQLIMIT)
> +               mlim = XDL_MAX_EQLIMIT;
>        for (i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart]; i <= xdf1->dend; i++, recs++) {
>                rcrec = cf->rcrecs[(*recs)->ha];
>                nm = rcrec ? rcrec->len2 : 0;
> -               dis1[i] = (nm == 0) ? 0: 1;
> +               dis1[i] = (nm == 0) ? 0: (nm >= mlim) ? 2: 1;
>        }
>
> +       if ((mlim = xdl_bogosqrt(xdf2->nrec)) > XDL_MAX_EQLIMIT)
> +               mlim = XDL_MAX_EQLIMIT;
>        for (i = xdf2->dstart, recs = &xdf2->recs[xdf2->dstart]; i <= xdf2->dend; i++, recs++) {
>                rcrec = cf->rcrecs[(*recs)->ha];
>                nm = rcrec ? rcrec->len1 : 0;
> -               dis2[i] = (nm == 0) ? 0: 1;
> +               dis2[i] = (nm == 0) ? 0: (nm >= mlim) ? 2: 1;
>        }
>
>        for (nreff = 0, i = xdf1->dstart, recs = &xdf1->recs[xdf1->dstart];
>
>
>
Thanks for the patch, René.
Sorry for not explaining that part of the change.

My understanding of mlim is that it "caps" how deep the for loop at around line 387 goes through a hash bucket/record chaing to find a matching record from side A in side B (and vice-versa in a later loop), probably to prevent running time from becoming too long.

But with 27af01d, this is no longer a concern. We can get an *exact*, pre-computed count of matching records in the other side, so we don't have go through the hash bucket. Thus mlim is no longer needed.

So re-introducing mlim doesn't seem right, even though it may fix this "bug" (ie restore the old behaviour).

-- 
Cheers,
Ray Chuan
Previous: René ScharfeNext: Tay Ray Chuan
Message 9 of 16 in “Bug: git log --numstat counts wrong”
  1. Alexander PepperSep 21, 2011
  2. Junio C HamanoSep 21, 2011
  3. Alexander PepperSep 21, 2011
  4. Alexander PepperSep 21, 2011
  5. Junio C HamanoSep 21, 2011
  6. Alexander PepperSep 22, 2011
  7. Junio C HamanoSep 22, 2011
  8. René ScharfeSep 22, 2011
  9. Tay Ray ChuanSep 23, 2011
  10. Revert removal of multi-match discard heuristic in 27af01Tay Ray Chuan, Sep 25, 2011
  11. René ScharfeSep 25, 2011
  12. Junio C HamanoSep 22, 2011
  13. Tay Ray ChuanSep 23, 2011
  14. Tay Ray ChuanSep 23, 2011
  15. Junio C HamanoSep 23, 2011
  16. Alexander PepperSep 23, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.