git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] blame.c: don't drop origin blobs as eagerly

From
Jeff King <peff@peff.net>
Date
Apr 3, 2019, 11:36 UTC
Message-ID
<20190403113604.GA2941@sigill.intra.peff.net>
In-Reply-To
<CACsJy8AbkmJ69ucCfGMdXHGvfko89SxH=DKjra6Ltwf7wpy-Og@mail.gmail.com>
On Wed, Apr 03, 2019 at 04:32:30PM +0700, Duy Nguyen wrote:
> That might explain why I could not see significant gain when blaming
> linux.git's MAINTAINERS file (0.5s was shaved out of 13s) even though
> the number of objects read was cut by half (8424 vs 15083).

I did a few timings, too, and managed to come up with similar improvements (only a small fraction, and only for large files). I think the main thing is simply that loading the blob from the object database is a fraction of the total work done. We still have to actually diff the blobs, which is at least as expensive as loading them from disk.

We also have to load commits and trees from disk as we traverse. Enabling the commit-graph would shrink that portion (and make improvements in the blob loading proportionally more impressive).

All that said, this seems like an easy and obvious win, and worth doing. 0.5s is still something.

I suspect we could do even better by storing and reusing not just the original blob between diffs, but the intermediate diff state (i.e., the hashes produced by xdl_prepare(), which should be usable between multiple diffs). That's quite a bit more complex, though, and I imagine would require some surgery to xdiff.

-Peff
Previous: Duy NguyenNext: Duy Nguyen
Message 4 of 8 in “blame.c: don't drop origin blobs as eagerly”
  1. blame.c: don't drop origin blobs as eagerlyDavid Kastrup, Apr 2, 2019
  2. Junio C HamanoApr 3, 2019
  3. Duy NguyenApr 3, 2019
  4. Jeff KingApr 3, 2019
  5. Duy NguyenApr 3, 2019
  6. Jeff KingApr 3, 2019
  7. David KastrupApr 3, 2019
  8. David KastrupApr 3, 2019

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.