Re: [PATCH 0/5] [RFC] diff: add diff.<driver>.process for external hunk providers
- From
Junio C Hamano <gitster@pobox.com>
- Date
- May 22, 2026, 05:29 UTC
- Message-ID
- <xmqq8q9cui5c.fsf@gitster.g>
- In-Reply-To
- <pull.2120.git.1779415884.gitgitgadget@gmail.com>
"Michael Montalbo via GitGitGadget" <gitgitgadget@gmail.com> writes:
Show 29 quoted lines
> This series adds diff.<driver>.process, a long-running subprocess protocol > that lets external tools provide hunks to git's diff and blame pipelines. > > Over the past 18 years, git's diff pipeline accumulated many features that > operate on hunks: word diff, function context, color-moved, indent > heuristic, blame. External tools can replace the pipeline entirely > (diff.<driver>.command) or select among builtin algorithms > (diff.<driver>.algorithm), but there is no way for a tool to provide > line-change information into the pipeline. Tools that understand code > structure (tree-sitter parsers, format-aware analyzers, tools like > Difftastic and Mergiraf) must bypass git's pipeline and lose access to > everything downstream. > > The protocol follows filter.<driver>.process: pkt-line over stdin/stdout, > capability negotiation, one tool invocation per git command. The tool > receives file pairs and returns hunk descriptors that git feeds into the > standard xdiff pipeline. All output features work normally. > > Zero hunks with status=success means the tool considers the files > equivalent. git diff shows no output for the file, and git blame skips the > commit, attributing lines to earlier commits. > > On error or tool crash, git falls back silently to the builtin diff > algorithm. The feature is opt-in via diff.<driver>.process and > .gitattributes; unconfigured files are unaffected. > > The series includes git diff-process-normalize, a built-in tool that > compares files line by line ignoring whitespace (same logic as "git diff -w" > via xdiff_compare_lines):
Interesting.
If the goal is purely to normalize content before comparison (e.g. stripping comments or canonicalizing formatting), we already have the `textconv` mechanism. While `textconv` is a "one-shot" per-file process, it is significantly simpler.
I suspect, however, that the primary focus here is to allow external tools to provide structural alignment (e.g. for AST- aware diffs like Difftastic or Mergiraf) without losing the original content in the display. Unlike `textconv`, which transforms the text the user sees, this protocol lets the display remain identical to the source while using a custom engine for the line-matching logic.
If that is the intent, it should be stated more explicitly in the documentation and commit messages. The "whitespace-normalize" demonstration in [PATCH 5/5] is misleading because it's exactly the case where `textconv` would be sufficient.
I am afraid that the use of a long-running subprocess for every diff/blame invocation adds significant complexity and overhead. In particular, wouldn't the `blame` implementation performs a round-trip to the subprocess for every commit in the history? Even with a persistent process, the overhead of serializing and deserializing the entire file content twice (old and new) for every commit could be prohibitive for large files or deep histories.
So, I dunno.