Re: [PATCH v2 1/2] blame: harden ignore-revs parser and tag peeling
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Oct 9, 2026, 05:07 UTC
- Message-ID
- <xmqqjynrwjg0.fsf@gitster.g>
- In-Reply-To
- <2e12486c0d5dd8b94393b08413a85a8d47f86edd.1791493644.git.gitgitgadget@gmail.com>
"Ravi Mistry via GitGitGadget" <gitgitgadget@gmail.com> writes:
Show 9 quoted lines
> - In oidset_parse_file_carefully(), strbuf_getline() reads up to the > next newline and records the full line length in sb.len, including > any embedded NUL bytes. However, strchr(sb.buf, '#') and > parse_oid_hex_algop(sb.buf, &oid, &p, algop) treat sb.buf as a > NUL-terminated string. If a line contains an embedded NUL byte after > a valid object name (such as "<oid>\0garbage" or "<oid>\0# comment"), > *p is '\0' and trailing bytes on the line are silently ignored. > Reject any line containing an embedded NUL byte via memchr() before > stripping comments and whitespace.
Maybe I am slow, but I do not immediately see why ignoring everything after the first NUL is a problem. A call to strbuf_trim() is ineffective at trimming whitespace that appears immediately before such a NUL. For example, while
cf9bdb1...f6092d # comment LF
would feed the leading 'cf9bdb1...f6092d' part (after stripping whitespace before '#') to parse_oid_hex_algop(), this
cf9bdb1...f6092d NUL comment LF
would keep the whitespace after '92d' and cause the parsing to fail. I do not see any security implications here.
On the other hand ...
Show 10 quoted lines
> - In peel_to_commit_oid(), odb_read_object_info() is called without > OBJECT_INFO_SKIP_FETCH_OBJECT or OBJECT_INFO_QUICK, and deref_tag() > calls parse_object() on tag targets without checking whether the > target object exists locally first. In a partial clone, any missing > commit OID or tag target listed in the ignore-revs file would trigger > lazy promisor fetches and pack directory rescans during git-blame(1). > Use odb_read_object_info_extended() with OBJECT_INFO_LOOKUP_REPLACE | > OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK and peel OBJ_TAG > objects one layer per iteration, verifying that each target object > exists locally and matches the tag's declared type before parsing it.
... this may be a very reasonable thing to do, I would think. In a shallow clone, if we are not auto-deepening the shallow boundary during a "git blame" session, we have no reason to lazy fetch entries in the ignore file that are older than the shallow boundary.
Show 10 quoted lines
> diff --git a/oidset.c b/oidset.c
> index c8ff0b385c..90d39204d3 100644
> --- a/oidset.c
> +++ b/oidset.c
> @@ -85,6 +85,9 @@ void oidset_parse_file_carefully(struct oidset *set, const char *path,
> const char *p;
> const char *name;
>
> + if (memchr(sb.buf, '\0', sb.len))
> + die("invalid object name: %s", sb.buf);A file with such an entry is rejected and the entire operation is aborted as suspected attack attempt, which feels like striking the balance between usability and security at a wrong place.
But a line with broken object name already is rejected with "die()" with the existing code, so it may be OK.
Show 47 quoted lines
> diff --git a/t/t8013-blame-ignore-revs.sh b/t/t8013-blame-ignore-revs.sh > index cace00ae8d..70fe509a64 100755 > --- a/t/t8013-blame-ignore-revs.sh > +++ b/t/t8013-blame-ignore-revs.sh > @@ -327,4 +327,42 @@ test_expect_success ignore_merge ' > test_cmp expect actual > ' > > +test_expect_success 'ignore-revs-file rejects lines with embedded NUL bytes' ' > + rev_b=$(git rev-parse B) && > + printf "%sQgarbage\n" "$rev_b" | q_to_nul >ignore_nul && > + test_must_fail git blame file --ignore-revs-file ignore_nul 2>err && > + test_grep "invalid object name:" err && > + > + printf "%sQ# comment\n" "$rev_b" | q_to_nul >ignore_nul_comment && > + test_must_fail git blame file --ignore-revs-file ignore_nul_comment 2>err && > + test_grep "invalid object name:" err > +' > + > +test_expect_success 'ignore-revs-file peels chained tags and skips missing tag targets' ' > + test_write_lines BB L2-modified L3 L4 L5 L6 L7 L8 CC >file && > + git add file && > + test_tick && > + git commit -m D && > + git tag -a -m "tag 1" D_TAG1 HEAD && > + git tag -a -m "tag 2" D_TAG2 D_TAG1 && > + git rev-parse D_TAG2 >ignore_tag_chain && > + git blame --line-porcelain file --ignore-revs-file ignore_tag_chain >blame_raw && > + sed -ne "/^[0-9a-f][0-9a-f]* [0-9][0-9]* 2/s/ .*//p" blame_raw >actual && > + git rev-parse A >expect && > + test_cmp expect actual && > + > + test_config extensions.partialClone origin && > + test_config remote.origin.promisor true && > + test_config remote.origin.url /nonexistent && > + missing_oid=$(test_oid deadbeef) && > + bad_tag=$(printf "object %s\ntype commit\ntag bad-tag\ntagger T <t@example.com> 0 +0000\n\nmsg\n" "$missing_oid" | > + git hash-object -t tag -w --stdin) && > + test_write_lines "$missing_oid" "$bad_tag" >ignore_bad_tag && > + git blame --line-porcelain file --ignore-revs-file ignore_bad_tag >blame_raw 2>err && > + test_must_be_empty err && > + sed -ne "/^[0-9a-f][0-9a-f]* [0-9][0-9]* 2/s/ .*//p" blame_raw >actual && > + git rev-parse HEAD >expect && > + test_cmp expect actual > +' > + > test_done