Re: [PATCH 8/9] fsck: avoid parse_timestamp() on buffer that isn't NUL-terminated
- From
Jeff King <peff@peff.net>
- Date
- Nov 15, 2025, 02:12 UTC
- Message-ID
- <20251115021248.GB3499607@coredump.intra.peff.net>
- In-Reply-To
- <aRRux2uBfORc214r@pks.im>
On Wed, Nov 12, 2025 at 12:25:59PM +0100, Patrick Steinhardt wrote:
Show 23 quoted lines
> On Wed, Nov 12, 2025 at 03:10:40AM -0500, Jeff King wrote: > > In fsck_ident(), we parse the timestamp with parse_timestamp(), which is > > really an alias for strtoumax(). But since our buffer may not be > > NUL-terminated, this can trigger a complaint from ASan's > > strict_string_checks mode. This is a false positive, since we know that > > the buffer contains a trailing newline (which we checked earlier in the > > function), and that strtoumax() would stop there. > > > > But it is worth working around ASan's complaint. One is because that > > will let us turn on strict_string_checks by default, which has helped > > catch other real problems. And two is that the safety of the current > > code is very hard to reason about (it subtly depends on distant code > > which could change). > > > > One option here is to just parse the number left-to-right ourselves. But > > we care about the size of a timestamp_t and detecting overflow, since > > that's part of the point of these checks. And doing that correctly is > > tricky. So we'll instead just pull the digits into a separate, > > NUL-terminated buffer, and use that to call parse_timestamp(). > > So this is another site that would benefit from having something like > `git_parse_int()` with an extra parameter indicating the number of > bytes available for parsing (and a way to disable unit factors).
Yes, but also no.
Yes, in the sense that if we had a robust global function to parse an integer from a buf we could use it here, as well as in cache-tree.
But there are lots of no's:
- We could not have one such function, because the implementation
would differ based on signedness and size of the integer type. And
cache-tree is a signed long, whereas this is a uintmax_t. We can
factor out some of the work with a helper that takes a max
parameter, but you can see we duplicate a bunch of code between
git_parse_signed() and git_parse_unsigned(). - The interface for git_parse_int() isn't quite a match. It wants to
parse every byte in the provided string and complains if there is
any extra cruft, rather than aiding in progressive parsing of a
buffer. So if you have a string "10 20\0", it will not just parse
"10" and then tell you how far it parsed; it will barf on the space.
If we added a new parameter for "this is how many bytes we have",
and you fed it the buffer "10 20" and the length 5, it would have
the same problem. So there's a fundamental mismatch between "parse this as an integer
and complain if there is anything else" versus "parse an integer,
advance the pointer, and we'll keep going". - The implementation for git_parse_int() isn't a match either. It is
just asking strtoimax() to do all of the work, which is the very
thing we need to avoid.So I think one _could_ write a strtoimax() replacement that handled everything we wanted, and then you could probably build git_parse_signed() etc around that. But it would be a lot more work than what's there (like checking overflow progressively as we multiply and add), and there are some decisions to be made (like handling leading whitespace or how +/- work on unsigned integers). I'd probably err more on the side of simplicity and strictness than strtol() does, but that also means that plugging in the new function would change user-visible behavior. Sometimes in a good or OK way (stricter parsing), but maybe sometimes in a bad or confusing one (rejecting inputs that used to work).
-Peff