Re: [PATCH v2 4/9] cache-tree: avoid strtol() on non-string buffer
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Nov 23, 2025, 18:06 UTC
- Message-ID
- <xmqqh5ukzkqt.fsf@gitster.g>
- In-Reply-To
- <633f4d92-c258-45a8-9d32-116c94838e68@gmail.com>
Phillip Wood <phillip.wood123@gmail.com> writes:
Show 14 quoted lines
> All we need to do to accept a single minus sign is s/while/if/ > ... > If we limit ourselves to accepting a single minus sign then this can become > if (s == *ptr + (sign == -1)) > > so we need very little in the way of extra code. > ... > A generic helper to replace strtol() that takes a length rather than > assuming the input is NUL terminated could be useful elsewhere but I'm > not sure we need something that complicated here. I do like the fact > that overflow does not cause undefined behavior though. Changing ret for > "int" to "unsigned" in peff's patch should fix that. > > Thanks
Perhaps.
By the way, an interesting tangent is this.
The only reason why these fields under discussion are stored in textual decimal is pretty much the same as the reason why the object header expresses the byte-length of the payload in textual decimal, i.e., to be independent from the platform natural implementation of "int" type (e.g., endiannness and width), but unlike object files, the index is a local matter (we are prepared for the same directory accessed over NFS from two platforms with different endianness, but we do not recommend network access to a repository in the first place). And a lot more importantly, the total number of the index entries contained within an index file is capped to 2^32-1 (the header has 32-bit count in the network byte order). The total number of subdirectories within a directory or the total number of entries for a level of directory hierarchy that would form a tree object from a slice of the index cannot exceed that number anyway.
And thanks to the design that made cache-tree an optional index extension, we can make cache-tree version 2 where the in-core representation is exactly the same as the current one, but only uses different serialization when writing to and reading from the index file. The new serialization can use 32-bit network byte order integers, or use our own varint.{c,h,rs}, to record these numbers.
A version of Git that knows about that extension could be taught to read from the current cache-tree and convert to a new version, but better yet, it can simply ignore the current cache-tree data in the file, and write the new version when we do need to write the index out with a cache-tree. When such a transparent auto conversion happens, one single invocation of write_index_as_tree() would become more expensive than usual (because the last invocation of the current Git left cache-tree data in the index and usually the next invocation of Git would take advantage of it when it writes a tree, but a new version of Git that uses the v2 format would behave as if there is no cache-tree data in the index and build the tree from scratch. After that happens, the cache-tree data in the new format will be reused and things will continue to work. You could use an older version of Git on such an index file and the same transparent auto conversion will take care of the transition.