git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 1/5] avoid parse_sha1_header() accessing memory out of bound

From
LYLiu Yubao <yubao.liu@gmail.com>
Date
Dec 3, 2008, 03:49 UTC
Message-ID
<493601DE.3050107@gmail.com>
In-Reply-To
<20081202154257.GK23984@spearce.org>
Shawn O. Pearce wrote:
Show 30 quoted lines
> Liu Yubao <yubao.liu@gmail.com> wrote:
>> diff --git a/sha1_file.c b/sha1_file.c
>> index 6c0e251..efe6967 100644
>> --- a/sha1_file.c
>> +++ b/sha1_file.c
>> @@ -1254,10 +1255,10 @@ static int parse_sha1_header(const char *hdr, unsigned long *sizep)
>>  	/*
>>  	 * The type can be at most ten bytes (including the
>>  	 * terminating '\0' that we add), and is followed by
>> -	 * a space.
>> +	 * a space, at least one byte for size, and a '\0'.
>>  	 */
>>  	i = 0;
>> -	for (;;) {
>> +	while (hdr < hdr_end - 2) {
>>  		char c = *hdr++;
>>  		if (c == ' ')
>>  			break;
>> @@ -1265,6 +1266,8 @@ static int parse_sha1_header(const char *hdr, unsigned long *sizep)
>>  		if (i >= sizeof(type))
>>  			return -1;
> 
> That first hunk I am citing is unnecessary, because of the lines
> right above.  All of the callers of this function pass in a buffer
> that is at least 32 bytes in size; this loop aborts if it does not
> find a ' ' within the first 10 bytes of the buffer.  We'll never
> access memory outside of the buffer during this loop.
> 
> So IMHO your first three hunks here aren't necessary.
> 

Seems you missed the cover letter sent as patch 0/5, all patches are explained in the cover letter, sorry I sent them as separate topics by mistake.

This bound check is mainly for uncompressed loose object, a loose object that just are uncompressed:

uncompressed loose object = inflate(loose object) loose object = deflate(typename + <space> + size + '\0' + data)

I'm doing a defensive programming, for uncompressed loose object the mmapped memory is passed to parse_sha1_header without being checked by inflateInit() first, so there may be a SIGSEGV crash for a corrupted uncompressed loose object.

Show 15 quoted lines
>> @@ -1275,7 +1278,7 @@ static int parse_sha1_header(const char *hdr, unsigned long *sizep)
>>  	if (size > 9)
>>  		return -1;
>>  	if (size) {
>> -		for (;;) {
>> +		while (hdr < hdr_end - 1) {
>>  			unsigned long c = *hdr - '0';
>>  			if (c > 9)
>>  				break;
> 
> OK, there's no promise here that we don't roll off the buffer.
> 
> This can be fixed in the caller, ensuring we always have the '\0'
> at some point in the initial header buffer we were asked to parse:
> 

Isn't it easier to solve this problem in one place and maintain it? Maybe someday someone forgets parse_sha1_header requires a null terminated buffer, and a corrupted uncompressed loose object even doesn't have to be null terminated (if there will be this kind of loose object).

Previous: Shawn O. PearceNext: Liu Yubao
Message 11 of 27 in “two questions about the format of loose object”
  1. Liu YubaoDec 1, 2008
  2. Junio C HamanoDec 1, 2008
  3. Liu YubaoDec 1, 2008
  4. Jakub NarebskiDec 1, 2008
  5. Liu YubaoDec 2, 2008
  6. Shawn O. PearceDec 1, 2008
  7. Liu YubaoDec 2, 2008
  8. 0/5 support reading and writing uncompressed loose objectLiu Yubao, Dec 2, 2008
  9. 1/5 avoid parse_sha1_header() accessing memory out of boundLiu Yubao, Dec 2, 2008
  10. Shawn O. PearceDec 2, 2008
  11. Liu YubaoDec 3, 2008
  12. 2/5 don't die immediately when convert an invalid type nameLiu Yubao, Dec 2, 2008
  13. 3/5 optimize parse_sha1_header() a little by detecting object typeLiu Yubao, Dec 2, 2008
  14. Shawn O. PearceDec 2, 2008
  15. Liu YubaoDec 3, 2008
  16. 4/5 support reading uncompressed loose objectLiu Yubao, Dec 2, 2008
  17. Shawn O. PearceDec 2, 2008
  18. Liu YubaoDec 3, 2008
  19. 5/5 support writing uncompressed loose objectLiu Yubao, Dec 2, 2008
  20. Shawn O. PearceDec 2, 2008
  21. Liu YubaoDec 3, 2008
  22. Liu YubaoDec 2, 2008
  23. Nick AndrewDec 1, 2008
  24. Liu YubaoDec 2, 2008
  25. Shawn O. PearceDec 1, 2008
  26. Liu YubaoDec 2, 2008
  27. Nicolas PitreDec 4, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.