git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v4 5/6] convert: add 'working-tree-encoding' attribute

From
Lars Schneider <larsxschneider@gmail.com>
Date
Jan 22, 2018, 12:35 UTC
Message-ID
<05265803-BD74-4667-ABB5-9752E55A5015@gmail.com>
In-Reply-To
<20180121142222.GA10248@ruderich.org>
Show 24 quoted lines
> On 21 Jan 2018, at 15:22, Simon Ruderich <simon@ruderich.org> wrote:
> 
> On Sat, Jan 20, 2018 at 04:24:17PM +0100, lars.schneider@autodesk.com wrote:
>> +static struct encoding *git_path_check_encoding(struct attr_check_item *check)
>> +{
>> +	const char *value = check->value;
>> +	struct encoding *enc;
>> +
>> +	if (ATTR_TRUE(value) || ATTR_FALSE(value) || ATTR_UNSET(value) ||
>> +	    !strlen(value))
>> +		return NULL;
>> +
>> +	for (enc = encoding; enc; enc = enc->next)
>> +		if (!strcasecmp(value, enc->name))
>> +			return enc;
>> +
>> +	/* Don't encode to the default encoding */
>> +	if (!strcasecmp(value, default_encoding))
>> +		return NULL;
>> +
>> +	enc = xcalloc(1, sizeof(struct convert_driver));
> 
> I think this should be "sizeof(struct encoding)" but I prefer
> "sizeof(*enc)" which prevents these kind of mistakes.
Great catch! Thank you!
Other code paths are at risk of this problem too. Consider this:
$ git grep 'sizeof(\*' | wc -l
     303
$ git grep 'sizeof(struct ' | wc -l
     208

E.g. even in the same file (likely where I got the code from): https://github.com/git/git/blob/59c276cf4da0705064c32c9dba54baefa282ea55/convert.c#L780

@Junio: Would you welcome a patch that replaces "struct foo" with "*foo" if applicable?

>> +	enc->name = xstrdup_toupper(value);  /* aways use upper case names! */
> 
> "aways" -> "always" and I think the comment should say why
> uppercase is important.
Would that be better?
	/* Aways use upper case names to simplify subsequent string comparison. */
	enc->name = xstrdup_toupper(value);

AFAIK uppercase and lowercase names are both valid. I just wanted to ensure that we use one consistent casing. That reads better in error messages and I don't need to check for the letter case in has_prohibited_utf_bom() and friends in utf8.c

Show 17 quoted lines
>> +test_expect_success 'ensure UTF-8 is stored in Git' '
>> +	git cat-file -p :test.utf16 >test.utf16.git &&
>> +	test_cmp_bin test.utf8.raw test.utf16.git &&
>> +	rm test.utf8.raw test.utf16.git
>> +'
>> +
>> +test_expect_success 're-encode to UTF-16 on checkout' '
>> +	rm test.utf16 &&
>> +	git checkout test.utf16 &&
>> +	test_cmp_bin test.utf16.raw test.utf16 &&
>> +
>> +	# cleanup
>> +	rm test.utf16.raw
> 
> Micro-nit: For consistency with the previous test, remove the
> empty line and comment (or just keep the files generated from the
> "setup test repo" phase and don't explicitly delete them)?

I would rather add a new line and a comment to the previous test to be consistent.

I know we could leave the files but these lingering files could always surprise writers of future tests (at least they surprised me in other tests).

Thank you very much for the review, Lars

Previous: Simon RuderichNext: Jeff King
Message 8 of 43 in “convert: add support for different encodings”
  1. 0/6 convert: add support for different encodingslars.schneider@autodesk.com, Jan 20, 2018
  2. 1/6 strbuf: remove unnecessary NUL assignment in xstrdup_tolower()lars.schneider@autodesk.com, Jan 20, 2018
  3. 2/6 strbuf: add xstrdup_toupper()lars.schneider@autodesk.com, Jan 20, 2018
  4. 3/6 utf8: add function to detect prohibited UTF-16/32 BOMlars.schneider@autodesk.com, Jan 20, 2018
  5. 4/6 utf8: add function to detect a missing UTF-16/32 BOMlars.schneider@autodesk.com, Jan 20, 2018
  6. 5/6 convert: add 'working-tree-encoding' attributelars.schneider@autodesk.com, Jan 20, 2018
  7. Simon RuderichJan 21, 2018
  8. Lars SchneiderJan 22, 2018
  9. Jeff KingJan 23, 2018
  10. Simon RuderichJan 23, 2018
  11. Jeff KingJan 23, 2018
  12. Junio C HamanoJan 23, 2018
  13. Simon RuderichJan 23, 2018
  14. SQUASH convert: add tracing for 'working-tree-encoding' attributelars.schneider@autodesk.com, Jan 22, 2018
  15. Eric SunshineJan 22, 2018
  16. SQUASH convert: add tracing for 'working-tree-encoding' attributelars.schneider@autodesk.com, Jan 23, 2018
  17. 6/6 convert: add tracing for 'working-tree-encoding' attributelars.schneider@autodesk.com, Jan 20, 2018
  18. Torsten BögershausenJan 23, 2018
  19. Junio C HamanoJan 23, 2018
  20. 0/7 convert: add support for different encodingstboegi@web.de, Jan 29, 2018
  21. 2/7 strbuf: add xstrdup_toupper()tboegi@web.de, Jan 29, 2018
  22. 3/7 utf8: add function to detect prohibited UTF-16/32 BOMtboegi@web.de, Jan 29, 2018
  23. 6/7 convert: add tracing for 'working-tree-encoding' attributetboegi@web.de, Jan 29, 2018
  24. 4/7 utf8: add function to detect a missing UTF-16/32 BOMtboegi@web.de, Jan 29, 2018
  25. Junio C HamanoJan 30, 2018
  26. Lars SchneiderJan 30, 2018
  27. Junio C HamanoJan 30, 2018
  28. 5/7 convert: add 'working-tree-encoding' attributetboegi@web.de, Jan 29, 2018
  29. Junio C HamanoJan 30, 2018
  30. Lars SchneiderJan 30, 2018
  31. Junio C HamanoJan 30, 2018
  32. Lars SchneiderJan 31, 2018
  33. Junio C HamanoJan 31, 2018
  34. 7/7 Careful with CRLF when using e.g. UTF-16 for working-tree-encodingtboegi@web.de, Jan 29, 2018
  35. Lars SchneiderJan 30, 2018
  36. Torsten BögershausenJan 30, 2018
  37. Lars SchneiderJan 30, 2018
  38. Torsten BögershausenJan 31, 2018
  39. Lars SchneiderJan 31, 2018
  40. Junio C HamanoFeb 2, 2018
  41. Torsten BögershausenFeb 7, 2018
  42. Junio C HamanoFeb 7, 2018
  43. 1/7 strbuf: remove unnecessary NUL assignment in xstrdup_tolower()tboegi@web.de, Jan 29, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.