git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids

From
EBEric W. Biederman <ebiederm@gmail.com>
Date
Feb 15, 2024, 06:22 UTC
Message-ID
<8734tumekr.fsf@gmail.froward.int.ebiederm.org>
In-Reply-To
<owlya5o4dbj1.fsf@fine.c.googlers.com>
Linus Arver <linusa@google.com> writes:
Show 52 quoted lines
> "Eric W. Biederman" <ebiederm@gmail.com> writes:
>
>> From: "Eric W. Biederman" <ebiederm@xmission.com>
>>
>> While looking at how to handle input of both SHA-1 and SHA-256 oids in
>> get_oid_with_context, I realized that the oid_array in
>> repo_for_each_abbrev might have more than one kind of oid stored in it
>> simultaneously.
>>
>> Update to oid_array_append to ensure that oids added to an oid array
>
> s/Update to/Update
>
>> always have an algorithm set.
>>
>> Update void_hashcmp to first verify two oids use the same hash algorithm
>> before comparing them to each other.
>>
>> With that oid-array should be safe to use with different kinds of
>
> s/oid-array/oid_array
>
>> oids simultaneously.
>>
>> Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
>> ---
>>  oid-array.c | 12 ++++++++++--
>>  1 file changed, 10 insertions(+), 2 deletions(-)
>>
>> diff --git a/oid-array.c b/oid-array.c
>> index 8e4717746c31..1f36651754ed 100644
>> --- a/oid-array.c
>> +++ b/oid-array.c
>> @@ -6,12 +6,20 @@ void oid_array_append(struct oid_array *array, const struct object_id *oid)
>>  {
>>  	ALLOC_GROW(array->oid, array->nr + 1, array->alloc);
>>  	oidcpy(&array->oid[array->nr++], oid);
>> +	if (!oid->algo)
>> +		oid_set_algo(&array->oid[array->nr - 1], the_hash_algo);
>
> How come we can't set oid->algo _before_ we call oidcpy()? It seems odd
> that we do the copy first and then modify what we just copied after the
> fact, instead of making sure that the thing we want to copy is correct
> before doing the copy.
>
> But also, if we are going to make the oid object "correct" before
> invoking oidcpy(), we might as well do it when the oid is first
> created/used (in the caller(s) of this function). I don't demand that
> you find/demonstrate where all these places are in this series (maybe
> that's a hairy problem to tackle?), but it seems cleaner in principle to
> fix the creation of oid objects instead of having to make oid users
> clean up their act like this after using them.
There is a hairy problem here.

I believe for reasons of simplicity when the algo field was added to struct object_id it was allowed to be zero for users that don't particularly care about the hash algorithm, and are happy to use the git default hash algorithm.

Me experience working on this set of change set showed that there are oids without their algo set in all kinds of places in the tree.

I could not think of any sure way to go through the entire tree and find those users, so I just made certain that oid array handled that case.

I need algo to be set properly in the oids in the oid array so I could extend oid_array to hold multiple kinds of oids at the same time. To allow multiple kinds of oids at the same time void_hashcmp needs a simple and reliable way to tell what the algorithm is of any given oid.

Show 23 quoted lines
>
>>  	array->sorted = 0;
>>  }
>>  
>> -static int void_hashcmp(const void *a, const void *b)
>> +static int void_hashcmp(const void *va, const void *vb)
>>  {
>> -	return oidcmp(a, b);
>> +	const struct object_id *a = va, *b = vb;
>> +	int ret;
>> +	if (a->algo == b->algo)
>> +		ret = oidcmp(a, b);
>
> This makes sense (per the commit message description) ...
>
>> +	else
>> +		ret = a->algo > b->algo ? 1 : -1;
>
> ... but this seems to go against it? I thought you wanted to only ever
> compare hashes if they were of the same algo? It would be good to add a
> comment explaining why this is OK (we are no longer doing a byte-by-byte
> comparison of these oids any more here like we do for oidcmp() above
> which boils down to calling memcmp()).

So the goal of this change is for oid_array to be able to hold hashes from multiple algorithms at the same time.

A key part of oid_array is oid_array_sort that allows functions such as oid_array_lookup and oid_array_for_each_unique.

To that end there needs to be a total ordering of oids.

The function oidcmp is only defined when two oids are of the same algorithm, it does not even test to detect the case of comparing mismatched algorithms.

Therefore to get a total ordering of oids. I must use oidcmp when the algorithm is the same (the common case) or simply order the oids by algorithm when the algorithms are different.

All of this is relevant to get_oid_with_context as get_oid_with_context and it's helper functions contain the logic that determines what we do when a hex string that is ambiguous is specified.

In the ambiguous case all of the possible candidates are placed in an oid_array, sorted and then displayed.

With a repository that can knows both the sha1 and the sha256 oid of it's objects it is possible for a short oid to match both some sha1 oids and some sha256 oids.

Show 13 quoted lines
>> +	return ret;
>
> Also, in terms of style I think the "early return for errors" style
> would be simpler to read. I.e.
>
>     if (a->algo > b->algo)
>         return 1;
>
>     if (a->algo < b->algo)
>         return -1;
>
>     return oidcmd(a, b);
>
I can see doing:
	if (a->algo == b->algo)
        	return oidcmp(a,b);
	if (a->algo > b->algo)
        	return 1;
        else
        	return -1;
Or even:
	if (a->algo == b->algo)
        	return oidcmp(a,b);
	return a->algo - b->algo;
Although I suspect using subtraction is a bit too clever.

Comparing for less than, and greater than, and then assuming the values are equal hides what is important before calling oidcmp which is that the algo values are equal.

Show 5 quoted lines
>>  }
>>  
>>  void oid_array_sort(struct oid_array *array)
>> -- 
>> 2.41.0
Eric
Previous: Linus ArverNext: Linus Arver
Message 55 of 104 in “Initial support for multiple hash functions”
  1. 00/30 Initial support for multiple hash functionsEric W. Biederman, Sep 27, 2023
  2. 01/30 object-file-convert: Stubs for converting from one object format to anotherEric W. Biederman, Sep 27, 2023
  3. Eric SunshineSep 27, 2023
  4. Eric W. BiedermanOct 2, 2023
  5. Eric SunshineOct 2, 2023
  6. 02/30 oid-array: Teach oid-array to handle multiple kinds of oidsEric W. Biederman, Sep 27, 2023
  7. Eric SunshineSep 27, 2023
  8. 04/30 repository: add a compatibility hash algorithmEric W. Biederman, Sep 27, 2023
  9. 03/30 object-names: Support input of oids in any supported hashEric W. Biederman, Sep 27, 2023
  10. Eric SunshineSep 27, 2023
  11. Eric W. BiedermanOct 2, 2023
  12. 05/30 loose: add a mapping between SHA-1 and SHA-256 for loose objectsEric W. Biederman, Sep 27, 2023
  13. Eric SunshineSep 28, 2023
  14. Eric W. BiedermanOct 2, 2023
  15. Eric SunshineOct 2, 2023
  16. 06/30 loose: Compatibilty short name supportEric W. Biederman, Sep 27, 2023
  17. 08/30 object-file: Add a compat_oid_in parameter to write_object_file_flagsEric W. Biederman, Sep 27, 2023
  18. 07/30 object-file: Update the loose object map when writing loose objectsEric W. Biederman, Sep 27, 2023
  19. 09/30 commit: write commits for both hashesEric W. Biederman, Sep 27, 2023
  20. 10/30 commit: Convert mergetag before computing the signature of a commitEric W. Biederman, Sep 27, 2023
  21. 11/30 commit: Export add_header_signature to support handling signatures on tagsEric W. Biederman, Sep 27, 2023
  22. 12/30 tag: sign both hashesEric W. Biederman, Sep 27, 2023
  23. 14/30 object: Factor out parse_mode out of fast-import and tree-walk into in object.hEric W. Biederman, Sep 27, 2023
  24. 13/30 cache: add a function to read an OID of a specific algorithmEric W. Biederman, Sep 27, 2023
  25. 15/30 object-file-convert: add a function to convert trees between algorithmsEric W. Biederman, Sep 27, 2023
  26. 16/30 object-file-convert: convert tag objects when writingEric W. Biederman, Sep 27, 2023
  27. 17/30 object-file-convert: Don't leak when converting tag objectsEric W. Biederman, Sep 27, 2023
  28. 18/30 object-file-convert: convert commit objects when writingEric W. Biederman, Sep 27, 2023
  29. 19/30 object-file-convert: Convert commits that embed signed tagsEric W. Biederman, Sep 27, 2023
  30. 20/30 object-file: Update object_info_extended to reencode objectsEric W. Biederman, Sep 27, 2023
  31. 22/30 rev-parse: Add an --output-object-format parameterEric W. Biederman, Sep 27, 2023
  32. 21/30 repository: Implement extensions.compatObjectFormatEric W. Biederman, Sep 27, 2023
  33. Junio C HamanoSep 27, 2023
  34. Junio C HamanoSep 28, 2023
  35. Eric BiedermanSep 29, 2023
  36. Eric W. BiedermanSep 29, 2023
  37. Junio C HamanoSep 29, 2023
  38. Eric W. BiedermanOct 2, 2023
  39. Eric W. BiedermanOct 2, 2023
  40. 23/30 builtin/cat-file: Let the oid determine the output algorithmEric W. Biederman, Sep 27, 2023
  41. 25/30 object-file: Handle compat objects in check_object_signatureEric W. Biederman, Sep 27, 2023
  42. 26/30 builtin/ls-tree: Let the oid determine the output algorithmEric W. Biederman, Sep 27, 2023
  43. 24/30 tree-walk: init_tree_desc take an oid to get the hash algorithmEric W. Biederman, Sep 27, 2023
  44. 27/30 test-lib: Compute the compatibility hash so tests may use itEric W. Biederman, Sep 27, 2023
  45. 29/30 t1006: Test oid compatibility with cat-fileEric W. Biederman, Sep 27, 2023
  46. 28/30 t1006: Rename sha1 to oidEric W. Biederman, Sep 27, 2023
  47. 30/30 t1016-compatObjectFormat: Add tests to verify the conversion between objectsEric W. Biederman, Sep 27, 2023
  48. Junio C HamanoSep 27, 2023
  49. 00/30 initial support for multiple hash functionsEric W. Biederman, Oct 2, 2023
  50. 01/30 object-file-convert: stubs for converting from one object format to anotherEric W. Biederman, Oct 2, 2023
  51. Linus ArverFeb 8, 2024
  52. Patrick SteinhardtFeb 15, 2024
  53. 02/30 oid-array: teach oid-array to handle multiple kinds of oidsEric W. Biederman, Oct 2, 2023
  54. Linus ArverFeb 13, 2024
  55. Eric W. BiedermanFeb 15, 2024
  56. Linus ArverFeb 16, 2024
  57. Eric W. BiedermanFeb 16, 2024
  58. Linus ArverFeb 17, 2024
  59. Kristoffer HaugsbakkFeb 13, 2024
  60. Eric W. BiedermanFeb 15, 2024
  61. Patrick SteinhardtFeb 15, 2024
  62. 03/30 object-names: support input of oids in any supported hashEric W. Biederman, Oct 2, 2023
  63. Linus ArverFeb 13, 2024
  64. Patrick SteinhardtFeb 15, 2024
  65. 04/30 repository: add a compatibility hash algorithmEric W. Biederman, Oct 2, 2023
  66. Linus ArverFeb 13, 2024
  67. Patrick SteinhardtFeb 15, 2024
  68. 06/30 loose: compatibilty short name supportEric W. Biederman, Oct 2, 2023
  69. Patrick SteinhardtFeb 15, 2024
  70. 05/30 loose: add a mapping between SHA-1 and SHA-256 for loose objectsEric W. Biederman, Oct 2, 2023
  71. Linus ArverFeb 14, 2024
  72. Eric W. BiedermanFeb 15, 2024
  73. Patrick SteinhardtFeb 15, 2024
  74. 07/30 object-file: update the loose object map when writing loose objectsEric W. Biederman, Oct 2, 2023
  75. Patrick SteinhardtFeb 15, 2024
  76. 08/30 object-file: add a compat_oid_in parameter to write_object_file_flagsEric W. Biederman, Oct 2, 2023
  77. 10/30 commit: convert mergetag before computing the signature of a commitEric W. Biederman, Oct 2, 2023
  78. 09/30 commit: write commits for both hashesEric W. Biederman, Oct 2, 2023
  79. 11/30 commit: export add_header_signature to support handling signatures on tagsEric W. Biederman, Oct 2, 2023
  80. 12/30 tag: sign both hashesEric W. Biederman, Oct 2, 2023
  81. 13/30 cache: add a function to read an OID of a specific algorithmEric W. Biederman, Oct 2, 2023
  82. 14/30 object: factor out parse_mode out of fast-import and tree-walk into in object.hEric W. Biederman, Oct 2, 2023
  83. 15/30 object-file-convert: add a function to convert trees between algorithmsEric W. Biederman, Oct 2, 2023
  84. 16/30 object-file-convert: convert tag objects when writingEric W. Biederman, Oct 2, 2023
  85. 17/30 object-file-convert: don't leak when converting tag objectsEric W. Biederman, Oct 2, 2023
  86. 19/30 object-file-convert: convert commits that embed signed tagsEric W. Biederman, Oct 2, 2023
  87. 18/30 object-file-convert: convert commit objects when writingEric W. Biederman, Oct 2, 2023
  88. 20/30 object-file: update object_info_extended to reencode objectsEric W. Biederman, Oct 2, 2023
  89. 21/30 repository: implement extensions.compatObjectFormatEric W. Biederman, Oct 2, 2023
  90. 22/30 rev-parse: add an --output-object-format parameterEric W. Biederman, Oct 2, 2023
  91. Jean-Noël AvilaFeb 8, 2024
  92. 23/30 builtin/cat-file: let the oid determine the output algorithmEric W. Biederman, Oct 2, 2023
  93. 25/30 object-file: handle compat objects in check_object_signatureEric W. Biederman, Oct 2, 2023
  94. 26/30 builtin/ls-tree: let the oid determine the output algorithmEric W. Biederman, Oct 2, 2023
  95. 24/30 tree-walk: init_tree_desc take an oid to get the hash algorithmEric W. Biederman, Oct 2, 2023
  96. 27/30 test-lib: compute the compatibility hash so tests may use itEric W. Biederman, Oct 2, 2023
  97. 29/30 t1006: test oid compatibility with cat-fileEric W. Biederman, Oct 2, 2023
  98. 28/30 t1006: rename sha1 to oidEric W. Biederman, Oct 2, 2023
  99. 30/30 t1016-compatObjectFormat: add tests to verify the conversion between objectsEric W. Biederman, Oct 2, 2023
  100. Junio C HamanoFeb 7, 2024
  101. Linus ArverFeb 8, 2024
  102. Patrick SteinhardtFeb 8, 2024
  103. Linus ArverFeb 14, 2024
  104. Patrick SteinhardtFeb 15, 2024

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.