git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 01/30] object-file-convert: Stubs for converting from one object format to another

From
EBEric W. Biederman <ebiederm@gmail.com>
Date
Oct 2, 2023, 01:22 UTC
Message-ID
<87o7hhbyxi.fsf@gmail.froward.int.ebiederm.org>
In-Reply-To
<CAPig+cRshiUNXfU=ZY4nZXgBgTJ_wF0WVDxWpqkEKPAT9pjX_w@mail.gmail.com>
Eric Sunshine <sunshine@sunshineco.com> writes:
Show 80 quoted lines
> On Wed, Sep 27, 2023 at 3:55 PM Eric W. Biederman <ebiederm@gmail.com> wrote:
>> Two basic functions are provided:
>> - convert_object_file Takes an object file it's type and hash algorithm
>>   and converts it into the equivalent object file that would
>>   have been generated with hash algorithm "to".
>>
>>   For blob objects there is no converstion to be done and it is an
>>   error to use this function on them.
>
> s/converstion/conversion/
>
>>   For commit, tree, and tag objects embedded oids are replaced by the
>>   oids of the objects they refer to with those objects and their
>>   object ids reencoded in with the hash algorithm "to".  Signatures
>>   are rearranged so that they remain valid after the object has
>>   been reencoded.
>>
>> - repo_oid_to_algop which takes an oid that refers to an object file
>>   and returns the oid of the equavalent object file generated
>>   with the target hash algorithm.
>
> s/equavalent/equivalent/
>
>> The pair of files object-file-convert.c and object-file-convert.h are
>> introduced to hold as much of this logic as possible to keep this
>> conversion logic cleanly separated from everything else and in the
>> hopes that someday the code will be clean enough git can support
>> compiling out support for sha1 and the various conversion functions.
>>
>> Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
>
> Just some minor comments below, many of which are subjective
> style-related observations, thus not necessarily actionable, but also
> one or two legitimate questions.
>
>> diff --git a/object-file-convert.c b/object-file-convert.c
>> @@ -0,0 +1,57 @@
>> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,
>> +                     const struct git_hash_algo *to, struct object_id *dest)
>> +{
>> +       /*
>> +        * If the source alogirthm is not set, then we're using the
>> +        * default hash algorithm for that object.
>> +        */
>
> s/alogirthm/algorithm/
>
>> +       const struct git_hash_algo *from =
>> +               src->algo ? &hash_algos[src->algo] : repo->hash_algo;
>> +
>> +       if (from == to) {
>> +               if (src != dest)
>> +                       oidcpy(dest, src);
>> +               return 0;
>> +       }
>> +       return -1;
>> +}
>
> On this project, we usually get the simple cases out of the way first,
> which often reduces the indentation level, making the code easier to
> digest at a glance. So, it would be typical to write this as:
>
>     if (from != to)
>         return -1
>     if (src != dest)
>         oidcpy(dest, src);
>     return 0;
>
> or even:
>
>     if (from != to)
>         return -1
>     if (src == dest)
>         return 0;
>     oidcpy(dest, src);
>     return 0;
>
> This way, for instance, the reader doesn't get to the end of the
> function and then have to scan backward to understand the condition of
> the `return -1`.

The "return -1" is there only because it is a stub, and it is there where the rest of the code needs to go.

As for simple cases the "if (from == to)" case is a simple case I am getting out of the way. It unfortunately is cluttered by the fact that "oidcpy(&oid, &oid)" is not valid so it has to guard the copy.

if (from == to) {
	if (src == dest)
        	return 0;
        oidcpy(dest, src);
        return 0;
}

Could be used but is wordier. And duplicates the return code for the same case so I am not enthusiastic about it.

Show 18 quoted lines
>> +int convert_object_file(struct strbuf *outbuf,
>> +                       const struct git_hash_algo *from,
>> +                       const struct git_hash_algo *to,
>> +                       const void *buf, size_t len,
>> +                       enum object_type type,
>> +                       int gentle)
>> +{
>> +       int ret;
>> +
>> +       /* Don't call this function when no conversion is necessary */
>> +       if ((from == to) || (type == OBJ_BLOB))
>> +               die("Refusing noop object file conversion");
>
> Several comments...
>
> Style: we usually reduce the noise level by dropping the extra parentheses:
>
>     if (from == to || type == OBJ_BLOB)
I honestly can not be confident of C code that does that.

The precedence of the operators in C has been wrong for longer than I have been programming, and I can never remember exactly how the precedence is wrong. So for the last 30 years I have been adding enough parenthesis that I don't have to remember.

> Does this condition represent a programming error or a runtime error
> triggerable by some input? If a programming error, then use BUG()
> rather than die().
Agreed BUG would be better there.
Show 33 quoted lines
> If a triggerable runtime error, then...
>
> * start user-facing messages with lowercase rather than capitalized word
>
> * make the user-facing message localizable so readers of other
> languages can digest it
>
>     die(_("refusing do-nothing object conversion"));
>
> On the other hand, don't make BUG() messages localizable.
>
>> +       switch (type) {
>> +       case OBJ_COMMIT:
>> +       case OBJ_TREE:
>> +       case OBJ_TAG:
>> +       default:
>> +               /* Not implemented yet, so fail. */
>> +               ret = -1;
>> +               break;
>> +       }
>> +       if (!ret)
>> +               return 0;
>> +       if (gentle) {
>> +               strbuf_release(outbuf);
>> +               return ret;
>> +       }
>
> This function appears to be a mere skeleton at the moment, so it's
> difficult to judge at this point whether you are using `outbuf` as a
> bag of bytes or as a legitimate string container. If the latter, then
> the API may be reasonable, but if you're using it as a bag-of-bytes,
> then it feels like you're leaking an implementation detail into the
> API.
It is a string that represents the entire object.

It might be arguable if tree objects are text given that trees represent oids in binary, but for tag and commit objects they are definitely one big text string.

This is an implementation detail that makes the code simpler, and less error prone.

I have not encountered anything where a string buffer would not be a reasonable fit.

>> +       die(_("Failed to convert object from %s to %s"),
>> +               from->name, to->name);
>
> s/Failed/failed/

I don't understand wanting to start a sentence with a lower case letter. Can you explain?

> For people trying to diagnose this problem, would it be helpful to
> present more information about the failed conversion, such as object
> type and perhaps even its OID?

I expect some of that context will come from the conversion functions themselves. Those messages are almost certain to give the type information one way or another because they are type specific.

That said it requires some version of corrupt repository to reach this error. Either missing mapping tables, or a corrupt object.

So I don't know how much it matters to get this perfect the first time. We can improve the error message we gain experience.

I would agree that including the OID makes sense. Unfortunately the code does not have the OID at this point.

Show 8 quoted lines
>> diff --git a/object-file-convert.h b/object-file-convert.h
>> @@ -0,0 +1,24 @@
>> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,
>> +                     const struct git_hash_algo *to, struct object_id *dest);
>
> I suppose the function name is pretty much self-explanatory to those
> familiar with the underlying concepts, but it might still be helpful
> to add a comment explaining what the function does.

I could use words that repeat what is in the function signature. But I don't think I could add anything.

I would have to say something like:

Look up the oid that an equivalent object would have in a repository whose object format is "to".

Is that helpful?
Eric
Previous: Eric SunshineNext: Eric Sunshine
Message 4 of 104 in “Initial support for multiple hash functions”
  1. 00/30 Initial support for multiple hash functionsEric W. Biederman, Sep 27, 2023
  2. 01/30 object-file-convert: Stubs for converting from one object format to anotherEric W. Biederman, Sep 27, 2023
  3. Eric SunshineSep 27, 2023
  4. Eric W. BiedermanOct 2, 2023
  5. Eric SunshineOct 2, 2023
  6. 02/30 oid-array: Teach oid-array to handle multiple kinds of oidsEric W. Biederman, Sep 27, 2023
  7. Eric SunshineSep 27, 2023
  8. 04/30 repository: add a compatibility hash algorithmEric W. Biederman, Sep 27, 2023
  9. 03/30 object-names: Support input of oids in any supported hashEric W. Biederman, Sep 27, 2023
  10. Eric SunshineSep 27, 2023
  11. Eric W. BiedermanOct 2, 2023
  12. 05/30 loose: add a mapping between SHA-1 and SHA-256 for loose objectsEric W. Biederman, Sep 27, 2023
  13. Eric SunshineSep 28, 2023
  14. Eric W. BiedermanOct 2, 2023
  15. Eric SunshineOct 2, 2023
  16. 06/30 loose: Compatibilty short name supportEric W. Biederman, Sep 27, 2023
  17. 08/30 object-file: Add a compat_oid_in parameter to write_object_file_flagsEric W. Biederman, Sep 27, 2023
  18. 07/30 object-file: Update the loose object map when writing loose objectsEric W. Biederman, Sep 27, 2023
  19. 09/30 commit: write commits for both hashesEric W. Biederman, Sep 27, 2023
  20. 10/30 commit: Convert mergetag before computing the signature of a commitEric W. Biederman, Sep 27, 2023
  21. 11/30 commit: Export add_header_signature to support handling signatures on tagsEric W. Biederman, Sep 27, 2023
  22. 12/30 tag: sign both hashesEric W. Biederman, Sep 27, 2023
  23. 14/30 object: Factor out parse_mode out of fast-import and tree-walk into in object.hEric W. Biederman, Sep 27, 2023
  24. 13/30 cache: add a function to read an OID of a specific algorithmEric W. Biederman, Sep 27, 2023
  25. 15/30 object-file-convert: add a function to convert trees between algorithmsEric W. Biederman, Sep 27, 2023
  26. 16/30 object-file-convert: convert tag objects when writingEric W. Biederman, Sep 27, 2023
  27. 17/30 object-file-convert: Don't leak when converting tag objectsEric W. Biederman, Sep 27, 2023
  28. 18/30 object-file-convert: convert commit objects when writingEric W. Biederman, Sep 27, 2023
  29. 19/30 object-file-convert: Convert commits that embed signed tagsEric W. Biederman, Sep 27, 2023
  30. 20/30 object-file: Update object_info_extended to reencode objectsEric W. Biederman, Sep 27, 2023
  31. 22/30 rev-parse: Add an --output-object-format parameterEric W. Biederman, Sep 27, 2023
  32. 21/30 repository: Implement extensions.compatObjectFormatEric W. Biederman, Sep 27, 2023
  33. Junio C HamanoSep 27, 2023
  34. Junio C HamanoSep 28, 2023
  35. Eric BiedermanSep 29, 2023
  36. Eric W. BiedermanSep 29, 2023
  37. Junio C HamanoSep 29, 2023
  38. Eric W. BiedermanOct 2, 2023
  39. Eric W. BiedermanOct 2, 2023
  40. 23/30 builtin/cat-file: Let the oid determine the output algorithmEric W. Biederman, Sep 27, 2023
  41. 25/30 object-file: Handle compat objects in check_object_signatureEric W. Biederman, Sep 27, 2023
  42. 26/30 builtin/ls-tree: Let the oid determine the output algorithmEric W. Biederman, Sep 27, 2023
  43. 24/30 tree-walk: init_tree_desc take an oid to get the hash algorithmEric W. Biederman, Sep 27, 2023
  44. 27/30 test-lib: Compute the compatibility hash so tests may use itEric W. Biederman, Sep 27, 2023
  45. 29/30 t1006: Test oid compatibility with cat-fileEric W. Biederman, Sep 27, 2023
  46. 28/30 t1006: Rename sha1 to oidEric W. Biederman, Sep 27, 2023
  47. 30/30 t1016-compatObjectFormat: Add tests to verify the conversion between objectsEric W. Biederman, Sep 27, 2023
  48. Junio C HamanoSep 27, 2023
  49. 00/30 initial support for multiple hash functionsEric W. Biederman, Oct 2, 2023
  50. 01/30 object-file-convert: stubs for converting from one object format to anotherEric W. Biederman, Oct 2, 2023
  51. Linus ArverFeb 8, 2024
  52. Patrick SteinhardtFeb 15, 2024
  53. 02/30 oid-array: teach oid-array to handle multiple kinds of oidsEric W. Biederman, Oct 2, 2023
  54. Linus ArverFeb 13, 2024
  55. Eric W. BiedermanFeb 15, 2024
  56. Linus ArverFeb 16, 2024
  57. Eric W. BiedermanFeb 16, 2024
  58. Linus ArverFeb 17, 2024
  59. Kristoffer HaugsbakkFeb 13, 2024
  60. Eric W. BiedermanFeb 15, 2024
  61. Patrick SteinhardtFeb 15, 2024
  62. 03/30 object-names: support input of oids in any supported hashEric W. Biederman, Oct 2, 2023
  63. Linus ArverFeb 13, 2024
  64. Patrick SteinhardtFeb 15, 2024
  65. 04/30 repository: add a compatibility hash algorithmEric W. Biederman, Oct 2, 2023
  66. Linus ArverFeb 13, 2024
  67. Patrick SteinhardtFeb 15, 2024
  68. 06/30 loose: compatibilty short name supportEric W. Biederman, Oct 2, 2023
  69. Patrick SteinhardtFeb 15, 2024
  70. 05/30 loose: add a mapping between SHA-1 and SHA-256 for loose objectsEric W. Biederman, Oct 2, 2023
  71. Linus ArverFeb 14, 2024
  72. Eric W. BiedermanFeb 15, 2024
  73. Patrick SteinhardtFeb 15, 2024
  74. 07/30 object-file: update the loose object map when writing loose objectsEric W. Biederman, Oct 2, 2023
  75. Patrick SteinhardtFeb 15, 2024
  76. 08/30 object-file: add a compat_oid_in parameter to write_object_file_flagsEric W. Biederman, Oct 2, 2023
  77. 10/30 commit: convert mergetag before computing the signature of a commitEric W. Biederman, Oct 2, 2023
  78. 09/30 commit: write commits for both hashesEric W. Biederman, Oct 2, 2023
  79. 11/30 commit: export add_header_signature to support handling signatures on tagsEric W. Biederman, Oct 2, 2023
  80. 12/30 tag: sign both hashesEric W. Biederman, Oct 2, 2023
  81. 13/30 cache: add a function to read an OID of a specific algorithmEric W. Biederman, Oct 2, 2023
  82. 14/30 object: factor out parse_mode out of fast-import and tree-walk into in object.hEric W. Biederman, Oct 2, 2023
  83. 15/30 object-file-convert: add a function to convert trees between algorithmsEric W. Biederman, Oct 2, 2023
  84. 16/30 object-file-convert: convert tag objects when writingEric W. Biederman, Oct 2, 2023
  85. 17/30 object-file-convert: don't leak when converting tag objectsEric W. Biederman, Oct 2, 2023
  86. 19/30 object-file-convert: convert commits that embed signed tagsEric W. Biederman, Oct 2, 2023
  87. 18/30 object-file-convert: convert commit objects when writingEric W. Biederman, Oct 2, 2023
  88. 20/30 object-file: update object_info_extended to reencode objectsEric W. Biederman, Oct 2, 2023
  89. 21/30 repository: implement extensions.compatObjectFormatEric W. Biederman, Oct 2, 2023
  90. 22/30 rev-parse: add an --output-object-format parameterEric W. Biederman, Oct 2, 2023
  91. Jean-Noël AvilaFeb 8, 2024
  92. 23/30 builtin/cat-file: let the oid determine the output algorithmEric W. Biederman, Oct 2, 2023
  93. 25/30 object-file: handle compat objects in check_object_signatureEric W. Biederman, Oct 2, 2023
  94. 26/30 builtin/ls-tree: let the oid determine the output algorithmEric W. Biederman, Oct 2, 2023
  95. 24/30 tree-walk: init_tree_desc take an oid to get the hash algorithmEric W. Biederman, Oct 2, 2023
  96. 27/30 test-lib: compute the compatibility hash so tests may use itEric W. Biederman, Oct 2, 2023
  97. 29/30 t1006: test oid compatibility with cat-fileEric W. Biederman, Oct 2, 2023
  98. 28/30 t1006: rename sha1 to oidEric W. Biederman, Oct 2, 2023
  99. 30/30 t1016-compatObjectFormat: add tests to verify the conversion between objectsEric W. Biederman, Oct 2, 2023
  100. Junio C HamanoFeb 7, 2024
  101. Linus ArverFeb 8, 2024
  102. Patrick SteinhardtFeb 8, 2024
  103. Linus ArverFeb 14, 2024
  104. Patrick SteinhardtFeb 15, 2024

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.