git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v2 01/30] object-file-convert: stubs for converting from one object format to another

From
LALinus Arver <linusa@google.com>
Date
Feb 8, 2024, 08:23 UTC
Message-ID
<owlyfry3cqkp.fsf@fine.c.googlers.com>
In-Reply-To
<20231002024034.2611-1-ebiederm@gmail.com>
"Eric W. Biederman" <ebiederm@gmail.com> writes:
Show 6 quoted lines
> From: "Eric W. Biederman" <ebiederm@xmission.com>
>
> Two basic functions are provided:
> - convert_object_file Takes an object file it's type and hash algorithm
>   and converts it into the equivalent object file that would
>   have been generated with hash algorithm "to".
Should probably be
    - convert_object_file takes an object file, its type, and hash algorithm
      and converts it into the equivalent object file using hash
      algorithm "to".

It would be nice if you gave the function name some clue of what sort of conversion is being done though. Something like "convert_object_file_hash" or "switch_object_file_hash". I like "switch" because the word "hash" itself means to convert some input bytes into another set of (aka "hashed") bytes, and the indirect shadowing of "convert" in this way can be avoided here by using a different word.The whole point of this function is to switch the hashing scheme from one to another, while still keeping everything else the same, so "switch" seems more appropriate.

Unless, of course, we already pervasively use "convert" this way elsewhere in the codebase (I have not checked).

>   For blob objects there is no conversation to be done and it is an
s/conversation/conversion
>   error to use this function on them.

Could you explain why no conversion is needed for blob objects, and also why it should be an error (and not just a NOP)?

In the code we can also call BUG() if the from/to algos are the same. It's probably worth mentioning in here as well?

Also for such detailed explanations, I think it's much better to place them as comments directly above the function (and only mention the important bits about these helper functions, other than the fact that they will come in handy in later patches, in the log message).

>   For commit, tree, and tag objects embedded oids are replaced by the
>   oids of the objects they refer to with those objects and their
>   object ids reencoded in with the hash algorithm "to".

That's a little wordy. I assume embedded oids just mean oids that these objects refer to (e.g., commit objects have oids, but for example the tree referred by a commit would be an example of an embedded oid). If so, then how about just

    For commit, tree, and tag objects both their oids and embedded
    (dependent) oids are converted using hash algorithm "to".

I dropped "reencoded" in favor of "converted" btecause that's the verb you use in your function name "convert_object_file()".

>   Signatures
>   are rearranged so that they remain valid after the object has
>   been reencoded.

Maybe s/reencoded/converted here as well? But also, it sounds odd to me that signatures are simply "rearranged" and not "regenerated" or "recreated" becaues "rearranged" means keeping most of the old stuff around but just repositioning them, which doesn't sound like it's doing justice to the meaning of a hash algo transition.

> - repo_oid_to_algop which takes an oid that refers to an object file
>   and returns the oid of the equivalent object file generated
>   with the target hash algorithm.

The name is odd to me because "repo_oid" doesn't make sense (a repo is not an object so it doesn't have an oid), but also the "to_algop" name (AFAICS "algop" just means "pointer to algorithm" in the codebase, for examle in <hash-ll.h>).

Here's a possible rewording:
    - switch_oid_hash takes an object file's oid and returns a new one
      using hash algorithm "to".
> The pair of files object-file-convert.c and object-file-convert.h are
> introduced to hold as much of this logic as possible to keep this
> conversion logic cleanly separated from everything else and in the
> hopes that someday the code will be clean enough git can support
Did you mean "clean enough so that Git ..."?
> compiling out support for sha1 and the various conversion functions.

FYI you can cut down this ambitious sentence into two for readability, like this:

    The new files object-file-convert.{c,h} hold as much of this logic
    as possible to keep this conversion logic cleanly separated from
    everything else. This separation will help us easily phase out SHA1
    support (perhaps as a compile-time flag) in the future.
Show 44 quoted lines
> Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
> ---
>  Makefile              |  1 +
>  object-file-convert.c | 57 +++++++++++++++++++++++++++++++++++++++++++
>  object-file-convert.h | 24 ++++++++++++++++++
>  3 files changed, 82 insertions(+)
>  create mode 100644 object-file-convert.c
>  create mode 100644 object-file-convert.h
>
> diff --git a/Makefile b/Makefile
> index 577630936535..f7e824f25cda 100644
> --- a/Makefile
> +++ b/Makefile
> @@ -1073,6 +1073,7 @@ LIB_OBJS += notes-cache.o
>  LIB_OBJS += notes-merge.o
>  LIB_OBJS += notes-utils.o
>  LIB_OBJS += notes.o
> +LIB_OBJS += object-file-convert.o
>  LIB_OBJS += object-file.o
>  LIB_OBJS += object-name.o
>  LIB_OBJS += object.o
> diff --git a/object-file-convert.c b/object-file-convert.c
> new file mode 100644
> index 000000000000..4777aba83636
> --- /dev/null
> +++ b/object-file-convert.c
> @@ -0,0 +1,57 @@
> +#include "git-compat-util.h"
> +#include "gettext.h"
> +#include "strbuf.h"
> +#include "repository.h"
> +#include "hash-ll.h"
> +#include "object.h"
> +#include "object-file-convert.h"
> +
> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,
> +		      const struct git_hash_algo *to, struct object_id *dest)
> +{
> +	/*
> +	 * If the source algorithm is not set, then we're using the
> +	 * default hash algorithm for that object.
> +	 */
> +	const struct git_hash_algo *from =
> +		src->algo ? &hash_algos[src->algo] : repo->hash_algo;

Hmm, if we're only using "repo" to grab its hash_algo member as a fallback, should this function be broken down into 2, one to get the fallback algo and another that has the meat of this function without the "repo" parameter?

I haven't read the rest of the series yet, let me keep reading.
Show 8 quoted lines
> +
> +	if (from == to) {
> +		if (src != dest)
> +			oidcpy(dest, src);
> +		return 0;
> +	}
> +	return -1;
> +}

It's curious to me that you treat same-hash-algo as a NOP here, but as a BUG() in the other helper. Why the difference (perhaps add a comment)?

Show 6 quoted lines
> +int convert_object_file(struct strbuf *outbuf,
> +			const struct git_hash_algo *from,
> +			const struct git_hash_algo *to,
> +			const void *buf, size_t len,
> +			enum object_type type,
> +			int gentle)

Is there a particular reason you chose to put the out parameter (outbuf) at the beginning, rather than at the end? In the other helper function the "dest" out param comes at the end. Style conflict...?

Also, you don't use buf or len, so you could have kept them out from this patch (so that you only add them later as needed).

> +{
> +	int ret;
> +
> +	/* Don't call this function when no conversion is necessary */
Please avoid double-negation. How about simply
    /* Refuse nonsensical conversion */

or simply drop the comment (as the BUG() description already serves the same purpose?)

> +	if ((from == to) || (type == OBJ_BLOB))
> +		BUG("Refusing noop object file conversion");

I think it would be better if you separated these out and gave them different BUG() messages. If we barf with a BUG() it would be so much more helpful if the message we get is as specific as possible, rather than leaving us guessing whether (as in this case) we passed in identical hash algos or whether we tried to handle a blob object. Since you already have the switch statement below, the OBJ_BLOB case could go there easily enough.

Nit: I think BUG() messages are not supposed to be capitalized.
Show 13 quoted lines
> +	switch (type) {
> +	case OBJ_COMMIT:
> +	case OBJ_TREE:
> +	case OBJ_TAG:
> +	default:
> +		/* Not implemented yet, so fail. */
> +		ret = -1;
> +		break;
> +	}
> +	if (!ret)
> +		return 0;
> +	if (gentle) {
> +		strbuf_release(outbuf);

What does "gentle" mean here? But also if you are just freeing the outbuf, then why not just name this param "free_outbuf"? But also it seems a bit odd that the caller (who presumably owns the outbuf) is asking a conversion function to possibly free it before returning.

Show 19 quoted lines
> +		return ret;
> +	}
> +	die(_("Failed to convert object from %s to %s"),
> +		from->name, to->name);
> +}
> diff --git a/object-file-convert.h b/object-file-convert.h
> new file mode 100644
> index 000000000000..a4f802aa8eea
> --- /dev/null
> +++ b/object-file-convert.h
> @@ -0,0 +1,24 @@
> +#ifndef OBJECT_CONVERT_H
> +#define OBJECT_CONVERT_H
> +
> +struct repository;
> +struct object_id;
> +struct git_hash_algo;
> +struct strbuf;
> +#include "object.h"

It looks a bit odd that the forward declarations of these structs come before the #include line. I don't think that's the pattern in our codebase (maybe I'm wrong?).

In hindsight, the log message could have added that switching back and forth between different hash algos is a fundamental (but currently missing) operation, and that this is the reason why these conversion functions (currently unfinished) are necessary as the first step for the patch series. I realize that you've stated as much in your cover letter, but it is always nice to have the intent embedded in the log message(s) where applicable (such as the case in this patch that introduces a brand new header file) to save future developers the hassle of looking up the relevant cover letter.

Thanks.
Show 17 quoted lines
> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,
> +		      const struct git_hash_algo *to, struct object_id *dest);
> +
> +/*
> + * Convert an object file from one hash algorithm to another algorithm.
> + * Return -1 on failure, 0 on success.
> + */
> +int convert_object_file(struct strbuf *outbuf,
> +			const struct git_hash_algo *from,
> +			const struct git_hash_algo *to,
> +			const void *buf, size_t len,
> +			enum object_type type,
> +			int gentle);
> +
> +#endif /* OBJECT_CONVERT_H */
> -- 
> 2.41.0
Previous: Eric W. BiedermanNext: Patrick Steinhardt
Message 51 of 104 in “Initial support for multiple hash functions”
  1. 00/30 Initial support for multiple hash functionsEric W. Biederman, Sep 27, 2023
  2. 01/30 object-file-convert: Stubs for converting from one object format to anotherEric W. Biederman, Sep 27, 2023
  3. Eric SunshineSep 27, 2023
  4. Eric W. BiedermanOct 2, 2023
  5. Eric SunshineOct 2, 2023
  6. 02/30 oid-array: Teach oid-array to handle multiple kinds of oidsEric W. Biederman, Sep 27, 2023
  7. Eric SunshineSep 27, 2023
  8. 04/30 repository: add a compatibility hash algorithmEric W. Biederman, Sep 27, 2023
  9. 03/30 object-names: Support input of oids in any supported hashEric W. Biederman, Sep 27, 2023
  10. Eric SunshineSep 27, 2023
  11. Eric W. BiedermanOct 2, 2023
  12. 05/30 loose: add a mapping between SHA-1 and SHA-256 for loose objectsEric W. Biederman, Sep 27, 2023
  13. Eric SunshineSep 28, 2023
  14. Eric W. BiedermanOct 2, 2023
  15. Eric SunshineOct 2, 2023
  16. 06/30 loose: Compatibilty short name supportEric W. Biederman, Sep 27, 2023
  17. 08/30 object-file: Add a compat_oid_in parameter to write_object_file_flagsEric W. Biederman, Sep 27, 2023
  18. 07/30 object-file: Update the loose object map when writing loose objectsEric W. Biederman, Sep 27, 2023
  19. 09/30 commit: write commits for both hashesEric W. Biederman, Sep 27, 2023
  20. 10/30 commit: Convert mergetag before computing the signature of a commitEric W. Biederman, Sep 27, 2023
  21. 11/30 commit: Export add_header_signature to support handling signatures on tagsEric W. Biederman, Sep 27, 2023
  22. 12/30 tag: sign both hashesEric W. Biederman, Sep 27, 2023
  23. 14/30 object: Factor out parse_mode out of fast-import and tree-walk into in object.hEric W. Biederman, Sep 27, 2023
  24. 13/30 cache: add a function to read an OID of a specific algorithmEric W. Biederman, Sep 27, 2023
  25. 15/30 object-file-convert: add a function to convert trees between algorithmsEric W. Biederman, Sep 27, 2023
  26. 16/30 object-file-convert: convert tag objects when writingEric W. Biederman, Sep 27, 2023
  27. 17/30 object-file-convert: Don't leak when converting tag objectsEric W. Biederman, Sep 27, 2023
  28. 18/30 object-file-convert: convert commit objects when writingEric W. Biederman, Sep 27, 2023
  29. 19/30 object-file-convert: Convert commits that embed signed tagsEric W. Biederman, Sep 27, 2023
  30. 20/30 object-file: Update object_info_extended to reencode objectsEric W. Biederman, Sep 27, 2023
  31. 22/30 rev-parse: Add an --output-object-format parameterEric W. Biederman, Sep 27, 2023
  32. 21/30 repository: Implement extensions.compatObjectFormatEric W. Biederman, Sep 27, 2023
  33. Junio C HamanoSep 27, 2023
  34. Junio C HamanoSep 28, 2023
  35. Eric BiedermanSep 29, 2023
  36. Eric W. BiedermanSep 29, 2023
  37. Junio C HamanoSep 29, 2023
  38. Eric W. BiedermanOct 2, 2023
  39. Eric W. BiedermanOct 2, 2023
  40. 23/30 builtin/cat-file: Let the oid determine the output algorithmEric W. Biederman, Sep 27, 2023
  41. 25/30 object-file: Handle compat objects in check_object_signatureEric W. Biederman, Sep 27, 2023
  42. 26/30 builtin/ls-tree: Let the oid determine the output algorithmEric W. Biederman, Sep 27, 2023
  43. 24/30 tree-walk: init_tree_desc take an oid to get the hash algorithmEric W. Biederman, Sep 27, 2023
  44. 27/30 test-lib: Compute the compatibility hash so tests may use itEric W. Biederman, Sep 27, 2023
  45. 29/30 t1006: Test oid compatibility with cat-fileEric W. Biederman, Sep 27, 2023
  46. 28/30 t1006: Rename sha1 to oidEric W. Biederman, Sep 27, 2023
  47. 30/30 t1016-compatObjectFormat: Add tests to verify the conversion between objectsEric W. Biederman, Sep 27, 2023
  48. Junio C HamanoSep 27, 2023
  49. 00/30 initial support for multiple hash functionsEric W. Biederman, Oct 2, 2023
  50. 01/30 object-file-convert: stubs for converting from one object format to anotherEric W. Biederman, Oct 2, 2023
  51. Linus ArverFeb 8, 2024
  52. Patrick SteinhardtFeb 15, 2024
  53. 02/30 oid-array: teach oid-array to handle multiple kinds of oidsEric W. Biederman, Oct 2, 2023
  54. Linus ArverFeb 13, 2024
  55. Eric W. BiedermanFeb 15, 2024
  56. Linus ArverFeb 16, 2024
  57. Eric W. BiedermanFeb 16, 2024
  58. Linus ArverFeb 17, 2024
  59. Kristoffer HaugsbakkFeb 13, 2024
  60. Eric W. BiedermanFeb 15, 2024
  61. Patrick SteinhardtFeb 15, 2024
  62. 03/30 object-names: support input of oids in any supported hashEric W. Biederman, Oct 2, 2023
  63. Linus ArverFeb 13, 2024
  64. Patrick SteinhardtFeb 15, 2024
  65. 04/30 repository: add a compatibility hash algorithmEric W. Biederman, Oct 2, 2023
  66. Linus ArverFeb 13, 2024
  67. Patrick SteinhardtFeb 15, 2024
  68. 06/30 loose: compatibilty short name supportEric W. Biederman, Oct 2, 2023
  69. Patrick SteinhardtFeb 15, 2024
  70. 05/30 loose: add a mapping between SHA-1 and SHA-256 for loose objectsEric W. Biederman, Oct 2, 2023
  71. Linus ArverFeb 14, 2024
  72. Eric W. BiedermanFeb 15, 2024
  73. Patrick SteinhardtFeb 15, 2024
  74. 07/30 object-file: update the loose object map when writing loose objectsEric W. Biederman, Oct 2, 2023
  75. Patrick SteinhardtFeb 15, 2024
  76. 08/30 object-file: add a compat_oid_in parameter to write_object_file_flagsEric W. Biederman, Oct 2, 2023
  77. 10/30 commit: convert mergetag before computing the signature of a commitEric W. Biederman, Oct 2, 2023
  78. 09/30 commit: write commits for both hashesEric W. Biederman, Oct 2, 2023
  79. 11/30 commit: export add_header_signature to support handling signatures on tagsEric W. Biederman, Oct 2, 2023
  80. 12/30 tag: sign both hashesEric W. Biederman, Oct 2, 2023
  81. 13/30 cache: add a function to read an OID of a specific algorithmEric W. Biederman, Oct 2, 2023
  82. 14/30 object: factor out parse_mode out of fast-import and tree-walk into in object.hEric W. Biederman, Oct 2, 2023
  83. 15/30 object-file-convert: add a function to convert trees between algorithmsEric W. Biederman, Oct 2, 2023
  84. 16/30 object-file-convert: convert tag objects when writingEric W. Biederman, Oct 2, 2023
  85. 17/30 object-file-convert: don't leak when converting tag objectsEric W. Biederman, Oct 2, 2023
  86. 19/30 object-file-convert: convert commits that embed signed tagsEric W. Biederman, Oct 2, 2023
  87. 18/30 object-file-convert: convert commit objects when writingEric W. Biederman, Oct 2, 2023
  88. 20/30 object-file: update object_info_extended to reencode objectsEric W. Biederman, Oct 2, 2023
  89. 21/30 repository: implement extensions.compatObjectFormatEric W. Biederman, Oct 2, 2023
  90. 22/30 rev-parse: add an --output-object-format parameterEric W. Biederman, Oct 2, 2023
  91. Jean-Noël AvilaFeb 8, 2024
  92. 23/30 builtin/cat-file: let the oid determine the output algorithmEric W. Biederman, Oct 2, 2023
  93. 25/30 object-file: handle compat objects in check_object_signatureEric W. Biederman, Oct 2, 2023
  94. 26/30 builtin/ls-tree: let the oid determine the output algorithmEric W. Biederman, Oct 2, 2023
  95. 24/30 tree-walk: init_tree_desc take an oid to get the hash algorithmEric W. Biederman, Oct 2, 2023
  96. 27/30 test-lib: compute the compatibility hash so tests may use itEric W. Biederman, Oct 2, 2023
  97. 29/30 t1006: test oid compatibility with cat-fileEric W. Biederman, Oct 2, 2023
  98. 28/30 t1006: rename sha1 to oidEric W. Biederman, Oct 2, 2023
  99. 30/30 t1016-compatObjectFormat: add tests to verify the conversion between objectsEric W. Biederman, Oct 2, 2023
  100. Junio C HamanoFeb 7, 2024
  101. Linus ArverFeb 8, 2024
  102. Patrick SteinhardtFeb 8, 2024
  103. Linus ArverFeb 14, 2024
  104. Patrick SteinhardtFeb 15, 2024

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.