git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 01/32] doc hash-file-transition: A map file for mapping between sha1 and sha256

From
Eric W. Biederman <ebiederm@xmission.com>
Date
Sep 12, 2023, 13:36 UTC
Message-ID
<87zg1r1pk9.fsf@email.froward.int.ebiederm.org>
In-Reply-To
<ZP+tTFK3Ly4sqlsq@tapette.crustytoothpaste.net>
"brian m. carlson" <sandals@crustytoothpaste.net> writes:
Show 54 quoted lines
> On 2023-09-08 at 23:10:18, Eric W. Biederman wrote:
>> The v3 pack index file as documented has a lot of complexity making it
>> difficult to implement correctly.  I worked with bryan's preliminary
>> implementation and it took several passes to get the bugs out.
>> 
>> The complexity also requires multiple table look-ups to find all of
>> the information that is needed to translate from one kind of oid to
>> another.  Which can't be good for cache locality.
>> 
>> Even worse coming up with a new index file version requires making
>> changes that have the potentialy to break anything that uses the index
>> of a pack file.
>> 
>> Instead of continuing to deal with the chance of braking things
>> besides the oid mapping functionality, the additional complexity in
>> the file format, and worry if the performance would be reasonable I
>> stripped down the problem to it's fundamental complexity and came up
>> with a file format that is exactly about mapping one kind of oid to
>> another, and only supports two kinds of oids.
>> 
>> Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
>> ---
>>  .../technical/hash-function-transition.txt    | 40 +++++++++++++++++++
>>  1 file changed, 40 insertions(+)
>> 
>> diff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt
>> index ed574810891c..4b937480848a 100644
>> --- a/Documentation/technical/hash-function-transition.txt
>> +++ b/Documentation/technical/hash-function-transition.txt
>> @@ -209,6 +209,46 @@ format described in linkgit:gitformat-pack[5], just like
>>  today. The content that is compressed and stored uses SHA-256 content
>>  instead of SHA-1 content.
>>  
>> +Per Pack Mapping Table
>> +~~~~~~~~~~~~~~~~~~~~~~
>> +A pack compat map file (.compat) files have the following format:
>> +
>> +HEADER:
>> +	4-byte signature:
>> +	    The signature is: {'C', 'M', 'A', 'P'}
>> +	1-byte version number:
>> +	    Git only writes or recognizes version 1.
>> +	1-byte First Object Id Version
>> +	    We infer the length of object IDs (OIDs) from this value:
>> +		1 => SHA-1
>> +		2 => SHA-256
>
> One thing I forgot to mention here, is that we have 32-bit format IDs
> for these in the structure, so we should use them here and below.  These
> are GIT_SHA1_FORMAT_ID and GIT_SHA256_FORMAT_ID.
>
> Not that I would encourage distributing such software, but it makes it
> much easier for people to experiment with additional hash algorithms (in
> terms of performance, etc.) if we make the space a little sparser.

Unfortunately that ship has already sailed. If you look at pack reverse indices, pack mtime files, multi-pack-index files, they all use an oid_version field. So to experiment with a new hash function a new number has to be picked.

The only use I can find of your 4 byte format_id's is in the reftable code.

Using a 4 byte magic number in this case also conflicts with basic simplicity. With a one byte field I can specify it easily, and read it back with no special tools, and understand what it means at a glance.

I admit I can only understand what a oid version field means at a glance because the variation of object id's is low, but that is fundamental. We require global agreement on names. Fundamentally git can not support many object id transitions. Names are just too expensive.

When I come to how the map file is specified a single byte has real advantages. A single byte never needs byte swapping. So it won't be misread. Using a single byte for each format allows me to keep the header for the file at 16 bytes. Which guarantees good alignment of everything in the file without having to be clever.

All of this is for a file that is strictly local and the entire function of the bytes is a sanity check to make certain that something weird is not going on, or to assist recover if something bad happens.

So in this case I don't see any additional agility provided by longer names helping.

Show 23 quoted lines
>> +	1-byte Second Object Id Version
>> +	    We infer the length of object IDs (OIDs) from this value:
>> +		1 => SHA-1
>> +		2 => SHA-256
>
> In your new patch for the next part, you consider that there might be
> multiple compatibility hash algorithms.  I had anticipated only one at
> a time in my series, but I'm not opposed to multiple if you want to
> support that.
>
> However, here you're making the assumption that there are only two.  If
> you want to support multiple values, we need to explicitly consider that
> both here (where we need a count of object ID version and multiple
> tables, one for each algorithm), and in the follow-up series.
>
> I had not considered more than two algorithms because it substantially
> complicates the code and requires us to develop n*(n-1) tables, but I'm
> not the one volunteering to do most of the work here, so I'll defer to
> your preference.  (I do intend to send a patch or two, though.)
>
> It's also possible we could be somewhat provident and define the on-disk
> formats for multiple algorithms and then punt on the code until later if
> you prefer that.

In the long term I anticipate people disabling compatObjectFormat and switching to readCompatMap so they still have access to their old objects by their original names, but they don't have the over head of computing a compatibility hash.

In a world where there is a transition to futureHash I anticipate the files associated with an old pack looking something like: pack-abcdefg.compat12 pack-abcdefg.compat32

For a repository still using hash version sha256 for storage, with a mapping to some sha1 names, and a mapping of everything to new names for compatibility with futureHash.

After transitioning to the futureHash those files would look like: pack-abcdefg.compat13 pack-abcdefg.compat23

I deeply and fundamentally care about having some way to look up old names because I do that all of the time.

In my work on the linux-kernel I have found myself frequently digging into old issues. I have on my hard drive tglx's git import of the old bitkeeper tree. I also have an import of all of the old kernel releases into git from before the code was stored in bitkeeper. I find myself actually using all of those trees when digging into issues.

So I think the idea that we will ever be able to get rid of the mapping for old converted repositories is unlikely. We have entirely too many references out there.

Which means that for every hash format conversion a repository goes through I am going to have another collection of old names.

I don't honestly anticipate ever needing to have multiple compatObjectFormat entries specified for a single repository. I do agree that if we are going to worry about forward and backward compatibility we should be robust and have a configuration file syntax that can handle the possibility.

I do very much anticipate needing to have multiple readCompatMap entries, and pretty much only using them in get_short_oid in object-name.c. It will make the loop in find_short_packed_compat_object a little longer but that is about all that will need to be implemented and maintained long term.

I view this compat map format a lot like the loose objects. It is simple and good enough to get us started. If it turns out we need to optimize it's simplicity means all of the interfaces in the code to use it have already been built, and we can just concentrate on optimizing.

Eric
Previous: brian m. carlsonNext: Eric W. Biederman
Message 35 of 59 in “SHA256 and SHA1 interoperability”
  1. Eric W. BiedermanSep 8, 2023
  2. 02/32 doc hash-function-transition: Replace compatObjectFormat with compatMapEric W. Biederman, Sep 8, 2023
  3. brian m. carlsonSep 10, 2023
  4. Eric W. BiedermanSep 10, 2023
  5. Junio C HamanoSep 11, 2023
  6. 02/32 doc hash-function-transition: Replace compatObjectFormat with mapObjectFormatEric W. Biederman, Sep 11, 2023
  7. 02/32 doc hash-function-transition: Augment compatObjectFormat with readCompatMapEric W. Biederman, Sep 11, 2023
  8. Oswald BuddenhagenSep 12, 2023
  9. Eric W. BiedermanSep 12, 2023
  10. Oswald BuddenhagenSep 13, 2023
  11. 04/32 object-name: Initial support for ^{sha1} and ^{sha256}Eric W. Biederman, Sep 8, 2023
  12. 06/32 repository: Implement core.compatMapEric W. Biederman, Sep 8, 2023
  13. 07/32 loose: add a mapping between SHA-1 and SHA-256 for loose objectsEric W. Biederman, Sep 8, 2023
  14. 19/32 object-file-convert: convert tag commits when writingEric W. Biederman, Sep 8, 2023
  15. 20/32 builtin/cat-file: Let the oid determine the output algorithmEric W. Biederman, Sep 8, 2023
  16. 22/32 object-file: Handle compat objects in check_object_signatureEric W. Biederman, Sep 8, 2023
  17. 26/32 object-file-convert: Implement convert_object_file_{begin,step,end}Eric W. Biederman, Sep 8, 2023
  18. Junio C HamanoSep 11, 2023
  19. 27/32 builtin/fast-import: compute compatibility hashs for imported objectsEric W. Biederman, Sep 8, 2023
  20. 29/32 builtin/index-pack: Compute the compatibility hashEric W. Biederman, Sep 8, 2023
  21. 31/32 unpack-objects: Update to compute and write the compatibility hashesEric W. Biederman, Sep 8, 2023
  22. 16/32 object: Factor out parse_mode out of fast-import and tree-walk into in object.hEric W. Biederman, Sep 8, 2023
  23. 10/32 bulk-checkin: Only accept blobsEric W. Biederman, Sep 8, 2023
  24. 23/32 builtin/ls-tree: Let the oid determine the output algorithmEric W. Biederman, Sep 8, 2023
  25. 12/32 bulk-checkin: hash object with compatibility algorithmEric W. Biederman, Sep 8, 2023
  26. Junio C HamanoSep 11, 2023
  27. 14/32 commit: write commits for both hashesEric W. Biederman, Sep 8, 2023
  28. Junio C HamanoSep 11, 2023
  29. 03/32 object-file-convert: Stubs for converting from one object format to anotherEric W. Biederman, Sep 8, 2023
  30. 08/32 loose: Compatibilty short name supportEric W. Biederman, Sep 8, 2023
  31. 01/32 doc hash-file-transition: A map file for mapping between sha1 and sha256Eric W. Biederman, Sep 8, 2023
  32. brian m. carlsonSep 10, 2023
  33. Eric W. BiedermanSep 10, 2023
  34. brian m. carlsonSep 12, 2023
  35. Eric W. BiedermanSep 12, 2023
  36. 15/32 cache: add a function to read an OID of a specific algorithmEric W. Biederman, Sep 8, 2023
  37. 32/32 object-file-convert: Implement repo_submodule_oid_to_algopEric W. Biederman, Sep 8, 2023
  38. 30/32 builtin/index-pack: Make the stack in compute_compat_oid explicitEric W. Biederman, Sep 8, 2023
  39. 28/32 builtin/index-pack: Add a simple oid indexEric W. Biederman, Sep 8, 2023
  40. 25/32 pack-compat-map: Add support for .compat files of a packfileEric W. Biederman, Sep 8, 2023
  41. Junio C HamanoSep 11, 2023
  42. Taylor BlauOct 5, 2023
  43. 21/32 tree-walk: init_tree_desc take an oid to get the hash algorithmEric W. Biederman, Sep 8, 2023
  44. 24/32 builtin/pack-objects: Communicate the compatibility hash through struct pack_idx_entryEric W. Biederman, Sep 8, 2023
  45. 18/32 object-file-convert: convert commit objects when writingEric W. Biederman, Sep 8, 2023
  46. 17/32 object-file-convert: add a function to convert trees between algorithmsEric W. Biederman, Sep 8, 2023
  47. 09/32 object-file: Update the loose object map when writing loose objectsEric W. Biederman, Sep 8, 2023
  48. 11/32 pack: Communicate the compat_oid through struct pack_idx_entryEric W. Biederman, Sep 8, 2023
  49. 05/32 repository: add a compatibility hash algorithmEric W. Biederman, Sep 8, 2023
  50. 13/32 object-file: Add a compat_oid_in parameter to write_object_file_flagsEric W. Biederman, Sep 8, 2023
  51. Eric W. BiedermanSep 9, 2023
  52. brian m. carlsonSep 10, 2023
  53. Eric W. BiedermanSep 10, 2023
  54. Junio C HamanoSep 11, 2023
  55. Eric W. BiedermanSep 11, 2023
  56. brian m. carlsonSep 11, 2023
  57. Eric W. BiedermanSep 12, 2023
  58. Junio C HamanoSep 12, 2023
  59. Eric W. BiedermanSep 14, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.