git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Including object type and size in object id (Re: Git Merge contributor summit notes)

From
JHJeff Hostetler <git@jeffhostetler.com>
Date
Mar 26, 2018, 21:42 UTC
Message-ID
<b7b6d617-1951-5934-5b1d-bb1a300006ef@jeffhostetler.com>
In-Reply-To
<20180326210039.GB21735@aiede.svl.corp.google.com>
On 3/26/2018 5:00 PM, Jonathan Nieder wrote:
Show 24 quoted lines
> Jeff Hostetler wrote:
> [long quote snipped]
> 
>> While we are converting to a new hash function, it would be nice
>> if we could add a couple of fields to the end of the OID:  the object
>> type and the raw uncompressed object size.
>>
>> If would be nice if we could extend the OID to include 6 bytes of data
>> (4 or 8 bits for the type and the rest for the raw object size), and
>> just say that an OID is a {hash,type,size} tuple.
>>
>> There are lots of places where we open an object to see what type it is
>> or how big it is.  This requires uncompressing/undeltafying the object
>> (or at least decoding enough to get the header).  In the case of missing
>> objects (partial clone or a gvfs-like projection) it requires either
>> dynamically fetching the object or asking an object-size-server for the
>> data.
>>
>> All of these cases could be eliminated if the type/size were available
>> in the OID.
> 
> This implies a limit on the object size (e.g. 5 bytes in your
> example).  What happens when someone wants to encode an object larger
> than that limit?

I could say add a full uint64 to the tail end of the hash, but we currently don't handle blobs/objects larger then 4GB right now anyway, right?

5 bytes for the size is just a compromise -- 1TB blobs would be
terrible to think about...
  
> 
> This also decreases the number of bits available for the hash, but
> that shouldn't be a big issue.

I was suggesting extending the OIDs by 6 bytes while we are changing the hash function.

Show 11 quoted lines
> Aside from those two, I don't see any downsides.  It would mean that
> tree objects contain information about the sizes of blobs contained
> there, which helps with virtual file systems.  It's also possible to
> do that without putting the size in the object id, but maybe having it
> in the object id is simpler.
> 
> Will think more about this.
> 
> Thanks for the idea,
> Jonathan
> 

Thanks Jeff

Previous: Jonathan NiederNext: Junio C Hamano
Message 14 of 17 in “Git Merge contributor summit notes”
  1. Alex VandiverMar 10, 2018
  2. Ævar Arnfjörð BjarmasonMar 10, 2018
  3. Junio C HamanoMar 11, 2018
  4. Jeff KingMar 12, 2018
  5. Brandon WilliamsMar 13, 2018
  6. Jeff KingMar 12, 2018
  7. Ævar Arnfjörð BjarmasonMar 25, 2018
  8. Jeff HostetlerMar 26, 2018
  9. Stefan BellerMar 26, 2018
  10. Jeff HostetlerMar 26, 2018
  11. Brandon WilliamsMar 26, 2018
  12. Jakub NarebskiApr 7, 2018
  13. Including object type and size in object id (Re: Git Merge contributor summit notes)Jonathan Nieder, Mar 26, 2018
  14. Jeff HostetlerMar 26, 2018
  15. Junio C HamanoMar 26, 2018
  16. Per-object encryption (Re: Git Merge contributor summit notes)Jonathan Nieder, Mar 26, 2018
  17. Ævar Arnfjörð BjarmasonMar 26, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.