git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Efficiently storing SHA-1 ↔ SHA-256 mappings in compatibility mode

From
brian m. carlson <sandals@crustytoothpaste.net>
Date
Aug 28, 2025, 21:43 UTC
Message-ID
<aLDNj5GPYA9nR3xR@fruit.crustytoothpaste.net>
In-Reply-To
<20250827190817.M36986@dcvr>
On 2025-08-27 at 19:08:16, Eric Wong wrote:
Show 32 quoted lines
> "brian m. carlson" <sandals@crustytoothpaste.net> wrote:
> > TL;DR: We need a different datastore than a flat file for storing
> > mappings between SHA-1 and SHA-256 in compatibility mode.  Advice and
> > opinions sought.
> 
> <snip>
> 
> > Our approach for mapping object IDs between algorithms uses data in pack
> > index v3 (outlined in the transition document), plus a flat file called
> > `loose-object-idx` for loose objects.  However, we didn't anticipate
> > that we'd need to handle mappings long-term for data that is neither a
> > loose object nor a packed object.
> > 
> > For instance, with shallow clones, we must store a mapping for the
> > shallows the server has sent us[1], since we lack the history to convert
> > objects otherwise.  Similarly, if there are submodules or we're using a
> > partial clone, we must store those mappings as well, since we cannot
> > convert trees without them.  We can store them in the
> > `loose-object-idx`, but since it's not sorted or easily searchable, it's
> > going to perform really terribly when we store enough of them.  Right
> > now, we read the entire file into two hashmaps (one in each direction)
> > and we sometimes need to re-read it when other processes add items, so
> > it won't take much to make it be slow and take a lot of memory.
> 
> This really seems ideal for SQLite, which has come a long way
> since 2005 when git started.
> 
> I really wish git would've relied on more on existing formats
> (e.g. LMDB refs) rather than introducing more one-off data
> formats that require more cognitive overhead to document and
> learn[1], especially when SQLite is extremely portable and works
> on tiny devices.

SQLite is not an option because it performs poorly with Java and we want our formats to work with other implementations, like JGit. That's why we created reftable instead of using SQLite.

Also, in general, I'm not interested in being tied to a single implementation. If the developers of SQLite decide to dramatically change the license of all their code like Oracle did with Berkeley DB, we're going to have a problem. Yes, we can use the older versions, but we'd still need people to maintain the library and update it.

-- 
brian m. carlson (they/them)
Toronto, Ontario, CA
Previous: Junio C HamanoNext: Eric Wong
Message 9 of 10 in “Efficiently storing SHA-1 ↔ SHA-256 mappings in compatibility mode”
  1. brian m. carlsonAug 14, 2025
  2. Junio C HamanoAug 14, 2025
  3. brian m. carlsonAug 14, 2025
  4. Junio C HamanoAug 14, 2025
  5. Derrick StoleeAug 15, 2025
  6. Patrick SteinhardtSep 3, 2025
  7. Eric WongAug 27, 2025
  8. Junio C HamanoAug 28, 2025
  9. brian m. carlsonAug 28, 2025
  10. Eric WongAug 29, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.