git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[RFC] Design of name-addressed data portion

From
Daniel Barkalow <barkalow@iabervon.org>
Date
Apr 24, 2005, 18:17 UTC
Message-ID
<Pine.LNX.4.21.0504241336250.30848-100000@iabervon.org>

I think it has gotten to be time to have a standard mechanism for name-addressed data in the .git directory. We currently have one agreed-upon item, HEAD, as well as a number of items in cogito: heads, tags, and, to a certain extent, remotes. (Even in core git, when we support tags, we'll want a mapping from tag names to tag objects, even if this mapping doesn't get transferred by push and pull operations; nobody's going to want to cut and paste a hash from email every time they want to refer to the tag, and fsck-cache could stand to know which tags you mean to have so that it can report the rest to git-prune-script)

It would be useful to have a bit more structure to the repository, such that there are a fixed number of paths that hold all of the information about the state of the repository, while the rest of the directory has information that is particular to a working directory's state (e.g., index).

I'd propose the following structure:
 objects/    the content-addressed repository portion
 references/ the name-addressed repository portion
   heads/    the heads that are being used out of this repository
     DEFAULT the head that people pulling this repository mean by default
     ...     other heads, by name, that fsck-cache should mark reachable
   tags/     the tags
     ...     files with the symbolic name of the tags, containing the hash
 info/       other per-repository information
   remotes   URLs of remote repositories
   complete  hashes that the repository contains all references from
   missing   hashes that the repository lacks but wants
   excluded  hashes that the repository doesn't want
 ...         other files are per .git directory, not shared on push/pull
 index       
 HEAD        symlink to the head that is the local default
 tracked     remote that this working directory tracks

All of the files in references/*/* contain hex for objects in the database, and are not synced between repositories in situ (but some sync operations will read some of them and write them under different names). fsck-cache would use as its reachability starting point $(cat references/*/*).

In info/ are, generically, other files that relate to operations which work on the repository rather than a working directory. Transfer programs would use and maintain this information.

I think we'd still eventually want some way of getting from a commit-id to any tags about it (I think git log would do well to mention any tags you have when it shows a commit), but I don't want to design this quite yet. It should also work for going from the real history to cached delta info, when we have comparison tools that are sufficiently smart, expensive, and intermediate-dependant to want to cache this.

	-Daniel
*This .sig left intentionally blank*
Next: Petr Baudis
Message 1 of 5 in “[RFC] Design of name-addressed data portion”
  1. Daniel BarkalowApr 24, 2005
  2. Petr BaudisApr 24, 2005
  3. Daniel BarkalowApr 24, 2005
  4. Fabian FranzApr 24, 2005
  5. Daniel BarkalowApr 24, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.