git/list[1] front-page[2] threads[3] people[4] search[5] about
 

The git repo format

From
BOBrian O'Mahoney <omb@khandalf.com>
Date
Apr 27, 2005, 19:47 UTC
Message-ID
<426FEC38.1060507@khandalf.com>
In-Reply-To
<Pine.LNX.4.58.0504271154470.18901@ppc970.osdl.org>

In understanding how to work with 'git' I had a number of initial difficulties which are mostly covered by the e-mail from Linus below.

Most of these are already covered in the README:

for objects, ie blob, commit, tag, tree: inflate, then <type>\s<size>\0<data>

where <data> is in the form, described by Linus below

when you look at them closely, all the formats are simple, un-ambiguous, and very easy to parse.

The index is also easy to parse, but there is a detail, after the 3-int header the records are padded to a multiple of 8 bytes. The detail is in cache.h.

Maybe the README needs to re-inforce this.
Brian
> I repeat: git does not do any free-form parsin AT ALL.
========================================================
 The links are in well-defined places, and you do not ever search for them.
And that's really very very important.
Show 64 quoted lines
> For a "commit", the format is
> 
>  - first line is exactly 46 bytes: five bytes of "tree ", 40 bytes of hex 
>    sha1, and one byte of "\n".
> 
>    NOTHING ELSE. Not extra spaces at the end, not extra spaces at the 
>    beginning or the middle. It's ASCII, but it's not free-format ASCII.
> 
>  - the next <n> (where 'n' can be 0 or more) lines are _exactly_ 48 bytes
>    each:  seven bytes of "parent ", 40 bytes of hex sha1, and one byte of 
>    "\n".
> 
>    NOTHING ELSE.
> 
>  - the next lines are "author " and "committer ". They have well-defined 
>    delimters for their fields, and no sha1's. The fields cannot contain 
>    '<', '>' or newlines, since those are the field/line delimeters.
> 
> There is no free-format text _anywhere_ that git parses. No room for 
> guesses, no room for mistakes, no room for anything half-way questionable.
> 
> And fsck actually enforces this. We do _not_ just use "gets()" to read one 
> line at a time. We literally verify that the lines are 46/48 bytes long, 
> and have the delimeters in the expected places.
> 
> Same goes for "tree" and "tag" objects. They all have fixed-format stuff. 
> A "tree" entry is always
> 
> 	"%o <space> %s" \0 [ 20 bytes of sha1 ]
> 
> with "%o" being "mode", and "%s" being "path". We don't guess. 
> 
> And this really is _important_. Exactly because we name things by the SHA1
> hash of the contents, we MUST NOT have flexible formats. Having a format
> which allows non-canonical representations (extra spaces etc) would mean
> that two trees that were identical would depend on how you happened to
> format them.
> 
> So there's really two issues:
>  - we don't guess or parse contents. We have strict rules, and that makes 
>    git more reliable. There are no gray areas. There's "right" and there 
>    is "wrong", and the right one works, and the wrong one gets flagged as 
>    being wrong and the tools refuse to touch it.
>  - there is only _one_ right way to do things, and that means that the 
>    the content is well-defined, and thus the SHA1 of the content is 
>    well-defined.
> 
> For example, another rule is that a "tree" object is always sorted by 
> the bytes in the filename (not by entry, btw: a directory called "foo" 
> will sort as "foo/", even though the _entry_ only shows "foo"). That rule 
> not only makes a lot of operations faster, but again, it means that there 
> is only _one_ way to represent a tree validly.
> 
> IOW, you _cannot_ represent a tree any other way (and I've been too lazy
> to check this in fsck, but it's alway sbeen my plan), and that is exactly 
> why we can just compare the hashes of the results - because there is no 
> random component of "layout" in the contents.
> 
> This really is important. It means that if you get to the same two tree
> contents in totally unrelated ways (you unpack a tar-file and encode it in
> git, or you have 5 years of git history and check it out), the "tree" will
> match _exactly_. There's no history. There's no "optional" stuff. Since
> the contents of the trees are the same, the SHA1 of the two trees will be
> the same. Exactly because git refuses to touch any free-format stuff.
Previous: Linus TorvaldsNext: H. Peter Anvin
Message 11 of 29 in “A shortcoming of the git repo format”
  1. H. Peter AnvinApr 27, 2005
  2. C. Scott AnanianApr 27, 2005
  3. Linus TorvaldsApr 27, 2005
  4. H. Peter AnvinApr 27, 2005
  5. Dave JonesApr 27, 2005
  6. H. Peter AnvinApr 27, 2005
  7. Jon SeymourApr 27, 2005
  8. Linus TorvaldsApr 27, 2005
  9. Petr BaudisApr 27, 2005
  10. Linus TorvaldsApr 27, 2005
  11. The git repo formatBrian O'Mahoney, Apr 27, 2005
  12. H. Peter AnvinApr 27, 2005
  13. Tom LordApr 27, 2005
  14. H. Peter AnvinApr 27, 2005
  15. Linus TorvaldsApr 28, 2005
  16. Paul JacksonApr 28, 2005
  17. Tom LordApr 28, 2005
  18. Ryan AndersonApr 28, 2005
  19. Morgan SchweersApr 28, 2005
  20. Barry SilvermanApr 28, 2005
  21. Linus TorvaldsApr 27, 2005
  22. David A. WheelerApr 28, 2005
  23. David LangApr 28, 2005
  24. Daniel BarkalowApr 27, 2005
  25. H. Peter AnvinApr 27, 2005
  26. Daniel BarkalowApr 28, 2005
  27. H. Peter AnvinApr 28, 2005
  28. David WoodhouseApr 28, 2005
  29. Gerhard SchrenkApr 27, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.