git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: A shortcoming of the git repo format

From
Linus Torvalds <torvalds@osdl.org>
Date
Apr 27, 2005, 19:11 UTC
Message-ID
<Pine.LNX.4.58.0504271154470.18901@ppc970.osdl.org>
In-Reply-To
<426FD3EE.5000404@zytor.com>
On Wed, 27 Apr 2005, H. Peter Anvin wrote:
> 
> That's true for email addresses, but the point was to distinguish links 
> to other git objects from any other kind of text.
No, that's definitely _not_ the point.

I repeat: git does not do any free-form parsin AT ALL. The links are in well-defined places, and you do not ever search for them. And that's really very very important.

> Currently there is no  such delimiter for that.
There absolutely is.
For a "commit", the format is
 - first line is exactly 46 bytes: five bytes of "tree ", 40 bytes of hex 
   sha1, and one byte of "\n".
   NOTHING ELSE. Not extra spaces at the end, not extra spaces at the 
   beginning or the middle. It's ASCII, but it's not free-format ASCII.
 - the next <n> (where 'n' can be 0 or more) lines are _exactly_ 48 bytes
   each:  seven bytes of "parent ", 40 bytes of hex sha1, and one byte of 
   "\n".
   NOTHING ELSE.
 - the next lines are "author " and "committer ". They have well-defined 
   delimters for their fields, and no sha1's. The fields cannot contain 
   '<', '>' or newlines, since those are the field/line delimeters.

There is no free-format text _anywhere_ that git parses. No room for guesses, no room for mistakes, no room for anything half-way questionable.

And fsck actually enforces this. We do _not_ just use "gets()" to read one line at a time. We literally verify that the lines are 46/48 bytes long, and have the delimeters in the expected places.

Same goes for "tree" and "tag" objects. They all have fixed-format stuff. A "tree" entry is always

	"%o <space> %s" \0 [ 20 bytes of sha1 ]
with "%o" being "mode", and "%s" being "path". We don't guess. 

And this really is _important_. Exactly because we name things by the SHA1 hash of the contents, we MUST NOT have flexible formats. Having a format which allows non-canonical representations (extra spaces etc) would mean that two trees that were identical would depend on how you happened to format them.

So there's really two issues:
 - we don't guess or parse contents. We have strict rules, and that makes 
   git more reliable. There are no gray areas. There's "right" and there 
   is "wrong", and the right one works, and the wrong one gets flagged as 
   being wrong and the tools refuse to touch it.
 - there is only _one_ right way to do things, and that means that the 
   the content is well-defined, and thus the SHA1 of the content is 
   well-defined.

For example, another rule is that a "tree" object is always sorted by the bytes in the filename (not by entry, btw: a directory called "foo" will sort as "foo/", even though the _entry_ only shows "foo"). That rule not only makes a lot of operations faster, but again, it means that there is only _one_ way to represent a tree validly.

IOW, you _cannot_ represent a tree any other way (and I've been too lazy to check this in fsck, but it's alway sbeen my plan), and that is exactly why we can just compare the hashes of the results - because there is no random component of "layout" in the contents.

This really is important. It means that if you get to the same two tree contents in totally unrelated ways (you unpack a tar-file and encode it in git, or you have 5 years of git history and check it out), the "tree" will match _exactly_. There's no history. There's no "optional" stuff. Since the contents of the trees are the same, the SHA1 of the two trees will be the same. Exactly because git refuses to touch any free-format stuff.

		Linus
Previous: Petr BaudisNext: Brian O'Mahoney
Message 10 of 29 in “A shortcoming of the git repo format”
  1. H. Peter AnvinApr 27, 2005
  2. C. Scott AnanianApr 27, 2005
  3. Linus TorvaldsApr 27, 2005
  4. H. Peter AnvinApr 27, 2005
  5. Dave JonesApr 27, 2005
  6. H. Peter AnvinApr 27, 2005
  7. Jon SeymourApr 27, 2005
  8. Linus TorvaldsApr 27, 2005
  9. Petr BaudisApr 27, 2005
  10. Linus TorvaldsApr 27, 2005
  11. The git repo formatBrian O'Mahoney, Apr 27, 2005
  12. H. Peter AnvinApr 27, 2005
  13. Tom LordApr 27, 2005
  14. H. Peter AnvinApr 27, 2005
  15. Linus TorvaldsApr 28, 2005
  16. Paul JacksonApr 28, 2005
  17. Tom LordApr 28, 2005
  18. Ryan AndersonApr 28, 2005
  19. Morgan SchweersApr 28, 2005
  20. Barry SilvermanApr 28, 2005
  21. Linus TorvaldsApr 27, 2005
  22. David A. WheelerApr 28, 2005
  23. David LangApr 28, 2005
  24. Daniel BarkalowApr 27, 2005
  25. H. Peter AnvinApr 27, 2005
  26. Daniel BarkalowApr 28, 2005
  27. H. Peter AnvinApr 28, 2005
  28. David WoodhouseApr 28, 2005
  29. Gerhard SchrenkApr 27, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.