git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC]: Pack-file object format for individual objects (Was: Revisiting large binary files issue.)

From
Linus Torvalds <torvalds@osdl.org>
Date
Jul 11, 2006, 18:00 UTC
Message-ID
<Pine.LNX.4.64.0607111053270.5623@g5.osdl.org>
In-Reply-To
<44B371FB.2070800@b-i-t.de>
[ I read my personal mailbox first, so I didn't see this one until after I 
  had already written my version.. ]
On Tue, 11 Jul 2006, sf wrote:
> 
> I just stumbled over the same fact and asked myself why there are still
> two formats. Wouldn't it make more sense to use the pack-file object
> format for individual objects as well?
Yes, see the git list for a series of patches that try to do this.
> As it happens individual objects all start with nibble 7 (deflated with
> default _zlib_ window size of 32K) whereas in the pack-file object
> format nibble 7 indicates delta entries which never occur as individual
> files.

I didn't actually do it that way, but it would be better to make the "parse_ascii_sha1_header()" more strict, and only accept the old names.

Right now my patch-series could in theory accept something that is _not_ an ASCII header (eg it would be a binary header that just happened to have the format "x n\0", where "n" was a valid number).

> Step 1. When reading individual objects from disk check the first nibble
> and decode accordingly (see above).

Check more than that, but yes, this should be tightened up in my series.

> Step 2. When writing individual objects to disk write them in pack-file
> object format. Make that optional (config-file parameter, command line
> option etc.)?
Done.
> Step 3. Remove code for (old) individual object disk format.

Well, I'm not sure how necessary that even is. We actually do have to generate the old header regardless, if for no other reason than the fact that we generate the SHA1 names based on it (even if we then write a new-style dense binary header to disk and discard the ASCII header).

Having it there means that you can always just get a new version of git, and never worry about how old the archive you're working with is.

(And then doing a "git repack -a -d" will make any archive also work with an old-style git, since the pack-file format didn't change, and a "git repack" thus ends up always creating something that is readable by anybody, including old clients).

		Linus
Previous: sfNext: sf
Message 6 of 37 in “Revisiting large binary files issue.”
  1. Carl BaldwinJul 10, 2006
  2. Junio C HamanoJul 10, 2006
  3. Peter BaumannJul 11, 2006
  4. Linus TorvaldsJul 10, 2006
  5. [RFC]: Pack-file object format for individual objects (Was: Revisiting large binary files issue.)sf, Jul 11, 2006
  6. Linus TorvaldsJul 11, 2006
  7. sfJul 11, 2006
  8. Linus TorvaldsJul 11, 2006
  9. Linus TorvaldsJul 11, 2006
  10. Carl BaldwinJul 11, 2006
  11. Linus TorvaldsJul 11, 2006
  12. 1/3 Make the unpacked object header functions static to sha1_file.cLinus Torvalds, Jul 11, 2006
  13. 2/3 sha1_file: add the ability to parse objects in "pack file format"Linus Torvalds, Jul 11, 2006
  14. Johannes SchindelinJul 11, 2006
  15. Linus TorvaldsJul 11, 2006
  16. Johannes SchindelinJul 11, 2006
  17. Linus TorvaldsJul 11, 2006
  18. Johannes SchindelinJul 11, 2006
  19. Junio C HamanoJul 11, 2006
  20. sfJul 11, 2006
  21. Linus TorvaldsJul 11, 2006
  22. sfJul 11, 2006
  23. Junio C HamanoJul 11, 2006
  24. Linus TorvaldsJul 12, 2006
  25. Johannes SchindelinJul 12, 2006
  26. Linus TorvaldsJul 12, 2006
  27. Linus TorvaldsJul 12, 2006
  28. Junio C HamanoJul 12, 2006
  29. Linus TorvaldsJul 12, 2006
  30. Junio C HamanoJul 12, 2006
  31. Linus TorvaldsJul 12, 2006
  32. Peter BaumannJul 12, 2006
  33. Junio C HamanoJul 12, 2006
  34. Peter BaumannJul 12, 2006
  35. Linus TorvaldsJul 12, 2006
  36. Junio C HamanoJul 12, 2006
  37. 3/3 Enable the new binary header format for unpacked objectsLinus Torvalds, Jul 11, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.