git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Fwd: Git and Large Binaries: A Proposed Solution

From
AMAlexander Miseler <alexander@miseler.de>
Date
Mar 13, 2011, 19:33 UTC
Message-ID
<4D7D1BFE.2030008@miseler.de>
In-Reply-To
<20110313025258.GA10452@sigill.intra.peff.net>
My thoughts on big file storage:
We want to store them as flat as possible. Ideally if we have a temp file with the content (e.g. the output of some filter) it should be possible to store it by simply doing a move/rename and updating some meta data external to the actual file.
Options:
1.) The loose file format is inherently unsuited for this. It has a header before the actual content and the whole file (header + content) is always compressed. Even if one changes this to compressing/decompressing header and content independently it is still unsuited by a) having the header within the same file and b) because the header has no flags or other means to indicate a different behavior (e.g. no compression) for the content. We could extend the header format or introduce a new object type (e.g. flatblob) but both would probably cause more trouble than other solutions. Another idea would be to keep the metadata in an external file (e.g. 84d7.header for the object 84d7). This would probably have a bad performance though since every object lookup would first need to check for the e
 xistence of a header file. A smarter variant would be to optionally keep the meta data directly in the filename (e.g. saving the object as 84d7.object_type.size.flag instead of just 84d7). 
This would only require special handling for cases where the normal lookup for 84d7 fails.

2.) The pack format fares a lot better. Content and meta data are already separated with the meta data describing how the content is stored. We would need a flag to mark the content as flat and that would pretty much be it. We would still need to include a virtual header when calculating the sha1 so it is guaranteed that the same content has always the same id. Thus i think we should simply forgo the loose object phase when storing big files and simply drop each big file flat as a individual pack file, with the idx file describing it as a pack file with one entry which is stored flat.

3.) Do some completely different handling for big files, as suggested by Eric:
>>   1.1 Perhaps a "binaries" directory, or structure of directories, within .git
> 
> I'd rather not do something so drastic.
My main issue with this approach (apart from the 'drastic' ^_^) is that the definition of big file may change at any time by e.g. changing a config value like core.bigFileThreshold. What has been stored as big file may suddenly be considered a normal blob and vice versa. Thus any storage variant that isn't well integrated in the normal object storage will probably be troublesome.
> There may also be code-paths for binary files where
> we accidentally load them (I just fixed one last week where we
> unnecessarily loaded them in the diffstat code path). Somebody will need
> to do some experimenting to shake out those code paths.
This is my main focus for now. They are easy to detect when your memory is small enough :D
Previous: Jeff KingNext: Jeff King
Message 16 of 20 in “Fwd: Git and Large Binaries: A Proposed Solution”
  1. Eric MontelleseJan 21, 2011
  2. Wesley J. LandakerJan 21, 2011
  3. Eric MontelleseJan 21, 2011
  4. Jeff KingJan 21, 2011
  5. Eric MontelleseJan 21, 2011
  6. Sverre RabbelierJan 22, 2011
  7. Pete WyckoffJan 23, 2011
  8. Scott ChaconJan 26, 2011
  9. Eric MontelleseJan 26, 2011
  10. Joey HessJan 26, 2011
  11. Jakub NarebskiJan 26, 2011
  12. Alexander MiselerMar 10, 2011
  13. Jeff KingMar 10, 2011
  14. Eric MontelleseMar 13, 2011
  15. Jeff KingMar 13, 2011
  16. Alexander MiselerMar 13, 2011
  17. Jeff KingMar 14, 2011
  18. Eric MontelleseMar 16, 2011
  19. Nguyen Thai Ngoc DuyMar 16, 2011
  20. Joey HessJan 22, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.