git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Fwd: Git and Large Binaries: A Proposed Solution

From
Joey Hess <joey@kitenet.net>
Date
Jan 22, 2011, 00:07 UTC
Message-ID
<20110122000712.GA7931@gnu.kitenet.net>
In-Reply-To
<AANLkTimPua_kz2w33BRPeTtOEWOKDCsJzf0sqxm=db68@mail.gmail.com>

Hi, I wrote git-annex, and pristine-tar, and etckeeper. I enjoy making git do things that I'm told it shouldn't be used for. :) I should have probably talked more about git-annex here, before.

Eric Montellese wrote:
Show 7 quoted lines
> 2. zipped tarballs of source code (that I will never need to modify)
> -- I could unpack these and then use git to track the source code.
> However, I prefer to track these deliverables as the tarballs
> themselves because it makes my customer happier to see the exact
> tarball that they delivered being used when I repackage updates.
> (Let's not discuss problems with this model - I understand that this
> is non-ideal).

In this specific case, you can use pristine-tar to recreate the original, exact tarballs from unpacked source files that you check into git. It accomplishes this without the overhead of duplicating compressed data in tarballs. I feel in this case, this is a better approach than generic large file support, since it stores all the data in git, just in a much more compressed form, and so fits in nicely with standard git-based source code management.

> The short version:
> ***Don't track binaries in git.  Track their hashes.***

That was my principle with git-annex. Although slightly generalized to: "Don't track large file contents in git. Track unique keys that an arbitrary backend can use to obtain the file contents."

Now, you mention in a followup that git-annex does not default to keeping a local copy of every binary referenced by a file in master. This is true, for the simple reason that a copy of every file in some of my git repos master would sum to multiple terabytes of data. :) I think that practically, anything that supports large files in git needs to support partial checkouts too.

But, git-annex can be run in eg, a post-merge hook, and asked to retrieve all current file contents, and drop outdated contents.

Show 10 quoted lines
> First the layout:
> my_git_project/binrepo/
> -- binaries/
> -- hashes/
> -- symlink_to_hashes_file
> -- symlink_to_another_hashes_file
> within the "binrepo" (binary repository) there is a subdirectory for
> binaries, and a subdirectory for hashes.  In the root of the 'binrepo'
> all of the files stored have a symlink to the current version of the
> hash.

Very similar to git-annex in the use of versioned symlinks here. It stores the binaries in .git/annex/objects to avoid needing to gitignore them.

> 3. In my setup, all of the binary files are in a single "binrepo"
> directory.  If done from within git, we would need a non-kludgey way
> to allow large binaries to exist anywhere within the git tree.

git-annex allows the symlinks to be mixed with regular git managed content throughout the repository. (This means that when symlinks are moved, they may need to be fixed, which is done at commit time.)

> 5. Command to purge all binaries in your "binrepo" that are not needed
> for the current revision (if you're running out of disk space
> locally).

Safely dropping data is really one of the complexities of this approach. Git-annex stores location tracking information in git, so it can know where it can retrieve file data *from*. I chose to make it very cautious about removing data, as location tracking data can fall out of date (if for example, a remote had the data, had dropped it, and has not pushed that information out). So it actively confirms that enough other copies of the data currently exist before dropping it. (Of course, these checks can be disabled.)

> 6. Automatically upload new versions of files to the "binrepo" (rather
> than needing to do this manually)

In git-annex, data transfer is done using rsync, so that interrupted transfers of large files can be resumed. I recently added a git-annex-shell to support locked-down access, similar to git-shell.

BTW, I have been meaning to look into using smudge filters with git-annex. I'm a bit worried about some of the potential overhead associated with smudge filters, and I'm not sure how a partial checkout would work with them.

-- 
see shy jo
Previous: Nguyen Thai Ngoc Duy
Message 20 of 20 in “Fwd: Git and Large Binaries: A Proposed Solution”
  1. Eric MontelleseJan 21, 2011
  2. Wesley J. LandakerJan 21, 2011
  3. Eric MontelleseJan 21, 2011
  4. Jeff KingJan 21, 2011
  5. Eric MontelleseJan 21, 2011
  6. Sverre RabbelierJan 22, 2011
  7. Pete WyckoffJan 23, 2011
  8. Scott ChaconJan 26, 2011
  9. Eric MontelleseJan 26, 2011
  10. Joey HessJan 26, 2011
  11. Jakub NarebskiJan 26, 2011
  12. Alexander MiselerMar 10, 2011
  13. Jeff KingMar 10, 2011
  14. Eric MontelleseMar 13, 2011
  15. Jeff KingMar 13, 2011
  16. Alexander MiselerMar 13, 2011
  17. Jeff KingMar 14, 2011
  18. Eric MontelleseMar 16, 2011
  19. Nguyen Thai Ngoc DuyMar 16, 2011
  20. Joey HessJan 22, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.