git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Fwd: Git and Large Binaries: A Proposed Solution

From
Jeff King <peff@peff.net>
Date
Mar 10, 2011, 22:24 UTC
Message-ID
<20110310222443.GC15828@sigill.intra.peff.net>
In-Reply-To
<4D793C7D.1000502@miseler.de>
On Thu, Mar 10, 2011 at 10:02:53PM +0100, Alexander Miseler wrote:
Show 22 quoted lines
> I've been debating whether to resurrect this thread, but since it has
> been referenced by the SoC2011Ideas wiki article I will just go ahead.
> I've spent a few hours trying to make this work to make git with big
> files usable under Windows.
> 
> > Just a quick aside.  Since (a2b665d, 2011-01-05) you can provide
> > the filename as an argument to the filter script:
> > 
> >     git config --global filter.huge.clean huge-clean %f
> > 
> > then use it in place:
> > 
> >     $ cat >huge-clean 
> >     #!/bin/sh
> >     f="$1"
> >     echo orig file is "$f" >&2
> >     sha1=`sha1sum "$f" | cut -d' ' -f1`
> >     cp "$f" /tmp/big_storage/$sha1
> >     rm -f "$f"
> >     echo $sha1
> > 
> > 		-- Pete

After thinking about this strategy more (the "convert big binary files into a hash via clean/smudge filter" strategy), it feels like a hack. That is, I don't see any reason that git can't give you the equivalent behavior without having to resort to bolted-on scripts.

For example, with this strategy you are giving up meaningful diffs in favor of just showing a diff of the hashes. But git _already_ can do this for binary diffs. The problem is that git unnecessarily uses a bunch of memory to come up with that answer because of assumptions in the diff code. So we should be fixing those assumptions. Any place that this smudge/clean filter solution could avoid looking at the blobs, we should be able to do the same inside git.

Of course that leaves the storage question; Scott's git-media script has pluggable storage that is backed by http, s3, or whatever. But again, that is a feature that might be worth putting into git (even if it is just a pluggable script at the object-db level).

-Peff
Previous: Alexander MiselerNext: Eric Montellese
Message 13 of 20 in “Fwd: Git and Large Binaries: A Proposed Solution”
  1. Eric MontelleseJan 21, 2011
  2. Wesley J. LandakerJan 21, 2011
  3. Eric MontelleseJan 21, 2011
  4. Jeff KingJan 21, 2011
  5. Eric MontelleseJan 21, 2011
  6. Sverre RabbelierJan 22, 2011
  7. Pete WyckoffJan 23, 2011
  8. Scott ChaconJan 26, 2011
  9. Eric MontelleseJan 26, 2011
  10. Joey HessJan 26, 2011
  11. Jakub NarebskiJan 26, 2011
  12. Alexander MiselerMar 10, 2011
  13. Jeff KingMar 10, 2011
  14. Eric MontelleseMar 13, 2011
  15. Jeff KingMar 13, 2011
  16. Alexander MiselerMar 13, 2011
  17. Jeff KingMar 14, 2011
  18. Eric MontelleseMar 16, 2011
  19. Nguyen Thai Ngoc DuyMar 16, 2011
  20. Joey HessJan 22, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.