git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Fwd: Git and Large Binaries: A Proposed Solution

From
AMAlexander Miseler <alexander@miseler.de>
Date
Mar 10, 2011, 21:02 UTC
Message-ID
<4D793C7D.1000502@miseler.de>
In-Reply-To
<20110123141417.GA6133@mew.padd.com>

I've been debating whether to resurrect this thread, but since it has been referenced by the SoC2011Ideas wiki article I will just go ahead. I've spent a few hours trying to make this work to make git with big files usable under Windows.

Show 17 quoted lines
> Just a quick aside.  Since (a2b665d, 2011-01-05) you can provide
> the filename as an argument to the filter script:
> 
>     git config --global filter.huge.clean huge-clean %f
> 
> then use it in place:
> 
>     $ cat >huge-clean 
>     #!/bin/sh
>     f="$1"
>     echo orig file is "$f" >&2
>     sha1=`sha1sum "$f" | cut -d' ' -f1`
>     cp "$f" /tmp/big_storage/$sha1
>     rm -f "$f"
>     echo $sha1
> 
> 		-- Pete
First off, the commit mentioned here is no help at all. This commit changes nothing about the input and output of filters. The file is still loaded completely into memory, still streamed to the filter via stdin, still streamed from the filter via stdout into yet another memory buffer. The two of which, IIRC, exist simultaneous for at least some time, thus doubling the memory requirements. This change only additionally provides the file name to the filter and nothing else. If one carefully rereads the commit message this apparently was the intention.
After this I started digging into the git source code. To change the filter input would be extremely trivial. However, the function that returns the filter output in a memory buffer is called from 8 places (all details from wetware memory and therefore unreliable). Most, maybe all, of the callers just dump the buffer into a file, which could easily be relocated into the filter calling function itself. But two callers detached the buffer from the strbuf and kept it beyond writing the file. I didn't track it any further since I decided to rather spend my time on improving big file handling in git itself, rather than targeting a workaround. Though of course a completely big-file-ready git should also provide a sane way to feed big files to and from filters.
If the two detached buffers are no complication this might be a trivial project. If they do it might become demanding though.
Previous: Jakub NarebskiNext: Jeff King
Message 12 of 20 in “Fwd: Git and Large Binaries: A Proposed Solution”
  1. Eric MontelleseJan 21, 2011
  2. Wesley J. LandakerJan 21, 2011
  3. Eric MontelleseJan 21, 2011
  4. Jeff KingJan 21, 2011
  5. Eric MontelleseJan 21, 2011
  6. Sverre RabbelierJan 22, 2011
  7. Pete WyckoffJan 23, 2011
  8. Scott ChaconJan 26, 2011
  9. Eric MontelleseJan 26, 2011
  10. Joey HessJan 26, 2011
  11. Jakub NarebskiJan 26, 2011
  12. Alexander MiselerMar 10, 2011
  13. Jeff KingMar 10, 2011
  14. Eric MontelleseMar 13, 2011
  15. Jeff KingMar 13, 2011
  16. Alexander MiselerMar 13, 2011
  17. Jeff KingMar 14, 2011
  18. Eric MontelleseMar 16, 2011
  19. Nguyen Thai Ngoc DuyMar 16, 2011
  20. Joey HessJan 22, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.