{"thread":{"id":"26320","subject":"Fwd: Git and Large Binaries: A Proposed Solution","startedAt":"2011-01-21T18:57:21Z","lastAt":"2011-03-16T14:40:51Z","messageCount":20,"participants":["Eric Montellese","Wesley J. Landaker","Jeff King","Joey Hess","Sverre Rabbelier","Pete Wyckoff","Scott Chacon","Jakub Narebski","Alexander Miseler","Nguyen Thai Ngoc Duy"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"159766","messageId":"AANLkTimPua_kz2w33BRPeTtOEWOKDCsJzf0sqxm=db68@mail.gmail.com","threadId":"26320","inReplyTo":"AANLkTin=UySutWLS0Y7OmuvkE=T=+YB8G8aUCxLH=GKa@mail.gmail.com","subject":"Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Eric Montellese","fromEmail":"emontellese@gmail.com","sentAt":"2011-01-21T18:57:21Z","receivedAt":"2011-01-21T18:57:21Z","isPatch":false,"sender":{"key":"emontellese@gmail.com","avatar":"https://gravatar.com/avatar/398aa9e755dc5d92d45ddbeab4e50c4d77608c41f6ff22afa11e24687ce3d209?d=mp&s=160"},"body":"I did a search for this issue before posting, but if I am touching on\nan old topic that already has a solution in progress, I apologize.  As\nfar as I know, this is still an open issue (last i saw in the\nkernel-trap archives was a \"git and binary files\" thread from Jan\n2008, and there are a couple of promising related works (git annex and\ngit bigfiles) -- but nothing that solves the complete problem).\n\nI'm interested in hearing your thoughts and suggestions, and\ninterested if there is community interest in adding this feature to\ngit.  I would be happy to be involved in making the changes, but I\nhave very limited time, so would prefer help and would like to know\nthat it has a strong chance of joining the main line before\nstarting...\nTo whet your appetite to read all of the below (I know it's long),\nthis is the root of the solution:\n\n---       Don't track binaries in git.  Track their hashes.       ---\n\n\nProblem Background:\nI work on embedded system software, with code for these products\ndelivered from multiple customers and in multiple formats, for\nexample:\n\n1. source code -- works great and is what git is designed for\n2. zipped tarballs of source code (that I will never need to modify)\n-- I could unpack these and then use git to track the source code.\nHowever, I prefer to track these deliverables as the tarballs\nthemselves because it makes my customer happier to see the exact\ntarball that they delivered being used when I repackage updates.\n(Let's not discuss problems with this model - I understand that this\nis non-ideal).\n3. large and/or many binaries.  (could be pictures, short videos,\npre-compiled binaries, etc)\n\nThe problem of course is that, as you know, git is not ideal for\nhandling large, or many, binaries.  It's better at just about\neverything else, of course, but not this, for largely these two\nreasons:\n\n1. git cannot diff the binaries to their previous iterations\neffectively to save space.  (and neither can any tool)\n2. git requires that all clones of that repository must therefore\ndownload all versions of all binaries -- which, if the binaries are\nlarge, or many, and are poorly compressed together (as stated in 1),\nthis will be a very expensive operation.\n\n\nProblem Statement:\nWe (the git user) want and \"need\" to be able to track large binaries\nfrom within our repository.  But putting them into git slows down git\nunnecessarily.\nThe only current alternative is to *not* check the large binaries into\ngit -- but now they are no longer tracked, which is unacceptable.  If\nI want to jump back in git to a point in the tree from 6 months ago, I\ndo not have any way to tell which version of the large binaries I\nneed.  I could keep track of this manually, of course, but that's what\ngit is for...\n\n\nSolution:\nThe short version:\n***Don't track binaries in git.  Track their hashes.***\n\nSolution:\nThe long version:\nFor my current project, I have this (the \"store the hashes\" idea)\nimplemented outside of git.  I am posting to this list because I would\nlike to see this functionality (well, something even better) become\nnative to git, and believe that it would remove one of the few\nremaining arguments that some projects have against adopting git.\nHere is how I have it implemented:\n\nFirst the layout:\nmy_git_project/binrepo/\n-- binaries/\n-- hashes/\n-- symlink_to_hashes_file\n-- symlink_to_another_hashes_file\nwithin the \"binrepo\" (binary repository) there is a subdirectory for\nbinaries, and a subdirectory for hashes.  In the root of the 'binrepo'\nall of the files stored have a symlink to the current version of the\nhash.\nThe \"binaries\" directory is .gitignore'd -- the hashes directory and\nthe symlinks to the current hashes are maintained by git.\nWhenever I receive a new version of a large binary file from a\ncustomer, I put it into \"binaries\" and I create a new hash for that\nfile in \"hashes\" and update the symlink to point to that hash.  I 'git\ncommit' and 'git push' those changes (this is fast since there is no\nlarge binary in the git repository).\nThe other important factor is that I must put this large binary file\nsomewhere accessible for others to download it.  In this example, it\nis:  my_git_server.net:/binrepo/\n\nThen I have a bash script (some psuedocode here to save space):\n\nfor (all BINFILE in binrepo) ; do\n  HASHFILE=$BINFILE\".md5\"\n  # check if the binary exists\n  if [[ -e binaries/$BINFILE ]] ; then\n    echo \"  $BINFILE available\"\n  else\n    echo \"  $BINFILE not available. Downloading...\"\n    wget http://my_git_server.net:/binrepo/$BINFILE\n  fi\n  # check md5sum\n  md5sum $BINFILE > temp.md5\n  if ! diff -q ../hashes/$HASHFILE temp.md5 >/dev/null ; then\n    echo \"ERROR! $BINFILE md5 does not match!\"\n    exit and/or redownload\n  fi\ndone\n\nThis confirms that I have the right version of all of the binaries --\nmy git repository is effectively tracking the large binaries, but\nwithout actually storing them internally to the git repo.  If someone\nelse updates the \"binrepo\" I will know it when I do a \"git pull\" and I\nwill automatically get the right version of the binary file so that my\nsandbox is up-to-date.  Now let's say I want to revert the version of\nthe large binary file to the previous version -- all I need to do is\nto to edit the symlink in \"binrepo\", commit, and push.  Other users\nwill automatically use the old version of the file as well after they\ndo their pull (and without needing to re-download that file)\n\n\nSummary of Big Advantages:\n\n1. Repository is unpolluted by large binary files.  git clone stays fast.\n2. User has access to any version of any binary file, but does not\nneed to store every version locally if they do not want to.\n3. Git does not need to worry about the big binaries - there are no\nslow attempts to calculate binary deltas or pack and unpack under the\nhood.\n\n\nImprovements:\n\nI imagine these features (among others):\n\n1. In my current setup, each large binary file has a different name (a\nrevision number).  This could be easily solved, however, by generating\nunique names under the hood and tracking this within git.\n2. A lot of the steps in my current setup are manual.  When I want to\nadd a new binary file, I need to manually create the hash and manually\nupload the binary to the joint server.  If done within git, this would\nbe automatic.\n3. In my setup, all of the binary files are in a single \"binrepo\"\ndirectory.  If done from within git, we would need a non-kludgey way\nto allow large binaries to exist anywhere within the git tree.  If git\nhandles the \"binrepo\" under the hood though, the user would never need\nto know about it -- instead git would just handle all binaries by\nchecking the internal \"binrepo\"  Instead of tracking symlinks, git\nwould track the file versions in the normal way -- it just wouldn't\nstore the binaries the same way (instead it would store the hash)\n4. User option to download all versions of all binaries, or only the\nversion necessary for the position on the current branch.  If you want\nto be able to run all versions of the repository when offline, you can\ndownload all versions of all binaries.  If you don't need to do this,\nyou can just download the versions you need.  Or perhaps have the\noption to download all binaries smaller than X-bytes, but skip the big\nones.\n5. Command to purge all binaries in your \"binrepo\" that are not needed\nfor the current revision (if you're running out of disk space\nlocally).\n6. Automatically upload new versions of files to the \"binrepo\" (rather\nthan needing to do this manually)\n\n\nRock on!\nEric\n"},{"id":"159771","messageId":"201101211436.32033.wjl@icecavern.net","threadId":"26320","inReplyTo":"AANLkTimPua_kz2w33BRPeTtOEWOKDCsJzf0sqxm=db68@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Wesley J. Landaker","fromEmail":"wjl@icecavern.net","sentAt":"2011-01-21T21:36:31Z","receivedAt":"2011-01-21T21:36:31Z","isPatch":false,"sender":{"key":"wjl@icecavern.net","avatar":"https://avatars.githubusercontent.com/u/67229?v=4"},"body":"On Friday, January 21, 2011 11:57:21 Eric Montellese wrote:\n> To whet your appetite to read all of the below (I know it's long),\n> this is the root of the solution:\n> \n> ---       Don't track binaries in git.  Track their hashes.       ---\n\nComment from the peanut gallery:\n\nI haven't read your approach in great detail, but just in case you are not \naware, there is a project call git-annex <http://git-annex.branchable.com/> \nby Joey Hess that I believe takes a similar approach.\n\nSince you've obviously given this a lot of thought, you might want to take a \npeek at that and see if it already does what you want, or if your proposal \ndoes something significantly different/better.\n"},{"id":"159772","messageId":"AANLkTikj-+bVW6P42Ejz+T=CtKTwUKBF6BmEDSBeZbL6@mail.gmail.com","threadId":"26320","inReplyTo":"201101211436.32033.wjl@icecavern.net","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Eric Montellese","fromEmail":"emontellese@gmail.com","sentAt":"2011-01-21T22:00:47Z","receivedAt":"2011-01-21T22:00:47Z","isPatch":false,"sender":{"key":"emontellese@gmail.com","avatar":"https://gravatar.com/avatar/398aa9e755dc5d92d45ddbeab4e50c4d77608c41f6ff22afa11e24687ce3d209?d=mp&s=160"},"body":"Thanks Wesely,\n\nI did take a look at git annex -- it looks to me as though that\nproject is more of a special-case, allowing users to use git to track\nthings like music and movies.  While it's possible this might be\nusable for the use case I described, what I'm really looking for is a\ntrue extension of git which allows binaries to be treated differently\n(if the user desires) when using git as a source management tool.\n\nThe major difference I see with git-annex is that the user must\nspecifically tell git-annex to download certain files.  Instead, I\nwant the user to always automatically have *all* of the files (both\nsource and binaries) for the current revision -- but not necessarily\nfor hundreds (thousands?) of past revisions (which as git is\nimplemented currently would take up many gigabytes)\n\ngit-annex does look like a neat piece of software, but I don't think\nit quite fits here -- thank you again for the comment though!\n\nEric\n\n\nOn Fri, Jan 21, 2011 at 4:36 PM, Wesley J. Landaker <wjl@icecavern.net> wrote:\n> On Friday, January 21, 2011 11:57:21 Eric Montellese wrote:\n>> To whet your appetite to read all of the below (I know it's long),\n>> this is the root of the solution:\n>>\n>> ---       Don't track binaries in git.  Track their hashes.       ---\n>\n> Comment from the peanut gallery:\n>\n> I haven't read your approach in great detail, but just in case you are not\n> aware, there is a project call git-annex <http://git-annex.branchable.com/>\n> by Joey Hess that I believe takes a similar approach.\n>\n> Since you've obviously given this a lot of thought, you might want to take a\n> peek at that and see if it already does what you want, or if your proposal\n> does something significantly different/better.\n>\n"},{"id":"159776","messageId":"20110121222440.GA1837@sigill.intra.peff.net","threadId":"26320","inReplyTo":"AANLkTimPua_kz2w33BRPeTtOEWOKDCsJzf0sqxm=db68@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-21T22:24:42Z","receivedAt":"2011-01-21T22:24:42Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 21, 2011 at 01:57:21PM -0500, Eric Montellese wrote:\n\n> I did a search for this issue before posting, but if I am touching on\n> an old topic that already has a solution in progress, I apologize.  As\n\nIt's been talked about a lot, but there is not exactly a solution in\nprogress. One promising direction is not very different from what you're\ndoing, though:\n\n> Solution:\n> The short version:\n> ***Don't track binaries in git.  Track their hashes.***\n\nYes, exactly. But what your solution lacks, I think, is more integration\ninto git. Specifically, using clean/smudge filters you can have git take\ncare of tracking the file contents automatically.\n\nAt the very simplest, it would look like:\n\n-- >8 --\ncat >$HOME/local/bin/huge-clean <<'EOF'\n#!/bin/sh\n\n# In an ideal world, we could actually\n# access the original file directly instead of\n# having to cat it to a new file.\ntemp=\"$(git rev-parse --git-dir)\"/huge.$$\ncat >\"$temp\"\nsha1=`sha1sum \"$temp\" | cut -d' ' -f1`\n\n# now move it to wherever your permanent storage is\n# scp \"$root/$sha1\" host:/path/to/big_storage/$sha1\ncp \"$temp\" /tmp/big_storage/$sha1\nrm -f \"$temp\"\n\necho $sha1\nEOF\n\ncat >$HOME/local/bin/huge-smudge <<'EOF'\n#!/bin/sh\n\n# Get sha1 from stored blob via stdin\nread sha1\n\n# Now retrieve blob. We could optionally do some caching here.\n# ssh host cat /path/to/big/storage/$sha1\ncat /tmp/big_storage/$sha1\nEOF\n-- 8< --\n\nObviously our storage mechanism (throwing things in /tmp) is simplistic,\nbut obviously you could store and retrieve via ssh, http, s3, or\nwhatever.\n\nYou can try it out like this:\n\n  # set up our filter config and fake storage area\n  mkdir /tmp/big_storage\n  git config --global filter.huge.clean huge-clean\n  git config --global filter.huge.smudge huge-smudge\n\n  # now make a repo, and make sure we mark *.bin files as huge\n  mkdir repo && cd repo && git init\n  echo '*.bin filter=huge' >.gitattributes\n  git add .gitattributes\n  git commit -m 'add attributes'\n\n  # let's do a moderate 20M file\n  perl -e 'print \"foo\\n\" for (1 .. 5000000)' >foo.bin\n  git add foo.bin\n  git commit -m 'add huge file (foo)'\n\n  # and then another revision\n  perl -e 'print \"bar\\n\" for (1 .. 5000000)' >foo.bin\n  git commit -a -m 'revise huge file (bar)'\n\nNotice that we just add and commit as normal.  And we can check that the\nspace usage is what you expect:\n\n  $ du -sh repo/.git\n  196K    repo/.git\n  $ du -sh /tmp/big_storage\n  39M     /tmp/big_storage\n\nDiffs obviously are going to be less interesting, as we just see the\nhash:\n\n  $ git log --oneline -p foo.bin\n  39e549c revise huge file (bar)\n  diff --git a/foo.bin b/foo.bin\n  index 281fd03..70874bd 100644\n  --- a/foo.bin\n  +++ b/foo.bin\n  @@ -1 +1 @@\n  -50a1ee265f4562721346566701fce1d06f54dd9e\n  +bbc2f7f191ad398fe3fcb57d885e1feacb4eae4e\n  845836e add huge file (foo)\n  diff --git a/foo.bin b/foo.bin\n  new file mode 100644\n  index 0000000..281fd03\n  --- /dev/null\n  +++ b/foo.bin\n  @@ -0,0 +1 @@\n  +50a1ee265f4562721346566701fce1d06f54dd9e\n\nbut if you wanted to, you could write a custom diff driver that does\nsomething more meaningful with your particular binary format (it would\nhave to grab from big_storage, though).\n\nChecking out other revisions works without extra action:\n\n  $ head -n 1 foo.bin\n  bar\n  $ git checkout HEAD^\n  HEAD is now at 845836e... add huge file (foo)\n  $ head -n 1 foo.bin\n  foo\n\nAnd since you have the filter config in your ~/.gitconfig, clones will\njust work:\n\n  $ git clone repo other\n  $ du -sh other/.git\n  204K    other/.git\n  $ du -sh other/foo.bin\n  20M\n\nSo conceptually it's pretty similar to yours, but the filter integration\nmeans that git takes care of putting the right files in place at the\nright time.\n\nIt would probably benefit a lot from caching the large binary files\ninstead of hitting big_storage all the time. And probably the\nputting/getting from storage should be factored out so you can plug in\ndifferent storage. And it should all be configurable. Different users of\nthe same repo might want different caching policies, or to access the\nbinary assets by different mechanisms or URLs.\n\n> I imagine these features (among others):\n> \n> 1. In my current setup, each large binary file has a different name (a\n> revision number).  This could be easily solved, however, by generating\n> unique names under the hood and tracking this within git.\n\nIn the scheme above, we just index by their hash. So you can easily fsck\nyour big_storage by making sure everything matches its hash (but you\ncan't know that you have _all_ of the blobs needed unless you\ncross-reference with the history).\n\n> 2. A lot of the steps in my current setup are manual.  When I want to\n> add a new binary file, I need to manually create the hash and manually\n> upload the binary to the joint server.  If done within git, this would\n> be automatic.\n\nI think the scheme above takes care of the manual bits.\n\n> 3. In my setup, all of the binary files are in a single \"binrepo\"\n> directory.  If done from within git, we would need a non-kludgey way\n> to allow large binaries to exist anywhere within the git tree.  If git\n\nAny scheme, whether it uses clean/smudge filters or not, should probably\ntie in via gitattributes.\n\n> 4. User option to download all versions of all binaries, or only the\n> version necessary for the position on the current branch.  If you want\n> to be able to run all versions of the repository when offline, you can\n> download all versions of all binaries.  If you don't need to do this,\n> you can just download the versions you need.  Or perhaps have the\n> option to download all binaries smaller than X-bytes, but skip the big\n> ones.\n\nThe scheme above will download on an as-needed basis. If caching were\nimplemented, you could just make the cache infinitely big and do a \"git\nlog -p\" which would download everything. :)\n\nProbably you would also want the smudge filter to return \"blob not\navailable\" when operating in some kind of offline mode.\n\n> 5. Command to purge all binaries in your \"binrepo\" that are not needed\n> for the current revision (if you're running out of disk space\n> locally).\n\nIn my scheme, just rm your cache directory (once it exists).\n\n> 6. Automatically upload new versions of files to the \"binrepo\" (rather\n> than needing to do this manually)\n\nHandled by the clean filter above.\n\n\nSo obviously this is not very complete. And there are a few changes to\ngit that could make it more efficient (e.g., letting the clean filter\ntouch the file directly instead of having to make a copy via stdin). But\nthe general idea is there, and it just needs somebody to make a nice\npolished script that is configurable, does caching, etc. I'll get to it\neventually, but if you'd like to work on it, be my guest.\n\n-Peff\n"},{"id":"159778","messageId":"AANLkTinKNtDDy6Pi4Tn+hpTrVw_DBoYpTn3ihCfN_fUd@mail.gmail.com","threadId":"26320","inReplyTo":"20110121222440.GA1837@sigill.intra.peff.net","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Eric Montellese","fromEmail":"emontellese@gmail.com","sentAt":"2011-01-21T23:15:37Z","receivedAt":"2011-01-21T23:15:37Z","isPatch":false,"sender":{"key":"emontellese@gmail.com","avatar":"https://gravatar.com/avatar/398aa9e755dc5d92d45ddbeab4e50c4d77608c41f6ff22afa11e24687ce3d209?d=mp&s=160"},"body":"Peff,\n\nThanks for your insight -- this looks great.\n\nOnce something like this is available and more polished, what's the\nprocess to request that it join the main line of git development?   (i\nknow functionally there's \"no main line\" in git... but you know what I\nmean)\n\nHas there already been discussion to this effect?  I do think that a\nfix like this would improve git adoption among certain groups.  (I\nknow I've heard the \"big binaries\" problem mentioned at least a few\ntimes)\n\n\nI haven't dug around in git code yet, so while I can get the gist of\nyour code, I'm unable to get the complete picture.  You wouldn't\nhappen to have a git patch, or a public repo somewhere that I can take\na look at?  Does there happen to be a git developers guide hidden away\nanywhere?  Though I have very limited time, I'd be happy to help out\nas much as I can.\n\n\nEric\n\n\n\n\n\n\n\nOn Fri, Jan 21, 2011 at 5:24 PM, Jeff King <peff@peff.net> wrote:\n> On Fri, Jan 21, 2011 at 01:57:21PM -0500, Eric Montellese wrote:\n>\n>> I did a search for this issue before posting, but if I am touching on\n>> an old topic that already has a solution in progress, I apologize.  As\n>\n> It's been talked about a lot, but there is not exactly a solution in\n> progress. One promising direction is not very different from what you're\n> doing, though:\n>\n>> Solution:\n>> The short version:\n>> ***Don't track binaries in git.  Track their hashes.***\n>\n> Yes, exactly. But what your solution lacks, I think, is more integration\n> into git. Specifically, using clean/smudge filters you can have git take\n> care of tracking the file contents automatically.\n>\n> At the very simplest, it would look like:\n>\n> -- >8 --\n> cat >$HOME/local/bin/huge-clean <<'EOF'\n> #!/bin/sh\n>\n> # In an ideal world, we could actually\n> # access the original file directly instead of\n> # having to cat it to a new file.\n> temp=\"$(git rev-parse --git-dir)\"/huge.$$\n> cat >\"$temp\"\n> sha1=`sha1sum \"$temp\" | cut -d' ' -f1`\n>\n> # now move it to wherever your permanent storage is\n> # scp \"$root/$sha1\" host:/path/to/big_storage/$sha1\n> cp \"$temp\" /tmp/big_storage/$sha1\n> rm -f \"$temp\"\n>\n> echo $sha1\n> EOF\n>\n> cat >$HOME/local/bin/huge-smudge <<'EOF'\n> #!/bin/sh\n>\n> # Get sha1 from stored blob via stdin\n> read sha1\n>\n> # Now retrieve blob. We could optionally do some caching here.\n> # ssh host cat /path/to/big/storage/$sha1\n> cat /tmp/big_storage/$sha1\n> EOF\n> -- 8< --\n>\n> Obviously our storage mechanism (throwing things in /tmp) is simplistic,\n> but obviously you could store and retrieve via ssh, http, s3, or\n> whatever.\n>\n> You can try it out like this:\n>\n>  # set up our filter config and fake storage area\n>  mkdir /tmp/big_storage\n>  git config --global filter.huge.clean huge-clean\n>  git config --global filter.huge.smudge huge-smudge\n>\n>  # now make a repo, and make sure we mark *.bin files as huge\n>  mkdir repo && cd repo && git init\n>  echo '*.bin filter=huge' >.gitattributes\n>  git add .gitattributes\n>  git commit -m 'add attributes'\n>\n>  # let's do a moderate 20M file\n>  perl -e 'print \"foo\\n\" for (1 .. 5000000)' >foo.bin\n>  git add foo.bin\n>  git commit -m 'add huge file (foo)'\n>\n>  # and then another revision\n>  perl -e 'print \"bar\\n\" for (1 .. 5000000)' >foo.bin\n>  git commit -a -m 'revise huge file (bar)'\n>\n> Notice that we just add and commit as normal.  And we can check that the\n> space usage is what you expect:\n>\n>  $ du -sh repo/.git\n>  196K    repo/.git\n>  $ du -sh /tmp/big_storage\n>  39M     /tmp/big_storage\n>\n> Diffs obviously are going to be less interesting, as we just see the\n> hash:\n>\n>  $ git log --oneline -p foo.bin\n>  39e549c revise huge file (bar)\n>  diff --git a/foo.bin b/foo.bin\n>  index 281fd03..70874bd 100644\n>  --- a/foo.bin\n>  +++ b/foo.bin\n>  @@ -1 +1 @@\n>  -50a1ee265f4562721346566701fce1d06f54dd9e\n>  +bbc2f7f191ad398fe3fcb57d885e1feacb4eae4e\n>  845836e add huge file (foo)\n>  diff --git a/foo.bin b/foo.bin\n>  new file mode 100644\n>  index 0000000..281fd03\n>  --- /dev/null\n>  +++ b/foo.bin\n>  @@ -0,0 +1 @@\n>  +50a1ee265f4562721346566701fce1d06f54dd9e\n>\n> but if you wanted to, you could write a custom diff driver that does\n> something more meaningful with your particular binary format (it would\n> have to grab from big_storage, though).\n>\n> Checking out other revisions works without extra action:\n>\n>  $ head -n 1 foo.bin\n>  bar\n>  $ git checkout HEAD^\n>  HEAD is now at 845836e... add huge file (foo)\n>  $ head -n 1 foo.bin\n>  foo\n>\n> And since you have the filter config in your ~/.gitconfig, clones will\n> just work:\n>\n>  $ git clone repo other\n>  $ du -sh other/.git\n>  204K    other/.git\n>  $ du -sh other/foo.bin\n>  20M\n>\n> So conceptually it's pretty similar to yours, but the filter integration\n> means that git takes care of putting the right files in place at the\n> right time.\n>\n> It would probably benefit a lot from caching the large binary files\n> instead of hitting big_storage all the time. And probably the\n> putting/getting from storage should be factored out so you can plug in\n> different storage. And it should all be configurable. Different users of\n> the same repo might want different caching policies, or to access the\n> binary assets by different mechanisms or URLs.\n>\n>> I imagine these features (among others):\n>>\n>> 1. In my current setup, each large binary file has a different name (a\n>> revision number).  This could be easily solved, however, by generating\n>> unique names under the hood and tracking this within git.\n>\n> In the scheme above, we just index by their hash. So you can easily fsck\n> your big_storage by making sure everything matches its hash (but you\n> can't know that you have _all_ of the blobs needed unless you\n> cross-reference with the history).\n>\n>> 2. A lot of the steps in my current setup are manual.  When I want to\n>> add a new binary file, I need to manually create the hash and manually\n>> upload the binary to the joint server.  If done within git, this would\n>> be automatic.\n>\n> I think the scheme above takes care of the manual bits.\n>\n>> 3. In my setup, all of the binary files are in a single \"binrepo\"\n>> directory.  If done from within git, we would need a non-kludgey way\n>> to allow large binaries to exist anywhere within the git tree.  If git\n>\n> Any scheme, whether it uses clean/smudge filters or not, should probably\n> tie in via gitattributes.\n>\n>> 4. User option to download all versions of all binaries, or only the\n>> version necessary for the position on the current branch.  If you want\n>> to be able to run all versions of the repository when offline, you can\n>> download all versions of all binaries.  If you don't need to do this,\n>> you can just download the versions you need.  Or perhaps have the\n>> option to download all binaries smaller than X-bytes, but skip the big\n>> ones.\n>\n> The scheme above will download on an as-needed basis. If caching were\n> implemented, you could just make the cache infinitely big and do a \"git\n> log -p\" which would download everything. :)\n>\n> Probably you would also want the smudge filter to return \"blob not\n> available\" when operating in some kind of offline mode.\n>\n>> 5. Command to purge all binaries in your \"binrepo\" that are not needed\n>> for the current revision (if you're running out of disk space\n>> locally).\n>\n> In my scheme, just rm your cache directory (once it exists).\n>\n>> 6. Automatically upload new versions of files to the \"binrepo\" (rather\n>> than needing to do this manually)\n>\n> Handled by the clean filter above.\n>\n>\n> So obviously this is not very complete. And there are a few changes to\n> git that could make it more efficient (e.g., letting the clean filter\n> touch the file directly instead of having to make a copy via stdin). But\n> the general idea is there, and it just needs somebody to make a nice\n> polished script that is configurable, does caching, etc. I'll get to it\n> eventually, but if you'd like to work on it, be my guest.\n>\n> -Peff\n>\n"},{"id":"159779","messageId":"20110122000712.GA7931@gnu.kitenet.net","threadId":"26320","inReplyTo":"AANLkTimPua_kz2w33BRPeTtOEWOKDCsJzf0sqxm=db68@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Joey Hess","fromEmail":"joey@kitenet.net","sentAt":"2011-01-22T00:07:12Z","receivedAt":"2011-01-22T00:07:12Z","isPatch":false,"sender":{"key":"joey@kitenet.net","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Hi, I wrote git-annex, and pristine-tar, and etckeeper. I enjoy making\ngit do things that I'm told it shouldn't be used for. :) I should have\nprobably talked more about git-annex here, before.\n\nEric Montellese wrote:\n> 2. zipped tarballs of source code (that I will never need to modify)\n> -- I could unpack these and then use git to track the source code.\n> However, I prefer to track these deliverables as the tarballs\n> themselves because it makes my customer happier to see the exact\n> tarball that they delivered being used when I repackage updates.\n> (Let's not discuss problems with this model - I understand that this\n> is non-ideal).\n\nIn this specific case, you can use pristine-tar to recreate the\noriginal, exact tarballs from unpacked source files that you check into\ngit. It accomplishes this without the overhead of duplicating compressed\ndata in tarballs. I feel in this case, this is a better approach than\ngeneric large file support, since it stores all the data in git, just in a\nmuch more compressed form, and so fits in nicely with standard git-based\nsource code management.\n\n> The short version:\n> ***Don't track binaries in git.  Track their hashes.***\n\nThat was my principle with git-annex. Although slightly generalized to:\n\"Don't track large file contents in git. Track unique keys that\nan arbitrary backend can use to obtain the file contents.\"\n\nNow, you mention in a followup that git-annex does not default to keeping\na local copy of every binary referenced by a file in master.\nThis is true, for the simple reason that a copy of every file in some of\nmy git repos master would sum to multiple terabytes of data. :) I think\nthat practically, anything that supports large files in git needs to\nsupport partial checkouts too.\n\nBut, git-annex can be run in eg, a post-merge hook, and asked to\nretrieve all current file contents, and drop outdated contents.\n\n> First the layout:\n> my_git_project/binrepo/\n> -- binaries/\n> -- hashes/\n> -- symlink_to_hashes_file\n> -- symlink_to_another_hashes_file\n> within the \"binrepo\" (binary repository) there is a subdirectory for\n> binaries, and a subdirectory for hashes.  In the root of the 'binrepo'\n> all of the files stored have a symlink to the current version of the\n> hash.\n\nVery similar to git-annex in the use of versioned symlinks here.\nIt stores the binaries in .git/annex/objects to avoid needing to\ngitignore them.\n\n> 3. In my setup, all of the binary files are in a single \"binrepo\"\n> directory.  If done from within git, we would need a non-kludgey way\n> to allow large binaries to exist anywhere within the git tree.\n\ngit-annex allows the symlinks to be mixed with regular git managed\ncontent throughout the repository. (This means that when symlinks\nare moved, they may need to be fixed, which is done at commit time.)\n\n> 5. Command to purge all binaries in your \"binrepo\" that are not needed\n> for the current revision (if you're running out of disk space\n> locally).\n\nSafely dropping data is really one of the complexities of this\napproach. Git-annex stores location tracking information in git,\nso it can know where it can retrieve file data *from*. I chose to make\nit very cautious about removing data, as location tracking data can \nfall out of date (if for example, a remote had the data, had dropped it,\nand has not pushed that information out). So it actively confirms that\nenough other copies of the data currently exist before dropping it.\n(Of course, these checks can be disabled.)\n\n> 6. Automatically upload new versions of files to the \"binrepo\" (rather\n> than needing to do this manually)\n\nIn git-annex, data transfer is done using rsync, so that interrupted\ntransfers of large files can be resumed. I recently added a git-annex-shell\nto support locked-down access, similar to git-shell.\n\n\nBTW, I have been meaning to look into using smudge filters with git-annex.\nI'm a bit worried about some of the potential overhead associated with\nsmudge filters, and I'm not sure how a partial checkout would work with\nthem.\n\n-- \nsee shy jo\n"},{"id":"159782","messageId":"AANLkTimurgnSFO=gR5Z-=GM27-QD00MCdXNc0x5Q-TQ4@mail.gmail.com","threadId":"26320","inReplyTo":"AANLkTinKNtDDy6Pi4Tn+hpTrVw_DBoYpTn3ihCfN_fUd@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-01-22T03:05:03Z","receivedAt":"2011-01-22T03:05:03Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\n[More or less separate from to the ongoing discussion, so no text quoted]\n\nEric, at the last GitTogether Avery presented his tool, bup, which\nimplements a number of solutions to the problem of large binary files.\nI think I remember that Jonathan is also interested in the topic.\nAvery, Jonathan, you can read up on the ongoing conversation at [0] if\nyou like :).\n\n[0] http://thread.gmane.org/gmane.comp.version-control.git/165389/focus=165401\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"159799","messageId":"20110123141417.GA6133@mew.padd.com","threadId":"26320","inReplyTo":"20110121222440.GA1837@sigill.intra.peff.net","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Pete Wyckoff","fromEmail":"pw@padd.com","sentAt":"2011-01-23T14:14:17Z","receivedAt":"2011-01-23T14:14:17Z","isPatch":false,"sender":{"key":"pw@padd.com","avatar":null},"body":"peff@peff.net wrote on Fri, 21 Jan 2011 17:24 -0500:\n> cat >$HOME/local/bin/huge-clean <<'EOF'\n> #!/bin/sh\n> \n> # In an ideal world, we could actually\n> # access the original file directly instead of\n> # having to cat it to a new file.\n> temp=\"$(git rev-parse --git-dir)\"/huge.$$\n> cat >\"$temp\"\n\nJust a quick aside.  Since (a2b665d, 2011-01-05) you can provide\nthe filename as an argument to the filter script:\n\n    git config --global filter.huge.clean huge-clean %f\n\nthen use it in place:\n\n    $ cat >huge-clean \n    #!/bin/sh\n    f=\"$1\"\n    echo orig file is \"$f\" >&2\n    sha1=`sha1sum \"$f\" | cut -d' ' -f1`\n    cp \"$f\" /tmp/big_storage/$sha1\n    rm -f \"$f\"\n    echo $sha1\n\n\t\t-- Pete\n"},{"id":"159870","messageId":"AANLkTimE+s81Xbj4snNX0WWxG8x=qSwaQWfK+08+1Zy+@mail.gmail.com","threadId":"26320","inReplyTo":"20110123141417.GA6133@mew.padd.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Scott Chacon","fromEmail":"schacon@gmail.com","sentAt":"2011-01-26T03:42:11Z","receivedAt":"2011-01-26T03:42:11Z","isPatch":false,"sender":{"key":"schacon@gmail.com","avatar":"https://gravatar.com/avatar/9b13a8a078e1dcf8588c4eea9554445d51ebed6c41b51f56f4d96738130b05c6?d=mp&s=160"},"body":"Hey,\n\nSorry to come in a bit late to this, but in addition to git-annex, I\nwrote something called 'git-media' a long time ago that works in a\nsimilar manner to what you both are discussing.\n\nMuch like what peff was talking about, it uses the smudge and clean\nfilters to automatically redirect content into a .git/media directory\ninstead of into Git itself while keeping the SHA in Git.  One of the\ncool thing is that it can use S3, scp or a local directory to transfer\nthe big files to and from.\n\nCheck it out if interested:\n\nhttps://github.com/schacon/git-media\n\nOn Sun, Jan 23, 2011 at 6:14 AM, Pete Wyckoff <pw@padd.com> wrote:\n> peff@peff.net wrote on Fri, 21 Jan 2011 17:24 -0500:\n>\n> Just a quick aside.  Since (a2b665d, 2011-01-05) you can provide\n> the filename as an argument to the filter script:\n>\n>    git config --global filter.huge.clean huge-clean %f\n>\n\nThis is amazing.  I absolutely did not know you could do this, and it\nwould make parts of git-media way better if I re-implemented it using\nthis.  Thanks for pointing this out.\n\nScott\n"},{"id":"159893","messageId":"AANLkTim8h7RGk59b7jsKjBEdEKaC77S7Nm2NYAE5R3i2@mail.gmail.com","threadId":"26320","inReplyTo":"AANLkTimE+s81Xbj4snNX0WWxG8x=qSwaQWfK+08+1Zy+@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Eric Montellese","fromEmail":"emontellese@gmail.com","sentAt":"2011-01-26T16:23:45Z","receivedAt":"2011-01-26T16:23:45Z","isPatch":false,"sender":{"key":"emontellese@gmail.com","avatar":"https://gravatar.com/avatar/398aa9e755dc5d92d45ddbeab4e50c4d77608c41f6ff22afa11e24687ce3d209?d=mp&s=160"},"body":"Good stuff!\n\nSo, it seems like there are at least a few decent ways to work around\nthe git binaries problem -- but my question is, will something like\nthis become part of mainline git?  (and how is such a decision made\nand by whom?)\n\nIt does seem that there is a real need for a solution like this, and a\nlot of the core code to handle it has already been written (perhaps\neven by multiple folks); it's just in need (as Peff said) of a\npolished and configurable script.  If one were to exist, would it\nbecome part of mainline git?\n\nThanks,\nEric\n\n\n\n\n\n\nOn Tue, Jan 25, 2011 at 10:42 PM, Scott Chacon <schacon@gmail.com> wrote:\n> Hey,\n>\n> Sorry to come in a bit late to this, but in addition to git-annex, I\n> wrote something called 'git-media' a long time ago that works in a\n> similar manner to what you both are discussing.\n>\n> Much like what peff was talking about, it uses the smudge and clean\n> filters to automatically redirect content into a .git/media directory\n> instead of into Git itself while keeping the SHA in Git.  One of the\n> cool thing is that it can use S3, scp or a local directory to transfer\n> the big files to and from.\n>\n> Check it out if interested:\n>\n> https://github.com/schacon/git-media\n>\n> On Sun, Jan 23, 2011 at 6:14 AM, Pete Wyckoff <pw@padd.com> wrote:\n>> peff@peff.net wrote on Fri, 21 Jan 2011 17:24 -0500:\n>>\n>> Just a quick aside.  Since (a2b665d, 2011-01-05) you can provide\n>> the filename as an argument to the filter script:\n>>\n>>    git config --global filter.huge.clean huge-clean %f\n>>\n>\n> This is amazing.  I absolutely did not know you could do this, and it\n> would make parts of git-media way better if I re-implemented it using\n> this.  Thanks for pointing this out.\n>\n> Scott\n>\n"},{"id":"159895","messageId":"20110126174226.GA14511@gnu.kitenet.net","threadId":"26320","inReplyTo":"AANLkTimE+s81Xbj4snNX0WWxG8x=qSwaQWfK+08+1Zy+@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Joey Hess","fromEmail":"joey@kitenet.net","sentAt":"2011-01-26T17:42:26Z","receivedAt":"2011-01-26T17:42:26Z","isPatch":false,"sender":{"key":"joey@kitenet.net","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Scott Chacon wrote:\n> Sorry to come in a bit late to this, but in addition to git-annex, I\n> wrote something called 'git-media' a long time ago that works in a\n> similar manner to what you both are discussing.\n\nHuh, if I had known about that I might not have written git-annex.\nAlthough probably I would have still, I needed a more distributed\napproach to storing data than git-media seems to support, and the\nability to partially check out only some files.\n\n> >    git config --global filter.huge.clean huge-clean %f\n> \n> This is amazing.  I absolutely did not know you could do this, and it\n> would make parts of git-media way better if I re-implemented it using\n> this.  Thanks for pointing this out.\n\nYeah, that's great, it should allow a smart clean filter to not waste\ntime reprocessing a known large file on every git status/git commit.\n\n-- \nsee shy jo\n"},{"id":"159923","messageId":"m3y6679wfy.fsf@localhost.localdomain","threadId":"26320","inReplyTo":"AANLkTimE+s81Xbj4snNX0WWxG8x=qSwaQWfK+08+1Zy+@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-01-26T21:40:09Z","receivedAt":"2011-01-26T21:40:09Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Scott Chacon <schacon@gmail.com> writes:\n\n> Sorry to come in a bit late to this, but in addition to git-annex, I\n> wrote something called 'git-media' a long time ago that works in a\n> similar manner to what you both are discussing.\n> \n> Much like what peff was talking about, it uses the smudge and clean\n> filters to automatically redirect content into a .git/media directory\n> instead of into Git itself while keeping the SHA in Git.  One of the\n> cool thing is that it can use S3, scp or a local directory to transfer\n> the big files to and from.\n> \n> Check it out if interested:\n> \n> https://github.com/schacon/git-media\n\nCould you please add short information about this project to the\nhttps://git.wiki.kernel.org/index.php/InterfacesFrontendsAndTools\npage, in the \"Backups, metadata, and large files\" subsection?\ngit-annex is there...\n\nThanks in advance.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"163173","messageId":"4D793C7D.1000502@miseler.de","threadId":"26320","inReplyTo":"20110123141417.GA6133@mew.padd.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Alexander Miseler","fromEmail":"alexander@miseler.de","sentAt":"2011-03-10T21:02:53Z","receivedAt":"2011-03-10T21:02:53Z","isPatch":false,"sender":{"key":"alexander@miseler.de","avatar":null},"body":"I've been debating whether to resurrect this thread, but since it has been referenced by the SoC2011Ideas wiki article I will just go ahead.\nI've spent a few hours trying to make this work to make git with big files usable under Windows.\n\n> Just a quick aside.  Since (a2b665d, 2011-01-05) you can provide\n> the filename as an argument to the filter script:\n> \n>     git config --global filter.huge.clean huge-clean %f\n> \n> then use it in place:\n> \n>     $ cat >huge-clean \n>     #!/bin/sh\n>     f=\"$1\"\n>     echo orig file is \"$f\" >&2\n>     sha1=`sha1sum \"$f\" | cut -d' ' -f1`\n>     cp \"$f\" /tmp/big_storage/$sha1\n>     rm -f \"$f\"\n>     echo $sha1\n> \n> \t\t-- Pete\n\nFirst off, the commit mentioned here is no help at all. This commit changes nothing about the input and output of filters. The file is still loaded completely into memory, still streamed to the filter via stdin, still streamed from the filter via stdout into yet another memory buffer. The two of which, IIRC, exist simultaneous for at least some time, thus doubling the memory requirements. This change only additionally provides the file name to the filter and nothing else. If one carefully rereads the commit message this apparently was the intention.\n\nAfter this I started digging into the git source code. To change the filter input would be extremely trivial. However, the function that returns the filter output in a memory buffer is called from 8 places (all details from wetware memory and therefore unreliable). Most, maybe all, of the callers just dump the buffer into a file, which could easily be relocated into the filter calling function itself. But two callers detached the buffer from the strbuf and kept it beyond writing the file. I didn't track it any further since I decided to rather spend my time on improving big file handling in git itself, rather than targeting a workaround. Though of course a completely big-file-ready git should also provide a sane way to feed big files to and from filters.\n\nIf the two detached buffers are no complication this might be a trivial project. If they do it might become demanding though.\n"},{"id":"163178","messageId":"20110310222443.GC15828@sigill.intra.peff.net","threadId":"26320","inReplyTo":"4D793C7D.1000502@miseler.de","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-03-10T22:24:43Z","receivedAt":"2011-03-10T22:24:43Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 10, 2011 at 10:02:53PM +0100, Alexander Miseler wrote:\n\n> I've been debating whether to resurrect this thread, but since it has\n> been referenced by the SoC2011Ideas wiki article I will just go ahead.\n> I've spent a few hours trying to make this work to make git with big\n> files usable under Windows.\n> \n> > Just a quick aside.  Since (a2b665d, 2011-01-05) you can provide\n> > the filename as an argument to the filter script:\n> > \n> >     git config --global filter.huge.clean huge-clean %f\n> > \n> > then use it in place:\n> > \n> >     $ cat >huge-clean \n> >     #!/bin/sh\n> >     f=\"$1\"\n> >     echo orig file is \"$f\" >&2\n> >     sha1=`sha1sum \"$f\" | cut -d' ' -f1`\n> >     cp \"$f\" /tmp/big_storage/$sha1\n> >     rm -f \"$f\"\n> >     echo $sha1\n> > \n> > \t\t-- Pete\n\nAfter thinking about this strategy more (the \"convert big binary files\ninto a hash via clean/smudge filter\" strategy), it feels like a hack.\nThat is, I don't see any reason that git can't give you the equivalent\nbehavior without having to resort to bolted-on scripts.\n\nFor example, with this strategy you are giving up meaningful diffs in\nfavor of just showing a diff of the hashes. But git _already_ can do\nthis for binary diffs.  The problem is that git unnecessarily uses a\nbunch of memory to come up with that answer because of assumptions in\nthe diff code. So we should be fixing those assumptions. Any place that\nthis smudge/clean filter solution could avoid looking at the blobs, we\nshould be able to do the same inside git.\n\nOf course that leaves the storage question; Scott's git-media script has\npluggable storage that is backed by http, s3, or whatever. But again,\nthat is a feature that might be worth putting into git (even if it is\njust a pluggable script at the object-db level).\n\n-Peff\n"},{"id":"163274","messageId":"AANLkTimpbhaGEfxW1wwRc14tpV6qnPDiZYnXp_tvA3Ft@mail.gmail.com","threadId":"26320","inReplyTo":"20110310222443.GC15828@sigill.intra.peff.net","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Eric Montellese","fromEmail":"emontellese@gmail.com","sentAt":"2011-03-13T01:53:53Z","receivedAt":"2011-03-13T01:53:53Z","isPatch":false,"sender":{"key":"emontellese@gmail.com","avatar":"https://gravatar.com/avatar/398aa9e755dc5d92d45ddbeab4e50c4d77608c41f6ff22afa11e24687ce3d209?d=mp&s=160"},"body":"This is a good point.\n\nThe best solution, it seems, has two parts:\n\n1. Clean up the way in which git considers, diffs, and stores binaries\nto cut down on the overhead of dealing with these files.\n  1.1 Perhaps a \"binaries\" directory, or structure of directories, within .git\n  1.2 Perhaps configurable options for when and how to try a binary\ndiff?  (allow user to decide if storage or speed is more important)\n2. Once (1) is accomplished, add an option to avoid copying binaries\nfrom all but the tip when doing a \"git clone.\"\n  2.1 The default behavior would be to copy everything, as users\ncurrently expect.\n  2.2 Core code would have hooks to allow a script to use a central\nlocation for the binary storage. (ssh, http, gmail-fs, whatever)\n\n(of course, the implementation of (1) should be friendly to the addition of (2))\n\nObviously, the major drawback to (2) without (2.2) is that if there is\ntruly distributed work going on, some clone-of-a-clone may not know\nwhere to get the binaries.\n\nBut, print a warning when turning on the non-default behavior (2.1),\nthen it's a user problem :-)\n\nEric\n\n\n\n\nOn Thu, Mar 10, 2011 at 5:24 PM, Jeff King <peff@peff.net> wrote:\n>\n> On Thu, Mar 10, 2011 at 10:02:53PM +0100, Alexander Miseler wrote:\n>\n> > I've been debating whether to resurrect this thread, but since it has\n> > been referenced by the SoC2011Ideas wiki article I will just go ahead.\n> > I've spent a few hours trying to make this work to make git with big\n> > files usable under Windows.\n> >\n> > > Just a quick aside.  Since (a2b665d, 2011-01-05) you can provide\n> > > the filename as an argument to the filter script:\n> > >\n> > >     git config --global filter.huge.clean huge-clean %f\n> > >\n> > > then use it in place:\n> > >\n> > >     $ cat >huge-clean\n> > >     #!/bin/sh\n> > >     f=\"$1\"\n> > >     echo orig file is \"$f\" >&2\n> > >     sha1=`sha1sum \"$f\" | cut -d' ' -f1`\n> > >     cp \"$f\" /tmp/big_storage/$sha1\n> > >     rm -f \"$f\"\n> > >     echo $sha1\n> > >\n> > >             -- Pete\n>\n> After thinking about this strategy more (the \"convert big binary files\n> into a hash via clean/smudge filter\" strategy), it feels like a hack.\n> That is, I don't see any reason that git can't give you the equivalent\n> behavior without having to resort to bolted-on scripts.\n>\n> For example, with this strategy you are giving up meaningful diffs in\n> favor of just showing a diff of the hashes. But git _already_ can do\n> this for binary diffs.  The problem is that git unnecessarily uses a\n> bunch of memory to come up with that answer because of assumptions in\n> the diff code. So we should be fixing those assumptions. Any place that\n> this smudge/clean filter solution could avoid looking at the blobs, we\n> should be able to do the same inside git.\n>\n> Of course that leaves the storage question; Scott's git-media script has\n> pluggable storage that is backed by http, s3, or whatever. But again,\n> that is a feature that might be worth putting into git (even if it is\n> just a pluggable script at the object-db level).\n>\n> -Peff\n"},{"id":"163275","messageId":"20110313025258.GA10452@sigill.intra.peff.net","threadId":"26320","inReplyTo":"AANLkTimpbhaGEfxW1wwRc14tpV6qnPDiZYnXp_tvA3Ft@mail.gmail.com","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-03-13T02:52:58Z","receivedAt":"2011-03-13T02:52:58Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Mar 12, 2011 at 08:53:53PM -0500, Eric Montellese wrote:\n\n> The best solution, it seems, has two parts:\n> \n> 1. Clean up the way in which git considers, diffs, and stores binaries\n> to cut down on the overhead of dealing with these files.\n\nThis is the easier half, I think.\n\n>   1.1 Perhaps a \"binaries\" directory, or structure of directories, within .git\n\nI'd rather not do something so drastic. We already have ways of marking\nfiles as binary and un-diffable within the tree. So you can already do\npretty well with marking them with gitattributes. I think we can do\nbetter by making them the binaryness auto-detection less expensive\n(right now we pull in the whole blob to check the first 1K or so for\nNULs or other patterns; this is fine in the common text case, where\nwe'll want the whole blob in a minute anyway, but for large files it's\nobviously wasteful). There may also be code-paths for binary files where\nwe accidentally load them (I just fixed one last week where we\nunnecessarily loaded them in the diffstat code path). Somebody will need\nto do some experimenting to shake out those code paths.\n\nFor packing, we have core.bigFileThreshold to turn off delta compression\nfor large files, but according to the documentation, it is only honored\nfor fast-import. I think we would want something similar to say \"for\nsome subset of files (indicated either by name or by minimum size),\ndon't bother with zlib-compression either, and always keep them loose\".\n\nThose are the two major ones, I think. There are probably a handful of\nother cases (like git-add, which really should be able to have a fixed\nmemory size). Again, the first step is figuring out where all of the\nproblems are (and I'm happy to just fix them one by one as they come up,\nbut I am also thinking of this in terms of a GSoC project).\n\n>   1.2 Perhaps configurable options for when and how to try a binary\n> diff?  (allow user to decide if storage or speed is more important)\n\nWe can already do that with gitattributes. But it would be nice to have\nit be fast in the binary auto-detection case.\n\n> 2. Once (1) is accomplished, add an option to avoid copying binaries\n> from all but the tip when doing a \"git clone.\"\n\nThis is much harder. :)\n\n>   2.1 The default behavior would be to copy everything, as users\n> currently expect.\n>   2.2 Core code would have hooks to allow a script to use a central\n> location for the binary storage. (ssh, http, gmail-fs, whatever)\n\nI think we would need a protocol extension for the fetching client to\nsay \"please don't bother sending me anything larger than N bytes; I will\nget it via alternate storage\". Although there are situations more\ncomplicated than that. Your alternate storage might have up to commit X,\nand you don't want large objects in X or its ancestors. But you _do_\nwant large objects in descendants of X, since you have no other way to\nget them.\n\nSo you need some way of saying which sets of large objects you need and\nwhich you don't. One implementation is that you could fetch from\nalternate storage (which would then need to be not just large-blob\nstorage, but actually have a full repo), and then afterwards fetch from\nthe remote (which would then send you all binaries, because by\ndefinition anything you are fetching is not something the alternate\nstorage has). That feels a bit hack-ish. Doing something more clever\nwould require a pretty major protocol extension, though.\n\nI haven't been paying attention to any sparse clone proposals. I know it\nhas come up but I don't know how mature the idea is. But this is\npotentially related.\n\n-Peff\n"},{"id":"163295","messageId":"4D7D1BFE.2030008@miseler.de","threadId":"26320","inReplyTo":"20110313025258.GA10452@sigill.intra.peff.net","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Alexander Miseler","fromEmail":"alexander@miseler.de","sentAt":"2011-03-13T19:33:18Z","receivedAt":"2011-03-13T19:33:18Z","isPatch":false,"sender":{"key":"alexander@miseler.de","avatar":null},"body":"\nMy thoughts on big file storage:\n\nWe want to store them as flat as possible. Ideally if we have a temp file with the content (e.g. the output of some filter) it should be possible to store it by simply doing a move/rename and updating some meta data external to the actual file.\n\nOptions:\n\n1.) The loose file format is inherently unsuited for this. It has a header before the actual content and the whole file (header + content) is always compressed. Even if one changes this to compressing/decompressing header and content independently it is still unsuited by a) having the header within the same file and b) because the header has no flags or other means to indicate a different behavior (e.g. no compression) for the content. We could extend the header format or introduce a new object type (e.g. flatblob) but both would probably cause more trouble than other solutions. Another idea would be to keep the metadata in an external file (e.g. 84d7.header for the object 84d7). This would probably have a bad performance though since every object lookup would first need to check for the e\n xistence of a header file. A smarter variant would be to optionally keep the meta data directly in the filename (e.g. saving the object as 84d7.object_type.size.flag instead of just 84d7). \nThis would only require special handling for cases where the normal lookup for 84d7 fails.\n\n2.) The pack format fares a lot better. Content and meta data are already separated with the meta data describing how the content is stored. We would need a flag to mark the content as flat and that would pretty much be it. We would still need to include a virtual header when calculating the sha1 so it is guaranteed that the same content has always the same id.\nThus i think we should simply forgo the loose object phase when storing big files and simply drop each big file flat as a individual pack file, with the idx file describing it as a pack file with one entry which is stored flat.\n\n3.) Do some completely different handling for big files, as suggested by Eric:\n>>   1.1 Perhaps a \"binaries\" directory, or structure of directories, within .git\n> \n> I'd rather not do something so drastic.\nMy main issue with this approach (apart from the 'drastic' ^_^) is that the definition of big file may change at any time by e.g. changing a config value like core.bigFileThreshold. What has been stored as big file may suddenly be considered a normal blob and vice versa. Thus any storage variant that isn't well integrated in the normal object storage will probably be troublesome.\n\n\n\n\n> There may also be code-paths for binary files where\n> we accidentally load them (I just fixed one last week where we\n> unnecessarily loaded them in the diffstat code path). Somebody will need\n> to do some experimenting to shake out those code paths.\n\nThis is my main focus for now. They are easy to detect when your memory is small enough :D\n"},{"id":"163333","messageId":"20110314193254.GA21581@sigill.intra.peff.net","threadId":"26320","inReplyTo":"4D7D1BFE.2030008@miseler.de","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-03-14T19:32:54Z","receivedAt":"2011-03-14T19:32:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Mar 13, 2011 at 08:33:18PM +0100, Alexander Miseler wrote:\n\n> We want to store them as flat as possible. Ideally if we have a temp\n> file with the content (e.g. the output of some filter) it should be\n> possible to store it by simply doing a move/rename and updating some\n> meta data external to the actual file.\n\nYeah, that would be a nice optimization.  But I'd rather do the easy\nstuff first and see if more advanced stuff is still worth doing.\n\nFor example, I spent some time a while back designing a faster textconv\ninterface (the current interface spools the blob to a tempfile, whereas\nin some cases a filter needs to only access the first couple kilobytes\nof the file to get metadata). But what I found was that an even better\nscheme was to cache textconv output in git-notes. Then it speeds up the\nslow case _and_ the already-fast case.\n\nNow after this, would my new textconv interface still speed up the\ninitial non-cached textconv? Absolutely. But I didn't really care\nanymore, because the small speed up on the first run was not worth the\ntrouble of maintaining two interfaces (at least for my datasets).\n\nAnd this may fall into the same category. Accessing big blobs is\nexpensive. One solution is to make it a bit faster. Another solution is\nto just do it less. So we may find that once we are doing it less, it is\nnot worth the complexity to make it faster.\n\nAnd note that I am not saying \"it definitely won't be worth it\"; only\nthat it is worth making the easy, big optimizations first and then\nseeing what's left to do.\n\n> 1.) The loose file format is inherently unsuited for this. It has a\n> header before the actual content and the whole file (header + content)\n> is always compressed. Even if one changes this to\n> compressing/decompressing header and content independently it is still\n> unsuited by a) having the header within the same file and b) because\n> the header has no flags or other means to indicate a different\n> behavior (e.g. no compression) for the content. We could extend the\n> header format or introduce a new object type (e.g. flatblob) but both\n> would probably cause more trouble than other solutions. Another idea\n> would be to keep the metadata in an external file (e.g. 84d7.header\n> for the object 84d7). This would probably have a bad performance\n> though since every object lookup would first need to check for the\n> existence of a header file. A smarter variant would be to optionally\n> keep the meta data directly in the filename (e.g. saving the object as\n> 84d7.object_type.size.flag instead of just 84d7).\n> This would only require special handling for cases where the normal lookup for 84d7 fails.\n\nA new object type is definitely a bad idea. It changes the sha1 of the\nresulting object, which means that our identical trees which differ only\nin the use of \"flatblob\" versus regular blob will have different sha1s.\n\nSo I think the right place to insert this would be at the object db\nlayer. The header just has the type and size. But I don't think anybody\nis having a problem with large objects that are _not_ blobs. So the\nsimplest implementation would be a special blob-only object db\ncontaining pristine files. We implicitly know that objects in this db\nare blobs, and we can get the size from the filesystem via stat().\nChecking their sha1 would involve prepending \"blob <size>\\0\" to the file\ndata. It does introduce an extra stat() into object lookup, so probably\nwe would have the lookup order of pack, regular loose object, flat blob\nobject. Then you pay the extra stat() only in the less-common case of\naccessing either a large blob or a non-existent object.\n\nThat being said, I'm not sure how much this optimization will buy us.\nThere are times when being able to mmap() the file directly, or point an\nexternal program directly at the original blob will be helpful. But we\nwill still have to copy, for example on checkout. It would be nice if\nthere was a way to make a copy-on-write link from the working tree to\nthe original file. But I don't think there is a portable way to do so,\nand we can't allow the user to accidentally munge the contents of the\nobject db, which are supposed to be immutable.\n\n-Peff\n"},{"id":"163412","messageId":"AANLkTikhARYW=UcvfGmHb4a8Jja6jmHw+_qut6LAMXpk@mail.gmail.com","threadId":"26320","inReplyTo":"20110314193254.GA21581@sigill.intra.peff.net","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Eric Montellese","fromEmail":"emontellese@gmail.com","sentAt":"2011-03-16T00:35:14Z","receivedAt":"2011-03-16T00:35:14Z","isPatch":false,"sender":{"key":"emontellese@gmail.com","avatar":"https://gravatar.com/avatar/398aa9e755dc5d92d45ddbeab4e50c4d77608c41f6ff22afa11e24687ce3d209?d=mp&s=160"},"body":"Makes a lot of sense --\n\nAs you said, the \"sparse clone\" idea and this one (not downloading all\nbinaries) probably have a similar or related solution...  In fact, I'd\nimagine that most of the reasons for needing a sparse clone are\nbecause of large binaries.    (since text files compress so nicely in\ngit).\n\nAnd actually, if the sparse-clone idea is limited to only binaries\nbeing \"sparse\" (i.e. not copied) that probably simplifies the\nspares-clone logic quite a bit since you don't need to split up the\nbits of patches to generate the resulting changesets you need?  (this\nis based on my very loose understanding of how git works)\n\n\nSo, if we simplify our requirements a bit (at least as a first cut),\nperhaps we've now simplified down to these tasks  (similar but\nmodified from before):\n\n1. clean up git's handling of binaries to improve efficiency.  In\ndoing so, see if it makes sense to separate somewhat the way that\nbinaries are stored (particularly because this would help (2))\n\n2.  Allow full clones to be \"sparsely cloned\" (that is, cloned with\nthe exception of the large binary files).\n   2.1 As a corollary, no clones of any kind can be made from a sparse\nclone (sparse clones are \"leaf\" nodes on a tree of descendants) --\nthat simplifies the complexity quite a bit, since the \"remote\" you\ncloned from will always have the files if you need 'em.\n\n\ndoing so limits some of the possible applications of \"binary sparse\"\nclones, but might yield a cleaner final solution -- thoughts?\n\n\nEric\n\n\n\n\nOn Mon, Mar 14, 2011 at 3:32 PM, Jeff King <peff@peff.net> wrote:\n> On Sun, Mar 13, 2011 at 08:33:18PM +0100, Alexander Miseler wrote:\n>\n>> We want to store them as flat as possible. Ideally if we have a temp\n>> file with the content (e.g. the output of some filter) it should be\n>> possible to store it by simply doing a move/rename and updating some\n>> meta data external to the actual file.\n>\n> Yeah, that would be a nice optimization.  But I'd rather do the easy\n> stuff first and see if more advanced stuff is still worth doing.\n>\n> For example, I spent some time a while back designing a faster textconv\n> interface (the current interface spools the blob to a tempfile, whereas\n> in some cases a filter needs to only access the first couple kilobytes\n> of the file to get metadata). But what I found was that an even better\n> scheme was to cache textconv output in git-notes. Then it speeds up the\n> slow case _and_ the already-fast case.\n>\n> Now after this, would my new textconv interface still speed up the\n> initial non-cached textconv? Absolutely. But I didn't really care\n> anymore, because the small speed up on the first run was not worth the\n> trouble of maintaining two interfaces (at least for my datasets).\n>\n> And this may fall into the same category. Accessing big blobs is\n> expensive. One solution is to make it a bit faster. Another solution is\n> to just do it less. So we may find that once we are doing it less, it is\n> not worth the complexity to make it faster.\n>\n> And note that I am not saying \"it definitely won't be worth it\"; only\n> that it is worth making the easy, big optimizations first and then\n> seeing what's left to do.\n>\n>> 1.) The loose file format is inherently unsuited for this. It has a\n>> header before the actual content and the whole file (header + content)\n>> is always compressed. Even if one changes this to\n>> compressing/decompressing header and content independently it is still\n>> unsuited by a) having the header within the same file and b) because\n>> the header has no flags or other means to indicate a different\n>> behavior (e.g. no compression) for the content. We could extend the\n>> header format or introduce a new object type (e.g. flatblob) but both\n>> would probably cause more trouble than other solutions. Another idea\n>> would be to keep the metadata in an external file (e.g. 84d7.header\n>> for the object 84d7). This would probably have a bad performance\n>> though since every object lookup would first need to check for the\n>> existence of a header file. A smarter variant would be to optionally\n>> keep the meta data directly in the filename (e.g. saving the object as\n>> 84d7.object_type.size.flag instead of just 84d7).\n>> This would only require special handling for cases where the normal lookup for 84d7 fails.\n>\n> A new object type is definitely a bad idea. It changes the sha1 of the\n> resulting object, which means that our identical trees which differ only\n> in the use of \"flatblob\" versus regular blob will have different sha1s.\n>\n> So I think the right place to insert this would be at the object db\n> layer. The header just has the type and size. But I don't think anybody\n> is having a problem with large objects that are _not_ blobs. So the\n> simplest implementation would be a special blob-only object db\n> containing pristine files. We implicitly know that objects in this db\n> are blobs, and we can get the size from the filesystem via stat().\n> Checking their sha1 would involve prepending \"blob <size>\\0\" to the file\n> data. It does introduce an extra stat() into object lookup, so probably\n> we would have the lookup order of pack, regular loose object, flat blob\n> object. Then you pay the extra stat() only in the less-common case of\n> accessing either a large blob or a non-existent object.\n>\n> That being said, I'm not sure how much this optimization will buy us.\n> There are times when being able to mmap() the file directly, or point an\n> external program directly at the original blob will be helpful. But we\n> will still have to copy, for example on checkout. It would be nice if\n> there was a way to make a copy-on-write link from the working tree to\n> the original file. But I don't think there is a portable way to do so,\n> and we can't allow the user to accidentally munge the contents of the\n> object db, which are supposed to be immutable.\n>\n> -Peff\n>\n"},{"id":"163476","messageId":"AANLkTintr3szKMhbegeUgu+KHGtBGTRyQ3Y4pOwfwhEr@mail.gmail.com","threadId":"26320","inReplyTo":"20110314193254.GA21581@sigill.intra.peff.net","subject":"Re: Fwd: Git and Large Binaries: A Proposed Solution","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2011-03-16T14:40:51Z","receivedAt":"2011-03-16T14:40:51Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Mar 15, 2011 at 2:32 AM, Jeff King <peff@peff.net> wrote:\n> That being said, I'm not sure how much this optimization will buy us.\n> There are times when being able to mmap() the file directly, or point an\n> external program directly at the original blob will be helpful. But we\n> will still have to copy, for example on checkout.\n\nSparse checkout code may help. If those large files are not always\nneeded, they can be marked skip-checkout based on\ncore.bigFileThreshold and won't be checked out until explictly\nrequested. This use may conflict with sparse checkout because it sets\nskip-checkout bits automatically from $GIT_DIR/info/sparsecheckout\nthough.\n-- \nDuy\n"}]}