{"thread":{"id":"29859","subject":"Replacing large blobs in git history","startedAt":"2012-03-06T16:09:44Z","lastAt":"2012-03-08T21:22:57Z","messageCount":6,"participants":["Barry Roberts","Neal Kreitzinger","Michael Haggerty","Ævar Arnfjörð Bjarmason","Holger Hellmuth","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"186225","messageId":"CAD-6W7byTiuE9MFZY1yG_ann-Ox7+wGjYduZ=Wwmw0ToF5Pynw@mail.gmail.com","threadId":"29859","inReplyTo":null,"subject":"Replacing large blobs in git history","fromName":"Barry Roberts","fromEmail":"blr@robertsr.us","sentAt":"2012-03-06T16:09:44Z","receivedAt":"2012-03-06T16:09:44Z","isPatch":false,"sender":{"key":"blr@robertsr.us","avatar":null},"body":"I started this question on #git last week, but this is getting long,\nand things have changed some, so I'm going to try here.\n\nI had a 3rd party jar file checked in to our git repository.  It was\nabout 4 mb, so no big deal.  Then about 17 months ago somebody checked\nin a 550 mb version.  There were several versions of the original file\nin several different directories.  The large version replaced the\nsmall version in some of those directories (but not all of them).\nThen somebody found a \"small\" version that was only 110 mb and\nreplaced some of the 550 mb files and some of the old 4 mb files.\nFinally several months after that we got the correct updated 5 mb\nlatest version.  But I'm still carrying around an extra 660 mb in my\nobject database, and we are adding developers and moving to an\noff-site location with lower bandwidth and higher latency, so I would\nlike to clean this up.\n\nMy first attempt just removed the blob (by hash ID).  It's been over a\nyear since the small correct file was checked in, so the odds of ever\nneeding to build anything that old are very slim. But after thinking\nabout it some, I came up with this to replace the blob with the\ncorrect one and wanted to see if this is a reasonable way to do this\nbefore I actually backup and then replace my central git repository.\n\ngit filter-branch --index-filter 'killem=$(git ls-files --stage  |\ngrep 7a36af54a6c47\\\\\\|abe809091bcb3 ) ; if [ -n \"$killem\" ] ; then git\nls-files --stage |grep 7a36af54a6c47\\\\\\|abe809091bcb3 | sed -f\n/home/blr/tmp/chgblob.sed |  git update-index --index-info ; fi'\n\nchgblob.sed looks like this:\ns/7a36af54a6c47a29eb9690caefa132489d39c4d0/8924ef0f78b3d09957a8697ca93cce6700771071/g\ns/abe809091bcb37a06284f8353366074622d72373/8924ef0f78b3d09957a8697ca93cce6700771071/g\n\n7a36af is the 550 mb blob, abe80909 is the 110 mb, and 8924ef0f is the\n5 mb new version.\n\nThis isn't extremely efficient since it does the 'git ls-filess\n--stage' twice (once to see if the blob is used, then again to change\nit ONLY if the blob is referenced in the current index).  But that\nonly adds a few seconds to the 28 minute runtime, so I'm not too\nworried about that.  And yes, I could just check for the return value\nof grep, but I did echo $killem while I was debugging and that was\nuseful, so I just left it like that.\n\nDoes this look like a reasonable way to accomplish what I'm trying to\ndo, or am I doing something that's going to cause grief later?\n\nThanks,\nBarry\n"},{"id":"186253","messageId":"4F56786D.60801@gmail.com","threadId":"29859","inReplyTo":"CAD-6W7byTiuE9MFZY1yG_ann-Ox7+wGjYduZ=Wwmw0ToF5Pynw@mail.gmail.com","subject":"Re: Replacing large blobs in git history","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-03-06T20:49:49Z","receivedAt":"2012-03-06T20:49:49Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 3/6/2012 10:09 AM, Barry Roberts wrote:\n> I started this question on #git last week, but this is getting long,\n> and things have changed some, so I'm going to try here.\n>\n> I had a 3rd party jar file checked in to our git repository.  It was\n> about 4 mb, so no big deal.  Then about 17 months ago somebody\n> checked in a 550 mb version.  There were several versions of the\n> original file in several different directories.  The large version\n> replaced the small version in some of those directories (but not all\n> of them). Then somebody found a \"small\" version that was only 110 mb\n> and replaced some of the 550 mb files and some of the old 4 mb\n> files. Finally several months after that we got the correct updated 5\n> mb latest version.  But I'm still carrying around an extra 660 mb in\n> my object database, and we are adding developers and moving to an\n> off-site location with lower bandwidth and higher latency, so I\n> would like to clean this up.\n>\n> My first attempt just removed the blob (by hash ID).  It's been over\n> a year since the small correct file was checked in, so the odds of\n> ever needing to build anything that old are very slim. But after\n> thinking about it some, I came up with this to replace the blob with\n> the correct one and wanted to see if this is a reasonable way to do\n> this before I actually backup and then replace my central git\n> repository.\n>\n> git filter-branch --index-filter 'killem=$(git ls-files --stage  |\n> grep 7a36af54a6c47\\\\\\|abe809091bcb3 ) ; if [ -n \"$killem\" ] ; then\n> git ls-files --stage |grep 7a36af54a6c47\\\\\\|abe809091bcb3 | sed -f\n> /home/blr/tmp/chgblob.sed |  git update-index --index-info ; fi'\n>\n> chgblob.sed looks like this:\n> s/7a36af54a6c47a29eb9690caefa132489d39c4d0/8924ef0f78b3d09957a8697ca93cce6700771071/g\n>\n>\ns/abe809091bcb37a06284f8353366074622d72373/8924ef0f78b3d09957a8697ca93cce6700771071/g\n>\n> 7a36af is the 550 mb blob, abe80909 is the 110 mb, and 8924ef0f is\n> the 5 mb new version.\n>\n> This isn't extremely efficient since it does the 'git ls-filess\n> --stage' twice (once to see if the blob is used, then again to\n> change it ONLY if the blob is referenced in the current index).  But\n> that only adds a few seconds to the 28 minute runtime, so I'm not\n> too worried about that.  And yes, I could just check for the return\n> value of grep, but I did echo $killem while I was debugging and that\n> was useful, so I just left it like that.\n>\n> Does this look like a reasonable way to accomplish what I'm trying\n> to do, or am I doing something that's going to cause grief later?\n>\nBe aware that you are rewriting history.  I assume this is published\nhistory that you are going to run filter-branch on.  That means everyone \nwho cloned from the old history (pre-filter-branch), not to mention \nthose who also have WIP based on the old history, will need to somehow \nadjust to the new history.  How do you plan on addressing that?  (see \ngit-rebase manpage section \"recovering from upstream rebase\" for more \ninfo on the implications of rewriting history.)\n\n(I have never done filter-branch, and am not an expert on git, but do \nfind this subject relevant to normal use of git.)\n\nv/r,\nneal\n"},{"id":"186293","messageId":"4F5724A5.7050405@alum.mit.edu","threadId":"29859","inReplyTo":"CAD-6W7byTiuE9MFZY1yG_ann-Ox7+wGjYduZ=Wwmw0ToF5Pynw@mail.gmail.com","subject":"Re: Replacing large blobs in git history","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2012-03-07T09:04:37Z","receivedAt":"2012-03-07T09:04:37Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 03/06/2012 05:09 PM, Barry Roberts wrote:\n> I started this question on #git last week, but this is getting long,\n> and things have changed some, so I'm going to try here.\n> \n> I had a 3rd party jar file checked in to our git repository.  It was\n> about 4 mb, so no big deal.  Then about 17 months ago somebody checked\n> in a 550 mb version.  There were several versions of the original file\n> in several different directories.  The large version replaced the\n> small version in some of those directories (but not all of them).\n> Then somebody found a \"small\" version that was only 110 mb and\n> replaced some of the 550 mb files and some of the old 4 mb files.\n> Finally several months after that we got the correct updated 5 mb\n> latest version.  But I'm still carrying around an extra 660 mb in my\n> object database, and we are adding developers and moving to an\n> off-site location with lower bandwidth and higher latency, so I would\n> like to clean this up.\n> \n> My first attempt just removed the blob (by hash ID).  It's been over a\n> year since the small correct file was checked in, so the odds of ever\n> needing to build anything that old are very slim. But after thinking\n> about it some, I came up with this to replace the blob with the\n> correct one and wanted to see if this is a reasonable way to do this\n> before I actually backup and then replace my central git repository.\n> \n> git filter-branch --index-filter 'killem=$(git ls-files --stage  |\n> grep 7a36af54a6c47\\\\\\|abe809091bcb3 ) ; if [ -n \"$killem\" ] ; then git\n> ls-files --stage |grep 7a36af54a6c47\\\\\\|abe809091bcb3 | sed -f\n> /home/blr/tmp/chgblob.sed |  git update-index --index-info ; fi'\n> \n> chgblob.sed looks like this:\n> s/7a36af54a6c47a29eb9690caefa132489d39c4d0/8924ef0f78b3d09957a8697ca93cce6700771071/g\n> s/abe809091bcb37a06284f8353366074622d72373/8924ef0f78b3d09957a8697ca93cce6700771071/g\n> \n> 7a36af is the 550 mb blob, abe80909 is the 110 mb, and 8924ef0f is the\n> 5 mb new version.\n\nYou could use \"git replace\" to cause the bad blobs to be replaced\neverywhere they appear:\n\n    $ git replace 7a36af54a6c47a29eb9690caefa132489d39c4d0 \\\n                  8924ef0f78b3d09957a8697ca93cce6700771071\n    $ git replace abe809091bcb37a06284f8353366074622d72373 \\\n                  8924ef0f78b3d09957a8697ca93cce6700771071\n\nThen you could use \"git filter-branch\" to \"bake in\" the substitutions\n(but please see the caveats mentioned by Neal).\n\nIt seems like an alternative to using \"git filter-branch\" would be to\nshare the \"git replace\" references across repositories.  This would make\nthe short versions of the file appear wherever they should without\nrequiring history to be rewritten entirely.  But I don't believe that\nthis approach would allow the long versions of the file to be discarded\nby the git garbage collector, so it would not help you reduce clone sizes.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"186339","messageId":"CACBZZX4hinV8vkebyNCLp_Ac6L80aNbdGOFqg1nSsCuRktFFrg@mail.gmail.com","threadId":"29859","inReplyTo":"4F56786D.60801@gmail.com","subject":"Re: Replacing large blobs in git history","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2012-03-07T21:27:11Z","receivedAt":"2012-03-07T21:27:11Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Tue, Mar 6, 2012 at 21:49, Neal Kreitzinger <nkreitzinger@gmail.com> wrote:\n> On 3/6/2012 10:09 AM, Barry Roberts wrote:\n> Be aware that you are rewriting history.  I assume this is published\n> history that you are going to run filter-branch on.  That means everyone who\n> cloned from the old history (pre-filter-branch), not to mention those who\n> also have WIP based on the old history, will need to somehow adjust to the\n> new history.\n\nDoes something other than git-fsck actually check whether the\ncollection of blobs you're getting from the remote when you clone have\nsensible sha1's?\n\nWhat'll happen if he replaces that 550MB blob with a 0 byte blob but\nhacks the object store so that it pretends to have the same sha1?\n\nOf course the real solution to this issue is to either rewrite\nhistory, or to change Git to support partially fetching the old blobs\nin your project.\n"},{"id":"186458","messageId":"4F58D2CD.2050502@ira.uka.de","threadId":"29859","inReplyTo":"CACBZZX4hinV8vkebyNCLp_Ac6L80aNbdGOFqg1nSsCuRktFFrg@mail.gmail.com","subject":"Re: Replacing large blobs in git history","fromName":"Holger Hellmuth","fromEmail":"hellmuth@ira.uka.de","sentAt":"2012-03-08T15:39:57Z","receivedAt":"2012-03-08T15:39:57Z","isPatch":false,"sender":{"key":"hellmuth@ira.uka.de","avatar":null},"body":"On 07.03.2012 22:27, Ævar Arnfjörð Bjarmason wrote:\n> Does something other than git-fsck actually check whether the\n> collection of blobs you're getting from the remote when you clone have\n> sensible sha1's?\n>\n> What'll happen if he replaces that 550MB blob with a 0 byte blob but\n> hacks the object store so that it pretends to have the same sha1?\n\nThis is something I tested once because of security concerns (i.e. what \nhappens if a malicious intruder just drops something else into the \nobject store) and if I remember correctly only git-fsck was able to spot \nthe switch. But I didn't test cloning, only a few local operations.\n"},{"id":"186498","messageId":"7vy5rabxe6.fsf@alter.siamese.dyndns.org","threadId":"29859","inReplyTo":"4F58D2CD.2050502@ira.uka.de","subject":"Re: Replacing large blobs in git history","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-08T21:22:57Z","receivedAt":"2012-03-08T21:22:57Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Holger Hellmuth <hellmuth@ira.uka.de> writes:\n\n> On 07.03.2012 22:27, Ævar Arnfjörð Bjarmason wrote:\n>> Does something other than git-fsck actually check whether the\n>> collection of blobs you're getting from the remote when you clone have\n>> sensible sha1's?\n>>\n>> What'll happen if he replaces that 550MB blob with a 0 byte blob but\n>> hacks the object store so that it pretends to have the same sha1?\n>\n> This is something I tested once because of security concerns\n> (i.e. what happens if a malicious intruder just drops something else\n> into the object store) and if I remember correctly only git-fsck was\n> able to spot the switch. But I didn't test cloning, only a few local\n> operations.\n\nLocal operation that do not have to look at such a corrupt blob will\nnot verify everything under the sun every time for obvious reasons.\n\nAn operation to transfer objects out of the repository (e.g. serving\nas the source of \"clone\" from elsewhere) will notice when it has to\nsend such a corrupt object and you will be prevented from spreading\nthe damage.\n\nThe same thing for a transfer in the reverse direction. When the\nother side tells us that it is giving us everything we asked, we\nstill look at all the objects we received to make sure.\n"}]}