{"thread":{"id":"6907","subject":"removing content from git history","startedAt":"2007-02-21T16:45:27Z","lastAt":"2007-10-10T14:41:21Z","messageCount":25,"participants":["Michael Hendricks","Shawn O. Pearce","Linus Torvalds","J. Bruce Fields","Nicolas Pitre","Junio C Hamano","Bill Lear","Johannes Schindelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"35182","messageId":"20070221164527.GA8513@ginosko.local","threadId":"6907","inReplyTo":null,"subject":"removing content from git history","fromName":"Michael Hendricks","fromEmail":"michael@ndrix.org","sentAt":"2007-02-21T16:45:27Z","receivedAt":"2007-02-21T16:45:27Z","isPatch":false,"sender":{"key":"michael@ndrix.org","avatar":"https://gravatar.com/avatar/315311e6daa79f24e5648f9534420c24ec48eada42efd4110f1d17167ff44fa8?d=mp&s=160"},"body":"I assume that this question has already been addressed on the mailing\nlist, but I wasn't able to find anything about it in the archives.\n\nIs it possible to remove content entirely from git's history?  I have a\nclient who does not use git for version control.  A couple months ago\nthey committed some sensitive client information which should never have\nbeen committed.  Recently, they realized the mistake and now want to\nremove all traces of the mistake from history.\n\nI would like to migrate them to git at some point.  However, if they had\nbeen using git for version control already, I'm not sure how I would\nsolved this particular problem.  Any suggestions?\n\n-- \nMichael\n"},{"id":"35183","messageId":"20070221165636.GH25559@spearce.org","threadId":"6907","inReplyTo":"20070221164527.GA8513@ginosko.local","subject":"Re: removing content from git history","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-02-21T16:56:36Z","receivedAt":"2007-02-21T16:56:36Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Michael Hendricks <michael@ndrix.org> wrote:\n> Is it possible to remove content entirely from git's history?\n\nNo, not once it has been published around to another repository.\nSince every developer has a copy of the repository its very difficult\nto remove something, as it must be removed from every developer's\nrepository, and each developer must perform an action to agree to\nthat removal.  So just one hold-out will keep the bad content around.\n\n> I have a\n> client who does not use git for version control.  A couple months ago\n> they committed some sensitive client information which should never have\n> been committed.  Recently, they realized the mistake and now want to\n> remove all traces of the mistake from history.\n> \n> I would like to migrate them to git at some point.  However, if they had\n> been using git for version control already, I'm not sure how I would\n> solved this particular problem.  Any suggestions?\n\nThe *only* way to do this in Git is to completely recreate every\ncommit after that point.  This changes all commit IDs and basically\nforks the project into two completely different histories: the\none with the bad thing in it, and the one without the bad thing.\nUsers who have the bad thing will continue to have the bad thing\nuntil they take explicit action to throw away all of that history\nand switch to the other one.\n\nNow this is actually not a huge deal if you do it on your local\nrepository and go \"whoops, I should not have committed that\".  If you\nhave not yet pushed the commit to another repository (and someone\nhas not yet fetched it from you either) you can use git-rebase to\ndiscard it.  But once its been pushed/fetched the genie is out of\nthe bottle, and its not going back in.\n\n-- \nShawn.\n"},{"id":"35187","messageId":"Pine.LNX.4.64.0702210904350.4043@woody.linux-foundation.org","threadId":"6907","inReplyTo":"20070221164527.GA8513@ginosko.local","subject":"Re: removing content from git history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-21T17:14:44Z","receivedAt":"2007-02-21T17:14:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Feb 2007, Michael Hendricks wrote:\n>\n> I assume that this question has already been addressed on the mailing\n> list, but I wasn't able to find anything about it in the archives.\n> \n> Is it possible to remove content entirely from git's history?\n\nIt's been discussed.\n\nThere are two options for doing it:\n\n - rewriting history. There are a few tools for this already, and for \n   specific needs it would be fairly easy to resurrect git-convert-objects \n   to do it for any kind of object.\n\n   See \"cg-admin-rewritehist\" from cogito for an example of a tool that \n   would do what you need done. In fact, it has this exact thing as the \n   first example.\n\n   (Btw, I think cg-admin-rewritehist is one of the few things that cogito \n   had that was really a good idea. Not that people probably _used_ it \n   much, but it's somethign that makes sense in the plumbing)\n\n - explicit support for \"missing objects\". We don't do it right now, but \n   we could add it. It was discussed for things like limited history etc \n   (the \"shallow clone\" kind of thing, before people actually added \n   shallow clones), and it would support the notion of \"we export all our \n   history, but for internal reasons we cannot make certain objects \n   available\" kinds of workflows.\n\nSo right now, rewriting history is an option that you can do. It will \neffectively create a totally new branch (which you can then make into a \nnew repository) which has nothing in common with the old branch from the \npoint where it was modified. So you can never really merge the two ever \nagain, and you need to make sure that everybody who had the old repo \ncontents will destroy it.\n\nBut at least in theory, it wouldn't be impossible to extend on the \n\".git/grafts\" kind of setup to say \"this object has been consciously \ndeleted\", and that could in some circumstances be a better model. The \nbiggest headache there would be the need to extend the native git protocol \nwith a way to add such objects.\n\n\t\t\tLinus\n"},{"id":"35188","messageId":"20070221171738.GA9112@fieldses.org","threadId":"6907","inReplyTo":"20070221165636.GH25559@spearce.org","subject":"Re: removing content from git history","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2007-02-21T17:17:38Z","receivedAt":"2007-02-21T17:17:38Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Wed, Feb 21, 2007 at 11:56:36AM -0500, Shawn O. Pearce wrote:\n> Now this is actually not a huge deal if you do it on your local\n> repository and go \"whoops, I should not have committed that\".  If you\n> have not yet pushed the commit to another repository (and someone\n> has not yet fetched it from you either) you can use git-rebase to\n> discard it.\n\nAlso it can't have done any (non-fast-forward) merges since then.\n\nReconstructing history with a bunch of merges seems like something that\ncould be a huge pain.  (Though with some tools it might be doable.)\n\n--b.\n"},{"id":"35191","messageId":"Pine.LNX.4.64.0702210934470.4043@woody.linux-foundation.org","threadId":"6907","inReplyTo":"20070221171738.GA9112@fieldses.org","subject":"Re: removing content from git history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-21T18:02:07Z","receivedAt":"2007-02-21T18:02:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Feb 2007, J. Bruce Fields wrote:\n> \n> Reconstructing history with a bunch of merges seems like something that\n> could be a huge pain.  (Though with some tools it might be doable.)\n\nIt's not actually that painful, but it *is* expensive.\n\nI wrote git-convert-cache (now \"git convert-objects\") back when we did the \nSHA1/compression switchover changes and the date format translation, so \nwe've actually had a tool that can do history rewriting pretty much since \nday 1 (well, \"day 14\", to be exact, but still.. April 2005).\n\nBUT:\n\n - I'm not guaranteeing that it works any more. We haven't changed the \n   fundamental object format since, so that particular program has never \n   gotten any testing. It still compiles, but does it work? I dunno.\n\n   I actually tested it on git itself. It converted the top of the git \n   tree successfully, and generated a *new* git history. Why? Because it \n   will actually rewrite the old git tree entries that have permission \n   0664 into 0644: the *data* will be identical (and no git tools except \n   for \"git fsck --pedantic\" will even notice the difference), but the \n   converted tree avoids one of the legacy decisions that we never fixed \n   in the git repository itself.\n\n   So it works at least to *some* degree, but I would suggest you be very \n   very careful!\n\n - it can be slow. For something like git, which isn't *that* big, and \n   where we actually don't need to do a lot of rewriting (ie all the blobs \n   stay the same, and only a few trees have to be rewritten, and so it's \n   really just rewriting commits), it's not that bad. It actyally \n   converted the whole git history in less than ten seconds for me.\n\n   But if you have a *huge* tree, and you actually convert objects too \n   (say, you started using git on Windows before the \"autocrlf\" thing, and \n   want to convert the old blobs from CRLF -> LF), it would\n\n    (a) require some extensions to convert-object.c to do the blob \n        conversion\n    (b) be *much* slower\n    (c) generate tons of unpacked objects (because git-convert-objects \n        doesn't know to pack in between, and doesn't use anything \n        newfangled like \"git-fast-import\" to do anything clever)\n\n   For the kernel, it took 2 minutes, but again, it was exactly the same \n   thing: just a few old tree objects that it rewrote, and as a result, \n   every single commit SHA1 changed. Still, it was almost _only_ commits \n   (it generated 49521 new objects, 49332 of which was the new commit \n   history)\n\n   If you want to rewrite a *lot* (ie somethign that exists in more than \n   just a few trees), and you have lots of history, it can be very \n   expensive indeed.\n\n - It currently doesn't convert the SHA1 numbers that show up in commit \n   messages. It could, and it should. But it doesn't. So once you convert \n   a git project, it doesn't do the nice \"gitk does links from the SHA1 \n   text in a commit message to the commit it talks about\" any more.\n\n   Somebody should fix that.\n\nAnyway, git-convert-objects does kind of give you a starting point. It \nshould be fixed to use \"git-fast-import\" or repack once in a while (so \nthat it doesn't leave tons and tons of unpacked objects), and it should be \nfixed to fix up any commit messages that mention SHA1's that it has \nalready converted to something else, but it seems to still work. It would \nnot be impossible at all to extend the tree-rewriting logic to remove some \nfile or a particular SHA1 object you want to replace.\n\n\t\t\tLinus\n"},{"id":"35190","messageId":"alpine.LRH.0.82.0702211236180.31945@xanadu.home","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702210904350.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-02-21T18:02:08Z","receivedAt":"2007-02-21T18:02:08Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 21 Feb 2007, Linus Torvalds wrote:\n\n> But at least in theory, it wouldn't be impossible to extend on the \n> \".git/grafts\" kind of setup to say \"this object has been consciously \n> deleted\", and that could in some circumstances be a better model. The \n> biggest headache there would be the need to extend the native git protocol \n> with a way to add such objects.\n\nI think that would be a big security issue.  Right now the GIT history \ncan be validated and more importantly trusted from a single commit \nsignature.  If poking holes in that model is allowed by the graft \nmechanism, it must remain a local thing and a very conscious one \notherwise the GIT trust model would be greatly weakened.\n\nIf your goal is to remove content froma repository then the only \nsensible way is to rewrite history before publishing.  It is pointless \nto add mechanisms to remove content after it has been distributed.\n\n\nNicolas\n"},{"id":"35192","messageId":"Pine.LNX.4.64.0702211009520.4043@woody.linux-foundation.org","threadId":"6907","inReplyTo":"alpine.LRH.0.82.0702211236180.31945@xanadu.home","subject":"Re: removing content from git history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-21T18:13:09Z","receivedAt":"2007-02-21T18:13:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Feb 2007, Nicolas Pitre wrote:\n> \n> If your goal is to remove content froma repository then the only \n> sensible way is to rewrite history before publishing.  It is pointless \n> to add mechanisms to remove content after it has been distributed.\n\nI'm not entirely in disagreement, but I can see the model where some \ncompany wants to make their work available (with the same history as their \nown internal stuff), but doesn't want to make a single file available for \nsome reason.\n\nSo they'd have an external thing that just has the file excised.\n\nNow, arguably, it's a lot better to use a \"supermodule\" approach for \nsomething like this: have two separate git trees, publish the public one, \nand use an internal supermodule that ties the public and internal trees \ntogether.\n\nSo supermodules might be a way to solve it in a better (and safer - the \n\"remove objects from the public tree\" thing is very error prone, since if \nyou *ever* expose the object by mistake, its now public) way. But I don't \nthink the \"filter out objects\" thing is necessarily fundamentally flawed \nas an approach.\n\n\t\t\tLinus\n"},{"id":"35193","messageId":"Pine.LNX.4.64.0702211016340.4043@woody.linux-foundation.org","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702210934470.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-21T18:24:12Z","receivedAt":"2007-02-21T18:24:12Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Feb 2007, Linus Torvalds wrote:\n> \n>    For the kernel, it took 2 minutes, but again, it was exactly the same \n>    thing: just a few old tree objects that it rewrote, and as a result, \n>    every single commit SHA1 changed. Still, it was almost _only_ commits \n>    (it generated 49521 new objects, 49332 of which was the new commit \n>    history)\n\nSide note: I wasn't entirelyaccurate. The kernel had trees with file mode \n0644 for all the early commits, because my umask is 0022. So everything up \nto commit 4bfa437cf1 is shared after the conversion.\n\nBut the next one (commit 5dfa9c1b4f) introduced the file \ninclude/asm-mips/vr41xx/pci.h with file mode 0664, and I'm not 100% sure \nwhy that one happened with that file mode, but as a result, every single \ncommit ever after will have a different SHA1, because the tree got \nrewritten (and subsequent commits - even if their trees did *not* get \nrewritten - will obviously have different parent SHA1's).\n\nSo 56 commits are shared, and \"only\" 49276 commits were rewritten (and \napparently 245 trees).\n\n\t\t\tLinus\n"},{"id":"35194","messageId":"20070221183028.GA9088@ginosko.local","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702210904350.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Michael Hendricks","fromEmail":"michael@ndrix.org","sentAt":"2007-02-21T18:30:29Z","receivedAt":"2007-02-21T18:30:29Z","isPatch":false,"sender":{"key":"michael@ndrix.org","avatar":"https://gravatar.com/avatar/315311e6daa79f24e5648f9534420c24ec48eada42efd4110f1d17167ff44fa8?d=mp&s=160"},"body":"On Wed, Feb 21, 2007 at 09:14:44AM -0800, Linus Torvalds wrote:\n> \n>    See \"cg-admin-rewritehist\" from cogito for an example of a tool that \n>    would do what you need done. In fact, it has this exact thing as the \n>    first example.\n\nThat's just what I was looking for.  Thanks.\n\n> So right now, rewriting history is an option that you can do. It will \n> effectively create a totally new branch (which you can then make into a \n> new repository) which has nothing in common with the old branch from the \n> point where it was modified. So you can never really merge the two ever \n> again, and you need to make sure that everybody who had the old repo \n> contents will destroy it.\n\nWhat's a decent way to make a branch into a new repository?  My first\ninclination is to \"cp -a\" the existing repository, checkout the branch,\ndelete all other branches and repack.  That seems to have worked in my\nquick test, but is there a better way?\n\n-- \nMichael\n"},{"id":"35195","messageId":"20070221183754.GK25559@spearce.org","threadId":"6907","inReplyTo":"20070221183028.GA9088@ginosko.local","subject":"Re: removing content from git history","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-02-21T18:37:54Z","receivedAt":"2007-02-21T18:37:54Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Michael Hendricks <michael@ndrix.org> wrote:\n> What's a decent way to make a branch into a new repository?  My first\n> inclination is to \"cp -a\" the existing repository, checkout the branch,\n> delete all other branches and repack.  That seems to have worked in my\n> quick test, but is there a better way?\n\nDon't \"cp -a\" the repository, use git-clone.\n\nAnd actually, if you just want to pull one branch out into its\nown repository you can do something like this:\n\n\tmkdir ../theonebranch\n\tcd ../theonebranch\n\tgit init\n\tgit fetch ../oldstuff theonebranch:master\n\nand you have just the content of `theonebranch` from ../oldstuff\nstored here, as master.\n\nOptionally if you now want to actually see the files, you would do:\n\n\tgit checkout\n\n-- \nShawn.\n"},{"id":"35196","messageId":"alpine.LRH.0.82.0702211321040.31945@xanadu.home","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702211009520.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-02-21T18:39:24Z","receivedAt":"2007-02-21T18:39:24Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 21 Feb 2007, Linus Torvalds wrote:\n\n> \n> \n> On Wed, 21 Feb 2007, Nicolas Pitre wrote:\n> So supermodules might be a way to solve it in a better (and safer - the \n> \"remove objects from the public tree\" thing is very error prone, since if \n> you *ever* expose the object by mistake, its now public) way. But I don't \n> think the \"filter out objects\" thing is necessarily fundamentally flawed \n> as an approach.\n\nWell if you really wanted to do such a thing then you could use a new \nobject type that only serves as a stub pretending to be another object \nwhich SHA1 would have been xyz.  When referenced this object would \ngenerate a warning indicating to the user that given object has been \nexcised out, but otherwise the whole reachability validation would still \nwork as usual.\n\nAnd since this object would be distributed through standard mechanisms \nthen there would be no need for protocol extensions.\n\nI don't know if this could help creating SHA1 collisions though.  We've \ndismissed them as highly improbable because the likelihood of a \ncollision to hide compromised material would most probably require a \nbinary blob somewhere to balance the hash and would hardly be \ncompilable/undetected.  But with object stubs with the ability to \npretend having any possible SHA1 is in fact a nice way to hide 20-byte \nbinary blobs in the hash chain possibly making it \"easier\" to create \n\"useful\" collisions.  This is where I see a weakening of the trust \nmodel.\n\n\nNicolas\n"},{"id":"35197","messageId":"Pine.LNX.4.64.0702211041500.4043@woody.linux-foundation.org","threadId":"6907","inReplyTo":"20070221183028.GA9088@ginosko.local","subject":"Re: removing content from git history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-21T18:47:29Z","receivedAt":"2007-02-21T18:47:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Feb 2007, Michael Hendricks wrote:\n> \n> What's a decent way to make a branch into a new repository?  My first\n> inclination is to \"cp -a\" the existing repository, checkout the branch,\n> delete all other branches and repack.  That seems to have worked in my\n> quick test, but is there a better way?\n\nThat works.\n\nAs does just \"clone repo, delete all unwanted branches, and prune\" (of \ncourse, if you don't want the old repo, you can skip the \"clone\" part, and \njust do the \"delete all unwanted branches and prune\" thing).\n\nIn some ways, a more straightforward approach may be to just create a new \nrepo, and populate it with just one branch (I say \"more straightforward\", \nnot \"easier\", because I just think it's conceptually simpler):\n\n\tmkdir new-repo\n\tcd new-repo\n\tgit init\n\tgit pull old-repo <branch>\n\n(add \"--bare\" and \"--shared\" to taste - with bare repos yu can also do it \nthe other way by doing a push into it from outside after you've created \nit, which can be the \"logical\" way to do it if you want to just publish \nthe end result on some shared site)\n\n\t\tLinus\n"},{"id":"35198","messageId":"alpine.LRH.0.82.0702211345350.31945@xanadu.home","threadId":"6907","inReplyTo":"20070221183028.GA9088@ginosko.local","subject":"Re: removing content from git history","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-02-21T18:52:40Z","receivedAt":"2007-02-21T18:52:40Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 21 Feb 2007, Michael Hendricks wrote:\n\n> What's a decent way to make a branch into a new repository?  My first\n> inclination is to \"cp -a\" the existing repository, checkout the branch,\n> delete all other branches and repack.  That seems to have worked in my\n> quick test, but is there a better way?\n\nLike Shawn said the better way is simply to fetch that branch into a new \nrepo.\n\nIf you do a cp -a and delete unwanted branches it'll work as well of \ncourse, but repacking won't get rid of all the data from the believed to \nbe deleted branches since some reflog, the HEAD reflog in particular, \nwill most probably have references to commits from the removed branches. \nTherefore the pack will still contain that data, at least untill the \nreflog entries expire and get pruned.\n\nOf course if you want to publish just the wanted branch and perform a \npush to a public place then only those objects for that branch will be \nsent like for the fetch case.\n\n\nNicolas\n"},{"id":"35199","messageId":"Pine.LNX.4.64.0702211048220.4043@woody.linux-foundation.org","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702211041500.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-21T18:56:11Z","receivedAt":"2007-02-21T18:56:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Feb 2007, Linus Torvalds wrote:\n\n> \n> \n> On Wed, 21 Feb 2007, Michael Hendricks wrote:\n> > \n> > What's a decent way to make a branch into a new repository?  My first\n> > inclination is to \"cp -a\" the existing repository, checkout the branch,\n> > delete all other branches and repack.  That seems to have worked in my\n> > quick test, but is there a better way?\n> \n> That works.\n\nBtw, when I say \"works\", I do mean that \"yeah, 'cp -a' works, but \ngenerally you're better off cloning\".\n\nWhen you use 'cp -a' you have to re-build the index at the very least. It \nso happens that since you checked out the branch explicitly, that will do \nit for you anyway, but it's still often a good idea to just *not* use the \nregular \"copy everything by hand\" approach.\n\nIf you want to be really efficient, there are actually better ways. For \nexample, since you want to avoid having any of the old objects even \nreachable by mistake), you're probably better off with an explicit pull of \nthe explicit branch, if only because that also involves a re-pack of only \nthe reachable objects, and you know that there won't be any reflogs etc \nthat might still make the object you try to remove be accessible to people \nwho can access the resulting repository directly.\n\n(Yeah, the \"cp -a\" is faster than the \"git pull\", but since you want to do \nthe packing that git pull does for you *anyway* to get rid of the old \nobjects, \"git pull\" actually ends up being better).\n\n\t\t\tLinus\n"},{"id":"35200","messageId":"7vy7mrnxlo.fsf@assigned-by-dhcp.cox.net","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702210904350.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-21T19:01:55Z","receivedAt":"2007-02-21T19:01:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n>  - explicit support for \"missing objects\". We don't do it right now, but \n>    we could add it. It was discussed for things like limited history etc \n>    (the \"shallow clone\" kind of thing, before people actually added \n>    shallow clones), and it would support the notion of \"we export all our \n>    history, but for internal reasons we cannot make certain objects \n>    available\" kinds of workflows.\n> ...\n> But at least in theory, it wouldn't be impossible to extend on the \n> \".git/grafts\" kind of setup to say \"this object has been consciously \n> deleted\", and that could in some circumstances be a better model. The \n> biggest headache there would be the need to extend the native git protocol \n> with a way to add such objects.\n\nWhile I agree in principle to the argument that there is no\ntaking it back what's already published, I've heard people\nwanting to just stop distributing further, without worrying\nabout copies already out there.  'missing objects' support would\nhelp us in such a situation.\n\nSupporting 'missing objects' in general would be painful, when\nthey contain pointers to other objects (i.e. tags, commits, and\ntrees).\n\nThinking aloud...\n\n * missing blob: we can have 'stub blob' objects.  Probably the\n   object header for such an object would look like:\n\n\tstub <length> NUL\n\t-----------------\n        object <object name of the real blob object>\n        type blob\n\n   Hashing a 'stub' object (along with its header as usual, in\n   write_sha1_file_prepare()) would instead just report the\n   object name recorded there.\n\n   When packing (this applies both to local repacking and\n   push/fetch object transfer to other repositories), the stub\n   object is included.  delta algorithm would probably not to\n   delta other objects with it.\n\n * missing commit and tag: 'stub object' needs to be extended to\n   include these object types, and we would also need 'stub\n   commit' and 'stub tag' objects, that copy the structural\n   fields from the corresponding true object.  So a stub commit\n   would probably look like:\n\n\tstub <length> NUL\n\t-----------------\n        object <object name of the real commit object>\n        type commit\n        tree <object name of the tree contained in the real commit object>\n        parent <object name of the first parent in the real commit object>\n        parent <object name of the first second in the real commit object>\n\n * missing tree would only be useful to conceal pathnames\n   recorded in the real tree object.  I am not sure if that is\n   needed.\n\n * fsck and verify-pack needs to be taught about 'stub' objects,\n   so that they know that their filenames (or the data pointed\n   at by pack .idx) do not match the result of hashing them.\n\nIf we were to do this, I suspect we can probably do nothing but\n'missing blob' first to cover a lot of ground, but we would\neventually need 'missing commit' to replace real commit objects\nthat has sensitive information in its log message.\n\nAs Nico pointed out, this has serious security implications.  We\nwould need a separate list of objects that are Ok to be stubbed\nout, with probably explanation of why they are stubbed out, and\nfsck should compare the stub objects found in the repository\nagainst that list.\n"},{"id":"35201","messageId":"alpine.LRH.0.82.0702211414410.31945@xanadu.home","threadId":"6907","inReplyTo":"7vy7mrnxlo.fsf@assigned-by-dhcp.cox.net","subject":"Re: removing content from git history","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-02-21T19:33:11Z","receivedAt":"2007-02-21T19:33:11Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 21 Feb 2007, Junio C Hamano wrote:\n\n> While I agree in principle to the argument that there is no\n> taking it back what's already published, I've heard people\n> wanting to just stop distributing further, without worrying\n> about copies already out there.  'missing objects' support would\n> help us in such a situation.\n\nI still think this is a \"put your head in the sand and pretend that some \nsensitive data never existed in the wild\" attitude.  And I really don't \nsee the point of supporting that illusion in GIT with technical means.\n\nEither you care about published data or you don't.\n\nIf you do then you are screwed anyway irrespective of any missing object \nsupport we might implement.  There will always be someone somewhere with \nthe real thing, and we all know how faster forbidden material does travel \non the Internet.\n\nIf you don't then it is just better to rewrite history and have a clean \nand unambiguous repository.  And because you don't care about existing \ncopies you shouldn't bother with the fact that the rewritten repo is not \ncompatible with the previously published one.\n\nSure rewriting history is a potentially expensive operation depending on \nthe size and nature of the change, but it is done only once.  And \nactually it can't be _that_ much expensive than a git-repack -a -f.\n\nI think it is much better to provide a tool to properly rewrite history \nthan adding support for missing objects and be stuck with them forever.\n\n\nNicolas\n"},{"id":"35203","messageId":"7vbqjnntut.fsf@assigned-by-dhcp.cox.net","threadId":"6907","inReplyTo":"alpine.LRH.0.82.0702211414410.31945@xanadu.home","subject":"Re: removing content from git history","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-21T20:22:50Z","receivedAt":"2007-02-21T20:22:50Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Wed, 21 Feb 2007, Junio C Hamano wrote:\n>\n>> While I agree in principle to the argument that there is no\n>> taking it back what's already published, I've heard people\n>> wanting to just stop distributing further, without worrying\n>> about copies already out there.  'missing objects' support would\n>> help us in such a situation.\n>\n> I still think this is a \"put your head in the sand and pretend that some \n> sensitive data never existed in the wild\" attitude.  And I really don't \n> see the point of supporting that illusion in GIT with technical means.\n\nWell, I think we are in agreement (and that is why I said \"I've\nheard people wanting\").\n\nBut it is entirely possible that somebody has a project that is\ninternal to a company managed for a long time with git, that he\nwants to go open source, with (almost) full history.  And the\nproject may have some proprietary add-on bit which cannot be\npublished, while building the public bits does not require that\npart.  Stubbing things out may help that kind of situation.  The\ndevelopment team can keep going forward, internally using the\nreal objects, while pushing stub objects out to the public\nrepository, without having to rewrite the history and re-partition\nthe project.\n\nBut after having thought about that, I think it would not buy us\nmuch.  You would want to re-partition the project sooner or\nlater in such a situation *anyway*, so our time is better spent\non giving better support to split existing projects.  It may\nalready be sufficient in the form of admin-rewritehist, in which\ncase we can worry about other things ;-).\n"},{"id":"35204","messageId":"alpine.LRH.0.82.0702211531060.31945@xanadu.home","threadId":"6907","inReplyTo":"7vbqjnntut.fsf@assigned-by-dhcp.cox.net","subject":"Re: removing content from git history","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-02-21T20:49:33Z","receivedAt":"2007-02-21T20:49:33Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 21 Feb 2007, Junio C Hamano wrote:\n\n> Well, I think we are in agreement (and that is why I said \"I've\n> heard people wanting\").\n> \n> But it is entirely possible that somebody has a project that is\n> internal to a company managed for a long time with git, that he\n> wants to go open source, with (almost) full history.  And the\n> project may have some proprietary add-on bit which cannot be\n> published, while building the public bits does not require that\n> part.  Stubbing things out may help that kind of situation.\n\nIt might help, or it might create a management nightmare.  It would be \nreally easy to accidentally push the real objects out since a repo with \nthem would be indistinguishable from a repo with stubs (that's the \npoint of stub objects isn't it?), and because of the distributed nature \nof GIT the leak could come from anyone with access to the private \nobjects.\n\nIn such a scenario I think it is still more sensible to rewrite the repo \nhistory before going open source.  You need only to worry about \nisolating the proprietary stuff once.\n\n\nNicolas\n"},{"id":"35205","messageId":"20070221210045.GB26525@spearce.org","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702210934470.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-02-21T21:00:45Z","receivedAt":"2007-02-21T21:00:45Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> Anyway, git-convert-objects does kind of give you a starting point. It \n> should be fixed to use \"git-fast-import\" or repack once in a while (so \n> that it doesn't leave tons and tons of unpacked objects), and it should be \n> fixed to fix up any commit messages that mention SHA1's that it has \n> already converted to something else, but it seems to still work. It would \n> not be impossible at all to extend the tree-rewriting logic to remove some \n> file or a particular SHA1 object you want to replace.\n\nOne idea Junio and I kicked around on #git a short while ago\nwas to arrange for a pipe between the current Git process\nand git-fast-import, where the pipe was used from within\nwrite_sha1_file() rather than creating the loose object.\n\nThis way an existing process like git-apply or git-convert-objects\ncould easily spew hundreds of thousands of objects without needing\nto worry about repacking in the middle; nor would we need to worry\nabout the complexity of trying to disentagle the multiobject packing\nparts of fast-import into some sort of library.\n\nObviously this is only a good idea if we are going to be making\nenough objects to warrant using a packfile; small 10-20 bursts\nof objects from a git-apply doesn't really justify a packfile.\nBut applying 100s of patches in a row might, if we could keep them\nall fed through the same git-fast-import backend (and thus into\nthe same packfile).\n\n-- \nShawn.\n"},{"id":"35206","messageId":"Pine.LNX.4.64.0702211306520.4043@woody.linux-foundation.org","threadId":"6907","inReplyTo":"20070221210045.GB26525@spearce.org","subject":"Re: removing content from git history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-21T21:11:50Z","receivedAt":"2007-02-21T21:11:50Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Feb 2007, Shawn O. Pearce wrote:\n> \n> One idea Junio and I kicked around on #git a short while ago\n> was to arrange for a pipe between the current Git process\n> and git-fast-import, where the pipe was used from within\n> write_sha1_file() rather than creating the loose object.\n\nThe probnlem there is that most conversion scripts that use \n\"write_sha1_file()\" will want to *read* that file later. If \ngit-fast-import hasn't generated the pack yet (because it's still waiting \nfor more data), that will not work at all.\n\nSo then you basically force the conversion script to keep remembering all \nthe old object data (using something like pretend_sha1_file), or you limit \nit to things that just always re-write the whole object and never need any \nold object references that they might have written.\n\nA lot of conversions tend to be incremental, ie they will depend on the \ndata they converted previously.\n\n\t\t\tLinus\n"},{"id":"35207","messageId":"20070221212129.GD26525@spearce.org","threadId":"6907","inReplyTo":"Pine.LNX.4.64.0702211306520.4043@woody.linux-foundation.org","subject":"Re: removing content from git history","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-02-21T21:21:30Z","receivedAt":"2007-02-21T21:21:30Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> The probnlem there is that most conversion scripts that use \n> \"write_sha1_file()\" will want to *read* that file later. If \n> git-fast-import hasn't generated the pack yet (because it's still waiting \n> for more data), that will not work at all.\n\nYes, indeed...\n \n> So then you basically force the conversion script to keep remembering all \n> the old object data (using something like pretend_sha1_file), or you limit \n> it to things that just always re-write the whole object and never need any \n> old object references that they might have written.\n> \n> A lot of conversions tend to be incremental, ie they will depend on the \n> data they converted previously.\n\nWhich is why I was actually thinking of flipping this on its head.\nLibify git-apply and embed that into fast-import, then one of the\nnative input formats might just be an mbox, or something close enough\nthat a simple C/perl/sed prefilter could make an mbox into the input.\n\nfast-import can (and does if necessary) go back to access the\npackfile it is writing.  It has the index data held in memory and\nuses only OBJ_OFS_REF so that sha1_file.c can unpack deltas just\nfine, even though we lack an index file and have not completely\nchecksummed the pack itself.\n\nSo although no other Git process can use the packfile, it is usuable\nfrom within fast-import...\n\n-- \nShawn.\n"},{"id":"55305","messageId":"18187.60305.613904.547916@lisa.zopyra.com","threadId":"6907","inReplyTo":"20070221212129.GD26525@spearce.org","subject":"Re: removing content from git history","fromName":"Bill Lear","fromEmail":"rael@zopyra.com","sentAt":"2007-10-09T20:58:57Z","receivedAt":"2007-10-09T20:58:57Z","isPatch":false,"sender":{"key":"rael@zopyra.com","avatar":"https://gravatar.com/avatar/c4f2d2790ca3828d3b4e7dfebabf61d2fe94fd82fa49cdac2a5295dd2d46a874?d=mp&s=160"},"body":"I'm resurrecting this old thread, as we have come across a similar need and\nI could not tell if this has been settled.  More below...\n\nOn Wednesday, February 21, 2007 at 16:21:30 (-0500) Shawn O. Pearce writes:\n>Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>> The probnlem there is that most conversion scripts that use \n>> \"write_sha1_file()\" will want to *read* that file later. If \n>> git-fast-import hasn't generated the pack yet (because it's still waiting \n>> for more data), that will not work at all.\n>\n>Yes, indeed...\n> \n>> So then you basically force the conversion script to keep remembering all \n>> the old object data (using something like pretend_sha1_file), or you limit \n>> it to things that just always re-write the whole object and never need any \n>> old object references that they might have written.\n>> \n>> A lot of conversions tend to be incremental, ie they will depend on the \n>> data they converted previously.\n>\n>Which is why I was actually thinking of flipping this on its head.\n>Libify git-apply and embed that into fast-import, then one of the\n>native input formats might just be an mbox, or something close enough\n>that a simple C/perl/sed prefilter could make an mbox into the input.\n>\n>fast-import can (and does if necessary) go back to access the\n>packfile it is writing.  It has the index data held in memory and\n>uses only OBJ_OFS_REF so that sha1_file.c can unpack deltas just\n>fine, even though we lack an index file and have not completely\n>checksummed the pack itself.\n>\n>So although no other Git process can use the packfile, it is usuable\n>from within fast-import...\n\nAs I understand this thread, it does not appear that a resolution\nwas reached.  Our company has content in our central git repository\nthat we need to remove per a contractual obligation.  I believe the\ncontent in question is limited to one sub-directory, that has existed\nsince (or near to) the beginning of the repo, if that matters.  We\nobviously would just like to issue a \"git nuke\" operation and be done\nwith it, if that is available.  Barring that, we could probably follow\nreasonably simple steps to purge the content and rebuild the repo.\n\nSo, what options do we have at present?\n\n\nBill\n"},{"id":"55307","messageId":"20071009210235.GB9633@fieldses.org","threadId":"6907","inReplyTo":"18187.60305.613904.547916@lisa.zopyra.com","subject":"Re: removing content from git history","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2007-10-09T21:02:35Z","receivedAt":"2007-10-09T21:02:35Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Tue, Oct 09, 2007 at 03:58:57PM -0500, Bill Lear wrote:\n> As I understand this thread, it does not appear that a resolution\n> was reached.  Our company has content in our central git repository\n> that we need to remove per a contractual obligation.  I believe the\n> content in question is limited to one sub-directory, that has existed\n> since (or near to) the beginning of the repo, if that matters.  We\n> obviously would just like to issue a \"git nuke\" operation and be done\n> with it, if that is available.  Barring that, we could probably follow\n> reasonably simple steps to purge the content and rebuild the repo.\n> \n> So, what options do we have at present?\n\nHave you looked at git-filter-branch in a recent version of git?  The\nman page has some good examples.\n\n--b.\n"},{"id":"55322","messageId":"18187.65491.551693.644510@lisa.zopyra.com","threadId":"6907","inReplyTo":"20071009210235.GB9633@fieldses.org","subject":"Re: removing content from git history","fromName":"Bill Lear","fromEmail":"rael@zopyra.com","sentAt":"2007-10-09T22:25:23Z","receivedAt":"2007-10-09T22:25:23Z","isPatch":false,"sender":{"key":"rael@zopyra.com","avatar":"https://gravatar.com/avatar/c4f2d2790ca3828d3b4e7dfebabf61d2fe94fd82fa49cdac2a5295dd2d46a874?d=mp&s=160"},"body":"On Tuesday, October 9, 2007 at 17:02:35 (-0400) J. Bruce Fields writes:\n>On Tue, Oct 09, 2007 at 03:58:57PM -0500, Bill Lear wrote:\n>> As I understand this thread, it does not appear that a resolution\n>> was reached.  Our company has content in our central git repository\n>> that we need to remove per a contractual obligation.  I believe the\n>> content in question is limited to one sub-directory, that has existed\n>> since (or near to) the beginning of the repo, if that matters.  We\n>> obviously would just like to issue a \"git nuke\" operation and be done\n>> with it, if that is available.  Barring that, we could probably follow\n>> reasonably simple steps to purge the content and rebuild the repo.\n>> \n>> So, what options do we have at present?\n>\n>Have you looked at git-filter-branch in a recent version of git?  The\n>man page has some good examples.\n\nAh, no, though I will do so.  It is apparently not in the version\nI have (1.5.2.4), but it is in 1.5.3.1.  We'll give this a shot\nand complain if we can't handle it.\n\nThank you.\n\n\nBill\n"},{"id":"55374","messageId":"Pine.LNX.4.64.0710101535560.4174@racer.site","threadId":"6907","inReplyTo":"18187.60305.613904.547916@lisa.zopyra.com","subject":"Re: removing content from git history","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-10-10T14:41:21Z","receivedAt":"2007-10-10T14:41:21Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 9 Oct 2007, Bill Lear wrote:\n\n> Our company has content in our central git repository that we need to \n> remove per a contractual obligation.  I believe the content in question \n> is limited to one sub-directory, that has existed since (or near to) the \n> beginning of the repo, if that matters.  We obviously would just like to \n> issue a \"git nuke\" operation and be done with it, if that is available.  \n> Barring that, we could probably follow reasonably simple steps to purge \n> the content and rebuild the repo.\n> \n> So, what options do we have at present?\n\ngit filter-branch.  I suggest using the index filter.  There is even a \nnice example in the man page of git filter-branch.\n\nWhich reminds me that I have some TODOs left in filter-branch...\n\nCiao,\nDscho\n"}]}