{"thread":{"id":"4527","subject":"Re: Security problem","startedAt":"2006-06-16T00:12:53Z","lastAt":"2006-06-16T08:18:46Z","messageCount":8,"participants":["Junio C Hamano","Linus Torvalds","Alexander Litvinov"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"21862","messageId":"7vbqsuc60q.fsf@assigned-by-dhcp.cox.net","threadId":"4527","inReplyTo":"200606151709.22752.lan@academsoft.ru","subject":"Re: Security problem","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-06-16T00:12:53Z","receivedAt":"2006-06-16T00:12:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Alexander Litvinov <lan@academsoft.ru> writes:\n\n> Why does not git-checkout check if file content match name of the object ?\n\nGood point.  We could do a few things:\n\n - entry.c:write_entry() could validate after read_sha1_file(). \n\n - read_sha1_file() could do the checking; this has performance\n   implications, though.\n\nCloning over git aware protocols validate the objects coming\nover the wire, so it may make sense to cheat and do the former,\nso that we do not have to pay the validation cost every time we\naccess any object.\n"},{"id":"21864","messageId":"Pine.LNX.4.64.0606151831470.5498@g5.osdl.org","threadId":"4527","inReplyTo":"7vbqsuc60q.fsf@assigned-by-dhcp.cox.net","subject":"Re: Security problem","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-16T02:28:43Z","receivedAt":"2006-06-16T02:28:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 15 Jun 2006, Junio C Hamano wrote:\n>\n> Alexander Litvinov <lan@academsoft.ru> writes:\n> \n> > Why does not git-checkout check if file content match name of the object ?\n> \n> Good point.  We could do a few things:\n\nI missed the original mail. What's the problem?\n\nIf this is about the remote end lying about the SHA1 name, it's a total \nnon-issue for any of the native protocols, since the native protocols \ndon't actually send SHA1 names at all, they just send the data (and we \nre-create the SHA1 name on receipt).\n\nSo there's no way to have the name of an object not match its content, \nunless you have actual corruption (which is for git-fsck-object to find, \nnot somethign that should slow down any normal operation), or if you use \none of the dumb protocols.\n\nAnd if you use the dumb protocols, the data should probably be validated \n_there_ (by fetch(), rather than anywhere else). And for \"rsync\", you \nreally don't have much choice apart from doing a full fsck, I suspect.\n\nSo I don't see the security issue, unless you don't trust the local \nfilesystem, in which case nothing git can do matters at all..\n\n\t\tLinus\n"},{"id":"21866","messageId":"Pine.LNX.4.64.0606151948230.5498@g5.osdl.org","threadId":"4527","inReplyTo":"200606160931.29553.lan@academsoft.ru","subject":"Re: Security problem","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-16T02:56:49Z","receivedAt":"2006-06-16T02:56:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 16 Jun 2006, Alexander Litvinov wrote:\n> \n> I have found the ability to hack git repo. After this hacking people will \n> checkout hacked files from the \"trusted\" commit. Only git-fsck-objects will \n> complain at this.\n\nRight.\n\nIf you can't trust your local filesystem, you are screwed. \n\ngit-fsck-objects will notice when somebody has done something bad, but \n\n> Why does not git-checkout check if file content match name of the object ?\n\nWhy would it? It really just slows things down, and if you don't trust \nyour local repo, people can \"hack\" you much more easily by just generating \na _proper_ tree with the _proper_ data, and git checkout checking the SHA1 \nwouldn't help at all.\n\nThe way to security lies in using git-fsck-objects, together with an \n_external_ source of trust. For example, that external source of trust may \nbe a signed tag, or, perhaps even more simply, just by saving off the top \ncommit name on some trusted medium.\n\nBut you do need a \"point of trust\" to start with. Without that, it's a lot \neasier to \"hack\" a git repo by doing\n\n\techo 'Hacked file' > a\n\tgit commit --amend a\n\tgit prune\n\nand now the file \"a\" has changed to \"Hacked file\", and even \ngit-fsck-objects can't tell that anything bad happened.\n\n(Btw, if you want to _hide_ the fact that \"a\" now contains \"Hacked file\", \nyou do so by faking it in the index. You can have the checked-out copy say \nwhat it should say - ie \"Usual file\" - and if you don't want git to show \nyou the difference to HEAD, you edit the .git/index file by hand so that \nthe timestamp, size and inode matches the real SHA1, even though the \n_contents_ match \"Usual file\").\n\nSee?\n\nYou do need to trust something. Normally you'd trust your own filesystem, \nbut git certainly supports other forms of trust through either the native \nsupport for signed certificates in the form of tags, or any other form of \nexternal trust.\n\n\t\t\tLinus\n"},{"id":"21867","messageId":"200606161054.46813.lan@academsoft.ru","threadId":"4527","inReplyTo":"Pine.LNX.4.64.0606151948230.5498@g5.osdl.org","subject":"Re: Security problem","fromName":"Alexander Litvinov","fromEmail":"lan@academsoft.ru","sentAt":"2006-06-16T03:54:46Z","receivedAt":"2006-06-16T03:54:46Z","isPatch":false,"sender":{"key":"lan@academsoft.ru","avatar":null},"body":"> If you can't trust your local filesystem, you are screwed.\n\nYou are right, I trust my file system. But if our team had central repo with \nssh access to that machine, every developer can hack central repo.\n\nWhould git-clone/git-fetch warn me about this ?\n\nMy own test with (another) local repo says:\nlan@lan:~/tmp/git/test> git clone 1 2\nGenerating pack...\nDone counting 3 objects.\nDeltifying 3 objects.\n 100% (3/3) done\nTotal 3, written 3 (delta 0), reused 0 (delta 0)\nerror: git-checkout-index: unable to read sha1 file of a \n(3609f20ebd357679b111783e8afaf36ec46427f3)\n\nIt can't checkout object (3609f20ebd357679b111783e8afaf36ec46427f3 is the \noriginal file). It seems packed repos are safe from this point.\n"},{"id":"21868","messageId":"Pine.LNX.4.64.0606152137410.5498@g5.osdl.org","threadId":"4527","inReplyTo":"200606161054.46813.lan@academsoft.ru","subject":"Re: Security problem","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-16T05:00:39Z","receivedAt":"2006-06-16T05:00:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 16 Jun 2006, Alexander Litvinov wrote:\n>\n> You are right, I trust my file system. But if our team had central repo with \n> ssh access to that machine, every developer can hack central repo.\n> \n> Whould git-clone/git-fetch warn me about this ?\n\nUsing the native protocol, yes. Using rsync, unless you explicitly fsck \nthe result, no.\n\n> It can't checkout object (3609f20ebd357679b111783e8afaf36ec46427f3 is the \n> original file). It seems packed repos are safe from this point.\n\nWell, they may not be \"safe\" - you just need to work a _lot_ harder to \ncorrupt a pack-file in any interesting manner. And again, git-fsck-objects \nwould pick up any such thing going on.\n\nAnyway, what it boils down to is that anybody who has write access to a \nparticular repository can certainly change the repo in \"interesting\" ways. \n\nHowever, there are various inherent safety valves in place that make it \nreally hard to corrupt on a bigger scale.\n\nThe first is that git-fsck-objects will definitely find any repository \ninconsistency, and to get around that, you either have to get around the \nbasic properties of SHA-1 (ie break the hash) _or_ you have to actually \nchange the repository so that it's still a valid repo, just with different \ncontent.\n\nSo let's take a look at those two cases:\n\n - if you corrupt the repository, subsequent clones (or even pulls) from \n   the corrupt repository simply won't work if you use the native \n   protocol, because the native protocol doesn't actually trust anything \n   but the actual contents (so if the contents won't match, then neither \n   will the SHA1 names). So the corruption is pretty strictly limited to \n   the _one_ repository that the attacker had write access to.\n\n   So there's a pretty fundamental \"corruption containment\" part there.\n\n   (Side note: there's no question that we might well be able to do \n   better. A _malicious_ server could actually send a corrupt pack, and \n   it's possible that a properly corrupted remote archive could cause even \n   a \"good\" git-send-pack to just silently send a corrupt pack, so that \n   you'd need to use \"git-fsck-objects\" on the receiving side to notice \n   that you are missing objects, for example)\n\n - if the repository is good (ie fsck is fine), then obviously a \"git \n   pull\" will also succeed. However, you can't _hide_ the data the way you \n   tried to do: when the receiver checks out the most recent version, it \n   will definitely use the data in the object, there's no way to get the \n   server to serve different data in objects and in the working tree \n   (because the server literally doesn't even send the working tree at \n   all).\n\n   So you can always convince somebody to pull from an \"evil repository\", \n   and that's no different from committing a bug by mistake. But at least \n   you can't try to hide the bug just in the object store and have it not \n   show up in diffs and in checked-out copies.\n\nThe latter case is true even with http and rsync, the actual pull event \nalways pulls just the database, never any checked-out state (in fact, \nthe common case is obviously to pull from a bare repository that doesn't \neven _have_ checked-out state). So you can't hide things in the index or \nin the checked-out state except in the filesystem that you have direct \nwrite access to.\n\nBut yeah, I actually still personally do a fair number of \n\"git-fsck-objects\". I've never found anything that way since very early on \n(and back then, the real problem was rsync getting objects that weren't \nreachable), but I still do it. It makes me feel happier.\n\nOf course, bugs always happen. But I can pretty much guarantee that git is \nfundamentally harder to corrupt than most things. We've had git-fsck-cache \nsince April 8th last year (or, put another way, literally since \"Day 2\" in \ngit terms - it's the eight commit in the whole git history).\n\nGit also has an almost total lack of redundant information. There's \nbasically no \"duplicate\" information in the repository format itself where \nyou could hide something so that it wouldn't be noticed.\n\nIn a checked-out project, the checked-out state itself is \"duplicate \ninformation\" (and that was where your \"attack\" tried to hide things), and \nthere's the index (which is actually a much better and subtle place to \nhide things ;). But neither of them have any life outside of that \nparticular repository.\n\n\t\t\tLinus\n"},{"id":"21869","messageId":"200606161237.21997.lan@academsoft.ru","threadId":"4527","inReplyTo":"Pine.LNX.4.64.0606152137410.5498@g5.osdl.org","subject":"Re: Security problem","fromName":"Alexander Litvinov","fromEmail":"lan@academsoft.ru","sentAt":"2006-06-16T05:37:21Z","receivedAt":"2006-06-16T05:37:21Z","isPatch":false,"sender":{"key":"lan@academsoft.ru","avatar":null},"body":"> Well, they may not be \"safe\" - you just need to work a _lot_ harder to\n> corrupt a pack-file in any interesting manner. And again, git-fsck-objects\n> would pick up any such thing going on.\nAs it shown in pack-objects.c, each object have stored sha1, almost the same \nas file rename.\n\n> The first is that git-fsck-objects will definitely find any repository\n> inconsistency, and to get around that, you either have to get around the\n> basic properties of SHA-1 (ie break the hash) _or_ you have to actually\n> change the repository so that it's still a valid repo, just with different\n> content.\nI still belive SHA-1 is good enouth to hash files - I did not hear about \ngeneration reasonable duplicate that can compile and work :-)\n\n>  - if you corrupt the repository, subsequent clones (or even pulls) from\n>    the corrupt repository simply won't work if you use the native\n>    protocol, because the native protocol doesn't actually trust anything\n>    but the actual contents (so if the contents won't match, then neither\n>    will the SHA1 names). So the corruption is pretty strictly limited to\n>    the _one_ repository that the attacker had write access to.\nAs I understand sent pack file will contains actial SHA-1 of objects. And any \nhack will be cleary visible.\n\n>    So there's a pretty fundamental \"corruption containment\" part there.\n...\nSituation with evil repo is clear to me: you can turst only to trusted commit \nidentified by SHA-1\n\n> But yeah, I actually still personally do a fair number of\n> \"git-fsck-objects\". I've never found anything that way since very early on\n> (and back then, the real problem was rsync getting objects that weren't\n> reachable), but I still do it. It makes me feel happier.\nAs the result: Always fsck repo after pull/clone !\n"},{"id":"21871","messageId":"Pine.LNX.4.64.0606152300460.5498@g5.osdl.org","threadId":"4527","inReplyTo":"200606161237.21997.lan@academsoft.ru","subject":"Re: Security problem","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-16T06:27:27Z","receivedAt":"2006-06-16T06:27:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 16 Jun 2006, Alexander Litvinov wrote:\n>\n> > Well, they may not be \"safe\" - you just need to work a _lot_ harder to\n> > corrupt a pack-file in any interesting manner. And again, git-fsck-objects\n> > would pick up any such thing going on.\n>\n> As it shown in pack-objects.c, each object have stored sha1, almost the same \n> as file rename.\n\nYes and no.\n\nThe index file has the stored sha1 (and in that sense you can do almost \nthe same thing as a file rename by just modifying the index file).\n\nBut when we actually transfer a pack over from one place to another (ie a \nclone or a push), we don't even transfer the index file. Instead, the \nindex file gets re-generated at the other end.\n\nThat's pretty much an on-going theme in most of git - trying to avoid \nhaving metadata, if that can instead of calculated directly.\n\nSo again, a \"rsync\" or a \"http\" thing that just gets the index and \npack-files directly _as_files_, will actually also download a corrupt \nfile. The git native protocol is much harder to fool.\n\ngit-fsck-objects actually verifies the pack-files and index files in \nseveral ways:\n\n - both the pack-file and the index-file actually contain a SHA1 checksum \n   of themselves, so any accidental corruption will be picked up (but if \n   somebody is able to get at the filesystem, they can obviously \n   re-calculate the SHA1 and update the checksum too)\n\n - the index file also contains the SHA-1 of the pack-file (and that is \n   then part of the checksum of the index file), again to avoid accidental \n   corruption or mixing of index and pack-files.\n\n - fsck checks all of these internal SHA-1 checksums, and verifies basic \n   information (ie number of objects must match etc)\n\n - each object in the index file is unpacked, and its SHA-1 is \n   re-calculated and checked against what the index file claimed.\n\nSo exactly as with individual objects, the pack-files are actually \nverified, and on (native-mode) transfer, the names of individual files are \nnever actually transferred, rather they are re-calculated from the raw \ncontents at the receiving end.\n\nThe pack-files then have a few additional sanity-checks of their own that \nshould help pinpoint at least the accidental kind of corruption.\n\nBut no, the SHA1 checksums of the pack-files are not checked by normal \noperations. That would be deadly - trying to check the SHA1 hash of a \npack-file obviously would involve reading it all in, something normal \noperations actually try to avoid (normal ops use the index exactly in \norder to only read the parts they need).\n\nPerhaps most importantly, after fsck has checked the SHA-1's of each \nindividual object, it will also do a full reachability check. That, in \nmany ways, is even more important than checking that each object name \nmatches its contents (ie there's no missing history either, and the \n\"tips\" of the repository end up basically validating all the rest).\n\nSo again, the thing is set up so that doing a full fsck actually does a \n_lot_ of integrity checking.\n\nBut in the absense of explicit fsck, we do trust the data, even if the \nactual _transfer_ of data will recalculate SHA-1's.\n\n> >  - if you corrupt the repository, subsequent clones (or even pulls) from\n> >    the corrupt repository simply won't work if you use the native\n> >    protocol, because the native protocol doesn't actually trust anything\n> >    but the actual contents (so if the contents won't match, then neither\n> >    will the SHA1 names). So the corruption is pretty strictly limited to\n> >    the _one_ repository that the attacker had write access to.\n>\n> As I understand sent pack file will contains actial SHA-1 of objects. And any \n> hack will be cleary visible.\n\nNo, as mentioned, the actual SHA-1's won't ever be sent, so what happens \nis that if the repository on the sending side was hacked, the _sending_ \nside may never even realize it (since it's not necessarily checking the \nSHA-1's), but the receiving side will only ever see the raw data, and as \nsuch, it won't ever even _see_ the \"false hidden names\", because it will \ngenerate a whole new index that purely depends on the data.\n\nAnd maybe that's exactly what you meant - yes, the hack will be clearly \nvisible, because the names will now be the \"real\" ones. You can't hide \nthings by using a false name.\n\n> >    So there's a pretty fundamental \"corruption containment\" part there.\n> ...\n> Situation with evil repo is clear to me: you can turst only to trusted commit \n> identified by SHA-1\n\nYes. Exactly.\n\nAnd once you have a reason to trust a commit, everything you can reach \nfrom that commit is also trustworthy, assuming it passes fsck. IOW, you \nonly really need to trust the head(s) in your repository.\n\n> > But yeah, I actually still personally do a fair number of\n> > \"git-fsck-objects\". I've never found anything that way since very early on\n> > (and back then, the real problem was rsync getting objects that weren't\n> > reachable), but I still do it. It makes me feel happier.\n>\n> As the result: Always fsck repo after pull/clone !\n\nWell, even better, try to avoid pulling from untrusted sources in the \nfirst place ;)\n\nBut yes, fsck is actually fairly fast if you do incremental pulls and \nrepack your repository. To help you do this, there's two modes to fsck: \nthere's the \"full mode\", which goes through _everything_, including \npack-files, and there's the \"fsck only lose objects\", which is the common \none.\n\nSo for example, let's say that you only ever repack your repository \nlocally when it's been \"known good\" (in fact, repacking in itself will \ngenerally find almost all of the problems that fsck can find, since a full \nrepack will obviously do the reachability analysis as part of just the \npreparatory work). That means that you only ever need to do the quick \ndefault \"light fsck\" after a pull, since an incremental pull (with the \nnative protocol) will have unpacked all the pulled objects.\n\nSo \"fsck after each pull\" is not something we do by default, but if you \nkeep your repo fairly packed, doing so manually (or by just scripting \nthings) won't even really slow you down, because it will only ever need to \ncheck incrementally - the stuff you've re-packed it doesn't need to check \n(assuming you can now trust your local filesystem).\n\nSo git certainly gives you the option to be really anal, and doesn't even \nmake it needlessly hard or expensive, even with large repositories.\n\n\t\t\tLinus\n"},{"id":"21874","messageId":"200606161518.47149.lan@academsoft.ru","threadId":"4527","inReplyTo":"Pine.LNX.4.64.0606152300460.5498@g5.osdl.org","subject":"Re: Security problem","fromName":"Alexander Litvinov","fromEmail":"lan@academsoft.ru","sentAt":"2006-06-16T08:18:46Z","receivedAt":"2006-06-16T08:18:46Z","isPatch":false,"sender":{"key":"lan@academsoft.ru","avatar":null},"body":"> So git certainly gives you the option to be really anal, and doesn't even\n> make it needlessly hard or expensive, even with large repositories.\n\nThanks for detailed description. Now I can sleep without any worry about my \nrepo :-)\n"}]}