{"thread":{"id":"20144","subject":"encrypted repositories?","startedAt":"2009-07-17T15:14:25Z","lastAt":"2012-08-02T14:52:52Z","messageCount":17,"participants":["Matthias Andree","Michael J Gruber","Matthias Kestenholz","Linus Torvalds","Jakub Narebski","John Tapsell","Thomas Koch","Jeff King","J-S-B"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"118177","messageId":"op.uw7wmbr41e62zd@balu.cs.uni-paderborn.de","threadId":"20144","inReplyTo":null,"subject":"encrypted repositories?","fromName":"Matthias Andree","fromEmail":"matthias.andree@gmx.de","sentAt":"2009-07-17T15:14:25Z","receivedAt":"2009-07-17T15:14:25Z","isPatch":false,"sender":{"key":"matthias.andree@gmx.de","avatar":null},"body":"Greetings,\n\nI have a rather special usage scenario.\n\nAssume you have a repository where you want to work on embargoed  \ninformation, so that not even system administrators of the server you're  \npushing to can get a hold of the cleartext data.\n\n\"Server\" would be a central reference repository that I can push to.\n\"Client\" would by my working computer that has a clone of the crypted  \nrepo, and an unencrypted checkout of it. Perhaps the client would also  \nneed an unencrypted copy of the repo (for performance reasons, I'm not  \nsure about that) that gets encrypted on the fly when pushing and decrypted  \nwhen fetching.\n\nExamples of use might be press releases of upcoming products, written  \nexams for students, whatever.\n\nRequirements:\n- \"client\" that is about to push must encrypt the data before pushing it  \nto the server.\n- all data (including file names, log messages,\n\nAllowed restrictions:\n- \"server\" limited to bare repositories\n- initial version limited to symmetric encryption with pre-shared secret\n\nIn a later step, some key management and asymmetric crypto would be  \nuseful, but that's not crucial now. In my current scenario, those who are  \nworking on the embargoed material would trust one another.\n\n\nHow would one go about this from the user side? I sincerely doubt I have  \nthe resources (time!) to actually implement this in Git.\n\nTIA\n\n-- \nMatthias Andree\n"},{"id":"118181","messageId":"4A60A168.2060105@drmicha.warpmail.net","threadId":"20144","inReplyTo":"op.uw7wmbr41e62zd@balu.cs.uni-paderborn.de","subject":"Re: encrypted repositories?","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2009-07-17T16:06:00Z","receivedAt":"2009-07-17T16:06:00Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Matthias Andree venit, vidit, dixit 17.07.2009 17:14:\n> Greetings,\n> \n> I have a rather special usage scenario.\n> \n> Assume you have a repository where you want to work on embargoed  \n> information, so that not even system administrators of the server you're  \n> pushing to can get a hold of the cleartext data.\n> \n> \"Server\" would be a central reference repository that I can push to.\n> \"Client\" would by my working computer that has a clone of the crypted  \n> repo, and an unencrypted checkout of it. Perhaps the client would also  \n> need an unencrypted copy of the repo (for performance reasons, I'm not  \n> sure about that) that gets encrypted on the fly when pushing and decrypted  \n> when fetching.\n> \n> Examples of use might be press releases of upcoming products, written  \n> exams for students, whatever.\n> \n> Requirements:\n> - \"client\" that is about to push must encrypt the data before pushing it  \n> to the server.\n> - all data (including file names, log messages,\n> \n> Allowed restrictions:\n> - \"server\" limited to bare repositories\n> - initial version limited to symmetric encryption with pre-shared secret\n> \n> In a later step, some key management and asymmetric crypto would be  \n> useful, but that's not crucial now. In my current scenario, those who are  \n> working on the embargoed material would trust one another.\n> \n> \n> How would one go about this from the user side? I sincerely doubt I have  \n> the resources (time!) to actually implement this in Git.\n\nIf the server can not decrypt anything then it can not serve anything,\nat least not as a git server. Note that if you're really fussy about\nsecurity then you should not allow the server to see even the DAG (which\nwould be the case if you encrypt blobs only), which makes it impossible\nto do any smart serving.\n\nSo, why not share some form of remote storage on which you have an\nencrypted luks partition? That way you can even set up multiple access\nkeys and revoke them when necessary.\n\nMichael\n"},{"id":"118184","messageId":"1f6632e50907170930w4860b841i3a6028ea47c4d522@mail.gmail.com","threadId":"20144","inReplyTo":"op.uw7wmbr41e62zd@balu.cs.uni-paderborn.de","subject":"Re: encrypted repositories?","fromName":"Matthias Kestenholz","fromEmail":"mk@feinheit.ch","sentAt":"2009-07-17T16:30:39Z","receivedAt":"2009-07-17T16:30:39Z","isPatch":false,"sender":{"key":"mk@feinheit.ch","avatar":"https://gravatar.com/avatar/f4f02a5336cf0e3d40b05498959e997f023cb5d8c83ab41545a3272268c67949?d=mp&s=160"},"body":"On Fri, Jul 17, 2009 at 3:14 PM, Matthias Andree<matthias.andree@gmx.de> wrote:\n>\n> How would one go about this from the user side? I sincerely doubt I have the\n> resources (time!) to actually implement this in Git.\n>\n\nMaybe you could send around packages created by git-bundle as\npgp-encrypted emails, and keep the original on your computer on an\nencrypted filesystem? You could also encrypt the bundles and put them\nonto a public server if that suits you better than email.\n\nThe biggest advantage of this approach would be that it can be done\nnow, and no changes to git itself are required.\n\n\nMatthias\n"},{"id":"118194","messageId":"alpine.LFD.2.01.0907171226460.13838@localhost.localdomain","threadId":"20144","inReplyTo":"op.uw7wmbr41e62zd@balu.cs.uni-paderborn.de","subject":"Re: encrypted repositories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-07-17T19:38:16Z","receivedAt":"2009-07-17T19:38:16Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 17 Jul 2009, Matthias Andree wrote:\n>\n> Assume you have a repository where you want to work on embargoed information,\n> so that not even system administrators of the server you're pushing to can get\n> a hold of the cleartext data.\n\nIf the server can't ever read it, you're basically limited to just one \nstory:\n\n - use rsync-like \"stupid\" transports to upload and download things.\n\n - a \"smart\" git server (eg the native git:// style protocol is not going \n   to be possible)\n\nand you strictly speaking need no real git changes, because you might as \nwell just do it by uploading an encrypted tar-file of the .git directory. \nAnd there is literally no upside in doing anything else - any native git \nsupport is almost entirely pointless.\n\nYou could make it a _bit_ more useful perhaps by adding some helper \nwrappers, probably by just implementing a new transport name (ie instead \nof using \"rsync://\", you'd just use \"crypt-tgz://\" or something).\n\nNow, that said, there are probably situations where maybe you'd allow the \nserver to decrypt things _temporarily_, but you don't want to be encrypted \non disk, and no persistent keys on the server, then that would open up a \nlot more possibilities.\n\nOf course, that still does require that you trust the server admin to \n_some_ degree - anybody who has root would be able to get the keys by \nrunning a debugger on the git upload/download sequence when you do a \nupload or download.\n\nMaybe that kind of security is still acceptable to you, though? \n\nIF that is the case, then at least in theory we could add support for \n\"encryption key exchange\" to the native git protocol, and then you could \nhave encryption over the network access (ssh obviously already does that, \nbut I'm including things like the anonymous git:// protocol too), and \nyou'd have encrypted data on disk, but git-upload-pack would be able to \ndecrypt things in order to do deltas etc.\n\nBut see above: in order for that to work, you do have to allow the pack \nupload and download processes on the server to decrypt things (in memory). \nSo it would not be \"absolutely secure\".\n\n\t\t\tLinus\n"},{"id":"118196","messageId":"m3skgvt5zi.fsf@localhost.localdomain","threadId":"20144","inReplyTo":"4A60A168.2060105@drmicha.warpmail.net","subject":"Re: encrypted repositories?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-07-17T20:22:15Z","receivedAt":"2009-07-17T20:22:15Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Michael J Gruber <git@drmicha.warpmail.net> writes:\n> Matthias Andree venit, vidit, dixit 17.07.2009 17:14:\n> > \n> > I have a rather special usage scenario.\n> > \n> > Assume you have a repository where you want to work on embargoed  \n> > information, so that not even system administrators of the server you're  \n> > pushing to can get a hold of the cleartext data.\n> > \n> > \"Server\" would be a central reference repository that I can push to.\n> > \"Client\" would by my working computer that has a clone of the crypted  \n> > repo, and an unencrypted checkout of it. Perhaps the client would also  \n> > need an unencrypted copy of the repo (for performance reasons, I'm not  \n> > sure about that) that gets encrypted on the fly when pushing and decrypted  \n> > when fetching.\n> > \n> > Examples of use might be press releases of upcoming products, written  \n> > exams for students, whatever.\n \n> If the server can not decrypt anything then it can not serve anything,\n> at least not as a git server. Note that if you're really fussy about\n> security then you should not allow the server to see even the DAG (which\n> would be the case if you encrypt blobs only), which makes it impossible\n> to do any smart serving.\n\nThere was shown here on git mailing list script which was meant to\nhelp in situation where you have repository with sensitive deta,\ndiscovered repository corruption or bug in git, and cannot be\nreproduced otherwise.  But I think it didn't encrypt repositry, but\njust emulate it's structure.\n\nAs to encrypting repository: you can encrypt blobs (content of files),\nyou can encrypt filenames (but the structure remains) or you can put\nfiles in a flat encrypted structure, and you can encrypt commit\nmessages and comitter and author info, and encrypt / rename branch\nnames.  Still some DAG structure will be visible, and need be visible\nfor \"smart\" git server (access via ssh and git protocols) to work.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"118197","messageId":"43d8ce650907171322y60aaa0f3na335b7a4a2fe32c1@mail.gmail.com","threadId":"20144","inReplyTo":"alpine.LFD.2.01.0907171226460.13838@localhost.localdomain","subject":"Re: encrypted repositories?","fromName":"John Tapsell","fromEmail":"johnflux@gmail.com","sentAt":"2009-07-17T20:22:47Z","receivedAt":"2009-07-17T20:22:47Z","isPatch":false,"sender":{"key":"johnflux@gmail.com","avatar":"https://gravatar.com/avatar/25f70d4c0f96396b84a2e34bcd9bdc233462c7b4be29b5fdca8266fc53f30b0c?d=mp&s=160"},"body":"2009/7/17 Linus Torvalds <torvalds@linux-foundation.org>:\n>\n>\n> On Fri, 17 Jul 2009, Matthias Andree wrote:\n>>\n>> Assume you have a repository where you want to work on embargoed information,\n>> so that not even system administrators of the server you're pushing to can get\n>> a hold of the cleartext data.\n>\n> If the server can't ever read it, you're basically limited to just one\n> story:\n\nWhy couldn't you have the actual code encrypted, but have the server\nstill know about the SHAs etc?  You would expose the actual commit\nstructure, but that might be acceptable?\n\nJohn\n"},{"id":"118198","messageId":"alpine.LFD.2.01.0907171337320.13838@localhost.localdomain","threadId":"20144","inReplyTo":"43d8ce650907171322y60aaa0f3na335b7a4a2fe32c1@mail.gmail.com","subject":"Re: encrypted repositories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-07-17T20:40:43Z","receivedAt":"2009-07-17T20:40:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 17 Jul 2009, John Tapsell wrote:\n> \n> Why couldn't you have the actual code encrypted, but have the server\n> still know about the SHAs etc?  You would expose the actual commit\n> structure, but that might be acceptable?\n\nEven that wouldn't really work, because you'd never be able to generate \nany deltas.\n\nSo there would be no real advantage. In fact, there would be only \ndisadvantages, because without any delta generation, you'd now have to \nactually transfer _more_ data.\n\n\t\t\tLinus\n"},{"id":"118199","messageId":"alpine.LFD.2.01.0907171341040.13838@localhost.localdomain","threadId":"20144","inReplyTo":"alpine.LFD.2.01.0907171337320.13838@localhost.localdomain","subject":"Re: encrypted repositories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-07-17T20:42:36Z","receivedAt":"2009-07-17T20:42:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 17 Jul 2009, Linus Torvalds wrote:\n> \n> On Fri, 17 Jul 2009, John Tapsell wrote:\n> > \n> > Why couldn't you have the actual code encrypted, but have the server\n> > still know about the SHAs etc?  You would expose the actual commit\n> > structure, but that might be acceptable?\n> \n> Even that wouldn't really work, because you'd never be able to generate \n> any deltas.\n> \n> So there would be no real advantage. In fact, there would be only \n> disadvantages, because without any delta generation, you'd now have to \n> actually transfer _more_ data.\n\nOh, if you let the server know all the SHA's at _all_ levels (ie down to \nthe blob itself), and then just make the blobs be encrypted, we'd be able \nto do some trivial optimizations, like only sending the actual blobs that \nchanged. HOWEVER. That would reveal absolutely tons of data about the \nrepository, and about the history. You'd have lost a _lot_ of security.\n\n\t\t\tLinus\n"},{"id":"118244","messageId":"200907182109.31275.thomas@koch.ro","threadId":"20144","inReplyTo":"alpine.LFD.2.01.0907171341040.13838@localhost.localdomain","subject":"Re: encrypted repositories? with git-torrent?","fromName":"Thomas Koch","fromEmail":"thomas@koch.ro","sentAt":"2009-07-18T19:09:30Z","receivedAt":"2009-07-18T19:09:30Z","isPatch":false,"sender":{"key":"thomas@koch.ro","avatar":null},"body":"Wouldn't this be a use case for git-torrent?\nhttp://code.google.com/p/gittorrent/\nhttp://repo.or.cz/w/VCS-Git-Torrent.git\n\nAs I understand it, all data would be stored decentraliced and the (optional?) \ncentral server only saves, who has which objects.\n\n(Just hearing a podcast on bittorrent while reading GIT mailinglist :-)\n\nThomas Koch, http://www.koch.ro\n"},{"id":"118306","messageId":"op.uxc712eh1e62zd@balu.cs.uni-paderborn.de","threadId":"20144","inReplyTo":"alpine.LFD.2.01.0907171226460.13838@localhost.localdomain","subject":"Re: encrypted repositories?","fromName":"Matthias Andree","fromEmail":"matthias.andree@gmx.de","sentAt":"2009-07-20T12:09:28Z","receivedAt":"2009-07-20T12:09:28Z","isPatch":false,"sender":{"key":"matthias.andree@gmx.de","avatar":null},"body":"Am 17.07.2009, 21:38 Uhr, schrieb Linus Torvalds  \n<torvalds@linux-foundation.org>:\n\n>\n>\n> On Fri, 17 Jul 2009, Matthias Andree wrote:\n>>\n>> Assume you have a repository where you want to work on embargoed  \n>> information,\n>> so that not even system administrators of the server you're pushing to  \n>> can get\n>> a hold of the cleartext data.\n>\n> If the server can't ever read it, you're basically limited to just one\n> story:\n>\n>  - use rsync-like \"stupid\" transports to upload and download things.\n>\n>  - a \"smart\" git server (eg the native git:// style protocol is not going\n>    to be possible)\n\nI don't know all its features, apparently it's online recompression - this  \nis no longer going to be available.\n\n> and you strictly speaking need no real git changes, because you might as\n> well just do it by uploading an encrypted tar-file of the .git directory.\n> And there is literally no upside in doing anything else - any native git\n> support is almost entirely pointless.\n>\n> You could make it a _bit_ more useful perhaps by adding some helper\n> wrappers, probably by just implementing a new transport name (ie instead\n> of using \"rsync://\", you'd just use \"crypt-tgz://\" or something).\n>\n> Now, that said, there are probably situations where maybe you'd allow the\n> server to decrypt things _temporarily_, but you don't want to be  \n> encrypted\n> on disk, and no persistent keys on the server, then that would open up a\n> lot more possibilities.\n\n> Of course, that still does require that you trust the server admin to\n> _some_ degree - anybody who has root would be able to get the keys by\n> running a debugger on the git upload/download sequence when you do a\n> upload or download.\n>\n> Maybe that kind of security is still acceptable to you, though?\n\nNo, the server can't be allowed access to the keys or decrypted data.\n\nI'm not sure about the graph, and if I should be concerned. Exposing the  \nDAG might be in order.\n\nIt would be ok if the disk storage and the over-the-wire format cannot use  \ndelta compression then. It would suffice to just send a set of objects  \nefficiently - and perhaps smaller revisions can be delta-compressed by the  \nclients when pushing.\n\nI admit haven't checked how the current git:// over-the-wire protocol[s]  \nwork[s]. I think client-side delta compression may require limiting the  \ngraph depths or delta size (when exceeded, the client must send the  \nstandalone self-contained object rather than a delta), so that the server  \ncan refuse patches when the delta nesting or size gets too deep/big.\n\nI think this would generate the git server to something like a storage  \ndevice for objects, perhaps with the DAG if exposed.\n\n\nOn a more general note, is someone looking into improving the http://  \nefficiency?  Perhaps there are synergies between my plan of (a) encryption  \nand (b) more efficient \"dumb\" (http/rsync/...) protocol use.\n\n-- \nMatthias Andree\n"},{"id":"118308","messageId":"op.uxc78tdp1e62zd@balu.cs.uni-paderborn.de","threadId":"20144","inReplyTo":"200907182109.31275.thomas@koch.ro","subject":"Re: encrypted repositories? with git-torrent?","fromName":"Matthias Andree","fromEmail":"matthias.andree@gmx.de","sentAt":"2009-07-20T12:13:31Z","receivedAt":"2009-07-20T12:13:31Z","isPatch":false,"sender":{"key":"matthias.andree@gmx.de","avatar":null},"body":"Am 18.07.2009, 21:09 Uhr, schrieb Thomas Koch <thomas@koch.ro>:\n\n> Wouldn't this be a use case for git-torrent?\n> http://code.google.com/p/gittorrent/\n> http://repo.or.cz/w/VCS-Git-Torrent.git\n>\n> As I understand it, all data would be stored decentraliced and the  \n> (optional?) central server only saves, who has which objects.\n\nI wonder about latency and accessibility here if clients are disconnected.  \nSeems this is more for transferring repositories to large numbers of  \ncustomers as sort of content distribution network for high load, rather  \nthan low connectivity - and the latter is my prime concern and also a  \ndetail of my scenario.\n\n-- \nMatthias Andree\n"},{"id":"118311","messageId":"m3zlazbh4e.fsf@localhost.localdomain","threadId":"20144","inReplyTo":"op.uxc712eh1e62zd@balu.cs.uni-paderborn.de","subject":"Re: encrypted repositories?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-07-20T13:48:03Z","receivedAt":"2009-07-20T13:48:03Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Matthias Andree\" <matthias.andree@gmx.de> writes:\n\n> On a more general note, is someone looking into improving the http://\n> efficiency?  Perhaps there are synergies between my plan of (a)\n> encryption  and (b) more efficient \"dumb\" (http/rsync/...) protocol\n> use.\n\nThere was idea about improving http:// efficiency, but it was via\ncrating git-over-HTTP aka. \"smart\" HTTP server, i.e. you would have to\nhave DAG exposed, like for git:// and ssh://\n\n\nOn the other hand for http:// server need only \"dumb\" web server, and\nadditional metadata generated by git-update-server-info.  It is client\nwho does \"walking\" the DAG, so all data including server metadata can\nbe encrypted, and decrypted on-the-fly by client.  \n\nI don't know though what information leakage you would get from\nexistence of loose objects and packfiles, and their sizes.  Probably\nnegligible...\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"118325","messageId":"20090720153024.GD5347@coredump.intra.peff.net","threadId":"20144","inReplyTo":"op.uxc712eh1e62zd@balu.cs.uni-paderborn.de","subject":"Re: encrypted repositories?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-07-20T15:30:24Z","receivedAt":"2009-07-20T15:30:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 20, 2009 at 02:09:28PM +0200, Matthias Andree wrote:\n\n> No, the server can't be allowed access to the keys or decrypted data.\n> \n> I'm not sure about the graph, and if I should be concerned. Exposing\n> the DAG might be in order.\n> \n> It would be ok if the disk storage and the over-the-wire format\n> cannot use delta compression then. It would suffice to just send a\n> set of objects efficiently - and perhaps smaller revisions can be\n> delta-compressed by the clients when pushing.\n\nThe problem is that you need to expose not just the DAG, but also the\nhashes of trees and blobs. Because if I know you have master^, and I want\nto send you master, then I need to know which objects are referenced by\nmaster that are not referenced by master^.\n\nSo now you have security implications, because I can do an offline\nguessing attack against your files (i.e., calculate git blob hashes for\nlikely candidates and see if you have them). Whether that is a problem\nreally depends on your data.\n\nNot to mention that it makes the protocol a lot more complex, as you\nwould be encrypting _parts_ of objects, like the filenames of a tree,\nand the commit message of a commit object.\n\nI suppose in theory you could obfuscate the sha1's in a way that\npreserved the object relationships but revealed no information. That is,\nthe server would have one \"fake\" set of sha1's, and the client would map\nits real sha1's to the fake ones when talking with the server. But that\nis again potentially getting complex.\n\n-Peff\n"},{"id":"118361","messageId":"op.uxesdcu41e62zd@merlin.emma.line.org","threadId":"20144","inReplyTo":"20090720153024.GD5347@coredump.intra.peff.net","subject":"Re: encrypted repositories?","fromName":"Matthias Andree","fromEmail":"matthias.andree@gmx.de","sentAt":"2009-07-21T08:25:50Z","receivedAt":"2009-07-21T08:25:50Z","isPatch":false,"sender":{"key":"matthias.andree@gmx.de","avatar":null},"body":"Am 20.07.2009, 17:30 Uhr, schrieb Jeff King <peff@peff.net>:\n\n> On Mon, Jul 20, 2009 at 02:09:28PM +0200, Matthias Andree wrote:\n>\n>> No, the server can't be allowed access to the keys or decrypted data.\n>>\n>> I'm not sure about the graph, and if I should be concerned. Exposing\n>> the DAG might be in order.\n>>\n>> It would be ok if the disk storage and the over-the-wire format\n>> cannot use delta compression then. It would suffice to just send a\n>> set of objects efficiently - and perhaps smaller revisions can be\n>> delta-compressed by the clients when pushing.\n>\n> The problem is that you need to expose not just the DAG, but also the\n> hashes of trees and blobs. Because if I know you have master^, and I want\n> to send you master, then I need to know which objects are referenced by\n> master that are not referenced by master^.\n\nYes, you need to know that.  Not all of the push logic needs to be  \nimplemented on the server though.\n\nIn my scenario, the server degenerates into sort of a general object store  \n- I really don't expect much smartness there. What is easily available  \n(clients providing deltas rather than full objects) could be exploited,  \nand that's it.\n\nWe can always have two local repositories, one reference and one checkout.  \nThe reference is a decrypted (unencrypted) copy of the set of objects on  \nthe server, and I could use that for tracking the server-side view (for  \ninstance, what are master^ and master pointing at so I can derive git  \nrev-list master^...master, what do I need to send to the server).\n\nI'm well aware that crypto requires more efforts on the client side if we  \ndon't trust the server, that's just natural.\n\nThe question is: which VCS can serve my scenario?\n\n> So now you have security implications, because I can do an offline\n> guessing attack against your files (i.e., calculate git blob hashes for\n> likely candidates and see if you have them). Whether that is a problem\n> really depends on your data.\n\nOr look at commit frequency and push sources. There's always a leak of  \ninformation even if I just upload a series of  \nblah-2009MMDD-NNN.tar.lzma.gpg files... The data is going to be obsolete,  \nsay, 3 months; students then write the exam and then it's sort of public  \nanyways. Even if your model does not entail not publishing exams (as  \nopposed to embargoed press releases under development), but you can't  \nprevent someone from writing their recollection of the problems from  \nmemory afterwards and sharing it with other students.\n\n> Not to mention that it makes the protocol a lot more complex, as you\n> would be encrypting _parts_ of objects, like the filenames of a tree,\n> and the commit message of a commit object.\n>\n> I suppose in theory you could obfuscate the sha1's in a way that\n> preserved the object relationships but revealed no information. That is,\n> the server would have one \"fake\" set of sha1's, and the client would map\n> its real sha1's to the fake ones when talking with the server. But that\n> is again potentially getting complex.\n\nIs your concern that the object name (SHA1) is derived from the  \nunencrypted version?\n\n-- \nMatthias Andree\n"},{"id":"118362","messageId":"op.uxesknlr1e62zd@merlin.emma.line.org","threadId":"20144","inReplyTo":"m3zlazbh4e.fsf@localhost.localdomain","subject":"Re: encrypted repositories?","fromName":"Matthias Andree","fromEmail":"matthias.andree@gmx.de","sentAt":"2009-07-21T08:30:13Z","receivedAt":"2009-07-21T08:30:13Z","isPatch":false,"sender":{"key":"matthias.andree@gmx.de","avatar":null},"body":"Am 20.07.2009, 15:48 Uhr, schrieb Jakub Narebski <jnareb@gmail.com>:\n\n> \"Matthias Andree\" <matthias.andree@gmx.de> writes:\n>\n>> On a more general note, is someone looking into improving the http://\n>> efficiency?  Perhaps there are synergies between my plan of (a)\n>> encryption  and (b) more efficient \"dumb\" (http/rsync/...) protocol\n>> use.\n>\n> There was idea about improving http:// efficiency, but it was via\n> crating git-over-HTTP aka. \"smart\" HTTP server, i.e. you would have to\n> have DAG exposed, like for git:// and ssh://\n>\n>\n> On the other hand for http:// server need only \"dumb\" web server, and\n> additional metadata generated by git-update-server-info.  It is client\n> who does \"walking\" the DAG, so all data including server metadata can\n> be encrypted, and decrypted on-the-fly by client.\n\nFine by me, and seems to be some \"minimal disclosure to server\".\n\n> I don't know though what information leakage you would get from\n> existence of loose objects and packfiles, and their sizes.  Probably\n> negligible...\n\nDunno. Given that it's just a collection of object sizes, you can't tell  \n from the SHA1 if the object in question is tree, tag, blob, or commit.\n\nI'm really not after on-the-fly delta re-compression on the server-side  \nfor crypto stuff. I'm more thinking along the lines of zsync/bsdiff/xdelta  \n(http://zsync.moria.org.uk/ for the least-known) - but zsync can't work on  \nencrypted data. Perhaps encrypting the diffs could work, but then what's  \nthe difference to using the http:// and update-server-info related  \nmaterial and combining that with client-side on-the-fly (de/en)cryption?\n\n-- \nMatthias Andree\n"},{"id":"118574","messageId":"20090723104032.GB4247@coredump.intra.peff.net","threadId":"20144","inReplyTo":"op.uxesdcu41e62zd@merlin.emma.line.org","subject":"Re: encrypted repositories?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-07-23T10:40:32Z","receivedAt":"2009-07-23T10:40:32Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jul 21, 2009 at 10:25:50AM +0200, Matthias Andree wrote:\n\n> >The problem is that you need to expose not just the DAG, but also the\n> >hashes of trees and blobs. Because if I know you have master^, and I want\n> >to send you master, then I need to know which objects are referenced by\n> >master that are not referenced by master^.\n> \n> Yes, you need to know that.  Not all of the push logic needs to be\n> implemented on the server though.\n\nYes, though fetching is much harder, since the server is the one holding\nthe information about how the transfer can be optimized. Still, you\nshould be able to achieve roughly the same performance as http fetching\nfrom a dumb server.\n\n> Or look at commit frequency and push sources. There's always a leak\n> of information even if I just upload a series of\n> blah-2009MMDD-NNN.tar.lzma.gpg files... The data is going to be\n> obsolete, say, 3 months; students then write the exam and then it's\n> sort of public anyways. Even if your model does not entail not\n> publishing exams (as opposed to embargoed press releases under\n> development), but you can't prevent someone from writing their\n> recollection of the problems from memory afterwards and sharing it\n> with other students.\n> [...]\n> Is your concern that the object name (SHA1) is derived from the\n> unencrypted version?\n\nYes. You are potentially leaking considerable information about the\nunencrypted contents which an attacker could use to guess those contents\n(especially if the file is mostly composed of low-entropy parts, like\ntext formatting).\n\n-Peff\n"},{"id":"196361","messageId":"1343919172934-7564308.post@n2.nabble.com","threadId":"20144","inReplyTo":"op.uw7wmbr41e62zd@balu.cs.uni-paderborn.de","subject":"Re: encrypted repositories?","fromName":"J-S-B","fromEmail":"john-s-brumbelow@hotmail.com","sentAt":"2012-08-02T14:52:52Z","receivedAt":"2012-08-02T14:52:52Z","isPatch":false,"sender":{"key":"john-s-brumbelow@hotmail.com","avatar":null},"body":"To Matthias.\n\nMy name is John Brumbelow, and I know how to solve your issue. Its requires\nclient-driven-encryption, such that the server never gets the data\nun-encrypted, ever, and thus can never decrypt it, but yet, on behalf of the\nclient, and other clients, users can still search, partially search,\nwild-char-search, document-reference-search, and most important, sort data\nstored on the server, that the server (and its administrators) can never\n\"see\" cause it never left the client making/editing it, un-encrypted.\n\nNote the key phrase is client-driven-encryption, not, client-encryption.\nThat is, the client drives the encryption process in conjunction with the\nserver(s).\n\nThis makes it so the data can never be stolen from the server, even if a\nvillan showed up at the server and held the administrators hostage, or at\nany point beyond the client.\n\nThere is more...\n\nI also know how to further obfuscate the data, as it is being made, to\nprotect the end user, so if the villan showed up at the client, and held\nthem hostage and/or hacked their computer(s), their data would still be\nprotected. This is not just a simple, extra-hash over the data, but\nsomething that gives the villan fake/false data, which could be\ntraced/tracked, by the server/authorities.\n\nPlease email me at John-S-Brumbelow@hotmail.com\n\n\n\n\n\n--\nView this message in context: http://git.661346.n2.nabble.com/encrypted-repositories-tp3275970p7564308.html\nSent from the git mailing list archive at Nabble.com.\n"}]}