{"thread":{"id":"19822","subject":"Using git for code deployment on webservers?","startedAt":"2009-06-15T23:11:47Z","lastAt":"2009-06-17T20:33:37Z","messageCount":10,"participants":["Ingo Oeser","Allan Wind","Thomas Koch","Daniel Barkalow","Alex Riesen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"116368","messageId":"200906160111.47325.ioe-git@rameria.de","threadId":"19822","inReplyTo":null,"subject":"Using git for code deployment on webservers?","fromName":"Ingo Oeser","fromEmail":"ioe-git@rameria.de","sentAt":"2009-06-15T23:11:47Z","receivedAt":"2009-06-15T23:11:47Z","isPatch":false,"sender":{"key":"ioe-git@rameria.de","avatar":null},"body":"[please CC me, as I'm not subscribed]\n\nHi there,\n\nI try to use git in a quite unusual way.\n\nI have a bunch of servers (hundreds), which get regular pulls of web developer code.\nThe code consists of images, flash files, scripting language files, you name it.\nAn exported repo (just the files, no SCM metadata) contains up to 4GB of files.\n\nNo I want to distribute changes the developers made in a tree like structure:\n\nmain server --> slave_1 --> webserver_0815\n            |-> slave_2 --> webserver_2342\n                        |-> webserver_4711\n\nBut with the following contraints:\n- Store as little as possible on the webservers.\n  One selected revision/tag is enough.\n- Transfer as little as possible data.\n  Cancel out addition and deletion on the fly.\n- Nearly atomic update of file tree (easy to implement outside git)\n\nNice to have:\n- Instead of copying the files to their proper names, \n  hardlink them to their git objects.\n\nAt the moment I always get more data than I need and have to store\nthe repository AND the checked out data.\n\nI couldn't find a way so far to get around this. Is this possible? \nAny ideas are welcome.\n\nMany Thanks in Advance!\n\nBest Regards\n\nIngo Oeser\n"},{"id":"116390","messageId":"20090616071328.GB6615@lifeintegrity.com","threadId":"19822","inReplyTo":"200906160111.47325.ioe-git@rameria.de","subject":"Re: Using git for code deployment on webservers?","fromName":"Allan Wind","fromEmail":"allan_wind@lifeintegrity.com","sentAt":"2009-06-16T07:13:28Z","receivedAt":"2009-06-16T07:13:28Z","isPatch":false,"sender":{"key":"allan_wind@lifeintegrity.com","avatar":null},"body":"On 2009-06-16T01:11:47, Ingo Oeser wrote:\n> - Transfer as little as possible data.\n>   Cancel out addition and deletion on the fly.\n\nI use `git diff` with the post-receive hook to distribute changes \nto my web server.  diff carries the previous content when you \ndelete a file, and in my case this was large mpeg files defeating \nthe purpose somewhat.\n\nIf you do not mind having a full repository on the web servers, \nthen pushing changes might work better.  This appears to be what \nyou are doing now though.\n\nIf I had to scale this I would probably build a master image \n(either locally or remotely) and use rsync to distribute the \ncontent instead of git.\n\n> - Nearly atomic update of file tree (easy to implement outside git)\n\nstow can be handy for this.\n\n\n/Allan\n-- \nAllan Wind\nLife Integrity, LLC\nhttp://lifeintegrity.com\n"},{"id":"116394","messageId":"200906161001.33678.thomas@koch.ro","threadId":"19822","inReplyTo":"200906160111.47325.ioe-git@rameria.de","subject":"Re: Using git for code deployment on webservers?","fromName":"Thomas Koch","fromEmail":"thomas@koch.ro","sentAt":"2009-06-16T08:01:33Z","receivedAt":"2009-06-16T08:01:33Z","isPatch":false,"sender":{"key":"thomas@koch.ro","avatar":null},"body":"Would it help, to share a read only GIT object store among all webservers via \nNFS?\n\nBest regards, Thomas Koch\n\n> [please CC me, as I'm not subscribed]\n>\n> Hi there,\n>\n> I try to use git in a quite unusual way.\n>\n> I have a bunch of servers (hundreds), which get regular pulls of web\n> developer code. The code consists of images, flash files, scripting\n> language files, you name it. An exported repo (just the files, no SCM\n> metadata) contains up to 4GB of files.\n>\n> No I want to distribute changes the developers made in a tree like\n> structure:\n>\n> main server --> slave_1 --> webserver_0815\n>\n>             |-> slave_2 --> webserver_2342\n>             |\n>                         |-> webserver_4711\n>\n> But with the following contraints:\n> - Store as little as possible on the webservers.\n>   One selected revision/tag is enough.\n> - Transfer as little as possible data.\n>   Cancel out addition and deletion on the fly.\n> - Nearly atomic update of file tree (easy to implement outside git)\n>\n> Nice to have:\n> - Instead of copying the files to their proper names,\n>   hardlink them to their git objects.\n>\n> At the moment I always get more data than I need and have to store\n> the repository AND the checked out data.\n>\n> I couldn't find a way so far to get around this. Is this possible?\n> Any ideas are welcome.\n>\n> Many Thanks in Advance!\n>\n> Best Regards\n>\n> Ingo Oeser\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\nThomas Koch, http://www.koch.ro\n"},{"id":"116426","messageId":"alpine.LNX.2.00.0906161332080.2147@iabervon.org","threadId":"19822","inReplyTo":"200906160111.47325.ioe-git@rameria.de","subject":"Re: Using git for code deployment on webservers?","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2009-06-16T17:49:01Z","receivedAt":"2009-06-16T17:49:01Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 16 Jun 2009, Ingo Oeser wrote:\n\n> [please CC me, as I'm not subscribed]\n> \n> Hi there,\n> \n> I try to use git in a quite unusual way.\n> \n> I have a bunch of servers (hundreds), which get regular pulls of web developer code.\n> The code consists of images, flash files, scripting language files, you name it.\n> An exported repo (just the files, no SCM metadata) contains up to 4GB of files.\n> \n> No I want to distribute changes the developers made in a tree like structure:\n> \n> main server --> slave_1 --> webserver_0815\n>             |-> slave_2 --> webserver_2342\n>                         |-> webserver_4711\n> \n> But with the following contraints:\n> - Store as little as possible on the webservers.\n>   One selected revision/tag is enough.\n> - Transfer as little as possible data.\n>   Cancel out addition and deletion on the fly.\n> - Nearly atomic update of file tree (easy to implement outside git)\n> \n> Nice to have:\n> - Instead of copying the files to their proper names, \n>   hardlink them to their git objects.\n> \n> At the moment I always get more data than I need and have to store\n> the repository AND the checked out data.\n\nYou should be able to have the slave repositories store tags for tree \nobjects (instead of commit objects), and have the webservers fetch those. \nYou'll still have the object database, but it will only contain stuff \nthat's been deployed to that webserver, not intermediate versions or \nhistorical versions. You'll still have to store both the repo and the \nchecked out data (but git stores the content delta-compressed against each \nother in one big file, normally, so there really aren't files to hard link \nto.\n\nOf course, the other possibility is to check out versions on the slaves, \nand rsync that to the webservers, which is probably the optimal method if \nyou're not in a situation where you benefit from anything git does in \ntransit.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"116490","messageId":"200906171923.08034.ioe-git@rameria.de","threadId":"19822","inReplyTo":"alpine.LNX.2.00.0906161332080.2147@iabervon.org","subject":"Re: Using git for code deployment on webservers?","fromName":"Ingo Oeser","fromEmail":"ioe-git@rameria.de","sentAt":"2009-06-17T17:23:07Z","receivedAt":"2009-06-17T17:23:07Z","isPatch":false,"sender":{"key":"ioe-git@rameria.de","avatar":null},"body":"Hi Daniel,\n\nOn Tuesday 16 June 2009, Daniel Barkalow wrote:\n> You should be able to have the slave repositories store tags for tree \n> objects (instead of commit objects), and have the webservers fetch those. \n> You'll still have the object database, but it will only contain stuff \n> that's been deployed to that webserver, not intermediate versions or \n> historical versions.\n\nAh, that sound like a great solution. I'll try that.\n\n> You'll still have to store both the repo and the checked out data \n> (but git stores the content delta-compressed against each \n> other in one big file, normally, so there really aren't files to hard link \n> to.\n\nOk. That was under the assumption, that the core of git is basically a \ncontent addressable file system. But that seems to be history :-)\n\n> Of course, the other possibility is to check out versions on the slaves, \n> and rsync that to the webservers, which is probably the optimal method if \n> you're not in a situation where you benefit from anything git does in \n> transit.\n\nI would benefit from noticing local changes. But simple rsync is what is tried now.\nProblem is, we get no de-duplication from rsync, which git could do.\n\nMany thanks for your suggestions!\n\n\nBest Regards \n\nIngo Oeser\n"},{"id":"116491","messageId":"200906171927.18435.ioe-lkml@rameria.de","threadId":"19822","inReplyTo":"200906161001.33678.thomas@koch.ro","subject":"Re: Using git for code deployment on webservers?","fromName":"Ingo Oeser","fromEmail":"ioe-lkml@rameria.de","sentAt":"2009-06-17T17:27:18Z","receivedAt":"2009-06-17T17:27:18Z","isPatch":false,"sender":{"key":"ioe-lkml@rameria.de","avatar":null},"body":"Hi Thomas,\n\nOn Tuesday 16 June 2009, Thomas Koch wrote:\n> Would it help, to share a read only GIT object store among all webservers via \n> NFS?\n\nNFS on hundreds of web servers has severe scaling problems. That is by design and is solved\nby alternative file systems or soon pNFS.\n\nWe tried such a setup already.\n\n\nBest Regards\n\nIngo Oeser\n"},{"id":"116492","messageId":"200906171942.43002.ioe-lkml@rameria.de","threadId":"19822","inReplyTo":"20090616071328.GB6615@lifeintegrity.com","subject":"Re: Using git for code deployment on webservers?","fromName":"Ingo Oeser","fromEmail":"ioe-lkml@rameria.de","sentAt":"2009-06-17T17:42:42Z","receivedAt":"2009-06-17T17:42:42Z","isPatch":false,"sender":{"key":"ioe-lkml@rameria.de","avatar":null},"body":"Hi Allan,\n\nOn Tuesday 16 June 2009, Allan Wind wrote:\n> If you do not mind having a full repository on the web servers, \n> then pushing changes might work better.  This appears to be what \n> you are doing now though.\n\nNo, at the moment we have built our own version of a content addressable \nfilesystem and are distributing changes to it. We have symlinks to real file names.\n\nI just thought, that git can do sth. similiar with its core, \nbefore trying to solve a solved problem :-)\n\n> If I had to scale this I would probably build a master image \n> (either locally or remotely) and use rsync to distribute the \n> content instead of git.\n\nWe do sth. similiar at the moment. De-duplication is important, because\nweb people copy lots of data for images and flash around when doing things.\n\n> > - Nearly atomic update of file tree (easy to implement outside git)\n> \n> stow can be handy for this.\n\nAh! Will have a look.\n\nMany Thanks!\n\n\nBest Regards\n\nIngo Oeser\n"},{"id":"116496","messageId":"alpine.LNX.2.00.0906171328080.2147@iabervon.org","threadId":"19822","inReplyTo":"200906171923.08034.ioe-git@rameria.de","subject":"Re: Using git for code deployment on webservers?","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2009-06-17T19:26:04Z","receivedAt":"2009-06-17T19:26:04Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 17 Jun 2009, Ingo Oeser wrote:\n\n> Hi Daniel,\n> \n> On Tuesday 16 June 2009, Daniel Barkalow wrote:\n> > You should be able to have the slave repositories store tags for tree \n> > objects (instead of commit objects), and have the webservers fetch those. \n> > You'll still have the object database, but it will only contain stuff \n> > that's been deployed to that webserver, not intermediate versions or \n> > historical versions.\n> \n> Ah, that sound like a great solution. I'll try that.\n> \n> > You'll still have to store both the repo and the checked out data \n> > (but git stores the content delta-compressed against each \n> > other in one big file, normally, so there really aren't files to hard link \n> > to.\n> \n> Ok. That was under the assumption, that the core of git is basically a \n> content addressable file system. But that seems to be history :-)\n\nIt is (based on) a content-addressable file system, but it's not a host \nfile system. It's a file system in the sense that you can put octet \nsequences into it and lookup them up by their names, but you can't mount \nit from the kernel and link to it. It's like a tar file, although it's \nmore limited in that it doesn't provide a \"list\" operation.\n\nThere's no fundamental reason there couldn't be a kernel driver (or, \nmore likely, FUSE helper) which could mount it, but that's not the normal \nmethod.\n\n> > Of course, the other possibility is to check out versions on the slaves, \n> > and rsync that to the webservers, which is probably the optimal method if \n> > you're not in a situation where you benefit from anything git does in \n> > transit.\n> \n> I would benefit from noticing local changes. But simple rsync is what is tried now.\n> Problem is, we get no de-duplication from rsync, which git could do.\n\nIn that case, fetching trees is probably the right thing; that should give \nyou a point-to-point de-duplication without any history (although you may \nalso turn up git bugs, since this isn't how git is normally used).\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"116498","messageId":"81b0412b0906171326y6821d511u5b93cda4a5c14458@mail.gmail.com","threadId":"19822","inReplyTo":"alpine.LNX.2.00.0906171328080.2147@iabervon.org","subject":"Re: Using git for code deployment on webservers?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2009-06-17T20:26:13Z","receivedAt":"2009-06-17T20:26:13Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"2009/6/17 Daniel Barkalow <barkalow@iabervon.org>:\n> On Wed, 17 Jun 2009, Ingo Oeser wrote:\n>> > Of course, the other possibility is to check out versions on the slaves,\n>> > and rsync that to the webservers, which is probably the optimal method if\n>> > you're not in a situation where you benefit from anything git does in\n>> > transit.\n>>\n>> I would benefit from noticing local changes. But simple rsync is what is tried now.\n>> Problem is, we get no de-duplication from rsync, which git could do.\n>\n> In that case, fetching trees is probably the right thing; that should give\n> you a point-to-point de-duplication without any history (although you may\n> also turn up git bugs, since this isn't how git is normally used).\n\nOr, you can just keep a namespace for each server in the intermediate\nrepositories, which records the version the server has and the version\nit should have. Then you can use git diff-tree to find you which files\nhave to be transferred. You wont be able to record changes on the servers,\nthough.\n"},{"id":"116499","messageId":"81b0412b0906171333l38b1c2e0y2e166bc0ae7a461d@mail.gmail.com","threadId":"19822","inReplyTo":"81b0412b0906171326y6821d511u5b93cda4a5c14458@mail.gmail.com","subject":"Re: Using git for code deployment on webservers?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2009-06-17T20:33:37Z","receivedAt":"2009-06-17T20:33:37Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"2009/6/17 Alex Riesen <raa.lkml@gmail.com>:\n> Or, you can just keep a namespace for each server in the intermediate\n\nI mean namespace of branches:\n\n  refs/heads/webserver_1/master (current)\n  refs/heads/webserver_1/next (to be updated to)\n\n> repositories, which records the version the server has and the version\n> it should have. Then you can use git diff-tree to find you which files\n> have to be transferred. You wont be able to record changes on the servers,\n> though.\n\nSomething like that:\n\ngit diff-tree --diff-filter=AM webserver1_/master..webserver_1/next |\nwhile read f; do scp \"$f\" webserver_1:\"$f\" || break; done\ngit diff-tree --diff-filter=D webserver1_/master..webserver_1/next |\nwhile read f; do ssh webserver_1 rm -f \"$f\" || break; done\n"}]}