{"thread":{"id":"32840","subject":"[Request] Git export with hardlinks","startedAt":"2013-02-06T15:19:07Z","lastAt":"2013-02-13T13:17:21Z","messageCount":5,"participants":["Thomas Koch","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"208817","messageId":"201302061619.07765.thomas@koch.ro","threadId":"32840","inReplyTo":null,"subject":"[Request] Git export with hardlinks","fromName":"Thomas Koch","fromEmail":"thomas@koch.ro","sentAt":"2013-02-06T15:19:07Z","receivedAt":"2013-02-06T15:19:07Z","isPatch":false,"sender":{"key":"thomas@koch.ro","avatar":null},"body":"Hi,\n\nI'd like to script a git export command that can be given a list of already \nexported worktrees and the tree SHA1s these worktrees correspond too. The git \nexport command should then for every file it wants to export lookup in the \nexisting worktrees whether an identical file is already present and in that \ncase hardlink to the new export location instead of writing the same file \nagain.\n\nUse Case: A git based web deployment system that exports git trees to be \nserved by a web server. Every new deployment is written to a new folder. After \nthe export the web server should start serving new requests from the new \nfolder.\n\nIt might be possible that this is premature optimization. But I'd like to \nlearn more Python and dulwich by hacking this.\n\nDo you have any additional thoughts or use cases about this?\n\nRegards,\n\nThomas Koch, http://www.koch.ro\n"},{"id":"208990","messageId":"20130208095819.GA17220@sigill.intra.peff.net","threadId":"32840","inReplyTo":"201302061619.07765.thomas@koch.ro","subject":"Re: [Request] Git export with hardlinks","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-02-08T09:58:19Z","receivedAt":"2013-02-08T09:58:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 06, 2013 at 04:19:07PM +0100, Thomas Koch wrote:\n\n> I'd like to script a git export command that can be given a list of already \n> exported worktrees and the tree SHA1s these worktrees correspond too. The git \n> export command should then for every file it wants to export lookup in the \n> existing worktrees whether an identical file is already present and in that \n> case hardlink to the new export location instead of writing the same file \n> again.\n> \n> Use Case: A git based web deployment system that exports git trees to be \n> served by a web server. Every new deployment is written to a new folder. After \n> the export the web server should start serving new requests from the new \n> folder.\n> \n> It might be possible that this is premature optimization. But I'd like to \n> learn more Python and dulwich by hacking this.\n> \n> Do you have any additional thoughts or use cases about this?\n\nIf you can handle losing the generality of N deployments, you can do it\nin a few lines of shell.\n\nLet's assume for a moment that you keep two trees at any given time:\nthe existing tree being used, and the tree you are setting up to deploy.\nTo save space, you want the new deployment to reuse (via hardlinks) as\nmany of the files from the old deployment as possible.\n\nSo imagine you have a bare repository storing the actual data:\n\n  $ git clone --bare /some/test/repo repo.git\n  $ du -sh *\n  49M     repo.git\n\nand then you have one deployment you've set up previously by checking\nout the repo contents:\n\n  $ export GIT_DIR=$PWD/repo.git\n  $ mkdir old\n  $ (cd old && GIT_WORK_TREE=$PWD git checkout HEAD)\n  $ du -sh *\n  24M     old\n  49M     repo.git\n\nSo a full checkout is 24M. For the next deploy, we'll start by asking\n\"cp\" to duplicate the old, using hard links:\n\n  $ cp -rl old new\n  $ du -sh *\n  24M     new\n  768K    old\n  49M     repo.git\n\nand we use hardly any extra space (it should just be directory inodes).\nAnd now we can ask git to make \"new\" look like some other commit. It\nwill only touch files which have changed, so the rest remain hardlinked,\nand we use only a small amount of extra space:\n\n  $ (cd new && GIT_WORK_TREE=$PWD git checkout HEAD~10)\n  $ du -sh *\n  24M     new\n  1.3M    old\n  49M     repo.git\n\nNow you point your deployment at \"new\", and you are free to leave \"old\"\nsitting around or remove it at your leisure. You save space while the\ntwo co-exist, and you saved the I/O of copying any files from \"old\" to\n\"new\".\n\nThis breaks down, of course, if you want to keep N trees around and\nhard-link to whichever one has the content you want. For that you'd have\nto write some custom code.\n\n-Peff\n"},{"id":"209116","messageId":"201302101133.28746.thomas@koch.ro","threadId":"32840","inReplyTo":"20130208095819.GA17220@sigill.intra.peff.net","subject":"Re: [Request] Git export with hardlinks","fromName":"Thomas Koch","fromEmail":"thomas@koch.ro","sentAt":"2013-02-10T10:33:26Z","receivedAt":"2013-02-10T10:33:26Z","isPatch":false,"sender":{"key":"thomas@koch.ro","avatar":null},"body":"Jeff King:\n> [...]\n> So a full checkout is 24M. For the next deploy, we'll start by asking\n> \"cp\" to duplicate the old, using hard links:\n\nHi Jeff,\n\nthank you very much for your idea! It's good and simple. It just breaks down \nfor the case when a large folder got renamed.\n\nBut I already hacked the basic layout of the algorithm and it's not \ncomplicated at all, I believe:\n\nhttps://github.com/thkoch2001/git_export_hardlinks/blob/master/git_export_hardlinks.py\n\nI had to interrupt work on this and could not yet finish and test it. But I \nthought you might be interested. Maybe something like this might one day be \nrewritten in C and become part of git core?\n\nRegards,\n\nThomas Koch, http://www.koch.ro\n"},{"id":"209271","messageId":"20130211171357.GF16402@sigill.intra.peff.net","threadId":"32840","inReplyTo":"201302101133.28746.thomas@koch.ro","subject":"Re: [Request] Git export with hardlinks","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-02-11T17:13:57Z","receivedAt":"2013-02-11T17:13:57Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 10, 2013 at 11:33:26AM +0100, Thomas Koch wrote:\n\n> thank you very much for your idea! It's good and simple. It just breaks down \n> for the case when a large folder got renamed.\n\nYes, it would never find renames, which a true sha1->path map could.\n\n> But I already hacked the basic layout of the algorithm and it's not \n> complicated at all, I believe:\n> \n> https://github.com/thkoch2001/git_export_hardlinks/blob/master/git_export_hardlinks.py\n\nIt looks like you create the sha1->path mapping by asking the user to\nprovide <tree_sha1>,<path> pairs, and then assuming that the exported\ntree at <path> exactly matches <tree_sha1>. Which it would in the\nworkflow you've proposed, but it is also easy for that not to be the\ncase (e.g., somebody munges a file in <path> after it has been\nexported).\n\nSo it's a bit dangerous as a general purpose tool, IMHO. It's also a\nslight pain in that you have to keep track of the tree sha1 for each\nexported path somehow.\n\nA safer and more convenient (but slightly less efficient) solution would\nbe to keep a git index file for each exported tree. Then we can just\nrefresh that index, which would check that our sha1 for each path is up\nto date (and in the common case of nothing changed, would only be as\nexpensive as stat()-ing each entry). And then we use that index as the\nsha1->path map.\n\nThe simplest way to have an index for each export would be to actually\ngive each one its own git repo (which does not have to use much space,\nif you use \"-s\" to share the objects with the master repo).\n\nThat's more complex, and uses more disk than what your script does, but\nI do think the added safety would be worth it for a general-purpose\ntool.\n\n> I had to interrupt work on this and could not yet finish and test it. But I \n> thought you might be interested. Maybe something like this might one day be \n> rewritten in C and become part of git core?\n\nI think if we had a `git export` command (and we do not, but there has\nbeen discussion in a nearby thread about whether such a thing might be a\ngood idea), having a `--hard-link-from` option to link with other\ncheckouts would make sense. It could also potentially be an option to\ngit-checkout-index, and you could script around it at that low level.\n\n-Peff\n"},{"id":"209469","messageId":"201302131417.22001.thomas@koch.ro","threadId":"32840","inReplyTo":"20130211171357.GF16402@sigill.intra.peff.net","subject":"[ANN] First beta: Git export with hardlinks","fromName":"Thomas Koch","fromEmail":"thomas@koch.ro","sentAt":"2013-02-13T13:17:21Z","receivedAt":"2013-02-13T13:17:21Z","isPatch":false,"sender":{"key":"thomas@koch.ro","avatar":null},"body":"Hi,\n\nmy git_export_hardlink command should now be in a usable state. I'd appreciate \nany feedback: https://github.com/thkoch2001/git_export_hardlinks\n\nI still have to choose a license: BSD/GPL/?\n\nJeff King:\n\n> It looks like you create the sha1->path mapping by asking the user to\n> provide <tree_sha1>,<path> pairs, and then assuming that the exported\n> tree at <path> exactly matches <tree_sha1>. Which it would in the\n> workflow you've proposed, but it is also easy for that not to be the\n> case (e.g., somebody munges a file in <path> after it has been\n> exported).\n> \n> So it's a bit dangerous as a general purpose tool, IMHO. It's also a\n> slight pain in that you have to keep track of the tree sha1 for each\n> exported path somehow.\nYou're right. I'd run a git reset --hard after each export to guarantee a \npristine export.\n\nThe tree sha1 of the exported tree might be part of the folder name of the \nexport or in some meta file related to the export, like\n\n/deployments\n  /2012-03-05_14-23-02_0b96bf5f72d2c282b31726b3fbff279a89220b15\n    /export <- exported tree goes here\n    /meta  <- git config file holding all relevant metadata: (who, when, tree,\n                  commit, ref)\n    /index <- git index file corresponding to the exported tree (maybe?)\n  \nRegards,\n\nThomas Koch, http://www.koch.ro\n"}]}