{"thread":{"id":"22592","subject":"[Announce] bup 0.09: git-based backup system for really huge datasets","startedAt":"2010-02-09T22:48:03Z","lastAt":"2010-02-12T17:51:35Z","messageCount":5,"participants":["Avery Pennarun","Jakub Narebski","Stephen R. van den Berg"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"134089","messageId":"32541b131002091448o6f809322x1d86d2d7f74a80ed@mail.gmail.com","threadId":"22592","inReplyTo":null,"subject":"[Announce] bup 0.09: git-based backup system for really huge datasets","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-02-09T22:48:03Z","receivedAt":"2010-02-09T22:48:03Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"Hi all,\n\nbup is a file backup tool based on the git packfile format.  If you're\ninterested in git, you might find bup interesting because:\n\n- It can handle really massive datasets (hundreds of gigabytes)\nwithout melting down.\n\n- It can handle huge individual files (hundreds of gigabytes), such as\nvirtual machine images or giant textual database dumps, while neither\nwasting disk space nor bogging down in xdelta.\n\n- It can backup files directly to a remote server, without creating\ngit objects on the local system first.\n\n- It uses a different format for its index file (.bup/bupindex) that\nallows you to search and iterate non-linearly.  Thus if you have a\nfilesystem with a million files and only one of them is marked dirty,\nbup can back it up near-instantly.\n\n- Like git, it separates the concept of indexing the filesystem from\nthe concept of actually making new commits.  Thus it would be easy to\nplugin an inotify-like system eventually, avoiding the slow filesystem\niteration every time you want to make a backup.\n\n- It introduces a \"multi-index\" file (midx) that has a sorted list of\nthe objects from multiple .pack files, so that checking for a\nnonexistent object only needs to swap in two pages at most.  (This is\nunimportant in git, but critical when most of your work is ingesting\nhuge files whose sha1sums haven't been seen before.)\n\n- It provides a FUSE-based filesystem so that you can easily browse\nyour backup history, including exporting it via samba if you want.\n\nbup doesn't yet back up extra file metadata (beyond what git already\ntracks).  Obviously this will be needed relatively soon.\n\nbup is still pretty experimental, but it's already a useful tool for\nbacking up your files, even if those files include millions of files\nand hundreds of gigs of VM images.\n\nYou can find the source code (and README) at github:\n\n    http://github.com/apenwarr/bup\n\nTo subscribe to the bup mailing list, send an email to:\n\n    bup-list+subscribe@googlegroups.com\n\nLooking forward to everyone's feedback.\n\nHave fun,\n\nAvery\n"},{"id":"134135","messageId":"m3sk998lhq.fsf@localhost.localdomain","threadId":"22592","inReplyTo":"32541b131002091448o6f809322x1d86d2d7f74a80ed@mail.gmail.com","subject":"Re: [Announce] bup 0.09: git-based backup system for really huge datasets","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-02-10T09:54:30Z","receivedAt":"2010-02-10T09:54:30Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> writes:\n\n> bup is a file backup tool based on the git packfile format.\n\n[...]\n> bup is still pretty experimental, but it's already a useful tool for\n> backing up your files, even if those files include millions of files\n> and hundreds of gigs of VM images.\n> \n> You can find the source code (and README) at github:\n> \n>     http://github.com/apenwarr/bup\n> \n> To subscribe to the bup mailing list, send an email to:\n> \n>     bup-list+subscribe@googlegroups.com\n> \n> Looking forward to everyone's feedback.\n\nWould you be adding short info about your project to\nhttp://git.wiki.kernel.org/index.php/InterfacesFrontendsAndTools\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"134175","messageId":"32541b131002101201n2d626152xe1aecd6697260b7f@mail.gmail.com","threadId":"22592","inReplyTo":"m3sk998lhq.fsf@localhost.localdomain","subject":"Re: [Announce] bup 0.09: git-based backup system for really huge datasets","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-02-10T20:01:08Z","receivedAt":"2010-02-10T20:01:08Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Feb 10, 2010 at 4:54 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n> Avery Pennarun <apenwarr@gmail.com> writes:\n>> bup is a file backup tool based on the git packfile format.\n> [...]\n>> bup is still pretty experimental, but it's already a useful tool for\n>> backing up your files, even if those files include millions of files\n>> and hundreds of gigs of VM images.\n>>\n>> You can find the source code (and README) at github:\n>>\n>>     http://github.com/apenwarr/bup\n>>\n>> To subscribe to the bup mailing list, send an email to:\n>>\n>>     bup-list+subscribe@googlegroups.com\n>>\n>> Looking forward to everyone's feedback.\n>\n> Would you be adding short info about your project to\n> http://git.wiki.kernel.org/index.php/InterfacesFrontendsAndTools\n\nDone.  Thanks for the reminder!\n\nHave fun,\n\nAvery\n"},{"id":"134223","messageId":"20100211135129.GA2988@cuci.nl","threadId":"22592","inReplyTo":"32541b131002091448o6f809322x1d86d2d7f74a80ed@mail.gmail.com","subject":"Re: [Announce] bup 0.09: git-based backup system for really huge datasets","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2010-02-11T13:51:29Z","receivedAt":"2010-02-11T13:51:29Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Avery Pennarun wrote:\n>bup is a file backup tool based on the git packfile format.  If you're\n>interested in git, you might find bup interesting because:\n\nInteresting concept.  It has some killer features which make it a good\ncompetitor to any of the existing solutions.\nThe only obvious thing missing for unattended backup-operation is a way\nto purge specific or old backups.\n-- \nSincerely,\n           Stephen R. van den Berg.\n"},{"id":"134365","messageId":"32541b131002120951h25368812w547e8dcbaf054fa1@mail.gmail.com","threadId":"22592","inReplyTo":"20100211135129.GA2988@cuci.nl","subject":"Re: [Announce] bup 0.09: git-based backup system for really huge datasets","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-02-12T17:51:35Z","receivedAt":"2010-02-12T17:51:35Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Thu, Feb 11, 2010 at 8:51 AM, Stephen R. van den Berg <srb@cuci.nl> wrote:\n> Avery Pennarun wrote:\n>>bup is a file backup tool based on the git packfile format.  If you're\n>>interested in git, you might find bup interesting because:\n>\n> Interesting concept.  It has some killer features which make it a good\n> competitor to any of the existing solutions.\n> The only obvious thing missing for unattended backup-operation is a way\n> to purge specific or old backups.\n\nThanks.  Oddly enough, pruning of old backups hasn't been a really\nhigh priority for me (or apparently any of the other users) because\nchunking-based deduplication is so efficient that my backup disk\nhasn't filled up yet :)  But it's clear that this will need to be\nadded eventually.\n\nUnfortunately git's normal pruning and gc stuff is inapplicable since\nit dies horribly when faced with hundreds of gigabytes of data.\nThat's to be expected, but it means I can't just cheat by running 'git\ngc' and hoping for magic.\n\nHave fun,\n\nAvery\n"}]}