git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[Announce] bup 0.09: git-based backup system for really huge datasets

From
Avery Pennarun <apenwarr@gmail.com>
Date
Feb 9, 2010, 22:48 UTC
Message-ID
<32541b131002091448o6f809322x1d86d2d7f74a80ed@mail.gmail.com>
Hi all,

bup is a file backup tool based on the git packfile format. If you're interested in git, you might find bup interesting because:

- It can handle really massive datasets (hundreds of gigabytes)
without melting down.
- It can handle huge individual files (hundreds of gigabytes), such as
virtual machine images or giant textual database dumps, while neither
wasting disk space nor bogging down in xdelta.
- It can backup files directly to a remote server, without creating
git objects on the local system first.
- It uses a different format for its index file (.bup/bupindex) that
allows you to search and iterate non-linearly.  Thus if you have a
filesystem with a million files and only one of them is marked dirty,
bup can back it up near-instantly.
- Like git, it separates the concept of indexing the filesystem from
the concept of actually making new commits.  Thus it would be easy to
plugin an inotify-like system eventually, avoiding the slow filesystem
iteration every time you want to make a backup.
- It introduces a "multi-index" file (midx) that has a sorted list of
the objects from multiple .pack files, so that checking for a
nonexistent object only needs to swap in two pages at most.  (This is
unimportant in git, but critical when most of your work is ingesting
huge files whose sha1sums haven't been seen before.)
- It provides a FUSE-based filesystem so that you can easily browse
your backup history, including exporting it via samba if you want.

bup doesn't yet back up extra file metadata (beyond what git already tracks). Obviously this will be needed relatively soon.

bup is still pretty experimental, but it's already a useful tool for backing up your files, even if those files include millions of files and hundreds of gigs of VM images.

You can find the source code (and README) at github:
    http://github.com/apenwarr/bup
To subscribe to the bup mailing list, send an email to:
    bup-list+subscribe@googlegroups.com
Looking forward to everyone's feedback.
Have fun,
Avery
Next: Jakub Narebski
Message 1 of 5 in “[Announce] bup 0.09: git-based backup system for really huge datasets”
  1. Avery PennarunFeb 9, 2010
  2. Jakub NarebskiFeb 10, 2010
  3. Avery PennarunFeb 10, 2010
  4. Stephen R. van den BergFeb 11, 2010
  5. Avery PennarunFeb 12, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.