{"thread":{"id":"8645","subject":"Re: Versioning file system","startedAt":"2007-06-19T03:10:42Z","lastAt":"2007-06-20T02:43:01Z","messageCount":6,"participants":["Kyle Moffett","Jack Stone","Bron Gondwana","Martin Langhoff","Jakub Narebski"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"45314","messageId":"6E9A6F9E-8948-40F2-9129-1F1491D49D83@mac.com","threadId":"8645","inReplyTo":"OF7FA807A1.64C0D5AF-ON882572FE.0061B34C-882572FE.00628322@us.ibm.com","subject":"Re: Versioning file system","fromName":"Kyle Moffett","fromEmail":"mrmacman_g4@mac.com","sentAt":"2007-06-19T03:10:42Z","receivedAt":"2007-06-19T03:10:42Z","isPatch":false,"sender":{"key":"mrmacman_g4@mac.com","avatar":null},"body":"On Jun 18, 2007, at 13:56:05, Bryan Henderson wrote:\n>> The question remains is where to implement versioning: directly in  \n>> individual filesystems or in the vfs code so all filesystems can  \n>> use it?\n>\n> Or not in the kernel at all.  I've been doing versioning of the  \n> types I described for years with user space code and I don't  \n> remember feeling that I compromised in order not to involve the  \n> kernel.\n>\n> Of course, if you want to do it with snapshots and COW, you'll have  \n> to ask where in the kernel to put that, but that's not a file  \n> versioning question; it's the larger snapshot question.\n\nWhat I think would be particularly interesting in this domain is  \nsomething similar in concept to GIT, except in a file-system:\n   1) Redundancy is easy, you just ensure that you have at least \"N\"  \ndistributed copies of each object, where \"N\" is some function of the  \nobject itself.\n   2) Network replication is easy, you look up objects based on the  \nSHA-1 stored in the parent directory entry and cache them where  \nneeded (IE: make the \"N\" function above dynamic based on frequency of  \naccess on a given computer).\n   3) Snapshots are easy and cheap; an RO snapshot is a tag and an RW  \nsnapshot is a branch.  These can be easily converted between.\n   4) Compression is easy; you can compress objects based on any  \narbitrary configurable criteria and the filesystem will record  \nwhether or not an object is compressed.  You can also compress  \ndifferently when archiving objects to secondary storage.\n   5) Streaming fsck-like verification is easy; ensure the hash name  \nfield matches the actual hash of the object.\n   6) Fsck is easy since rollback is trivial, you can always revert  \nto an older tree to boot and start up services before attempting  \nresurrection of lost objects and trees in the background.\n   7) Multiple-drive or multiple-host storage pools are easy:  Think  \nthe git \"alternates\" file.\n   8) Network filesystem load-balancing is easy; SHA-1s are  \nessentially random so you can just assign SHA-1 prefixes to different  \nsystems for data storage and your data is automatically split up.\n\n\nOther issues:\n\nQ. How do you deal with block allocation?\nA. Same way other filesystems deal with block allocation.  Reference- \ncounting gets tricky, especially across a network, but it's easy to  \nplay it safe with simple cross-network refcount-journalling.  Since  \nthe _only_ thing that needs journalling is the refcounts and block- \nfree data, you need at most a megabyte or two of journal.  If in  \ndoubt, it's easy to play it safe and keep an extra refcount around  \nfor an in-the-background consistency check later on.  When networked- \ngitfs systems crash, you just assume they still have all the  \nrefcounts they had at the moment they died, and compare notes when  \nthey start back up again.  If a node has a cached copy of data on its  \nlocal disk then it can just nonatomically increment the refcount for  \nthat object in its own RAM (ordered with respect to disk-flushes, of  \ncourse) and tell its peers at some point.  A node should probably  \ncache most of its working set on local disk for efficiency; it's  \ntrivially verified against updates from other nodes and provides an  \neasy way to keep refcounts for such data.  If a node increments the  \nrefcount on such data and dies before getting that info out to its  \npeers, then when it starts up again its peers will just be told that  \nit has a \"new\" node with insufficient replication and they will clone  \nit out again properly.  For networked refcount-increments you can do  \none of 2 things: (1) Tell at least X many peers and wait for them to  \nsync the update out to disk, or (2) Get the object from any peer (at  \nleast one of whom hopefully has it in RAM) and save it to local disk  \nwith an increased refcount.\n\nQ. How do you actually delete things?\nA. Just replace all the to-be-erased tree and commit objects before a  \nspecified point with \"History erased\" objects with their SHA-1's  \nmagically set to that of the erased objects.  If you want you may  \ndelete only the \"tree\" objects and leave the commits intact.  If you  \ndelete a whole linear segment of history then you can just use a  \nsingle \"History erased\" commit object with its parent pointed to the  \nobject before the erased segment.  Probably needs some form of back- \nreference storage to make it efficient; not sure how expensive that  \nwould be.  This would allow making a bunch of snapshots and purging  \nthem logarithmically based on passage of time.  For instance, you  \nmight have snapshots of every 5 minutes for the last hour, every 30  \nminutes for the last day, every 4 hours for the last week, every day  \nfor the last month, once per week for the last year, once per month  \nfor the last 5 years, and once per year beyond that.\n\nThat's pretty impressive data-recovery resolution, and it accounts  \nfor only 200 unique commits after it's been running for 10 years.\n\nQ. How do you archive data?\nA. Same as deleting, except instead of a \"History erased\" object you  \nwould use a \"History archived\" object with a little bit of string  \ndata to indicate which volume it's stored on (and where on the  \nvolume).  When you stick that volume into the system you could easily  \ntell the kernel to use it as an alternate for the given storage group.\n\nQ. What enforces data integrity?\nA. Ensure that a new tree object and its associated sub objects are  \non disk before you delete the old one.  Doesn't need any actual full  \nsyncs at all, just barriers.  If you replace the tree object before  \nwrite-out is complete then just skip writing the old one and write  \nthe new one in its place.\n\nQ. What consists of a \"commit\"?\nA. Anything the administrator wants to define it as.  Useful  \nalgorithms include: \"Once per x Mbyte of page dirtying\", \"Once per 5  \nmin\", \"Only when sync() or fsync() are called\", \"Only when gitfs- \ncommit is called\".  You could even combine them:  \"Every x Mbyte of  \npage dirtying or every 5 minutes, whichever is shorter (or longer,  \ndepending on admin requirements)\".  There would also be appropriate  \nsyscalls to trigger appropriate git-like behavior.  Network- \naccessible gitfs would want to have mechanisms to trigger commits  \nbased on activity on other systems (needs more thought).\n\nQ. How do you access old versions?\nA. Mount another instance of the filesystem with an SHA-1 ID, a tag- \nname, or a branch-name in a special mount option.  Should be user  \naccessible with some restrictions (needs more thought).\n\nQ. How do you deal with conflicts on networked filesystems.\nA. Once again, however the administrator wants to deal with them.   \nOptions:\n    1)  Forcibly create a new branch for the conflicted tree.\n    2)  Attempt to merge changes using the standard git-merge semantics\n    3)  Merge independent changes to different files and pick one for  \nchanges to the same file\n    4)  Your Algorithm Here(TM).  GIT makes it easy to extend  \nconflict-resolution.\n\nQ. How do you deal with little scattered changes in big (or sparse)  \nfiles?\nA. Two questions, two answers:  For sparse files, git would need  \nextending to understand (and hash) the nature of the sparse-ness.   \nFor big files, you should be able to introduce a \"compound-file\"  \ndatatype and configure git to deal with specific X-Mbyte chunks of it  \nindependently.  This might not be a bad idea for native git as well.   \nWould need system-specific configuration.\n\nQ. How do you prevent massive data consumption by spurious tiny changes\nA. You have a few options:\n    1)  Configure your commit algorithm as above to not commit so often\n    2)  Configure a stepped commit-discard algorithm as described  \nabove in the \"How do you delete things\" question\n    3)  Archive unused data to secondary storage more often\n\nQ. What about all the unanswered questions?\nA. These are all the ones I could think of off the top of my head but  \nthere are at least a hundred more.  I'm pretty sure these are some of  \nthe most significant ones.\n\nQ. That's a great idea and I'll implement it right away!\nA. Yay!  (but that's not a question :-D)  Good luck and happy hacking.\n\nQ. That's a stupid idea and would never ever work!\nA. Thanks for your useful input! (but that's not a question either)   \nI'm sure anybody who takes up a project like this will consider such  \nopinions.\n\nQ. *flamage*\nA. I'm glad you have such strong opinions, feel free to to continue  \nto spam my /dev/null device (and that's also not a question).\n\nAll opinions and comments welcomed.\n\nCheers,\nKyle Moffett\n"},{"id":"45325","messageId":"46778A7A.1080403@hawkeye.stone.uk.eu.org","threadId":"8645","inReplyTo":"6E9A6F9E-8948-40F2-9129-1F1491D49D83@mac.com","subject":"Re: Versioning file system","fromName":"Jack Stone","fromEmail":"jack@hawkeye.stone.uk.eu.org","sentAt":"2007-06-19T07:49:14Z","receivedAt":"2007-06-19T07:49:14Z","isPatch":false,"sender":{"key":"jack@hawkeye.stone.uk.eu.org","avatar":null},"body":"Kyle Moffett wrote:\n> On Jun 18, 2007, at 13:56:05, Bryan Henderson wrote:\n>>> The question remains is where to implement versioning: directly in\n>>> individual filesystems or in the vfs code so all filesystems can use it?\n>>\n>> Or not in the kernel at all.  I've been doing versioning of the types\n>> I described for years with user space code and I don't remember\n>> feeling that I compromised in order not to involve the kernel.\n>>\n>> Of course, if you want to do it with snapshots and COW, you'll have to\n>> ask where in the kernel to put that, but that's not a file versioning\n>> question; it's the larger snapshot question.\n> \n> What I think would be particularly interesting in this domain is\n> something similar in concept to GIT, except in a file-system:\n>   1) Redundancy is easy, you just ensure that you have at least \"N\"\n> distributed copies of each object, where \"N\" is some function of the\n> object itself.\n>   2) Network replication is easy, you look up objects based on the SHA-1\n> stored in the parent directory entry and cache them where needed (IE:\n> make the \"N\" function above dynamic based on frequency of access on a\n> given computer).\n>   3) Snapshots are easy and cheap; an RO snapshot is a tag and an RW\n> snapshot is a branch.  These can be easily converted between.\n>   4) Compression is easy; you can compress objects based on any\n> arbitrary configurable criteria and the filesystem will record whether\n> or not an object is compressed.  You can also compress differently when\n> archiving objects to secondary storage.\n>   5) Streaming fsck-like verification is easy; ensure the hash name\n> field matches the actual hash of the object.\n>   6) Fsck is easy since rollback is trivial, you can always revert to an\n> older tree to boot and start up services before attempting resurrection\n> of lost objects and trees in the background.\n>   7) Multiple-drive or multiple-host storage pools are easy:  Think the\n> git \"alternates\" file.\n>   8) Network filesystem load-balancing is easy; SHA-1s are essentially\n> random so you can just assign SHA-1 prefixes to different systems for\n> data storage and your data is automatically split up.\n> \n> \n> Other issues:\n> \n> Q. How do you deal with block allocation?\n> A. Same way other filesystems deal with block allocation. \n> Reference-counting gets tricky, especially across a network, but it's\n> easy to play it safe with simple cross-network refcount-journalling. \n> Since the _only_ thing that needs journalling is the refcounts and\n> block-free data, you need at most a megabyte or two of journal.  If in\n> doubt, it's easy to play it safe and keep an extra refcount around for\n> an in-the-background consistency check later on.  When networked-gitfs\n> systems crash, you just assume they still have all the refcounts they\n> had at the moment they died, and compare notes when they start back up\n> again.  If a node has a cached copy of data on its local disk then it\n> can just nonatomically increment the refcount for that object in its own\n> RAM (ordered with respect to disk-flushes, of course) and tell its peers\n> at some point.  A node should probably cache most of its working set on\n> local disk for efficiency; it's trivially verified against updates from\n> other nodes and provides an easy way to keep refcounts for such data. \n> If a node increments the refcount on such data and dies before getting\n> that info out to its peers, then when it starts up again its peers will\n> just be told that it has a \"new\" node with insufficient replication and\n> they will clone it out again properly.  For networked\n> refcount-increments you can do one of 2 things: (1) Tell at least X many\n> peers and wait for them to sync the update out to disk, or (2) Get the\n> object from any peer (at least one of whom hopefully has it in RAM) and\n> save it to local disk with an increased refcount.\n> \n> Q. How do you actually delete things?\n> A. Just replace all the to-be-erased tree and commit objects before a\n> specified point with \"History erased\" objects with their SHA-1's\n> magically set to that of the erased objects.  If you want you may delete\n> only the \"tree\" objects and leave the commits intact.  If you delete a\n> whole linear segment of history then you can just use a single \"History\n> erased\" commit object with its parent pointed to the object before the\n> erased segment.  Probably needs some form of back-reference storage to\n> make it efficient; not sure how expensive that would be.  This would\n> allow making a bunch of snapshots and purging them logarithmically based\n> on passage of time.  For instance, you might have snapshots of every 5\n> minutes for the last hour, every 30 minutes for the last day, every 4\n> hours for the last week, every day for the last month, once per week for\n> the last year, once per month for the last 5 years, and once per year\n> beyond that.\n> \n> That's pretty impressive data-recovery resolution, and it accounts for\n> only 200 unique commits after it's been running for 10 years.\n> \n> Q. How do you archive data?\n> A. Same as deleting, except instead of a \"History erased\" object you\n> would use a \"History archived\" object with a little bit of string data\n> to indicate which volume it's stored on (and where on the volume).  When\n> you stick that volume into the system you could easily tell the kernel\n> to use it as an alternate for the given storage group.\n> \n> Q. What enforces data integrity?\n> A. Ensure that a new tree object and its associated sub objects are on\n> disk before you delete the old one.  Doesn't need any actual full syncs\n> at all, just barriers.  If you replace the tree object before write-out\n> is complete then just skip writing the old one and write the new one in\n> its place.\n> \n> Q. What consists of a \"commit\"?\n> A. Anything the administrator wants to define it as.  Useful algorithms\n> include: \"Once per x Mbyte of page dirtying\", \"Once per 5 min\", \"Only\n> when sync() or fsync() are called\", \"Only when gitfs-commit is called\". \n> You could even combine them:  \"Every x Mbyte of page dirtying or every 5\n> minutes, whichever is shorter (or longer, depending on admin\n> requirements)\".  There would also be appropriate syscalls to trigger\n> appropriate git-like behavior.  Network-accessible gitfs would want to\n> have mechanisms to trigger commits based on activity on other systems\n> (needs more thought).\n> \n> Q. How do you access old versions?\n> A. Mount another instance of the filesystem with an SHA-1 ID, a\n> tag-name, or a branch-name in a special mount option.  Should be user\n> accessible with some restrictions (needs more thought).\n> \n> Q. How do you deal with conflicts on networked filesystems.\n> A. Once again, however the administrator wants to deal with them.  Options:\n>    1)  Forcibly create a new branch for the conflicted tree.\n>    2)  Attempt to merge changes using the standard git-merge semantics\n>    3)  Merge independent changes to different files and pick one for\n> changes to the same file\n>    4)  Your Algorithm Here(TM).  GIT makes it easy to extend\n> conflict-resolution.\n> \n> Q. How do you deal with little scattered changes in big (or sparse) files?\n> A. Two questions, two answers:  For sparse files, git would need\n> extending to understand (and hash) the nature of the sparse-ness.  For\n> big files, you should be able to introduce a \"compound-file\" datatype\n> and configure git to deal with specific X-Mbyte chunks of it\n> independently.  This might not be a bad idea for native git as well. \n> Would need system-specific configuration.\n> \n> Q. How do you prevent massive data consumption by spurious tiny changes\n> A. You have a few options:\n>    1)  Configure your commit algorithm as above to not commit so often\n>    2)  Configure a stepped commit-discard algorithm as described above\n> in the \"How do you delete things\" question\n>    3)  Archive unused data to secondary storage more often\n> \n> Q. What about all the unanswered questions?\n> A. These are all the ones I could think of off the top of my head but\n> there are at least a hundred more.  I'm pretty sure these are some of\n> the most significant ones.\n> \n> Q. That's a great idea and I'll implement it right away!\n> A. Yay!  (but that's not a question :-D)  Good luck and happy hacking.\n> \n> Q. That's a stupid idea and would never ever work!\n> A. Thanks for your useful input! (but that's not a question either)  I'm\n> sure anybody who takes up a project like this will consider such opinions.\n> \n> Q. *flamage*\n> A. I'm glad you have such strong opinions, feel free to to continue to\n> spam my /dev/null device (and that's also not a question).\n> \n> All opinions and comments welcomed.\n> \n> Cheers,\n> Kyle Moffett\n> \n> \n\nIt sounds brilliant and I'd love to have a got at implementing it but I\ndon't know enough (yet :-D) about how git works, a little research is\ncalled for I think.\n\nJack\n"},{"id":"45328","messageId":"20070619075857.GA2944@brong.net","threadId":"8645","inReplyTo":"6E9A6F9E-8948-40F2-9129-1F1491D49D83@mac.com","subject":"Re: Versioning file system","fromName":"Bron Gondwana","fromEmail":"brong@fastmail.fm","sentAt":"2007-06-19T07:58:57Z","receivedAt":"2007-06-19T07:58:57Z","isPatch":false,"sender":{"key":"brong@fastmail.fm","avatar":"https://gravatar.com/avatar/9b7b59d9f9f104cf348df27c705c0ab5537bd40f291b0adf7885d5d1cea10d00?d=mp&s=160"},"body":"On Mon, Jun 18, 2007 at 11:10:42PM -0400, Kyle Moffett wrote:\n> On Jun 18, 2007, at 13:56:05, Bryan Henderson wrote:\n>>> The question remains is where to implement versioning: directly in \n>>> individual filesystems or in the vfs code so all filesystems can use it?\n>>\n>> Or not in the kernel at all.  I've been doing versioning of the types I \n>> described for years with user space code and I don't remember feeling that \n>> I compromised in order not to involve the kernel.\n>\n> What I think would be particularly interesting in this domain is something \n> similar in concept to GIT, except in a file-system:\n\nI've written a couple of user-space things very much like this - one\nbeing a purely database (blobs in database, yeah I know) system for\nmanaging medical data, where signatures and auditability were the most\nimportant part of the system.  Performance really wasn't a\nconsideration.\n\nThe other one is my current job, FastMail - we have a virtual filesystem\nwhich uses files stored by sha1 on ordainary filesystems for data\nstorage and a database for metadata (filename to sha1 mappings, mtime,\nmimetype, directory structure, etc).\n\nMultiple machine distribution is handled by a daemon on each machine\nwhich can be asked to make sure the file gets sent out to every machine\nthat matches the prefix and will only return success once it's written\nto at least one other machine.  Database replication is a different\nbeast.\n\n\nIt can work, but there's one big pain at the file level: no mmap.\n\nIf you don't want to support mmap it can work reasonably happily, though\nyou may want to keep your sha1 (or other digest) state as well as the\nfinal digest so you can cheaply calculate the digest for a small append\nwithout walking the entire file.  You may also want to keep state\ncheckpoints every so often along a big file so that truncates don't cost\ntoo much to recalculate.\n\nLuckily in a userspace VFS that's only accessed via FTP and DAV we can\nsupport a limited set of operations (basically create, append, read,\ndelete)  You don't get that luxury for a general purpose filesystem, and\nthat's the problem.  There will always be particular usage patterns\n(especially something that mmaps or seeks and touches all over the place\nlike a loopback mounted filesystem or a database file) that just dodn't\nwork for file-level sha1s.\n\n\nIt does have some lovely properties though.  I'd enjoy working in an\nenvionment that didn't look much like POSIX but had the strong\nguarantees and auditability that addressing by sha1 buys you.\n\nBron.\n\n\n"},{"id":"45332","messageId":"46a038f90706190209h7cbde42h34d6ac819711c3d3@mail.gmail.com","threadId":"8645","inReplyTo":"6E9A6F9E-8948-40F2-9129-1F1491D49D83@mac.com","subject":"Re: Versioning file system","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-06-19T09:09:58Z","receivedAt":"2007-06-19T09:09:58Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/19/07, Kyle Moffett <mrmacman_g4@mac.com> wrote:\n> What I think would be particularly interesting in this domain is\n> something similar in concept to GIT, except in a file-system:\n\nperhaps stating the blindingly obvious, but there was an early\nimplementation of a FUSE-based gitfs --\nhttp://www.sfgoth.com/~mitch/linux/gitfs/\n\ncheers,\n\n\nmartin\n"},{"id":"45369","messageId":"f591k5$odb$2@sea.gmane.org","threadId":"8645","inReplyTo":"6E9A6F9E-8948-40F2-9129-1F1491D49D83@mac.com","subject":"Re: Versioning file system","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-06-19T16:52:22Z","receivedAt":"2007-06-19T16:52:22Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Kyle Moffett wrote:\n> On Jun 18, 2007, at 13:56:05, Bryan Henderson wrote:\n\n>>> The question remains is where to implement versioning: directly in  \n>>> individual filesystems or in the vfs code so all filesystems can  \n>>> use it?\n>>\n>> Or not in the kernel at all.  I've been doing versioning of the  \n>> types I described for years with user space code and I don't  \n>> remember feeling that I compromised in order not to involve the  \n>> kernel.\n>>\n>> Of course, if you want to do it with snapshots and COW, you'll have  \n>> to ask where in the kernel to put that, but that's not a file  \n>> versioning question; it's the larger snapshot question.\n> \n> What I think would be particularly interesting in this domain is  \n> something similar in concept to GIT, except in a file-system\n[cut]\n\nHow it relates to ext3cow versioning (snapshotting) filesystem,\nfor example? ext3cow assumes linear history, which simplifies things\na bit.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"45395","messageId":"91F471D5-DFDB-4C4B-8374-292B2DF35F10@mac.com","threadId":"8645","inReplyTo":"20070619075857.GA2944@brong.net","subject":"Re: Versioning file system","fromName":"Kyle Moffett","fromEmail":"mrmacman_g4@mac.com","sentAt":"2007-06-20T02:43:01Z","receivedAt":"2007-06-20T02:43:01Z","isPatch":false,"sender":{"key":"mrmacman_g4@mac.com","avatar":null},"body":"On Jun 19, 2007, at 03:58:57, Bron Gondwana wrote:\n> On Mon, Jun 18, 2007 at 11:10:42PM -0400, Kyle Moffett wrote:\n>> On Jun 18, 2007, at 13:56:05, Bryan Henderson wrote:\n>>>> The question remains is where to implement versioning: directly  \n>>>> in individual filesystems or in the vfs code so all filesystems  \n>>>> can use it?\n>>>\n>>> Or not in the kernel at all.  I've been doing versioning of the  \n>>> types I described for years with user space code and I don't  \n>>> remember feeling that I compromised in order not to involve the  \n>>> kernel.\n>>\n>> What I think would be particularly interesting in this domain is  \n>> something similar in concept to GIT, except in a file-system:\n>\n> [...snip...]\n>\n> It can work, but there's one big pain at the file level: no mmap.\n\nIMHO it's actually not that bad.  The \"gitfs\" would divide larger  \nfiles up into manageable chunks (say 4MB) which could be quickly  \nSHA-1ed.  When a file is mmapped and partially modified, the SHA-1  \nwould be marked as locally invalid, but since mmap() loses most  \nconsistency guarantees that's OK.  A time or writeout based \"commit\"  \nscheme might still freeze, SHA-1, and write-out the page at regular  \nintervals without the program's knowledge, but since you only have to  \nSHA-1 the relatively-small 4MB chunk (which is about to hit disk  \nanyways), it's not a significant time penalty.  Even if under memory  \npressure and swapping data out to disk you don't have to update the  \nSHA-1 and create a new commit as long as you keep a reference to the  \nobject stored in the volume header somewhere and maintain the \"SHA-1  \nout-of-date\" bit.\n\nA program which carefully uses msync() would be fine, of course (with  \nproper configuration) as that would create a new commit as appropriate.\n\nSince mmap() is poorly defined on network filesystems in the absence  \nof msync(), I don't see that such behaviour would be a problem.  And  \nit certainly would be fine on local filesystems as there you can just  \nstuff the \"SHA-1 out-of-date\" bit and a reference to the parent  \ncommit and path in the object itself.  Then you just need to keep a  \nuseful reference to that object in a table somewhere in the volume  \nand you're set.\n\n> If you don't want to support mmap it can work reasonably happily,  \n> though you may want to keep your sha1 (or other digest) state as  \n> well as the final digest so you can cheaply calculate the digest  \n> for a small append without walking the entire file.  You may also  \n> want to keep state checkpoints every so often along a big file so  \n> that truncates don't cost too much to recalculate.\n\nThat may be worth it even if the file is divided into 4MB chunks (or  \nother configurable value), but it would need benchmarking.\n\n> Luckily in a userspace VFS that's only accessed via FTP and DAV we  \n> can support a limited set of operations (basically create, append,  \n> read, delete)  You don't get that luxury for a general purpose  \n> filesystem, and that's the problem.  There will always be  \n> particular usage patterns (especially something that mmaps or seeks  \n> and touches all over the place like a loopback mounted filesystem  \n> or a database file) that just dodn't work for file-level sha1s.\n\nI'd think that loopback-mounted filesystems wouldn't be that difficult\n   1)  Set the SHA-1 block size appropriately to divide the big file  \ninto a bunch of little manageable files.  Could conceivably be multi- \nlayered like directories, depending on the size of the file.\n   2)  Mark the file as exempt from normal commits (IE: without  \nspecial syscalls or fsync/msync() on the file itself, it is never  \nupdated in the tree objects.\n   3)  Set up the loopback device to call the gitfs commit code when  \nit receives barriers or flushes from the parent filesystem.\n\nAnd database files aren't a big issue.  I have yet to see a networked  \nfilesystem which you could stick a MySQL database on it from one node  \nand expect to get useful/recent read results from other nodes.  If  \nyou really wanted something like that for such a \"gitfs\", you could  \njust add code to MySQL to create a gitfs commit every N transactions  \nand not otherwise.  The best part is: that would make online MySQL  \nbackups from another node trivial!  Just pick any arbitrary  \nappropriate commit object and mount that object, then \"cp -a  \nmysql_db_dir mysql_backup_dir\".  That's not to say it wouldn't have a  \nperformance penalty, but for some people the performance penalty  \nmight be worth it.\n\nOh, and for those programs which want multi-master replication, this  \nmakes it ten times easier:\n   1)  Put each master-server on a different gitfs branch\n   2)  Write your program as gitfs aware.  Make it create gitfs  \ncommits at appropriate times (so the data is accessible from other  \nnodes).\n   3)  Come up with a useful non-interactive database-file merge  \nalgorithm.  Useful examples of different kinds of merge engines may  \nbe found in the git project.  This should take $BASE_VERSION,  \n$NEWVERSION1, $NEWVERSION2, and produce a $MERGEDVERSION.  A good  \nalgorithm should probably pick a safe default and save a \"conflict\"  \nentry in the face of conflicting changes.\n   4)  Hook your merge algorithm into the gitfs mechanics using some  \nto-be-defined API.\n   5)  Whenever your software does a database-file commit it sends  \nout a little notification to the other nodes (maybe using a gitfs API?)\n   6)  Run a periodic (as defined by the admin yet again) thread on  \neach node which does branch merging.  When two or more branches have  \ndifferent SHA-1 sums the servers will rotate the merging task between  \nthem.  The thus-selected server will merge changes from the other  \nserver(s) into its current working copy.  With 2 servers this means  \nthat the maximum delay between one server making a change and the  \nother server seeing it will be 2 times the merge interval.\n   7)  For small pools of servers a simple rotated-merge-master  \nalgorithm would work.  For larger pools you would need to come up  \nwith some logarithmic rotating-merge-node algorithm to evenly divide  \nthe work of propagating changes across all nodes.\n\n> It does have some lovely properties though.  I'd enjoy working in  \n> an envionment that didn't look much like POSIX but had the strong  \n> guarantees and auditability that addressing by sha1 buys you.\n\nI'd like to think we can have our cake and eat it too :-D.  POSIX  \nrequirements should be doable on the local system and can be mimiced  \nwell enough on networked filesystems (albeit with update latency)  \nthat most programs won't care.  If you're the only person modifying  \nfiles on gitfs, regardless of what node they are stored on, it should  \nhave the same behavior as local files (since with gitfs caching they  \nwould *become* local files too :-D).  The few programs that do care  \nabout POSIX atomicity across networked filesystems (which is already  \nmostly implementation defined) could probably be updated to map gitfs  \ncommits and merges into their own internal transactions and do just  \nfine.\n\nCheers,\nKyle Moffett\n"}]}