{"thread":{"id":"240","subject":"proposal: delta based git archival","startedAt":"2005-04-22T09:03:41Z","lastAt":"2005-04-22T09:49:20Z","messageCount":3,"participants":["Michel Lespinasse","Jeffrey E. Hundstad","Jaime Medrano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1249","messageId":"20050422090341.GC22479@zoy.org","threadId":"240","inReplyTo":null,"subject":"proposal: delta based git archival","fromName":"Michel Lespinasse","fromEmail":"walken@zoy.org","sentAt":"2005-04-22T09:03:41Z","receivedAt":"2005-04-22T09:03:41Z","isPatch":false,"sender":{"key":"walken@zoy.org","avatar":null},"body":"I noticed people on this mailing list start talking about using blob deltas\nfor compression, and the basic issue that the resulting files are too small\nfor efficient filesystem storage. I thought about this a little and decided\nI should send out my ideas for discussion.\n\nIn my proposal, the current git object storage model (one compressed object\nper file) remains as the primary storage mechanism, however there would be\nsome kind of backup mechanism based on multiple deltas grouped in one file.\n\nFor example, suppose you're looking for an object with a hash of\neab75ce51622aa312bb0b03572d43769f420c347\n\nFirst you'd look at .git/objects/ea/b75ce51622aa312bb0b03572d43769f420c347 -\nif the file exists, that's your object.\n\nIf the file does not exist, you'd then look for .git/deltas/ea/b,\n.git/deltas/ea/b7, .git/deltas/ea/b75, .git/deltas/ea/b75c, ...\nup to some maximum search path lenght. You stop at the first file you can\nfind.\n\nSupposing that file is .git/deltas/ea/b7, it would contain a diff\n(let's assume unified format for now, though ideally it'd be better to\nhave something that allows binary file deltas too) of many archived\nobjects with hashes starting with eab7, compared to a different object\n(presumably some direct or indirect ancestor):\n\ndiff -u 8f5ba0203e31204c5c052d995a5b4449226bcfb5 eab75ce51622aa312bb0b03572d43769f420c347\n--- 8f5ba0203e31204c5c052d995a5b4449226bcfb5\n+++ eab75ce51622aa312bb0b03572d43769f420c347\n@@ -522,7 +522,7 @@\n....\ndiff -u 77dc2cb94930017f62b55b9706cbadda8c90f650 eab71c51dbc62797d6c903203de44cc6a734c05c\n--- 77dc2cb94930017f62b55b9706cbadda8c90f650\n+++ eab71c51dbc62797d6c903203de44cc6a734c05c\n@@ -560,13 +563,17 @@\n...\n\nBased on this delta file, we'd then look for the object\n8f5ba0203e31204c5c052d995a5b4449226bcfb5 (this process could require\nrecursively rebuilding that object) and try to build\neab75ce51622aa312bb0b03572d43769f420c347 by applying the delta and then\ndouble checking the hash.\n\nTo me the strenghts of this proposal would be:\n* It does not muddy the git object model - it just acts independently of it,\n  as a way to rebuild git objects from deltas\n* Old objects can be compressed by creating a delta with a close ancestor,\n  then erasing the original file storage for that object. The object delta\n  can be appended to an existing delta file (which avoids the small-file\n  storage issue), or if the delta file gets too big, it can be split off\n  into 16 smaller files based on the hashes of the objects this file stores\n  deltas for.\n* The system is flexible enough to explore different delta\n  strategies. For example one could decide to keep one object every 10\n  in the database and store other 9 as deltas based on the immediate\n  object ancestor, or any other tradeoff - and the system would still\n  work the same (with different performance tradeoffs though).\n\nDoes this sound insane ? Too complicated maybe ?\n\nIs there any kind of semi-standard binary-capable multiple-file diff format\nthat could be used for this application instead of unified diffs ?\n\n-- \nMichel \"Walken\" Lespinasse\n\"Bill Gates is a monocle and a Persian cat away from being the villain\nin a James Bond movie.\" -- Dennis Miller\n"},{"id":"1250","messageId":"4268C00C.7040308@mnsu.edu","threadId":"240","inReplyTo":"20050422090341.GC22479@zoy.org","subject":"Re: proposal: delta based git archival","fromName":"Jeffrey E. Hundstad","fromEmail":"jeffrey.hundstad@mnsu.edu","sentAt":"2005-04-22T09:12:44Z","receivedAt":"2005-04-22T09:12:44Z","isPatch":false,"sender":{"key":"jeffrey.hundstad@mnsu.edu","avatar":null},"body":"Michel Lespinasse wrote:\n\n>Does this sound insane ? Too complicated maybe ?\n>  \n>\nMy vote is YES on both counts.\n\nSimplicity and flexibility is what makes git a good thing; and imho this \nworks against that quite aggressively.\n\n-- \nJeffrey Hundstad\n\n"},{"id":"1254","messageId":"b008c6a40504220249477e70ae@mail.gmail.com","threadId":"240","inReplyTo":"20050422090341.GC22479@zoy.org","subject":"Re: proposal: delta based git archival","fromName":"Jaime Medrano","fromEmail":"overflow5@gmail.com","sentAt":"2005-04-22T09:49:20Z","receivedAt":"2005-04-22T09:49:20Z","isPatch":false,"sender":{"key":"overflow5@gmail.com","avatar":null},"body":"On 4/22/05, Michel Lespinasse <walken@zoy.org> wrote:\n> I noticed people on this mailing list start talking about using blob deltas\n> for compression, and the basic issue that the resulting files are too small\n> for efficient filesystem storage. I thought about this a little and decided\n> I should send out my ideas for discussion.\n> \n\nI've been thinking in another simpler approach.\n\nThe main benefit of using deltas is reducing the bandwith use in\npull/push. My idea is leaving the blob storage as it is by now and\nadding a new kind of object (remote) that acts as a link to an object\nin another repository.\n\nSo that, when you rsync, you don't have to get all the blobs (which\ncan be a lot of data), but only the sha1 of the new objects created.\nThen a remote object is created for each new object in the local\nrepository pointing to its location in the external repository.\n\nOnce the rsync is done, when git has to access any of the new objects\nthey can be fetched from the original location, so that only necessary\nobjects are transfered.\n\nThis way, the cost of a sync in terms of bandwith is nearly zero.\n\nI've been working on this, so if you think it to be a good idea, I can\nsend a patch when I get it fully working.\n\nRegards,\nJaime Medrano.\nhttp://jmedrano.sl-form.com\n"}]}