{"thread":{"id":"17062","subject":"Curious about details of optimization of object database...","startedAt":"2009-01-09T17:46:23Z","lastAt":"2009-01-09T19:07:21Z","messageCount":5,"participants":["chris@seberino.org","David Brown","Matthieu Moy","Nicolas Pitre","Boyd Stephen Smith Jr."],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"99801","messageId":"20090109174623.GC12552@seberino.org","threadId":"17062","inReplyTo":null,"subject":"Curious about details of optimization of object database...","fromName":"","fromEmail":"chris@seberino.org","sentAt":"2009-01-09T17:46:23Z","receivedAt":"2009-01-09T17:46:23Z","isPatch":false,"sender":{"key":"chris@seberino.org","avatar":null},"body":"I'm told a commit is *not* a patch (diff), but, rather a copy of the entire\ntree.\n\nCan anyone say, in a few sentences, how git avoids needing to keep multiple\nslightly different copies of entire files without just storing lots of\npatches/diffs?\n\ncs\n"},{"id":"99803","messageId":"vpqzli01hzl.fsf@bauges.imag.fr","threadId":"17062","inReplyTo":"20090109174623.GC12552@seberino.org","subject":"Re: Curious about details of optimization of object database...","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2009-01-09T17:55:58Z","receivedAt":"2009-01-09T17:55:58Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"chris@seberino.org writes:\n\n> I'm told a commit is *not* a patch (diff), but, rather a copy of the entire\n> tree.\n\nConceptually, yes. But obviously, the storage format (pack) does what\npeople usually call \"delta-compression\", which is basically storing\nonly the diff against another, similar object.\n\n-- \nMatthieu\n"},{"id":"99802","messageId":"20090109175619.GA807@linode.davidb.org","threadId":"17062","inReplyTo":"20090109174623.GC12552@seberino.org","subject":"Re: Curious about details of optimization of object database...","fromName":"David Brown","fromEmail":"git@davidb.org","sentAt":"2009-01-09T17:56:19Z","receivedAt":"2009-01-09T17:56:19Z","isPatch":false,"sender":{"key":"git@davidb.org","avatar":"https://gravatar.com/avatar/94c86a2938470a74c2eac5e2b69afc0871f79a660295c02219597aba8cb101c1?d=mp&s=160"},"body":"On Fri, Jan 09, 2009 at 09:46:23AM -0800, chris@seberino.org wrote:\n>I'm told a commit is *not* a patch (diff), but, rather a copy of the entire\n>tree.\n>\n>Can anyone say, in a few sentences, how git avoids needing to keep multiple\n>slightly different copies of entire files without just storing lots of\n>patches/diffs?\n\n   Documentation/technical/pack-heuristics.txt\n\nDavid\n"},{"id":"99806","messageId":"alpine.LFD.2.00.0901091330010.9524@xanadu.home","threadId":"17062","inReplyTo":"vpqzli01hzl.fsf@bauges.imag.fr","subject":"Re: Curious about details of optimization of object database...","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-01-09T18:34:03Z","receivedAt":"2009-01-09T18:34:03Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 9 Jan 2009, Matthieu Moy wrote:\n\n> chris@seberino.org writes:\n> \n> > I'm told a commit is *not* a patch (diff), but, rather a copy of the entire\n> > tree.\n> \n> Conceptually, yes. But obviously, the storage format (pack) does what\n> people usually call \"delta-compression\", which is basically storing\n> only the diff against another, similar object.\n\nAlso, since objects representing files and directories are named after \ntheir actual content, having two commits with identical files and \ndirectories will of course share the same blob and tree objects for \nthose identical parts.\n\n\nNicolas\n"},{"id":"99810","messageId":"200901091307.33483.bss@iguanasuicide.net","threadId":"17062","inReplyTo":"20090109174623.GC12552@seberino.org","subject":"Re: Curious about details of optimization of object database...","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-09T19:07:21Z","receivedAt":"2009-01-09T19:07:21Z","isPatch":false,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Friday 2009 January 09 11:46:23 chris@seberino.org wrote:\n>I'm told a commit is *not* a patch (diff), but, rather a copy of the entire\n>tree.\n\nIt's even more than that.  A commit object contains its message, the SHA of \nthe tree, and zero or more SHAs for its parents.\n\n>Can anyone say, in a few sentences, how git avoids needing to keep multiple\n>slightly different copies of entire files without just storing lots of\n>patches/diffs?\n\nLoose objects can have large swaths of duplicated data.  However, git also \nsupports storing objects in a packed format, which uses delta compression to \nreduce the duplication to close to nothing.\n\nSome examples:\nSizes are from \"du -sh .git .\"; The .git directory stores all the objects as \nwell as the repository configuration, refs, reflogs, etc.  The . directory \nhas .git and a clean checkout of master.\n\nThe LinuxPMI (http://linuxpmi.org/) tree:\n41M     .git\n83M     .\n(So, the storage is actually a bit smaller than the checkout; 984 objects; 140 \ncommits)\n\nA small project between me an my flatmates:\n309K    .git\n3.6M    .\n(Here, the storage is significantly smaller than the checkout; 786 objects; \n155 commits)\n\nMy repository that tracks my dotfiles:\n124K    .git\n176K    .\n(113 objects; 28 commits)\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"}]}