{"thread":{"id":"26084","subject":"How to unpack recent objects?","startedAt":"2010-12-16T20:33:43Z","lastAt":"2010-12-16T23:12:52Z","messageCount":6,"participants":["Phillip Susi","Jonathan Nieder","Nicolas Pitre","Jakub Narebski"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"158214","messageId":"4D0A77A7.9080702@cfl.rr.com","threadId":"26084","inReplyTo":null,"subject":"How to unpack recent objects?","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2010-12-16T20:33:43Z","receivedAt":"2010-12-16T20:33:43Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"It looks like you can use git-unpack-objects to unpack ALL objects, but\nhow can you unpack only recent ones that you are likely to use while\nleaving the ancient stuff packed?  Ideally I want to unpack all file\nobjects from the current commit, and a reasonable number of commit\nobjects going back into the history so accessing them with checkout,\ndiff, log, etc will be fast.\n"},{"id":"158215","messageId":"20101216204059.GA3830@burratino","threadId":"26084","inReplyTo":"4D0A77A7.9080702@cfl.rr.com","subject":"Re: How to unpack recent objects?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-12-16T20:40:59Z","receivedAt":"2010-12-16T20:40:59Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi Phillip,\n\nPhillip Susi wrote:\n\n> It looks like you can use git-unpack-objects to unpack ALL objects, but\n> how can you unpack only recent ones that you are likely to use while\n> leaving the ancient stuff packed?  Ideally I want to unpack all file\n> objects from the current commit, and a reasonable number of commit\n> objects going back into the history so accessing them with checkout,\n> diff, log, etc will be fast.\n\nHave you tried the experiment?  You can pack all objects and then make\na few commits that do not reuse any blobs from before on top of that;\nthen \"cp -a\" the repository and use \"git gc --aggressive\" to get one\nbig pack as a control.  Then it should be possible to time checkout,\ndiff, log, etc[1].\n\nIt would also be interesting to know what the nature of these objects\nare, in case it is possible to speed things up some other way.\n\nJonathan\n\n[1] My uninformed guess is that the packed version will be faster,\nbecause of cache effects among other reasons.  The point of loose\nobjects is to speed up writing objects rather than reading them.\nBut I'd be happy to be surprised.\n"},{"id":"158216","messageId":"alpine.LFD.2.00.1012161616170.10437@xanadu.home","threadId":"26084","inReplyTo":"4D0A77A7.9080702@cfl.rr.com","subject":"Re: How to unpack recent objects?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-12-16T21:19:01Z","receivedAt":"2010-12-16T21:19:01Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 16 Dec 2010, Phillip Susi wrote:\n\n> It looks like you can use git-unpack-objects to unpack ALL objects, but\n> how can you unpack only recent ones that you are likely to use while\n> leaving the ancient stuff packed?  Ideally I want to unpack all file\n> objects from the current commit, and a reasonable number of commit\n> objects going back into the history so accessing them with checkout,\n> diff, log, etc will be fast.\n\nWhat makes you think that unpacking them will actually make the access \nto them faster?  Instead, you should consider _repacking_ them, \nultimately using the --aggressive parameter with the gc command, if you \nwant faster accesses.\n\n\nNicolas\n"},{"id":"158223","messageId":"4D0A8D83.9080705@cfl.rr.com","threadId":"26084","inReplyTo":"alpine.LFD.2.00.1012161616170.10437@xanadu.home","subject":"Re: How to unpack recent objects?","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2010-12-16T22:06:59Z","receivedAt":"2010-12-16T22:06:59Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 12/16/2010 4:19 PM, Nicolas Pitre wrote:\n> What makes you think that unpacking them will actually make the access \n> to them faster?  Instead, you should consider _repacking_ them, \n> ultimately using the --aggressive parameter with the gc command, if you \n> want faster accesses.\n\nBecause decompressing and undeltifying the objects in the pack file\ntakes a fair amount of cpu time.  It seems a waste to do this for the\nsame set of objects repeatedly rather than just keeping them loose.\n"},{"id":"158224","messageId":"m339r0ncai.fsf@localhost.localdomain","threadId":"26084","inReplyTo":"4D0A8D83.9080705@cfl.rr.com","subject":"Re: How to unpack recent objects?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-12-16T22:18:50Z","receivedAt":"2010-12-16T22:18:50Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Phillip Susi <psusi@cfl.rr.com> writes:\n> On 12/16/2010 4:19 PM, Nicolas Pitre wrote:\n\n> > What makes you think that unpacking them will actually make the access \n> > to them faster?  Instead, you should consider _repacking_ them, \n> > ultimately using the --aggressive parameter with the gc command, if you \n> > want faster accesses.\n> \n> Because decompressing and undeltifying the objects in the pack file\n> takes a fair amount of cpu time.  It seems a waste to do this for the\n> same set of objects repeatedly rather than just keeping them loose.\n\nLoose objects are also compressed.  \n\nBesides git has some kind of delta cache, so when you are accessing a\nfew objects (like e.g. when doing 'git log -p' - log + diff) you don't\nneed to undeltify and uncompress the same objects repeatedly.\n\nAlso in practice it is IO that is bottleneck, not CPU.  And having\nmany files is bad for filesystem cache.  Originally packfiles were for\nthe network transfer, but it turned out that they are better also as\non-disk format.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"158241","messageId":"alpine.LFD.2.00.1012161743530.10437@xanadu.home","threadId":"26084","inReplyTo":"4D0A8D83.9080705@cfl.rr.com","subject":"Re: How to unpack recent objects?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-12-16T23:12:52Z","receivedAt":"2010-12-16T23:12:52Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 16 Dec 2010, Phillip Susi wrote:\n\n> On 12/16/2010 4:19 PM, Nicolas Pitre wrote:\n> > What makes you think that unpacking them will actually make the access \n> > to them faster?  Instead, you should consider _repacking_ them, \n> > ultimately using the --aggressive parameter with the gc command, if you \n> > want faster accesses.\n> \n> Because decompressing and undeltifying the objects in the pack file\n> takes a fair amount of cpu time.  It seems a waste to do this for the\n> same set of objects repeatedly rather than just keeping them loose.\n\nWell, here are a couple implementation details you might not know about:\n\n1) Loose objects are compressed too.  So you gain nothing on that front \n   by keeping objects loose.\n\n2) Delta ordering is so that recent objects, i.e. those belonging to \n   most recent commits, are not delta compressed but rather used as base \n   objects for \"older\" objects to delta against.  So in practice, the \n   cost of undeltifying objects is pushed towards objects that you're \n   most unlikely to access frequently.\n\n3) Object placement within the pack is also optimized so that \n   objects belonging to recent commits are close together, and walking \n   them creates a linear IO access pattern which is much faster than \n   accessing random individual files as loose objects are.\n\n4) Packed objects take considerably less space than loose ones which \n   makes for much better usage of the file system cache in the operating \n   system.  This largely outweights the cost of undeltifying objects.\n\n5) Git also keeps a cache of most frequently referenced objects when \n   replaying delta chains so deep deltas don't bring exponential costs.\n\nAnd, in some cases, Git does even pick up the content of an object by \nusing its checked out form in the working directory directly instead of \nlocating and decompressing the object data.\n\nSo you shouldn't have to worry on that front.  Git is not the fastest \nSCM out there just by luck.\n\n\nNicolas\n"}]}