threads / discuss / 26084

How to unpack recent objects?

Subject: How to unpack recent objects?

## tl;dr

6 messages between Dec 16, 2010 and Dec 16, 2010.

replies: 5people: 4as markdown or json

Phillip Susi· Dec 16, 2010, 20:33 UTC · lore

It looks like you can use git-unpack-objects to unpack ALL objects, but how can you unpack only recent ones that you are likely to use while leaving the ancient stuff packed? Ideally I want to unpack all file objects from the current commit, and a reasonable number of commit objects going back into the history so accessing them with checkout, diff, log, etc will be fast.

Jonathan Nieder· Dec 16, 2010, 20:40 UTC · re: Phillip Susi · lore

Re: How to unpack recent objects?

Hi Phillip,
Phillip Susi wrote:
Show 6 quoted lines
> It looks like you can use git-unpack-objects to unpack ALL objects, but
> how can you unpack only recent ones that you are likely to use while
> leaving the ancient stuff packed?  Ideally I want to unpack all file
> objects from the current commit, and a reasonable number of commit
> objects going back into the history so accessing them with checkout,
> diff, log, etc will be fast.

Have you tried the experiment? You can pack all objects and then make a few commits that do not reuse any blobs from before on top of that; then "cp -a" the repository and use "git gc --aggressive" to get one big pack as a control. Then it should be possible to time checkout, diff, log, etc[1].

It would also be interesting to know what the nature of these objects are, in case it is possible to speed things up some other way.

Jonathan

[1] My uninformed guess is that the packed version will be faster, because of cache effects among other reasons. The point of loose objects is to speed up writing objects rather than reading them. But I'd be happy to be surprised.

Nicolas Pitre· Dec 16, 2010, 21:19 UTC · re: Phillip Susi · lore

Re: How to unpack recent objects?

On Thu, 16 Dec 2010, Phillip Susi wrote:
Show 6 quoted lines
> It looks like you can use git-unpack-objects to unpack ALL objects, but
> how can you unpack only recent ones that you are likely to use while
> leaving the ancient stuff packed?  Ideally I want to unpack all file
> objects from the current commit, and a reasonable number of commit
> objects going back into the history so accessing them with checkout,
> diff, log, etc will be fast.

What makes you think that unpacking them will actually make the access to them faster? Instead, you should consider _repacking_ them, ultimately using the --aggressive parameter with the gc command, if you want faster accesses.

Nicolas
Phillip Susi· Dec 16, 2010, 22:06 UTC · re: Nicolas Pitre · lore

Re: How to unpack recent objects?

On 12/16/2010 4:19 PM, Nicolas Pitre wrote:
> What makes you think that unpacking them will actually make the access 
> to them faster?  Instead, you should consider _repacking_ them, 
> ultimately using the --aggressive parameter with the gc command, if you 
> want faster accesses.

Because decompressing and undeltifying the objects in the pack file takes a fair amount of cpu time. It seems a waste to do this for the same set of objects repeatedly rather than just keeping them loose.

Jakub Narebski· Dec 16, 2010, 22:18 UTC · re: Phillip Susi · lore

Re: How to unpack recent objects?

Phillip Susi <psusi@cfl.rr.com> writes:
> On 12/16/2010 4:19 PM, Nicolas Pitre wrote:
Show 8 quoted lines
> > What makes you think that unpacking them will actually make the access 
> > to them faster?  Instead, you should consider _repacking_ them, 
> > ultimately using the --aggressive parameter with the gc command, if you 
> > want faster accesses.
> 
> Because decompressing and undeltifying the objects in the pack file
> takes a fair amount of cpu time.  It seems a waste to do this for the
> same set of objects repeatedly rather than just keeping them loose.
Loose objects are also compressed.  

Besides git has some kind of delta cache, so when you are accessing a few objects (like e.g. when doing 'git log -p' - log + diff) you don't need to undeltify and uncompress the same objects repeatedly.

Also in practice it is IO that is bottleneck, not CPU. And having many files is bad for filesystem cache. Originally packfiles were for the network transfer, but it turned out that they are better also as on-disk format.

-- 
Jakub Narebski
Poland
ShadeHawk on #git
Nicolas Pitre· Dec 16, 2010, 23:12 UTC · re: Phillip Susi · lore

Re: How to unpack recent objects?

On Thu, 16 Dec 2010, Phillip Susi wrote:
Show 9 quoted lines
> On 12/16/2010 4:19 PM, Nicolas Pitre wrote:
> > What makes you think that unpacking them will actually make the access 
> > to them faster?  Instead, you should consider _repacking_ them, 
> > ultimately using the --aggressive parameter with the gc command, if you 
> > want faster accesses.
> 
> Because decompressing and undeltifying the objects in the pack file
> takes a fair amount of cpu time.  It seems a waste to do this for the
> same set of objects repeatedly rather than just keeping them loose.
Well, here are a couple implementation details you might not know about:
1) Loose objects are compressed too.  So you gain nothing on that front 
   by keeping objects loose.
2) Delta ordering is so that recent objects, i.e. those belonging to 
   most recent commits, are not delta compressed but rather used as base 
   objects for "older" objects to delta against.  So in practice, the 
   cost of undeltifying objects is pushed towards objects that you're 
   most unlikely to access frequently.
3) Object placement within the pack is also optimized so that 
   objects belonging to recent commits are close together, and walking 
   them creates a linear IO access pattern which is much faster than 
   accessing random individual files as loose objects are.
4) Packed objects take considerably less space than loose ones which 
   makes for much better usage of the file system cache in the operating 
   system.  This largely outweights the cost of undeltifying objects.
5) Git also keeps a cache of most frequently referenced objects when 
   replaying delta chains so deep deltas don't bring exponential costs.

And, in some cases, Git does even pick up the content of an object by using its checked out form in the working directory directly instead of locating and decompressing the object data.

So you shouldn't have to worry on that front. Git is not the fastest SCM out there just by luck.

Nicolas

← back to recent threads