threads / discuss / 17062

Curious about details of optimization of object database...

Subject: Curious about details of optimization of object database...

## tl;dr

5 messages between Jan 9, 2009 and Jan 9, 2009.

replies: 4people: 5as markdown or json

chris@seberino.org· Jan 9, 2009, 17:46 UTC · lore

I'm told a commit is *not* a patch (diff), but, rather a copy of the entire tree.

Can anyone say, in a few sentences, how git avoids needing to keep multiple slightly different copies of entire files without just storing lots of patches/diffs?

cs
Matthieu Moy· Jan 9, 2009, 17:55 UTC · re: chris@seberino.org · lore

Re: Curious about details of optimization of object database...

chris@seberino.org writes:
> I'm told a commit is *not* a patch (diff), but, rather a copy of the entire
> tree.

Conceptually, yes. But obviously, the storage format (pack) does what people usually call "delta-compression", which is basically storing only the diff against another, similar object.

-- 
Matthieu
Nicolas Pitre· Jan 9, 2009, 18:34 UTC · re: Matthieu Moy · lore

Re: Curious about details of optimization of object database...

On Fri, 9 Jan 2009, Matthieu Moy wrote:
Show 8 quoted lines
> chris@seberino.org writes:
> 
> > I'm told a commit is *not* a patch (diff), but, rather a copy of the entire
> > tree.
> 
> Conceptually, yes. But obviously, the storage format (pack) does what
> people usually call "delta-compression", which is basically storing
> only the diff against another, similar object.

Also, since objects representing files and directories are named after their actual content, having two commits with identical files and directories will of course share the same blob and tree objects for those identical parts.

Nicolas
David Brown· Jan 9, 2009, 17:56 UTC · re: chris@seberino.org · lore

Re: Curious about details of optimization of object database...

On Fri, Jan 09, 2009 at 09:46:23AM -0800, chris@seberino.org wrote:
Show 6 quoted lines
>I'm told a commit is *not* a patch (diff), but, rather a copy of the entire
>tree.
>
>Can anyone say, in a few sentences, how git avoids needing to keep multiple
>slightly different copies of entire files without just storing lots of
>patches/diffs?
   Documentation/technical/pack-heuristics.txt
David
Boyd Stephen Smith Jr.· Jan 9, 2009, 19:07 UTC · re: chris@seberino.org · lore

Re: Curious about details of optimization of object database...

On Friday 2009 January 09 11:46:23 chris@seberino.org wrote:
>I'm told a commit is *not* a patch (diff), but, rather a copy of the entire
>tree.

It's even more than that. A commit object contains its message, the SHA of the tree, and zero or more SHAs for its parents.

>Can anyone say, in a few sentences, how git avoids needing to keep multiple
>slightly different copies of entire files without just storing lots of
>patches/diffs?

Loose objects can have large swaths of duplicated data. However, git also supports storing objects in a packed format, which uses delta compression to reduce the duplication to close to nothing.

Some examples: Sizes are from "du -sh .git ."; The .git directory stores all the objects as well as the repository configuration, refs, reflogs, etc. The . directory has .git and a clean checkout of master.

The LinuxPMI (http://linuxpmi.org/) tree: 41M .git 83M . (So, the storage is actually a bit smaller than the checkout; 984 objects; 140 commits)

A small project between me an my flatmates: 309K .git 3.6M . (Here, the storage is significantly smaller than the checkout; 786 objects; 155 commits)

My repository that tracks my dotfiles: 124K .git 176K . (113 objects; 28 commits)

-- 
Boyd Stephen Smith Jr.                     ,= ,-_-. =. 
bss@iguanasuicide.net                     ((_/)o o(\_))
ICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' 
http://iguanasuicide.net/                      \_/     

← back to recent threads