# Re: git-fetching from a big repository is slow

11 messages from 2006-12-14 to 2006-12-16. Participants: Shawn Pearce, Nicolas Pitre, Horst H. von Brand, Geert Bosch, Johannes Schindelin, Pazu, Robin Rosenberg.
Thread: https://gitlist.dev/t/43382

## Shawn Pearce, 2006-12-14 19:46

Subject: Re: git-fetching from a big repository is slow
Message-ID: <20061214194636.GO1747@spearce.org>
URL: https://gitlist.dev/e/20061214194636.GO1747%40spearce.org
In-Reply-To: <C287764F-6755-4291-A87A-3E8816E90B49@adacore.com>

```
Geert Bosch <bosch@adacore.com> wrote:
> Such special magic based on filenames is always a bad idea. Tomorrow  
> somebody
> comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
> compressed content. In the end git will be doing lots of magic and  
> still perform
> badly on unknown compressed content.
> 
> There is a very simple way of detecting compressed files: just look  
> at the
> size of the compressed blob and compare against the size of the  
> expanded blob.
> If the compressed blob has a non-trivial size which is close to the  
> expanded
> size, assume the file is not interesting as source or target for deltas.
> 
> Example:
>    if (compressed_size > expanded_size / 4 * 3 + 1024) {
>      /* don't try to deltify if blob doesn't compress well */
>      return ...;
>    }

And yet I get good delta compression on a number of ZIP formatted
files which don't get good additional zlib compression (<3%).
Doing the above would cause those packfiles to explode to about
10x their current size.

-- 

```

## Horst H. von Brand, 2006-12-14 22:12

Subject: Re: git-fetching from a big repository is slow
Message-ID: <200612142212.kBEMCVeu032626@laptop13.inf.utfsm.cl>
URL: https://gitlist.dev/e/200612142212.kBEMCVeu032626%40laptop13.inf.utfsm.cl
In-Reply-To: <spearce@spearce.org>

```
Shawn Pearce <spearce@spearce.org> wrote:

[...]

> And yet I get good delta compression on a number of ZIP formatted
> files which don't get good additional zlib compression (<3%).

.zip is something like a tar of the compressed files, if the files inside
the archive don't change, the deltas will be small.
-- 
Dr. Horst H. von Brand                   User #22616 counter.li.org
Departamento de Informatica                    Fono: +56 32 2654431
Universidad Tecnica Federico Santa Maria             +56 32 2654239

```

## Shawn Pearce, 2006-12-14 22:38

Subject: Re: git-fetching from a big repository is slow
Message-ID: <20061214223813.GC26202@spearce.org>
URL: https://gitlist.dev/e/20061214223813.GC26202%40spearce.org
In-Reply-To: <200612142212.kBEMCVeu032626@laptop13.inf.utfsm.cl>

```
"Horst H. von Brand" <vonbrand@inf.utfsm.cl> wrote:
> Shawn Pearce <spearce@spearce.org> wrote:
> 
> [...]
> 
> > And yet I get good delta compression on a number of ZIP formatted
> > files which don't get good additional zlib compression (<3%).
> 
> .zip is something like a tar of the compressed files, if the files inside
> the archive don't change, the deltas will be small.

Yes, especially when the new zip is made using the exact same
software with the same parameters, so the resulting compressed file
stream is identical for files whose content has not changed.  :-)

Since this is actually a JAR full of Java classes which have
been recompiled, its even more interesting that javac produced an
identical class file given the same input.  I've seen times where
it doesn't thanks to the automatic serialVersionUID field being
somewhat randomly generated.

-- 

```

## Geert Bosch, 2006-12-14 23:01

Subject: Re: git-fetching from a big repository is slow
Message-ID: <E30DCF6F-5D3E-4CA3-85D7-CD2847B86F86@adacore.com>
URL: https://gitlist.dev/e/E30DCF6F-5D3E-4CA3-85D7-CD2847B86F86%40adacore.com
In-Reply-To: <20061214194636.GO1747@spearce.org>

```

On Dec 14, 2006, at 14:46, Shawn Pearce wrote:
> And yet I get good delta compression on a number of ZIP formatted
> files which don't get good additional zlib compression (<3%).
> Doing the above would cause those packfiles to explode to about
> 10x their current size.

Yes, that's because for zip files each file in the archive is
compressed independently. Similar things might happen when
checking in uncompressed tar files with JPG's. The question
is whether you prefer bad time usage or bad space usage when
handling large binary blobs. Maybe we should use a faster,
less precise algorithm instead of giving up.

Still, I think doing anything based on filename is a mistake.
If we want to have a heuristic to prevent spending too much time
on deltifying large compressed files, the heuristic should be
based on content, not filename.

Maybe we could some "magic" as used by the file(1) command
that allows git to say a bit more about the content of blobs.
This could be used both for ordering files during deltification
and to determine wether to try deltification at all.

   -Geert


```

## Johannes Schindelin, 2006-12-14 23:15

Subject: Re: git-fetching from a big repository is slow
Message-ID: <Pine.LNX.4.63.0612150013390.3635@wbgn013.biozentrum.uni-wuerzburg.de>
URL: https://gitlist.dev/e/Pine.LNX.4.63.0612150013390.3635%40wbgn013.biozentrum.uni-wuerzburg.de
In-Reply-To: <20061214194636.GO1747@spearce.org>

```
Hi,

On Thu, 14 Dec 2006, Shawn Pearce wrote:

> Geert Bosch <bosch@adacore.com> wrote:
> > Such special magic based on filenames is always a bad idea. Tomorrow  
> > somebody
> > comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
> > compressed content. In the end git will be doing lots of magic and  
> > still perform
> > badly on unknown compressed content.
> > 
> > There is a very simple way of detecting compressed files: just look  
> > at the
> > size of the compressed blob and compare against the size of the  
> > expanded blob.
> > If the compressed blob has a non-trivial size which is close to the  
> > expanded
> > size, assume the file is not interesting as source or target for deltas.
> > 
> > Example:
> >    if (compressed_size > expanded_size / 4 * 3 + 1024) {
> >      /* don't try to deltify if blob doesn't compress well */
> >      return ...;
> >    }
> 
> And yet I get good delta compression on a number of ZIP formatted files 
> which don't get good additional zlib compression (<3%). Doing the above 
> would cause those packfiles to explode to about 10x their current size.

A pity. Geert's proposition sounded good to me.

However, there's got to be a way to cut short the search for a delta 
base/deltification when a certain (maybe even configurable) amount of time 
has been spent on it.

Ciao,
Dscho

```

## Shawn Pearce, 2006-12-14 23:29

Subject: Re: git-fetching from a big repository is slow
Message-ID: <20061214232936.GH26202@spearce.org>
URL: https://gitlist.dev/e/20061214232936.GH26202%40spearce.org
In-Reply-To: <Pine.LNX.4.63.0612150013390.3635@wbgn013.biozentrum.uni-wuerzburg.de>

```
Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
> On Thu, 14 Dec 2006, Shawn Pearce wrote:
> > Geert Bosch <bosch@adacore.com> wrote:
> > >    if (compressed_size > expanded_size / 4 * 3 + 1024) {
> > >      /* don't try to deltify if blob doesn't compress well */
> > >      return ...;
> > >    }
> > 
> > And yet I get good delta compression on a number of ZIP formatted files 
> > which don't get good additional zlib compression (<3%). Doing the above 
> > would cause those packfiles to explode to about 10x their current size.
> 
> A pity. Geert's proposition sounded good to me.
> 
> However, there's got to be a way to cut short the search for a delta 
> base/deltification when a certain (maybe even configurable) amount of time 
> has been spent on it.

I'm not sure time is the best rule there.

Maybe if the object is large (e.g. over 512 KiB or some configured
limit) and did not compress well when we last deflated it
(e.g. Geert's rule above) then only try to delta it against another
object whose hinted filename is very close/exactly matches and
whose size is very close, and don't make nearly as many attempts
on the matching hunks within any two files if the file appears to
be binary and not text.

I'm OK with a small increase in packfile size as a result of slightly
less optimal delta base selection on the really large binary files
due to something like the above, but 10x is insane.

-- 

```

## Johannes Schindelin, 2006-12-15 00:07

Subject: Re: git-fetching from a big repository is slow
Message-ID: <Pine.LNX.4.63.0612150105450.3635@wbgn013.biozentrum.uni-wuerzburg.de>
URL: https://gitlist.dev/e/Pine.LNX.4.63.0612150105450.3635%40wbgn013.biozentrum.uni-wuerzburg.de
In-Reply-To: <20061214232936.GH26202@spearce.org>

```
Hi,

On Thu, 14 Dec 2006, Shawn Pearce wrote:

> I'm OK with a small increase in packfile size as a result of slightly 
> less optimal delta base selection on the really large binary files due 
> to something like the above, but 10x is insane.

Not if it is a server having to do all the work. Along with all the work 
for all other clients. When you do a fetch, you really should be nice to 
the serving side.

Ciao,
Dscho

```

## Shawn Pearce, 2006-12-15 00:42

Subject: Re: git-fetching from a big repository is slow
Message-ID: <20061215004225.GJ26202@spearce.org>
URL: https://gitlist.dev/e/20061215004225.GJ26202%40spearce.org
In-Reply-To: <Pine.LNX.4.63.0612150105450.3635@wbgn013.biozentrum.uni-wuerzburg.de>

```
Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
> On Thu, 14 Dec 2006, Shawn Pearce wrote:
> 
> > I'm OK with a small increase in packfile size as a result of slightly 
> > less optimal delta base selection on the really large binary files due 
> > to something like the above, but 10x is insane.
> 
> Not if it is a server having to do all the work. Along with all the work 
> for all other clients. When you do a fetch, you really should be nice to 
> the serving side.

Yes, that's true.

But I fail to see what that has to do with the part you quoted above.
A 1% increase in transfer bandwidth may be better for a server if
it halves the CPU usage or disk IO usage if the server has more
bandwidth than those available; likewise a 1% decrease in transfer
bandwidth may be better for a server if it has lots of CPU to spare
but very little network bandwidth available.

Since every server is different its not like we can tune for just
one of those cases and cross our fingers.

-- 

```

## Nicolas Pitre, 2006-12-15 02:26

Subject: Re: git-fetching from a big repository is slow
Message-ID: <Pine.LNX.4.64.0612142125460.18171@xanadu.home>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0612142125460.18171%40xanadu.home
In-Reply-To: <Pine.LNX.4.63.0612150013390.3635@wbgn013.biozentrum.uni-wuerzburg.de>

```
On Fri, 15 Dec 2006, Johannes Schindelin wrote:

> Hi,
> 
> On Thu, 14 Dec 2006, Shawn Pearce wrote:
> 
> > Geert Bosch <bosch@adacore.com> wrote:
> > > Such special magic based on filenames is always a bad idea. Tomorrow  
> > > somebody
> > > comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
> > > compressed content. In the end git will be doing lots of magic and  
> > > still perform
> > > badly on unknown compressed content.
> > > 
> > > There is a very simple way of detecting compressed files: just look  
> > > at the
> > > size of the compressed blob and compare against the size of the  
> > > expanded blob.
> > > If the compressed blob has a non-trivial size which is close to the  
> > > expanded
> > > size, assume the file is not interesting as source or target for deltas.
> > > 
> > > Example:
> > >    if (compressed_size > expanded_size / 4 * 3 + 1024) {
> > >      /* don't try to deltify if blob doesn't compress well */
> > >      return ...;
> > >    }
> > 
> > And yet I get good delta compression on a number of ZIP formatted files 
> > which don't get good additional zlib compression (<3%). Doing the above 
> > would cause those packfiles to explode to about 10x their current size.
> 
> A pity. Geert's proposition sounded good to me.
> 
> However, there's got to be a way to cut short the search for a delta 
> base/deltification when a certain (maybe even configurable) amount of time 
> has been spent on it.

Yes! Run git-repack -a -d on the remote repository.



```

## Pazu, 2006-12-15 21:49

Subject: Re: git-fetching from a big repository is slow
Message-ID: <loom.20061215T223909-156@post.gmane.org>
URL: https://gitlist.dev/e/loom.20061215T223909-156%40post.gmane.org
In-Reply-To: <20061214223813.GC26202@spearce.org>

```
Shawn Pearce <spearce <at> spearce.org> writes:

> identical class file given the same input.  I've seen times where
> it doesn't thanks to the automatic serialVersionUID field being
> somewhat randomly generated.

Probably offline, but… serialVersionUID isn't randomly generated. It's
calculated using the types of fields in the class, recursively. The actual
algorithm is quite arbitrary, but not random. The automatically generated
serialVersionUID should change only if you add/remove class fields (either on
the class itself, or to the class of nested objects).

*sigh* Java chases me. 8+ hours of java work everyday, and when I finally get
home… there it is, looking at me again. *sob*

-- Pazu

```

## Robin Rosenberg, 2006-12-16 13:32

Subject: Re: git-fetching from a big repository is slow
Message-ID: <200612161433.00030.robin.rosenberg.lists@dewire.com>
URL: https://gitlist.dev/e/200612161433.00030.robin.rosenberg.lists%40dewire.com
In-Reply-To: <loom.20061215T223909-156@post.gmane.org>

```
fredag 15 december 2006 22:49 skrev Pazu:
> Shawn Pearce <spearce <at> spearce.org> writes:
> > identical class file given the same input.  I've seen times where
> > it doesn't thanks to the automatic serialVersionUID field being
> > somewhat randomly generated.
>
> Probably offline, but… serialVersionUID isn't randomly generated. It's
> calculated using the types of fields in the class, recursively. The actual
> algorithm is quite arbitrary, but not random. The automatically generated
> serialVersionUID should change only if you add/remove class fields (either
> on the class itself, or to the class of nested objects).

Different java compilers (e.g. SUN's javac and Eclipse) generate slipghtly 
different code for some cases, including somee synthetic member fields. that 
get involved in the UID calculation. Neither compiler is wrong. The java 
specifications don't cover all cases.


```
