threads / discuss / 43382

Re: git-fetching from a big repository is slow

Subject: Re: git-fetching from a big repository is slow

## tl;dr

11 messages between Dec 14, 2006 and Dec 16, 2006.

replies: 10people: 7as markdown or json

Shawn Pearce· Dec 14, 2006, 19:46 UTC · lore
Geert Bosch <bosch@adacore.com> wrote:
Show 20 quoted lines
> Such special magic based on filenames is always a bad idea. Tomorrow  
> somebody
> comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
> compressed content. In the end git will be doing lots of magic and  
> still perform
> badly on unknown compressed content.
> 
> There is a very simple way of detecting compressed files: just look  
> at the
> size of the compressed blob and compare against the size of the  
> expanded blob.
> If the compressed blob has a non-trivial size which is close to the  
> expanded
> size, assume the file is not interesting as source or target for deltas.
> 
> Example:
>    if (compressed_size > expanded_size / 4 * 3 + 1024) {
>      /* don't try to deltify if blob doesn't compress well */
>      return ...;
>    }

And yet I get good delta compression on a number of ZIP formatted files which don't get good additional zlib compression (<3%). Doing the above would cause those packfiles to explode to about 10x their current size.

Geert Bosch· Dec 14, 2006, 23:01 UTC · re: Shawn Pearce · lore
On Dec 14, 2006, at 14:46, Shawn Pearce wrote:
> And yet I get good delta compression on a number of ZIP formatted
> files which don't get good additional zlib compression (<3%).
> Doing the above would cause those packfiles to explode to about
> 10x their current size.

Yes, that's because for zip files each file in the archive is compressed independently. Similar things might happen when checking in uncompressed tar files with JPG's. The question is whether you prefer bad time usage or bad space usage when handling large binary blobs. Maybe we should use a faster, less precise algorithm instead of giving up.

Still, I think doing anything based on filename is a mistake. If we want to have a heuristic to prevent spending too much time on deltifying large compressed files, the heuristic should be based on content, not filename.

Maybe we could some "magic" as used by the file(1) command that allows git to say a bit more about the content of blobs. This could be used both for ordering files during deltification and to determine wether to try deltification at all.

   -Geert
Johannes Schindelin· Dec 14, 2006, 23:15 UTC · re: Shawn Pearce · lore
Hi,
On Thu, 14 Dec 2006, Shawn Pearce wrote:
Show 25 quoted lines
> Geert Bosch <bosch@adacore.com> wrote:
> > Such special magic based on filenames is always a bad idea. Tomorrow  
> > somebody
> > comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
> > compressed content. In the end git will be doing lots of magic and  
> > still perform
> > badly on unknown compressed content.
> > 
> > There is a very simple way of detecting compressed files: just look  
> > at the
> > size of the compressed blob and compare against the size of the  
> > expanded blob.
> > If the compressed blob has a non-trivial size which is close to the  
> > expanded
> > size, assume the file is not interesting as source or target for deltas.
> > 
> > Example:
> >    if (compressed_size > expanded_size / 4 * 3 + 1024) {
> >      /* don't try to deltify if blob doesn't compress well */
> >      return ...;
> >    }
> 
> And yet I get good delta compression on a number of ZIP formatted files 
> which don't get good additional zlib compression (<3%). Doing the above 
> would cause those packfiles to explode to about 10x their current size.
A pity. Geert's proposition sounded good to me.

However, there's got to be a way to cut short the search for a delta base/deltification when a certain (maybe even configurable) amount of time has been spent on it.

Ciao, Dscho

Shawn Pearce· Dec 14, 2006, 23:29 UTC · re: Johannes Schindelin · lore
Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
Show 16 quoted lines
> On Thu, 14 Dec 2006, Shawn Pearce wrote:
> > Geert Bosch <bosch@adacore.com> wrote:
> > >    if (compressed_size > expanded_size / 4 * 3 + 1024) {
> > >      /* don't try to deltify if blob doesn't compress well */
> > >      return ...;
> > >    }
> > 
> > And yet I get good delta compression on a number of ZIP formatted files 
> > which don't get good additional zlib compression (<3%). Doing the above 
> > would cause those packfiles to explode to about 10x their current size.
> 
> A pity. Geert's proposition sounded good to me.
> 
> However, there's got to be a way to cut short the search for a delta 
> base/deltification when a certain (maybe even configurable) amount of time 
> has been spent on it.
I'm not sure time is the best rule there.

Maybe if the object is large (e.g. over 512 KiB or some configured limit) and did not compress well when we last deflated it (e.g. Geert's rule above) then only try to delta it against another object whose hinted filename is very close/exactly matches and whose size is very close, and don't make nearly as many attempts on the matching hunks within any two files if the file appears to be binary and not text.

I'm OK with a small increase in packfile size as a result of slightly less optimal delta base selection on the really large binary files due to something like the above, but 10x is insane.

Johannes Schindelin· Dec 15, 2006, 00:07 UTC · re: Shawn Pearce · lore
Hi,
On Thu, 14 Dec 2006, Shawn Pearce wrote:
> I'm OK with a small increase in packfile size as a result of slightly 
> less optimal delta base selection on the really large binary files due 
> to something like the above, but 10x is insane.

Not if it is a server having to do all the work. Along with all the work for all other clients. When you do a fetch, you really should be nice to the serving side.

Ciao, Dscho

Shawn Pearce· Dec 15, 2006, 00:42 UTC · re: Johannes Schindelin · lore
Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
Show 9 quoted lines
> On Thu, 14 Dec 2006, Shawn Pearce wrote:
> 
> > I'm OK with a small increase in packfile size as a result of slightly 
> > less optimal delta base selection on the really large binary files due 
> > to something like the above, but 10x is insane.
> 
> Not if it is a server having to do all the work. Along with all the work 
> for all other clients. When you do a fetch, you really should be nice to 
> the serving side.
Yes, that's true.

But I fail to see what that has to do with the part you quoted above. A 1% increase in transfer bandwidth may be better for a server if it halves the CPU usage or disk IO usage if the server has more bandwidth than those available; likewise a 1% decrease in transfer bandwidth may be better for a server if it has lots of CPU to spare but very little network bandwidth available.

Since every server is different its not like we can tune for just one of those cases and cross our fingers.

Nicolas Pitre· Dec 15, 2006, 02:26 UTC · re: Johannes Schindelin · lore
On Fri, 15 Dec 2006, Johannes Schindelin wrote:
Show 35 quoted lines
> Hi,
> 
> On Thu, 14 Dec 2006, Shawn Pearce wrote:
> 
> > Geert Bosch <bosch@adacore.com> wrote:
> > > Such special magic based on filenames is always a bad idea. Tomorrow  
> > > somebody
> > > comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
> > > compressed content. In the end git will be doing lots of magic and  
> > > still perform
> > > badly on unknown compressed content.
> > > 
> > > There is a very simple way of detecting compressed files: just look  
> > > at the
> > > size of the compressed blob and compare against the size of the  
> > > expanded blob.
> > > If the compressed blob has a non-trivial size which is close to the  
> > > expanded
> > > size, assume the file is not interesting as source or target for deltas.
> > > 
> > > Example:
> > >    if (compressed_size > expanded_size / 4 * 3 + 1024) {
> > >      /* don't try to deltify if blob doesn't compress well */
> > >      return ...;
> > >    }
> > 
> > And yet I get good delta compression on a number of ZIP formatted files 
> > which don't get good additional zlib compression (<3%). Doing the above 
> > would cause those packfiles to explode to about 10x their current size.
> 
> A pity. Geert's proposition sounded good to me.
> 
> However, there's got to be a way to cut short the search for a delta 
> base/deltification when a certain (maybe even configurable) amount of time 
> has been spent on it.
Yes! Run git-repack -a -d on the remote repository.
Horst H. von Brand· Dec 14, 2006, 22:12 UTC · lore
Shawn Pearce <spearce@spearce.org> wrote:
[...]
> And yet I get good delta compression on a number of ZIP formatted
> files which don't get good additional zlib compression (<3%).

.zip is something like a tar of the compressed files, if the files inside the archive don't change, the deltas will be small.

-- 
Dr. Horst H. von Brand                   User #22616 counter.li.org
Departamento de Informatica                    Fono: +56 32 2654431
Universidad Tecnica Federico Santa Maria             +56 32 2654239
Shawn Pearce· Dec 14, 2006, 22:38 UTC · re: Horst H. von Brand · lore
"Horst H. von Brand" <vonbrand@inf.utfsm.cl> wrote:
Show 9 quoted lines
> Shawn Pearce <spearce@spearce.org> wrote:
> 
> [...]
> 
> > And yet I get good delta compression on a number of ZIP formatted
> > files which don't get good additional zlib compression (<3%).
> 
> .zip is something like a tar of the compressed files, if the files inside
> the archive don't change, the deltas will be small.

Yes, especially when the new zip is made using the exact same software with the same parameters, so the resulting compressed file stream is identical for files whose content has not changed. :-)

Since this is actually a JAR full of Java classes which have been recompiled, its even more interesting that javac produced an identical class file given the same input. I've seen times where it doesn't thanks to the automatic serialVersionUID field being somewhat randomly generated.

Pazu· Dec 15, 2006, 21:49 UTC · re: Shawn Pearce · lore
Shawn Pearce <spearce <at> spearce.org> writes:
> identical class file given the same input.  I've seen times where
> it doesn't thanks to the automatic serialVersionUID field being
> somewhat randomly generated.

Probably offline, but… serialVersionUID isn't randomly generated. It's calculated using the types of fields in the class, recursively. The actual algorithm is quite arbitrary, but not random. The automatically generated serialVersionUID should change only if you add/remove class fields (either on the class itself, or to the class of nested objects).

*sigh* Java chases me. 8+ hours of java work everyday, and when I finally get home… there it is, looking at me again. *sob*

-- Pazu
Robin Rosenberg· Dec 16, 2006, 13:32 UTC · re: Pazu · lore
fredag 15 december 2006 22:49 skrev Pazu:
Show 10 quoted lines
> Shawn Pearce <spearce <at> spearce.org> writes:
> > identical class file given the same input.  I've seen times where
> > it doesn't thanks to the automatic serialVersionUID field being
> > somewhat randomly generated.
>
> Probably offline, but… serialVersionUID isn't randomly generated. It's
> calculated using the types of fields in the class, recursively. The actual
> algorithm is quite arbitrary, but not random. The automatically generated
> serialVersionUID should change only if you add/remove class fields (either
> on the class itself, or to the class of nested objects).

Different java compilers (e.g. SUN's javac and Eclipse) generate slipghtly different code for some cases, including somee synthetic member fields. that get involved in the UID calculation. Neither compiler is wrong. The java specifications don't cover all cases.

← back to recent threads