{"thread":{"id":"43382","subject":"Re: git-fetching from a big repository is slow","startedAt":"2006-12-14T19:46:36Z","lastAt":"2006-12-16T13:32:59Z","messageCount":11,"participants":["Shawn Pearce","Nicolas Pitre","Horst H. von Brand","Geert Bosch","Johannes Schindelin","Pazu","Robin Rosenberg"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"297685","messageId":"20061214194636.GO1747@spearce.org","threadId":"43382","inReplyTo":"C287764F-6755-4291-A87A-3E8816E90B49@adacore.com","subject":"Re: git-fetching from a big repository is slow","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-14T19:46:36Z","receivedAt":"2006-12-14T19:46:36Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Geert Bosch <bosch@adacore.com> wrote:\n> Such special magic based on filenames is always a bad idea. Tomorrow  \n> somebody\n> comes with .zip files (oh, and of course .ZIP), then it's .jpg's other\n> compressed content. In the end git will be doing lots of magic and  \n> still perform\n> badly on unknown compressed content.\n> \n> There is a very simple way of detecting compressed files: just look  \n> at the\n> size of the compressed blob and compare against the size of the  \n> expanded blob.\n> If the compressed blob has a non-trivial size which is close to the  \n> expanded\n> size, assume the file is not interesting as source or target for deltas.\n> \n> Example:\n>    if (compressed_size > expanded_size / 4 * 3 + 1024) {\n>      /* don't try to deltify if blob doesn't compress well */\n>      return ...;\n>    }\n\nAnd yet I get good delta compression on a number of ZIP formatted\nfiles which don't get good additional zlib compression (<3%).\nDoing the above would cause those packfiles to explode to about\n10x their current size.\n\n-- \n"},{"id":"295134","messageId":"200612142212.kBEMCVeu032626@laptop13.inf.utfsm.cl","threadId":"43382","inReplyTo":"spearce@spearce.org","subject":"Re: git-fetching from a big repository is slow","fromName":"Horst H. von Brand","fromEmail":"vonbrand@inf.utfsm.cl","sentAt":"2006-12-14T22:12:31Z","receivedAt":"2006-12-14T22:12:31Z","isPatch":false,"sender":{"key":"vonbrand@inf.utfsm.cl","avatar":"https://avatars.githubusercontent.com/u/211384?v=4"},"body":"Shawn Pearce <spearce@spearce.org> wrote:\n\n[...]\n\n> And yet I get good delta compression on a number of ZIP formatted\n> files which don't get good additional zlib compression (<3%).\n\n.zip is something like a tar of the compressed files, if the files inside\nthe archive don't change, the deltas will be small.\n-- \nDr. Horst H. von Brand                   User #22616 counter.li.org\nDepartamento de Informatica                    Fono: +56 32 2654431\nUniversidad Tecnica Federico Santa Maria             +56 32 2654239\n"},{"id":"294569","messageId":"20061214223813.GC26202@spearce.org","threadId":"43382","inReplyTo":"200612142212.kBEMCVeu032626@laptop13.inf.utfsm.cl","subject":"Re: git-fetching from a big repository is slow","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-14T22:38:13Z","receivedAt":"2006-12-14T22:38:13Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Horst H. von Brand\" <vonbrand@inf.utfsm.cl> wrote:\n> Shawn Pearce <spearce@spearce.org> wrote:\n> \n> [...]\n> \n> > And yet I get good delta compression on a number of ZIP formatted\n> > files which don't get good additional zlib compression (<3%).\n> \n> .zip is something like a tar of the compressed files, if the files inside\n> the archive don't change, the deltas will be small.\n\nYes, especially when the new zip is made using the exact same\nsoftware with the same parameters, so the resulting compressed file\nstream is identical for files whose content has not changed.  :-)\n\nSince this is actually a JAR full of Java classes which have\nbeen recompiled, its even more interesting that javac produced an\nidentical class file given the same input.  I've seen times where\nit doesn't thanks to the automatic serialVersionUID field being\nsomewhat randomly generated.\n\n-- \n"},{"id":"295533","messageId":"E30DCF6F-5D3E-4CA3-85D7-CD2847B86F86@adacore.com","threadId":"43382","inReplyTo":"20061214194636.GO1747@spearce.org","subject":"Re: git-fetching from a big repository is slow","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2006-12-14T23:01:46Z","receivedAt":"2006-12-14T23:01:46Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"\nOn Dec 14, 2006, at 14:46, Shawn Pearce wrote:\n> And yet I get good delta compression on a number of ZIP formatted\n> files which don't get good additional zlib compression (<3%).\n> Doing the above would cause those packfiles to explode to about\n> 10x their current size.\n\nYes, that's because for zip files each file in the archive is\ncompressed independently. Similar things might happen when\nchecking in uncompressed tar files with JPG's. The question\nis whether you prefer bad time usage or bad space usage when\nhandling large binary blobs. Maybe we should use a faster,\nless precise algorithm instead of giving up.\n\nStill, I think doing anything based on filename is a mistake.\nIf we want to have a heuristic to prevent spending too much time\non deltifying large compressed files, the heuristic should be\nbased on content, not filename.\n\nMaybe we could some \"magic\" as used by the file(1) command\nthat allows git to say a bit more about the content of blobs.\nThis could be used both for ordering files during deltification\nand to determine wether to try deltification at all.\n\n   -Geert\n\n"},{"id":"296151","messageId":"Pine.LNX.4.63.0612150013390.3635@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43382","inReplyTo":"20061214194636.GO1747@spearce.org","subject":"Re: git-fetching from a big repository is slow","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-14T23:15:25Z","receivedAt":"2006-12-14T23:15:25Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 14 Dec 2006, Shawn Pearce wrote:\n\n> Geert Bosch <bosch@adacore.com> wrote:\n> > Such special magic based on filenames is always a bad idea. Tomorrow  \n> > somebody\n> > comes with .zip files (oh, and of course .ZIP), then it's .jpg's other\n> > compressed content. In the end git will be doing lots of magic and  \n> > still perform\n> > badly on unknown compressed content.\n> > \n> > There is a very simple way of detecting compressed files: just look  \n> > at the\n> > size of the compressed blob and compare against the size of the  \n> > expanded blob.\n> > If the compressed blob has a non-trivial size which is close to the  \n> > expanded\n> > size, assume the file is not interesting as source or target for deltas.\n> > \n> > Example:\n> >    if (compressed_size > expanded_size / 4 * 3 + 1024) {\n> >      /* don't try to deltify if blob doesn't compress well */\n> >      return ...;\n> >    }\n> \n> And yet I get good delta compression on a number of ZIP formatted files \n> which don't get good additional zlib compression (<3%). Doing the above \n> would cause those packfiles to explode to about 10x their current size.\n\nA pity. Geert's proposition sounded good to me.\n\nHowever, there's got to be a way to cut short the search for a delta \nbase/deltification when a certain (maybe even configurable) amount of time \nhas been spent on it.\n\nCiao,\nDscho\n"},{"id":"296637","messageId":"20061214232936.GH26202@spearce.org","threadId":"43382","inReplyTo":"Pine.LNX.4.63.0612150013390.3635@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git-fetching from a big repository is slow","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-14T23:29:36Z","receivedAt":"2006-12-14T23:29:36Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Thu, 14 Dec 2006, Shawn Pearce wrote:\n> > Geert Bosch <bosch@adacore.com> wrote:\n> > >    if (compressed_size > expanded_size / 4 * 3 + 1024) {\n> > >      /* don't try to deltify if blob doesn't compress well */\n> > >      return ...;\n> > >    }\n> > \n> > And yet I get good delta compression on a number of ZIP formatted files \n> > which don't get good additional zlib compression (<3%). Doing the above \n> > would cause those packfiles to explode to about 10x their current size.\n> \n> A pity. Geert's proposition sounded good to me.\n> \n> However, there's got to be a way to cut short the search for a delta \n> base/deltification when a certain (maybe even configurable) amount of time \n> has been spent on it.\n\nI'm not sure time is the best rule there.\n\nMaybe if the object is large (e.g. over 512 KiB or some configured\nlimit) and did not compress well when we last deflated it\n(e.g. Geert's rule above) then only try to delta it against another\nobject whose hinted filename is very close/exactly matches and\nwhose size is very close, and don't make nearly as many attempts\non the matching hunks within any two files if the file appears to\nbe binary and not text.\n\nI'm OK with a small increase in packfile size as a result of slightly\nless optimal delta base selection on the really large binary files\ndue to something like the above, but 10x is insane.\n\n-- \n"},{"id":"297461","messageId":"Pine.LNX.4.63.0612150105450.3635@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43382","inReplyTo":"20061214232936.GH26202@spearce.org","subject":"Re: git-fetching from a big repository is slow","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-15T00:07:24Z","receivedAt":"2006-12-15T00:07:24Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 14 Dec 2006, Shawn Pearce wrote:\n\n> I'm OK with a small increase in packfile size as a result of slightly \n> less optimal delta base selection on the really large binary files due \n> to something like the above, but 10x is insane.\n\nNot if it is a server having to do all the work. Along with all the work \nfor all other clients. When you do a fetch, you really should be nice to \nthe serving side.\n\nCiao,\nDscho\n"},{"id":"296516","messageId":"20061215004225.GJ26202@spearce.org","threadId":"43382","inReplyTo":"Pine.LNX.4.63.0612150105450.3635@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git-fetching from a big repository is slow","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-15T00:42:25Z","receivedAt":"2006-12-15T00:42:25Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Thu, 14 Dec 2006, Shawn Pearce wrote:\n> \n> > I'm OK with a small increase in packfile size as a result of slightly \n> > less optimal delta base selection on the really large binary files due \n> > to something like the above, but 10x is insane.\n> \n> Not if it is a server having to do all the work. Along with all the work \n> for all other clients. When you do a fetch, you really should be nice to \n> the serving side.\n\nYes, that's true.\n\nBut I fail to see what that has to do with the part you quoted above.\nA 1% increase in transfer bandwidth may be better for a server if\nit halves the CPU usage or disk IO usage if the server has more\nbandwidth than those available; likewise a 1% decrease in transfer\nbandwidth may be better for a server if it has lots of CPU to spare\nbut very little network bandwidth available.\n\nSince every server is different its not like we can tune for just\none of those cases and cross our fingers.\n\n-- \n"},{"id":"295081","messageId":"Pine.LNX.4.64.0612142125460.18171@xanadu.home","threadId":"43382","inReplyTo":"Pine.LNX.4.63.0612150013390.3635@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git-fetching from a big repository is slow","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-12-15T02:26:38Z","receivedAt":"2006-12-15T02:26:38Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 15 Dec 2006, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Thu, 14 Dec 2006, Shawn Pearce wrote:\n> \n> > Geert Bosch <bosch@adacore.com> wrote:\n> > > Such special magic based on filenames is always a bad idea. Tomorrow  \n> > > somebody\n> > > comes with .zip files (oh, and of course .ZIP), then it's .jpg's other\n> > > compressed content. In the end git will be doing lots of magic and  \n> > > still perform\n> > > badly on unknown compressed content.\n> > > \n> > > There is a very simple way of detecting compressed files: just look  \n> > > at the\n> > > size of the compressed blob and compare against the size of the  \n> > > expanded blob.\n> > > If the compressed blob has a non-trivial size which is close to the  \n> > > expanded\n> > > size, assume the file is not interesting as source or target for deltas.\n> > > \n> > > Example:\n> > >    if (compressed_size > expanded_size / 4 * 3 + 1024) {\n> > >      /* don't try to deltify if blob doesn't compress well */\n> > >      return ...;\n> > >    }\n> > \n> > And yet I get good delta compression on a number of ZIP formatted files \n> > which don't get good additional zlib compression (<3%). Doing the above \n> > would cause those packfiles to explode to about 10x their current size.\n> \n> A pity. Geert's proposition sounded good to me.\n> \n> However, there's got to be a way to cut short the search for a delta \n> base/deltification when a certain (maybe even configurable) amount of time \n> has been spent on it.\n\nYes! Run git-repack -a -d on the remote repository.\n\n\n"},{"id":"297885","messageId":"loom.20061215T223909-156@post.gmane.org","threadId":"43382","inReplyTo":"20061214223813.GC26202@spearce.org","subject":"Re: git-fetching from a big repository is slow","fromName":"Pazu","fromEmail":"pazu@pazu.com.br","sentAt":"2006-12-15T21:49:13Z","receivedAt":"2006-12-15T21:49:13Z","isPatch":false,"sender":{"key":"pazu@pazu.com.br","avatar":null},"body":"Shawn Pearce <spearce <at> spearce.org> writes:\n\n> identical class file given the same input.  I've seen times where\n> it doesn't thanks to the automatic serialVersionUID field being\n> somewhat randomly generated.\n\nProbably offline, but… serialVersionUID isn't randomly generated. It's\ncalculated using the types of fields in the class, recursively. The actual\nalgorithm is quite arbitrary, but not random. The automatically generated\nserialVersionUID should change only if you add/remove class fields (either on\nthe class itself, or to the class of nested objects).\n\n*sigh* Java chases me. 8+ hours of java work everyday, and when I finally get\nhome… there it is, looking at me again. *sob*\n\n-- Pazu\n"},{"id":"298631","messageId":"200612161433.00030.robin.rosenberg.lists@dewire.com","threadId":"43382","inReplyTo":"loom.20061215T223909-156@post.gmane.org","subject":"Re: git-fetching from a big repository is slow","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2006-12-16T13:32:59Z","receivedAt":"2006-12-16T13:32:59Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"fredag 15 december 2006 22:49 skrev Pazu:\n> Shawn Pearce <spearce <at> spearce.org> writes:\n> > identical class file given the same input.  I've seen times where\n> > it doesn't thanks to the automatic serialVersionUID field being\n> > somewhat randomly generated.\n>\n> Probably offline, but… serialVersionUID isn't randomly generated. It's\n> calculated using the types of fields in the class, recursively. The actual\n> algorithm is quite arbitrary, but not random. The automatically generated\n> serialVersionUID should change only if you add/remove class fields (either\n> on the class itself, or to the class of nested objects).\n\nDifferent java compilers (e.g. SUN's javac and Eclipse) generate slipghtly \ndifferent code for some cases, including somee synthetic member fields. that \nget involved in the UID calculation. Neither compiler is wrong. The java \nspecifications don't cover all cases.\n\n"}]}