git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] diff-delta: produce optimal pack data

From
Linus Torvalds <torvalds@osdl.org>
Date
Feb 24, 2006, 20:02 UTC
Message-ID
<Pine.LNX.4.64.0602241152290.22647@g5.osdl.org>
In-Reply-To
<20060224192354.GC387@hpsvcnb.fc.hp.com>
On Fri, 24 Feb 2006, Carl Baldwin wrote:
Show 5 quoted lines
> 
> I meant that there is no benefit in disk space usage.  Packing may
> actually increase my disk space usage in this case.  Refer to what I
> said about experimentally running gzip and xdelta on the files
> independantly of git.
Yes. The deltas tend to compress a lot less well than "normal" files.
Show 5 quoted lines
> I see what you're saying about this data reuse helping to speed up
> subsequent cloning operations.  However, if packing takes this long and
> doesn't give me any disk space savings then I don't want to pay the very
> heavy price of packing these files even the first time nor do I want to
> pay the price incrementally.

I would look at tuning the heuristics in "try_delta()" (pack-objects.c) a bit. That's the place that decides whether to even bother trying to make a delta, and how big a delta is acceptable. For example, looking at them, I already see one bug:

	..
        sizediff = oldsize > size ? oldsize - size : size - oldsize;
        if (sizediff > size / 8)
                return -1;
	..

we really should compare sizediff to the _smaller_ of the two sizes, and skip the delta if the difference in sizes is bound to be bigger than that.

However, the "size / 8" thing isn't a very strict limit anyway, so this probably doesn't matter (and I think Nico already removed it as part of his patches: the heuristic can make us avoid some deltas that would be ok).

The other thing to look at is "max_size": right now it initializes that to "size / 2 - 20", which just says that we don't ever want a delta that is larger than about half the result (plus the 20 byte overhead for pointing to the thing we delta against). Again, if you feel that normal compression compresses better than half, you could try changing that to

	..
	max_size = size / 4 - 20;
	..

or something like that instead (but then you need to check that it's still positive - otherwise the comparisons with unsigned later on are screwed up. Right now that value is guaranteed to be positive if only because we already checked

	..
	if (size < 50)
		return -1;
	..
earlier).

NOTE! Every SINGLE one of those heuristics are just totally made up by yours truly, and have no testing behind them. They're more of the type "that sounds about right" than "this is how it must be". As mentioned, Nico has already been playing with the heuristics - but he wanted better packs, not better CPU usage, so he went the other way from what you would want to try..

		Linus
Previous: Nicolas PitreNext: Nicolas Pitre
Message 15 of 35 in “diff-delta: produce optimal pack data”
  1. diff-delta: produce optimal pack dataNicolas Pitre, Feb 22, 2006
  2. Junio C HamanoFeb 24, 2006
  3. Nicolas PitreFeb 24, 2006
  4. Junio C HamanoFeb 24, 2006
  5. Carl BaldwinFeb 24, 2006
  6. Nicolas PitreFeb 24, 2006
  7. Carl BaldwinFeb 24, 2006
  8. Nicolas PitreFeb 24, 2006
  9. Carl BaldwinFeb 24, 2006
  10. Nicolas PitreFeb 24, 2006
  11. Carl BaldwinFeb 24, 2006
  12. Nicolas PitreFeb 24, 2006
  13. Carl BaldwinFeb 24, 2006
  14. Nicolas PitreFeb 25, 2006
  15. Linus TorvaldsFeb 24, 2006
  16. Nicolas PitreFeb 24, 2006
  17. Junio C HamanoFeb 24, 2006
  18. Nicolas PitreFeb 24, 2006
  19. Nicolas PitreFeb 24, 2006
  20. Linus TorvaldsFeb 25, 2006
  21. Nicolas PitreFeb 25, 2006
  22. Linus TorvaldsFeb 25, 2006
  23. Nicolas PitreFeb 25, 2006
  24. Nicolas PitreFeb 25, 2006
  25. [RFH] zlib gurus out there?Junio C Hamano, Mar 7, 2006
  26. Linus TorvaldsMar 8, 2006
  27. Junio C HamanoMar 8, 2006
  28. Linus TorvaldsMar 8, 2006
  29. Johannes SchindelinMar 8, 2006
  30. write_sha1_file(): Perform Z_FULL_FLUSH between header and dataSergey Vlasov, Mar 8, 2006
  31. Junio C HamanoMar 8, 2006
  32. Sergey VlasovMar 8, 2006
  33. Linus TorvaldsFeb 25, 2006
  34. Carl BaldwinFeb 24, 2006
  35. Nicolas PitreFeb 24, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.