git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: GSoC - Some questions on the idea of

From
Jeff King <peff@peff.net>
Date
Apr 2, 2012, 21:07 UTC
Message-ID
<20120402210708.GA28926@sigill.intra.peff.net>
In-Reply-To
<4F772E48.3030708@gmail.com>
On Sat, Mar 31, 2012 at 11:18:16AM -0500, Neal Kreitzinger wrote:
Show 15 quoted lines
> On 3/31/2012 6:02 AM, Sergio Callegari wrote:
> >I wonder if it could make sense to have some pluggable mechanism for
> > file splitting. Something under the lines of filters, so to say.
> >Bupsplit can be a rather general mechanism, but large binaries that
> >are containers (zip, jar, docx, tgz, pdf - seen as a collection of
> >streams) may possibly be more conveniently split by their inherent
> >components.
> >
> 
> gitattributes or gitconfig could configure the big-file handler for
> specified files.  Known/supported filetypes like gif, png, zip, pdf,
> etc., could be auto-configured by git.  Any
> yet-unknown/yet-unsupported filetypes could be configured manually by
> the user, e.g.
> *.zgp=bigcontainer

This is a tempting route (and one I've even suggested myself before), but I think ultimately it is a bad way to go. The problem is that splitting is only half of the equation. Once you have split contents, you have to use them intelligently, which means looking at the sha1s of each split chunk and discarding whole chunks as "the same" without even looking at the contents.

Which means that it is very important that your chunking algorithm remain stable from version to version. A change in the algorithm is going to completely negate the benefits of chunking in the first place. So something configurable, or something that is not applied consistently (because it depends on each user's git config, or even on the specific version of a tool used) can end up being no help at all.

Properly applied, I think a content-aware chunking algorithm could out-perform a generic one. But I think we need to first find out exactly how well the generic algorithm can perform. It may be "good enough" compared to the hassle that inconsistent application of a content-aware algorithm will cause. So I wouldn't rule it out, but I'd rather try the bup-style splitting first, and see how good (or bad) it is.

-Peff
Previous: Neal KreitzingerNext: Sergio Callegari
Message 10 of 43 in “GSoC - Some questions on the idea of "Better big-file support".”
  1. Bo ChenMar 28, 2012
  2. Nguyen Thai Ngoc DuyMar 28, 2012
  3. SergioMar 28, 2012
  4. Bo ChenMar 30, 2012
  5. Bo ChenMar 30, 2012
  6. Jeff KingMar 30, 2012
  7. Bo ChenMar 30, 2012
  8. Sergio CallegariMar 31, 2012
  9. Neal KreitzingerMar 31, 2012
  10. Jeff KingApr 2, 2012
  11. Sergio CallegariApr 3, 2012
  12. Neal KreitzingerApr 11, 2012
  13. Jonathan NiederApr 11, 2012
  14. Neal KreitzingerApr 11, 2012
  15. Jeff KingApr 11, 2012
  16. Neal KreitzingerApr 11, 2012
  17. Neal KreitzingerApr 11, 2012
  18. Jonathan NiederApr 11, 2012
  19. Junio C HamanoApr 11, 2012
  20. Jonathan NiederApr 11, 2012
  21. Neal KreitzingerApr 11, 2012
  22. Jeff KingApr 11, 2012
  23. Neal KreitzingerApr 12, 2012
  24. Jeff KingApr 12, 2012
  25. Neal KreitzingerApr 12, 2012
  26. Bo ChenApr 13, 2012
  27. Neal KreitzingerMar 31, 2012
  28. Jeff KingApr 2, 2012
  29. Junio C HamanoApr 2, 2012
  30. Jeff KingApr 3, 2012
  31. Neal KreitzingerMar 31, 2012
  32. Neal KreitzingerMar 31, 2012
  33. Bo ChenMar 31, 2012
  34. Nguyen Thai Ngoc DuyApr 1, 2012
  35. Bo ChenApr 1, 2012
  36. Nguyen Thai Ngoc DuyApr 2, 2012
  37. Bo ChenMar 30, 2012
  38. Jeff KingMar 30, 2012
  39. Jeff KingApr 15, 2012
  40. Neal KreitzingerApr 15, 2012
  41. Jeff KingApr 16, 2012
  42. Neal KreitzingerMay 10, 2012
  43. Jeff KingMay 10, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.