git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: GSoC - Some questions on the idea of

From
BCBo Chen <chen@chenirvine.org>
Date
Mar 30, 2012, 19:44 UTC
Message-ID
<CA+M5ThT47twke7xkeAdDbk0c_J_=U6t1swDVexD6WDrQjG9_-w@mail.gmail.com>
In-Reply-To
<loom.20120328T131530-717@post.gmane.org>

The following is the list of sub-problems according to my understanding of the "big file support" problem. Can anyone give some feed back and help refine it. Thanks.

            ---- text file (always delta well? need to be confirmed)
             |
                                               --- delta well (ok)
large file-|                    ----    general binary file (without
encryption, compression. Other cases which definitely can not delta
well)  -|
             |                     |
                                               --- does not delta well
(improvement?)
            ---- binary file   -|---   encrypted file (improvement?
one straightforward method is to decrypt the file before delta-ing it,
however, we don't always have the key for decryption. Other?)
                                   |
                                  ---    compressed file (improvement?
Decompress before delta-ing it? Other?)
Bo
On Wed, Mar 28, 2012 at 7:33 AM, Sergio <sergio.callegari@gmail.com> wrote:
Show 51 quoted lines
> Nguyen Thai Ngoc Duy <pclouds <at> gmail.com> writes:
>
>>
>> On Wed, Mar 28, 2012 at 11:38 AM, Bo Chen <chen <at> chenirvine.org> wrote:
>> > Hi, Everyone. This is Bo Chen. I am interested in the idea of "Better
>> > big-file support".
>> >
>> > As it is described in the idea page,
>> > "Many large files (like media) do not delta very well. However, some
>> > do (like VM disk images). Git could split large objects into smaller
>> > chunks, similar to bup, and find deltas between these much more
>> > manageable chunks. There are some preliminary patches in this
>> > direction, but they are in need of review and expansion."
>> >
>> > Can anyone elaborate a little bit why many large files do not delta
>> > very well?
>>
>> Large files are usually binary. Depends on the type of binary, they
>> may or may not delta well. Those that are compressed/encrypted
>> obviously don't delta well because one change can make the final
>> result completely different.
>
> I would add that the larger a file, the larger the temptation to use a
> compressed format for it, so that large files are often compressed binaries.
>
> For these, a trick to obtain good deltas can be to decompress before splitting
> in chunks with the rsync algorithm. Git filters can already be used for this,
> but it can be tricky to assure that the decompress - recompress roundtrip
> re-creates the original compressed file.
>
> Furhermore, some compressed binaries are internally composed by multiple streams
> (think of a zip archive containing multiple files, but this is by no means
> limited to zip). In this case, it is frequent to have many possible orderings of
> the streams. If so, the best deltas can be obtained by sorting the streams in
> some 'canonical' order and decompressing. Even without decompressing, sorting
> alone can obtain good results as long as changes are only due to changes in a
> single stream of the container. Personally, I know no example of git filters
> used to perform this sorting which can be extremely tricky in assuring the
> possibility of recovering the file in the original stream order.
>
> Maybe (but this is just speculation), once the bup-inspired file chunking
> support is in place, people will start contributing filters to improve the
> management of many types of standard files (obviously 'improve' in terms of
> space efficiency as filters can be quite slow).
>
> Sergio
>
> --
> To unsubscribe from this list: send the line "unsubscribe git" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
Previous: SergioNext: Bo Chen
Message 4 of 43 in “GSoC - Some questions on the idea of "Better big-file support".”
  1. Bo ChenMar 28, 2012
  2. Nguyen Thai Ngoc DuyMar 28, 2012
  3. SergioMar 28, 2012
  4. Bo ChenMar 30, 2012
  5. Bo ChenMar 30, 2012
  6. Jeff KingMar 30, 2012
  7. Bo ChenMar 30, 2012
  8. Sergio CallegariMar 31, 2012
  9. Neal KreitzingerMar 31, 2012
  10. Jeff KingApr 2, 2012
  11. Sergio CallegariApr 3, 2012
  12. Neal KreitzingerApr 11, 2012
  13. Jonathan NiederApr 11, 2012
  14. Neal KreitzingerApr 11, 2012
  15. Jeff KingApr 11, 2012
  16. Neal KreitzingerApr 11, 2012
  17. Neal KreitzingerApr 11, 2012
  18. Jonathan NiederApr 11, 2012
  19. Junio C HamanoApr 11, 2012
  20. Jonathan NiederApr 11, 2012
  21. Neal KreitzingerApr 11, 2012
  22. Jeff KingApr 11, 2012
  23. Neal KreitzingerApr 12, 2012
  24. Jeff KingApr 12, 2012
  25. Neal KreitzingerApr 12, 2012
  26. Bo ChenApr 13, 2012
  27. Neal KreitzingerMar 31, 2012
  28. Jeff KingApr 2, 2012
  29. Junio C HamanoApr 2, 2012
  30. Jeff KingApr 3, 2012
  31. Neal KreitzingerMar 31, 2012
  32. Neal KreitzingerMar 31, 2012
  33. Bo ChenMar 31, 2012
  34. Nguyen Thai Ngoc DuyApr 1, 2012
  35. Bo ChenApr 1, 2012
  36. Nguyen Thai Ngoc DuyApr 2, 2012
  37. Bo ChenMar 30, 2012
  38. Jeff KingMar 30, 2012
  39. Jeff KingApr 15, 2012
  40. Neal KreitzingerApr 15, 2012
  41. Jeff KingApr 16, 2012
  42. Neal KreitzingerMay 10, 2012
  43. Jeff KingMay 10, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.