git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: GSoC - Some questions on the idea of

From
NKNeal Kreitzinger <nkreitzinger@gmail.com>
Date
Apr 11, 2012, 01:24 UTC
Message-ID
<4F84DD60.20903@gmail.com>
In-Reply-To
<20120402210708.GA28926@sigill.intra.peff.net>
On 4/2/2012 4:07 PM, Jeff King wrote:
Show 6 quoted lines
> ...I think we need to first find out exactly
> how well the generic algorithm can perform. It may be "good enough"
> compared to the hassle that inconsistent application of a content-aware
> algorithm will cause.  So I wouldn't rule it out, but I'd rather try the
> bup-style splitting first, and see how good (or bad) it is.
>

(I read bup DESIGN doc to see what bup-style splitting is.) When you use bup delta technology in git.git I take it that you will use it for big-worktree-files *and* big-history-files (not-big-worktree-files that are not xdelta delta-friendly)? IOW, all binaries plus big-text-worktree-files. Otherwise, small binaries will become large histories.

If small binaries are not going to be bup-delta-compressed, then what about using xxd to convert the binary to text and then xdelta compressing the hex dump to achieve efficient delta compression in the pack file? You could convert the hexdump back to binary with xxd for checkout and such.

Maybe small binaries do xdelta well and the above is a moot point. This is all theory to me, but the reality is looming over my head since most of the components I should be tracking are binaries small (large history?) and big (but am not yet because of "big-file" concerns -- I don't want to have to refactor my vast git ecosystem with filter branch later because I slammed binaries into the main project or superproject without proper systems programming (I'm not sure what the c/linux term is for 'systems programming', but in the mainframe world it meant making sure everything was configured for efficient performance)).

Now that I say that out loud I guess a superproject with binaries in separate repos could be easily refactored by creating new efficient repos and making a new commit that points to them instead of the old inefficient repos. That way, when someone checks out the binary repo (submodule) into their worktree they get the new efficiency instead of the old inefficiency. Over time, as folks are less likely to check out old stuff the old inefficiency goes away on its own. I think. (Submodules are mostly theory to me at this point also.)

v/r, neal

Previous: Sergio CallegariNext: Jonathan Nieder
Message 12 of 43 in “GSoC - Some questions on the idea of "Better big-file support".”
  1. Bo ChenMar 28, 2012
  2. Nguyen Thai Ngoc DuyMar 28, 2012
  3. SergioMar 28, 2012
  4. Bo ChenMar 30, 2012
  5. Bo ChenMar 30, 2012
  6. Jeff KingMar 30, 2012
  7. Bo ChenMar 30, 2012
  8. Sergio CallegariMar 31, 2012
  9. Neal KreitzingerMar 31, 2012
  10. Jeff KingApr 2, 2012
  11. Sergio CallegariApr 3, 2012
  12. Neal KreitzingerApr 11, 2012
  13. Jonathan NiederApr 11, 2012
  14. Neal KreitzingerApr 11, 2012
  15. Jeff KingApr 11, 2012
  16. Neal KreitzingerApr 11, 2012
  17. Neal KreitzingerApr 11, 2012
  18. Jonathan NiederApr 11, 2012
  19. Junio C HamanoApr 11, 2012
  20. Jonathan NiederApr 11, 2012
  21. Neal KreitzingerApr 11, 2012
  22. Jeff KingApr 11, 2012
  23. Neal KreitzingerApr 12, 2012
  24. Jeff KingApr 12, 2012
  25. Neal KreitzingerApr 12, 2012
  26. Bo ChenApr 13, 2012
  27. Neal KreitzingerMar 31, 2012
  28. Jeff KingApr 2, 2012
  29. Junio C HamanoApr 2, 2012
  30. Jeff KingApr 3, 2012
  31. Neal KreitzingerMar 31, 2012
  32. Neal KreitzingerMar 31, 2012
  33. Bo ChenMar 31, 2012
  34. Nguyen Thai Ngoc DuyApr 1, 2012
  35. Bo ChenApr 1, 2012
  36. Nguyen Thai Ngoc DuyApr 2, 2012
  37. Bo ChenMar 30, 2012
  38. Jeff KingMar 30, 2012
  39. Jeff KingApr 15, 2012
  40. Neal KreitzingerApr 15, 2012
  41. Jeff KingApr 16, 2012
  42. Neal KreitzingerMay 10, 2012
  43. Jeff KingMay 10, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.