git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] don't use mmap() to hash files

From
Avery Pennarun <apenwarr@gmail.com>
Date
Feb 15, 2010, 05:01 UTC
Message-ID
<32541b131002142101i226663cfk90d1ba14f1031788@mail.gmail.com>
In-Reply-To
<alpine.LFD.2.00.1002142252020.1946@xanadu.home>
On Sun, Feb 14, 2010 at 11:16 PM, Nicolas Pitre <nico@fluxnic.net> wrote:
Show 8 quoted lines
> On Sun, 14 Feb 2010, Avery Pennarun wrote:
>> In fact, arguably you should prevent git-add from adding large files
>> at all, because at least then you don't get the repository into a
>> hard-to-recover-from state with huge files.  (This happened at work a
>> few months ago; most people have no idea what to do in such a
>> situation.)
>
> Git needs to be fixed in that case, not be crippled.

That would be ideal, but is more work than disabling imports for large files by default (for example), which would be easy. In any case, my solution at work was to say "if it hurts, don't do that" and it seems to have worked out okay for now.

Show 8 quoted lines
>> For my own situation, I think I'm more likely to (and I know people
>> who are more likely to) try storing huge files in git than I am likely
>> to modify a file *while* I'm trying to store it in git.
>
> And fancy operations on huge files are pretty unlikely.  Blame, diff,
> etc, are suited for text file which are by nature relatively small.
> And if your source code is all pasted in one single huge file that Git
> can't handle right now, then the compiler is unlikely to cope either.

Well, I'm thinking of things like textual database dumps, such as those produced by mysqldump. It would be nice to be able to diff those efficiently, even if they're several gigs in size. bup's hierarchical chunking allows this.

Show 9 quoted lines
>> > The other way to handle huge files is to split them into chunks.
>> > http://article.gmane.org/gmane.comp.version-control.git/120112
>
> No.  The chunk idea doesn't fit the Git model well enough without many
> corner cases all over the place which is a major drawback.  I think that
> was discussed in that thread already.
>
>> I have a bit of experience splitting files into chunks:
>> http://groups.google.com/group/bup-list/browse_thread/thread/812031efd4c5f7e4

Note that bup's rolling-checksum-based hierarchical chunking is not the same as the chunking that was discussed in that thread, and it resolves most of the problems. Unless I'm missing something.

Also note that bup just uses normal tree objects (for better or worse) instead of introducing a new object type.

Show 7 quoted lines
>> It works.  Also note that the speed gain from mmap'ing packs appears
>> to be much less than the gain from mmap'ing indexes.  You could
>> probably sacrifice most or all of the former and never really notice.
>> Caching expanded deltas can be pretty valuable, though.  (bup
>> presently avoids that whole question by not using deltas.)
>
> We do have a cache of expanded deltas already.

Yes, sorry to have implied otherwise. I was just comparing the performance advantage of the delta expansion cache (which should be a lot) with that of mmaping packfiles (which probably isn't much since the packfile data is typically needed in expanded form anyway).

Show 11 quoted lines
>> I can also confirm that streaming objects directly into packs is a
>> massive performance increase when dealing with big files.  However,
>> you then start to run into git's heuristics that often assume (for
>> example) that if an object is in a pack, it should never (or rarely)
>> be pruned.  This is normally a fine assumption, because if it was
>> likely to get pruned, it probably never would have been put into a
>> pack in the first place.
>
> Would you please for my own sanity tell me where we do such thing.  I
> thought I had a firm grip on the pack model but you're casting a shadow
> of doubts on some code I might have written myself.

Sorry, I didn't hunt down the code, but I ran into it while experimenting before. The rules are something like:

- git-prune only prunes unpacked objects
- git-repack claims to be willing to explode unreachable objects back
into loose objects with -A, but I'm not quite sure if its definition
of "unreachable" is the same as mine.  And I'm not sure rewriting a
pack with -A makes the old pack reliably unreachable according to -d.
It's possible I was just being dense.
- there seems to be no documented situation in which you can ever
delete unused objects from a pack without using repack -a or -A, which
can be amazingly slow if your packs are huge.  (Ideally you'd only
repack the particular packs that you want to shrink.)  For example, my
bup repo is currently 200 GB.

Anyway, I didn't have much luck when playing with it earlier, but didn't investigate since I assumed it's just a workflow that nobody much cares about. Which I think is a reasonable position for git developers to take anyway.

Have fun,
Avery
Previous: Nicolas PitreNext: Nicolas Pitre
Message 33 of 84 in “Re: Bug#569505: git-core: 'git add' corrupts repository if the working directory is modified as it runs”
  1. Jonathan NiederFeb 12, 2010
  2. Zygo BlaxellFeb 12, 2010
  3. Jonathan NiederFeb 13, 2010
  4. Ilari LiusvaaraFeb 13, 2010
  5. Thomas RastFeb 13, 2010
  6. Ilari LiusvaaraFeb 13, 2010
  7. Dmitry PotapovFeb 13, 2010
  8. Zygo BlaxellFeb 13, 2010
  9. don't use mmap() to hash filesDmitry Potapov, Feb 14, 2010
  10. Junio C HamanoFeb 14, 2010
  11. Dmitry PotapovFeb 14, 2010
  12. Junio C HamanoFeb 14, 2010
  13. Thomas RastFeb 14, 2010
  14. Junio C HamanoFeb 14, 2010
  15. Johannes SchindelinFeb 14, 2010
  16. Junio C HamanoFeb 14, 2010
  17. Dmitry PotapovFeb 14, 2010
  18. Jakub NarebskiFeb 14, 2010
  19. Paolo BonziniFeb 14, 2010
  20. Johannes SchindelinFeb 14, 2010
  21. Dmitry PotapovFeb 14, 2010
  22. Johannes SchindelinFeb 14, 2010
  23. Johannes SchindelinFeb 14, 2010
  24. Dmitry PotapovFeb 14, 2010
  25. Zygo BlaxellFeb 14, 2010
  26. Nicolas PitreFeb 15, 2010
  27. Dmitry PotapovFeb 15, 2010
  28. Paolo BonziniFeb 15, 2010
  29. Dmitry PotapovFeb 15, 2010
  30. Dmitry PotapovFeb 14, 2010
  31. Avery PennarunFeb 14, 2010
  32. Nicolas PitreFeb 15, 2010
  33. Avery PennarunFeb 15, 2010
  34. Nicolas PitreFeb 15, 2010
  35. Avery PennarunFeb 15, 2010
  36. Nicolas PitreFeb 15, 2010
  37. don't use mmap() to hash filesDmitry Potapov, Feb 14, 2010
  38. Teach "git add" and friends to be paranoidJunio C Hamano, Feb 18, 2010
  39. Junio C HamanoFeb 18, 2010
  40. Zygo BlaxellFeb 18, 2010
  41. Junio C HamanoFeb 19, 2010
  42. Jeff KingFeb 18, 2010
  43. Nicolas PitreFeb 18, 2010
  44. Junio C HamanoFeb 18, 2010
  45. Wincent ColaiutaFeb 18, 2010
  46. Zygo BlaxellFeb 18, 2010
  47. Jonathan NiederFeb 18, 2010
  48. Junio C HamanoFeb 18, 2010
  49. Paolo BonziniFeb 22, 2010
  50. Dmitry PotapovFeb 22, 2010
  51. Thomas RastFeb 18, 2010
  52. Junio C HamanoFeb 18, 2010
  53. Nicolas PitreFeb 18, 2010
  54. 16 gig, 350,000 file repositoryBill Lear, Feb 18, 2010
  55. Nicolas PitreFeb 18, 2010
  56. Erik Faye-LundFeb 19, 2010
  57. Bill LearFeb 22, 2010
  58. Nicolas PitreFeb 22, 2010
  59. Peter HarrisFeb 18, 2010
  60. Junio C HamanoFeb 18, 2010
  61. Nicolas PitreFeb 18, 2010
  62. Jonathan NiederFeb 19, 2010
  63. Zygo BlaxellFeb 19, 2010
  64. Junio C HamanoFeb 19, 2010
  65. Zygo BlaxellFeb 19, 2010
  66. Dmitry PotapovFeb 19, 2010
  67. Junio C HamanoFeb 19, 2010
  68. Junio C HamanoFeb 20, 2010
  69. Dmitry PotapovFeb 21, 2010
  70. Junio C HamanoFeb 21, 2010
  71. Dmitry PotapovFeb 22, 2010
  72. Junio C HamanoFeb 22, 2010
  73. Dmitry PotapovFeb 22, 2010
  74. Nicolas PitreFeb 22, 2010
  75. Dmitry PotapovFeb 22, 2010
  76. Zygo BlaxellFeb 22, 2010
  77. Nicolas PitreFeb 22, 2010
  78. Junio C HamanoFeb 22, 2010
  79. Nicolas PitreFeb 22, 2010
  80. Dmitry PotapovFeb 22, 2010
  81. Nicolas PitreFeb 22, 2010
  82. mmap with MAP_PRIVATE is useless (was Re: Bug#569505: git-core: 'git add' corrupts repository if the working directory is modified as it runs)Paolo Bonzini, Feb 14, 2010
  83. Junio C HamanoFeb 14, 2010
  84. Paolo BonziniFeb 14, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.