git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: cloning the kernel - why long time in "Resolving 313037 deltas"

From
Shawn Pearce <spearce@spearce.org>
Date
Dec 20, 2006, 01:54 UTC
Message-ID
<20061220015431.GA27638@spearce.org>
In-Reply-To
<Pine.LNX.4.64.0612190855300.3479@woody.osdl.org>
Linus Torvalds <torvalds@osdl.org> wrote:
Show 15 quoted lines
> On Tue, 19 Dec 2006, Theodore Tso wrote:
> > 
> > So the main reason to use mamp, as Linus puts it, is if the management
> > overhead of needing to read lots of small bits of the file makes the
> > use of malloc/read to be a pain in the *ss, then go for it.
> 
> An example of this in git is the regular pack-file accesses. We're MUCH 
> better off just mmap'ing the whole pack-file (or at least big chunks of 
> it) and not having to maintain difficult structures of "this is where I 
> read that part of the file into memory", or read _big_ chunks when 
> quite often we just use a few kB of it.
> 
> So mmap for pack-files does make sense, but probably only when you can 
> mmap big chunks, and are going to access much smaller (random) parts of 
> it.
Yes, exactly.

git-fast-import mmaps the pack file for this very reason. It every so often needs to go back and reread a tree object which has expired from its own in-memory LRU cache. This usually doesn't happen very often, but when it does we don't know where we are going to jump to get data from. mmaping a huge segment of the pack file (or the whole thing if its reasonably small) works for this case as the OS buffer cache can just take care of it for us. But as Linus pointed out mmap and write() aren't safe on some systems. Arrrgh.

However git-fast-import would probably work just as well (or maybe slightly better) with pread(). I really should port that code forward to current Git, use pread() instead, and submit the patch to Junio. But nobody really showed a lot of interest.

My sliding window pack-file access implementation (that I'm currently rewriting on top of current Git) tries to work in very large chunks, by default its 32 MiB per chunk, but its user/repository configurable so kernel hackers may just set it to 256 MiB and continue to get one large mmap for quite some time to come. Of course I would also like to get that to autoselect the window size rather than just hardcode it. :-)

The implementation would prefer a very small number (<8) of very large chunks (>32 MiB), but is designed to more gracefully degrade on huge packs on limited address space systems (e.g. Windows 32 bit) then the current code does.

Previous: Linus TorvaldsNext: Shawn Pearce
Message 30 of 51 in “Re: [PATCH] fetch-pack: avoid fixing thin packs when unnecessary”
  1. Johannes SchindelinDec 18, 2006
  2. Nicolas PitreDec 18, 2006
  3. Randal L. SchwartzDec 18, 2006
  4. Nicolas PitreDec 18, 2006
  5. Randal L. SchwartzDec 18, 2006
  6. Nicolas PitreDec 18, 2006
  7. Linus TorvaldsDec 18, 2006
  8. Randal L. SchwartzDec 18, 2006
  9. Martin LanghoffDec 18, 2006
  10. Kyle MoffettDec 22, 2006
  11. Shawn PearceDec 22, 2006
  12. Marco RoelandDec 22, 2006
  13. Andreas EricssonJan 3, 2007
  14. Linus TorvaldsDec 18, 2006
  15. Nicolas PitreDec 19, 2006
  16. Theodore TsoDec 19, 2006
  17. Shawn PearceDec 19, 2006
  18. Linus TorvaldsDec 19, 2006
  19. Shawn PearceDec 19, 2006
  20. Marco RoelandDec 19, 2006
  21. Shawn PearceDec 19, 2006
  22. Shawn PearceDec 19, 2006
  23. Marco RoelandDec 19, 2006
  24. Shawn PearceDec 19, 2006
  25. Marco RoelandDec 19, 2006
  26. Alex RiesenDec 19, 2006
  27. Juergen RuehleDec 21, 2006
  28. Theodore TsoDec 19, 2006
  29. Linus TorvaldsDec 19, 2006
  30. Shawn PearceDec 20, 2006
  31. Shawn PearceDec 20, 2006
  32. Linus TorvaldsDec 19, 2006
  33. Johannes SchindelinDec 19, 2006
  34. Junio C HamanoDec 19, 2006
  35. Jeff KingDec 19, 2006
  36. Andy WhitcroftDec 19, 2006
  37. index-pack usage of mmap() is unacceptably slower on many OSes other than LinuxNicolas Pitre, Dec 19, 2006
  38. Junio C HamanoDec 19, 2006
  39. Nicolas PitreDec 19, 2006
  40. Linus TorvaldsDec 19, 2006
  41. Randal L. SchwartzDec 19, 2006
  42. Randal L. SchwartzDec 19, 2006
  43. Jeff GarzikDec 19, 2006
  44. Junio C HamanoDec 20, 2006
  45. Linus TorvaldsDec 20, 2006
  46. Jeff GarzikDec 20, 2006
  47. Junio C HamanoDec 20, 2006
  48. Junio C HamanoDec 20, 2006
  49. Linus TorvaldsDec 20, 2006
  50. Junio C HamanoDec 20, 2006
  51. Nikolai WeibullDec 20, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.