git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git-svnimport failed and now git-repack hates me

From
Linus Torvalds <torvalds@osdl.org>
Date
Jan 4, 2007, 21:12 UTC
Message-ID
<Pine.LNX.4.64.0701041300410.3661@woody.osdl.org>
In-Reply-To
<204011cb0701041124g40440fd4udf1088ab1341c031@mail.gmail.com>
On Thu, 4 Jan 2007, Chris Lee wrote:
> 
> Seems like *something* was definitely lost there. The 'used' number
> didn't go down at all when I started doing other things; it went up as
> the new programs started

The 'used' number basically _never_ goes down as long as there is memory free. The kernel simply doesn't have any reason to free any of its caches, even if those caches end up not being very useful.

What happened is almost certainly that with your big unpacked repository, the kernel ended up using a lot of memory on filename caching. In other words, I'd have expected that if you were to do

	cat /proc/slabinfo

you'd have seen a _lot_ of memory being used for dentries ("dentry_cache") and inodes ("ext3_inode_cache" assuming you're an ext3 user).

The kernel can easily drop those caches on demand, but "free" isn't quite smart enough to know about them as being caches, so they will just show up as "used".

That said, since you didn't want them, dropping them by hand with sysctl certainly didn't hurt. Manual control can often be better than automatic heuristics..

So the reason why repacking is so useful is that it gets rid of all these millions of individual files. They all take up space on the disk, but they also do end up having a lot of caches associated with them.

Btw, you may find that despite your 4GB of RAM, you might still be better off with a swapfile. It gives the kernel a certain amount of freedom in choosing how to allocate memory, and perhaps more importantly, even when the kernel doesn't actively use it, it means that IF the kernel runs out of totally free memory (because it has decided to keep a lot of stuff in the dentry cache), it gives the kernel choices, and a certain "buffer" for making the right decision.

What often happens is that the memory management heuristics don't make the "perfect" choice (partly because it's theoretically impossible anyway, but largely just because it's just a damn hard problem to even get all that *close* to perfect), and having a swap partition or even a swap file just allows the kernel to make some mistakes without it hitting a hard wall of "oh, I can't do anything at all about this particular page".

So that buffer zone can be helpful in avoiding bad situations, but it can actually also end up improving performance - it doesn't sound like the case in this particular situation, but in some other loads there really are a lot of dirty pages that aren't all that useful and where the memory really could be better used for other things if the largely unused dirty page could just be written to disk.

			Linus
Previous: Chris LeeNext: Sasha Khapyorsky
Message 49 of 55 in “git-svnimport failed and now git-repack hates me”
  1. Chris LeeJan 3, 2007
  2. Linus TorvaldsJan 4, 2007
  3. Shawn O. PearceJan 4, 2007
  4. Shawn O. PearceJan 4, 2007
  5. Chris LeeJan 4, 2007
  6. Shawn O. PearceJan 4, 2007
  7. Chris LeeJan 4, 2007
  8. Shawn O. PearceJan 4, 2007
  9. Chris LeeJan 4, 2007
  10. Shawn O. PearceJan 4, 2007
  11. Chris LeeJan 4, 2007
  12. Chris LeeJan 4, 2007
  13. Chris LeeJan 4, 2007
  14. Linus TorvaldsJan 4, 2007
  15. Chris LeeJan 4, 2007
  16. Eric WongJan 4, 2007
  17. Randal L. SchwartzJan 4, 2007
  18. Eric WongJan 4, 2007
  19. git-svn: make --repack work consistently between fetch and multi-fetchEric Wong, Jan 5, 2007
  20. Junio C HamanoJan 4, 2007
  21. pack-check.c::verify_packfile(): don't run SHA-1 update on huge dataJunio C Hamano, Jan 4, 2007
  22. Chris LeeJan 4, 2007
  23. Junio C HamanoJan 4, 2007
  24. Chris LeeJan 5, 2007
  25. Junio C HamanoJan 5, 2007
  26. Chris LeeJan 5, 2007
  27. Shawn O. PearceJan 5, 2007
  28. Chris LeeJan 5, 2007
  29. Junio C HamanoJan 5, 2007
  30. Linus TorvaldsJan 5, 2007
  31. alanJan 5, 2007
  32. Eric WongJan 7, 2007
  33. Linus TorvaldsJan 5, 2007
  34. Junio C HamanoJan 5, 2007
  35. Linus TorvaldsJan 5, 2007
  36. Linus TorvaldsJan 5, 2007
  37. Junio C HamanoJan 5, 2007
  38. Linus TorvaldsJan 5, 2007
  39. Johannes SchindelinJan 6, 2007
  40. Chris LeeJan 5, 2007
  41. Junio C HamanoJan 5, 2007
  42. Linus TorvaldsJan 5, 2007
  43. Junio C HamanoJan 5, 2007
  44. Linus TorvaldsJan 6, 2007
  45. Linus TorvaldsJan 6, 2007
  46. Junio C HamanoJan 6, 2007
  47. Linus TorvaldsJan 6, 2007
  48. Chris LeeJan 4, 2007
  49. Linus TorvaldsJan 4, 2007
  50. Sasha KhapyorskyJan 4, 2007
  51. Chris LeeJan 4, 2007
  52. git-svnimport: support for incremental importSasha Khapyorsky, Jan 7, 2007
  53. Chris LeeJan 7, 2007
  54. Sasha KhapyorskyJan 7, 2007
  55. git-svnimport: fix edge revisions double importingSasha Khapyorsky, Jan 8, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.