git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 2/2] Implement a simple delta_base cache

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Mar 18, 2007, 17:23 UTC
Message-ID
<Pine.LNX.4.64.0703181012520.6730@woody.linux-foundation.org>
In-Reply-To
<Pine.LNX.4.64.0703180517360.24626@beast.quantumfyre.co.uk>
On Sun, 18 Mar 2007, Julian Phillips wrote:
> 
> (This is a rather unrealistic repository consisting of a long series of
> commits of new binary files, but I don't have access to the repository that is
> being approximated until I get back to work on Monday ...)
This is a *horrible* test repo.

Is this actually really trying to approximate anything you work with? If so, please check whether you have cyanide or some other effective poison to kill all your cow-orkers - it's really doing them a favor - and then do the honorable thing yourself? Use something especially painful on whoever came up with the idea to track 25000 files in a single directory.

I'll see what the profile is, but even without the repo full generated yet, I can already tell you that you should *not* put tens of thousands of files in a single directory like this.

It's not only usually horribly bad quite independently of any SCM issues (ie most filesystems will have some bad performance behaviour with things like this - if only because "readdir()" will inevitably be slow).

And for git it means that you lose all ability to efficiently prune away the parts of the tree that you don't care about. git will always end up working with a full linear filemanifest instead of a nice collection of recursive trees, and a lot of the nice tree-walking optimizations that git has will just end up being no-ops: each tree is always one *huge* manifest.

So it's not that git cannot handle it, it's that a lot of the nice things that make git really efficient simply won't trigger for your repository.

In short: avoiding tens of thousands of files in a single directory is *always* a good idea. With or without git.

(Again, SCM's that are really just "one file at a time" like CVS, won't care as much. They never really track all files anyway, so while they are limited by potential filesystem performance bottlenecks, they won't have the fundamental issue of tracking 25,000 files..)

		Linus
Previous: Julian PhillipsNext: Robin Rosenberg
Message 40 of 79 in “cleaner/better zlib sources?”
  1. Linus TorvaldsMar 16, 2007
  2. Shawn O. PearceMar 16, 2007
  3. Jeff GarzikMar 16, 2007
  4. Matt MackallMar 16, 2007
  5. Linus TorvaldsMar 16, 2007
  6. Linus TorvaldsMar 16, 2007
  7. Davide LibenziMar 16, 2007
  8. Linus TorvaldsMar 16, 2007
  9. Davide LibenziMar 16, 2007
  10. Linus TorvaldsMar 16, 2007
  11. Davide LibenziMar 16, 2007
  12. Linus TorvaldsMar 16, 2007
  13. Davide LibenziMar 16, 2007
  14. Linus TorvaldsMar 17, 2007
  15. Linus TorvaldsMar 17, 2007
  16. Nicolas PitreMar 17, 2007
  17. Shawn O. PearceMar 17, 2007
  18. Linus TorvaldsMar 17, 2007
  19. Linus TorvaldsMar 17, 2007
  20. 1/2 Make trivial wrapper functions around delta base generation and freeingLinus Torvalds, Mar 17, 2007
  21. 2/2 Implement a simple delta_base cacheLinus Torvalds, Mar 17, 2007
  22. Linus TorvaldsMar 17, 2007
  23. Junio C HamanoMar 17, 2007
  24. Linus TorvaldsMar 17, 2007
  25. Linus TorvaldsMar 17, 2007
  26. Nicolas PitreMar 18, 2007
  27. Junio C HamanoMar 18, 2007
  28. Junio C HamanoMar 17, 2007
  29. Linus TorvaldsMar 17, 2007
  30. Jon SmirlMar 17, 2007
  31. Morten WelinderMar 18, 2007
  32. Linus TorvaldsMar 18, 2007
  33. Nicolas PitreMar 18, 2007
  34. Linus TorvaldsMar 18, 2007
  35. Nicolas PitreMar 18, 2007
  36. Linus TorvaldsMar 18, 2007
  37. Nicolas PitreMar 18, 2007
  38. Linus TorvaldsMar 18, 2007
  39. Julian PhillipsMar 18, 2007
  40. Linus TorvaldsMar 18, 2007
  41. Robin RosenbergMar 18, 2007
  42. Linus TorvaldsMar 18, 2007
  43. Robin RosenbergMar 18, 2007
  44. Shawn O. PearceMar 18, 2007
  45. David BrodskyMar 19, 2007
  46. Robin RosenbergMar 20, 2007
  47. David BrodskyMar 20, 2007
  48. Linus TorvaldsMar 21, 2007
  49. Nicolas PitreMar 21, 2007
  50. 3/2 Avoid unnecessary strlen() callsLinus Torvalds, Mar 18, 2007
  51. Junio C HamanoMar 18, 2007
  52. Linus TorvaldsMar 18, 2007
  53. Linus TorvaldsMar 18, 2007
  54. Shawn O. PearceMar 18, 2007
  55. Linus TorvaldsMar 18, 2007
  56. Johannes SchindelinMar 20, 2007
  57. Shawn O. PearceMar 20, 2007
  58. Shawn O. PearceMar 20, 2007
  59. Linus TorvaldsMar 20, 2007
  60. Shawn O. PearceMar 20, 2007
  61. Linus TorvaldsMar 20, 2007
  62. Junio C HamanoMar 20, 2007
  63. Junio C HamanoMar 20, 2007
  64. Linus TorvaldsMar 20, 2007
  65. Shawn O. PearceMar 20, 2007
  66. Linus TorvaldsMar 20, 2007
  67. Linus TorvaldsMar 18, 2007
  68. Avi KivityMar 18, 2007
  69. Linus TorvaldsMar 17, 2007
  70. Jeff GarzikMar 16, 2007
  71. Matt MackallMar 16, 2007
  72. Linus TorvaldsMar 16, 2007
  73. Nicolas PitreMar 16, 2007
  74. Shawn O. PearceMar 16, 2007
  75. Nicolas PitreMar 16, 2007
  76. Linus TorvaldsMar 16, 2007
  77. Nicolas PitreMar 16, 2007
  78. Davide LibenziMar 16, 2007
  79. Davide LibenziMar 16, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.