git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Switching from CVS to GIT

From
Andreas Ericsson <ae@op5.se>
Date
Oct 16, 2007, 05:14 UTC
Message-ID
<471448D0.6080200@op5.se>
In-Reply-To
<uodezisvg.fsf@gnu.org>
Eli Zaretskii wrote:
Show 12 quoted lines
>> Date: Mon, 15 Oct 2007 20:45:02 -0400 (EDT)
>> From: Daniel Barkalow <barkalow@iabervon.org>
>> cc: Alex Riesen <raa.lkml@gmail.com>, Johannes.Schindelin@gmx.de, ae@op5.se, 
>>     tsuna@lrde.epita.fr, git@vger.kernel.org, make-w32@gnu.org
>>
>> I believe the hassle is that readdir doesn't necessarily report a README in 
>> a directory which is supposed to have a README, when it has a readme 
>> instead.
> 
> Sorry I'm asking potentially stupid questions out of ignorance: why
> would you want readdir to return `README' when you have `readme'?
> 

Because it might have been checked in as README, and since git is case sensitive that is what it'll think should be there when it reads the directories. If it's not, users get to see

	removed: README
	untracked: readme

and there's really no easy way out of this one, since users on a case- sensitive filesystem might be involved in this project too, so it could be an intentional rename, but we don't know for sure. Just clobbering the in-git file is wrong, but overwriting a file on disk is wrong too. git tries hard to not ever lose any data for the user.

Show 26 quoted lines
> 
>>>> - no acceptable level of performance in filesystem and VFS (readdir,
>>>>   stat, open and read/write are annoyingly slow)
>>> With what libraries?  Native `stat' and `readdir' are quite fast.
>>> Perhaps you mean the ported glibc (libgw32c), where `readdir' is
>>> indeed painfully slow, but then you don't need to use it.
>> We want getting stat info, using readdir to figure out what files exist, 
>> for 106083 files in 1603 directories with a hot cache to take under 1s; 
>> otherwise "git status" takes a noticeable amount of time with a medium-big 
>> project, and we want people to be able to get info on what's changed 
>> effectively instantly. My impression is that Windows' native stat and 
>> readdir are plenty fast for what normal Windows programs want, but we 
>> actually expect reasonable performance on an unreasonably-big 
>> metadata-heavy input.
> 
> If that's the issue, then it's not a good idea to call `stat' and
> `readdir' on Windows at all.  `stat' is a single system call on Posix
> systems, while on Windows it usually needs to go out of its way
> calling half a dozen system services to gather the `struct stat' info.
> You need to call something like FindFirstFile, which can do the job of
> `stat' and `readdir' together (and of `fnmatch', if you need to filter
> only some files) in one go.  I don't know whether this will scan 100K
> files under one second (maybe I will try it one of these days), but it
> will definitely be faster than `readdir'+`stat' by maybe as much as an
> order of magnitude.
> 

To be honest though, there are so many places which do the readdir+stat that I don't think it'd be worth factoring it out, especially since it *works* on windows. It's just slow, and only slow compared to various unices. I *think* (correct me if I'm wrong) that git is still faster than a whole bunch of other scm's on windows, but to one who's used to its performance on Linux that waiting several seconds to scan 10k files just feels wrong.

Show 6 quoted lines
>> We also expect to be able to make a sequence of file system operations 
>> such that programs starting at any time see the same database as the files 
>> containing the database get restructured.
> 
> Sorry, I don't understand this; please tell more about the operations,
> ``the same database'' issue (what database?)
The object database, located under .git/objects.
> and what do you mean by
> ``the files containing the database get restructured''.
> 
/* I'm on a limb here. Nicolas Pitre knows the git packfile format, so
 * perhaps he'll be kind enough to correct me if I'm wrong */

The mmap() stuff is primarily convenient when reading huge packfiles. As far as I understand it, they're ordered by some sort of delta similarity score, so mmap()'ing 100MiB or so of a certain packfile will most likely mean we have a couple of thousand "connected" revisions in memory. That database gets sort of restructured as the memory-chunk that's mmap()'ed get moved to read in the next couple of thousand revisions.

In all honesty, this doesn't matter much for already fully packed projects unless they're significantly larger than the Linux kernel, since git is so amazingly good at compressing large repos to a small size. Linux is ~180 MiB fully packed, and most developer's systems could just read() that entire packfile into memory without much problem. But then again, no-one's ever had problems supporting the "normal" cases.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231
Previous: Eli ZaretskiiNext: Eli Zaretskii
Message 63 of 120 in “Re: Switching from CVS to GIT”
  1. Benoit SIGOUREOct 14, 2007
  2. Marco CostalbaOct 14, 2007
  3. Johannes SchindelinOct 14, 2007
  4. Martin LanghoffOct 15, 2007
  5. Andreas EricssonOct 14, 2007
  6. Johannes SchindelinOct 14, 2007
  7. Andreas EricssonOct 14, 2007
  8. Johannes SchindelinOct 14, 2007
  9. Alex RiesenOct 14, 2007
  10. Eli ZaretskiiOct 14, 2007
  11. Johannes SchindelinOct 14, 2007
  12. Brian DessentOct 15, 2007
  13. Johannes SchindelinOct 15, 2007
  14. Johannes SchindelinOct 15, 2007
  15. Eli ZaretskiiOct 15, 2007
  16. Steffen ProhaskaOct 15, 2007
  17. Eli ZaretskiiOct 15, 2007
  18. Johannes SchindelinOct 15, 2007
  19. Eli ZaretskiiOct 15, 2007
  20. Johannes SixtOct 15, 2007
  21. Eli ZaretskiiOct 15, 2007
  22. Paul SmithOct 15, 2007
  23. Steffen ProhaskaOct 15, 2007
  24. Eli ZaretskiiOct 15, 2007
  25. Eli ZaretskiiOct 15, 2007
  26. Johannes SchindelinOct 15, 2007
  27. Benoit SIGOUREOct 15, 2007
  28. Alex RiesenOct 15, 2007
  29. Brian DessentOct 15, 2007
  30. Johannes SchindelinOct 15, 2007
  31. Brian DessentOct 15, 2007
  32. Johannes SchindelinOct 15, 2007
  33. Linus TorvaldsOct 15, 2007
  34. Johannes SchindelinOct 15, 2007
  35. Alex RiesenOct 15, 2007
  36. Eli ZaretskiiOct 15, 2007
  37. Johannes SchindelinOct 15, 2007
  38. Eli ZaretskiiOct 15, 2007
  39. Brian DessentOct 15, 2007
  40. Johannes SchindelinOct 15, 2007
  41. Steffen ProhaskaOct 15, 2007
  42. Johannes SchindelinOct 15, 2007
  43. Nguyen Thai Ngoc DuyOct 16, 2007
  44. Eli ZaretskiiOct 16, 2007
  45. Nguyen Thai Ngoc DuyOct 16, 2007
  46. Eli ZaretskiiOct 16, 2007
  47. Steffen ProhaskaOct 16, 2007
  48. Eli ZaretskiiOct 15, 2007
  49. Mark WattsOct 15, 2007
  50. Eli ZaretskiiOct 15, 2007
  51. Eli ZaretskiiOct 15, 2007
  52. Johannes SchindelinOct 15, 2007
  53. David KastrupOct 15, 2007
  54. David KastrupOct 15, 2007
  55. Alex RiesenOct 15, 2007
  56. Dave KornOct 15, 2007
  57. Johannes SchindelinOct 15, 2007
  58. Alex RiesenOct 15, 2007
  59. Alex RiesenOct 15, 2007
  60. Andreas EricssonOct 14, 2007
  61. Daniel BarkalowOct 16, 2007
  62. Eli ZaretskiiOct 16, 2007
  63. Andreas EricssonOct 16, 2007
  64. Eli ZaretskiiOct 16, 2007
  65. Daniel BarkalowOct 16, 2007
  66. Johannes SchindelinOct 16, 2007
  67. Peter KarlssonOct 16, 2007
  68. Eli ZaretskiiOct 16, 2007
  69. Eli ZaretskiiOct 16, 2007
  70. David KastrupOct 16, 2007
  71. Johannes SchindelinOct 16, 2007
  72. Dave KornOct 16, 2007
  73. David BrownOct 16, 2007
  74. Nicolas PitreOct 16, 2007
  75. Dave KornOct 16, 2007
  76. Christopher FaylorOct 16, 2007
  77. Andreas EricssonOct 16, 2007
  78. Steffen ProhaskaOct 16, 2007
  79. Johannes SchindelinOct 16, 2007
  80. Steffen ProhaskaOct 16, 2007
  81. Johannes SchindelinOct 16, 2007
  82. Steffen ProhaskaOct 16, 2007
  83. Johannes SchindelinOct 16, 2007
  84. Steffen ProhaskaOct 16, 2007
  85. Eli ZaretskiiOct 16, 2007
  86. Robin RosenbergOct 17, 2007
  87. Daniel BarkalowOct 16, 2007
  88. Eli ZaretskiiOct 16, 2007
  89. Johannes SchindelinOct 16, 2007
  90. David KastrupOct 16, 2007
  91. Eli ZaretskiiOct 16, 2007
  92. Johannes SchindelinOct 16, 2007
  93. Eli ZaretskiiOct 16, 2007
  94. Johannes SchindelinOct 16, 2007
  95. Eli ZaretskiiOct 16, 2007
  96. Daniel BarkalowOct 16, 2007
  97. David KastrupOct 16, 2007
  98. Johannes SixtOct 16, 2007
  99. Eli ZaretskiiOct 16, 2007
  100. Dave KornOct 14, 2007
  101. Johannes SchindelinOct 15, 2007
  102. Alex RiesenOct 15, 2007
  103. David BrownOct 15, 2007
  104. Eli ZaretskiiOct 15, 2007
  105. Andreas EricssonOct 15, 2007
  106. Johannes SixtOct 15, 2007
  107. Andreas EricssonOct 15, 2007
  108. Dave KornOct 15, 2007
  109. Michael GebetsroitherOct 15, 2007
  110. Alex RiesenOct 15, 2007
  111. David KastrupOct 15, 2007
  112. Alex RiesenOct 15, 2007
  113. Peter KarlssonOct 16, 2007
  114. Martin LanghoffOct 15, 2007
  115. Johannes SixtOct 15, 2007
  116. Shawn O. PearceOct 15, 2007
  117. Johannes SixtOct 16, 2007
  118. Shawn O. PearceOct 16, 2007
  119. Johannes SixtOct 16, 2007
  120. Johannes SchindelinOct 16, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.