git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Switching from CVS to GIT

From
Daniel Barkalow <barkalow@iabervon.org>
Date
Oct 16, 2007, 05:56 UTC
Message-ID
<Pine.LNX.4.64.0710160032020.7638@iabervon.org>
In-Reply-To
<uodezisvg.fsf@gnu.org>
On Tue, 16 Oct 2007, Eli Zaretskii wrote:
Show 11 quoted lines
> > Date: Mon, 15 Oct 2007 20:45:02 -0400 (EDT)
> > From: Daniel Barkalow <barkalow@iabervon.org>
> > cc: Alex Riesen <raa.lkml@gmail.com>, Johannes.Schindelin@gmx.de, ae@op5.se, 
> >     tsuna@lrde.epita.fr, git@vger.kernel.org, make-w32@gnu.org
> > 
> > I believe the hassle is that readdir doesn't necessarily report a README in 
> > a directory which is supposed to have a README, when it has a readme 
> > instead.
> 
> Sorry I'm asking potentially stupid questions out of ignorance: why
> would you want readdir to return `README' when you have `readme'?

Say the project upstream has the file being "README", but, for some reason, it has ended up checked out as "readme" in your directory. Since your filesystem is case insensitive, it's supposed to be the same file, but when git goes through the list of files in the directory, it sees "readme", and there's nothing between reachable.h and read-cache.c in the list of tracked files. We've got a sorted list of filenames we're tracking along with their most-recently-seen content, and we want to merge the results of readdir with them, and this is obviously more straightforward if the filename that's the match for "README" is provided byte-for-byte the same, and therefore sorts the same.

Show 5 quoted lines
> > I think we want O(n) comparison of sorted lists, which doesn't 
> > work if equivalent names don't sort the same.
> 
> You comparison function should be case-insensitive on Windows, or am I
> missing something?

We want both lists sorted, so that we can step through the pair together and always reach matches together. This requires that the equivalent names sort together, as well as comparing equal.

Show 26 quoted lines
> > > > - no acceptable level of performance in filesystem and VFS (readdir,
> > > >   stat, open and read/write are annoyingly slow)
> > > 
> > > With what libraries?  Native `stat' and `readdir' are quite fast.
> > > Perhaps you mean the ported glibc (libgw32c), where `readdir' is
> > > indeed painfully slow, but then you don't need to use it.
> > 
> > We want getting stat info, using readdir to figure out what files exist, 
> > for 106083 files in 1603 directories with a hot cache to take under 1s; 
> > otherwise "git status" takes a noticeable amount of time with a medium-big 
> > project, and we want people to be able to get info on what's changed 
> > effectively instantly. My impression is that Windows' native stat and 
> > readdir are plenty fast for what normal Windows programs want, but we 
> > actually expect reasonable performance on an unreasonably-big 
> > metadata-heavy input.
> 
> If that's the issue, then it's not a good idea to call `stat' and
> `readdir' on Windows at all.  `stat' is a single system call on Posix
> systems, while on Windows it usually needs to go out of its way
> calling half a dozen system services to gather the `struct stat' info.
> You need to call something like FindFirstFile, which can do the job of
> `stat' and `readdir' together (and of `fnmatch', if you need to filter
> only some files) in one go.  I don't know whether this will scan 100K
> files under one second (maybe I will try it one of these days), but it
> will definitely be faster than `readdir'+`stat' by maybe as much as an
> order of magnitude.

Ah, that's helpful. We don't actually care too much about the particular info in stat; we just want to know quickly if the file has changed, so we can hash only the ones that have been touched and get the actual content changes.

Show 10 quoted lines
> > > > - no real "mmap" (which kills perfomance and complicates code)
> > > 
> > > You only need mmap because you are accustomed to use it on GNU/Linux.
> > 
> > I believe the need here is quick setup and fast access to sparse portions 
> > of several 100M files. It's hard to beat a page fault for read speed.
> 
> If you need memory-mapped files, they are available on Windows.  I
> thought the original comment about `mmap' was because it was used to
> allocate memory, not read files into memory.

No, we get our memory with malloc like normal people. The mmap is because we want to feed files and parts of files to zlib, and mmap makes that easy.

Show 7 quoted lines
> > We also expect to be able to make a sequence of file system operations 
> > such that programs starting at any time see the same database as the files 
> > containing the database get restructured.
> 
> Sorry, I don't understand this; please tell more about the operations,
> ``the same database'' issue (what database?) and what do you mean by
> ``the files containing the database get restructured''.

Git is built around a database of objects, which includes "blobs" (file content), "trees" (directory structure), "commits" (history linkage), and "tags" (additional annotations). Each of these objects gets hashed, and is referenced by hash. So we need to be able to get the object with a given hash quickly, and write an object and take its hash (ideally, stream the write and find out the hash at the end, with the database key set at that point). Also, this database should be compressed effectively, because it ought to compress really well, since a lot of the blobs and trees are only slightly different from other blobs or trees (by whatever changes were made between that revision and other revisions).

The current implementation of the persistant storage of this database is a bit complicated, with the goal being that creating objects is really fast, and looking up objects doesn't degrade too quickly, and there are optimization operations available that take some time and speed up future lookups and reduce the storage overhead (especially so that data can be transferred efficiently). The tricky thing is that, while the optimization process is running, other programs may be reading the database, so (1) the files that are no longer needed, because better-optimized versions are in place, may be open in another task, and (2) complete and correct new files have to appear and be such that pre-existing tasks will find them before old files can be removed. The optimization creates "pack files" and "pack indices", where the pack file has a lot of objects with delta compression between them and zlib compression of them, and the index files tell where everything in the pack file is. So we mmap the index files to search through, and mmap portions of the pack files to get the data out of, and we may be using them as they're replaced with more comprehensive pack files by another task.

Now, it's entirely possible that a completely different database implementation would be better on Windows, but our current one does a lot of creating files under different names, moving them to names where they'll be seen (since this is atomic under POSIX, and partial files are never seen by other tasks). Also, once we have new files in place, we unlink the files that they replace, so that new tasks will use the new ones and tasks that already have old ones open can still get the data out of them. Also, the files generally get mmaped,

> > A unixy pipeline was convenient
> 
> Windows supports pipelines with almost 100% the same functionality as
> Posix.  Again, perhaps I'm missing something.

I'm probably the one missing something here; I don't really know anything about Windows, and I only know what code other people have had problems porting. Mostly what we use for IPC is pipelines, so, if they work well, I don't know what the problem is.

	-Daniel
*This .sig left intentionally blank*
Previous: Robin RosenbergNext: Eli Zaretskii
Message 87 of 120 in “Re: Switching from CVS to GIT”
  1. Benoit SIGOUREOct 14, 2007
  2. Marco CostalbaOct 14, 2007
  3. Johannes SchindelinOct 14, 2007
  4. Martin LanghoffOct 15, 2007
  5. Andreas EricssonOct 14, 2007
  6. Johannes SchindelinOct 14, 2007
  7. Andreas EricssonOct 14, 2007
  8. Johannes SchindelinOct 14, 2007
  9. Alex RiesenOct 14, 2007
  10. Eli ZaretskiiOct 14, 2007
  11. Johannes SchindelinOct 14, 2007
  12. Brian DessentOct 15, 2007
  13. Johannes SchindelinOct 15, 2007
  14. Johannes SchindelinOct 15, 2007
  15. Eli ZaretskiiOct 15, 2007
  16. Steffen ProhaskaOct 15, 2007
  17. Eli ZaretskiiOct 15, 2007
  18. Johannes SchindelinOct 15, 2007
  19. Eli ZaretskiiOct 15, 2007
  20. Johannes SixtOct 15, 2007
  21. Eli ZaretskiiOct 15, 2007
  22. Paul SmithOct 15, 2007
  23. Steffen ProhaskaOct 15, 2007
  24. Eli ZaretskiiOct 15, 2007
  25. Eli ZaretskiiOct 15, 2007
  26. Johannes SchindelinOct 15, 2007
  27. Benoit SIGOUREOct 15, 2007
  28. Alex RiesenOct 15, 2007
  29. Brian DessentOct 15, 2007
  30. Johannes SchindelinOct 15, 2007
  31. Brian DessentOct 15, 2007
  32. Johannes SchindelinOct 15, 2007
  33. Linus TorvaldsOct 15, 2007
  34. Johannes SchindelinOct 15, 2007
  35. Alex RiesenOct 15, 2007
  36. Eli ZaretskiiOct 15, 2007
  37. Johannes SchindelinOct 15, 2007
  38. Eli ZaretskiiOct 15, 2007
  39. Brian DessentOct 15, 2007
  40. Johannes SchindelinOct 15, 2007
  41. Steffen ProhaskaOct 15, 2007
  42. Johannes SchindelinOct 15, 2007
  43. Nguyen Thai Ngoc DuyOct 16, 2007
  44. Eli ZaretskiiOct 16, 2007
  45. Nguyen Thai Ngoc DuyOct 16, 2007
  46. Eli ZaretskiiOct 16, 2007
  47. Steffen ProhaskaOct 16, 2007
  48. Eli ZaretskiiOct 15, 2007
  49. Mark WattsOct 15, 2007
  50. Eli ZaretskiiOct 15, 2007
  51. Eli ZaretskiiOct 15, 2007
  52. Johannes SchindelinOct 15, 2007
  53. David KastrupOct 15, 2007
  54. David KastrupOct 15, 2007
  55. Alex RiesenOct 15, 2007
  56. Dave KornOct 15, 2007
  57. Johannes SchindelinOct 15, 2007
  58. Alex RiesenOct 15, 2007
  59. Alex RiesenOct 15, 2007
  60. Andreas EricssonOct 14, 2007
  61. Daniel BarkalowOct 16, 2007
  62. Eli ZaretskiiOct 16, 2007
  63. Andreas EricssonOct 16, 2007
  64. Eli ZaretskiiOct 16, 2007
  65. Daniel BarkalowOct 16, 2007
  66. Johannes SchindelinOct 16, 2007
  67. Peter KarlssonOct 16, 2007
  68. Eli ZaretskiiOct 16, 2007
  69. Eli ZaretskiiOct 16, 2007
  70. David KastrupOct 16, 2007
  71. Johannes SchindelinOct 16, 2007
  72. Dave KornOct 16, 2007
  73. David BrownOct 16, 2007
  74. Nicolas PitreOct 16, 2007
  75. Dave KornOct 16, 2007
  76. Christopher FaylorOct 16, 2007
  77. Andreas EricssonOct 16, 2007
  78. Steffen ProhaskaOct 16, 2007
  79. Johannes SchindelinOct 16, 2007
  80. Steffen ProhaskaOct 16, 2007
  81. Johannes SchindelinOct 16, 2007
  82. Steffen ProhaskaOct 16, 2007
  83. Johannes SchindelinOct 16, 2007
  84. Steffen ProhaskaOct 16, 2007
  85. Eli ZaretskiiOct 16, 2007
  86. Robin RosenbergOct 17, 2007
  87. Daniel BarkalowOct 16, 2007
  88. Eli ZaretskiiOct 16, 2007
  89. Johannes SchindelinOct 16, 2007
  90. David KastrupOct 16, 2007
  91. Eli ZaretskiiOct 16, 2007
  92. Johannes SchindelinOct 16, 2007
  93. Eli ZaretskiiOct 16, 2007
  94. Johannes SchindelinOct 16, 2007
  95. Eli ZaretskiiOct 16, 2007
  96. Daniel BarkalowOct 16, 2007
  97. David KastrupOct 16, 2007
  98. Johannes SixtOct 16, 2007
  99. Eli ZaretskiiOct 16, 2007
  100. Dave KornOct 14, 2007
  101. Johannes SchindelinOct 15, 2007
  102. Alex RiesenOct 15, 2007
  103. David BrownOct 15, 2007
  104. Eli ZaretskiiOct 15, 2007
  105. Andreas EricssonOct 15, 2007
  106. Johannes SixtOct 15, 2007
  107. Andreas EricssonOct 15, 2007
  108. Dave KornOct 15, 2007
  109. Michael GebetsroitherOct 15, 2007
  110. Alex RiesenOct 15, 2007
  111. David KastrupOct 15, 2007
  112. Alex RiesenOct 15, 2007
  113. Peter KarlssonOct 16, 2007
  114. Martin LanghoffOct 15, 2007
  115. Johannes SixtOct 15, 2007
  116. Shawn O. PearceOct 15, 2007
  117. Johannes SixtOct 16, 2007
  118. Shawn O. PearceOct 16, 2007
  119. Johannes SixtOct 16, 2007
  120. Johannes SchindelinOct 16, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.