git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v4 2/6] core.fsyncobjectfiles: batched disk flushes

From
Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Date
Sep 22, 2021, 02:02 UTC
Message-ID
<87sfxx73vm.fsf@evledraar.gmail.com>
In-Reply-To
<CANQDOdc1bNwDYhJ8ck2cwUfKmr3064uBHFDACphW+cGZRd-6EQ@mail.gmail.com>
On Tue, Sep 21 2021, Neeraj Singh wrote:
Show 67 quoted lines
> On Tue, Sep 21, 2021 at 4:41 PM Ævar Arnfjörð Bjarmason
> <avarab@gmail.com> wrote:
>>
>>
>> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:
>>
>> > When the new mode is enabled we do the following for new objects:
>> >
>> > 1. Create a tmp_obj_XXXX file and write the object data to it.
>> > 2. Issue a pagecache writeback request and wait for it to complete.
>> > 3. Record the tmp name and the final name in the bulk-checkin state for
>> >    later rename.
>> >
>> > At the end of the entire transaction we:
>> > 1. Issue a fsync against the lock file to flush the hardware writeback
>> >    cache, which should by now have processed the tmp file writes.
>> > 2. Rename all of the temp files to their final names.
>> > 3. When updating the index and/or refs, we assume that Git will issue
>> >    another fsync internal to that operation.
>>
>> Perhaps note too that:
>>
>> 4. For loose objects, refs etc. we may or may not create directories,
>>    and most certainly will be updating metadata on the immediate
>>    directory containing the file, but none of that's fsync()'d.
>>
>> > On a filesystem with a singular journal that is updated during name
>> > operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we
>> > would expect the fsync to trigger a journal writeout so that this
>> > sequence is enough to ensure that the user's data is durable by the time
>> > the git command returns.
>> >
>> > This change also updates the macOS code to trigger a real hardware flush
>> > via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on
>> > macOS there was no guarantee of durability since a simple fsync(2) call
>> > does not flush any hardware caches.
>>
>> There's no discussion of whether this is or isn't known to also work
>> some Linux FS's, and for these OS's where this does work is this only
>> for the object files themselves, or does metadata also "ride along"?
>>
>
> I unfortunately can't examine Linux kernel source code and the details
> of metadata
> consistency behavior across files is not something that anyone in that
> group wants
> to pin down. As far as I can tell, the only thing that's really
> guaranteed is fsyncing
> every single file you write down and its parent directory if you're
> creating a new file
> (which we always are).  As came up in conversation with Christoph
> Hellwig elsewhere
> on thread, Linux doesn't have any set of syscalls to make batch mode
> safe.  It does look
> like XFS would be safe if sync_file_ranges actually promised to wait
> for all pagecache
> writeback definitively, since it would do a "log force" to push all
> the dirty metadata to
> disk when we do our final fsync.
>
> I really didn't want to say something definitive about what Linux can
> or will do, since I'm
> not in a position to really know or influence them.  Christoph did say
> that he would be
> interested in contributing a variant to this patch that would be
> definitively safe on filesystems
> that honor syncfs.
*nod*, it's fine if it's omitted. Just wondering if we knew but weren't
 saying etc.
Show 35 quoted lines
>> > _Performance numbers_:
>> >
>> > Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.
>> > Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.
>> > Windows - Same host as Linux, a preview version of Windows 11.
>> >         This number is from a patch later in the series.
>> >
>> > Adding 500 files to the repo with 'git add' Times reported in seconds.
>> >
>> > core.fsyncObjectFiles | Linux | Mac   | Windows
>> > ----------------------|-------|-------|--------
>> >                 false | 0.06  |  0.35 | 0.61
>> >                 true  | 1.88  | 11.18 | 2.47
>> >                 batch | 0.15  |  0.41 | 1.53
>>
>> Per my https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com
>> and 6/6 in this series we've got perf tests for add/stash, but it would
>> be really interesting to see how this is impacted by
>> transfer.unpackLimit in cases where we may be writing packs or loose
>> objects.
>
> I'm having trouble understanding how unpackLimit is related to 'git stash'
> or 'git add'. From code inspection, it doesn't look like we're using
> those settings
> for adding objects except from across a transport.
>
> Are you proposing that we have a similar setting for adding objects
> via 'add' using
> a packfile?  I think that would be a good goal, but it might be a bit
> tricky since we've
> likely done a lot of the work to buffer the input objects in order to
> compute their OIDs,
> before we know how many objects there are to add. If the policy were
> to "always add to
> a packfile", it would be easier.

No, just that in the documentation that we should be explaining to the reader that this mode that optimizes for loose object writing benefits particular commands, but e.g. on the server-side that we'll probably never write 500 objects, but stream them to one pack.

Which might also inform next steps for the commands this does help with, i.e. can we make more things stream to packs? I think having this mode is at worst a good transitory thing to have, but perhaps longer term we'll want to simply write fewer individual loose objects.

In any case, pushing to a server with this configured and scaling that by transfer.unpackLimit should nicely demonstrate the pack v.s. loose object scenario at different fsck-settings.

Show 67 quoted lines
>>
>> > [...]
>> >  core.fsyncObjectFiles::
>> > -     This boolean will enable 'fsync()' when writing object files.
>> > -+
>> > -This is a total waste of time and effort on a filesystem that orders
>> > -data writes properly, but can be useful for filesystems that do not use
>> > -journalling (traditional UNIX filesystems) or that only journal metadata
>> > -and not file contents (OS X's HFS+, or Linux ext3 with "data=writeback").
>> > +     A value indicating the level of effort Git will expend in
>> > +     trying to make objects added to the repo durable in the event
>> > +     of an unclean system shutdown. This setting currently only
>> > +     controls the object store, so updates to any refs or the
>> > +     index may not be equally durable.
>>
>> All these mentions of "object" should really clarify that it's "loose
>> objects", i.e. we always fsync pack files.
>>
>> > +* `false` allows data to remain in file system caches according to
>> > +  operating system policy, whence it may be lost if the system loses power
>> > +  or crashes.
>>
>> As noted in point #4 of
>> https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com/ while
>> this direction is overall an improvement over the previously flippant
>> docs, they at least alluded to the context that the assumption behind
>> "false" is that you don't really care about loose objects, you care
>> about loose objects *and* the ref update or whatever.
>>
>> As I think (this is from memory) we've covered already this may have
>> been all based on some old ext3 assumption, but it's probably worth
>> summarizing that here, i.e. if you've got an FS with global ordered
>> operations you can probably skip this, but probably not etc.
>>
>> > +* `true` triggers a data integrity flush for each object added to the
>> > +  object store. This is the safest setting that is likely to ensure durability
>> > +  across all operating systems and file systems that honor the 'fsync' system
>> > +  call. However, this setting comes with a significant performance cost on
>> > +  common hardware.
>>
>> This is really overpromising things by omitting the fact that eve if
>> we're getting this feature you've hacked up right, we're still not
>> fsyncing dir entries etc (also noted above).
>>
>> So something that describes the narrow scope here, along with "loose
>> objects" etc....
>>
>> > +* `batch` enables an experimental mode that uses interfaces available in some
>> > +  operating systems to write object data with a minimal set of FLUSH CACHE
>> > +  (or equivalent) commands sent to the storage controller. If the operating
>> > +  system interfaces are not available, this mode behaves the same as `true`.
>> > +  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS
>> > +  filesystems and on Windows for repos stored on NTFS or ReFS.
>>
>> Again, even if it's called "core.fsyncObjectFiles" if we're going to say
>> "safe" we really need to say safe in what sense. Having written and
>> fsync()'d the file is helping nobody if the metadata never arrives....
>>
>
> My concern with your feedback here is that this is user-facing documentation.
> I'd assume that people who are not intimately familiar with both their
> filesystem
> and Git's internals would just be completely mystified by a long commentary on
> the specifics in the Config documentation. I think over time Git should focus on
> making this setting really guarantee durability in a meaningful way
> across the entire
> repository.

Yeah, this setting though is probably going to be tweaked only by fairly expert-level users of git.

I think it's fine if it just explicitly punts and says something like 'this is what it does, this may or may not work on your FS' etc., my main issue with the current docs is that they give off this vibe of knowing a lot more than they're telling you.

Show 27 quoted lines
>> > +static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)
>> > +{
>> > +     if (fsync_state->nr) {
>>
>> I think less indentation here would be nice:
>>
>>     if (!fsync_state->nr)
>>         return;
>>     /* rest of unindented body */
>>
>
> Will fix.
>
>> Or better yet do this check in unplug_bulk_checkin(), then here:
>>
>>     fsync_or_die();
>>     for_each_string_list_item() { ...}
>>     string_list_clear(....);
>>
>>
>
> I'd prefer to put it in the callee for reasons of
> separation-of-concerns.  I don't want
> to have the caller and callee partially implement the contract. The
> compiler should
> do a good enough job, since it's only one caller and will probably get
> totally inilined.
*nod*

For what it's worth I meant the "inlined" just in terms of avoiding the indirection for human readers, it won't matter to the machine, especially since this is all I/O bound...

Previous: Neeraj SinghNext: Neeraj Singh
Message 73 of 160 in “[RFC] Implement a bulk-checkin option for core.fsyncObjectFiles”
  1. 0/2 [RFC] Implement a bulk-checkin option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Aug 25, 2021
  2. 1/2 object-file: use futimes rather than utimeNeeraj Singh via GitGitGadget, Aug 25, 2021
  3. Johannes SchindelinAug 25, 2021
  4. Neeraj SinghAug 25, 2021
  5. 2/2 core.fsyncobjectfiles: batch disk flushesNeeraj Singh via GitGitGadget, Aug 25, 2021
  6. Christoph HellwigAug 25, 2021
  7. Neeraj SinghAug 25, 2021
  8. Christoph HellwigAug 26, 2021
  9. Ævar Arnfjörð BjarmasonAug 25, 2021
  10. Neeraj SinghAug 26, 2021
  11. Christoph HellwigAug 26, 2021
  12. Neeraj SinghAug 28, 2021
  13. Christoph HellwigAug 28, 2021
  14. Neeraj SinghAug 31, 2021
  15. Christoph HellwigSep 1, 2021
  16. Christoph HellwigAug 26, 2021
  17. Johannes SchindelinAug 25, 2021
  18. Junio C HamanoAug 25, 2021
  19. Neeraj SinghAug 26, 2021
  20. Neeraj SinghAug 25, 2021
  21. 0/6 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Aug 27, 2021
  22. 1/6 object-file: use futimens rather than utimeNeeraj Singh via GitGitGadget, Aug 27, 2021
  23. 2/6 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Aug 27, 2021
  24. 3/6 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Aug 27, 2021
  25. 5/6 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Aug 27, 2021
  26. 6/6 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Aug 27, 2021
  27. 4/6 core.fsyncobjectfiles: add windows support for batch modeNeeraj Singh via GitGitGadget, Aug 27, 2021
  28. Neeraj SinghSep 7, 2021
  29. Ævar Arnfjörð BjarmasonSep 7, 2021
  30. Randall S. BeckerSep 7, 2021
  31. Neeraj SinghSep 8, 2021
  32. Ævar Arnfjörð BjarmasonSep 8, 2021
  33. Randall S. BeckerSep 8, 2021
  34. Neeraj SinghSep 8, 2021
  35. Neeraj SinghSep 8, 2021
  36. Junio C HamanoSep 8, 2021
  37. Christoph HellwigSep 8, 2021
  38. Randall S. BeckerSep 8, 2021
  39. 'Christoph Hellwig'Sep 8, 2021
  40. Randall S. BeckerSep 8, 2021
  41. Neeraj SinghSep 8, 2021
  42. Junio C HamanoSep 8, 2021
  43. Neeraj SinghSep 8, 2021
  44. Ævar Arnfjörð BjarmasonSep 8, 2021
  45. 0/6 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Sep 14, 2021
  46. 1/6 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Sep 14, 2021
  47. 2/6 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Sep 14, 2021
  48. Bagas SanjayaSep 14, 2021
  49. Neeraj SinghSep 14, 2021
  50. Junio C HamanoSep 14, 2021
  51. Junio C HamanoSep 14, 2021
  52. Neeraj SinghSep 15, 2021
  53. 3/6 core.fsyncobjectfiles: add windows support for batch modeNeeraj Singh via GitGitGadget, Sep 14, 2021
  54. 4/6 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 14, 2021
  55. Junio C HamanoSep 14, 2021
  56. 5/6 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Sep 14, 2021
  57. 6/6 core.fsyncobjectfiles: enable batch mode for testingNeeraj Singh via GitGitGadget, Sep 14, 2021
  58. Junio C HamanoSep 15, 2021
  59. Neeraj SinghSep 15, 2021
  60. Junio C HamanoSep 15, 2021
  61. Junio C HamanoSep 16, 2021
  62. Christoph HellwigSep 14, 2021
  63. 0/6 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Sep 20, 2021
  64. 5/6 core.fsyncobjectfiles: tests for batch modeNeeraj Singh via GitGitGadget, Sep 20, 2021
  65. Ævar Arnfjörð BjarmasonSep 21, 2021
  66. Neeraj SinghSep 22, 2021
  67. Ævar Arnfjörð BjarmasonSep 22, 2021
  68. Neeraj SinghSep 22, 2021
  69. Ævar Arnfjörð BjarmasonSep 22, 2021
  70. 2/6 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Sep 20, 2021
  71. Ævar Arnfjörð BjarmasonSep 21, 2021
  72. Neeraj SinghSep 22, 2021
  73. Ævar Arnfjörð BjarmasonSep 22, 2021
  74. Neeraj SinghSep 22, 2021
  75. 1/6 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Sep 20, 2021
  76. 4/6 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 20, 2021
  77. Ævar Arnfjörð BjarmasonSep 21, 2021
  78. Neeraj SinghSep 22, 2021
  79. Neeraj SinghSep 23, 2021
  80. 3/6 core.fsyncobjectfiles: add windows support for batch modeNeeraj Singh via GitGitGadget, Sep 20, 2021
  81. Ævar Arnfjörð BjarmasonSep 21, 2021
  82. Neeraj SinghSep 22, 2021
  83. 6/6 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Sep 20, 2021
  84. 0/7 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Sep 24, 2021
  85. 1/7 object-file.c: do not rename in a temp odbNeeraj Singh via GitGitGadget, Sep 24, 2021
  86. 2/7 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Sep 24, 2021
  87. 3/7 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Sep 24, 2021
  88. Neeraj SinghSep 24, 2021
  89. 4/7 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 24, 2021
  90. Neeraj SinghSep 24, 2021
  91. 5/7 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 24, 2021
  92. 6/7 core.fsyncobjectfiles: tests for batch modeNeeraj Singh via GitGitGadget, Sep 24, 2021
  93. 7/7 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Sep 24, 2021
  94. Neeraj SinghSep 24, 2021
  95. 0/8 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Sep 24, 2021
  96. 1/8 object-file.c: do not rename in a temp odbNeeraj Singh via GitGitGadget, Sep 24, 2021
  97. 2/8 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Sep 24, 2021
  98. 3/8 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Sep 24, 2021
  99. Bagas SanjayaSep 25, 2021
  100. Neeraj SinghSep 27, 2021
  101. 5/8 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 24, 2021
  102. 4/8 core.fsyncobjectfiles: add windows support for batch modeNeeraj Singh via GitGitGadget, Sep 24, 2021
  103. Junio C HamanoSep 27, 2021
  104. Neeraj SinghSep 27, 2021
  105. Neeraj SinghSep 27, 2021
  106. Junio C HamanoSep 27, 2021
  107. 6/8 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 24, 2021
  108. 8/8 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Sep 24, 2021
  109. 7/8 core.fsyncobjectfiles: tests for batch modeNeeraj Singh via GitGitGadget, Sep 24, 2021
  110. 0/9 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Sep 28, 2021
  111. 1/9 object-file.c: do not rename in a temp odbNeeraj Singh via GitGitGadget, Sep 28, 2021
  112. Jeff KingSep 28, 2021
  113. Neeraj SinghSep 29, 2021
  114. 3/9 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Sep 28, 2021
  115. 2/9 tmp-objdir: new API for creating temporary writable databasesNeeraj Singh via GitGitGadget, Sep 28, 2021
  116. Elijah NewrenSep 29, 2021
  117. Neeraj SinghSep 29, 2021
  118. 4/9 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Sep 28, 2021
  119. 5/9 core.fsyncobjectfiles: add windows support for batch modeNeeraj Singh via GitGitGadget, Sep 28, 2021
  120. 6/9 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 28, 2021
  121. 8/9 core.fsyncobjectfiles: tests for batch modeNeeraj Singh via GitGitGadget, Sep 28, 2021
  122. 7/9 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Sep 28, 2021
  123. 9/9 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Sep 28, 2021
  124. 0/9 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Oct 4, 2021
  125. 1/9 tmp-objdir: new API for creating temporary writable databasesNeeraj Singh via GitGitGadget, Oct 4, 2021
  126. 2/9 tmp-objdir: disable ref updates when replacing the primary odbNeeraj Singh via GitGitGadget, Oct 4, 2021
  127. 3/9 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Oct 4, 2021
  128. 4/9 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Oct 4, 2021
  129. 5/9 core.fsyncobjectfiles: add windows support for batch modeNeeraj Singh via GitGitGadget, Oct 4, 2021
  130. 6/9 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Oct 4, 2021
  131. 7/9 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Oct 4, 2021
  132. 8/9 core.fsyncobjectfiles: tests for batch modeNeeraj Singh via GitGitGadget, Oct 4, 2021
  133. 9/9 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Oct 4, 2021
  134. 0/9 Implement a batched fsync option for core.fsyncObjectFilesNeeraj K. Singh via GitGitGadget, Nov 15, 2021
  135. 2/9 tmp-objdir: disable ref updates when replacing the primary odbNeeraj Singh via GitGitGadget, Nov 15, 2021
  136. Ævar Arnfjörð BjarmasonNov 16, 2021
  137. Neeraj SinghNov 16, 2021
  138. 1/9 tmp-objdir: new API for creating temporary writable databasesNeeraj Singh via GitGitGadget, Nov 15, 2021
  139. Elijah NewrenNov 30, 2021
  140. Neeraj SinghNov 30, 2021
  141. Elijah NewrenNov 30, 2021
  142. 3/9 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Nov 15, 2021
  143. 4/9 core.fsyncobjectfiles: batched disk flushesNeeraj Singh via GitGitGadget, Nov 15, 2021
  144. 5/9 core.fsyncobjectfiles: add windows support for batch modeNeeraj Singh via GitGitGadget, Nov 15, 2021
  145. 6/9 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Nov 15, 2021
  146. 7/9 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Nov 15, 2021
  147. 8/9 core.fsyncobjectfiles: tests for batch modeNeeraj Singh via GitGitGadget, Nov 15, 2021
  148. 9/9 core.fsyncobjectfiles: performance tests for add and stashNeeraj Singh via GitGitGadget, Nov 15, 2021
  149. Ævar Arnfjörð BjarmasonNov 16, 2021
  150. Neeraj SinghNov 17, 2021
  151. Ævar Arnfjörð BjarmasonNov 17, 2021
  152. Neeraj SinghNov 18, 2021
  153. Ævar Arnfjörð BjarmasonDec 1, 2021
  154. Ævar Arnfjörð BjarmasonMar 9, 2022
  155. Neeraj SinghMar 10, 2022
  156. Ævar Arnfjörð BjarmasonMar 10, 2022
  157. Neeraj SinghMar 10, 2022
  158. rsbecker@nexbridge.comMar 10, 2022
  159. Neeraj SinghMar 10, 2022
  160. rsbecker@nexbridge.comMar 10, 2022

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.