{"thread":{"id":"56870","subject":"RFC: A configuration design for future-proofing fsync() configuration","startedAt":"2021-11-10T15:56:38Z","lastAt":"2021-11-18T19:47:14Z","messageCount":9,"participants":["Ævar Arnfjörð Bjarmason","Neeraj Singh","Junio C Hamano","Christoph Hellwig"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"440850","messageId":"211110.86r1bogg27.gmgdl@evledraar.gmail.com","threadId":"56870","inReplyTo":null,"subject":"RFC: A configuration design for future-proofing fsync() configuration","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-11-10T15:09:33Z","receivedAt":"2021-11-10T15:56:38Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"As a follow-up to various fsync topics in-flight I've been encouraging\nthose involved to come up with some way to configure fsync() in a way\nthat'll make holistic sense in the end-state.\n\nContinuing a discussion from [1] currently we have:\n\n    ; Defaults to 'false'\n    core.fsyncObjectFiles = [true|false]\n\nIn master..next this has been extended to this by Neeraj:\n\n   core.fsyncObjectFiles = [true|false|batch]\n\nWhich, as an aside I hadn't considered before and I think we need to\nchange before it lands on \"master\", we really don't want config users\nwant to enable that makes older versions hard die. It's annoying to want\nto configure a new thing and not being able to put it in .gitconfig\nbecause older versions die on it:\n\n    $ git -c core.fsyncObjectFiles=batch status; echo $?\n    fatal: bad boolean config value 'batch' for 'core.fsyncobjectfiles'\n    128\n\nThen there's Eric Wong's proposed[2]:\n\n    core.fsync = <bool>\n\nAnd now Patrick Steinhardt has a proposal to extend Neeraj's with[3]:\n\n    ; Like core.fsyncObjectFiles, but apparently for .git/refs, not\n    ; .git/objects (but see my confusion on that topic in [1])\n    core.fsyncRefFiles = [<bool>|batch]\n\nI think this sort of config schema would make everyone above happy\n\nIt would:\n\n A) Be easy to extend for any future fsync behavior we'd reasonably\n    implement\n \n B) Not make older git versions die. It's fine if they warn(), but not die.\n\n C) Has some pretty contrived key names, but I'm trying to maintain the\n    constraint that you can set both fsck.X=Y and\n    e.g. fetch.fsck.X=Y. I.e. we should be able to configure things\n    globally *and* per-command, like color.*, fsck.* etc.\n\nProposal:\n\n  ; Turns on/off all fsync, whatever the method is. I.e. allows you to\n  ; never make any fsync() calls whatsoever (which we have another\n  ; in-flight topic for).\n\n  ; The \"false\" was controversial, and we could just leave it\n  ; unimplemented\n  core.fsync = <bool>\n\n  ; Optional, by default we'd use the most pedantic (I'd call our\n  ; current \"loose\", whether we want to forward-support it is another\n  ; matter.\n  ;\n  ; Whatever names we pick an option like this should ignore (or at most\n  ; warn about) values it doesn't know about, not hard die on it.\n  ;\n  ; Here \"bach\" is what Neeraj and Patrick are pursuing, a hypothetical\n  ; POSIX would be a pedantic way of exhaustively fsyncing everything.\n  ; \n  ; We'd leave door open to e.g. setting it to \"linux:ext4\" or whatever,\n  ; to do only the work needed on some specific popular FS\n  core.fsyncMethod = loose | POSIX | batch | linux:ext4 | NTFS | ...\n\n  ; Turn on or off entire categories of files we'd like to sync. This\n  ; way Neeraj's and Patrick's approach would be to set\n  ; core.fsyncMethod=batch, and then core.fsyncGroup=files &\n  ; core.fsyncGroup=refs.\n\n  ; If we learn about a new core.fsyncGroup = xyz in the future a <bool>\n  ; in \"core.fsyncGroupDefault\" will prevail. I.e. if true it's\n  ; included, if false not.\n  ;\n  ; Whether \"false\" or \"true\" is the default depends on\n  ; core.fsyncMethod. For POSIX it would be true, for \"loose\" it's\n  ; false.\n  core.fsyncGroup = files\n  core.fsyncGroup = refs\n  core.fsyncGroup = objects\n\nI'm not sure I like calling it \"group\". Maybe \"class\", \"category\"? Doing\nit with this structure is extensible to the two-level keys, as noted\nabove.\n\n  ; Our existing config knob. When \"false\" synonymous with:\n  ;\n  ;     core.fsync = true\n  ;     core.fsyncMethod = loose\n  ;     core.fsyncGroup = pack\n  ;\n  ; When \"true\" synonymous with the same as the above, plus:\n  ;     core.fsyncGroup = loose\n  ;\n  : Or something like that. I.e. we'll fsync *.pack, *.bitmap etc, and ;\n  ; probably some other stuff, but not loose objects etc.\n  ;\n  ; Whatever we fsync now exactly this schema should be generic enough\n  ; to support it.\n  core.fsyncObjectFiles = <bool>\n\n  ; A namespace for core.fsyncMethod = <X>. Specific methods will\n  ; own this namespace and can configure whatever they want.\n  fsyncMethod.<x>.<a> = <b>\n\nE.g. we might have:\n\n  fsyncMethod.POSIX.content = true\n  fsyncMethod.POSIX.metadata = false\n\nIf we know we'd like to (depending on other config) to fsync things\nexhaustively or not, but do different things depending on file content\nor metadata. I.e. maybe your FS's fsync() on a file fd always implies a\nsync of the metadata, and maybe not.\n\n  ; Change whatever fsync configuration you want per-command, similar to\n  ; fsck.* and fetch.fsck.*\n  transfer.fsyncGroup=*\n  fetch.fsyncGroup=*\n  ...\n\n1. https://lore.kernel.org/git/211110.86v910gi9a.gmgdl@evledraar.gmail.com/\n2. https://lore.kernel.org/git/20211028002102.19384-1-e@80x24.org/\n3. https://lore.kernel.org/git/cover.1636544377.git.ps@pks.im/\n"},{"id":"440893","messageId":"20211111004724.GA839@neerajsi-x1.localdomain","threadId":"56870","inReplyTo":"211110.86r1bogg27.gmgdl@evledraar.gmail.com","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-11T00:47:24Z","receivedAt":"2021-11-11T00:47:31Z","isPatch":false,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Nov 10, 2021 at 04:09:33PM +0100, Ævar Arnfjörð Bjarmason wrote:\n> As a follow-up to various fsync topics in-flight I've been encouraging\n> those involved to come up with some way to configure fsync() in a way\n> that'll make holistic sense in the end-state.\n> \n> Continuing a discussion from [1] currently we have:\n> \n>     ; Defaults to 'false'\n>     core.fsyncObjectFiles = [true|false]\n> \n> In master..next this has been extended to this by Neeraj:\n> \n>    core.fsyncObjectFiles = [true|false|batch]\n> \n> Which, as an aside I hadn't considered before and I think we need to\n> change before it lands on \"master\", we really don't want config users\n> want to enable that makes older versions hard die. It's annoying to want\n> to configure a new thing and not being able to put it in .gitconfig\n> because older versions die on it:\n> \n>     $ git -c core.fsyncObjectFiles=batch status; echo $?\n>     fatal: bad boolean config value 'batch' for 'core.fsyncobjectfiles'\n>     128\n> \n> Then there's Eric Wong's proposed[2]:\n> \n>     core.fsync = <bool>\n> \n> And now Patrick Steinhardt has a proposal to extend Neeraj's with[3]:\n> \n>     ; Like core.fsyncObjectFiles, but apparently for .git/refs, not\n>     ; .git/objects (but see my confusion on that topic in [1])\n>     core.fsyncRefFiles = [<bool>|batch]\n> \n> I think this sort of config schema would make everyone above happy\n> \n> It would:\n> \n>  A) Be easy to extend for any future fsync behavior we'd reasonably\n>     implement\n>  \n>  B) Not make older git versions die. It's fine if they warn(), but not die.\n> \n>  C) Has some pretty contrived key names, but I'm trying to maintain the\n>     constraint that you can set both fsck.X=Y and\n>     e.g. fetch.fsck.X=Y. I.e. we should be able to configure things\n>     globally *and* per-command, like color.*, fsck.* etc.\n> \n> Proposal:\n> \n>   ; Turns on/off all fsync, whatever the method is. I.e. allows you to\n>   ; never make any fsync() calls whatsoever (which we have another\n>   ; in-flight topic for).\n> \n>   ; The \"false\" was controversial, and we could just leave it\n>   ; unimplemented\n>   core.fsync = <bool>\n> \n>   ; Optional, by default we'd use the most pedantic (I'd call our\n>   ; current \"loose\", whether we want to forward-support it is another\n>   ; matter.\n>   ;\n>   ; Whatever names we pick an option like this should ignore (or at most\n>   ; warn about) values it doesn't know about, not hard die on it.\n>   ;\n>   ; Here \"bach\" is what Neeraj and Patrick are pursuing, a hypothetical\n>   ; POSIX would be a pedantic way of exhaustively fsyncing everything.\n>   ; \n>   ; We'd leave door open to e.g. setting it to \"linux:ext4\" or whatever,\n>   ; to do only the work needed on some specific popular FS\n>   core.fsyncMethod = loose | POSIX | batch | linux:ext4 | NTFS | ...\n> \n>   ; Turn on or off entire categories of files we'd like to sync. This\n>   ; way Neeraj's and Patrick's approach would be to set\n>   ; core.fsyncMethod=batch, and then core.fsyncGroup=files &\n>   ; core.fsyncGroup=refs.\n> \n>   ; If we learn about a new core.fsyncGroup = xyz in the future a <bool>\n>   ; in \"core.fsyncGroupDefault\" will prevail. I.e. if true it's\n>   ; included, if false not.\n>   ;\n>   ; Whether \"false\" or \"true\" is the default depends on\n>   ; core.fsyncMethod. For POSIX it would be true, for \"loose\" it's\n>   ; false.\n>   core.fsyncGroup = files\n>   core.fsyncGroup = refs\n>   core.fsyncGroup = objects\n> \n> I'm not sure I like calling it \"group\". Maybe \"class\", \"category\"? Doing\n> it with this structure is extensible to the two-level keys, as noted\n> above.\n> \n>   ; Our existing config knob. When \"false\" synonymous with:\n>   ;\n>   ;     core.fsync = true\n>   ;     core.fsyncMethod = loose\n>   ;     core.fsyncGroup = pack\n>   ;\n>   ; When \"true\" synonymous with the same as the above, plus:\n>   ;     core.fsyncGroup = loose\n>   ;\n>   : Or something like that. I.e. we'll fsync *.pack, *.bitmap etc, and ;\n>   ; probably some other stuff, but not loose objects etc.\n>   ;\n>   ; Whatever we fsync now exactly this schema should be generic enough\n>   ; to support it.\n>   core.fsyncObjectFiles = <bool>\n> \n>   ; A namespace for core.fsyncMethod = <X>. Specific methods will\n>   ; own this namespace and can configure whatever they want.\n>   fsyncMethod.<x>.<a> = <b>\n> \n> E.g. we might have:\n> \n>   fsyncMethod.POSIX.content = true\n>   fsyncMethod.POSIX.metadata = false\n> \n> If we know we'd like to (depending on other config) to fsync things\n> exhaustively or not, but do different things depending on file content\n> or metadata. I.e. maybe your FS's fsync() on a file fd always implies a\n> sync of the metadata, and maybe not.\n> \n>   ; Change whatever fsync configuration you want per-command, similar to\n>   ; fsck.* and fetch.fsck.*\n>   transfer.fsyncGroup=*\n>   fetch.fsyncGroup=*\n>   ...\n> \n> 1. https://lore.kernel.org/git/211110.86v910gi9a.gmgdl@evledraar.gmail.com/\n> 2. https://lore.kernel.org/git/20211028002102.19384-1-e@80x24.org/\n> 3. https://lore.kernel.org/git/cover.1636544377.git.ps@pks.im/\nHi Ævar,\n\nThanks for noticing the backwards compatibility issue with the 'batch' flag. I\nagree that we need to fix that before committing my changes to master.\n\nI'm hoping that we can agree to a version of what you're proposing, but my\npreference would be to cut out the more granular controls. I'd prefer to see\njust:\n\tcore.fsync = [bool]   \t\t- Turn fsyncing on or off.\n\tcore.fsyncMethod = [string] \t- Controls how it's done (with a non-fatal warn on unrecognized values).\n\tcore.fsyncObjectFiles = [bool]  - Sets core.fsync if that setting doesn't already have a value. For back-compat.\n\nI don't think either we or the users should have to reason about what it means\nfor some parts of the repo to be fsynced and others not to be. If core.fsync is\n'false' and someone gets a weird state after a system crash, no one should be\nsurprised. If core.fsync is 'true', and people are running on a reasonable\ncommon filesystem, we should be trying to give decent performance and good\ndurability.\n\nIt would be nice to loop in some Linux fs developers to find out what can be\ndone on current implementations to get the durability without terrible\nperformance. From reading the docs and mailing threads it looks like the\nsync_file_range + bulk fsync approach should actually work on the current XFS\nimplementation.\n\nThanks,\nNeeraj\nWindows Core Filesystem Dev\n"},{"id":"440894","messageId":"211111.86pmr7pk9o.gmgdl@evledraar.gmail.com","threadId":"56870","inReplyTo":"20211111004724.GA839@neerajsi-x1.localdomain","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-11-11T00:57:49Z","receivedAt":"2021-11-11T01:13:15Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Nov 10 2021, Neeraj Singh wrote:\n\n> On Wed, Nov 10, 2021 at 04:09:33PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>> As a follow-up to various fsync topics in-flight I've been encouraging\n>> those involved to come up with some way to configure fsync() in a way\n>> that'll make holistic sense in the end-state.\n>> \n>> Continuing a discussion from [1] currently we have:\n>> \n>>     ; Defaults to 'false'\n>>     core.fsyncObjectFiles = [true|false]\n>> \n>> In master..next this has been extended to this by Neeraj:\n>> \n>>    core.fsyncObjectFiles = [true|false|batch]\n>> \n>> Which, as an aside I hadn't considered before and I think we need to\n>> change before it lands on \"master\", we really don't want config users\n>> want to enable that makes older versions hard die. It's annoying to want\n>> to configure a new thing and not being able to put it in .gitconfig\n>> because older versions die on it:\n>> \n>>     $ git -c core.fsyncObjectFiles=batch status; echo $?\n>>     fatal: bad boolean config value 'batch' for 'core.fsyncobjectfiles'\n>>     128\n>> \n>> Then there's Eric Wong's proposed[2]:\n>> \n>>     core.fsync = <bool>\n>> \n>> And now Patrick Steinhardt has a proposal to extend Neeraj's with[3]:\n>> \n>>     ; Like core.fsyncObjectFiles, but apparently for .git/refs, not\n>>     ; .git/objects (but see my confusion on that topic in [1])\n>>     core.fsyncRefFiles = [<bool>|batch]\n>> \n>> I think this sort of config schema would make everyone above happy\n>> \n>> It would:\n>> \n>>  A) Be easy to extend for any future fsync behavior we'd reasonably\n>>     implement\n>>  \n>>  B) Not make older git versions die. It's fine if they warn(), but not die.\n>> \n>>  C) Has some pretty contrived key names, but I'm trying to maintain the\n>>     constraint that you can set both fsck.X=Y and\n>>     e.g. fetch.fsck.X=Y. I.e. we should be able to configure things\n>>     globally *and* per-command, like color.*, fsck.* etc.\n>> \n>> Proposal:\n>> \n>>   ; Turns on/off all fsync, whatever the method is. I.e. allows you to\n>>   ; never make any fsync() calls whatsoever (which we have another\n>>   ; in-flight topic for).\n>> \n>>   ; The \"false\" was controversial, and we could just leave it\n>>   ; unimplemented\n>>   core.fsync = <bool>\n>> \n>>   ; Optional, by default we'd use the most pedantic (I'd call our\n>>   ; current \"loose\", whether we want to forward-support it is another\n>>   ; matter.\n>>   ;\n>>   ; Whatever names we pick an option like this should ignore (or at most\n>>   ; warn about) values it doesn't know about, not hard die on it.\n>>   ;\n>>   ; Here \"bach\" is what Neeraj and Patrick are pursuing, a hypothetical\n>>   ; POSIX would be a pedantic way of exhaustively fsyncing everything.\n>>   ; \n>>   ; We'd leave door open to e.g. setting it to \"linux:ext4\" or whatever,\n>>   ; to do only the work needed on some specific popular FS\n>>   core.fsyncMethod = loose | POSIX | batch | linux:ext4 | NTFS | ...\n>> \n>>   ; Turn on or off entire categories of files we'd like to sync. This\n>>   ; way Neeraj's and Patrick's approach would be to set\n>>   ; core.fsyncMethod=batch, and then core.fsyncGroup=files &\n>>   ; core.fsyncGroup=refs.\n>> \n>>   ; If we learn about a new core.fsyncGroup = xyz in the future a <bool>\n>>   ; in \"core.fsyncGroupDefault\" will prevail. I.e. if true it's\n>>   ; included, if false not.\n>>   ;\n>>   ; Whether \"false\" or \"true\" is the default depends on\n>>   ; core.fsyncMethod. For POSIX it would be true, for \"loose\" it's\n>>   ; false.\n>>   core.fsyncGroup = files\n>>   core.fsyncGroup = refs\n>>   core.fsyncGroup = objects\n>> \n>> I'm not sure I like calling it \"group\". Maybe \"class\", \"category\"? Doing\n>> it with this structure is extensible to the two-level keys, as noted\n>> above.\n>> \n>>   ; Our existing config knob. When \"false\" synonymous with:\n>>   ;\n>>   ;     core.fsync = true\n>>   ;     core.fsyncMethod = loose\n>>   ;     core.fsyncGroup = pack\n>>   ;\n>>   ; When \"true\" synonymous with the same as the above, plus:\n>>   ;     core.fsyncGroup = loose\n>>   ;\n>>   : Or something like that. I.e. we'll fsync *.pack, *.bitmap etc, and ;\n>>   ; probably some other stuff, but not loose objects etc.\n>>   ;\n>>   ; Whatever we fsync now exactly this schema should be generic enough\n>>   ; to support it.\n>>   core.fsyncObjectFiles = <bool>\n>> \n>>   ; A namespace for core.fsyncMethod = <X>. Specific methods will\n>>   ; own this namespace and can configure whatever they want.\n>>   fsyncMethod.<x>.<a> = <b>\n>> \n>> E.g. we might have:\n>> \n>>   fsyncMethod.POSIX.content = true\n>>   fsyncMethod.POSIX.metadata = false\n>> \n>> If we know we'd like to (depending on other config) to fsync things\n>> exhaustively or not, but do different things depending on file content\n>> or metadata. I.e. maybe your FS's fsync() on a file fd always implies a\n>> sync of the metadata, and maybe not.\n>> \n>>   ; Change whatever fsync configuration you want per-command, similar to\n>>   ; fsck.* and fetch.fsck.*\n>>   transfer.fsyncGroup=*\n>>   fetch.fsyncGroup=*\n>>   ...\n>> \n>> 1. https://lore.kernel.org/git/211110.86v910gi9a.gmgdl@evledraar.gmail.com/\n>> 2. https://lore.kernel.org/git/20211028002102.19384-1-e@80x24.org/\n>> 3. https://lore.kernel.org/git/cover.1636544377.git.ps@pks.im/\n> Hi Ævar,\n>\n> Thanks for noticing the backwards compatibility issue with the 'batch' flag. I\n> agree that we need to fix that before committing my changes to master.\n>\n> I'm hoping that we can agree to a version of what you're proposing, but my\n> preference would be to cut out the more granular controls. I'd prefer to see\n> just:\n> \tcore.fsync = [bool]   \t\t- Turn fsyncing on or off.\n> \tcore.fsyncMethod = [string] \t- Controls how it's done (with a non-fatal warn on unrecognized values).\n> \tcore.fsyncObjectFiles = [bool]  - Sets core.fsync if that setting doesn't already have a value. For back-compat.\n\nI'm fine with something simpler as long as we don't think we'll\nplausibly start painting ourselves into a corner.\n\nBut core.fsyncObjectFiles is *not* a setting of a \"core.fsync\" in the\nsense that Eric suggested we have.\n\nI.e. it's effectively a sort of early and partial Linux-only version of\nwhat your \"batch\" mode is. I.e. to skip fsyncing the loose object files,\nand only fsync() the eventual refs we write.\n\n\"Sort of\" because we'd e.g. fsync packs unconditionally etc, but if we\nmake core.fsyncObjectFiles=false be core.fsync=false then we can't have\na \"real\" core.fsync=false, i.e. one that guarantees no fsync() calls at\nall.\n\nWe could also simply decide that it's a bad setting and we're going to\ndeprecate it, but another way is having a generic config layout that can\nexpress what it's doing and more.\n\n> I don't think either we or the users should have to reason about what it means\n> for some parts of the repo to be fsynced and others not to be. If core.fsync is\n> 'false' and someone gets a weird state after a system crash, no one should be\n> surprised.\n\nYes. I'm fine with leaving this on the table. I should have be more\nexplicit that I'm not suggesting we implement all this exhaustive config\nsupport, but if we imagine a sensible config schema that is extensible\n(my proposal may or may not be that) then we can implement just 1-2\nvariables in it and know that we have room to grow in the future.\n\n> If core.fsync is 'true', and people are running on a reasonable\n> common filesystem, we should be trying to give decent performance and good\n> durability.\n\nYeah, I just wonder if we can easily provide config to have people\ndecide that trade-off themselves.\n\nE.g. from the performance numbers in [1] I might turn off fsyncing when\nwe write anything in the working tree.\n\nWe don't do that particular thing now, but if we're being pulled in one\ndirection of always being fsync-safe by default...\n\nI can also see it being useful to e.g. do:\n\n    gc.fsync = false\n\nOr blacklist other similar batch operations, although with a global knob\nthat can also rather easily be:\n\n    git -c core.fsync gc\n\nSo maybe the whole \"fsck\" rationale doesn't apply here.\n\n> It would be nice to loop in some Linux fs developers to find out what can be\n> done on current implementations to get the durability without terrible\n> performance. From reading the docs and mailing threads it looks like the\n> sync_file_range + bulk fsync approach should actually work on the current XFS\n> implementation.\n\nI CC'd Linus on the topic of core.fsyncObjectFiles on this thread, and\nChristoph Hellwig who chimed in on the topic of the behavior of Linux\nFS's on recent threads, I don't know where we'd find a focused set of\nLinux devs who might be interested (I'm not going to just spam LKML). If\nanyone does pointing them to this thread would be most welcome.\n\n1. https://lore.kernel.org/git/YYwvVy6AWDjkWazn@coredump.intra.peff.net/\n"},{"id":"440925","messageId":"xmqqh7cimuxt.fsf@gitster.g","threadId":"56870","inReplyTo":"211110.86r1bogg27.gmgdl@evledraar.gmail.com","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-11-11T18:03:10Z","receivedAt":"2021-11-11T18:03:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> Continuing a discussion from [1] currently we have:\n>\n>     ; Defaults to 'false'\n>     core.fsyncObjectFiles = [true|false]\n>\n> In master..next this has been extended to this by Neeraj:\n>\n>    core.fsyncObjectFiles = [true|false|batch]\n>\n> Which, as an aside I hadn't considered before and I think we need to\n> change before it lands on \"master\", we really don't want config users\n> want to enable that makes older versions hard die. It's annoying to want\n> to configure a new thing and not being able to put it in .gitconfig\n> because older versions die on it:\n>\n>     $ git -c core.fsyncObjectFiles=batch status; echo $?\n>     fatal: bad boolean config value 'batch' for 'core.fsyncobjectfiles'\n>     128\n\nBut then it is also annoying to find out that the shiny new toy you\nthought you configured silently is not kicking in.  I actually think\nNeeraj's \"if you are in a mixed environment, you need to be aware of\nwhich copies of Git you use are prepared to use it\" would be better\nfor end users.\n\nFor us Git developers, it would be less convenient, though.\n"},{"id":"440960","messageId":"20211112055421.GA27823@lst.de","threadId":"56870","inReplyTo":"20211111004724.GA839@neerajsi-x1.localdomain","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-11-12T05:54:21Z","receivedAt":"2021-11-12T05:54:35Z","isPatch":false,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Wed, Nov 10, 2021 at 04:47:24PM -0800, Neeraj Singh wrote:\n> It would be nice to loop in some Linux fs developers to find out what can be\n> done on current implementations to get the durability without terrible\n> performance. From reading the docs and mailing threads it looks like the\n> sync_file_range + bulk fsync approach should actually work on the current XFS\n> implementation.\n\nIf you want more than just my advice linux-fsdevel@vger.kernel.org is\na good place to find a wide range of opinions.\n\nAnyway, I think syncfs is the biggest band for the buck as it will give\nyou very efficient syncing with very little overhead in git, but it does\nhave a huge noisy neighbor problem that might make it unattractive\nfor multi-tenant file systems or git hosting.\n"},{"id":"441519","messageId":"CANQDOdedAoOvPHra0e8PuOO68xt+gOSbbV3tHzGxcyJy5nTm_A@mail.gmail.com","threadId":"56870","inReplyTo":"20211112055421.GA27823@lst.de","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-17T18:49:20Z","receivedAt":"2021-11-17T18:49:35Z","isPatch":false,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Nov 11, 2021 at 9:54 PM Christoph Hellwig <hch@lst.de> wrote:\n>\n> On Wed, Nov 10, 2021 at 04:47:24PM -0800, Neeraj Singh wrote:\n> > It would be nice to loop in some Linux fs developers to find out what can be\n> > done on current implementations to get the durability without terrible\n> > performance. From reading the docs and mailing threads it looks like the\n> > sync_file_range + bulk fsync approach should actually work on the current XFS\n> > implementation.\n>\n> If you want more than just my advice linux-fsdevel@vger.kernel.org is\n> a good place to find a wide range of opinions.\n>\n> Anyway, I think syncfs is the biggest band for the buck as it will give\n> you very efficient syncing with very little overhead in git, but it does\n> have a huge noisy neighbor problem that might make it unattractive\n> for multi-tenant file systems or git hosting.\n\nTo summarize where we are at for linux-fsdevel:\nWe're working on making Git preserve data added to the repo even if\nthe system crashes or loses power at some point soon after a Git\ncommand completes. The default behavior of git-for-windows is to set\ncore.fsyncobjectfiles=true, which at least ensures durability for\nloose object files.\n\nThe current implementation of core.fsyncobjectfiles inserts an fsync\nbetween writing each new object to a temp name and renaming it to its\nfinal hash-based name. This approach is slow when adding hundreds of\nfiles to the repo [1]. The main cost on the hardware we tested is\nactually the CACHE_FLUSH request sent down to\nthe storage hardware. There is also work in-flight by Patrick\nSteinhardt to sync ref files [2].\n\nIn a patch series at [3], I implemented a batch mode that issues\npagecache writeback for each object file when it's being written and\nthen before any of the files are renamed to their final destination we\ndo an fsync to a dummy file on the same filesystem.  On linux, this is\nusing the sync_file_range(fd,0,0,  SYNC_FILE_RANGE_WRITE_AND_WAIT) to\ndo the pagecache writeback.  According to Amir's thread at [4] this\nflag combo should actually trigger the desired writeback. The\nexpectation is that the fsync of the dummy file should trigger a log\nwriteback and one or more CACHE_FLUSH commands to harden the block\nmapping metadata and directory entries such that the data would be\nretrievable after the fsync completes.\n\nThe equivalent sequence is specified to work on the common Windows\nfilesystems [5]. The question I have for the Linux community is\nwhether the same sequence will work on any of the common extant Linux\nfilesystems such that it can provide value to Git users on Linux. My\nunderstanding from Christoph Hellwig's comments is that on XFS at\nleast the sync_file_range, fsync, and rename sequence would allow us\nto guarantee that the complete written contents of the file would be\nvisible if the new name is visible.  I also expect that additional\nfsync to a dummy file after the renames would also ensure that the log\nis forced again, which should ensure that all of the renames are\nvisible before a ref file could be written that points at one of the\nobject names.\n\nI wasn't able to find any clear semantics about the ext4 filesystem,\nand I gather from what I've read that the btrfs filesystem does not\nsupport the desired semantics.  Christoph mentioned that syncfs would\nefficiently provide a batched CACHE_FLUSH with the cost of picking up\ndirty cached data unrelated to Git.\n\nAre there any opinions on the Linux side about what APIs we should use\nto provide durability across multiple Git files while not completely\ntanking performance by adding one CACHE_FLUSH per file modified?  What\nare the semantics of the ext4 log (when it is enabled) with regards to\ncreating a temp file, populating its contents and then renaming it?\nAre they similar enough to XFS's 'log force' such that our batch mode\nwould work there?\n\nThanks,\nNeeraj\nWindows Core Filesystem Dev\n\n[1] https://docs.google.com/spreadsheets/d/1uxMBkEXFFnQ1Y3lXKqcKpw6Mq44BzhpCAcPex14T-QQ/edit#gid=1898936117\n[2] https://lore.kernel.org/git/cover.1636544377.git.ps@pks.im/\n[3] https://lore.kernel.org/git/b9d3d87443266767f00e77c967bd77357fe50484.1633366667.git.gitgitgadget@gmail.com/\n[4] https://lore.kernel.org/linux-fsdevel/20190419072938.31320-1-amir73il@gmail.com/\n[5] See FLUSH_FLAGS_NO_SYNC -\nhttps://docs.microsoft.com/en-us/windows-hardware/drivers/ddi/ntifs/nf-ntifs-ntflushbuffersfileex\n"},{"id":"441528","messageId":"CANQDOdcdhfGtPg0PxpXQA5gQ4x9VknKDKCCi4HEB0Z1xgnjKzg@mail.gmail.com","threadId":"56870","inReplyTo":"211111.86pmr7pk9o.gmgdl@evledraar.gmail.com","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-17T22:16:35Z","receivedAt":"2021-11-17T22:16:52Z","isPatch":false,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Nov 10, 2021 at 5:13 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> On Wed, Nov 10 2021, Neeraj Singh wrote:\n>\n> > On Wed, Nov 10, 2021 at 04:09:33PM +0100, Ævar Arnfjörð Bjarmason wrote:\n> >> I think this sort of config schema would make everyone above happy\n> >>\n> >> It would:\n> >>\n> >>  A) Be easy to extend for any future fsync behavior we'd reasonably\n> >>     implement\n> >>\n> >>  B) Not make older git versions die. It's fine if they warn(), but not die.\n> >>\n> >>  C) Has some pretty contrived key names, but I'm trying to maintain the\n> >>     constraint that you can set both fsck.X=Y and\n> >>     e.g. fetch.fsck.X=Y. I.e. we should be able to configure things\n> >>     globally *and* per-command, like color.*, fsck.* etc.\n> >>\n> >> Proposal:\n> >>\n> >>   ; Turns on/off all fsync, whatever the method is. I.e. allows you to\n> >>   ; never make any fsync() calls whatsoever (which we have another\n> >>   ; in-flight topic for).\n> >>\n> >>   ; The \"false\" was controversial, and we could just leave it\n> >>   ; unimplemented\n> >>   core.fsync = <bool>\n> >>\n> >>   ; Optional, by default we'd use the most pedantic (I'd call our\n> >>   ; current \"loose\", whether we want to forward-support it is another\n> >>   ; matter.\n> >>   ;\n> >>   ; Whatever names we pick an option like this should ignore (or at most\n> >>   ; warn about) values it doesn't know about, not hard die on it.\n> >>   ;\n> >>   ; Here \"bach\" is what Neeraj and Patrick are pursuing, a hypothetical\n> >>   ; POSIX would be a pedantic way of exhaustively fsyncing everything.\n> >>   ;\n> >>   ; We'd leave door open to e.g. setting it to \"linux:ext4\" or whatever,\n> >>   ; to do only the work needed on some specific popular FS\n> >>   core.fsyncMethod = loose | POSIX | batch | linux:ext4 | NTFS | ...\n> >>\n> >>   ; Turn on or off entire categories of files we'd like to sync. This\n> >>   ; way Neeraj's and Patrick's approach would be to set\n> >>   ; core.fsyncMethod=batch, and then core.fsyncGroup=files &\n> >>   ; core.fsyncGroup=refs.\n> >>\n> >>   ; If we learn about a new core.fsyncGroup = xyz in the future a <bool>\n> >>   ; in \"core.fsyncGroupDefault\" will prevail. I.e. if true it's\n> >>   ; included, if false not.\n> >>   ;\n> >>   ; Whether \"false\" or \"true\" is the default depends on\n> >>   ; core.fsyncMethod. For POSIX it would be true, for \"loose\" it's\n> >>   ; false.\n> >>   core.fsyncGroup = files\n> >>   core.fsyncGroup = refs\n> >>   core.fsyncGroup = objects\n> >>\n> >> I'm not sure I like calling it \"group\". Maybe \"class\", \"category\"? Doing\n> >> it with this structure is extensible to the two-level keys, as noted\n> >> above.\n> >>\n> >>   ; Our existing config knob. When \"false\" synonymous with:\n> >>   ;\n> >>   ;     core.fsync = true\n> >>   ;     core.fsyncMethod = loose\n> >>   ;     core.fsyncGroup = pack\n> >>   ;\n> >>   ; When \"true\" synonymous with the same as the above, plus:\n> >>   ;     core.fsyncGroup = loose\n> >>   ;\n> >>   : Or something like that. I.e. we'll fsync *.pack, *.bitmap etc, and ;\n> >>   ; probably some other stuff, but not loose objects etc.\n> >>   ;\n> >>   ; Whatever we fsync now exactly this schema should be generic enough\n> >>   ; to support it.\n> >>   core.fsyncObjectFiles = <bool>\n> >>\n> >>   ; A namespace for core.fsyncMethod = <X>. Specific methods will\n> >>   ; own this namespace and can configure whatever they want.\n> >>   fsyncMethod.<x>.<a> = <b>\n> >>\n> >> E.g. we might have:\n> >>\n> >>   fsyncMethod.POSIX.content = true\n> >>   fsyncMethod.POSIX.metadata = false\n> >>\n> >> If we know we'd like to (depending on other config) to fsync things\n> >> exhaustively or not, but do different things depending on file content\n> >> or metadata. I.e. maybe your FS's fsync() on a file fd always implies a\n> >> sync of the metadata, and maybe not.\n> >>\n> >>   ; Change whatever fsync configuration you want per-command, similar to\n> >>   ; fsck.* and fetch.fsck.*\n> >>   transfer.fsyncGroup=*\n> >>   fetch.fsyncGroup=*\n> >>   ...\n> >>\n> >> 1. https://lore.kernel.org/git/211110.86v910gi9a.gmgdl@evledraar.gmail.com/\n> >> 2. https://lore.kernel.org/git/20211028002102.19384-1-e@80x24.org/\n> >> 3. https://lore.kernel.org/git/cover.1636544377.git.ps@pks.im/\n> > Hi Ævar,\n> >\n> > Thanks for noticing the backwards compatibility issue with the 'batch' flag. I\n> > agree that we need to fix that before committing my changes to master.\n> >\n> > I'm hoping that we can agree to a version of what you're proposing, but my\n> > preference would be to cut out the more granular controls. I'd prefer to see\n> > just:\n> >       core.fsync = [bool]             - Turn fsyncing on or off.\n> >       core.fsyncMethod = [string]     - Controls how it's done (with a non-fatal warn on unrecognized values).\n> >       core.fsyncObjectFiles = [bool]  - Sets core.fsync if that setting doesn't already have a value. For back-compat.\n>\n> I'm fine with something simpler as long as we don't think we'll\n> plausibly start painting ourselves into a corner.\n>\n> But core.fsyncObjectFiles is *not* a setting of a \"core.fsync\" in the\n> sense that Eric suggested we have.\n>\n> I.e. it's effectively a sort of early and partial Linux-only version of\n> what your \"batch\" mode is. I.e. to skip fsyncing the loose object files,\n> and only fsync() the eventual refs we write.\n>\n> \"Sort of\" because we'd e.g. fsync packs unconditionally etc, but if we\n> make core.fsyncObjectFiles=false be core.fsync=false then we can't have\n> a \"real\" core.fsync=false, i.e. one that guarantees no fsync() calls at\n> all.\n>\n> We could also simply decide that it's a bad setting and we're going to\n> deprecate it, but another way is having a generic config layout that can\n> express what it's doing and more.\n>\n> > I don't think either we or the users should have to reason about what it means\n> > for some parts of the repo to be fsynced and others not to be. If core.fsync is\n> > 'false' and someone gets a weird state after a system crash, no one should be\n> > surprised.\n>\n> Yes. I'm fine with leaving this on the table. I should have be more\n> explicit that I'm not suggesting we implement all this exhaustive config\n> support, but if we imagine a sensible config schema that is extensible\n> (my proposal may or may not be that) then we can implement just 1-2\n> variables in it and know that we have room to grow in the future.\n>\n> > If core.fsync is 'true', and people are running on a reasonable\n> > common filesystem, we should be trying to give decent performance and good\n> > durability.\n>\n> Yeah, I just wonder if we can easily provide config to have people\n> decide that trade-off themselves.\n>\n> E.g. from the performance numbers in [1] I might turn off fsyncing when\n> we write anything in the working tree.\n>\n> We don't do that particular thing now, but if we're being pulled in one\n> direction of always being fsync-safe by default...\n>\n> I can also see it being useful to e.g. do:\n>\n>     gc.fsync = false\n>\n> Or blacklist other similar batch operations, although with a global knob\n> that can also rather easily be:\n>\n>     git -c core.fsync gc\n>\n> So maybe the whole \"fsck\" rationale doesn't apply here.\n>\n> > It would be nice to loop in some Linux fs developers to find out what can be\n> > done on current implementations to get the durability without terrible\n> > performance. From reading the docs and mailing threads it looks like the\n> > sync_file_range + bulk fsync approach should actually work on the current XFS\n> > implementation.\n>\n\nAfter sleeping on it for a while, I'm willing to consolidate the\nconfiguration along the lines that you've specified, but I'd like to\nreduce the number of degrees of freedom.\n\nMy proposal in Documentation form:\n\ncore.fsync::\nA comma-separated list of parts of the repository which should be hardened by\ncalling fsync when created or modified. When an aggregate option is\nspecified, a subcomponent can be overriden by prefixing it with a '-'. For\nexample, `core.fsync=all,-index` means \"fsync everything except the index\".\nItems which are not fsync'ed may be lost in the even of an unclean system\nshutdown. This setting defaults to `objects,-loose-objects`\n+\n* `loose-objects` hardens objects added to the repo in loose-object form.\n* `packs` hardens objects added to the repo in packfile form and the related\n  bitmap and index files.\n* `commit-graph` hardens the commit graph file.\n* `refs` (future) hardens references when they are modified.\n* `index` (future) hardens the index when it is modified.\n* `objects` is an aggregate option that includes `loose-objects`, `packs`, and\n  `commit-graph`.\n* `all` is an aggregate option that syncs all individual components above.\n* `none` is an aggregate option that disables fsync completely.\n\ncore.fsyncMethod::\nA value indicating the strategy Git will use to harden repository data using\nfsync and related primitives.\n+\n* 'default' uses the fsync(2) system call or platform equivalents.\n* 'batch' uses APIs such as sync_file_range or equivalent to reduce the number\n  of hardware FLUSH CACHE requests sent to the storage hardware.\n* 'writeout-only' (future) issues requests to send the writes to the storage\n* hardware, but does not send any FLUSH CACHE request.\n* 'syncfs' (future) uses the syncfs API, where available, to sync all of the\n  files on the same filesystem as the Git repo.\n\ncore.fsyncObjectFiles::\nIf `true`, this legacy setting is equivalent to `core.fsync=objects`. If\n`core.fsync` is explicitly specified, then this setting is ignored.\n\nThanks,\nNeeraj\n"},{"id":"441651","messageId":"xmqq35ntxp9y.fsf@gitster.g","threadId":"56870","inReplyTo":"CANQDOdcdhfGtPg0PxpXQA5gQ4x9VknKDKCCi4HEB0Z1xgnjKzg@mail.gmail.com","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-11-18T19:00:25Z","receivedAt":"2021-11-18T19:00:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> After sleeping on it for a while, I'm willing to consolidate the\n> configuration along the lines that you've specified, but I'd like to\n> reduce the number of degrees of freedom.\n>\n> My proposal in Documentation form:\n>\n> core.fsync::\n> A comma-separated list of parts of the repository which should be hardened by\n> calling fsync when created or modified. When an aggregate option is\n> specified, a subcomponent can be overriden by prefixing it with a '-'. For\n> example, `core.fsync=all,-index` means \"fsync everything except the index\".\n> Items which are not fsync'ed may be lost in the even of an unclean system\n> shutdown. This setting defaults to `objects,-loose-objects`\n> +\n> * `loose-objects` hardens objects added to the repo in loose-object form.\n> * `packs` hardens objects added to the repo in packfile form and the related\n>   bitmap and index files.\n> * `commit-graph` hardens the commit graph file.\n> * `refs` (future) hardens references when they are modified.\n> * `index` (future) hardens the index when it is modified.\n> * `objects` is an aggregate option that includes `loose-objects`, `packs`, and\n>   `commit-graph`.\n> * `all` is an aggregate option that syncs all individual components above.\n> * `none` is an aggregate option that disables fsync completely.\n\nI wasn't closely following the discussion at all, but the above\nsimplification may still even be too fine-grained?  For example,\nwhat does it mean to care less about the robustness of loose objects\nthan packs or ref updates?  How does an existing fine-grained\nclassification interact with new classes of filesystem entity we\nwill introduce under .git in the future?  Imagine that we didn't\nhave .midx and multi-pack bitmap yet; since 'loose-objects',\n'packs', and 'commit-graph' are the only three groups we can choose\nto place any \"objects and reachability\" related data in, we need to\npick one, and choosing 'packs' class may be the choice of least\nresistance, the default kitchen-sync category for anything related\nto \"object\".  Or just like 'commit-graph' has its own category,\nwould we invent a new class and call it 'multi-pack'?\n\nI cannot shake the feeling that these are making everything\nunnecessarily complex and adding more things that we need to explain\nto the end-user---and the worst part is I doubt it would help the\nend-users very much tot understand what gets explained.\n\n> core.fsyncMethod::\n> A value indicating the strategy Git will use to harden repository data using\n> fsync and related primitives.\n> +\n> * 'default' uses the fsync(2) system call or platform equivalents.\n> * 'batch' uses APIs such as sync_file_range or equivalent to reduce the number\n>   of hardware FLUSH CACHE requests sent to the storage hardware.\n> * 'writeout-only' (future) issues requests to send the writes to the storage\n> * hardware, but does not send any FLUSH CACHE request.\n> * 'syncfs' (future) uses the syncfs API, where available, to sync all of the\n>   files on the same filesystem as the Git repo.\n\nHow would an end-user choose among these?  If they assume that the\nversion of Git they use is bug-free, is there a reason why they\nshould ever pick 'default' over 'batch', for example?  Shouldn't we\nbe the one to choose the best approach on the underlying filesystem\nfor the users, instead of forcing them to choose?\n\nAs implementors, these choices may be of interest and give you a\nhandy way to compare different design, but I am not sure if we want\nto give anything more complex than a binary choice, \"default\" and\n\"eatmydata\".\n\n> core.fsyncObjectFiles::\n> If `true`, this legacy setting is equivalent to `core.fsync=objects`. If\n> `core.fsync` is explicitly specified, then this setting is ignored.\n\nI think deprecating this very-specific knob is a good idea,\nregardless of how complex we'd want to make the alternative.\n\nThanks.\n"},{"id":"441654","messageId":"CANQDOdc+j0PgRJw0bzTLoAnSq=taabXSE4r9jrYczNaBHX9XuQ@mail.gmail.com","threadId":"56870","inReplyTo":"xmqq35ntxp9y.fsf@gitster.g","subject":"Re: RFC: A configuration design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-18T19:46:59Z","receivedAt":"2021-11-18T19:47:14Z","isPatch":false,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Nov 18, 2021 at 11:00 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > After sleeping on it for a while, I'm willing to consolidate the\n> > configuration along the lines that you've specified, but I'd like to\n> > reduce the number of degrees of freedom.\n> >\n> > My proposal in Documentation form:\n> >\n> > core.fsync::\n> > A comma-separated list of parts of the repository which should be hardened by\n> > calling fsync when created or modified. When an aggregate option is\n> > specified, a subcomponent can be overriden by prefixing it with a '-'. For\n> > example, `core.fsync=all,-index` means \"fsync everything except the index\".\n> > Items which are not fsync'ed may be lost in the even of an unclean system\n> > shutdown. This setting defaults to `objects,-loose-objects`\n> > +\n> > * `loose-objects` hardens objects added to the repo in loose-object form.\n> > * `packs` hardens objects added to the repo in packfile form and the related\n> >   bitmap and index files.\n> > * `commit-graph` hardens the commit graph file.\n> > * `refs` (future) hardens references when they are modified.\n> > * `index` (future) hardens the index when it is modified.\n> > * `objects` is an aggregate option that includes `loose-objects`, `packs`, and\n> >   `commit-graph`.\n> > * `all` is an aggregate option that syncs all individual components above.\n> > * `none` is an aggregate option that disables fsync completely.\n>\n> I wasn't closely following the discussion at all, but the above\n> simplification may still even be too fine-grained?  For example,\n> what does it mean to care less about the robustness of loose objects\n> than packs or ref updates?  How does an existing fine-grained\n> classification interact with new classes of filesystem entity we\n> will introduce under .git in the future?  Imagine that we didn't\n> have .midx and multi-pack bitmap yet; since 'loose-objects',\n> 'packs', and 'commit-graph' are the only three groups we can choose\n> to place any \"objects and reachability\" related data in, we need to\n> pick one, and choosing 'packs' class may be the choice of least\n> resistance, the default kitchen-sync category for anything related\n> to \"object\".  Or just like 'commit-graph' has its own category,\n> would we invent a new class and call it 'multi-pack'?\n>\n> I cannot shake the feeling that these are making everything\n> unnecessarily complex and adding more things that we need to explain\n> to the end-user---and the worst part is I doubt it would help the\n> end-users very much tot understand what gets explained.\n I agree with you that this is fairly complex. I did two things to\ncome up with this specific list:\n1) Looked at what we're fsyncing today (most of these items are being\nsynced through the CSUM_FSYNC flag).  Some of these things should\nperhaps not be fsynced if they are derived metadata that can be\nreconstructed easily.  For instance, can the commit-graph file be\nrecomputed easily enough from the ODB and the refs? How hard is it to\nreconstruct the pack-indexes? Maybe they should get their own item.\n2) Thought about what's necessary to be able to retrieve data out of\ngit after a series of commands followed by a system crash.  If I fetch\na repo, modify some worktree files, add the modifications, and then\ncommit them, what does Git need to persist to not lose unique work?\nWe need to sync the packfiles since they form the base of the commit\ngraph and trees and it may be difficult to construct a sane repo if we\nhave objects without their dependencies. We obviously need to sync the\nnew objects the user is adding. We need to sync the index after 'add'\nin case we crash between 'add' and 'commit', so the user can find\ntheir new objects.  Lastly we need to sync the refs after a commit so\nthat the user can find the added objects after switching branches.\n\nI also wanted to make sure we could express the current state of\nfsyncing so that users could go back to that if the cost of syncing\nsome particular thing is too high in their workload.\n\nI expect that people should really specify `objects` to sync all of\nthe things that comprise the object store and leave it to us to decide\nwhich subcomponents are considered derived metadata.\n\n>\n> > core.fsyncMethod::\n> > A value indicating the strategy Git will use to harden repository data using\n> > fsync and related primitives.\n> > +\n> > * 'default' uses the fsync(2) system call or platform equivalents.\n> > * 'batch' uses APIs such as sync_file_range or equivalent to reduce the number\n> >   of hardware FLUSH CACHE requests sent to the storage hardware.\n> > * 'writeout-only' (future) issues requests to send the writes to the storage\n> > * hardware, but does not send any FLUSH CACHE request.\n> > * 'syncfs' (future) uses the syncfs API, where available, to sync all of the\n> >   files on the same filesystem as the Git repo.\n>\n> How would an end-user choose among these?  If they assume that the\n> version of Git they use is bug-free, is there a reason why they\n> should ever pick 'default' over 'batch', for example?  Shouldn't we\n> be the one to choose the best approach on the underlying filesystem\n> for the users, instead of forcing them to choose?\n>\n> As implementors, these choices may be of interest and give you a\n> handy way to compare different design, but I am not sure if we want\n> to give anything more complex than a binary choice, \"default\" and\n> \"eatmydata\".\n\nMaybe `default` should be renamed `fsync` to indicate the specific\naction to be performed and no value can be \"let git decide\".  I think\nthe \"eatmydata\" option would be core.fsync=none and core.fsyncMethod\nwould then be ignored.  One reason to have this option would be to\nallow distributors and organizations to choose a good config based on\nthe actual filesystem them deploy on.  On Windows, it would be clear\nthat we should use 'batch' because we're committed to making sure that\nsetting is actually safe on our filesystems (we'll fix bugs in the FS\nif we find out that people are reporting corrupt repos).  Given that\nthe Linux and POSIX durability situation is really murky, it's hard to\nsee how Git can give any useful guarantee on those platforms.\n\n>\n> > core.fsyncObjectFiles::\n> > If `true`, this legacy setting is equivalent to `core.fsync=objects`. If\n> > `core.fsync` is explicitly specified, then this setting is ignored.\n>\n> I think deprecating this very-specific knob is a good idea,\n> regardless of how complex we'd want to make the alternative.\n>\n\nGlad you agree.  I hope we see others weigh in on the tradeoff between\ncomplexity and control.\n\nThanks,\nNeeraj\n"}]}