{"thread":{"id":"57030","subject":"[PATCH 0/2] A design for future-proofing fsync() configuration","startedAt":"2021-12-04T03:28:29Z","lastAt":"2022-03-30T16:59:55Z","messageCount":122,"participants":["Neeraj K. Singh via GitGitGadget","Neeraj Singh via GitGitGadget","Neeraj Singh","Patrick Steinhardt","Ævar Arnfjörð Bjarmason","Junio C Hamano","rsbecker@nexbridge.com","brian m. carlson","Johannes Schindelin","SZEDER Gábor","Jiang Xin"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"443060","messageId":"pull.1093.git.1638588503.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":null,"subject":"[PATCH 0/2] A design for future-proofing fsync() configuration","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-04T03:28:21Z","receivedAt":"2021-12-04T03:28:29Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"This is an implementation of an extensible configuration mechanism for\nfsyncing persistent components of a repo.\n\nThe main goals are to separate the \"what\" to sync from the \"how\". There are\nnow two settings: core.fsync - Control the 'what', including the index.\ncore.fsyncMethod - Control the 'how'. Currently we support writeout-only and\nfull fsync.\n\nSyncing of refs can be layered on top of core.fsync. And batch mode will be\nlayered on core.fsyncMethod.\n\ncore.fsyncobjectfiles is removed and will issue a deprecation warning if\nit's seen.\n\nI'd like to get agreement on this direction before redoing batch mode on top\nof this change.\n\nAs of this writing, the change isn't tested in detail. I'll be making sure\nwith strace that each component has the desired effect.\n\nPlease see [1], [2], and [3] for discussions that led to this series.\n\n[1] https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n[2]\nhttps://lore.kernel.org/git/dd65718814011eb93ccc4428f9882e0f025224a6.1636029491.git.ps@pks.im/\n[3]\nhttps://lore.kernel.org/git/pull.1076.git.git.1629856292.gitgitgadget@gmail.com/\n\nNeeraj Singh (2):\n  fsync: add writeout-only mode for fsyncing repo data\n  core.fsync: introduce granular fsync control\n\n Documentation/config/core.txt       | 35 +++++++++---\n builtin/fast-import.c               |  2 +-\n builtin/index-pack.c                |  4 +-\n builtin/pack-objects.c              |  8 ++-\n bulk-checkin.c                      |  5 +-\n cache.h                             | 48 +++++++++++++++-\n commit-graph.c                      |  3 +-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n config.c                            | 89 ++++++++++++++++++++++++++++-\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n csum-file.c                         |  5 +-\n csum-file.h                         |  2 +-\n environment.c                       |  3 +-\n git-compat-util.h                   | 24 ++++++++\n midx.c                              |  3 +-\n object-file.c                       |  3 +-\n pack-bitmap-write.c                 |  3 +-\n pack-write.c                        | 12 ++--\n read-cache.c                        | 19 ++++--\n wrapper.c                           | 56 ++++++++++++++++++\n write-or-die.c                      | 10 ++--\n 24 files changed, 337 insertions(+), 42 deletions(-)\n create mode 100644 compat/win32/flush.c\n\n\nbase-commit: abe6bb3905392d5eb6b01fa6e54d7e784e0522aa\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1093%2Fneerajsi-msft%2Fns%2Fcore-fsync-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1093/neerajsi-msft/ns/core-fsync-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/1093\n-- \ngitgitgadget\n"},{"id":"443062","messageId":"527380ddc3fe8b2fac8c8512de8fcdee6c96e65a.1638588503.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.git.1638588503.gitgitgadget@gmail.com","subject":"[PATCH 1/2] fsync: add writeout-only mode for fsyncing repo data","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-04T03:28:22Z","receivedAt":"2021-12-04T03:28:31Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe new writeout-only mode attempts to tell the operating system to\nflush its in-memory page cache to the storage hardware without issuing a\nCACHE_FLUSH command to the storage controller.\n\nWriteout-only fsync is significantly faster than a vanilla fsync on\ncommon hardware, since data is written to a disk-side cache rather than\nall the way to a durable medium. Later changes in this patch series will\ntake advantage of this primitive to implement batching of hardware\nflushes.\n\nWhen git_fsync is called with FSYNC_WRITEOUT_ONLY, it may fail and the\ncaller is expected to do an ordinary fsync as needed.\n\nOn Apple platforms, the fsync system call does not issue a CACHE_FLUSH\ndirective to the storage controller. This change updates fsync to do\nfcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\nwith existing behavior on Apple platforms by setting the default value\nof the new core.fsyncmethod option.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt       |  9 +++++\n cache.h                             |  7 ++++\n compat/mingw.h                      |  3 ++\n compat/win32/flush.c                | 28 +++++++++++++++\n config.c                            | 12 +++++++\n config.mak.uname                    |  3 ++\n configure.ac                        |  8 +++++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n environment.c                       |  2 +-\n git-compat-util.h                   | 24 +++++++++++++\n wrapper.c                           | 56 +++++++++++++++++++++++++++++\n write-or-die.c                      | 10 +++---\n 12 files changed, 159 insertions(+), 6 deletions(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..c91eccea598 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,15 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsyncMethod::\n+\tA value indicating the strategy Git will use to harden repository data\n+\tusing fsync and related primitives.\n++\n+* `fsync` uses the fsync() system call or platform equivalents.\n+* `writeout-only` issues requests to send the writes to the storage\n+  hardware, but does not send any FLUSH CACHE request. If the operating system\n+  does not support the required interfaces, this falls back to fsync().\n+\n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n +\ndiff --git a/cache.h b/cache.h\nindex eba12487b99..9cd60d94952 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -986,6 +986,13 @@ extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n extern int fsync_object_files;\n+\n+enum fsync_method {\n+\tFSYNC_METHOD_FSYNC,\n+\tFSYNC_METHOD_WRITEOUT_ONLY\n+};\n+\n+extern enum fsync_method fsync_method;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..75324c24ee7\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.c b/config.c\nindex c5873f3a706..c3410b8a868 100644\n--- a/config.c\n+++ b/config.c\n@@ -1490,6 +1490,18 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsyncmethod\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcmp(value, \"fsync\"))\n+\t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n+\t\telse if (!strcmp(value, \"writeout-only\"))\n+\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse\n+\t\t\twarning(_(\"unknown %s value '%s'\"), var, value);\n+\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\ndiff --git a/config.mak.uname b/config.mak.uname\nindex d0701f9beb0..774a09622d2 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -57,6 +57,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tBASIC_CFLAGS += -DHAVE_SYSINFO\n@@ -453,6 +454,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -628,6 +630,7 @@ ifeq ($(uname_S),MINGW)\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/configure.ac b/configure.ac\nindex 5ee25ec95c8..6bd6bef1c44 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1082,6 +1082,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 86b46114464..6d7bc16d054 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/environment.c b/environment.c\nindex 9da7f3c1a19..f9140e842cf 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -41,7 +41,7 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex c6bd2a84e55..cb9abd7a08c 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1239,6 +1239,30 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+#ifdef __APPLE__\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n+#else\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n+#endif\n+\n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+/*\n+ * Issues an fsync against the specified file according to the specified mode.\n+ *\n+ * FSYNC_WRITEOUT_ONLY attempts to use interfaces available on some operating\n+ * systems to flush the OS cache without issuing a flush command to the storage\n+ * controller. If those interfaces are unavailable, the function fails with\n+ * ENOSYS.\n+ *\n+ * FSYNC_HARDWARE_FLUSH does an OS writeout and hardware flush to ensure that\n+ * changes are durable. It is not expected to fail.\n+ */\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/wrapper.c b/wrapper.c\nindex 36e12119d76..1c5f2c87791 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -546,6 +546,62 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\t\t/*\n+\t\t * On some platforms fsync may return EINTR. Try again in this\n+\t\t * case, since callers asking for a hardware flush may die if\n+\t\t * this function returns an error.\n+\t\t */\n+\t\tfor (;;) {\n+\t\t\tint err;\n+#ifdef __APPLE__\n+\t\t\terr = fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\t\terr = fsync(fd);\n+#endif\n+\t\t\tif (err >= 0 || errno != EINTR)\n+\t\t\t\treturn err;\n+\t\t}\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 0b1ec8190b6..0702acdd5e8 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,10 +57,12 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n-\t\tif (errno != EINTR)\n-\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n-\t}\n+\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n+\t\treturn;\n+\n+\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n+\t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n void write_or_die(int fd, const void *buf, size_t count)\n-- \ngitgitgadget\n\n"},{"id":"443061","messageId":"23311a1014226b74a9d313552cd2de886db907a5.1638588503.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.git.1638588503.gitgitgadget@gmail.com","subject":"[PATCH 2/2] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-04T03:28:23Z","receivedAt":"2021-12-04T03:28:33Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsync` configuration\nknob which can be used to control how components of the\nrepository are made durable on disk.\n\nThis setting allows future extensibility of components\nthat could be synced in two ways:\n* We issue a warning rather than an error for unrecognized\n  components, so new configs can be used with old Git versions.\n* We support negation, so users can choose one of the default\n  aggregate options and then remove components that they don't\n  want.\n\nThis also support the common request of doing absolutely no\nfysncing with the `core.fsync=none` value, which is expected\nto make the test suite faster.\n\nThis commit introduces the new ability for the user to harden\nthe index, which is a requirement for being able to actually\nfind a file that has been added to the repo and then deleted\nfrom the working tree.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 28 +++++++++----\n builtin/fast-import.c         |  2 +-\n builtin/index-pack.c          |  4 +-\n builtin/pack-objects.c        |  8 ++--\n bulk-checkin.c                |  5 ++-\n cache.h                       | 41 ++++++++++++++++++-\n commit-graph.c                |  3 +-\n config.c                      | 77 ++++++++++++++++++++++++++++++++++-\n csum-file.c                   |  5 ++-\n csum-file.h                   |  2 +-\n environment.c                 |  1 +\n midx.c                        |  3 +-\n object-file.c                 |  3 +-\n pack-bitmap-write.c           |  3 +-\n pack-write.c                  | 12 +++---\n read-cache.c                  | 19 ++++++---\n 16 files changed, 179 insertions(+), 37 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c91eccea598..d502f8a1bf5 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,26 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsync::\n+\tA comma-separated list of parts of the repository which should be\n+\thardened via the core.fsyncMethod when created or modified. You can\n+\tdisable hardening of any component by prefixing it with a '-'. Later\n+\titems take precedence over earlier ones in the list. For example,\n+\t`core.fsync=all,-index` means \"harden everything except the index\".\n+\tItems that are not hardened may be lost in the event of an unclean\n+\tsystem shutdown.\n++\n+* `none` disables fsync completely. This must be specified alone.\n+* `loose-object` hardens objects added to the repo in loose-object form.\n+* `pack` hardens objects added to the repo in packfile form.\n+* `pack-metadata` hardens packfile bitmaps and indexes.\n+* `commit-graph` hardens the commit graph file.\n+* `index` hardens the index when it is modified.\n+* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n+  `pack-metadata`, and `commit-graph`.\n+* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n+* `all` is an aggregate option that syncs all individual components above.\n+\n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\n@@ -556,14 +576,6 @@ core.fsyncMethod::\n   hardware, but does not send any FLUSH CACHE request. If the operating system\n   does not support the required interfaces, this falls back to fsync().\n \n-core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n-\n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\n +\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 20406f67754..0ae17d63618 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -856,7 +856,7 @@ static void end_packfile(void)\n \t\tstruct tag *t;\n \n \t\tclose_pack_windows(pack_data);\n-\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, REPO_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n \t\t\t\t\t pack_data->pack_name, object_count,\n \t\t\t\t\t cur_pack_oid.hash, pack_size);\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c23d01de7dc..9157a955de7 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1286,7 +1286,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\t\t    nr_objects - nr_objects_initial);\n \t\tstop_progress_msg(&progress, msg.buf);\n \t\tstrbuf_release(&msg);\n-\t\tfinalize_hashfile(f, tail_hash, 0);\n+\t\tfinalize_hashfile(f, tail_hash, REPO_COMPONENT_PACK, 0);\n \t\thashcpy(read_hash, pack_hash);\n \t\tfixup_pack_header_footer(output_fd, pack_hash,\n \t\t\t\t\t curr_pack, nr_objects,\n@@ -1508,7 +1508,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tif (!from_stdin) {\n \t\tclose(input_fd);\n \t} else {\n-\t\tfsync_or_die(output_fd, curr_pack_name);\n+\t\tfsync_component_or_die(REPO_COMPONENT_PACK, output_fd, curr_pack_name);\n \t\terr = close(output_fd);\n \t\tif (err)\n \t\t\tdie_errno(_(\"error while closing pack file\"));\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 857be7826f3..48c2f9f3847 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n \t\t * If so, rewrite it like in fast-import\n \t\t */\n \t\tif (pack_to_stdout) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, REPO_COMPONENT_NONE,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n \t\t} else if (nr_written == nr_remaining) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, REPO_COMPONENT_PACK,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t\t} else {\n-\t\t\tint fd = finalize_hashfile(f, hash, 0);\n+\t\t\tint fd = finalize_hashfile(f, hash, REPO_COMPONENT_PACK, 0);\n \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n \t\t\t\t\t\t nr_written, hash, offset);\n \t\t\tclose(fd);\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..b9f3d315334 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -53,9 +53,10 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \t\tunlink(state->pack_tmp_name);\n \t\tgoto clear_exit;\n \t} else if (state->nr_written == 1) {\n-\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\tfinalize_hashfile(state->f, hash, REPO_COMPONENT_PACK,\n+\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t} else {\n-\t\tint fd = finalize_hashfile(state->f, hash, 0);\n+\t\tint fd = finalize_hashfile(state->f, hash, REPO_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n \t\t\t\t\t state->nr_written, hash,\n \t\t\t\t\t state->offset);\ndiff --git a/cache.h b/cache.h\nindex 9cd60d94952..b2966352440 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,40 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+/*\n+ * These values are used to help identify parts of a repository to fsync.\n+ * REPO_COMPONENT_NONE identifies data that will not be a persistent part of the\n+ * repository and so shouldn't be fsynced.\n+ */\n+enum repo_component {\n+\tREPO_COMPONENT_NONE\t\t\t= 0,\n+\tREPO_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n+\tREPO_COMPONENT_PACK\t\t\t= 1 << 1,\n+\tREPO_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n+\tREPO_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+\tREPO_COMPONENT_INDEX\t\t\t= 1 << 4,\n+};\n+\n+#define FSYNC_COMPONENTS_DEFAULT (REPO_COMPONENT_PACK | \\\n+\t\t\t\t  REPO_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  REPO_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_OBJECTS (REPO_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t\t  REPO_COMPONENT_PACK | \\\n+\t\t\t\t  REPO_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  REPO_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_ALL (REPO_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t      REPO_COMPONENT_PACK | \\\n+\t\t\t      REPO_COMPONENT_PACK_METADATA | \\\n+\t\t\t      REPO_COMPONENT_COMMIT_GRAPH | \\\n+\t\t\t      REPO_COMPONENT_INDEX)\n+\n+\n+/*\n+ * A bitmask indicating which components of the repo should be fsynced.\n+ */\n+extern enum repo_component fsync_components;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n@@ -1747,6 +1780,12 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n \n+inline void fsync_component_or_die(enum repo_component component, int fd, const char *msg)\n+{\n+\tif (fsync_components & component)\n+\t\tfsync_or_die(fd, msg);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 2706683acfe..4bed4175ab2 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1939,7 +1939,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t}\n \n \tclose_commit_graph(ctx->r->objects);\n-\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n+\tfinalize_hashfile(f, file_hash, REPO_COMPONENT_COMMIT_GRAPH,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n \tfree_chunkfile(cf);\n \n \tif (ctx->split) {\ndiff --git a/config.c b/config.c\nindex c3410b8a868..6c8b102ed7a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1213,6 +1213,74 @@ static int git_parse_maybe_bool_text(const char *value)\n \treturn -1;\n }\n \n+static const struct fsync_component_entry {\n+\tconst char *name;\n+\tenum repo_component component_bits;\n+} fsync_component_table[] = {\n+\t{ \"loose-object\", REPO_COMPONENT_LOOSE_OBJECT },\n+\t{ \"pack\", REPO_COMPONENT_PACK },\n+\t{ \"pack-metadata\", REPO_COMPONENT_PACK_METADATA },\n+\t{ \"commit-graph\", REPO_COMPONENT_COMMIT_GRAPH },\n+\t{ \"index\", REPO_COMPONENT_INDEX },\n+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n+};\n+\n+static enum repo_component parse_fsync_components(const char *var, const char *string)\n+{\n+\tenum repo_component output = 0;\n+\n+\tif (!strcmp(string, \"none\"))\n+\t\treturn output;\n+\n+\twhile (string) {\n+\t\tint i;\n+\t\tsize_t len;\n+\t\tconst char *ep;\n+\t\tint negated = 0;\n+\t\tint found = 0;\n+\n+\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n+\t\tep = strchrnul(string, ',');\n+\t\tlen = ep - string;\n+\n+\t\tif (*string == '-') {\n+\t\t\tnegated = 1;\n+\t\t\tstring++;\n+\t\t\tlen--;\n+\t\t\tif (!len)\n+\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n+\t\t}\n+\n+\t\tif (!len)\n+\t\t\tbreak;\n+\n+\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n+\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n+\n+\t\t\tif (strncmp(entry->name, string, len))\n+\t\t\t\tcontinue;\n+\n+\t\t\tfound = 1;\n+\t\t\tif (negated)\n+\t\t\t\toutput &= ~entry->component_bits;\n+\t\t\telse\n+\t\t\t\toutput |= entry->component_bits;\n+\t\t}\n+\n+\t\tif (!found) {\n+\t\t\tchar *component = xstrndup(string, len);\n+\t\t\twarning(_(\"unknown %s value '%s'\"), var, component);\n+\t\t\tfree(component);\n+\t\t}\n+\n+\t\tstring = ep;\n+\t}\n+\n+\treturn output;\n+}\n+\n int git_parse_maybe_bool(const char *value)\n {\n \tint v = git_parse_maybe_bool_text(value);\n@@ -1490,6 +1558,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsync\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tfsync_components = parse_fsync_components(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncmethod\")) {\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n@@ -1503,7 +1578,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n \t\treturn 0;\n \t}\n \ndiff --git a/csum-file.c b/csum-file.c\nindex 26e8a6df44e..adc8023d5af 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -58,7 +58,8 @@ static void free_hashfile(struct hashfile *f)\n \tfree(f);\n }\n \n-int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n+int finalize_hashfile(struct hashfile *f, unsigned char *result,\n+\t\t      enum repo_component component, unsigned int flags)\n {\n \tint fd;\n \n@@ -69,7 +70,7 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int fl\n \tif (flags & CSUM_HASH_IN_STREAM)\n \t\tflush(f, f->buffer, the_hash_algo->rawsz);\n \tif (flags & CSUM_FSYNC)\n-\t\tfsync_or_die(f->fd, f->name);\n+\t\tfsync_component_or_die(component, f->fd, f->name);\n \tif (flags & CSUM_CLOSE) {\n \t\tif (close(f->fd))\n \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\ndiff --git a/csum-file.h b/csum-file.h\nindex 291215b34eb..3820b4a0e94 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -38,7 +38,7 @@ int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n struct hashfile *hashfd(int fd, const char *name);\n struct hashfile *hashfd_check(const char *name);\n struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n-int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n+int finalize_hashfile(struct hashfile *, unsigned char *, enum repo_component, unsigned int);\n void hashwrite(struct hashfile *, const void *, unsigned int);\n void hashflush(struct hashfile *f);\n void crc32_begin(struct hashfile *);\ndiff --git a/environment.c b/environment.c\nindex f9140e842cf..190df463475 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,6 +42,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n+enum repo_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/midx.c b/midx.c\nindex 837b46b2af5..6e9510ab0dc 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -1406,7 +1406,8 @@ static int write_midx_internal(const char *object_dir,\n \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n \twrite_chunkfile(cf, &ctx);\n \n-\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, midx_hash, REPO_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n \tfree_chunkfile(cf);\n \n \tif (flags & (MIDX_WRITE_REV_INDEX | MIDX_WRITE_BITMAP))\ndiff --git a/object-file.c b/object-file.c\nindex eb972cdccd2..871ea6bcd53 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1809,8 +1809,7 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\tfsync_component_or_die(REPO_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex 9c55c1531e1..e82e87af996 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -719,7 +719,8 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \tif (options & BITMAP_OPT_HASH_CACHE)\n \t\twrite_hash_cache(f, index, index_nr);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \n \tif (adjust_shared_perm(tmp_file.buf))\n \t\tdie_errno(\"unable to make temporary bitmap file readable\");\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a346..d9c37803e98 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -159,8 +159,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t}\n \n \thashwrite(f, sha1, the_hash_algo->rawsz);\n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY)\n \t\t\t\t    ? 0 : CSUM_FSYNC));\n \treturn index_name;\n }\n@@ -281,8 +282,9 @@ const char *write_rev_file_order(const char *rev_name,\n \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n \t\tdie(_(\"failed to make %s readable\"), rev_name);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \n \treturn rev_name;\n }\n@@ -390,7 +392,7 @@ void fixup_pack_header_footer(int pack_fd,\n \t\tthe_hash_algo->final_fn(partial_pack_hash, &old_hash_ctx);\n \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n-\tfsync_or_die(pack_fd, pack_name);\n+\tfsync_component_or_die(REPO_COMPONENT_PACK, pack_fd, pack_name);\n }\n \n char *index_pack_lockfile(int ip_out, int *is_well_formed)\ndiff --git a/read-cache.c b/read-cache.c\nindex f3986596623..883d0c0019a 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2816,7 +2816,7 @@ static int record_ieot(void)\n  * rely on it.\n  */\n static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n-\t\t\t  int strip_extensions)\n+\t\t\t  int strip_extensions, unsigned flags)\n {\n \tuint64_t start = getnanotime();\n \tstruct hashfile *f;\n@@ -2830,6 +2830,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint csum_fsync_flag;\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n@@ -3060,7 +3061,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n+\tcsum_fsync_flag = 0;\n+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n+\t\tcsum_fsync_flag = CSUM_FSYNC;\n+\n+\tfinalize_hashfile(f, istate->oid.hash, REPO_COMPONENT_INDEX,\n+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n+\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n@@ -3115,7 +3122,7 @@ static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n \t */\n \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n-\tret = do_write_index(istate, lock->tempfile, 0);\n+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n \n@@ -3209,7 +3216,7 @@ static int clean_shared_index_files(const char *current_hex)\n }\n \n static int write_shared_index(struct index_state *istate,\n-\t\t\t      struct tempfile **temp)\n+\t\t\t      struct tempfile **temp, unsigned flags)\n {\n \tstruct split_index *si = istate->split_index;\n \tint ret, was_full = !istate->sparse_index;\n@@ -3219,7 +3226,7 @@ static int write_shared_index(struct index_state *istate,\n \n \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n-\tret = do_write_index(si->base, *temp, 1);\n+\tret = do_write_index(si->base, *temp, 1, flags);\n \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n \n@@ -3328,7 +3335,7 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n \t\t\tret = do_write_locked_index(istate, lock, flags);\n \t\t\tgoto out;\n \t\t}\n-\t\tret = write_shared_index(istate, &temp);\n+\t\tret = write_shared_index(istate, &temp, flags);\n \n \t\tsaved_errno = errno;\n \t\tif (is_tempfile_active(temp))\n-- \ngitgitgadget\n"},{"id":"443141","messageId":"20211206075453.GA26639@neerajsi-x1.localdomain","threadId":"57030","inReplyTo":"527380ddc3fe8b2fac8c8512de8fcdee6c96e65a.1638588503.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/2] fsync: add writeout-only mode for fsyncing repo data","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-06T07:54:53Z","receivedAt":"2021-12-06T07:54:59Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"Context for reviewers: The mechanism part of this patch is equivalent to [1]\nand [2]. With the following changes:\n\n* The configuration code is a bit different and we now expose a\n  `writeout-only` mode to the user. This mode is the default on macOS\n  to prevent a change in end-user behavior.\n\n* git_fsync now contains the EINTR retry loop internally.\n\n[1] https://lore.kernel.org/git/e1747ce00af7ab3170a69955b07d995d5321d6f3.1637020263.git.gitgitgadget@gmail.com/\n[2] https://lore.kernel.org/git/546ad9c82e8e0c2eb4683f9f360d8f30e2136020.1630108177.git.gitgitgadget@gmail.com/\n\n"},{"id":"443222","messageId":"pull.1093.v2.git.1638845211.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.git.1638588503.gitgitgadget@gmail.com","subject":"[PATCH v2 0/3] A design for future-proofing fsync() configuration","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-07T02:46:48Z","receivedAt":"2021-12-07T02:46:56Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"This is an implementation of an extensible configuration mechanism for\nfsyncing persistent components of a repo.\n\nThe main goals are to separate the \"what\" to sync from the \"how\". There are\nnow two settings: core.fsync - Control the 'what', including the index.\ncore.fsyncMethod - Control the 'how'. Currently we support writeout-only and\nfull fsync.\n\nSyncing of refs can be layered on top of core.fsync. And batch mode will be\nlayered on core.fsyncMethod.\n\ncore.fsyncobjectfiles is removed and will issue a deprecation warning if\nit's seen.\n\nI'd like to get agreement on this direction before submitting batch mode to\nthe list. The batch mode series is available to view at\nhttps://github.com/neerajsi-msft/git/pull/1.\n\nPlease see [1], [2], and [3] for discussions that led to this series.\n\nV2 changes:\n\n * Updated the documentation for core.fsyncmethod to be less certain.\n   writeout-only probably does not do the right thing on Linux.\n * Split out the core.fsync=index change into its own commit.\n * Rename REPO_COMPONENT to FSYNC_COMPONENT. This is really specific to\n   fsyncing, so the name should reflect that.\n * Re-add missing Makefile change for SYNC_FILE_RANGE.\n * Tested writeout-only mode, index syncing, and general config settings.\n\n[1] https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n[2]\nhttps://lore.kernel.org/git/dd65718814011eb93ccc4428f9882e0f025224a6.1636029491.git.ps@pks.im/\n[3]\nhttps://lore.kernel.org/git/pull.1076.git.git.1629856292.gitgitgadget@gmail.com/\n\nNeeraj Singh (3):\n  core.fsyncmethod: add writeout-only mode\n  core.fsync: introduce granular fsync control\n  core.fsync: new option to harden the index\n\n Documentation/config/core.txt       | 35 +++++++++---\n Makefile                            |  6 ++\n builtin/fast-import.c               |  2 +-\n builtin/index-pack.c                |  4 +-\n builtin/pack-objects.c              |  8 ++-\n bulk-checkin.c                      |  5 +-\n cache.h                             | 48 +++++++++++++++-\n commit-graph.c                      |  3 +-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n config.c                            | 89 ++++++++++++++++++++++++++++-\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n csum-file.c                         |  5 +-\n csum-file.h                         |  3 +-\n environment.c                       |  3 +-\n git-compat-util.h                   | 24 ++++++++\n midx.c                              |  3 +-\n object-file.c                       |  3 +-\n pack-bitmap-write.c                 |  3 +-\n pack-write.c                        | 13 +++--\n read-cache.c                        | 19 ++++--\n wrapper.c                           | 56 ++++++++++++++++++\n write-or-die.c                      | 10 ++--\n 25 files changed, 344 insertions(+), 43 deletions(-)\n create mode 100644 compat/win32/flush.c\n\n\nbase-commit: abe6bb3905392d5eb6b01fa6e54d7e784e0522aa\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1093%2Fneerajsi-msft%2Fns%2Fcore-fsync-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1093/neerajsi-msft/ns/core-fsync-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/1093\n\nRange-diff vs v1:\n\n 1:  527380ddc3f ! 1:  e79522cbdd4 fsync: add writeout-only mode for fsyncing repo data\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    fsync: add writeout-only mode for fsyncing repo data\n     +    core.fsyncmethod: add writeout-only mode\n     +\n     +    This commit introduces the `core.fsyncmethod` configuration\n     +    knob, which can currently be set to `fsync` or `writeout-only`.\n      \n          The new writeout-only mode attempts to tell the operating system to\n          flush its in-memory page cache to the storage hardware without issuing a\n     @@ Documentation/config/core.txt: core.whitespace::\n      +\tusing fsync and related primitives.\n      ++\n      +* `fsync` uses the fsync() system call or platform equivalents.\n     -+* `writeout-only` issues requests to send the writes to the storage\n     -+  hardware, but does not send any FLUSH CACHE request. If the operating system\n     -+  does not support the required interfaces, this falls back to fsync().\n     ++* `writeout-only` issues pagecache writeback requests, but depending on the\n     ++  filesystem and storage hardware, data added to the repository may not be\n     ++  durable in the event of a system crash. This is the default mode on macOS.\n      +\n       core.fsyncObjectFiles::\n       \tThis boolean will enable 'fsync()' when writing object files.\n       +\n      \n     + ## Makefile ##\n     +@@ Makefile: all::\n     + #\n     + # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n     + #\n     ++# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n     ++#\n     + # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n     + # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n     + #\n     +@@ Makefile: ifdef HAVE_CLOCK_MONOTONIC\n     + \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n     + endif\n     + \n     ++ifdef HAVE_SYNC_FILE_RANGE\n     ++\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n     ++endif\n     ++\n     + ifdef NEEDS_LIBRT\n     + \tEXTLIBS += -lrt\n     + endif\n     +\n       ## cache.h ##\n      @@ cache.h: extern int read_replace_refs;\n       extern char *git_replace_ref_base;\n 2:  23311a10142 ! 2:  ff80a94bf9a core.fsync: introduce granular fsync control\n     @@ Commit message\n          knob which can be used to control how components of the\n          repository are made durable on disk.\n      \n     -    This setting allows future extensibility of components\n     -    that could be synced in two ways:\n     +    This setting allows future extensibility of the list of\n     +    syncable components:\n          * We issue a warning rather than an error for unrecognized\n            components, so new configs can be used with old Git versions.\n          * We support negation, so users can choose one of the default\n            aggregate options and then remove components that they don't\n     -      want.\n     +      want. The user would then harden any new components added in\n     +      a Git version update.\n      \n     -    This also support the common request of doing absolutely no\n     +    This also supports the common request of doing absolutely no\n          fysncing with the `core.fsync=none` value, which is expected\n          to make the test suite faster.\n      \n     -    This commit introduces the new ability for the user to harden\n     -    the index, which is a requirement for being able to actually\n     -    find a file that has been added to the repo and then deleted\n     -    from the working tree.\n     -\n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Documentation/config/core.txt ##\n     @@ Documentation/config/core.txt: core.whitespace::\n      +\thardened via the core.fsyncMethod when created or modified. You can\n      +\tdisable hardening of any component by prefixing it with a '-'. Later\n      +\titems take precedence over earlier ones in the list. For example,\n     -+\t`core.fsync=all,-index` means \"harden everything except the index\".\n     -+\tItems that are not hardened may be lost in the event of an unclean\n     -+\tsystem shutdown.\n     ++\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n     ++\tmetadata.\" Items that are not hardened may be lost in the event of an\n     ++\tunclean system shutdown.\n      ++\n      +* `none` disables fsync completely. This must be specified alone.\n      +* `loose-object` hardens objects added to the repo in loose-object form.\n      +* `pack` hardens objects added to the repo in packfile form.\n      +* `pack-metadata` hardens packfile bitmaps and indexes.\n      +* `commit-graph` hardens the commit graph file.\n     -+* `index` hardens the index when it is modified.\n      +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n      +  `pack-metadata`, and `commit-graph`.\n      +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n     @@ Documentation/config/core.txt: core.whitespace::\n       \tA value indicating the strategy Git will use to harden repository data\n       \tusing fsync and related primitives.\n      @@ Documentation/config/core.txt: core.fsyncMethod::\n     -   hardware, but does not send any FLUSH CACHE request. If the operating system\n     -   does not support the required interfaces, this falls back to fsync().\n     +   filesystem and storage hardware, data added to the repository may not be\n     +   durable in the event of a system crash. This is the default mode on macOS.\n       \n      -core.fsyncObjectFiles::\n      -\tThis boolean will enable 'fsync()' when writing object files.\n     @@ builtin/fast-import.c: static void end_packfile(void)\n       \n       \t\tclose_pack_windows(pack_data);\n      -\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n     -+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, REPO_COMPONENT_PACK, 0);\n     ++\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, FSYNC_COMPONENT_PACK, 0);\n       \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n       \t\t\t\t\t pack_data->pack_name, object_count,\n       \t\t\t\t\t cur_pack_oid.hash, pack_size);\n     @@ builtin/index-pack.c: static void conclude_pack(int fix_thin_pack, const char *c\n       \t\tstop_progress_msg(&progress, msg.buf);\n       \t\tstrbuf_release(&msg);\n      -\t\tfinalize_hashfile(f, tail_hash, 0);\n     -+\t\tfinalize_hashfile(f, tail_hash, REPO_COMPONENT_PACK, 0);\n     ++\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n       \t\thashcpy(read_hash, pack_hash);\n       \t\tfixup_pack_header_footer(output_fd, pack_hash,\n       \t\t\t\t\t curr_pack, nr_objects,\n     @@ builtin/index-pack.c: static void final(const char *final_pack_name, const char\n       \t\tclose(input_fd);\n       \t} else {\n      -\t\tfsync_or_die(output_fd, curr_pack_name);\n     -+\t\tfsync_component_or_die(REPO_COMPONENT_PACK, output_fd, curr_pack_name);\n     ++\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n       \t\terr = close(output_fd);\n       \t\tif (err)\n       \t\t\tdie_errno(_(\"error while closing pack file\"));\n     @@ builtin/pack-objects.c: static void write_pack_file(void)\n       \t\t */\n       \t\tif (pack_to_stdout) {\n      -\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n     -+\t\t\tfinalize_hashfile(f, hash, REPO_COMPONENT_NONE,\n     ++\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n      +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n       \t\t} else if (nr_written == nr_remaining) {\n      -\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n     -+\t\t\tfinalize_hashfile(f, hash, REPO_COMPONENT_PACK,\n     ++\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_PACK,\n      +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n       \t\t} else {\n      -\t\t\tint fd = finalize_hashfile(f, hash, 0);\n     -+\t\t\tint fd = finalize_hashfile(f, hash, REPO_COMPONENT_PACK, 0);\n     ++\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n       \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n       \t\t\t\t\t\t nr_written, hash, offset);\n       \t\t\tclose(fd);\n     @@ bulk-checkin.c: static void finish_bulk_checkin(struct bulk_checkin_state *state\n       \t\tgoto clear_exit;\n       \t} else if (state->nr_written == 1) {\n      -\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n     -+\t\tfinalize_hashfile(state->f, hash, REPO_COMPONENT_PACK,\n     ++\t\tfinalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK,\n      +\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n       \t} else {\n      -\t\tint fd = finalize_hashfile(state->f, hash, 0);\n     -+\t\tint fd = finalize_hashfile(state->f, hash, REPO_COMPONENT_PACK, 0);\n     ++\t\tint fd = finalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK, 0);\n       \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n       \t\t\t\t\t state->nr_written, hash,\n       \t\t\t\t\t state->offset);\n     @@ cache.h: void reset_shared_repository(void);\n      -extern int fsync_object_files;\n      +/*\n      + * These values are used to help identify parts of a repository to fsync.\n     -+ * REPO_COMPONENT_NONE identifies data that will not be a persistent part of the\n     ++ * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n      + * repository and so shouldn't be fsynced.\n      + */\n     -+enum repo_component {\n     -+\tREPO_COMPONENT_NONE\t\t\t= 0,\n     -+\tREPO_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n     -+\tREPO_COMPONENT_PACK\t\t\t= 1 << 1,\n     -+\tREPO_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n     -+\tREPO_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n     -+\tREPO_COMPONENT_INDEX\t\t\t= 1 << 4,\n     ++enum fsync_component {\n     ++\tFSYNC_COMPONENT_NONE\t\t\t= 0,\n     ++\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n     ++\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n     ++\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n     ++\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n      +};\n      +\n     -+#define FSYNC_COMPONENTS_DEFAULT (REPO_COMPONENT_PACK | \\\n     -+\t\t\t\t  REPO_COMPONENT_PACK_METADATA | \\\n     -+\t\t\t\t  REPO_COMPONENT_COMMIT_GRAPH)\n     ++#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n     ++\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n     ++\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n      +\n     -+#define FSYNC_COMPONENTS_OBJECTS (REPO_COMPONENT_LOOSE_OBJECT | \\\n     -+\t\t\t\t  REPO_COMPONENT_PACK | \\\n     -+\t\t\t\t  REPO_COMPONENT_PACK_METADATA | \\\n     -+\t\t\t\t  REPO_COMPONENT_COMMIT_GRAPH)\n     ++#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     ++\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n     ++\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n     ++\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n      +\n     -+#define FSYNC_COMPONENTS_ALL (REPO_COMPONENT_LOOSE_OBJECT | \\\n     -+\t\t\t      REPO_COMPONENT_PACK | \\\n     -+\t\t\t      REPO_COMPONENT_PACK_METADATA | \\\n     -+\t\t\t      REPO_COMPONENT_COMMIT_GRAPH | \\\n     -+\t\t\t      REPO_COMPONENT_INDEX)\n     ++#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     ++\t\t\t      FSYNC_COMPONENT_PACK | \\\n     ++\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n     ++\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n      +\n      +\n      +/*\n      + * A bitmask indicating which components of the repo should be fsynced.\n      + */\n     -+extern enum repo_component fsync_components;\n     ++extern enum fsync_component fsync_components;\n       \n       enum fsync_method {\n       \tFSYNC_METHOD_FSYNC,\n     @@ cache.h: int copy_file_with_time(const char *dst, const char *src, int mode);\n       void write_or_die(int fd, const void *buf, size_t count);\n       void fsync_or_die(int fd, const char *);\n       \n     -+inline void fsync_component_or_die(enum repo_component component, int fd, const char *msg)\n     ++inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n      +{\n      +\tif (fsync_components & component)\n      +\t\tfsync_or_die(fd, msg);\n     @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_con\n       \n       \tclose_commit_graph(ctx->r->objects);\n      -\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n     -+\tfinalize_hashfile(f, file_hash, REPO_COMPONENT_COMMIT_GRAPH,\n     ++\tfinalize_hashfile(f, file_hash, FSYNC_COMPONENT_COMMIT_GRAPH,\n      +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n       \tfree_chunkfile(cf);\n       \n     @@ config.c: static int git_parse_maybe_bool_text(const char *value)\n       \n      +static const struct fsync_component_entry {\n      +\tconst char *name;\n     -+\tenum repo_component component_bits;\n     ++\tenum fsync_component component_bits;\n      +} fsync_component_table[] = {\n     -+\t{ \"loose-object\", REPO_COMPONENT_LOOSE_OBJECT },\n     -+\t{ \"pack\", REPO_COMPONENT_PACK },\n     -+\t{ \"pack-metadata\", REPO_COMPONENT_PACK_METADATA },\n     -+\t{ \"commit-graph\", REPO_COMPONENT_COMMIT_GRAPH },\n     -+\t{ \"index\", REPO_COMPONENT_INDEX },\n     ++\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n     ++\t{ \"pack\", FSYNC_COMPONENT_PACK },\n     ++\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n     ++\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n      +\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n      +\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n      +\t{ \"all\", FSYNC_COMPONENTS_ALL },\n      +};\n      +\n     -+static enum repo_component parse_fsync_components(const char *var, const char *string)\n     ++static enum fsync_component parse_fsync_components(const char *var, const char *string)\n      +{\n     -+\tenum repo_component output = 0;\n     ++\tenum fsync_component output = 0;\n      +\n      +\tif (!strcmp(string, \"none\"))\n      +\t\treturn output;\n     @@ csum-file.c: static void free_hashfile(struct hashfile *f)\n       \n      -int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n      +int finalize_hashfile(struct hashfile *f, unsigned char *result,\n     -+\t\t      enum repo_component component, unsigned int flags)\n     ++\t\t      enum fsync_component component, unsigned int flags)\n       {\n       \tint fd;\n       \n     @@ csum-file.c: int finalize_hashfile(struct hashfile *f, unsigned char *result, un\n       \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\n      \n       ## csum-file.h ##\n     +@@\n     + #ifndef CSUM_FILE_H\n     + #define CSUM_FILE_H\n     + \n     ++#include \"cache.h\"\n     + #include \"hash.h\"\n     + \n     + struct progress;\n      @@ csum-file.h: int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n       struct hashfile *hashfd(int fd, const char *name);\n       struct hashfile *hashfd_check(const char *name);\n       struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n      -int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n     -+int finalize_hashfile(struct hashfile *, unsigned char *, enum repo_component, unsigned int);\n     ++int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n       void hashwrite(struct hashfile *, const void *, unsigned int);\n       void hashflush(struct hashfile *f);\n       void crc32_begin(struct hashfile *);\n     @@ environment.c: const char *git_hooks_path;\n       int zlib_compression_level = Z_BEST_SPEED;\n       int pack_compression_level = Z_DEFAULT_COMPRESSION;\n       enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n     -+enum repo_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n     ++enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n       size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n       size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n       size_t delta_base_cache_limit = 96 * 1024 * 1024;\n     @@ midx.c: static int write_midx_internal(const char *object_dir,\n       \twrite_chunkfile(cf, &ctx);\n       \n      -\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n     -+\tfinalize_hashfile(f, midx_hash, REPO_COMPONENT_PACK_METADATA,\n     ++\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n      +\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n       \tfree_chunkfile(cf);\n       \n     @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void\n       {\n      -\tif (fsync_object_files)\n      -\t\tfsync_or_die(fd, \"loose object file\");\n     -+\tfsync_component_or_die(REPO_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n     ++\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n       \tif (close(fd) != 0)\n       \t\tdie_errno(_(\"error when closing loose object file\"));\n       }\n     @@ pack-bitmap-write.c: void bitmap_writer_finish(struct pack_idx_entry **index,\n       \t\twrite_hash_cache(f, index, index_nr);\n       \n      -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n     -+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n     ++\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n      +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n       \n       \tif (adjust_shared_perm(tmp_file.buf))\n     @@ pack-write.c: const char *write_idx_file(const char *index_name, struct pack_idx\n       \thashwrite(f, sha1, the_hash_algo->rawsz);\n      -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n      -\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n     -+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n     +-\t\t\t\t    ? 0 : CSUM_FSYNC));\n     ++\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n      +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n     -+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY)\n     - \t\t\t\t    ? 0 : CSUM_FSYNC));\n     ++\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n       \treturn index_name;\n       }\n     + \n      @@ pack-write.c: const char *write_rev_file_order(const char *rev_name,\n       \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n       \t\tdie(_(\"failed to make %s readable\"), rev_name);\n       \n      -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n      -\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n     -+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n     ++\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n      +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n      +\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n       \n     @@ pack-write.c: void fixup_pack_header_footer(int pack_fd,\n       \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n       \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n      -\tfsync_or_die(pack_fd, pack_name);\n     -+\tfsync_component_or_die(REPO_COMPONENT_PACK, pack_fd, pack_name);\n     ++\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n       }\n       \n       char *index_pack_lockfile(int ip_out, int *is_well_formed)\n      \n       ## read-cache.c ##\n     -@@ read-cache.c: static int record_ieot(void)\n     -  * rely on it.\n     -  */\n     - static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n     --\t\t\t  int strip_extensions)\n     -+\t\t\t  int strip_extensions, unsigned flags)\n     - {\n     - \tuint64_t start = getnanotime();\n     - \tstruct hashfile *f;\n     -@@ read-cache.c: static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n     - \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n     - \tint drop_cache_tree = istate->drop_cache_tree;\n     - \toff_t offset;\n     -+\tint csum_fsync_flag;\n     - \tint ieot_entries = 1;\n     - \tstruct index_entry_offset_table *ieot = NULL;\n     - \tint nr, nr_threads;\n      @@ read-cache.c: static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n       \t\t\treturn -1;\n       \t}\n       \n      -\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n     -+\tcsum_fsync_flag = 0;\n     -+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n     -+\t\tcsum_fsync_flag = CSUM_FSYNC;\n     -+\n     -+\tfinalize_hashfile(f, istate->oid.hash, REPO_COMPONENT_INDEX,\n     -+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n     -+\n     ++\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n       \tif (close_tempfile_gently(tempfile)) {\n       \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n       \t\treturn -1;\n     -@@ read-cache.c: static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n     - \t */\n     - \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n     - \t\t\t\t   \"%s\", get_lock_file_path(lock));\n     --\tret = do_write_index(istate, lock->tempfile, 0);\n     -+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n     - \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n     - \t\t\t\t   \"%s\", get_lock_file_path(lock));\n     - \n     -@@ read-cache.c: static int clean_shared_index_files(const char *current_hex)\n     - }\n     - \n     - static int write_shared_index(struct index_state *istate,\n     --\t\t\t      struct tempfile **temp)\n     -+\t\t\t      struct tempfile **temp, unsigned flags)\n     - {\n     - \tstruct split_index *si = istate->split_index;\n     - \tint ret, was_full = !istate->sparse_index;\n     -@@ read-cache.c: static int write_shared_index(struct index_state *istate,\n     - \n     - \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n     - \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n     --\tret = do_write_index(si->base, *temp, 1);\n     -+\tret = do_write_index(si->base, *temp, 1, flags);\n     - \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n     - \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n     - \n     -@@ read-cache.c: int write_locked_index(struct index_state *istate, struct lock_file *lock,\n     - \t\t\tret = do_write_locked_index(istate, lock, flags);\n     - \t\t\tgoto out;\n     - \t\t}\n     --\t\tret = write_shared_index(istate, &temp);\n     -+\t\tret = write_shared_index(istate, &temp, flags);\n     - \n     - \t\tsaved_errno = errno;\n     - \t\tif (is_tempfile_active(temp))\n -:  ----------- > 3:  86e39b8f8d1 core.fsync: new option to harden the index\n\n-- \ngitgitgadget\n"},{"id":"443223","messageId":"e79522cbdd4feb45b062b75225475f34039d1866.1638845211.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v2.git.1638845211.gitgitgadget@gmail.com","subject":"[PATCH v2 1/3] core.fsyncmethod: add writeout-only mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-07T02:46:49Z","receivedAt":"2021-12-07T02:46:58Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsyncmethod` configuration\nknob, which can currently be set to `fsync` or `writeout-only`.\n\nThe new writeout-only mode attempts to tell the operating system to\nflush its in-memory page cache to the storage hardware without issuing a\nCACHE_FLUSH command to the storage controller.\n\nWriteout-only fsync is significantly faster than a vanilla fsync on\ncommon hardware, since data is written to a disk-side cache rather than\nall the way to a durable medium. Later changes in this patch series will\ntake advantage of this primitive to implement batching of hardware\nflushes.\n\nWhen git_fsync is called with FSYNC_WRITEOUT_ONLY, it may fail and the\ncaller is expected to do an ordinary fsync as needed.\n\nOn Apple platforms, the fsync system call does not issue a CACHE_FLUSH\ndirective to the storage controller. This change updates fsync to do\nfcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\nwith existing behavior on Apple platforms by setting the default value\nof the new core.fsyncmethod option.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt       |  9 +++++\n Makefile                            |  6 ++++\n cache.h                             |  7 ++++\n compat/mingw.h                      |  3 ++\n compat/win32/flush.c                | 28 +++++++++++++++\n config.c                            | 12 +++++++\n config.mak.uname                    |  3 ++\n configure.ac                        |  8 +++++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n environment.c                       |  2 +-\n git-compat-util.h                   | 24 +++++++++++++\n wrapper.c                           | 56 +++++++++++++++++++++++++++++\n write-or-die.c                      | 10 +++---\n 13 files changed, 165 insertions(+), 6 deletions(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..dbb134f7136 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,15 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsyncMethod::\n+\tA value indicating the strategy Git will use to harden repository data\n+\tusing fsync and related primitives.\n++\n+* `fsync` uses the fsync() system call or platform equivalents.\n+* `writeout-only` issues pagecache writeback requests, but depending on the\n+  filesystem and storage hardware, data added to the repository may not be\n+  durable in the event of a system crash. This is the default mode on macOS.\n+\n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n +\ndiff --git a/Makefile b/Makefile\nindex d56c0e4aadc..cba024615c9 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -403,6 +403,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1881,6 +1883,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/cache.h b/cache.h\nindex eba12487b99..9cd60d94952 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -986,6 +986,13 @@ extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n extern int fsync_object_files;\n+\n+enum fsync_method {\n+\tFSYNC_METHOD_FSYNC,\n+\tFSYNC_METHOD_WRITEOUT_ONLY\n+};\n+\n+extern enum fsync_method fsync_method;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..75324c24ee7\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.c b/config.c\nindex c5873f3a706..c3410b8a868 100644\n--- a/config.c\n+++ b/config.c\n@@ -1490,6 +1490,18 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsyncmethod\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcmp(value, \"fsync\"))\n+\t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n+\t\telse if (!strcmp(value, \"writeout-only\"))\n+\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse\n+\t\t\twarning(_(\"unknown %s value '%s'\"), var, value);\n+\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\ndiff --git a/config.mak.uname b/config.mak.uname\nindex d0701f9beb0..774a09622d2 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -57,6 +57,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tBASIC_CFLAGS += -DHAVE_SYSINFO\n@@ -453,6 +454,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -628,6 +630,7 @@ ifeq ($(uname_S),MINGW)\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/configure.ac b/configure.ac\nindex 5ee25ec95c8..6bd6bef1c44 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1082,6 +1082,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 86b46114464..6d7bc16d054 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/environment.c b/environment.c\nindex 9da7f3c1a19..f9140e842cf 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -41,7 +41,7 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex c6bd2a84e55..cb9abd7a08c 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1239,6 +1239,30 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+#ifdef __APPLE__\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n+#else\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n+#endif\n+\n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+/*\n+ * Issues an fsync against the specified file according to the specified mode.\n+ *\n+ * FSYNC_WRITEOUT_ONLY attempts to use interfaces available on some operating\n+ * systems to flush the OS cache without issuing a flush command to the storage\n+ * controller. If those interfaces are unavailable, the function fails with\n+ * ENOSYS.\n+ *\n+ * FSYNC_HARDWARE_FLUSH does an OS writeout and hardware flush to ensure that\n+ * changes are durable. It is not expected to fail.\n+ */\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/wrapper.c b/wrapper.c\nindex 36e12119d76..1c5f2c87791 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -546,6 +546,62 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\t\t/*\n+\t\t * On some platforms fsync may return EINTR. Try again in this\n+\t\t * case, since callers asking for a hardware flush may die if\n+\t\t * this function returns an error.\n+\t\t */\n+\t\tfor (;;) {\n+\t\t\tint err;\n+#ifdef __APPLE__\n+\t\t\terr = fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\t\terr = fsync(fd);\n+#endif\n+\t\t\tif (err >= 0 || errno != EINTR)\n+\t\t\t\treturn err;\n+\t\t}\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 0b1ec8190b6..0702acdd5e8 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,10 +57,12 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n-\t\tif (errno != EINTR)\n-\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n-\t}\n+\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n+\t\treturn;\n+\n+\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n+\t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n void write_or_die(int fd, const void *buf, size_t count)\n-- \ngitgitgadget\n\n"},{"id":"443225","messageId":"ff80a94bf9add8a6fabcd5146e5177edf5e35e49.1638845211.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v2.git.1638845211.gitgitgadget@gmail.com","subject":"[PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-07T02:46:50Z","receivedAt":"2021-12-07T02:47:01Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsync` configuration\nknob which can be used to control how components of the\nrepository are made durable on disk.\n\nThis setting allows future extensibility of the list of\nsyncable components:\n* We issue a warning rather than an error for unrecognized\n  components, so new configs can be used with old Git versions.\n* We support negation, so users can choose one of the default\n  aggregate options and then remove components that they don't\n  want. The user would then harden any new components added in\n  a Git version update.\n\nThis also supports the common request of doing absolutely no\nfysncing with the `core.fsync=none` value, which is expected\nto make the test suite faster.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 27 +++++++++----\n builtin/fast-import.c         |  2 +-\n builtin/index-pack.c          |  4 +-\n builtin/pack-objects.c        |  8 ++--\n bulk-checkin.c                |  5 ++-\n cache.h                       | 39 +++++++++++++++++-\n commit-graph.c                |  3 +-\n config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n csum-file.c                   |  5 ++-\n csum-file.h                   |  3 +-\n environment.c                 |  1 +\n midx.c                        |  3 +-\n object-file.c                 |  3 +-\n pack-bitmap-write.c           |  3 +-\n pack-write.c                  | 13 +++---\n read-cache.c                  |  2 +-\n 16 files changed, 164 insertions(+), 33 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex dbb134f7136..4f1747ec871 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,25 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsync::\n+\tA comma-separated list of parts of the repository which should be\n+\thardened via the core.fsyncMethod when created or modified. You can\n+\tdisable hardening of any component by prefixing it with a '-'. Later\n+\titems take precedence over earlier ones in the list. For example,\n+\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n+\tmetadata.\" Items that are not hardened may be lost in the event of an\n+\tunclean system shutdown.\n++\n+* `none` disables fsync completely. This must be specified alone.\n+* `loose-object` hardens objects added to the repo in loose-object form.\n+* `pack` hardens objects added to the repo in packfile form.\n+* `pack-metadata` hardens packfile bitmaps and indexes.\n+* `commit-graph` hardens the commit graph file.\n+* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n+  `pack-metadata`, and `commit-graph`.\n+* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n+* `all` is an aggregate option that syncs all individual components above.\n+\n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\n@@ -556,14 +575,6 @@ core.fsyncMethod::\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n \n-core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n-\n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\n +\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 20406f67754..e27a4580f85 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -856,7 +856,7 @@ static void end_packfile(void)\n \t\tstruct tag *t;\n \n \t\tclose_pack_windows(pack_data);\n-\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n \t\t\t\t\t pack_data->pack_name, object_count,\n \t\t\t\t\t cur_pack_oid.hash, pack_size);\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c23d01de7dc..c32534c13b4 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1286,7 +1286,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\t\t    nr_objects - nr_objects_initial);\n \t\tstop_progress_msg(&progress, msg.buf);\n \t\tstrbuf_release(&msg);\n-\t\tfinalize_hashfile(f, tail_hash, 0);\n+\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n \t\thashcpy(read_hash, pack_hash);\n \t\tfixup_pack_header_footer(output_fd, pack_hash,\n \t\t\t\t\t curr_pack, nr_objects,\n@@ -1508,7 +1508,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tif (!from_stdin) {\n \t\tclose(input_fd);\n \t} else {\n-\t\tfsync_or_die(output_fd, curr_pack_name);\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n \t\terr = close(output_fd);\n \t\tif (err)\n \t\t\tdie_errno(_(\"error while closing pack file\"));\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 857be7826f3..916c55d6ce9 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n \t\t * If so, rewrite it like in fast-import\n \t\t */\n \t\tif (pack_to_stdout) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n \t\t} else if (nr_written == nr_remaining) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t\t} else {\n-\t\t\tint fd = finalize_hashfile(f, hash, 0);\n+\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n \t\t\t\t\t\t nr_written, hash, offset);\n \t\t\tclose(fd);\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..a2cf9dcbc8d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -53,9 +53,10 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \t\tunlink(state->pack_tmp_name);\n \t\tgoto clear_exit;\n \t} else if (state->nr_written == 1) {\n-\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\tfinalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t} else {\n-\t\tint fd = finalize_hashfile(state->f, hash, 0);\n+\t\tint fd = finalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n \t\t\t\t\t state->nr_written, hash,\n \t\t\t\t\t state->offset);\ndiff --git a/cache.h b/cache.h\nindex 9cd60d94952..d83fbaf2619 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,38 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+/*\n+ * These values are used to help identify parts of a repository to fsync.\n+ * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n+ * repository and so shouldn't be fsynced.\n+ */\n+enum fsync_component {\n+\tFSYNC_COMPONENT_NONE\t\t\t= 0,\n+\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n+\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n+\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n+\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+};\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t      FSYNC_COMPONENT_PACK | \\\n+\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+\n+/*\n+ * A bitmask indicating which components of the repo should be fsynced.\n+ */\n+extern enum fsync_component fsync_components;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n@@ -1747,6 +1778,12 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n \n+inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n+{\n+\tif (fsync_components & component)\n+\t\tfsync_or_die(fd, msg);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 2706683acfe..c8a5dea4541 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1939,7 +1939,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t}\n \n \tclose_commit_graph(ctx->r->objects);\n-\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n+\tfinalize_hashfile(f, file_hash, FSYNC_COMPONENT_COMMIT_GRAPH,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n \tfree_chunkfile(cf);\n \n \tif (ctx->split) {\ndiff --git a/config.c b/config.c\nindex c3410b8a868..29c867aab03 100644\n--- a/config.c\n+++ b/config.c\n@@ -1213,6 +1213,73 @@ static int git_parse_maybe_bool_text(const char *value)\n \treturn -1;\n }\n \n+static const struct fsync_component_entry {\n+\tconst char *name;\n+\tenum fsync_component component_bits;\n+} fsync_component_table[] = {\n+\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n+\t{ \"pack\", FSYNC_COMPONENT_PACK },\n+\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n+\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n+};\n+\n+static enum fsync_component parse_fsync_components(const char *var, const char *string)\n+{\n+\tenum fsync_component output = 0;\n+\n+\tif (!strcmp(string, \"none\"))\n+\t\treturn output;\n+\n+\twhile (string) {\n+\t\tint i;\n+\t\tsize_t len;\n+\t\tconst char *ep;\n+\t\tint negated = 0;\n+\t\tint found = 0;\n+\n+\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n+\t\tep = strchrnul(string, ',');\n+\t\tlen = ep - string;\n+\n+\t\tif (*string == '-') {\n+\t\t\tnegated = 1;\n+\t\t\tstring++;\n+\t\t\tlen--;\n+\t\t\tif (!len)\n+\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n+\t\t}\n+\n+\t\tif (!len)\n+\t\t\tbreak;\n+\n+\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n+\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n+\n+\t\t\tif (strncmp(entry->name, string, len))\n+\t\t\t\tcontinue;\n+\n+\t\t\tfound = 1;\n+\t\t\tif (negated)\n+\t\t\t\toutput &= ~entry->component_bits;\n+\t\t\telse\n+\t\t\t\toutput |= entry->component_bits;\n+\t\t}\n+\n+\t\tif (!found) {\n+\t\t\tchar *component = xstrndup(string, len);\n+\t\t\twarning(_(\"unknown %s value '%s'\"), var, component);\n+\t\t\tfree(component);\n+\t\t}\n+\n+\t\tstring = ep;\n+\t}\n+\n+\treturn output;\n+}\n+\n int git_parse_maybe_bool(const char *value)\n {\n \tint v = git_parse_maybe_bool_text(value);\n@@ -1490,6 +1557,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsync\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tfsync_components = parse_fsync_components(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncmethod\")) {\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n@@ -1503,7 +1577,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n \t\treturn 0;\n \t}\n \ndiff --git a/csum-file.c b/csum-file.c\nindex 26e8a6df44e..59ef3398ca2 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -58,7 +58,8 @@ static void free_hashfile(struct hashfile *f)\n \tfree(f);\n }\n \n-int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n+int finalize_hashfile(struct hashfile *f, unsigned char *result,\n+\t\t      enum fsync_component component, unsigned int flags)\n {\n \tint fd;\n \n@@ -69,7 +70,7 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int fl\n \tif (flags & CSUM_HASH_IN_STREAM)\n \t\tflush(f, f->buffer, the_hash_algo->rawsz);\n \tif (flags & CSUM_FSYNC)\n-\t\tfsync_or_die(f->fd, f->name);\n+\t\tfsync_component_or_die(component, f->fd, f->name);\n \tif (flags & CSUM_CLOSE) {\n \t\tif (close(f->fd))\n \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\ndiff --git a/csum-file.h b/csum-file.h\nindex 291215b34eb..0d29f528fbc 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -1,6 +1,7 @@\n #ifndef CSUM_FILE_H\n #define CSUM_FILE_H\n \n+#include \"cache.h\"\n #include \"hash.h\"\n \n struct progress;\n@@ -38,7 +39,7 @@ int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n struct hashfile *hashfd(int fd, const char *name);\n struct hashfile *hashfd_check(const char *name);\n struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n-int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n+int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n void hashwrite(struct hashfile *, const void *, unsigned int);\n void hashflush(struct hashfile *f);\n void crc32_begin(struct hashfile *);\ndiff --git a/environment.c b/environment.c\nindex f9140e842cf..09905adecf9 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,6 +42,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n+enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/midx.c b/midx.c\nindex 837b46b2af5..882f91f7d57 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -1406,7 +1406,8 @@ static int write_midx_internal(const char *object_dir,\n \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n \twrite_chunkfile(cf, &ctx);\n \n-\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n \tfree_chunkfile(cf);\n \n \tif (flags & (MIDX_WRITE_REV_INDEX | MIDX_WRITE_BITMAP))\ndiff --git a/object-file.c b/object-file.c\nindex eb972cdccd2..9d9c4a39e85 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1809,8 +1809,7 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex 9c55c1531e1..c16e43d1669 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -719,7 +719,8 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \tif (options & BITMAP_OPT_HASH_CACHE)\n \t\twrite_hash_cache(f, index, index_nr);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \n \tif (adjust_shared_perm(tmp_file.buf))\n \t\tdie_errno(\"unable to make temporary bitmap file readable\");\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a346..51812cb1299 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -159,9 +159,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t}\n \n \thashwrite(f, sha1, the_hash_algo->rawsz);\n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n-\t\t\t\t    ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \treturn index_name;\n }\n \n@@ -281,8 +281,9 @@ const char *write_rev_file_order(const char *rev_name,\n \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n \t\tdie(_(\"failed to make %s readable\"), rev_name);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \n \treturn rev_name;\n }\n@@ -390,7 +391,7 @@ void fixup_pack_header_footer(int pack_fd,\n \t\tthe_hash_algo->final_fn(partial_pack_hash, &old_hash_ctx);\n \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n-\tfsync_or_die(pack_fd, pack_name);\n+\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n }\n \n char *index_pack_lockfile(int ip_out, int *is_well_formed)\ndiff --git a/read-cache.c b/read-cache.c\nindex f3986596623..f3539681f49 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -3060,7 +3060,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"443224","messageId":"86e39b8f8d16100b33346d16e7f722064ab9615c.1638845211.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v2.git.1638845211.gitgitgadget@gmail.com","subject":"[PATCH v2 3/3] core.fsync: new option to harden the index","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-07T02:46:51Z","receivedAt":"2021-12-07T02:47:03Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the new ability for the user to harden\nthe index. In the event of a system crash, the index must be\ndurable for the user to actually find a file that has been added\nto the repo and then deleted from the working tree.\n\nWe use the presence of the COMMIT_LOCK flag and absence of the\nalternate_index_output as a proxy for determining whether we're\nupdating the persistent index of the repo or some temporary\nindex. We don't sync these temporary indexes.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  1 +\n cache.h                       |  4 +++-\n config.c                      |  1 +\n read-cache.c                  | 19 +++++++++++++------\n 4 files changed, 18 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 4f1747ec871..8e5b7a795ab 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -561,6 +561,7 @@ core.fsync::\n * `pack` hardens objects added to the repo in packfile form.\n * `pack-metadata` hardens packfile bitmaps and indexes.\n * `commit-graph` hardens the commit graph file.\n+* `index` hardens the index when it is modified.\n * `objects` is an aggregate option that includes `loose-objects`, `pack`,\n   `pack-metadata`, and `commit-graph`.\n * `default` is an aggregate option that is equivalent to `objects,-loose-object`\ndiff --git a/cache.h b/cache.h\nindex d83fbaf2619..4dc26d7b2c9 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -996,6 +996,7 @@ enum fsync_component {\n \tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+\tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n #define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n@@ -1010,7 +1011,8 @@ enum fsync_component {\n #define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n \t\t\t      FSYNC_COMPONENT_PACK | \\\n \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n+\t\t\t      FSYNC_COMPONENT_INDEX)\n \n \n /*\ndiff --git a/config.c b/config.c\nindex 29c867aab03..17039fa9c10 100644\n--- a/config.c\n+++ b/config.c\n@@ -1221,6 +1221,7 @@ static const struct fsync_component_entry {\n \t{ \"pack\", FSYNC_COMPONENT_PACK },\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"index\", FSYNC_COMPONENT_INDEX },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n \t{ \"all\", FSYNC_COMPONENTS_ALL },\ndiff --git a/read-cache.c b/read-cache.c\nindex f3539681f49..783cb3ea5db 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2816,7 +2816,7 @@ static int record_ieot(void)\n  * rely on it.\n  */\n static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n-\t\t\t  int strip_extensions)\n+\t\t\t  int strip_extensions, unsigned flags)\n {\n \tuint64_t start = getnanotime();\n \tstruct hashfile *f;\n@@ -2830,6 +2830,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint csum_fsync_flag;\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n@@ -3060,7 +3061,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n+\tcsum_fsync_flag = 0;\n+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n+\t\tcsum_fsync_flag = CSUM_FSYNC;\n+\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n+\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n@@ -3115,7 +3122,7 @@ static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n \t */\n \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n-\tret = do_write_index(istate, lock->tempfile, 0);\n+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n \n@@ -3209,7 +3216,7 @@ static int clean_shared_index_files(const char *current_hex)\n }\n \n static int write_shared_index(struct index_state *istate,\n-\t\t\t      struct tempfile **temp)\n+\t\t\t      struct tempfile **temp, unsigned flags)\n {\n \tstruct split_index *si = istate->split_index;\n \tint ret, was_full = !istate->sparse_index;\n@@ -3219,7 +3226,7 @@ static int write_shared_index(struct index_state *istate,\n \n \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n-\tret = do_write_index(si->base, *temp, 1);\n+\tret = do_write_index(si->base, *temp, 1, flags);\n \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n \n@@ -3328,7 +3335,7 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n \t\t\tret = do_write_locked_index(istate, lock, flags);\n \t\t\tgoto out;\n \t\t}\n-\t\tret = write_shared_index(istate, &temp);\n+\t\tret = write_shared_index(istate, &temp, flags);\n \n \t\tsaved_errno = errno;\n \t\tif (is_tempfile_active(temp))\n-- \ngitgitgadget\n"},{"id":"443268","messageId":"Ya9JJlItvDJCLHqj@ncase","threadId":"57030","inReplyTo":"e79522cbdd4feb45b062b75225475f34039d1866.1638845211.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 1/3] core.fsyncmethod: add writeout-only mode","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2021-12-07T11:44:38Z","receivedAt":"2021-12-07T11:45:25Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 07, 2021 at 02:46:49AM +0000, Neeraj Singh via GitGitGadget wrote:\n> From: Neeraj Singh <neerajsi@microsoft.com>\n[snip]\n> --- a/compat/mingw.h\n> +++ b/compat/mingw.h\n> @@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n>  #define getpagesize mingw_getpagesize\n>  #endif\n>  \n> +int win32_fsync_no_flush(int fd);\n> +#define fsync_no_flush win32_fsync_no_flush\n> +\n>  struct rlimit {\n>  \tunsigned int rlim_cur;\n>  };\n> diff --git a/compat/win32/flush.c b/compat/win32/flush.c\n> new file mode 100644\n> index 00000000000..75324c24ee7\n> --- /dev/null\n> +++ b/compat/win32/flush.c\n> @@ -0,0 +1,28 @@\n> +#include \"../../git-compat-util.h\"\n> +#include <winternl.h>\n> +#include \"lazyload.h\"\n> +\n> +int win32_fsync_no_flush(int fd)\n> +{\n> +       IO_STATUS_BLOCK io_status;\n> +\n> +#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n> +\n> +       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n> +\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n> +\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n> +\n> +       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n> +\t\terrno = ENOSYS;\n> +\t\treturn -1;\n> +       }\n\nI'm wondering whether it would make sense to fall back to fsync(3P) in\ncase we cannot use writeout-only, but I see that were doing essentially\nthat in `fsync_or_die()`. There is no indicator to the user though that\nwriteout-only doesn't work -- do we want to print a one-time warning?\n\n> +       memset(&io_status, 0, sizeof(io_status));\n> +       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n> +\t\t\t\tNULL, 0, &io_status)) {\n> +\t\terrno = EINVAL;\n> +\t\treturn -1;\n> +       }\n> +\n> +       return 0;\n> +}\n\n[snip]\n> diff --git a/wrapper.c b/wrapper.c\n> index 36e12119d76..1c5f2c87791 100644\n> --- a/wrapper.c\n> +++ b/wrapper.c\n> @@ -546,6 +546,62 @@ int xmkstemp_mode(char *filename_template, int mode)\n>  \treturn fd;\n>  }\n>  \n> +int git_fsync(int fd, enum fsync_action action)\n> +{\n> +\tswitch (action) {\n> +\tcase FSYNC_WRITEOUT_ONLY:\n> +\n> +#ifdef __APPLE__\n> +\t\t/*\n> +\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n> +\t\t * flush hardware caches.\n> +\t\t */\n> +\t\treturn fsync(fd);\n\nBelow we're looping around `EINTR` -- are Apple systems never returning\nit?\n\nPatrick\n\n> +#endif\n> +\n> +#ifdef HAVE_SYNC_FILE_RANGE\n> +\t\t/*\n> +\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n> +\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n> +\t\t * indicates writeout of the entire file and the wait flags ensure that all\n> +\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n> +\t\t * before we continue.\n> +\t\t */\n> +\n> +\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n> +#endif\n> +\n> +#ifdef fsync_no_flush\n> +\t\treturn fsync_no_flush(fd);\n> +#endif\n> +\n> +\t\terrno = ENOSYS;\n> +\t\treturn -1;\n> +\n> +\tcase FSYNC_HARDWARE_FLUSH:\n> +\t\t/*\n> +\t\t * On some platforms fsync may return EINTR. Try again in this\n> +\t\t * case, since callers asking for a hardware flush may die if\n> +\t\t * this function returns an error.\n> +\t\t */\n> +\t\tfor (;;) {\n> +\t\t\tint err;\n> +#ifdef __APPLE__\n> +\t\t\terr = fcntl(fd, F_FULLFSYNC);\n> +#else\n> +\t\t\terr = fsync(fd);\n> +#endif\n> +\t\t\tif (err >= 0 || errno != EINTR)\n> +\t\t\t\treturn err;\n> +\t\t}\n> +\n> +\tdefault:\n> +\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n> +\t}\n> +}\n> +\n>  static int warn_if_unremovable(const char *op, const char *file, int rc)\n>  {\n>  \tint err;\n> diff --git a/write-or-die.c b/write-or-die.c\n> index 0b1ec8190b6..0702acdd5e8 100644\n> --- a/write-or-die.c\n> +++ b/write-or-die.c\n> @@ -57,10 +57,12 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n>  \n>  void fsync_or_die(int fd, const char *msg)\n>  {\n> -\twhile (fsync(fd) < 0) {\n> -\t\tif (errno != EINTR)\n> -\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n> -\t}\n> +\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n> +\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n> +\t\treturn;\n> +\n> +\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n> +\t\tdie_errno(\"fsync error on '%s'\", msg);\n>  }\n>  \n>  void write_or_die(int fd, const void *buf, size_t count)\n> -- \n> gitgitgadget\n> \n"},{"id":"443269","messageId":"Ya9LPOseu8geBi4v@ncase","threadId":"57030","inReplyTo":"ff80a94bf9add8a6fabcd5146e5177edf5e35e49.1638845211.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2021-12-07T11:53:32Z","receivedAt":"2021-12-07T11:54:17Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 07, 2021 at 02:46:50AM +0000, Neeraj Singh via GitGitGadget wrote:\n> From: Neeraj Singh <neerajsi@microsoft.com>\n[snip]\n> diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> index c23d01de7dc..c32534c13b4 100644\n> --- a/builtin/index-pack.c\n> +++ b/builtin/index-pack.c\n> @@ -1286,7 +1286,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n>  \t\t\t    nr_objects - nr_objects_initial);\n>  \t\tstop_progress_msg(&progress, msg.buf);\n>  \t\tstrbuf_release(&msg);\n> -\t\tfinalize_hashfile(f, tail_hash, 0);\n> +\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n>  \t\thashcpy(read_hash, pack_hash);\n>  \t\tfixup_pack_header_footer(output_fd, pack_hash,\n>  \t\t\t\t\t curr_pack, nr_objects,\n> @@ -1508,7 +1508,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n>  \tif (!from_stdin) {\n>  \t\tclose(input_fd);\n>  \t} else {\n> -\t\tfsync_or_die(output_fd, curr_pack_name);\n> +\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n>  \t\terr = close(output_fd);\n>  \t\tif (err)\n>  \t\t\tdie_errno(_(\"error while closing pack file\"));\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> index 857be7826f3..916c55d6ce9 100644\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n> @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n>  \t\t * If so, rewrite it like in fast-import\n>  \t\t */\n>  \t\tif (pack_to_stdout) {\n> -\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> +\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n> +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n\nIt doesn't have any effect here given that we don't sync at all when\nwriting to stdout, but I wonder whether we should set up the component\ncorrectly regardless of that such that it makes for a less confusing\nread.\n\n[snip]\n> diff --git a/config.c b/config.c\n> index c3410b8a868..29c867aab03 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1213,6 +1213,73 @@ static int git_parse_maybe_bool_text(const char *value)\n>  \treturn -1;\n>  }\n>  \n> +static const struct fsync_component_entry {\n> +\tconst char *name;\n> +\tenum fsync_component component_bits;\n> +} fsync_component_table[] = {\n> +\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> +\t{ \"pack\", FSYNC_COMPONENT_PACK },\n> +\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> +\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> +\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> +\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n> +\t{ \"all\", FSYNC_COMPONENTS_ALL },\n> +};\n> +\n> +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> +{\n> +\tenum fsync_component output = 0;\n> +\n> +\tif (!strcmp(string, \"none\"))\n> +\t\treturn output;\n> +\n> +\twhile (string) {\n> +\t\tint i;\n> +\t\tsize_t len;\n> +\t\tconst char *ep;\n> +\t\tint negated = 0;\n> +\t\tint found = 0;\n> +\n> +\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n> +\t\tep = strchrnul(string, ',');\n> +\t\tlen = ep - string;\n> +\n> +\t\tif (*string == '-') {\n> +\t\t\tnegated = 1;\n> +\t\t\tstring++;\n> +\t\t\tlen--;\n> +\t\t\tif (!len)\n> +\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n> +\t\t}\n> +\n> +\t\tif (!len)\n> +\t\t\tbreak;\n> +\n> +\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n> +\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n> +\n> +\t\t\tif (strncmp(entry->name, string, len))\n> +\t\t\t\tcontinue;\n> +\n> +\t\t\tfound = 1;\n> +\t\t\tif (negated)\n> +\t\t\t\toutput &= ~entry->component_bits;\n> +\t\t\telse\n> +\t\t\t\toutput |= entry->component_bits;\n> +\t\t}\n> +\n> +\t\tif (!found) {\n> +\t\t\tchar *component = xstrndup(string, len);\n> +\t\t\twarning(_(\"unknown %s value '%s'\"), var, component);\n> +\t\t\tfree(component);\n> +\t\t}\n> +\n> +\t\tstring = ep;\n> +\t}\n> +\n> +\treturn output;\n> +}\n> +\n>  int git_parse_maybe_bool(const char *value)\n>  {\n>  \tint v = git_parse_maybe_bool_text(value);\n> @@ -1490,6 +1557,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t\treturn 0;\n>  \t}\n>  \n> +\tif (!strcmp(var, \"core.fsync\")) {\n> +\t\tif (!value)\n> +\t\t\treturn config_error_nonbool(var);\n> +\t\tfsync_components = parse_fsync_components(var, value);\n> +\t\treturn 0;\n> +\t}\n> +\n>  \tif (!strcmp(var, \"core.fsyncmethod\")) {\n>  \t\tif (!value)\n>  \t\t\treturn config_error_nonbool(var);\n> @@ -1503,7 +1577,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t}\n>  \n>  \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> -\t\tfsync_object_files = git_config_bool(var, value);\n> +\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n>  \t\treturn 0;\n>  \t}\n\nShouldn't we continue to support this for now such that users can\nmigrate from the old, deprecated value first before we start to ignore\nit?\n\nPatrick\n\n> diff --git a/csum-file.c b/csum-file.c\n> index 26e8a6df44e..59ef3398ca2 100644\n> --- a/csum-file.c\n> +++ b/csum-file.c\n> @@ -58,7 +58,8 @@ static void free_hashfile(struct hashfile *f)\n>  \tfree(f);\n>  }\n>  \n> -int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n> +int finalize_hashfile(struct hashfile *f, unsigned char *result,\n> +\t\t      enum fsync_component component, unsigned int flags)\n>  {\n>  \tint fd;\n>  \n> @@ -69,7 +70,7 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int fl\n>  \tif (flags & CSUM_HASH_IN_STREAM)\n>  \t\tflush(f, f->buffer, the_hash_algo->rawsz);\n>  \tif (flags & CSUM_FSYNC)\n> -\t\tfsync_or_die(f->fd, f->name);\n> +\t\tfsync_component_or_die(component, f->fd, f->name);\n>  \tif (flags & CSUM_CLOSE) {\n>  \t\tif (close(f->fd))\n>  \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\n> diff --git a/csum-file.h b/csum-file.h\n> index 291215b34eb..0d29f528fbc 100644\n> --- a/csum-file.h\n> +++ b/csum-file.h\n> @@ -1,6 +1,7 @@\n>  #ifndef CSUM_FILE_H\n>  #define CSUM_FILE_H\n>  \n> +#include \"cache.h\"\n>  #include \"hash.h\"\n>  \n>  struct progress;\n> @@ -38,7 +39,7 @@ int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n>  struct hashfile *hashfd(int fd, const char *name);\n>  struct hashfile *hashfd_check(const char *name);\n>  struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n> -int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n> +int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n>  void hashwrite(struct hashfile *, const void *, unsigned int);\n>  void hashflush(struct hashfile *f);\n>  void crc32_begin(struct hashfile *);\n> diff --git a/environment.c b/environment.c\n> index f9140e842cf..09905adecf9 100644\n> --- a/environment.c\n> +++ b/environment.c\n> @@ -42,6 +42,7 @@ const char *git_hooks_path;\n>  int zlib_compression_level = Z_BEST_SPEED;\n>  int pack_compression_level = Z_DEFAULT_COMPRESSION;\n>  enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n> +enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n>  size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n>  size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n>  size_t delta_base_cache_limit = 96 * 1024 * 1024;\n> diff --git a/midx.c b/midx.c\n> index 837b46b2af5..882f91f7d57 100644\n> --- a/midx.c\n> +++ b/midx.c\n> @@ -1406,7 +1406,8 @@ static int write_midx_internal(const char *object_dir,\n>  \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n>  \twrite_chunkfile(cf, &ctx);\n>  \n> -\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n> +\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n> +\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n>  \tfree_chunkfile(cf);\n>  \n>  \tif (flags & (MIDX_WRITE_REV_INDEX | MIDX_WRITE_BITMAP))\n> diff --git a/object-file.c b/object-file.c\n> index eb972cdccd2..9d9c4a39e85 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1809,8 +1809,7 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  /* Finalize a file on disk, and close it. */\n>  static void close_loose_object(int fd)\n>  {\n> -\tif (fsync_object_files)\n> -\t\tfsync_or_die(fd, \"loose object file\");\n> +\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n>  \tif (close(fd) != 0)\n>  \t\tdie_errno(_(\"error when closing loose object file\"));\n>  }\n> diff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\n> index 9c55c1531e1..c16e43d1669 100644\n> --- a/pack-bitmap-write.c\n> +++ b/pack-bitmap-write.c\n> @@ -719,7 +719,8 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n>  \tif (options & BITMAP_OPT_HASH_CACHE)\n>  \t\twrite_hash_cache(f, index, index_nr);\n>  \n> -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n> +\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n> +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n>  \n>  \tif (adjust_shared_perm(tmp_file.buf))\n>  \t\tdie_errno(\"unable to make temporary bitmap file readable\");\n> diff --git a/pack-write.c b/pack-write.c\n> index a5846f3a346..51812cb1299 100644\n> --- a/pack-write.c\n> +++ b/pack-write.c\n> @@ -159,9 +159,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n>  \t}\n>  \n>  \thashwrite(f, sha1, the_hash_algo->rawsz);\n> -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n> -\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n> -\t\t\t\t    ? 0 : CSUM_FSYNC));\n> +\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n> +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n> +\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n>  \treturn index_name;\n>  }\n>  \n> @@ -281,8 +281,9 @@ const char *write_rev_file_order(const char *rev_name,\n>  \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n>  \t\tdie(_(\"failed to make %s readable\"), rev_name);\n>  \n> -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n> -\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n> +\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n> +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n> +\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n>  \n>  \treturn rev_name;\n>  }\n> @@ -390,7 +391,7 @@ void fixup_pack_header_footer(int pack_fd,\n>  \t\tthe_hash_algo->final_fn(partial_pack_hash, &old_hash_ctx);\n>  \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n>  \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n> -\tfsync_or_die(pack_fd, pack_name);\n> +\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n>  }\n>  \n>  char *index_pack_lockfile(int ip_out, int *is_well_formed)\n> diff --git a/read-cache.c b/read-cache.c\n> index f3986596623..f3539681f49 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -3060,7 +3060,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>  \t\t\treturn -1;\n>  \t}\n>  \n> -\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n> +\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n>  \tif (close_tempfile_gently(tempfile)) {\n>  \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n>  \t\treturn -1;\n> -- \n> gitgitgadget\n> \n"},{"id":"443270","messageId":"Ya9L5GJqlBF1YEk2@ncase","threadId":"57030","inReplyTo":"pull.1093.v2.git.1638845211.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 0/3] A design for future-proofing fsync() configuration","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2021-12-07T11:56:20Z","receivedAt":"2021-12-07T11:57:12Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Dec 07, 2021 at 02:46:48AM +0000, Neeraj K. Singh via GitGitGadget wrote:\n> This is an implementation of an extensible configuration mechanism for\n> fsyncing persistent components of a repo.\n> \n> The main goals are to separate the \"what\" to sync from the \"how\". There are\n> now two settings: core.fsync - Control the 'what', including the index.\n> core.fsyncMethod - Control the 'how'. Currently we support writeout-only and\n> full fsync.\n> \n> Syncing of refs can be layered on top of core.fsync. And batch mode will be\n> layered on core.fsyncMethod.\n> \n> core.fsyncobjectfiles is removed and will issue a deprecation warning if\n> it's seen.\n> \n> I'd like to get agreement on this direction before submitting batch mode to\n> the list. The batch mode series is available to view at\n> https://github.com/neerajsi-msft/git/pull/1.\n> \n> Please see [1], [2], and [3] for discussions that led to this series.\n> \n> V2 changes:\n> \n>  * Updated the documentation for core.fsyncmethod to be less certain.\n>    writeout-only probably does not do the right thing on Linux.\n>  * Split out the core.fsync=index change into its own commit.\n>  * Rename REPO_COMPONENT to FSYNC_COMPONENT. This is really specific to\n>    fsyncing, so the name should reflect that.\n>  * Re-add missing Makefile change for SYNC_FILE_RANGE.\n>  * Tested writeout-only mode, index syncing, and general config settings.\n> \n> [1] https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n> [2]\n> https://lore.kernel.org/git/dd65718814011eb93ccc4428f9882e0f025224a6.1636029491.git.ps@pks.im/\n> [3]\n> https://lore.kernel.org/git/pull.1076.git.git.1629856292.gitgitgadget@gmail.com/\n\nWhile I bail from the question of whether we want to grant as much\nconfigurability to the user as this patch series does, I quite like the\nimplementation. It feels rather straight-forward and it's easy to see\nhow to extend it to support syncing of other subsystems like the loose\nrefs.\n\nThanks!\n\nPatrick\n\n> Neeraj Singh (3):\n>   core.fsyncmethod: add writeout-only mode\n>   core.fsync: introduce granular fsync control\n>   core.fsync: new option to harden the index\n> \n>  Documentation/config/core.txt       | 35 +++++++++---\n>  Makefile                            |  6 ++\n>  builtin/fast-import.c               |  2 +-\n>  builtin/index-pack.c                |  4 +-\n>  builtin/pack-objects.c              |  8 ++-\n>  bulk-checkin.c                      |  5 +-\n>  cache.h                             | 48 +++++++++++++++-\n>  commit-graph.c                      |  3 +-\n>  compat/mingw.h                      |  3 +\n>  compat/win32/flush.c                | 28 +++++++++\n>  config.c                            | 89 ++++++++++++++++++++++++++++-\n>  config.mak.uname                    |  3 +\n>  configure.ac                        |  8 +++\n>  contrib/buildsystems/CMakeLists.txt |  3 +-\n>  csum-file.c                         |  5 +-\n>  csum-file.h                         |  3 +-\n>  environment.c                       |  3 +-\n>  git-compat-util.h                   | 24 ++++++++\n>  midx.c                              |  3 +-\n>  object-file.c                       |  3 +-\n>  pack-bitmap-write.c                 |  3 +-\n>  pack-write.c                        | 13 +++--\n>  read-cache.c                        | 19 ++++--\n>  wrapper.c                           | 56 ++++++++++++++++++\n>  write-or-die.c                      | 10 ++--\n>  25 files changed, 344 insertions(+), 43 deletions(-)\n>  create mode 100644 compat/win32/flush.c\n> \n> \n> base-commit: abe6bb3905392d5eb6b01fa6e54d7e784e0522aa\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-1093%2Fneerajsi-msft%2Fns%2Fcore-fsync-v2\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1093/neerajsi-msft/ns/core-fsync-v2\n> Pull-Request: https://github.com/gitgitgadget/git/pull/1093\n> \n> Range-diff vs v1:\n> \n>  1:  527380ddc3f ! 1:  e79522cbdd4 fsync: add writeout-only mode for fsyncing repo data\n>      @@ Metadata\n>       Author: Neeraj Singh <neerajsi@microsoft.com>\n>       \n>        ## Commit message ##\n>      -    fsync: add writeout-only mode for fsyncing repo data\n>      +    core.fsyncmethod: add writeout-only mode\n>      +\n>      +    This commit introduces the `core.fsyncmethod` configuration\n>      +    knob, which can currently be set to `fsync` or `writeout-only`.\n>       \n>           The new writeout-only mode attempts to tell the operating system to\n>           flush its in-memory page cache to the storage hardware without issuing a\n>      @@ Documentation/config/core.txt: core.whitespace::\n>       +\tusing fsync and related primitives.\n>       ++\n>       +* `fsync` uses the fsync() system call or platform equivalents.\n>      -+* `writeout-only` issues requests to send the writes to the storage\n>      -+  hardware, but does not send any FLUSH CACHE request. If the operating system\n>      -+  does not support the required interfaces, this falls back to fsync().\n>      ++* `writeout-only` issues pagecache writeback requests, but depending on the\n>      ++  filesystem and storage hardware, data added to the repository may not be\n>      ++  durable in the event of a system crash. This is the default mode on macOS.\n>       +\n>        core.fsyncObjectFiles::\n>        \tThis boolean will enable 'fsync()' when writing object files.\n>        +\n>       \n>      + ## Makefile ##\n>      +@@ Makefile: all::\n>      + #\n>      + # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n>      + #\n>      ++# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n>      ++#\n>      + # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n>      + # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n>      + #\n>      +@@ Makefile: ifdef HAVE_CLOCK_MONOTONIC\n>      + \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n>      + endif\n>      + \n>      ++ifdef HAVE_SYNC_FILE_RANGE\n>      ++\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n>      ++endif\n>      ++\n>      + ifdef NEEDS_LIBRT\n>      + \tEXTLIBS += -lrt\n>      + endif\n>      +\n>        ## cache.h ##\n>       @@ cache.h: extern int read_replace_refs;\n>        extern char *git_replace_ref_base;\n>  2:  23311a10142 ! 2:  ff80a94bf9a core.fsync: introduce granular fsync control\n>      @@ Commit message\n>           knob which can be used to control how components of the\n>           repository are made durable on disk.\n>       \n>      -    This setting allows future extensibility of components\n>      -    that could be synced in two ways:\n>      +    This setting allows future extensibility of the list of\n>      +    syncable components:\n>           * We issue a warning rather than an error for unrecognized\n>             components, so new configs can be used with old Git versions.\n>           * We support negation, so users can choose one of the default\n>             aggregate options and then remove components that they don't\n>      -      want.\n>      +      want. The user would then harden any new components added in\n>      +      a Git version update.\n>       \n>      -    This also support the common request of doing absolutely no\n>      +    This also supports the common request of doing absolutely no\n>           fysncing with the `core.fsync=none` value, which is expected\n>           to make the test suite faster.\n>       \n>      -    This commit introduces the new ability for the user to harden\n>      -    the index, which is a requirement for being able to actually\n>      -    find a file that has been added to the repo and then deleted\n>      -    from the working tree.\n>      -\n>           Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>       \n>        ## Documentation/config/core.txt ##\n>      @@ Documentation/config/core.txt: core.whitespace::\n>       +\thardened via the core.fsyncMethod when created or modified. You can\n>       +\tdisable hardening of any component by prefixing it with a '-'. Later\n>       +\titems take precedence over earlier ones in the list. For example,\n>      -+\t`core.fsync=all,-index` means \"harden everything except the index\".\n>      -+\tItems that are not hardened may be lost in the event of an unclean\n>      -+\tsystem shutdown.\n>      ++\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n>      ++\tmetadata.\" Items that are not hardened may be lost in the event of an\n>      ++\tunclean system shutdown.\n>       ++\n>       +* `none` disables fsync completely. This must be specified alone.\n>       +* `loose-object` hardens objects added to the repo in loose-object form.\n>       +* `pack` hardens objects added to the repo in packfile form.\n>       +* `pack-metadata` hardens packfile bitmaps and indexes.\n>       +* `commit-graph` hardens the commit graph file.\n>      -+* `index` hardens the index when it is modified.\n>       +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n>       +  `pack-metadata`, and `commit-graph`.\n>       +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n>      @@ Documentation/config/core.txt: core.whitespace::\n>        \tA value indicating the strategy Git will use to harden repository data\n>        \tusing fsync and related primitives.\n>       @@ Documentation/config/core.txt: core.fsyncMethod::\n>      -   hardware, but does not send any FLUSH CACHE request. If the operating system\n>      -   does not support the required interfaces, this falls back to fsync().\n>      +   filesystem and storage hardware, data added to the repository may not be\n>      +   durable in the event of a system crash. This is the default mode on macOS.\n>        \n>       -core.fsyncObjectFiles::\n>       -\tThis boolean will enable 'fsync()' when writing object files.\n>      @@ builtin/fast-import.c: static void end_packfile(void)\n>        \n>        \t\tclose_pack_windows(pack_data);\n>       -\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n>      -+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, REPO_COMPONENT_PACK, 0);\n>      ++\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, FSYNC_COMPONENT_PACK, 0);\n>        \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n>        \t\t\t\t\t pack_data->pack_name, object_count,\n>        \t\t\t\t\t cur_pack_oid.hash, pack_size);\n>      @@ builtin/index-pack.c: static void conclude_pack(int fix_thin_pack, const char *c\n>        \t\tstop_progress_msg(&progress, msg.buf);\n>        \t\tstrbuf_release(&msg);\n>       -\t\tfinalize_hashfile(f, tail_hash, 0);\n>      -+\t\tfinalize_hashfile(f, tail_hash, REPO_COMPONENT_PACK, 0);\n>      ++\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n>        \t\thashcpy(read_hash, pack_hash);\n>        \t\tfixup_pack_header_footer(output_fd, pack_hash,\n>        \t\t\t\t\t curr_pack, nr_objects,\n>      @@ builtin/index-pack.c: static void final(const char *final_pack_name, const char\n>        \t\tclose(input_fd);\n>        \t} else {\n>       -\t\tfsync_or_die(output_fd, curr_pack_name);\n>      -+\t\tfsync_component_or_die(REPO_COMPONENT_PACK, output_fd, curr_pack_name);\n>      ++\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n>        \t\terr = close(output_fd);\n>        \t\tif (err)\n>        \t\t\tdie_errno(_(\"error while closing pack file\"));\n>      @@ builtin/pack-objects.c: static void write_pack_file(void)\n>        \t\t */\n>        \t\tif (pack_to_stdout) {\n>       -\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>      -+\t\t\tfinalize_hashfile(f, hash, REPO_COMPONENT_NONE,\n>      ++\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n>       +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>        \t\t} else if (nr_written == nr_remaining) {\n>       -\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n>      -+\t\t\tfinalize_hashfile(f, hash, REPO_COMPONENT_PACK,\n>      ++\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_PACK,\n>       +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n>        \t\t} else {\n>       -\t\t\tint fd = finalize_hashfile(f, hash, 0);\n>      -+\t\t\tint fd = finalize_hashfile(f, hash, REPO_COMPONENT_PACK, 0);\n>      ++\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n>        \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n>        \t\t\t\t\t\t nr_written, hash, offset);\n>        \t\t\tclose(fd);\n>      @@ bulk-checkin.c: static void finish_bulk_checkin(struct bulk_checkin_state *state\n>        \t\tgoto clear_exit;\n>        \t} else if (state->nr_written == 1) {\n>       -\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n>      -+\t\tfinalize_hashfile(state->f, hash, REPO_COMPONENT_PACK,\n>      ++\t\tfinalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK,\n>       +\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n>        \t} else {\n>       -\t\tint fd = finalize_hashfile(state->f, hash, 0);\n>      -+\t\tint fd = finalize_hashfile(state->f, hash, REPO_COMPONENT_PACK, 0);\n>      ++\t\tint fd = finalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK, 0);\n>        \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n>        \t\t\t\t\t state->nr_written, hash,\n>        \t\t\t\t\t state->offset);\n>      @@ cache.h: void reset_shared_repository(void);\n>       -extern int fsync_object_files;\n>       +/*\n>       + * These values are used to help identify parts of a repository to fsync.\n>      -+ * REPO_COMPONENT_NONE identifies data that will not be a persistent part of the\n>      ++ * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n>       + * repository and so shouldn't be fsynced.\n>       + */\n>      -+enum repo_component {\n>      -+\tREPO_COMPONENT_NONE\t\t\t= 0,\n>      -+\tREPO_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n>      -+\tREPO_COMPONENT_PACK\t\t\t= 1 << 1,\n>      -+\tREPO_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n>      -+\tREPO_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n>      -+\tREPO_COMPONENT_INDEX\t\t\t= 1 << 4,\n>      ++enum fsync_component {\n>      ++\tFSYNC_COMPONENT_NONE\t\t\t= 0,\n>      ++\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n>      ++\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n>      ++\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n>      ++\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n>       +};\n>       +\n>      -+#define FSYNC_COMPONENTS_DEFAULT (REPO_COMPONENT_PACK | \\\n>      -+\t\t\t\t  REPO_COMPONENT_PACK_METADATA | \\\n>      -+\t\t\t\t  REPO_COMPONENT_COMMIT_GRAPH)\n>      ++#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n>      ++\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n>      ++\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n>       +\n>      -+#define FSYNC_COMPONENTS_OBJECTS (REPO_COMPONENT_LOOSE_OBJECT | \\\n>      -+\t\t\t\t  REPO_COMPONENT_PACK | \\\n>      -+\t\t\t\t  REPO_COMPONENT_PACK_METADATA | \\\n>      -+\t\t\t\t  REPO_COMPONENT_COMMIT_GRAPH)\n>      ++#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n>      ++\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n>      ++\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n>      ++\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n>       +\n>      -+#define FSYNC_COMPONENTS_ALL (REPO_COMPONENT_LOOSE_OBJECT | \\\n>      -+\t\t\t      REPO_COMPONENT_PACK | \\\n>      -+\t\t\t      REPO_COMPONENT_PACK_METADATA | \\\n>      -+\t\t\t      REPO_COMPONENT_COMMIT_GRAPH | \\\n>      -+\t\t\t      REPO_COMPONENT_INDEX)\n>      ++#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n>      ++\t\t\t      FSYNC_COMPONENT_PACK | \\\n>      ++\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n>      ++\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n>       +\n>       +\n>       +/*\n>       + * A bitmask indicating which components of the repo should be fsynced.\n>       + */\n>      -+extern enum repo_component fsync_components;\n>      ++extern enum fsync_component fsync_components;\n>        \n>        enum fsync_method {\n>        \tFSYNC_METHOD_FSYNC,\n>      @@ cache.h: int copy_file_with_time(const char *dst, const char *src, int mode);\n>        void write_or_die(int fd, const void *buf, size_t count);\n>        void fsync_or_die(int fd, const char *);\n>        \n>      -+inline void fsync_component_or_die(enum repo_component component, int fd, const char *msg)\n>      ++inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n>       +{\n>       +\tif (fsync_components & component)\n>       +\t\tfsync_or_die(fd, msg);\n>      @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_con\n>        \n>        \tclose_commit_graph(ctx->r->objects);\n>       -\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n>      -+\tfinalize_hashfile(f, file_hash, REPO_COMPONENT_COMMIT_GRAPH,\n>      ++\tfinalize_hashfile(f, file_hash, FSYNC_COMPONENT_COMMIT_GRAPH,\n>       +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n>        \tfree_chunkfile(cf);\n>        \n>      @@ config.c: static int git_parse_maybe_bool_text(const char *value)\n>        \n>       +static const struct fsync_component_entry {\n>       +\tconst char *name;\n>      -+\tenum repo_component component_bits;\n>      ++\tenum fsync_component component_bits;\n>       +} fsync_component_table[] = {\n>      -+\t{ \"loose-object\", REPO_COMPONENT_LOOSE_OBJECT },\n>      -+\t{ \"pack\", REPO_COMPONENT_PACK },\n>      -+\t{ \"pack-metadata\", REPO_COMPONENT_PACK_METADATA },\n>      -+\t{ \"commit-graph\", REPO_COMPONENT_COMMIT_GRAPH },\n>      -+\t{ \"index\", REPO_COMPONENT_INDEX },\n>      ++\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n>      ++\t{ \"pack\", FSYNC_COMPONENT_PACK },\n>      ++\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n>      ++\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n>       +\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n>       +\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n>       +\t{ \"all\", FSYNC_COMPONENTS_ALL },\n>       +};\n>       +\n>      -+static enum repo_component parse_fsync_components(const char *var, const char *string)\n>      ++static enum fsync_component parse_fsync_components(const char *var, const char *string)\n>       +{\n>      -+\tenum repo_component output = 0;\n>      ++\tenum fsync_component output = 0;\n>       +\n>       +\tif (!strcmp(string, \"none\"))\n>       +\t\treturn output;\n>      @@ csum-file.c: static void free_hashfile(struct hashfile *f)\n>        \n>       -int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n>       +int finalize_hashfile(struct hashfile *f, unsigned char *result,\n>      -+\t\t      enum repo_component component, unsigned int flags)\n>      ++\t\t      enum fsync_component component, unsigned int flags)\n>        {\n>        \tint fd;\n>        \n>      @@ csum-file.c: int finalize_hashfile(struct hashfile *f, unsigned char *result, un\n>        \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\n>       \n>        ## csum-file.h ##\n>      +@@\n>      + #ifndef CSUM_FILE_H\n>      + #define CSUM_FILE_H\n>      + \n>      ++#include \"cache.h\"\n>      + #include \"hash.h\"\n>      + \n>      + struct progress;\n>       @@ csum-file.h: int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n>        struct hashfile *hashfd(int fd, const char *name);\n>        struct hashfile *hashfd_check(const char *name);\n>        struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n>       -int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n>      -+int finalize_hashfile(struct hashfile *, unsigned char *, enum repo_component, unsigned int);\n>      ++int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n>        void hashwrite(struct hashfile *, const void *, unsigned int);\n>        void hashflush(struct hashfile *f);\n>        void crc32_begin(struct hashfile *);\n>      @@ environment.c: const char *git_hooks_path;\n>        int zlib_compression_level = Z_BEST_SPEED;\n>        int pack_compression_level = Z_DEFAULT_COMPRESSION;\n>        enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n>      -+enum repo_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n>      ++enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n>        size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n>        size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n>        size_t delta_base_cache_limit = 96 * 1024 * 1024;\n>      @@ midx.c: static int write_midx_internal(const char *object_dir,\n>        \twrite_chunkfile(cf, &ctx);\n>        \n>       -\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n>      -+\tfinalize_hashfile(f, midx_hash, REPO_COMPONENT_PACK_METADATA,\n>      ++\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n>       +\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n>        \tfree_chunkfile(cf);\n>        \n>      @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void\n>        {\n>       -\tif (fsync_object_files)\n>       -\t\tfsync_or_die(fd, \"loose object file\");\n>      -+\tfsync_component_or_die(REPO_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n>      ++\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n>        \tif (close(fd) != 0)\n>        \t\tdie_errno(_(\"error when closing loose object file\"));\n>        }\n>      @@ pack-bitmap-write.c: void bitmap_writer_finish(struct pack_idx_entry **index,\n>        \t\twrite_hash_cache(f, index, index_nr);\n>        \n>       -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n>      -+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n>      ++\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n>       +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n>        \n>        \tif (adjust_shared_perm(tmp_file.buf))\n>      @@ pack-write.c: const char *write_idx_file(const char *index_name, struct pack_idx\n>        \thashwrite(f, sha1, the_hash_algo->rawsz);\n>       -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n>       -\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n>      -+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n>      +-\t\t\t\t    ? 0 : CSUM_FSYNC));\n>      ++\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n>       +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n>      -+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY)\n>      - \t\t\t\t    ? 0 : CSUM_FSYNC));\n>      ++\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n>        \treturn index_name;\n>        }\n>      + \n>       @@ pack-write.c: const char *write_rev_file_order(const char *rev_name,\n>        \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n>        \t\tdie(_(\"failed to make %s readable\"), rev_name);\n>        \n>       -\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n>       -\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n>      -+\tfinalize_hashfile(f, NULL, REPO_COMPONENT_PACK_METADATA,\n>      ++\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n>       +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n>       +\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n>        \n>      @@ pack-write.c: void fixup_pack_header_footer(int pack_fd,\n>        \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n>        \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n>       -\tfsync_or_die(pack_fd, pack_name);\n>      -+\tfsync_component_or_die(REPO_COMPONENT_PACK, pack_fd, pack_name);\n>      ++\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n>        }\n>        \n>        char *index_pack_lockfile(int ip_out, int *is_well_formed)\n>       \n>        ## read-cache.c ##\n>      -@@ read-cache.c: static int record_ieot(void)\n>      -  * rely on it.\n>      -  */\n>      - static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>      --\t\t\t  int strip_extensions)\n>      -+\t\t\t  int strip_extensions, unsigned flags)\n>      - {\n>      - \tuint64_t start = getnanotime();\n>      - \tstruct hashfile *f;\n>      -@@ read-cache.c: static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>      - \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>      - \tint drop_cache_tree = istate->drop_cache_tree;\n>      - \toff_t offset;\n>      -+\tint csum_fsync_flag;\n>      - \tint ieot_entries = 1;\n>      - \tstruct index_entry_offset_table *ieot = NULL;\n>      - \tint nr, nr_threads;\n>       @@ read-cache.c: static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>        \t\t\treturn -1;\n>        \t}\n>        \n>       -\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n>      -+\tcsum_fsync_flag = 0;\n>      -+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n>      -+\t\tcsum_fsync_flag = CSUM_FSYNC;\n>      -+\n>      -+\tfinalize_hashfile(f, istate->oid.hash, REPO_COMPONENT_INDEX,\n>      -+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n>      -+\n>      ++\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n>        \tif (close_tempfile_gently(tempfile)) {\n>        \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n>        \t\treturn -1;\n>      -@@ read-cache.c: static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n>      - \t */\n>      - \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n>      - \t\t\t\t   \"%s\", get_lock_file_path(lock));\n>      --\tret = do_write_index(istate, lock->tempfile, 0);\n>      -+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n>      - \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n>      - \t\t\t\t   \"%s\", get_lock_file_path(lock));\n>      - \n>      -@@ read-cache.c: static int clean_shared_index_files(const char *current_hex)\n>      - }\n>      - \n>      - static int write_shared_index(struct index_state *istate,\n>      --\t\t\t      struct tempfile **temp)\n>      -+\t\t\t      struct tempfile **temp, unsigned flags)\n>      - {\n>      - \tstruct split_index *si = istate->split_index;\n>      - \tint ret, was_full = !istate->sparse_index;\n>      -@@ read-cache.c: static int write_shared_index(struct index_state *istate,\n>      - \n>      - \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n>      - \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n>      --\tret = do_write_index(si->base, *temp, 1);\n>      -+\tret = do_write_index(si->base, *temp, 1, flags);\n>      - \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n>      - \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n>      - \n>      -@@ read-cache.c: int write_locked_index(struct index_state *istate, struct lock_file *lock,\n>      - \t\t\tret = do_write_locked_index(istate, lock, flags);\n>      - \t\t\tgoto out;\n>      - \t\t}\n>      --\t\tret = write_shared_index(istate, &temp);\n>      -+\t\tret = write_shared_index(istate, &temp, flags);\n>      - \n>      - \t\tsaved_errno = errno;\n>      - \t\tif (is_tempfile_active(temp))\n>  -:  ----------- > 3:  86e39b8f8d1 core.fsync: new option to harden the index\n> \n> -- \n> gitgitgadget\n"},{"id":"443271","messageId":"211207.865ys0pq4x.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"Ya9JJlItvDJCLHqj@ncase","subject":"Re: [PATCH v2 1/3] core.fsyncmethod: add writeout-only mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-07T12:14:45Z","receivedAt":"2021-12-07T12:15:46Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 07 2021, Patrick Steinhardt wrote:\n\n> [[PGP Signed Part:Undecided]]\n> On Tue, Dec 07, 2021 at 02:46:49AM +0000, Neeraj Singh via GitGitGadget wrote:\n>> From: Neeraj Singh <neerajsi@microsoft.com>\n> [...]\n> [snip]\n>> diff --git a/wrapper.c b/wrapper.c\n>> index 36e12119d76..1c5f2c87791 100644\n>> --- a/wrapper.c\n>> +++ b/wrapper.c\n>> @@ -546,6 +546,62 @@ int xmkstemp_mode(char *filename_template, int mode)\n>>  \treturn fd;\n>>  }\n>>  \n>> +int git_fsync(int fd, enum fsync_action action)\n>> +{\n>> +\tswitch (action) {\n>> +\tcase FSYNC_WRITEOUT_ONLY:\n>> +\n>> +#ifdef __APPLE__\n>> +\t\t/*\n>> +\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n>> +\t\t * flush hardware caches.\n>> +\t\t */\n>> +\t\treturn fsync(fd);\n>\n> Below we're looping around `EINTR` -- are Apple systems never returning\n> it?\n\nI think so per cccdfd22436 (fsync(): be prepared to see EINTR,\n2021-06-04), but I'm not sure, but in any case it would make sense for\nthis to just call the same loop we've been calling elsewhere. It doesn't\nseem to hurt, so we can just do that part in the portable portion of the\ncode.\n"},{"id":"443272","messageId":"211207.861r2opplg.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"e79522cbdd4feb45b062b75225475f34039d1866.1638845211.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 1/3] core.fsyncmethod: add writeout-only mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-07T12:18:31Z","receivedAt":"2021-12-07T12:27:29Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> This commit introduces the `core.fsyncmethod` configuration\n\nJust a commit msg nit: core.fsyncMethod (I see the docs etc. are using\nit camelCased, good..\n\n> diff --git a/compat/win32/flush.c b/compat/win32/flush.c\n> new file mode 100644\n> index 00000000000..75324c24ee7\n> --- /dev/null\n> +++ b/compat/win32/flush.c\n> @@ -0,0 +1,28 @@\n> +#include \"../../git-compat-util.h\"\n\nnit: Just FWIW I think the better thing is '#include\n\"git-compat-util.h\"', i.e. we're compiling at the top-level and have\nadded it to -I.\n\n(I know a lot of compat/ and contrib/ and even main-tree stuff does\nthat, but just FWIW it's not needed).\n\n> +\tif (!strcmp(var, \"core.fsyncmethod\")) {\n> +\t\tif (!value)\n> +\t\t\treturn config_error_nonbool(var);\n> +\t\tif (!strcmp(value, \"fsync\"))\n> +\t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n> +\t\telse if (!strcmp(value, \"writeout-only\"))\n> +\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n> +\t\telse\n\nAs a non-nit comment I think this config schema looks great so far.\n\n> +\t\t\twarning(_(\"unknown %s value '%s'\"), var, value);\n\nJust a suggestion maybe something slightly scarier like:\n\n    \"unknown core.fsyncMethod value '%s'; config from future git version? ignoring requested fsync strategy\"\n\nAlso using the nicer camelCased version instead of \"var\" (also helps\ntranslators with context...)\n\n> +int git_fsync(int fd, enum fsync_action action)\n> +{\n> +\tswitch (action) {\n> +\tcase FSYNC_WRITEOUT_ONLY:\n> +\n> +#ifdef __APPLE__\n> +\t\t/*\n> +\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n> +\t\t * flush hardware caches.\n> +\t\t */\n> +\t\treturn fsync(fd);\n> +#endif\n> +\n> +#ifdef HAVE_SYNC_FILE_RANGE\n> +\t\t/*\n> +\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n> +\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n> +\t\t * indicates writeout of the entire file and the wait flags ensure that all\n> +\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n> +\t\t * before we continue.\n> +\t\t */\n> +\n> +\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n> +#endif\n> +\n> +#ifdef fsync_no_flush\n> +\t\treturn fsync_no_flush(fd);\n> +#endif\n> +\n> +\t\terrno = ENOSYS;\n> +\t\treturn -1;\n> +\n> +\tcase FSYNC_HARDWARE_FLUSH:\n> +\t\t/*\n> +\t\t * On some platforms fsync may return EINTR. Try again in this\n> +\t\t * case, since callers asking for a hardware flush may die if\n> +\t\t * this function returns an error.\n> +\t\t */\n> +\t\tfor (;;) {\n> +\t\t\tint err;\n> +#ifdef __APPLE__\n> +\t\t\terr = fcntl(fd, F_FULLFSYNC);\n> +#else\n> +\t\t\terr = fsync(fd);\n> +#endif\n> +\t\t\tif (err >= 0 || errno != EINTR)\n> +\t\t\t\treturn err;\n> +\t\t}\n> +\n> +\tdefault:\n> +\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n\nDon't include such \"default\" cases, you have an exhaustive \"enum\", if\nyou skip it the compiler will check this for you.\n\n> +\t}\n> +}\n> +\n>  static int warn_if_unremovable(const char *op, const char *file, int rc)\n\nJust a code nit: I think it's very much preferred if possible to have as\nmuch of code like this compile on all platforms. See the series at\n4002e87cb25 (grep: remove #ifdef NO_PTHREADS, 2018-11-03) is part of for\na good example.\n\nMaybe not worth it in this case since they're not nested ifdef's.\n\nI'm basically thinking of something (also re Patrick's comment on the\n2nd patch) where we have a platform_fsync() whose return\nvalue/arguments/whatever capture this \"I want to return now\" or \"you\nshould be looping\" and takes the enum_fsync_action\" strategy.\n\nThen the git_fsync() would be the platform-independent looping etc., and\nanother funciton would do the \"one fsync at a time, maybe call me\nagain\".\n\nMaybe it would suck more, just food for thought... :)\n"},{"id":"443273","messageId":"211207.86wnkgo9fv.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"ff80a94bf9add8a6fabcd5146e5177edf5e35e49.1638845211.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-07T12:29:04Z","receivedAt":"2021-12-07T13:01:49Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> This commit introduces the `core.fsync` configuration\n> knob which can be used to control how components of the\n> repository are made durable on disk.\n>\n> This setting allows future extensibility of the list of\n> syncable components:\n> * We issue a warning rather than an error for unrecognized\n>   components, so new configs can be used with old Git versions.\n\nLooks good!\n\n> * We support negation, so users can choose one of the default\n>   aggregate options and then remove components that they don't\n>   want. The user would then harden any new components added in\n>   a Git version update.\n\nI think this config schema makes sense, but just a (I think important)\ncomment on the \"how\" not \"what\" of it. It's really much better to define\nconfig as:\n\n    [some]\n    key = value\n    key = value2\n\nThan:\n\n    [some]\n    key = value,value2\n\nThe reason is that \"git config\" has good support for working with\nmulti-valued stuff, so you can do e.g.:\n\n    git config --get-all -z some.key\n\nAnd you can easily (re)set such config e.g. with --replace-all etc., but\nfor comma-delimited you (and users) need to do all that work themselves.\n\nSimilarly instead of:\n\n    some.key = want-this\n    some.key = -not-this\n    some.key = but-want-this\n\nI think it's better to just have two lists, one inclusive another\nexclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n\"transfer.hideRefs\"\n\nWhich would mean:\n\n    core.fsync = want-this\n    core.fsyncExcludes = -not-this\n\nFor some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\nsuggestion on making this easier for users & the implementation.\n\n> This also supports the common request of doing absolutely no\n> fysncing with the `core.fsync=none` value, which is expected\n> to make the test suite faster.\n\nLet's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\nso we'll accept \"false\", \"off\", \"no\" like most other such config?\n\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  Documentation/config/core.txt | 27 +++++++++----\n>  builtin/fast-import.c         |  2 +-\n>  builtin/index-pack.c          |  4 +-\n>  builtin/pack-objects.c        |  8 ++--\n>  bulk-checkin.c                |  5 ++-\n>  cache.h                       | 39 +++++++++++++++++-\n>  commit-graph.c                |  3 +-\n>  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n>  csum-file.c                   |  5 ++-\n>  csum-file.h                   |  3 +-\n>  environment.c                 |  1 +\n>  midx.c                        |  3 +-\n>  object-file.c                 |  3 +-\n>  pack-bitmap-write.c           |  3 +-\n>  pack-write.c                  | 13 +++---\n>  read-cache.c                  |  2 +-\n>  16 files changed, 164 insertions(+), 33 deletions(-)\n>\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index dbb134f7136..4f1747ec871 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -547,6 +547,25 @@ core.whitespace::\n>    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n>    errors. The default tab width is 8. Allowed values are 1 to 63.\n>  \n> +core.fsync::\n> +\tA comma-separated list of parts of the repository which should be\n> +\thardened via the core.fsyncMethod when created or modified. You can\n> +\tdisable hardening of any component by prefixing it with a '-'. Later\n> +\titems take precedence over earlier ones in the list. For example,\n> +\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n> +\tmetadata.\" Items that are not hardened may be lost in the event of an\n> +\tunclean system shutdown.\n> ++\n> +* `none` disables fsync completely. This must be specified alone.\n> +* `loose-object` hardens objects added to the repo in loose-object form.\n> +* `pack` hardens objects added to the repo in packfile form.\n> +* `pack-metadata` hardens packfile bitmaps and indexes.\n> +* `commit-graph` hardens the commit graph file.\n> +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n> +  `pack-metadata`, and `commit-graph`.\n> +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n> +* `all` is an aggregate option that syncs all individual components above.\n> +\n\nIt's probably a *bit* more work to set up, but I wonder if this wouldn't\nbe simpler if we just said (and this is partially going against what I\nnoted above):\n\n== BEGIN DOC\n\ncore.fsync is a multi-value config variable where each item is a\npathspec that'll get matched the same way as 'git-ls-files' et al.\n\nWhen we sync pretend that a path like .git/objects/de/adbeef... is\nrelative to the top-level of the git\ndirectory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n\nYou can then supply a list of wildcards and exclusions to configure\nsyncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\nwith:\n\n    ; same as \"false\"\n    core.fsync = \":!*\"\n\nOr:\n\n    ; same as \"true\"\n    core.fsync = \"*\"\n\nOr, to selectively sync some things and not others:\n\n    ;; Sync objects, but not \"info\"\n    core.fsync = \":!objects/info/**\"\n    core.fsync = \"objects/**\"\n\nSee gitrepository-layout(5) for details about what sort of paths you\nmight be expected to match. Not all paths listed there will go through\nthis mechanism (e.g. currently objects do, but nothing to do with config\ndoes).\n\nWe can and will match this against \"fake paths\", e.g. when writing out\npacks we may match against just the string \"objects/pack\", we're not\ngoing to re-check if every packfile we're writing matches your globs,\nditto for loose objects. Be reasonable!\n\nThis metharism is intended as a shorthand that provides some flexibility\nwhen fsyncing, while not forcing git to come up with labels for all\npaths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n\nMore paths may be added or removed in the future, and we make no\npromises that we won't move things around, so if in doubt use\ne.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\nstick to the golden path of examples provided in this documentation.\n\n== END DOC\n\n\nIt's a tad more complex to set up, but I wonder if that isn't worth\nit. It nicely gets around any current and future issues of deciding what\nlabels such as \"loose-object\" etc. to pick, as well as slotting into an\nexisting method of doing exclude/include lists.\n\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> index 857be7826f3..916c55d6ce9 100644\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n> @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n>  \t\t * If so, rewrite it like in fast-import\n>  \t\t */\n>  \t\tif (pack_to_stdout) {\n> -\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> +\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n> +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n\nNot really related to this per-se, but since you're touching the API\neverything goes through I wonder if callers should just always try to\nfsync, and we can just catch EROFS and EINVAL in the wrapper if someone\ntries to flush stdout, or catch the fd at that lower level.\n\nOr maybe there's a good reason for this...\n\n> [...]\n> +/*\n> + * These values are used to help identify parts of a repository to fsync.\n> + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n> + * repository and so shouldn't be fsynced.\n> + */\n> +enum fsync_component {\n> +\tFSYNC_COMPONENT_NONE\t\t\t= 0,\n\nI haven't read ahead much but in most other such cases we don't define\nthe \"= 0\", just start at 1<<0, then check the flags elsewhere...\n\n> +static const struct fsync_component_entry {\n> +\tconst char *name;\n> +\tenum fsync_component component_bits;\n> +} fsync_component_table[] = {\n> +\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> +\t{ \"pack\", FSYNC_COMPONENT_PACK },\n> +\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> +\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> +\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> +\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n> +\t{ \"all\", FSYNC_COMPONENTS_ALL },\n> +};\n> +\n> +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> +{\n> +\tenum fsync_component output = 0;\n> +\n> +\tif (!strcmp(string, \"none\"))\n> +\t\treturn output;\n> +\n> +\twhile (string) {\n> +\t\tint i;\n> +\t\tsize_t len;\n> +\t\tconst char *ep;\n> +\t\tint negated = 0;\n> +\t\tint found = 0;\n> +\n> +\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n\nAside from the \"use a list\" isn't this hardcoding some windows-specific\nassumptions with \\n\\r? Maybe not...\n"},{"id":"443366","messageId":"CANQDOddgwc4X2snAGvoz0KFaCWcx+k2U3qfoLywcUS=6yF=Htg@mail.gmail.com","threadId":"57030","inReplyTo":"Ya9LPOseu8geBi4v@ncase","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-07T20:46:54Z","receivedAt":"2021-12-07T20:47:08Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Dec 7, 2021 at 3:54 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Tue, Dec 07, 2021 at 02:46:50AM +0000, Neeraj Singh via GitGitGadget wrote:\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> [snip]\n> > diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> > index c23d01de7dc..c32534c13b4 100644\n> > --- a/builtin/index-pack.c\n> > +++ b/builtin/index-pack.c\n> > @@ -1286,7 +1286,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n> >                           nr_objects - nr_objects_initial);\n> >               stop_progress_msg(&progress, msg.buf);\n> >               strbuf_release(&msg);\n> > -             finalize_hashfile(f, tail_hash, 0);\n> > +             finalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n> >               hashcpy(read_hash, pack_hash);\n> >               fixup_pack_header_footer(output_fd, pack_hash,\n> >                                        curr_pack, nr_objects,\n> > @@ -1508,7 +1508,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n> >       if (!from_stdin) {\n> >               close(input_fd);\n> >       } else {\n> > -             fsync_or_die(output_fd, curr_pack_name);\n> > +             fsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n> >               err = close(output_fd);\n> >               if (err)\n> >                       die_errno(_(\"error while closing pack file\"));\n> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> > index 857be7826f3..916c55d6ce9 100644\n> > --- a/builtin/pack-objects.c\n> > +++ b/builtin/pack-objects.c\n> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n> >                * If so, rewrite it like in fast-import\n> >                */\n> >               if (pack_to_stdout) {\n> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>\n> It doesn't have any effect here given that we don't sync at all when\n> writing to stdout, but I wonder whether we should set up the component\n> correctly regardless of that such that it makes for a less confusing\n> read.\n>\n\nIf it's not actually a file with some name known to git, is it really\na component of the repository? I'd like to leave this one as-is.\n\n> [snip]\n> > diff --git a/config.c b/config.c\n> > index c3410b8a868..29c867aab03 100644\n> > --- a/config.c\n> > +++ b/config.c\n> > @@ -1213,6 +1213,73 @@ static int git_parse_maybe_bool_text(const char *value)\n> >       return -1;\n> >  }\n> >\n> > +static const struct fsync_component_entry {\n> > +     const char *name;\n> > +     enum fsync_component component_bits;\n> > +} fsync_component_table[] = {\n> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n> > +};\n> > +\n> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> > +{\n> > +     enum fsync_component output = 0;\n> > +\n> > +     if (!strcmp(string, \"none\"))\n> > +             return output;\n> > +\n> > +     while (string) {\n> > +             int i;\n> > +             size_t len;\n> > +             const char *ep;\n> > +             int negated = 0;\n> > +             int found = 0;\n> > +\n> > +             string = string + strspn(string, \", \\t\\n\\r\");\n> > +             ep = strchrnul(string, ',');\n> > +             len = ep - string;\n> > +\n> > +             if (*string == '-') {\n> > +                     negated = 1;\n> > +                     string++;\n> > +                     len--;\n> > +                     if (!len)\n> > +                             warning(_(\"invalid value for variable %s\"), var);\n> > +             }\n> > +\n> > +             if (!len)\n> > +                     break;\n> > +\n> > +             for (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n> > +                     const struct fsync_component_entry *entry = &fsync_component_table[i];\n> > +\n> > +                     if (strncmp(entry->name, string, len))\n> > +                             continue;\n> > +\n> > +                     found = 1;\n> > +                     if (negated)\n> > +                             output &= ~entry->component_bits;\n> > +                     else\n> > +                             output |= entry->component_bits;\n> > +             }\n> > +\n> > +             if (!found) {\n> > +                     char *component = xstrndup(string, len);\n> > +                     warning(_(\"unknown %s value '%s'\"), var, component);\n> > +                     free(component);\n> > +             }\n> > +\n> > +             string = ep;\n> > +     }\n> > +\n> > +     return output;\n> > +}\n> > +\n> >  int git_parse_maybe_bool(const char *value)\n> >  {\n> >       int v = git_parse_maybe_bool_text(value);\n> > @@ -1490,6 +1557,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >               return 0;\n> >       }\n> >\n> > +     if (!strcmp(var, \"core.fsync\")) {\n> > +             if (!value)\n> > +                     return config_error_nonbool(var);\n> > +             fsync_components = parse_fsync_components(var, value);\n> > +             return 0;\n> > +     }\n> > +\n> >       if (!strcmp(var, \"core.fsyncmethod\")) {\n> >               if (!value)\n> >                       return config_error_nonbool(var);\n> > @@ -1503,7 +1577,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >       }\n> >\n> >       if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> > -             fsync_object_files = git_config_bool(var, value);\n> > +             warning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n> >               return 0;\n> >       }\n>\n> Shouldn't we continue to support this for now such that users can\n> migrate from the old, deprecated value first before we start to ignore\n> it?\n>\n> Patrick\n>\n\nThat's a good question and one I was hoping to answer through this\ndiscussion.  I'm guessing that most users do not have this setting\nexplicitly set, and it's largely a non-functional change. Git only\nbehaves differently in the extreme corner case of a system crash after\ngit exits.  That's why I believe it's okay to deprecate and remove in\none release.\n\nIf we choose to keep supporting the setting, it would introduce a\nlittle complexity in the configuration code, but it should be doable.\nI think the right semantics would be to ignore core.fsyncobjectfiles\nif core.fsync is specified with any value. Otherwise,\ncore.fsyncobjectfiles would be equivalent to `default,-loose-object`\nor `default,loose-object` for `false` and `true` respectively. I'd\nprefer not to do this if I can get away with it :).\n\nThanks,\nNeeraj\n"},{"id":"443375","messageId":"CANQDOdfX2KaosPwLM4hS4rp+FH9V7VUxUh_md43FfZ9NG4iroQ@mail.gmail.com","threadId":"57030","inReplyTo":"211207.86wnkgo9fv.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-07T21:44:11Z","receivedAt":"2021-12-07T21:44:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Dec 7, 2021 at 5:01 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > This commit introduces the `core.fsync` configuration\n> > knob which can be used to control how components of the\n> > repository are made durable on disk.\n> >\n> > This setting allows future extensibility of the list of\n> > syncable components:\n> > * We issue a warning rather than an error for unrecognized\n> >   components, so new configs can be used with old Git versions.\n>\n> Looks good!\n>\n> > * We support negation, so users can choose one of the default\n> >   aggregate options and then remove components that they don't\n> >   want. The user would then harden any new components added in\n> >   a Git version update.\n>\n> I think this config schema makes sense, but just a (I think important)\n> comment on the \"how\" not \"what\" of it. It's really much better to define\n> config as:\n>\n>     [some]\n>     key = value\n>     key = value2\n>\n> Than:\n>\n>     [some]\n>     key = value,value2\n>\n> The reason is that \"git config\" has good support for working with\n> multi-valued stuff, so you can do e.g.:\n>\n>     git config --get-all -z some.key\n>\n> And you can easily (re)set such config e.g. with --replace-all etc., but\n> for comma-delimited you (and users) need to do all that work themselves.\n>\n> Similarly instead of:\n>\n>     some.key = want-this\n>     some.key = -not-this\n>     some.key = but-want-this\n>\n> I think it's better to just have two lists, one inclusive another\n> exclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n> \"transfer.hideRefs\"\n>\n> Which would mean:\n>\n>     core.fsync = want-this\n>     core.fsyncExcludes = -not-this\n>\n> For some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\n> suggestion on making this easier for users & the implementation.\n>\n\nMaybe there's some way to handle this I'm unaware of, but a\ndisadvantage of your multi-valued config proposal is that it's harder,\nfor example, for a per-repo config store to reasonably override a\nper-user config store.  With the configuration scheme as-is, I can\nhave a per-user setting like `core.fsync=all` which covers my typical\nrepos, but then have a maintainer repo with a private setting of\n`core.fsync=none` to speed up cases where I'm mostly working with\nother people's changes that are backed up in email or server-side\nrepos.  The latter setting conveniently overrides the former setting\nin all aspects.\n\nAlso, with the core.fsync and core.fsyncExcludes, how would you spell\n\"don't sync anything\"? Would you still have the aggregate options.?\n\n> > This also supports the common request of doing absolutely no\n> > fysncing with the `core.fsync=none` value, which is expected\n> > to make the test suite faster.\n>\n> Let's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\n> so we'll accept \"false\", \"off\", \"no\" like most other such config?\n\nJunio's previous feedback when discussing batch mode [1] was to offer\nless flexibility when parsing new values of these configuration\noptions. I agree with his statement that \"making machine-readable\ntokens be spelled in different ways is a 'disease'.\"  I'd like to\nleave this as-is so that the documentation can clearly state the exact\nset of allowable values.\n\n[1] https://lore.kernel.org/git/xmqqr1dqzyl7.fsf@gitster.g/\n\n>\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  Documentation/config/core.txt | 27 +++++++++----\n> >  builtin/fast-import.c         |  2 +-\n> >  builtin/index-pack.c          |  4 +-\n> >  builtin/pack-objects.c        |  8 ++--\n> >  bulk-checkin.c                |  5 ++-\n> >  cache.h                       | 39 +++++++++++++++++-\n> >  commit-graph.c                |  3 +-\n> >  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n> >  csum-file.c                   |  5 ++-\n> >  csum-file.h                   |  3 +-\n> >  environment.c                 |  1 +\n> >  midx.c                        |  3 +-\n> >  object-file.c                 |  3 +-\n> >  pack-bitmap-write.c           |  3 +-\n> >  pack-write.c                  | 13 +++---\n> >  read-cache.c                  |  2 +-\n> >  16 files changed, 164 insertions(+), 33 deletions(-)\n> >\n> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> > index dbb134f7136..4f1747ec871 100644\n> > --- a/Documentation/config/core.txt\n> > +++ b/Documentation/config/core.txt\n> > @@ -547,6 +547,25 @@ core.whitespace::\n> >    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >\n> > +core.fsync::\n> > +     A comma-separated list of parts of the repository which should be\n> > +     hardened via the core.fsyncMethod when created or modified. You can\n> > +     disable hardening of any component by prefixing it with a '-'. Later\n> > +     items take precedence over earlier ones in the list. For example,\n> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n> > +     metadata.\" Items that are not hardened may be lost in the event of an\n> > +     unclean system shutdown.\n> > ++\n> > +* `none` disables fsync completely. This must be specified alone.\n> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> > +* `pack` hardens objects added to the repo in packfile form.\n> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> > +* `commit-graph` hardens the commit graph file.\n> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n> > +  `pack-metadata`, and `commit-graph`.\n> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n> > +* `all` is an aggregate option that syncs all individual components above.\n> > +\n>\n> It's probably a *bit* more work to set up, but I wonder if this wouldn't\n> be simpler if we just said (and this is partially going against what I\n> noted above):\n>\n> == BEGIN DOC\n>\n> core.fsync is a multi-value config variable where each item is a\n> pathspec that'll get matched the same way as 'git-ls-files' et al.\n>\n> When we sync pretend that a path like .git/objects/de/adbeef... is\n> relative to the top-level of the git\n> directory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n>\n> You can then supply a list of wildcards and exclusions to configure\n> syncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\n> with:\n>\n>     ; same as \"false\"\n>     core.fsync = \":!*\"\n>\n> Or:\n>\n>     ; same as \"true\"\n>     core.fsync = \"*\"\n>\n> Or, to selectively sync some things and not others:\n>\n>     ;; Sync objects, but not \"info\"\n>     core.fsync = \":!objects/info/**\"\n>     core.fsync = \"objects/**\"\n>\n> See gitrepository-layout(5) for details about what sort of paths you\n> might be expected to match. Not all paths listed there will go through\n> this mechanism (e.g. currently objects do, but nothing to do with config\n> does).\n>\n> We can and will match this against \"fake paths\", e.g. when writing out\n> packs we may match against just the string \"objects/pack\", we're not\n> going to re-check if every packfile we're writing matches your globs,\n> ditto for loose objects. Be reasonable!\n>\n> This metharism is intended as a shorthand that provides some flexibility\n> when fsyncing, while not forcing git to come up with labels for all\n> paths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n>\n> More paths may be added or removed in the future, and we make no\n> promises that we won't move things around, so if in doubt use\n> e.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\n> stick to the golden path of examples provided in this documentation.\n>\n> == END DOC\n>\n>\n> It's a tad more complex to set up, but I wonder if that isn't worth\n> it. It nicely gets around any current and future issues of deciding what\n> labels such as \"loose-object\" etc. to pick, as well as slotting into an\n> existing method of doing exclude/include lists.\n>\n\nI think this proposal is a lot of complexity to avoid coming up with a\nnew name for syncable things as they are added to Git.  A path based\nmechanism makes it hard to document for the (advanced) user what the\nfull set of things is and how it might change from release to release.\nI think the current core.fsync scheme is a bit easier to understand,\nquery, and extend.\n\n> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> > index 857be7826f3..916c55d6ce9 100644\n> > --- a/builtin/pack-objects.c\n> > +++ b/builtin/pack-objects.c\n> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n> >                * If so, rewrite it like in fast-import\n> >                */\n> >               if (pack_to_stdout) {\n> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>\n> Not really related to this per-se, but since you're touching the API\n> everything goes through I wonder if callers should just always try to\n> fsync, and we can just catch EROFS and EINVAL in the wrapper if someone\n> tries to flush stdout, or catch the fd at that lower level.\n>\n> Or maybe there's a good reason for this...\n\nIt's platform dependent, but I'd expect fsync would do something for\npipes or stdout redirected to a file.  In these cases we really don't\nwant to fsync since we have no idea what we're talking to and we're\npotentially worsening performance for probably no benefit.\n\n> > [...]\n> > +/*\n> > + * These values are used to help identify parts of a repository to fsync.\n> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n> > + * repository and so shouldn't be fsynced.\n> > + */\n> > +enum fsync_component {\n> > +     FSYNC_COMPONENT_NONE                    = 0,\n>\n> I haven't read ahead much but in most other such cases we don't define\n> the \"= 0\", just start at 1<<0, then check the flags elsewhere...\n>\n> > +static const struct fsync_component_entry {\n> > +     const char *name;\n> > +     enum fsync_component component_bits;\n> > +} fsync_component_table[] = {\n> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n> > +};\n> > +\n> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> > +{\n> > +     enum fsync_component output = 0;\n> > +\n> > +     if (!strcmp(string, \"none\"))\n> > +             return output;\n> > +\n> > +     while (string) {\n> > +             int i;\n> > +             size_t len;\n> > +             const char *ep;\n> > +             int negated = 0;\n> > +             int found = 0;\n> > +\n> > +             string = string + strspn(string, \", \\t\\n\\r\");\n>\n> Aside from the \"use a list\" isn't this hardcoding some windows-specific\n> assumptions with \\n\\r? Maybe not...\n\nI shamelessly stole this code from parse_whitespace_rule. I thought\nabout making a helper to be called by both functions, but the amount\nof state going into and out of the wrapper via arguments was\nsubstantial and seemed to negate the benefit of deduplication.\n\nThanks for the review,\nNeeraj\n"},{"id":"443387","messageId":"CANQDOdd3uWh7SN=GDdG+=rtJA+ja+YdRLsodwSBYMH4JMKj9Zg@mail.gmail.com","threadId":"57030","inReplyTo":"Ya9JJlItvDJCLHqj@ncase","subject":"Re: [PATCH v2 1/3] core.fsyncmethod: add writeout-only mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-07T23:29:10Z","receivedAt":"2021-12-07T23:29:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Dec 7, 2021 at 3:45 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Tue, Dec 07, 2021 at 02:46:49AM +0000, Neeraj Singh via GitGitGadget wrote:\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> [snip]\n> > --- a/compat/mingw.h\n> > +++ b/compat/mingw.h\n> > @@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n> >  #define getpagesize mingw_getpagesize\n> >  #endif\n> >\n> > +int win32_fsync_no_flush(int fd);\n> > +#define fsync_no_flush win32_fsync_no_flush\n> > +\n> >  struct rlimit {\n> >       unsigned int rlim_cur;\n> >  };\n> > diff --git a/compat/win32/flush.c b/compat/win32/flush.c\n> > new file mode 100644\n> > index 00000000000..75324c24ee7\n> > --- /dev/null\n> > +++ b/compat/win32/flush.c\n> > @@ -0,0 +1,28 @@\n> > +#include \"../../git-compat-util.h\"\n> > +#include <winternl.h>\n> > +#include \"lazyload.h\"\n> > +\n> > +int win32_fsync_no_flush(int fd)\n> > +{\n> > +       IO_STATUS_BLOCK io_status;\n> > +\n> > +#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n> > +\n> > +       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n> > +                      HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n> > +                      PIO_STATUS_BLOCK IoStatusBlock);\n> > +\n> > +       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n> > +             errno = ENOSYS;\n> > +             return -1;\n> > +       }\n>\n> I'm wondering whether it would make sense to fall back to fsync(3P) in\n> case we cannot use writeout-only, but I see that were doing essentially\n> that in `fsync_or_die()`. There is no indicator to the user though that\n> writeout-only doesn't work -- do we want to print a one-time warning?\n>\n\nI wanted to leave the fallback to the caller so that the algorithm can\nbe adjusted in some way based on whether writeout-only succeeded.  For\nbatched fsync object files, we refrain from doing the last fsync if we\nwere doing real fsyncs all along.\n\nI didn't want to issue a warning, since this writeout-only codepath\nmight be invoked from multiple subprocesses, which would each\npotentially issue their one warning.  The consequence of failing\nwriteout only is worse performance, but should not be compromised\nsafety, so I'm not sure the user gets enough from the warning to\njustify something that's potentially spammy.  In practice, when batch\nmode is adopted on Windows (by default), some older pre-Win8 systems\nwill experience fsyncs and equivalent performance to what they're\nseeing today. I don't want these users to have a warning too.\n\n> > +       memset(&io_status, 0, sizeof(io_status));\n> > +       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n> > +                             NULL, 0, &io_status)) {\n> > +             errno = EINVAL;\n> > +             return -1;\n> > +       }\n> > +\n> > +       return 0;\n> > +}\n>\n> [snip]\n> > diff --git a/wrapper.c b/wrapper.c\n> > index 36e12119d76..1c5f2c87791 100644\n> > --- a/wrapper.c\n> > +++ b/wrapper.c\n> > @@ -546,6 +546,62 @@ int xmkstemp_mode(char *filename_template, int mode)\n> >       return fd;\n> >  }\n> >\n> > +int git_fsync(int fd, enum fsync_action action)\n> > +{\n> > +     switch (action) {\n> > +     case FSYNC_WRITEOUT_ONLY:\n> > +\n> > +#ifdef __APPLE__\n> > +             /*\n> > +              * on macOS, fsync just causes filesystem cache writeback but does not\n> > +              * flush hardware caches.\n> > +              */\n> > +             return fsync(fd);\n>\n> Below we're looping around `EINTR` -- are Apple systems never returning\n> it?\n>\n\nThe EINTR check was added due to a test failure on HP NonStop.  I\ndon't believe any other platform has actually complained about that.\n\nThanks again for the code review!\n-Neeraj\n"},{"id":"443388","messageId":"CANQDOdf8C4-haK9=Q_J4Cid8bQALnmGDm=SvatRbaVf+tkzqLw@mail.gmail.com","threadId":"57030","inReplyTo":"211207.861r2opplg.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 1/3] core.fsyncmethod: add writeout-only mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-07T23:58:57Z","receivedAt":"2021-12-07T23:59:13Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Dec 7, 2021 at 4:27 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > This commit introduces the `core.fsyncmethod` configuration\n>\n> Just a commit msg nit: core.fsyncMethod (I see the docs etc. are using\n> it camelCased, good..\n\nWill fix.\n\n> > diff --git a/compat/win32/flush.c b/compat/win32/flush.c\n> > new file mode 100644\n> > index 00000000000..75324c24ee7\n> > --- /dev/null\n> > +++ b/compat/win32/flush.c\n> > @@ -0,0 +1,28 @@\n> > +#include \"../../git-compat-util.h\"\n>\n> nit: Just FWIW I think the better thing is '#include\n> \"git-compat-util.h\"', i.e. we're compiling at the top-level and have\n> added it to -I.\n>\n> (I know a lot of compat/ and contrib/ and even main-tree stuff does\n> that, but just FWIW it's not needed).\n>\n\nWill fix.\n\n> > +     if (!strcmp(var, \"core.fsyncmethod\")) {\n> > +             if (!value)\n> > +                     return config_error_nonbool(var);\n> > +             if (!strcmp(value, \"fsync\"))\n> > +                     fsync_method = FSYNC_METHOD_FSYNC;\n> > +             else if (!strcmp(value, \"writeout-only\"))\n> > +                     fsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n> > +             else\n>\n> As a non-nit comment I think this config schema looks great so far.\n>\n> > +                     warning(_(\"unknown %s value '%s'\"), var, value);\n>\n> Just a suggestion maybe something slightly scarier like:\n>\n>     \"unknown core.fsyncMethod value '%s'; config from future git version? ignoring requested fsync strategy\"\n>\n> Also using the nicer camelCased version instead of \"var\" (also helps\n> translators with context...)\n>\n\nWill fix.  The motivation for this scheme was to 'factor' the messages\nso there would be less to translate. But I see now that the message\ndoesn't have enough context to translate reasonably.\n\n> > +int git_fsync(int fd, enum fsync_action action)\n> > +{\n> > +     switch (action) {\n> > +     case FSYNC_WRITEOUT_ONLY:\n> > +\n> > +#ifdef __APPLE__\n> > +             /*\n> > +              * on macOS, fsync just causes filesystem cache writeback but does not\n> > +              * flush hardware caches.\n> > +              */\n> > +             return fsync(fd);\n> > +#endif\n> > +\n> > +#ifdef HAVE_SYNC_FILE_RANGE\n> > +             /*\n> > +              * On linux 2.6.17 and above, sync_file_range is the way to issue\n> > +              * a writeback without a hardware flush. An offset of 0 and size of 0\n> > +              * indicates writeout of the entire file and the wait flags ensure that all\n> > +              * dirty data is written to the disk (potentially in a disk-side cache)\n> > +              * before we continue.\n> > +              */\n> > +\n> > +             return sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n> > +                                              SYNC_FILE_RANGE_WRITE |\n> > +                                              SYNC_FILE_RANGE_WAIT_AFTER);\n> > +#endif\n> > +\n> > +#ifdef fsync_no_flush\n> > +             return fsync_no_flush(fd);\n> > +#endif\n> > +\n> > +             errno = ENOSYS;\n> > +             return -1;\n> > +\n> > +     case FSYNC_HARDWARE_FLUSH:\n> > +             /*\n> > +              * On some platforms fsync may return EINTR. Try again in this\n> > +              * case, since callers asking for a hardware flush may die if\n> > +              * this function returns an error.\n> > +              */\n> > +             for (;;) {\n> > +                     int err;\n> > +#ifdef __APPLE__\n> > +                     err = fcntl(fd, F_FULLFSYNC);\n> > +#else\n> > +                     err = fsync(fd);\n> > +#endif\n> > +                     if (err >= 0 || errno != EINTR)\n> > +                             return err;\n> > +             }\n> > +\n> > +     default:\n> > +             BUG(\"unexpected git_fsync(%d) call\", action);\n>\n> Don't include such \"default\" cases, you have an exhaustive \"enum\", if\n> you skip it the compiler will check this for you.\n>\n\nJunio gave the feedback to include this \"default:\" case in the switch\n[1].  Removing the default leads to the \"error: control reaches end of\nnon-void function\" on gcc. Fixing that error and adding a trial option\ndoes give the exhaustiveness error that you're talking about.  I'd\nrather just leave this as-is since the BUG() obviates the need for an\nin-practice-unreachable return statement.\n\n[1] https://lore.kernel.org/git/xmqqfsu70x58.fsf@gitster.g/\n\n> > +     }\n> > +}\n> > +\n> >  static int warn_if_unremovable(const char *op, const char *file, int rc)\n>\n> Just a code nit: I think it's very much preferred if possible to have as\n> much of code like this compile on all platforms. See the series at\n> 4002e87cb25 (grep: remove #ifdef NO_PTHREADS, 2018-11-03) is part of for\n> a good example.\n>\n> Maybe not worth it in this case since they're not nested ifdef's.\n>\n> I'm basically thinking of something (also re Patrick's comment on the\n> 2nd patch) where we have a platform_fsync() whose return\n> value/arguments/whatever capture this \"I want to return now\" or \"you\n> should be looping\" and takes the enum_fsync_action\" strategy.\n>\n> Then the git_fsync() would be the platform-independent looping etc., and\n> another funciton would do the \"one fsync at a time, maybe call me\n> again\".\n>\n> Maybe it would suck more, just food for thought... :)\n\nI'm going to introduce a new static function called fsync_loop which\ndoes the looping and I'll call it from git_fsync.  That appears to be\nthe cleanest to me to address your and Patrick's feedback.\n\nThanks,\nNeeraj\n"},{"id":"443392","messageId":"CANQDOdeL5LUV-2pffSyqZ9kGAhS4V4wg5--6ERvMALEkfsUCTg@mail.gmail.com","threadId":"57030","inReplyTo":"Ya9L5GJqlBF1YEk2@ncase","subject":"Re: [PATCH v2 0/3] A design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-08T00:44:09Z","receivedAt":"2021-12-08T00:44:23Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Dec 7, 2021 at 3:57 AM Patrick Steinhardt <ps@pks.im> wrote:\n\n> While I bail from the question of whether we want to grant as much\n> configurability to the user as this patch series does, I quite like the\n> implementation. It feels rather straight-forward and it's easy to see\n> how to extend it to support syncing of other subsystems like the loose\n> refs.\n>\n> Thanks!\n>\n> Patrick\n\nThanks for the positive comment.  I'm assuming that a major Git\nservices like GitHub or GitLab would be able to take advantage of the\ngranular options and knowledge of their hosting environment to choose\nthe right values for any server-side git deployments.  I'd probably\nturn off syncing for derived stuff like the commit-graph file and pack\nmetadata.\n\nMy underlying interest in all of these changes is to make Windows stop\nlooking so bad (we're defaulting to core.fsyncobjectfiles=true).\nBatch mode should give similar safety and much more optimizable\nperformance in our environment.\n\nThanks,\nNeeraj\n"},{"id":"443412","messageId":"211208.86ee6nmme5.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"CANQDOdfX2KaosPwLM4hS4rp+FH9V7VUxUh_md43FfZ9NG4iroQ@mail.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-08T10:05:30Z","receivedAt":"2021-12-08T10:17:11Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 07 2021, Neeraj Singh wrote:\n\n> On Tue, Dec 7, 2021 at 5:01 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>\n>>\n>> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n>>\n>> > From: Neeraj Singh <neerajsi@microsoft.com>\n>> >\n>> > This commit introduces the `core.fsync` configuration\n>> > knob which can be used to control how components of the\n>> > repository are made durable on disk.\n>> >\n>> > This setting allows future extensibility of the list of\n>> > syncable components:\n>> > * We issue a warning rather than an error for unrecognized\n>> >   components, so new configs can be used with old Git versions.\n>>\n>> Looks good!\n>>\n>> > * We support negation, so users can choose one of the default\n>> >   aggregate options and then remove components that they don't\n>> >   want. The user would then harden any new components added in\n>> >   a Git version update.\n>>\n>> I think this config schema makes sense, but just a (I think important)\n>> comment on the \"how\" not \"what\" of it. It's really much better to define\n>> config as:\n>>\n>>     [some]\n>>     key = value\n>>     key = value2\n>>\n>> Than:\n>>\n>>     [some]\n>>     key = value,value2\n>>\n>> The reason is that \"git config\" has good support for working with\n>> multi-valued stuff, so you can do e.g.:\n>>\n>>     git config --get-all -z some.key\n>>\n>> And you can easily (re)set such config e.g. with --replace-all etc., but\n>> for comma-delimited you (and users) need to do all that work themselves.\n>>\n>> Similarly instead of:\n>>\n>>     some.key = want-this\n>>     some.key = -not-this\n>>     some.key = but-want-this\n>>\n>> I think it's better to just have two lists, one inclusive another\n>> exclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n>> \"transfer.hideRefs\"\n>>\n>> Which would mean:\n>>\n>>     core.fsync = want-this\n>>     core.fsyncExcludes = -not-this\n>>\n>> For some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\n>> suggestion on making this easier for users & the implementation.\n>>\n>\n> Maybe there's some way to handle this I'm unaware of, but a\n> disadvantage of your multi-valued config proposal is that it's harder,\n> for example, for a per-repo config store to reasonably override a\n> per-user config store.  With the configuration scheme as-is, I can\n> have a per-user setting like `core.fsync=all` which covers my typical\n> repos, but then have a maintainer repo with a private setting of\n> `core.fsync=none` to speed up cases where I'm mostly working with\n> other people's changes that are backed up in email or server-side\n> repos.  The latter setting conveniently overrides the former setting\n> in all aspects.\n\nEven if you turn just your comma-delimited proposal into a list proposal\ncan't we just say that the last one wins? Then it won't matter what cmae\nbefore, you'd specify \"core.fsync=none\" in your local .git/config.\n\nBut this is also a general issue with a bunch of things in git's config\nspace. I'd rather see us use the list-based values and just come up with\nsome general way to reset them that works for all keys, rather than\nregretting having comma-delimited values that'll be harder to work with\n& parse, which will be a legacy wart if/when we come up with a way to\nsay \"reset all previous settings\".\n\n> Also, with the core.fsync and core.fsyncExcludes, how would you spell\n> \"don't sync anything\"? Would you still have the aggregate options.?\n\n    core.fsyncExcludes = *\n\nI.e. the same as the core.fsync=none above, anyway I still like the\nwildcard thing below a bit more...\n\n>> > This also supports the common request of doing absolutely no\n>> > fysncing with the `core.fsync=none` value, which is expected\n>> > to make the test suite faster.\n>>\n>> Let's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\n>> so we'll accept \"false\", \"off\", \"no\" like most other such config?\n>\n> Junio's previous feedback when discussing batch mode [1] was to offer\n> less flexibility when parsing new values of these configuration\n> options. I agree with his statement that \"making machine-readable\n> tokens be spelled in different ways is a 'disease'.\"  I'd like to\n> leave this as-is so that the documentation can clearly state the exact\n> set of allowable values.\n>\n> [1] https://lore.kernel.org/git/xmqqr1dqzyl7.fsf@gitster.g/\n\nI think he's talking about batch, Batch, BATCH, bAtCh etc. there. But\nthe \"maybe bool\" is a stanard pattern we use.\n\nI don't think we'd call one of these 0, off, no or false etc. to avoid\nconfusion, so then you can use git_parse_maybe_...()\n\n>>\n>> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>> > ---\n>> >  Documentation/config/core.txt | 27 +++++++++----\n>> >  builtin/fast-import.c         |  2 +-\n>> >  builtin/index-pack.c          |  4 +-\n>> >  builtin/pack-objects.c        |  8 ++--\n>> >  bulk-checkin.c                |  5 ++-\n>> >  cache.h                       | 39 +++++++++++++++++-\n>> >  commit-graph.c                |  3 +-\n>> >  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n>> >  csum-file.c                   |  5 ++-\n>> >  csum-file.h                   |  3 +-\n>> >  environment.c                 |  1 +\n>> >  midx.c                        |  3 +-\n>> >  object-file.c                 |  3 +-\n>> >  pack-bitmap-write.c           |  3 +-\n>> >  pack-write.c                  | 13 +++---\n>> >  read-cache.c                  |  2 +-\n>> >  16 files changed, 164 insertions(+), 33 deletions(-)\n>> >\n>> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n>> > index dbb134f7136..4f1747ec871 100644\n>> > --- a/Documentation/config/core.txt\n>> > +++ b/Documentation/config/core.txt\n>> > @@ -547,6 +547,25 @@ core.whitespace::\n>> >    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n>> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n>> >\n>> > +core.fsync::\n>> > +     A comma-separated list of parts of the repository which should be\n>> > +     hardened via the core.fsyncMethod when created or modified. You can\n>> > +     disable hardening of any component by prefixing it with a '-'. Later\n>> > +     items take precedence over earlier ones in the list. For example,\n>> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n>> > +     metadata.\" Items that are not hardened may be lost in the event of an\n>> > +     unclean system shutdown.\n>> > ++\n>> > +* `none` disables fsync completely. This must be specified alone.\n>> > +* `loose-object` hardens objects added to the repo in loose-object form.\n>> > +* `pack` hardens objects added to the repo in packfile form.\n>> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n>> > +* `commit-graph` hardens the commit graph file.\n>> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n>> > +  `pack-metadata`, and `commit-graph`.\n>> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n>> > +* `all` is an aggregate option that syncs all individual components above.\n>> > +\n>>\n>> It's probably a *bit* more work to set up, but I wonder if this wouldn't\n>> be simpler if we just said (and this is partially going against what I\n>> noted above):\n>>\n>> == BEGIN DOC\n>>\n>> core.fsync is a multi-value config variable where each item is a\n>> pathspec that'll get matched the same way as 'git-ls-files' et al.\n>>\n>> When we sync pretend that a path like .git/objects/de/adbeef... is\n>> relative to the top-level of the git\n>> directory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n>>\n>> You can then supply a list of wildcards and exclusions to configure\n>> syncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\n>> with:\n>>\n>>     ; same as \"false\"\n>>     core.fsync = \":!*\"\n>>\n>> Or:\n>>\n>>     ; same as \"true\"\n>>     core.fsync = \"*\"\n>>\n>> Or, to selectively sync some things and not others:\n>>\n>>     ;; Sync objects, but not \"info\"\n>>     core.fsync = \":!objects/info/**\"\n>>     core.fsync = \"objects/**\"\n>>\n>> See gitrepository-layout(5) for details about what sort of paths you\n>> might be expected to match. Not all paths listed there will go through\n>> this mechanism (e.g. currently objects do, but nothing to do with config\n>> does).\n>>\n>> We can and will match this against \"fake paths\", e.g. when writing out\n>> packs we may match against just the string \"objects/pack\", we're not\n>> going to re-check if every packfile we're writing matches your globs,\n>> ditto for loose objects. Be reasonable!\n>>\n>> This metharism is intended as a shorthand that provides some flexibility\n>> when fsyncing, while not forcing git to come up with labels for all\n>> paths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n>>\n>> More paths may be added or removed in the future, and we make no\n>> promises that we won't move things around, so if in doubt use\n>> e.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\n>> stick to the golden path of examples provided in this documentation.\n>>\n>> == END DOC\n>>\n>>\n>> It's a tad more complex to set up, but I wonder if that isn't worth\n>> it. It nicely gets around any current and future issues of deciding what\n>> labels such as \"loose-object\" etc. to pick, as well as slotting into an\n>> existing method of doing exclude/include lists.\n>>\n>\n> I think this proposal is a lot of complexity to avoid coming up with a\n> new name for syncable things as they are added to Git.  A path based\n> mechanism makes it hard to document for the (advanced) user what the\n> full set of things is and how it might change from release to release.\n> I think the current core.fsync scheme is a bit easier to understand,\n> query, and extend.\n\nWe document it in gitrepository-layout(5). Yeah it has some\ndisadvantages, but one advantage is that you could make the\ncomposability easy. I.e. if last exclude wins then a setting of:\n\n    core.fsync = \":!*\"\n    core.fsync = \"objects/**\"\n\nWould reset all previous matches & only match objects/**.\n\n>> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n>> > index 857be7826f3..916c55d6ce9 100644\n>> > --- a/builtin/pack-objects.c\n>> > +++ b/builtin/pack-objects.c\n>> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n>> >                * If so, rewrite it like in fast-import\n>> >                */\n>> >               if (pack_to_stdout) {\n>> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n>> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>>\n>> Not really related to this per-se, but since you're touching the API\n>> everything goes through I wonder if callers should just always try to\n>> fsync, and we can just catch EROFS and EINVAL in the wrapper if someone\n>> tries to flush stdout, or catch the fd at that lower level.\n>>\n>> Or maybe there's a good reason for this...\n>\n> It's platform dependent, but I'd expect fsync would do something for\n> pipes or stdout redirected to a file.  In these cases we really don't\n> want to fsync since we have no idea what we're talking to and we're\n> potentially worsening performance for probably no benefit.\n\nYeah maybe we should just leave it be.\n\nI'd think the C library returning EINVAL would be a trivial performance\ncost though.\n\nIt just seemed odd to hardcode assumptions about what can and can't be\nsynced when the POSIX defined function will also tell us that.\n\nAnyway...\n\n>> > [...]\n>> > +/*\n>> > + * These values are used to help identify parts of a repository to fsync.\n>> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n>> > + * repository and so shouldn't be fsynced.\n>> > + */\n>> > +enum fsync_component {\n>> > +     FSYNC_COMPONENT_NONE                    = 0,\n>>\n>> I haven't read ahead much but in most other such cases we don't define\n>> the \"= 0\", just start at 1<<0, then check the flags elsewhere...\n>>\n>> > +static const struct fsync_component_entry {\n>> > +     const char *name;\n>> > +     enum fsync_component component_bits;\n>> > +} fsync_component_table[] = {\n>> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n>> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n>> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n>> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n>> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n>> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n>> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n>> > +};\n>> > +\n>> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n>> > +{\n>> > +     enum fsync_component output = 0;\n>> > +\n>> > +     if (!strcmp(string, \"none\"))\n>> > +             return output;\n>> > +\n>> > +     while (string) {\n>> > +             int i;\n>> > +             size_t len;\n>> > +             const char *ep;\n>> > +             int negated = 0;\n>> > +             int found = 0;\n>> > +\n>> > +             string = string + strspn(string, \", \\t\\n\\r\");\n>>\n>> Aside from the \"use a list\" isn't this hardcoding some windows-specific\n>> assumptions with \\n\\r? Maybe not...\n>\n> I shamelessly stole this code from parse_whitespace_rule. I thought\n> about making a helper to be called by both functions, but the amount\n> of state going into and out of the wrapper via arguments was\n> substantial and seemed to negate the benefit of deduplication.\n\nFWIW string_list_split() is easier to work with in those cases, or at\nleast I think so...\n"},{"id":"443549","messageId":"CANQDOddkKbUC-g97JOf40nS28Yv1KACvbjW9gtQZemfBMutPCw@mail.gmail.com","threadId":"57030","inReplyTo":"211208.86ee6nmme5.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-09T00:14:40Z","receivedAt":"2021-12-09T00:14:55Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Dec 8, 2021 at 2:17 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Tue, Dec 07 2021, Neeraj Singh wrote:\n>\n> > On Tue, Dec 7, 2021 at 5:01 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n> >>\n> >> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >> >\n> >> > This commit introduces the `core.fsync` configuration\n> >> > knob which can be used to control how components of the\n> >> > repository are made durable on disk.\n> >> >\n> >> > This setting allows future extensibility of the list of\n> >> > syncable components:\n> >> > * We issue a warning rather than an error for unrecognized\n> >> >   components, so new configs can be used with old Git versions.\n> >>\n> >> Looks good!\n> >>\n> >> > * We support negation, so users can choose one of the default\n> >> >   aggregate options and then remove components that they don't\n> >> >   want. The user would then harden any new components added in\n> >> >   a Git version update.\n> >>\n> >> I think this config schema makes sense, but just a (I think important)\n> >> comment on the \"how\" not \"what\" of it. It's really much better to define\n> >> config as:\n> >>\n> >>     [some]\n> >>     key = value\n> >>     key = value2\n> >>\n> >> Than:\n> >>\n> >>     [some]\n> >>     key = value,value2\n> >>\n> >> The reason is that \"git config\" has good support for working with\n> >> multi-valued stuff, so you can do e.g.:\n> >>\n> >>     git config --get-all -z some.key\n> >>\n> >> And you can easily (re)set such config e.g. with --replace-all etc., but\n> >> for comma-delimited you (and users) need to do all that work themselves.\n> >>\n> >> Similarly instead of:\n> >>\n> >>     some.key = want-this\n> >>     some.key = -not-this\n> >>     some.key = but-want-this\n> >>\n> >> I think it's better to just have two lists, one inclusive another\n> >> exclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n> >> \"transfer.hideRefs\"\n> >>\n> >> Which would mean:\n> >>\n> >>     core.fsync = want-this\n> >>     core.fsyncExcludes = -not-this\n> >>\n> >> For some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\n> >> suggestion on making this easier for users & the implementation.\n> >>\n> >\n> > Maybe there's some way to handle this I'm unaware of, but a\n> > disadvantage of your multi-valued config proposal is that it's harder,\n> > for example, for a per-repo config store to reasonably override a\n> > per-user config store.  With the configuration scheme as-is, I can\n> > have a per-user setting like `core.fsync=all` which covers my typical\n> > repos, but then have a maintainer repo with a private setting of\n> > `core.fsync=none` to speed up cases where I'm mostly working with\n> > other people's changes that are backed up in email or server-side\n> > repos.  The latter setting conveniently overrides the former setting\n> > in all aspects.\n>\n> Even if you turn just your comma-delimited proposal into a list proposal\n> can't we just say that the last one wins? Then it won't matter what cmae\n> before, you'd specify \"core.fsync=none\" in your local .git/config.\n>\n> But this is also a general issue with a bunch of things in git's config\n> space. I'd rather see us use the list-based values and just come up with\n> some general way to reset them that works for all keys, rather than\n> regretting having comma-delimited values that'll be harder to work with\n> & parse, which will be a legacy wart if/when we come up with a way to\n> say \"reset all previous settings\".\n>\n> > Also, with the core.fsync and core.fsyncExcludes, how would you spell\n> > \"don't sync anything\"? Would you still have the aggregate options.?\n>\n>     core.fsyncExcludes = *\n>\n> I.e. the same as the core.fsync=none above, anyway I still like the\n> wildcard thing below a bit more...\n\nI'm not going to take this feedback unless there are additional votes\nfrom the Git community in this direction.  I make the claim that\nsingle-valued comma-separated config lists are easier to work with in\nthe existing Git infrastructure.  We already use essentially the same\nparsing code for the core.whitespace variable and users are used to\nthis syntax there. There are several other comma-separated lists in\nthe config space, so this construct has precedence and will be with\nGit for some time.  Also, fsync configurations aren't composable like\nsome other configurations may be. It makes sense to have a holistic\nsingular fsync configuration, which is best represented by a single\nvariable.\n\n> >> > This also supports the common request of doing absolutely no\n> >> > fysncing with the `core.fsync=none` value, which is expected\n> >> > to make the test suite faster.\n> >>\n> >> Let's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\n> >> so we'll accept \"false\", \"off\", \"no\" like most other such config?\n> >\n> > Junio's previous feedback when discussing batch mode [1] was to offer\n> > less flexibility when parsing new values of these configuration\n> > options. I agree with his statement that \"making machine-readable\n> > tokens be spelled in different ways is a 'disease'.\"  I'd like to\n> > leave this as-is so that the documentation can clearly state the exact\n> > set of allowable values.\n> >\n> > [1] https://lore.kernel.org/git/xmqqr1dqzyl7.fsf@gitster.g/\n>\n> I think he's talking about batch, Batch, BATCH, bAtCh etc. there. But\n> the \"maybe bool\" is a stanard pattern we use.\n>\n> I don't think we'd call one of these 0, off, no or false etc. to avoid\n> confusion, so then you can use git_parse_maybe_...()\n\nI don't see the advantage of having multiple ways of specifying\n\"none\".  The user can read the doc and know exactly what to write.  If\nthey write something unallowable, they get a clear warning and they\ncan read the doc again to figure out what to write.  This isn't a\nboolean options at all, so why should we entertain bool-like ways of\nspelling it?\n\n> >>\n> >> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> >> > ---\n> >> >  Documentation/config/core.txt | 27 +++++++++----\n> >> >  builtin/fast-import.c         |  2 +-\n> >> >  builtin/index-pack.c          |  4 +-\n> >> >  builtin/pack-objects.c        |  8 ++--\n> >> >  bulk-checkin.c                |  5 ++-\n> >> >  cache.h                       | 39 +++++++++++++++++-\n> >> >  commit-graph.c                |  3 +-\n> >> >  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n> >> >  csum-file.c                   |  5 ++-\n> >> >  csum-file.h                   |  3 +-\n> >> >  environment.c                 |  1 +\n> >> >  midx.c                        |  3 +-\n> >> >  object-file.c                 |  3 +-\n> >> >  pack-bitmap-write.c           |  3 +-\n> >> >  pack-write.c                  | 13 +++---\n> >> >  read-cache.c                  |  2 +-\n> >> >  16 files changed, 164 insertions(+), 33 deletions(-)\n> >> >\n> >> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> >> > index dbb134f7136..4f1747ec871 100644\n> >> > --- a/Documentation/config/core.txt\n> >> > +++ b/Documentation/config/core.txt\n> >> > @@ -547,6 +547,25 @@ core.whitespace::\n> >> >    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n> >> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >> >\n> >> > +core.fsync::\n> >> > +     A comma-separated list of parts of the repository which should be\n> >> > +     hardened via the core.fsyncMethod when created or modified. You can\n> >> > +     disable hardening of any component by prefixing it with a '-'. Later\n> >> > +     items take precedence over earlier ones in the list. For example,\n> >> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n> >> > +     metadata.\" Items that are not hardened may be lost in the event of an\n> >> > +     unclean system shutdown.\n> >> > ++\n> >> > +* `none` disables fsync completely. This must be specified alone.\n> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> >> > +* `pack` hardens objects added to the repo in packfile form.\n> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> >> > +* `commit-graph` hardens the commit graph file.\n> >> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n> >> > +  `pack-metadata`, and `commit-graph`.\n> >> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n> >> > +* `all` is an aggregate option that syncs all individual components above.\n> >> > +\n> >>\n> >> It's probably a *bit* more work to set up, but I wonder if this wouldn't\n> >> be simpler if we just said (and this is partially going against what I\n> >> noted above):\n> >>\n> >> == BEGIN DOC\n> >>\n> >> core.fsync is a multi-value config variable where each item is a\n> >> pathspec that'll get matched the same way as 'git-ls-files' et al.\n> >>\n> >> When we sync pretend that a path like .git/objects/de/adbeef... is\n> >> relative to the top-level of the git\n> >> directory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n> >>\n> >> You can then supply a list of wildcards and exclusions to configure\n> >> syncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\n> >> with:\n> >>\n> >>     ; same as \"false\"\n> >>     core.fsync = \":!*\"\n> >>\n> >> Or:\n> >>\n> >>     ; same as \"true\"\n> >>     core.fsync = \"*\"\n> >>\n> >> Or, to selectively sync some things and not others:\n> >>\n> >>     ;; Sync objects, but not \"info\"\n> >>     core.fsync = \":!objects/info/**\"\n> >>     core.fsync = \"objects/**\"\n> >>\n> >> See gitrepository-layout(5) for details about what sort of paths you\n> >> might be expected to match. Not all paths listed there will go through\n> >> this mechanism (e.g. currently objects do, but nothing to do with config\n> >> does).\n> >>\n> >> We can and will match this against \"fake paths\", e.g. when writing out\n> >> packs we may match against just the string \"objects/pack\", we're not\n> >> going to re-check if every packfile we're writing matches your globs,\n> >> ditto for loose objects. Be reasonable!\n> >>\n> >> This metharism is intended as a shorthand that provides some flexibility\n> >> when fsyncing, while not forcing git to come up with labels for all\n> >> paths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n> >>\n> >> More paths may be added or removed in the future, and we make no\n> >> promises that we won't move things around, so if in doubt use\n> >> e.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\n> >> stick to the golden path of examples provided in this documentation.\n> >>\n> >> == END DOC\n> >>\n> >>\n> >> It's a tad more complex to set up, but I wonder if that isn't worth\n> >> it. It nicely gets around any current and future issues of deciding what\n> >> labels such as \"loose-object\" etc. to pick, as well as slotting into an\n> >> existing method of doing exclude/include lists.\n> >>\n> >\n> > I think this proposal is a lot of complexity to avoid coming up with a\n> > new name for syncable things as they are added to Git.  A path based\n> > mechanism makes it hard to document for the (advanced) user what the\n> > full set of things is and how it might change from release to release.\n> > I think the current core.fsync scheme is a bit easier to understand,\n> > query, and extend.\n>\n> We document it in gitrepository-layout(5). Yeah it has some\n> disadvantages, but one advantage is that you could make the\n> composability easy. I.e. if last exclude wins then a setting of:\n>\n>     core.fsync = \":!*\"\n>     core.fsync = \"objects/**\"\n>\n> Would reset all previous matches & only match objects/**.\n>\n\nThe value of changing this is predicated on taking your previous\nmulti-valued config proposal, which I'm still not at all convinced\nabout.  The schema in the current (v1-v2) version of the patch already\nincludes an example of extending the list of syncable things, and\nPatrick Steinhardt made it clear that he feels comfortable adding\n'refs' to the same schema in a future change.\n\nI'll also emphasize that we're talking about a non-functional,\nrelatively corner-case behavioral configuration.  These values don't\nchange how git's interface behaves except when the system crashes\nduring a git command or shortly after one completes.\n\nWhile you may not personally love the proposed configuration\ninterface, I'd want your view on some questions:\n1. Is it easy for the (advanced) user to set a configuration?\n2. Is it easy for the (advanced) user to see what was configured?\n3. Is it easy for the Git community to build on this as we want to add\nthings to the list of things to sync?\n    a) Is there a good best practice configuration so that people can\navoid losing integrity for new stuff that they are intending to sync.\n    b) If someone has a custom configuration, can that custom\nconfiguration do something reasonable as they upgrade versions of Git?\n             ** In response to this question, I might see some value\nin adding a 'derived-metadata' aggregate that can be disabled so that\na custom configuration can exclude those as they change version to\nversion.\n    c) Is it too much maintenance overhead to consider how to present\nthis configuration knob for any new hashfile or other datafile in the\ngit repo?\n4. Is there a good path forward to change the default syncable set,\nboth in git-for-windows and in Git for other platforms?\n\n> >> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> >> > index 857be7826f3..916c55d6ce9 100644\n> >> > --- a/builtin/pack-objects.c\n> >> > +++ b/builtin/pack-objects.c\n> >> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n> >> >                * If so, rewrite it like in fast-import\n> >> >                */\n> >> >               if (pack_to_stdout) {\n> >> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> >> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n> >> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> >>\n> >> Not really related to this per-se, but since you're touching the API\n> >> everything goes through I wonder if callers should just always try to\n> >> fsync, and we can just catch EROFS and EINVAL in the wrapper if someone\n> >> tries to flush stdout, or catch the fd at that lower level.\n> >>\n> >> Or maybe there's a good reason for this...\n> >\n> > It's platform dependent, but I'd expect fsync would do something for\n> > pipes or stdout redirected to a file.  In these cases we really don't\n> > want to fsync since we have no idea what we're talking to and we're\n> > potentially worsening performance for probably no benefit.\n>\n> Yeah maybe we should just leave it be.\n>\n> I'd think the C library returning EINVAL would be a trivial performance\n> cost though.\n>\n> It just seemed odd to hardcode assumptions about what can and can't be\n> synced when the POSIX defined function will also tell us that.\n>\n\nRedirecting stdout to a file seems like a common usage for this\ncommand. That would definitely be fsyncable, but Git has no idea what\nits proper category is since there's no way to know the purpose or\nlifetime of the packfile.  I'm going to leave this be, because I'd\nposit that \"can it be fsynced?\" is not the same as \"should it be\nfsynced?\".  The latter question can't be answered for stdout.\n\n>\n> >> > [...]\n> >> > +/*\n> >> > + * These values are used to help identify parts of a repository to fsync.\n> >> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n> >> > + * repository and so shouldn't be fsynced.\n> >> > + */\n> >> > +enum fsync_component {\n> >> > +     FSYNC_COMPONENT_NONE                    = 0,\n> >>\n> >> I haven't read ahead much but in most other such cases we don't define\n> >> the \"= 0\", just start at 1<<0, then check the flags elsewhere...\n> >>\n> >> > +static const struct fsync_component_entry {\n> >> > +     const char *name;\n> >> > +     enum fsync_component component_bits;\n> >> > +} fsync_component_table[] = {\n> >> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> >> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n> >> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> >> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> >> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> >> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n> >> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n> >> > +};\n> >> > +\n> >> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> >> > +{\n> >> > +     enum fsync_component output = 0;\n> >> > +\n> >> > +     if (!strcmp(string, \"none\"))\n> >> > +             return output;\n> >> > +\n> >> > +     while (string) {\n> >> > +             int i;\n> >> > +             size_t len;\n> >> > +             const char *ep;\n> >> > +             int negated = 0;\n> >> > +             int found = 0;\n> >> > +\n> >> > +             string = string + strspn(string, \", \\t\\n\\r\");\n> >>\n> >> Aside from the \"use a list\" isn't this hardcoding some windows-specific\n> >> assumptions with \\n\\r? Maybe not...\n> >\n> > I shamelessly stole this code from parse_whitespace_rule. I thought\n> > about making a helper to be called by both functions, but the amount\n> > of state going into and out of the wrapper via arguments was\n> > substantial and seemed to negate the benefit of deduplication.\n>\n> FWIW string_list_split() is easier to work with in those cases, or at\n> least I think so...\n\nThis code runs at startup for a variable that may be present on some\ninstallations.  The nice property of the current patch's code is that\nit's already a well-tested pattern that doesn't do any allocations as\nit's working, unlike string_list_split().\n\nI hope you know that I appreciate your review feedback, even though\nI'm pushing back on most of it so far this round. I'll be sending v3\nto the list soon after giving it another look over.\n\nThanks,\nNeeraj\n"},{"id":"443551","messageId":"xmqq5yrywqsu.fsf@gitster.g","threadId":"57030","inReplyTo":"CANQDOddkKbUC-g97JOf40nS28Yv1KACvbjW9gtQZemfBMutPCw@mail.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-09T00:44:01Z","receivedAt":"2021-12-09T00:44:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> I'm not going to take this feedback unless there are additional votes\n> from the Git community in this direction.  I make the claim that\n> single-valued comma-separated config lists are easier to work with in\n> the existing Git infrastructure.  We already use essentially the same\n> parsing code for the core.whitespace variable and users are used to\n> this syntax there. There are several other comma-separated lists in\n> the config space, so this construct has precedence and will be with\n> Git for some time.  Also, fsync configurations aren't composable like\n> some other configurations may be. It makes sense to have a holistic\n> singular fsync configuration, which is best represented by a single\n> variable.\n\nI haven't caught up with the discussion in this thread, even though\nI have been meaning to think about it for some time---I just haven't\ngot around to it (sorry).  So I'll stop at giving a general guidance\nand leave the decision if it applies to this particular discussion\nto readers.\n\nAs the inventor of core.whitespace \"list of values and its parser, I\nam biased, but I would say that it works well for simple things that\ndo not need too much overriding.  The other side of the coin is that\nit can become very awkward going forward if we use it to things that\nhave more complex needs than answering a simple question like \"what\nwhitespace errors should be checked?\".\n\nMore specifically, core.whitespace is pecuriar in a few ways.\n\n * It does follow the usual \"the last one wins\" rule, but in a\n   strange way.  Notice the \"unsigned rule = WS_DEFAULT_RULE\"\n   assignment at the beginning of ws.c::parse_whitespace_rule()?\n   For each configuration \"core.whitespace=<list>\" we encounter,\n   we start from the default, discarding everything we saw so far,\n   and tweak that default value with tokens found on the list.\n\n * You cannot do \"The system config gives one set of values, which\n   you tweak with the personal config, which is further tweaked with\n   the repository config\" as the consequence.  This is blessing and\n   is curse at the same time, as it makes inspection simpler when\n   things go wrong (you only need to check whta the last one does),\n   but it is harder to share the common things in more common file.\n\n * Its design relies on the choices being strings chosen from a\n   fixed vocabulary to allow you to say \"the value of the\n   configuration variable is a list of tokens separated by a comma\"\n   and \"the default has X bit set, but we disable that with -X\".\n   For a configuration variable whose value is an arbitrary string\n   or a number, obviously that approach would not work.\n\nIf the need of the topic is simple enough that the above limitation\ncore.whitespace does not pose a problem going forward, it would be\nfine, but we may regret the choice we make today if that is not the\ncase.\n\nThanks.\n"},{"id":"443552","messageId":"pull.1093.v3.git.1639011433.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v2.git.1638845211.gitgitgadget@gmail.com","subject":"[PATCH v3 0/4] A design for future-proofing fsync() configuration","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-09T00:57:09Z","receivedAt":"2021-12-09T00:57:20Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"This is an implementation of an extensible configuration mechanism for\nfsyncing persistent components of a repo.\n\nThe main goals are to separate the \"what\" to sync from the \"how\". There are\nnow two settings: core.fsync - Control the 'what', including the index.\ncore.fsyncMethod - Control the 'how'. Currently we support writeout-only and\nfull fsync.\n\nSyncing of refs can be layered on top of core.fsync. And batch mode will be\nlayered on core.fsyncMethod.\n\ncore.fsyncobjectfiles is removed and will issue a deprecation warning if\nit's seen.\n\nI'd like to get agreement on this direction before submitting batch mode to\nthe list. The batch mode series is available to view at\nhttps://github.com/neerajsi-msft/git/pull/1.\n\nPlease see [1], [2], and [3] for discussions that led to this series.\n\nOne major concern I'd voice that is adverse to this change: when a new\npersistent file is added to the Git repo, the person adding that file will\nneed to update this configuration code and the documentation. Maybe this is\nthe right thing to always think about, but the FSYNC_COMPONENT lists will be\na single place that will receive a number of updates over time.\n\nNote: There's a minor conflict with ns/tmp-objdir. In\nobject-file.c:close_loose_object we need to resolve it like this:\n\nif (!the_repository->objects->odb->will_destroy)\nfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object\nfile\");\n\nV3 changes:\n\n * Remove relative path from git-compat-util.h include [4].\n * Updated newly added warning texts to have more context for localization\n   [4].\n * Fixed tab spacing in enum fsync_action\n * Moved the fsync looping out to a helper and do it consistently. [4]\n * Changed commit description to use camelCase for config names. [5]\n * Add an optional fourth patch with derived-metadata so that the user can\n   exclude a forward-compatible set of things that should be recomputable\n   given existing data.\n\nV2 changes:\n\n * Updated the documentation for core.fsyncmethod to be less certain.\n   writeout-only probably does not do the right thing on Linux.\n * Split out the core.fsync=index change into its own commit.\n * Rename REPO_COMPONENT to FSYNC_COMPONENT. This is really specific to\n   fsyncing, so the name should reflect that.\n * Re-add missing Makefile change for SYNC_FILE_RANGE.\n * Tested writeout-only mode, index syncing, and general config settings.\n\n[1] https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n[2]\nhttps://lore.kernel.org/git/dd65718814011eb93ccc4428f9882e0f025224a6.1636029491.git.ps@pks.im/\n[3]\nhttps://lore.kernel.org/git/pull.1076.git.git.1629856292.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/CANQDOdf8C4-haK9=Q_J4Cid8bQALnmGDm=SvatRbaVf+tkzqLw@mail.gmail.com/\n[5] https://lore.kernel.org/git/211207.861r2opplg.gmgdl@evledraar.gmail.com/\n\nNeeraj Singh (4):\n  core.fsyncmethod: add writeout-only mode\n  core.fsync: introduce granular fsync control\n  core.fsync: new option to harden the index\n  core.fsync: add a `derived-metadata` aggregate option\n\n Documentation/config/core.txt       | 35 ++++++++---\n Makefile                            |  6 ++\n builtin/fast-import.c               |  2 +-\n builtin/index-pack.c                |  4 +-\n builtin/pack-objects.c              |  8 ++-\n bulk-checkin.c                      |  5 +-\n cache.h                             | 49 +++++++++++++++-\n commit-graph.c                      |  3 +-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n config.c                            | 90 ++++++++++++++++++++++++++++-\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n csum-file.c                         |  5 +-\n csum-file.h                         |  3 +-\n environment.c                       |  3 +-\n git-compat-util.h                   | 24 ++++++++\n midx.c                              |  3 +-\n object-file.c                       |  3 +-\n pack-bitmap-write.c                 |  3 +-\n pack-write.c                        | 13 +++--\n read-cache.c                        | 19 ++++--\n wrapper.c                           | 64 ++++++++++++++++++++\n write-or-die.c                      | 10 ++--\n 25 files changed, 354 insertions(+), 43 deletions(-)\n create mode 100644 compat/win32/flush.c\n\n\nbase-commit: abe6bb3905392d5eb6b01fa6e54d7e784e0522aa\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1093%2Fneerajsi-msft%2Fns%2Fcore-fsync-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1093/neerajsi-msft/ns/core-fsync-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/1093\n\nRange-diff vs v2:\n\n 1:  e79522cbdd4 ! 1:  15edfe51509 core.fsyncmethod: add writeout-only mode\n     @@ Metadata\n       ## Commit message ##\n          core.fsyncmethod: add writeout-only mode\n      \n     -    This commit introduces the `core.fsyncmethod` configuration\n     +    This commit introduces the `core.fsyncMethod` configuration\n          knob, which can currently be set to `fsync` or `writeout-only`.\n      \n          The new writeout-only mode attempts to tell the operating system to\n     @@ Commit message\n          directive to the storage controller. This change updates fsync to do\n          fcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\n          with existing behavior on Apple platforms by setting the default value\n     -    of the new core.fsyncmethod option.\n     +    of the new core.fsyncMethod option.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ compat/mingw.h: int mingw_getpagesize(void);\n      \n       ## compat/win32/flush.c (new) ##\n      @@\n     -+#include \"../../git-compat-util.h\"\n     ++#include \"git-compat-util.h\"\n      +#include <winternl.h>\n      +#include \"lazyload.h\"\n      +\n     @@ config.c: static int git_default_core_config(const char *var, const char *value,\n      +\t\telse if (!strcmp(value, \"writeout-only\"))\n      +\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n      +\t\telse\n     -+\t\t\twarning(_(\"unknown %s value '%s'\"), var, value);\n     ++\t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n      +\n      +\t}\n      +\n     @@ git-compat-util.h: __attribute__((format (printf, 1, 2))) NORETURN\n      +#endif\n      +\n      +enum fsync_action {\n     -+    FSYNC_WRITEOUT_ONLY,\n     -+    FSYNC_HARDWARE_FLUSH\n     ++\tFSYNC_WRITEOUT_ONLY,\n     ++\tFSYNC_HARDWARE_FLUSH\n      +};\n      +\n      +/*\n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n       \treturn fd;\n       }\n       \n     ++/*\n     ++ * Some platforms return EINTR from fsync. Since fsync is invoked in some\n     ++ * cases by a wrapper that dies on failure, do not expose EINTR to callers.\n     ++ */\n     ++static int fsync_loop(int fd)\n     ++{\n     ++\tint err;\n     ++\n     ++\tdo {\n     ++\t\terr = fsync(fd);\n     ++\t} while (err < 0 && errno == EINTR);\n     ++\treturn err;\n     ++}\n     ++\n      +int git_fsync(int fd, enum fsync_action action)\n      +{\n      +\tswitch (action) {\n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n      +\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n      +\t\t * flush hardware caches.\n      +\t\t */\n     -+\t\treturn fsync(fd);\n     ++\t\treturn fsync_loop(fd);\n      +#endif\n      +\n      +#ifdef HAVE_SYNC_FILE_RANGE\n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n      +\t\t * case, since callers asking for a hardware flush may die if\n      +\t\t * this function returns an error.\n      +\t\t */\n     -+\t\tfor (;;) {\n     -+\t\t\tint err;\n      +#ifdef __APPLE__\n     -+\t\t\terr = fcntl(fd, F_FULLFSYNC);\n     ++\t\treturn fcntl(fd, F_FULLFSYNC);\n      +#else\n     -+\t\t\terr = fsync(fd);\n     ++\t\treturn fsync_loop(fd);\n      +#endif\n     -+\t\t\tif (err >= 0 || errno != EINTR)\n     -+\t\t\t\treturn err;\n     -+\t\t}\n     -+\n      +\tdefault:\n      +\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n      +\t}\n 2:  ff80a94bf9a ! 2:  080be1a6f64 core.fsync: introduce granular fsync control\n     @@ config.c: static int git_parse_maybe_bool_text(const char *value)\n      +\n      +\t\tif (!found) {\n      +\t\t\tchar *component = xstrndup(string, len);\n     -+\t\t\twarning(_(\"unknown %s value '%s'\"), var, component);\n     ++\t\t\twarning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n      +\t\t\tfree(component);\n      +\t\t}\n      +\n 3:  86e39b8f8d1 = 3:  2207950beba core.fsync: new option to harden the index\n -:  ----------- > 4:  a830d177d4c core.fsync: add a `derived-metadata` aggregate option\n\n-- \ngitgitgadget\n"},{"id":"443553","messageId":"15edfe5150961e8ec34f880a8edc44c7754af444.1639011433.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v3.git.1639011433.gitgitgadget@gmail.com","subject":"[PATCH v3 1/4] core.fsyncmethod: add writeout-only mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-09T00:57:10Z","receivedAt":"2021-12-09T00:57:20Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsyncMethod` configuration\nknob, which can currently be set to `fsync` or `writeout-only`.\n\nThe new writeout-only mode attempts to tell the operating system to\nflush its in-memory page cache to the storage hardware without issuing a\nCACHE_FLUSH command to the storage controller.\n\nWriteout-only fsync is significantly faster than a vanilla fsync on\ncommon hardware, since data is written to a disk-side cache rather than\nall the way to a durable medium. Later changes in this patch series will\ntake advantage of this primitive to implement batching of hardware\nflushes.\n\nWhen git_fsync is called with FSYNC_WRITEOUT_ONLY, it may fail and the\ncaller is expected to do an ordinary fsync as needed.\n\nOn Apple platforms, the fsync system call does not issue a CACHE_FLUSH\ndirective to the storage controller. This change updates fsync to do\nfcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\nwith existing behavior on Apple platforms by setting the default value\nof the new core.fsyncMethod option.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt       |  9 ++++\n Makefile                            |  6 +++\n cache.h                             |  7 ++++\n compat/mingw.h                      |  3 ++\n compat/win32/flush.c                | 28 +++++++++++++\n config.c                            | 12 ++++++\n config.mak.uname                    |  3 ++\n configure.ac                        |  8 ++++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n environment.c                       |  2 +-\n git-compat-util.h                   | 24 +++++++++++\n wrapper.c                           | 64 +++++++++++++++++++++++++++++\n write-or-die.c                      | 10 +++--\n 13 files changed, 173 insertions(+), 6 deletions(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..dbb134f7136 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,15 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsyncMethod::\n+\tA value indicating the strategy Git will use to harden repository data\n+\tusing fsync and related primitives.\n++\n+* `fsync` uses the fsync() system call or platform equivalents.\n+* `writeout-only` issues pagecache writeback requests, but depending on the\n+  filesystem and storage hardware, data added to the repository may not be\n+  durable in the event of a system crash. This is the default mode on macOS.\n+\n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n +\ndiff --git a/Makefile b/Makefile\nindex d56c0e4aadc..cba024615c9 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -403,6 +403,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1881,6 +1883,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/cache.h b/cache.h\nindex eba12487b99..9cd60d94952 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -986,6 +986,13 @@ extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n extern int fsync_object_files;\n+\n+enum fsync_method {\n+\tFSYNC_METHOD_FSYNC,\n+\tFSYNC_METHOD_WRITEOUT_ONLY\n+};\n+\n+extern enum fsync_method fsync_method;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..6ca82f1dae5\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.c b/config.c\nindex c5873f3a706..139df71ba17 100644\n--- a/config.c\n+++ b/config.c\n@@ -1490,6 +1490,18 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsyncmethod\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcmp(value, \"fsync\"))\n+\t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n+\t\telse if (!strcmp(value, \"writeout-only\"))\n+\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse\n+\t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n+\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\ndiff --git a/config.mak.uname b/config.mak.uname\nindex d0701f9beb0..774a09622d2 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -57,6 +57,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tBASIC_CFLAGS += -DHAVE_SYSINFO\n@@ -453,6 +454,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -628,6 +630,7 @@ ifeq ($(uname_S),MINGW)\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/configure.ac b/configure.ac\nindex 5ee25ec95c8..6bd6bef1c44 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1082,6 +1082,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 86b46114464..6d7bc16d054 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/environment.c b/environment.c\nindex 9da7f3c1a19..f9140e842cf 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -41,7 +41,7 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex c6bd2a84e55..50db85a8610 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1239,6 +1239,30 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+#ifdef __APPLE__\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n+#else\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n+#endif\n+\n+enum fsync_action {\n+\tFSYNC_WRITEOUT_ONLY,\n+\tFSYNC_HARDWARE_FLUSH\n+};\n+\n+/*\n+ * Issues an fsync against the specified file according to the specified mode.\n+ *\n+ * FSYNC_WRITEOUT_ONLY attempts to use interfaces available on some operating\n+ * systems to flush the OS cache without issuing a flush command to the storage\n+ * controller. If those interfaces are unavailable, the function fails with\n+ * ENOSYS.\n+ *\n+ * FSYNC_HARDWARE_FLUSH does an OS writeout and hardware flush to ensure that\n+ * changes are durable. It is not expected to fail.\n+ */\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/wrapper.c b/wrapper.c\nindex 36e12119d76..572f28f14ff 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -546,6 +546,70 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+/*\n+ * Some platforms return EINTR from fsync. Since fsync is invoked in some\n+ * cases by a wrapper that dies on failure, do not expose EINTR to callers.\n+ */\n+static int fsync_loop(int fd)\n+{\n+\tint err;\n+\n+\tdo {\n+\t\terr = fsync(fd);\n+\t} while (err < 0 && errno == EINTR);\n+\treturn err;\n+}\n+\n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync_loop(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\t\t/*\n+\t\t * On some platforms fsync may return EINTR. Try again in this\n+\t\t * case, since callers asking for a hardware flush may die if\n+\t\t * this function returns an error.\n+\t\t */\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync_loop(fd);\n+#endif\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 0b1ec8190b6..0702acdd5e8 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,10 +57,12 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n-\t\tif (errno != EINTR)\n-\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n-\t}\n+\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n+\t\treturn;\n+\n+\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n+\t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n void write_or_die(int fd, const void *buf, size_t count)\n-- \ngitgitgadget\n\n"},{"id":"443554","messageId":"080be1a6f642c64b5cbd74c612027245026833d3.1639011434.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v3.git.1639011433.gitgitgadget@gmail.com","subject":"[PATCH v3 2/4] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-09T00:57:11Z","receivedAt":"2021-12-09T00:57:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsync` configuration\nknob which can be used to control how components of the\nrepository are made durable on disk.\n\nThis setting allows future extensibility of the list of\nsyncable components:\n* We issue a warning rather than an error for unrecognized\n  components, so new configs can be used with old Git versions.\n* We support negation, so users can choose one of the default\n  aggregate options and then remove components that they don't\n  want. The user would then harden any new components added in\n  a Git version update.\n\nThis also supports the common request of doing absolutely no\nfysncing with the `core.fsync=none` value, which is expected\nto make the test suite faster.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 27 +++++++++----\n builtin/fast-import.c         |  2 +-\n builtin/index-pack.c          |  4 +-\n builtin/pack-objects.c        |  8 ++--\n bulk-checkin.c                |  5 ++-\n cache.h                       | 39 +++++++++++++++++-\n commit-graph.c                |  3 +-\n config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n csum-file.c                   |  5 ++-\n csum-file.h                   |  3 +-\n environment.c                 |  1 +\n midx.c                        |  3 +-\n object-file.c                 |  3 +-\n pack-bitmap-write.c           |  3 +-\n pack-write.c                  | 13 +++---\n read-cache.c                  |  2 +-\n 16 files changed, 164 insertions(+), 33 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex dbb134f7136..4f1747ec871 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,25 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsync::\n+\tA comma-separated list of parts of the repository which should be\n+\thardened via the core.fsyncMethod when created or modified. You can\n+\tdisable hardening of any component by prefixing it with a '-'. Later\n+\titems take precedence over earlier ones in the list. For example,\n+\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n+\tmetadata.\" Items that are not hardened may be lost in the event of an\n+\tunclean system shutdown.\n++\n+* `none` disables fsync completely. This must be specified alone.\n+* `loose-object` hardens objects added to the repo in loose-object form.\n+* `pack` hardens objects added to the repo in packfile form.\n+* `pack-metadata` hardens packfile bitmaps and indexes.\n+* `commit-graph` hardens the commit graph file.\n+* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n+  `pack-metadata`, and `commit-graph`.\n+* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n+* `all` is an aggregate option that syncs all individual components above.\n+\n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\n@@ -556,14 +575,6 @@ core.fsyncMethod::\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n \n-core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n-\n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\n +\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 20406f67754..e27a4580f85 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -856,7 +856,7 @@ static void end_packfile(void)\n \t\tstruct tag *t;\n \n \t\tclose_pack_windows(pack_data);\n-\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n \t\t\t\t\t pack_data->pack_name, object_count,\n \t\t\t\t\t cur_pack_oid.hash, pack_size);\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c23d01de7dc..c32534c13b4 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1286,7 +1286,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\t\t    nr_objects - nr_objects_initial);\n \t\tstop_progress_msg(&progress, msg.buf);\n \t\tstrbuf_release(&msg);\n-\t\tfinalize_hashfile(f, tail_hash, 0);\n+\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n \t\thashcpy(read_hash, pack_hash);\n \t\tfixup_pack_header_footer(output_fd, pack_hash,\n \t\t\t\t\t curr_pack, nr_objects,\n@@ -1508,7 +1508,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tif (!from_stdin) {\n \t\tclose(input_fd);\n \t} else {\n-\t\tfsync_or_die(output_fd, curr_pack_name);\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n \t\terr = close(output_fd);\n \t\tif (err)\n \t\t\tdie_errno(_(\"error while closing pack file\"));\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 857be7826f3..916c55d6ce9 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n \t\t * If so, rewrite it like in fast-import\n \t\t */\n \t\tif (pack_to_stdout) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n \t\t} else if (nr_written == nr_remaining) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t\t} else {\n-\t\t\tint fd = finalize_hashfile(f, hash, 0);\n+\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n \t\t\t\t\t\t nr_written, hash, offset);\n \t\t\tclose(fd);\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..a2cf9dcbc8d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -53,9 +53,10 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \t\tunlink(state->pack_tmp_name);\n \t\tgoto clear_exit;\n \t} else if (state->nr_written == 1) {\n-\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\tfinalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t} else {\n-\t\tint fd = finalize_hashfile(state->f, hash, 0);\n+\t\tint fd = finalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n \t\t\t\t\t state->nr_written, hash,\n \t\t\t\t\t state->offset);\ndiff --git a/cache.h b/cache.h\nindex 9cd60d94952..d83fbaf2619 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,38 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+/*\n+ * These values are used to help identify parts of a repository to fsync.\n+ * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n+ * repository and so shouldn't be fsynced.\n+ */\n+enum fsync_component {\n+\tFSYNC_COMPONENT_NONE\t\t\t= 0,\n+\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n+\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n+\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n+\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+};\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t      FSYNC_COMPONENT_PACK | \\\n+\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+\n+/*\n+ * A bitmask indicating which components of the repo should be fsynced.\n+ */\n+extern enum fsync_component fsync_components;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n@@ -1747,6 +1778,12 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n \n+inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n+{\n+\tif (fsync_components & component)\n+\t\tfsync_or_die(fd, msg);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 2706683acfe..c8a5dea4541 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1939,7 +1939,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t}\n \n \tclose_commit_graph(ctx->r->objects);\n-\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n+\tfinalize_hashfile(f, file_hash, FSYNC_COMPONENT_COMMIT_GRAPH,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n \tfree_chunkfile(cf);\n \n \tif (ctx->split) {\ndiff --git a/config.c b/config.c\nindex 139df71ba17..5ab381388f9 100644\n--- a/config.c\n+++ b/config.c\n@@ -1213,6 +1213,73 @@ static int git_parse_maybe_bool_text(const char *value)\n \treturn -1;\n }\n \n+static const struct fsync_component_entry {\n+\tconst char *name;\n+\tenum fsync_component component_bits;\n+} fsync_component_table[] = {\n+\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n+\t{ \"pack\", FSYNC_COMPONENT_PACK },\n+\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n+\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n+};\n+\n+static enum fsync_component parse_fsync_components(const char *var, const char *string)\n+{\n+\tenum fsync_component output = 0;\n+\n+\tif (!strcmp(string, \"none\"))\n+\t\treturn output;\n+\n+\twhile (string) {\n+\t\tint i;\n+\t\tsize_t len;\n+\t\tconst char *ep;\n+\t\tint negated = 0;\n+\t\tint found = 0;\n+\n+\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n+\t\tep = strchrnul(string, ',');\n+\t\tlen = ep - string;\n+\n+\t\tif (*string == '-') {\n+\t\t\tnegated = 1;\n+\t\t\tstring++;\n+\t\t\tlen--;\n+\t\t\tif (!len)\n+\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n+\t\t}\n+\n+\t\tif (!len)\n+\t\t\tbreak;\n+\n+\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n+\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n+\n+\t\t\tif (strncmp(entry->name, string, len))\n+\t\t\t\tcontinue;\n+\n+\t\t\tfound = 1;\n+\t\t\tif (negated)\n+\t\t\t\toutput &= ~entry->component_bits;\n+\t\t\telse\n+\t\t\t\toutput |= entry->component_bits;\n+\t\t}\n+\n+\t\tif (!found) {\n+\t\t\tchar *component = xstrndup(string, len);\n+\t\t\twarning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n+\t\t\tfree(component);\n+\t\t}\n+\n+\t\tstring = ep;\n+\t}\n+\n+\treturn output;\n+}\n+\n int git_parse_maybe_bool(const char *value)\n {\n \tint v = git_parse_maybe_bool_text(value);\n@@ -1490,6 +1557,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsync\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tfsync_components = parse_fsync_components(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncmethod\")) {\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n@@ -1503,7 +1577,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n \t\treturn 0;\n \t}\n \ndiff --git a/csum-file.c b/csum-file.c\nindex 26e8a6df44e..59ef3398ca2 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -58,7 +58,8 @@ static void free_hashfile(struct hashfile *f)\n \tfree(f);\n }\n \n-int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n+int finalize_hashfile(struct hashfile *f, unsigned char *result,\n+\t\t      enum fsync_component component, unsigned int flags)\n {\n \tint fd;\n \n@@ -69,7 +70,7 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int fl\n \tif (flags & CSUM_HASH_IN_STREAM)\n \t\tflush(f, f->buffer, the_hash_algo->rawsz);\n \tif (flags & CSUM_FSYNC)\n-\t\tfsync_or_die(f->fd, f->name);\n+\t\tfsync_component_or_die(component, f->fd, f->name);\n \tif (flags & CSUM_CLOSE) {\n \t\tif (close(f->fd))\n \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\ndiff --git a/csum-file.h b/csum-file.h\nindex 291215b34eb..0d29f528fbc 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -1,6 +1,7 @@\n #ifndef CSUM_FILE_H\n #define CSUM_FILE_H\n \n+#include \"cache.h\"\n #include \"hash.h\"\n \n struct progress;\n@@ -38,7 +39,7 @@ int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n struct hashfile *hashfd(int fd, const char *name);\n struct hashfile *hashfd_check(const char *name);\n struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n-int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n+int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n void hashwrite(struct hashfile *, const void *, unsigned int);\n void hashflush(struct hashfile *f);\n void crc32_begin(struct hashfile *);\ndiff --git a/environment.c b/environment.c\nindex f9140e842cf..09905adecf9 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,6 +42,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n+enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/midx.c b/midx.c\nindex 837b46b2af5..882f91f7d57 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -1406,7 +1406,8 @@ static int write_midx_internal(const char *object_dir,\n \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n \twrite_chunkfile(cf, &ctx);\n \n-\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n \tfree_chunkfile(cf);\n \n \tif (flags & (MIDX_WRITE_REV_INDEX | MIDX_WRITE_BITMAP))\ndiff --git a/object-file.c b/object-file.c\nindex eb972cdccd2..9d9c4a39e85 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1809,8 +1809,7 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex 9c55c1531e1..c16e43d1669 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -719,7 +719,8 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \tif (options & BITMAP_OPT_HASH_CACHE)\n \t\twrite_hash_cache(f, index, index_nr);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \n \tif (adjust_shared_perm(tmp_file.buf))\n \t\tdie_errno(\"unable to make temporary bitmap file readable\");\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a346..51812cb1299 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -159,9 +159,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t}\n \n \thashwrite(f, sha1, the_hash_algo->rawsz);\n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n-\t\t\t\t    ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \treturn index_name;\n }\n \n@@ -281,8 +281,9 @@ const char *write_rev_file_order(const char *rev_name,\n \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n \t\tdie(_(\"failed to make %s readable\"), rev_name);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \n \treturn rev_name;\n }\n@@ -390,7 +391,7 @@ void fixup_pack_header_footer(int pack_fd,\n \t\tthe_hash_algo->final_fn(partial_pack_hash, &old_hash_ctx);\n \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n-\tfsync_or_die(pack_fd, pack_name);\n+\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n }\n \n char *index_pack_lockfile(int ip_out, int *is_well_formed)\ndiff --git a/read-cache.c b/read-cache.c\nindex f3986596623..f3539681f49 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -3060,7 +3060,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"443555","messageId":"2207950beba89b690690f98c77761c27c5da8dcc.1639011434.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v3.git.1639011433.gitgitgadget@gmail.com","subject":"[PATCH v3 3/4] core.fsync: new option to harden the index","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-09T00:57:12Z","receivedAt":"2021-12-09T00:57:26Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the new ability for the user to harden\nthe index. In the event of a system crash, the index must be\ndurable for the user to actually find a file that has been added\nto the repo and then deleted from the working tree.\n\nWe use the presence of the COMMIT_LOCK flag and absence of the\nalternate_index_output as a proxy for determining whether we're\nupdating the persistent index of the repo or some temporary\nindex. We don't sync these temporary indexes.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  1 +\n cache.h                       |  4 +++-\n config.c                      |  1 +\n read-cache.c                  | 19 +++++++++++++------\n 4 files changed, 18 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 4f1747ec871..8e5b7a795ab 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -561,6 +561,7 @@ core.fsync::\n * `pack` hardens objects added to the repo in packfile form.\n * `pack-metadata` hardens packfile bitmaps and indexes.\n * `commit-graph` hardens the commit graph file.\n+* `index` hardens the index when it is modified.\n * `objects` is an aggregate option that includes `loose-objects`, `pack`,\n   `pack-metadata`, and `commit-graph`.\n * `default` is an aggregate option that is equivalent to `objects,-loose-object`\ndiff --git a/cache.h b/cache.h\nindex d83fbaf2619..4dc26d7b2c9 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -996,6 +996,7 @@ enum fsync_component {\n \tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+\tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n #define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n@@ -1010,7 +1011,8 @@ enum fsync_component {\n #define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n \t\t\t      FSYNC_COMPONENT_PACK | \\\n \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n+\t\t\t      FSYNC_COMPONENT_INDEX)\n \n \n /*\ndiff --git a/config.c b/config.c\nindex 5ab381388f9..b3e7006c68e 100644\n--- a/config.c\n+++ b/config.c\n@@ -1221,6 +1221,7 @@ static const struct fsync_component_entry {\n \t{ \"pack\", FSYNC_COMPONENT_PACK },\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"index\", FSYNC_COMPONENT_INDEX },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n \t{ \"all\", FSYNC_COMPONENTS_ALL },\ndiff --git a/read-cache.c b/read-cache.c\nindex f3539681f49..783cb3ea5db 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2816,7 +2816,7 @@ static int record_ieot(void)\n  * rely on it.\n  */\n static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n-\t\t\t  int strip_extensions)\n+\t\t\t  int strip_extensions, unsigned flags)\n {\n \tuint64_t start = getnanotime();\n \tstruct hashfile *f;\n@@ -2830,6 +2830,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint csum_fsync_flag;\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n@@ -3060,7 +3061,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n+\tcsum_fsync_flag = 0;\n+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n+\t\tcsum_fsync_flag = CSUM_FSYNC;\n+\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n+\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n@@ -3115,7 +3122,7 @@ static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n \t */\n \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n-\tret = do_write_index(istate, lock->tempfile, 0);\n+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n \n@@ -3209,7 +3216,7 @@ static int clean_shared_index_files(const char *current_hex)\n }\n \n static int write_shared_index(struct index_state *istate,\n-\t\t\t      struct tempfile **temp)\n+\t\t\t      struct tempfile **temp, unsigned flags)\n {\n \tstruct split_index *si = istate->split_index;\n \tint ret, was_full = !istate->sparse_index;\n@@ -3219,7 +3226,7 @@ static int write_shared_index(struct index_state *istate,\n \n \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n-\tret = do_write_index(si->base, *temp, 1);\n+\tret = do_write_index(si->base, *temp, 1, flags);\n \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n \n@@ -3328,7 +3335,7 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n \t\t\tret = do_write_locked_index(istate, lock, flags);\n \t\t\tgoto out;\n \t\t}\n-\t\tret = write_shared_index(istate, &temp);\n+\t\tret = write_shared_index(istate, &temp, flags);\n \n \t\tsaved_errno = errno;\n \t\tif (is_tempfile_active(temp))\n-- \ngitgitgadget\n\n"},{"id":"443556","messageId":"a830d177d4cd6cebde24e2c51c2185801b6d57dc.1639011434.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v3.git.1639011433.gitgitgadget@gmail.com","subject":"[PATCH v3 4/4] core.fsync: add a `derived-metadata` aggregate option","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-12-09T00:57:13Z","receivedAt":"2021-12-09T00:57:27Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds an aggregate option that currently includes the\ncommit-graph file and pack metadata (indexes and bitmaps).\n\nThe user may want to exclude this set from durability since they can be\nrecomputed from other data if they wind up corrupt or missing.\n\nThis is split out from the other patches in the series since it is\nan optional nice-to-have that might be controversial.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 6 +++---\n cache.h                       | 7 ++++---\n config.c                      | 1 +\n 3 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 8e5b7a795ab..21092f3a4d1 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -562,9 +562,9 @@ core.fsync::\n * `pack-metadata` hardens packfile bitmaps and indexes.\n * `commit-graph` hardens the commit graph file.\n * `index` hardens the index when it is modified.\n-* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n-  `pack-metadata`, and `commit-graph`.\n-* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n+* `objects` is an aggregate option that includes `loose-objects` and `pack`.\n+* `derived-metadata` is an aggregate option that includes `pack-metadata` and `commit-graph`.\n+* `default` is an aggregate option that is equivalent to `objects,derived-metadata,-loose-object`\n * `all` is an aggregate option that syncs all individual components above.\n \n core.fsyncMethod::\ndiff --git a/cache.h b/cache.h\nindex 4dc26d7b2c9..cc1c084242e 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1004,9 +1004,10 @@ enum fsync_component {\n \t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n \n #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n-\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n-\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\t\t\t\t  FSYNC_COMPONENT_PACK)\n+\n+#define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n \n #define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n \t\t\t      FSYNC_COMPONENT_PACK | \\\ndiff --git a/config.c b/config.c\nindex b3e7006c68e..d9ef3ef0060 100644\n--- a/config.c\n+++ b/config.c\n@@ -1223,6 +1223,7 @@ static const struct fsync_component_entry {\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n \t{ \"all\", FSYNC_COMPONENTS_ALL },\n };\n-- \ngitgitgadget\n"},{"id":"443564","messageId":"211209.86bl1ql718.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"CANQDOddkKbUC-g97JOf40nS28Yv1KACvbjW9gtQZemfBMutPCw@mail.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-09T04:08:50Z","receivedAt":"2021-12-09T04:46:35Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Dec 08 2021, Neeraj Singh wrote:\n\n> On Wed, Dec 8, 2021 at 2:17 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>\n>>\n>> On Tue, Dec 07 2021, Neeraj Singh wrote:\n>>\n>> > On Tue, Dec 7, 2021 at 5:01 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>> >>\n>> >>\n>> >> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n>> >>\n>> >> > From: Neeraj Singh <neerajsi@microsoft.com>\n>> >> >\n>> >> > This commit introduces the `core.fsync` configuration\n>> >> > knob which can be used to control how components of the\n>> >> > repository are made durable on disk.\n>> >> >\n>> >> > This setting allows future extensibility of the list of\n>> >> > syncable components:\n>> >> > * We issue a warning rather than an error for unrecognized\n>> >> >   components, so new configs can be used with old Git versions.\n>> >>\n>> >> Looks good!\n>> >>\n>> >> > * We support negation, so users can choose one of the default\n>> >> >   aggregate options and then remove components that they don't\n>> >> >   want. The user would then harden any new components added in\n>> >> >   a Git version update.\n>> >>\n>> >> I think this config schema makes sense, but just a (I think important)\n>> >> comment on the \"how\" not \"what\" of it. It's really much better to define\n>> >> config as:\n>> >>\n>> >>     [some]\n>> >>     key = value\n>> >>     key = value2\n>> >>\n>> >> Than:\n>> >>\n>> >>     [some]\n>> >>     key = value,value2\n>> >>\n>> >> The reason is that \"git config\" has good support for working with\n>> >> multi-valued stuff, so you can do e.g.:\n>> >>\n>> >>     git config --get-all -z some.key\n>> >>\n>> >> And you can easily (re)set such config e.g. with --replace-all etc., but\n>> >> for comma-delimited you (and users) need to do all that work themselves.\n>> >>\n>> >> Similarly instead of:\n>> >>\n>> >>     some.key = want-this\n>> >>     some.key = -not-this\n>> >>     some.key = but-want-this\n>> >>\n>> >> I think it's better to just have two lists, one inclusive another\n>> >> exclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n>> >> \"transfer.hideRefs\"\n>> >>\n>> >> Which would mean:\n>> >>\n>> >>     core.fsync = want-this\n>> >>     core.fsyncExcludes = -not-this\n>> >>\n>> >> For some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\n>> >> suggestion on making this easier for users & the implementation.\n>> >>\n>> >\n>> > Maybe there's some way to handle this I'm unaware of, but a\n>> > disadvantage of your multi-valued config proposal is that it's harder,\n>> > for example, for a per-repo config store to reasonably override a\n>> > per-user config store.  With the configuration scheme as-is, I can\n>> > have a per-user setting like `core.fsync=all` which covers my typical\n>> > repos, but then have a maintainer repo with a private setting of\n>> > `core.fsync=none` to speed up cases where I'm mostly working with\n>> > other people's changes that are backed up in email or server-side\n>> > repos.  The latter setting conveniently overrides the former setting\n>> > in all aspects.\n>>\n>> Even if you turn just your comma-delimited proposal into a list proposal\n>> can't we just say that the last one wins? Then it won't matter what cmae\n>> before, you'd specify \"core.fsync=none\" in your local .git/config.\n>>\n>> But this is also a general issue with a bunch of things in git's config\n>> space. I'd rather see us use the list-based values and just come up with\n>> some general way to reset them that works for all keys, rather than\n>> regretting having comma-delimited values that'll be harder to work with\n>> & parse, which will be a legacy wart if/when we come up with a way to\n>> say \"reset all previous settings\".\n>>\n>> > Also, with the core.fsync and core.fsyncExcludes, how would you spell\n>> > \"don't sync anything\"? Would you still have the aggregate options.?\n>>\n>>     core.fsyncExcludes = *\n>>\n>> I.e. the same as the core.fsync=none above, anyway I still like the\n>> wildcard thing below a bit more...\n>\n> I'm not going to take this feedback unless there are additional votes\n> from the Git community in this direction.  I make the claim that\n> single-valued comma-separated config lists are easier to work with in\n> the existing Git infrastructure. \n\nEasier in what sense? I showed examples of how \"git-config\" trivially\nworks with multi-valued config, but for comma-delimited you'll need to\nwrite your own shellscript boilerplate around simple things like adding\nvalues, removing existing ones etc.\n\nI.e. you can use --add, --unset, --unset-all, --get, --get-all etc.\n\n> parsing code for the core.whitespace variable and users are used to\n> this syntax there. There are several other comma-separated lists in\n> the config space, so this construct has precedence and will be with\n> Git for some time.\n\nThat's not really an argument either way for why we'd pick X over Y for\nsomething new. We've got some comma-delimited, some multi-value (I'm\nfairly sure we have more multi-value).\n\n> Also, fsync configurations aren't composable like\n> some other configurations may be. It makes sense to have a holistic\n> singular fsync configuration, which is best represented by a single\n> variable.\n\nWhat's a \"variable\" here? We call these \"keys\", you can have a\nsingle-value key like user.name that you get with --get, or a\nmulti-value key like say branch.<name>.merge or push.pushOption that\nyou'd get with --get-all.\n\nI think you may be referring to either not wanting these to be\n\"inherited\" (which is not really a think we do for anything else in\nconfig), or lacking the ability to \"reset\".\n\nFor the latter if that's absolutely needed we could just use the same\ntrick as \"diff.wsErrorHighlight\" uses of making an empty value \"reset\"\nthe list, and you'd get better \"git config\" support for editing it.\n\n>> >> > This also supports the common request of doing absolutely no\n>> >> > fysncing with the `core.fsync=none` value, which is expected\n>> >> > to make the test suite faster.\n>> >>\n>> >> Let's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\n>> >> so we'll accept \"false\", \"off\", \"no\" like most other such config?\n>> >\n>> > Junio's previous feedback when discussing batch mode [1] was to offer\n>> > less flexibility when parsing new values of these configuration\n>> > options. I agree with his statement that \"making machine-readable\n>> > tokens be spelled in different ways is a 'disease'.\"  I'd like to\n>> > leave this as-is so that the documentation can clearly state the exact\n>> > set of allowable values.\n>> >\n>> > [1] https://lore.kernel.org/git/xmqqr1dqzyl7.fsf@gitster.g/\n>>\n>> I think he's talking about batch, Batch, BATCH, bAtCh etc. there. But\n>> the \"maybe bool\" is a stanard pattern we use.\n>>\n>> I don't think we'd call one of these 0, off, no or false etc. to avoid\n>> confusion, so then you can use git_parse_maybe_...()\n>\n> I don't see the advantage of having multiple ways of specifying\n> \"none\".  The user can read the doc and know exactly what to write.  If\n> they write something unallowable, they get a clear warning and they\n> can read the doc again to figure out what to write.  This isn't a\n> boolean options at all, so why should we entertain bool-like ways of\n> spelling it?\n\nIt's not boolean, it's multi-value and one of the values includes a true\nor false boolean value. You just spell it \"none\".\n\nI think both this and your comment above suggest that you think there's\nno point in this because you haven't interacted with/used \"git config\"\nas a command line or API mechanism, but have just hand-crafted config\nfiles.\n\nThat's fair enough, but there's a lot of tooling that benefits from the\nlatter.\n\nE.g.:\n    \n    $ git -c core.fsync=off config --type=bool core.fsync\n    false\n    $ git -c core.fsync=blah config --type=bool core.fsync\n    fatal: bad boolean config value 'blah' for 'core.fsync'\n\nHere we can get 'git config' to normalize what you call 'none', and you\ncan tell via exit codes/normalization if it's \"false\". But if you invent\na new term for \"false\" you can't do that as easily.\n\nWe have various historical keys that take odd exceptions to that,\ne.g. \"never\", but unless we have a good reason to let's not invent more\nexceptions.\n\n>> >> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>> >> > ---\n>> >> >  Documentation/config/core.txt | 27 +++++++++----\n>> >> >  builtin/fast-import.c         |  2 +-\n>> >> >  builtin/index-pack.c          |  4 +-\n>> >> >  builtin/pack-objects.c        |  8 ++--\n>> >> >  bulk-checkin.c                |  5 ++-\n>> >> >  cache.h                       | 39 +++++++++++++++++-\n>> >> >  commit-graph.c                |  3 +-\n>> >> >  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n>> >> >  csum-file.c                   |  5 ++-\n>> >> >  csum-file.h                   |  3 +-\n>> >> >  environment.c                 |  1 +\n>> >> >  midx.c                        |  3 +-\n>> >> >  object-file.c                 |  3 +-\n>> >> >  pack-bitmap-write.c           |  3 +-\n>> >> >  pack-write.c                  | 13 +++---\n>> >> >  read-cache.c                  |  2 +-\n>> >> >  16 files changed, 164 insertions(+), 33 deletions(-)\n>> >> >\n>> >> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n>> >> > index dbb134f7136..4f1747ec871 100644\n>> >> > --- a/Documentation/config/core.txt\n>> >> > +++ b/Documentation/config/core.txt\n>> >> > @@ -547,6 +547,25 @@ core.whitespace::\n>> >> >    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n>> >> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n>> >> >\n>> >> > +core.fsync::\n>> >> > +     A comma-separated list of parts of the repository which should be\n>> >> > +     hardened via the core.fsyncMethod when created or modified. You can\n>> >> > +     disable hardening of any component by prefixing it with a '-'. Later\n>> >> > +     items take precedence over earlier ones in the list. For example,\n>> >> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n>> >> > +     metadata.\" Items that are not hardened may be lost in the event of an\n>> >> > +     unclean system shutdown.\n>> >> > ++\n>> >> > +* `none` disables fsync completely. This must be specified alone.\n>> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n>> >> > +* `pack` hardens objects added to the repo in packfile form.\n>> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n>> >> > +* `commit-graph` hardens the commit graph file.\n>> >> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n>> >> > +  `pack-metadata`, and `commit-graph`.\n>> >> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n>> >> > +* `all` is an aggregate option that syncs all individual components above.\n>> >> > +\n>> >>\n>> >> It's probably a *bit* more work to set up, but I wonder if this wouldn't\n>> >> be simpler if we just said (and this is partially going against what I\n>> >> noted above):\n>> >>\n>> >> == BEGIN DOC\n>> >>\n>> >> core.fsync is a multi-value config variable where each item is a\n>> >> pathspec that'll get matched the same way as 'git-ls-files' et al.\n>> >>\n>> >> When we sync pretend that a path like .git/objects/de/adbeef... is\n>> >> relative to the top-level of the git\n>> >> directory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n>> >>\n>> >> You can then supply a list of wildcards and exclusions to configure\n>> >> syncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\n>> >> with:\n>> >>\n>> >>     ; same as \"false\"\n>> >>     core.fsync = \":!*\"\n>> >>\n>> >> Or:\n>> >>\n>> >>     ; same as \"true\"\n>> >>     core.fsync = \"*\"\n>> >>\n>> >> Or, to selectively sync some things and not others:\n>> >>\n>> >>     ;; Sync objects, but not \"info\"\n>> >>     core.fsync = \":!objects/info/**\"\n>> >>     core.fsync = \"objects/**\"\n>> >>\n>> >> See gitrepository-layout(5) for details about what sort of paths you\n>> >> might be expected to match. Not all paths listed there will go through\n>> >> this mechanism (e.g. currently objects do, but nothing to do with config\n>> >> does).\n>> >>\n>> >> We can and will match this against \"fake paths\", e.g. when writing out\n>> >> packs we may match against just the string \"objects/pack\", we're not\n>> >> going to re-check if every packfile we're writing matches your globs,\n>> >> ditto for loose objects. Be reasonable!\n>> >>\n>> >> This metharism is intended as a shorthand that provides some flexibility\n>> >> when fsyncing, while not forcing git to come up with labels for all\n>> >> paths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n>> >>\n>> >> More paths may be added or removed in the future, and we make no\n>> >> promises that we won't move things around, so if in doubt use\n>> >> e.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\n>> >> stick to the golden path of examples provided in this documentation.\n>> >>\n>> >> == END DOC\n>> >>\n>> >>\n>> >> It's a tad more complex to set up, but I wonder if that isn't worth\n>> >> it. It nicely gets around any current and future issues of deciding what\n>> >> labels such as \"loose-object\" etc. to pick, as well as slotting into an\n>> >> existing method of doing exclude/include lists.\n>> >>\n>> >\n>> > I think this proposal is a lot of complexity to avoid coming up with a\n>> > new name for syncable things as they are added to Git.  A path based\n>> > mechanism makes it hard to document for the (advanced) user what the\n>> > full set of things is and how it might change from release to release.\n>> > I think the current core.fsync scheme is a bit easier to understand,\n>> > query, and extend.\n>>\n>> We document it in gitrepository-layout(5). Yeah it has some\n>> disadvantages, but one advantage is that you could make the\n>> composability easy. I.e. if last exclude wins then a setting of:\n>>\n>>     core.fsync = \":!*\"\n>>     core.fsync = \"objects/**\"\n>>\n>> Would reset all previous matches & only match objects/**.\n>>\n>\n> The value of changing this is predicated on taking your previous\n> multi-valued config proposal, which I'm still not at all convinced\n> about.\n\nThey're orthagonal. I.e. you get benefits from multi-value with or\nwithout this globbing mechanism.\n\nIn any case, I don't feel strongly about/am really advocating this\nglobbing mechanism. I just wondered if it wouldn't make things simpler\nsince it would sidestep the need to create any sort of categories for\nsubsets of gitrepository-layout(5), but maybe not...\n\n> The schema in the current (v1-v2) version of the patch already\n> includes an example of extending the list of syncable things, and\n> Patrick Steinhardt made it clear that he feels comfortable adding\n> 'refs' to the same schema in a future change.\n>\n> I'll also emphasize that we're talking about a non-functional,\n> relatively corner-case behavioral configuration.  These values don't\n> change how git's interface behaves except when the system crashes\n> during a git command or shortly after one completes.\n\nThat may be something some OS's promise, but it's not something fsync()\nor POSIX promises. I.e. you might write a ref, but unless you fsync and\nthe relevant dir entries another process might not see the update, crash\nor not.\n\nThat's an aside from these config design questions, and I think\nmost/(all?) OS's anyone cares about these days tend to make that\nimplicit promise as part of their VFS behavior, but we should probably\ndesign such an interface to fsync() with such pedantic portability in\nmind.\n\n> While you may not personally love the proposed configuration\n> interface, I'd want your view on some questions:\n> 1. Is it easy for the (advanced) user to set a configuration?\n> 2. Is it easy for the (advanced) user to see what was configured?\n> 3. Is it easy for the Git community to build on this as we want to add\n> things to the list of things to sync?\n>     a) Is there a good best practice configuration so that people can\n> avoid losing integrity for new stuff that they are intending to sync.\n>     b) If someone has a custom configuration, can that custom\n> configuration do something reasonable as they upgrade versions of Git?\n>              ** In response to this question, I might see some value\n> in adding a 'derived-metadata' aggregate that can be disabled so that\n> a custom configuration can exclude those as they change version to\n> version.\n>     c) Is it too much maintenance overhead to consider how to present\n> this configuration knob for any new hashfile or other datafile in the\n> git repo?\n> 4. Is there a good path forward to change the default syncable set,\n> both in git-for-windows and in Git for other platforms?\n\nI'm not really sure this globbing this is a good idea, as noted above\njust a suggestion etc.\n\nAs noted there it just gets you out of the business of re-defining\ngitrepository-layout(5), and assuming too much in advance about certain\nuse-cases.\n\nE.g. even \"refs\" might be too broad for some. I don't tend to be I/O\nlimited, but I could see how someone who would be would care about\nrefs/heads but not refs/remotes, or want to exclude logs/* but not the\nrefs updates themselves etc.\n\n>> >> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n>> >> > index 857be7826f3..916c55d6ce9 100644\n>> >> > --- a/builtin/pack-objects.c\n>> >> > +++ b/builtin/pack-objects.c\n>> >> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n>> >> >                * If so, rewrite it like in fast-import\n>> >> >                */\n>> >> >               if (pack_to_stdout) {\n>> >> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>> >> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n>> >> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>> >>\n>> >> Not really related to this per-se, but since you're touching the API\n>> >> everything goes through I wonder if callers should just always try to\n>> >> fsync, and we can just catch EROFS and EINVAL in the wrapper if someone\n>> >> tries to flush stdout, or catch the fd at that lower level.\n>> >>\n>> >> Or maybe there's a good reason for this...\n>> >\n>> > It's platform dependent, but I'd expect fsync would do something for\n>> > pipes or stdout redirected to a file.  In these cases we really don't\n>> > want to fsync since we have no idea what we're talking to and we're\n>> > potentially worsening performance for probably no benefit.\n>>\n>> Yeah maybe we should just leave it be.\n>>\n>> I'd think the C library returning EINVAL would be a trivial performance\n>> cost though.\n>>\n>> It just seemed odd to hardcode assumptions about what can and can't be\n>> synced when the POSIX defined function will also tell us that.\n>>\n>\n> Redirecting stdout to a file seems like a common usage for this\n> command. That would definitely be fsyncable, but Git has no idea what\n> its proper category is since there's no way to know the purpose or\n> lifetime of the packfile.  I'm going to leave this be, because I'd\n> posit that \"can it be fsynced?\" is not the same as \"should it be\n> fsynced?\".  The latter question can't be answered for stdout.\n\nAs noted this was just an aside, and I don't even know if any OS would\ndo anything meaningful with an fsync() of such a FD anyway.\n\nI just don't see why we wouldn't say:\n\n 1. We're syncing this category of thing\n 2. Try it\n 3. If fsync returns \"can't fsync that sort of thing\" move on\n\nAs opposed to trying to shortcut #3 by doing the detection ourselves.\n\nI.e. maybe there was a good reason, but it seemed to be some easy\npotential win for more simplification, since you were re-doing and\nsimplifying some of the interface anyway...\n\n>>\n>> >> > [...]\n>> >> > +/*\n>> >> > + * These values are used to help identify parts of a repository to fsync.\n>> >> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n>> >> > + * repository and so shouldn't be fsynced.\n>> >> > + */\n>> >> > +enum fsync_component {\n>> >> > +     FSYNC_COMPONENT_NONE                    = 0,\n>> >>\n>> >> I haven't read ahead much but in most other such cases we don't define\n>> >> the \"= 0\", just start at 1<<0, then check the flags elsewhere...\n>> >>\n>> >> > +static const struct fsync_component_entry {\n>> >> > +     const char *name;\n>> >> > +     enum fsync_component component_bits;\n>> >> > +} fsync_component_table[] = {\n>> >> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n>> >> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n>> >> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n>> >> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n>> >> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n>> >> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n>> >> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n>> >> > +};\n>> >> > +\n>> >> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n>> >> > +{\n>> >> > +     enum fsync_component output = 0;\n>> >> > +\n>> >> > +     if (!strcmp(string, \"none\"))\n>> >> > +             return output;\n>> >> > +\n>> >> > +     while (string) {\n>> >> > +             int i;\n>> >> > +             size_t len;\n>> >> > +             const char *ep;\n>> >> > +             int negated = 0;\n>> >> > +             int found = 0;\n>> >> > +\n>> >> > +             string = string + strspn(string, \", \\t\\n\\r\");\n>> >>\n>> >> Aside from the \"use a list\" isn't this hardcoding some windows-specific\n>> >> assumptions with \\n\\r? Maybe not...\n>> >\n>> > I shamelessly stole this code from parse_whitespace_rule. I thought\n>> > about making a helper to be called by both functions, but the amount\n>> > of state going into and out of the wrapper via arguments was\n>> > substantial and seemed to negate the benefit of deduplication.\n>>\n>> FWIW string_list_split() is easier to work with in those cases, or at\n>> least I think so...\n>\n> This code runs at startup for a variable that may be present on some\n> installations.  The nice property of the current patch's code is that\n> it's already a well-tested pattern that doesn't do any allocations as\n> it's working, unlike string_list_split().\n\nMulti-value config would also get you fewer allocations :)\n\nAnyway, I mainly meant to point out that for stuff like this it's fine\nto optimize it for ease rather than micro-optimize allocations. Those\nreally aren't a bottleneck on this scale.\n\nEven in that case there's string_list_split_in_place(), which can be a\nbit nicer than manual C-string fiddling.\n\n> I hope you know that I appreciate your review feedback, even though\n> I'm pushing back on most of it so far this round. I'll be sending v3\n> to the list soon after giving it another look over.\n\nSure, no worries. Just hoping to help. If you go for something different\netc. that's fine. Just hoping to bridge the gap in some knowledge /\noffer potentially interesting suggestions (some of which may be dead\nends, like the config glob thing...).\n"},{"id":"443606","messageId":"CANQDOdchh3mfC8S6ouWAQbtWzZUkmTzF1p5D9dg4muoBu4N1Fg@mail.gmail.com","threadId":"57030","inReplyTo":"211209.86bl1ql718.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-12-09T06:18:03Z","receivedAt":"2021-12-09T06:18:20Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Dec 8, 2021 at 8:46 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Wed, Dec 08 2021, Neeraj Singh wrote:\n>\n> > On Wed, Dec 8, 2021 at 2:17 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Tue, Dec 07 2021, Neeraj Singh wrote:\n> >>\n> >> > On Tue, Dec 7, 2021 at 5:01 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> >> >>\n> >> >>\n> >> >> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n> >> >>\n> >> >> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >> >> >\n> >> >> > This commit introduces the `core.fsync` configuration\n> >> >> > knob which can be used to control how components of the\n> >> >> > repository are made durable on disk.\n> >> >> >\n> >> >> > This setting allows future extensibility of the list of\n> >> >> > syncable components:\n> >> >> > * We issue a warning rather than an error for unrecognized\n> >> >> >   components, so new configs can be used with old Git versions.\n> >> >>\n> >> >> Looks good!\n> >> >>\n> >> >> > * We support negation, so users can choose one of the default\n> >> >> >   aggregate options and then remove components that they don't\n> >> >> >   want. The user would then harden any new components added in\n> >> >> >   a Git version update.\n> >> >>\n> >> >> I think this config schema makes sense, but just a (I think important)\n> >> >> comment on the \"how\" not \"what\" of it. It's really much better to define\n> >> >> config as:\n> >> >>\n> >> >>     [some]\n> >> >>     key = value\n> >> >>     key = value2\n> >> >>\n> >> >> Than:\n> >> >>\n> >> >>     [some]\n> >> >>     key = value,value2\n> >> >>\n> >> >> The reason is that \"git config\" has good support for working with\n> >> >> multi-valued stuff, so you can do e.g.:\n> >> >>\n> >> >>     git config --get-all -z some.key\n> >> >>\n> >> >> And you can easily (re)set such config e.g. with --replace-all etc., but\n> >> >> for comma-delimited you (and users) need to do all that work themselves.\n> >> >>\n> >> >> Similarly instead of:\n> >> >>\n> >> >>     some.key = want-this\n> >> >>     some.key = -not-this\n> >> >>     some.key = but-want-this\n> >> >>\n> >> >> I think it's better to just have two lists, one inclusive another\n> >> >> exclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n> >> >> \"transfer.hideRefs\"\n> >> >>\n> >> >> Which would mean:\n> >> >>\n> >> >>     core.fsync = want-this\n> >> >>     core.fsyncExcludes = -not-this\n> >> >>\n> >> >> For some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\n> >> >> suggestion on making this easier for users & the implementation.\n> >> >>\n> >> >\n> >> > Maybe there's some way to handle this I'm unaware of, but a\n> >> > disadvantage of your multi-valued config proposal is that it's harder,\n> >> > for example, for a per-repo config store to reasonably override a\n> >> > per-user config store.  With the configuration scheme as-is, I can\n> >> > have a per-user setting like `core.fsync=all` which covers my typical\n> >> > repos, but then have a maintainer repo with a private setting of\n> >> > `core.fsync=none` to speed up cases where I'm mostly working with\n> >> > other people's changes that are backed up in email or server-side\n> >> > repos.  The latter setting conveniently overrides the former setting\n> >> > in all aspects.\n> >>\n> >> Even if you turn just your comma-delimited proposal into a list proposal\n> >> can't we just say that the last one wins? Then it won't matter what cmae\n> >> before, you'd specify \"core.fsync=none\" in your local .git/config.\n> >>\n> >> But this is also a general issue with a bunch of things in git's config\n> >> space. I'd rather see us use the list-based values and just come up with\n> >> some general way to reset them that works for all keys, rather than\n> >> regretting having comma-delimited values that'll be harder to work with\n> >> & parse, which will be a legacy wart if/when we come up with a way to\n> >> say \"reset all previous settings\".\n> >>\n> >> > Also, with the core.fsync and core.fsyncExcludes, how would you spell\n> >> > \"don't sync anything\"? Would you still have the aggregate options.?\n> >>\n> >>     core.fsyncExcludes = *\n> >>\n> >> I.e. the same as the core.fsync=none above, anyway I still like the\n> >> wildcard thing below a bit more...\n> >\n> > I'm not going to take this feedback unless there are additional votes\n> > from the Git community in this direction.  I make the claim that\n> > single-valued comma-separated config lists are easier to work with in\n> > the existing Git infrastructure.\n>\n> Easier in what sense? I showed examples of how \"git-config\" trivially\n> works with multi-valued config, but for comma-delimited you'll need to\n> write your own shellscript boilerplate around simple things like adding\n> values, removing existing ones etc.\n>\n> I.e. you can use --add, --unset, --unset-all, --get, --get-all etc.\n>\n\nI see what you're saying for cases where someone would want to set a\ncore.fsync setting that's derived from the user's current config.  But\nI'm guessing that the dominant use case is someone setting a new fsync\nconfiguration that achieves some atomic goal with respect to a given\nrepository. Like \"this is a throwaway, so sync nothing\" or \"this is\nreally important, so sync all objects and refs and the index\".\n\n> > parsing code for the core.whitespace variable and users are used to\n> > this syntax there. There are several other comma-separated lists in\n> > the config space, so this construct has precedence and will be with\n> > Git for some time.\n>\n> That's not really an argument either way for why we'd pick X over Y for\n> something new. We've got some comma-delimited, some multi-value (I'm\n> fairly sure we have more multi-value).\n>\n\nMy main point here is that there's precedent for patch's current exact\nschema for a config in the same core config leaf of the Documentation.\nIt seems from your comments that we'd have to invent and document some\nnew convention for \"reset\" of a multi-valued config.  So you're\nsuggesting that I solve an extra set of problems to get this change\nin.  Just want to remind you that my personal itch to scratch is to\nget the underlying mechanism in so that git-for-windows can set its\ndefault setting to batch mode. I'm not expecting many users to\nactually configure this setting to any non-default value.\n\n> > Also, fsync configurations aren't composable like\n> > some other configurations may be. It makes sense to have a holistic\n> > singular fsync configuration, which is best represented by a single\n> > variable.\n>\n> What's a \"variable\" here? We call these \"keys\", you can have a\n> single-value key like user.name that you get with --get, or a\n> multi-value key like say branch.<name>.merge or push.pushOption that\n> you'd get with --get-all.\n\nYeah, I meant \"key\".  I conflated the config key with the underlying\nglobal variable in git.\n\n> I think you may be referring to either not wanting these to be\n> \"inherited\" (which is not really a think we do for anything else in\n> config), or lacking the ability to \"reset\".\n>\n> For the latter if that's absolutely needed we could just use the same\n> trick as \"diff.wsErrorHighlight\" uses of making an empty value \"reset\"\n> the list, and you'd get better \"git config\" support for editing it.\n>\n\nMy reading of the code is that diff.wsErrorHighlight is a comma\nseparated list and not a multi-valued config.  Actually I haven't yet\nfound an existing multi-valued config (not sure how to grep for it).\n\n> >> >> > This also supports the common request of doing absolutely no\n> >> >> > fysncing with the `core.fsync=none` value, which is expected\n> >> >> > to make the test suite faster.\n> >> >>\n> >> >> Let's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\n> >> >> so we'll accept \"false\", \"off\", \"no\" like most other such config?\n> >> >\n> >> > Junio's previous feedback when discussing batch mode [1] was to offer\n> >> > less flexibility when parsing new values of these configuration\n> >> > options. I agree with his statement that \"making machine-readable\n> >> > tokens be spelled in different ways is a 'disease'.\"  I'd like to\n> >> > leave this as-is so that the documentation can clearly state the exact\n> >> > set of allowable values.\n> >> >\n> >> > [1] https://lore.kernel.org/git/xmqqr1dqzyl7.fsf@gitster.g/\n> >>\n> >> I think he's talking about batch, Batch, BATCH, bAtCh etc. there. But\n> >> the \"maybe bool\" is a stanard pattern we use.\n> >>\n> >> I don't think we'd call one of these 0, off, no or false etc. to avoid\n> >> confusion, so then you can use git_parse_maybe_...()\n> >\n> > I don't see the advantage of having multiple ways of specifying\n> > \"none\".  The user can read the doc and know exactly what to write.  If\n> > they write something unallowable, they get a clear warning and they\n> > can read the doc again to figure out what to write.  This isn't a\n> > boolean options at all, so why should we entertain bool-like ways of\n> > spelling it?\n>\n> It's not boolean, it's multi-value and one of the values includes a true\n> or false boolean value. You just spell it \"none\".\n>\n> I think both this and your comment above suggest that you think there's\n> no point in this because you haven't interacted with/used \"git config\"\n> as a command line or API mechanism, but have just hand-crafted config\n> files.\n>\n> That's fair enough, but there's a lot of tooling that benefits from the\n> latter.\n\nMy batch mode perf tests (on github, not yet submitted to the list)\nuse `git -c core.fsync=<foo>` to set up a per-process config.  I\nhaven't used the `git config` writing support in a while, so I haven't\ndeeply thought about that.  However, I favor simplifying the use case\nof \"atomically\" setting a new holistic core.fsync value versus the use\ncase of deriving a new core.fsync value from the preexisting value.\n\n> E.g.:\n>\n>     $ git -c core.fsync=off config --type=bool core.fsync\n>     false\n>     $ git -c core.fsync=blah config --type=bool core.fsync\n>     fatal: bad boolean config value 'blah' for 'core.fsync'\n>\n> Here we can get 'git config' to normalize what you call 'none', and you\n> can tell via exit codes/normalization if it's \"false\". But if you invent\n> a new term for \"false\" you can't do that as easily.\n>\n> We have various historical keys that take odd exceptions to that,\n> e.g. \"never\", but unless we have a good reason to let's not invent more\n> exceptions.\n>\n> >> >> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> >> >> > ---\n> >> >> >  Documentation/config/core.txt | 27 +++++++++----\n> >> >> >  builtin/fast-import.c         |  2 +-\n> >> >> >  builtin/index-pack.c          |  4 +-\n> >> >> >  builtin/pack-objects.c        |  8 ++--\n> >> >> >  bulk-checkin.c                |  5 ++-\n> >> >> >  cache.h                       | 39 +++++++++++++++++-\n> >> >> >  commit-graph.c                |  3 +-\n> >> >> >  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n> >> >> >  csum-file.c                   |  5 ++-\n> >> >> >  csum-file.h                   |  3 +-\n> >> >> >  environment.c                 |  1 +\n> >> >> >  midx.c                        |  3 +-\n> >> >> >  object-file.c                 |  3 +-\n> >> >> >  pack-bitmap-write.c           |  3 +-\n> >> >> >  pack-write.c                  | 13 +++---\n> >> >> >  read-cache.c                  |  2 +-\n> >> >> >  16 files changed, 164 insertions(+), 33 deletions(-)\n> >> >> >\n> >> >> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> >> >> > index dbb134f7136..4f1747ec871 100644\n> >> >> > --- a/Documentation/config/core.txt\n> >> >> > +++ b/Documentation/config/core.txt\n> >> >> > @@ -547,6 +547,25 @@ core.whitespace::\n> >> >> >    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n> >> >> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >> >> >\n> >> >> > +core.fsync::\n> >> >> > +     A comma-separated list of parts of the repository which should be\n> >> >> > +     hardened via the core.fsyncMethod when created or modified. You can\n> >> >> > +     disable hardening of any component by prefixing it with a '-'. Later\n> >> >> > +     items take precedence over earlier ones in the list. For example,\n> >> >> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n> >> >> > +     metadata.\" Items that are not hardened may be lost in the event of an\n> >> >> > +     unclean system shutdown.\n> >> >> > ++\n> >> >> > +* `none` disables fsync completely. This must be specified alone.\n> >> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> >> >> > +* `pack` hardens objects added to the repo in packfile form.\n> >> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> >> >> > +* `commit-graph` hardens the commit graph file.\n> >> >> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n> >> >> > +  `pack-metadata`, and `commit-graph`.\n> >> >> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n> >> >> > +* `all` is an aggregate option that syncs all individual components above.\n> >> >> > +\n> >> >>\n> >> >> It's probably a *bit* more work to set up, but I wonder if this wouldn't\n> >> >> be simpler if we just said (and this is partially going against what I\n> >> >> noted above):\n> >> >>\n> >> >> == BEGIN DOC\n> >> >>\n> >> >> core.fsync is a multi-value config variable where each item is a\n> >> >> pathspec that'll get matched the same way as 'git-ls-files' et al.\n> >> >>\n> >> >> When we sync pretend that a path like .git/objects/de/adbeef... is\n> >> >> relative to the top-level of the git\n> >> >> directory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n> >> >>\n> >> >> You can then supply a list of wildcards and exclusions to configure\n> >> >> syncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\n> >> >> with:\n> >> >>\n> >> >>     ; same as \"false\"\n> >> >>     core.fsync = \":!*\"\n> >> >>\n> >> >> Or:\n> >> >>\n> >> >>     ; same as \"true\"\n> >> >>     core.fsync = \"*\"\n> >> >>\n> >> >> Or, to selectively sync some things and not others:\n> >> >>\n> >> >>     ;; Sync objects, but not \"info\"\n> >> >>     core.fsync = \":!objects/info/**\"\n> >> >>     core.fsync = \"objects/**\"\n> >> >>\n> >> >> See gitrepository-layout(5) for details about what sort of paths you\n> >> >> might be expected to match. Not all paths listed there will go through\n> >> >> this mechanism (e.g. currently objects do, but nothing to do with config\n> >> >> does).\n> >> >>\n> >> >> We can and will match this against \"fake paths\", e.g. when writing out\n> >> >> packs we may match against just the string \"objects/pack\", we're not\n> >> >> going to re-check if every packfile we're writing matches your globs,\n> >> >> ditto for loose objects. Be reasonable!\n> >> >>\n> >> >> This metharism is intended as a shorthand that provides some flexibility\n> >> >> when fsyncing, while not forcing git to come up with labels for all\n> >> >> paths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n> >> >>\n> >> >> More paths may be added or removed in the future, and we make no\n> >> >> promises that we won't move things around, so if in doubt use\n> >> >> e.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\n> >> >> stick to the golden path of examples provided in this documentation.\n> >> >>\n> >> >> == END DOC\n> >> >>\n> >> >>\n> >> >> It's a tad more complex to set up, but I wonder if that isn't worth\n> >> >> it. It nicely gets around any current and future issues of deciding what\n> >> >> labels such as \"loose-object\" etc. to pick, as well as slotting into an\n> >> >> existing method of doing exclude/include lists.\n> >> >>\n> >> >\n> >> > I think this proposal is a lot of complexity to avoid coming up with a\n> >> > new name for syncable things as they are added to Git.  A path based\n> >> > mechanism makes it hard to document for the (advanced) user what the\n> >> > full set of things is and how it might change from release to release.\n> >> > I think the current core.fsync scheme is a bit easier to understand,\n> >> > query, and extend.\n> >>\n> >> We document it in gitrepository-layout(5). Yeah it has some\n> >> disadvantages, but one advantage is that you could make the\n> >> composability easy. I.e. if last exclude wins then a setting of:\n> >>\n> >>     core.fsync = \":!*\"\n> >>     core.fsync = \"objects/**\"\n> >>\n> >> Would reset all previous matches & only match objects/**.\n> >>\n> >\n> > The value of changing this is predicated on taking your previous\n> > multi-valued config proposal, which I'm still not at all convinced\n> > about.\n>\n> They're orthagonal. I.e. you get benefits from multi-value with or\n> without this globbing mechanism.\n>\n> In any case, I don't feel strongly about/am really advocating this\n> globbing mechanism. I just wondered if it wouldn't make things simpler\n> since it would sidestep the need to create any sort of categories for\n> subsets of gitrepository-layout(5), but maybe not...\n>\n> > The schema in the current (v1-v2) version of the patch already\n> > includes an example of extending the list of syncable things, and\n> > Patrick Steinhardt made it clear that he feels comfortable adding\n> > 'refs' to the same schema in a future change.\n> >\n> > I'll also emphasize that we're talking about a non-functional,\n> > relatively corner-case behavioral configuration.  These values don't\n> > change how git's interface behaves except when the system crashes\n> > during a git command or shortly after one completes.\n>\n> That may be something some OS's promise, but it's not something fsync()\n> or POSIX promises. I.e. you might write a ref, but unless you fsync and\n> the relevant dir entries another process might not see the update, crash\n> or not.\n>\n\nI haven't seen any indication that POSIX requires an fsync for\nvisiblity within a running system.  I looked at the doc for open,\nwrite, and fsync, and saw no indication that it's posix compliant to\nrequire an fsync for visibility.  I think an OS that required fsync\nfor cross-process visiblity would fail to run Git for a myriad of\nother reasons and would likely lose all its users.  I'm curious where\nyou've seen documentation that allows such unhelpful behavior?\n\n> That's an aside from these config design questions, and I think\n> most/(all?) OS's anyone cares about these days tend to make that\n> implicit promise as part of their VFS behavior, but we should probably\n> design such an interface to fsync() with such pedantic portability in\n> mind.\n\nWhy? To be rude to such a hypothetical system, if a system were so\ninsanely designed, it would be nuts to support it.\n\n> > While you may not personally love the proposed configuration\n> > interface, I'd want your view on some questions:\n> > 1. Is it easy for the (advanced) user to set a configuration?\n> > 2. Is it easy for the (advanced) user to see what was configured?\n> > 3. Is it easy for the Git community to build on this as we want to add\n> > things to the list of things to sync?\n> >     a) Is there a good best practice configuration so that people can\n> > avoid losing integrity for new stuff that they are intending to sync.\n> >     b) If someone has a custom configuration, can that custom\n> > configuration do something reasonable as they upgrade versions of Git?\n> >              ** In response to this question, I might see some value\n> > in adding a 'derived-metadata' aggregate that can be disabled so that\n> > a custom configuration can exclude those as they change version to\n> > version.\n> >     c) Is it too much maintenance overhead to consider how to present\n> > this configuration knob for any new hashfile or other datafile in the\n> > git repo?\n> > 4. Is there a good path forward to change the default syncable set,\n> > both in git-for-windows and in Git for other platforms?\n>\n> I'm not really sure this globbing this is a good idea, as noted above\n> just a suggestion etc.\n>\n> As noted there it just gets you out of the business of re-defining\n> gitrepository-layout(5), and assuming too much in advance about certain\n> use-cases.\n>\n> E.g. even \"refs\" might be too broad for some. I don't tend to be I/O\n> limited, but I could see how someone who would be would care about\n> refs/heads but not refs/remotes, or want to exclude logs/* but not the\n> refs updates themselves etc.\n\nThis use-case is interesting (distinguishing remote refs from local\nrefs).  I think the difficulty of verifying (for even an advanced\nuser) that the right fsyncing is actually happening still puts me on\nthe side of having a carefully curated and documented set of syncable\nthings rather than a file-path-based mechanism.\n\nIs this meaningful in the presumably nearby future world of the refsdb\nbackend?  Is that somehow split by remote versus local?\n\n> >> >> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> >> >> > index 857be7826f3..916c55d6ce9 100644\n> >> >> > --- a/builtin/pack-objects.c\n> >> >> > +++ b/builtin/pack-objects.c\n> >> >> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n> >> >> >                * If so, rewrite it like in fast-import\n> >> >> >                */\n> >> >> >               if (pack_to_stdout) {\n> >> >> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> >> >> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n> >> >> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> >> >>\n> >> >> Not really related to this per-se, but since you're touching the API\n> >> >> everything goes through I wonder if callers should just always try to\n> >> >> fsync, and we can just catch EROFS and EINVAL in the wrapper if someone\n> >> >> tries to flush stdout, or catch the fd at that lower level.\n> >> >>\n> >> >> Or maybe there's a good reason for this...\n> >> >\n> >> > It's platform dependent, but I'd expect fsync would do something for\n> >> > pipes or stdout redirected to a file.  In these cases we really don't\n> >> > want to fsync since we have no idea what we're talking to and we're\n> >> > potentially worsening performance for probably no benefit.\n> >>\n> >> Yeah maybe we should just leave it be.\n> >>\n> >> I'd think the C library returning EINVAL would be a trivial performance\n> >> cost though.\n> >>\n> >> It just seemed odd to hardcode assumptions about what can and can't be\n> >> synced when the POSIX defined function will also tell us that.\n> >>\n> >\n> > Redirecting stdout to a file seems like a common usage for this\n> > command. That would definitely be fsyncable, but Git has no idea what\n> > its proper category is since there's no way to know the purpose or\n> > lifetime of the packfile.  I'm going to leave this be, because I'd\n> > posit that \"can it be fsynced?\" is not the same as \"should it be\n> > fsynced?\".  The latter question can't be answered for stdout.\n>\n> As noted this was just an aside, and I don't even know if any OS would\n> do anything meaningful with an fsync() of such a FD anyway.\n>\n\nThe underlying fsync primitive does have a meaning on Windows for\npipes, but it's certainly not what Git would want to do. Also if\nstdout is redirected to a file, I'm pretty sure that UNIX OSes would\nrespect the fsync call.  However it's not meaningful in the sense of\nthe git repository, since we don't know what the packfile is or why it\nwas created.\n\n> I just don't see why we wouldn't say:\n>\n>  1. We're syncing this category of thing\n>  2. Try it\n>  3. If fsync returns \"can't fsync that sort of thing\" move on\n>\n> As opposed to trying to shortcut #3 by doing the detection ourselves.\n>\n> I.e. maybe there was a good reason, but it seemed to be some easy\n> potential win for more simplification, since you were re-doing and\n> simplifying some of the interface anyway...\n\nWe're trying to be deliberate about what we're fsyncing.  Fsyncing an\nunknown file created by the packfile code doesn't move us in that\ndirection.  In your taxonomy we don't know (1), \"what is this category\nof thing?\"  Sure it's got the packfile format, but is not known to be\nan actual packfile that's part of the repository.\n\n> >>\n> >> >> > [...]\n> >> >> > +/*\n> >> >> > + * These values are used to help identify parts of a repository to fsync.\n> >> >> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n> >> >> > + * repository and so shouldn't be fsynced.\n> >> >> > + */\n> >> >> > +enum fsync_component {\n> >> >> > +     FSYNC_COMPONENT_NONE                    = 0,\n> >> >>\n> >> >> I haven't read ahead much but in most other such cases we don't define\n> >> >> the \"= 0\", just start at 1<<0, then check the flags elsewhere...\n> >> >>\n> >> >> > +static const struct fsync_component_entry {\n> >> >> > +     const char *name;\n> >> >> > +     enum fsync_component component_bits;\n> >> >> > +} fsync_component_table[] = {\n> >> >> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> >> >> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n> >> >> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> >> >> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> >> >> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> >> >> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n> >> >> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n> >> >> > +};\n> >> >> > +\n> >> >> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> >> >> > +{\n> >> >> > +     enum fsync_component output = 0;\n> >> >> > +\n> >> >> > +     if (!strcmp(string, \"none\"))\n> >> >> > +             return output;\n> >> >> > +\n> >> >> > +     while (string) {\n> >> >> > +             int i;\n> >> >> > +             size_t len;\n> >> >> > +             const char *ep;\n> >> >> > +             int negated = 0;\n> >> >> > +             int found = 0;\n> >> >> > +\n> >> >> > +             string = string + strspn(string, \", \\t\\n\\r\");\n> >> >>\n> >> >> Aside from the \"use a list\" isn't this hardcoding some windows-specific\n> >> >> assumptions with \\n\\r? Maybe not...\n> >> >\n> >> > I shamelessly stole this code from parse_whitespace_rule. I thought\n> >> > about making a helper to be called by both functions, but the amount\n> >> > of state going into and out of the wrapper via arguments was\n> >> > substantial and seemed to negate the benefit of deduplication.\n> >>\n> >> FWIW string_list_split() is easier to work with in those cases, or at\n> >> least I think so...\n> >\n> > This code runs at startup for a variable that may be present on some\n> > installations.  The nice property of the current patch's code is that\n> > it's already a well-tested pattern that doesn't do any allocations as\n> > it's working, unlike string_list_split().\n>\n> Multi-value config would also get you fewer allocations :)\n>\n> Anyway, I mainly meant to point out that for stuff like this it's fine\n> to optimize it for ease rather than micro-optimize allocations. Those\n> really aren't a bottleneck on this scale.\n>\n> Even in that case there's string_list_split_in_place(), which can be a\n> bit nicer than manual C-string fiddling.\n>\n\nAm I allowed to change the config value string in place? The\ncore.whitespace code is careful not to modify the string. I kind of\nlike the parse_ws_error_highlight code a little better now that I've\nseen it, but I think the current code is fine too.\n\n> > I hope you know that I appreciate your review feedback, even though\n> > I'm pushing back on most of it so far this round. I'll be sending v3\n> > to the list soon after giving it another look over.\n>\n> Sure, no worries. Just hoping to help. If you go for something different\n> etc. that's fine. Just hoping to bridge the gap in some knowledge /\n> offer potentially interesting suggestions (some of which may be dead\n> ends, like the config glob thing...).\n\nThanks,\nNeeraj\n"},{"id":"445772","messageId":"CANQDOddJzgyfQHi8hVMPU5iLwyt4GmpGt5qob0ZrjqTax6K0tw@mail.gmail.com","threadId":"57030","inReplyTo":"pull.1093.v3.git.1639011433.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 0/4] A design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-01-08T01:13:59Z","receivedAt":"2022-01-08T01:14:13Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"Hello Everyone,\nI wanted to revive this thread in the new year.\n\nTo summarize the current state of affairs:\n* The current fsync patch series implements two new configuration options:\n       core.fsync = <comma-separate list> -- select which repo\ncomponents will be fsynced\n       core.fsyncMethod = fsync|writeout-only  -- select what form of\nfsyncing will be done\n\n* This patch series now ignores core.fsyncObjectFiles with a\ndeprecation warning pointing the user at core.fsync.\n\n* There is a follow-on series that will extend the core.fsyncMethod to\nalso include a `batch` mode that speeds up bulk operations by avoiding\nrepeated disk cache flushes.\n\n* I developed the current mechanism after Ævar pointed out that the\noriginal `core.fsyncObjectFiles=batch` change would cause older\nversions of Git to die() when exposed to a new configuration. There\nwere also several fsync changes floating around, including Patrick\nSteinhardts `core.fsyncRefFiles` change [1] and Eric Wong's\n`core.fsync = false` change [2].\n\n* The biggest sticking points are in [3].  The fundamental\ndisagreement is about whether core.fsync should look like:\n      A) core.fsync = objects,commit-graph   [current patch implementation]\n      or\n      B) core.fsync = objects\n          core.fsync = commit-graph    [Ævar's multivalued proposal].\nI prefer sticking with (A) for reasons spelled out in the thread. I'm\nhappy to re-litigate this discussion though.\n\n* There's also a sticking point about whether we should fsync when\ninvoking pack-objects against stdout.  I think that mostly reflects a\nmissing comment in the code rather than a real disagreement.\n\n* Now that ew/test-wo-fsync has been integrated, there's some\nredundancy between core.fsync=none and Eric's patch.\n\nOpen questions:\n1) What format should we use for the core.fsync configuration to\nselect individual repo components to sync?\n2) Are we okay with deprecating core.fsyncObjectFiles in a single\nrelease with a warning?\n3) Is it reasonable to expect people adding new persistent files to\nadd and document new values of the core.fsync settings?\n\nThanks,\nNeeraj\n\n[1]  https://lore.kernel.org/git/20211030103950.M489266@dcvr/\n[2] https://lore.kernel.org/git/20211028002102.19384-1-e@80x24.org/\n[3] https://lore.kernel.org/git/211207.86wnkgo9fv.gmgdl@evledraar.gmail.com/\n"},{"id":"445797","messageId":"007001d804f3$a2e573a0$e8b05ae0$@nexbridge.com","threadId":"57030","inReplyTo":"CANQDOddJzgyfQHi8hVMPU5iLwyt4GmpGt5qob0ZrjqTax6K0tw@mail.gmail.com","subject":"RE: [PATCH v3 0/4] A design for future-proofing fsync() configuration","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-01-09T00:55:40Z","receivedAt":"2022-01-09T00:55:49Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On January 7, 2022 8:14 PM, Neeraj Singh wrote:\n> Hello Everyone,\n> I wanted to revive this thread in the new year.\n> \n> To summarize the current state of affairs:\n> * The current fsync patch series implements two new configuration options:\n>        core.fsync = <comma-separate list> -- select which repo components will\n> be fsynced\n>        core.fsyncMethod = fsync|writeout-only  -- select what form of fsyncing\n> will be done\n> \n> * This patch series now ignores core.fsyncObjectFiles with a deprecation\n> warning pointing the user at core.fsync.\n> \n> * There is a follow-on series that will extend the core.fsyncMethod to also\n> include a `batch` mode that speeds up bulk operations by avoiding repeated\n> disk cache flushes.\n> \n> * I developed the current mechanism after Ævar pointed out that the original\n> `core.fsyncObjectFiles=batch` change would cause older versions of Git to\n> die() when exposed to a new configuration. There were also several fsync\n> changes floating around, including Patrick Steinhardts `core.fsyncRefFiles`\n> change [1] and Eric Wong's `core.fsync = false` change [2].\n> \n> * The biggest sticking points are in [3].  The fundamental disagreement is\n> about whether core.fsync should look like:\n>       A) core.fsync = objects,commit-graph   [current patch implementation]\n>       or\n>       B) core.fsync = objects\n>           core.fsync = commit-graph    [Ævar's multivalued proposal].\n> I prefer sticking with (A) for reasons spelled out in the thread. I'm happy to re-\n> litigate this discussion though.\n> \n> * There's also a sticking point about whether we should fsync when invoking\n> pack-objects against stdout.  I think that mostly reflects a missing comment in\n> the code rather than a real disagreement.\n> \n> * Now that ew/test-wo-fsync has been integrated, there's some redundancy\n> between core.fsync=none and Eric's patch.\n> \n> Open questions:\n> 1) What format should we use for the core.fsync configuration to select\n> individual repo components to sync?\n> 2) Are we okay with deprecating core.fsyncObjectFiles in a single release with\n> a warning?\n> 3) Is it reasonable to expect people adding new persistent files to add and\n> document new values of the core.fsync settings?\n> \n> Thanks,\n> Neeraj\n> \n> [1]  https://lore.kernel.org/git/20211030103950.M489266@dcvr/\n> [2] https://lore.kernel.org/git/20211028002102.19384-1-e@80x24.org/\n> [3]\n> https://lore.kernel.org/git/211207.86wnkgo9fv.gmgdl@evledraar.gmail.com/\n\nNeeraj,\n\nPlease remember that fsync() is operating system and version specific. You cannot make any assumptions about what is supported and what is not. I have recently had issues with git built on a recent operating system not running on a version from 2020. The proposed patches do not work, as I recall, in a portable manner, so caution is required making this change. You can expect this not to work on some platforms and some versions. Please account for that. Requiring users who are not aware of OS details to configure git to function at all is a bad move, in my view - which has not changed since last time.\n\nThanks,\nRandall\n\n"},{"id":"445863","messageId":"CANQDOdf4k5K8U-hs+Cv5W9DZFDonab6eMO4XFX4UFHkJ2RzQQg@mail.gmail.com","threadId":"57030","inReplyTo":"007001d804f3$a2e573a0$e8b05ae0$@nexbridge.com","subject":"Re: [PATCH v3 0/4] A design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-01-10T19:00:15Z","receivedAt":"2022-01-10T19:00:30Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Sat, Jan 8, 2022 at 4:55 PM <rsbecker@nexbridge.com> wrote:\n>\n> Please remember that fsync() is operating system and version specific. You cannot make any assumptions about what is supported and what is not. I have recently had issues with git built on a recent operating system not running on a version from 2020. The proposed patches do not work, as I recall, in a portable manner, so caution is required making this change. You can expect this not to work on some platforms and some versions. Please account for that. Requiring users who are not aware of OS details to configure git to function at all is a bad move, in my view - which has not changed since last time.\n>\n\nThere was already an implied configuration of fsync in the Git\ncodebase.  None of the defaults are changing--assuming that a user\ndoes not explicitly configure the core.fsync setting, Git should work\nthe same as it always has.  I don't believe the current patch series\nintroduces any new incompatibilities.\n\nThanks,\nNeeraj\n"},{"id":"446464","messageId":"CANQDOdf97WK7kWZHM7xr_fMS1aep5hYYcwC2Pfr_7w4h_8ybfw@mail.gmail.com","threadId":"57030","inReplyTo":"CANQDOdchh3mfC8S6ouWAQbtWzZUkmTzF1p5D9dg4muoBu4N1Fg@mail.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-01-18T23:50:23Z","receivedAt":"2022-01-18T23:50:37Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"Hi Ævar,\nCould you please respond to the parent message?\nTo summarize the main points and questions:\n1) The current patch series design of core.fsync favors ease of\nmodification of the complete configuration as a single atomic config\nstep, at the cost of making it somewhat harder to derive a new\nconfiguration from an existing one. See [1] where this facility is\nused.\n\n2) Is there any existing configuration that uses the multi-value\nschema you proposed? The diff.wsErrorHighlight setting is actually\ncomma-separated.\n\n3) Is there any system you can point at or any support in the POSIX\nspec for requiring fsync for anything other than durability of data\nacross system crash?\n\n4) I think string_list_split_in_place is not valid for splitting a\ncomma-separated list in the config code, since the value coming from\nthe configset code is in some global data structure that might be used\nagain. It seems like we could have subtle problems down the line if we\nmutate it.\n\n[1] https://github.com/neerajsi-msft/git/commit/7e9749a7e94d26c88789459588997329c5130792#diff-ee0307657f5a76b723c8973db0dbd5a2ca62e14b02711b897418b35d78fc6023R1327\n\nThanks,\nNeeraj\n"},{"id":"446496","messageId":"220119.8635ljoidt.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"CANQDOdchh3mfC8S6ouWAQbtWzZUkmTzF1p5D9dg4muoBu4N1Fg@mail.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-01-19T14:52:15Z","receivedAt":"2022-01-19T15:27:48Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Dec 08 2021, Neeraj Singh wrote:\n\n[Sorry about the late reply, and thanks for the downthread prodding]\n\n> On Wed, Dec 8, 2021 at 8:46 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>\n>>\n>> On Wed, Dec 08 2021, Neeraj Singh wrote:\n>>\n>> > On Wed, Dec 8, 2021 at 2:17 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>> >>\n>> >>\n>> >> On Tue, Dec 07 2021, Neeraj Singh wrote:\n>> >>\n>> >> > On Tue, Dec 7, 2021 at 5:01 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>> >> >>\n>> >> >>\n>> >> >> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n>> >> >>\n>> >> >> > From: Neeraj Singh <neerajsi@microsoft.com>\n>> >> >> >\n>> >> >> > This commit introduces the `core.fsync` configuration\n>> >> >> > knob which can be used to control how components of the\n>> >> >> > repository are made durable on disk.\n>> >> >> >\n>> >> >> > This setting allows future extensibility of the list of\n>> >> >> > syncable components:\n>> >> >> > * We issue a warning rather than an error for unrecognized\n>> >> >> >   components, so new configs can be used with old Git versions.\n>> >> >>\n>> >> >> Looks good!\n>> >> >>\n>> >> >> > * We support negation, so users can choose one of the default\n>> >> >> >   aggregate options and then remove components that they don't\n>> >> >> >   want. The user would then harden any new components added in\n>> >> >> >   a Git version update.\n>> >> >>\n>> >> >> I think this config schema makes sense, but just a (I think important)\n>> >> >> comment on the \"how\" not \"what\" of it. It's really much better to define\n>> >> >> config as:\n>> >> >>\n>> >> >>     [some]\n>> >> >>     key = value\n>> >> >>     key = value2\n>> >> >>\n>> >> >> Than:\n>> >> >>\n>> >> >>     [some]\n>> >> >>     key = value,value2\n>> >> >>\n>> >> >> The reason is that \"git config\" has good support for working with\n>> >> >> multi-valued stuff, so you can do e.g.:\n>> >> >>\n>> >> >>     git config --get-all -z some.key\n>> >> >>\n>> >> >> And you can easily (re)set such config e.g. with --replace-all etc., but\n>> >> >> for comma-delimited you (and users) need to do all that work themselves.\n>> >> >>\n>> >> >> Similarly instead of:\n>> >> >>\n>> >> >>     some.key = want-this\n>> >> >>     some.key = -not-this\n>> >> >>     some.key = but-want-this\n>> >> >>\n>> >> >> I think it's better to just have two lists, one inclusive another\n>> >> >> exclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n>> >> >> \"transfer.hideRefs\"\n>> >> >>\n>> >> >> Which would mean:\n>> >> >>\n>> >> >>     core.fsync = want-this\n>> >> >>     core.fsyncExcludes = -not-this\n>> >> >>\n>> >> >> For some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\n>> >> >> suggestion on making this easier for users & the implementation.\n>> >> >>\n>> >> >\n>> >> > Maybe there's some way to handle this I'm unaware of, but a\n>> >> > disadvantage of your multi-valued config proposal is that it's harder,\n>> >> > for example, for a per-repo config store to reasonably override a\n>> >> > per-user config store.  With the configuration scheme as-is, I can\n>> >> > have a per-user setting like `core.fsync=all` which covers my typical\n>> >> > repos, but then have a maintainer repo with a private setting of\n>> >> > `core.fsync=none` to speed up cases where I'm mostly working with\n>> >> > other people's changes that are backed up in email or server-side\n>> >> > repos.  The latter setting conveniently overrides the former setting\n>> >> > in all aspects.\n>> >>\n>> >> Even if you turn just your comma-delimited proposal into a list proposal\n>> >> can't we just say that the last one wins? Then it won't matter what cmae\n>> >> before, you'd specify \"core.fsync=none\" in your local .git/config.\n>> >>\n>> >> But this is also a general issue with a bunch of things in git's config\n>> >> space. I'd rather see us use the list-based values and just come up with\n>> >> some general way to reset them that works for all keys, rather than\n>> >> regretting having comma-delimited values that'll be harder to work with\n>> >> & parse, which will be a legacy wart if/when we come up with a way to\n>> >> say \"reset all previous settings\".\n>> >>\n>> >> > Also, with the core.fsync and core.fsyncExcludes, how would you spell\n>> >> > \"don't sync anything\"? Would you still have the aggregate options.?\n>> >>\n>> >>     core.fsyncExcludes = *\n>> >>\n>> >> I.e. the same as the core.fsync=none above, anyway I still like the\n>> >> wildcard thing below a bit more...\n>> >\n>> > I'm not going to take this feedback unless there are additional votes\n>> > from the Git community in this direction.  I make the claim that\n>> > single-valued comma-separated config lists are easier to work with in\n>> > the existing Git infrastructure.\n>>\n>> Easier in what sense? I showed examples of how \"git-config\" trivially\n>> works with multi-valued config, but for comma-delimited you'll need to\n>> write your own shellscript boilerplate around simple things like adding\n>> values, removing existing ones etc.\n>>\n>> I.e. you can use --add, --unset, --unset-all, --get, --get-all etc.\n>>\n>\n> I see what you're saying for cases where someone would want to set a\n> core.fsync setting that's derived from the user's current config.  But\n> I'm guessing that the dominant use case is someone setting a new fsync\n> configuration that achieves some atomic goal with respect to a given\n> repository. Like \"this is a throwaway, so sync nothing\" or \"this is\n> really important, so sync all objects and refs and the index\".\n\nWhether it's multi-value or comma-separated you could do:\n\n    -c core.fsync=[none|false]\n\nTo sync nothing, i.e. if we see \"none/false\" it doesn't matter if we saw\ncore.fsync=loose-object or whatever before, it would override it, ditto\nfor \"all\".\n\n>> > parsing code for the core.whitespace variable and users are used to\n>> > this syntax there. There are several other comma-separated lists in\n>> > the config space, so this construct has precedence and will be with\n>> > Git for some time.\n>>\n>> That's not really an argument either way for why we'd pick X over Y for\n>> something new. We've got some comma-delimited, some multi-value (I'm\n>> fairly sure we have more multi-value).\n>>\n>\n> My main point here is that there's precedent for patch's current exact\n> schema for a config in the same core config leaf of the Documentation.\n> It seems from your comments that we'd have to invent and document some\n> new convention for \"reset\" of a multi-valued config.  So you're\n> suggesting that I solve an extra set of problems to get this change\n> in.  Just want to remind you that my personal itch to scratch is to\n> get the underlying mechanism in so that git-for-windows can set its\n> default setting to batch mode. I'm not expecting many users to\n> actually configure this setting to any non-default value.\n\nMe neither. I think people will most likely set this once in\n~/.gitconfig or /etc/gitconfig.\n\nWe have some config keys that are multi-value and either comma-separated\nor space-separated, e.g. core.alternateRefsPrefixes\n\nThen we have e.g. blame.ignoreRevsFile which is multi-value, and further\nhas the convention that setting it to an empty value clears the\nlist. which would scratch the \"override existing\" itch.\n\nformat.notes, versionsort.suffix, transfer.hideRefs, branch.<name>.merge\nare exmples of existing multi-value config.\n\n>> > Also, fsync configurations aren't composable like\n>> > some other configurations may be. It makes sense to have a holistic\n>> > singular fsync configuration, which is best represented by a single\n>> > variable.\n>>\n>> What's a \"variable\" here? We call these \"keys\", you can have a\n>> single-value key like user.name that you get with --get, or a\n>> multi-value key like say branch.<name>.merge or push.pushOption that\n>> you'd get with --get-all.\n>\n> Yeah, I meant \"key\".  I conflated the config key with the underlying\n> global variable in git.\n\n*nod*\n\n>> I think you may be referring to either not wanting these to be\n>> \"inherited\" (which is not really a think we do for anything else in\n>> config), or lacking the ability to \"reset\".\n>>\n>> For the latter if that's absolutely needed we could just use the same\n>> trick as \"diff.wsErrorHighlight\" uses of making an empty value \"reset\"\n>> the list, and you'd get better \"git config\" support for editing it.\n>>\n>\n> My reading of the code is that diff.wsErrorHighlight is a comma\n> separated list and not a multi-valued config.  Actually I haven't yet\n> found an existing multi-valued config (not sure how to grep for it).\n\nYes, I think I conflated it with one of the ones above when I wrote\nthis.\n\n>> >> >> > This also supports the common request of doing absolutely no\n>> >> >> > fysncing with the `core.fsync=none` value, which is expected\n>> >> >> > to make the test suite faster.\n>> >> >>\n>> >> >> Let's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\n>> >> >> so we'll accept \"false\", \"off\", \"no\" like most other such config?\n>> >> >\n>> >> > Junio's previous feedback when discussing batch mode [1] was to offer\n>> >> > less flexibility when parsing new values of these configuration\n>> >> > options. I agree with his statement that \"making machine-readable\n>> >> > tokens be spelled in different ways is a 'disease'.\"  I'd like to\n>> >> > leave this as-is so that the documentation can clearly state the exact\n>> >> > set of allowable values.\n>> >> >\n>> >> > [1] https://lore.kernel.org/git/xmqqr1dqzyl7.fsf@gitster.g/\n>> >>\n>> >> I think he's talking about batch, Batch, BATCH, bAtCh etc. there. But\n>> >> the \"maybe bool\" is a stanard pattern we use.\n>> >>\n>> >> I don't think we'd call one of these 0, off, no or false etc. to avoid\n>> >> confusion, so then you can use git_parse_maybe_...()\n>> >\n>> > I don't see the advantage of having multiple ways of specifying\n>> > \"none\".  The user can read the doc and know exactly what to write.  If\n>> > they write something unallowable, they get a clear warning and they\n>> > can read the doc again to figure out what to write.  This isn't a\n>> > boolean options at all, so why should we entertain bool-like ways of\n>> > spelling it?\n>>\n>> It's not boolean, it's multi-value and one of the values includes a true\n>> or false boolean value. You just spell it \"none\".\n>>\n>> I think both this and your comment above suggest that you think there's\n>> no point in this because you haven't interacted with/used \"git config\"\n>> as a command line or API mechanism, but have just hand-crafted config\n>> files.\n>>\n>> That's fair enough, but there's a lot of tooling that benefits from the\n>> latter.\n>\n> My batch mode perf tests (on github, not yet submitted to the list)\n> use `git -c core.fsync=<foo>` to set up a per-process config.  I\n> haven't used the `git config` writing support in a while, so I haven't\n> deeply thought about that.  However, I favor simplifying the use case\n> of \"atomically\" setting a new holistic core.fsync value versus the use\n> case of deriving a new core.fsync value from the preexisting value.\n\nIf you implement it like blame.ignoreRevsFile you can have your cake and\neat it too, i.e.:\n\n    -c core.fsync= core.fsync=loose-object\n\nensures only loose objects are synced, as with your single-value config,\nbut I'd think what you'd be more likely to actually mean would be:\n\n    # With \"core.fsync=pack\" set in ~/.gitconfig\n    -c core.fsync=loose-object\n\nI.e. that the common case is \"I want this to be synced here\", not that\nyou'd like to declare sync policy from scratch every time.\n\nIn any case, on this general topic my main point is that the\ngit-config(1) command has pretty integration for multi-value if you do\nit that way, and not for comma-delimited. I.e. you get --add, --unset,\n--unset-all, --get, --get-all etc.\n\nSo I think for anything new it makes sense to lean into that, I think\nmost of these existing comma-delimited ones are ones we'd do differently\ntoday on reflection.\n\nAnd if you suppor the \"empty resets\" like blame.ignoreRevsFile it seems\nto me you'll have your cake & eat it too.\n\n>> E.g.:\n>>\n>>     $ git -c core.fsync=off config --type=bool core.fsync\n>>     false\n>>     $ git -c core.fsync=blah config --type=bool core.fsync\n>>     fatal: bad boolean config value 'blah' for 'core.fsync'\n>>\n>> Here we can get 'git config' to normalize what you call 'none', and you\n>> can tell via exit codes/normalization if it's \"false\". But if you invent\n>> a new term for \"false\" you can't do that as easily.\n>>\n>> We have various historical keys that take odd exceptions to that,\n>> e.g. \"never\", but unless we have a good reason to let's not invent more\n>> exceptions.\n>>\n>> >> >> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>> >> >> > ---\n>> >> >> >  Documentation/config/core.txt | 27 +++++++++----\n>> >> >> >  builtin/fast-import.c         |  2 +-\n>> >> >> >  builtin/index-pack.c          |  4 +-\n>> >> >> >  builtin/pack-objects.c        |  8 ++--\n>> >> >> >  bulk-checkin.c                |  5 ++-\n>> >> >> >  cache.h                       | 39 +++++++++++++++++-\n>> >> >> >  commit-graph.c                |  3 +-\n>> >> >> >  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n>> >> >> >  csum-file.c                   |  5 ++-\n>> >> >> >  csum-file.h                   |  3 +-\n>> >> >> >  environment.c                 |  1 +\n>> >> >> >  midx.c                        |  3 +-\n>> >> >> >  object-file.c                 |  3 +-\n>> >> >> >  pack-bitmap-write.c           |  3 +-\n>> >> >> >  pack-write.c                  | 13 +++---\n>> >> >> >  read-cache.c                  |  2 +-\n>> >> >> >  16 files changed, 164 insertions(+), 33 deletions(-)\n>> >> >> >\n>> >> >> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n>> >> >> > index dbb134f7136..4f1747ec871 100644\n>> >> >> > --- a/Documentation/config/core.txt\n>> >> >> > +++ b/Documentation/config/core.txt\n>> >> >> > @@ -547,6 +547,25 @@ core.whitespace::\n>> >> >> >    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n>> >> >> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n>> >> >> >\n>> >> >> > +core.fsync::\n>> >> >> > +     A comma-separated list of parts of the repository which should be\n>> >> >> > +     hardened via the core.fsyncMethod when created or modified. You can\n>> >> >> > +     disable hardening of any component by prefixing it with a '-'. Later\n>> >> >> > +     items take precedence over earlier ones in the list. For example,\n>> >> >> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n>> >> >> > +     metadata.\" Items that are not hardened may be lost in the event of an\n>> >> >> > +     unclean system shutdown.\n>> >> >> > ++\n>> >> >> > +* `none` disables fsync completely. This must be specified alone.\n>> >> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n>> >> >> > +* `pack` hardens objects added to the repo in packfile form.\n>> >> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n>> >> >> > +* `commit-graph` hardens the commit graph file.\n>> >> >> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n>> >> >> > +  `pack-metadata`, and `commit-graph`.\n>> >> >> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n>> >> >> > +* `all` is an aggregate option that syncs all individual components above.\n>> >> >> > +\n>> >> >>\n>> >> >> It's probably a *bit* more work to set up, but I wonder if this wouldn't\n>> >> >> be simpler if we just said (and this is partially going against what I\n>> >> >> noted above):\n>> >> >>\n>> >> >> == BEGIN DOC\n>> >> >>\n>> >> >> core.fsync is a multi-value config variable where each item is a\n>> >> >> pathspec that'll get matched the same way as 'git-ls-files' et al.\n>> >> >>\n>> >> >> When we sync pretend that a path like .git/objects/de/adbeef... is\n>> >> >> relative to the top-level of the git\n>> >> >> directory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n>> >> >>\n>> >> >> You can then supply a list of wildcards and exclusions to configure\n>> >> >> syncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\n>> >> >> with:\n>> >> >>\n>> >> >>     ; same as \"false\"\n>> >> >>     core.fsync = \":!*\"\n>> >> >>\n>> >> >> Or:\n>> >> >>\n>> >> >>     ; same as \"true\"\n>> >> >>     core.fsync = \"*\"\n>> >> >>\n>> >> >> Or, to selectively sync some things and not others:\n>> >> >>\n>> >> >>     ;; Sync objects, but not \"info\"\n>> >> >>     core.fsync = \":!objects/info/**\"\n>> >> >>     core.fsync = \"objects/**\"\n>> >> >>\n>> >> >> See gitrepository-layout(5) for details about what sort of paths you\n>> >> >> might be expected to match. Not all paths listed there will go through\n>> >> >> this mechanism (e.g. currently objects do, but nothing to do with config\n>> >> >> does).\n>> >> >>\n>> >> >> We can and will match this against \"fake paths\", e.g. when writing out\n>> >> >> packs we may match against just the string \"objects/pack\", we're not\n>> >> >> going to re-check if every packfile we're writing matches your globs,\n>> >> >> ditto for loose objects. Be reasonable!\n>> >> >>\n>> >> >> This metharism is intended as a shorthand that provides some flexibility\n>> >> >> when fsyncing, while not forcing git to come up with labels for all\n>> >> >> paths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n>> >> >>\n>> >> >> More paths may be added or removed in the future, and we make no\n>> >> >> promises that we won't move things around, so if in doubt use\n>> >> >> e.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\n>> >> >> stick to the golden path of examples provided in this documentation.\n>> >> >>\n>> >> >> == END DOC\n>> >> >>\n>> >> >>\n>> >> >> It's a tad more complex to set up, but I wonder if that isn't worth\n>> >> >> it. It nicely gets around any current and future issues of deciding what\n>> >> >> labels such as \"loose-object\" etc. to pick, as well as slotting into an\n>> >> >> existing method of doing exclude/include lists.\n>> >> >>\n>> >> >\n>> >> > I think this proposal is a lot of complexity to avoid coming up with a\n>> >> > new name for syncable things as they are added to Git.  A path based\n>> >> > mechanism makes it hard to document for the (advanced) user what the\n>> >> > full set of things is and how it might change from release to release.\n>> >> > I think the current core.fsync scheme is a bit easier to understand,\n>> >> > query, and extend.\n>> >>\n>> >> We document it in gitrepository-layout(5). Yeah it has some\n>> >> disadvantages, but one advantage is that you could make the\n>> >> composability easy. I.e. if last exclude wins then a setting of:\n>> >>\n>> >>     core.fsync = \":!*\"\n>> >>     core.fsync = \"objects/**\"\n>> >>\n>> >> Would reset all previous matches & only match objects/**.\n>> >>\n>> >\n>> > The value of changing this is predicated on taking your previous\n>> > multi-valued config proposal, which I'm still not at all convinced\n>> > about.\n>>\n>> They're orthagonal. I.e. you get benefits from multi-value with or\n>> without this globbing mechanism.\n>>\n>> In any case, I don't feel strongly about/am really advocating this\n>> globbing mechanism. I just wondered if it wouldn't make things simpler\n>> since it would sidestep the need to create any sort of categories for\n>> subsets of gitrepository-layout(5), but maybe not...\n>>\n>> > The schema in the current (v1-v2) version of the patch already\n>> > includes an example of extending the list of syncable things, and\n>> > Patrick Steinhardt made it clear that he feels comfortable adding\n>> > 'refs' to the same schema in a future change.\n>> >\n>> > I'll also emphasize that we're talking about a non-functional,\n>> > relatively corner-case behavioral configuration.  These values don't\n>> > change how git's interface behaves except when the system crashes\n>> > during a git command or shortly after one completes.\n>>\n>> That may be something some OS's promise, but it's not something fsync()\n>> or POSIX promises. I.e. you might write a ref, but unless you fsync and\n>> the relevant dir entries another process might not see the update, crash\n>> or not.\n>>\n>\n> I haven't seen any indication that POSIX requires an fsync for\n> visiblity within a running system.  I looked at the doc for open,\n> write, and fsync, and saw no indication that it's posix compliant to\n> require an fsync for visibility.  I think an OS that required fsync\n> for cross-process visiblity would fail to run Git for a myriad of\n> other reasons and would likely lose all its users.  I'm curious where\n> you've seen documentation that allows such unhelpful behavior?\n\nThere's multiple unrelated and related things in this area. One is a\ncase where you'll e.g. write a file \"foo\" using stdio, spawn a program\nto work on it in the same program, but it might not see it at all, or\nsee empty content, the latter being because you haven't flushed your I/O\nbuffers (which you can do via fsync()).\n\nThe former is that on *nix systems you're generally only guaranteed to\nwrite to a fd, but not to have the associated metadata be synced for\nyou.\n\nThat is spelled out e.g. in the DESCRIPTION section of linux's fsync()\nmanpage: https://man7.org/linux/man-pages/man2/fdatasync.2.html\n\nI don't know how much you follow non-Windows FS development, but there\nwas also a very well known \"incident\" early in ext4 where it leaned into\nsome permissive-by-POSIX behavior that caused data loss in practice on\nbuggy programs that didn't correctly use fsync() , since various tooling\nhad come to expect the stricter behavior of ext3:\nhttps://lwn.net/Articles/328363/\n\nThat was explicitly in the area of fs metadata being discussed here.\n\nGenerally you can expect your VFS layer to be forgiving when it comes to\nIPC, but even that is out the window when it comes to networked\nfilesystems, e.g. a shared git repository hosted on NFS.\n\n>> That's an aside from these config design questions, and I think\n>> most/(all?) OS's anyone cares about these days tend to make that\n>> implicit promise as part of their VFS behavior, but we should probably\n>> design such an interface to fsync() with such pedantic portability in\n>> mind.\n>\n> Why? To be rude to such a hypothetical system, if a system were so\n> insanely designed, it would be nuts to support it.\n\nBecause we know that right now the system calls we're invoking aren't\nguaranteed to store data persistently to disk portably, although they do\nso in practice on most modern OS's.\n\nWe're portably to a lot of platforms, and also need to keep e.g. NFS in\nmind, so being able to ask for a pedantic mode when you care about data\nretention at the cost of performance would be nice.\n\nAnd because the fsync config mode you're proposing is thoroughly\nnon-standard, but is known to me much faster by leaning into known\nattributes of specific FS's on specific OS's, if we're not running on\nthose it would be sensible to fall back to a stricter mode of\noperation. E.g. syncing all 100 loose objects we just wrote, not just\nthe last one.\n\n>> > While you may not personally love the proposed configuration\n>> > interface, I'd want your view on some questions:\n>> > 1. Is it easy for the (advanced) user to set a configuration?\n>> > 2. Is it easy for the (advanced) user to see what was configured?\n>> > 3. Is it easy for the Git community to build on this as we want to add\n>> > things to the list of things to sync?\n>> >     a) Is there a good best practice configuration so that people can\n>> > avoid losing integrity for new stuff that they are intending to sync.\n>> >     b) If someone has a custom configuration, can that custom\n>> > configuration do something reasonable as they upgrade versions of Git?\n>> >              ** In response to this question, I might see some value\n>> > in adding a 'derived-metadata' aggregate that can be disabled so that\n>> > a custom configuration can exclude those as they change version to\n>> > version.\n>> >     c) Is it too much maintenance overhead to consider how to present\n>> > this configuration knob for any new hashfile or other datafile in the\n>> > git repo?\n>> > 4. Is there a good path forward to change the default syncable set,\n>> > both in git-for-windows and in Git for other platforms?\n>>\n>> I'm not really sure this globbing this is a good idea, as noted above\n>> just a suggestion etc.\n>>\n>> As noted there it just gets you out of the business of re-defining\n>> gitrepository-layout(5), and assuming too much in advance about certain\n>> use-cases.\n>>\n>> E.g. even \"refs\" might be too broad for some. I don't tend to be I/O\n>> limited, but I could see how someone who would be would care about\n>> refs/heads but not refs/remotes, or want to exclude logs/* but not the\n>> refs updates themselves etc.\n>\n> This use-case is interesting (distinguishing remote refs from local\n> refs).  I think the difficulty of verifying (for even an advanced\n> user) that the right fsyncing is actually happening still puts me on\n> the side of having a carefully curated and documented set of syncable\n> things rather than a file-path-based mechanism.\n>\n> Is this meaningful in the presumably nearby future world of the refsdb\n> backend?  Is that somehow split by remote versus local?\n\nThere is the upcoming \"reftable\" work, but that's probably 2-3 years out\nat the earliest for series production workloads in git.git.\n\n>> >> >> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n>> >> >> > index 857be7826f3..916c55d6ce9 100644\n>> >> >> > --- a/builtin/pack-objects.c\n>> >> >> > +++ b/builtin/pack-objects.c\n>> >> >> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n>> >> >> >                * If so, rewrite it like in fast-import\n>> >> >> >                */\n>> >> >> >               if (pack_to_stdout) {\n>> >> >> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>> >> >> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n>> >> >> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n>> >> >>\n>> >> >> Not really related to this per-se, but since you're touching the API\n>> >> >> everything goes through I wonder if callers should just always try to\n>> >> >> fsync, and we can just catch EROFS and EINVAL in the wrapper if someone\n>> >> >> tries to flush stdout, or catch the fd at that lower level.\n>> >> >>\n>> >> >> Or maybe there's a good reason for this...\n>> >> >\n>> >> > It's platform dependent, but I'd expect fsync would do something for\n>> >> > pipes or stdout redirected to a file.  In these cases we really don't\n>> >> > want to fsync since we have no idea what we're talking to and we're\n>> >> > potentially worsening performance for probably no benefit.\n>> >>\n>> >> Yeah maybe we should just leave it be.\n>> >>\n>> >> I'd think the C library returning EINVAL would be a trivial performance\n>> >> cost though.\n>> >>\n>> >> It just seemed odd to hardcode assumptions about what can and can't be\n>> >> synced when the POSIX defined function will also tell us that.\n>> >>\n>> >\n>> > Redirecting stdout to a file seems like a common usage for this\n>> > command. That would definitely be fsyncable, but Git has no idea what\n>> > its proper category is since there's no way to know the purpose or\n>> > lifetime of the packfile.  I'm going to leave this be, because I'd\n>> > posit that \"can it be fsynced?\" is not the same as \"should it be\n>> > fsynced?\".  The latter question can't be answered for stdout.\n>>\n>> As noted this was just an aside, and I don't even know if any OS would\n>> do anything meaningful with an fsync() of such a FD anyway.\n>>\n>\n> The underlying fsync primitive does have a meaning on Windows for\n> pipes, but it's certainly not what Git would want to do. Also if\n> stdout is redirected to a file, I'm pretty sure that UNIX OSes would\n> respect the fsync call.  However it's not meaningful in the sense of\n> the git repository, since we don't know what the packfile is or why it\n> was created.\n\nI suggested that because I think it's probably nonsensical, but it's\nnonsense that POSIX seems to explicitly tell us that it'll handle\n(probably by silently doing nothing). So in terms of our interface we\ncould lean into that and avoid our own special-casing.\n\n>> I just don't see why we wouldn't say:\n>>\n>>  1. We're syncing this category of thing\n>>  2. Try it\n>>  3. If fsync returns \"can't fsync that sort of thing\" move on\n>>\n>> As opposed to trying to shortcut #3 by doing the detection ourselves.\n>>\n>> I.e. maybe there was a good reason, but it seemed to be some easy\n>> potential win for more simplification, since you were re-doing and\n>> simplifying some of the interface anyway...\n>\n> We're trying to be deliberate about what we're fsyncing.  Fsyncing an\n> unknown file created by the packfile code doesn't move us in that\n> direction.  In your taxonomy we don't know (1), \"what is this category\n> of thing?\"  Sure it's got the packfile format, but is not known to be\n> an actual packfile that's part of the repository.\n\nWe know it's a fd, isn't that sufficient? In any case, I'm fine with\nalso keeping it as is, I don't mean to split hairs here.\n\nIt just stuck out as an odd part of the interface, why treat some fd's\nspecially, instead of just throwing it all at the OS. Presumably the\nfirst thing the OS will do is figure out if it's a syncable fd or not,\nand act appropriately.\n\n>> >>\n>> >> >> > [...]\n>> >> >> > +/*\n>> >> >> > + * These values are used to help identify parts of a repository to fsync.\n>> >> >> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n>> >> >> > + * repository and so shouldn't be fsynced.\n>> >> >> > + */\n>> >> >> > +enum fsync_component {\n>> >> >> > +     FSYNC_COMPONENT_NONE                    = 0,\n>> >> >>\n>> >> >> I haven't read ahead much but in most other such cases we don't define\n>> >> >> the \"= 0\", just start at 1<<0, then check the flags elsewhere...\n>> >> >>\n>> >> >> > +static const struct fsync_component_entry {\n>> >> >> > +     const char *name;\n>> >> >> > +     enum fsync_component component_bits;\n>> >> >> > +} fsync_component_table[] = {\n>> >> >> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n>> >> >> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n>> >> >> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n>> >> >> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n>> >> >> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n>> >> >> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n>> >> >> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n>> >> >> > +};\n>> >> >> > +\n>> >> >> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n>> >> >> > +{\n>> >> >> > +     enum fsync_component output = 0;\n>> >> >> > +\n>> >> >> > +     if (!strcmp(string, \"none\"))\n>> >> >> > +             return output;\n>> >> >> > +\n>> >> >> > +     while (string) {\n>> >> >> > +             int i;\n>> >> >> > +             size_t len;\n>> >> >> > +             const char *ep;\n>> >> >> > +             int negated = 0;\n>> >> >> > +             int found = 0;\n>> >> >> > +\n>> >> >> > +             string = string + strspn(string, \", \\t\\n\\r\");\n>> >> >>\n>> >> >> Aside from the \"use a list\" isn't this hardcoding some windows-specific\n>> >> >> assumptions with \\n\\r? Maybe not...\n>> >> >\n>> >> > I shamelessly stole this code from parse_whitespace_rule. I thought\n>> >> > about making a helper to be called by both functions, but the amount\n>> >> > of state going into and out of the wrapper via arguments was\n>> >> > substantial and seemed to negate the benefit of deduplication.\n>> >>\n>> >> FWIW string_list_split() is easier to work with in those cases, or at\n>> >> least I think so...\n>> >\n>> > This code runs at startup for a variable that may be present on some\n>> > installations.  The nice property of the current patch's code is that\n>> > it's already a well-tested pattern that doesn't do any allocations as\n>> > it's working, unlike string_list_split().\n>>\n>> Multi-value config would also get you fewer allocations :)\n>>\n>> Anyway, I mainly meant to point out that for stuff like this it's fine\n>> to optimize it for ease rather than micro-optimize allocations. Those\n>> really aren't a bottleneck on this scale.\n>>\n>> Even in that case there's string_list_split_in_place(), which can be a\n>> bit nicer than manual C-string fiddling.\n>>\n>\n> Am I allowed to change the config value string in place? The\n> core.whitespace code is careful not to modify the string. I kind of\n> like the parse_ws_error_highlight code a little better now that I've\n> seen it, but I think the current code is fine too.\n\nI don't remember offhand if that's safe, probably not. So you'll need a\ncopy here.\n\n>> > I hope you know that I appreciate your review feedback, even though\n>> > I'm pushing back on most of it so far this round. I'll be sending v3\n>> > to the list soon after giving it another look over.\n>>\n>> Sure, no worries. Just hoping to help. If you go for something different\n>> etc. that's fine. Just hoping to bridge the gap in some knowledge /\n>> offer potentially interesting suggestions (some of which may be dead\n>> ends, like the config glob thing...).\n"},{"id":"446498","messageId":"220119.86y23bn3no.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"CANQDOdf97WK7kWZHM7xr_fMS1aep5hYYcwC2Pfr_7w4h_8ybfw@mail.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-01-19T15:28:04Z","receivedAt":"2022-01-19T15:31:30Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jan 18 2022, Neeraj Singh wrote:\n\n> Hi Ævar,\n> Could you please respond to the parent message?\n\nI did just now in\nhttps://lore.kernel.org/git/220119.8635ljoidt.gmgdl@evledraar.gmail.com/;\nsorry about the delay. This thread fell off my radar.\n\n> To summarize the main points and questions:\n> 1) The current patch series design of core.fsync favors ease of\n> modification of the complete configuration as a single atomic config\n> step, at the cost of making it somewhat harder to derive a new\n> configuration from an existing one. See [1] where this facility is\n> used.\n>\n> 2) Is there any existing configuration that uses the multi-value\n> schema you proposed? The diff.wsErrorHighlight setting is actually\n> comma-separated.\n\nI replied to that. To add a bit to that the comma-delimited thing isn't\nany sort of a \"blocker\" or whatever in my mind.\n\nI just wanted to point out that you could get the same with multi-value\nwith some better integration, and we might have our cake & eat it too.\n\nBut at the end of the day if you disagree you're doing the work, and\nthat's ultimately bikeshedding of the config interface. I'm fine with it\neither way.\n\n> 3) Is there any system you can point at or any support in the POSIX\n> spec for requiring fsync for anything other than durability of data\n> across system crash?\n\nSome examples in the linked...\n\n> 4) I think string_list_split_in_place is not valid for splitting a\n> comma-separated list in the config code, since the value coming from\n> the configset code is in some global data structure that might be used\n> again. It seems like we could have subtle problems down the line if we\n> mutate it.\n\nReplied to that too, hope it's all useful. Thanks again for working on\nthis, it's really nice to have someone move the state of data integrity\nin git forward.\n\n> [1] https://github.com/neerajsi-msft/git/commit/7e9749a7e94d26c88789459588997329c5130792#diff-ee0307657f5a76b723c8973db0dbd5a2ca62e14b02711b897418b35d78fc6023R1327\n"},{"id":"447164","messageId":"CANQDOdep6VAGbK-Wh+MFOCHV4pgxgCU55=KV5Jux26=OZKfWfw@mail.gmail.com","threadId":"57030","inReplyTo":"220119.8635ljoidt.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/3] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-01-28T01:28:31Z","receivedAt":"2022-01-28T01:28:47Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Jan 19, 2022 at 10:27 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Wed, Dec 08 2021, Neeraj Singh wrote:\n>\n> [Sorry about the late reply, and thanks for the downthread prodding]\n\nI also apologize for the late reply here.  I've been dealing with a\nhospitalized parent this week.\n\n>\n> > On Wed, Dec 8, 2021 at 8:46 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Wed, Dec 08 2021, Neeraj Singh wrote:\n> >>\n> >> > On Wed, Dec 8, 2021 at 2:17 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> >> >>\n> >> >>\n> >> >> On Tue, Dec 07 2021, Neeraj Singh wrote:\n> >> >>\n> >> >> > On Tue, Dec 7, 2021 at 5:01 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> >> >> >>\n> >> >> >>\n> >> >> >> On Tue, Dec 07 2021, Neeraj Singh via GitGitGadget wrote:\n> >> >> >>\n> >> >> >> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >> >> >> >\n> >> >> >> > This commit introduces the `core.fsync` configuration\n> >> >> >> > knob which can be used to control how components of the\n> >> >> >> > repository are made durable on disk.\n> >> >> >> >\n> >> >> >> > This setting allows future extensibility of the list of\n> >> >> >> > syncable components:\n> >> >> >> > * We issue a warning rather than an error for unrecognized\n> >> >> >> >   components, so new configs can be used with old Git versions.\n> >> >> >>\n> >> >> >> Looks good!\n> >> >> >>\n> >> >> >> > * We support negation, so users can choose one of the default\n> >> >> >> >   aggregate options and then remove components that they don't\n> >> >> >> >   want. The user would then harden any new components added in\n> >> >> >> >   a Git version update.\n> >> >> >>\n> >> >> >> I think this config schema makes sense, but just a (I think important)\n> >> >> >> comment on the \"how\" not \"what\" of it. It's really much better to define\n> >> >> >> config as:\n> >> >> >>\n> >> >> >>     [some]\n> >> >> >>     key = value\n> >> >> >>     key = value2\n> >> >> >>\n> >> >> >> Than:\n> >> >> >>\n> >> >> >>     [some]\n> >> >> >>     key = value,value2\n> >> >> >>\n> >> >> >> The reason is that \"git config\" has good support for working with\n> >> >> >> multi-valued stuff, so you can do e.g.:\n> >> >> >>\n> >> >> >>     git config --get-all -z some.key\n> >> >> >>\n> >> >> >> And you can easily (re)set such config e.g. with --replace-all etc., but\n> >> >> >> for comma-delimited you (and users) need to do all that work themselves.\n> >> >> >>\n> >> >> >> Similarly instead of:\n> >> >> >>\n> >> >> >>     some.key = want-this\n> >> >> >>     some.key = -not-this\n> >> >> >>     some.key = but-want-this\n> >> >> >>\n> >> >> >> I think it's better to just have two lists, one inclusive another\n> >> >> >> exclusive. E.g. see \"log.decorate\" and \"log.excludeDecoration\",\n> >> >> >> \"transfer.hideRefs\"\n> >> >> >>\n> >> >> >> Which would mean:\n> >> >> >>\n> >> >> >>     core.fsync = want-this\n> >> >> >>     core.fsyncExcludes = -not-this\n> >> >> >>\n> >> >> >> For some value of \"fsyncExcludes\", maybe \"noFsync\"? Anyway, just a\n> >> >> >> suggestion on making this easier for users & the implementation.\n> >> >> >>\n> >> >> >\n> >> >> > Maybe there's some way to handle this I'm unaware of, but a\n> >> >> > disadvantage of your multi-valued config proposal is that it's harder,\n> >> >> > for example, for a per-repo config store to reasonably override a\n> >> >> > per-user config store.  With the configuration scheme as-is, I can\n> >> >> > have a per-user setting like `core.fsync=all` which covers my typical\n> >> >> > repos, but then have a maintainer repo with a private setting of\n> >> >> > `core.fsync=none` to speed up cases where I'm mostly working with\n> >> >> > other people's changes that are backed up in email or server-side\n> >> >> > repos.  The latter setting conveniently overrides the former setting\n> >> >> > in all aspects.\n> >> >>\n> >> >> Even if you turn just your comma-delimited proposal into a list proposal\n> >> >> can't we just say that the last one wins? Then it won't matter what cmae\n> >> >> before, you'd specify \"core.fsync=none\" in your local .git/config.\n> >> >>\n> >> >> But this is also a general issue with a bunch of things in git's config\n> >> >> space. I'd rather see us use the list-based values and just come up with\n> >> >> some general way to reset them that works for all keys, rather than\n> >> >> regretting having comma-delimited values that'll be harder to work with\n> >> >> & parse, which will be a legacy wart if/when we come up with a way to\n> >> >> say \"reset all previous settings\".\n> >> >>\n> >> >> > Also, with the core.fsync and core.fsyncExcludes, how would you spell\n> >> >> > \"don't sync anything\"? Would you still have the aggregate options.?\n> >> >>\n> >> >>     core.fsyncExcludes = *\n> >> >>\n> >> >> I.e. the same as the core.fsync=none above, anyway I still like the\n> >> >> wildcard thing below a bit more...\n> >> >\n> >> > I'm not going to take this feedback unless there are additional votes\n> >> > from the Git community in this direction.  I make the claim that\n> >> > single-valued comma-separated config lists are easier to work with in\n> >> > the existing Git infrastructure.\n> >>\n> >> Easier in what sense? I showed examples of how \"git-config\" trivially\n> >> works with multi-valued config, but for comma-delimited you'll need to\n> >> write your own shellscript boilerplate around simple things like adding\n> >> values, removing existing ones etc.\n> >>\n> >> I.e. you can use --add, --unset, --unset-all, --get, --get-all etc.\n> >>\n> >\n> > I see what you're saying for cases where someone would want to set a\n> > core.fsync setting that's derived from the user's current config.  But\n> > I'm guessing that the dominant use case is someone setting a new fsync\n> > configuration that achieves some atomic goal with respect to a given\n> > repository. Like \"this is a throwaway, so sync nothing\" or \"this is\n> > really important, so sync all objects and refs and the index\".\n>\n> Whether it's multi-value or comma-separated you could do:\n>\n>     -c core.fsync=[none|false]\n>\n> To sync nothing, i.e. if we see \"none/false\" it doesn't matter if we saw\n> core.fsync=loose-object or whatever before, it would override it, ditto\n> for \"all\".\n>\n> >> > parsing code for the core.whitespace variable and users are used to\n> >> > this syntax there. There are several other comma-separated lists in\n> >> > the config space, so this construct has precedence and will be with\n> >> > Git for some time.\n> >>\n> >> That's not really an argument either way for why we'd pick X over Y for\n> >> something new. We've got some comma-delimited, some multi-value (I'm\n> >> fairly sure we have more multi-value).\n> >>\n> >\n> > My main point here is that there's precedent for patch's current exact\n> > schema for a config in the same core config leaf of the Documentation.\n> > It seems from your comments that we'd have to invent and document some\n> > new convention for \"reset\" of a multi-valued config.  So you're\n> > suggesting that I solve an extra set of problems to get this change\n> > in.  Just want to remind you that my personal itch to scratch is to\n> > get the underlying mechanism in so that git-for-windows can set its\n> > default setting to batch mode. I'm not expecting many users to\n> > actually configure this setting to any non-default value.\n>\n> Me neither. I think people will most likely set this once in\n> ~/.gitconfig or /etc/gitconfig.\n>\n> We have some config keys that are multi-value and either comma-separated\n> or space-separated, e.g. core.alternateRefsPrefixes\n>\n> Then we have e.g. blame.ignoreRevsFile which is multi-value, and further\n> has the convention that setting it to an empty value clears the\n> list. which would scratch the \"override existing\" itch.\n>\n> format.notes, versionsort.suffix, transfer.hideRefs, branch.<name>.merge\n> are exmples of existing multi-value config.\n>\n\nThanks for the examples.  I can see the benefit of mutli-value for\nmost of those settings, but for versionsort.suffix, I'd personally\nhave wanted a comma-separated list.\n\n> >> > Also, fsync configurations aren't composable like\n> >> > some other configurations may be. It makes sense to have a holistic\n> >> > singular fsync configuration, which is best represented by a single\n> >> > variable.\n> >>\n> >> What's a \"variable\" here? We call these \"keys\", you can have a\n> >> single-value key like user.name that you get with --get, or a\n> >> multi-value key like say branch.<name>.merge or push.pushOption that\n> >> you'd get with --get-all.\n> >\n> > Yeah, I meant \"key\".  I conflated the config key with the underlying\n> > global variable in git.\n>\n> *nod*\n>\n> >> I think you may be referring to either not wanting these to be\n> >> \"inherited\" (which is not really a think we do for anything else in\n> >> config), or lacking the ability to \"reset\".\n> >>\n> >> For the latter if that's absolutely needed we could just use the same\n> >> trick as \"diff.wsErrorHighlight\" uses of making an empty value \"reset\"\n> >> the list, and you'd get better \"git config\" support for editing it.\n> >>\n> >\n> > My reading of the code is that diff.wsErrorHighlight is a comma\n> > separated list and not a multi-valued config.  Actually I haven't yet\n> > found an existing multi-valued config (not sure how to grep for it).\n>\n> Yes, I think I conflated it with one of the ones above when I wrote\n> this.\n>\n> >> >> >> > This also supports the common request of doing absolutely no\n> >> >> >> > fysncing with the `core.fsync=none` value, which is expected\n> >> >> >> > to make the test suite faster.\n> >> >> >>\n> >> >> >> Let's just use the git_parse_maybe_bool() or git_parse_maybe_bool_text()\n> >> >> >> so we'll accept \"false\", \"off\", \"no\" like most other such config?\n> >> >> >\n> >> >> > Junio's previous feedback when discussing batch mode [1] was to offer\n> >> >> > less flexibility when parsing new values of these configuration\n> >> >> > options. I agree with his statement that \"making machine-readable\n> >> >> > tokens be spelled in different ways is a 'disease'.\"  I'd like to\n> >> >> > leave this as-is so that the documentation can clearly state the exact\n> >> >> > set of allowable values.\n> >> >> >\n> >> >> > [1] https://lore.kernel.org/git/xmqqr1dqzyl7.fsf@gitster.g/\n> >> >>\n> >> >> I think he's talking about batch, Batch, BATCH, bAtCh etc. there. But\n> >> >> the \"maybe bool\" is a stanard pattern we use.\n> >> >>\n> >> >> I don't think we'd call one of these 0, off, no or false etc. to avoid\n> >> >> confusion, so then you can use git_parse_maybe_...()\n> >> >\n> >> > I don't see the advantage of having multiple ways of specifying\n> >> > \"none\".  The user can read the doc and know exactly what to write.  If\n> >> > they write something unallowable, they get a clear warning and they\n> >> > can read the doc again to figure out what to write.  This isn't a\n> >> > boolean options at all, so why should we entertain bool-like ways of\n> >> > spelling it?\n> >>\n> >> It's not boolean, it's multi-value and one of the values includes a true\n> >> or false boolean value. You just spell it \"none\".\n> >>\n> >> I think both this and your comment above suggest that you think there's\n> >> no point in this because you haven't interacted with/used \"git config\"\n> >> as a command line or API mechanism, but have just hand-crafted config\n> >> files.\n> >>\n> >> That's fair enough, but there's a lot of tooling that benefits from the\n> >> latter.\n> >\n> > My batch mode perf tests (on github, not yet submitted to the list)\n> > use `git -c core.fsync=<foo>` to set up a per-process config.  I\n> > haven't used the `git config` writing support in a while, so I haven't\n> > deeply thought about that.  However, I favor simplifying the use case\n> > of \"atomically\" setting a new holistic core.fsync value versus the use\n> > case of deriving a new core.fsync value from the preexisting value.\n>\n> If you implement it like blame.ignoreRevsFile you can have your cake and\n> eat it too, i.e.:\n>\n>     -c core.fsync= core.fsync=loose-object\n>\n> ensures only loose objects are synced, as with your single-value config,\n> but I'd think what you'd be more likely to actually mean would be:\n>\n>     # With \"core.fsync=pack\" set in ~/.gitconfig\n>     -c core.fsync=loose-object\n>\n> I.e. that the common case is \"I want this to be synced here\", not that\n> you'd like to declare sync policy from scratch every time.\n>\n> In any case, on this general topic my main point is that the\n> git-config(1) command has pretty integration for multi-value if you do\n> it that way, and not for comma-delimited. I.e. you get --add, --unset,\n> --unset-all, --get, --get-all etc.\n>\n> So I think for anything new it makes sense to lean into that, I think\n> most of these existing comma-delimited ones are ones we'd do differently\n> today on reflection.\n>\n> And if you suppor the \"empty resets\" like blame.ignoreRevsFile it seems\n> to me you'll have your cake & eat it too.\n>\n\nI really don't like the multiple `-c core.fsync=` clauses.  If I want\nto do packs and loose-objects as a single config, I'd have to do:\n        * multi-value: `git -c core.fsync= -c core.fsync=pack -c\ncore.fsync=loose-object`\n        * comma-sep `git -c core.fsync=pack,loose-object`\n\nMulti-valued configs are stateful and more verbose to configure.\nLast-one-wins with comma-separated values has an advantage for\nachieving a desired final configuration without regard for the\nprevious configuration, which is the way I expect the feature to be\nused.\n\n> >> E.g.:\n> >>\n> >>     $ git -c core.fsync=off config --type=bool core.fsync\n> >>     false\n> >>     $ git -c core.fsync=blah config --type=bool core.fsync\n> >>     fatal: bad boolean config value 'blah' for 'core.fsync'\n> >>\n> >> Here we can get 'git config' to normalize what you call 'none', and you\n> >> can tell via exit codes/normalization if it's \"false\". But if you invent\n> >> a new term for \"false\" you can't do that as easily.\n> >>\n> >> We have various historical keys that take odd exceptions to that,\n> >> e.g. \"never\", but unless we have a good reason to let's not invent more\n> >> exceptions.\n> >>\n> >> >> >> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> >> >> >> > ---\n> >> >> >> >  Documentation/config/core.txt | 27 +++++++++----\n> >> >> >> >  builtin/fast-import.c         |  2 +-\n> >> >> >> >  builtin/index-pack.c          |  4 +-\n> >> >> >> >  builtin/pack-objects.c        |  8 ++--\n> >> >> >> >  bulk-checkin.c                |  5 ++-\n> >> >> >> >  cache.h                       | 39 +++++++++++++++++-\n> >> >> >> >  commit-graph.c                |  3 +-\n> >> >> >> >  config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n> >> >> >> >  csum-file.c                   |  5 ++-\n> >> >> >> >  csum-file.h                   |  3 +-\n> >> >> >> >  environment.c                 |  1 +\n> >> >> >> >  midx.c                        |  3 +-\n> >> >> >> >  object-file.c                 |  3 +-\n> >> >> >> >  pack-bitmap-write.c           |  3 +-\n> >> >> >> >  pack-write.c                  | 13 +++---\n> >> >> >> >  read-cache.c                  |  2 +-\n> >> >> >> >  16 files changed, 164 insertions(+), 33 deletions(-)\n> >> >> >> >\n> >> >> >> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> >> >> >> > index dbb134f7136..4f1747ec871 100644\n> >> >> >> > --- a/Documentation/config/core.txt\n> >> >> >> > +++ b/Documentation/config/core.txt\n> >> >> >> > @@ -547,6 +547,25 @@ core.whitespace::\n> >> >> >> >    is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n> >> >> >> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >> >> >> >\n> >> >> >> > +core.fsync::\n> >> >> >> > +     A comma-separated list of parts of the repository which should be\n> >> >> >> > +     hardened via the core.fsyncMethod when created or modified. You can\n> >> >> >> > +     disable hardening of any component by prefixing it with a '-'. Later\n> >> >> >> > +     items take precedence over earlier ones in the list. For example,\n> >> >> >> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n> >> >> >> > +     metadata.\" Items that are not hardened may be lost in the event of an\n> >> >> >> > +     unclean system shutdown.\n> >> >> >> > ++\n> >> >> >> > +* `none` disables fsync completely. This must be specified alone.\n> >> >> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> >> >> >> > +* `pack` hardens objects added to the repo in packfile form.\n> >> >> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> >> >> >> > +* `commit-graph` hardens the commit graph file.\n> >> >> >> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n> >> >> >> > +  `pack-metadata`, and `commit-graph`.\n> >> >> >> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n> >> >> >> > +* `all` is an aggregate option that syncs all individual components above.\n> >> >> >> > +\n> >> >> >>\n> >> >> >> It's probably a *bit* more work to set up, but I wonder if this wouldn't\n> >> >> >> be simpler if we just said (and this is partially going against what I\n> >> >> >> noted above):\n> >> >> >>\n> >> >> >> == BEGIN DOC\n> >> >> >>\n> >> >> >> core.fsync is a multi-value config variable where each item is a\n> >> >> >> pathspec that'll get matched the same way as 'git-ls-files' et al.\n> >> >> >>\n> >> >> >> When we sync pretend that a path like .git/objects/de/adbeef... is\n> >> >> >> relative to the top-level of the git\n> >> >> >> directory. E.g. \"objects/de/adbeaf..\" or \"objects/pack/...\".\n> >> >> >>\n> >> >> >> You can then supply a list of wildcards and exclusions to configure\n> >> >> >> syncing.  or \"false\", \"off\" etc. to turn it off. These are synonymous\n> >> >> >> with:\n> >> >> >>\n> >> >> >>     ; same as \"false\"\n> >> >> >>     core.fsync = \":!*\"\n> >> >> >>\n> >> >> >> Or:\n> >> >> >>\n> >> >> >>     ; same as \"true\"\n> >> >> >>     core.fsync = \"*\"\n> >> >> >>\n> >> >> >> Or, to selectively sync some things and not others:\n> >> >> >>\n> >> >> >>     ;; Sync objects, but not \"info\"\n> >> >> >>     core.fsync = \":!objects/info/**\"\n> >> >> >>     core.fsync = \"objects/**\"\n> >> >> >>\n> >> >> >> See gitrepository-layout(5) for details about what sort of paths you\n> >> >> >> might be expected to match. Not all paths listed there will go through\n> >> >> >> this mechanism (e.g. currently objects do, but nothing to do with config\n> >> >> >> does).\n> >> >> >>\n> >> >> >> We can and will match this against \"fake paths\", e.g. when writing out\n> >> >> >> packs we may match against just the string \"objects/pack\", we're not\n> >> >> >> going to re-check if every packfile we're writing matches your globs,\n> >> >> >> ditto for loose objects. Be reasonable!\n> >> >> >>\n> >> >> >> This metharism is intended as a shorthand that provides some flexibility\n> >> >> >> when fsyncing, while not forcing git to come up with labels for all\n> >> >> >> paths the git dir, or to support crazyness like \"objects/de/adbeef*\"\n> >> >> >>\n> >> >> >> More paths may be added or removed in the future, and we make no\n> >> >> >> promises that we won't move things around, so if in doubt use\n> >> >> >> e.g. \"true\" or a wide pattern match like \"objects/**\". When in doubt\n> >> >> >> stick to the golden path of examples provided in this documentation.\n> >> >> >>\n> >> >> >> == END DOC\n> >> >> >>\n> >> >> >>\n> >> >> >> It's a tad more complex to set up, but I wonder if that isn't worth\n> >> >> >> it. It nicely gets around any current and future issues of deciding what\n> >> >> >> labels such as \"loose-object\" etc. to pick, as well as slotting into an\n> >> >> >> existing method of doing exclude/include lists.\n> >> >> >>\n> >> >> >\n> >> >> > I think this proposal is a lot of complexity to avoid coming up with a\n> >> >> > new name for syncable things as they are added to Git.  A path based\n> >> >> > mechanism makes it hard to document for the (advanced) user what the\n> >> >> > full set of things is and how it might change from release to release.\n> >> >> > I think the current core.fsync scheme is a bit easier to understand,\n> >> >> > query, and extend.\n> >> >>\n> >> >> We document it in gitrepository-layout(5). Yeah it has some\n> >> >> disadvantages, but one advantage is that you could make the\n> >> >> composability easy. I.e. if last exclude wins then a setting of:\n> >> >>\n> >> >>     core.fsync = \":!*\"\n> >> >>     core.fsync = \"objects/**\"\n> >> >>\n> >> >> Would reset all previous matches & only match objects/**.\n> >> >>\n> >> >\n> >> > The value of changing this is predicated on taking your previous\n> >> > multi-valued config proposal, which I'm still not at all convinced\n> >> > about.\n> >>\n> >> They're orthagonal. I.e. you get benefits from multi-value with or\n> >> without this globbing mechanism.\n> >>\n> >> In any case, I don't feel strongly about/am really advocating this\n> >> globbing mechanism. I just wondered if it wouldn't make things simpler\n> >> since it would sidestep the need to create any sort of categories for\n> >> subsets of gitrepository-layout(5), but maybe not...\n> >>\n> >> > The schema in the current (v1-v2) version of the patch already\n> >> > includes an example of extending the list of syncable things, and\n> >> > Patrick Steinhardt made it clear that he feels comfortable adding\n> >> > 'refs' to the same schema in a future change.\n> >> >\n> >> > I'll also emphasize that we're talking about a non-functional,\n> >> > relatively corner-case behavioral configuration.  These values don't\n> >> > change how git's interface behaves except when the system crashes\n> >> > during a git command or shortly after one completes.\n> >>\n> >> That may be something some OS's promise, but it's not something fsync()\n> >> or POSIX promises. I.e. you might write a ref, but unless you fsync and\n> >> the relevant dir entries another process might not see the update, crash\n> >> or not.\n> >>\n> >\n> > I haven't seen any indication that POSIX requires an fsync for\n> > visiblity within a running system.  I looked at the doc for open,\n> > write, and fsync, and saw no indication that it's posix compliant to\n> > require an fsync for visibility.  I think an OS that required fsync\n> > for cross-process visiblity would fail to run Git for a myriad of\n> > other reasons and would likely lose all its users.  I'm curious where\n> > you've seen documentation that allows such unhelpful behavior?\n>\n> There's multiple unrelated and related things in this area. One is a\n> case where you'll e.g. write a file \"foo\" using stdio, spawn a program\n> to work on it in the same program, but it might not see it at all, or\n> see empty content, the latter being because you haven't flushed your I/O\n> buffers (which you can do via fsync()).\n>\n\nFor stdio you need to use fflush(3), which just flushes the C\nruntime's internal buffers.  You need to do the following to do a full\ndurable write using stdio:\n```\nFILE *fp;\n...\nfflush(fp);\nfsync(fileno(fp))\n```\n\n> The former is that on *nix systems you're generally only guaranteed to\n> write to a fd, but not to have the associated metadata be synced for\n> you.\n>\n> That is spelled out e.g. in the DESCRIPTION section of linux's fsync()\n> manpage: https://man7.org/linux/man-pages/man2/fdatasync.2.html\n>\n> I don't know how much you follow non-Windows FS development, but there\n> was also a very well known \"incident\" early in ext4 where it leaned into\n> some permissive-by-POSIX behavior that caused data loss in practice on\n> buggy programs that didn't correctly use fsync() , since various tooling\n> had come to expect the stricter behavior of ext3:\n> https://lwn.net/Articles/328363/\n>\n> That was explicitly in the area of fs metadata being discussed here.\n>\n> Generally you can expect your VFS layer to be forgiving when it comes to\n> IPC, but even that is out the window when it comes to networked\n> filesystems, e.g. a shared git repository hosted on NFS.\n>\n\nEverything in the fsync(2) DESCRIPTION section is about what data and\nmetadata reaches the disk (versus just being cached in-memory).  I've\nbecome a bit familiar with the ext3 vs ext4 (and delayed alloc)\nbehavior while researching this feature.  These behaviors are all\naround the durability you get in the case of kernel crash,\npower-failure, or other forms of dirty dismount.\n\nNFS is a complex story.  I'm not intimately familiar with its\nparticular pitfalls, but from looking at the Linux NFS faq\n(http://nfs.sourceforge.net/), it appears that a given single NFS\nclient will remain coherent with itself. Multiple NFS clients\naccessing a single Git repo concurrently are probably going to see\nsome inconsistency.  In that kind of case, fsync would help, perhaps,\nsince it would force NFS clients to issue a COMMIT command to the\nserver.\n\n> >> That's an aside from these config design questions, and I think\n> >> most/(all?) OS's anyone cares about these days tend to make that\n> >> implicit promise as part of their VFS behavior, but we should probably\n> >> design such an interface to fsync() with such pedantic portability in\n> >> mind.\n> >\n> > Why? To be rude to such a hypothetical system, if a system were so\n> > insanely designed, it would be nuts to support it.\n>\n> Because we know that right now the system calls we're invoking aren't\n> guaranteed to store data persistently to disk portably, although they do\n> so in practice on most modern OS's.\n>\n> We're portably to a lot of platforms, and also need to keep e.g. NFS in\n> mind, so being able to ask for a pedantic mode when you care about data\n> retention at the cost of performance would be nice.\n>\n> And because the fsync config mode you're proposing is thoroughly\n> non-standard, but is known to me much faster by leaning into known\n> attributes of specific FS's on specific OS's, if we're not running on\n> those it would be sensible to fall back to a stricter mode of\n> operation. E.g. syncing all 100 loose objects we just wrote, not just\n> the last one.\n>\n> >> > While you may not personally love the proposed configuration\n> >> > interface, I'd want your view on some questions:\n> >> > 1. Is it easy for the (advanced) user to set a configuration?\n> >> > 2. Is it easy for the (advanced) user to see what was configured?\n> >> > 3. Is it easy for the Git community to build on this as we want to add\n> >> > things to the list of things to sync?\n> >> >     a) Is there a good best practice configuration so that people can\n> >> > avoid losing integrity for new stuff that they are intending to sync.\n> >> >     b) If someone has a custom configuration, can that custom\n> >> > configuration do something reasonable as they upgrade versions of Git?\n> >> >              ** In response to this question, I might see some value\n> >> > in adding a 'derived-metadata' aggregate that can be disabled so that\n> >> > a custom configuration can exclude those as they change version to\n> >> > version.\n> >> >     c) Is it too much maintenance overhead to consider how to present\n> >> > this configuration knob for any new hashfile or other datafile in the\n> >> > git repo?\n> >> > 4. Is there a good path forward to change the default syncable set,\n> >> > both in git-for-windows and in Git for other platforms?\n> >>\n> >> I'm not really sure this globbing this is a good idea, as noted above\n> >> just a suggestion etc.\n> >>\n> >> As noted there it just gets you out of the business of re-defining\n> >> gitrepository-layout(5), and assuming too much in advance about certain\n> >> use-cases.\n> >>\n> >> E.g. even \"refs\" might be too broad for some. I don't tend to be I/O\n> >> limited, but I could see how someone who would be would care about\n> >> refs/heads but not refs/remotes, or want to exclude logs/* but not the\n> >> refs updates themselves etc.\n> >\n> > This use-case is interesting (distinguishing remote refs from local\n> > refs).  I think the difficulty of verifying (for even an advanced\n> > user) that the right fsyncing is actually happening still puts me on\n> > the side of having a carefully curated and documented set of syncable\n> > things rather than a file-path-based mechanism.\n> >\n> > Is this meaningful in the presumably nearby future world of the refsdb\n> > backend?  Is that somehow split by remote versus local?\n>\n> There is the upcoming \"reftable\" work, but that's probably 2-3 years out\n> at the earliest for series production workloads in git.git.\n>\n> >> >> >> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> >> >> >> > index 857be7826f3..916c55d6ce9 100644\n> >> >> >> > --- a/builtin/pack-objects.c\n> >> >> >> > +++ b/builtin/pack-objects.c\n> >> >> >> > @@ -1204,11 +1204,13 @@ static void write_pack_file(void)\n> >> >> >> >                * If so, rewrite it like in fast-import\n> >> >> >> >                */\n> >> >> >> >               if (pack_to_stdout) {\n> >> >> >> > -                     finalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> >> >> >> > +                     finalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n> >> >> >> > +                                       CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n> >> >> >>\n> >> >> >> Not really related to this per-se, but since you're touching the API\n> >> >> >> everything goes through I wonder if callers should just always try to\n> >> >> >> fsync, and we can just catch EROFS and EINVAL in the wrapper if someone\n> >> >> >> tries to flush stdout, or catch the fd at that lower level.\n> >> >> >>\n> >> >> >> Or maybe there's a good reason for this...\n> >> >> >\n> >> >> > It's platform dependent, but I'd expect fsync would do something for\n> >> >> > pipes or stdout redirected to a file.  In these cases we really don't\n> >> >> > want to fsync since we have no idea what we're talking to and we're\n> >> >> > potentially worsening performance for probably no benefit.\n> >> >>\n> >> >> Yeah maybe we should just leave it be.\n> >> >>\n> >> >> I'd think the C library returning EINVAL would be a trivial performance\n> >> >> cost though.\n> >> >>\n> >> >> It just seemed odd to hardcode assumptions about what can and can't be\n> >> >> synced when the POSIX defined function will also tell us that.\n> >> >>\n> >> >\n> >> > Redirecting stdout to a file seems like a common usage for this\n> >> > command. That would definitely be fsyncable, but Git has no idea what\n> >> > its proper category is since there's no way to know the purpose or\n> >> > lifetime of the packfile.  I'm going to leave this be, because I'd\n> >> > posit that \"can it be fsynced?\" is not the same as \"should it be\n> >> > fsynced?\".  The latter question can't be answered for stdout.\n> >>\n> >> As noted this was just an aside, and I don't even know if any OS would\n> >> do anything meaningful with an fsync() of such a FD anyway.\n> >>\n> >\n> > The underlying fsync primitive does have a meaning on Windows for\n> > pipes, but it's certainly not what Git would want to do. Also if\n> > stdout is redirected to a file, I'm pretty sure that UNIX OSes would\n> > respect the fsync call.  However it's not meaningful in the sense of\n> > the git repository, since we don't know what the packfile is or why it\n> > was created.\n>\n> I suggested that because I think it's probably nonsensical, but it's\n> nonsense that POSIX seems to explicitly tell us that it'll handle\n> (probably by silently doing nothing). So in terms of our interface we\n> could lean into that and avoid our own special-casing.\n>\n> >> I just don't see why we wouldn't say:\n> >>\n> >>  1. We're syncing this category of thing\n> >>  2. Try it\n> >>  3. If fsync returns \"can't fsync that sort of thing\" move on\n> >>\n> >> As opposed to trying to shortcut #3 by doing the detection ourselves.\n> >>\n> >> I.e. maybe there was a good reason, but it seemed to be some easy\n> >> potential win for more simplification, since you were re-doing and\n> >> simplifying some of the interface anyway...\n> >\n> > We're trying to be deliberate about what we're fsyncing.  Fsyncing an\n> > unknown file created by the packfile code doesn't move us in that\n> > direction.  In your taxonomy we don't know (1), \"what is this category\n> > of thing?\"  Sure it's got the packfile format, but is not known to be\n> > an actual packfile that's part of the repository.\n>\n> We know it's a fd, isn't that sufficient? In any case, I'm fine with\n> also keeping it as is, I don't mean to split hairs here.\n>\n> It just stuck out as an odd part of the interface, why treat some fd's\n> specially, instead of just throwing it all at the OS. Presumably the\n> first thing the OS will do is figure out if it's a syncable fd or not,\n> and act appropriately.\n>\n\nI'll put the following comment in pack-objects.c:\n/*\n* We never fsync when writing to stdout since we may\n* not be writing to a specific file. For instance, the\n* upload-pack code passes a pipe here. Calling fsync\n* on a pipe results in unnecessary synchronization with\n* the reader on some platforms.\n*/\n\n> >> >>\n> >> >> >> > [...]\n> >> >> >> > +/*\n> >> >> >> > + * These values are used to help identify parts of a repository to fsync.\n> >> >> >> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n> >> >> >> > + * repository and so shouldn't be fsynced.\n> >> >> >> > + */\n> >> >> >> > +enum fsync_component {\n> >> >> >> > +     FSYNC_COMPONENT_NONE                    = 0,\n> >> >> >>\n> >> >> >> I haven't read ahead much but in most other such cases we don't define\n> >> >> >> the \"= 0\", just start at 1<<0, then check the flags elsewhere...\n> >> >> >>\n> >> >> >> > +static const struct fsync_component_entry {\n> >> >> >> > +     const char *name;\n> >> >> >> > +     enum fsync_component component_bits;\n> >> >> >> > +} fsync_component_table[] = {\n> >> >> >> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> >> >> >> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n> >> >> >> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> >> >> >> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> >> >> >> > +     { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> >> >> >> > +     { \"default\", FSYNC_COMPONENTS_DEFAULT },\n> >> >> >> > +     { \"all\", FSYNC_COMPONENTS_ALL },\n> >> >> >> > +};\n> >> >> >> > +\n> >> >> >> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> >> >> >> > +{\n> >> >> >> > +     enum fsync_component output = 0;\n> >> >> >> > +\n> >> >> >> > +     if (!strcmp(string, \"none\"))\n> >> >> >> > +             return output;\n> >> >> >> > +\n> >> >> >> > +     while (string) {\n> >> >> >> > +             int i;\n> >> >> >> > +             size_t len;\n> >> >> >> > +             const char *ep;\n> >> >> >> > +             int negated = 0;\n> >> >> >> > +             int found = 0;\n> >> >> >> > +\n> >> >> >> > +             string = string + strspn(string, \", \\t\\n\\r\");\n> >> >> >>\n> >> >> >> Aside from the \"use a list\" isn't this hardcoding some windows-specific\n> >> >> >> assumptions with \\n\\r? Maybe not...\n> >> >> >\n> >> >> > I shamelessly stole this code from parse_whitespace_rule. I thought\n> >> >> > about making a helper to be called by both functions, but the amount\n> >> >> > of state going into and out of the wrapper via arguments was\n> >> >> > substantial and seemed to negate the benefit of deduplication.\n> >> >>\n> >> >> FWIW string_list_split() is easier to work with in those cases, or at\n> >> >> least I think so...\n> >> >\n> >> > This code runs at startup for a variable that may be present on some\n> >> > installations.  The nice property of the current patch's code is that\n> >> > it's already a well-tested pattern that doesn't do any allocations as\n> >> > it's working, unlike string_list_split().\n> >>\n> >> Multi-value config would also get you fewer allocations :)\n> >>\n> >> Anyway, I mainly meant to point out that for stuff like this it's fine\n> >> to optimize it for ease rather than micro-optimize allocations. Those\n> >> really aren't a bottleneck on this scale.\n> >>\n> >> Even in that case there's string_list_split_in_place(), which can be a\n> >> bit nicer than manual C-string fiddling.\n> >>\n> >\n> > Am I allowed to change the config value string in place? The\n> > core.whitespace code is careful not to modify the string. I kind of\n> > like the parse_ws_error_highlight code a little better now that I've\n> > seen it, but I think the current code is fine too.\n>\n> I don't remember offhand if that's safe, probably not. So you'll need a\n> copy here.\n>\n> >> > I hope you know that I appreciate your review feedback, even though\n> >> > I'm pushing back on most of it so far this round. I'll be sending v3\n> >> > to the list soon after giving it another look over.\n> >>\n> >> Sure, no worries. Just hoping to help. If you go for something different\n> >> etc. that's fine. Just hoping to bridge the gap in some knowledge /\n> >> offer potentially interesting suggestions (some of which may be dead\n> >> ends, like the config glob thing...).\n\nThanks again for the review, I'll send an updated PR soon.\n\nThanks,\nNeeraj\n"},{"id":"447390","messageId":"pull.1093.v4.git.1643686424.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v3.git.1639011433.gitgitgadget@gmail.com","subject":"[PATCH v4 0/4] A design for future-proofing fsync() configuration","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-01T03:33:40Z","receivedAt":"2022-02-01T03:33:50Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"This is an implementation of an extensible configuration mechanism for\nfsyncing persistent components of a repo.\n\nThe main goals are to separate the \"what\" to sync from the \"how\". There are\nnow two settings: core.fsync - Control the 'what', including the index.\ncore.fsyncMethod - Control the 'how'. Currently we support writeout-only and\nfull fsync.\n\nSyncing of refs can be layered on top of core.fsync. And batch mode will be\nlayered on core.fsyncMethod.\n\ncore.fsyncObjectfiles is removed and will issue a deprecation warning if\nit's seen.\n\nI'd like to get agreement on this direction before submitting batch mode to\nthe list. The batch mode series is available to view at\nhttps://github.com/gitgitgadget/git/pull/1134\n\nPlease see [1], [2], and [3] for discussions that led to this series.\n\nAfter this change, new persistent data files added to the repo will need to\nbe added to the fsync_component enum and documented in the\nDocumentation/config/core.txt text.\n\nV4 changes:\n\n * Rebase onto master at b23dac905bd.\n * Add a comment to write_pack_file indicating why we don't fsync when\n   writing to stdout.\n * I kept the configuration schema as-is rather than switching to\n   multi-value. The thinking here is that a stateless last-one-wins config\n   schema (comma separated) will make it easier to achieve some holistic\n   self-consistent fsync configuration for a particular repo.\n\nV3 changes:\n\n * Remove relative path from git-compat-util.h include [4].\n * Updated newly added warning texts to have more context for localization\n   [4].\n * Fixed tab spacing in enum fsync_action\n * Moved the fsync looping out to a helper and do it consistently. [4]\n * Changed commit description to use camelCase for config names. [5]\n * Add an optional fourth patch with derived-metadata so that the user can\n   exclude a forward-compatible set of things that should be recomputable\n   given existing data.\n\nV2 changes:\n\n * Updated the documentation for core.fsyncmethod to be less certain.\n   writeout-only probably does not do the right thing on Linux.\n * Split out the core.fsync=index change into its own commit.\n * Rename REPO_COMPONENT to FSYNC_COMPONENT. This is really specific to\n   fsyncing, so the name should reflect that.\n * Re-add missing Makefile change for SYNC_FILE_RANGE.\n * Tested writeout-only mode, index syncing, and general config settings.\n\n[1] https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n[2]\nhttps://lore.kernel.org/git/dd65718814011eb93ccc4428f9882e0f025224a6.1636029491.git.ps@pks.im/\n[3]\nhttps://lore.kernel.org/git/pull.1076.git.git.1629856292.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/CANQDOdf8C4-haK9=Q_J4Cid8bQALnmGDm=SvatRbaVf+tkzqLw@mail.gmail.com/\n[5] https://lore.kernel.org/git/211207.861r2opplg.gmgdl@evledraar.gmail.com/\n\nNeeraj Singh (4):\n  core.fsyncmethod: add writeout-only mode\n  core.fsync: introduce granular fsync control\n  core.fsync: new option to harden the index\n  core.fsync: add a `derived-metadata` aggregate option\n\n Documentation/config/core.txt       | 35 ++++++++---\n Makefile                            |  6 ++\n builtin/fast-import.c               |  2 +-\n builtin/index-pack.c                |  4 +-\n builtin/pack-objects.c              | 24 +++++---\n bulk-checkin.c                      |  5 +-\n cache.h                             | 49 +++++++++++++++-\n commit-graph.c                      |  3 +-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n config.c                            | 90 ++++++++++++++++++++++++++++-\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n csum-file.c                         |  5 +-\n csum-file.h                         |  3 +-\n environment.c                       |  3 +-\n git-compat-util.h                   | 24 ++++++++\n midx.c                              |  3 +-\n object-file.c                       |  3 +-\n pack-bitmap-write.c                 |  3 +-\n pack-write.c                        | 13 +++--\n read-cache.c                        | 19 ++++--\n wrapper.c                           | 64 ++++++++++++++++++++\n write-or-die.c                      | 11 ++--\n 25 files changed, 367 insertions(+), 47 deletions(-)\n create mode 100644 compat/win32/flush.c\n\n\nbase-commit: b23dac905bde28da47543484320db16312c87551\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1093%2Fneerajsi-msft%2Fns%2Fcore-fsync-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1093/neerajsi-msft/ns/core-fsync-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/1093\n\nRange-diff vs v3:\n\n 1:  15edfe51509 ! 1:  51a218d100d core.fsyncmethod: add writeout-only mode\n     @@ Makefile: ifdef HAVE_CLOCK_MONOTONIC\n       endif\n      \n       ## cache.h ##\n     -@@ cache.h: extern int read_replace_refs;\n     - extern char *git_replace_ref_base;\n     +@@ cache.h: extern char *git_replace_ref_base;\n       \n       extern int fsync_object_files;\n     + extern int use_fsync;\n      +\n      +enum fsync_method {\n      +\tFSYNC_METHOD_FSYNC,\n     @@ compat/win32/flush.c (new)\n      +\n      +#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n      +\n     -+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n     ++       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NTAPI, NtFlushBuffersFileEx,\n      +\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n      +\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n      +\n     @@ contrib/buildsystems/CMakeLists.txt: if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n       \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\n      \n       ## environment.c ##\n     -@@ environment.c: const char *git_attributes_file;\n     - const char *git_hooks_path;\n     - int zlib_compression_level = Z_BEST_SPEED;\n     +@@ environment.c: int zlib_compression_level = Z_BEST_SPEED;\n       int pack_compression_level = Z_DEFAULT_COMPRESSION;\n     --int fsync_object_files;\n     + int fsync_object_files;\n     + int use_fsync = -1;\n      +enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n       size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n       size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n       \tint err;\n      \n       ## write-or-die.c ##\n     -@@ write-or-die.c: void fprintf_or_die(FILE *f, const char *fmt, ...)\n     - \n     - void fsync_or_die(int fd, const char *msg)\n     - {\n     +@@ write-or-die.c: void fsync_or_die(int fd, const char *msg)\n     + \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n     + \tif (!use_fsync)\n     + \t\treturn;\n      -\twhile (fsync(fd) < 0) {\n      -\t\tif (errno != EINTR)\n      -\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n      -\t}\n     ++\n      +\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n      +\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n      +\t\treturn;\n 2:  080be1a6f64 ! 2:  7a164ba9571 core.fsync: introduce granular fsync control\n     @@ builtin/index-pack.c: static void final(const char *final_pack_name, const char\n      \n       ## builtin/pack-objects.c ##\n      @@ builtin/pack-objects.c: static void write_pack_file(void)\n     - \t\t * If so, rewrite it like in fast-import\n     - \t\t */\n     + \t\t\tdisplay_progress(progress_state, written);\n     + \t\t}\n     + \n     +-\t\t/*\n     +-\t\t * Did we write the wrong # entries in the header?\n     +-\t\t * If so, rewrite it like in fast-import\n     +-\t\t */\n       \t\tif (pack_to_stdout) {\n      -\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n     ++\t\t\t/*\n     ++\t\t\t * We never fsync when writing to stdout since we may\n     ++\t\t\t * not be writing to an actual pack file. For instance,\n     ++\t\t\t * the upload-pack code passes a pipe here. Calling\n     ++\t\t\t * fsync on a pipe results in unnecessary\n     ++\t\t\t * synchronization with the reader on some platforms.\n     ++\t\t\t */\n      +\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n      +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n       \t\t} else if (nr_written == nr_remaining) {\n     @@ builtin/pack-objects.c: static void write_pack_file(void)\n      +\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n       \t\t} else {\n      -\t\t\tint fd = finalize_hashfile(f, hash, 0);\n     ++\t\t\t/*\n     ++\t\t\t * If we wrote the wrong number of entries in the\n     ++\t\t\t * header, rewrite it like in fast-import.\n     ++\t\t\t */\n     ++\n      +\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n       \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n       \t\t\t\t\t\t nr_written, hash, offset);\n     @@ cache.h: void reset_shared_repository(void);\n       extern char *git_replace_ref_base;\n       \n      -extern int fsync_object_files;\n     +-extern int use_fsync;\n      +/*\n      + * These values are used to help identify parts of a repository to fsync.\n      + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n      + * repository and so shouldn't be fsynced.\n      + */\n      +enum fsync_component {\n     -+\tFSYNC_COMPONENT_NONE\t\t\t= 0,\n     ++\tFSYNC_COMPONENT_NONE,\n      +\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n      +\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n      +\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n     @@ cache.h: void reset_shared_repository(void);\n       \n       enum fsync_method {\n       \tFSYNC_METHOD_FSYNC,\n     +@@ cache.h: enum fsync_method {\n     + };\n     + \n     + extern enum fsync_method fsync_method;\n     ++extern int use_fsync;\n     + extern int core_preload_index;\n     + extern int precomposed_unicode;\n     + extern int protect_hfs;\n      @@ cache.h: int copy_file_with_time(const char *dst, const char *src, int mode);\n       void write_or_die(int fd, const void *buf, size_t count);\n       void fsync_or_die(int fd, const char *);\n     @@ csum-file.h: int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint\n       void crc32_begin(struct hashfile *);\n      \n       ## environment.c ##\n     -@@ environment.c: const char *git_hooks_path;\n     +@@ environment.c: const char *git_attributes_file;\n     + const char *git_hooks_path;\n       int zlib_compression_level = Z_BEST_SPEED;\n       int pack_compression_level = Z_DEFAULT_COMPRESSION;\n     +-int fsync_object_files;\n     + int use_fsync = -1;\n       enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n      +enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n       size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n     @@ midx.c: static int write_midx_internal(const char *object_dir,\n      \n       ## object-file.c ##\n      @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n     - /* Finalize a file on disk, and close it. */\n       static void close_loose_object(int fd)\n       {\n     --\tif (fsync_object_files)\n     --\t\tfsync_or_die(fd, \"loose object file\");\n     -+\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n     + \tif (!the_repository->objects->odb->will_destroy) {\n     +-\t\tif (fsync_object_files)\n     +-\t\t\tfsync_or_die(fd, \"loose object file\");\n     ++\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n     + \t}\n     + \n       \tif (close(fd) != 0)\n     - \t\tdie_errno(_(\"error when closing loose object file\"));\n     - }\n      \n       ## pack-bitmap-write.c ##\n      @@ pack-bitmap-write.c: void bitmap_writer_finish(struct pack_idx_entry **index,\n 3:  2207950beba = 3:  f217dba77a1 core.fsync: new option to harden the index\n 4:  a830d177d4c = 4:  5c22a41c1f3 core.fsync: add a `derived-metadata` aggregate option\n\n-- \ngitgitgadget\n"},{"id":"447391","messageId":"51a218d100db2b8282b6a2b5c68033b532bcd19e.1643686425.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v4.git.1643686424.gitgitgadget@gmail.com","subject":"[PATCH v4 1/4] core.fsyncmethod: add writeout-only mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-01T03:33:41Z","receivedAt":"2022-02-01T03:33:51Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsyncMethod` configuration\nknob, which can currently be set to `fsync` or `writeout-only`.\n\nThe new writeout-only mode attempts to tell the operating system to\nflush its in-memory page cache to the storage hardware without issuing a\nCACHE_FLUSH command to the storage controller.\n\nWriteout-only fsync is significantly faster than a vanilla fsync on\ncommon hardware, since data is written to a disk-side cache rather than\nall the way to a durable medium. Later changes in this patch series will\ntake advantage of this primitive to implement batching of hardware\nflushes.\n\nWhen git_fsync is called with FSYNC_WRITEOUT_ONLY, it may fail and the\ncaller is expected to do an ordinary fsync as needed.\n\nOn Apple platforms, the fsync system call does not issue a CACHE_FLUSH\ndirective to the storage controller. This change updates fsync to do\nfcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\nwith existing behavior on Apple platforms by setting the default value\nof the new core.fsyncMethod option.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt       |  9 ++++\n Makefile                            |  6 +++\n cache.h                             |  7 ++++\n compat/mingw.h                      |  3 ++\n compat/win32/flush.c                | 28 +++++++++++++\n config.c                            | 12 ++++++\n config.mak.uname                    |  3 ++\n configure.ac                        |  8 ++++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n environment.c                       |  1 +\n git-compat-util.h                   | 24 +++++++++++\n wrapper.c                           | 64 +++++++++++++++++++++++++++++\n write-or-die.c                      | 11 +++--\n 13 files changed, 174 insertions(+), 5 deletions(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..dbb134f7136 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,15 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsyncMethod::\n+\tA value indicating the strategy Git will use to harden repository data\n+\tusing fsync and related primitives.\n++\n+* `fsync` uses the fsync() system call or platform equivalents.\n+* `writeout-only` issues pagecache writeback requests, but depending on the\n+  filesystem and storage hardware, data added to the repository may not be\n+  durable in the event of a system crash. This is the default mode on macOS.\n+\n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n +\ndiff --git a/Makefile b/Makefile\nindex 5580859afdb..1eff9953280 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -405,6 +405,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1892,6 +1894,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/cache.h b/cache.h\nindex 281f00ab1b1..37a32034b2f 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -995,6 +995,13 @@ extern char *git_replace_ref_base;\n \n extern int fsync_object_files;\n extern int use_fsync;\n+\n+enum fsync_method {\n+\tFSYNC_METHOD_FSYNC,\n+\tFSYNC_METHOD_WRITEOUT_ONLY\n+};\n+\n+extern enum fsync_method fsync_method;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..291f90ea940\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NTAPI, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.c b/config.c\nindex 2bffa8d4a01..f67f545f839 100644\n--- a/config.c\n+++ b/config.c\n@@ -1490,6 +1490,18 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsyncmethod\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcmp(value, \"fsync\"))\n+\t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n+\t\telse if (!strcmp(value, \"writeout-only\"))\n+\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse\n+\t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n+\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\ndiff --git a/config.mak.uname b/config.mak.uname\nindex c48db45106c..2c67b3b93ce 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -57,6 +57,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tBASIC_CFLAGS += -DHAVE_SYSINFO\n@@ -462,6 +463,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -639,6 +641,7 @@ ifeq ($(uname_S),MINGW)\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/configure.ac b/configure.ac\nindex d60d494ee4c..660b91f90b4 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1095,6 +1095,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 5100f56bb37..276e74c1d54 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/environment.c b/environment.c\nindex fd0501e77a5..3e3620d759f 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -44,6 +44,7 @@ int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n int fsync_object_files;\n int use_fsync = -1;\n+enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 1229c8296b9..76cb85a56b1 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1265,6 +1265,30 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+#ifdef __APPLE__\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n+#else\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n+#endif\n+\n+enum fsync_action {\n+\tFSYNC_WRITEOUT_ONLY,\n+\tFSYNC_HARDWARE_FLUSH\n+};\n+\n+/*\n+ * Issues an fsync against the specified file according to the specified mode.\n+ *\n+ * FSYNC_WRITEOUT_ONLY attempts to use interfaces available on some operating\n+ * systems to flush the OS cache without issuing a flush command to the storage\n+ * controller. If those interfaces are unavailable, the function fails with\n+ * ENOSYS.\n+ *\n+ * FSYNC_HARDWARE_FLUSH does an OS writeout and hardware flush to ensure that\n+ * changes are durable. It is not expected to fail.\n+ */\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/wrapper.c b/wrapper.c\nindex 36e12119d76..572f28f14ff 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -546,6 +546,70 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+/*\n+ * Some platforms return EINTR from fsync. Since fsync is invoked in some\n+ * cases by a wrapper that dies on failure, do not expose EINTR to callers.\n+ */\n+static int fsync_loop(int fd)\n+{\n+\tint err;\n+\n+\tdo {\n+\t\terr = fsync(fd);\n+\t} while (err < 0 && errno == EINTR);\n+\treturn err;\n+}\n+\n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync_loop(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\t\t/*\n+\t\t * On some platforms fsync may return EINTR. Try again in this\n+\t\t * case, since callers asking for a hardware flush may die if\n+\t\t * this function returns an error.\n+\t\t */\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync_loop(fd);\n+#endif\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex a3d5784cec9..9faa5f9f563 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -62,10 +62,13 @@ void fsync_or_die(int fd, const char *msg)\n \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n \tif (!use_fsync)\n \t\treturn;\n-\twhile (fsync(fd) < 0) {\n-\t\tif (errno != EINTR)\n-\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n-\t}\n+\n+\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n+\t\treturn;\n+\n+\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n+\t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n void write_or_die(int fd, const void *buf, size_t count)\n-- \ngitgitgadget\n\n"},{"id":"447392","messageId":"7a164ba95710b4231d07982fd27ec51022929b81.1643686425.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v4.git.1643686424.gitgitgadget@gmail.com","subject":"[PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-01T03:33:42Z","receivedAt":"2022-02-01T03:33:56Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsync` configuration\nknob which can be used to control how components of the\nrepository are made durable on disk.\n\nThis setting allows future extensibility of the list of\nsyncable components:\n* We issue a warning rather than an error for unrecognized\n  components, so new configs can be used with old Git versions.\n* We support negation, so users can choose one of the default\n  aggregate options and then remove components that they don't\n  want. The user would then harden any new components added in\n  a Git version update.\n\nThis also supports the common request of doing absolutely no\nfysncing with the `core.fsync=none` value, which is expected\nto make the test suite faster.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 27 +++++++++----\n builtin/fast-import.c         |  2 +-\n builtin/index-pack.c          |  4 +-\n builtin/pack-objects.c        | 24 +++++++----\n bulk-checkin.c                |  5 ++-\n cache.h                       | 41 ++++++++++++++++++-\n commit-graph.c                |  3 +-\n config.c                      | 76 ++++++++++++++++++++++++++++++++++-\n csum-file.c                   |  5 ++-\n csum-file.h                   |  3 +-\n environment.c                 |  2 +-\n midx.c                        |  3 +-\n object-file.c                 |  3 +-\n pack-bitmap-write.c           |  3 +-\n pack-write.c                  | 13 +++---\n read-cache.c                  |  2 +-\n 16 files changed, 177 insertions(+), 39 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex dbb134f7136..4f1747ec871 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,25 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsync::\n+\tA comma-separated list of parts of the repository which should be\n+\thardened via the core.fsyncMethod when created or modified. You can\n+\tdisable hardening of any component by prefixing it with a '-'. Later\n+\titems take precedence over earlier ones in the list. For example,\n+\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n+\tmetadata.\" Items that are not hardened may be lost in the event of an\n+\tunclean system shutdown.\n++\n+* `none` disables fsync completely. This must be specified alone.\n+* `loose-object` hardens objects added to the repo in loose-object form.\n+* `pack` hardens objects added to the repo in packfile form.\n+* `pack-metadata` hardens packfile bitmaps and indexes.\n+* `commit-graph` hardens the commit graph file.\n+* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n+  `pack-metadata`, and `commit-graph`.\n+* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n+* `all` is an aggregate option that syncs all individual components above.\n+\n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\n@@ -556,14 +575,6 @@ core.fsyncMethod::\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n \n-core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n-\n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\n +\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 2b2e28bad79..ff70aeb1a0e 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -858,7 +858,7 @@ static void end_packfile(void)\n \t\tstruct tag *t;\n \n \t\tclose_pack_windows(pack_data);\n-\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n \t\t\t\t\t pack_data->pack_name, object_count,\n \t\t\t\t\t cur_pack_oid.hash, pack_size);\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 3c2e6aee3cc..b871d721e95 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1286,7 +1286,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\t\t    nr_objects - nr_objects_initial);\n \t\tstop_progress_msg(&progress, msg.buf);\n \t\tstrbuf_release(&msg);\n-\t\tfinalize_hashfile(f, tail_hash, 0);\n+\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n \t\thashcpy(read_hash, pack_hash);\n \t\tfixup_pack_header_footer(output_fd, pack_hash,\n \t\t\t\t\t curr_pack, nr_objects,\n@@ -1508,7 +1508,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tif (!from_stdin) {\n \t\tclose(input_fd);\n \t} else {\n-\t\tfsync_or_die(output_fd, curr_pack_name);\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n \t\terr = close(output_fd);\n \t\tif (err)\n \t\t\tdie_errno(_(\"error while closing pack file\"));\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex ba2006f2212..b483ba65adb 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1199,16 +1199,26 @@ static void write_pack_file(void)\n \t\t\tdisplay_progress(progress_state, written);\n \t\t}\n \n-\t\t/*\n-\t\t * Did we write the wrong # entries in the header?\n-\t\t * If so, rewrite it like in fast-import\n-\t\t */\n \t\tif (pack_to_stdout) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n+\t\t\t/*\n+\t\t\t * We never fsync when writing to stdout since we may\n+\t\t\t * not be writing to an actual pack file. For instance,\n+\t\t\t * the upload-pack code passes a pipe here. Calling\n+\t\t\t * fsync on a pipe results in unnecessary\n+\t\t\t * synchronization with the reader on some platforms.\n+\t\t\t */\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n \t\t} else if (nr_written == nr_remaining) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t\t} else {\n-\t\t\tint fd = finalize_hashfile(f, hash, 0);\n+\t\t\t/*\n+\t\t\t * If we wrote the wrong number of entries in the\n+\t\t\t * header, rewrite it like in fast-import.\n+\t\t\t */\n+\n+\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n \t\t\t\t\t\t nr_written, hash, offset);\n \t\t\tclose(fd);\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..a2cf9dcbc8d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -53,9 +53,10 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \t\tunlink(state->pack_tmp_name);\n \t\tgoto clear_exit;\n \t} else if (state->nr_written == 1) {\n-\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\tfinalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t} else {\n-\t\tint fd = finalize_hashfile(state->f, hash, 0);\n+\t\tint fd = finalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n \t\t\t\t\t state->nr_written, hash,\n \t\t\t\t\t state->offset);\ndiff --git a/cache.h b/cache.h\nindex 37a32034b2f..b3cd7d928de 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -993,8 +993,38 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n-extern int use_fsync;\n+/*\n+ * These values are used to help identify parts of a repository to fsync.\n+ * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n+ * repository and so shouldn't be fsynced.\n+ */\n+enum fsync_component {\n+\tFSYNC_COMPONENT_NONE,\n+\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n+\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n+\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n+\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+};\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t      FSYNC_COMPONENT_PACK | \\\n+\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+\n+/*\n+ * A bitmask indicating which components of the repo should be fsynced.\n+ */\n+extern enum fsync_component fsync_components;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n@@ -1002,6 +1032,7 @@ enum fsync_method {\n };\n \n extern enum fsync_method fsync_method;\n+extern int use_fsync;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\n@@ -1757,6 +1788,12 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n \n+inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n+{\n+\tif (fsync_components & component)\n+\t\tfsync_or_die(fd, msg);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 265c010122e..64897f57d9f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1942,7 +1942,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t}\n \n \tclose_commit_graph(ctx->r->objects);\n-\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n+\tfinalize_hashfile(f, file_hash, FSYNC_COMPONENT_COMMIT_GRAPH,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n \tfree_chunkfile(cf);\n \n \tif (ctx->split) {\ndiff --git a/config.c b/config.c\nindex f67f545f839..224563c7b3e 100644\n--- a/config.c\n+++ b/config.c\n@@ -1213,6 +1213,73 @@ static int git_parse_maybe_bool_text(const char *value)\n \treturn -1;\n }\n \n+static const struct fsync_component_entry {\n+\tconst char *name;\n+\tenum fsync_component component_bits;\n+} fsync_component_table[] = {\n+\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n+\t{ \"pack\", FSYNC_COMPONENT_PACK },\n+\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n+\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n+};\n+\n+static enum fsync_component parse_fsync_components(const char *var, const char *string)\n+{\n+\tenum fsync_component output = 0;\n+\n+\tif (!strcmp(string, \"none\"))\n+\t\treturn output;\n+\n+\twhile (string) {\n+\t\tint i;\n+\t\tsize_t len;\n+\t\tconst char *ep;\n+\t\tint negated = 0;\n+\t\tint found = 0;\n+\n+\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n+\t\tep = strchrnul(string, ',');\n+\t\tlen = ep - string;\n+\n+\t\tif (*string == '-') {\n+\t\t\tnegated = 1;\n+\t\t\tstring++;\n+\t\t\tlen--;\n+\t\t\tif (!len)\n+\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n+\t\t}\n+\n+\t\tif (!len)\n+\t\t\tbreak;\n+\n+\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n+\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n+\n+\t\t\tif (strncmp(entry->name, string, len))\n+\t\t\t\tcontinue;\n+\n+\t\t\tfound = 1;\n+\t\t\tif (negated)\n+\t\t\t\toutput &= ~entry->component_bits;\n+\t\t\telse\n+\t\t\t\toutput |= entry->component_bits;\n+\t\t}\n+\n+\t\tif (!found) {\n+\t\t\tchar *component = xstrndup(string, len);\n+\t\t\twarning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n+\t\t\tfree(component);\n+\t\t}\n+\n+\t\tstring = ep;\n+\t}\n+\n+\treturn output;\n+}\n+\n int git_parse_maybe_bool(const char *value)\n {\n \tint v = git_parse_maybe_bool_text(value);\n@@ -1490,6 +1557,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsync\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tfsync_components = parse_fsync_components(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncmethod\")) {\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n@@ -1503,7 +1577,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n \t\treturn 0;\n \t}\n \ndiff --git a/csum-file.c b/csum-file.c\nindex 26e8a6df44e..59ef3398ca2 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -58,7 +58,8 @@ static void free_hashfile(struct hashfile *f)\n \tfree(f);\n }\n \n-int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n+int finalize_hashfile(struct hashfile *f, unsigned char *result,\n+\t\t      enum fsync_component component, unsigned int flags)\n {\n \tint fd;\n \n@@ -69,7 +70,7 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int fl\n \tif (flags & CSUM_HASH_IN_STREAM)\n \t\tflush(f, f->buffer, the_hash_algo->rawsz);\n \tif (flags & CSUM_FSYNC)\n-\t\tfsync_or_die(f->fd, f->name);\n+\t\tfsync_component_or_die(component, f->fd, f->name);\n \tif (flags & CSUM_CLOSE) {\n \t\tif (close(f->fd))\n \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\ndiff --git a/csum-file.h b/csum-file.h\nindex 291215b34eb..0d29f528fbc 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -1,6 +1,7 @@\n #ifndef CSUM_FILE_H\n #define CSUM_FILE_H\n \n+#include \"cache.h\"\n #include \"hash.h\"\n \n struct progress;\n@@ -38,7 +39,7 @@ int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n struct hashfile *hashfd(int fd, const char *name);\n struct hashfile *hashfd_check(const char *name);\n struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n-int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n+int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n void hashwrite(struct hashfile *, const void *, unsigned int);\n void hashflush(struct hashfile *f);\n void crc32_begin(struct hashfile *);\ndiff --git a/environment.c b/environment.c\nindex 3e3620d759f..378424b9af5 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,9 +42,9 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n int use_fsync = -1;\n enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n+enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/midx.c b/midx.c\nindex 837b46b2af5..882f91f7d57 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -1406,7 +1406,8 @@ static int write_midx_internal(const char *object_dir,\n \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n \twrite_chunkfile(cf, &ctx);\n \n-\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n \tfree_chunkfile(cf);\n \n \tif (flags & (MIDX_WRITE_REV_INDEX | MIDX_WRITE_BITMAP))\ndiff --git a/object-file.c b/object-file.c\nindex 8be57f48de7..ff29e678204 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1850,8 +1850,7 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n static void close_loose_object(int fd)\n {\n \tif (!the_repository->objects->odb->will_destroy) {\n-\t\tif (fsync_object_files)\n-\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n \t}\n \n \tif (close(fd) != 0)\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex 9c55c1531e1..c16e43d1669 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -719,7 +719,8 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \tif (options & BITMAP_OPT_HASH_CACHE)\n \t\twrite_hash_cache(f, index, index_nr);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \n \tif (adjust_shared_perm(tmp_file.buf))\n \t\tdie_errno(\"unable to make temporary bitmap file readable\");\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a346..51812cb1299 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -159,9 +159,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t}\n \n \thashwrite(f, sha1, the_hash_algo->rawsz);\n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n-\t\t\t\t    ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \treturn index_name;\n }\n \n@@ -281,8 +281,9 @@ const char *write_rev_file_order(const char *rev_name,\n \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n \t\tdie(_(\"failed to make %s readable\"), rev_name);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \n \treturn rev_name;\n }\n@@ -390,7 +391,7 @@ void fixup_pack_header_footer(int pack_fd,\n \t\tthe_hash_algo->final_fn(partial_pack_hash, &old_hash_ctx);\n \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n-\tfsync_or_die(pack_fd, pack_name);\n+\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n }\n \n char *index_pack_lockfile(int ip_out, int *is_well_formed)\ndiff --git a/read-cache.c b/read-cache.c\nindex cbe73f14e5e..a0de70195c8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -3081,7 +3081,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"447393","messageId":"f217dba77a19714668f352825e5c91ee24f46779.1643686425.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v4.git.1643686424.gitgitgadget@gmail.com","subject":"[PATCH v4 3/4] core.fsync: new option to harden the index","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-01T03:33:43Z","receivedAt":"2022-02-01T03:33:57Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the new ability for the user to harden\nthe index. In the event of a system crash, the index must be\ndurable for the user to actually find a file that has been added\nto the repo and then deleted from the working tree.\n\nWe use the presence of the COMMIT_LOCK flag and absence of the\nalternate_index_output as a proxy for determining whether we're\nupdating the persistent index of the repo or some temporary\nindex. We don't sync these temporary indexes.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  1 +\n cache.h                       |  4 +++-\n config.c                      |  1 +\n read-cache.c                  | 19 +++++++++++++------\n 4 files changed, 18 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 4f1747ec871..8e5b7a795ab 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -561,6 +561,7 @@ core.fsync::\n * `pack` hardens objects added to the repo in packfile form.\n * `pack-metadata` hardens packfile bitmaps and indexes.\n * `commit-graph` hardens the commit graph file.\n+* `index` hardens the index when it is modified.\n * `objects` is an aggregate option that includes `loose-objects`, `pack`,\n   `pack-metadata`, and `commit-graph`.\n * `default` is an aggregate option that is equivalent to `objects,-loose-object`\ndiff --git a/cache.h b/cache.h\nindex b3cd7d928de..9f3c1ec4c42 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1004,6 +1004,7 @@ enum fsync_component {\n \tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+\tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n #define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n@@ -1018,7 +1019,8 @@ enum fsync_component {\n #define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n \t\t\t      FSYNC_COMPONENT_PACK | \\\n \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n+\t\t\t      FSYNC_COMPONENT_INDEX)\n \n \n /*\ndiff --git a/config.c b/config.c\nindex 224563c7b3e..325644e3c2c 100644\n--- a/config.c\n+++ b/config.c\n@@ -1221,6 +1221,7 @@ static const struct fsync_component_entry {\n \t{ \"pack\", FSYNC_COMPONENT_PACK },\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"index\", FSYNC_COMPONENT_INDEX },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n \t{ \"all\", FSYNC_COMPONENTS_ALL },\ndiff --git a/read-cache.c b/read-cache.c\nindex a0de70195c8..eb02439ab4b 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2837,7 +2837,7 @@ static int record_ieot(void)\n  * rely on it.\n  */\n static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n-\t\t\t  int strip_extensions)\n+\t\t\t  int strip_extensions, unsigned flags)\n {\n \tuint64_t start = getnanotime();\n \tstruct hashfile *f;\n@@ -2851,6 +2851,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint csum_fsync_flag;\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n@@ -3081,7 +3082,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n+\tcsum_fsync_flag = 0;\n+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n+\t\tcsum_fsync_flag = CSUM_FSYNC;\n+\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n+\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n@@ -3136,7 +3143,7 @@ static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n \t */\n \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n-\tret = do_write_index(istate, lock->tempfile, 0);\n+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n \n@@ -3230,7 +3237,7 @@ static int clean_shared_index_files(const char *current_hex)\n }\n \n static int write_shared_index(struct index_state *istate,\n-\t\t\t      struct tempfile **temp)\n+\t\t\t      struct tempfile **temp, unsigned flags)\n {\n \tstruct split_index *si = istate->split_index;\n \tint ret, was_full = !istate->sparse_index;\n@@ -3240,7 +3247,7 @@ static int write_shared_index(struct index_state *istate,\n \n \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n-\tret = do_write_index(si->base, *temp, 1);\n+\tret = do_write_index(si->base, *temp, 1, flags);\n \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n \n@@ -3349,7 +3356,7 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n \t\t\tret = do_write_locked_index(istate, lock, flags);\n \t\t\tgoto out;\n \t\t}\n-\t\tret = write_shared_index(istate, &temp);\n+\t\tret = write_shared_index(istate, &temp, flags);\n \n \t\tsaved_errno = errno;\n \t\tif (is_tempfile_active(temp))\n-- \ngitgitgadget\n\n"},{"id":"447394","messageId":"5c22a41c1f3b6edc941e8bcd0d77df87f1087f85.1643686425.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v4.git.1643686424.gitgitgadget@gmail.com","subject":"[PATCH v4 4/4] core.fsync: add a `derived-metadata` aggregate option","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-01T03:33:44Z","receivedAt":"2022-02-01T03:33:58Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds an aggregate option that currently includes the\ncommit-graph file and pack metadata (indexes and bitmaps).\n\nThe user may want to exclude this set from durability since they can be\nrecomputed from other data if they wind up corrupt or missing.\n\nThis is split out from the other patches in the series since it is\nan optional nice-to-have that might be controversial.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 6 +++---\n cache.h                       | 7 ++++---\n config.c                      | 1 +\n 3 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 8e5b7a795ab..21092f3a4d1 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -562,9 +562,9 @@ core.fsync::\n * `pack-metadata` hardens packfile bitmaps and indexes.\n * `commit-graph` hardens the commit graph file.\n * `index` hardens the index when it is modified.\n-* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n-  `pack-metadata`, and `commit-graph`.\n-* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n+* `objects` is an aggregate option that includes `loose-objects` and `pack`.\n+* `derived-metadata` is an aggregate option that includes `pack-metadata` and `commit-graph`.\n+* `default` is an aggregate option that is equivalent to `objects,derived-metadata,-loose-object`\n * `all` is an aggregate option that syncs all individual components above.\n \n core.fsyncMethod::\ndiff --git a/cache.h b/cache.h\nindex 9f3c1ec4c42..3327cf6af0b 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1012,9 +1012,10 @@ enum fsync_component {\n \t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n \n #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n-\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n-\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\t\t\t\t  FSYNC_COMPONENT_PACK)\n+\n+#define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n \n #define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n \t\t\t      FSYNC_COMPONENT_PACK | \\\ndiff --git a/config.c b/config.c\nindex 325644e3c2c..64a8a4d7c2a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1223,6 +1223,7 @@ static const struct fsync_component_entry {\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n \t{ \"all\", FSYNC_COMPONENTS_ALL },\n };\n-- \ngitgitgadget\n"},{"id":"447482","messageId":"xmqqr18m8514.fsf@gitster.g","threadId":"57030","inReplyTo":"7a164ba95710b4231d07982fd27ec51022929b81.1643686425.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-02T00:51:19Z","receivedAt":"2022-02-02T00:51:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +core.fsync::\n> +\tA comma-separated list of parts of the repository which should be\n> +\thardened via the core.fsyncMethod when created or modified. You can\n> +\tdisable hardening of any component by prefixing it with a '-'. Later\n> +\titems take precedence over earlier ones in the list. For example,\n> +\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n> +\tmetadata.\" Items that are not hardened may be lost in the event of an\n> +\tunclean system shutdown.\n> ++\n> +* `none` disables fsync completely. This must be specified alone.\n> +* `loose-object` hardens objects added to the repo in loose-object form.\n> +* `pack` hardens objects added to the repo in packfile form.\n> +* `pack-metadata` hardens packfile bitmaps and indexes.\n> +* `commit-graph` hardens the commit graph file.\n> +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n> +  `pack-metadata`, and `commit-graph`.\n> +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n> +* `all` is an aggregate option that syncs all individual components above.\n\nI am not quite sure if this is way too complex (e.g. what does it\nmean that we do not care much about loose-object safety while we do\ncare about commit-graph files?) and at the same time it is too\nlimited (e.g. if it makes sense to say a class of items deserve more\nprotection than another class of items, don't we want to be able to\nsay \"class X is ultra-precious so use method A on them, while class\nY is mildly precious and use method B on them, everything else are\nnot that important and doing the default thing is just fine\").\n\nIf we wanted to allow the \"matrix\" kind of flexibility, I think the\nway to do so would be\n\n\tfsync.<class>.method = <value>\n\ne.g.\n\n\t[fsync \"default\"] method = none\n\t[fsync \"loose-object\"] method = fsync\n\t[fsync \"pack-metadata\"] method = writeout-only\n\nWhere do we expect users to take the core.fsync settings from?  Per\nrepository?  If it is from per user (i.e. $HOME/.gitconfig), do\npeople tend to share it across systems (not necessarily over NFS)\nwith the same contents?  If so, I am not sure if fsync.method that\nis way too close to the actual \"implementation\" is a good idea to\nbegin with.  From end-user's point of view, it may be easier to\nexpress \"class X is ultra-precious, and class Y and Z are mildly\nso\", with something like fsync.<class>.level = <how-precious> and\nlet the Git implementation on each platform choose the appropriate\nfsync method to protect the stuff at that precious-ness.\n\nThanks.\n\n\n\n\n\t\n"},{"id":"447483","messageId":"xmqqy22u6o3d.fsf@gitster.g","threadId":"57030","inReplyTo":"xmqqr18m8514.fsf@gitster.g","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-02T01:42:30Z","receivedAt":"2022-02-02T01:42:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> I am not quite sure if this is way too complex (e.g. what does it\n> mean that we do not care much about loose-object safety while we do\n> care about commit-graph files?) and at the same time it is too\n> limited (e.g. if it makes sense to say a class of items deserve more\n> protection than another class of items, don't we want to be able to\n> say \"class X is ultra-precious so use method A on them, while class\n> Y is mildly precious and use method B on them, everything else are\n> not that important and doing the default thing is just fine\").\n>\n> If we wanted to allow the \"matrix\" kind of flexibility,...\n\nTo continue with the thinking aloud...\n\nSometimes configuration flexibility is truly needed, but often it is\njust a sign of designer being lazy and not thinking it through as an\nend-user facing problem.  In other words, \"I am giving enough knobs\nto you, so it is up to you to express your policy in whatever way\nyou want with the knobs provided\" is a very irresponsible thing to\ntell end-users.\n\nAnd this one smells like the case of a lazy design.\n\nIt may be that it makes sense in some workflows to protect\ncommit-graph files less than object files and pack.idx files can be\ncorrupted as long as pack.pack files are adequately protected\nbecause the former can be recomputed from the latter, but in no\nworkflows, the reverse would be true.  Yet the design gives such\nneedless flexibility, which makes it hard for lay end-users to\nchoose the best combination and allows them to protect .idx files\nmore than .pack files by mistake, for example.\n\nI am wondering if the classification itself introduced by this step\nactually can form a natural and linear progression of safe-ness.  By\ndefault, we'd want _all_ classes of things to be equally safe, but\nat one level down, there is \"protect things that are not\nrecomputable, but recomputable things can be left to the system\"\nlevel, and there would be even riskier \"protect packs as it would\nhurt a _lot_ to lose them, but losing loose ones will typically lose\nonly the most recent work, and they are less valuable\" level.\n\nIf we, as the Git experts, spend extra brain cycles to come up with\nan easy to understand spectrum of performance vs durability\ntrade-off, end-users won't have to learn the full flexibility and\neasily take the advice from experts.  They just need to say what\nlevel of durability they want (or how much durability they can risk\nin exchange for an additional throughput), and leave the rest to us.\n\nOn the core.fsyncMethod side, the same suggestion applies.\n\nOnce we know the desired level of performance vs durability\ntrade-off from the user, we, as the impolementors, should know the\nbest method, for each class of items, to achieve that durability on\neach platform when writing it to the storage, without exposing the\nlow level details of the implementation that only the Git folks need\nto be aware of.\n\nSo, from the end-user UI perspective, I'd very much prefer if we can\njust come up with a single scalar variable, (say \"fsync.durability\"\nthat ranges from \"conservative\" to \"performance\") that lets our\nusers express the level of durability desired.  The combination of\ncore.fsyncMethod and core.fsync are one variable too many, and the\nlatter being a variable that takes a list of things as its value\nmakes it even worse to sell to the end users.\n\n\n"},{"id":"448251","messageId":"CANQDOdcXssVBAbUU_+_0GKEUey0c9tdC3Oq8Kky_kLdW6uLeYQ@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqr18m8514.fsf@gitster.g","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-02-11T20:38:02Z","receivedAt":"2022-02-11T20:38:23Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"Apologies in advance for the delayed reply.  I've finally been able to\nreturn to Git after an absence.\n\nOn Tue, Feb 1, 2022 at 4:51 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > +core.fsync::\n> > +     A comma-separated list of parts of the repository which should be\n> > +     hardened via the core.fsyncMethod when created or modified. You can\n> > +     disable hardening of any component by prefixing it with a '-'. Later\n> > +     items take precedence over earlier ones in the list. For example,\n> > +     `core.fsync=all,-pack-metadata` means \"harden everything except pack\n> > +     metadata.\" Items that are not hardened may be lost in the event of an\n> > +     unclean system shutdown.\n> > ++\n> > +* `none` disables fsync completely. This must be specified alone.\n> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> > +* `pack` hardens objects added to the repo in packfile form.\n> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> > +* `commit-graph` hardens the commit graph file.\n> > +* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n> > +  `pack-metadata`, and `commit-graph`.\n> > +* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n> > +* `all` is an aggregate option that syncs all individual components above.\n>\n> I am not quite sure if this is way too complex (e.g. what does it\n> mean that we do not care much about loose-object safety while we do\n> care about commit-graph files?) and at the same time it is too\n> limited (e.g. if it makes sense to say a class of items deserve more\n> protection than another class of items, don't we want to be able to\n> say \"class X is ultra-precious so use method A on them, while class\n> Y is mildly precious and use method B on them, everything else are\n> not that important and doing the default thing is just fine\").\n>\n> If we wanted to allow the \"matrix\" kind of flexibility, I think the\n> way to do so would be\n>\n>         fsync.<class>.method = <value>\n>\n> e.g.\n>\n>         [fsync \"default\"] method = none\n>         [fsync \"loose-object\"] method = fsync\n>         [fsync \"pack-metadata\"] method = writeout-only\n>\n\nI don't believe it makes sense to offer a full matrix of what to fsync\nand what method to use, since the method is a property of the\nfilesystem and OS the repo is running on, while the list of things to\nfsync is more a selection of what the user values. So if I'm hosting\non APFS on macOS or NTFS on Windows, I'd want to set the fsyncMethod\nto batch so that I can get good performance at the safety level I\nchoose.  If I'm working on my maintainer repo, I'd maybe not want to\nfsync anything, but I'd want to fsync everything when working on my\ndeveloper repo.\n\n> Where do we expect users to take the core.fsync settings from?  Per\n> repository?  If it is from per user (i.e. $HOME/.gitconfig), do\n> people tend to share it across systems (not necessarily over NFS)\n> with the same contents?  If so, I am not sure if fsync.method that\n> is way too close to the actual \"implementation\" is a good idea to\n> begin with.  From end-user's point of view, it may be easier to\n> express \"class X is ultra-precious, and class Y and Z are mildly\n> so\", with something like fsync.<class>.level = <how-precious> and\n> let the Git implementation on each platform choose the appropriate\n> fsync method to protect the stuff at that precious-ness.\n>\n\nI expect the vast majority of users to have whatever setting is baked\ninto their build of Git.  For the users that want to do something\ndifferent, I expect them to have core.fsyncMethod and core.fsync\nconfigured per-user for the majority of their repos. Some repos might\nhave custom settings that override the per-user settings: 1) Ephemeral\nrepos that don't contain unique data would probably want to set\ncore.fsync=none. 2) Repos hosting on NFS or on a different FS may have\na stricter core.fsyncmethod setting.\n\n(More more text to follow in reply to your next email).\n"},{"id":"448284","messageId":"CANQDOdfVg4e=nLLAynm261_R5z+rjZV3QgE8nLwGEmj1wQm_uA@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqy22u6o3d.fsf@gitster.g","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-02-11T21:18:09Z","receivedAt":"2022-02-11T21:18:24Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Feb 1, 2022 at 5:42 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > I am not quite sure if this is way too complex (e.g. what does it\n> > mean that we do not care much about loose-object safety while we do\n> > care about commit-graph files?) and at the same time it is too\n> > limited (e.g. if it makes sense to say a class of items deserve more\n> > protection than another class of items, don't we want to be able to\n> > say \"class X is ultra-precious so use method A on them, while class\n> > Y is mildly precious and use method B on them, everything else are\n> > not that important and doing the default thing is just fine\").\n> >\n> > If we wanted to allow the \"matrix\" kind of flexibility,...\n>\n> To continue with the thinking aloud...\n>\n> Sometimes configuration flexibility is truly needed, but often it is\n> just a sign of designer being lazy and not thinking it through as an\n> end-user facing problem.  In other words, \"I am giving enough knobs\n> to you, so it is up to you to express your policy in whatever way\n> you want with the knobs provided\" is a very irresponsible thing to\n> tell end-users.\n>\n> And this one smells like the case of a lazy design.\n>\n> It may be that it makes sense in some workflows to protect\n> commit-graph files less than object files and pack.idx files can be\n> corrupted as long as pack.pack files are adequately protected\n> because the former can be recomputed from the latter, but in no\n> workflows, the reverse would be true.  Yet the design gives such\n> needless flexibility, which makes it hard for lay end-users to\n> choose the best combination and allows them to protect .idx files\n> more than .pack files by mistake, for example.\n>\n> I am wondering if the classification itself introduced by this step\n> actually can form a natural and linear progression of safe-ness.  By\n> default, we'd want _all_ classes of things to be equally safe, but\n> at one level down, there is \"protect things that are not\n> recomputable, but recomputable things can be left to the system\"\n> level, and there would be even riskier \"protect packs as it would\n> hurt a _lot_ to lose them, but losing loose ones will typically lose\n> only the most recent work, and they are less valuable\" level.\n>\n> If we, as the Git experts, spend extra brain cycles to come up with\n> an easy to understand spectrum of performance vs durability\n> trade-off, end-users won't have to learn the full flexibility and\n> easily take the advice from experts.  They just need to say what\n> level of durability they want (or how much durability they can risk\n> in exchange for an additional throughput), and leave the rest to us.\n>\n> On the core.fsyncMethod side, the same suggestion applies.\n>\n> Once we know the desired level of performance vs durability\n> trade-off from the user, we, as the impolementors, should know the\n> best method, for each class of items, to achieve that durability on\n> each platform when writing it to the storage, without exposing the\n> low level details of the implementation that only the Git folks need\n> to be aware of.\n>\n> So, from the end-user UI perspective, I'd very much prefer if we can\n> just come up with a single scalar variable, (say \"fsync.durability\"\n> that ranges from \"conservative\" to \"performance\") that lets our\n> users express the level of durability desired.  The combination of\n> core.fsyncMethod and core.fsync are one variable too many, and the\n> latter being a variable that takes a list of things as its value\n> makes it even worse to sell to the end users.\n\nI see the value in simplifying the core.fsync configuration to a\nsingle scalar knob of preciousness. The main motivation for this more\ngranular scheme is that I didn't think the current configuration\nfollows a sensible principle. We should be fsyncing the loose objects,\nindex, refs, and config files in addition to what we're already\nsyncing today.  On macOS, we should be doing a full hardware flush if\nwe've said we want to fsync.  But you expressed the notion in [1] that\nwe don't want to degrade the performance of the vast majority of users\nwho are happy with the current \"unprincipled but mostly works\"\nconfiguration.  I agree with that sentiment, but it leads to a design\nwhere we can express the current configuration, which does not follow\na scalar hierarchy.\n\nThe aggregate core.fsync options are meant to provide a way for us to\nrecommend a sensible configuration to the user without having to get\ninto the intricacies of repo layout. Maybe we can define and document\naggregate options that make sense in terms of a scalar level of\npreciousness.\n\nOne reason to keep core.fsyncMethod separate from the core.fsync knob\nis that it's more a property of the system and FS the repo is on,\nrather than the user's value of the repo.  We could try to auto-detect\nsome known filesystems that might support batch mode using 'statfs',\nbut having a table like that in Git would go out of date over time.\n\nPlease let me know what you think about these justifications for the\ncurrent design.  I'd be happy to make a change if the current\nconstraint of \"keep the default config the same\" can be relaxed in\nsome way.  I'd also be happy to go back to some variation of\nexpressing 'core.fsyncObjectFiles = batch' and leaving the rest of\nfsync story alone.\n\nThanks,\nNeeraj\n\n[1] https://lore.kernel.org/git/xmqqtuilyfls.fsf@gitster.g/\n"},{"id":"448287","messageId":"xmqqczjt9hbz.fsf@gitster.g","threadId":"57030","inReplyTo":"CANQDOdfVg4e=nLLAynm261_R5z+rjZV3QgE8nLwGEmj1wQm_uA@mail.gmail.com","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-11T22:19:44Z","receivedAt":"2022-02-11T22:19:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> The main motivation for this more granular scheme is that I didn't\n> think the current configuration follows a sensible principle. We\n> should be fsyncing the loose objects, index, refs, and config\n> files in addition to what we're already syncing today.  On macOS,\n> we should be doing a full hardware flush if we've said we want to\n> fsync.\n\nIf the \"robustness vs performance\" trade-off is unevenly made in the\ncurrent code, then that is a very good problem to address first, and\nsuch a change is very much justified on its own.\n\nPerhaps \"this is not a primary work repository but is used only to\nfollow external site to build, hence no fsync is fine\" folks, who do\nnot have core.fsyncObjectFiles set to true, may appreciate if we\nstopped doing fsync for packs and other things.  As the Boolean\ncore.fsyncObjectFiles is the only end-user visible knob to express\nhow the end-users express the trade-off, a good first step would be\nto align other file accesses to the preference expressed by it, i.e.\nothers who say they want fsync in the current scheme would\nappreciate if we start fsync in places like ref-files backend.\n\nMaking the choice more granular, from Boolean \"yes/no\", to linear\nlevels, would also be a good idea.  Doing both at the same time may\nmake it harder to explain and justify, but as long as at the end, if\n\"very performant\" choice uniformly does not do any fsync while\n\"ultra durable\" choice makes a uniform effort across subsystems to\nmake sure bits hit the platter, it would be a very good idea to do\nthem.\n\nThanks.\n"},{"id":"448291","messageId":"CANQDOdcRM-GdxQ6iiV6pSBZifzpn+vJrBi0f88um9Rk4YJMFng@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqczjt9hbz.fsf@gitster.g","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-02-11T23:04:23Z","receivedAt":"2022-02-11T23:04:38Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Feb 11, 2022 at 2:19 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > The main motivation for this more granular scheme is that I didn't\n> > think the current configuration follows a sensible principle. We\n> > should be fsyncing the loose objects, index, refs, and config\n> > files in addition to what we're already syncing today.  On macOS,\n> > we should be doing a full hardware flush if we've said we want to\n> > fsync.\n>\n> If the \"robustness vs performance\" trade-off is unevenly made in the\n> current code, then that is a very good problem to address first, and\n> such a change is very much justified on its own.\n>\n> Perhaps \"this is not a primary work repository but is used only to\n> follow external site to build, hence no fsync is fine\" folks, who do\n> not have core.fsyncObjectFiles set to true, may appreciate if we\n> stopped doing fsync for packs and other things.  As the Boolean\n> core.fsyncObjectFiles is the only end-user visible knob to express\n> how the end-users express the trade-off, a good first step would be\n> to align other file accesses to the preference expressed by it, i.e.\n> others who say they want fsync in the current scheme would\n> appreciate if we start fsync in places like ref-files backend.\n>\n\nIn practice, almost all users have core.fsyncObjectFiles set to the\nplatform default, which is 'false' everywhere besides Windows.  So at\nminimum, we have to take default to mean that we maintain behavior no\nweaker than the current version of Git, otherwise users will start\nlosing their data. The advantage of introducing a new knob is that we\ndon't have to try to divine why the user set a particular value of\ncore.fsyncObjectFiles.\n\n> Making the choice more granular, from Boolean \"yes/no\", to linear\n> levels, would also be a good idea.  Doing both at the same time may\n> make it harder to explain and justify, but as long as at the end, if\n> \"very performant\" choice uniformly does not do any fsync while\n> \"ultra durable\" choice makes a uniform effort across subsystems to\n> make sure bits hit the platter, it would be a very good idea to do\n> them.\n>\n\nOne path to get to your suggestion from the current patch series would\nbe to remove the component-specific options and only provide aggregate\noptions.  Alternatively, we could just not document the\ncomponent-specific options and leave them available to be people who\nread source code. So if I rename the aggregate options in terms of\n'levels of durability', and only document those, would that be\nacceptable?\n\nThanks for the review!\n-Neeraj\n"},{"id":"448292","messageId":"xmqq35kp806v.fsf@gitster.g","threadId":"57030","inReplyTo":"CANQDOdcRM-GdxQ6iiV6pSBZifzpn+vJrBi0f88um9Rk4YJMFng@mail.gmail.com","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-11T23:15:20Z","receivedAt":"2022-02-11T23:15:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> In practice, almost all users have core.fsyncObjectFiles set to the\n> platform default, which is 'false' everywhere besides Windows.  So at\n> minimum, we have to take default to mean that we maintain behavior no\n> weaker than the current version of Git, otherwise users will start\n> losing their data.\n\nWould they?   If they think platform default is performant and safe\nenough for their use, as long as our adjustment is out outrageously\nmore dangerous or less performant, I do not think \"no weaker than\"\nis a strict requirement.  If we were overly conservative in some\nareas than the \"platform default\", making it less conservative in\nthose areas to match the looseness of other areas should be OK and\nvice versa.\n\n> One path to get to your suggestion from the current patch series would\n> be to remove the component-specific options and only provide aggregate\n> options.  Alternatively, we could just not document the\n> component-specific options and leave them available to be people who\n> read source code. So if I rename the aggregate options in terms of\n> 'levels of durability', and only document those, would that be\n> acceptable?\n\nIn any case, if others who reviewed the series in the past are happy\nwith the \"two knobs\" approach and are willing to jump in to help new\nusers who will be confused with one knob too many, I actually am OK\nwith the series that I called \"overly complex\".  Let me let them\nweigh in before I can answer that question.\n\nThanks.\n\n\n"},{"id":"448293","messageId":"010e01d81fa8$f22bc070$d6834150$@nexbridge.com","threadId":"57030","inReplyTo":"xmqq35kp806v.fsf@gitster.g","subject":"RE: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-02-12T00:39:02Z","receivedAt":"2022-02-12T00:39:12Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 11, 2022 6:15 PM, Junio C Hamano wrote:\n> Neeraj Singh <nksingh85@gmail.com> writes:\n> \n> > In practice, almost all users have core.fsyncObjectFiles set to the\n> > platform default, which is 'false' everywhere besides Windows.  So at\n> > minimum, we have to take default to mean that we maintain behavior no\n> > weaker than the current version of Git, otherwise users will start\n> > losing their data.\n> \n> Would they?   If they think platform default is performant and safe\n> enough for their use, as long as our adjustment is out outrageously more\n> dangerous or less performant, I do not think \"no weaker than\"\n> is a strict requirement.  If we were overly conservative in some areas than the\n> \"platform default\", making it less conservative in those areas to match the\n> looseness of other areas should be OK and vice versa.\n> \n> > One path to get to your suggestion from the current patch series would\n> > be to remove the component-specific options and only provide aggregate\n> > options.  Alternatively, we could just not document the\n> > component-specific options and leave them available to be people who\n> > read source code. So if I rename the aggregate options in terms of\n> > 'levels of durability', and only document those, would that be\n> > acceptable?\n> \n> In any case, if others who reviewed the series in the past are happy with the\n> \"two knobs\" approach and are willing to jump in to help new users who will be\n> confused with one knob too many, I actually am OK with the series that I called\n> \"overly complex\".  Let me let them weigh in before I can answer that question.\n\nOn behalf of those who are likely to set fsync to true in all cases, because a SIGSEGV or some other early abort will cause changes to be lost, I am not happy with excessive knobs, as we will have to ensure that the are all set to \"write this out to disk as quickly as soon as possible or else\". I end up having to teach large numbers of people about these settings and think that excessive controls in this area are overwhelmingly bad. This is not a \"windows vs. everything else\" situation. Not all platforms write buffered by default. fwrite and write behave different on NonStop (and other POSIXy things, and some variants are fully virtualized, so who knows what the hypervisor will do - it's bad enough that there is even an option to tell the hypervisor to keep things in memory until convenient to write even on RAID 0/1 emulation). Teaching people how to use the knobs correctly is going to be a challenge. For me, I'm always willing to sacrifice performance for reliability - 100% of the time. I'm sure this whole series is going to be problematic. I'm sorry but that's my position on it. The default position of all knobs must be to force the write and keeping it simple so that variations on knob settings give different results is not going to end well.\n\nSincerely,\nRandall\n\n"},{"id":"448349","messageId":"Ygn/GvLEjbCxN3Cc@ncase","threadId":"57030","inReplyTo":"xmqq35kp806v.fsf@gitster.g","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-02-14T07:04:58Z","receivedAt":"2022-02-14T07:05:11Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Feb 11, 2022 at 03:15:20PM -0800, Junio C Hamano wrote:\n> Neeraj Singh <nksingh85@gmail.com> writes:\n> \n> > In practice, almost all users have core.fsyncObjectFiles set to the\n> > platform default, which is 'false' everywhere besides Windows.  So at\n> > minimum, we have to take default to mean that we maintain behavior no\n> > weaker than the current version of Git, otherwise users will start\n> > losing their data.\n> \n> Would they?   If they think platform default is performant and safe\n> enough for their use, as long as our adjustment is out outrageously\n> more dangerous or less performant, I do not think \"no weaker than\"\n> is a strict requirement.  If we were overly conservative in some\n> areas than the \"platform default\", making it less conservative in\n> those areas to match the looseness of other areas should be OK and\n> vice versa.\n> \n> > One path to get to your suggestion from the current patch series would\n> > be to remove the component-specific options and only provide aggregate\n> > options.  Alternatively, we could just not document the\n> > component-specific options and leave them available to be people who\n> > read source code. So if I rename the aggregate options in terms of\n> > 'levels of durability', and only document those, would that be\n> > acceptable?\n> \n> In any case, if others who reviewed the series in the past are happy\n> with the \"two knobs\" approach and are willing to jump in to help new\n> users who will be confused with one knob too many, I actually am OK\n> with the series that I called \"overly complex\".  Let me let them\n> weigh in before I can answer that question.\n> \n> Thanks.\n\nI wonder whether it makes sense to distinguish client- and server-side\nrequirements. While we probably want to \"do the right thing\" on the\nclient-side by default so that we don't have to teach users that \"We may\nuse data in some cases on some systems, but not in other cases on other\nsystems by default.\" On the server-side folks who implement the Git\nhosting are typically a lot more knowledgeable with regards to how Git\nbehaves, and it can be expected of them to dig a lot deeper than we can\nand should reasonably expect from a user.\n\nOne point I really care about and that I think is extremely important in\nthe context of both client and server is data consistency. While it is\nbad to lose data, the reality is that it can simply happen when systems\nhit exceptional cases like a hard-reset. But what we really must not let\nhappen is that as a result, a repository gets corrupted because of this\nhard reset. In the context of client-side repositories a user will be\nleft to wonder how to fix the repository, while on the server-side we\nneed to go in and somehow repair the repository fast or otherwise users\nwill come to us and complain their repository stopped working.\n\nOur current defaults can end up with repository corruption though. We\ndon't sync object files before renaming them into place by default, and\nfurthermore we don't even give a knob to fsync loose references before\nrenaming them. On gitlab.com we enable \"core.fsyncObjectFiles\" and thus\nto the best of my knowledge never hit a corrupted ODB until now. But\nwhat we do regularly hit is corrupted references because there is no\nknob to fsync them.\n\nFurthermore, I really think that the advice in git-config(1) with\nregards to loose syncing loose objects that \"This is a total waste of\ntime and effort\" is wrong. Filesystem developers have repeatedly stated\nthat for proper atomicity guarantees, a file must be synced to disk\nbefore renaming it into place [1][2][3]. So from the perspective of the\npeople who write Linux-based filesystems our application is \"broken\"\nright now, even though there are mechanisms in place which try to work\naround our brokenness by using heuristics to detect what we're doing.\n\nTo summarize my take: while the degree of durability may be something\nthat's up for discussions, I think that the current defaults for\natomicity are bad for users because they can and do lead to repository\ncorruption.\n\nPatrick\n\n[1]: https://thunk.org/tytso/blog/2009/03/15/dont-fear-the-fsync/\n[2]: https://btrfs.wiki.kernel.org/index.php/FAQ (What are the crash guarantees of overwrite-by-rename)\n[3]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/Documentation/admin-guide/ext4.rst (see auto_da_alloc)\n"},{"id":"448366","messageId":"xmqqh7914bbo.fsf@gitster.g","threadId":"57030","inReplyTo":"Ygn/GvLEjbCxN3Cc@ncase","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-14T17:17:31Z","receivedAt":"2022-02-14T17:17:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> To summarize my take: while the degree of durability may be something\n> that's up for discussions, I think that the current defaults for\n> atomicity are bad for users because they can and do lead to repository\n> corruption.\n\nGood summary.\n\nIf the user cares about fsynching loose object files in the right\nway, we shouldn't leave loose ref files not following the safe\nsafety level, regardless of how this new core.fsync knobs would look\nlike.\n\nI think we three are in agreement on that.\n"},{"id":"450869","messageId":"YiiuqK/tCnQOXrSV@ncase","threadId":"57030","inReplyTo":"xmqqh7914bbo.fsf@gitster.g","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-09T13:42:00Z","receivedAt":"2022-03-09T13:42:17Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Mon, Feb 14, 2022 at 09:17:31AM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > To summarize my take: while the degree of durability may be something\n> > that's up for discussions, I think that the current defaults for\n> > atomicity are bad for users because they can and do lead to repository\n> > corruption.\n> \n> Good summary.\n> \n> If the user cares about fsynching loose object files in the right\n> way, we shouldn't leave loose ref files not following the safe\n> safety level, regardless of how this new core.fsync knobs would look\n> like.\n> \n> I think we three are in agreement on that.\n\nIs there anything I can specifically do to help out with this topic? We\nhave again hit data loss in production because we don't sync loose refs\nto disk before renaming them into place, so I'd really love to sort out\nthis issue somehow so that I can revive my patch series which fixes the\nknown repository corruption [1].\n\nAlternatively, can we maybe find a way forward with applying a version\nof my patch series without first settling the bigger question of how we\nwant the overall design to look like? In my opinion repository\ncorruption is a severe bug that needs to be fixed, and it doesn't feel\nsensible to block such a fix over a discussion that potentially will\ntake a long time to settle.\n\nPatrick\n\n[1]: http://public-inbox.org/git/cover.1636544377.git.ps@pks.im/\n"},{"id":"450905","messageId":"220309.867d93lztw.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"YiiuqK/tCnQOXrSV@ncase","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-09T18:50:15Z","receivedAt":"2022-03-09T18:56:33Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Mar 09 2022, Patrick Steinhardt wrote:\n\n> [[PGP Signed Part:Undecided]]\n> On Mon, Feb 14, 2022 at 09:17:31AM -0800, Junio C Hamano wrote:\n>> Patrick Steinhardt <ps@pks.im> writes:\n>> \n>> > To summarize my take: while the degree of durability may be something\n>> > that's up for discussions, I think that the current defaults for\n>> > atomicity are bad for users because they can and do lead to repository\n>> > corruption.\n>> \n>> Good summary.\n>> \n>> If the user cares about fsynching loose object files in the right\n>> way, we shouldn't leave loose ref files not following the safe\n>> safety level, regardless of how this new core.fsync knobs would look\n>> like.\n>> \n>> I think we three are in agreement on that.\n>\n> Is there anything I can specifically do to help out with this topic? We\n> have again hit data loss in production because we don't sync loose refs\n> to disk before renaming them into place, so I'd really love to sort out\n> this issue somehow so that I can revive my patch series which fixes the\n> known repository corruption [1].\n>\n> Alternatively, can we maybe find a way forward with applying a version\n> of my patch series without first settling the bigger question of how we\n> want the overall design to look like? In my opinion repository\n> corruption is a severe bug that needs to be fixed, and it doesn't feel\n> sensible to block such a fix over a discussion that potentially will\n> take a long time to settle.\n>\n> Patrick\n>\n> [1]: http://public-inbox.org/git/cover.1636544377.git.ps@pks.im/\n\nI share that view. I was wondering how this topic fizzled out the other\nday, but then promptly forgot about it.\n\nI think the best thing at this point (hint hint!) would be for someone\nin the know to (re-)submit the various patches appropriate to move this\nforward. Whether that's just this series, part of it, or some/both of\nthose + patches from you and Eric and this point I don't know/remember.\n\nBut just to be explicitly clear, as probably the person most responsible\nfor pushing this towards the \"bigger question of [...] overall\ndesign\".\n\nI just wanted to facilitate a discussion that would result in the\nvarious stakeholders who wanted to add some fsync-related config coming\nup with something that's mutually compatible, and I think the design\nfrom Neeraj in this series fits that purpose, is Good Enough etc.\n\nI.e. the actually important and IMO blockers were all resolved, e.g. not\nhaving an fsync configuration that older git versions would needlessly\ndie on, and not painting ourselves into a corner where\ne.g. core.fsync=false or something was squatted on by something other\nthan a \"no fsync, whatsoever\" etc.\n\n(But I haven't looked at it again just now, so...)\n\nAnyway, just trying to be explicit that to whatever extent this was held\nup by questions/comments of mine I'm very happy to see this go forward.\nAs you (basically) say we shouldn't lose sight of ongoing data loss in\nthis area because of some config bikeshedding :)\n"},{"id":"450924","messageId":"xmqqpmmuki68.fsf@gitster.g","threadId":"57030","inReplyTo":"YiiuqK/tCnQOXrSV@ncase","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-09T20:03:11Z","receivedAt":"2022-03-09T20:03:16Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> If the user cares about fsynching loose object files in the right\n>> way, we shouldn't leave loose ref files not following the safe\n>> safety level, regardless of how this new core.fsync knobs would look\n>> like.\n>> \n>> I think we three are in agreement on that.\n>\n> Is there anything I can specifically do to help out with this topic? We\n> have again hit data loss in production because we don't sync loose refs\n> to disk before renaming them into place, so I'd really love to sort out\n> this issue somehow so that I can revive my patch series which fixes the\n> known repository corruption [1].\n\nHow about doing a series to unconditionally sync loose ref creation\nand modification?\n\nAlternatively, we could link it to the existing configuration to\ncontrol synching of object files.\n\nI do not think core.fsyncObjectFiles having \"object\" in its name is\na good reason not to think those who set it to true only care about\nthe loose object files and nothing else.  It is more sensible to\nconsider that those who set it to true cares about the repository\nintegrity more than those who set it to false, I would think.\n\nBut that (i.e. doing it conditionally and choose which knob to use)\nis one extra thing that needs justification, so starting from\nunconditional fsync_or_die() may be the best way to ease it in.\n\n"},{"id":"450926","messageId":"CANQDOdfHYnKKFfQ6ptPcLacx=HX5cnXGRuL55Ttfr+zKCSBQJg@mail.gmail.com","threadId":"57030","inReplyTo":"YiiuqK/tCnQOXrSV@ncase","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-09T20:05:09Z","receivedAt":"2022-03-09T20:05:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 9, 2022 at 5:42 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Mon, Feb 14, 2022 at 09:17:31AM -0800, Junio C Hamano wrote:\n> > Patrick Steinhardt <ps@pks.im> writes:\n> >\n> > > To summarize my take: while the degree of durability may be something\n> > > that's up for discussions, I think that the current defaults for\n> > > atomicity are bad for users because they can and do lead to repository\n> > > corruption.\n> >\n> > Good summary.\n> >\n> > If the user cares about fsynching loose object files in the right\n> > way, we shouldn't leave loose ref files not following the safe\n> > safety level, regardless of how this new core.fsync knobs would look\n> > like.\n> >\n> > I think we three are in agreement on that.\n>\n> Is there anything I can specifically do to help out with this topic? We\n> have again hit data loss in production because we don't sync loose refs\n> to disk before renaming them into place, so I'd really love to sort out\n> this issue somehow so that I can revive my patch series which fixes the\n> known repository corruption [1].\n>\n> Alternatively, can we maybe find a way forward with applying a version\n> of my patch series without first settling the bigger question of how we\n> want the overall design to look like? In my opinion repository\n> corruption is a severe bug that needs to be fixed, and it doesn't feel\n> sensible to block such a fix over a discussion that potentially will\n> take a long time to settle.\n>\n> Patrick\n>\n> [1]: http://public-inbox.org/git/cover.1636544377.git.ps@pks.im/\n\nHi Patrick,\nThanks for reviving this discussion.  I've updated the PR on\nGitGitGadget with a rebase\nonto the current 'main' branch and some minor build fixes. I've also\nrevamped the aggregate options\nand documentation to be more inline with Junio's suggestion of having\n'levels of safety' that we steer\nthe user towards. I'm still keeping the detailed options, but\nhopefully the guidance is clear enough to\navoid confusion.\n\nI'd be happy to make any point fixes as necessary to get that branch\ninto proper shape for\nupstream, if we've gotten to the point where we don't want to change\nthe fundamental design.\n\nI agree with Patrick that the detailed knobs are primarily for use by\nhosters like GitLab and GitHub.\n\nPlease expect a v5 today.\n\nThanks,\nNeeraj\n"},{"id":"450948","messageId":"685b1db888079c83573cfd984ae64f46284544af.1646866998.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH v5 1/5] wrapper: move inclusion of CSPRNG headers the wrapper.c file","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-09T23:03:14Z","receivedAt":"2022-03-09T23:03:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIncluding NTSecAPI.h in git-compat-util.h causes build errors in any\nother file that includes winternl.h. That file was included in order to\nget access to the RtlGenRandom cryptographically secure PRNG. This\nchange scopes the inclusion of all PRNG headers to just the wrapper.c\nfile, which is the only place it is really needed.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/winansi.c  |  5 -----\n git-compat-util.h | 12 ------------\n wrapper.c         | 14 ++++++++++++++\n 3 files changed, 14 insertions(+), 17 deletions(-)\n\ndiff --git a/compat/winansi.c b/compat/winansi.c\nindex 936a80a5f00..3abe8dd5a27 100644\n--- a/compat/winansi.c\n+++ b/compat/winansi.c\n@@ -4,11 +4,6 @@\n \n #undef NOGDI\n \n-/*\n- * Including the appropriate header file for RtlGenRandom causes MSVC to see a\n- * redefinition of types in an incompatible way when including headers below.\n- */\n-#undef HAVE_RTLGENRANDOM\n #include \"../git-compat-util.h\"\n #include <wingdi.h>\n #include <winreg.h>\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 876907b9df4..a25ebb822ee 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -197,12 +197,6 @@\n #endif\n #include <windows.h>\n #define GIT_WINDOWS_NATIVE\n-#ifdef HAVE_RTLGENRANDOM\n-/* This is required to get access to RtlGenRandom. */\n-#define SystemFunction036 NTAPI SystemFunction036\n-#include <NTSecAPI.h>\n-#undef SystemFunction036\n-#endif\n #endif\n \n #include <unistd.h>\n@@ -273,12 +267,6 @@\n #else\n #include <stdint.h>\n #endif\n-#ifdef HAVE_ARC4RANDOM_LIBBSD\n-#include <bsd/stdlib.h>\n-#endif\n-#ifdef HAVE_GETRANDOM\n-#include <sys/random.h>\n-#endif\n #ifdef NO_INTPTR_T\n /*\n  * On I16LP32, ILP32 and LP64 \"long\" is the safe bet, however\ndiff --git a/wrapper.c b/wrapper.c\nindex 3258cdb171f..2a1aade473b 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -4,6 +4,20 @@\n #include \"cache.h\"\n #include \"config.h\"\n \n+#ifdef HAVE_RTLGENRANDOM\n+/* This is required to get access to RtlGenRandom. */\n+#define SystemFunction036 NTAPI SystemFunction036\n+#include <NTSecAPI.h>\n+#undef SystemFunction036\n+#endif\n+\n+#ifdef HAVE_ARC4RANDOM_LIBBSD\n+#include <bsd/stdlib.h>\n+#endif\n+#ifdef HAVE_GETRANDOM\n+#include <sys/random.h>\n+#endif\n+\n static int memory_limit_check(size_t size, int gentle)\n {\n \tstatic size_t limit = 0;\n-- \ngitgitgadget\n\n"},{"id":"450949","messageId":"da8cfc10bb4bfa473619d8d737c3aa160643ccf7.1646866998.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH v5 2/5] core.fsyncmethod: add writeout-only mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-09T23:03:15Z","receivedAt":"2022-03-09T23:03:30Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsyncMethod` configuration\nknob, which can currently be set to `fsync` or `writeout-only`.\n\nThe new writeout-only mode attempts to tell the operating system to\nflush its in-memory page cache to the storage hardware without issuing a\nCACHE_FLUSH command to the storage controller.\n\nWriteout-only fsync is significantly faster than a vanilla fsync on\ncommon hardware, since data is written to a disk-side cache rather than\nall the way to a durable medium. Later changes in this patch series will\ntake advantage of this primitive to implement batching of hardware\nflushes.\n\nWhen git_fsync is called with FSYNC_WRITEOUT_ONLY, it may fail and the\ncaller is expected to do an ordinary fsync as needed.\n\nOn Apple platforms, the fsync system call does not issue a CACHE_FLUSH\ndirective to the storage controller. This change updates fsync to do\nfcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\nwith existing behavior on Apple platforms by setting the default value\nof the new core.fsyncMethod option.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt       |  9 ++++\n Makefile                            |  6 +++\n cache.h                             |  7 ++++\n compat/mingw.h                      |  3 ++\n compat/win32/flush.c                | 28 +++++++++++++\n config.c                            | 12 ++++++\n config.mak.uname                    |  3 ++\n configure.ac                        |  8 ++++\n contrib/buildsystems/CMakeLists.txt | 16 ++++++--\n environment.c                       |  1 +\n git-compat-util.h                   | 24 +++++++++++\n wrapper.c                           | 64 +++++++++++++++++++++++++++++\n write-or-die.c                      | 11 +++--\n 13 files changed, 184 insertions(+), 8 deletions(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..dbb134f7136 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,15 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsyncMethod::\n+\tA value indicating the strategy Git will use to harden repository data\n+\tusing fsync and related primitives.\n++\n+* `fsync` uses the fsync() system call or platform equivalents.\n+* `writeout-only` issues pagecache writeback requests, but depending on the\n+  filesystem and storage hardware, data added to the repository may not be\n+  durable in the event of a system crash. This is the default mode on macOS.\n+\n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n +\ndiff --git a/Makefile b/Makefile\nindex 6f0b4b775fe..17fd9b023a4 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -411,6 +411,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1897,6 +1899,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/cache.h b/cache.h\nindex 04d4d2db25c..82f0194a3dd 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -995,6 +995,13 @@ extern char *git_replace_ref_base;\n \n extern int fsync_object_files;\n extern int use_fsync;\n+\n+enum fsync_method {\n+\tFSYNC_METHOD_FSYNC,\n+\tFSYNC_METHOD_WRITEOUT_ONLY\n+};\n+\n+extern enum fsync_method fsync_method;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..291f90ea940\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NTAPI, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.c b/config.c\nindex 383b1a4885b..f3ff80b01c9 100644\n--- a/config.c\n+++ b/config.c\n@@ -1600,6 +1600,18 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsyncmethod\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcmp(value, \"fsync\"))\n+\t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n+\t\telse if (!strcmp(value, \"writeout-only\"))\n+\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse\n+\t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n+\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 4352ea39e9b..404fff5dd04 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -57,6 +57,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tBASIC_CFLAGS += -DHAVE_SYSINFO\n@@ -463,6 +464,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -640,6 +642,7 @@ ifeq ($(uname_S),MINGW)\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/configure.ac b/configure.ac\nindex 5ee25ec95c8..6bd6bef1c44 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1082,6 +1082,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex e44232f85d3..3a9e6241660 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,10 +261,18 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET HAVE_RTLGENRANDOM)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n-\t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n-\t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n-\t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\n+\tlist(APPEND compat_SOURCES\n+\t\tcompat/mingw.c\n+\t\tcompat/winansi.c\n+\t\tcompat/win32/flush.c\n+\t\tcompat/win32/path-utils.c\n+\t\tcompat/win32/pthread.c\n+\t\tcompat/win32mmap.c\n+\t\tcompat/win32/syslog.c\n+\t\tcompat/win32/trace2_win32_process_info.c\n+\t\tcompat/win32/dirent.c\n+\t\tcompat/nedmalloc/nedmalloc.c\n+\t\tcompat/strdup.c)\n \tset(NO_UNIX_SOCKETS 1)\n \n elseif(CMAKE_SYSTEM_NAME STREQUAL \"Linux\")\ndiff --git a/environment.c b/environment.c\nindex fd0501e77a5..3e3620d759f 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -44,6 +44,7 @@ int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n int fsync_object_files;\n int use_fsync = -1;\n+enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex a25ebb822ee..41b97027457 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1265,6 +1265,30 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+#ifdef __APPLE__\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n+#else\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n+#endif\n+\n+enum fsync_action {\n+\tFSYNC_WRITEOUT_ONLY,\n+\tFSYNC_HARDWARE_FLUSH\n+};\n+\n+/*\n+ * Issues an fsync against the specified file according to the specified mode.\n+ *\n+ * FSYNC_WRITEOUT_ONLY attempts to use interfaces available on some operating\n+ * systems to flush the OS cache without issuing a flush command to the storage\n+ * controller. If those interfaces are unavailable, the function fails with\n+ * ENOSYS.\n+ *\n+ * FSYNC_HARDWARE_FLUSH does an OS writeout and hardware flush to ensure that\n+ * changes are durable. It is not expected to fail.\n+ */\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/wrapper.c b/wrapper.c\nindex 2a1aade473b..2c44f71da42 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -553,6 +553,70 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+/*\n+ * Some platforms return EINTR from fsync. Since fsync is invoked in some\n+ * cases by a wrapper that dies on failure, do not expose EINTR to callers.\n+ */\n+static int fsync_loop(int fd)\n+{\n+\tint err;\n+\n+\tdo {\n+\t\terr = fsync(fd);\n+\t} while (err < 0 && errno == EINTR);\n+\treturn err;\n+}\n+\n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * On macOS, fsync just causes filesystem cache writeback but\n+\t\t * does not flush hardware caches.\n+\t\t */\n+\t\treturn fsync_loop(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to\n+\t\t * issue a writeback without a hardware flush. An offset of\n+\t\t * 0 and size of 0 indicates writeout of the entire file and the\n+\t\t * wait flags ensure that all dirty data is written to the disk\n+\t\t * (potentially in a disk-side cache) before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\t\t/*\n+\t\t * On macOS, a special fcntl is required to really flush the\n+\t\t * caches within the storage controller. As of this writing,\n+\t\t * this is a very expensive operation on Apple SSDs.\n+\t\t */\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync_loop(fd);\n+#endif\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex a3d5784cec9..9faa5f9f563 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -62,10 +62,13 @@ void fsync_or_die(int fd, const char *msg)\n \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n \tif (!use_fsync)\n \t\treturn;\n-\twhile (fsync(fd) < 0) {\n-\t\tif (errno != EINTR)\n-\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n-\t}\n+\n+\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n+\t\treturn;\n+\n+\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n+\t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n void write_or_die(int fd, const void *buf, size_t count)\n-- \ngitgitgadget\n\n"},{"id":"450951","messageId":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v4.git.1643686424.gitgitgadget@gmail.com","subject":"[PATCH v5 0/5] A design for future-proofing fsync() configuration","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-09T23:03:13Z","receivedAt":"2022-03-09T23:03:34Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"This is an implementation of an extensible configuration mechanism for\nfsyncing persistent components of a repo.\n\nThe main goals are to separate the \"what\" to sync from the \"how\". There are\nnow two settings: core.fsync - Control the 'what', including the index.\ncore.fsyncMethod - Control the 'how'. Currently we support writeout-only and\nfull fsync.\n\nSyncing of refs can be layered on top of core.fsync. And batch mode will be\nlayered on core.fsyncMethod. Once this series reaches 'seen', I'll submit\nns/batched-fsync to introduce batch mode. Please see\nhttps://github.com/gitgitgadget/git/pull/1134.\n\ncore.fsyncObjectfiles is removed and we will issue a deprecation warning if\nit's seen.\n\nI'd like to get agreement on this direction before submitting batch mode to\nthe list. The batch mode series is available to view at\n\nPlease see [1], [2], and [3] for discussions that led to this series.\n\nAfter this change, new persistent data files added to the repo will need to\nbe added to the fsync_component enum and documented in the\nDocumentation/config/core.txt text.\n\nV5 changes:\n\n * Rebase onto main at c2162907e9\n * Add a patch to move CSPRNG platform includes to wrapper.c. This avoids\n   build errors in compat/win32/flush.c and other files.\n * Move the documentation and aggregate options to the final patch in the\n   series.\n * Define new aggregate options and guidance in line with Junio's suggestion\n   to present the user with 'levels of safety' rather than a morass of\n   detailed options.\n\nV4 changes:\n\n * Rebase onto master at b23dac905bd.\n * Add a comment to write_pack_file indicating why we don't fsync when\n   writing to stdout.\n * I kept the configuration schema as-is rather than switching to\n   multi-value. The thinking here is that a stateless last-one-wins config\n   schema (comma separated) will make it easier to achieve some holistic\n   self-consistent fsync configuration for a particular repo.\n\nV3 changes:\n\n * Remove relative path from git-compat-util.h include [4].\n * Updated newly added warning texts to have more context for localization\n   [4].\n * Fixed tab spacing in enum fsync_action\n * Moved the fsync looping out to a helper and do it consistently. [4]\n * Changed commit description to use camelCase for config names. [5]\n * Add an optional fourth patch with derived-metadata so that the user can\n   exclude a forward-compatible set of things that should be recomputable\n   given existing data.\n\nV2 changes:\n\n * Updated the documentation for core.fsyncmethod to be less certain.\n   writeout-only probably does not do the right thing on Linux.\n * Split out the core.fsync=index change into its own commit.\n * Rename REPO_COMPONENT to FSYNC_COMPONENT. This is really specific to\n   fsyncing, so the name should reflect that.\n * Re-add missing Makefile change for SYNC_FILE_RANGE.\n * Tested writeout-only mode, index syncing, and general config settings.\n\n[1] https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n[2]\nhttps://lore.kernel.org/git/dd65718814011eb93ccc4428f9882e0f025224a6.1636029491.git.ps@pks.im/\n[3]\nhttps://lore.kernel.org/git/pull.1076.git.git.1629856292.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/CANQDOdf8C4-haK9=Q_J4Cid8bQALnmGDm=SvatRbaVf+tkzqLw@mail.gmail.com/\n[5] https://lore.kernel.org/git/211207.861r2opplg.gmgdl@evledraar.gmail.com/\n\nNeeraj Singh (5):\n  wrapper: move inclusion of CSPRNG headers the wrapper.c file\n  core.fsyncmethod: add writeout-only mode\n  core.fsync: introduce granular fsync control\n  core.fsync: new option to harden the index\n  core.fsync: documentation and user-friendly aggregate options\n\n Documentation/config/core.txt       | 51 +++++++++++++---\n Makefile                            |  6 ++\n builtin/fast-import.c               |  2 +-\n builtin/index-pack.c                |  4 +-\n builtin/pack-objects.c              | 24 +++++---\n bulk-checkin.c                      |  5 +-\n cache.h                             | 53 ++++++++++++++++-\n commit-graph.c                      |  3 +-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n compat/winansi.c                    |  5 --\n config.c                            | 92 ++++++++++++++++++++++++++++-\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt | 16 +++--\n csum-file.c                         |  5 +-\n csum-file.h                         |  3 +-\n environment.c                       |  3 +-\n git-compat-util.h                   | 36 +++++++----\n midx.c                              |  3 +-\n object-file.c                       |  3 +-\n pack-bitmap-write.c                 |  3 +-\n pack-write.c                        | 13 ++--\n read-cache.c                        | 19 ++++--\n wrapper.c                           | 78 ++++++++++++++++++++++++\n write-or-die.c                      | 11 ++--\n 26 files changed, 413 insertions(+), 67 deletions(-)\n create mode 100644 compat/win32/flush.c\n\n\nbase-commit: c2162907e9aa884bdb70208389cb99b181620d51\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1093%2Fneerajsi-msft%2Fns%2Fcore-fsync-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1093/neerajsi-msft/ns/core-fsync-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/1093\n\nRange-diff vs v4:\n\n -:  ----------- > 1:  685b1db8880 wrapper: move inclusion of CSPRNG headers the wrapper.c file\n 1:  51a218d100d ! 2:  da8cfc10bb4 core.fsyncmethod: add writeout-only mode\n     @@ contrib/buildsystems/CMakeLists.txt\n      @@ contrib/buildsystems/CMakeLists.txt: if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n       \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n       \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n     - \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n     + \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET HAVE_RTLGENRANDOM)\n      -\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n     -+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n     -+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n     - \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n     - \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n     - \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\n     +-\t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n     +-\t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n     +-\t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\n     ++\tlist(APPEND compat_SOURCES\n     ++\t\tcompat/mingw.c\n     ++\t\tcompat/winansi.c\n     ++\t\tcompat/win32/flush.c\n     ++\t\tcompat/win32/path-utils.c\n     ++\t\tcompat/win32/pthread.c\n     ++\t\tcompat/win32mmap.c\n     ++\t\tcompat/win32/syslog.c\n     ++\t\tcompat/win32/trace2_win32_process_info.c\n     ++\t\tcompat/win32/dirent.c\n     ++\t\tcompat/nedmalloc/nedmalloc.c\n     ++\t\tcompat/strdup.c)\n     + \tset(NO_UNIX_SOCKETS 1)\n     + \n     + elseif(CMAKE_SYSTEM_NAME STREQUAL \"Linux\")\n      \n       ## environment.c ##\n      @@ environment.c: int zlib_compression_level = Z_BEST_SPEED;\n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n      +\n      +#ifdef __APPLE__\n      +\t\t/*\n     -+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n     -+\t\t * flush hardware caches.\n     ++\t\t * On macOS, fsync just causes filesystem cache writeback but\n     ++\t\t * does not flush hardware caches.\n      +\t\t */\n      +\t\treturn fsync_loop(fd);\n      +#endif\n      +\n      +#ifdef HAVE_SYNC_FILE_RANGE\n      +\t\t/*\n     -+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n     -+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n     -+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n     -+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n     -+\t\t * before we continue.\n     ++\t\t * On linux 2.6.17 and above, sync_file_range is the way to\n     ++\t\t * issue a writeback without a hardware flush. An offset of\n     ++\t\t * 0 and size of 0 indicates writeout of the entire file and the\n     ++\t\t * wait flags ensure that all dirty data is written to the disk\n     ++\t\t * (potentially in a disk-side cache) before we continue.\n      +\t\t */\n      +\n      +\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n      +\n      +\tcase FSYNC_HARDWARE_FLUSH:\n      +\t\t/*\n     -+\t\t * On some platforms fsync may return EINTR. Try again in this\n     -+\t\t * case, since callers asking for a hardware flush may die if\n     -+\t\t * this function returns an error.\n     ++\t\t * On macOS, a special fcntl is required to really flush the\n     ++\t\t * caches within the storage controller. As of this writing,\n     ++\t\t * this is a very expensive operation on Apple SSDs.\n      +\t\t */\n      +#ifdef __APPLE__\n      +\t\treturn fcntl(fd, F_FULLFSYNC);\n 2:  7a164ba9571 ! 3:  e31886717b4 core.fsync: introduce granular fsync control\n     @@ Commit message\n          syncable components:\n          * We issue a warning rather than an error for unrecognized\n            components, so new configs can be used with old Git versions.\n     -    * We support negation, so users can choose one of the default\n     -      aggregate options and then remove components that they don't\n     -      want. The user would then harden any new components added in\n     -      a Git version update.\n     +    * We support negation, so users can choose one of the aggregate\n     +      options and then remove components that they don't want.\n     +      Aggregate options are defined in a later patch in this series.\n      \n          This also supports the common request of doing absolutely no\n          fysncing with the `core.fsync=none` value, which is expected\n          to make the test suite faster.\n      \n     +    Complete documentation for the new setting is included in a later patch\n     +    in the series so that it can be reviewed in final form.\n     +\n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Documentation/config/core.txt ##\n     -@@ Documentation/config/core.txt: core.whitespace::\n     -   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n     -   errors. The default tab width is 8. Allowed values are 1 to 63.\n     - \n     -+core.fsync::\n     -+\tA comma-separated list of parts of the repository which should be\n     -+\thardened via the core.fsyncMethod when created or modified. You can\n     -+\tdisable hardening of any component by prefixing it with a '-'. Later\n     -+\titems take precedence over earlier ones in the list. For example,\n     -+\t`core.fsync=all,-pack-metadata` means \"harden everything except pack\n     -+\tmetadata.\" Items that are not hardened may be lost in the event of an\n     -+\tunclean system shutdown.\n     -++\n     -+* `none` disables fsync completely. This must be specified alone.\n     -+* `loose-object` hardens objects added to the repo in loose-object form.\n     -+* `pack` hardens objects added to the repo in packfile form.\n     -+* `pack-metadata` hardens packfile bitmaps and indexes.\n     -+* `commit-graph` hardens the commit graph file.\n     -+* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n     -+  `pack-metadata`, and `commit-graph`.\n     -+* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n     -+* `all` is an aggregate option that syncs all individual components above.\n     -+\n     - core.fsyncMethod::\n     - \tA value indicating the strategy Git will use to harden repository data\n     - \tusing fsync and related primitives.\n      @@ Documentation/config/core.txt: core.fsyncMethod::\n         filesystem and storage hardware, data added to the repository may not be\n         durable in the event of a system crash. This is the default mode on macOS.\n     @@ cache.h: void reset_shared_repository(void);\n      +\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n      +\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n      +\n     -+#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     -+\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n     -+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n     -+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n     -+\n     -+#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     -+\t\t\t      FSYNC_COMPONENT_PACK | \\\n     -+\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n     -+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n     -+\n     -+\n      +/*\n      + * A bitmask indicating which components of the repo should be fsynced.\n      + */\n     @@ cache.h: int copy_file_with_time(const char *dst, const char *src, int mode);\n       void write_or_die(int fd, const void *buf, size_t count);\n       void fsync_or_die(int fd, const char *);\n       \n     -+inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n     ++static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n      +{\n      +\tif (fsync_components & component)\n      +\t\tfsync_or_die(fd, msg);\n     @@ config.c: static int git_parse_maybe_bool_text(const char *value)\n      +\t{ \"pack\", FSYNC_COMPONENT_PACK },\n      +\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n      +\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n     -+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n     -+\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n     -+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n      +};\n      +\n      +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n     @@ config.c: static int git_parse_maybe_bool_text(const char *value)\n      +\tenum fsync_component output = 0;\n      +\n      +\tif (!strcmp(string, \"none\"))\n     -+\t\treturn output;\n     ++\t\treturn FSYNC_COMPONENT_NONE;\n      +\n      +\twhile (string) {\n      +\t\tint i;\n     @@ midx.c: static int write_midx_internal(const char *object_dir,\n      +\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n       \tfree_chunkfile(cf);\n       \n     - \tif (flags & (MIDX_WRITE_REV_INDEX | MIDX_WRITE_BITMAP))\n     + \tif (flags & MIDX_WRITE_REV_INDEX &&\n      \n       ## object-file.c ##\n      @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n 3:  f217dba77a1 ! 4:  9da808ba743 core.fsync: new option to harden the index\n     @@ Commit message\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     - ## Documentation/config/core.txt ##\n     -@@ Documentation/config/core.txt: core.fsync::\n     - * `pack` hardens objects added to the repo in packfile form.\n     - * `pack-metadata` hardens packfile bitmaps and indexes.\n     - * `commit-graph` hardens the commit graph file.\n     -+* `index` hardens the index when it is modified.\n     - * `objects` is an aggregate option that includes `loose-objects`, `pack`,\n     -   `pack-metadata`, and `commit-graph`.\n     - * `default` is an aggregate option that is equivalent to `objects,-loose-object`\n     -\n       ## cache.h ##\n      @@ cache.h: enum fsync_component {\n       \tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n     @@ cache.h: enum fsync_component {\n       };\n       \n       #define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n     -@@ cache.h: enum fsync_component {\n     - #define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     - \t\t\t      FSYNC_COMPONENT_PACK | \\\n     - \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n     --\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH)\n     -+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n     -+\t\t\t      FSYNC_COMPONENT_INDEX)\n     - \n     - \n     - /*\n      \n       ## config.c ##\n      @@ config.c: static const struct fsync_component_entry {\n     @@ config.c: static const struct fsync_component_entry {\n       \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n       \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n      +\t{ \"index\", FSYNC_COMPONENT_INDEX },\n     - \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n     - \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n     - \t{ \"all\", FSYNC_COMPONENTS_ALL },\n     + };\n     + \n     + static enum fsync_component parse_fsync_components(const char *var, const char *string)\n      \n       ## read-cache.c ##\n      @@ read-cache.c: static int record_ieot(void)\n 4:  5c22a41c1f3 ! 5:  2d71346b10e core.fsync: add a `derived-metadata` aggregate option\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    core.fsync: add a `derived-metadata` aggregate option\n     +    core.fsync: documentation and user-friendly aggregate options\n      \n     -    This commit adds an aggregate option that currently includes the\n     -    commit-graph file and pack metadata (indexes and bitmaps).\n     +    This commit adds aggregate options for the core.fsync setting that are\n     +    more user-friendly. These options are specified in terms of 'levels of\n     +    safety', indicating which Git operations are considered to be sync\n     +    points for durability.\n      \n     -    The user may want to exclude this set from durability since they can be\n     -    recomputed from other data if they wind up corrupt or missing.\n     -\n     -    This is split out from the other patches in the series since it is\n     -    an optional nice-to-have that might be controversial.\n     +    The new documentation is also included here in its entirety for ease of\n     +    review.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Documentation/config/core.txt ##\n     -@@ Documentation/config/core.txt: core.fsync::\n     - * `pack-metadata` hardens packfile bitmaps and indexes.\n     - * `commit-graph` hardens the commit graph file.\n     - * `index` hardens the index when it is modified.\n     --* `objects` is an aggregate option that includes `loose-objects`, `pack`,\n     --  `pack-metadata`, and `commit-graph`.\n     --* `default` is an aggregate option that is equivalent to `objects,-loose-object`\n     -+* `objects` is an aggregate option that includes `loose-objects` and `pack`.\n     -+* `derived-metadata` is an aggregate option that includes `pack-metadata` and `commit-graph`.\n     -+* `default` is an aggregate option that is equivalent to `objects,derived-metadata,-loose-object`\n     - * `all` is an aggregate option that syncs all individual components above.\n     +@@ Documentation/config/core.txt: core.whitespace::\n     +   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n     +   errors. The default tab width is 8. Allowed values are 1 to 63.\n       \n     ++core.fsync::\n     ++\tA comma-separated list of parts of the repository which should be\n     ++\thardened via the core.fsyncMethod when created or modified. You can\n     ++\tdisable hardening of any component by prefixing it with a '-'. Later\n     ++\titems take precedence over earlier ones in the comma-separated list.\n     ++\tFor example, `core.fsync=all,-pack-metadata` means \"harden everything\n     ++\texcept pack metadata.\" Items that are not hardened may be lost in the\n     ++\tevent of an unclean system shutdown. Unless you have special\n     ++\trequirements, it is recommended that you leave this option as default\n     ++\tor pick one of `committed`, `added`, or `all`.\n     +++\n     ++* `none` disables fsync completely. This value must be specified alone.\n     ++* `loose-object` hardens objects added to the repo in loose-object form.\n     ++* `pack` hardens objects added to the repo in packfile form.\n     ++* `pack-metadata` hardens packfile bitmaps and indexes.\n     ++* `commit-graph` hardens the commit graph file.\n     ++* `index` hardens the index when it is modified.\n     ++* `objects` is an aggregate option that is equivalent to\n     ++  `loose-object,pack`.\n     ++* `derived-metadata` is an aggregate option that is equivalent to\n     ++  `pack-metadata,commit-graph`.\n     ++* `default` is an aggregate option that is equivalent to\n     ++  `objects,derived-metadata,-loose-object`. This mode is enabled by default.\n     ++  It has good performance, but risks losing recent work if the system shuts\n     ++  down uncleanly, since commits, trees, and blobs in loose-object form may be\n     ++  lost.\n     ++* `committed` is an aggregate option that is currently equivalent to\n     ++  `objects`. This mode sacrifices some performance to ensure that all work\n     ++  that is committed to the repository with `git commit` or similar commands\n     ++  is preserved.\n     ++* `added` is an aggregate option that is currently equivalent to\n     ++  `committed,index`. This mode sacrifices additional performance to\n     ++  ensure that the results of commands like `git add` and similar operations\n     ++  are preserved.\n     ++* `all` is an aggregate option that syncs all individual components above.\n     ++\n       core.fsyncMethod::\n     + \tA value indicating the strategy Git will use to harden repository data\n     + \tusing fsync and related primitives.\n      \n       ## cache.h ##\n      @@ cache.h: enum fsync_component {\n     - \t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n     + \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n     + };\n       \n     - #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     --\t\t\t\t  FSYNC_COMPONENT_PACK | \\\n     +-#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n      -\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n      -\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n     ++#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n      +\t\t\t\t  FSYNC_COMPONENT_PACK)\n      +\n      +#define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n      +\t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n     ++\n     ++#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENTS_OBJECTS | \\\n     ++\t\t\t\t  FSYNC_COMPONENTS_DERIVED_METADATA | \\\n     ++\t\t\t\t  ~FSYNC_COMPONENT_LOOSE_OBJECT)\n     ++\n     ++#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS)\n     ++\n     ++#define FSYNC_COMPONENTS_ADDED (FSYNC_COMPONENTS_COMMITTED | \\\n     ++\t\t\t\tFSYNC_COMPONENT_INDEX)\n     ++\n     ++#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     ++\t\t\t      FSYNC_COMPONENT_PACK | \\\n     ++\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n     ++\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n     ++\t\t\t      FSYNC_COMPONENT_INDEX)\n       \n     - #define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n     - \t\t\t      FSYNC_COMPONENT_PACK | \\\n     + /*\n     +  * A bitmask indicating which components of the repo should be fsynced.\n      \n       ## config.c ##\n      @@ config.c: static const struct fsync_component_entry {\n     + \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n       \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n       \t{ \"index\", FSYNC_COMPONENT_INDEX },\n     - \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n     ++\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n      +\t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n     - \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n     - \t{ \"all\", FSYNC_COMPONENTS_ALL },\n     ++\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n     ++\t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\n     ++\t{ \"added\", FSYNC_COMPONENTS_ADDED },\n     ++\t{ \"all\", FSYNC_COMPONENTS_ALL },\n       };\n     + \n     + static enum fsync_component parse_fsync_components(const char *var, const char *string)\n\n-- \ngitgitgadget\n"},{"id":"450950","messageId":"e31886717b42837f4e1538a13c8954aa07865af5.1646866998.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-09T23:03:16Z","receivedAt":"2022-03-09T23:03:36Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsync` configuration\nknob which can be used to control how components of the\nrepository are made durable on disk.\n\nThis setting allows future extensibility of the list of\nsyncable components:\n* We issue a warning rather than an error for unrecognized\n  components, so new configs can be used with old Git versions.\n* We support negation, so users can choose one of the aggregate\n  options and then remove components that they don't want.\n  Aggregate options are defined in a later patch in this series.\n\nThis also supports the common request of doing absolutely no\nfysncing with the `core.fsync=none` value, which is expected\nto make the test suite faster.\n\nComplete documentation for the new setting is included in a later patch\nin the series so that it can be reviewed in final form.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  8 ----\n builtin/fast-import.c         |  2 +-\n builtin/index-pack.c          |  4 +-\n builtin/pack-objects.c        | 24 ++++++++----\n bulk-checkin.c                |  5 ++-\n cache.h                       | 30 +++++++++++++-\n commit-graph.c                |  3 +-\n config.c                      | 73 ++++++++++++++++++++++++++++++++++-\n csum-file.c                   |  5 ++-\n csum-file.h                   |  3 +-\n environment.c                 |  2 +-\n midx.c                        |  3 +-\n object-file.c                 |  3 +-\n pack-bitmap-write.c           |  3 +-\n pack-write.c                  | 13 ++++---\n read-cache.c                  |  2 +-\n 16 files changed, 144 insertions(+), 39 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex dbb134f7136..74399072843 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -556,14 +556,6 @@ core.fsyncMethod::\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n \n-core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n-\n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\n +\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex b7105fcad9b..f2c036a8955 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -865,7 +865,7 @@ static void end_packfile(void)\n \t\tstruct tag *t;\n \n \t\tclose_pack_windows(pack_data);\n-\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n \t\t\t\t\t pack_data->pack_name, object_count,\n \t\t\t\t\t cur_pack_oid.hash, pack_size);\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c45273de3b1..c5f12f14df5 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1290,7 +1290,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\t\t    nr_objects - nr_objects_initial);\n \t\tstop_progress_msg(&progress, msg.buf);\n \t\tstrbuf_release(&msg);\n-\t\tfinalize_hashfile(f, tail_hash, 0);\n+\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n \t\thashcpy(read_hash, pack_hash);\n \t\tfixup_pack_header_footer(output_fd, pack_hash,\n \t\t\t\t\t curr_pack, nr_objects,\n@@ -1512,7 +1512,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tif (!from_stdin) {\n \t\tclose(input_fd);\n \t} else {\n-\t\tfsync_or_die(output_fd, curr_pack_name);\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n \t\terr = close(output_fd);\n \t\tif (err)\n \t\t\tdie_errno(_(\"error while closing pack file\"));\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 178e611f09d..c14fee8e99f 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1199,16 +1199,26 @@ static void write_pack_file(void)\n \t\t\tdisplay_progress(progress_state, written);\n \t\t}\n \n-\t\t/*\n-\t\t * Did we write the wrong # entries in the header?\n-\t\t * If so, rewrite it like in fast-import\n-\t\t */\n \t\tif (pack_to_stdout) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n+\t\t\t/*\n+\t\t\t * We never fsync when writing to stdout since we may\n+\t\t\t * not be writing to an actual pack file. For instance,\n+\t\t\t * the upload-pack code passes a pipe here. Calling\n+\t\t\t * fsync on a pipe results in unnecessary\n+\t\t\t * synchronization with the reader on some platforms.\n+\t\t\t */\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n \t\t} else if (nr_written == nr_remaining) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t\t} else {\n-\t\t\tint fd = finalize_hashfile(f, hash, 0);\n+\t\t\t/*\n+\t\t\t * If we wrote the wrong number of entries in the\n+\t\t\t * header, rewrite it like in fast-import.\n+\t\t\t */\n+\n+\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n \t\t\t\t\t\t nr_written, hash, offset);\n \t\t\tclose(fd);\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..a2cf9dcbc8d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -53,9 +53,10 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \t\tunlink(state->pack_tmp_name);\n \t\tgoto clear_exit;\n \t} else if (state->nr_written == 1) {\n-\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\tfinalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t} else {\n-\t\tint fd = finalize_hashfile(state->f, hash, 0);\n+\t\tint fd = finalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n \t\t\t\t\t state->nr_written, hash,\n \t\t\t\t\t state->offset);\ndiff --git a/cache.h b/cache.h\nindex 82f0194a3dd..a5eaa60a7a8 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -993,8 +993,27 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n-extern int use_fsync;\n+/*\n+ * These values are used to help identify parts of a repository to fsync.\n+ * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n+ * repository and so shouldn't be fsynced.\n+ */\n+enum fsync_component {\n+\tFSYNC_COMPONENT_NONE,\n+\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n+\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n+\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n+\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+};\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+/*\n+ * A bitmask indicating which components of the repo should be fsynced.\n+ */\n+extern enum fsync_component fsync_components;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n@@ -1002,6 +1021,7 @@ enum fsync_method {\n };\n \n extern enum fsync_method fsync_method;\n+extern int use_fsync;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\n@@ -1708,6 +1728,12 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n \n+static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n+{\n+\tif (fsync_components & component)\n+\t\tfsync_or_die(fd, msg);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 265c010122e..64897f57d9f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1942,7 +1942,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t}\n \n \tclose_commit_graph(ctx->r->objects);\n-\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n+\tfinalize_hashfile(f, file_hash, FSYNC_COMPONENT_COMMIT_GRAPH,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n \tfree_chunkfile(cf);\n \n \tif (ctx->split) {\ndiff --git a/config.c b/config.c\nindex f3ff80b01c9..51a35715642 100644\n--- a/config.c\n+++ b/config.c\n@@ -1323,6 +1323,70 @@ static int git_parse_maybe_bool_text(const char *value)\n \treturn -1;\n }\n \n+static const struct fsync_component_entry {\n+\tconst char *name;\n+\tenum fsync_component component_bits;\n+} fsync_component_table[] = {\n+\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n+\t{ \"pack\", FSYNC_COMPONENT_PACK },\n+\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n+\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+};\n+\n+static enum fsync_component parse_fsync_components(const char *var, const char *string)\n+{\n+\tenum fsync_component output = 0;\n+\n+\tif (!strcmp(string, \"none\"))\n+\t\treturn FSYNC_COMPONENT_NONE;\n+\n+\twhile (string) {\n+\t\tint i;\n+\t\tsize_t len;\n+\t\tconst char *ep;\n+\t\tint negated = 0;\n+\t\tint found = 0;\n+\n+\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n+\t\tep = strchrnul(string, ',');\n+\t\tlen = ep - string;\n+\n+\t\tif (*string == '-') {\n+\t\t\tnegated = 1;\n+\t\t\tstring++;\n+\t\t\tlen--;\n+\t\t\tif (!len)\n+\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n+\t\t}\n+\n+\t\tif (!len)\n+\t\t\tbreak;\n+\n+\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n+\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n+\n+\t\t\tif (strncmp(entry->name, string, len))\n+\t\t\t\tcontinue;\n+\n+\t\t\tfound = 1;\n+\t\t\tif (negated)\n+\t\t\t\toutput &= ~entry->component_bits;\n+\t\t\telse\n+\t\t\t\toutput |= entry->component_bits;\n+\t\t}\n+\n+\t\tif (!found) {\n+\t\t\tchar *component = xstrndup(string, len);\n+\t\t\twarning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n+\t\t\tfree(component);\n+\t\t}\n+\n+\t\tstring = ep;\n+\t}\n+\n+\treturn output;\n+}\n+\n int git_parse_maybe_bool(const char *value)\n {\n \tint v = git_parse_maybe_bool_text(value);\n@@ -1600,6 +1664,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsync\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tfsync_components = parse_fsync_components(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncmethod\")) {\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n@@ -1613,7 +1684,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n \t\treturn 0;\n \t}\n \ndiff --git a/csum-file.c b/csum-file.c\nindex 26e8a6df44e..59ef3398ca2 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -58,7 +58,8 @@ static void free_hashfile(struct hashfile *f)\n \tfree(f);\n }\n \n-int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n+int finalize_hashfile(struct hashfile *f, unsigned char *result,\n+\t\t      enum fsync_component component, unsigned int flags)\n {\n \tint fd;\n \n@@ -69,7 +70,7 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int fl\n \tif (flags & CSUM_HASH_IN_STREAM)\n \t\tflush(f, f->buffer, the_hash_algo->rawsz);\n \tif (flags & CSUM_FSYNC)\n-\t\tfsync_or_die(f->fd, f->name);\n+\t\tfsync_component_or_die(component, f->fd, f->name);\n \tif (flags & CSUM_CLOSE) {\n \t\tif (close(f->fd))\n \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\ndiff --git a/csum-file.h b/csum-file.h\nindex 291215b34eb..0d29f528fbc 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -1,6 +1,7 @@\n #ifndef CSUM_FILE_H\n #define CSUM_FILE_H\n \n+#include \"cache.h\"\n #include \"hash.h\"\n \n struct progress;\n@@ -38,7 +39,7 @@ int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n struct hashfile *hashfd(int fd, const char *name);\n struct hashfile *hashfd_check(const char *name);\n struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n-int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n+int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n void hashwrite(struct hashfile *, const void *, unsigned int);\n void hashflush(struct hashfile *f);\n void crc32_begin(struct hashfile *);\ndiff --git a/environment.c b/environment.c\nindex 3e3620d759f..378424b9af5 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,9 +42,9 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n int use_fsync = -1;\n enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n+enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/midx.c b/midx.c\nindex 865170bad05..107365d2114 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -1438,7 +1438,8 @@ static int write_midx_internal(const char *object_dir,\n \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n \twrite_chunkfile(cf, &ctx);\n \n-\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n \tfree_chunkfile(cf);\n \n \tif (flags & MIDX_WRITE_REV_INDEX &&\ndiff --git a/object-file.c b/object-file.c\nindex 03bd6a3baf3..c9de912faca 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1850,8 +1850,7 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n static void close_loose_object(int fd)\n {\n \tif (!the_repository->objects->odb->will_destroy) {\n-\t\tif (fsync_object_files)\n-\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n \t}\n \n \tif (close(fd) != 0)\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex cab3eaa2acd..cf681547f2e 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -719,7 +719,8 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \tif (options & BITMAP_OPT_HASH_CACHE)\n \t\twrite_hash_cache(f, index, index_nr);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \n \tif (adjust_shared_perm(tmp_file.buf))\n \t\tdie_errno(\"unable to make temporary bitmap file readable\");\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a346..51812cb1299 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -159,9 +159,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t}\n \n \thashwrite(f, sha1, the_hash_algo->rawsz);\n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n-\t\t\t\t    ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \treturn index_name;\n }\n \n@@ -281,8 +281,9 @@ const char *write_rev_file_order(const char *rev_name,\n \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n \t\tdie(_(\"failed to make %s readable\"), rev_name);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \n \treturn rev_name;\n }\n@@ -390,7 +391,7 @@ void fixup_pack_header_footer(int pack_fd,\n \t\tthe_hash_algo->final_fn(partial_pack_hash, &old_hash_ctx);\n \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n-\tfsync_or_die(pack_fd, pack_name);\n+\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n }\n \n char *index_pack_lockfile(int ip_out, int *is_well_formed)\ndiff --git a/read-cache.c b/read-cache.c\nindex 79b9b99ebf7..df869691fd4 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -3089,7 +3089,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"450952","messageId":"9da808ba743673bb0c40b127fed7dd34c87d232d.1646866998.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH v5 4/5] core.fsync: new option to harden the index","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-09T23:03:17Z","receivedAt":"2022-03-09T23:03:39Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the new ability for the user to harden\nthe index. In the event of a system crash, the index must be\ndurable for the user to actually find a file that has been added\nto the repo and then deleted from the working tree.\n\nWe use the presence of the COMMIT_LOCK flag and absence of the\nalternate_index_output as a proxy for determining whether we're\nupdating the persistent index of the repo or some temporary\nindex. We don't sync these temporary indexes.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache.h      |  1 +\n config.c     |  1 +\n read-cache.c | 19 +++++++++++++------\n 3 files changed, 15 insertions(+), 6 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex a5eaa60a7a8..3fefe8f6f60 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1004,6 +1004,7 @@ enum fsync_component {\n \tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+\tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n #define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\ndiff --git a/config.c b/config.c\nindex 51a35715642..bb44f6f506d 100644\n--- a/config.c\n+++ b/config.c\n@@ -1331,6 +1331,7 @@ static const struct fsync_component_entry {\n \t{ \"pack\", FSYNC_COMPONENT_PACK },\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"index\", FSYNC_COMPONENT_INDEX },\n };\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\ndiff --git a/read-cache.c b/read-cache.c\nindex df869691fd4..7683b679258 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2842,7 +2842,7 @@ static int record_ieot(void)\n  * rely on it.\n  */\n static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n-\t\t\t  int strip_extensions)\n+\t\t\t  int strip_extensions, unsigned flags)\n {\n \tuint64_t start = getnanotime();\n \tstruct hashfile *f;\n@@ -2856,6 +2856,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint csum_fsync_flag;\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n@@ -3089,7 +3090,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n+\tcsum_fsync_flag = 0;\n+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n+\t\tcsum_fsync_flag = CSUM_FSYNC;\n+\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n+\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n@@ -3144,7 +3151,7 @@ static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n \t */\n \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n-\tret = do_write_index(istate, lock->tempfile, 0);\n+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n \n@@ -3238,7 +3245,7 @@ static int clean_shared_index_files(const char *current_hex)\n }\n \n static int write_shared_index(struct index_state *istate,\n-\t\t\t      struct tempfile **temp)\n+\t\t\t      struct tempfile **temp, unsigned flags)\n {\n \tstruct split_index *si = istate->split_index;\n \tint ret, was_full = !istate->sparse_index;\n@@ -3248,7 +3255,7 @@ static int write_shared_index(struct index_state *istate,\n \n \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n-\tret = do_write_index(si->base, *temp, 1);\n+\tret = do_write_index(si->base, *temp, 1, flags);\n \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n \n@@ -3357,7 +3364,7 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n \t\t\tret = do_write_locked_index(istate, lock, flags);\n \t\t\tgoto out;\n \t\t}\n-\t\tret = write_shared_index(istate, &temp);\n+\t\tret = write_shared_index(istate, &temp, flags);\n \n \t\tsaved_errno = errno;\n \t\tif (is_tempfile_active(temp))\n-- \ngitgitgadget\n\n"},{"id":"450953","messageId":"2d71346b10e8beb3c44bdf8e06694e4dafc657d5.1646866998.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH v5 5/5] core.fsync: documentation and user-friendly aggregate options","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-09T23:03:18Z","receivedAt":"2022-03-09T23:03:41Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds aggregate options for the core.fsync setting that are\nmore user-friendly. These options are specified in terms of 'levels of\nsafety', indicating which Git operations are considered to be sync\npoints for durability.\n\nThe new documentation is also included here in its entirety for ease of\nreview.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 36 +++++++++++++++++++++++++++++++++++\n cache.h                       | 23 +++++++++++++++++++---\n config.c                      |  6 ++++++\n 3 files changed, 62 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 74399072843..973805e8a98 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,42 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsync::\n+\tA comma-separated list of parts of the repository which should be\n+\thardened via the core.fsyncMethod when created or modified. You can\n+\tdisable hardening of any component by prefixing it with a '-'. Later\n+\titems take precedence over earlier ones in the comma-separated list.\n+\tFor example, `core.fsync=all,-pack-metadata` means \"harden everything\n+\texcept pack metadata.\" Items that are not hardened may be lost in the\n+\tevent of an unclean system shutdown. Unless you have special\n+\trequirements, it is recommended that you leave this option as default\n+\tor pick one of `committed`, `added`, or `all`.\n++\n+* `none` disables fsync completely. This value must be specified alone.\n+* `loose-object` hardens objects added to the repo in loose-object form.\n+* `pack` hardens objects added to the repo in packfile form.\n+* `pack-metadata` hardens packfile bitmaps and indexes.\n+* `commit-graph` hardens the commit graph file.\n+* `index` hardens the index when it is modified.\n+* `objects` is an aggregate option that is equivalent to\n+  `loose-object,pack`.\n+* `derived-metadata` is an aggregate option that is equivalent to\n+  `pack-metadata,commit-graph`.\n+* `default` is an aggregate option that is equivalent to\n+  `objects,derived-metadata,-loose-object`. This mode is enabled by default.\n+  It has good performance, but risks losing recent work if the system shuts\n+  down uncleanly, since commits, trees, and blobs in loose-object form may be\n+  lost.\n+* `committed` is an aggregate option that is currently equivalent to\n+  `objects`. This mode sacrifices some performance to ensure that all work\n+  that is committed to the repository with `git commit` or similar commands\n+  is preserved.\n+* `added` is an aggregate option that is currently equivalent to\n+  `committed,index`. This mode sacrifices additional performance to\n+  ensure that the results of commands like `git add` and similar operations\n+  are preserved.\n+* `all` is an aggregate option that syncs all individual components above.\n+\n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\ndiff --git a/cache.h b/cache.h\nindex 3fefe8f6f60..833f0236e68 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1007,9 +1007,26 @@ enum fsync_component {\n \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n-#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n-\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK)\n+\n+#define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENTS_OBJECTS | \\\n+\t\t\t\t  FSYNC_COMPONENTS_DERIVED_METADATA | \\\n+\t\t\t\t  ~FSYNC_COMPONENT_LOOSE_OBJECT)\n+\n+#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS)\n+\n+#define FSYNC_COMPONENTS_ADDED (FSYNC_COMPONENTS_COMMITTED | \\\n+\t\t\t\tFSYNC_COMPONENT_INDEX)\n+\n+#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t      FSYNC_COMPONENT_PACK | \\\n+\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n+\t\t\t      FSYNC_COMPONENT_INDEX)\n \n /*\n  * A bitmask indicating which components of the repo should be fsynced.\ndiff --git a/config.c b/config.c\nindex bb44f6f506d..3976ec74fd4 100644\n--- a/config.c\n+++ b/config.c\n@@ -1332,6 +1332,12 @@ static const struct fsync_component_entry {\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n+\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n+\t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\n+\t{ \"added\", FSYNC_COMPONENTS_ADDED },\n+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n };\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n-- \ngitgitgadget\n"},{"id":"450956","messageId":"xmqq8rtiln6a.fsf@gitster.g","threadId":"57030","inReplyTo":"685b1db888079c83573cfd984ae64f46284544af.1646866998.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 1/5] wrapper: move inclusion of CSPRNG headers the wrapper.c file","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-09T23:29:49Z","receivedAt":"2022-03-09T23:29:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Including NTSecAPI.h in git-compat-util.h causes build errors in any\n> other file that includes winternl.h. That file was included in order to\n> get access to the RtlGenRandom cryptographically secure PRNG. This\n> change scopes the inclusion of all PRNG headers to just the wrapper.c\n> file, which is the only place it is really needed.\n\nIt is true that wrapper.c is the only thing that needs these headers\nincluded as part of its implementation detail of csprng_bytes(), and\nI think I very much like this change for that reason.\n\nHaving said that, if it true that including these two header files\nin the same file will lead to compilation failure?  That sounds like\neither (1) they represent two mutually incompatible APIs that will\ncause breakage when code that use them, perhaps in two separate\nfiles to avoid compilation failures, are linked together, or (2)\nthese system header files are simply broken.  Or something else?\n\n> -/*\n> - * Including the appropriate header file for RtlGenRandom causes MSVC to see a\n> - * redefinition of types in an incompatible way when including headers below.\n> - */\n> -#undef HAVE_RTLGENRANDOM\n\nThis comment hints it is more like (1) above?  A type used in one\npart of the system is defined differently in other parts of the\nsystem?  I cannot imagine anything but bad things happen when a\npiece of code uses one definition of the type to declare a function,\nand another piece of code uses the other definition of the same type\nto declare a variable and passes it as a parameter to that function.\n\nI do not know this patch makes the situation worse, and I am not a\nWindows person with boxes to dig deeper to begin with.  Hence I do\nnot mind the change itself, but justifying the change primarily as a\nworkaround for some obscure header type clashes on a single system\nleaves a bad taste.  If the first sentence weren't there, I wouldn't\nhave spent this many minutes wondering if this is a good change ;-)\n"},{"id":"450958","messageId":"xmqqv8wmk7rd.fsf@gitster.g","threadId":"57030","inReplyTo":"da8cfc10bb4bfa473619d8d737c3aa160643ccf7.1646866998.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 2/5] core.fsyncmethod: add writeout-only mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-09T23:48:06Z","receivedAt":"2022-03-09T23:48:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> This commit introduces the `core.fsyncMethod` configuration\n> knob, which can currently be set to `fsync` or `writeout-only`.\n>\n> The new writeout-only mode attempts to tell the operating system to\n> flush its in-memory page cache to the storage hardware without issuing a\n> CACHE_FLUSH command to the storage controller.\n>\n> Writeout-only fsync is significantly faster than a vanilla fsync on\n> common hardware, since data is written to a disk-side cache rather than\n> all the way to a durable medium. Later changes in this patch series will\n> take advantage of this primitive to implement batching of hardware\n> flushes.\n>\n> When git_fsync is called with FSYNC_WRITEOUT_ONLY, it may fail and the\n> caller is expected to do an ordinary fsync as needed.\n>\n> On Apple platforms, the fsync system call does not issue a CACHE_FLUSH\n> directive to the storage controller. This change updates fsync to do\n> fcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\n> with existing behavior on Apple platforms by setting the default value\n> of the new core.fsyncMethod option.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n\nOK.  This seems to be quite reasonable in that the pieces of code\nthat use fsync_or_die() do not have to change at all, and all of\nthem will keep behaving the same way.  In other words, the \"how\" of\nfsync is very much well isolated.\n\nNice.\n"},{"id":"450962","messageId":"xmqqo82eirnv.fsf@gitster.g","threadId":"57030","inReplyTo":"e31886717b42837f4e1538a13c8954aa07865af5.1646866998.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T00:21:08Z","receivedAt":"2022-03-10T00:21:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +/*\n> + * These values are used to help identify parts of a repository to fsync.\n> + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n> + * repository and so shouldn't be fsynced.\n> + */\n> +enum fsync_component {\n> +\tFSYNC_COMPONENT_NONE,\n> +\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n> +\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n> +\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n> +\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n> +};\n\nOK, so the idea is that Patrick's \"we need to fsync refs\" will be\ndone by adding a new component to this list, and sprinkling a call\nto fsync_component_or_die() in the code of ref-files backend?\n\nI am wondering if fsync_or_die() interface is abstracted well\nenough, or we need things like \"the fd is inside this directory; in\naddition to doing the fsync of the fd, please sync the parent\ndirectory as well\" support before we start adding more components\n(if there is such a need, perhaps it comes before this step).\n\n> +#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n> +\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n> +\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n\nIOW, everything other than loose object, which already has a\nseparate core.fsyncObjectFiles knob to loosen.  Everything else we\ncurrently sync unconditionally and the default keeps that\narrangement?\n\n> +static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n> +{\n> +\tif (fsync_components & component)\n> +\t\tfsync_or_die(fd, msg);\n> +}\n\nDo we have a compelling reason to have this as a static inline\nfunction?  We are talking about concluding an I/O operation and\nI doubt there is a good performance argument for it.\n\n> +static const struct fsync_component_entry {\n> +\tconst char *name;\n> +\tenum fsync_component component_bits;\n> +} fsync_component_table[] = {\n\nthing[] is an array of \"thing\" (and thing[4] is the \"fourth\" such\nthing), but this is not an array of a table (it is a name-to-bit\nmapping).\n\nI wonder if this array works without \"_table\" suffix in its name.\n\n> +\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> +\t{ \"pack\", FSYNC_COMPONENT_PACK },\n> +\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> +\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> +};\n> +\n> +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> +{\n> +\tenum fsync_component output = 0;\n> +\n> +\tif (!strcmp(string, \"none\"))\n> +\t\treturn FSYNC_COMPONENT_NONE;\n> +\n> +\twhile (string) {\n> +\t\tint i;\n> +\t\tsize_t len;\n> +\t\tconst char *ep;\n> +\t\tint negated = 0;\n> +\t\tint found = 0;\n> +\n> +\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n> +\t\tep = strchrnul(string, ',');\n> +\t\tlen = ep - string;\n> +\n> +\t\tif (*string == '-') {\n> +\t\t\tnegated = 1;\n> +\t\t\tstring++;\n> +\t\t\tlen--;\n> +\t\t\tif (!len)\n> +\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n> +\t\t}\n> +\n> +\t\tif (!len)\n> +\t\t\tbreak;\n> +\n> +\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n> +\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n> +\n> +\t\t\tif (strncmp(entry->name, string, len))\n> +\t\t\t\tcontinue;\n> +\n> +\t\t\tfound = 1;\n> +\t\t\tif (negated)\n> +\t\t\t\toutput &= ~entry->component_bits;\n> +\t\t\telse\n> +\t\t\t\toutput |= entry->component_bits;\n> +\t\t}\n> +\n> +\t\tif (!found) {\n> +\t\t\tchar *component = xstrndup(string, len);\n> +\t\t\twarning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n> +\t\t\tfree(component);\n> +\t\t}\n> +\n> +\t\tstring = ep;\n> +\t}\n> +\n> +\treturn output;\n> +}\n\nHmph.  I would have expected, with built-in default of\npack,pack-metadata,commit-graph,\n\n - \"none,pack\" would choose only \"pack\" by first clearing the\n   built-in default (or whatever was set in configuration files that\n   are lower precedence than what we are reading) and then OR'ing\n   the \"pack\" bit in.\n\n - \"-pack\" would choose \"pack-metadata,commit-graph\" by first\n   starting from the built-in default and then CLR'ing the \"pack\"\n   bit out.  If there were already changes made by the lower\n   precedence configuration files like /etc/gitconfig, the result\n   might be different and the only definite thing we can say is that\n   the pack bit is cleared.\n\n - \"loose-object\" would choose all of the bits by first starting\n   from the built-in default and then OR'ing the \"loose-object\" bit\n   in.\n\nOtherwise, parsing \"none\" is more or less pointless, as the above\nparser always start from 0 and OR's in or CLR's out the named bit.\nWhoever writes \"none\" can just write an empty string, no?\n\nI wonder you'd rather want to do it this way?\n\nparse_fsync_components(var, value, current) {\n\tenum fsync_component positive = 0, negative = 0;\n\n\twhile (string) {\n\t\tint negated = 0;\n\t\tenum fsync_component bits;\n\n\t\tparse out a single component into <negated, bits>;\n\n\t\tif (bits == 0) { /* \"none\" given */\n                \tcurrent = 0;\n\t\t} else if (negated) {\n\t\t\tnegative |= bits;\n\t\t} else {\n\t\t\tpositive |= bits;\n\t\t}\n\t\tadvance <string> pointer;\n\t}\n\n\treturn (current | positive) & ~negative;\n}\n\nAnd then ...\n\n> +\tif (!strcmp(var, \"core.fsync\")) {\n> +\t\tif (!value)\n> +\t\t\treturn config_error_nonbool(var);\n> +\t\tfsync_components = parse_fsync_components(var, value);\n> +\t\treturn 0;\n> +\t}\n> +\n\n... this part would pass the current value of fsync_components as\nthe third parameter to the parse_fsync_components().  The variable\nwould be initialized to the FSYNC_COMPONENTS_DEFAULT we saw earlier.\n\n\n> @@ -1613,7 +1684,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t}\n>  \n>  \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> -\t\tfsync_object_files = git_config_bool(var, value);\n> +\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n\nThis is not deprecating but removing the support, which I am not\nsure is a sensible thing to do.  Rather we should pretend that\ncore.fsync = \"loose-object\" (or \"-loose-object\") were found in the\nconfiguration, shouldn't we?\n\n"},{"id":"450972","messageId":"CANQDOde4e6434mDE3rra9VQMfhfksmJT48jvgNk6M4XDqtthrg@mail.gmail.com","threadId":"57030","inReplyTo":"xmqq8rtiln6a.fsf@gitster.g","subject":"Re: [PATCH v5 1/5] wrapper: move inclusion of CSPRNG headers the wrapper.c file","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T01:21:43Z","receivedAt":"2022-03-10T01:22:00Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 9, 2022 at 3:29 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > Including NTSecAPI.h in git-compat-util.h causes build errors in any\n> > other file that includes winternl.h. That file was included in order to\n> > get access to the RtlGenRandom cryptographically secure PRNG. This\n> > change scopes the inclusion of all PRNG headers to just the wrapper.c\n> > file, which is the only place it is really needed.\n>\n> It is true that wrapper.c is the only thing that needs these headers\n> included as part of its implementation detail of csprng_bytes(), and\n> I think I very much like this change for that reason.\n>\n> Having said that, if it true that including these two header files\n> in the same file will lead to compilation failure?  That sounds like\n> either (1) they represent two mutually incompatible APIs that will\n> cause breakage when code that use them, perhaps in two separate\n> files to avoid compilation failures, are linked together, or (2)\n> these system header files are simply broken.  Or something else?\n>\n> > -/*\n> > - * Including the appropriate header file for RtlGenRandom causes MSVC to see a\n> > - * redefinition of types in an incompatible way when including headers below.\n> > - */\n> > -#undef HAVE_RTLGENRANDOM\n>\n> This comment hints it is more like (1) above?  A type used in one\n> part of the system is defined differently in other parts of the\n> system?  I cannot imagine anything but bad things happen when a\n> piece of code uses one definition of the type to declare a function,\n> and another piece of code uses the other definition of the same type\n> to declare a variable and passes it as a parameter to that function.\n>\n> I do not know this patch makes the situation worse, and I am not a\n> Windows person with boxes to dig deeper to begin with.  Hence I do\n> not mind the change itself, but justifying the change primarily as a\n> workaround for some obscure header type clashes on a single system\n> leaves a bad taste.  If the first sentence weren't there, I wouldn't\n> have spent this many minutes wondering if this is a good change ;-)\n\nThis is (2), these system header files are simply broken.  I've been\nlooking deeper into why, but haven't bottomed out yet.\n\nThanks,\nNeeraj\n"},{"id":"450974","messageId":"YilTug0iH/N2Fbpb@camp.crustytoothpaste.net","threadId":"57030","inReplyTo":"685b1db888079c83573cfd984ae64f46284544af.1646866998.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 1/5] wrapper: move inclusion of CSPRNG headers the wrapper.c file","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2022-03-10T01:26:18Z","receivedAt":"2022-03-10T01:26:30Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2022-03-09 at 23:03:14, Neeraj Singh via GitGitGadget wrote:\n> From: Neeraj Singh <neerajsi@microsoft.com>\n> \n> Including NTSecAPI.h in git-compat-util.h causes build errors in any\n> other file that includes winternl.h. That file was included in order to\n> get access to the RtlGenRandom cryptographically secure PRNG. This\n> change scopes the inclusion of all PRNG headers to just the wrapper.c\n> file, which is the only place it is really needed.\n\nWe generally prefer to do system includes in git-compat-util.h because\nit allows us to paper over platform incompatibilities in one place and\nto deal with the various ordering problems that can happen on certain\nsystems.\n\nIt may be that Windows needs additional help here; I don't know, because\nI don't use Windows.  I personally find it unsavoury that Windows ships\nwith multiple incompatible header files like this, since such problems\nare typically avoided by suitable include guards, whose utility has been\nwell known for several decades.  However, if that's the case, let's move\nonly the Windows changes there, and leave the Unix systems, which lack\nthis problem, alone.\n\nIt would also be helpful to explain the problem that Windows has in more\ndetail here, including any references to documentation that explains\nthis incompatibility, so those of us who are not Windows users can more\naccurately reason about why we need to be so careful when including\nheader files there and why this is the best solution (and not, say,\nproviding our own include guards in a compat file).\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"450977","messageId":"CANQDOdfZbOHZQt9Ah0t1AamTO2T7Gq0tmWX1jLqL6njE0LF6DA@mail.gmail.com","threadId":"57030","inReplyTo":"YilTug0iH/N2Fbpb@camp.crustytoothpaste.net","subject":"Re: [PATCH v5 1/5] wrapper: move inclusion of CSPRNG headers the wrapper.c file","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T01:56:09Z","receivedAt":"2022-03-10T01:56:24Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 9, 2022 at 5:26 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> On 2022-03-09 at 23:03:14, Neeraj Singh via GitGitGadget wrote:\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Including NTSecAPI.h in git-compat-util.h causes build errors in any\n> > other file that includes winternl.h. That file was included in order to\n> > get access to the RtlGenRandom cryptographically secure PRNG. This\n> > change scopes the inclusion of all PRNG headers to just the wrapper.c\n> > file, which is the only place it is really needed.\n>\n> We generally prefer to do system includes in git-compat-util.h because\n> it allows us to paper over platform incompatibilities in one place and\n> to deal with the various ordering problems that can happen on certain\n> systems.\n>\n> It may be that Windows needs additional help here; I don't know, because\n> I don't use Windows.  I personally find it unsavoury that Windows ships\n> with multiple incompatible header files like this, since such problems\n> are typically avoided by suitable include guards, whose utility has been\n> well known for several decades.  However, if that's the case, let's move\n> only the Windows changes there, and leave the Unix systems, which lack\n> this problem, alone.\n>\n> It would also be helpful to explain the problem that Windows has in more\n> detail here, including any references to documentation that explains\n> this incompatibility, so those of us who are not Windows users can more\n> accurately reason about why we need to be so careful when including\n> header files there and why this is the best solution (and not, say,\n> providing our own include guards in a compat file).\n> --\n> brian m. carlson (he/him or they/them)\n> Toronto, Ontario, CA\n\nI wasn't able to find any documentation from other people who hit this problem.\n\nThe root cause is that NtSecAPI.h has a typedef like this:\n```\n#ifndef _NTDEF_\ntypedef LSA_UNICODE_STRING UNICODE_STRING, *PUNICODE_STRING;\ntypedef LSA_STRING STRING, *PSTRING ;\n#endif\n```\nThat's not really appropriate since NtSecAPI.h isn't the correct place\nto define this core primitive NT type.  It should be including\nwinternl.h or a similar header.\n\nI'll update the change to only move the Windows definition to the .c file.\n"},{"id":"450982","messageId":"CANQDOddU_WXD-6ncDGBrgpsuKT-XDGC=SeaaQTNQFdODFZ7TkQ@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqo82eirnv.fsf@gitster.g","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T02:53:45Z","receivedAt":"2022-03-10T02:54:02Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 9, 2022 at 4:21 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > +/*\n> > + * These values are used to help identify parts of a repository to fsync.\n> > + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n> > + * repository and so shouldn't be fsynced.\n> > + */\n> > +enum fsync_component {\n> > +     FSYNC_COMPONENT_NONE,\n> > +     FSYNC_COMPONENT_LOOSE_OBJECT            = 1 << 0,\n> > +     FSYNC_COMPONENT_PACK                    = 1 << 1,\n> > +     FSYNC_COMPONENT_PACK_METADATA           = 1 << 2,\n> > +     FSYNC_COMPONENT_COMMIT_GRAPH            = 1 << 3,\n> > +};\n>\n> OK, so the idea is that Patrick's \"we need to fsync refs\" will be\n> done by adding a new component to this list, and sprinkling a call\n> to fsync_component_or_die() in the code of ref-files backend?\n>\n\nYes. Patrick will need to add fsync_component_or_die wherever his\npatch series has already added fsync_or_die.\n\nIf he follows Ævar's suggestion of treating remote refs differently\nfrom local refs, he might want to define multiple components.\n\n> I am wondering if fsync_or_die() interface is abstracted well\n> enough, or we need things like \"the fd is inside this directory; in\n> addition to doing the fsync of the fd, please sync the parent\n> directory as well\" support before we start adding more components\n> (if there is such a need, perhaps it comes before this step).\n>\n\nI think syncing the parent directory is a separate fsyncMethod that\nwould require changes across the codebase to obtain an appropriate\ndirectory fd. I'd prefer to treat that as a separable concern.\n\n> > +#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n> > +                               FSYNC_COMPONENT_PACK_METADATA | \\\n> > +                               FSYNC_COMPONENT_COMMIT_GRAPH)\n>\n> IOW, everything other than loose object, which already has a\n> separate core.fsyncObjectFiles knob to loosen.  Everything else we\n> currently sync unconditionally and the default keeps that\n> arrangement?\n>\n\nYes, trying to keep default behavior identical on non-Windows\nplatforms.  Windows will be expected to adopt batch mode, and have\nloose objects in this set.\n\n> > +static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n> > +{\n> > +     if (fsync_components & component)\n> > +             fsync_or_die(fd, msg);\n> > +}\n>\n> Do we have a compelling reason to have this as a static inline\n> function?  We are talking about concluding an I/O operation and\n> I doubt there is a good performance argument for it.\n>\n\nYeah, this is meant to optimize the case where the component isn't\nbeing fsynced. I'll move this function to write-or-die.c below\nfsync_or_die.\n\n> > +static const struct fsync_component_entry {\n> > +     const char *name;\n> > +     enum fsync_component component_bits;\n> > +} fsync_component_table[] = {\n>\n> thing[] is an array of \"thing\" (and thing[4] is the \"fourth\" such\n> thing), but this is not an array of a table (it is a name-to-bit\n> mapping).\n>\n> I wonder if this array works without \"_table\" suffix in its name.\n\nThis is modeled after whitespace_rule_names.  What if I change this to\nthe following?\nstatic const struct fsync_component_name {\n...\n} fsync_component_names[]\n\n>\n> > +     { \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n> > +     { \"pack\", FSYNC_COMPONENT_PACK },\n> > +     { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n> > +     { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n> > +};\n> > +\n> > +static enum fsync_component parse_fsync_components(const char *var, const char *string)\n> > +{\n> > +     enum fsync_component output = 0;\n> > +\n> > +     if (!strcmp(string, \"none\"))\n> > +             return FSYNC_COMPONENT_NONE;\n> > +\n> > +     while (string) {\n> > +             int i;\n> > +             size_t len;\n> > +             const char *ep;\n> > +             int negated = 0;\n> > +             int found = 0;\n> > +\n> > +             string = string + strspn(string, \", \\t\\n\\r\");\n> > +             ep = strchrnul(string, ',');\n> > +             len = ep - string;\n> > +\n> > +             if (*string == '-') {\n> > +                     negated = 1;\n> > +                     string++;\n> > +                     len--;\n> > +                     if (!len)\n> > +                             warning(_(\"invalid value for variable %s\"), var);\n> > +             }\n> > +\n> > +             if (!len)\n> > +                     break;\n> > +\n> > +             for (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n> > +                     const struct fsync_component_entry *entry = &fsync_component_table[i];\n> > +\n> > +                     if (strncmp(entry->name, string, len))\n> > +                             continue;\n> > +\n> > +                     found = 1;\n> > +                     if (negated)\n> > +                             output &= ~entry->component_bits;\n> > +                     else\n> > +                             output |= entry->component_bits;\n> > +             }\n> > +\n> > +             if (!found) {\n> > +                     char *component = xstrndup(string, len);\n> > +                     warning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n> > +                     free(component);\n> > +             }\n> > +\n> > +             string = ep;\n> > +     }\n> > +\n> > +     return output;\n> > +}\n>\n> Hmph.  I would have expected, with built-in default of\n> pack,pack-metadata,commit-graph,\n>\n\nAt the conclusion of this series, I defined 'default' as an aggregate\noption that includes\nthe platform default.  I'd prefer not to have any statefulness of the\ncore.fsync setting so\nthat there is less confusion about the final fsync configuration. My\ncolleagues had a fair\namount of confusion internally when testing Git performance internally\nwith regards to\nthe core.fsyncObjectFiles setting.  Inline this is how your configs\nwould be written:\n\n>  - \"none,pack\" would choose only \"pack\" by first clearing the\n>    built-in default (or whatever was set in configuration files that\n>    are lower precedence than what we are reading) and then OR'ing\n>    the \"pack\" bit in.\n>\n\n\"pack\" would choose only \"pack\"\n\n>  - \"-pack\" would choose \"pack-metadata,commit-graph\" by first\n>    starting from the built-in default and then CLR'ing the \"pack\"\n>    bit out.  If there were already changes made by the lower\n>    precedence configuration files like /etc/gitconfig, the result\n>    might be different and the only definite thing we can say is that\n>    the pack bit is cleared.\n>\n\n\"default,-pack\" would be the platform default, but not packfiles.\n\n>  - \"loose-object\" would choose all of the bits by first starting\n>    from the built-in default and then OR'ing the \"loose-object\" bit\n>    in.\n>\n\n\"default,loose-object\" would add loose objects to the platform config.\n\n> Otherwise, parsing \"none\" is more or less pointless, as the above\n> parser always start from 0 and OR's in or CLR's out the named bit.\n> Whoever writes \"none\" can just write an empty string, no?\n\nI think the empty string would be disallowed since \"core.fsync=\" would\nbe entirely\nmissing a value. But on testing this doesn't seem to be the case. I'll change\nthis to be more strict in that the user has to pass an explicit value,\nsuch as 'none'.\n\n>\n> I wonder you'd rather want to do it this way?\n>\n> parse_fsync_components(var, value, current) {\n>         enum fsync_component positive = 0, negative = 0;\n>\n>         while (string) {\n>                 int negated = 0;\n>                 enum fsync_component bits;\n>\n>                 parse out a single component into <negated, bits>;\n>\n>                 if (bits == 0) { /* \"none\" given */\n>                         current = 0;\n>                 } else if (negated) {\n>                         negative |= bits;\n>                 } else {\n>                         positive |= bits;\n>                 }\n>                 advance <string> pointer;\n>         }\n>\n>         return (current | positive) & ~negative;\n> }\n>\n> And then ...\n>\n> > +     if (!strcmp(var, \"core.fsync\")) {\n> > +             if (!value)\n> > +                     return config_error_nonbool(var);\n> > +             fsync_components = parse_fsync_components(var, value);\n> > +             return 0;\n> > +     }\n> > +\n>\n> ... this part would pass the current value of fsync_components as\n> the third parameter to the parse_fsync_components().  The variable\n> would be initialized to the FSYNC_COMPONENTS_DEFAULT we saw earlier.\n>\n\nI'm afraid that this method would lead to statefulness between\ndifferent levels configuring\ncore.fsync.  I'd prefer that the user could know what will happen by\njust inspecting the value\nreturned by `git config core.fsync`.\n\n>\n> > @@ -1613,7 +1684,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >       }\n> >\n> >       if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> > -             fsync_object_files = git_config_bool(var, value);\n> > +             warning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n>\n> This is not deprecating but removing the support, which I am not\n> sure is a sensible thing to do.  Rather we should pretend that\n> core.fsync = \"loose-object\" (or \"-loose-object\") were found in the\n> configuration, shouldn't we?\n>\n\nThe problem I anticipate is that figuring out the final configuration\nbecomes pretty\ncomplex in the face of conflicting configurations of fsyncObjectFiles\nand core.fsync.\nThe user won't know what will happen without reading the Git code or\ndoing performance\nexperiments.  I thought we can avoid all of this complexity by just\nhaving a simple warning\nthat pushes users toward the new configuration value.  Aside from\nseeing a warning, a\nuser's actual usage of Git functionality shouldn't be affected.\n\nAlternatively, what if we just silently ignore the old\ncore.fsyncObjectFiles setting?\nIf neither of these is an option, I'll put back some support for the\nold setting.\n\nThanks,\nNeeraj\n"},{"id":"450985","messageId":"xmqqcziugtpw.fsf@gitster.g","threadId":"57030","inReplyTo":"CANQDOddU_WXD-6ncDGBrgpsuKT-XDGC=SeaaQTNQFdODFZ7TkQ@mail.gmail.com","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T07:19:39Z","receivedAt":"2022-03-10T07:19:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> This is modeled after whitespace_rule_names.  What if I change this to\n> the following?\n> static const struct fsync_component_name {\n> ...\n> } fsync_component_names[]\n\nThat's better.\n\n> At the conclusion of this series, I defined 'default' as an aggregate\n> option that includes\n> the platform default.  I'd prefer not to have any statefulness of the\n> core.fsync setting so\n> that there is less confusion about the final fsync configuration.\n\nThen scratch your preference ;-) \n\nOur configuration files are designed to be \"hierarchical\" in that\nsystem-wide default can be set in /etc/gitconfig, which can be\noverridden by per-user default in $HOME/.gitconfig, which can in\nturn be overridden by per-repository setting in .git/config, so\nstarting from the compiled-in default, reading/augmenting \"the value\nwe tentatively decided based on what we have read so far\" with what\nwe read from lower-precedence files to higher-precedence files is a\nnorm.\n\nDon't make this little corner of the system different from\neverything else; that will only confuse users.\n\nThe git_config() callback should expect to see the same var with\ndifferent values for that reason.  Always restarting from zero will\ndefeat it.\n\nAnd always restarting from zero will mean \"none\" is meaningless,\nwhile it would be a quite nice way to say \"forget everything we have\nread so far and start from scratch\" when you really want to refuse\nwhat the system default wants to give you.\n\n>> > @@ -1613,7 +1684,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>> >       }\n>> >\n>> >       if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n>> > -             fsync_object_files = git_config_bool(var, value);\n>> > +             warning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n>>\n>> This is not deprecating but removing the support, which I am not\n>> sure is a sensible thing to do.  Rather we should pretend that\n>> core.fsync = \"loose-object\" (or \"-loose-object\") were found in the\n>> configuration, shouldn't we?\n>>\n>\n> The problem I anticipate is that figuring out the final configuration\n> becomes pretty\n> complex in the face of conflicting configurations of fsyncObjectFiles\n> and core.fsync.\n\nDon't start your thinking from too complex configuration that mixes\nand matches.  Instead, think what happens to folks who are *NOT*\ninterested in the new way of doing this.  They aren't interested in\nsetting core.fsync, and they already have core.fsyncObjectFiles set.\nYou want to make sure their experience does not suck.\n\nThe quoted code simply _IGNORES_ their wish and forces whatever\ndefault configuration core.fsync codepath happens to choose, which\nis a grave regression from their point of view.\n\nOne way to handle this more gracefully is to delay the final\ndecision until the end of the configuraiton file processing, and\nkeep track of core.fsyncObjectFiles and core.fsync separately.  If\nthe latter is never set but the former is, then you are dealing with\nsuch a user who hasn't migrated.  Give them a warning (the text\nabove is fine---we can tell them \"that's deprecated and you should\nuse the other one instead\"), but in the meantime, until deprecation\nturns into removal of support, keep honoring their original wish.\nIf core.fsync is set to something, you can still give them a warning\nwhen you see core.fsyncObjectFiles, saying \"that's deprecated and\nbecause you have core.fsync, we'll ignore the old one\", and use the\nnew method exclusively, without having to worry about mixing.\n"},{"id":"450987","messageId":"cover.1646905589.git.ps@pks.im","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"Future-proofed syncing of refs","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-10T09:53:13Z","receivedAt":"2022-03-10T09:53:55Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi,\n\nthese three patches apply on top of Neeraj's v5 of his \"A design for\nfuture-proofing fsync() configuration\". I'm sending this as a reply to\nhis v5 to keep the discussion in one place -- I think ultimately, we may\nwant to merge both series into a single one anyway. But please shout at\nme if this is considered bad style and I'll split it out into a separate\nthread.\n\nIn any case, these three patches implement fsyncing for loose and packed\nreferences using the proposed `core.fsync` option, with three additional\nknobs:\n\n    - \"loose-ref\" will fsync loose references.\n    - \"packed-refs\" will fsync packed references.\n    - \"refs\" will fsync all references, which should ideally also\n      include all new backends like the reftable backend.\n\nI think this extension demonstrates that the design proposed by Neeraj\nis quite easy to extend without too much boilerplate.\n\nPatrick\n\nPatrick Steinhardt (3):\n  core.fsync: add `fsync_component()` wrapper which doesn't die\n  core.fsync: new option to harden loose references\n  core.fsync: new option to harden packed references\n\n Documentation/config/core.txt |  3 +++\n cache.h                       | 22 ++++++++++++++++++----\n config.c                      |  3 +++\n refs/files-backend.c          | 29 +++++++++++++++++++++++++++++\n refs/packed-backend.c         |  3 ++-\n write-or-die.c                | 10 ++++++----\n 6 files changed, 61 insertions(+), 9 deletions(-)\n\n-- \n2.35.1\n\n"},{"id":"450988","messageId":"50e39f698a7c0cc06d3bc060e6dbc539ea693241.1646905589.git.ps@pks.im","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH 6/8] core.fsync: add `fsync_component()` wrapper which doesn't die","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-10T09:53:17Z","receivedAt":"2022-03-10T09:53:57Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"We have a `fsync_component_or_die()` helper function which only syncs\nchanges to disk in case the corresponding config is enabled by the user.\nThis wrapper will always die on an error though, which makes it\ninsufficient for new callsites we are about to add.\n\nAdd a new `fsync_component()` wrapper which returns an error code\ninstead of dying.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n cache.h        | 13 ++++++++++---\n write-or-die.c | 10 ++++++----\n 2 files changed, 16 insertions(+), 7 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex f307e89516..63a95d1977 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1745,12 +1745,19 @@ int copy_file(const char *dst, const char *src, int mode);\n int copy_file_with_time(const char *dst, const char *src, int mode);\n \n void write_or_die(int fd, const void *buf, size_t count);\n-void fsync_or_die(int fd, const char *);\n+int maybe_fsync(int fd);\n+\n+static inline int fsync_component(enum fsync_component component, int fd)\n+{\n+\tif (fsync_components & component)\n+\t\treturn maybe_fsync(fd);\n+\treturn 0;\n+}\n \n static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n {\n-\tif (fsync_components & component)\n-\t\tfsync_or_die(fd, msg);\n+\tif (fsync_component(component, fd) < 0)\n+\t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n ssize_t read_in_full(int fd, void *buf, size_t count);\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 9faa5f9f56..4a5455ce46 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -56,19 +56,21 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \t}\n }\n \n-void fsync_or_die(int fd, const char *msg)\n+int maybe_fsync(int fd)\n {\n \tif (use_fsync < 0)\n \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n \tif (!use_fsync)\n-\t\treturn;\n+\t\treturn 0;\n \n \tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n \t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n-\t\treturn;\n+\t\treturn 0;\n \n \tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n-\t\tdie_errno(\"fsync error on '%s'\", msg);\n+\t\treturn -1;\n+\n+\treturn 0;\n }\n \n void write_or_die(int fd, const void *buf, size_t count)\n-- \n2.35.1\n\n"},{"id":"450989","messageId":"f1e8a7bb3bf0f4c0414819cb1d5579dc08fd2a4f.1646905589.git.ps@pks.im","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH 7/8] core.fsync: new option to harden loose references","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-10T09:53:21Z","receivedAt":"2022-03-10T09:54:00Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When writing loose references to disk we first create a lockfile, write\nthe updated value of the reference into that lockfile, and on commit we\nrename the file into place. According to filesystem developers, this\nbehaviour is broken because applications should always sync data to disk\nbefore doing the final rename to ensure data consistency [1][2][3]. If\napplications fail to do this correctly, a hard crash of the machine can\neasily result in corrupted on-disk data.\n\nThis kind of corruption can in fact be easily observed with Git when the\nmachine hard-crashes shortly after writing loose references to disk. On\nmachines with ext4, this will likely lead to the \"empty files\" problem:\nthe file has been renamed, but its data has not been synced to disk. The\nresult is that the references is corrupt, and in the worst case this can\nlead to data loss.\n\nImplement a new option to harden loose references so that users and\nadmins can avoid this scenario by syncing locked loose references to\ndisk before we rename them into place.\n\n[1]: https://thunk.org/tytso/blog/2009/03/15/dont-fear-the-fsync/\n[2]: https://btrfs.wiki.kernel.org/index.php/FAQ (What are the crash guarantees of overwrite-by-rename)\n[3]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/Documentation/admin-guide/ext4.rst (see auto_da_alloc)\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Documentation/config/core.txt |  2 ++\n cache.h                       |  6 +++++-\n config.c                      |  2 ++\n refs/files-backend.c          | 29 +++++++++++++++++++++++++++++\n 4 files changed, 38 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 973805e8a9..b67d3c340e 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -564,8 +564,10 @@ core.fsync::\n * `pack-metadata` hardens packfile bitmaps and indexes.\n * `commit-graph` hardens the commit graph file.\n * `index` hardens the index when it is modified.\n+* `loose-ref` hardens references modified in the repo in loose-ref form.\n * `objects` is an aggregate option that is equivalent to\n   `loose-object,pack`.\n+* `refs` is an aggregate option that is equivalent to `loose-ref`.\n * `derived-metadata` is an aggregate option that is equivalent to\n   `pack-metadata,commit-graph`.\n * `default` is an aggregate option that is equivalent to\ndiff --git a/cache.h b/cache.h\nindex 63a95d1977..b56a56f539 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1005,11 +1005,14 @@ enum fsync_component {\n \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n+\tFSYNC_COMPONENT_LOOSE_REF\t\t= 1 << 5,\n };\n \n #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n \t\t\t\t  FSYNC_COMPONENT_PACK)\n \n+#define FSYNC_COMPONENTS_REFS (FSYNC_COMPONENT_LOOSE_REF)\n+\n #define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n \t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n \n@@ -1026,7 +1029,8 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_PACK | \\\n \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n \t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n-\t\t\t      FSYNC_COMPONENT_INDEX)\n+\t\t\t      FSYNC_COMPONENT_INDEX | \\\n+\t\t\t      FSYNC_COMPONENT_LOOSE_REF)\n \n /*\n  * A bitmask indicating which components of the repo should be fsynced.\ndiff --git a/config.c b/config.c\nindex f03f27c3de..b5d3e6e404 100644\n--- a/config.c\n+++ b/config.c\n@@ -1332,7 +1332,9 @@ static const struct fsync_component_entry {\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n+\t{ \"loose-ref\", FSYNC_COMPONENT_LOOSE_REF },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"refs\", FSYNC_COMPONENTS_REFS },\n \t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n \t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n \t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\ndiff --git a/refs/files-backend.c b/refs/files-backend.c\nindex f59589d6cc..279316de45 100644\n--- a/refs/files-backend.c\n+++ b/refs/files-backend.c\n@@ -1392,6 +1392,15 @@ static int refs_rename_ref_available(struct ref_store *refs,\n \treturn ok;\n }\n \n+static int files_sync_loose_ref(struct ref_lock *lock, struct strbuf *err)\n+{\n+\tint ret = fsync_component(FSYNC_COMPONENT_LOOSE_REF, get_lock_file_fd(&lock->lk));\n+\tif (ret)\n+\t\tstrbuf_addf(err, \"could not sync loose ref '%s': %s\", lock->ref_name,\n+\t\t\t    strerror(errno));\n+\treturn ret;\n+}\n+\n static int files_copy_or_rename_ref(struct ref_store *ref_store,\n \t\t\t    const char *oldrefname, const char *newrefname,\n \t\t\t    const char *logmsg, int copy)\n@@ -1504,6 +1513,7 @@ static int files_copy_or_rename_ref(struct ref_store *ref_store,\n \toidcpy(&lock->old_oid, &orig_oid);\n \n \tif (write_ref_to_lockfile(lock, &orig_oid, 0, &err) ||\n+\t    files_sync_loose_ref(lock, &err) ||\n \t    commit_ref_update(refs, lock, &orig_oid, logmsg, &err)) {\n \t\terror(\"unable to write current sha1 into %s: %s\", newrefname, err.buf);\n \t\tstrbuf_release(&err);\n@@ -1524,6 +1534,7 @@ static int files_copy_or_rename_ref(struct ref_store *ref_store,\n \tflag = log_all_ref_updates;\n \tlog_all_ref_updates = LOG_REFS_NONE;\n \tif (write_ref_to_lockfile(lock, &orig_oid, 0, &err) ||\n+\t    files_sync_loose_ref(lock, &err) ||\n \t    commit_ref_update(refs, lock, &orig_oid, NULL, &err)) {\n \t\terror(\"unable to write current sha1 into %s: %s\", oldrefname, err.buf);\n \t\tstrbuf_release(&err);\n@@ -2819,6 +2830,24 @@ static int files_transaction_prepare(struct ref_store *ref_store,\n \t\t}\n \t}\n \n+\t/*\n+\t * Sync all lockfiles to disk to ensure data consistency. We do this in\n+\t * a separate step such that we can sync all modified refs in a single\n+\t * step, which may be more efficient on some filesystems.\n+\t */\n+\tfor (i = 0; i < transaction->nr; i++) {\n+\t\tstruct ref_update *update = transaction->updates[i];\n+\t\tstruct ref_lock *lock = update->backend_data;\n+\n+\t\tif (!(update->flags & REF_NEEDS_COMMIT))\n+\t\t\tcontinue;\n+\n+\t\tif (files_sync_loose_ref(lock, err)) {\n+\t\t\tret  = TRANSACTION_GENERIC_ERROR;\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t}\n+\n cleanup:\n \tfree(head_ref);\n \tstring_list_clear(&affected_refnames, 0);\n-- \n2.35.1\n\n"},{"id":"450990","messageId":"3b81d8f5aeffb73a32b0bff0da947f023a3df517.1646905589.git.ps@pks.im","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH 8/8] core.fsync: new option to harden packed references","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-10T09:53:25Z","receivedAt":"2022-03-10T09:54:05Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Similar to the preceding commit, this commit adds a new option to harden\npacked references so that users and admins can avoid data loss when we\ncommit a new packed-refs file.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Documentation/config/core.txt | 3 ++-\n cache.h                       | 7 +++++--\n config.c                      | 1 +\n refs/packed-backend.c         | 3 ++-\n 4 files changed, 10 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex b67d3c340e..3fd466f955 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -565,9 +565,10 @@ core.fsync::\n * `commit-graph` hardens the commit graph file.\n * `index` hardens the index when it is modified.\n * `loose-ref` hardens references modified in the repo in loose-ref form.\n+* `packed-refs` hardens references modified in the repo in packed-refs form.\n * `objects` is an aggregate option that is equivalent to\n   `loose-object,pack`.\n-* `refs` is an aggregate option that is equivalent to `loose-ref`.\n+* `refs` is an aggregate option that is equivalent to `loose-ref,packed-refs`.\n * `derived-metadata` is an aggregate option that is equivalent to\n   `pack-metadata,commit-graph`.\n * `default` is an aggregate option that is equivalent to\ndiff --git a/cache.h b/cache.h\nindex b56a56f539..9b7c282fa5 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1006,12 +1006,14 @@ enum fsync_component {\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n \tFSYNC_COMPONENT_LOOSE_REF\t\t= 1 << 5,\n+\tFSYNC_COMPONENT_PACKED_REFS\t\t= 1 << 6,\n };\n \n #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n \t\t\t\t  FSYNC_COMPONENT_PACK)\n \n-#define FSYNC_COMPONENTS_REFS (FSYNC_COMPONENT_LOOSE_REF)\n+#define FSYNC_COMPONENTS_REFS (FSYNC_COMPONENT_LOOSE_REF | \\\n+\t\t\t       FSYNC_COMPONENT_PACKED_REFS)\n \n #define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n \t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n@@ -1030,7 +1032,8 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n \t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n \t\t\t      FSYNC_COMPONENT_INDEX | \\\n-\t\t\t      FSYNC_COMPONENT_LOOSE_REF)\n+\t\t\t      FSYNC_COMPONENT_LOOSE_REF | \\\n+\t\t\t      FSYNC_COMPONENT_PACKED_REFS)\n \n /*\n  * A bitmask indicating which components of the repo should be fsynced.\ndiff --git a/config.c b/config.c\nindex b5d3e6e404..b4a2ee3a8c 100644\n--- a/config.c\n+++ b/config.c\n@@ -1333,6 +1333,7 @@ static const struct fsync_component_entry {\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n \t{ \"loose-ref\", FSYNC_COMPONENT_LOOSE_REF },\n+\t{ \"packed-refs\", FSYNC_COMPONENT_PACKED_REFS },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n \t{ \"refs\", FSYNC_COMPONENTS_REFS },\n \t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 27dd8c3922..32d6635969 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -1262,7 +1262,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\tgoto error;\n \t}\n \n-\tif (close_tempfile_gently(refs->tempfile)) {\n+\tif (fsync_component(FSYNC_COMPONENT_PACKED_REFS, get_tempfile_fd(refs->tempfile)) ||\n+\t    close_tempfile_gently(refs->tempfile)) {\n \t\tstrbuf_addf(err, \"error closing file %s: %s\",\n \t\t\t    get_tempfile_path(refs->tempfile),\n \t\t\t    strerror(errno));\n-- \n2.35.1\n\n"},{"id":"450992","messageId":"YinwACkic3X1DKdr@ncase","threadId":"57030","inReplyTo":"xmqqpmmuki68.fsf@gitster.g","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-10T12:33:04Z","receivedAt":"2022-03-10T12:33:20Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Mar 09, 2022 at 12:03:11PM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> >> If the user cares about fsynching loose object files in the right\n> >> way, we shouldn't leave loose ref files not following the safe\n> >> safety level, regardless of how this new core.fsync knobs would look\n> >> like.\n> >> \n> >> I think we three are in agreement on that.\n> >\n> > Is there anything I can specifically do to help out with this topic? We\n> > have again hit data loss in production because we don't sync loose refs\n> > to disk before renaming them into place, so I'd really love to sort out\n> > this issue somehow so that I can revive my patch series which fixes the\n> > known repository corruption [1].\n> \n> How about doing a series to unconditionally sync loose ref creation\n> and modification?\n> \n> Alternatively, we could link it to the existing configuration to\n> control synching of object files.\n> \n> I do not think core.fsyncObjectFiles having \"object\" in its name is\n> a good reason not to think those who set it to true only care about\n> the loose object files and nothing else.  It is more sensible to\n> consider that those who set it to true cares about the repository\n> integrity more than those who set it to false, I would think.\n> \n> But that (i.e. doing it conditionally and choose which knob to use)\n> is one extra thing that needs justification, so starting from\n> unconditional fsync_or_die() may be the best way to ease it in.\n\nI'd be happy to revive my old patch series, but this kind of depends on\nwhere we're standing with the other patch series by Neeraj. If you say\nthat we'll likely not land his patch series for the upcoming release,\nbut a small patch series which only starts to sync loose refs may have a\nchance, then I'd like to go down that path as a stop-gap solution.\nOtherwise it probably wouldn't make a lot of sense.\n\nPatrick\n"},{"id":"450993","messageId":"nycvar.QRO.7.76.6.2203101406000.357@tvgsbejvaqbjf.bet","threadId":"57030","inReplyTo":"CANQDOddU_WXD-6ncDGBrgpsuKT-XDGC=SeaaQTNQFdODFZ7TkQ@mail.gmail.com","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-03-10T13:11:25Z","receivedAt":"2022-03-10T13:11:59Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Neeraj,\n\nOn Wed, 9 Mar 2022, Neeraj Singh wrote:\n\n> On Wed, Mar 9, 2022 at 4:21 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> > I am wondering if fsync_or_die() interface is abstracted well enough,\n> > or we need things like \"the fd is inside this directory; in addition\n> > to doing the fsync of the fd, please sync the parent directory as\n> > well\" support before we start adding more components (if there is such\n> > a need, perhaps it comes before this step).\n> >\n>\n> I think syncing the parent directory is a separate fsyncMethod that\n> would require changes across the codebase to obtain an appropriate\n> directory fd. I'd prefer to treat that as a separable concern.\n\nThat makes sense to me because I expect further abstraction to be\nnecessary here because Unix/Linux semantics differ quite a bit more from\nWindows semantics when it comes to directory \"file\" descriptors than when\ntalking about files' file descriptors.\n\n> > > +#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n> > > +                               FSYNC_COMPONENT_PACK_METADATA | \\\n> > > +                               FSYNC_COMPONENT_COMMIT_GRAPH)\n> >\n> > IOW, everything other than loose object, which already has a\n> > separate core.fsyncObjectFiles knob to loosen.  Everything else we\n> > currently sync unconditionally and the default keeps that\n> > arrangement?\n> >\n>\n> Yes, trying to keep default behavior identical on non-Windows\n> platforms.  Windows will be expected to adopt batch mode, and have\n> loose objects in this set.\n\nWe already adopted an early version of this patch series:\nhttps://github.com/git-for-windows/git/commit/98209a5f6e4\n\nAnd yes, I will gladly adapt that to whatever lands in core Git.\n\nThank you _so_ much for working on this!\nDscho\n"},{"id":"451038","messageId":"xmqq8rthhgpx.fsf@gitster.g","threadId":"57030","inReplyTo":"YinwACkic3X1DKdr@ncase","subject":"Re: [PATCH v4 2/4] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T17:15:06Z","receivedAt":"2022-03-10T17:15:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> How about doing a series to unconditionally sync loose ref creation\n>> and modification?\n>> \n>> Alternatively, we could link it to the existing configuration to\n>> control synching of object files.\n>> \n>> I do not think core.fsyncObjectFiles having \"object\" in its name is\n>> a good reason not to think those who set it to true only care about\n>> the loose object files and nothing else.  It is more sensible to\n>> consider that those who set it to true cares about the repository\n>> integrity more than those who set it to false, I would think.\n>> \n>> But that (i.e. doing it conditionally and choose which knob to use)\n>> is one extra thing that needs justification, so starting from\n>> unconditional fsync_or_die() may be the best way to ease it in.\n>\n> I'd be happy to revive my old patch series, but this kind of depends on\n> where we're standing with the other patch series by Neeraj. If you say\n> that we'll likely not land his patch series for the upcoming release,\n> but a small patch series which only starts to sync loose refs may have a\n> chance, then I'd like to go down that path as a stop-gap solution.\n> Otherwise it probably wouldn't make a lot of sense.\n\nThe above was what I wrote before the revived series from Neeraj.\nNow I've seen it, and more importantly, the most recent one from you\non top of that to add ref hardening as a new \"component\" or two, I\nlike the overall shape of the end result (except for the semantics\nimplemented by the configuration parser, which can be fixed without\naffecting how the hardening components are implemented).  Hopefully\nthe base series of Neeraj can be solidified soon enough to make it\nunnecessary for a stop-gap measure.  We'll see.\n\nThanks.\n"},{"id":"451040","messageId":"xmqq4k45hgk3.fsf@gitster.g","threadId":"57030","inReplyTo":"CANQDOddU_WXD-6ncDGBrgpsuKT-XDGC=SeaaQTNQFdODFZ7TkQ@mail.gmail.com","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T17:18:36Z","receivedAt":"2022-03-10T17:18:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n>> I am wondering if fsync_or_die() interface is abstracted well\n>> enough, or we need things like \"the fd is inside this directory; in\n>> addition to doing the fsync of the fd, please sync the parent\n>> directory as well\" support before we start adding more components\n>> (if there is such a need, perhaps it comes before this step).\n>>\n>\n> I think syncing the parent directory is a separate fsyncMethod that\n> would require changes across the codebase to obtain an appropriate\n> directory fd. I'd prefer to treat that as a separable concern.\n\nYeah, that would be a sensible direction to go.  If we never did the\n\"sync the parent\" thing, we do not need it in the fsyncMethod world\nimmediately.  It can be added later.\n\nThanks.\n\n"},{"id":"451041","messageId":"xmqqwnh1g19q.fsf@gitster.g","threadId":"57030","inReplyTo":"50e39f698a7c0cc06d3bc060e6dbc539ea693241.1646905589.git.ps@pks.im","subject":"Re: [PATCH 6/8] core.fsync: add `fsync_component()` wrapper which doesn't die","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T17:34:09Z","receivedAt":"2022-03-10T17:34:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> We have a `fsync_component_or_die()` helper function which only syncs\n> changes to disk in case the corresponding config is enabled by the user.\n> This wrapper will always die on an error though, which makes it\n> insufficient for new callsites we are about to add.\n\nYou can replace \"which makes it ...\" part with a bit more concrete\ndescription to save suspense from the readers.\n\n    fsync_component_or_die() that dies upon an error is not useful\n    for callers with their own error handling or recovery logic,\n    like ref transaction API.\n\n    Split fsync_component() out that returns an error to help them.\n\n> -void fsync_or_die(int fd, const char *);\n> +int maybe_fsync(int fd);\n> ...\n> +static inline int fsync_component(enum fsync_component component, int fd)\n> +{\n> +\tif (fsync_components & component)\n> +\t\treturn maybe_fsync(fd);\n> +\treturn 0;\n> +}\n>  \n>  static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n>  {\n> -\tif (fsync_components & component)\n> -\t\tfsync_or_die(fd, msg);\n> +\tif (fsync_component(component, fd) < 0)\n> +\t\tdie_errno(\"fsync error on '%s'\", msg);\n>  }\n\nI think in the eventuall reroll, these \"static inline\" functions on\nthe I/O code path will become real functions in write-or-die.c but\nother than that this reorganization looks sensible.\n\nThanks.\n\n> diff --git a/write-or-die.c b/write-or-die.c\n> index 9faa5f9f56..4a5455ce46 100644\n> --- a/write-or-die.c\n> +++ b/write-or-die.c\n> @@ -56,19 +56,21 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n>  \t}\n>  }\n>  \n> -void fsync_or_die(int fd, const char *msg)\n> +int maybe_fsync(int fd)\n>  {\n>  \tif (use_fsync < 0)\n>  \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n>  \tif (!use_fsync)\n> -\t\treturn;\n> +\t\treturn 0;\n>  \n>  \tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n>  \t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n> -\t\treturn;\n> +\t\treturn 0;\n>  \n>  \tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n> -\t\tdie_errno(\"fsync error on '%s'\", msg);\n> +\t\treturn -1;\n> +\n> +\treturn 0;\n>  }\n>  \n>  void write_or_die(int fd, const void *buf, size_t count)\n"},{"id":"451055","messageId":"xmqq4k45fyvc.fsf@gitster.g","threadId":"57030","inReplyTo":"f1e8a7bb3bf0f4c0414819cb1d5579dc08fd2a4f.1646905589.git.ps@pks.im","subject":"Re: [PATCH 7/8] core.fsync: new option to harden loose references","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T18:25:59Z","receivedAt":"2022-03-10T18:26:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index 973805e8a9..b67d3c340e 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -564,8 +564,10 @@ core.fsync::\n>  * `pack-metadata` hardens packfile bitmaps and indexes.\n>  * `commit-graph` hardens the commit graph file.\n>  * `index` hardens the index when it is modified.\n> +* `loose-ref` hardens references modified in the repo in loose-ref form.\n>  * `objects` is an aggregate option that is equivalent to\n>    `loose-object,pack`.\n> +* `refs` is an aggregate option that is equivalent to `loose-ref`.\n\nAggregate of one feels strange.  I do not see a strong reason to\nhave two separate classes loose vs packed and allow them be\nprotected to different robustness, but that aside, if we were to\nhave two, this \"aggregate\" is better added when the second one is.\n\nHaving said that, given that a separate ref backend that has no\ndistinction between loose or packed is on horizen, I think we would\nrather prefer to see a single \"ref\" component that governs all\nbackends.\n\n> @@ -1026,7 +1029,8 @@ enum fsync_component {\n>  \t\t\t      FSYNC_COMPONENT_PACK | \\\n>  \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n>  \t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n> -\t\t\t      FSYNC_COMPONENT_INDEX)\n> +\t\t\t      FSYNC_COMPONENT_INDEX | \\\n> +\t\t\t      FSYNC_COMPONENT_LOOSE_REF)\n\nOK.\n\n> +static int files_sync_loose_ref(struct ref_lock *lock, struct strbuf *err)\n\nThis file-scope static function will not be in the vtable (in other\nwords, it is not like \"sync_loose_ref\" method must be defined across\nall ref backends and this function is called as the implementation\nof the method for the files backend), so we do not have to give it\n\"files_\" prefix if we do not want to.\n\nsync_loose_ref() may be easier to read, perhaps?\n\n> +{\n> +\tint ret = fsync_component(FSYNC_COMPONENT_LOOSE_REF, get_lock_file_fd(&lock->lk));\n> +\tif (ret)\n> +\t\tstrbuf_addf(err, \"could not sync loose ref '%s': %s\", lock->ref_name,\n> +\t\t\t    strerror(errno));\n> +\treturn ret;\n> +}\n\nOK.\n\nGood illustration how the new helper in 6/8 is useful.  It would be\nnice if the reroll of the base topic by Neeraj reorders the patches\nto introduce fsync_component() much earlier, at the same time it\nintroduces the fsync_component_or_die().\n\n> @@ -1504,6 +1513,7 @@ static int files_copy_or_rename_ref(struct ref_store *ref_store,\n>  \toidcpy(&lock->old_oid, &orig_oid);\n>  \n>  \tif (write_ref_to_lockfile(lock, &orig_oid, 0, &err) ||\n> +\t    files_sync_loose_ref(lock, &err) ||\n>  \t    commit_ref_update(refs, lock, &orig_oid, logmsg, &err)) {\n>  \t\terror(\"unable to write current sha1 into %s: %s\", newrefname, err.buf);\n>  \t\tstrbuf_release(&err);\n> @@ -1524,6 +1534,7 @@ static int files_copy_or_rename_ref(struct ref_store *ref_store,\n>  \tflag = log_all_ref_updates;\n>  \tlog_all_ref_updates = LOG_REFS_NONE;\n>  \tif (write_ref_to_lockfile(lock, &orig_oid, 0, &err) ||\n> +\t    files_sync_loose_ref(lock, &err) ||\n>  \t    commit_ref_update(refs, lock, &orig_oid, NULL, &err)) {\n>  \t\terror(\"unable to write current sha1 into %s: %s\", oldrefname, err.buf);\n>  \t\tstrbuf_release(&err);\n\nWe used to skip commit_ref_update() and gone to the error code path\nupon write failure, and we do the same upon fsync failure now.  The\nerror code path will do the rollback the same way as before.\n\nOK, these look sensible.\n\n> @@ -2819,6 +2830,24 @@ static int files_transaction_prepare(struct ref_store *ref_store,\n>  \t\t}\n>  \t}\n>  \n> +\t/*\n> +\t * Sync all lockfiles to disk to ensure data consistency. We do this in\n> +\t * a separate step such that we can sync all modified refs in a single\n> +\t * step, which may be more efficient on some filesystems.\n> +\t */\n> +\tfor (i = 0; i < transaction->nr; i++) {\n> +\t\tstruct ref_update *update = transaction->updates[i];\n> +\t\tstruct ref_lock *lock = update->backend_data;\n> +\n> +\t\tif (!(update->flags & REF_NEEDS_COMMIT))\n> +\t\t\tcontinue;\n> +\n> +\t\tif (files_sync_loose_ref(lock, err)) {\n> +\t\t\tret  = TRANSACTION_GENERIC_ERROR;\n> +\t\t\tgoto cleanup;\n> +\t\t}\n> +\t}\n\nAn obvious alternative that naïvely comes to mind is to keep going\nafter seeing the first failure to sync and attempt to sync all the\nrest, but remember that we had an error.  But I think what this\npatch does makes a lot more sense, as the error code path will just\ncleans the transaction up.  If any of them fails, there is no reason\nto spend more effort.\n\nThanks.\n\n>  cleanup:\n>  \tfree(head_ref);\n>  \tstring_list_clear(&affected_refnames, 0);\n"},{"id":"451056","messageId":"xmqqzglxek5y.fsf@gitster.g","threadId":"57030","inReplyTo":"3b81d8f5aeffb73a32b0bff0da947f023a3df517.1646905589.git.ps@pks.im","subject":"Re: [PATCH 8/8] core.fsync: new option to harden packed references","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T18:28:57Z","receivedAt":"2022-03-10T18:29:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> diff --git a/refs/packed-backend.c b/refs/packed-backend.c\n> index 27dd8c3922..32d6635969 100644\n> --- a/refs/packed-backend.c\n> +++ b/refs/packed-backend.c\n> @@ -1262,7 +1262,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n>  \t\tgoto error;\n>  \t}\n>  \n> -\tif (close_tempfile_gently(refs->tempfile)) {\n> +\tif (fsync_component(FSYNC_COMPONENT_PACKED_REFS, get_tempfile_fd(refs->tempfile)) ||\n> +\t    close_tempfile_gently(refs->tempfile)) {\n>  \t\tstrbuf_addf(err, \"error closing file %s: %s\",\n>  \t\t\t    get_tempfile_path(refs->tempfile),\n>  \t\t\t    strerror(errno));\n\nI do not necessarily agree with the organization to have it as a\ncomponent that is separate from other ref backends, but it is\nvery pleasing to see that there is only one fsync necessary for the\npacked backend.\n\nNice.\n"},{"id":"451057","messageId":"CANQDOdcDbYHyRuJj0hV_LcYPJdkoJjF_EGN4CXpndc4VQ9dVAA@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqcziugtpw.fsf@gitster.g","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T18:38:45Z","receivedAt":"2022-03-10T18:39:01Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 9, 2022 at 11:19 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > This is modeled after whitespace_rule_names.  What if I change this to\n> > the following?\n> > static const struct fsync_component_name {\n> > ...\n> > } fsync_component_names[]\n>\n> That's better.\n>\n> > At the conclusion of this series, I defined 'default' as an aggregate\n> > option that includes\n> > the platform default.  I'd prefer not to have any statefulness of the\n> > core.fsync setting so\n> > that there is less confusion about the final fsync configuration.\n>\n> Then scratch your preference ;-)\n\nJust to clarify, linguistically, by 'scratch' do you mean that I should drop\nmy preference (which I can do to get the important parts of the\nseries in), or are you saying that that your suggestion is in line with\nmy preference, and I'm just not seeing it properly?\n\n> Our configuration files are designed to be \"hierarchical\" in that\n> system-wide default can be set in /etc/gitconfig, which can be\n> overridden by per-user default in $HOME/.gitconfig, which can in\n> turn be overridden by per-repository setting in .git/config, so\n> starting from the compiled-in default, reading/augmenting \"the value\n> we tentatively decided based on what we have read so far\" with what\n> we read from lower-precedence files to higher-precedence files is a\n> norm.\n>\n> Don't make this little corner of the system different from\n> everything else; that will only confuse users.\n>\n> The git_config() callback should expect to see the same var with\n> different values for that reason.  Always restarting from zero will\n> defeat it.\n>\n\nConsider core.whitespace. The parse_whitespace_rule code starts with\nthe compiled-in default every time it encounters a config value, and then\nmodifies it according to what the user passed in.  So the user could figure\nout what's going to happen by just looking at the value returned by\n`git config --get core.whitespace`.  The user doesn't need to call\n`git config --get-all core.whitespace` and then reason about the entries\nfrom top to bottom to figure out the actual state that Git will use.\n\n> And always restarting from zero will mean \"none\" is meaningless,\n> while it would be a quite nice way to say \"forget everything we have\n> read so far and start from scratch\" when you really want to refuse\n> what the system default wants to give you.\n>\n\nThe intention, which I've implemented in my local v6 changes, is for an\nempty list or empty string to be an illegal value of core.fsync.  It should be\nset to one or more legal values.  But I see the advantage in always resetting to\nthe system default, like core.whitespace does, so that a set of unrecognized\nvalues results in at least default behavior. An empty string would mean\n'unconfigured', which will be meaningful when we integrate core.fsync with\ncore.fsyncObjectFiles.\n\nI'll update the change to start from default, with none as a reset. I'm still in\nfavor of making it so that the most recent value of core.fsync in the\nhierarchical configuration store stands alone without picking up state from\nprior values.\n\n> >> > @@ -1613,7 +1684,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >> >       }\n> >> >\n> >> >       if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> >> > -             fsync_object_files = git_config_bool(var, value);\n> >> > +             warning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n> >>\n> >> This is not deprecating but removing the support, which I am not\n> >> sure is a sensible thing to do.  Rather we should pretend that\n> >> core.fsync = \"loose-object\" (or \"-loose-object\") were found in the\n> >> configuration, shouldn't we?\n> >>\n> >\n> > The problem I anticipate is that figuring out the final configuration\n> > becomes pretty\n> > complex in the face of conflicting configurations of fsyncObjectFiles\n> > and core.fsync.\n>\n> Don't start your thinking from too complex configuration that mixes\n> and matches.  Instead, think what happens to folks who are *NOT*\n> interested in the new way of doing this.  They aren't interested in\n> setting core.fsync, and they already have core.fsyncObjectFiles set.\n> You want to make sure their experience does not suck.\n>\n> The quoted code simply _IGNORES_ their wish and forces whatever\n> default configuration core.fsync codepath happens to choose, which\n> is a grave regression from their point of view.\n>\n> One way to handle this more gracefully is to delay the final\n> decision until the end of the configuraiton file processing, and\n> keep track of core.fsyncObjectFiles and core.fsync separately.  If\n> the latter is never set but the former is, then you are dealing with\n> such a user who hasn't migrated.  Give them a warning (the text\n> above is fine---we can tell them \"that's deprecated and you should\n> use the other one instead\"), but in the meantime, until deprecation\n> turns into removal of support, keep honoring their original wish.\n> If core.fsync is set to something, you can still give them a warning\n> when you see core.fsyncObjectFiles, saying \"that's deprecated and\n> because you have core.fsync, we'll ignore the old one\", and use the\n> new method exclusively, without having to worry about mixing.\n\nIs there a well-defined place where we know that configuration processing\nis complete?  The most obvious spot to me to integrate these two values would\nbe the first time we need to figure out the fsync state. Another spot would be\nprepare_repo_settings.  Are there any other good candidates?\n\nOnce the right spot is picked, I'll implement the integration of the\nsettings as you\nsuggested.  For now I'll stick it in prepare_repo_settings.\n\nThanks for the review.  Please expect a v6 today.\n"},{"id":"451058","messageId":"CANQDOdeK8CZmBqaWZgY17qrfPbwMHvq+=CYS_nsRHdg6aarwEA@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqwnh1g19q.fsf@gitster.g","subject":"Re: [PATCH 6/8] core.fsync: add `fsync_component()` wrapper which doesn't die","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T18:40:29Z","receivedAt":"2022-03-10T18:40:45Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 10, 2022 at 9:34 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Patrick Steinhardt <ps@pks.im> writes:\n>\n> > We have a `fsync_component_or_die()` helper function which only syncs\n> > changes to disk in case the corresponding config is enabled by the user.\n> > This wrapper will always die on an error though, which makes it\n> > insufficient for new callsites we are about to add.\n>\n> You can replace \"which makes it ...\" part with a bit more concrete\n> description to save suspense from the readers.\n>\n>     fsync_component_or_die() that dies upon an error is not useful\n>     for callers with their own error handling or recovery logic,\n>     like ref transaction API.\n>\n>     Split fsync_component() out that returns an error to help them.\n>\n> > -void fsync_or_die(int fd, const char *);\n> > +int maybe_fsync(int fd);\n> > ...\n> > +static inline int fsync_component(enum fsync_component component, int fd)\n> > +{\n> > +     if (fsync_components & component)\n> > +             return maybe_fsync(fd);\n> > +     return 0;\n> > +}\n> >\n> >  static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n> >  {\n> > -     if (fsync_components & component)\n> > -             fsync_or_die(fd, msg);\n> > +     if (fsync_component(component, fd) < 0)\n> > +             die_errno(\"fsync error on '%s'\", msg);\n> >  }\n>\n> I think in the eventuall reroll, these \"static inline\" functions on\n> the I/O code path will become real functions in write-or-die.c but\n> other than that this reorganization looks sensible.\n>\n> Thanks.\n>\n\nYes, that will be part of v6 of the base changeset.\n\n> > diff --git a/write-or-die.c b/write-or-die.c\n> > index 9faa5f9f56..4a5455ce46 100644\n> > --- a/write-or-die.c\n> > +++ b/write-or-die.c\n> > @@ -56,19 +56,21 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n> >       }\n> >  }\n> >\n> > -void fsync_or_die(int fd, const char *msg)\n> > +int maybe_fsync(int fd)\n> >  {\n> >       if (use_fsync < 0)\n> >               use_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n> >       if (!use_fsync)\n> > -             return;\n> > +             return 0;\n> >\n> >       if (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n> >           git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n> > -             return;\n> > +             return 0;\n> >\n> >       if (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n> > -             die_errno(\"fsync error on '%s'\", msg);\n> > +             return -1;\n> > +\n> > +     return 0;\n> >  }\n> >\n> >  void write_or_die(int fd, const void *buf, size_t count)\n"},{"id":"451060","messageId":"xmqqv8wlejgc.fsf@gitster.g","threadId":"57030","inReplyTo":"CANQDOdcDbYHyRuJj0hV_LcYPJdkoJjF_EGN4CXpndc4VQ9dVAA@mail.gmail.com","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T18:44:19Z","receivedAt":"2022-03-10T18:44:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n>> > At the conclusion of this series, I defined 'default' as an aggregate\n>> > option that includes\n>> > the platform default.  I'd prefer not to have any statefulness of the\n>> > core.fsync setting so\n>> > that there is less confusion about the final fsync configuration.\n>>\n>> Then scratch your preference ;-)\n>\n> Just to clarify, linguistically, by 'scratch' do you mean that I should drop\n> my preference\n\nYes.\n\n> Is there a well-defined place where we know that configuration processing\n> is complete?  The most obvious spot to me to integrate these two values would\n> be the first time we need to figure out the fsync state.\n\nThat sounds like a good place.\n"},{"id":"451066","messageId":"CANQDOdeOu3iWqQr_m0vL0DAfDAaGgFc6eeHP2112Veh8yNu=Gw@mail.gmail.com","threadId":"57030","inReplyTo":"xmqq4k45fyvc.fsf@gitster.g","subject":"Re: [PATCH 7/8] core.fsync: new option to harden loose references","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T19:03:47Z","receivedAt":"2022-03-10T19:04:04Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"> Good illustration how the new helper in 6/8 is useful.  It would be\n> nice if the reroll of the base topic by Neeraj reorders the patches\n> to introduce fsync_component() much earlier, at the same time it\n> introduces the fsync_component_or_die().\n\nWill do.\n"},{"id":"451073","messageId":"xmqqtuc5d1hp.fsf@gitster.g","threadId":"57030","inReplyTo":"xmqqv8wlejgc.fsf@gitster.g","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T19:57:38Z","receivedAt":"2022-03-10T19:57:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n>>> > At the conclusion of this series, I defined 'default' as an aggregate\n>>> > option that includes\n>>> > the platform default.  I'd prefer not to have any statefulness of the\n>>> > core.fsync setting so\n>>> > that there is less confusion about the final fsync configuration.\n>>>\n>>> Then scratch your preference ;-)\n>>\n>> Just to clarify, linguistically, by 'scratch' do you mean that I should drop\n>> my preference\n>\n> Yes.\n\nLet me take this part back.\n\nI do not mind too deeply if this were \"each occurrence of core.fsync\nas a whole replaces whatever we saw earlier, i.e. last-one-wins\".\n\nBut if we were going that route, instead of starting from an empty\nset, I'd prefer to see it begin with the built-in default (i.e. the\none you defined to mimic the traditional behaviour before core.fsync\nwas introduced) and added or deleted by each (possibly '-' prefixed)\nelement on the comma-separated list, with an explicit way to clear\nthe built-in default.  E.g. \"none,refs\" would clear the components\ntraditionally fsync'ed by default and choose only \"refs\" component,\nwhile \"-pack-metadata\" would mean the default ones minus\n\"pack-metadata\" component are subject for fsync'ing.  An empty\nstring would naturally mean \"By having this core.fsync entry, I am\ntelling you not to pay any attention to what lower-precedence\nconfiguration files said.  But I want the built-in default, without\nany additions or subtractions made by this entry, just the default,\nplease\" in such a scheme, so do not forbid it.\n\nOr, we can inherit from the previous configuration file to allow\n/etc/gitconfig and the ones shipped by Git for Windows to augment\nthe built-in default before letting end-user configuration to\nfurther customize the preference.\n\nEither is fine by me.\n\nThanks.\n"},{"id":"451075","messageId":"CANQDOdfvADcZ5Ng876EO=W2w+6ROvkWAm=XaAOtSUnxV5GGAXA@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqtuc5d1hp.fsf@gitster.g","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T20:25:20Z","receivedAt":"2022-03-10T20:25:37Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 10, 2022 at 11:57 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > Neeraj Singh <nksingh85@gmail.com> writes:\n> >\n> >>> > At the conclusion of this series, I defined 'default' as an aggregate\n> >>> > option that includes\n> >>> > the platform default.  I'd prefer not to have any statefulness of the\n> >>> > core.fsync setting so\n> >>> > that there is less confusion about the final fsync configuration.\n> >>>\n> >>> Then scratch your preference ;-)\n> >>\n> >> Just to clarify, linguistically, by 'scratch' do you mean that I should drop\n> >> my preference\n> >\n> > Yes.\n>\n> Let me take this part back.\n>\n> I do not mind too deeply if this were \"each occurrence of core.fsync\n> as a whole replaces whatever we saw earlier, i.e. last-one-wins\".\n>\n> But if we were going that route, instead of starting from an empty\n> set, I'd prefer to see it begin with the built-in default (i.e. the\n> one you defined to mimic the traditional behaviour before core.fsync\n> was introduced) and added or deleted by each (possibly '-' prefixed)\n> element on the comma-separated list, with an explicit way to clear\n> the built-in default.  E.g. \"none,refs\" would clear the components\n> traditionally fsync'ed by default and choose only \"refs\" component,\n> while \"-pack-metadata\" would mean the default ones minus\n> \"pack-metadata\" component are subject for fsync'ing.  An empty\n> string would naturally mean \"By having this core.fsync entry, I am\n> telling you not to pay any attention to what lower-precedence\n> configuration files said.  But I want the built-in default, without\n> any additions or subtractions made by this entry, just the default,\n> please\" in such a scheme, so do not forbid it.\n>\n> Or, we can inherit from the previous configuration file to allow\n> /etc/gitconfig and the ones shipped by Git for Windows to augment\n> the built-in default before letting end-user configuration to\n> further customize the preference.\n>\n> Either is fine by me.\n>\n> Thanks.\n\nOkay, I'll implement this version, since this is close to my preference.\n\nUnder this schema, 'default' isn't useful as an aggregate option, so\nI'll eliminate\nthat.\n"},{"id":"451078","messageId":"xmqqk0d1cxsv.fsf@gitster.g","threadId":"57030","inReplyTo":"CANQDOdfvADcZ5Ng876EO=W2w+6ROvkWAm=XaAOtSUnxV5GGAXA@mail.gmail.com","subject":"Re: [PATCH v5 3/5] core.fsync: introduce granular fsync control","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T21:17:20Z","receivedAt":"2022-03-10T21:17:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> Under this schema, 'default' isn't useful as an aggregate option, so\n> I'll eliminate\n> that.\n\nYeah, the only difference is the starting point.  Either start with\ndefault set of bits and give an option to clear, or start with an\nempty set and give an option to set default.  The former may be a\nbit less cumbersome to users but there isn't a huge difference.\n\n"},{"id":"451088","messageId":"825079b6aa14c3975254a336e0e313b963bc9ab4.1646952205.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"[PATCH v6 1/6] wrapper: make inclusion of Windows csprng header tightly scoped","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-10T22:43:19Z","receivedAt":"2022-03-10T22:43:33Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIncluding NTSecAPI.h in git-compat-util.h causes build errors in any\nother file that includes winternl.h. NTSecAPI.h was included in order to\nget access to the RtlGenRandom cryptographically secure PRNG. This\nchange scopes the inclusion of ntsecapi.h to wrapper.c, which is the only\nplace that it's actually needed.\n\nThe build breakage is due to the definition of UNICODE_STRING in\nNtSecApi.h:\n    #ifndef _NTDEF_\n    typedef LSA_UNICODE_STRING UNICODE_STRING, *PUNICODE_STRING;\n    typedef LSA_STRING STRING, *PSTRING ;\n    #endif\n\nLsaLookup.h:\n    typedef struct _LSA_UNICODE_STRING {\n        USHORT Length;\n        USHORT MaximumLength;\n    #ifdef MIDL_PASS\n        [size_is(MaximumLength/2), length_is(Length/2)]\n    #endif // MIDL_PASS\n        PWSTR  Buffer;\n    } LSA_UNICODE_STRING, *PLSA_UNICODE_STRING;\n\nwinternl.h also defines UNICODE_STRING:\n    typedef struct _UNICODE_STRING {\n        USHORT Length;\n        USHORT MaximumLength;\n        PWSTR  Buffer;\n    } UNICODE_STRING;\n    typedef UNICODE_STRING *PUNICODE_STRING;\n\nBoth definitions have equivalent layouts. Apparently these internal\nWindows headers aren't designed to be included together. This is\nan oversight in the headers and does not represent an incompatibility\nbetween the APIs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/winansi.c  | 5 -----\n git-compat-util.h | 6 ------\n wrapper.c         | 7 +++++++\n 3 files changed, 7 insertions(+), 11 deletions(-)\n\ndiff --git a/compat/winansi.c b/compat/winansi.c\nindex 936a80a5f00..3abe8dd5a27 100644\n--- a/compat/winansi.c\n+++ b/compat/winansi.c\n@@ -4,11 +4,6 @@\n \n #undef NOGDI\n \n-/*\n- * Including the appropriate header file for RtlGenRandom causes MSVC to see a\n- * redefinition of types in an incompatible way when including headers below.\n- */\n-#undef HAVE_RTLGENRANDOM\n #include \"../git-compat-util.h\"\n #include <wingdi.h>\n #include <winreg.h>\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 876907b9df4..d210cff058c 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -197,12 +197,6 @@\n #endif\n #include <windows.h>\n #define GIT_WINDOWS_NATIVE\n-#ifdef HAVE_RTLGENRANDOM\n-/* This is required to get access to RtlGenRandom. */\n-#define SystemFunction036 NTAPI SystemFunction036\n-#include <NTSecAPI.h>\n-#undef SystemFunction036\n-#endif\n #endif\n \n #include <unistd.h>\ndiff --git a/wrapper.c b/wrapper.c\nindex 3258cdb171f..1108e4840a4 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -4,6 +4,13 @@\n #include \"cache.h\"\n #include \"config.h\"\n \n+#ifdef HAVE_RTLGENRANDOM\n+/* This is required to get access to RtlGenRandom. */\n+#define SystemFunction036 NTAPI SystemFunction036\n+#include <NTSecAPI.h>\n+#undef SystemFunction036\n+#endif\n+\n static int memory_limit_check(size_t size, int gentle)\n {\n \tstatic size_t limit = 0;\n-- \ngitgitgadget\n\n"},{"id":"451089","messageId":"64e2bdcdfd92ba643cee52605155c21d116710a7.1646952205.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"[PATCH v6 3/6] core.fsync: introduce granular fsync control infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-10T22:43:21Z","receivedAt":"2022-03-10T22:43:35Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the infrastructure for the core.fsync\nconfiguration knob. The repository components we want to sync\nare identified by flags so that we can turn on or off syncing\nfor specific components.\n\nIf core.fsyncObjectFiles is set and the core.fsync configuration\nalso includes FSYNC_COMPONENT_LOOSE_OBJECT, we will fsync any\nloose objects. This picks the strictest data integrity behavior\nif core.fsync and core.fsyncObjectFiles are set to conflicting values.\n\nThis change introduces the currently unused fsync_component\nhelper, which will be used by a later patch that adds fsyncing to\nthe refs backend.\n\nActual configuration and documentation of the fsync components\nlist are in other patches in the series to separate review of\nthe underlying mechanism from the policy of how it's configured.\n\nHelped-by: Patrick Steinhardt <ps@pks.im>\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/fast-import.c  |  2 +-\n builtin/index-pack.c   |  4 ++--\n builtin/pack-objects.c | 24 +++++++++++++++++-------\n bulk-checkin.c         |  5 +++--\n cache.h                | 23 +++++++++++++++++++++++\n commit-graph.c         |  3 ++-\n csum-file.c            |  5 +++--\n csum-file.h            |  3 ++-\n environment.c          |  1 +\n midx.c                 |  3 ++-\n object-file.c          | 13 +++++++++----\n pack-bitmap-write.c    |  3 ++-\n pack-write.c           | 13 +++++++------\n read-cache.c           |  2 +-\n write-or-die.c         | 26 ++++++++++++++++++++++----\n 15 files changed, 97 insertions(+), 33 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex b7105fcad9b..f2c036a8955 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -865,7 +865,7 @@ static void end_packfile(void)\n \t\tstruct tag *t;\n \n \t\tclose_pack_windows(pack_data);\n-\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, 0);\n+\t\tfinalize_hashfile(pack_file, cur_pack_oid.hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->hash,\n \t\t\t\t\t pack_data->pack_name, object_count,\n \t\t\t\t\t cur_pack_oid.hash, pack_size);\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c45273de3b1..c5f12f14df5 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1290,7 +1290,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\t\t    nr_objects - nr_objects_initial);\n \t\tstop_progress_msg(&progress, msg.buf);\n \t\tstrbuf_release(&msg);\n-\t\tfinalize_hashfile(f, tail_hash, 0);\n+\t\tfinalize_hashfile(f, tail_hash, FSYNC_COMPONENT_PACK, 0);\n \t\thashcpy(read_hash, pack_hash);\n \t\tfixup_pack_header_footer(output_fd, pack_hash,\n \t\t\t\t\t curr_pack, nr_objects,\n@@ -1512,7 +1512,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tif (!from_stdin) {\n \t\tclose(input_fd);\n \t} else {\n-\t\tfsync_or_die(output_fd, curr_pack_name);\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_PACK, output_fd, curr_pack_name);\n \t\terr = close(output_fd);\n \t\tif (err)\n \t\t\tdie_errno(_(\"error while closing pack file\"));\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 178e611f09d..c14fee8e99f 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1199,16 +1199,26 @@ static void write_pack_file(void)\n \t\t\tdisplay_progress(progress_state, written);\n \t\t}\n \n-\t\t/*\n-\t\t * Did we write the wrong # entries in the header?\n-\t\t * If so, rewrite it like in fast-import\n-\t\t */\n \t\tif (pack_to_stdout) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n+\t\t\t/*\n+\t\t\t * We never fsync when writing to stdout since we may\n+\t\t\t * not be writing to an actual pack file. For instance,\n+\t\t\t * the upload-pack code passes a pipe here. Calling\n+\t\t\t * fsync on a pipe results in unnecessary\n+\t\t\t * synchronization with the reader on some platforms.\n+\t\t\t */\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_NONE,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE);\n \t\t} else if (nr_written == nr_remaining) {\n-\t\t\tfinalize_hashfile(f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\t\tfinalize_hashfile(f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t\t} else {\n-\t\t\tint fd = finalize_hashfile(f, hash, 0);\n+\t\t\t/*\n+\t\t\t * If we wrote the wrong number of entries in the\n+\t\t\t * header, rewrite it like in fast-import.\n+\t\t\t */\n+\n+\t\t\tint fd = finalize_hashfile(f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\t\tfixup_pack_header_footer(fd, hash, pack_tmp_name,\n \t\t\t\t\t\t nr_written, hash, offset);\n \t\t\tclose(fd);\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..a2cf9dcbc8d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -53,9 +53,10 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \t\tunlink(state->pack_tmp_name);\n \t\tgoto clear_exit;\n \t} else if (state->nr_written == 1) {\n-\t\tfinalize_hashfile(state->f, hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\t\tfinalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK,\n+\t\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \t} else {\n-\t\tint fd = finalize_hashfile(state->f, hash, 0);\n+\t\tint fd = finalize_hashfile(state->f, hash, FSYNC_COMPONENT_PACK, 0);\n \t\tfixup_pack_header_footer(fd, hash, state->pack_tmp_name,\n \t\t\t\t\t state->nr_written, hash,\n \t\t\t\t\t state->offset);\ndiff --git a/cache.h b/cache.h\nindex 82f0194a3dd..7ac1959258d 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -993,6 +993,27 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n+/*\n+ * These values are used to help identify parts of a repository to fsync.\n+ * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n+ * repository and so shouldn't be fsynced.\n+ */\n+enum fsync_component {\n+\tFSYNC_COMPONENT_NONE,\n+\tFSYNC_COMPONENT_LOOSE_OBJECT\t\t= 1 << 0,\n+\tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n+\tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n+\tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+};\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+/*\n+ * A bitmask indicating which components of the repo should be fsynced.\n+ */\n+extern enum fsync_component fsync_components;\n extern int fsync_object_files;\n extern int use_fsync;\n \n@@ -1707,6 +1728,8 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n \n void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n+int fsync_component(enum fsync_component component, int fd);\n+void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 265c010122e..64897f57d9f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1942,7 +1942,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t}\n \n \tclose_commit_graph(ctx->r->objects);\n-\tfinalize_hashfile(f, file_hash, CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n+\tfinalize_hashfile(f, file_hash, FSYNC_COMPONENT_COMMIT_GRAPH,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n \tfree_chunkfile(cf);\n \n \tif (ctx->split) {\ndiff --git a/csum-file.c b/csum-file.c\nindex 26e8a6df44e..59ef3398ca2 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -58,7 +58,8 @@ static void free_hashfile(struct hashfile *f)\n \tfree(f);\n }\n \n-int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int flags)\n+int finalize_hashfile(struct hashfile *f, unsigned char *result,\n+\t\t      enum fsync_component component, unsigned int flags)\n {\n \tint fd;\n \n@@ -69,7 +70,7 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result, unsigned int fl\n \tif (flags & CSUM_HASH_IN_STREAM)\n \t\tflush(f, f->buffer, the_hash_algo->rawsz);\n \tif (flags & CSUM_FSYNC)\n-\t\tfsync_or_die(f->fd, f->name);\n+\t\tfsync_component_or_die(component, f->fd, f->name);\n \tif (flags & CSUM_CLOSE) {\n \t\tif (close(f->fd))\n \t\t\tdie_errno(\"%s: sha1 file error on close\", f->name);\ndiff --git a/csum-file.h b/csum-file.h\nindex 291215b34eb..0d29f528fbc 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -1,6 +1,7 @@\n #ifndef CSUM_FILE_H\n #define CSUM_FILE_H\n \n+#include \"cache.h\"\n #include \"hash.h\"\n \n struct progress;\n@@ -38,7 +39,7 @@ int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint *);\n struct hashfile *hashfd(int fd, const char *name);\n struct hashfile *hashfd_check(const char *name);\n struct hashfile *hashfd_throughput(int fd, const char *name, struct progress *tp);\n-int finalize_hashfile(struct hashfile *, unsigned char *, unsigned int);\n+int finalize_hashfile(struct hashfile *, unsigned char *, enum fsync_component, unsigned int);\n void hashwrite(struct hashfile *, const void *, unsigned int);\n void hashflush(struct hashfile *f);\n void crc32_begin(struct hashfile *);\ndiff --git a/environment.c b/environment.c\nindex 3e3620d759f..36ca5fb2e77 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -45,6 +45,7 @@ int pack_compression_level = Z_DEFAULT_COMPRESSION;\n int fsync_object_files;\n int use_fsync = -1;\n enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n+enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/midx.c b/midx.c\nindex 865170bad05..107365d2114 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -1438,7 +1438,8 @@ static int write_midx_internal(const char *object_dir,\n \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n \twrite_chunkfile(cf, &ctx);\n \n-\tfinalize_hashfile(f, midx_hash, CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n \tfree_chunkfile(cf);\n \n \tif (flags & MIDX_WRITE_REV_INDEX &&\ndiff --git a/object-file.c b/object-file.c\nindex 03bd6a3baf3..e3f0bf27ff1 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1849,11 +1849,16 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (!the_repository->objects->odb->will_destroy) {\n-\t\tif (fsync_object_files)\n-\t\t\tfsync_or_die(fd, \"loose object file\");\n-\t}\n+\tif (the_repository->objects->odb->will_destroy)\n+\t\tgoto out;\n \n+\tif (fsync_object_files > 0)\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\telse\n+\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n+\t\t\t\t       \"loose object file\");\n+\n+out:\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex cab3eaa2acd..cf681547f2e 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -719,7 +719,8 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \tif (options & BITMAP_OPT_HASH_CACHE)\n \t\twrite_hash_cache(f, index, index_nr);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC | CSUM_CLOSE);\n \n \tif (adjust_shared_perm(tmp_file.buf))\n \t\tdie_errno(\"unable to make temporary bitmap file readable\");\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a346..51812cb1299 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -159,9 +159,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t}\n \n \thashwrite(f, sha1, the_hash_algo->rawsz);\n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((opts->flags & WRITE_IDX_VERIFY)\n-\t\t\t\t    ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((opts->flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \treturn index_name;\n }\n \n@@ -281,8 +281,9 @@ const char *write_rev_file_order(const char *rev_name,\n \tif (rev_name && adjust_shared_perm(rev_name) < 0)\n \t\tdie(_(\"failed to make %s readable\"), rev_name);\n \n-\tfinalize_hashfile(f, NULL, CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n-\t\t\t\t    ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE |\n+\t\t\t  ((flags & WRITE_IDX_VERIFY) ? 0 : CSUM_FSYNC));\n \n \treturn rev_name;\n }\n@@ -390,7 +391,7 @@ void fixup_pack_header_footer(int pack_fd,\n \t\tthe_hash_algo->final_fn(partial_pack_hash, &old_hash_ctx);\n \tthe_hash_algo->final_fn(new_pack_hash, &new_hash_ctx);\n \twrite_or_die(pack_fd, new_pack_hash, the_hash_algo->rawsz);\n-\tfsync_or_die(pack_fd, pack_name);\n+\tfsync_component_or_die(FSYNC_COMPONENT_PACK, pack_fd, pack_name);\n }\n \n char *index_pack_lockfile(int ip_out, int *is_well_formed)\ndiff --git a/read-cache.c b/read-cache.c\nindex 79b9b99ebf7..df869691fd4 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -3089,7 +3089,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, CSUM_HASH_IN_STREAM);\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 9faa5f9f563..c4fd91b5b43 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -56,21 +56,39 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \t}\n }\n \n-void fsync_or_die(int fd, const char *msg)\n+static int maybe_fsync(int fd)\n {\n \tif (use_fsync < 0)\n \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n \tif (!use_fsync)\n-\t\treturn;\n+\t\treturn 0;\n \n \tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n \t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n-\t\treturn;\n+\t\treturn 0;\n+\n+\treturn git_fsync(fd, FSYNC_HARDWARE_FLUSH);\n+}\n \n-\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n+void fsync_or_die(int fd, const char *msg)\n+{\n+\tif (maybe_fsync(fd) < 0)\n \t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n+int fsync_component(enum fsync_component component, int fd)\n+{\n+\tif (fsync_components & component)\n+\t\treturn maybe_fsync(fd);\n+\treturn 0;\n+}\n+\n+void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n+{\n+\tif (fsync_components & component)\n+\t\tfsync_or_die(fd, msg);\n+}\n+\n void write_or_die(int fd, const void *buf, size_t count)\n {\n \tif (write_in_full(fd, buf, count) < 0) {\n-- \ngitgitgadget\n\n"},{"id":"451090","messageId":"6adc8dc13852c219763a9830f848fbc8663f2fa9.1646952205.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"[PATCH v6 4/6] core.fsync: add configuration parsing","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-10T22:43:22Z","receivedAt":"2022-03-10T22:43:42Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis change introduces code to parse the core.fsync setting and\nconfigure the fsync_components variable.\n\ncore.fsync is configured as a comma-separated list of component names to\nsync. Each time a core.fsync variable is encountered in the\nconfiguration heirarchy, we start off with a clean state with the\nplatform default value. Passing 'none' resets the value to indicate\nnothing will be synced. We gather all negative and positive entries from\nthe comma separated list and then compute the new value by removing all\nthe negative entries and adding all of the positive entries.\n\nWe issue a warning for components that are not recognized so that the\nconfiguration code is compatible with configs from future versions of\nGit with more repo components.\n\nComplete documentation for the new setting is included in a later patch\nin the series so that it can be reviewed once in final form.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  9 +++--\n config.c                      | 76 +++++++++++++++++++++++++++++++++++\n environment.c                 |  2 +-\n 3 files changed, 82 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex dbb134f7136..ab911d6e269 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -558,11 +558,12 @@ core.fsyncMethod::\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n+\tThis setting is deprecated. Use core.fsync instead.\n +\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+This setting affects data added to the Git repository in loose-object\n+form. When set to true, Git will issue an fsync or similar system call\n+to flush caches so that loose-objects remain consistent in the face\n+of a unclean system shutdown.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/config.c b/config.c\nindex f3ff80b01c9..94a4598b5fc 100644\n--- a/config.c\n+++ b/config.c\n@@ -1323,6 +1323,73 @@ static int git_parse_maybe_bool_text(const char *value)\n \treturn -1;\n }\n \n+static const struct fsync_component_name {\n+\tconst char *name;\n+\tenum fsync_component component_bits;\n+} fsync_component_names[] = {\n+\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n+\t{ \"pack\", FSYNC_COMPONENT_PACK },\n+\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n+\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+};\n+\n+static enum fsync_component parse_fsync_components(const char *var, const char *string)\n+{\n+\tenum fsync_component current = FSYNC_COMPONENTS_DEFAULT;\n+\tenum fsync_component positive = 0, negative = 0;\n+\n+\twhile (string) {\n+\t\tint i;\n+\t\tsize_t len;\n+\t\tconst char *ep;\n+\t\tint negated = 0;\n+\t\tint found = 0;\n+\n+\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n+\t\tep = strchrnul(string, ',');\n+\t\tlen = ep - string;\n+\t\tif (!strcmp(string, \"none\")) {\n+\t\t\tcurrent = FSYNC_COMPONENT_NONE;\n+\t\t\tgoto next_name;\n+\t\t}\n+\n+\t\tif (*string == '-') {\n+\t\t\tnegated = 1;\n+\t\t\tstring++;\n+\t\t\tlen--;\n+\t\t\tif (!len)\n+\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n+\t\t}\n+\n+\t\tif (!len)\n+\t\t\tbreak;\n+\n+\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_names); ++i) {\n+\t\t\tconst struct fsync_component_name *n = &fsync_component_names[i];\n+\n+\t\t\tif (strncmp(n->name, string, len))\n+\t\t\t\tcontinue;\n+\n+\t\t\tfound = 1;\n+\t\t\tif (negated)\n+\t\t\t\tnegative |= n->component_bits;\n+\t\t\telse\n+\t\t\t\tpositive |= n->component_bits;\n+\t\t}\n+\n+\t\tif (!found) {\n+\t\t\tchar *component = xstrndup(string, len);\n+\t\t\twarning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n+\t\t\tfree(component);\n+\t\t}\n+\n+next_name:\n+\t\tstring = ep;\n+\t}\n+\n+\treturn (current & ~negative) | positive;\n+}\n+\n int git_parse_maybe_bool(const char *value)\n {\n \tint v = git_parse_maybe_bool_text(value);\n@@ -1600,6 +1667,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsync\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tfsync_components = parse_fsync_components(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncmethod\")) {\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n@@ -1613,6 +1687,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n+\t\tif (fsync_object_files < 0)\n+\t\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\n \t}\ndiff --git a/environment.c b/environment.c\nindex 36ca5fb2e77..698f03a2f47 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,7 +42,7 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+int fsync_object_files = -1;\n int use_fsync = -1;\n enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n-- \ngitgitgadget\n\n"},{"id":"451091","messageId":"757f6d0bbd21550a9e8f4783eb0b3c351b12f930.1646952205.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"[PATCH v6 5/6] core.fsync: new option to harden the index","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-10T22:43:23Z","receivedAt":"2022-03-10T22:43:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the new ability for the user to harden\nthe index. In the event of a system crash, the index must be\ndurable for the user to actually find a file that has been added\nto the repo and then deleted from the working tree.\n\nWe use the presence of the COMMIT_LOCK flag and absence of the\nalternate_index_output as a proxy for determining whether we're\nupdating the persistent index of the repo or some temporary\nindex. We don't sync these temporary indexes.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache.h      |  1 +\n config.c     |  1 +\n read-cache.c | 19 +++++++++++++------\n 3 files changed, 15 insertions(+), 6 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 7ac1959258d..e08eeac6c15 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1004,6 +1004,7 @@ enum fsync_component {\n \tFSYNC_COMPONENT_PACK\t\t\t= 1 << 1,\n \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n+\tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n #define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\ndiff --git a/config.c b/config.c\nindex 94a4598b5fc..80f33c91982 100644\n--- a/config.c\n+++ b/config.c\n@@ -1331,6 +1331,7 @@ static const struct fsync_component_name {\n \t{ \"pack\", FSYNC_COMPONENT_PACK },\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n+\t{ \"index\", FSYNC_COMPONENT_INDEX },\n };\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\ndiff --git a/read-cache.c b/read-cache.c\nindex df869691fd4..7683b679258 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2842,7 +2842,7 @@ static int record_ieot(void)\n  * rely on it.\n  */\n static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n-\t\t\t  int strip_extensions)\n+\t\t\t  int strip_extensions, unsigned flags)\n {\n \tuint64_t start = getnanotime();\n \tstruct hashfile *f;\n@@ -2856,6 +2856,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint csum_fsync_flag;\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n@@ -3089,7 +3090,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\treturn -1;\n \t}\n \n-\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_NONE, CSUM_HASH_IN_STREAM);\n+\tcsum_fsync_flag = 0;\n+\tif (!alternate_index_output && (flags & COMMIT_LOCK))\n+\t\tcsum_fsync_flag = CSUM_FSYNC;\n+\n+\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n+\t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n+\n \tif (close_tempfile_gently(tempfile)) {\n \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n \t\treturn -1;\n@@ -3144,7 +3151,7 @@ static int do_write_locked_index(struct index_state *istate, struct lock_file *l\n \t */\n \ttrace2_region_enter_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n-\tret = do_write_index(istate, lock->tempfile, 0);\n+\tret = do_write_index(istate, lock->tempfile, 0, flags);\n \ttrace2_region_leave_printf(\"index\", \"do_write_index\", the_repository,\n \t\t\t\t   \"%s\", get_lock_file_path(lock));\n \n@@ -3238,7 +3245,7 @@ static int clean_shared_index_files(const char *current_hex)\n }\n \n static int write_shared_index(struct index_state *istate,\n-\t\t\t      struct tempfile **temp)\n+\t\t\t      struct tempfile **temp, unsigned flags)\n {\n \tstruct split_index *si = istate->split_index;\n \tint ret, was_full = !istate->sparse_index;\n@@ -3248,7 +3255,7 @@ static int write_shared_index(struct index_state *istate,\n \n \ttrace2_region_enter_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n-\tret = do_write_index(si->base, *temp, 1);\n+\tret = do_write_index(si->base, *temp, 1, flags);\n \ttrace2_region_leave_printf(\"index\", \"shared/do_write_index\",\n \t\t\t\t   the_repository, \"%s\", get_tempfile_path(*temp));\n \n@@ -3357,7 +3364,7 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n \t\t\tret = do_write_locked_index(istate, lock, flags);\n \t\t\tgoto out;\n \t\t}\n-\t\tret = write_shared_index(istate, &temp);\n+\t\tret = write_shared_index(istate, &temp, flags);\n \n \t\tsaved_errno = errno;\n \t\tif (is_tempfile_active(temp))\n-- \ngitgitgadget\n\n"},{"id":"451092","messageId":"a41bd4a06afcf9878d5941841519bf35ece60eac.1646952205.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"[PATCH v6 2/6] core.fsyncmethod: add writeout-only mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-10T22:43:20Z","receivedAt":"2022-03-10T22:43:47Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit introduces the `core.fsyncMethod` configuration\nknob, which can currently be set to `fsync` or `writeout-only`.\n\nThe new writeout-only mode attempts to tell the operating system to\nflush its in-memory page cache to the storage hardware without issuing a\nCACHE_FLUSH command to the storage controller.\n\nWriteout-only fsync is significantly faster than a vanilla fsync on\ncommon hardware, since data is written to a disk-side cache rather than\nall the way to a durable medium. Later changes in this patch series will\ntake advantage of this primitive to implement batching of hardware\nflushes.\n\nWhen git_fsync is called with FSYNC_WRITEOUT_ONLY, it may fail and the\ncaller is expected to do an ordinary fsync as needed.\n\nOn Apple platforms, the fsync system call does not issue a CACHE_FLUSH\ndirective to the storage controller. This change updates fsync to do\nfcntl(F_FULLFSYNC) to make fsync actually durable. We maintain parity\nwith existing behavior on Apple platforms by setting the default value\nof the new core.fsyncMethod option.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt       |  9 ++++\n Makefile                            |  6 +++\n cache.h                             |  7 ++++\n compat/mingw.h                      |  3 ++\n compat/win32/flush.c                | 28 +++++++++++++\n config.c                            | 12 ++++++\n config.mak.uname                    |  3 ++\n configure.ac                        |  8 ++++\n contrib/buildsystems/CMakeLists.txt | 16 ++++++--\n environment.c                       |  1 +\n git-compat-util.h                   | 24 +++++++++++\n wrapper.c                           | 64 +++++++++++++++++++++++++++++\n write-or-die.c                      | 11 +++--\n 13 files changed, 184 insertions(+), 8 deletions(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..dbb134f7136 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,15 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsyncMethod::\n+\tA value indicating the strategy Git will use to harden repository data\n+\tusing fsync and related primitives.\n++\n+* `fsync` uses the fsync() system call or platform equivalents.\n+* `writeout-only` issues pagecache writeback requests, but depending on the\n+  filesystem and storage hardware, data added to the repository may not be\n+  durable in the event of a system crash. This is the default mode on macOS.\n+\n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n +\ndiff --git a/Makefile b/Makefile\nindex 6f0b4b775fe..17fd9b023a4 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -411,6 +411,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1897,6 +1899,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/cache.h b/cache.h\nindex 04d4d2db25c..82f0194a3dd 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -995,6 +995,13 @@ extern char *git_replace_ref_base;\n \n extern int fsync_object_files;\n extern int use_fsync;\n+\n+enum fsync_method {\n+\tFSYNC_METHOD_FSYNC,\n+\tFSYNC_METHOD_WRITEOUT_ONLY\n+};\n+\n+extern enum fsync_method fsync_method;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..291f90ea940\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NTAPI, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.c b/config.c\nindex 383b1a4885b..f3ff80b01c9 100644\n--- a/config.c\n+++ b/config.c\n@@ -1600,6 +1600,18 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.fsyncmethod\")) {\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcmp(value, \"fsync\"))\n+\t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n+\t\telse if (!strcmp(value, \"writeout-only\"))\n+\t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse\n+\t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n+\n+\t}\n+\n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 4352ea39e9b..404fff5dd04 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -57,6 +57,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tBASIC_CFLAGS += -DHAVE_SYSINFO\n@@ -463,6 +464,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -640,6 +642,7 @@ ifeq ($(uname_S),MINGW)\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/configure.ac b/configure.ac\nindex 5ee25ec95c8..6bd6bef1c44 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1082,6 +1082,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex e44232f85d3..3a9e6241660 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,10 +261,18 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET HAVE_RTLGENRANDOM)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n-\t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n-\t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n-\t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\n+\tlist(APPEND compat_SOURCES\n+\t\tcompat/mingw.c\n+\t\tcompat/winansi.c\n+\t\tcompat/win32/flush.c\n+\t\tcompat/win32/path-utils.c\n+\t\tcompat/win32/pthread.c\n+\t\tcompat/win32mmap.c\n+\t\tcompat/win32/syslog.c\n+\t\tcompat/win32/trace2_win32_process_info.c\n+\t\tcompat/win32/dirent.c\n+\t\tcompat/nedmalloc/nedmalloc.c\n+\t\tcompat/strdup.c)\n \tset(NO_UNIX_SOCKETS 1)\n \n elseif(CMAKE_SYSTEM_NAME STREQUAL \"Linux\")\ndiff --git a/environment.c b/environment.c\nindex fd0501e77a5..3e3620d759f 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -44,6 +44,7 @@ int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n int fsync_object_files;\n int use_fsync = -1;\n+enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex d210cff058c..00356476a9d 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1271,6 +1271,30 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+#ifdef __APPLE__\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n+#else\n+#define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n+#endif\n+\n+enum fsync_action {\n+\tFSYNC_WRITEOUT_ONLY,\n+\tFSYNC_HARDWARE_FLUSH\n+};\n+\n+/*\n+ * Issues an fsync against the specified file according to the specified mode.\n+ *\n+ * FSYNC_WRITEOUT_ONLY attempts to use interfaces available on some operating\n+ * systems to flush the OS cache without issuing a flush command to the storage\n+ * controller. If those interfaces are unavailable, the function fails with\n+ * ENOSYS.\n+ *\n+ * FSYNC_HARDWARE_FLUSH does an OS writeout and hardware flush to ensure that\n+ * changes are durable. It is not expected to fail.\n+ */\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/wrapper.c b/wrapper.c\nindex 1108e4840a4..354d784c034 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -546,6 +546,70 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+/*\n+ * Some platforms return EINTR from fsync. Since fsync is invoked in some\n+ * cases by a wrapper that dies on failure, do not expose EINTR to callers.\n+ */\n+static int fsync_loop(int fd)\n+{\n+\tint err;\n+\n+\tdo {\n+\t\terr = fsync(fd);\n+\t} while (err < 0 && errno == EINTR);\n+\treturn err;\n+}\n+\n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * On macOS, fsync just causes filesystem cache writeback but\n+\t\t * does not flush hardware caches.\n+\t\t */\n+\t\treturn fsync_loop(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to\n+\t\t * issue a writeback without a hardware flush. An offset of\n+\t\t * 0 and size of 0 indicates writeout of the entire file and the\n+\t\t * wait flags ensure that all dirty data is written to the disk\n+\t\t * (potentially in a disk-side cache) before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\t\t/*\n+\t\t * On macOS, a special fcntl is required to really flush the\n+\t\t * caches within the storage controller. As of this writing,\n+\t\t * this is a very expensive operation on Apple SSDs.\n+\t\t */\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync_loop(fd);\n+#endif\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex a3d5784cec9..9faa5f9f563 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -62,10 +62,13 @@ void fsync_or_die(int fd, const char *msg)\n \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n \tif (!use_fsync)\n \t\treturn;\n-\twhile (fsync(fd) < 0) {\n-\t\tif (errno != EINTR)\n-\t\t\tdie_errno(\"fsync error on '%s'\", msg);\n-\t}\n+\n+\tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n+\t\treturn;\n+\n+\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n+\t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n void write_or_die(int fd, const void *buf, size_t count)\n-- \ngitgitgadget\n\n"},{"id":"451093","messageId":"7e4cc6e10a5d88f4c6c44efaa68f2325007fd935.1646952205.git.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"[PATCH v6 6/6] core.fsync: documentation and user-friendly aggregate options","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-10T22:43:24Z","receivedAt":"2022-03-10T22:43:50Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds aggregate options for the core.fsync setting that are\nmore user-friendly. These options are specified in terms of 'levels of\nsafety', indicating which Git operations are considered to be sync\npoints for durability.\n\nThe new documentation is also included here in its entirety for ease of\nreview.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 40 +++++++++++++++++++++++++++++++++++\n cache.h                       | 23 +++++++++++++++++---\n config.c                      |  5 +++++\n 3 files changed, 65 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex ab911d6e269..37105a7be40 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,46 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsync::\n+\tA comma-separated list of components of the repository that\n+\tshould be hardened via the core.fsyncMethod when created or\n+\tmodified.  You can disable hardening of any component by\n+\tprefixing it with a '-'.  Items that are not hardened may be\n+\tlost in the event of an unclean\tsystem shutdown. Unless you\n+\thave special requirements, it is recommended that you leave\n+\tthis option empty or pick one of `committed`, `added`,\n+\tor `all`.\n++\n+When this configuration is encountered, the set of components starts with\n+the platform default value, disabled components are removed, and additional\n+components are added. `none` resets the state so that the platform default\n+is ignored.\n++\n+The empty string resets the fsync configuration to the platform\n+default. The platform default on most platform is equivalent to\n+`core.fsync=committed,-loose-object`, which has good performance,\n+but risks losing recent work in the event of an unclean system shutdown.\n++\n+* `none` clears the set of fsynced components.\n+* `loose-object` hardens objects added to the repo in loose-object form.\n+* `pack` hardens objects added to the repo in packfile form.\n+* `pack-metadata` hardens packfile bitmaps and indexes.\n+* `commit-graph` hardens the commit graph file.\n+* `index` hardens the index when it is modified.\n+* `objects` is an aggregate option that is equivalent to\n+  `loose-object,pack`.\n+* `derived-metadata` is an aggregate option that is equivalent to\n+  `pack-metadata,commit-graph`.\n+* `committed` is an aggregate option that is currently equivalent to\n+  `objects`. This mode sacrifices some performance to ensure that work\n+  that is committed to the repository with `git commit` or similar commands\n+  is hardened.\n+* `added` is an aggregate option that is currently equivalent to\n+  `committed,index`. This mode sacrifices additional performance to\n+  ensure that the results of commands like `git add` and similar operations\n+  are hardened.\n+* `all` is an aggregate option that syncs all individual components above.\n+\n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\ndiff --git a/cache.h b/cache.h\nindex e08eeac6c15..86680f144ec 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1007,9 +1007,26 @@ enum fsync_component {\n \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n-#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n-\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK)\n+\n+#define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENTS_OBJECTS | \\\n+\t\t\t\t  FSYNC_COMPONENTS_DERIVED_METADATA | \\\n+\t\t\t\t  ~FSYNC_COMPONENT_LOOSE_OBJECT)\n+\n+#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS)\n+\n+#define FSYNC_COMPONENTS_ADDED (FSYNC_COMPONENTS_COMMITTED | \\\n+\t\t\t\tFSYNC_COMPONENT_INDEX)\n+\n+#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t      FSYNC_COMPONENT_PACK | \\\n+\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n+\t\t\t      FSYNC_COMPONENT_INDEX)\n \n /*\n  * A bitmask indicating which components of the repo should be fsynced.\ndiff --git a/config.c b/config.c\nindex 80f33c91982..fd8e1659312 100644\n--- a/config.c\n+++ b/config.c\n@@ -1332,6 +1332,11 @@ static const struct fsync_component_name {\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n+\t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\n+\t{ \"added\", FSYNC_COMPONENTS_ADDED },\n+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n };\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n-- \ngitgitgadget\n"},{"id":"451094","messageId":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","threadId":"57030","inReplyTo":"pull.1093.v5.git.1646866998.gitgitgadget@gmail.com","subject":"[PATCH v6 0/6] A design for future-proofing fsync() configuration","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-10T22:43:18Z","receivedAt":"2022-03-10T22:43:52Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"This is an implementation of an extensible configuration mechanism for\nfsyncing persistent components of a repo.\n\nThe main goals are to separate the \"what\" to sync from the \"how\". There are\nnow two settings: core.fsync - Control the 'what', including the index.\ncore.fsyncMethod - Control the 'how'. Currently we support writeout-only and\nfull fsync.\n\nSyncing of refs can be layered on top of core.fsync. And batch mode will be\nlayered on core.fsyncMethod. Once this series reaches 'seen', I'll submit\nns/batched-fsync to introduce batch mode. Please see\nhttps://github.com/gitgitgadget/git/pull/1134.\n\ncore.fsyncObjectfiles is removed and we will issue a deprecation warning if\nit's seen.\n\nI'd like to get agreement on this direction before submitting batch mode to\nthe list. The batch mode series is available to view at\n\nPlease see [1], [2], and [3] for discussions that led to this series.\n\nAfter this change, new persistent data files added to the repo will need to\nbe added to the fsync_component enum and documented in the\nDocumentation/config/core.txt text.\n\nV6 changes:\n\n * Move only the windows csprng includes into wrapper.c rather than all of\n   them. This fixes the specific build issue due to broken Windows headers.\n   [6]\n * Split the configuration parsing of core.fsync from the mechanism to focus\n   the review.\n * Incorporate Patrick's patch at [7] into the core.fsync mechanism patch.\n * Pick the stricter one of core.fsyncObjectFiles and (fsync_components &\n   FSYNC_COMPONENT_LOOSE_OBJECTS), to respect the older setting.\n * Issue a deprecation warning but keep parsing and honoring\n   core.fsyncObjectFiles.\n * Change configuration parsing of core.fsync to always start with the\n   platform default. none resets to the empty set. The comma separated list\n   implies a set without regards to ordering now. This follows Junio's\n   suggestion in [8].\n * Change the documentation of the core.fsync option to reflect the way the\n   new parsing code works.\n * The patch 7 and 8 of Patrick's series at [7] can be cherry-picked after\n   being applied to ns/core-fsyncmethod.\n\nV5 changes:\n\n * Rebase onto main at c2162907e9\n * Add a patch to move CSPRNG platform includes to wrapper.c. This avoids\n   build errors in compat/win32/flush.c and other files.\n * Move the documentation and aggregate options to the final patch in the\n   series.\n * Define new aggregate options and guidance in line with Junio's suggestion\n   to present the user with 'levels of safety' rather than a morass of\n   detailed options.\n\nV4 changes:\n\n * Rebase onto master at b23dac905bd.\n * Add a comment to write_pack_file indicating why we don't fsync when\n   writing to stdout.\n * I kept the configuration schema as-is rather than switching to\n   multi-value. The thinking here is that a stateless last-one-wins config\n   schema (comma separated) will make it easier to achieve some holistic\n   self-consistent fsync configuration for a particular repo.\n\nV3 changes:\n\n * Remove relative path from git-compat-util.h include [4].\n * Updated newly added warning texts to have more context for localization\n   [4].\n * Fixed tab spacing in enum fsync_action\n * Moved the fsync looping out to a helper and do it consistently. [4]\n * Changed commit description to use camelCase for config names. [5]\n * Add an optional fourth patch with derived-metadata so that the user can\n   exclude a forward-compatible set of things that should be recomputable\n   given existing data.\n\nV2 changes:\n\n * Updated the documentation for core.fsyncmethod to be less certain.\n   writeout-only probably does not do the right thing on Linux.\n * Split out the core.fsync=index change into its own commit.\n * Rename REPO_COMPONENT to FSYNC_COMPONENT. This is really specific to\n   fsyncing, so the name should reflect that.\n * Re-add missing Makefile change for SYNC_FILE_RANGE.\n * Tested writeout-only mode, index syncing, and general config settings.\n\n[1] https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n[2]\nhttps://lore.kernel.org/git/dd65718814011eb93ccc4428f9882e0f025224a6.1636029491.git.ps@pks.im/\n[3]\nhttps://lore.kernel.org/git/pull.1076.git.git.1629856292.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/CANQDOdf8C4-haK9=Q_J4Cid8bQALnmGDm=SvatRbaVf+tkzqLw@mail.gmail.com/\n[5] https://lore.kernel.org/git/211207.861r2opplg.gmgdl@evledraar.gmail.com/\n[6]\nhttps://lore.kernel.org/git/CANQDOdfZbOHZQt9Ah0t1AamTO2T7Gq0tmWX1jLqL6njE0LF6DA@mail.gmail.com/\n[7]\nhttps://lore.kernel.org/git/50e39f698a7c0cc06d3bc060e6dbc539ea693241.1646905589.git.ps@pks.im/\n[8] https://lore.kernel.org/git/xmqqk0d1cxsv.fsf@gitster.g/\n\nNeeraj Singh (6):\n  wrapper: make inclusion of Windows csprng header tightly scoped\n  core.fsyncmethod: add writeout-only mode\n  core.fsync: introduce granular fsync control infrastructure\n  core.fsync: add configuration parsing\n  core.fsync: new option to harden the index\n  core.fsync: documentation and user-friendly aggregate options\n\n Documentation/config/core.txt       | 58 ++++++++++++++++--\n Makefile                            |  6 ++\n builtin/fast-import.c               |  2 +-\n builtin/index-pack.c                |  4 +-\n builtin/pack-objects.c              | 24 +++++---\n bulk-checkin.c                      |  5 +-\n cache.h                             | 48 +++++++++++++++\n commit-graph.c                      |  3 +-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n compat/winansi.c                    |  5 --\n config.c                            | 94 +++++++++++++++++++++++++++++\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt | 16 +++--\n csum-file.c                         |  5 +-\n csum-file.h                         |  3 +-\n environment.c                       |  4 +-\n git-compat-util.h                   | 30 +++++++--\n midx.c                              |  3 +-\n object-file.c                       | 13 ++--\n pack-bitmap-write.c                 |  3 +-\n pack-write.c                        | 13 ++--\n read-cache.c                        | 19 ++++--\n wrapper.c                           | 71 ++++++++++++++++++++++\n write-or-die.c                      | 33 ++++++++--\n 26 files changed, 444 insertions(+), 60 deletions(-)\n create mode 100644 compat/win32/flush.c\n\n\nbase-commit: c2162907e9aa884bdb70208389cb99b181620d51\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1093%2Fneerajsi-msft%2Fns%2Fcore-fsync-v6\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1093/neerajsi-msft/ns/core-fsync-v6\nPull-Request: https://github.com/gitgitgadget/git/pull/1093\n\nRange-diff vs v5:\n\n 1:  685b1db8880 ! 1:  825079b6aa1 wrapper: move inclusion of CSPRNG headers the wrapper.c file\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    wrapper: move inclusion of CSPRNG headers the wrapper.c file\n     +    wrapper: make inclusion of Windows csprng header tightly scoped\n      \n          Including NTSecAPI.h in git-compat-util.h causes build errors in any\n     -    other file that includes winternl.h. That file was included in order to\n     +    other file that includes winternl.h. NTSecAPI.h was included in order to\n          get access to the RtlGenRandom cryptographically secure PRNG. This\n     -    change scopes the inclusion of all PRNG headers to just the wrapper.c\n     -    file, which is the only place it is really needed.\n     +    change scopes the inclusion of ntsecapi.h to wrapper.c, which is the only\n     +    place that it's actually needed.\n     +\n     +    The build breakage is due to the definition of UNICODE_STRING in\n     +    NtSecApi.h:\n     +        #ifndef _NTDEF_\n     +        typedef LSA_UNICODE_STRING UNICODE_STRING, *PUNICODE_STRING;\n     +        typedef LSA_STRING STRING, *PSTRING ;\n     +        #endif\n     +\n     +    LsaLookup.h:\n     +        typedef struct _LSA_UNICODE_STRING {\n     +            USHORT Length;\n     +            USHORT MaximumLength;\n     +        #ifdef MIDL_PASS\n     +            [size_is(MaximumLength/2), length_is(Length/2)]\n     +        #endif // MIDL_PASS\n     +            PWSTR  Buffer;\n     +        } LSA_UNICODE_STRING, *PLSA_UNICODE_STRING;\n     +\n     +    winternl.h also defines UNICODE_STRING:\n     +        typedef struct _UNICODE_STRING {\n     +            USHORT Length;\n     +            USHORT MaximumLength;\n     +            PWSTR  Buffer;\n     +        } UNICODE_STRING;\n     +        typedef UNICODE_STRING *PUNICODE_STRING;\n     +\n     +    Both definitions have equivalent layouts. Apparently these internal\n     +    Windows headers aren't designed to be included together. This is\n     +    an oversight in the headers and does not represent an incompatibility\n     +    between the APIs.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ git-compat-util.h\n       #endif\n       \n       #include <unistd.h>\n     -@@\n     - #else\n     - #include <stdint.h>\n     - #endif\n     --#ifdef HAVE_ARC4RANDOM_LIBBSD\n     --#include <bsd/stdlib.h>\n     --#endif\n     --#ifdef HAVE_GETRANDOM\n     --#include <sys/random.h>\n     --#endif\n     - #ifdef NO_INTPTR_T\n     - /*\n     -  * On I16LP32, ILP32 and LP64 \"long\" is the safe bet, however\n      \n       ## wrapper.c ##\n      @@\n     @@ wrapper.c\n      +#include <NTSecAPI.h>\n      +#undef SystemFunction036\n      +#endif\n     -+\n     -+#ifdef HAVE_ARC4RANDOM_LIBBSD\n     -+#include <bsd/stdlib.h>\n     -+#endif\n     -+#ifdef HAVE_GETRANDOM\n     -+#include <sys/random.h>\n     -+#endif\n      +\n       static int memory_limit_check(size_t size, int gentle)\n       {\n 2:  da8cfc10bb4 = 2:  a41bd4a06af core.fsyncmethod: add writeout-only mode\n 3:  e31886717b4 ! 3:  64e2bdcdfd9 core.fsync: introduce granular fsync control\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    core.fsync: introduce granular fsync control\n     +    core.fsync: introduce granular fsync control infrastructure\n      \n     -    This commit introduces the `core.fsync` configuration\n     -    knob which can be used to control how components of the\n     -    repository are made durable on disk.\n     +    This commit introduces the infrastructure for the core.fsync\n     +    configuration knob. The repository components we want to sync\n     +    are identified by flags so that we can turn on or off syncing\n     +    for specific components.\n      \n     -    This setting allows future extensibility of the list of\n     -    syncable components:\n     -    * We issue a warning rather than an error for unrecognized\n     -      components, so new configs can be used with old Git versions.\n     -    * We support negation, so users can choose one of the aggregate\n     -      options and then remove components that they don't want.\n     -      Aggregate options are defined in a later patch in this series.\n     +    If core.fsyncObjectFiles is set and the core.fsync configuration\n     +    also includes FSYNC_COMPONENT_LOOSE_OBJECT, we will fsync any\n     +    loose objects. This picks the strictest data integrity behavior\n     +    if core.fsync and core.fsyncObjectFiles are set to conflicting values.\n      \n     -    This also supports the common request of doing absolutely no\n     -    fysncing with the `core.fsync=none` value, which is expected\n     -    to make the test suite faster.\n     +    This change introduces the currently unused fsync_component\n     +    helper, which will be used by a later patch that adds fsyncing to\n     +    the refs backend.\n      \n     -    Complete documentation for the new setting is included in a later patch\n     -    in the series so that it can be reviewed in final form.\n     +    Actual configuration and documentation of the fsync components\n     +    list are in other patches in the series to separate review of\n     +    the underlying mechanism from the policy of how it's configured.\n      \n     +    Helped-by: Patrick Steinhardt <ps@pks.im>\n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     - ## Documentation/config/core.txt ##\n     -@@ Documentation/config/core.txt: core.fsyncMethod::\n     -   filesystem and storage hardware, data added to the repository may not be\n     -   durable in the event of a system crash. This is the default mode on macOS.\n     - \n     --core.fsyncObjectFiles::\n     --\tThis boolean will enable 'fsync()' when writing object files.\n     --+\n     --This is a total waste of time and effort on a filesystem that orders\n     --data writes properly, but can be useful for filesystems that do not use\n     --journalling (traditional UNIX filesystems) or that only journal metadata\n     --and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n     --\n     - core.preloadIndex::\n     - \tEnable parallel index preload for operations like 'git diff'\n     - +\n     -\n       ## builtin/fast-import.c ##\n      @@ builtin/fast-import.c: static void end_packfile(void)\n       \t\tstruct tag *t;\n     @@ cache.h: void reset_shared_repository(void);\n       extern int read_replace_refs;\n       extern char *git_replace_ref_base;\n       \n     --extern int fsync_object_files;\n     --extern int use_fsync;\n      +/*\n      + * These values are used to help identify parts of a repository to fsync.\n      + * FSYNC_COMPONENT_NONE identifies data that will not be a persistent part of the\n     @@ cache.h: void reset_shared_repository(void);\n      + * A bitmask indicating which components of the repo should be fsynced.\n      + */\n      +extern enum fsync_component fsync_components;\n     + extern int fsync_object_files;\n     + extern int use_fsync;\n       \n     - enum fsync_method {\n     - \tFSYNC_METHOD_FSYNC,\n     -@@ cache.h: enum fsync_method {\n     - };\n     - \n     - extern enum fsync_method fsync_method;\n     -+extern int use_fsync;\n     - extern int core_preload_index;\n     - extern int precomposed_unicode;\n     - extern int protect_hfs;\n      @@ cache.h: int copy_file_with_time(const char *dst, const char *src, int mode);\n     + \n       void write_or_die(int fd, const void *buf, size_t count);\n       void fsync_or_die(int fd, const char *);\n     ++int fsync_component(enum fsync_component component, int fd);\n     ++void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n       \n     -+static inline void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n     -+{\n     -+\tif (fsync_components & component)\n     -+\t\tfsync_or_die(fd, msg);\n     -+}\n     -+\n       ssize_t read_in_full(int fd, void *buf, size_t count);\n       ssize_t write_in_full(int fd, const void *buf, size_t count);\n     - ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\n      \n       ## commit-graph.c ##\n      @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n     @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_con\n       \n       \tif (ctx->split) {\n      \n     - ## config.c ##\n     -@@ config.c: static int git_parse_maybe_bool_text(const char *value)\n     - \treturn -1;\n     - }\n     - \n     -+static const struct fsync_component_entry {\n     -+\tconst char *name;\n     -+\tenum fsync_component component_bits;\n     -+} fsync_component_table[] = {\n     -+\t{ \"loose-object\", FSYNC_COMPONENT_LOOSE_OBJECT },\n     -+\t{ \"pack\", FSYNC_COMPONENT_PACK },\n     -+\t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n     -+\t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n     -+};\n     -+\n     -+static enum fsync_component parse_fsync_components(const char *var, const char *string)\n     -+{\n     -+\tenum fsync_component output = 0;\n     -+\n     -+\tif (!strcmp(string, \"none\"))\n     -+\t\treturn FSYNC_COMPONENT_NONE;\n     -+\n     -+\twhile (string) {\n     -+\t\tint i;\n     -+\t\tsize_t len;\n     -+\t\tconst char *ep;\n     -+\t\tint negated = 0;\n     -+\t\tint found = 0;\n     -+\n     -+\t\tstring = string + strspn(string, \", \\t\\n\\r\");\n     -+\t\tep = strchrnul(string, ',');\n     -+\t\tlen = ep - string;\n     -+\n     -+\t\tif (*string == '-') {\n     -+\t\t\tnegated = 1;\n     -+\t\t\tstring++;\n     -+\t\t\tlen--;\n     -+\t\t\tif (!len)\n     -+\t\t\t\twarning(_(\"invalid value for variable %s\"), var);\n     -+\t\t}\n     -+\n     -+\t\tif (!len)\n     -+\t\t\tbreak;\n     -+\n     -+\t\tfor (i = 0; i < ARRAY_SIZE(fsync_component_table); ++i) {\n     -+\t\t\tconst struct fsync_component_entry *entry = &fsync_component_table[i];\n     -+\n     -+\t\t\tif (strncmp(entry->name, string, len))\n     -+\t\t\t\tcontinue;\n     -+\n     -+\t\t\tfound = 1;\n     -+\t\t\tif (negated)\n     -+\t\t\t\toutput &= ~entry->component_bits;\n     -+\t\t\telse\n     -+\t\t\t\toutput |= entry->component_bits;\n     -+\t\t}\n     -+\n     -+\t\tif (!found) {\n     -+\t\t\tchar *component = xstrndup(string, len);\n     -+\t\t\twarning(_(\"ignoring unknown core.fsync component '%s'\"), component);\n     -+\t\t\tfree(component);\n     -+\t\t}\n     -+\n     -+\t\tstring = ep;\n     -+\t}\n     -+\n     -+\treturn output;\n     -+}\n     -+\n     - int git_parse_maybe_bool(const char *value)\n     - {\n     - \tint v = git_parse_maybe_bool_text(value);\n     -@@ config.c: static int git_default_core_config(const char *var, const char *value, void *cb)\n     - \t\treturn 0;\n     - \t}\n     - \n     -+\tif (!strcmp(var, \"core.fsync\")) {\n     -+\t\tif (!value)\n     -+\t\t\treturn config_error_nonbool(var);\n     -+\t\tfsync_components = parse_fsync_components(var, value);\n     -+\t\treturn 0;\n     -+\t}\n     -+\n     - \tif (!strcmp(var, \"core.fsyncmethod\")) {\n     - \t\tif (!value)\n     - \t\t\treturn config_error_nonbool(var);\n     -@@ config.c: static int git_default_core_config(const char *var, const char *value, void *cb)\n     - \t}\n     - \n     - \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n     --\t\tfsync_object_files = git_config_bool(var, value);\n     -+\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n     - \t\treturn 0;\n     - \t}\n     - \n     -\n       ## csum-file.c ##\n      @@ csum-file.c: static void free_hashfile(struct hashfile *f)\n       \tfree(f);\n     @@ csum-file.h: int hashfile_truncate(struct hashfile *, struct hashfile_checkpoint\n       void crc32_begin(struct hashfile *);\n      \n       ## environment.c ##\n     -@@ environment.c: const char *git_attributes_file;\n     - const char *git_hooks_path;\n     - int zlib_compression_level = Z_BEST_SPEED;\n     - int pack_compression_level = Z_DEFAULT_COMPRESSION;\n     --int fsync_object_files;\n     +@@ environment.c: int pack_compression_level = Z_DEFAULT_COMPRESSION;\n     + int fsync_object_files;\n       int use_fsync = -1;\n       enum fsync_method fsync_method = FSYNC_METHOD_DEFAULT;\n      +enum fsync_component fsync_components = FSYNC_COMPONENTS_DEFAULT;\n     @@ midx.c: static int write_midx_internal(const char *object_dir,\n      \n       ## object-file.c ##\n      @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n     + /* Finalize a file on disk, and close it. */\n       static void close_loose_object(int fd)\n       {\n     - \tif (!the_repository->objects->odb->will_destroy) {\n     +-\tif (!the_repository->objects->odb->will_destroy) {\n      -\t\tif (fsync_object_files)\n      -\t\t\tfsync_or_die(fd, \"loose object file\");\n     -+\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd, \"loose object file\");\n     - \t}\n     +-\t}\n     ++\tif (the_repository->objects->odb->will_destroy)\n     ++\t\tgoto out;\n       \n     ++\tif (fsync_object_files > 0)\n     ++\t\tfsync_or_die(fd, \"loose object file\");\n     ++\telse\n     ++\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n     ++\t\t\t\t       \"loose object file\");\n     ++\n     ++out:\n       \tif (close(fd) != 0)\n     + \t\tdie_errno(_(\"error when closing loose object file\"));\n     + }\n      \n       ## pack-bitmap-write.c ##\n      @@ pack-bitmap-write.c: void bitmap_writer_finish(struct pack_idx_entry **index,\n     @@ read-cache.c: static int do_write_index(struct index_state *istate, struct tempf\n       \tif (close_tempfile_gently(tempfile)) {\n       \t\terror(_(\"could not close '%s'\"), get_tempfile_path(tempfile));\n       \t\treturn -1;\n     +\n     + ## write-or-die.c ##\n     +@@ write-or-die.c: void fprintf_or_die(FILE *f, const char *fmt, ...)\n     + \t}\n     + }\n     + \n     +-void fsync_or_die(int fd, const char *msg)\n     ++static int maybe_fsync(int fd)\n     + {\n     + \tif (use_fsync < 0)\n     + \t\tuse_fsync = git_env_bool(\"GIT_TEST_FSYNC\", 1);\n     + \tif (!use_fsync)\n     +-\t\treturn;\n     ++\t\treturn 0;\n     + \n     + \tif (fsync_method == FSYNC_METHOD_WRITEOUT_ONLY &&\n     + \t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0)\n     +-\t\treturn;\n     ++\t\treturn 0;\n     ++\n     ++\treturn git_fsync(fd, FSYNC_HARDWARE_FLUSH);\n     ++}\n     + \n     +-\tif (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0)\n     ++void fsync_or_die(int fd, const char *msg)\n     ++{\n     ++\tif (maybe_fsync(fd) < 0)\n     + \t\tdie_errno(\"fsync error on '%s'\", msg);\n     + }\n     + \n     ++int fsync_component(enum fsync_component component, int fd)\n     ++{\n     ++\tif (fsync_components & component)\n     ++\t\treturn maybe_fsync(fd);\n     ++\treturn 0;\n     ++}\n     ++\n     ++void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n     ++{\n     ++\tif (fsync_components & component)\n     ++\t\tfsync_or_die(fd, msg);\n     ++}\n     ++\n     + void write_or_die(int fd, const void *buf, size_t count)\n     + {\n     + \tif (write_in_full(fd, buf, count) < 0) {\n -:  ----------- > 4:  6adc8dc1385 core.fsync: add configuration parsing\n 4:  9da808ba743 ! 5:  757f6d0bbd2 core.fsync: new option to harden the index\n     @@ cache.h: enum fsync_component {\n       #define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n      \n       ## config.c ##\n     -@@ config.c: static const struct fsync_component_entry {\n     +@@ config.c: static const struct fsync_component_name {\n       \t{ \"pack\", FSYNC_COMPONENT_PACK },\n       \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n       \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n 5:  2d71346b10e ! 6:  7e4cc6e10a5 core.fsync: documentation and user-friendly aggregate options\n     @@ Documentation/config/core.txt: core.whitespace::\n         errors. The default tab width is 8. Allowed values are 1 to 63.\n       \n      +core.fsync::\n     -+\tA comma-separated list of parts of the repository which should be\n     -+\thardened via the core.fsyncMethod when created or modified. You can\n     -+\tdisable hardening of any component by prefixing it with a '-'. Later\n     -+\titems take precedence over earlier ones in the comma-separated list.\n     -+\tFor example, `core.fsync=all,-pack-metadata` means \"harden everything\n     -+\texcept pack metadata.\" Items that are not hardened may be lost in the\n     -+\tevent of an unclean system shutdown. Unless you have special\n     -+\trequirements, it is recommended that you leave this option as default\n     -+\tor pick one of `committed`, `added`, or `all`.\n     ++\tA comma-separated list of components of the repository that\n     ++\tshould be hardened via the core.fsyncMethod when created or\n     ++\tmodified.  You can disable hardening of any component by\n     ++\tprefixing it with a '-'.  Items that are not hardened may be\n     ++\tlost in the event of an unclean\tsystem shutdown. Unless you\n     ++\thave special requirements, it is recommended that you leave\n     ++\tthis option empty or pick one of `committed`, `added`,\n     ++\tor `all`.\n      ++\n     -+* `none` disables fsync completely. This value must be specified alone.\n     ++When this configuration is encountered, the set of components starts with\n     ++the platform default value, disabled components are removed, and additional\n     ++components are added. `none` resets the state so that the platform default\n     ++is ignored.\n     +++\n     ++The empty string resets the fsync configuration to the platform\n     ++default. The platform default on most platform is equivalent to\n     ++`core.fsync=committed,-loose-object`, which has good performance,\n     ++but risks losing recent work in the event of an unclean system shutdown.\n     +++\n     ++* `none` clears the set of fsynced components.\n      +* `loose-object` hardens objects added to the repo in loose-object form.\n      +* `pack` hardens objects added to the repo in packfile form.\n      +* `pack-metadata` hardens packfile bitmaps and indexes.\n     @@ Documentation/config/core.txt: core.whitespace::\n      +  `loose-object,pack`.\n      +* `derived-metadata` is an aggregate option that is equivalent to\n      +  `pack-metadata,commit-graph`.\n     -+* `default` is an aggregate option that is equivalent to\n     -+  `objects,derived-metadata,-loose-object`. This mode is enabled by default.\n     -+  It has good performance, but risks losing recent work if the system shuts\n     -+  down uncleanly, since commits, trees, and blobs in loose-object form may be\n     -+  lost.\n      +* `committed` is an aggregate option that is currently equivalent to\n     -+  `objects`. This mode sacrifices some performance to ensure that all work\n     ++  `objects`. This mode sacrifices some performance to ensure that work\n      +  that is committed to the repository with `git commit` or similar commands\n     -+  is preserved.\n     ++  is hardened.\n      +* `added` is an aggregate option that is currently equivalent to\n      +  `committed,index`. This mode sacrifices additional performance to\n      +  ensure that the results of commands like `git add` and similar operations\n     -+  are preserved.\n     ++  are hardened.\n      +* `all` is an aggregate option that syncs all individual components above.\n      +\n       core.fsyncMethod::\n     @@ cache.h: enum fsync_component {\n        * A bitmask indicating which components of the repo should be fsynced.\n      \n       ## config.c ##\n     -@@ config.c: static const struct fsync_component_entry {\n     +@@ config.c: static const struct fsync_component_name {\n       \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n       \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n       \t{ \"index\", FSYNC_COMPONENT_INDEX },\n      +\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n      +\t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n     -+\t{ \"default\", FSYNC_COMPONENTS_DEFAULT },\n      +\t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\n      +\t{ \"added\", FSYNC_COMPONENTS_ADDED },\n      +\t{ \"all\", FSYNC_COMPONENTS_ALL },\n\n-- \ngitgitgadget\n"},{"id":"451096","messageId":"CANQDOddJYNiGY7raZFaMdX-ySSV7xfop9iMXoNG9jKBf5PnqTQ@mail.gmail.com","threadId":"57030","inReplyTo":"f1e8a7bb3bf0f4c0414819cb1d5579dc08fd2a4f.1646905589.git.ps@pks.im","subject":"Re: [PATCH 7/8] core.fsync: new option to harden loose references","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T22:54:35Z","receivedAt":"2022-03-10T22:54:51Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 10, 2022 at 1:53 AM Patrick Steinhardt <ps@pks.im> wrote:\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index 973805e8a9..b67d3c340e 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -564,8 +564,10 @@ core.fsync::\n>  * `pack-metadata` hardens packfile bitmaps and indexes.\n>  * `commit-graph` hardens the commit graph file.\n>  * `index` hardens the index when it is modified.\n> +* `loose-ref` hardens references modified in the repo in loose-ref form.\n>  * `objects` is an aggregate option that is equivalent to\n>    `loose-object,pack`.\n> +* `refs` is an aggregate option that is equivalent to `loose-ref`.\n>  * `derived-metadata` is an aggregate option that is equivalent to\n>    `pack-metadata,commit-graph`.\n>  * `default` is an aggregate option that is equivalent to\n> diff --git a/cache.h b/cache.h\n> index 63a95d1977..b56a56f539 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1005,11 +1005,14 @@ enum fsync_component {\n>         FSYNC_COMPONENT_PACK_METADATA           = 1 << 2,\n>         FSYNC_COMPONENT_COMMIT_GRAPH            = 1 << 3,\n>         FSYNC_COMPONENT_INDEX                   = 1 << 4,\n> +       FSYNC_COMPONENT_LOOSE_REF               = 1 << 5,\n>  };\n>\n>  #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n>                                   FSYNC_COMPONENT_PACK)\n>\n> +#define FSYNC_COMPONENTS_REFS (FSYNC_COMPONENT_LOOSE_REF)\n> +\n>  #define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n>                                            FSYNC_COMPONENT_COMMIT_GRAPH)\n>\n> @@ -1026,7 +1029,8 @@ enum fsync_component {\n>                               FSYNC_COMPONENT_PACK | \\\n>                               FSYNC_COMPONENT_PACK_METADATA | \\\n>                               FSYNC_COMPONENT_COMMIT_GRAPH | \\\n> -                             FSYNC_COMPONENT_INDEX)\n> +                             FSYNC_COMPONENT_INDEX | \\\n> +                             FSYNC_COMPONENT_LOOSE_REF)\n>\n>  /*\n>   * A bitmask indicating which components of the repo should be fsynced.\n> diff --git a/config.c b/config.c\n> index f03f27c3de..b5d3e6e404 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1332,7 +1332,9 @@ static const struct fsync_component_entry {\n>         { \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n>         { \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n>         { \"index\", FSYNC_COMPONENT_INDEX },\n> +       { \"loose-ref\", FSYNC_COMPONENT_LOOSE_REF },\n>         { \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> +       { \"refs\", FSYNC_COMPONENTS_REFS },\n>         { \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n>         { \"default\", FSYNC_COMPONENTS_DEFAULT },\n>         { \"committed\", FSYNC_COMPONENTS_COMMITTED },\n\nIn terms of the 'preciousness-levels', refs should be included in\nFSYNC_COMPONENTS_COMMITTED,\nfrom which it will also be included in _ADDED.\n"},{"id":"451100","messageId":"xmqqmthxbcv2.fsf@gitster.g","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 0/6] A design for future-proofing fsync() configuration","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-10T23:34:57Z","receivedAt":"2022-03-10T23:35:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj K. Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> After this change, new persistent data files added to the repo will need to\n> be added to the fsync_component enum and documented in the\n> Documentation/config/core.txt text.\n>\n> V6 changes:\n>\n>  * Move only the windows csprng includes into wrapper.c rather than all of\n>    them. This fixes the specific build issue due to broken Windows headers.\n>    [6]\n>  * Split the configuration parsing of core.fsync from the mechanism to focus\n>    the review.\n>  * Incorporate Patrick's patch at [7] into the core.fsync mechanism patch.\n>  * Pick the stricter one of core.fsyncObjectFiles and (fsync_components &\n>    FSYNC_COMPONENT_LOOSE_OBJECTS), to respect the older setting.\n>  * Issue a deprecation warning but keep parsing and honoring\n>    core.fsyncObjectFiles.\n>  * Change configuration parsing of core.fsync to always start with the\n>    platform default. none resets to the empty set. The comma separated list\n>    implies a set without regards to ordering now. This follows Junio's\n>    suggestion in [8].\n>  * Change the documentation of the core.fsync option to reflect the way the\n>    new parsing code works.\n\nHmph, this seems to make one test fail.\n\nt5801-remote-helpers.sh (Wstat: 256 Tests: 31 Failed: 4)\n  Failed tests:  14-16, 31\n    Non-zero exit status: 1\nFiles=1, Tests=31,  2 wallclock secs ( 0.04 usr  0.00 sys + 1.40 cusr  1.62 csys =  3.06 CPU)\nResult: FAIL\n"},{"id":"451104","messageId":"CANQDOdfX4tXSy_9r8mnXpYD=BoJtEcydAXBTff_ekf=6ryoebQ@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqmthxbcv2.fsf@gitster.g","subject":"Re: [PATCH v6 0/6] A design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-11T00:03:18Z","receivedAt":"2022-03-11T00:03:34Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 10, 2022 at 3:35 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj K. Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > After this change, new persistent data files added to the repo will need to\n> > be added to the fsync_component enum and documented in the\n> > Documentation/config/core.txt text.\n> >\n> > V6 changes:\n> >\n> >  * Move only the windows csprng includes into wrapper.c rather than all of\n> >    them. This fixes the specific build issue due to broken Windows headers.\n> >    [6]\n> >  * Split the configuration parsing of core.fsync from the mechanism to focus\n> >    the review.\n> >  * Incorporate Patrick's patch at [7] into the core.fsync mechanism patch.\n> >  * Pick the stricter one of core.fsyncObjectFiles and (fsync_components &\n> >    FSYNC_COMPONENT_LOOSE_OBJECTS), to respect the older setting.\n> >  * Issue a deprecation warning but keep parsing and honoring\n> >    core.fsyncObjectFiles.\n> >  * Change configuration parsing of core.fsync to always start with the\n> >    platform default. none resets to the empty set. The comma separated list\n> >    implies a set without regards to ordering now. This follows Junio's\n> >    suggestion in [8].\n> >  * Change the documentation of the core.fsync option to reflect the way the\n> >    new parsing code works.\n>\n> Hmph, this seems to make one test fail.\n>\n> t5801-remote-helpers.sh (Wstat: 256 Tests: 31 Failed: 4)\n>   Failed tests:  14-16, 31\n>     Non-zero exit status: 1\n> Files=1, Tests=31,  2 wallclock secs ( 0.04 usr  0.00 sys + 1.40 cusr  1.62 csys =  3.06 CPU)\n> Result: FAIL\n\nThanks for reporting this.  I didn't see a failure in CI, nor when\nrunning that specific test in mingw.  I also munged my config to\ninclude core.fsyncObjectFiles and didn't see a failure.\n\nCould you please share some more verbose output of the test, so I can\nlook a bit deeper?  In parallel, I'm trying again after merging my\nchanges onto seen.\n\nThanks,\nNeeraj\n"},{"id":"451111","messageId":"xmqqzglx9em0.fsf@gitster.g","threadId":"57030","inReplyTo":"f1e8a7bb3bf0f4c0414819cb1d5579dc08fd2a4f.1646905589.git.ps@pks.im","subject":"Re: [PATCH 7/8] core.fsync: new option to harden loose references","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-11T06:40:07Z","receivedAt":"2022-03-11T06:40:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> @@ -1504,6 +1513,7 @@ static int files_copy_or_rename_ref(struct ref_store *ref_store,\n>  \toidcpy(&lock->old_oid, &orig_oid);\n>  \n>  \tif (write_ref_to_lockfile(lock, &orig_oid, 0, &err) ||\n> +\t    files_sync_loose_ref(lock, &err) ||\n>  \t    commit_ref_update(refs, lock, &orig_oid, logmsg, &err)) {\n>  \t\terror(\"unable to write current sha1 into %s: %s\", newrefname, err.buf);\n>  \t\tstrbuf_release(&err);\n\nGiven that write_ref_to_lockfile() on the success code path does this:\n\n\tfd = get_lock_file_fd(&lock->lk);\n\tif (write_in_full(fd, oid_to_hex(oid), the_hash_algo->hexsz) < 0 ||\n\t    write_in_full(fd, &term, 1) < 0 ||\n\t    close_ref_gently(lock) < 0) {\n\t\tstrbuf_addf(err,\n\t\t\t    \"couldn't write '%s'\", get_lock_file_path(&lock->lk));\n\t\tunlock_ref(lock);\n\t\treturn -1;\n\t}\n\treturn 0;\n\nthe above unfortunately does not work.  By the time the new call to\nfiles_sync_loose_ref() is made, lock->fd is closed by the call to\nclose_lock_file_gently() made in close_ref_gently(), and because of\nthat, you'll get an error like this:\n\n    Writing objects: 100% (3/3), 279 bytes | 279.00 KiB/s, done.\n    Total 3 (delta 0), reused 0 (delta 0), pack-reused 0\n    remote: error: could not sync loose ref 'refs/heads/client_branch':\n    Bad file descriptor     \n\nwhen running \"make test\" (the above is from t5702 but I wouldn't be\nsurprised if this broke ALL ref updates).\n\nJust before write_ref_to_lockfile() calls close_ref_gently() would\nbe a good place to make the fsync_loose_ref() call, perhaps?\n\n\nThanks.\n"},{"id":"451116","messageId":"YisR92nI81SHlOYb@ncase","threadId":"57030","inReplyTo":"xmqqzglxek5y.fsf@gitster.g","subject":"Re: [PATCH 8/8] core.fsync: new option to harden packed references","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-11T09:10:15Z","receivedAt":"2022-03-11T09:10:31Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Thu, Mar 10, 2022 at 10:28:57AM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > diff --git a/refs/packed-backend.c b/refs/packed-backend.c\n> > index 27dd8c3922..32d6635969 100644\n> > --- a/refs/packed-backend.c\n> > +++ b/refs/packed-backend.c\n> > @@ -1262,7 +1262,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n> >  \t\tgoto error;\n> >  \t}\n> >  \n> > -\tif (close_tempfile_gently(refs->tempfile)) {\n> > +\tif (fsync_component(FSYNC_COMPONENT_PACKED_REFS, get_tempfile_fd(refs->tempfile)) ||\n> > +\t    close_tempfile_gently(refs->tempfile)) {\n> >  \t\tstrbuf_addf(err, \"error closing file %s: %s\",\n> >  \t\t\t    get_tempfile_path(refs->tempfile),\n> >  \t\t\t    strerror(errno));\n> \n> I do not necessarily agree with the organization to have it as a\n> component that is separate from other ref backends, but it is\n> very pleasing to see that there is only one fsync necessary for the\n> packed backend.\n> \n> Nice.\n\nI was mostly adapting to the precedent set by Neeraj, where we also\ndistinguish loose objects and packed objects. Personally I don't mind\nmuch whether we want to discern those two cases, and I'd be happy to\njust merge them into a single \"refs\" knob. We can still split these up\nat a later point if the need ever arises.\n\nPatrick\n"},{"id":"451117","messageId":"YisTPSOqKkQQ1RbQ@ncase","threadId":"57030","inReplyTo":"xmqqzglx9em0.fsf@gitster.g","subject":"Re: [PATCH 7/8] core.fsync: new option to harden loose references","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-11T09:15:41Z","receivedAt":"2022-03-11T09:15:49Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Thu, Mar 10, 2022 at 10:40:07PM -0800, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > @@ -1504,6 +1513,7 @@ static int files_copy_or_rename_ref(struct ref_store *ref_store,\n> >  \toidcpy(&lock->old_oid, &orig_oid);\n> >  \n> >  \tif (write_ref_to_lockfile(lock, &orig_oid, 0, &err) ||\n> > +\t    files_sync_loose_ref(lock, &err) ||\n> >  \t    commit_ref_update(refs, lock, &orig_oid, logmsg, &err)) {\n> >  \t\terror(\"unable to write current sha1 into %s: %s\", newrefname, err.buf);\n> >  \t\tstrbuf_release(&err);\n> \n> Given that write_ref_to_lockfile() on the success code path does this:\n> \n> \tfd = get_lock_file_fd(&lock->lk);\n> \tif (write_in_full(fd, oid_to_hex(oid), the_hash_algo->hexsz) < 0 ||\n> \t    write_in_full(fd, &term, 1) < 0 ||\n> \t    close_ref_gently(lock) < 0) {\n> \t\tstrbuf_addf(err,\n> \t\t\t    \"couldn't write '%s'\", get_lock_file_path(&lock->lk));\n> \t\tunlock_ref(lock);\n> \t\treturn -1;\n> \t}\n> \treturn 0;\n> \n> the above unfortunately does not work.  By the time the new call to\n> files_sync_loose_ref() is made, lock->fd is closed by the call to\n> close_lock_file_gently() made in close_ref_gently(), and because of\n> that, you'll get an error like this:\n> \n>     Writing objects: 100% (3/3), 279 bytes | 279.00 KiB/s, done.\n>     Total 3 (delta 0), reused 0 (delta 0), pack-reused 0\n>     remote: error: could not sync loose ref 'refs/heads/client_branch':\n>     Bad file descriptor     \n> \n> when running \"make test\" (the above is from t5702 but I wouldn't be\n> surprised if this broke ALL ref updates).\n> \n> Just before write_ref_to_lockfile() calls close_ref_gently() would\n> be a good place to make the fsync_loose_ref() call, perhaps?\n> \n> \n> Thanks.\n\nYeah, that thought indeed occurred to me this night, too. I was hoping\nthat I could fix this before anybody noticed ;)\n\nIt's a bit unfortunate that we can't just defer this to a later place to\nhopefully implement this more efficiently, but so be it. The alternative\nwould be to re-open all locked loose refs and then sync them to disk,\nbut this would likely be a lot more painful than just syncing them to\ndisk before closing it.\n\nWill fix, thanks.\n\nPatrick\n"},{"id":"451119","messageId":"220311.86y21ghlln.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"YisTPSOqKkQQ1RbQ@ncase","subject":"Re: [PATCH 7/8] core.fsync: new option to harden loose references","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-11T09:36:26Z","receivedAt":"2022-03-11T09:42:02Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Mar 11 2022, Patrick Steinhardt wrote:\n\n> [[PGP Signed Part:Undecided]]\n> On Thu, Mar 10, 2022 at 10:40:07PM -0800, Junio C Hamano wrote:\n>> Patrick Steinhardt <ps@pks.im> writes:\n>> \n>> > @@ -1504,6 +1513,7 @@ static int files_copy_or_rename_ref(struct ref_store *ref_store,\n>> >  \toidcpy(&lock->old_oid, &orig_oid);\n>> >  \n>> >  \tif (write_ref_to_lockfile(lock, &orig_oid, 0, &err) ||\n>> > +\t    files_sync_loose_ref(lock, &err) ||\n>> >  \t    commit_ref_update(refs, lock, &orig_oid, logmsg, &err)) {\n>> >  \t\terror(\"unable to write current sha1 into %s: %s\", newrefname, err.buf);\n>> >  \t\tstrbuf_release(&err);\n>> \n>> Given that write_ref_to_lockfile() on the success code path does this:\n>> \n>> \tfd = get_lock_file_fd(&lock->lk);\n>> \tif (write_in_full(fd, oid_to_hex(oid), the_hash_algo->hexsz) < 0 ||\n>> \t    write_in_full(fd, &term, 1) < 0 ||\n>> \t    close_ref_gently(lock) < 0) {\n>> \t\tstrbuf_addf(err,\n>> \t\t\t    \"couldn't write '%s'\", get_lock_file_path(&lock->lk));\n>> \t\tunlock_ref(lock);\n>> \t\treturn -1;\n>> \t}\n>> \treturn 0;\n>> \n>> the above unfortunately does not work.  By the time the new call to\n>> files_sync_loose_ref() is made, lock->fd is closed by the call to\n>> close_lock_file_gently() made in close_ref_gently(), and because of\n>> that, you'll get an error like this:\n>> \n>>     Writing objects: 100% (3/3), 279 bytes | 279.00 KiB/s, done.\n>>     Total 3 (delta 0), reused 0 (delta 0), pack-reused 0\n>>     remote: error: could not sync loose ref 'refs/heads/client_branch':\n>>     Bad file descriptor     \n>> \n>> when running \"make test\" (the above is from t5702 but I wouldn't be\n>> surprised if this broke ALL ref updates).\n>> \n>> Just before write_ref_to_lockfile() calls close_ref_gently() would\n>> be a good place to make the fsync_loose_ref() call, perhaps?\n>> \n>> \n>> Thanks.\n>\n> Yeah, that thought indeed occurred to me this night, too. I was hoping\n> that I could fix this before anybody noticed ;)\n>\n> It's a bit unfortunate that we can't just defer this to a later place to\n> hopefully implement this more efficiently, but so be it. The alternative\n> would be to re-open all locked loose refs and then sync them to disk,\n> but this would likely be a lot more painful than just syncing them to\n> disk before closing it.\n\nAside: is open/write/close followed by open/fsync/close on the same file\nportably guaranteed to yield the same end result as a single\nopen/write/fsync/close?\n\nI think in practice nobody would be insane enough to implement a system\nto do otherwise, but on the other hand I've seen some really insane\nbehavior :)\n\nI could see it being different e.g. in some NFS cases/configurations\nwhere the fsync() for an open FD syncs to the remote storage, and the\nsecond open() might therefore get the old version and noop-sync that.\n\nMost implementations would guard against that in the common case by\nhaving a local cache of outstanding data to flush, but if you're talking\nto some sharded storage array for each request...\n\nAnyway, I *think* it should be OK, just an aside to check the assumption\nfor any future work... :)\n"},{"id":"451121","messageId":"47dd79106b93bb81750320d50ccaa74c24aacd28.1646992380.git.ps@pks.im","threadId":"57030","inReplyTo":"pull.1093.v6.git.1646952204.gitgitgadget@gmail.com","subject":"[PATCH v2] core.fsync: new option to harden references","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-11T09:58:59Z","receivedAt":"2022-03-11T09:59:07Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"When writing both loose and packed references to disk we first create a\nlockfile, write the updated values into that lockfile, and on commit we\nrename the file into place. According to filesystem developers, this\nbehaviour is broken because applications should always sync data to disk\nbefore doing the final rename to ensure data consistency [1][2][3]. If\napplications fail to do this correctly, a hard crash of the machine can\neasily result in corrupted on-disk data.\n\nThis kind of corruption can in fact be easily observed with Git when the\nmachine hard-resets shortly after writing references to disk. On\nmachines with ext4, this will likely lead to the \"empty files\" problem:\nthe file has been renamed, but its data has not been synced to disk. The\nresult is that the reference is corrupt, and in the worst case this can\nlead to data loss.\n\nImplement a new option to harden references so that users and admins can\navoid this scenario by syncing locked loose and packed references to\ndisk before we rename them into place.\n\n[1]: https://thunk.org/tytso/blog/2009/03/15/dont-fear-the-fsync/\n[2]: https://btrfs.wiki.kernel.org/index.php/FAQ (What are the crash guarantees of overwrite-by-rename)\n[3]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/Documentation/admin-guide/ext4.rst (see auto_da_alloc)\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n\nHi,\n\nhere's my updated patch series which implements syncing of refs. It\napplies on top of Neeraj's v6 of \"A design for future-proofing fsync()\nconfiguration\".\n\nI've simplified these patches a bit:\n\n    - I don't distinguishing between \"loose\" and \"packed\" refs anymore.\n      I agree with Junio that it's probably not worth it, but we can\n      still reintroduce the split at a later point without breaking\n      backwards compatibility if the need comes up.\n\n    - I've simplified the way loose refs are written to disk so that we\n      now sync them when before we close their files. The previous\n      implementation I had was broken because we tried to sync after\n      closing.\n\nBecause this really only changes a few lines of code I've also decided\nto squash together the patches into a single one.\n\nPatrick\n\n Documentation/config/core.txt | 1 +\n cache.h                       | 7 +++++--\n config.c                      | 1 +\n refs/files-backend.c          | 1 +\n refs/packed-backend.c         | 3 ++-\n 5 files changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 37105a7be4..812cca7de7 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -575,6 +575,7 @@ but risks losing recent work in the event of an unclean system shutdown.\n * `index` hardens the index when it is modified.\n * `objects` is an aggregate option that is equivalent to\n   `loose-object,pack`.\n+* `reference` hardens references modified in the repo.\n * `derived-metadata` is an aggregate option that is equivalent to\n   `pack-metadata,commit-graph`.\n * `committed` is an aggregate option that is currently equivalent to\ndiff --git a/cache.h b/cache.h\nindex cde0900d05..033e5b0779 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1005,6 +1005,7 @@ enum fsync_component {\n \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n+\tFSYNC_COMPONENT_REFERENCE\t\t= 1 << 5,\n };\n \n #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n@@ -1017,7 +1018,8 @@ enum fsync_component {\n \t\t\t\t  FSYNC_COMPONENTS_DERIVED_METADATA | \\\n \t\t\t\t  ~FSYNC_COMPONENT_LOOSE_OBJECT)\n \n-#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS)\n+#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS | \\\n+\t\t\t\t    FSYNC_COMPONENT_REFERENCE)\n \n #define FSYNC_COMPONENTS_ADDED (FSYNC_COMPONENTS_COMMITTED | \\\n \t\t\t\tFSYNC_COMPONENT_INDEX)\n@@ -1026,7 +1028,8 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_PACK | \\\n \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n \t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n-\t\t\t      FSYNC_COMPONENT_INDEX)\n+\t\t\t      FSYNC_COMPONENT_INDEX | \\\n+\t\t\t      FSYNC_COMPONENT_REFERENCE)\n \n /*\n  * A bitmask indicating which components of the repo should be fsynced.\ndiff --git a/config.c b/config.c\nindex eb75f65338..3c9b6b589a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1333,6 +1333,7 @@ static const struct fsync_component_name {\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"reference\", FSYNC_COMPONENT_REFERENCE },\n \t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n \t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\n \t{ \"added\", FSYNC_COMPONENTS_ADDED },\ndiff --git a/refs/files-backend.c b/refs/files-backend.c\nindex f59589d6cc..6521ee8af5 100644\n--- a/refs/files-backend.c\n+++ b/refs/files-backend.c\n@@ -1787,6 +1787,7 @@ static int write_ref_to_lockfile(struct ref_lock *lock,\n \tfd = get_lock_file_fd(&lock->lk);\n \tif (write_in_full(fd, oid_to_hex(oid), the_hash_algo->hexsz) < 0 ||\n \t    write_in_full(fd, &term, 1) < 0 ||\n+\t    fsync_component(FSYNC_COMPONENT_REFERENCE, get_lock_file_fd(&lock->lk)) < 0 ||\n \t    close_ref_gently(lock) < 0) {\n \t\tstrbuf_addf(err,\n \t\t\t    \"couldn't write '%s'\", get_lock_file_path(&lock->lk));\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 27dd8c3922..9d704ccd3e 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -1262,7 +1262,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\tgoto error;\n \t}\n \n-\tif (close_tempfile_gently(refs->tempfile)) {\n+\tif (fsync_component(FSYNC_COMPONENT_REFERENCE, get_tempfile_fd(refs->tempfile)) ||\n+\t    close_tempfile_gently(refs->tempfile)) {\n \t\tstrbuf_addf(err, \"error closing file %s: %s\",\n \t\t\t    get_tempfile_path(refs->tempfile),\n \t\t\t    strerror(errno));\n-- \n2.35.1\n\n"},{"id":"451157","messageId":"CANQDOdcp3gA+Uro9qzfPOtusni2j4tHqT2wCHqyvQHcwtVHQCg@mail.gmail.com","threadId":"57030","inReplyTo":"CANQDOdfX4tXSy_9r8mnXpYD=BoJtEcydAXBTff_ekf=6ryoebQ@mail.gmail.com","subject":"Re: [PATCH v6 0/6] A design for future-proofing fsync() configuration","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-11T18:50:42Z","receivedAt":"2022-03-11T18:51:00Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 10, 2022 at 4:03 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> On Thu, Mar 10, 2022 at 3:35 PM Junio C Hamano <gitster@pobox.com> wrote:\n> >\n> > \"Neeraj K. Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> >\n> > > After this change, new persistent data files added to the repo will need to\n> > > be added to the fsync_component enum and documented in the\n> > > Documentation/config/core.txt text.\n> > >\n> > > V6 changes:\n> > >\n> > >  * Move only the windows csprng includes into wrapper.c rather than all of\n> > >    them. This fixes the specific build issue due to broken Windows headers.\n> > >    [6]\n> > >  * Split the configuration parsing of core.fsync from the mechanism to focus\n> > >    the review.\n> > >  * Incorporate Patrick's patch at [7] into the core.fsync mechanism patch.\n> > >  * Pick the stricter one of core.fsyncObjectFiles and (fsync_components &\n> > >    FSYNC_COMPONENT_LOOSE_OBJECTS), to respect the older setting.\n> > >  * Issue a deprecation warning but keep parsing and honoring\n> > >    core.fsyncObjectFiles.\n> > >  * Change configuration parsing of core.fsync to always start with the\n> > >    platform default. none resets to the empty set. The comma separated list\n> > >    implies a set without regards to ordering now. This follows Junio's\n> > >    suggestion in [8].\n> > >  * Change the documentation of the core.fsync option to reflect the way the\n> > >    new parsing code works.\n> >\n> > Hmph, this seems to make one test fail.\n> >\n> > t5801-remote-helpers.sh (Wstat: 256 Tests: 31 Failed: 4)\n> >   Failed tests:  14-16, 31\n> >     Non-zero exit status: 1\n> > Files=1, Tests=31,  2 wallclock secs ( 0.04 usr  0.00 sys + 1.40 cusr  1.62 csys =  3.06 CPU)\n> > Result: FAIL\n>\n> Thanks for reporting this.  I didn't see a failure in CI, nor when\n> running that specific test in mingw.  I also munged my config to\n> include core.fsyncObjectFiles and didn't see a failure.\n>\n> Could you please share some more verbose output of the test, so I can\n> look a bit deeper?  In parallel, I'm trying again after merging my\n> changes onto seen.\n>\n> Thanks,\n> Neeraj\n\nHi Junio,\nI've also tested v6-on-seen on Linux and I'm still not seeing the\nfailure. Does the\nfailure still happen on your end?\nThanks,\nNeeraj\n"},{"id":"451276","messageId":"xmqqpmmptns6.fsf@gitster.g","threadId":"57030","inReplyTo":"xmqqmthxbcv2.fsf@gitster.g","subject":"Re: [PATCH v6 0/6] A design for future-proofing fsync() configuration","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-13T23:50:49Z","receivedAt":"2022-03-13T23:50:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Hmph, this seems to make one test fail.\n>\n> t5801-remote-helpers.sh (Wstat: 256 Tests: 31 Failed: 4)\n>   Failed tests:  14-16, 31\n>     Non-zero exit status: 1\n> Files=1, Tests=31,  2 wallclock secs ( 0.04 usr  0.00 sys + 1.40 cusr  1.62 csys =  3.06 CPU)\n> Result: FAIL\n\nFalse alarm.  This byitself, or merged to 'seen' with other random\ntopics, no longer seem to break these tests.\n\n"},{"id":"451417","messageId":"20220315191245.17990-1-neerajsi@microsoft.com","threadId":"57030","inReplyTo":"7e4cc6e10a5d88f4c6c44efaa68f2325007fd935.1646952205.git.gitgitgadget@gmail.com","subject":"[PATCH v7] core.fsync: documentation and user-friendly aggregate options","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-15T19:12:45Z","receivedAt":"2022-03-15T19:13:30Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"This commit adds aggregate options for the core.fsync setting that are\nmore user-friendly. These options are specified in terms of 'levels of\nsafety', indicating which Git operations are considered to be sync\npoints for durability.\n\nThe new documentation is also included here in its entirety for ease of\nreview.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\nThis revision fixes a grammatical mistake in the core.fsync documentation.\n\n Documentation/config/core.txt | 40 +++++++++++++++++++++++++++++++++++\n cache.h                       | 23 +++++++++++++++++---\n config.c                      |  5 +++++\n 3 files changed, 65 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex ab911d6e26..9a3ad71e9e 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -547,6 +547,46 @@ core.whitespace::\n   is relevant for `indent-with-non-tab` and when Git fixes `tab-in-indent`\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n+core.fsync::\n+\tA comma-separated list of components of the repository that\n+\tshould be hardened via the core.fsyncMethod when created or\n+\tmodified.  You can disable hardening of any component by\n+\tprefixing it with a '-'.  Items that are not hardened may be\n+\tlost in the event of an unclean\tsystem shutdown. Unless you\n+\thave special requirements, it is recommended that you leave\n+\tthis option empty or pick one of `committed`, `added`,\n+\tor `all`.\n++\n+When this configuration is encountered, the set of components starts with\n+the platform default value, disabled components are removed, and additional\n+components are added. `none` resets the state so that the platform default\n+is ignored.\n++\n+The empty string resets the fsync configuration to the platform\n+default. The default on most platforms is equivalent to\n+`core.fsync=committed,-loose-object`, which has good performance,\n+but risks losing recent work in the event of an unclean system shutdown.\n++\n+* `none` clears the set of fsynced components.\n+* `loose-object` hardens objects added to the repo in loose-object form.\n+* `pack` hardens objects added to the repo in packfile form.\n+* `pack-metadata` hardens packfile bitmaps and indexes.\n+* `commit-graph` hardens the commit graph file.\n+* `index` hardens the index when it is modified.\n+* `objects` is an aggregate option that is equivalent to\n+  `loose-object,pack`.\n+* `derived-metadata` is an aggregate option that is equivalent to\n+  `pack-metadata,commit-graph`.\n+* `committed` is an aggregate option that is currently equivalent to\n+  `objects`. This mode sacrifices some performance to ensure that work\n+  that is committed to the repository with `git commit` or similar commands\n+  is hardened.\n+* `added` is an aggregate option that is currently equivalent to\n+  `committed,index`. This mode sacrifices additional performance to\n+  ensure that the results of commands like `git add` and similar operations\n+  are hardened.\n+* `all` is an aggregate option that syncs all individual components above.\n+\n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\ndiff --git a/cache.h b/cache.h\nindex e08eeac6c1..86680f144e 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1007,9 +1007,26 @@ enum fsync_component {\n \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n };\n \n-#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENT_PACK | \\\n-\t\t\t\t  FSYNC_COMPONENT_PACK_METADATA | \\\n-\t\t\t\t  FSYNC_COMPONENT_COMMIT_GRAPH)\n+#define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t\t  FSYNC_COMPONENT_PACK)\n+\n+#define FSYNC_COMPONENTS_DERIVED_METADATA (FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t\t\t   FSYNC_COMPONENT_COMMIT_GRAPH)\n+\n+#define FSYNC_COMPONENTS_DEFAULT (FSYNC_COMPONENTS_OBJECTS | \\\n+\t\t\t\t  FSYNC_COMPONENTS_DERIVED_METADATA | \\\n+\t\t\t\t  ~FSYNC_COMPONENT_LOOSE_OBJECT)\n+\n+#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS)\n+\n+#define FSYNC_COMPONENTS_ADDED (FSYNC_COMPONENTS_COMMITTED | \\\n+\t\t\t\tFSYNC_COMPONENT_INDEX)\n+\n+#define FSYNC_COMPONENTS_ALL (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n+\t\t\t      FSYNC_COMPONENT_PACK | \\\n+\t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n+\t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n+\t\t\t      FSYNC_COMPONENT_INDEX)\n \n /*\n  * A bitmask indicating which components of the repo should be fsynced.\ndiff --git a/config.c b/config.c\nindex 80f33c9198..fd8e165931 100644\n--- a/config.c\n+++ b/config.c\n@@ -1332,6 +1332,11 @@ static const struct fsync_component_name {\n \t{ \"pack-metadata\", FSYNC_COMPONENT_PACK_METADATA },\n \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n \t{ \"index\", FSYNC_COMPONENT_INDEX },\n+\t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n+\t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n+\t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\n+\t{ \"added\", FSYNC_COMPONENTS_ADDED },\n+\t{ \"all\", FSYNC_COMPONENTS_ALL },\n };\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n-- \n2.33.0.334.g55a40fc8fd\n\n"},{"id":"451420","messageId":"xmqqv8wfc8qb.fsf@gitster.g","threadId":"57030","inReplyTo":"20220315191245.17990-1-neerajsi@microsoft.com","subject":"Re: [PATCH v7] core.fsync: documentation and user-friendly aggregate options","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-15T19:32:28Z","receivedAt":"2022-03-15T19:32:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> This commit adds aggregate options for the core.fsync setting that are\n> more user-friendly. These options are specified in terms of 'levels of\n> safety', indicating which Git operations are considered to be sync\n> points for durability.\n>\n> The new documentation is also included here in its entirety for ease of\n> review.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n> This revision fixes a grammatical mistake in the core.fsync documentation.\n\nIs this meant to be [PATCH v7 6/6] where 1-5/6 of v7 are supposed to\nbe identical to their counterparts of v6 and therefor not sent?  Not\ncomplaining, just double-checking, as I do not want to assume and\nend up missing the final updates to the other 5.\n\nIn the meantime, I'd assume that it is the case, and will fix the\nauthor ident (you sent this from your gmail address) before\nreplacing the last bit.  \n\nThe only change from the previous round is \"the platform default on\nmost platform\" -> \"the default on most platforms\", which looks\nsensible.\n\nThanks.\n\n1:  39f4b94c2c ! 1:  dfeab99d23 core.fsync: documentation and user-friendly aggregate options\n    @@\n      ## Metadata ##\n    -Author: Neeraj Singh <neerajsi@microsoft.com>\n    +Author: Neeraj Singh <nksingh85@gmail.com>\n     \n      ## Commit message ##\n         core.fsync: documentation and user-friendly aggregate options\n    @@ Documentation/config/core.txt: core.whitespace::\n     +is ignored.\n     ++\n     +The empty string resets the fsync configuration to the platform\n    -+default. The platform default on most platform is equivalent to\n    ++default. The default on most platforms is equivalent to\n     +`core.fsync=committed,-loose-object`, which has good performance,\n     +but risks losing recent work in the event of an unclean system shutdown.\n     ++\n\nThanks.\n"},{"id":"451424","messageId":"CANQDOdcL7N=GOCf2iz0Pofxd=aR2qUD5uUarnMGO578edUPV_Q@mail.gmail.com","threadId":"57030","inReplyTo":"xmqqv8wfc8qb.fsf@gitster.g","subject":"Re: [PATCH v7] core.fsync: documentation and user-friendly aggregate options","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-15T19:56:23Z","receivedAt":"2022-03-15T19:57:15Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 15, 2022 at 12:32 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > This commit adds aggregate options for the core.fsync setting that are\n> > more user-friendly. These options are specified in terms of 'levels of\n> > safety', indicating which Git operations are considered to be sync\n> > points for durability.\n> >\n> > The new documentation is also included here in its entirety for ease of\n> > review.\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> > This revision fixes a grammatical mistake in the core.fsync documentation.\n>\n> Is this meant to be [PATCH v7 6/6] where 1-5/6 of v7 are supposed to\n> be identical to their counterparts of v6 and therefor not sent?  Not\n> complaining, just double-checking, as I do not want to assume and\n> end up missing the final updates to the other 5.\n>\n> In the meantime, I'd assume that it is the case, and will fix the\n> author ident (you sent this from your gmail address) before\n> replacing the last bit.\n>\n> The only change from the previous round is \"the platform default on\n> most platform\" -> \"the default on most platforms\", which looks\n> sensible.\n>\n> Thanks.\n\nYes, that's right.  I was trying out git-send-email for the first\ntime.   I didn't want to spam the list with a whole new series where\nonly one patch had a change.\n\n>\n> 1:  39f4b94c2c ! 1:  dfeab99d23 core.fsync: documentation and user-friendly aggregate options\n>     @@\n>       ## Metadata ##\n>     -Author: Neeraj Singh <neerajsi@microsoft.com>\n>     +Author: Neeraj Singh <nksingh85@gmail.com>\n>\n>       ## Commit message ##\n>          core.fsync: documentation and user-friendly aggregate options\n>     @@ Documentation/config/core.txt: core.whitespace::\n>      +is ignored.\n>      ++\n>      +The empty string resets the fsync configuration to the platform\n>     -+default. The platform default on most platform is equivalent to\n>     ++default. The default on most platforms is equivalent to\n>      +`core.fsync=committed,-loose-object`, which has good performance,\n>      +but risks losing recent work in the event of an unclean system shutdown.\n>      ++\n>\n> Thanks.\n"},{"id":"451992","messageId":"220323.86fsn8ohg8.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"20220315191245.17990-1-neerajsi@microsoft.com","subject":"do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:20:52Z","receivedAt":"2022-03-23T14:46:06Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Mar 15 2022, Neeraj Singh wrote:\n\nI know this is probably 80% my fault by egging you on about initially\nadding the wildmatch() based thing you didn't go for.\n\nBut having looked at this with fresh eyes quite deeply I really think\nwe're severely over-configuring things here:\n\n> +core.fsync::\n> +\tA comma-separated list of components of the repository that\n> +\tshould be hardened via the core.fsyncMethod when created or\n> +\tmodified.  You can disable hardening of any component by\n> +\tprefixing it with a '-'.  Items that are not hardened may be\n> +\tlost in the event of an unclean\tsystem shutdown. Unless you\n> +\thave special requirements, it is recommended that you leave\n> +\tthis option empty or pick one of `committed`, `added`,\n> +\tor `all`.\n> ++\n> +When this configuration is encountered, the set of components starts with\n> +the platform default value, disabled components are removed, and additional\n> +components are added. `none` resets the state so that the platform default\n> +is ignored.\n> ++\n> +The empty string resets the fsync configuration to the platform\n> +default. The default on most platforms is equivalent to\n> +`core.fsync=committed,-loose-object`, which has good performance,\n> +but risks losing recent work in the event of an unclean system shutdown.\n> ++\n> +* `none` clears the set of fsynced components.\n> +* `loose-object` hardens objects added to the repo in loose-object form.\n> +* `pack` hardens objects added to the repo in packfile form.\n> +* `pack-metadata` hardens packfile bitmaps and indexes.\n> +* `commit-graph` hardens the commit graph file.\n> +* `index` hardens the index when it is modified.\n> +* `objects` is an aggregate option that is equivalent to\n> +  `loose-object,pack`.\n> +* `derived-metadata` is an aggregate option that is equivalent to\n> +  `pack-metadata,commit-graph`.\n> +* `committed` is an aggregate option that is currently equivalent to\n> +  `objects`. This mode sacrifices some performance to ensure that work\n> +  that is committed to the repository with `git commit` or similar commands\n> +  is hardened.\n> +* `added` is an aggregate option that is currently equivalent to\n> +  `committed,index`. This mode sacrifices additional performance to\n> +  ensure that the results of commands like `git add` and similar operations\n> +  are hardened.\n> +* `all` is an aggregate option that syncs all individual components above.\n> +\n>  core.fsyncMethod::\n>  \tA value indicating the strategy Git will use to harden repository data\n>  \tusing fsync and related primitives.\n\nOn top of my\nhttps://lore.kernel.org/git/RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com/\nwhich makes the tmp-objdir part of your not-in-next-just-seen follow-up\nseries configurable via \"fsyncMethod.batch.quarantine\" I really think we\nshould just go for something like the belwo patch (note that\nmisspelled/mistook \"bulk\" for \"batch\" in that linked-t patch, fixed\nbelow.\n\nI.e. I think we should just do our default fsync() of everything, and\nprobably SOON make the fsync-ing of loose objects the default. Those who\ncare about performance will have \"batch\" (or \"writeout-only\"), which we\ncan have OS-specific detections for.\n\nBut really, all of the rest of this is unduly boxing us into\noverconfiguration that I think nobody really needs.\n\nIf someone really needs this level of detail they can LD_PRELOAD\nsomething to have fsync intercept fd's and paths, and act appropriately.\n\nWorse, as the RFC series I sent\n(https://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/)\nshows we can and should \"batch\" up fsync() operations across these\nconfiguration boundaries, which this level of configuration would seem\nto preclude.\n\nOr, we'd need to explain why \"core.fsync=loose-object\" won't *actually*\ncall fsync() on a single loose object's fd under \"batch\" as I had to do\non top of this in\nhttps://lore.kernel.org/git/RFC-patch-v2-6.7-c20301d7967-20220323T140753Z-avarab@gmail.com/\n\nThe same is going to apply for almost all of the rest of these\nconfiguration categories.\n\nI.e. a natural follow-up to e.g. batching across objects & index as I'm\ndoing in\nhttps://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\nis to do likewise for all the PACK-related stuff before we rename it\nin-place. Or even have \"git gc\" issue only a single fsync() for all of\nPACKs, their metadata files, commit-graph etc., and then rename() things\nin-place as appropriate afterwards.\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 365a12dc7ae..536238e209b 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,49 +548,35 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsync::\n-\tA comma-separated list of components of the repository that\n-\tshould be hardened via the core.fsyncMethod when created or\n-\tmodified.  You can disable hardening of any component by\n-\tprefixing it with a '-'.  Items that are not hardened may be\n-\tlost in the event of an unclean\tsystem shutdown. Unless you\n-\thave special requirements, it is recommended that you leave\n-\tthis option empty or pick one of `committed`, `added`,\n-\tor `all`.\n-+\n-When this configuration is encountered, the set of components starts with\n-the platform default value, disabled components are removed, and additional\n-components are added. `none` resets the state so that the platform default\n-is ignored.\n-+\n-The empty string resets the fsync configuration to the platform\n-default. The default on most platforms is equivalent to\n-`core.fsync=committed,-loose-object`, which has good performance,\n-but risks losing recent work in the event of an unclean system shutdown.\n-+\n-* `none` clears the set of fsynced components.\n-* `loose-object` hardens objects added to the repo in loose-object form.\n-* `pack` hardens objects added to the repo in packfile form.\n-* `pack-metadata` hardens packfile bitmaps and indexes.\n-* `commit-graph` hardens the commit graph file.\n-* `index` hardens the index when it is modified.\n-* `objects` is an aggregate option that is equivalent to\n-  `loose-object,pack`.\n-* `derived-metadata` is an aggregate option that is equivalent to\n-  `pack-metadata,commit-graph`.\n-* `committed` is an aggregate option that is currently equivalent to\n-  `objects`. This mode sacrifices some performance to ensure that work\n-  that is committed to the repository with `git commit` or similar commands\n-  is hardened.\n-* `added` is an aggregate option that is currently equivalent to\n-  `committed,index`. This mode sacrifices additional performance to\n-  ensure that the results of commands like `git add` and similar operations\n-  are hardened.\n-* `all` is an aggregate option that syncs all individual components above.\n+\tA boolen defaulting to `true`. To ensure data integrity git\n+\twill fsync() its objects, index and refu updates etc. This can\n+\tbe set to `false` to disable `fsync()`-ing.\n++\n+Only set this to `false` if you know what you're doing, and are\n+prepared to deal with data corruption. Valid use-cases include\n+throwaway uses of repositories on ramdisks, one-off mass-imports\n+followed by calling `sync(1)` etc.\n++\n+Note that the syncing of loose objects is currently excluded from\n+`core.fsync=true`. To turn on all fsync-ing you'll need\n+`core.fsync=true` and `core.fsyncObjectFiles=true`, but see\n+`core.fsyncMethod=batch` below for a much faster alternative that's\n+just as safe on various modern OS's.\n++\n+The default is in flux and may change in the future, in particular the\n+equivalent of the already-deprecated `core.fsyncObjectFiles` setting\n+might soon default to `true`, and `core.fsyncMethod`'s default of\n+`fsync` might default to a setting deemed to be safe on the local OS,\n+suc has `batch` or `writeout-only`\n \n core.fsyncMethod::\n \tA value indicating the strategy Git will use to harden repository data\n \tusing fsync and related primitives.\n +\n+Defaults to `fsync`, but as discussed for `core.fsync` above might\n+change to use one of the values below taking advantage of\n+platform-specific \"faster `fsync()`\".\n++\n * `fsync` uses the fsync() system call or platform equivalents.\n * `writeout-only` issues pagecache writeback requests, but depending on the\n   filesystem and storage hardware, data added to the repository may not be\n@@ -680,8 +666,8 @@ backed up by any standard (e.g. POSIX), but worked in practice on some\n Linux setups.\n +\n Nowadays you should almost certainly want to use\n-`core.fsync=loose-object` instead in combination with\n-`core.fsyncMethod=bulk`, and possibly with\n+`core.fsync=true` instead in combination with\n+`core.fsyncMethod=batch`, and possibly with\n `fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n OSX, Windows) that gives you most of the performance benefit of\n `core.fsyncObjectFiles=false` with all of the safety of the old\n"},{"id":"452255","messageId":"20220325061149.GA2571@szeder.dev","threadId":"57030","inReplyTo":"47dd79106b93bb81750320d50ccaa74c24aacd28.1646992380.git.ps@pks.im","subject":"Re: [PATCH v2] core.fsync: new option to harden references","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2022-03-25T06:11:49Z","receivedAt":"2022-03-25T06:12:00Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Fri, Mar 11, 2022 at 10:58:59AM +0100, Patrick Steinhardt wrote:\n> When writing both loose and packed references to disk we first create a\n> lockfile, write the updated values into that lockfile, and on commit we\n> rename the file into place. According to filesystem developers, this\n> behaviour is broken because applications should always sync data to disk\n> before doing the final rename to ensure data consistency [1][2][3]. If\n> applications fail to do this correctly, a hard crash of the machine can\n> easily result in corrupted on-disk data.\n> \n> This kind of corruption can in fact be easily observed with Git when the\n> machine hard-resets shortly after writing references to disk. On\n> machines with ext4, this will likely lead to the \"empty files\" problem:\n> the file has been renamed, but its data has not been synced to disk. The\n> result is that the reference is corrupt, and in the worst case this can\n> lead to data loss.\n> \n> Implement a new option to harden references so that users and admins can\n> avoid this scenario by syncing locked loose and packed references to\n> disk before we rename them into place.\n\nIn 't5541-http-push-smart.sh' there is a test case called 'push 2000\ntags over http', which does pretty much what it's title says.  This\npatch makes that test case significantly slower.\n\n\ndiff --git a/t/t5541-http-push-smart.sh b/t/t5541-http-push-smart.sh\nindex 8ca50f8b18..d7e94cb791 100755\n--- a/t/t5541-http-push-smart.sh\n+++ b/t/t5541-http-push-smart.sh\n@@ -415,7 +415,7 @@ test_expect_success CMDLINE_LIMIT 'push 2000 tags over http' '\n \t  sort |\n \t  sed \"s|.*|$sha1 refs/tags/really-long-tag-name-&|\" \\\n \t  >.git/packed-refs &&\n-\trun_with_limited_cmdline git push --mirror\n+\trun_with_limited_cmdline /usr/bin/time git push --mirror\n '\n \n test_expect_success GPG 'push with post-receive to inspect certificate' '\n\nBefore this patch (bc22d845c4^) 'time' reports:\n\n  3.62user 0.03system 0:03.83elapsed 95%CPU (0avgtext+0avgdata 11904maxresident)k\n  0inputs+312outputs (0major+4597minor)pagefaults 0swaps\n\nWith this patch (bc22d845c4):\n\n  3.56user 0.04system 0:33.60elapsed 10%CPU (0avgtext+0avgdata 11832maxresident)k\n  0inputs+320outputs (0major+4578minor)pagefaults 0swaps\n\nAnd the total runtime of the whole test script increases from 8-9s to\n37-39s.\n\n\nI wonder whether we should relax the fsync options for this test case.\n\n> \n> [1]: https://thunk.org/tytso/blog/2009/03/15/dont-fear-the-fsync/\n> [2]: https://btrfs.wiki.kernel.org/index.php/FAQ (What are the crash guarantees of overwrite-by-rename)\n> [3]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/Documentation/admin-guide/ext4.rst (see auto_da_alloc)\n> \n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n> \n> Hi,\n> \n> here's my updated patch series which implements syncing of refs. It\n> applies on top of Neeraj's v6 of \"A design for future-proofing fsync()\n> configuration\".\n> \n> I've simplified these patches a bit:\n> \n>     - I don't distinguishing between \"loose\" and \"packed\" refs anymore.\n>       I agree with Junio that it's probably not worth it, but we can\n>       still reintroduce the split at a later point without breaking\n>       backwards compatibility if the need comes up.\n> \n>     - I've simplified the way loose refs are written to disk so that we\n>       now sync them when before we close their files. The previous\n>       implementation I had was broken because we tried to sync after\n>       closing.\n> \n> Because this really only changes a few lines of code I've also decided\n> to squash together the patches into a single one.\n> \n> Patrick\n> \n>  Documentation/config/core.txt | 1 +\n>  cache.h                       | 7 +++++--\n>  config.c                      | 1 +\n>  refs/files-backend.c          | 1 +\n>  refs/packed-backend.c         | 3 ++-\n>  5 files changed, 10 insertions(+), 3 deletions(-)\n> \n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index 37105a7be4..812cca7de7 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -575,6 +575,7 @@ but risks losing recent work in the event of an unclean system shutdown.\n>  * `index` hardens the index when it is modified.\n>  * `objects` is an aggregate option that is equivalent to\n>    `loose-object,pack`.\n> +* `reference` hardens references modified in the repo.\n>  * `derived-metadata` is an aggregate option that is equivalent to\n>    `pack-metadata,commit-graph`.\n>  * `committed` is an aggregate option that is currently equivalent to\n> diff --git a/cache.h b/cache.h\n> index cde0900d05..033e5b0779 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1005,6 +1005,7 @@ enum fsync_component {\n>  \tFSYNC_COMPONENT_PACK_METADATA\t\t= 1 << 2,\n>  \tFSYNC_COMPONENT_COMMIT_GRAPH\t\t= 1 << 3,\n>  \tFSYNC_COMPONENT_INDEX\t\t\t= 1 << 4,\n> +\tFSYNC_COMPONENT_REFERENCE\t\t= 1 << 5,\n>  };\n>  \n>  #define FSYNC_COMPONENTS_OBJECTS (FSYNC_COMPONENT_LOOSE_OBJECT | \\\n> @@ -1017,7 +1018,8 @@ enum fsync_component {\n>  \t\t\t\t  FSYNC_COMPONENTS_DERIVED_METADATA | \\\n>  \t\t\t\t  ~FSYNC_COMPONENT_LOOSE_OBJECT)\n>  \n> -#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS)\n> +#define FSYNC_COMPONENTS_COMMITTED (FSYNC_COMPONENTS_OBJECTS | \\\n> +\t\t\t\t    FSYNC_COMPONENT_REFERENCE)\n>  \n>  #define FSYNC_COMPONENTS_ADDED (FSYNC_COMPONENTS_COMMITTED | \\\n>  \t\t\t\tFSYNC_COMPONENT_INDEX)\n> @@ -1026,7 +1028,8 @@ enum fsync_component {\n>  \t\t\t      FSYNC_COMPONENT_PACK | \\\n>  \t\t\t      FSYNC_COMPONENT_PACK_METADATA | \\\n>  \t\t\t      FSYNC_COMPONENT_COMMIT_GRAPH | \\\n> -\t\t\t      FSYNC_COMPONENT_INDEX)\n> +\t\t\t      FSYNC_COMPONENT_INDEX | \\\n> +\t\t\t      FSYNC_COMPONENT_REFERENCE)\n>  \n>  /*\n>   * A bitmask indicating which components of the repo should be fsynced.\n> diff --git a/config.c b/config.c\n> index eb75f65338..3c9b6b589a 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1333,6 +1333,7 @@ static const struct fsync_component_name {\n>  \t{ \"commit-graph\", FSYNC_COMPONENT_COMMIT_GRAPH },\n>  \t{ \"index\", FSYNC_COMPONENT_INDEX },\n>  \t{ \"objects\", FSYNC_COMPONENTS_OBJECTS },\n> +\t{ \"reference\", FSYNC_COMPONENT_REFERENCE },\n>  \t{ \"derived-metadata\", FSYNC_COMPONENTS_DERIVED_METADATA },\n>  \t{ \"committed\", FSYNC_COMPONENTS_COMMITTED },\n>  \t{ \"added\", FSYNC_COMPONENTS_ADDED },\n> diff --git a/refs/files-backend.c b/refs/files-backend.c\n> index f59589d6cc..6521ee8af5 100644\n> --- a/refs/files-backend.c\n> +++ b/refs/files-backend.c\n> @@ -1787,6 +1787,7 @@ static int write_ref_to_lockfile(struct ref_lock *lock,\n>  \tfd = get_lock_file_fd(&lock->lk);\n>  \tif (write_in_full(fd, oid_to_hex(oid), the_hash_algo->hexsz) < 0 ||\n>  \t    write_in_full(fd, &term, 1) < 0 ||\n> +\t    fsync_component(FSYNC_COMPONENT_REFERENCE, get_lock_file_fd(&lock->lk)) < 0 ||\n>  \t    close_ref_gently(lock) < 0) {\n>  \t\tstrbuf_addf(err,\n>  \t\t\t    \"couldn't write '%s'\", get_lock_file_path(&lock->lk));\n> diff --git a/refs/packed-backend.c b/refs/packed-backend.c\n> index 27dd8c3922..9d704ccd3e 100644\n> --- a/refs/packed-backend.c\n> +++ b/refs/packed-backend.c\n> @@ -1262,7 +1262,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n>  \t\tgoto error;\n>  \t}\n>  \n> -\tif (close_tempfile_gently(refs->tempfile)) {\n> +\tif (fsync_component(FSYNC_COMPONENT_REFERENCE, get_tempfile_fd(refs->tempfile)) ||\n> +\t    close_tempfile_gently(refs->tempfile)) {\n>  \t\tstrbuf_addf(err, \"error closing file %s: %s\",\n>  \t\t\t    get_tempfile_path(refs->tempfile),\n>  \t\t\t    strerror(errno));\n> -- \n> 2.35.1\n> \n\n\n"},{"id":"452390","messageId":"CANQDOdeeP8opTQj-j_j3=KnU99nYTnNYhyQmAojj=FZtZEkCZQ@mail.gmail.com","threadId":"57030","inReplyTo":"220323.86fsn8ohg8.gmgdl@evledraar.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-25T21:24:57Z","receivedAt":"2022-03-25T21:25:15Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 23, 2022 at 7:46 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Tue, Mar 15 2022, Neeraj Singh wrote:\n>\n> I know this is probably 80% my fault by egging you on about initially\n> adding the wildmatch() based thing you didn't go for.\n>\n> But having looked at this with fresh eyes quite deeply I really think\n> we're severely over-configuring things here:\n>\n> > +core.fsync::\n> > +     A comma-separated list of components of the repository that\n> > +     should be hardened via the core.fsyncMethod when created or\n> > +     modified.  You can disable hardening of any component by\n> > +     prefixing it with a '-'.  Items that are not hardened may be\n> > +     lost in the event of an unclean system shutdown. Unless you\n> > +     have special requirements, it is recommended that you leave\n> > +     this option empty or pick one of `committed`, `added`,\n> > +     or `all`.\n> > ++\n> > +When this configuration is encountered, the set of components starts with\n> > +the platform default value, disabled components are removed, and additional\n> > +components are added. `none` resets the state so that the platform default\n> > +is ignored.\n> > ++\n> > +The empty string resets the fsync configuration to the platform\n> > +default. The default on most platforms is equivalent to\n> > +`core.fsync=committed,-loose-object`, which has good performance,\n> > +but risks losing recent work in the event of an unclean system shutdown.\n> > ++\n> > +* `none` clears the set of fsynced components.\n> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> > +* `pack` hardens objects added to the repo in packfile form.\n> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> > +* `commit-graph` hardens the commit graph file.\n> > +* `index` hardens the index when it is modified.\n> > +* `objects` is an aggregate option that is equivalent to\n> > +  `loose-object,pack`.\n> > +* `derived-metadata` is an aggregate option that is equivalent to\n> > +  `pack-metadata,commit-graph`.\n> > +* `committed` is an aggregate option that is currently equivalent to\n> > +  `objects`. This mode sacrifices some performance to ensure that work\n> > +  that is committed to the repository with `git commit` or similar commands\n> > +  is hardened.\n> > +* `added` is an aggregate option that is currently equivalent to\n> > +  `committed,index`. This mode sacrifices additional performance to\n> > +  ensure that the results of commands like `git add` and similar operations\n> > +  are hardened.\n> > +* `all` is an aggregate option that syncs all individual components above.\n> > +\n> >  core.fsyncMethod::\n> >       A value indicating the strategy Git will use to harden repository data\n> >       using fsync and related primitives.\n>\n> On top of my\n> https://lore.kernel.org/git/RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com/\n> which makes the tmp-objdir part of your not-in-next-just-seen follow-up\n> series configurable via \"fsyncMethod.batch.quarantine\" I really think we\n> should just go for something like the belwo patch (note that\n> misspelled/mistook \"bulk\" for \"batch\" in that linked-t patch, fixed\n> below.\n>\n> I.e. I think we should just do our default fsync() of everything, and\n> probably SOON make the fsync-ing of loose objects the default. Those who\n> care about performance will have \"batch\" (or \"writeout-only\"), which we\n> can have OS-specific detections for.\n>\n> But really, all of the rest of this is unduly boxing us into\n> overconfiguration that I think nobody really needs.\n>\n\nWe've gone over this a few times already, but just wanted to state it\nagain.  The really detailed settings are really there for Git hosters\nlike GitLab or GitHub. I'd be happy to remove the per-component\ncore.fsync values from the documentation and leave just the ones we\npoint the user to.\n\n> If someone really needs this level of detail they can LD_PRELOAD\n> something to have fsync intercept fd's and paths, and act appropriately.\n>\n> Worse, as the RFC series I sent\n> (https://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/)\n> shows we can and should \"batch\" up fsync() operations across these\n> configuration boundaries, which this level of configuration would seem\n> to preclude.\n>\n> Or, we'd need to explain why \"core.fsync=loose-object\" won't *actually*\n> call fsync() on a single loose object's fd under \"batch\" as I had to do\n> on top of this in\n> https://lore.kernel.org/git/RFC-patch-v2-6.7-c20301d7967-20220323T140753Z-avarab@gmail.com/\n>\n\n99.9% of users don't care and won't look.  The ones who do look deeper\nand understand the issues have source code and access to this ML\ndiscussion to understand why this works this way.\n\n> The same is going to apply for almost all of the rest of these\n> configuration categories.\n>\n> I.e. a natural follow-up to e.g. batching across objects & index as I'm\n> doing in\n> https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n> is to do likewise for all the PACK-related stuff before we rename it\n> in-place. Or even have \"git gc\" issue only a single fsync() for all of\n> PACKs, their metadata files, commit-graph etc., and then rename() things\n> in-place as appropriate afterwards.\n>\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index 365a12dc7ae..536238e209b 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -548,49 +548,35 @@ core.whitespace::\n>    errors. The default tab width is 8. Allowed values are 1 to 63.\n>\n>  core.fsync::\n> -       A comma-separated list of components of the repository that\n> -       should be hardened via the core.fsyncMethod when created or\n> -       modified.  You can disable hardening of any component by\n> -       prefixing it with a '-'.  Items that are not hardened may be\n> -       lost in the event of an unclean system shutdown. Unless you\n> -       have special requirements, it is recommended that you leave\n> -       this option empty or pick one of `committed`, `added`,\n> -       or `all`.\n> -+\n> -When this configuration is encountered, the set of components starts with\n> -the platform default value, disabled components are removed, and additional\n> -components are added. `none` resets the state so that the platform default\n> -is ignored.\n> -+\n> -The empty string resets the fsync configuration to the platform\n> -default. The default on most platforms is equivalent to\n> -`core.fsync=committed,-loose-object`, which has good performance,\n> -but risks losing recent work in the event of an unclean system shutdown.\n> -+\n> -* `none` clears the set of fsynced components.\n> -* `loose-object` hardens objects added to the repo in loose-object form.\n> -* `pack` hardens objects added to the repo in packfile form.\n> -* `pack-metadata` hardens packfile bitmaps and indexes.\n> -* `commit-graph` hardens the commit graph file.\n> -* `index` hardens the index when it is modified.\n> -* `objects` is an aggregate option that is equivalent to\n> -  `loose-object,pack`.\n> -* `derived-metadata` is an aggregate option that is equivalent to\n> -  `pack-metadata,commit-graph`.\n> -* `committed` is an aggregate option that is currently equivalent to\n> -  `objects`. This mode sacrifices some performance to ensure that work\n> -  that is committed to the repository with `git commit` or similar commands\n> -  is hardened.\n> -* `added` is an aggregate option that is currently equivalent to\n> -  `committed,index`. This mode sacrifices additional performance to\n> -  ensure that the results of commands like `git add` and similar operations\n> -  are hardened.\n> -* `all` is an aggregate option that syncs all individual components above.\n> +       A boolen defaulting to `true`. To ensure data integrity git\n> +       will fsync() its objects, index and refu updates etc. This can\n> +       be set to `false` to disable `fsync()`-ing.\n> ++\n> +Only set this to `false` if you know what you're doing, and are\n> +prepared to deal with data corruption. Valid use-cases include\n> +throwaway uses of repositories on ramdisks, one-off mass-imports\n> +followed by calling `sync(1)` etc.\n> ++\n> +Note that the syncing of loose objects is currently excluded from\n> +`core.fsync=true`. To turn on all fsync-ing you'll need\n> +`core.fsync=true` and `core.fsyncObjectFiles=true`, but see\n> +`core.fsyncMethod=batch` below for a much faster alternative that's\n> +just as safe on various modern OS's.\n> ++\n> +The default is in flux and may change in the future, in particular the\n> +equivalent of the already-deprecated `core.fsyncObjectFiles` setting\n> +might soon default to `true`, and `core.fsyncMethod`'s default of\n> +`fsync` might default to a setting deemed to be safe on the local OS,\n> +suc has `batch` or `writeout-only`\n>\n>  core.fsyncMethod::\n>         A value indicating the strategy Git will use to harden repository data\n>         using fsync and related primitives.\n>  +\n> +Defaults to `fsync`, but as discussed for `core.fsync` above might\n> +change to use one of the values below taking advantage of\n> +platform-specific \"faster `fsync()`\".\n> ++\n>  * `fsync` uses the fsync() system call or platform equivalents.\n>  * `writeout-only` issues pagecache writeback requests, but depending on the\n>    filesystem and storage hardware, data added to the repository may not be\n> @@ -680,8 +666,8 @@ backed up by any standard (e.g. POSIX), but worked in practice on some\n>  Linux setups.\n>  +\n>  Nowadays you should almost certainly want to use\n> -`core.fsync=loose-object` instead in combination with\n> -`core.fsyncMethod=bulk`, and possibly with\n> +`core.fsync=true` instead in combination with\n> +`core.fsyncMethod=batch`, and possibly with\n>  `fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n>  OSX, Windows) that gives you most of the performance benefit of\n>  `core.fsyncObjectFiles=false` with all of the safety of the old\n\nI'm at the point where I don't want to endlessly revisit this discussion.\n\n-Neeraj\n"},{"id":"452396","messageId":"220326.86ils1lfho.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"CANQDOdeeP8opTQj-j_j3=KnU99nYTnNYhyQmAojj=FZtZEkCZQ@mail.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-26T00:24:57Z","receivedAt":"2022-03-26T00:33:30Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Mar 25 2022, Neeraj Singh wrote:\n\n> On Wed, Mar 23, 2022 at 7:46 AM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Tue, Mar 15 2022, Neeraj Singh wrote:\n>>\n>> I know this is probably 80% my fault by egging you on about initially\n>> adding the wildmatch() based thing you didn't go for.\n>>\n>> But having looked at this with fresh eyes quite deeply I really think\n>> we're severely over-configuring things here:\n>>\n>> > +core.fsync::\n>> > +     A comma-separated list of components of the repository that\n>> > +     should be hardened via the core.fsyncMethod when created or\n>> > +     modified.  You can disable hardening of any component by\n>> > +     prefixing it with a '-'.  Items that are not hardened may be\n>> > +     lost in the event of an unclean system shutdown. Unless you\n>> > +     have special requirements, it is recommended that you leave\n>> > +     this option empty or pick one of `committed`, `added`,\n>> > +     or `all`.\n>> > ++\n>> > +When this configuration is encountered, the set of components starts with\n>> > +the platform default value, disabled components are removed, and additional\n>> > +components are added. `none` resets the state so that the platform default\n>> > +is ignored.\n>> > ++\n>> > +The empty string resets the fsync configuration to the platform\n>> > +default. The default on most platforms is equivalent to\n>> > +`core.fsync=committed,-loose-object`, which has good performance,\n>> > +but risks losing recent work in the event of an unclean system shutdown.\n>> > ++\n>> > +* `none` clears the set of fsynced components.\n>> > +* `loose-object` hardens objects added to the repo in loose-object form.\n>> > +* `pack` hardens objects added to the repo in packfile form.\n>> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n>> > +* `commit-graph` hardens the commit graph file.\n>> > +* `index` hardens the index when it is modified.\n>> > +* `objects` is an aggregate option that is equivalent to\n>> > +  `loose-object,pack`.\n>> > +* `derived-metadata` is an aggregate option that is equivalent to\n>> > +  `pack-metadata,commit-graph`.\n>> > +* `committed` is an aggregate option that is currently equivalent to\n>> > +  `objects`. This mode sacrifices some performance to ensure that work\n>> > +  that is committed to the repository with `git commit` or similar commands\n>> > +  is hardened.\n>> > +* `added` is an aggregate option that is currently equivalent to\n>> > +  `committed,index`. This mode sacrifices additional performance to\n>> > +  ensure that the results of commands like `git add` and similar operations\n>> > +  are hardened.\n>> > +* `all` is an aggregate option that syncs all individual components above.\n>> > +\n>> >  core.fsyncMethod::\n>> >       A value indicating the strategy Git will use to harden repository data\n>> >       using fsync and related primitives.\n>>\n>> On top of my\n>> https://lore.kernel.org/git/RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com/\n>> which makes the tmp-objdir part of your not-in-next-just-seen follow-up\n>> series configurable via \"fsyncMethod.batch.quarantine\" I really think we\n>> should just go for something like the belwo patch (note that\n>> misspelled/mistook \"bulk\" for \"batch\" in that linked-t patch, fixed\n>> below.\n>>\n>> I.e. I think we should just do our default fsync() of everything, and\n>> probably SOON make the fsync-ing of loose objects the default. Those who\n>> care about performance will have \"batch\" (or \"writeout-only\"), which we\n>> can have OS-specific detections for.\n>>\n>> But really, all of the rest of this is unduly boxing us into\n>> overconfiguration that I think nobody really needs.\n>>\n>\n> We've gone over this a few times already, but just wanted to state it\n> again.  The really detailed settings are really there for Git hosters\n> like GitLab or GitHub. I'd be happy to remove the per-component\n> core.fsync values from the documentation and leave just the ones we\n> point the user to.\n\nI'm prettty sure (but Patrick knows more) that GitLab's plan for this is\nto keep it at whatever the safest setting is, presumably GitHub's as\nwell (but I don't know at all on that front).\n\n>> If someone really needs this level of detail they can LD_PRELOAD\n>> something to have fsync intercept fd's and paths, and act appropriately.\n>>\n>> Worse, as the RFC series I sent\n>> (https://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/)\n>> shows we can and should \"batch\" up fsync() operations across these\n>> configuration boundaries, which this level of configuration would seem\n>> to preclude.\n>>\n>> Or, we'd need to explain why \"core.fsync=loose-object\" won't *actually*\n>> call fsync() on a single loose object's fd under \"batch\" as I had to do\n>> on top of this in\n>> https://lore.kernel.org/git/RFC-patch-v2-6.7-c20301d7967-20220323T140753Z-avarab@gmail.com/\n>>\n>\n> 99.9% of users don't care and won't look.  The ones who do look deeper\n> and understand the issues have source code and access to this ML\n> discussion to understand why this works this way.\n\nExactly, so we can hopefully have a simpler interface.\n\n>> The same is going to apply for almost all of the rest of these\n>> configuration categories.\n>>\n>> I.e. a natural follow-up to e.g. batching across objects & index as I'm\n>> doing in\n>> https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n>> is to do likewise for all the PACK-related stuff before we rename it\n>> in-place. Or even have \"git gc\" issue only a single fsync() for all of\n>> PACKs, their metadata files, commit-graph etc., and then rename() things\n>> in-place as appropriate afterwards.\n>>\n>> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n>> index 365a12dc7ae..536238e209b 100644\n>> --- a/Documentation/config/core.txt\n>> +++ b/Documentation/config/core.txt\n>> @@ -548,49 +548,35 @@ core.whitespace::\n>>    errors. The default tab width is 8. Allowed values are 1 to 63.\n>>\n>>  core.fsync::\n>> -       A comma-separated list of components of the repository that\n>> -       should be hardened via the core.fsyncMethod when created or\n>> -       modified.  You can disable hardening of any component by\n>> -       prefixing it with a '-'.  Items that are not hardened may be\n>> -       lost in the event of an unclean system shutdown. Unless you\n>> -       have special requirements, it is recommended that you leave\n>> -       this option empty or pick one of `committed`, `added`,\n>> -       or `all`.\n>> -+\n>> -When this configuration is encountered, the set of components starts with\n>> -the platform default value, disabled components are removed, and additional\n>> -components are added. `none` resets the state so that the platform default\n>> -is ignored.\n>> -+\n>> -The empty string resets the fsync configuration to the platform\n>> -default. The default on most platforms is equivalent to\n>> -`core.fsync=committed,-loose-object`, which has good performance,\n>> -but risks losing recent work in the event of an unclean system shutdown.\n>> -+\n>> -* `none` clears the set of fsynced components.\n>> -* `loose-object` hardens objects added to the repo in loose-object form.\n>> -* `pack` hardens objects added to the repo in packfile form.\n>> -* `pack-metadata` hardens packfile bitmaps and indexes.\n>> -* `commit-graph` hardens the commit graph file.\n>> -* `index` hardens the index when it is modified.\n>> -* `objects` is an aggregate option that is equivalent to\n>> -  `loose-object,pack`.\n>> -* `derived-metadata` is an aggregate option that is equivalent to\n>> -  `pack-metadata,commit-graph`.\n>> -* `committed` is an aggregate option that is currently equivalent to\n>> -  `objects`. This mode sacrifices some performance to ensure that work\n>> -  that is committed to the repository with `git commit` or similar commands\n>> -  is hardened.\n>> -* `added` is an aggregate option that is currently equivalent to\n>> -  `committed,index`. This mode sacrifices additional performance to\n>> -  ensure that the results of commands like `git add` and similar operations\n>> -  are hardened.\n>> -* `all` is an aggregate option that syncs all individual components above.\n>> +       A boolen defaulting to `true`. To ensure data integrity git\n>> +       will fsync() its objects, index and refu updates etc. This can\n>> +       be set to `false` to disable `fsync()`-ing.\n>> ++\n>> +Only set this to `false` if you know what you're doing, and are\n>> +prepared to deal with data corruption. Valid use-cases include\n>> +throwaway uses of repositories on ramdisks, one-off mass-imports\n>> +followed by calling `sync(1)` etc.\n>> ++\n>> +Note that the syncing of loose objects is currently excluded from\n>> +`core.fsync=true`. To turn on all fsync-ing you'll need\n>> +`core.fsync=true` and `core.fsyncObjectFiles=true`, but see\n>> +`core.fsyncMethod=batch` below for a much faster alternative that's\n>> +just as safe on various modern OS's.\n>> ++\n>> +The default is in flux and may change in the future, in particular the\n>> +equivalent of the already-deprecated `core.fsyncObjectFiles` setting\n>> +might soon default to `true`, and `core.fsyncMethod`'s default of\n>> +`fsync` might default to a setting deemed to be safe on the local OS,\n>> +suc has `batch` or `writeout-only`\n>>\n>>  core.fsyncMethod::\n>>         A value indicating the strategy Git will use to harden repository data\n>>         using fsync and related primitives.\n>>  +\n>> +Defaults to `fsync`, but as discussed for `core.fsync` above might\n>> +change to use one of the values below taking advantage of\n>> +platform-specific \"faster `fsync()`\".\n>> ++\n>>  * `fsync` uses the fsync() system call or platform equivalents.\n>>  * `writeout-only` issues pagecache writeback requests, but depending on the\n>>    filesystem and storage hardware, data added to the repository may not be\n>> @@ -680,8 +666,8 @@ backed up by any standard (e.g. POSIX), but worked in practice on some\n>>  Linux setups.\n>>  +\n>>  Nowadays you should almost certainly want to use\n>> -`core.fsync=loose-object` instead in combination with\n>> -`core.fsyncMethod=bulk`, and possibly with\n>> +`core.fsync=true` instead in combination with\n>> +`core.fsyncMethod=batch`, and possibly with\n>>  `fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n>>  OSX, Windows) that gives you most of the performance benefit of\n>>  `core.fsyncObjectFiles=false` with all of the safety of the old\n>\n> I'm at the point where I don't want to endlessly revisit this discussion.\n\nSorry, my intention isn't to frustrate you, but I do think it's\nimportant to get this right.\n\nParticularly since this is now in \"next\", and we're getting closer to a\nrelease. We can either talk about this now and decide on something, or\nit'll be in a release, and then publicly documented promises will be\nharder to back out of.\n\nI think your suggestion of just hiding the relevant documentation would\nbe a good band-aid solution to that.\n\nBut I also think that given how I was altering this in my RFC series\nthat the premise of how this could be structured has been called into\nquestion in a way that we didn't (or I don't recall) us having discussed\nbefore.\n\nI.e. that we can say \"sync loose, but not index\", or \"sync index, but\nnot loose\" with this config schema. When with \"bulk\" we it really isn't\nany more expensive to do both if one is true (even cheaper, actually).\n\n"},{"id":"452402","messageId":"xmqqk0ch7bhi.fsf@gitster.g","threadId":"57030","inReplyTo":"220326.86ils1lfho.gmgdl@evledraar.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-26T01:23:37Z","receivedAt":"2022-03-26T01:23:44Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> We've gone over this a few times already, but just wanted to state it\n>> again.  The really detailed settings are really there for Git hosters\n>> like GitLab or GitHub. I'd be happy to remove the per-component\n>> core.fsync values from the documentation and leave just the ones we\n>> point the user to.\n>\n> I'm prettty sure (but Patrick knows more) that GitLab's plan for this is\n> to keep it at whatever the safest setting is, presumably GitHub's as\n> well (but I don't know at all on that front).\n\nI thought we've already settled it long ago, and it wasn't like you\nwere taking vacation from the list for a few weeks.  Why are we\ntalking about this again now?  Did we discover any new argument\nagainst it?\n\nPuzzled.\n"},{"id":"452404","messageId":"CANQDOdeduc8bFA_=R-kXmkM+nb__oTxVhjBfFYj70vCFew1EyA@mail.gmail.com","threadId":"57030","inReplyTo":"220326.86ils1lfho.gmgdl@evledraar.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-26T01:25:38Z","receivedAt":"2022-03-26T01:25:54Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Mar 25, 2022 at 5:33 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Fri, Mar 25 2022, Neeraj Singh wrote:\n>\n> > On Wed, Mar 23, 2022 at 7:46 AM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Tue, Mar 15 2022, Neeraj Singh wrote:\n> >>\n> >> I know this is probably 80% my fault by egging you on about initially\n> >> adding the wildmatch() based thing you didn't go for.\n> >>\n> >> But having looked at this with fresh eyes quite deeply I really think\n> >> we're severely over-configuring things here:\n> >>\n> >> > +core.fsync::\n> >> > +     A comma-separated list of components of the repository that\n> >> > +     should be hardened via the core.fsyncMethod when created or\n> >> > +     modified.  You can disable hardening of any component by\n> >> > +     prefixing it with a '-'.  Items that are not hardened may be\n> >> > +     lost in the event of an unclean system shutdown. Unless you\n> >> > +     have special requirements, it is recommended that you leave\n> >> > +     this option empty or pick one of `committed`, `added`,\n> >> > +     or `all`.\n> >> > ++\n> >> > +When this configuration is encountered, the set of components starts with\n> >> > +the platform default value, disabled components are removed, and additional\n> >> > +components are added. `none` resets the state so that the platform default\n> >> > +is ignored.\n> >> > ++\n> >> > +The empty string resets the fsync configuration to the platform\n> >> > +default. The default on most platforms is equivalent to\n> >> > +`core.fsync=committed,-loose-object`, which has good performance,\n> >> > +but risks losing recent work in the event of an unclean system shutdown.\n> >> > ++\n> >> > +* `none` clears the set of fsynced components.\n> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> >> > +* `pack` hardens objects added to the repo in packfile form.\n> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> >> > +* `commit-graph` hardens the commit graph file.\n> >> > +* `index` hardens the index when it is modified.\n> >> > +* `objects` is an aggregate option that is equivalent to\n> >> > +  `loose-object,pack`.\n> >> > +* `derived-metadata` is an aggregate option that is equivalent to\n> >> > +  `pack-metadata,commit-graph`.\n> >> > +* `committed` is an aggregate option that is currently equivalent to\n> >> > +  `objects`. This mode sacrifices some performance to ensure that work\n> >> > +  that is committed to the repository with `git commit` or similar commands\n> >> > +  is hardened.\n> >> > +* `added` is an aggregate option that is currently equivalent to\n> >> > +  `committed,index`. This mode sacrifices additional performance to\n> >> > +  ensure that the results of commands like `git add` and similar operations\n> >> > +  are hardened.\n> >> > +* `all` is an aggregate option that syncs all individual components above.\n> >> > +\n> >> >  core.fsyncMethod::\n> >> >       A value indicating the strategy Git will use to harden repository data\n> >> >       using fsync and related primitives.\n> >>\n> >> On top of my\n> >> https://lore.kernel.org/git/RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com/\n> >> which makes the tmp-objdir part of your not-in-next-just-seen follow-up\n> >> series configurable via \"fsyncMethod.batch.quarantine\" I really think we\n> >> should just go for something like the belwo patch (note that\n> >> misspelled/mistook \"bulk\" for \"batch\" in that linked-t patch, fixed\n> >> below.\n> >>\n> >> I.e. I think we should just do our default fsync() of everything, and\n> >> probably SOON make the fsync-ing of loose objects the default. Those who\n> >> care about performance will have \"batch\" (or \"writeout-only\"), which we\n> >> can have OS-specific detections for.\n> >>\n> >> But really, all of the rest of this is unduly boxing us into\n> >> overconfiguration that I think nobody really needs.\n> >>\n> >\n> > We've gone over this a few times already, but just wanted to state it\n> > again.  The really detailed settings are really there for Git hosters\n> > like GitLab or GitHub. I'd be happy to remove the per-component\n> > core.fsync values from the documentation and leave just the ones we\n> > point the user to.\n>\n> I'm prettty sure (but Patrick knows more) that GitLab's plan for this is\n> to keep it at whatever the safest setting is, presumably GitHub's as\n> well (but I don't know at all on that front).\n>\n> >> If someone really needs this level of detail they can LD_PRELOAD\n> >> something to have fsync intercept fd's and paths, and act appropriately.\n> >>\n> >> Worse, as the RFC series I sent\n> >> (https://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/)\n> >> shows we can and should \"batch\" up fsync() operations across these\n> >> configuration boundaries, which this level of configuration would seem\n> >> to preclude.\n> >>\n> >> Or, we'd need to explain why \"core.fsync=loose-object\" won't *actually*\n> >> call fsync() on a single loose object's fd under \"batch\" as I had to do\n> >> on top of this in\n> >> https://lore.kernel.org/git/RFC-patch-v2-6.7-c20301d7967-20220323T140753Z-avarab@gmail.com/\n> >>\n> >\n> > 99.9% of users don't care and won't look.  The ones who do look deeper\n> > and understand the issues have source code and access to this ML\n> > discussion to understand why this works this way.\n>\n> Exactly, so we can hopefully have a simpler interface.\n>\n> >> The same is going to apply for almost all of the rest of these\n> >> configuration categories.\n> >>\n> >> I.e. a natural follow-up to e.g. batching across objects & index as I'm\n> >> doing in\n> >> https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n> >> is to do likewise for all the PACK-related stuff before we rename it\n> >> in-place. Or even have \"git gc\" issue only a single fsync() for all of\n> >> PACKs, their metadata files, commit-graph etc., and then rename() things\n> >> in-place as appropriate afterwards.\n> >>\n> >> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> >> index 365a12dc7ae..536238e209b 100644\n> >> --- a/Documentation/config/core.txt\n> >> +++ b/Documentation/config/core.txt\n> >> @@ -548,49 +548,35 @@ core.whitespace::\n> >>    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >>\n> >>  core.fsync::\n> >> -       A comma-separated list of components of the repository that\n> >> -       should be hardened via the core.fsyncMethod when created or\n> >> -       modified.  You can disable hardening of any component by\n> >> -       prefixing it with a '-'.  Items that are not hardened may be\n> >> -       lost in the event of an unclean system shutdown. Unless you\n> >> -       have special requirements, it is recommended that you leave\n> >> -       this option empty or pick one of `committed`, `added`,\n> >> -       or `all`.\n> >> -+\n> >> -When this configuration is encountered, the set of components starts with\n> >> -the platform default value, disabled components are removed, and additional\n> >> -components are added. `none` resets the state so that the platform default\n> >> -is ignored.\n> >> -+\n> >> -The empty string resets the fsync configuration to the platform\n> >> -default. The default on most platforms is equivalent to\n> >> -`core.fsync=committed,-loose-object`, which has good performance,\n> >> -but risks losing recent work in the event of an unclean system shutdown.\n> >> -+\n> >> -* `none` clears the set of fsynced components.\n> >> -* `loose-object` hardens objects added to the repo in loose-object form.\n> >> -* `pack` hardens objects added to the repo in packfile form.\n> >> -* `pack-metadata` hardens packfile bitmaps and indexes.\n> >> -* `commit-graph` hardens the commit graph file.\n> >> -* `index` hardens the index when it is modified.\n> >> -* `objects` is an aggregate option that is equivalent to\n> >> -  `loose-object,pack`.\n> >> -* `derived-metadata` is an aggregate option that is equivalent to\n> >> -  `pack-metadata,commit-graph`.\n> >> -* `committed` is an aggregate option that is currently equivalent to\n> >> -  `objects`. This mode sacrifices some performance to ensure that work\n> >> -  that is committed to the repository with `git commit` or similar commands\n> >> -  is hardened.\n> >> -* `added` is an aggregate option that is currently equivalent to\n> >> -  `committed,index`. This mode sacrifices additional performance to\n> >> -  ensure that the results of commands like `git add` and similar operations\n> >> -  are hardened.\n> >> -* `all` is an aggregate option that syncs all individual components above.\n> >> +       A boolen defaulting to `true`. To ensure data integrity git\n> >> +       will fsync() its objects, index and refu updates etc. This can\n> >> +       be set to `false` to disable `fsync()`-ing.\n> >> ++\n> >> +Only set this to `false` if you know what you're doing, and are\n> >> +prepared to deal with data corruption. Valid use-cases include\n> >> +throwaway uses of repositories on ramdisks, one-off mass-imports\n> >> +followed by calling `sync(1)` etc.\n> >> ++\n> >> +Note that the syncing of loose objects is currently excluded from\n> >> +`core.fsync=true`. To turn on all fsync-ing you'll need\n> >> +`core.fsync=true` and `core.fsyncObjectFiles=true`, but see\n> >> +`core.fsyncMethod=batch` below for a much faster alternative that's\n> >> +just as safe on various modern OS's.\n> >> ++\n> >> +The default is in flux and may change in the future, in particular the\n> >> +equivalent of the already-deprecated `core.fsyncObjectFiles` setting\n> >> +might soon default to `true`, and `core.fsyncMethod`'s default of\n> >> +`fsync` might default to a setting deemed to be safe on the local OS,\n> >> +suc has `batch` or `writeout-only`\n> >>\n> >>  core.fsyncMethod::\n> >>         A value indicating the strategy Git will use to harden repository data\n> >>         using fsync and related primitives.\n> >>  +\n> >> +Defaults to `fsync`, but as discussed for `core.fsync` above might\n> >> +change to use one of the values below taking advantage of\n> >> +platform-specific \"faster `fsync()`\".\n> >> ++\n> >>  * `fsync` uses the fsync() system call or platform equivalents.\n> >>  * `writeout-only` issues pagecache writeback requests, but depending on the\n> >>    filesystem and storage hardware, data added to the repository may not be\n> >> @@ -680,8 +666,8 @@ backed up by any standard (e.g. POSIX), but worked in practice on some\n> >>  Linux setups.\n> >>  +\n> >>  Nowadays you should almost certainly want to use\n> >> -`core.fsync=loose-object` instead in combination with\n> >> -`core.fsyncMethod=bulk`, and possibly with\n> >> +`core.fsync=true` instead in combination with\n> >> +`core.fsyncMethod=batch`, and possibly with\n> >>  `fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n> >>  OSX, Windows) that gives you most of the performance benefit of\n> >>  `core.fsyncObjectFiles=false` with all of the safety of the old\n> >\n> > I'm at the point where I don't want to endlessly revisit this discussion.\n>\n> Sorry, my intention isn't to frustrate you, but I do think it's\n> important to get this right.\n>\n> Particularly since this is now in \"next\", and we're getting closer to a\n> release. We can either talk about this now and decide on something, or\n> it'll be in a release, and then publicly documented promises will be\n> harder to back out of.\n>\n> I think your suggestion of just hiding the relevant documentation would\n> be a good band-aid solution to that.\n>\n> But I also think that given how I was altering this in my RFC series\n> that the premise of how this could be structured has been called into\n> question in a way that we didn't (or I don't recall) us having discussed\n> before.\n>\n> I.e. that we can say \"sync loose, but not index\", or \"sync index, but\n> not loose\" with this config schema. When with \"bulk\" we it really isn't\n> any more expensive to do both if one is true (even cheaper, actually).\n>\n\nI want to make a comment about the Index here.  Syncing the index is\nstrictly required for the \"added\" level of consistency, so that we\ndon't lose stuff that leaves the work tree but is staged.  But my\nWindows enlistment has an index that's 266MB, which would be painful\nto sync even with all the optimizations.  Maybe with split-index, this\nwouldn't be so bad, but I just wanted to call out that some advanced\nusers may really care about the configurability.\n\nAs Git's various database implementations improve, the fsync stuff\nwill hopefully be more optimal and self-tuning.  But as that happens,\nGit could just start ignoring settings that lose meaning without tying\nanyones hands.\n"},{"id":"452410","messageId":"220326.86sfr4k9rm.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"CANQDOdeduc8bFA_=R-kXmkM+nb__oTxVhjBfFYj70vCFew1EyA@mail.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-26T15:31:28Z","receivedAt":"2022-03-26T15:34:44Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Mar 25 2022, Neeraj Singh wrote:\n\n> On Fri, Mar 25, 2022 at 5:33 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Fri, Mar 25 2022, Neeraj Singh wrote:\n>>\n>> > On Wed, Mar 23, 2022 at 7:46 AM Ævar Arnfjörð Bjarmason\n>> > <avarab@gmail.com> wrote:\n>> >>\n>> >>\n>> >> On Tue, Mar 15 2022, Neeraj Singh wrote:\n>> >>\n>> >> I know this is probably 80% my fault by egging you on about initially\n>> >> adding the wildmatch() based thing you didn't go for.\n>> >>\n>> >> But having looked at this with fresh eyes quite deeply I really think\n>> >> we're severely over-configuring things here:\n>> >>\n>> >> > +core.fsync::\n>> >> > +     A comma-separated list of components of the repository that\n>> >> > +     should be hardened via the core.fsyncMethod when created or\n>> >> > +     modified.  You can disable hardening of any component by\n>> >> > +     prefixing it with a '-'.  Items that are not hardened may be\n>> >> > +     lost in the event of an unclean system shutdown. Unless you\n>> >> > +     have special requirements, it is recommended that you leave\n>> >> > +     this option empty or pick one of `committed`, `added`,\n>> >> > +     or `all`.\n>> >> > ++\n>> >> > +When this configuration is encountered, the set of components starts with\n>> >> > +the platform default value, disabled components are removed, and additional\n>> >> > +components are added. `none` resets the state so that the platform default\n>> >> > +is ignored.\n>> >> > ++\n>> >> > +The empty string resets the fsync configuration to the platform\n>> >> > +default. The default on most platforms is equivalent to\n>> >> > +`core.fsync=committed,-loose-object`, which has good performance,\n>> >> > +but risks losing recent work in the event of an unclean system shutdown.\n>> >> > ++\n>> >> > +* `none` clears the set of fsynced components.\n>> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n>> >> > +* `pack` hardens objects added to the repo in packfile form.\n>> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n>> >> > +* `commit-graph` hardens the commit graph file.\n>> >> > +* `index` hardens the index when it is modified.\n>> >> > +* `objects` is an aggregate option that is equivalent to\n>> >> > +  `loose-object,pack`.\n>> >> > +* `derived-metadata` is an aggregate option that is equivalent to\n>> >> > +  `pack-metadata,commit-graph`.\n>> >> > +* `committed` is an aggregate option that is currently equivalent to\n>> >> > +  `objects`. This mode sacrifices some performance to ensure that work\n>> >> > +  that is committed to the repository with `git commit` or similar commands\n>> >> > +  is hardened.\n>> >> > +* `added` is an aggregate option that is currently equivalent to\n>> >> > +  `committed,index`. This mode sacrifices additional performance to\n>> >> > +  ensure that the results of commands like `git add` and similar operations\n>> >> > +  are hardened.\n>> >> > +* `all` is an aggregate option that syncs all individual components above.\n>> >> > +\n>> >> >  core.fsyncMethod::\n>> >> >       A value indicating the strategy Git will use to harden repository data\n>> >> >       using fsync and related primitives.\n>> >>\n>> >> On top of my\n>> >> https://lore.kernel.org/git/RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com/\n>> >> which makes the tmp-objdir part of your not-in-next-just-seen follow-up\n>> >> series configurable via \"fsyncMethod.batch.quarantine\" I really think we\n>> >> should just go for something like the belwo patch (note that\n>> >> misspelled/mistook \"bulk\" for \"batch\" in that linked-t patch, fixed\n>> >> below.\n>> >>\n>> >> I.e. I think we should just do our default fsync() of everything, and\n>> >> probably SOON make the fsync-ing of loose objects the default. Those who\n>> >> care about performance will have \"batch\" (or \"writeout-only\"), which we\n>> >> can have OS-specific detections for.\n>> >>\n>> >> But really, all of the rest of this is unduly boxing us into\n>> >> overconfiguration that I think nobody really needs.\n>> >>\n>> >\n>> > We've gone over this a few times already, but just wanted to state it\n>> > again.  The really detailed settings are really there for Git hosters\n>> > like GitLab or GitHub. I'd be happy to remove the per-component\n>> > core.fsync values from the documentation and leave just the ones we\n>> > point the user to.\n>>\n>> I'm prettty sure (but Patrick knows more) that GitLab's plan for this is\n>> to keep it at whatever the safest setting is, presumably GitHub's as\n>> well (but I don't know at all on that front).\n>>\n>> >> If someone really needs this level of detail they can LD_PRELOAD\n>> >> something to have fsync intercept fd's and paths, and act appropriately.\n>> >>\n>> >> Worse, as the RFC series I sent\n>> >> (https://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/)\n>> >> shows we can and should \"batch\" up fsync() operations across these\n>> >> configuration boundaries, which this level of configuration would seem\n>> >> to preclude.\n>> >>\n>> >> Or, we'd need to explain why \"core.fsync=loose-object\" won't *actually*\n>> >> call fsync() on a single loose object's fd under \"batch\" as I had to do\n>> >> on top of this in\n>> >> https://lore.kernel.org/git/RFC-patch-v2-6.7-c20301d7967-20220323T140753Z-avarab@gmail.com/\n>> >>\n>> >\n>> > 99.9% of users don't care and won't look.  The ones who do look deeper\n>> > and understand the issues have source code and access to this ML\n>> > discussion to understand why this works this way.\n>>\n>> Exactly, so we can hopefully have a simpler interface.\n>>\n>> >> The same is going to apply for almost all of the rest of these\n>> >> configuration categories.\n>> >>\n>> >> I.e. a natural follow-up to e.g. batching across objects & index as I'm\n>> >> doing in\n>> >> https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n>> >> is to do likewise for all the PACK-related stuff before we rename it\n>> >> in-place. Or even have \"git gc\" issue only a single fsync() for all of\n>> >> PACKs, their metadata files, commit-graph etc., and then rename() things\n>> >> in-place as appropriate afterwards.\n>> >>\n>> >> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n>> >> index 365a12dc7ae..536238e209b 100644\n>> >> --- a/Documentation/config/core.txt\n>> >> +++ b/Documentation/config/core.txt\n>> >> @@ -548,49 +548,35 @@ core.whitespace::\n>> >>    errors. The default tab width is 8. Allowed values are 1 to 63.\n>> >>\n>> >>  core.fsync::\n>> >> -       A comma-separated list of components of the repository that\n>> >> -       should be hardened via the core.fsyncMethod when created or\n>> >> -       modified.  You can disable hardening of any component by\n>> >> -       prefixing it with a '-'.  Items that are not hardened may be\n>> >> -       lost in the event of an unclean system shutdown. Unless you\n>> >> -       have special requirements, it is recommended that you leave\n>> >> -       this option empty or pick one of `committed`, `added`,\n>> >> -       or `all`.\n>> >> -+\n>> >> -When this configuration is encountered, the set of components starts with\n>> >> -the platform default value, disabled components are removed, and additional\n>> >> -components are added. `none` resets the state so that the platform default\n>> >> -is ignored.\n>> >> -+\n>> >> -The empty string resets the fsync configuration to the platform\n>> >> -default. The default on most platforms is equivalent to\n>> >> -`core.fsync=committed,-loose-object`, which has good performance,\n>> >> -but risks losing recent work in the event of an unclean system shutdown.\n>> >> -+\n>> >> -* `none` clears the set of fsynced components.\n>> >> -* `loose-object` hardens objects added to the repo in loose-object form.\n>> >> -* `pack` hardens objects added to the repo in packfile form.\n>> >> -* `pack-metadata` hardens packfile bitmaps and indexes.\n>> >> -* `commit-graph` hardens the commit graph file.\n>> >> -* `index` hardens the index when it is modified.\n>> >> -* `objects` is an aggregate option that is equivalent to\n>> >> -  `loose-object,pack`.\n>> >> -* `derived-metadata` is an aggregate option that is equivalent to\n>> >> -  `pack-metadata,commit-graph`.\n>> >> -* `committed` is an aggregate option that is currently equivalent to\n>> >> -  `objects`. This mode sacrifices some performance to ensure that work\n>> >> -  that is committed to the repository with `git commit` or similar commands\n>> >> -  is hardened.\n>> >> -* `added` is an aggregate option that is currently equivalent to\n>> >> -  `committed,index`. This mode sacrifices additional performance to\n>> >> -  ensure that the results of commands like `git add` and similar operations\n>> >> -  are hardened.\n>> >> -* `all` is an aggregate option that syncs all individual components above.\n>> >> +       A boolen defaulting to `true`. To ensure data integrity git\n>> >> +       will fsync() its objects, index and refu updates etc. This can\n>> >> +       be set to `false` to disable `fsync()`-ing.\n>> >> ++\n>> >> +Only set this to `false` if you know what you're doing, and are\n>> >> +prepared to deal with data corruption. Valid use-cases include\n>> >> +throwaway uses of repositories on ramdisks, one-off mass-imports\n>> >> +followed by calling `sync(1)` etc.\n>> >> ++\n>> >> +Note that the syncing of loose objects is currently excluded from\n>> >> +`core.fsync=true`. To turn on all fsync-ing you'll need\n>> >> +`core.fsync=true` and `core.fsyncObjectFiles=true`, but see\n>> >> +`core.fsyncMethod=batch` below for a much faster alternative that's\n>> >> +just as safe on various modern OS's.\n>> >> ++\n>> >> +The default is in flux and may change in the future, in particular the\n>> >> +equivalent of the already-deprecated `core.fsyncObjectFiles` setting\n>> >> +might soon default to `true`, and `core.fsyncMethod`'s default of\n>> >> +`fsync` might default to a setting deemed to be safe on the local OS,\n>> >> +suc has `batch` or `writeout-only`\n>> >>\n>> >>  core.fsyncMethod::\n>> >>         A value indicating the strategy Git will use to harden repository data\n>> >>         using fsync and related primitives.\n>> >>  +\n>> >> +Defaults to `fsync`, but as discussed for `core.fsync` above might\n>> >> +change to use one of the values below taking advantage of\n>> >> +platform-specific \"faster `fsync()`\".\n>> >> ++\n>> >>  * `fsync` uses the fsync() system call or platform equivalents.\n>> >>  * `writeout-only` issues pagecache writeback requests, but depending on the\n>> >>    filesystem and storage hardware, data added to the repository may not be\n>> >> @@ -680,8 +666,8 @@ backed up by any standard (e.g. POSIX), but worked in practice on some\n>> >>  Linux setups.\n>> >>  +\n>> >>  Nowadays you should almost certainly want to use\n>> >> -`core.fsync=loose-object` instead in combination with\n>> >> -`core.fsyncMethod=bulk`, and possibly with\n>> >> +`core.fsync=true` instead in combination with\n>> >> +`core.fsyncMethod=batch`, and possibly with\n>> >>  `fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n>> >>  OSX, Windows) that gives you most of the performance benefit of\n>> >>  `core.fsyncObjectFiles=false` with all of the safety of the old\n>> >\n>> > I'm at the point where I don't want to endlessly revisit this discussion.\n>>\n>> Sorry, my intention isn't to frustrate you, but I do think it's\n>> important to get this right.\n>>\n>> Particularly since this is now in \"next\", and we're getting closer to a\n>> release. We can either talk about this now and decide on something, or\n>> it'll be in a release, and then publicly documented promises will be\n>> harder to back out of.\n>>\n>> I think your suggestion of just hiding the relevant documentation would\n>> be a good band-aid solution to that.\n>>\n>> But I also think that given how I was altering this in my RFC series\n>> that the premise of how this could be structured has been called into\n>> question in a way that we didn't (or I don't recall) us having discussed\n>> before.\n>>\n>> I.e. that we can say \"sync loose, but not index\", or \"sync index, but\n>> not loose\" with this config schema. When with \"bulk\" we it really isn't\n>> any more expensive to do both if one is true (even cheaper, actually).\n>>\n>\n> I want to make a comment about the Index here.  Syncing the index is\n> strictly required for the \"added\" level of consistency, so that we\n> don't lose stuff that leaves the work tree but is staged.  But my\n> Windows enlistment has an index that's 266MB, which would be painful\n> to sync even with all the optimizations.  Maybe with split-index, this\n> wouldn't be so bad, but I just wanted to call out that some advanced\n> users may really care about the configurability.\n\nSo for that use-case you'd like to fsync the loose objects (if any), but\nnot the index? So the FS will \"flush\" up to the index, and then queue\nthe index for later syncing to platter?\n\n\nBut even in that case don't the settings need to be tied to one another,\nbecause in the method=bulk sync=index && sync=!loose case wouldn't we be\nsyncing \"loose\" in any case?\n\n> As Git's various database implementations improve, the fsync stuff\n> will hopefully be more optimal and self-tuning.  But as that happens,\n> Git could just start ignoring settings that lose meaning without tying\n> anyones hands.\n\nYeah that would alleviate most of my concerns here, but the docs aren't\nsaying anything like that. Since you added them & they just landed, do\nyou mind doing a small follow-up where we e.g. say that these new\nsettings are \"EXPERIMENTAL\" or whatever, and subject to drastic change?\n"},{"id":"452425","messageId":"CANQDOdfWh5aO9cuJVuUccKyD9Cj+NndisokiewBH9Sq4oSUp5A@mail.gmail.com","threadId":"57030","inReplyTo":"220326.86sfr4k9rm.gmgdl@evledraar.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-27T05:27:52Z","receivedAt":"2022-03-27T05:28:15Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Sat, Mar 26, 2022 at 8:34 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Fri, Mar 25 2022, Neeraj Singh wrote:\n>\n> > On Fri, Mar 25, 2022 at 5:33 PM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Fri, Mar 25 2022, Neeraj Singh wrote:\n> >>\n> >> > On Wed, Mar 23, 2022 at 7:46 AM Ævar Arnfjörð Bjarmason\n> >> > <avarab@gmail.com> wrote:\n> >> >>\n> >> >>\n> >> >> On Tue, Mar 15 2022, Neeraj Singh wrote:\n> >> >>\n> >> >> I know this is probably 80% my fault by egging you on about initially\n> >> >> adding the wildmatch() based thing you didn't go for.\n> >> >>\n> >> >> But having looked at this with fresh eyes quite deeply I really think\n> >> >> we're severely over-configuring things here:\n> >> >>\n> >> >> > +core.fsync::\n> >> >> > +     A comma-separated list of components of the repository that\n> >> >> > +     should be hardened via the core.fsyncMethod when created or\n> >> >> > +     modified.  You can disable hardening of any component by\n> >> >> > +     prefixing it with a '-'.  Items that are not hardened may be\n> >> >> > +     lost in the event of an unclean system shutdown. Unless you\n> >> >> > +     have special requirements, it is recommended that you leave\n> >> >> > +     this option empty or pick one of `committed`, `added`,\n> >> >> > +     or `all`.\n> >> >> > ++\n> >> >> > +When this configuration is encountered, the set of components starts with\n> >> >> > +the platform default value, disabled components are removed, and additional\n> >> >> > +components are added. `none` resets the state so that the platform default\n> >> >> > +is ignored.\n> >> >> > ++\n> >> >> > +The empty string resets the fsync configuration to the platform\n> >> >> > +default. The default on most platforms is equivalent to\n> >> >> > +`core.fsync=committed,-loose-object`, which has good performance,\n> >> >> > +but risks losing recent work in the event of an unclean system shutdown.\n> >> >> > ++\n> >> >> > +* `none` clears the set of fsynced components.\n> >> >> > +* `loose-object` hardens objects added to the repo in loose-object form.\n> >> >> > +* `pack` hardens objects added to the repo in packfile form.\n> >> >> > +* `pack-metadata` hardens packfile bitmaps and indexes.\n> >> >> > +* `commit-graph` hardens the commit graph file.\n> >> >> > +* `index` hardens the index when it is modified.\n> >> >> > +* `objects` is an aggregate option that is equivalent to\n> >> >> > +  `loose-object,pack`.\n> >> >> > +* `derived-metadata` is an aggregate option that is equivalent to\n> >> >> > +  `pack-metadata,commit-graph`.\n> >> >> > +* `committed` is an aggregate option that is currently equivalent to\n> >> >> > +  `objects`. This mode sacrifices some performance to ensure that work\n> >> >> > +  that is committed to the repository with `git commit` or similar commands\n> >> >> > +  is hardened.\n> >> >> > +* `added` is an aggregate option that is currently equivalent to\n> >> >> > +  `committed,index`. This mode sacrifices additional performance to\n> >> >> > +  ensure that the results of commands like `git add` and similar operations\n> >> >> > +  are hardened.\n> >> >> > +* `all` is an aggregate option that syncs all individual components above.\n> >> >> > +\n> >> >> >  core.fsyncMethod::\n> >> >> >       A value indicating the strategy Git will use to harden repository data\n> >> >> >       using fsync and related primitives.\n> >> >>\n> >> >> On top of my\n> >> >> https://lore.kernel.org/git/RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com/\n> >> >> which makes the tmp-objdir part of your not-in-next-just-seen follow-up\n> >> >> series configurable via \"fsyncMethod.batch.quarantine\" I really think we\n> >> >> should just go for something like the belwo patch (note that\n> >> >> misspelled/mistook \"bulk\" for \"batch\" in that linked-t patch, fixed\n> >> >> below.\n> >> >>\n> >> >> I.e. I think we should just do our default fsync() of everything, and\n> >> >> probably SOON make the fsync-ing of loose objects the default. Those who\n> >> >> care about performance will have \"batch\" (or \"writeout-only\"), which we\n> >> >> can have OS-specific detections for.\n> >> >>\n> >> >> But really, all of the rest of this is unduly boxing us into\n> >> >> overconfiguration that I think nobody really needs.\n> >> >>\n> >> >\n> >> > We've gone over this a few times already, but just wanted to state it\n> >> > again.  The really detailed settings are really there for Git hosters\n> >> > like GitLab or GitHub. I'd be happy to remove the per-component\n> >> > core.fsync values from the documentation and leave just the ones we\n> >> > point the user to.\n> >>\n> >> I'm prettty sure (but Patrick knows more) that GitLab's plan for this is\n> >> to keep it at whatever the safest setting is, presumably GitHub's as\n> >> well (but I don't know at all on that front).\n> >>\n> >> >> If someone really needs this level of detail they can LD_PRELOAD\n> >> >> something to have fsync intercept fd's and paths, and act appropriately.\n> >> >>\n> >> >> Worse, as the RFC series I sent\n> >> >> (https://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/)\n> >> >> shows we can and should \"batch\" up fsync() operations across these\n> >> >> configuration boundaries, which this level of configuration would seem\n> >> >> to preclude.\n> >> >>\n> >> >> Or, we'd need to explain why \"core.fsync=loose-object\" won't *actually*\n> >> >> call fsync() on a single loose object's fd under \"batch\" as I had to do\n> >> >> on top of this in\n> >> >> https://lore.kernel.org/git/RFC-patch-v2-6.7-c20301d7967-20220323T140753Z-avarab@gmail.com/\n> >> >>\n> >> >\n> >> > 99.9% of users don't care and won't look.  The ones who do look deeper\n> >> > and understand the issues have source code and access to this ML\n> >> > discussion to understand why this works this way.\n> >>\n> >> Exactly, so we can hopefully have a simpler interface.\n> >>\n> >> >> The same is going to apply for almost all of the rest of these\n> >> >> configuration categories.\n> >> >>\n> >> >> I.e. a natural follow-up to e.g. batching across objects & index as I'm\n> >> >> doing in\n> >> >> https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n> >> >> is to do likewise for all the PACK-related stuff before we rename it\n> >> >> in-place. Or even have \"git gc\" issue only a single fsync() for all of\n> >> >> PACKs, their metadata files, commit-graph etc., and then rename() things\n> >> >> in-place as appropriate afterwards.\n> >> >>\n> >> >> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> >> >> index 365a12dc7ae..536238e209b 100644\n> >> >> --- a/Documentation/config/core.txt\n> >> >> +++ b/Documentation/config/core.txt\n> >> >> @@ -548,49 +548,35 @@ core.whitespace::\n> >> >>    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >> >>\n> >> >>  core.fsync::\n> >> >> -       A comma-separated list of components of the repository that\n> >> >> -       should be hardened via the core.fsyncMethod when created or\n> >> >> -       modified.  You can disable hardening of any component by\n> >> >> -       prefixing it with a '-'.  Items that are not hardened may be\n> >> >> -       lost in the event of an unclean system shutdown. Unless you\n> >> >> -       have special requirements, it is recommended that you leave\n> >> >> -       this option empty or pick one of `committed`, `added`,\n> >> >> -       or `all`.\n> >> >> -+\n> >> >> -When this configuration is encountered, the set of components starts with\n> >> >> -the platform default value, disabled components are removed, and additional\n> >> >> -components are added. `none` resets the state so that the platform default\n> >> >> -is ignored.\n> >> >> -+\n> >> >> -The empty string resets the fsync configuration to the platform\n> >> >> -default. The default on most platforms is equivalent to\n> >> >> -`core.fsync=committed,-loose-object`, which has good performance,\n> >> >> -but risks losing recent work in the event of an unclean system shutdown.\n> >> >> -+\n> >> >> -* `none` clears the set of fsynced components.\n> >> >> -* `loose-object` hardens objects added to the repo in loose-object form.\n> >> >> -* `pack` hardens objects added to the repo in packfile form.\n> >> >> -* `pack-metadata` hardens packfile bitmaps and indexes.\n> >> >> -* `commit-graph` hardens the commit graph file.\n> >> >> -* `index` hardens the index when it is modified.\n> >> >> -* `objects` is an aggregate option that is equivalent to\n> >> >> -  `loose-object,pack`.\n> >> >> -* `derived-metadata` is an aggregate option that is equivalent to\n> >> >> -  `pack-metadata,commit-graph`.\n> >> >> -* `committed` is an aggregate option that is currently equivalent to\n> >> >> -  `objects`. This mode sacrifices some performance to ensure that work\n> >> >> -  that is committed to the repository with `git commit` or similar commands\n> >> >> -  is hardened.\n> >> >> -* `added` is an aggregate option that is currently equivalent to\n> >> >> -  `committed,index`. This mode sacrifices additional performance to\n> >> >> -  ensure that the results of commands like `git add` and similar operations\n> >> >> -  are hardened.\n> >> >> -* `all` is an aggregate option that syncs all individual components above.\n> >> >> +       A boolen defaulting to `true`. To ensure data integrity git\n> >> >> +       will fsync() its objects, index and refu updates etc. This can\n> >> >> +       be set to `false` to disable `fsync()`-ing.\n> >> >> ++\n> >> >> +Only set this to `false` if you know what you're doing, and are\n> >> >> +prepared to deal with data corruption. Valid use-cases include\n> >> >> +throwaway uses of repositories on ramdisks, one-off mass-imports\n> >> >> +followed by calling `sync(1)` etc.\n> >> >> ++\n> >> >> +Note that the syncing of loose objects is currently excluded from\n> >> >> +`core.fsync=true`. To turn on all fsync-ing you'll need\n> >> >> +`core.fsync=true` and `core.fsyncObjectFiles=true`, but see\n> >> >> +`core.fsyncMethod=batch` below for a much faster alternative that's\n> >> >> +just as safe on various modern OS's.\n> >> >> ++\n> >> >> +The default is in flux and may change in the future, in particular the\n> >> >> +equivalent of the already-deprecated `core.fsyncObjectFiles` setting\n> >> >> +might soon default to `true`, and `core.fsyncMethod`'s default of\n> >> >> +`fsync` might default to a setting deemed to be safe on the local OS,\n> >> >> +suc has `batch` or `writeout-only`\n> >> >>\n> >> >>  core.fsyncMethod::\n> >> >>         A value indicating the strategy Git will use to harden repository data\n> >> >>         using fsync and related primitives.\n> >> >>  +\n> >> >> +Defaults to `fsync`, but as discussed for `core.fsync` above might\n> >> >> +change to use one of the values below taking advantage of\n> >> >> +platform-specific \"faster `fsync()`\".\n> >> >> ++\n> >> >>  * `fsync` uses the fsync() system call or platform equivalents.\n> >> >>  * `writeout-only` issues pagecache writeback requests, but depending on the\n> >> >>    filesystem and storage hardware, data added to the repository may not be\n> >> >> @@ -680,8 +666,8 @@ backed up by any standard (e.g. POSIX), but worked in practice on some\n> >> >>  Linux setups.\n> >> >>  +\n> >> >>  Nowadays you should almost certainly want to use\n> >> >> -`core.fsync=loose-object` instead in combination with\n> >> >> -`core.fsyncMethod=bulk`, and possibly with\n> >> >> +`core.fsync=true` instead in combination with\n> >> >> +`core.fsyncMethod=batch`, and possibly with\n> >> >>  `fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n> >> >>  OSX, Windows) that gives you most of the performance benefit of\n> >> >>  `core.fsyncObjectFiles=false` with all of the safety of the old\n> >> >\n> >> > I'm at the point where I don't want to endlessly revisit this discussion.\n> >>\n> >> Sorry, my intention isn't to frustrate you, but I do think it's\n> >> important to get this right.\n> >>\n> >> Particularly since this is now in \"next\", and we're getting closer to a\n> >> release. We can either talk about this now and decide on something, or\n> >> it'll be in a release, and then publicly documented promises will be\n> >> harder to back out of.\n> >>\n> >> I think your suggestion of just hiding the relevant documentation would\n> >> be a good band-aid solution to that.\n> >>\n> >> But I also think that given how I was altering this in my RFC series\n> >> that the premise of how this could be structured has been called into\n> >> question in a way that we didn't (or I don't recall) us having discussed\n> >> before.\n> >>\n> >> I.e. that we can say \"sync loose, but not index\", or \"sync index, but\n> >> not loose\" with this config schema. When with \"bulk\" we it really isn't\n> >> any more expensive to do both if one is true (even cheaper, actually).\n> >>\n> >\n> > I want to make a comment about the Index here.  Syncing the index is\n> > strictly required for the \"added\" level of consistency, so that we\n> > don't lose stuff that leaves the work tree but is staged.  But my\n> > Windows enlistment has an index that's 266MB, which would be painful\n> > to sync even with all the optimizations.  Maybe with split-index, this\n> > wouldn't be so bad, but I just wanted to call out that some advanced\n> > users may really care about the configurability.\n>\n> So for that use-case you'd like to fsync the loose objects (if any), but\n> not the index? So the FS will \"flush\" up to the index, and then queue\n> the index for later syncing to platter?\n>\n>\n> But even in that case don't the settings need to be tied to one another,\n> because in the method=bulk sync=index && sync=!loose case wouldn't we be\n> syncing \"loose\" in any case?\n>\n> > As Git's various database implementations improve, the fsync stuff\n> > will hopefully be more optimal and self-tuning.  But as that happens,\n> > Git could just start ignoring settings that lose meaning without tying\n> > anyones hands.\n>\n> Yeah that would alleviate most of my concerns here, but the docs aren't\n> saying anything like that. Since you added them & they just landed, do\n> you mind doing a small follow-up where we e.g. say that these new\n> settings are \"EXPERIMENTAL\" or whatever, and subject to drastic change?\n\nThe doc is already pretty prescriptive.  It has this line at the end\nof the first  paragraph:\n\"Unless you\nhave special requirements, it is recommended that you leave\nthis option empty or pick one of `committed`, `added`,\nor `all`.\"\n\nThose values are already designed to change as Git changes.\n"},{"id":"452434","messageId":"220327.86y20veeua.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"CANQDOdfWh5aO9cuJVuUccKyD9Cj+NndisokiewBH9Sq4oSUp5A@mail.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-27T12:43:48Z","receivedAt":"2022-03-27T12:53:56Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Mar 26 2022, Neeraj Singh wrote:\n\n> On Sat, Mar 26, 2022 at 8:34 AM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Fri, Mar 25 2022, Neeraj Singh wrote:\n>>\n>> > On Fri, Mar 25, 2022 at 5:33 PM Ævar Arnfjörð Bjarmason\n>> > <avarab@gmail.com> wrote:\n[...]\n>> > I want to make a comment about the Index here.  Syncing the index is\n>> > strictly required for the \"added\" level of consistency, so that we\n>> > don't lose stuff that leaves the work tree but is staged.  But my\n>> > Windows enlistment has an index that's 266MB, which would be painful\n>> > to sync even with all the optimizations.  Maybe with split-index, this\n>> > wouldn't be so bad, but I just wanted to call out that some advanced\n>> > users may really care about the configurability.\n>>\n>> So for that use-case you'd like to fsync the loose objects (if any), but\n>> not the index? So the FS will \"flush\" up to the index, and then queue\n>> the index for later syncing to platter?\n>>\n>>\n>> But even in that case don't the settings need to be tied to one another,\n>> because in the method=bulk sync=index && sync=!loose case wouldn't we be\n>> syncing \"loose\" in any case?\n>>\n>> > As Git's various database implementations improve, the fsync stuff\n>> > will hopefully be more optimal and self-tuning.  But as that happens,\n>> > Git could just start ignoring settings that lose meaning without tying\n>> > anyones hands.\n>>\n>> Yeah that would alleviate most of my concerns here, but the docs aren't\n>> saying anything like that. Since you added them & they just landed, do\n>> you mind doing a small follow-up where we e.g. say that these new\n>> settings are \"EXPERIMENTAL\" or whatever, and subject to drastic change?\n>\n> The doc is already pretty prescriptive.  It has this line at the end\n> of the first  paragraph:\n> \"Unless you\n> have special requirements, it is recommended that you leave\n> this option empty or pick one of `committed`, `added`,\n> or `all`.\"\n>\n> Those values are already designed to change as Git changes.\n\nI'm referring to the documentation as it stands not being marked as\nexperimental in the sense that we might decide to re-do this to a large\nextent, i.e. something like the diff I suggested upthread in\nhttps://lore.kernel.org/git/220323.86fsn8ohg8.gmgdl@evledraar.gmail.com/\n\nSo yes, I agree that it e.g. clearly states that you can add a new\ncore.git=foobar or whatever down the line, but it clearly doesn't\nsuggest that e.g. core.fsync might have boolean semantics in some later\nversion, or that the rest might simply be ignored, even if that\ne.g. means that we wouldn't sync loose objects on\ncore.fsync=loose-object, as we'd just warn with a \"we don't provide this\nanymore\".\n\nOr do you disagree with that? IOW I mean that we'd do something like\nthis, either in docs or code:\n\ndiff --git a/config.c b/config.c\nindex 3c9b6b589ab..94548566073 100644\n--- a/config.c\n+++ b/config.c\n@@ -1675,6 +1675,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsync\")) {\n+\t\tif (!the_repository->settings.feature_experimental)\n+\t\t\twarning(_(\"the '%s' configuration option is EXPERIMENTAL. opt-in to use it with feature.experimental=true\"),\n+\t\t\t\tvar);\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n \t\tfsync_components = parse_fsync_components(var, value);\n@@ -1682,6 +1685,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncmethod\")) {\n+\t\tif (!the_repository->settings.feature_experimental)\n+\t\t\twarning(_(\"the '%s' configuration option is EXPERIMENTAL. opt-in to use it with feature.experimental=true\"),\n+\t\t\t\tvar);\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n \t\tif (!strcmp(value, \"fsync\"))\ndiff --git a/repo-settings.c b/repo-settings.c\nindex b4fbd16cdcc..f949b65b91e 100644\n--- a/repo-settings.c\n+++ b/repo-settings.c\n@@ -31,6 +31,7 @@ void prepare_repo_settings(struct repository *r)\n \t/* Booleans config or default, cascades to other settings */\n \trepo_cfg_bool(r, \"feature.manyfiles\", &manyfiles, 0);\n \trepo_cfg_bool(r, \"feature.experimental\", &experimental, 0);\n+\tr->settings.feature_experimental = experimental;\n \n \t/* Defaults modified by feature.* */\n \tif (experimental) {\ndiff --git a/repository.h b/repository.h\nindex e29f361703d..db8f99a8989 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -28,6 +28,7 @@ enum fetch_negotiation_setting {\n struct repo_settings {\n \tint initialized;\n \n+\tint feature_experimental;\n \tint core_commit_graph;\n \tint commit_graph_read_changed_paths;\n \tint gc_write_commit_graph;\n"},{"id":"452461","messageId":"YkGUeQH4y1KIAdCc@ncase","threadId":"57030","inReplyTo":"220327.86y20veeua.gmgdl@evledraar.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-28T10:56:57Z","receivedAt":"2022-03-28T10:57:10Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Mar 27, 2022 at 02:43:48PM +0200, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Sat, Mar 26 2022, Neeraj Singh wrote:\n> \n> > On Sat, Mar 26, 2022 at 8:34 AM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Fri, Mar 25 2022, Neeraj Singh wrote:\n> >>\n> >> > On Fri, Mar 25, 2022 at 5:33 PM Ævar Arnfjörð Bjarmason\n> >> > <avarab@gmail.com> wrote:\n> [...]\n> >> > I want to make a comment about the Index here.  Syncing the index is\n> >> > strictly required for the \"added\" level of consistency, so that we\n> >> > don't lose stuff that leaves the work tree but is staged.  But my\n> >> > Windows enlistment has an index that's 266MB, which would be painful\n> >> > to sync even with all the optimizations.  Maybe with split-index, this\n> >> > wouldn't be so bad, but I just wanted to call out that some advanced\n> >> > users may really care about the configurability.\n> >>\n> >> So for that use-case you'd like to fsync the loose objects (if any), but\n> >> not the index? So the FS will \"flush\" up to the index, and then queue\n> >> the index for later syncing to platter?\n> >>\n> >>\n> >> But even in that case don't the settings need to be tied to one another,\n> >> because in the method=bulk sync=index && sync=!loose case wouldn't we be\n> >> syncing \"loose\" in any case?\n> >>\n> >> > As Git's various database implementations improve, the fsync stuff\n> >> > will hopefully be more optimal and self-tuning.  But as that happens,\n> >> > Git could just start ignoring settings that lose meaning without tying\n> >> > anyones hands.\n> >>\n> >> Yeah that would alleviate most of my concerns here, but the docs aren't\n> >> saying anything like that. Since you added them & they just landed, do\n> >> you mind doing a small follow-up where we e.g. say that these new\n> >> settings are \"EXPERIMENTAL\" or whatever, and subject to drastic change?\n> >\n> > The doc is already pretty prescriptive.  It has this line at the end\n> > of the first  paragraph:\n> > \"Unless you\n> > have special requirements, it is recommended that you leave\n> > this option empty or pick one of `committed`, `added`,\n> > or `all`.\"\n> >\n> > Those values are already designed to change as Git changes.\n> \n> I'm referring to the documentation as it stands not being marked as\n> experimental in the sense that we might decide to re-do this to a large\n> extent, i.e. something like the diff I suggested upthread in\n> https://lore.kernel.org/git/220323.86fsn8ohg8.gmgdl@evledraar.gmail.com/\n> \n> So yes, I agree that it e.g. clearly states that you can add a new\n> core.git=foobar or whatever down the line, but it clearly doesn't\n> suggest that e.g. core.fsync might have boolean semantics in some later\n> version, or that the rest might simply be ignored, even if that\n> e.g. means that we wouldn't sync loose objects on\n> core.fsync=loose-object, as we'd just warn with a \"we don't provide this\n> anymore\".\n> \n> Or do you disagree with that? IOW I mean that we'd do something like\n> this, either in docs or code:\n> \n> diff --git a/config.c b/config.c\n> index 3c9b6b589ab..94548566073 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1675,6 +1675,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t}\n>  \n>  \tif (!strcmp(var, \"core.fsync\")) {\n> +\t\tif (!the_repository->settings.feature_experimental)\n> +\t\t\twarning(_(\"the '%s' configuration option is EXPERIMENTAL. opt-in to use it with feature.experimental=true\"),\n> +\t\t\t\tvar);\n>  \t\tif (!value)\n>  \t\t\treturn config_error_nonbool(var);\n>  \t\tfsync_components = parse_fsync_components(var, value);\n> @@ -1682,6 +1685,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t}\n>  \n>  \tif (!strcmp(var, \"core.fsyncmethod\")) {\n> +\t\tif (!the_repository->settings.feature_experimental)\n> +\t\t\twarning(_(\"the '%s' configuration option is EXPERIMENTAL. opt-in to use it with feature.experimental=true\"),\n> +\t\t\t\tvar);\n>  \t\tif (!value)\n>  \t\t\treturn config_error_nonbool(var);\n>  \t\tif (!strcmp(value, \"fsync\"))\n\nLet's please not tie this to `feature.experimental=true`. Setting that\noption has unintended sideeffects and will also change defaults which we\nmay not want to have in production. I don't mind adding a warning in the\ndocs though that the specific items which can be configured may be\nsubject to change in the future.\n\nAt GitLab, we've got a three-step plan:\n\n    1. We need to migrate to `core.fsync` in the first place. In order\n       to not migrate and change behaviour at the same point in time we\n       already benefit from the fine-grainedness of this config because\n       we can simply say `core.fsync=loose-objects` and have the same\n       behaviour as before with `core.fsyncLooseObjects=true`.\n\n    2. We'll want to enable syncing of packfiles, which I think wasn't\n       previously covered by `core.fsyncLooseobjects`.\n\n    3. We'll add `refs` to also sync loose refs to disk.\n\nSo while the end result will be the same as `committed`, having this\nlevel of control helps us to assess the impact in a nicer way by being\nable to do this step by step with feature flags.\n\nOn the other hand, many of the other parts we don't really care about.\nAuxiliary metadata like the commit-graph or pack indices are data that\ncan in the worst case be regenerated by us, so it's not clear to me\nwhether it makes to also enable fsyncing those in production.\n\nSo altogether, I agree with Neeraj: having the fine-grainedness greatly\nhelps us to roll out changes like this and be able to pick what we deem\nto be important. Personally I would be fine with explicitly pointing out\nthat there are two groups of this config in our docs though:\n\n    1. The \"porcelain\" group: \"committed\", \"added\", \"all\", \"none\". These\n       are abstract groups whose behaviour should adapt as we change\n       implementations, and are those that should typically be set by a\n       user, if intended.\n\n    2. The \"plumbing\" or \"expert\" group: these are fine-grained options\n       which shouldn't typically be used by Git users. They still have\n       merit though in hosting environments, where requirements are\n       typically a lot more specific.\n\nWe may also provide different guarantees for both groups. The first one\nshould definitely be stable, but we might state that the second group is\nsubject to change in the future.\n\nPatrick\n"},{"id":"452462","messageId":"CANYiYbFjRMV-_opvFn78mq7tgtZFMrfPyDjDa+kyaZZfk_LmWQ@mail.gmail.com","threadId":"57030","inReplyTo":"6adc8dc13852c219763a9830f848fbc8663f2fa9.1646952205.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 4/6] core.fsync: add configuration parsing","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2022-03-28T11:06:54Z","receivedAt":"2022-03-28T11:07:11Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Sat, Mar 12, 2022 at 6:25 AM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n> @@ -1613,6 +1687,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>         }\n>\n>         if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> +               if (fsync_object_files < 0)\n> +                       warning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n\ns/core.fsyncobjectfiles/core.fsyncObjectFiles/  to use bumpyCaps for\nconfig variable in documentation.\n"},{"id":"452464","messageId":"220328.86pmm6e0i0.gmgdl@evledraar.gmail.com","threadId":"57030","inReplyTo":"YkGUeQH4y1KIAdCc@ncase","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-28T11:25:02Z","receivedAt":"2022-03-28T12:15:58Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Mar 28 2022, Patrick Steinhardt wrote:\n\n> [[PGP Signed Part:Undecided]]\n> On Sun, Mar 27, 2022 at 02:43:48PM +0200, Ævar Arnfjörð Bjarmason wrote:\n>> \n>> On Sat, Mar 26 2022, Neeraj Singh wrote:\n>> \n>> > On Sat, Mar 26, 2022 at 8:34 AM Ævar Arnfjörð Bjarmason\n>> > <avarab@gmail.com> wrote:\n>> >>\n>> >>\n>> >> On Fri, Mar 25 2022, Neeraj Singh wrote:\n>> >>\n>> >> > On Fri, Mar 25, 2022 at 5:33 PM Ævar Arnfjörð Bjarmason\n>> >> > <avarab@gmail.com> wrote:\n>> [...]\n>> >> > I want to make a comment about the Index here.  Syncing the index is\n>> >> > strictly required for the \"added\" level of consistency, so that we\n>> >> > don't lose stuff that leaves the work tree but is staged.  But my\n>> >> > Windows enlistment has an index that's 266MB, which would be painful\n>> >> > to sync even with all the optimizations.  Maybe with split-index, this\n>> >> > wouldn't be so bad, but I just wanted to call out that some advanced\n>> >> > users may really care about the configurability.\n>> >>\n>> >> So for that use-case you'd like to fsync the loose objects (if any), but\n>> >> not the index? So the FS will \"flush\" up to the index, and then queue\n>> >> the index for later syncing to platter?\n>> >>\n>> >>\n>> >> But even in that case don't the settings need to be tied to one another,\n>> >> because in the method=bulk sync=index && sync=!loose case wouldn't we be\n>> >> syncing \"loose\" in any case?\n>> >>\n>> >> > As Git's various database implementations improve, the fsync stuff\n>> >> > will hopefully be more optimal and self-tuning.  But as that happens,\n>> >> > Git could just start ignoring settings that lose meaning without tying\n>> >> > anyones hands.\n>> >>\n>> >> Yeah that would alleviate most of my concerns here, but the docs aren't\n>> >> saying anything like that. Since you added them & they just landed, do\n>> >> you mind doing a small follow-up where we e.g. say that these new\n>> >> settings are \"EXPERIMENTAL\" or whatever, and subject to drastic change?\n>> >\n>> > The doc is already pretty prescriptive.  It has this line at the end\n>> > of the first  paragraph:\n>> > \"Unless you\n>> > have special requirements, it is recommended that you leave\n>> > this option empty or pick one of `committed`, `added`,\n>> > or `all`.\"\n>> >\n>> > Those values are already designed to change as Git changes.\n>> \n>> I'm referring to the documentation as it stands not being marked as\n>> experimental in the sense that we might decide to re-do this to a large\n>> extent, i.e. something like the diff I suggested upthread in\n>> https://lore.kernel.org/git/220323.86fsn8ohg8.gmgdl@evledraar.gmail.com/\n>> \n>> So yes, I agree that it e.g. clearly states that you can add a new\n>> core.git=foobar or whatever down the line, but it clearly doesn't\n>> suggest that e.g. core.fsync might have boolean semantics in some later\n>> version, or that the rest might simply be ignored, even if that\n>> e.g. means that we wouldn't sync loose objects on\n>> core.fsync=loose-object, as we'd just warn with a \"we don't provide this\n>> anymore\".\n>> \n>> Or do you disagree with that? IOW I mean that we'd do something like\n>> this, either in docs or code:\n>> \n>> diff --git a/config.c b/config.c\n>> index 3c9b6b589ab..94548566073 100644\n>> --- a/config.c\n>> +++ b/config.c\n>> @@ -1675,6 +1675,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>>  \t}\n>>  \n>>  \tif (!strcmp(var, \"core.fsync\")) {\n>> +\t\tif (!the_repository->settings.feature_experimental)\n>> +\t\t\twarning(_(\"the '%s' configuration option is EXPERIMENTAL. opt-in to use it with feature.experimental=true\"),\n>> +\t\t\t\tvar);\n>>  \t\tif (!value)\n>>  \t\t\treturn config_error_nonbool(var);\n>>  \t\tfsync_components = parse_fsync_components(var, value);\n>> @@ -1682,6 +1685,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>>  \t}\n>>  \n>>  \tif (!strcmp(var, \"core.fsyncmethod\")) {\n>> +\t\tif (!the_repository->settings.feature_experimental)\n>> +\t\t\twarning(_(\"the '%s' configuration option is EXPERIMENTAL. opt-in to use it with feature.experimental=true\"),\n>> +\t\t\t\tvar);\n>>  \t\tif (!value)\n>>  \t\t\treturn config_error_nonbool(var);\n>>  \t\tif (!strcmp(value, \"fsync\"))\n>\n> Let's please not tie this to `feature.experimental=true`. Setting that\n> option has unintended sideeffects and will also change defaults which we\n> may not want to have in production. I don't mind adding a warning in the\n> docs though that the specific items which can be configured may be\n> subject to change in the future.\n\nYes, that was a bad (throwaway) idea. I think probably any sort of\nwarning is over-doing it, but having the same in the docs would be good,\nas in:\n\n    git gre EXPERIMENTAL -- Documentation\n\n> At GitLab, we've got a three-step plan:\n>\n>     1. We need to migrate to `core.fsync` in the first place. In order\n>        to not migrate and change behaviour at the same point in time we\n>        already benefit from the fine-grainedness of this config because\n>        we can simply say `core.fsync=loose-objects` and have the same\n>        behaviour as before with `core.fsyncLooseObjects=true`.\n\n*nod*.\n\n>     2. We'll want to enable syncing of packfiles, which I think wasn't\n>        previously covered by `core.fsyncLooseobjects`.\n\nWe've always fsynced packfiles and other things that use the\nfinalize_hashfile() API. I.e. the pack metadata (idx,midx,bitmap etc.),\ncommit-graph etc.\n\nWhich is one thing I find a bit uncomfortable about the proposed config\nschema, i.e. it's allowing *granular* unsafe behavior we didn't allow\nbefore.\n\nI think we *should* allow disabling fsync() entirely via:\n\n    core.fsync=false\n\nPer Eric Wong's [added to CC] proposal here:\nhttps://lore.kernel.org/git/20211028002102.19384-1-e@80x24.org/; I think\nthat's useful for e.g. running git in CI, one off scripted mass-imports\nwhere you run \"sync(1)\" after (or not...).\n\nBut I don't really see the use-case for turning off say \"index\" or\n\"pack-metadata\", but otherwise keeping the default fsync().\n\n>     3. We'll add `refs` to also sync loose refs to disk.\n\nMaybe I'm reading this wrong, but AFAICT if you upgrade from pre-v2.36.0\nto v2.36.0 you'll have no way to do fsync()-ing as it was done before\nwith your bc22d845c43 (core.fsync: new option to harden references,\n2022-03-11).\n\nI.e. I first thought you meant to start with:\n\n    core.fsync=-loose-objects\n    core.fsync=-reference\n\nAnd then remove that \"core.fsync=-reference\" line to get the behavior\nthat'll be new in v2.36.0, but that won't do that. The new \"reference\"\ncategory doesn't just affect loose refs, but all ref updates.\n\nSo we don't have any way to move to v2.36.0 and get exactly the fsync()\nbehavior we did before, or have I misread the code?\n\nNow, I don't think we need it to be configurable at all.\n\nI.e. I think your is a bc22d845c43 good change, but it feels weird to\nmake it and leave the default of core.fsyncLooseObjects=false on the\ntable. I.e. what you're summarizing there is true, but it's also true of\nthe loose objects.\n\nTo the extent that we've had any reason at all not to sync those by\ndefault (which has really mostly been \"we didn't re-visit it for a\nwhile\") it's been performance.\n\nAnd both loose refs & loose objects will suffer from the same\ndegradation in performance from many fsync()'s.\n\nIOW I think it's perfectly fine not to add a config knob for it other\nthan core.fsync=false, and VERY SOON turn on\ncore.fsyncLooseobjects=true, especially if we can get most of the\nperformance benefits with the \"bulk\" mode.\n\nBut why half-way with bc22d845c43? I mean, *that change* should be\nnarrow, but in terms of where we go next what do you think of the above?\n\n> So while the end result will be the same as `committed`, having this\n> level of control helps us to assess the impact in a nicer way by being\n> able to do this step by step with feature flags.\n\n*Nod*, although leaving aside the new syncing of loose refs the plan you\n outlined above could be done with the proposal of the simpler:\n\n    core.fsync=true\n    core.fsyncLooseObjects=false\n\n> On the other hand, many of the other parts we don't really care about.\n> Auxiliary metadata like the commit-graph or pack indices are data that\n> can in the worst case be regenerated by us, so it's not clear to me\n> whether it makes to also enable fsyncing those in production.\n\nI'm not familiar with all of those in detail, i.e. how we behave\nspecifically in the face of them being corrupt.\n\nThe commit-graph I am, we *should* recover \"gracefully\" there, but you\nmight have an incident shortly there after due to \"for-each-ref\n--contains\" slowdowns by 1000x or whatever.\n\nFor e.g. *.idx we'd be hosed until manual recovery.\n\nSo I think those area all in the same bucket as core.fsync=false, and\nthat we don't need the granularity.\n\n> So altogether, I agree with Neeraj: having the fine-grainedness greatly\n> helps us to roll out changes like this and be able to pick what we deem\n> to be important. Personally I would be fine with explicitly pointing out\n> that there are two groups of this config in our docs though:\n\nYes, that's fair. But pending replies to the above I think the main\npoint & proposal of us having too much config stands. I.e. depending on\nwhat you want to do with loose object refs we'd just need this:\n\n    core.fsync=[<bool>] # true by default\n    core.fsyncLooseObjects=[<bool>] # false by default\n    core.fsyncLooseRefs=[<bool>] # true by default?\n\nAnd then the \"bulk\" config, which would be orthagonal to this.\n\nI.e. do we really have a use-case for the rest of the kitchen sink?\n\n>     1. The \"porcelain\" group: \"committed\", \"added\", \"all\", \"none\". These\n>        are abstract groups whose behaviour should adapt as we change\n>        implementations, and are those that should typically be set by a\n>        user, if intended.\n>\n>     2. The \"plumbing\" or \"expert\" group: these are fine-grained options\n>        which shouldn't typically be used by Git users. They still have\n>        merit though in hosting environments, where requirements are\n>        typically a lot more specific.\n>\n> We may also provide different guarantees for both groups. The first one\n> should definitely be stable, but we might state that the second group is\n> subject to change in the future.\n\nI hope we can work something out :)\n\nOverall: I think you've left one of the the main things I brought up[1]\nunaddressed, i.e. that the core.fsync config schema in its current form\nassumes that we can sync A or B, and configure those separately.\n\nWhich AFAIKT is because Neeraj's initial implementation & the discussion\nwas focused on finishing A or B with a per-\"group\" \"cookie\" to flush the\nfiles.\n\nBut as [2] shows it's more performant for us to simply defer the fsync\nof A until the committing of B.\n\nWhich is the main reason I think we should be re-visiting this. Sure, if\nwe were just syncing A, B or C having per-[ABC] config options might be\na bit overdoing it, but would be relatively simple.\n\nBut once we start using a more optimized version of the \"bulk\" mode the\nconfig schema will be making promises about individual steps in a\ntransaction that I think we'll want to leave opaque, and only promise\nthat when git returns it will have synced all the relevant assets as\nefficiently as possible.\n\n1. https://lore.kernel.org/git/220323.86fsn8ohg8.gmgdl@evledraar.gmail.com/\n2. https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n"},{"id":"452502","messageId":"CANQDOdfJVdCzqk-98i7E7XZN25+0ZLYPjKWOVs9YNUPwx1JDWQ@mail.gmail.com","threadId":"57030","inReplyTo":"CANYiYbFjRMV-_opvFn78mq7tgtZFMrfPyDjDa+kyaZZfk_LmWQ@mail.gmail.com","subject":"Re: [PATCH v6 4/6] core.fsync: add configuration parsing","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-28T19:45:20Z","receivedAt":"2022-03-28T19:49:02Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 28, 2022 at 4:07 AM Jiang Xin <worldhello.net@gmail.com> wrote:\n>\n> On Sat, Mar 12, 2022 at 6:25 AM Neeraj Singh via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> > @@ -1613,6 +1687,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >         }\n> >\n> >         if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> > +               if (fsync_object_files < 0)\n> > +                       warning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n>\n> s/core.fsyncobjectfiles/core.fsyncObjectFiles/  to use bumpyCaps for\n> config variable in documentation.\n\nTHanks for pointing this out.  I'll fix it.\n"},{"id":"452504","messageId":"CANQDOddSnSByqMrU+b11zywwWjOS+A5W0BSXa0rZURWn5zi2Tg@mail.gmail.com","threadId":"57030","inReplyTo":"220328.86pmm6e0i0.gmgdl@evledraar.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-28T19:56:46Z","receivedAt":"2022-03-28T19:57:21Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 28, 2022 at 5:15 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> I hope we can work something out :)\n>\n> Overall: I think you've left one of the the main things I brought up[1]\n> unaddressed, i.e. that the core.fsync config schema in its current form\n> assumes that we can sync A or B, and configure those separately.\n>\n> Which AFAIKT is because Neeraj's initial implementation & the discussion\n> was focused on finishing A or B with a per-\"group\" \"cookie\" to flush the\n> files.\n>\n> But as [2] shows it's more performant for us to simply defer the fsync\n> of A until the committing of B.\n>\n> Which is the main reason I think we should be re-visiting this. Sure, if\n> we were just syncing A, B or C having per-[ABC] config options might be\n> a bit overdoing it, but would be relatively simple.\n>\n> But once we start using a more optimized version of the \"bulk\" mode the\n> config schema will be making promises about individual steps in a\n> transaction that I think we'll want to leave opaque, and only promise\n> that when git returns it will have synced all the relevant assets as\n> efficiently as possible.\n>\n> 1. https://lore.kernel.org/git/220323.86fsn8ohg8.gmgdl@evledraar.gmail.com/\n> 2. https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n\nI think the current documentation is fine (obviously since I wrote\nit).  Let's reproduce the first part again:\n---\ncore.fsync::\nA comma-separated list of components of the repository that\nshould be hardened via the core.fsyncMethod when created or\nmodified.  You can disable hardening of any component by\nprefixing it with a '-'.  Items that are not hardened may be\nlost in the event of an unclean system shutdown. Unless you\nhave special requirements, it is recommended that you leave\nthis option empty or pick one of `committed`, `added`,\nor `all`.\n+\nWhen this configuration is encountered, the set of components starts with\nthe platform default value, disabled components are removed, and additional\ncomponents are added. `none` resets the state so that the platform default\nis ignored.\n+\nThe empty string resets the fsync configuration to the platform\ndefault. The default on most platforms is equivalent to\n`core.fsync=committed,-loose-object`, which has good performance,\nbut risks losing recent work in the event of an unclean system shutdown.\n+\n---\n\nWe're only talking about \"hardening\" parts of the repository, and we\nsay we'll do it using the \"fsyncMethod\".  If you don't harden\nsomething, you could lose it if the system dies.  All of these\nstatements are true and don't say so much about the implementation of\nhow the hardening happens. It's perfectly valid to not force any\ncomponent out of the disk cache, and a straightforward implementation\nof repo transactions can put the sync point anywhere. We also\nexplicitly point the user at a \"porcelain\" setting in the first\nparagraph.\n\nSo I think you're alone in thinking that anything needs to change here.\n\nThanks,\nNeeraj\n"},{"id":"452669","messageId":"CANQDOddVLOZJZtvTE3zDizwWdu3RsxiE6dCuNc3D=RyGsLRRqA@mail.gmail.com","threadId":"57030","inReplyTo":"CANQDOddSnSByqMrU+b11zywwWjOS+A5W0BSXa0rZURWn5zi2Tg@mail.gmail.com","subject":"Re: do we have too much fsync() configuration in 'next'? (was: [PATCH v7] core.fsync: documentation and user-friendly aggregate options)","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-30T16:59:26Z","receivedAt":"2022-03-30T16:59:55Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 28, 2022 at 12:56 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> On Mon, Mar 28, 2022 at 5:15 AM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>\n> So I think you're alone in thinking that anything needs to change here.\n>\n\nÆvar,\nI apologize for the negative tone and content of this comment.\n\nThanks,\nNeeraj\n"}]}