{"thread":{"id":"56371","subject":"[PATCH 1/2] object-file: use futimes rather than utime","startedAt":"2021-08-25T01:51:37Z","lastAt":"2022-03-10T18:48:53Z","messageCount":160,"participants":["Neeraj Singh via GitGitGadget","Neeraj K. Singh via GitGitGadget","Christoph Hellwig","Johannes Schindelin","Ævar Arnfjörð Bjarmason","Neeraj Singh","Junio C Hamano","Randall S. Becker","'Christoph Hellwig'","Bagas Sanjaya","Jeff King","Elijah Newren","rsbecker@nexbridge.com"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"433689","messageId":"pull.1076.git.git.1629856292.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":null,"subject":"[PATCH 0/2] [RFC] Implement a bulk-checkin option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-25T01:51:30Z","receivedAt":"2021-08-25T01:51:37Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Git for Windows has had fsyncing of object files enabled since \"409cae91eb\n(mingw: change core.fsyncObjectFiles = 1 by default, 2017-09-04)\".\n\nThere have been requests to make core.fsyncObjectFiles the default\neverywhere, but there are concerns about its performance cost (perf results\nbelow). There's a long and gory thread here:\nhttps://lore.kernel.org/git/87a7xcw8sa.fsf@linux-m68k.org/t/.\n\nMy change introduces the new 'core.fsyncobjectFiles = 2' setting, which\nbatches the data-integrity FLUSH command sent to the disk across multiple\nloose object files added to the object database.\n\nWe take advantage of the bulk-checkin hooks already in the add command and\nadd some hooks to the update-index (which is used internally by stash).\nDetails are in the last patch of the series.\n\nHere's a simple performance test script:\n\n    #!/bin/sh\n    git clone https://github.com/nodejs/node.git node-repo-cache\n    git clone node-repo-cache node-repo\n    cd node-repo\n    git --version\n    \n    find . -name \"*.c\" -exec sh -c 'echo foo1 >> $1' -- {} \\;\n    echo \"----GIT stash fsync\"\n    time git -c core.fsyncObjectFiles=true stash push\n    \n    find . -name \"*.c\" -exec sh -c 'echo foo2 >> $1' -- {} \\;\n    echo \"----GIT stash fsync_defer\"\n    time git -c core.fsyncObjectFiles=2 stash push\n    \n    find . -name \"*.c\" -exec sh -c 'echo foo3 >> $1' -- {} \\;\n    echo \"----GIT stash no_fsync\"\n    time git -c core.fsyncObjectFiles=false stash push\n    \n    cd ..\n    rm -r -f node-repo\n\n\nHardware:\n\n * Mac - Mac Mini 2018 running MacOS 11.5.1, APFS with a 1TB Apple NMVE SSD,\n * Linux - Ubuntu 20.04 - ext4 running on a Hyper-V VM with a fixed VHDX\n   backed by a Samsung PM981.\n * Win - Windows NTFS - Same Hyper-V host as Linux. Operation | Mac | Linux\n   | Windows\n\n---------------- |---------|-------|---------- git fsync | 40.6 s | 7.8 s |\n6.9s git fsync_defer | 6.5 s | 2.1 s | 3.8s git no_fsync | 1.7 s | 1.0 s |\n2.6s\n\nThe windows version of git is slightly different:\nhttps://github.com/git-for-windows/git/pull/3391. I also used a\nWindows-specific test script.\n\nI hope I'm CC'ing a reasonable set of people on this patch, based on the\nlast discussion.\n\nThanks, Neeraj Singh Windows Core File Systems.\n\nNeeraj Singh (2):\n  object-file: use futimes rather than utime\n  core.fsyncobjectfiles: batch disk flushes\n\n Documentation/config/core.txt |  17 ++++--\n Makefile                      |   4 ++\n builtin/add.c                 |   3 +-\n builtin/update-index.c        |   3 +\n bulk-checkin.c                | 105 +++++++++++++++++++++++++++++++---\n bulk-checkin.h                |   4 +-\n compat/mingw.c                |  42 +++++++++-----\n compat/mingw.h                |   2 +\n config.c                      |   4 +-\n config.mak.uname              |   2 +\n configure.ac                  |   8 +++\n git-compat-util.h             |   7 +++\n object-file.c                 |  23 ++------\n wrapper.c                     |  36 ++++++++++++\n write-or-die.c                |   2 +-\n 15 files changed, 213 insertions(+), 49 deletions(-)\n\n\nbase-commit: 225bc32a989d7a22fa6addafd4ce7dcd04675dbf\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v1\nPull-Request: https://github.com/git/git/pull/1076\n-- \ngitgitgadget\n"},{"id":"433688","messageId":"2c1ddef6057157d85da74a7274e03eacf0374e45.1629856293.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.git.git.1629856292.gitgitgadget@gmail.com","subject":"[PATCH 1/2] object-file: use futimes rather than utime","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-25T01:51:31Z","receivedAt":"2021-08-25T01:51:39Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nRefactor the loose object file creation code and use the futimes(2) API\nrather than utime. This should be slightly faster given that we already\nhave an FD to work with.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.c | 42 +++++++++++++++++++++++++++++-------------\n compat/mingw.h |  2 ++\n object-file.c  | 17 ++++++++---------\n 3 files changed, 39 insertions(+), 22 deletions(-)\n\ndiff --git a/compat/mingw.c b/compat/mingw.c\nindex 9e0cd1e097f..948f4c3428b 100644\n--- a/compat/mingw.c\n+++ b/compat/mingw.c\n@@ -949,19 +949,40 @@ int mingw_fstat(int fd, struct stat *buf)\n \t}\n }\n \n-static inline void time_t_to_filetime(time_t t, FILETIME *ft)\n+static inline void timeval_to_filetime(const struct timeval *t, FILETIME *ft)\n {\n-\tlong long winTime = t * 10000000LL + 116444736000000000LL;\n+\tlong long winTime = t->tv_sec * 10000000LL + t->tv_usec * 10 + 116444736000000000LL;\n \tft->dwLowDateTime = winTime;\n \tft->dwHighDateTime = winTime >> 32;\n }\n \n-int mingw_utime (const char *file_name, const struct utimbuf *times)\n+int mingw_futimes(int fd, const struct timeval times[2])\n {\n \tFILETIME mft, aft;\n+\n+\tif (times) {\n+\t\ttimeval_to_filetime(&times[0], &aft);\n+\t\ttimeval_to_filetime(&times[1], &mft);\n+\t} else {\n+\t\tGetSystemTimeAsFileTime(&mft);\n+\t\taft = mft;\n+\t}\n+\n+\tif (!SetFileTime((HANDLE)_get_osfhandle(fd), NULL, &aft, &mft)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+int mingw_utime (const char *file_name, const struct utimbuf *times)\n+{\n \tint fh, rc;\n \tDWORD attrs;\n \twchar_t wfilename[MAX_PATH];\n+\tstruct timeval tvs[2];\n+\n \tif (xutftowcs_path(wfilename, file_name) < 0)\n \t\treturn -1;\n \n@@ -979,17 +1000,12 @@ int mingw_utime (const char *file_name, const struct utimbuf *times)\n \t}\n \n \tif (times) {\n-\t\ttime_t_to_filetime(times->modtime, &mft);\n-\t\ttime_t_to_filetime(times->actime, &aft);\n-\t} else {\n-\t\tGetSystemTimeAsFileTime(&mft);\n-\t\taft = mft;\n+\t\tmemset(tvs, 0, sizeof(tvs));\n+\t\ttvs[0].tv_sec = times->actime;\n+\t\ttvs[1].tv_sec = times->modtime;\n \t}\n-\tif (!SetFileTime((HANDLE)_get_osfhandle(fh), NULL, &aft, &mft)) {\n-\t\terrno = EINVAL;\n-\t\trc = -1;\n-\t} else\n-\t\trc = 0;\n+\n+\trc = mingw_futimes(fh, times ? tvs : NULL);\n \tclose(fh);\n \n revert_attrs:\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..1eb14edb2ed 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -398,6 +398,8 @@ int mingw_fstat(int fd, struct stat *buf);\n \n int mingw_utime(const char *file_name, const struct utimbuf *times);\n #define utime mingw_utime\n+int mingw_futimes(int fd, const struct timeval times[2]);\n+#define futimes mingw_futimes\n size_t mingw_strftime(char *s, size_t max,\n \t\t   const char *format, const struct tm *tm);\n #define strftime mingw_strftime\ndiff --git a/object-file.c b/object-file.c\nindex a8be8994814..607e9e2f80b 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1860,12 +1860,13 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n }\n \n /* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd)\n+static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n {\n \tif (fsync_object_files)\n \t\tfsync_or_die(fd, \"loose object file\");\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n+\treturn finalize_object_file(tmpfile, filename);\n }\n \n /* Size of directory component, including the ending '/' */\n@@ -1973,17 +1974,15 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n-\n \tif (mtime) {\n-\t\tstruct utimbuf utb;\n-\t\tutb.actime = mtime;\n-\t\tutb.modtime = mtime;\n-\t\tif (utime(tmp_file.buf, &utb) < 0)\n-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n+\t\tstruct timeval tvs[2] = {0};\n+\t\ttvs[0].tv_sec = mtime;\n+\t\ttvs[1].tv_sec = mtime;\n+\t\tif (futimes(fd, tvs) < 0)\n+\t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n \t}\n \n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\treturn close_loose_object(fd, tmp_file.buf, filename.buf);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\n-- \ngitgitgadget\n\n"},{"id":"433690","messageId":"d1e68d4a2afc1d0ba74af64680bea09f412f21cc.1629856293.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.git.git.1629856292.gitgitgadget@gmail.com","subject":"[PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-25T01:51:32Z","receivedAt":"2021-08-25T01:51:40Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nMacOS, and Linux each offer mechanisms to write data from the filesystem\npage cache without initiating a hardware flush.\n\nThis patch introduces a new 'core.fsyncObjectFiles = 2' option that\ntakes advantage of the bulk-checkin infrastructure to batch up hardware\nflushes.\n\nWhen the new mode is enabled we do the following for new objects:\n\n1. Create a tmp_obj_XXXX file and write the object data to it.\n2. Issue a pagecache writeback request and wait for it to complete.\n3. Record the tmp name and the final name in the bulk-checkin state for\n   later name.\n\nAt the end of the entire transaction we:\n1. Issue a fsync against the lock file to flush the hardware writeback\n   cache, which should by now have processed the tmp file writes.\n2. Rename all of the temp files to their final names.\n3. When updating the index and/or refs, we will issue another fsync\n   internal to that operation.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\nwould expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the MacOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nMacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  17 ++++--\n Makefile                      |   4 ++\n builtin/add.c                 |   3 +-\n builtin/update-index.c        |   3 +\n bulk-checkin.c                | 105 +++++++++++++++++++++++++++++++---\n bulk-checkin.h                |   4 +-\n config.c                      |   4 +-\n config.mak.uname              |   2 +\n configure.ac                  |   8 +++\n git-compat-util.h             |   7 +++\n object-file.c                 |  12 +---\n wrapper.c                     |  36 ++++++++++++\n write-or-die.c                |   2 +-\n 13 files changed, 177 insertions(+), 30 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..3b672c2db67 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,17 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA boolean value or the number '2', indicating the level of durability\n+\tapplied to object files.\n++\n+This setting controls how much effort Git makes to ensure that data added to\n+the object store are durable in the case of an unclean system shutdown. If\n+'false', Git allows data to remain in file system caches according to operating\n+system policy, whence they may be lost if the system loses power or crashes. A\n+value of 'true' instructs Git to force objects to stable storage immediately\n+when they are added to the object store. The number '2' is an experimental\n+value that also preserves durability but tries to perform hardware flushes in a\n+batch.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex 9573190f1d7..cb950ee43d3 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1896,6 +1896,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 09e684585d9..c58dfcd4bc3 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -670,7 +670,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\n+\tunplug_bulk_checkin(&lock_file);\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex f1f16f2de52..64d025cf49e 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1152,6 +1153,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n \n \t\tsetup_work_tree();\n+\t\tplug_bulk_checkin();\n \t\twhile (getline_fn(&buf, stdin) != EOF) {\n \t\t\tchar *p;\n \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n@@ -1166,6 +1168,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n \t\t}\n+\t\tunplug_bulk_checkin(&lock_file);\n \t\tstrbuf_release(&unquoted);\n \t\tstrbuf_release(&buf);\n \t}\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b023d9959aa..71004db863e 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,6 +3,7 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n@@ -10,6 +11,17 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n+struct object_rename {\n+\tchar *src;\n+\tchar *dst;\n+};\n+\n+static struct bulk_rename_state {\n+\tstruct object_rename *renames;\n+\tuint32_t alloc_renames;\n+\tuint32_t nr_renames;\n+} bulk_rename_state;\n+\n static struct bulk_checkin_state {\n \tunsigned plugged:1;\n \n@@ -21,13 +33,15 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+\n+} bulk_checkin_state;\n \n static void finish_bulk_checkin(struct bulk_checkin_state *state)\n {\n \tstruct object_id oid;\n \tstruct strbuf packname = STRBUF_INIT;\n \tint i;\n+\tunsigned old_plugged;\n \n \tif (!state->f)\n \t\treturn;\n@@ -55,13 +69,42 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \n clear_exit:\n \tfree(state->written);\n+\told_plugged = state->plugged;\n \tmemset(state, 0, sizeof(*state));\n+\tstate->plugged = old_plugged;\n \n \tstrbuf_release(&packname);\n \t/* Make objects we just wrote available to ourselves */\n \treprepare_packed_git(the_repository);\n }\n \n+static void do_sync_and_rename(struct bulk_rename_state *state, struct lock_file *lock_file)\n+{\n+\tif (state->nr_renames) {\n+\t\tint i;\n+\n+\t\t/*\n+\t\t * Issue a full hardware flush against the lock file to ensure\n+\t\t * that all objects are durable before any renames occur.\n+\t\t * The code in fsync_and_close_loose_object_bulk_checkin has\n+\t\t * already ensured that writeout has occurred, but it has not\n+\t\t * flushed any writeback cache in the storage hardware.\n+\t\t */\n+\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n+\n+\t\tfor (i = 0; i < state->nr_renames; i++) {\n+\t\t\tif (finalize_object_file(state->renames[i].src, state->renames[i].dst))\n+\t\t\t\tdie_errno(_(\"could not rename '%s'\"), state->renames[i].src);\n+\n+\t\t\tfree(state->renames[i].src);\n+\t\t\tfree(state->renames[i].dst);\n+\t\t}\n+\n+\t\tfree(state->renames);\n+\t\tmemset(state, 0, sizeof(*state));\n+\t}\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -256,25 +299,69 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+static void add_rename_bulk_checkin(struct bulk_rename_state *state,\n+\t\t\t\t    const char *src, const char *dst)\n+{\n+\tstruct object_rename *rename;\n+\n+\tALLOC_GROW(state->renames, state->nr_renames + 1, state->alloc_renames);\n+\n+\trename = &state->renames[state->nr_renames++];\n+\trename->src = xstrdup(src);\n+\trename->dst = xstrdup(dst);\n+}\n+\n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n+\t\t\t\t\t      const char *filename)\n+{\n+\tif (fsync_object_files) {\n+\t\t/*\n+\t\t * If we have a plugged bulk checkin, we issue a call that\n+\t\t * cleans the filesystem page cache but avoids a hardware flush\n+\t\t * command. Later on we will issue a single hardware flush\n+\t\t * before renaming files as part of do_sync_and_rename.\n+\t\t */\n+\t\tif (bulk_checkin_state.plugged &&\n+\t\t    fsync_object_files == 2 &&\n+\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\t\tadd_rename_bulk_checkin(&bulk_rename_state, tmpfile, filename);\n+\t\t\tif (close(fd))\n+\t\t\t\tdie_errno(_(\"error when closing loose object file\"));\n+\n+\t\t\treturn 0;\n+\n+\t\t} else {\n+\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\t}\n+\t}\n+\n+\tif (close(fd))\n+\t\tdie_errno(_(\"error when closing loose object file\"));\n+\n+\treturn finalize_object_file(tmpfile, filename);\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_state.plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tbulk_checkin_state.plugged = 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void unplug_bulk_checkin(struct lock_file *lock_file)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tbulk_checkin_state.plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_sync_and_rename(&bulk_rename_state, lock_file);\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..8efb01ed669 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,11 +6,13 @@\n \n #include \"cache.h\"\n \n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile, const char *filename);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+void unplug_bulk_checkin(struct lock_file *);\n \n #endif\ndiff --git a/config.c b/config.c\nindex f33abeab851..375bdb24b0a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1509,7 +1509,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tint is_bool;\n+\n+\t\tfsync_object_files = git_config_bool_or_int(var, value, &is_bool);\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 69413fb3dc0..8c07f2265a8 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n@@ -133,6 +134,7 @@ ifeq ($(uname_S),Darwin)\n \tCOMPAT_OBJS += compat/precompose_utf8.o\n \tBASIC_CFLAGS += -DPRECOMPOSE_UNICODE\n \tBASIC_CFLAGS += -DPROTECT_HFS_DEFAULT=1\n+\tBASIC_CFLAGS += -DFSYNC_DOESNT_FLUSH=1\n \tHAVE_BSD_SYSCTL = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tHAVE_NS_GET_EXECUTABLE_PATH = YesPlease\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b46605300ab..d14e2436276 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex 607e9e2f80b..5f04143dde0 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1859,16 +1859,6 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \treturn 0;\n }\n \n-/* Finalize a file on disk, and close it. */\n-static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n-{\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n-\tif (close(fd) != 0)\n-\t\tdie_errno(_(\"error when closing loose object file\"));\n-\treturn finalize_object_file(tmpfile, filename);\n-}\n-\n /* Size of directory component, including the ending '/' */\n static inline int directory_size(const char *filename)\n {\n@@ -1982,7 +1972,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n \t}\n \n-\treturn close_loose_object(fd, tmp_file.buf, filename.buf);\n+\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf, filename.buf);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\ndiff --git a/wrapper.c b/wrapper.c\nindex 563ad590df1..37a8b61a7df 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -538,6 +538,42 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tif (action == FSYNC_WRITEOUT_ONLY) {\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on Mac OS X, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\t}\n+\n+#ifdef __APPLE__\n+\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\treturn fsync(fd);\n+#endif\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex d33e68f6abb..8f53953d4ab 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n"},{"id":"433708","messageId":"20210825053839.GA27037@lst.de","threadId":"56371","inReplyTo":"d1e68d4a2afc1d0ba74af64680bea09f412f21cc.1629856293.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-08-25T05:38:39Z","receivedAt":"2021-08-25T05:38:43Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Wed, Aug 25, 2021 at 01:51:32AM +0000, Neeraj Singh via GitGitGadget wrote:\n> From: Neeraj Singh <neerajsi@microsoft.com>\n> \n> When adding many objects to a repo with core.fsyncObjectFiles set to\n> true, the cost of fsync'ing each object file can become prohibitive.\n> \n> One major source of the cost of fsync is the implied flush of the\n> hardware writeback cache within the disk drive. Fortunately, Windows,\n> MacOS, and Linux each offer mechanisms to write data from the filesystem\n> page cache without initiating a hardware flush.\n> \n> This patch introduces a new 'core.fsyncObjectFiles = 2' option that\n> takes advantage of the bulk-checkin infrastructure to batch up hardware\n> flushes.\n\nAnother interesting way to flush on linux would be the syncfs call,\nwhich syncs all files on a file system.  Once you write more than\nhandful or two of files that tends to win out over a batch of fsync\ncalls.\n"},{"id":"433734","messageId":"nycvar.QRO.7.76.6.2108251530210.55@tvgsbejvaqbjf.bet","threadId":"56371","inReplyTo":"2c1ddef6057157d85da74a7274e03eacf0374e45.1629856293.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/2] object-file: use futimes rather than utime","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2021-08-25T13:51:01Z","receivedAt":"2021-08-25T13:51:13Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Neeraj,\n\nThank you so much for this patch series! Overall, I am very happy with the\ndirection this is going.\n\nI will offer a couple of suggestions below, inlined.\n\nOn Wed, 25 Aug 2021, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Refactor the loose object file creation code and use the futimes(2) API\n> rather than utime. This should be slightly faster given that we already\n> have an FD to work with.\n\nIf I were you, I would spell out \"file descriptor\" here.\n\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  compat/mingw.c | 42 +++++++++++++++++++++++++++++-------------\n>  compat/mingw.h |  2 ++\n>  object-file.c  | 17 ++++++++---------\n>  3 files changed, 39 insertions(+), 22 deletions(-)\n>\n> diff --git a/compat/mingw.c b/compat/mingw.c\n> index 9e0cd1e097f..948f4c3428b 100644\n> --- a/compat/mingw.c\n> +++ b/compat/mingw.c\n> @@ -949,19 +949,40 @@ int mingw_fstat(int fd, struct stat *buf)\n>  \t}\n>  }\n>\n> -static inline void time_t_to_filetime(time_t t, FILETIME *ft)\n> +static inline void timeval_to_filetime(const struct timeval *t, FILETIME *ft)\n>  {\n> -\tlong long winTime = t * 10000000LL + 116444736000000000LL;\n> +\tlong long winTime = t->tv_sec * 10000000LL + t->tv_usec * 10 + 116444736000000000LL;\n\nTechnically, this is a change in behavior, right? We did not use to use\nnanosecond precision. But I don't think that we actually make use of this\nin this patch.\n\n>  \tft->dwLowDateTime = winTime;\n>  \tft->dwHighDateTime = winTime >> 32;\n>  }\n>\n> -int mingw_utime (const char *file_name, const struct utimbuf *times)\n> +int mingw_futimes(int fd, const struct timeval times[2])\n\nAt first, I wondered whether it would make sense to pass the access time\nand the modified time separately, as pointers. I don't think that we pass\naround arrays as function parameters in Git anywhere else.\n\nBut then I realized that `futimes()` is available in this precise form on\nLinux and on the BSDs. Therefore, it is not up to us to decide the\nfunction's signature.\n\nHowever, now that I looked at the manual page, I noticed that this\nfunction is not part of any POSIX standard.\n\nWhich makes me think that we will have to do a bit more than just define\nit on Windows: we will have to introduce a `Makefile` knob (just like you\ndid with `HAVE_SYNC_FILE_RANGE` in patch 2/2) and set that specifically\nfor Linux and the BSDs, and use `futimes()` only if it is available\n(otherwise fall back to `utime()`).\n\nThen, as a separate patch, we should introduce this Windows-specific shim\nand declare that it is available via `config.mak.uname`.\n\nI am a _huge_ fan of patches that are so clear and obvious that bugs have\na hard time creeping in without being spotted immediately. And I think\nthat this organization would help achieve this goal.\n\n>  {\n>  \tFILETIME mft, aft;\n> +\n> +\tif (times) {\n> +\t\ttimeval_to_filetime(&times[0], &aft);\n> +\t\ttimeval_to_filetime(&times[1], &mft);\n> +\t} else {\n> +\t\tGetSystemTimeAsFileTime(&mft);\n> +\t\taft = mft;\n> +\t}\n> +\n> +\tif (!SetFileTime((HANDLE)_get_osfhandle(fd), NULL, &aft, &mft)) {\n> +\t\terrno = EINVAL;\n> +\t\treturn -1;\n> +\t}\n> +\n> +\treturn 0;\n> +}\n> +\n> +int mingw_utime (const char *file_name, const struct utimbuf *times)\n\nPlease lose the space between the function name and the opening\nparenthesis. I know, the preimage of this diff has it, but that was an\noversight and definitely disagrees with our current coding style.\n\n> +{\n>  \tint fh, rc;\n>  \tDWORD attrs;\n>  \twchar_t wfilename[MAX_PATH];\n> +\tstruct timeval tvs[2];\n> +\n>  \tif (xutftowcs_path(wfilename, file_name) < 0)\n>  \t\treturn -1;\n>\n> @@ -979,17 +1000,12 @@ int mingw_utime (const char *file_name, const struct utimbuf *times)\n>  \t}\n>\n>  \tif (times) {\n> -\t\ttime_t_to_filetime(times->modtime, &mft);\n> -\t\ttime_t_to_filetime(times->actime, &aft);\n> -\t} else {\n> -\t\tGetSystemTimeAsFileTime(&mft);\n> -\t\taft = mft;\n> +\t\tmemset(tvs, 0, sizeof(tvs));\n> +\t\ttvs[0].tv_sec = times->actime;\n> +\t\ttvs[1].tv_sec = times->modtime;\n\nIt is too bad that we have to copy around those values just to convert\nthem, but I cannot think of any better way, either. And it's not like\nwe're in a hot loop: this code will be dominated by I/O anyways.\n\n>  \t}\n> -\tif (!SetFileTime((HANDLE)_get_osfhandle(fh), NULL, &aft, &mft)) {\n> -\t\terrno = EINVAL;\n> -\t\trc = -1;\n> -\t} else\n> -\t\trc = 0;\n> +\n> +\trc = mingw_futimes(fh, times ? tvs : NULL);\n>  \tclose(fh);\n>\n>  revert_attrs:\n> diff --git a/compat/mingw.h b/compat/mingw.h\n> index c9a52ad64a6..1eb14edb2ed 100644\n> --- a/compat/mingw.h\n> +++ b/compat/mingw.h\n> @@ -398,6 +398,8 @@ int mingw_fstat(int fd, struct stat *buf);\n>\n>  int mingw_utime(const char *file_name, const struct utimbuf *times);\n>  #define utime mingw_utime\n> +int mingw_futimes(int fd, const struct timeval times[2]);\n> +#define futimes mingw_futimes\n>  size_t mingw_strftime(char *s, size_t max,\n>  \t\t   const char *format, const struct tm *tm);\n>  #define strftime mingw_strftime\n> diff --git a/object-file.c b/object-file.c\n> index a8be8994814..607e9e2f80b 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1860,12 +1860,13 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  }\n>\n>  /* Finalize a file on disk, and close it. */\n> -static void close_loose_object(int fd)\n> +static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n>  {\n>  \tif (fsync_object_files)\n>  \t\tfsync_or_die(fd, \"loose object file\");\n>  \tif (close(fd) != 0)\n>  \t\tdie_errno(_(\"error when closing loose object file\"));\n> +\treturn finalize_object_file(tmpfile, filename);\n\nWhile this is a clear change of behavior, this function has only one\ncaller, and that caller is adjusted accordingly.\n\nCould you add this clarification of context to the commit message? I know\nit will help me in the future, when I have to get up to speed again by\nreading the commit history.\n\nThank you,\nJohannes\n\n>  }\n>\n>  /* Size of directory component, including the ending '/' */\n> @@ -1973,17 +1974,15 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>\n> -\tclose_loose_object(fd);\n> -\n>  \tif (mtime) {\n> -\t\tstruct utimbuf utb;\n> -\t\tutb.actime = mtime;\n> -\t\tutb.modtime = mtime;\n> -\t\tif (utime(tmp_file.buf, &utb) < 0)\n> -\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n> +\t\tstruct timeval tvs[2] = {0};\n> +\t\ttvs[0].tv_sec = mtime;\n> +\t\ttvs[1].tv_sec = mtime;\n> +\t\tif (futimes(fd, tvs) < 0)\n> +\t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n>  \t}\n>\n> -\treturn finalize_object_file(tmp_file.buf, filename.buf);\n> +\treturn close_loose_object(fd, tmp_file.buf, filename.buf);\n>  }\n>\n>  static int freshen_loose_object(const struct object_id *oid)\n> --\n> gitgitgadget\n>\n>\n"},{"id":"433750","messageId":"87mtp5cwpn.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"d1e68d4a2afc1d0ba74af64680bea09f412f21cc.1629856293.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-08-25T16:11:13Z","receivedAt":"2021-08-25T16:31:39Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Aug 25 2021, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> When adding many objects to a repo with core.fsyncObjectFiles set to\n> true, the cost of fsync'ing each object file can become prohibitive.\n>\n> One major source of the cost of fsync is the implied flush of the\n> hardware writeback cache within the disk drive. Fortunately, Windows,\n> MacOS, and Linux each offer mechanisms to write data from the filesystem\n> page cache without initiating a hardware flush.\n>\n> This patch introduces a new 'core.fsyncObjectFiles = 2' option that\n> takes advantage of the bulk-checkin infrastructure to batch up hardware\n> flushes.\n>\n> When the new mode is enabled we do the following for new objects:\n>\n> 1. Create a tmp_obj_XXXX file and write the object data to it.\n> 2. Issue a pagecache writeback request and wait for it to complete.\n> 3. Record the tmp name and the final name in the bulk-checkin state for\n>    later name.\n>\n> At the end of the entire transaction we:\n> 1. Issue a fsync against the lock file to flush the hardware writeback\n>    cache, which should by now have processed the tmp file writes.\n> 2. Rename all of the temp files to their final names.\n> 3. When updating the index and/or refs, we will issue another fsync\n>    internal to that operation.\n>\n> On a filesystem with a singular journal that is updated during name\n> operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n> would expect the fsync to trigger a journal writeout so that this\n> sequence is enough to ensure that the user's data is durable by the time\n> the git command returns.\n>\n> This change also updates the MacOS code to trigger a real hardware flush\n> via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n> MacOS there was no guarantee of durability since a simple fsync(2) call\n> does not flush any hardware caches.\n\nThanks for working on this, good to see fsck issues picked up after some\non-list pause.\n\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index c04f62a54a1..3b672c2db67 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -548,12 +548,17 @@ core.whitespace::\n>    errors. The default tab width is 8. Allowed values are 1 to 63.\n>  \n>  core.fsyncObjectFiles::\n> -\tThis boolean will enable 'fsync()' when writing object files.\n> -+\n> -This is a total waste of time and effort on a filesystem that orders\n> -data writes properly, but can be useful for filesystems that do not use\n> -journalling (traditional UNIX filesystems) or that only journal metadata\n> -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n> +\tA boolean value or the number '2', indicating the level of durability\n> +\tapplied to object files.\n> ++\n> +This setting controls how much effort Git makes to ensure that data added to\n> +the object store are durable in the case of an unclean system shutdown. If\n> +'false', Git allows data to remain in file system caches according to operating\n> +system policy, whence they may be lost if the system loses power or crashes. A\n> +value of 'true' instructs Git to force objects to stable storage immediately\n> +when they are added to the object store. The number '2' is an experimental\n> +value that also preserves durability but tries to perform hardware flushes in a\n> +batch.\n\nSome feedback/thoughts:\n\n0) Let's not expose \"2\" to users, but give it some friendly config name\nand just translate this to the enum internally.\n\n1) Your commit message says \"When updating the index and/or refs[...]\"\nbut we're changing core.fsyncObjectFiles here, I assume that's\nsummarizing existing behavior then\n\n2) You say \"when adding many [loose] objects to a repo[...]\", and the\ntest-case is \"git stash push\", but for e.g. accepting pushes we have\ntransfer.unpackLimit.\n\nIt would be interesting to see if/how this impacts performance there,\nand also if not that should at least be called out in\ndocumentation. I.e. users might want to set this differently on servers\nv.s. checkouts.\n\nBut also, is this sort of thing something we could mitigate even more in\ncommands like \"git stash push\" by just writing a pack instead of N loose\nobjects?\n\nI don't think such questions should preclude changing the fsync\napproach, or offering more options, but they should inform our\nlonger-term goals.\n\n3) Re some of the musings about fsync() recently in\nhttps://lore.kernel.org/git/877dhs20x3.fsf@evledraar.gmail.com/; is this\nmethod of doing not-quite-an-fsync guaranteed by some OS's / POSIX etc,\nor is it more like the initial approach before core.fsyncObjectFiles,\ni.e. the happy-go-lucky approach described in the \"[...]that orders data\nwrites properly[...]\" documentation you're removing.\n\n4) While that documentation written by Linus long ago is rather\nflippant, I think just removing it and not replacing it with some\ndiscussion about how this is a trade-off v.s. real-world filesystem\nsemantics isn't a good change.\n\n5) On a similar thought as transfer.unpackLimit in #2, I wonder if this\nfsync() setting shouldn't really be something we should be splitting\nup. I.e. maybe handle batch loose object writes one way, ref updates\nanother way etc. I think moving core.fsync* to a setting like what we\nhave for fsck.* and <cmd>.fsck.* is probably a better thing to do in the\nlonger term.\n\nI.e. being able to do things like:\n\n    fsync.objectFiles = none\n    fsync.refFiles = cache # or \"hardware\"\n    receive.fsync.objectFiles = hardware\n    receive.fsync.refFiles = hardware\n\nOr whatever, i.e. we're using one hammer for all of these now, but I\nsuspect most users who care about fsync care about /some/ fsync, not\neverything.\n\n6) Inline comments below.\n\n> +struct object_rename {\n> +\tchar *src;\n> +\tchar *dst;\n> +};\n> +\n> +static struct bulk_rename_state {\n> +\tstruct object_rename *renames;\n> +\tuint32_t alloc_renames;\n> +\tuint32_t nr_renames;\n> +} bulk_rename_state;\n\nIn a crash of git itself it seems we're going to leave some litter\nbehind in the object dir now, and \"git gc\" won't know how to clean it\nup. I think this is going to want to just use the tmp-objdir.[ch] API,\nwhich might or might not need to be extended for loose objects / some\noddities of this use-case.\n\nAlso, if you have a pair of things like this the string-list API is much\nmore pleasing to use than coming up with your own encapsulation.\n\n>  static struct bulk_checkin_state {\n>  \tunsigned plugged:1;\n>  \n> @@ -21,13 +33,15 @@ static struct bulk_checkin_state {\n>  \tstruct pack_idx_entry **written;\n>  \tuint32_t alloc_written;\n>  \tuint32_t nr_written;\n> -} state;\n\n\n> +\n> +\t\tfree(state->renames);\n> +\t\tmemset(state, 0, sizeof(*state));\n\nSo with this and other use of the \"state\" variable is this part of\nbulk-checkin going to become thread-unsafe, was that already the case?\n\n> +static void add_rename_bulk_checkin(struct bulk_rename_state *state,\n> +\t\t\t\t    const char *src, const char *dst)\n> +{\n> +\tstruct object_rename *rename;\n> +\n> +\tALLOC_GROW(state->renames, state->nr_renames + 1, state->alloc_renames);\n> +\n> +\trename = &state->renames[state->nr_renames++];\n> +\trename->src = xstrdup(src);\n> +\trename->dst = xstrdup(dst);\n> +}\n\nAll boilerplate duplicating things you'd get with a string-list for free...\n\n> +\t\t/*\n> +\t\t * If we have a plugged bulk checkin, we issue a call that\n> +\t\t * cleans the filesystem page cache but avoids a hardware flush\n> +\t\t * command. Later on we will issue a single hardware flush\n> +\t\t * before renaming files as part of do_sync_and_rename.\n> +\t\t */\n\nSo this is the sort of thing I meant by extending Linus's docs, I know\nsome FS's work this way, but not all do.\n\nAlso there's no guarantee in git that your .git is on one FS, so I think\neven for the FS's you have in mind this might not be an absolute\nguarantee...\n"},{"id":"433753","messageId":"CANQDOdcxsR4n9SCtueRBUO4Ea98NBoT17m5srurpvg=CVft51g@mail.gmail.com","threadId":"56371","inReplyTo":"pull.1076.git.git.1629856292.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/2] [RFC] Implement a bulk-checkin option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-08-25T16:58:00Z","receivedAt":"2021-08-25T16:58:13Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Aug 24, 2021 at 6:51 PM Neeraj K. Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n> Hardware:\n>\n>  * Mac - Mac Mini 2018 running MacOS 11.5.1, APFS with a 1TB Apple NMVE SSD,\n>  * Linux - Ubuntu 20.04 - ext4 running on a Hyper-V VM with a fixed VHDX\n>    backed by a Samsung PM981.\n>  * Win - Windows NTFS - Same Hyper-V host as Linux. Operation | Mac | Linux\n>    | Windows\n>\n> ---------------- |---------|-------|---------- git fsync | 40.6 s | 7.8 s |\n> 6.9s git fsync_defer | 6.5 s | 2.1 s | 3.8s git no_fsync | 1.7 s | 1.0 s |\n> 2.6s\n>\nI just wanted to fix this performance test table so that it is readable.\nOperation       | Mac     | Linux | Windows\n----------------|---------|-------|----------\ngit fsync       | 40.6 s  | 7.8 s | 6.9 s\ngit fsync_defer | 6.5 s   | 2.1 s | 3.8 s\ngit no_fsync    | 1.7 s   | 1.0 s | 2.6 s\n\nHere's the graphical version:\nhttps://docs.google.com/spreadsheets/d/18HWXSUVAVqqKATsuVvgxDF6ftX_5qG1UGgNjGtwOuu8/edit?usp=sharing\n"},{"id":"433756","messageId":"CANQDOdf7rMyT4Swriw9=Ei7KN1iLv_dGDWSSck22Zu6AztOyjg@mail.gmail.com","threadId":"56371","inReplyTo":"20210825053839.GA27037@lst.de","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-08-25T17:40:53Z","receivedAt":"2021-08-25T17:41:06Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Aug 24, 2021 at 10:38 PM Christoph Hellwig <hch@lst.de> wrote:\n>\n> On Wed, Aug 25, 2021 at 01:51:32AM +0000, Neeraj Singh via GitGitGadget wrote:\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > When adding many objects to a repo with core.fsyncObjectFiles set to\n> > true, the cost of fsync'ing each object file can become prohibitive.\n> >\n> > One major source of the cost of fsync is the implied flush of the\n> > hardware writeback cache within the disk drive. Fortunately, Windows,\n> > MacOS, and Linux each offer mechanisms to write data from the filesystem\n> > page cache without initiating a hardware flush.\n> >\n> > This patch introduces a new 'core.fsyncObjectFiles = 2' option that\n> > takes advantage of the bulk-checkin infrastructure to batch up hardware\n> > flushes.\n>\n> Another interesting way to flush on linux would be the syncfs call,\n> which syncs all files on a file system.  Once you write more than\n> handful or two of files that tends to win out over a batch of fsync\n> calls.\n\nI'd expect syncfs to suffer from the noisy-neighbor problem that Linus\nalluded to on the big\nthread you kicked off.  The equivalent call on Windows currently\nrequires administrative\nprivileges (I'm really not sure exactly why, perhaps we should change that).\n\nIf someone adds a more targeted bulk sync interface to the Linux\nkernel, I'm sure Git could be\nchanged to use it. Maybe an fcntl(2) interface that initiates\nwriteback and registers completion with an\neventfd.\n"},{"id":"433761","messageId":"nycvar.QRO.7.76.6.2108251551100.55@tvgsbejvaqbjf.bet","threadId":"56371","inReplyTo":"d1e68d4a2afc1d0ba74af64680bea09f412f21cc.1629856293.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2021-08-25T18:52:26Z","receivedAt":"2021-08-25T18:52:42Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Neeraj,\n\ncontinuing my review here, inlined.\n\nOn Wed, 25 Aug 2021, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> When adding many objects to a repo with core.fsyncObjectFiles set to\n> true, the cost of fsync'ing each object file can become prohibitive.\n>\n> One major source of the cost of fsync is the implied flush of the\n> hardware writeback cache within the disk drive. Fortunately, Windows,\n> MacOS, and Linux each offer mechanisms to write data from the filesystem\n> page cache without initiating a hardware flush.\n>\n> This patch introduces a new 'core.fsyncObjectFiles = 2' option that\n> takes advantage of the bulk-checkin infrastructure to batch up hardware\n> flushes.\n\nIt makes sense, but I would recommend using a more easily explained value\nthan `2`. Maybe `delayed`? Or `bulk` or `batched`?\n\nThe way this would be implemented would look somewhat like the\nimplementation for `core.abbrev`, which also accepts a string (\"auto\") or\na Boolean (or even an integral number), see\nhttps://github.com/git/git/blob/v2.33.0/config.c#L1367-L1381:\n\n\tif (!strcmp(var, \"core.abbrev\")) {\n\t\tif (!value)\n\t\t\treturn config_error_nonbool(var);\n\t\tif (!strcasecmp(value, \"auto\"))\n\t\t\tdefault_abbrev = -1;\n\t\telse if (!git_parse_maybe_bool_text(value))\n\t\t\tdefault_abbrev = the_hash_algo->hexsz;\n\t\telse {\n\t\t\tint abbrev = git_config_int(var, value);\n\t\t\tif (abbrev < minimum_abbrev || abbrev > the_hash_algo->hexsz)\n\t\t\t\treturn error(_(\"abbrev length out of range: %d\"), abbrev);\n\t\t\tdefault_abbrev = abbrev;\n\t\t}\n\t\treturn 0;\n\t}\n\n> When the new mode is enabled we do the following for new objects:\n>\n> 1. Create a tmp_obj_XXXX file and write the object data to it.\n> 2. Issue a pagecache writeback request and wait for it to complete.\n> 3. Record the tmp name and the final name in the bulk-checkin state for\n>    later name.\n>\n> At the end of the entire transaction we:\n> 1. Issue a fsync against the lock file to flush the hardware writeback\n>    cache, which should by now have processed the tmp file writes.\n> 2. Rename all of the temp files to their final names.\n> 3. When updating the index and/or refs, we will issue another fsync\n>    internal to that operation.\n>\n> On a filesystem with a singular journal that is updated during name\n> operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n> would expect the fsync to trigger a journal writeout so that this\n> sequence is enough to ensure that the user's data is durable by the time\n> the git command returns.\n>\n> This change also updates the MacOS code to trigger a real hardware flush\n> via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n> MacOS there was no guarantee of durability since a simple fsync(2) call\n> does not flush any hardware caches.\n\nYou included a very nice table with performance numbers in the cover\nletter. Maybe include that here, in the commit message?\n\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  Documentation/config/core.txt |  17 ++++--\n>  Makefile                      |   4 ++\n>  builtin/add.c                 |   3 +-\n>  builtin/update-index.c        |   3 +\n>  bulk-checkin.c                | 105 +++++++++++++++++++++++++++++++---\n>  bulk-checkin.h                |   4 +-\n>  config.c                      |   4 +-\n>  config.mak.uname              |   2 +\n>  configure.ac                  |   8 +++\n>  git-compat-util.h             |   7 +++\n>  object-file.c                 |  12 +---\n>  wrapper.c                     |  36 ++++++++++++\n>  write-or-die.c                |   2 +-\n>  13 files changed, 177 insertions(+), 30 deletions(-)\n>\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index c04f62a54a1..3b672c2db67 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -548,12 +548,17 @@ core.whitespace::\n>    errors. The default tab width is 8. Allowed values are 1 to 63.\n>\n>  core.fsyncObjectFiles::\n> -\tThis boolean will enable 'fsync()' when writing object files.\n> -+\n> -This is a total waste of time and effort on a filesystem that orders\n> -data writes properly, but can be useful for filesystems that do not use\n> -journalling (traditional UNIX filesystems) or that only journal metadata\n> -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n> +\tA boolean value or the number '2', indicating the level of durability\n> +\tapplied to object files.\n> ++\n> +This setting controls how much effort Git makes to ensure that data added to\n> +the object store are durable in the case of an unclean system shutdown. If\n\nIn addition to the content, I also like a lot that this tempers down the\nlanguage to be a lot more agreeable to read.\n\n> +'false', Git allows data to remain in file system caches according to operating\n> +system policy, whence they may be lost if the system loses power or crashes. A\n> +value of 'true' instructs Git to force objects to stable storage immediately\n> +when they are added to the object store. The number '2' is an experimental\n> +value that also preserves durability but tries to perform hardware flushes in a\n> +batch.\n>\n>  core.preloadIndex::\n>  \tEnable parallel index preload for operations like 'git diff'\n> diff --git a/Makefile b/Makefile\n> index 9573190f1d7..cb950ee43d3 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -1896,6 +1896,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n>  \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n>  endif\n>\n> +ifdef HAVE_SYNC_FILE_RANGE\n> +\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n> +endif\n> +\n>  ifdef NEEDS_LIBRT\n>  \tEXTLIBS += -lrt\n>  endif\n> diff --git a/builtin/add.c b/builtin/add.c\n> index 09e684585d9..c58dfcd4bc3 100644\n> --- a/builtin/add.c\n> +++ b/builtin/add.c\n> @@ -670,7 +670,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n>\n>  \tif (chmod_arg && pathspec.nr)\n>  \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n> -\tunplug_bulk_checkin();\n> +\n> +\tunplug_bulk_checkin(&lock_file);\n>\n>  finish:\n>  \tif (write_locked_index(&the_index, &lock_file,\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index f1f16f2de52..64d025cf49e 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -5,6 +5,7 @@\n>   */\n>  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n>  #include \"cache.h\"\n> +#include \"bulk-checkin.h\"\n>  #include \"config.h\"\n>  #include \"lockfile.h\"\n>  #include \"quote.h\"\n> @@ -1152,6 +1153,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t\tstruct strbuf unquoted = STRBUF_INIT;\n>\n>  \t\tsetup_work_tree();\n> +\t\tplug_bulk_checkin();\n>  \t\twhile (getline_fn(&buf, stdin) != EOF) {\n>  \t\t\tchar *p;\n>  \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n> @@ -1166,6 +1168,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t\t\t\tchmod_path(set_executable_bit, p);\n>  \t\t\tfree(p);\n>  \t\t}\n> +\t\tunplug_bulk_checkin(&lock_file);\n>  \t\tstrbuf_release(&unquoted);\n>  \t\tstrbuf_release(&buf);\n>  \t}\n\nThis change to `cmd_update_index()`, would it make sense to separate it\nout into its own commit? I think it would, as it is a slight change of\nbehavior of the `--stdin` mode, no?\n\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index b023d9959aa..71004db863e 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -3,6 +3,7 @@\n>   */\n>  #include \"cache.h\"\n>  #include \"bulk-checkin.h\"\n> +#include \"lockfile.h\"\n>  #include \"repository.h\"\n>  #include \"csum-file.h\"\n>  #include \"pack.h\"\n> @@ -10,6 +11,17 @@\n>  #include \"packfile.h\"\n>  #include \"object-store.h\"\n>\n> +struct object_rename {\n> +\tchar *src;\n> +\tchar *dst;\n> +};\n> +\n> +static struct bulk_rename_state {\n> +\tstruct object_rename *renames;\n> +\tuint32_t alloc_renames;\n> +\tuint32_t nr_renames;\n> +} bulk_rename_state;\n> +\n>  static struct bulk_checkin_state {\n>  \tunsigned plugged:1;\n>\n> @@ -21,13 +33,15 @@ static struct bulk_checkin_state {\n>  \tstruct pack_idx_entry **written;\n>  \tuint32_t alloc_written;\n>  \tuint32_t nr_written;\n> -} state;\n> +\n> +} bulk_checkin_state;\n\nWhile it definitely looks better after this patch, having the new code\n_and_ the rename in the same set of changes makes it a bit harder to\nreview and to spot bugs.\n\nCould I ask you to split this rename out into its own, preparatory patch\n(\"preparatory\" meaning that it should be ordered before the patch that\nadds support for the new fsync mode)?\n\n>\n>  static void finish_bulk_checkin(struct bulk_checkin_state *state)\n>  {\n>  \tstruct object_id oid;\n>  \tstruct strbuf packname = STRBUF_INIT;\n>  \tint i;\n> +\tunsigned old_plugged;\n\nSince this variable is designed to hold the value of the `plugged` field\nof the `bulk_checkin_state`, which is declared as `unsigned plugged:1;`,\nwe probably want a `:1` here, too.\n\nAlso: is it really \"old\", rather than \"orig\"? I would have expected the\nname `orig_plugged` or `save_plugged`.\n\n>\n>  \tif (!state->f)\n>  \t\treturn;\n> @@ -55,13 +69,42 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n>\n>  clear_exit:\n>  \tfree(state->written);\n> +\told_plugged = state->plugged;\n>  \tmemset(state, 0, sizeof(*state));\n> +\tstate->plugged = old_plugged;\n\nUnfortunately, I lack the context to understand the purpose of this. Is\nthe idea that `plugged` gives an indication whether we're still within\nthat batch that should be fsync'ed all at once?\n\nI only see one caller where this would make a difference, and that caller\nis `deflate_to_pack()`. Maybe we should just start that function with\n`unsigned save_plugged:1 = state->plugged;` and restore it after the\n`while (1)` loop?\n\n>\n>  \tstrbuf_release(&packname);\n>  \t/* Make objects we just wrote available to ourselves */\n>  \treprepare_packed_git(the_repository);\n>  }\n>\n> +static void do_sync_and_rename(struct bulk_rename_state *state, struct lock_file *lock_file)\n> +{\n> +\tif (state->nr_renames) {\n> +\t\tint i;\n> +\n> +\t\t/*\n> +\t\t * Issue a full hardware flush against the lock file to ensure\n> +\t\t * that all objects are durable before any renames occur.\n> +\t\t * The code in fsync_and_close_loose_object_bulk_checkin has\n> +\t\t * already ensured that writeout has occurred, but it has not\n> +\t\t * flushed any writeback cache in the storage hardware.\n> +\t\t */\n> +\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n> +\n> +\t\tfor (i = 0; i < state->nr_renames; i++) {\n> +\t\t\tif (finalize_object_file(state->renames[i].src, state->renames[i].dst))\n> +\t\t\t\tdie_errno(_(\"could not rename '%s'\"), state->renames[i].src);\n> +\n> +\t\t\tfree(state->renames[i].src);\n> +\t\t\tfree(state->renames[i].dst);\n> +\t\t}\n> +\n> +\t\tfree(state->renames);\n> +\t\tmemset(state, 0, sizeof(*state));\n\nHmm. There is a lot of `memset()`ing going on, and I am not quite sure\nthat I like what I am seeing. It does not help that there are now two very\neasily-confused structs: `bulk_rename_state` and `bulk_checkin_state`.\nWhich made me worried at first that we might be resetting the `renames`\nfield inadvertently in `finish_bulk_checkin()`.\n\nMaybe we can do this instead?\n\n\t\tFREE_AND_NULL(state->renames);\n\t\tstate->nr_renames = state->alloc_renames = 0;\n\n> +\t}\n> +}\n> +\n>  static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n>  {\n>  \tint i;\n> @@ -256,25 +299,69 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>  \treturn 0;\n>  }\n>\n> +static void add_rename_bulk_checkin(struct bulk_rename_state *state,\n> +\t\t\t\t    const char *src, const char *dst)\n> +{\n> +\tstruct object_rename *rename;\n> +\n> +\tALLOC_GROW(state->renames, state->nr_renames + 1, state->alloc_renames);\n> +\n> +\trename = &state->renames[state->nr_renames++];\n> +\trename->src = xstrdup(src);\n> +\trename->dst = xstrdup(dst);\n> +}\n> +\n> +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n> +\t\t\t\t\t      const char *filename)\n> +{\n> +\tif (fsync_object_files) {\n> +\t\t/*\n> +\t\t * If we have a plugged bulk checkin, we issue a call that\n> +\t\t * cleans the filesystem page cache but avoids a hardware flush\n> +\t\t * command. Later on we will issue a single hardware flush\n> +\t\t * before renaming files as part of do_sync_and_rename.\n> +\t\t */\n> +\t\tif (bulk_checkin_state.plugged &&\n> +\t\t    fsync_object_files == 2 &&\n> +\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n> +\t\t\tadd_rename_bulk_checkin(&bulk_rename_state, tmpfile, filename);\n> +\t\t\tif (close(fd))\n> +\t\t\t\tdie_errno(_(\"error when closing loose object file\"));\n> +\n> +\t\t\treturn 0;\n> +\n> +\t\t} else {\n> +\t\t\tfsync_or_die(fd, \"loose object file\");\n> +\t\t}\n> +\t}\n> +\n> +\tif (close(fd))\n> +\t\tdie_errno(_(\"error when closing loose object file\"));\n> +\n> +\treturn finalize_object_file(tmpfile, filename);\n> +}\n> +\n>  int index_bulk_checkin(struct object_id *oid,\n>  \t\t       int fd, size_t size, enum object_type type,\n>  \t\t       const char *path, unsigned flags)\n>  {\n> -\tint status = deflate_to_pack(&state, oid, fd, size, type,\n> +\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n>  \t\t\t\t     path, flags);\n> -\tif (!state.plugged)\n> -\t\tfinish_bulk_checkin(&state);\n> +\tif (!bulk_checkin_state.plugged)\n> +\t\tfinish_bulk_checkin(&bulk_checkin_state);\n>  \treturn status;\n>  }\n>\n>  void plug_bulk_checkin(void)\n>  {\n> -\tstate.plugged = 1;\n> +\tbulk_checkin_state.plugged = 1;\n>  }\n>\n> -void unplug_bulk_checkin(void)\n> +void unplug_bulk_checkin(struct lock_file *lock_file)\n>  {\n> -\tstate.plugged = 0;\n> -\tif (state.f)\n> -\t\tfinish_bulk_checkin(&state);\n> +\tbulk_checkin_state.plugged = 0;\n> +\tif (bulk_checkin_state.f)\n> +\t\tfinish_bulk_checkin(&bulk_checkin_state);\n> +\n> +\tdo_sync_and_rename(&bulk_rename_state, lock_file);\n>  }\n> diff --git a/bulk-checkin.h b/bulk-checkin.h\n> index b26f3dc3b74..8efb01ed669 100644\n> --- a/bulk-checkin.h\n> +++ b/bulk-checkin.h\n> @@ -6,11 +6,13 @@\n>\n>  #include \"cache.h\"\n>\n> +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile, const char *filename);\n> +\n>  int index_bulk_checkin(struct object_id *oid,\n>  \t\t       int fd, size_t size, enum object_type type,\n>  \t\t       const char *path, unsigned flags);\n>\n>  void plug_bulk_checkin(void);\n> -void unplug_bulk_checkin(void);\n> +void unplug_bulk_checkin(struct lock_file *);\n>\n>  #endif\n> diff --git a/config.c b/config.c\n> index f33abeab851..375bdb24b0a 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1509,7 +1509,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t}\n>\n>  \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> -\t\tfsync_object_files = git_config_bool(var, value);\n> +\t\tint is_bool;\n> +\n> +\t\tfsync_object_files = git_config_bool_or_int(var, value, &is_bool);\n>  \t\treturn 0;\n>  \t}\n>\n> diff --git a/config.mak.uname b/config.mak.uname\n> index 69413fb3dc0..8c07f2265a8 100644\n> --- a/config.mak.uname\n> +++ b/config.mak.uname\n> @@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n>  \tHAVE_CLOCK_MONOTONIC = YesPlease\n>  \t# -lrt is needed for clock_gettime on glibc <= 2.16\n>  \tNEEDS_LIBRT = YesPlease\n> +\tHAVE_SYNC_FILE_RANGE = YesPlease\n>  \tHAVE_GETDELIM = YesPlease\n>  \tSANE_TEXT_GREP=-a\n>  \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n> @@ -133,6 +134,7 @@ ifeq ($(uname_S),Darwin)\n>  \tCOMPAT_OBJS += compat/precompose_utf8.o\n>  \tBASIC_CFLAGS += -DPRECOMPOSE_UNICODE\n>  \tBASIC_CFLAGS += -DPROTECT_HFS_DEFAULT=1\n> +\tBASIC_CFLAGS += -DFSYNC_DOESNT_FLUSH=1\n>  \tHAVE_BSD_SYSCTL = YesPlease\n>  \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n>  \tHAVE_NS_GET_EXECUTABLE_PATH = YesPlease\n> diff --git a/configure.ac b/configure.ac\n> index 031e8d3fee8..c711037d625 100644\n> --- a/configure.ac\n> +++ b/configure.ac\n> @@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n>  \t[AC_MSG_RESULT([no])\n>  \tHAVE_CLOCK_MONOTONIC=])\n>  GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n> +\n> +#\n> +# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n> +GIT_CHECK_FUNC(sync_file_range,\n> +\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n> +\t[HAVE_SYNC_FILE_RANGE])\n> +GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n> +\n>  #\n>  # Define NO_SETITIMER if you don't have setitimer.\n>  GIT_CHECK_FUNC(setitimer,\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index b46605300ab..d14e2436276 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n> @@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n>  void BUG(const char *fmt, ...);\n>  #endif\n>\n> +enum fsync_action {\n> +    FSYNC_WRITEOUT_ONLY,\n> +    FSYNC_HARDWARE_FLUSH\n> +};\n> +\n> +int git_fsync(int fd, enum fsync_action action);\n> +\n>  /*\n>   * Preserves errno, prints a message, but gives no warning for ENOENT.\n>   * Returns 0 on success, which includes trying to unlink an object that does\n> diff --git a/object-file.c b/object-file.c\n> index 607e9e2f80b..5f04143dde0 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1859,16 +1859,6 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  \treturn 0;\n>  }\n>\n> -/* Finalize a file on disk, and close it. */\n> -static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n> -{\n> -\tif (fsync_object_files)\n> -\t\tfsync_or_die(fd, \"loose object file\");\n> -\tif (close(fd) != 0)\n> -\t\tdie_errno(_(\"error when closing loose object file\"));\n> -\treturn finalize_object_file(tmpfile, filename);\n> -}\n> -\n>  /* Size of directory component, including the ending '/' */\n>  static inline int directory_size(const char *filename)\n>  {\n> @@ -1982,7 +1972,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n>  \t}\n>\n> -\treturn close_loose_object(fd, tmp_file.buf, filename.buf);\n> +\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf, filename.buf);\n>  }\n>\n>  static int freshen_loose_object(const struct object_id *oid)\n> diff --git a/wrapper.c b/wrapper.c\n> index 563ad590df1..37a8b61a7df 100644\n> --- a/wrapper.c\n> +++ b/wrapper.c\n> @@ -538,6 +538,42 @@ int xmkstemp_mode(char *filename_template, int mode)\n>  \treturn fd;\n>  }\n>\n> +int git_fsync(int fd, enum fsync_action action)\n> +{\n> +\tif (action == FSYNC_WRITEOUT_ONLY) {\n> +#ifdef __APPLE__\n> +\t\t/*\n> +\t\t * on Mac OS X, fsync just causes filesystem cache writeback but does not\n> +\t\t * flush hardware caches.\n> +\t\t */\n> +\t\treturn fsync(fd);\n> +#endif\n> +\n> +#ifdef HAVE_SYNC_FILE_RANGE\n> +\t\t/*\n> +\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n> +\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n> +\t\t * indicates writeout of the entire file and the wait flags ensure that all\n> +\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n> +\t\t * before we continue.\n> +\t\t */\n> +\n> +\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n> +#endif\n> +\n> +\t\terrno = ENOSYS;\n> +\t\treturn -1;\n> +\t}\n\nHmm. I wonder whether we can do this more consistently with how Git\nusually does platform-specific things.\n\nIn the 3rd patch, the one where you implemented Windows-specific support,\nin the Git for Windows PR at\nhttps://github.com/git-for-windows/git/pull/3391, you introduce a\n`mingw_fsync_no_flush()` function and define `fsync_no_flush` to expand to\nthat function name.\n\nThis is very similar to how Git does things. Take for example the\n`offset_1st_component` macro:\nhttps://github.com/git/git/blob/v2.33.0/git-compat-util.h#L386-L392\n\nUnless defined in a platform-specific manner, it is defined in\n`git-compat-util.h`:\n\n\t#ifndef offset_1st_component\n\tstatic inline int git_offset_1st_component(const char *path)\n\t{\n\t\treturn is_dir_sep(path[0]);\n\t}\n\t#define offset_1st_component git_offset_1st_component\n\t#endif\n\nAnd on Windows, it is defined as following\n(https://github.com/git/git/blob/v2.33.0/compat/win32/path-utils.h#L34-L35),\nbefore the lines quoted above:\n\n\tint win32_offset_1st_component(const char *path);\n\t#define offset_1st_component win32_offset_1st_component\n\nWe could do the exact same thing here. Define a platform-specific\n`mingw_fsync_no_flush()` in `compat/mingw.h` and define the macro\n`fsync_no_flush` to point to it. In `git-compat-util.h`, in the\n`__APPLE__`-specific part, implement it via `fsync()`. And later, in the\nplatform-independent part, _iff_ the macro has not yet been defined,\nimplement an inline function that does that `HAVE_SYNC_FILE_RANGE` dance\nand falls back to `ENOSYS`.\n\nThat would contain the platform-specific `#ifdef` blocks to\n`git-compat-util.h`, which is exactly where we want them.\n\n> +\n> +#ifdef __APPLE__\n> +\treturn fcntl(fd, F_FULLFSYNC);\n> +#else\n> +\treturn fsync(fd);\n> +#endif\n\nSame thing here. We would probably want something like `fsync_with_flush`\nhere.\n\nIt is my hope that you find my comments and suggestions helpful.\n\nThank you,\nJohannes\n\n> +}\n> +\n>  static int warn_if_unremovable(const char *op, const char *file, int rc)\n>  {\n>  \tint err;\n> diff --git a/write-or-die.c b/write-or-die.c\n> index d33e68f6abb..8f53953d4ab 100644\n> --- a/write-or-die.c\n> +++ b/write-or-die.c\n> @@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n>\n>  void fsync_or_die(int fd, const char *msg)\n>  {\n> -\twhile (fsync(fd) < 0) {\n> +\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n>  \t\tif (errno != EINTR)\n>  \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n>  \t}\n> --\n> gitgitgadget\n>\n"},{"id":"433769","messageId":"xmqqv93tgqqu.fsf@gitster.g","threadId":"56371","inReplyTo":"nycvar.QRO.7.76.6.2108251551100.55@tvgsbejvaqbjf.bet","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-08-25T21:26:49Z","receivedAt":"2021-08-25T21:26:52Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> It makes sense, but I would recommend using a more easily explained value\n> than `2`. Maybe `delayed`? Or `bulk` or `batched`?\n\nWhile we have less than 100% confidence in the implementation, it\nmay make sense to have such a knob to choose between \"do we fsync\nthe old, known-safe but slow way, or do we fsync in batch\"\nbehaviours, and I agree that the knob should not be called cryptic\n\"2\".\n\nBut in a distant future when this new way of flushing proves to be\nstable, it would make sense if the enw behaviour were triggered by\nthe plain vanilla 'true', no?  In a sense, running fsync in a batch\n(or using syncfs) is an implementation detail of \"we sync after\nwriting out object files and before declaring success\".\n\nThanks.\n"},{"id":"433771","messageId":"CANQDOdfXBNABCgVsiqc2TJpctEgsJ2rJKHC5=u=q2P+mJWDYxQ@mail.gmail.com","threadId":"56371","inReplyTo":"nycvar.QRO.7.76.6.2108251530210.55@tvgsbejvaqbjf.bet","subject":"Re: [PATCH 1/2] object-file: use futimes rather than utime","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-08-25T22:08:25Z","receivedAt":"2021-08-25T22:08:39Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Aug 25, 2021 at 6:51 AM Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> If I were you, I would spell out \"file descriptor\" here.\nWill do.\n\n> > diff --git a/compat/mingw.c b/compat/mingw.c\n> > index 9e0cd1e097f..948f4c3428b 100644\n> > --- a/compat/mingw.c\n> > +++ b/compat/mingw.c\n> > @@ -949,19 +949,40 @@ int mingw_fstat(int fd, struct stat *buf)\n> >       }\n> >  }\n> >\n> > -static inline void time_t_to_filetime(time_t t, FILETIME *ft)\n> > +static inline void timeval_to_filetime(const struct timeval *t, FILETIME *ft)\n> >  {\n> > -     long long winTime = t * 10000000LL + 116444736000000000LL;\n> > +     long long winTime = t->tv_sec * 10000000LL + t->tv_usec * 10 + 116444736000000000LL;\n>\n> Technically, this is a change in behavior, right? We did not use to use\n> nanosecond precision. But I don't think that we actually make use of this\n> in this patch.\n>\n> >       ft->dwLowDateTime = winTime;\n> >       ft->dwHighDateTime = winTime >> 32;\n> >  }\n> >\n> > -int mingw_utime (const char *file_name, const struct utimbuf *times)\n> > +int mingw_futimes(int fd, const struct timeval times[2])\n>\n> At first, I wondered whether it would make sense to pass the access time\n> and the modified time separately, as pointers. I don't think that we pass\n> around arrays as function parameters in Git anywhere else.\n>\n> But then I realized that `futimes()` is available in this precise form on\n> Linux and on the BSDs. Therefore, it is not up to us to decide the\n> function's signature.\n>\n> However, now that I looked at the manual page, I noticed that this\n> function is not part of any POSIX standard.\n>\n> Which makes me think that we will have to do a bit more than just define\n> it on Windows: we will have to introduce a `Makefile` knob (just like you\n> did with `HAVE_SYNC_FILE_RANGE` in patch 2/2) and set that specifically\n> for Linux and the BSDs, and use `futimes()` only if it is available\n> (otherwise fall back to `utime()`).\n>\n> Then, as a separate patch, we should introduce this Windows-specific shim\n> and declare that it is available via `config.mak.uname`.\n\nThanks for taking another look at the man pages. I looked again too and saw\nthat futimens is part of POSIX.1-2008:\nhttps://pubs.opengroup.org/onlinepubs/9699919799/functions/futimens.html.\nIf I switch to futimens and implement the Windows shim at the same\ntime, is that sufficient to\naddress your feedback? I'd rather not ifdef this one since the\ncodeflow is quite different\ndepending on the presence of the API.\n\n> > +}\n> > +\n> > +int mingw_utime (const char *file_name, const struct utimbuf *times)\n>\n> Please lose the space between the function name and the opening\n> parenthesis. I know, the preimage of this diff has it, but that was an\n> oversight and definitely disagrees with our current coding style.\nWill do.\n\n>\n> > +{\n> >       int fh, rc;\n> >       DWORD attrs;\n> >       wchar_t wfilename[MAX_PATH];\n> > +     struct timeval tvs[2];\n> > +\n> >       if (xutftowcs_path(wfilename, file_name) < 0)\n> >               return -1;\n> >\n> > @@ -979,17 +1000,12 @@ int mingw_utime (const char *file_name, const struct utimbuf *times)\n> >       }\n> >\n> >       if (times) {\n> > -             time_t_to_filetime(times->modtime, &mft);\n> > -             time_t_to_filetime(times->actime, &aft);\n> > -     } else {\n> > -             GetSystemTimeAsFileTime(&mft);\n> > -             aft = mft;\n> > +             memset(tvs, 0, sizeof(tvs));\n> > +             tvs[0].tv_sec = times->actime;\n> > +             tvs[1].tv_sec = times->modtime;\n>\n> It is too bad that we have to copy around those values just to convert\n> them, but I cannot think of any better way, either. And it's not like\n> we're in a hot loop: this code will be dominated by I/O anyways.\n\nYeah, the cost of this is approximately 3-4 cycles (load-to-use\nlatency), so no one will notice relative to the system call overhead.\n\n> > diff --git a/object-file.c b/object-file.c\n> > index a8be8994814..607e9e2f80b 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1860,12 +1860,13 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n> >  }\n> >\n> >  /* Finalize a file on disk, and close it. */\n> > -static void close_loose_object(int fd)\n> > +static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n> >  {\n> >       if (fsync_object_files)\n> >               fsync_or_die(fd, \"loose object file\");\n> >       if (close(fd) != 0)\n> >               die_errno(_(\"error when closing loose object file\"));\n> > +     return finalize_object_file(tmpfile, filename);\n>\n> While this is a clear change of behavior, this function has only one\n> caller, and that caller is adjusted accordingly.\n>\n> Could you add this clarification of context to the commit message? I know\n> it will help me in the future, when I have to get up to speed again by\n> reading the commit history.\n\nHow does the following revised wording sound?\n```\n    object-file: use futimens rather than utime\n\n    Make close_loose_object do all of the steps for syncing and correctly\n    naming a new loose object so that it can be reimplemented in the\n    upcoming bulk-fsync mode.\n\n    Use futimens, which is available in POSIX.1-2008 to update the file\n    timestamps. This should be slightly faster than utime, since we have\n    a file descriptor already available. This change allows us to update\n    the time before closing, renaming, and potentially fsyincing the file\n    being refreshed. This code is currently only invoked by git-pack-objects\n    via force_object_loose.\n\n    Implement a futimens shim for the Windows port of Git.\n\n    Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n```\n\nThanks for the detailed and quick feedback!\n-Neeraj\n"},{"id":"433795","messageId":"CANQDOdd2FDNXnXLdm2FSmxUTk3oi+mQtiW2rf3YG7MJayrexPQ@mail.gmail.com","threadId":"56371","inReplyTo":"87mtp5cwpn.fsf@evledraar.gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-08-26T00:49:45Z","receivedAt":"2021-08-26T00:49:59Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Aug 25, 2021 at 9:31 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> On Wed, Aug 25 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > When adding many objects to a repo with core.fsyncObjectFiles set to\n> > true, the cost of fsync'ing each object file can become prohibitive.\n> >\n> > One major source of the cost of fsync is the implied flush of the\n> > hardware writeback cache within the disk drive. Fortunately, Windows,\n> > MacOS, and Linux each offer mechanisms to write data from the filesystem\n> > page cache without initiating a hardware flush.\n> >\n> > This patch introduces a new 'core.fsyncObjectFiles = 2' option that\n> > takes advantage of the bulk-checkin infrastructure to batch up hardware\n> > flushes.\n> >\n> > When the new mode is enabled we do the following for new objects:\n> >\n> > 1. Create a tmp_obj_XXXX file and write the object data to it.\n> > 2. Issue a pagecache writeback request and wait for it to complete.\n> > 3. Record the tmp name and the final name in the bulk-checkin state for\n> >    later name.\n> >\n> > At the end of the entire transaction we:\n> > 1. Issue a fsync against the lock file to flush the hardware writeback\n> >    cache, which should by now have processed the tmp file writes.\n> > 2. Rename all of the temp files to their final names.\n> > 3. When updating the index and/or refs, we will issue another fsync\n> >    internal to that operation.\n> >\n> > On a filesystem with a singular journal that is updated during name\n> > operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n> > would expect the fsync to trigger a journal writeout so that this\n> > sequence is enough to ensure that the user's data is durable by the time\n> > the git command returns.\n> >\n> > This change also updates the MacOS code to trigger a real hardware flush\n> > via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n> > MacOS there was no guarantee of durability since a simple fsync(2) call\n> > does not flush any hardware caches.\n>\n> Thanks for working on this, good to see fsck issues picked up after some\n> on-list pause.\n>\n> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> > index c04f62a54a1..3b672c2db67 100644\n> > --- a/Documentation/config/core.txt\n> > +++ b/Documentation/config/core.txt\n> > @@ -548,12 +548,17 @@ core.whitespace::\n> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >\n> >  core.fsyncObjectFiles::\n> > -     This boolean will enable 'fsync()' when writing object files.\n> > -+\n> > -This is a total waste of time and effort on a filesystem that orders\n> > -data writes properly, but can be useful for filesystems that do not use\n> > -journalling (traditional UNIX filesystems) or that only journal metadata\n> > -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n> > +     A boolean value or the number '2', indicating the level of durability\n> > +     applied to object files.\n> > ++\n> > +This setting controls how much effort Git makes to ensure that data added to\n> > +the object store are durable in the case of an unclean system shutdown. If\n> > +'false', Git allows data to remain in file system caches according to operating\n> > +system policy, whence they may be lost if the system loses power or crashes. A\n> > +value of 'true' instructs Git to force objects to stable storage immediately\n> > +when they are added to the object store. The number '2' is an experimental\n> > +value that also preserves durability but tries to perform hardware flushes in a\n> > +batch.\n>\n> Some feedback/thoughts:\n>\n> 0) Let's not expose \"2\" to users, but give it some friendly config name\n> and just translate this to the enum internally.\n\nAgreed. I'll follow suggestions from here and elsewhere to make this a\nhuman-readable\nstring. Is \"core.fsyncObjectFiles=batch\" acceptable?\n\n>\n> 1) Your commit message says \"When updating the index and/or refs[...]\"\n> but we're changing core.fsyncObjectFiles here, I assume that's\n> summarizing existing behavior then\nThat's what I intended. I'll make the patch description more explicit\nabout that.\nIn general, this patch is only concerned with loose object files. It's\nmy assumption\nthat other parts of the system (like the refs db) need to perform their own data\nconsistency and must have their own durability.\n\nI do think that long-term, the Git community should think about having\na general transaction\nmechanism with redo logging to have a consistent method for achieving\ndurability.\n\n>\n> 2) You say \"when adding many [loose] objects to a repo[...]\", and the\n> test-case is \"git stash push\", but for e.g. accepting pushes we have\n> transfer.unpackLimit.\n>\n> It would be interesting to see if/how this impacts performance there,\n> and also if not that should at least be called out in\n> documentation. I.e. users might want to set this differently on servers\n> v.s. checkouts.\n>\n> But also, is this sort of thing something we could mitigate even more in\n> commands like \"git stash push\" by just writing a pack instead of N loose\n> objects?\n>\n> I don't think such questions should preclude changing the fsync\n> approach, or offering more options, but they should inform our\n> longer-term goals.\n\nDealing only/mostly in packfiles would be a great approach. I'd hope that\nthis fsyncing work would mostly be superseded if such a change is rolled out.\nI just read about the geometric repacking stuff, and it looks reminiscent of the\nLSM-tree approach to databases.\n\n>\n> 3) Re some of the musings about fsync() recently in\n> https://lore.kernel.org/git/877dhs20x3.fsf@evledraar.gmail.com/; is this\n> method of doing not-quite-an-fsync guaranteed by some OS's / POSIX etc,\n> or is it more like the initial approach before core.fsyncObjectFiles,\n> i.e. the happy-go-lucky approach described in the \"[...]that orders data\n> writes properly[...]\" documentation you're removing.\n\nI am confident about the validity of the batched approach on Windows when\nrunning on NTFS and ReFS, given my background as a Windows FS developer.\nWe are unlikely to change any mainstream data consistency behavior of\nour filesystems\nto be weaker, given the type of errors such changes would cause.\n\nIn Windows, we call the FS requirements to support this change\n\"external metadata consistency\",\nwhich states that all metadata operations that could have led to a\nstate later FSYNCed must be\nvisible after the fsync.\n\nmacOS's fsync documentation indicates that they are likely to\nimplement the required guarantees.\nThey specifically say that fsync triggers writeback or data and\nmetadata, but does not issue a hardware\nflush.  Please see the doc at\nhttps://developer.apple.com/library/archive/documentation/System/Conceptual/ManPages_iPhoneOS/man2/fsync.2.html.\nIt is also notable that Apple SSDs have particularly bad performance\nfor flush operations (we noticed this\nwhen booting Windows through BootCamp as well).\n\nUnfortunately my perusal of the man pages and documentation I could find doesn't\ngive me this level of confidence on typical Linux filesystems. For\ninstance, the notion of having to\nfsync the parent directory in order to render an inode's link findable\neliminates a lot of the\nadvantage of this change, though we could batch those and would have\nto do at most 256.\n\nThis thread is somewhat instructive, but inconclusive:\nhttps://lwn.net/ml/linux-fsdevel/1552418820-18102-1-git-send-email-jaya@cs.utexas.edu/.\nOne conclusion from reviewing that thread is that as of then,\nsync_file_ranges isn't actually enough\nto make a hard guarantee about writeout occurring. See\nhttps://lore.kernel.org/linux-fsdevel/20190319204330.GY26298@dastard/.\nMy hope is that the Linux FS developers have rectified that shortcoming by now.\n\n>\n> 4) While that documentation written by Linus long ago is rather\n> flippant, I think just removing it and not replacing it with some\n> discussion about how this is a trade-off v.s. real-world filesystem\n> semantics isn't a good change.\n\nI don't think the replaced documentation applies anymore to an ext4 or xfs\nsystem with delayed allocation. Those filesystems effectively have\ndata=writeback\nsemantics because they don't need data ordering to avoid exposing unwritten\ndata, and so don't write the data at any particular syscall boundary\nor with any particular\nordering.\n\nI think my updated version of the documentation for \"= false\" is\naccurate and more helpful\nfrom a user perspective (\"up to OS policy when your data becomes durable in\nthe event of an unclean shutdown\").  \"= true\" also has a reasonable\ndescription, though I\nmight add some verbiage indicating that this setting could be costly.\n\nI'll take a crack at improving the batched mode documentation.\n\n>\n> 5) On a similar thought as transfer.unpackLimit in #2, I wonder if this\n> fsync() setting shouldn't really be something we should be splitting\n> up. I.e. maybe handle batch loose object writes one way, ref updates\n> another way etc. I think moving core.fsync* to a setting like what we\n> have for fsck.* and <cmd>.fsck.* is probably a better thing to do in the\n> longer term.\n>\n> I.e. being able to do things like:\n>\n>     fsync.objectFiles = none\n>     fsync.refFiles = cache # or \"hardware\"\n>     receive.fsync.objectFiles = hardware\n>     receive.fsync.refFiles = hardware\n>\n> Or whatever, i.e. we're using one hammer for all of these now, but I\n> suspect most users who care about fsync care about /some/ fsync, not\n> everything.\n>\n\nI disagree. I believe Git should offer a consolidated config setting\nwith two overall goals:\n\n1) A consistent high-integrity setting across the entire git\nindex/object/ref state, primarily\nfor people using a repo for active development of changes. This should\nroughly guarantee\nthat when a git command that adds data to the repo completes, the data\nis durable within git,\nincluding the refs needed to find it.\n\n2) A lower-integrity setting useful for build/CI, maintainers who are\napplying lots of patches, etc,\nwhere it is expected that the data is available elsewhere and can be\nrecovered easily.\n\n> 6) Inline comments below.\n>\n> > +struct object_rename {\n> > +     char *src;\n> > +     char *dst;\n> > +};\n> > +\n> > +static struct bulk_rename_state {\n> > +     struct object_rename *renames;\n> > +     uint32_t alloc_renames;\n> > +     uint32_t nr_renames;\n> > +} bulk_rename_state;\n>\n> In a crash of git itself it seems we're going to leave some litter\n> behind in the object dir now, and \"git gc\" won't know how to clean it\n> up. I think this is going to want to just use the tmp-objdir.[ch] API,\n> which might or might not need to be extended for loose objects / some\n> oddities of this use-case.\n>\n\nIt appears that \"git prune\" would take care of these files.\n\n> Also, if you have a pair of things like this the string-list API is much\n> more pleasing to use than coming up with your own encapsulation.\n\nThanks, I'll update to use the string-list API.\n\n> >  static struct bulk_checkin_state {\n> >       unsigned plugged:1;\n> >\n> > @@ -21,13 +33,15 @@ static struct bulk_checkin_state {\n> >       struct pack_idx_entry **written;\n> >       uint32_t alloc_written;\n> >       uint32_t nr_written;\n> > -} state;\n>\n>\n> > +\n> > +             free(state->renames);\n> > +             memset(state, 0, sizeof(*state));\n>\n> So with this and other use of the \"state\" variable is this part of\n> bulk-checkin going to become thread-unsafe, was that already the case?\n\nYes, this code was already thread-unsafe if we're in the \"bulk checkin\nplugged\" mode and\nthat hasn't changed. Is it worth fixing this right now? Is there a\npreexisting example of code\nthat uses thread-local-storage inside a library function and then\nmerges the state later? Bonus\npoints for only doing the thread-local stuff if alternate threads are\nactually active.\n\n> > +static void add_rename_bulk_checkin(struct bulk_rename_state *state,\n> > +                                 const char *src, const char *dst)\n> > +{\n> > +     struct object_rename *rename;\n> > +\n> > +     ALLOC_GROW(state->renames, state->nr_renames + 1, state->alloc_renames);\n> > +\n> > +     rename = &state->renames[state->nr_renames++];\n> > +     rename->src = xstrdup(src);\n> > +     rename->dst = xstrdup(dst);\n> > +}\n>\n> All boilerplate duplicating things you'd get with a string-list for free...\n\nWill fix.Thanks.\n\n>\n> > +             /*\n> > +              * If we have a plugged bulk checkin, we issue a call that\n> > +              * cleans the filesystem page cache but avoids a hardware flush\n> > +              * command. Later on we will issue a single hardware flush\n> > +              * before renaming files as part of do_sync_and_rename.\n> > +              */\n>\n> So this is the sort of thing I meant by extending Linus's docs, I know\n> some FS's work this way, but not all do.\n>\n> Also there's no guarantee in git that your .git is on one FS, so I think\n> even for the FS's you have in mind this might not be an absolute\n> guarantee...\n\nThis is unfortunately an issue that isn't resolvable within Git. I think there's\nvalue in supporting batched mode for the common systems and filesystems\nthat do support the guarantee.  I'd like to be able to set batched mode as the\ndefault on Windows eventually.  It also looks like it can be default on macOS.\n\nI think linux might be able to get to the desired semantics for default distro\nfilesystems with minimal changes. Perhaps this new mode in Git would provide\nthem with a motivation to do so.\n"},{"id":"433796","messageId":"CANQDOdfxMx-GZE3OOZny8VtvDtvJuuid=DZvR4Tj8Ltev0bYnA@mail.gmail.com","threadId":"56371","inReplyTo":"nycvar.QRO.7.76.6.2108251551100.55@tvgsbejvaqbjf.bet","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-08-26T01:19:20Z","receivedAt":"2021-08-26T01:19:35Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Aug 25, 2021 at 11:52 AM Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> On Wed, 25 Aug 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > When adding many objects to a repo with core.fsyncObjectFiles set to\n> > true, the cost of fsync'ing each object file can become prohibitive.\n> >\n> > One major source of the cost of fsync is the implied flush of the\n> > hardware writeback cache within the disk drive. Fortunately, Windows,\n> > MacOS, and Linux each offer mechanisms to write data from the filesystem\n> > page cache without initiating a hardware flush.\n> >\n> > This patch introduces a new 'core.fsyncObjectFiles = 2' option that\n> > takes advantage of the bulk-checkin infrastructure to batch up hardware\n> > flushes.\n>\n> It makes sense, but I would recommend using a more easily explained value\n> than `2`. Maybe `delayed`? Or `bulk` or `batched`?\n>\n> The way this would be implemented would look somewhat like the\n> implementation for `core.abbrev`, which also accepts a string (\"auto\") or\n> a Boolean (or even an integral number), see\n> https://github.com/git/git/blob/v2.33.0/config.c#L1367-L1381:\n>\n>         if (!strcmp(var, \"core.abbrev\")) {\n>                 if (!value)\n>                         return config_error_nonbool(var);\n>                 if (!strcasecmp(value, \"auto\"))\n>                         default_abbrev = -1;\n>                 else if (!git_parse_maybe_bool_text(value))\n>                         default_abbrev = the_hash_algo->hexsz;\n>                 else {\n>                         int abbrev = git_config_int(var, value);\n>                         if (abbrev < minimum_abbrev || abbrev > the_hash_algo->hexsz)\n>                                 return error(_(\"abbrev length out of range: %d\"), abbrev);\n>                         default_abbrev = abbrev;\n>                 }\n>                 return 0;\n>         }\n>\n\nThanks for the code example. I'll follow something similar. I'll prefer the name\n\"batch\" for the new fsync mode.\n\n> > When the new mode is enabled we do the following for new objects:\n> >\n> > 1. Create a tmp_obj_XXXX file and write the object data to it.\n> > 2. Issue a pagecache writeback request and wait for it to complete.\n> > 3. Record the tmp name and the final name in the bulk-checkin state for\n> >    later name.\n> >\n> > At the end of the entire transaction we:\n> > 1. Issue a fsync against the lock file to flush the hardware writeback\n> >    cache, which should by now have processed the tmp file writes.\n> > 2. Rename all of the temp files to their final names.\n> > 3. When updating the index and/or refs, we will issue another fsync\n> >    internal to that operation.\n> >\n> > On a filesystem with a singular journal that is updated during name\n> > operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n> > would expect the fsync to trigger a journal writeout so that this\n> > sequence is enough to ensure that the user's data is durable by the time\n> > the git command returns.\n> >\n> > This change also updates the MacOS code to trigger a real hardware flush\n> > via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n> > MacOS there was no guarantee of durability since a simple fsync(2) call\n> > does not flush any hardware caches.\n>\n> You included a very nice table with performance numbers in the cover\n> letter. Maybe include that here, in the commit message?\n\nWill do.\n\n>\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  Documentation/config/core.txt |  17 ++++--\n> >  Makefile                      |   4 ++\n> >  builtin/add.c                 |   3 +-\n> >  builtin/update-index.c        |   3 +\n> >  bulk-checkin.c                | 105 +++++++++++++++++++++++++++++++---\n> >  bulk-checkin.h                |   4 +-\n> >  config.c                      |   4 +-\n> >  config.mak.uname              |   2 +\n> >  configure.ac                  |   8 +++\n> >  git-compat-util.h             |   7 +++\n> >  object-file.c                 |  12 +---\n> >  wrapper.c                     |  36 ++++++++++++\n> >  write-or-die.c                |   2 +-\n> >  13 files changed, 177 insertions(+), 30 deletions(-)\n> >\n> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> > index c04f62a54a1..3b672c2db67 100644\n> > --- a/Documentation/config/core.txt\n> > +++ b/Documentation/config/core.txt\n> > @@ -548,12 +548,17 @@ core.whitespace::\n> >    errors. The default tab width is 8. Allowed values are 1 to 63.\n> >\n> >  core.fsyncObjectFiles::\n> > -     This boolean will enable 'fsync()' when writing object files.\n> > -+\n> > -This is a total waste of time and effort on a filesystem that orders\n> > -data writes properly, but can be useful for filesystems that do not use\n> > -journalling (traditional UNIX filesystems) or that only journal metadata\n> > -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n> > +     A boolean value or the number '2', indicating the level of durability\n> > +     applied to object files.\n> > ++\n> > +This setting controls how much effort Git makes to ensure that data added to\n> > +the object store are durable in the case of an unclean system shutdown. If\n>\n> In addition to the content, I also like a lot that this tempers down the\n> language to be a lot more agreeable to read.\n>\n> > +'false', Git allows data to remain in file system caches according to operating\n> > +system policy, whence they may be lost if the system loses power or crashes. A\n> > +value of 'true' instructs Git to force objects to stable storage immediately\n> > +when they are added to the object store. The number '2' is an experimental\n> > +value that also preserves durability but tries to perform hardware flushes in a\n> > +batch.\n\nI'll be revising this text a little bit to respond to avarab's\nfeedback.  I'm guessing\nthat this will need a few more rounds of tweaking.\n\n> >\n> >  core.preloadIndex::\n> >       Enable parallel index preload for operations like 'git diff'\n> > diff --git a/Makefile b/Makefile\n> > index 9573190f1d7..cb950ee43d3 100644\n> > --- a/Makefile\n> > +++ b/Makefile\n> > @@ -1896,6 +1896,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n> >       BASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n> >  endif\n> >\n> > +ifdef HAVE_SYNC_FILE_RANGE\n> > +     BASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n> > +endif\n> > +\n> >  ifdef NEEDS_LIBRT\n> >       EXTLIBS += -lrt\n> >  endif\n> > diff --git a/builtin/add.c b/builtin/add.c\n> > index 09e684585d9..c58dfcd4bc3 100644\n> > --- a/builtin/add.c\n> > +++ b/builtin/add.c\n> > @@ -670,7 +670,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n> >\n> >       if (chmod_arg && pathspec.nr)\n> >               exit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n> > -     unplug_bulk_checkin();\n> > +\n> > +     unplug_bulk_checkin(&lock_file);\n> >\n> >  finish:\n> >       if (write_locked_index(&the_index, &lock_file,\n> > diff --git a/builtin/update-index.c b/builtin/update-index.c\n> > index f1f16f2de52..64d025cf49e 100644\n> > --- a/builtin/update-index.c\n> > +++ b/builtin/update-index.c\n> > @@ -5,6 +5,7 @@\n> >   */\n> >  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n> >  #include \"cache.h\"\n> > +#include \"bulk-checkin.h\"\n> >  #include \"config.h\"\n> >  #include \"lockfile.h\"\n> >  #include \"quote.h\"\n> > @@ -1152,6 +1153,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >               struct strbuf unquoted = STRBUF_INIT;\n> >\n> >               setup_work_tree();\n> > +             plug_bulk_checkin();\n> >               while (getline_fn(&buf, stdin) != EOF) {\n> >                       char *p;\n> >                       if (!nul_term_line && buf.buf[0] == '\"') {\n> > @@ -1166,6 +1168,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >                               chmod_path(set_executable_bit, p);\n> >                       free(p);\n> >               }\n> > +             unplug_bulk_checkin(&lock_file);\n> >               strbuf_release(&unquoted);\n> >               strbuf_release(&buf);\n> >       }\n>\n> This change to `cmd_update_index()`, would it make sense to separate it\n> out into its own commit? I think it would, as it is a slight change of\n> behavior of the `--stdin` mode, no?\n\nWill do.\n\nThis makes me think that there is some risk here if someone launches an\nupdate-index and then attempts to use the index to find a newly-added\nOID without\ncompleting the update-index invocation.  It's possible to construct\nsuch a scenario\nif someone's using the \"--verbose\" scenario to find out about actions\ntaken by the\nsubprocess.  Do you think I should add a flag to conditionally enable\nbulk-checkin to\navoid regression for this (I expect unlikely) case?\n\n\n> > diff --git a/bulk-checkin.c b/bulk-checkin.c\n> > index b023d9959aa..71004db863e 100644\n> > --- a/bulk-checkin.c\n> > +++ b/bulk-checkin.c\n> > @@ -3,6 +3,7 @@\n> >   */\n> >  #include \"cache.h\"\n> >  #include \"bulk-checkin.h\"\n> > +#include \"lockfile.h\"\n> >  #include \"repository.h\"\n> >  #include \"csum-file.h\"\n> >  #include \"pack.h\"\n> > @@ -10,6 +11,17 @@\n> >  #include \"packfile.h\"\n> >  #include \"object-store.h\"\n> >\n> > +struct object_rename {\n> > +     char *src;\n> > +     char *dst;\n> > +};\n> > +\n> > +static struct bulk_rename_state {\n> > +     struct object_rename *renames;\n> > +     uint32_t alloc_renames;\n> > +     uint32_t nr_renames;\n> > +} bulk_rename_state;\n> > +\n> >  static struct bulk_checkin_state {\n> >       unsigned plugged:1;\n> >\n> > @@ -21,13 +33,15 @@ static struct bulk_checkin_state {\n> >       struct pack_idx_entry **written;\n> >       uint32_t alloc_written;\n> >       uint32_t nr_written;\n> > -} state;\n> > +\n> > +} bulk_checkin_state;\n>\n> While it definitely looks better after this patch, having the new code\n> _and_ the rename in the same set of changes makes it a bit harder to\n> review and to spot bugs.\n>\n> Could I ask you to split this rename out into its own, preparatory patch\n> (\"preparatory\" meaning that it should be ordered before the patch that\n> adds support for the new fsync mode)?\n\nWill do.  I'm going to change a few things as I'll mention below.\n\n>\n> >\n> >  static void finish_bulk_checkin(struct bulk_checkin_state *state)\n> >  {\n> >       struct object_id oid;\n> >       struct strbuf packname = STRBUF_INIT;\n> >       int i;\n> > +     unsigned old_plugged;\n>\n> Since this variable is designed to hold the value of the `plugged` field\n> of the `bulk_checkin_state`, which is declared as `unsigned plugged:1;`,\n> we probably want a `:1` here, too.\n>\n> Also: is it really \"old\", rather than \"orig\"? I would have expected the\n> name `orig_plugged` or `save_plugged`.\n>\n> >\n> >       if (!state->f)\n> >               return;\n> > @@ -55,13 +69,42 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n> >\n> >  clear_exit:\n> >       free(state->written);\n> > +     old_plugged = state->plugged;\n> >       memset(state, 0, sizeof(*state));\n> > +     state->plugged = old_plugged;\n>\n> Unfortunately, I lack the context to understand the purpose of this. Is\n> the idea that `plugged` gives an indication whether we're still within\n> that batch that should be fsync'ed all at once?\n>\n> I only see one caller where this would make a difference, and that caller\n> is `deflate_to_pack()`. Maybe we should just start that function with\n> `unsigned save_plugged:1 = state->plugged;` and restore it after the\n> `while (1)` loop?\n>\n\nThe problem being solved here is that in some (rare) circumstances it looks like\nwe could unplug the state before an explicit unplug call.  That would\nbe fatal for the\nbulk-fsync code, and is probably suboptimal for the bulk checkin code.\nI'm going to separate\nthis out the boolean to a different variable in my preparatory patch\nso that this fragility goes away.\n\n> >\n> >       strbuf_release(&packname);\n> >       /* Make objects we just wrote available to ourselves */\n> >       reprepare_packed_git(the_repository);\n> >  }\n> >\n> > +static void do_sync_and_rename(struct bulk_rename_state *state, struct lock_file *lock_file)\n> > +{\n> > +     if (state->nr_renames) {\n> > +             int i;\n> > +\n> > +             /*\n> > +              * Issue a full hardware flush against the lock file to ensure\n> > +              * that all objects are durable before any renames occur.\n> > +              * The code in fsync_and_close_loose_object_bulk_checkin has\n> > +              * already ensured that writeout has occurred, but it has not\n> > +              * flushed any writeback cache in the storage hardware.\n> > +              */\n> > +             fsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n> > +\n> > +             for (i = 0; i < state->nr_renames; i++) {\n> > +                     if (finalize_object_file(state->renames[i].src, state->renames[i].dst))\n> > +                             die_errno(_(\"could not rename '%s'\"), state->renames[i].src);\n> > +\n> > +                     free(state->renames[i].src);\n> > +                     free(state->renames[i].dst);\n> > +             }\n> > +\n> > +             free(state->renames);\n> > +             memset(state, 0, sizeof(*state));\n>\n> Hmm. There is a lot of `memset()`ing going on, and I am not quite sure\n> that I like what I am seeing. It does not help that there are now two very\n> easily-confused structs: `bulk_rename_state` and `bulk_checkin_state`.\n> Which made me worried at first that we might be resetting the `renames`\n> field inadvertently in `finish_bulk_checkin()`.\n\nYeah, the problem I was trying to avoid was an issue with resetting the rename\nstate during the \"too big pack\" case.  What do you think about a\n\"bulk_pack_state\" and\na \"bulk_fsync_state\" variable?  One side advantage of the\n\"s/state/bulk_rename_state\" change\nis that the variable name is no longer ambiguous in the debugger\nacross different object files.\n\n>\n> Maybe we can do this instead?\n>\n>                 FREE_AND_NULL(state->renames);\n>                 state->nr_renames = state->alloc_renames = 0;\n>\n\nI wrote it that way originally, but saw an error with coccinelle and\nthen decided to follow\nthe convention from elsewhere in the file.  That's when I noticed the\nnasty \"early unplug\" case,\nso the Github actions certainly saved me from an unfortunate bug.\nThanks for setting them up!\n\nI'll go toward your suggestion.\n\n> > +     }\n> > +}\n> > +\n> >  static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n> >  {\n> >       int i;\n> > @@ -256,25 +299,69 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n> >       return 0;\n> >  }\n> >\n> > +static void add_rename_bulk_checkin(struct bulk_rename_state *state,\n> > +                                 const char *src, const char *dst)\n> > +{\n> > +     struct object_rename *rename;\n> > +\n> > +     ALLOC_GROW(state->renames, state->nr_renames + 1, state->alloc_renames);\n> > +\n> > +     rename = &state->renames[state->nr_renames++];\n> > +     rename->src = xstrdup(src);\n> > +     rename->dst = xstrdup(dst);\n> > +}\n> > +\n> > +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n> > +                                           const char *filename)\n> > +{\n> > +     if (fsync_object_files) {\n> > +             /*\n> > +              * If we have a plugged bulk checkin, we issue a call that\n> > +              * cleans the filesystem page cache but avoids a hardware flush\n> > +              * command. Later on we will issue a single hardware flush\n> > +              * before renaming files as part of do_sync_and_rename.\n> > +              */\n> > +             if (bulk_checkin_state.plugged &&\n> > +                 fsync_object_files == 2 &&\n> > +                 git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n> > +                     add_rename_bulk_checkin(&bulk_rename_state, tmpfile, filename);\n> > +                     if (close(fd))\n> > +                             die_errno(_(\"error when closing loose object file\"));\n> > +\n> > +                     return 0;\n> > +\n> > +             } else {\n> > +                     fsync_or_die(fd, \"loose object file\");\n> > +             }\n> > +     }\n> > +\n> > +     if (close(fd))\n> > +             die_errno(_(\"error when closing loose object file\"));\n> > +\n> > +     return finalize_object_file(tmpfile, filename);\n> > +}\n> > +\n> >  int index_bulk_checkin(struct object_id *oid,\n> >                      int fd, size_t size, enum object_type type,\n> >                      const char *path, unsigned flags)\n> >  {\n> > -     int status = deflate_to_pack(&state, oid, fd, size, type,\n> > +     int status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n> >                                    path, flags);\n> > -     if (!state.plugged)\n> > -             finish_bulk_checkin(&state);\n> > +     if (!bulk_checkin_state.plugged)\n> > +             finish_bulk_checkin(&bulk_checkin_state);\n> >       return status;\n> >  }\n> >\n> >  void plug_bulk_checkin(void)\n> >  {\n> > -     state.plugged = 1;\n> > +     bulk_checkin_state.plugged = 1;\n> >  }\n> >\n> > -void unplug_bulk_checkin(void)\n> > +void unplug_bulk_checkin(struct lock_file *lock_file)\n> >  {\n> > -     state.plugged = 0;\n> > -     if (state.f)\n> > -             finish_bulk_checkin(&state);\n> > +     bulk_checkin_state.plugged = 0;\n> > +     if (bulk_checkin_state.f)\n> > +             finish_bulk_checkin(&bulk_checkin_state);\n> > +\n> > +     do_sync_and_rename(&bulk_rename_state, lock_file);\n> >  }\n> > diff --git a/bulk-checkin.h b/bulk-checkin.h\n> > index b26f3dc3b74..8efb01ed669 100644\n> > --- a/bulk-checkin.h\n> > +++ b/bulk-checkin.h\n> > @@ -6,11 +6,13 @@\n> >\n> >  #include \"cache.h\"\n> >\n> > +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile, const char *filename);\n> > +\n> >  int index_bulk_checkin(struct object_id *oid,\n> >                      int fd, size_t size, enum object_type type,\n> >                      const char *path, unsigned flags);\n> >\n> >  void plug_bulk_checkin(void);\n> > -void unplug_bulk_checkin(void);\n> > +void unplug_bulk_checkin(struct lock_file *);\n> >\n> >  #endif\n> > diff --git a/config.c b/config.c\n> > index f33abeab851..375bdb24b0a 100644\n> > --- a/config.c\n> > +++ b/config.c\n> > @@ -1509,7 +1509,9 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >       }\n> >\n> >       if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> > -             fsync_object_files = git_config_bool(var, value);\n> > +             int is_bool;\n> > +\n> > +             fsync_object_files = git_config_bool_or_int(var, value, &is_bool);\n> >               return 0;\n> >       }\n> >\n> > diff --git a/config.mak.uname b/config.mak.uname\n> > index 69413fb3dc0..8c07f2265a8 100644\n> > --- a/config.mak.uname\n> > +++ b/config.mak.uname\n> > @@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n> >       HAVE_CLOCK_MONOTONIC = YesPlease\n> >       # -lrt is needed for clock_gettime on glibc <= 2.16\n> >       NEEDS_LIBRT = YesPlease\n> > +     HAVE_SYNC_FILE_RANGE = YesPlease\n> >       HAVE_GETDELIM = YesPlease\n> >       SANE_TEXT_GREP=-a\n> >       FREAD_READS_DIRECTORIES = UnfortunatelyYes\n> > @@ -133,6 +134,7 @@ ifeq ($(uname_S),Darwin)\n> >       COMPAT_OBJS += compat/precompose_utf8.o\n> >       BASIC_CFLAGS += -DPRECOMPOSE_UNICODE\n> >       BASIC_CFLAGS += -DPROTECT_HFS_DEFAULT=1\n> > +     BASIC_CFLAGS += -DFSYNC_DOESNT_FLUSH=1\n> >       HAVE_BSD_SYSCTL = YesPlease\n> >       FREAD_READS_DIRECTORIES = UnfortunatelyYes\n> >       HAVE_NS_GET_EXECUTABLE_PATH = YesPlease\n> > diff --git a/configure.ac b/configure.ac\n> > index 031e8d3fee8..c711037d625 100644\n> > --- a/configure.ac\n> > +++ b/configure.ac\n> > @@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n> >       [AC_MSG_RESULT([no])\n> >       HAVE_CLOCK_MONOTONIC=])\n> >  GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n> > +\n> > +#\n> > +# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n> > +GIT_CHECK_FUNC(sync_file_range,\n> > +     [HAVE_SYNC_FILE_RANGE=YesPlease],\n> > +     [HAVE_SYNC_FILE_RANGE])\n> > +GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n> > +\n> >  #\n> >  # Define NO_SETITIMER if you don't have setitimer.\n> >  GIT_CHECK_FUNC(setitimer,\n> > diff --git a/git-compat-util.h b/git-compat-util.h\n> > index b46605300ab..d14e2436276 100644\n> > --- a/git-compat-util.h\n> > +++ b/git-compat-util.h\n> > @@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n> >  void BUG(const char *fmt, ...);\n> >  #endif\n> >\n> > +enum fsync_action {\n> > +    FSYNC_WRITEOUT_ONLY,\n> > +    FSYNC_HARDWARE_FLUSH\n> > +};\n> > +\n> > +int git_fsync(int fd, enum fsync_action action);\n> > +\n> >  /*\n> >   * Preserves errno, prints a message, but gives no warning for ENOENT.\n> >   * Returns 0 on success, which includes trying to unlink an object that does\n> > diff --git a/object-file.c b/object-file.c\n> > index 607e9e2f80b..5f04143dde0 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1859,16 +1859,6 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n> >       return 0;\n> >  }\n> >\n> > -/* Finalize a file on disk, and close it. */\n> > -static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n> > -{\n> > -     if (fsync_object_files)\n> > -             fsync_or_die(fd, \"loose object file\");\n> > -     if (close(fd) != 0)\n> > -             die_errno(_(\"error when closing loose object file\"));\n> > -     return finalize_object_file(tmpfile, filename);\n> > -}\n> > -\n> >  /* Size of directory component, including the ending '/' */\n> >  static inline int directory_size(const char *filename)\n> >  {\n> > @@ -1982,7 +1972,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >                       warning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n> >       }\n> >\n> > -     return close_loose_object(fd, tmp_file.buf, filename.buf);\n> > +     return fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf, filename.buf);\n> >  }\n> >\n> >  static int freshen_loose_object(const struct object_id *oid)\n> > diff --git a/wrapper.c b/wrapper.c\n> > index 563ad590df1..37a8b61a7df 100644\n> > --- a/wrapper.c\n> > +++ b/wrapper.c\n> > @@ -538,6 +538,42 @@ int xmkstemp_mode(char *filename_template, int mode)\n> >       return fd;\n> >  }\n> >\n> > +int git_fsync(int fd, enum fsync_action action)\n> > +{\n> > +     if (action == FSYNC_WRITEOUT_ONLY) {\n> > +#ifdef __APPLE__\n> > +             /*\n> > +              * on Mac OS X, fsync just causes filesystem cache writeback but does not\n> > +              * flush hardware caches.\n> > +              */\n> > +             return fsync(fd);\n> > +#endif\n> > +\n> > +#ifdef HAVE_SYNC_FILE_RANGE\n> > +             /*\n> > +              * On linux 2.6.17 and above, sync_file_range is the way to issue\n> > +              * a writeback without a hardware flush. An offset of 0 and size of 0\n> > +              * indicates writeout of the entire file and the wait flags ensure that all\n> > +              * dirty data is written to the disk (potentially in a disk-side cache)\n> > +              * before we continue.\n> > +              */\n> > +\n> > +             return sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n> > +                                              SYNC_FILE_RANGE_WRITE |\n> > +                                              SYNC_FILE_RANGE_WAIT_AFTER);\n> > +#endif\n> > +\n> > +             errno = ENOSYS;\n> > +             return -1;\n> > +     }\n>\n> Hmm. I wonder whether we can do this more consistently with how Git\n> usually does platform-specific things.\n>\n> In the 3rd patch, the one where you implemented Windows-specific support,\n> in the Git for Windows PR at\n> https://github.com/git-for-windows/git/pull/3391, you introduce a\n> `mingw_fsync_no_flush()` function and define `fsync_no_flush` to expand to\n> that function name.\n>\n> This is very similar to how Git does things. Take for example the\n> `offset_1st_component` macro:\n> https://github.com/git/git/blob/v2.33.0/git-compat-util.h#L386-L392\n>\n> Unless defined in a platform-specific manner, it is defined in\n> `git-compat-util.h`:\n>\n>         #ifndef offset_1st_component\n>         static inline int git_offset_1st_component(const char *path)\n>         {\n>                 return is_dir_sep(path[0]);\n>         }\n>         #define offset_1st_component git_offset_1st_component\n>         #endif\n>\n> And on Windows, it is defined as following\n> (https://github.com/git/git/blob/v2.33.0/compat/win32/path-utils.h#L34-L35),\n> before the lines quoted above:\n>\n>         int win32_offset_1st_component(const char *path);\n>         #define offset_1st_component win32_offset_1st_component\n>\n> We could do the exact same thing here. Define a platform-specific\n> `mingw_fsync_no_flush()` in `compat/mingw.h` and define the macro\n> `fsync_no_flush` to point to it. In `git-compat-util.h`, in the\n> `__APPLE__`-specific part, implement it via `fsync()`. And later, in the\n> platform-independent part, _iff_ the macro has not yet been defined,\n> implement an inline function that does that `HAVE_SYNC_FILE_RANGE` dance\n> and falls back to `ENOSYS`.\n>\n> That would contain the platform-specific `#ifdef` blocks to\n> `git-compat-util.h`, which is exactly where we want them.\n>\n> > +\n> > +#ifdef __APPLE__\n> > +     return fcntl(fd, F_FULLFSYNC);\n> > +#else\n> > +     return fsync(fd);\n> > +#endif\n>\n> Same thing here. We would probably want something like `fsync_with_flush`\n> here.\n\nI thought about doing it that way originally.  But there's the\nunfortunate fact that\nI'd have to alias fsync_no_flush to fsync on macOS and then fsync to\nfnctl(F_FULLFSYNC),\nI felt that it would be clearer to someone reviewing this\nfunctionality if we provide a\nvery explicit git_fsync API with a well-named flag and document the\nOS-specific craziness\nin the C file rather than through a layer of macros in the header file.\n\nGiven that, are you okay with keeping this code layout in the C file,\npotentially with\nmore local modifications?\n\n>\n> It is my hope that you find my comments and suggestions helpful.\n>\n> Thank you,\n> Johannes\n>\n\nVery much so! Again thanks for the review.\n\n-Neeraj\n"},{"id":"433805","messageId":"20210826055024.GA17178@lst.de","threadId":"56371","inReplyTo":"CANQDOdd2FDNXnXLdm2FSmxUTk3oi+mQtiW2rf3YG7MJayrexPQ@mail.gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-08-26T05:50:24Z","receivedAt":"2021-08-26T05:50:29Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Wed, Aug 25, 2021 at 05:49:45PM -0700, Neeraj Singh wrote:\n> Unfortunately my perusal of the man pages and documentation I could find doesn't\n> give me this level of confidence on typical Linux filesystems. For\n> instance, the notion of having to\n> fsync the parent directory in order to render an inode's link findable\n> eliminates a lot of the\n> advantage of this change, though we could batch those and would have\n> to do at most 256.\n> \n> This thread is somewhat instructive, but inconclusive:\n> https://lwn.net/ml/linux-fsdevel/1552418820-18102-1-git-send-email-jaya@cs.utexas.edu/.\n\nfsync/fdatasync only guarantees consistency for the file handle they\nare called on.  The first linked document mentioned an implementation\nartifact that file systems with metadata logging tend to force their\nlog out until the last modified transaction and thus force out metadata\nchanges done earlier.  This won't help with actual data writes at all,\nas for them the fact of writing back data will often generate new\nmetadata changes., and in general is not a property to rely on if you\ncare about data integrity.  It is nice to optimize the order of the\nfsync calls for metadata only workloads, as then often the later fsync\ncalls on earlier modified file handles will be no-ops.\n\n> One conclusion from reviewing that thread is that as of then,\n> sync_file_ranges isn't actually enough\n> to make a hard guarantee about writeout occurring. See\n> https://lore.kernel.org/linux-fsdevel/20190319204330.GY26298@dastard/.\n> My hope is that the Linux FS developers have rectified that shortcoming by now.\n\nI'm not sure what shortcoming you mean.  sync_file_ranges is a system\ncall that only causes data writeback.  It never performs metadata write\nback and thus is not an integrity operation at all.  That is also very\nclearly documented in the man page.\n\n> I think my updated version of the documentation for \"= false\" is\n> accurate and more helpful\n> from a user perspective (\"up to OS policy when your data becomes durable in\n> the event of an unclean shutdown\").  \"= true\" also has a reasonable\n> description, though I\n> might add some verbiage indicating that this setting could be costly.\n\nYour version is much better.  In fact it almost still too nice as in\ngeneral it will not be durable and you do end up with a corrupted\nrepository in that case.  Note that even for bad old ext3 that was\nusually the case.\n"},{"id":"433806","messageId":"20210826055439.GA17560@lst.de","threadId":"56371","inReplyTo":"CANQDOdf7rMyT4Swriw9=Ei7KN1iLv_dGDWSSck22Zu6AztOyjg@mail.gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-08-26T05:54:39Z","receivedAt":"2021-08-26T05:54:43Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Wed, Aug 25, 2021 at 10:40:53AM -0700, Neeraj Singh wrote:\n> I'd expect syncfs to suffer from the noisy-neighbor problem that Linus\n> alluded to on the big\n> thread you kicked off.\n\nIt does.  That being said I suspect in most developer workstation\nuse cases it will still be a win.  Maybe I'll look into implemeting\nit after your series lands.\n\n> If someone adds a more targeted bulk sync interface to the Linux\n> kernel, I'm sure Git could be\n> changed to use it. Maybe an fcntl(2) interface that initiates\n> writeback and registers completion with an\n> eventfd.\n\nThat is in general very hard to do with how the VM-level writeback\noccurs.  In the file system itself it could work much better, e.g.\nfor XFS we write the log up to a specific sequence number and could\nnotify when doing that.\n"},{"id":"433807","messageId":"20210826055719.GB17560@lst.de","threadId":"56371","inReplyTo":"87mtp5cwpn.fsf@evledraar.gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-08-26T05:57:19Z","receivedAt":"2021-08-26T05:57:28Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Wed, Aug 25, 2021 at 06:11:13PM +0200, Ævar Arnfjörð Bjarmason wrote:\n> 3) Re some of the musings about fsync() recently in\n> https://lore.kernel.org/git/877dhs20x3.fsf@evledraar.gmail.com/; is this\n> method of doing not-quite-an-fsync guaranteed by some OS's / POSIX etc,\n> or is it more like the initial approach before core.fsyncObjectFiles,\n> i.e. the happy-go-lucky approach described in the \"[...]that orders data\n> writes properly[...]\" documentation you're removing.\n\nExcept for the now removed ext3 filesystem in Linux that basically turned\nevery fsync into syncfs, that is a file system-wide sync I've never\nheard about such behavior for data writeback.  Many file systems will\nsometimes or always behave like that for metadata writeback, but there\nis no guarantees you could rely on for that.\n"},{"id":"433982","messageId":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.git.git.1629856292.gitgitgadget@gmail.com","subject":"[PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-27T23:49:31Z","receivedAt":"2021-08-27T23:49:42Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far! I've responded to the previous\nfeedback and changed the patch series a bit.\n\nChanges since v1:\n\n * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n   to dscho's suggestion, I'm still implementing the Windows version in the\n   same patch and I'm not doing autoconf detection since this is a POSIX\n   function.\n\n * Introduce a separate preparatory patch to the bulk-checkin infrastructure\n   to separate the 'plugged' variable and rename the 'state' variable, as\n   suggested by dscho.\n\n * Add performance numbers to the commit message of the main bulk fsync\n   patch, as suggested by dscho.\n\n * Add a comment about the non-thread-safety of the bulk-checkin\n   infrastructure, as suggested by avarab.\n\n * Rename the experimental mode to core.fsyncobjectfiles=batch, as suggested\n   by dscho and avarab and others.\n\n * Add more details to Documentation/config/core.txt about the various\n   settings and their intended effects, as suggested by avarab.\n\n * Switch to the string-list API to hold the rename state, as suggested by\n   avarab.\n\n * Create a separate update-index patch to use bulk-checkin as suggested by\n   dscho.\n\n * Add Windows support in the upstream git. This is done in a way that\n   should not conflict with git-for-windows.\n\n * Add new performance tests that shows the delta based on fsync mode.\n\nNOTE: Based on Christoph Hellwig's comments, the 'batch' mode is not correct\non Linux, since sync_file_range does not provide data integrity guarantees.\nThere is currently no kernel interface suitable to achieve disk flush\nbatching as is, but he suggested that he might implement a 'syncfs' variant\non top of this patchset. This code is still useful on macOS and Windows, and\nthe config documentation makes that clear.\n\nNeeraj Singh (6):\n  object-file: use futimens rather than utime\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  core.fsyncobjectfiles: add windows support for batch mode\n  update-index: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: performance tests for add and stash\n\n Documentation/config/core.txt       | 26 ++++++--\n Makefile                            |  6 ++\n builtin/add.c                       |  3 +-\n builtin/update-index.c              |  3 +\n bulk-checkin.c                      | 92 +++++++++++++++++++++++++----\n bulk-checkin.h                      |  4 +-\n cache.h                             |  8 ++-\n compat/mingw.c                      | 53 +++++++++++------\n compat/mingw.h                      |  5 ++\n compat/win32/flush.c                | 29 +++++++++\n config.c                            |  8 ++-\n config.mak.uname                    |  4 ++\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n environment.c                       |  2 +-\n git-compat-util.h                   |  7 +++\n object-file.c                       | 23 ++------\n t/perf/lib-unique-files.sh          | 32 ++++++++++\n t/perf/p3700-add.sh                 | 43 ++++++++++++++\n t/perf/p3900-stash.sh               | 46 +++++++++++++++\n wrapper.c                           | 40 +++++++++++++\n write-or-die.c                      |  2 +-\n 22 files changed, 389 insertions(+), 58 deletions(-)\n create mode 100644 compat/win32/flush.c\n create mode 100644 t/perf/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: 225bc32a989d7a22fa6addafd4ce7dcd04675dbf\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v2\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v1:\n\n 1:  2c1ddef6057 ! 1:  fc3d5a7b635 object-file: use futimes rather than utime\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    object-file: use futimes rather than utime\n     +    object-file: use futimens rather than utime\n      \n     -    Refactor the loose object file creation code and use the futimes(2) API\n     -    rather than utime. This should be slightly faster given that we already\n     -    have an FD to work with.\n     +    Make close_loose_object do all of the steps for syncing and correctly\n     +    naming a new loose object so that it can be reimplemented in the\n     +    upcoming bulk-fsync mode.\n     +\n     +    Use futimens, which is available in POSIX.1-2008 to update the file\n     +    timestamps. This should be slightly faster than utime, since we have\n     +    a file descriptor already available. This change allows us to update\n     +    the time before closing, renaming, and potentially fsyincing the file\n     +    being refreshed. This code is currently only invoked by git-pack-objects\n     +    via force_object_loose.\n     +\n     +    Implement a futimens shim for the Windows port of Git.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## compat/mingw.c ##\n     +@@ compat/mingw.c: int mingw_chmod(const char *filename, int mode)\n     +  * The unit of FILETIME is 100-nanoseconds since January 1, 1601, UTC.\n     +  * Returns the 100-nanoseconds (\"hekto nanoseconds\") since the epoch.\n     +  */\n     ++\n     ++#define UNIX_EPOCH_FILETIME 116444736000000000LL\n     ++\n     + static inline long long filetime_to_hnsec(const FILETIME *ft)\n     + {\n     + \tlong long winTime = ((long long)ft->dwHighDateTime << 32) + ft->dwLowDateTime;\n     + \t/* Windows to Unix Epoch conversion */\n     +-\treturn winTime - 116444736000000000LL;\n     ++\treturn winTime - UNIX_EPOCH_FILETIME;\n     + }\n     + \n     + static inline void filetime_to_timespec(const FILETIME *ft, struct timespec *ts)\n     +@@ compat/mingw.c: static inline void filetime_to_timespec(const FILETIME *ft, struct timespec *ts)\n     + \tts->tv_nsec = (hnsec % 10000000) * 100;\n     + }\n     + \n     ++static inline void timespec_to_filetime(const struct timespec *t, FILETIME *ft)\n     ++{\n     ++\tlong long winTime = t->tv_sec * 10000000LL + t->tv_nsec / 100 + UNIX_EPOCH_FILETIME;\n     ++\tft->dwLowDateTime = winTime;\n     ++\tft->dwHighDateTime = winTime >> 32;\n     ++}\n     ++\n     + /**\n     +  * Verifies that safe_create_leading_directories() would succeed.\n     +  */\n      @@ compat/mingw.c: int mingw_fstat(int fd, struct stat *buf)\n       \t}\n       }\n       \n      -static inline void time_t_to_filetime(time_t t, FILETIME *ft)\n     -+static inline void timeval_to_filetime(const struct timeval *t, FILETIME *ft)\n     ++int mingw_futimens(int fd, const struct timespec times[2])\n       {\n      -\tlong long winTime = t * 10000000LL + 116444736000000000LL;\n     -+\tlong long winTime = t->tv_sec * 10000000LL + t->tv_usec * 10 + 116444736000000000LL;\n     - \tft->dwLowDateTime = winTime;\n     - \tft->dwHighDateTime = winTime >> 32;\n     - }\n     - \n     --int mingw_utime (const char *file_name, const struct utimbuf *times)\n     -+int mingw_futimes(int fd, const struct timeval times[2])\n     - {\n     - \tFILETIME mft, aft;\n     +-\tft->dwLowDateTime = winTime;\n     +-\tft->dwHighDateTime = winTime >> 32;\n     ++\tFILETIME mft, aft;\n      +\n      +\tif (times) {\n     -+\t\ttimeval_to_filetime(&times[0], &aft);\n     -+\t\ttimeval_to_filetime(&times[1], &mft);\n     ++\t\ttimespec_to_filetime(&times[0], &aft);\n     ++\t\ttimespec_to_filetime(&times[1], &mft);\n      +\t} else {\n      +\t\tGetSystemTimeAsFileTime(&mft);\n      +\t\taft = mft;\n     @@ compat/mingw.c: int mingw_fstat(int fd, struct stat *buf)\n      +\t}\n      +\n      +\treturn 0;\n     -+}\n     -+\n     -+int mingw_utime (const char *file_name, const struct utimbuf *times)\n     -+{\n     + }\n     + \n     +-int mingw_utime (const char *file_name, const struct utimbuf *times)\n     ++int mingw_utime(const char *file_name, const struct utimbuf *times)\n     + {\n     +-\tFILETIME mft, aft;\n       \tint fh, rc;\n       \tDWORD attrs;\n       \twchar_t wfilename[MAX_PATH];\n     -+\tstruct timeval tvs[2];\n     ++\tstruct timespec ts[2];\n      +\n       \tif (xutftowcs_path(wfilename, file_name) < 0)\n       \t\treturn -1;\n     @@ compat/mingw.c: int mingw_utime (const char *file_name, const struct utimbuf *ti\n      -\t} else {\n      -\t\tGetSystemTimeAsFileTime(&mft);\n      -\t\taft = mft;\n     -+\t\tmemset(tvs, 0, sizeof(tvs));\n     -+\t\ttvs[0].tv_sec = times->actime;\n     -+\t\ttvs[1].tv_sec = times->modtime;\n     ++\t\tmemset(ts, 0, sizeof(ts));\n     ++\t\tts[0].tv_sec = times->actime;\n     ++\t\tts[1].tv_sec = times->modtime;\n       \t}\n      -\tif (!SetFileTime((HANDLE)_get_osfhandle(fh), NULL, &aft, &mft)) {\n      -\t\terrno = EINVAL;\n     @@ compat/mingw.c: int mingw_utime (const char *file_name, const struct utimbuf *ti\n      -\t} else\n      -\t\trc = 0;\n      +\n     -+\trc = mingw_futimes(fh, times ? tvs : NULL);\n     ++\trc = mingw_futimens(fh, times ? ts : NULL);\n       \tclose(fh);\n       \n       revert_attrs:\n     @@ compat/mingw.h: int mingw_fstat(int fd, struct stat *buf);\n       \n       int mingw_utime(const char *file_name, const struct utimbuf *times);\n       #define utime mingw_utime\n     -+int mingw_futimes(int fd, const struct timeval times[2]);\n     -+#define futimes mingw_futimes\n     ++int mingw_futimens(int fd, const struct timespec times[2]);\n     ++#define futimens mingw_futimens\n       size_t mingw_strftime(char *s, size_t max,\n       \t\t   const char *format, const struct tm *tm);\n       #define strftime mingw_strftime\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      -\t\tutb.modtime = mtime;\n      -\t\tif (utime(tmp_file.buf, &utb) < 0)\n      -\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n     -+\t\tstruct timeval tvs[2] = {0};\n     -+\t\ttvs[0].tv_sec = mtime;\n     -+\t\ttvs[1].tv_sec = mtime;\n     -+\t\tif (futimes(fd, tvs) < 0)\n     ++\t\tstruct timespec ts[2] = {0};\n     ++\t\tts[0].tv_sec = mtime;\n     ++\t\tts[1].tv_sec = mtime;\n     ++\t\tif (futimens(fd, ts) < 0)\n      +\t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n       \t}\n       \n -:  ----------- > 2:  49f72800bfb bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n 2:  d1e68d4a2af ! 3:  2c1c907b12a core.fsyncobjectfiles: batch disk flushes\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    core.fsyncobjectfiles: batch disk flushes\n     +    core.fsyncobjectfiles: batched disk flushes\n      \n          When adding many objects to a repo with core.fsyncObjectFiles set to\n          true, the cost of fsync'ing each object file can become prohibitive.\n      \n          One major source of the cost of fsync is the implied flush of the\n          hardware writeback cache within the disk drive. Fortunately, Windows,\n     -    MacOS, and Linux each offer mechanisms to write data from the filesystem\n     +    macOS, and Linux each offer mechanisms to write data from the filesystem\n          page cache without initiating a hardware flush.\n      \n     -    This patch introduces a new 'core.fsyncObjectFiles = 2' option that\n     +    This patch introduces a new 'core.fsyncObjectFiles = batch' option that\n          takes advantage of the bulk-checkin infrastructure to batch up hardware\n          flushes.\n      \n     @@ Commit message\n          1. Create a tmp_obj_XXXX file and write the object data to it.\n          2. Issue a pagecache writeback request and wait for it to complete.\n          3. Record the tmp name and the final name in the bulk-checkin state for\n     -       later name.\n     +       later rename.\n      \n          At the end of the entire transaction we:\n          1. Issue a fsync against the lock file to flush the hardware writeback\n             cache, which should by now have processed the tmp file writes.\n          2. Rename all of the temp files to their final names.\n     -    3. When updating the index and/or refs, we will issue another fsync\n     -       internal to that operation.\n     +    3. When updating the index and/or refs, we assume that Git will issue\n     +       another fsync internal to that operation.\n      \n          On a filesystem with a singular journal that is updated during name\n          operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n     @@ Commit message\n          sequence is enough to ensure that the user's data is durable by the time\n          the git command returns.\n      \n     -    This change also updates the MacOS code to trigger a real hardware flush\n     +    This change also updates the macOS code to trigger a real hardware flush\n          via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n     -    MacOS there was no guarantee of durability since a simple fsync(2) call\n     +    macOS there was no guarantee of durability since a simple fsync(2) call\n          does not flush any hardware caches.\n      \n     +    _Performance numbers_:\n     +\n     +    Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\n     +    Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\n     +    Windows - Same host as Linux, a preview version of Windows 11.\n     +              This number is from a patch later in the series.\n     +\n     +    Adding 500 files to the repo with 'git add' Times reported in seconds.\n     +\n     +    core.fsyncObjectFiles | Linux | Mac   | Windows\n     +    ----------------------|-------|-------|--------\n     +                    false | 0.06  |  0.35 | 0.61\n     +                    true  | 1.88  | 11.18 | 2.47\n     +                    batch | 0.15  |  0.41 | 1.53\n     +\n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Documentation/config/core.txt ##\n     @@ Documentation/config/core.txt: core.whitespace::\n      -data writes properly, but can be useful for filesystems that do not use\n      -journalling (traditional UNIX filesystems) or that only journal metadata\n      -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n     -+\tA boolean value or the number '2', indicating the level of durability\n     -+\tapplied to object files.\n     ++\tA value indicating the level of effort Git will expend in\n     ++\ttrying to make objects added to the repo durable in the event\n     ++\tof an unclean system shutdown. This setting currently only\n     ++\tcontrols the object store, so updates to any refs or the\n     ++\tindex may not be equally durable.\n      ++\n     -+This setting controls how much effort Git makes to ensure that data added to\n     -+the object store are durable in the case of an unclean system shutdown. If\n     -+'false', Git allows data to remain in file system caches according to operating\n     -+system policy, whence they may be lost if the system loses power or crashes. A\n     -+value of 'true' instructs Git to force objects to stable storage immediately\n     -+when they are added to the object store. The number '2' is an experimental\n     -+value that also preserves durability but tries to perform hardware flushes in a\n     -+batch.\n     ++* `false` allows data to remain in file system caches according to\n     ++  operating system policy, whence it may be lost if the system loses power\n     ++  or crashes.\n     ++* `true` triggers a data integrity flush for each object added to the\n     ++  object store. This is the safest setting that is likely to ensure durability\n     ++  across all operating systems and file systems that honor the 'fsync' system\n     ++  call. However, this setting comes with a significant performance cost on\n     ++  common hardware.\n     ++* `batch` enables an experimental mode that uses interfaces available in some\n     ++  operating systems to write object data with a minimal set of FLUSH CACHE\n     ++  (or equivalent) commands sent to the storage controller. If the operating\n     ++  system interfaces are not available, this mode behaves the same as `true`.\n     ++  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n     ++  filesystems and on Windows for repos stored on NTFS or ReFS.\n       \n       core.preloadIndex::\n       \tEnable parallel index preload for operations like 'git diff'\n      \n       ## Makefile ##\n     +@@ Makefile: all::\n     + #\n     + # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n     + #\n     ++# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n     ++#\n     + # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n     + # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n     + #\n      @@ Makefile: ifdef HAVE_CLOCK_MONOTONIC\n       \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n       endif\n     @@ builtin/add.c: int cmd_add(int argc, const char **argv, const char *prefix)\n       finish:\n       \tif (write_locked_index(&the_index, &lock_file,\n      \n     - ## builtin/update-index.c ##\n     -@@\n     -  */\n     - #define USE_THE_INDEX_COMPATIBILITY_MACROS\n     - #include \"cache.h\"\n     -+#include \"bulk-checkin.h\"\n     - #include \"config.h\"\n     - #include \"lockfile.h\"\n     - #include \"quote.h\"\n     -@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n     - \t\tstruct strbuf unquoted = STRBUF_INIT;\n     - \n     - \t\tsetup_work_tree();\n     -+\t\tplug_bulk_checkin();\n     - \t\twhile (getline_fn(&buf, stdin) != EOF) {\n     - \t\t\tchar *p;\n     - \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n     -@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n     - \t\t\t\tchmod_path(set_executable_bit, p);\n     - \t\t\tfree(p);\n     - \t\t}\n     -+\t\tunplug_bulk_checkin(&lock_file);\n     - \t\tstrbuf_release(&unquoted);\n     - \t\tstrbuf_release(&buf);\n     - \t}\n     -\n       ## bulk-checkin.c ##\n      @@\n        */\n     @@ bulk-checkin.c\n       #include \"repository.h\"\n       #include \"csum-file.h\"\n       #include \"pack.h\"\n     -@@\n     + #include \"strbuf.h\"\n     ++#include \"string-list.h\"\n       #include \"packfile.h\"\n       #include \"object-store.h\"\n       \n     -+struct object_rename {\n     -+\tchar *src;\n     -+\tchar *dst;\n     -+};\n     -+\n     -+static struct bulk_rename_state {\n     -+\tstruct object_rename *renames;\n     -+\tuint32_t alloc_renames;\n     -+\tuint32_t nr_renames;\n     -+} bulk_rename_state;\n     -+\n     - static struct bulk_checkin_state {\n     - \tunsigned plugged:1;\n     + static int bulk_checkin_plugged;\n       \n     -@@ bulk-checkin.c: static struct bulk_checkin_state {\n     - \tstruct pack_idx_entry **written;\n     - \tuint32_t alloc_written;\n     - \tuint32_t nr_written;\n     --} state;\n     ++static struct string_list bulk_fsync_state = STRING_LIST_INIT_DUP;\n      +\n     -+} bulk_checkin_state;\n     - \n     - static void finish_bulk_checkin(struct bulk_checkin_state *state)\n     - {\n     - \tstruct object_id oid;\n     - \tstruct strbuf packname = STRBUF_INIT;\n     - \tint i;\n     -+\tunsigned old_plugged;\n     - \n     - \tif (!state->f)\n     - \t\treturn;\n     -@@ bulk-checkin.c: static void finish_bulk_checkin(struct bulk_checkin_state *state)\n     - \n     - clear_exit:\n     - \tfree(state->written);\n     -+\told_plugged = state->plugged;\n     - \tmemset(state, 0, sizeof(*state));\n     -+\tstate->plugged = old_plugged;\n     - \n     - \tstrbuf_release(&packname);\n     - \t/* Make objects we just wrote available to ourselves */\n     + static struct bulk_checkin_state {\n     + \tchar *pack_tmp_name;\n     + \tstruct hashfile *f;\n     +@@ bulk-checkin.c: clear_exit:\n       \treprepare_packed_git(the_repository);\n       }\n       \n     -+static void do_sync_and_rename(struct bulk_rename_state *state, struct lock_file *lock_file)\n     ++static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n      +{\n     -+\tif (state->nr_renames) {\n     -+\t\tint i;\n     ++\tif (fsync_state->nr) {\n     ++\t\tstruct string_list_item *rename;\n      +\n      +\t\t/*\n      +\t\t * Issue a full hardware flush against the lock file to ensure\n     @@ bulk-checkin.c: static void finish_bulk_checkin(struct bulk_checkin_state *state\n      +\t\t */\n      +\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n      +\n     -+\t\tfor (i = 0; i < state->nr_renames; i++) {\n     -+\t\t\tif (finalize_object_file(state->renames[i].src, state->renames[i].dst))\n     -+\t\t\t\tdie_errno(_(\"could not rename '%s'\"), state->renames[i].src);\n     ++\t\tfor_each_string_list_item(rename, fsync_state) {\n     ++\t\t\tconst char *src = rename->string;\n     ++\t\t\tconst char *dst = rename->util;\n      +\n     -+\t\t\tfree(state->renames[i].src);\n     -+\t\t\tfree(state->renames[i].dst);\n     ++\t\t\tif (finalize_object_file(src, dst))\n     ++\t\t\t\tdie_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n      +\t\t}\n      +\n     -+\t\tfree(state->renames);\n     -+\t\tmemset(state, 0, sizeof(*state));\n     ++\t\tstring_list_clear(fsync_state, 1);\n      +\t}\n      +}\n      +\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n       \treturn 0;\n       }\n       \n     -+static void add_rename_bulk_checkin(struct bulk_rename_state *state,\n     ++static void add_rename_bulk_checkin(struct string_list *fsync_state,\n      +\t\t\t\t    const char *src, const char *dst)\n      +{\n     -+\tstruct object_rename *rename;\n     -+\n     -+\tALLOC_GROW(state->renames, state->nr_renames + 1, state->alloc_renames);\n     -+\n     -+\trename = &state->renames[state->nr_renames++];\n     -+\trename->src = xstrdup(src);\n     -+\trename->dst = xstrdup(dst);\n     ++\tstring_list_insert(fsync_state, src)->util = xstrdup(dst);\n      +}\n      +\n      +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n      +\t\t\t\t\t      const char *filename)\n      +{\n     -+\tif (fsync_object_files) {\n     ++\tif (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n      +\t\t/*\n      +\t\t * If we have a plugged bulk checkin, we issue a call that\n      +\t\t * cleans the filesystem page cache but avoids a hardware flush\n      +\t\t * command. Later on we will issue a single hardware flush\n      +\t\t * before renaming files as part of do_sync_and_rename.\n      +\t\t */\n     -+\t\tif (bulk_checkin_state.plugged &&\n     -+\t\t    fsync_object_files == 2 &&\n     ++\t\tif (bulk_checkin_plugged &&\n     ++\t\t    fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n      +\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n     -+\t\t\tadd_rename_bulk_checkin(&bulk_rename_state, tmpfile, filename);\n     ++\t\t\tadd_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n      +\t\t\tif (close(fd))\n      +\t\t\t\tdie_errno(_(\"error when closing loose object file\"));\n      +\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n       int index_bulk_checkin(struct object_id *oid,\n       \t\t       int fd, size_t size, enum object_type type,\n       \t\t       const char *path, unsigned flags)\n     - {\n     --\tint status = deflate_to_pack(&state, oid, fd, size, type,\n     -+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n     - \t\t\t\t     path, flags);\n     --\tif (!state.plugged)\n     --\t\tfinish_bulk_checkin(&state);\n     -+\tif (!bulk_checkin_state.plugged)\n     -+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n     - \treturn status;\n     - }\n     - \n     - void plug_bulk_checkin(void)\n     - {\n     --\tstate.plugged = 1;\n     -+\tbulk_checkin_state.plugged = 1;\n     +@@ bulk-checkin.c: void plug_bulk_checkin(void)\n     + \tbulk_checkin_plugged = 1;\n       }\n       \n      -void unplug_bulk_checkin(void)\n      +void unplug_bulk_checkin(struct lock_file *lock_file)\n       {\n     --\tstate.plugged = 0;\n     --\tif (state.f)\n     --\t\tfinish_bulk_checkin(&state);\n     -+\tbulk_checkin_state.plugged = 0;\n     -+\tif (bulk_checkin_state.f)\n     -+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n     + \tassert(bulk_checkin_plugged);\n     + \tbulk_checkin_plugged = 0;\n     + \tif (bulk_checkin_state.f)\n     + \t\tfinish_bulk_checkin(&bulk_checkin_state);\n      +\n     -+\tdo_sync_and_rename(&bulk_rename_state, lock_file);\n     ++\tdo_sync_and_rename(&bulk_fsync_state, lock_file);\n       }\n      \n       ## bulk-checkin.h ##\n     @@ bulk-checkin.h\n       \n       #endif\n      \n     + ## cache.h ##\n     +@@ cache.h: void reset_shared_repository(void);\n     + extern int read_replace_refs;\n     + extern char *git_replace_ref_base;\n     + \n     +-extern int fsync_object_files;\n     ++enum FSYNC_OBJECT_FILES_MODE {\n     ++    FSYNC_OBJECT_FILES_OFF,\n     ++    FSYNC_OBJECT_FILES_ON,\n     ++    FSYNC_OBJECT_FILES_BATCH\n     ++};\n     ++\n     ++extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n     + extern int core_preload_index;\n     + extern int precomposed_unicode;\n     + extern int protect_hfs;\n     +\n       ## config.c ##\n      @@ config.c: static int git_default_core_config(const char *var, const char *value, void *cb)\n       \t}\n       \n       \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n      -\t\tfsync_object_files = git_config_bool(var, value);\n     -+\t\tint is_bool;\n     -+\n     -+\t\tfsync_object_files = git_config_bool_or_int(var, value, &is_bool);\n     ++\t\tif (!value)\n     ++\t\t\treturn config_error_nonbool(var);\n     ++\t\tif (!strcasecmp(value, \"batch\"))\n     ++\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n     ++\t\telse\n     ++\t\t\tfsync_object_files = git_config_bool(var, value)\n     ++\t\t\t\t? FSYNC_OBJECT_FILES_ON : FSYNC_OBJECT_FILES_OFF;\n       \t\treturn 0;\n       \t}\n       \n     @@ configure.ac: AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n       # Define NO_SETITIMER if you don't have setitimer.\n       GIT_CHECK_FUNC(setitimer,\n      \n     + ## environment.c ##\n     +@@ environment.c: const char *git_hooks_path;\n     + int zlib_compression_level = Z_BEST_SPEED;\n     + int core_compression_level;\n     + int pack_compression_level = Z_DEFAULT_COMPRESSION;\n     +-int fsync_object_files;\n     ++enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n     + size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n     + size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n     + size_t delta_base_cache_limit = 96 * 1024 * 1024;\n     +\n       ## git-compat-util.h ##\n      @@ git-compat-util.h: __attribute__((format (printf, 1, 2))) NORETURN\n       void BUG(const char *fmt, ...);\n -:  ----------- > 4:  546ad9c82e8 core.fsyncobjectfiles: add windows support for batch mode\n -:  ----------- > 5:  d8843185fe4 update-index: use the bulk-checkin infrastructure\n -:  ----------- > 6:  73b5d41be94 core.fsyncobjectfiles: performance tests for add and stash\n\n-- \ngitgitgadget\n"},{"id":"433980","messageId":"fc3d5a7b63524647c4a0de53e41772a7eede4f2d.1630108177.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"[PATCH v2 1/6] object-file: use futimens rather than utime","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-27T23:49:32Z","receivedAt":"2021-08-27T23:49:43Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nMake close_loose_object do all of the steps for syncing and correctly\nnaming a new loose object so that it can be reimplemented in the\nupcoming bulk-fsync mode.\n\nUse futimens, which is available in POSIX.1-2008 to update the file\ntimestamps. This should be slightly faster than utime, since we have\na file descriptor already available. This change allows us to update\nthe time before closing, renaming, and potentially fsyincing the file\nbeing refreshed. This code is currently only invoked by git-pack-objects\nvia force_object_loose.\n\nImplement a futimens shim for the Windows port of Git.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.c | 53 ++++++++++++++++++++++++++++++++++----------------\n compat/mingw.h |  2 ++\n object-file.c  | 17 ++++++++--------\n 3 files changed, 46 insertions(+), 26 deletions(-)\n\ndiff --git a/compat/mingw.c b/compat/mingw.c\nindex 9e0cd1e097f..ce14b21c182 100644\n--- a/compat/mingw.c\n+++ b/compat/mingw.c\n@@ -734,11 +734,14 @@ int mingw_chmod(const char *filename, int mode)\n  * The unit of FILETIME is 100-nanoseconds since January 1, 1601, UTC.\n  * Returns the 100-nanoseconds (\"hekto nanoseconds\") since the epoch.\n  */\n+\n+#define UNIX_EPOCH_FILETIME 116444736000000000LL\n+\n static inline long long filetime_to_hnsec(const FILETIME *ft)\n {\n \tlong long winTime = ((long long)ft->dwHighDateTime << 32) + ft->dwLowDateTime;\n \t/* Windows to Unix Epoch conversion */\n-\treturn winTime - 116444736000000000LL;\n+\treturn winTime - UNIX_EPOCH_FILETIME;\n }\n \n static inline void filetime_to_timespec(const FILETIME *ft, struct timespec *ts)\n@@ -748,6 +751,13 @@ static inline void filetime_to_timespec(const FILETIME *ft, struct timespec *ts)\n \tts->tv_nsec = (hnsec % 10000000) * 100;\n }\n \n+static inline void timespec_to_filetime(const struct timespec *t, FILETIME *ft)\n+{\n+\tlong long winTime = t->tv_sec * 10000000LL + t->tv_nsec / 100 + UNIX_EPOCH_FILETIME;\n+\tft->dwLowDateTime = winTime;\n+\tft->dwHighDateTime = winTime >> 32;\n+}\n+\n /**\n  * Verifies that safe_create_leading_directories() would succeed.\n  */\n@@ -949,19 +959,33 @@ int mingw_fstat(int fd, struct stat *buf)\n \t}\n }\n \n-static inline void time_t_to_filetime(time_t t, FILETIME *ft)\n+int mingw_futimens(int fd, const struct timespec times[2])\n {\n-\tlong long winTime = t * 10000000LL + 116444736000000000LL;\n-\tft->dwLowDateTime = winTime;\n-\tft->dwHighDateTime = winTime >> 32;\n+\tFILETIME mft, aft;\n+\n+\tif (times) {\n+\t\ttimespec_to_filetime(&times[0], &aft);\n+\t\ttimespec_to_filetime(&times[1], &mft);\n+\t} else {\n+\t\tGetSystemTimeAsFileTime(&mft);\n+\t\taft = mft;\n+\t}\n+\n+\tif (!SetFileTime((HANDLE)_get_osfhandle(fd), NULL, &aft, &mft)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+\t}\n+\n+\treturn 0;\n }\n \n-int mingw_utime (const char *file_name, const struct utimbuf *times)\n+int mingw_utime(const char *file_name, const struct utimbuf *times)\n {\n-\tFILETIME mft, aft;\n \tint fh, rc;\n \tDWORD attrs;\n \twchar_t wfilename[MAX_PATH];\n+\tstruct timespec ts[2];\n+\n \tif (xutftowcs_path(wfilename, file_name) < 0)\n \t\treturn -1;\n \n@@ -979,17 +1003,12 @@ int mingw_utime (const char *file_name, const struct utimbuf *times)\n \t}\n \n \tif (times) {\n-\t\ttime_t_to_filetime(times->modtime, &mft);\n-\t\ttime_t_to_filetime(times->actime, &aft);\n-\t} else {\n-\t\tGetSystemTimeAsFileTime(&mft);\n-\t\taft = mft;\n+\t\tmemset(ts, 0, sizeof(ts));\n+\t\tts[0].tv_sec = times->actime;\n+\t\tts[1].tv_sec = times->modtime;\n \t}\n-\tif (!SetFileTime((HANDLE)_get_osfhandle(fh), NULL, &aft, &mft)) {\n-\t\terrno = EINVAL;\n-\t\trc = -1;\n-\t} else\n-\t\trc = 0;\n+\n+\trc = mingw_futimens(fh, times ? ts : NULL);\n \tclose(fh);\n \n revert_attrs:\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..87944dfec72 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -398,6 +398,8 @@ int mingw_fstat(int fd, struct stat *buf);\n \n int mingw_utime(const char *file_name, const struct utimbuf *times);\n #define utime mingw_utime\n+int mingw_futimens(int fd, const struct timespec times[2]);\n+#define futimens mingw_futimens\n size_t mingw_strftime(char *s, size_t max,\n \t\t   const char *format, const struct tm *tm);\n #define strftime mingw_strftime\ndiff --git a/object-file.c b/object-file.c\nindex a8be8994814..5421811273e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1860,12 +1860,13 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n }\n \n /* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd)\n+static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n {\n \tif (fsync_object_files)\n \t\tfsync_or_die(fd, \"loose object file\");\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n+\treturn finalize_object_file(tmpfile, filename);\n }\n \n /* Size of directory component, including the ending '/' */\n@@ -1973,17 +1974,15 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n-\n \tif (mtime) {\n-\t\tstruct utimbuf utb;\n-\t\tutb.actime = mtime;\n-\t\tutb.modtime = mtime;\n-\t\tif (utime(tmp_file.buf, &utb) < 0)\n-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n+\t\tstruct timespec ts[2] = {0};\n+\t\tts[0].tv_sec = mtime;\n+\t\tts[1].tv_sec = mtime;\n+\t\tif (futimens(fd, ts) < 0)\n+\t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n \t}\n \n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\treturn close_loose_object(fd, tmp_file.buf, filename.buf);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\n-- \ngitgitgadget\n\n"},{"id":"433981","messageId":"49f72800bfb9d23d5f42005f494a56bbabdf4aca.1630108177.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"[PATCH v2 2/6] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-27T23:49:33Z","receivedAt":"2021-08-27T23:49:46Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b023d9959aa..f117d62c908 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_bulk_checkin(struct bulk_checkin_state *state)\n {\n@@ -260,21 +260,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"433985","messageId":"2c1c907b12a5a8592136dcf775c9c61fac986952.1630108177.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"[PATCH v2 3/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-27T23:49:34Z","receivedAt":"2021-08-27T23:49:47Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nmacOS, and Linux each offer mechanisms to write data from the filesystem\npage cache without initiating a hardware flush.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\ntakes advantage of the bulk-checkin infrastructure to batch up hardware\nflushes.\n\nWhen the new mode is enabled we do the following for new objects:\n\n1. Create a tmp_obj_XXXX file and write the object data to it.\n2. Issue a pagecache writeback request and wait for it to complete.\n3. Record the tmp name and the final name in the bulk-checkin state for\n   later rename.\n\nAt the end of the entire transaction we:\n1. Issue a fsync against the lock file to flush the hardware writeback\n   cache, which should by now have processed the tmp file writes.\n2. Rename all of the temp files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\nwould expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 26 ++++++++++---\n Makefile                      |  6 +++\n builtin/add.c                 |  3 +-\n bulk-checkin.c                | 70 ++++++++++++++++++++++++++++++++++-\n bulk-checkin.h                |  4 +-\n cache.h                       |  8 +++-\n config.c                      |  8 +++-\n config.mak.uname              |  2 +\n configure.ac                  |  8 ++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 ++++\n object-file.c                 | 12 +-----\n wrapper.c                     | 36 ++++++++++++++++++\n write-or-die.c                |  2 +-\n 14 files changed, 170 insertions(+), 24 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..0006d90980d 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,26 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols the object store, so updates to any refs or the\n+\tindex may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write object data with a minimal set of FLUSH CACHE\n+  (or equivalent) commands sent to the storage controller. If the operating\n+  system interfaces are not available, this mode behaves the same as `true`.\n+  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n+  filesystems and on Windows for repos stored on NTFS or ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex 9573190f1d7..143d30f04cb 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1896,6 +1898,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 09e684585d9..c58dfcd4bc3 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -670,7 +670,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\n+\tunplug_bulk_checkin(&lock_file);\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f117d62c908..47b42f610c0 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,15 +3,19 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n \n+static struct string_list bulk_fsync_state = STRING_LIST_INIT_DUP;\n+\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -62,6 +66,32 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n+{\n+\tif (fsync_state->nr) {\n+\t\tstruct string_list_item *rename;\n+\n+\t\t/*\n+\t\t * Issue a full hardware flush against the lock file to ensure\n+\t\t * that all objects are durable before any renames occur.\n+\t\t * The code in fsync_and_close_loose_object_bulk_checkin has\n+\t\t * already ensured that writeout has occurred, but it has not\n+\t\t * flushed any writeback cache in the storage hardware.\n+\t\t */\n+\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n+\n+\t\tfor_each_string_list_item(rename, fsync_state) {\n+\t\t\tconst char *src = rename->string;\n+\t\t\tconst char *dst = rename->util;\n+\n+\t\t\tif (finalize_object_file(src, dst))\n+\t\t\t\tdie_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n+\t\t}\n+\n+\t\tstring_list_clear(fsync_state, 1);\n+\t}\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -256,6 +286,42 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+static void add_rename_bulk_checkin(struct string_list *fsync_state,\n+\t\t\t\t    const char *src, const char *dst)\n+{\n+\tstring_list_insert(fsync_state, src)->util = xstrdup(dst);\n+}\n+\n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n+\t\t\t\t\t      const char *filename)\n+{\n+\tif (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n+\t\t/*\n+\t\t * If we have a plugged bulk checkin, we issue a call that\n+\t\t * cleans the filesystem page cache but avoids a hardware flush\n+\t\t * command. Later on we will issue a single hardware flush\n+\t\t * before renaming files as part of do_sync_and_rename.\n+\t\t */\n+\t\tif (bulk_checkin_plugged &&\n+\t\t    fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n+\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\t\tadd_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n+\t\t\tif (close(fd))\n+\t\t\t\tdie_errno(_(\"error when closing loose object file\"));\n+\n+\t\t\treturn 0;\n+\n+\t\t} else {\n+\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\t}\n+\t}\n+\n+\tif (close(fd))\n+\t\tdie_errno(_(\"error when closing loose object file\"));\n+\n+\treturn finalize_object_file(tmpfile, filename);\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -273,10 +339,12 @@ void plug_bulk_checkin(void)\n \tbulk_checkin_plugged = 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void unplug_bulk_checkin(struct lock_file *lock_file)\n {\n \tassert(bulk_checkin_plugged);\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_sync_and_rename(&bulk_fsync_state, lock_file);\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..8efb01ed669 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,11 +6,13 @@\n \n #include \"cache.h\"\n \n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile, const char *filename);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+void unplug_bulk_checkin(struct lock_file *);\n \n #endif\ndiff --git a/cache.h b/cache.h\nindex bd4869beee4..cde6c6ae6b1 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum FSYNC_OBJECT_FILES_MODE {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex f33abeab851..ab1980f8fec 100644\n--- a/config.c\n+++ b/config.c\n@@ -1509,7 +1509,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcasecmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse\n+\t\t\tfsync_object_files = git_config_bool(var, value)\n+\t\t\t\t? FSYNC_OBJECT_FILES_ON : FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 69413fb3dc0..8c07f2265a8 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n@@ -133,6 +134,7 @@ ifeq ($(uname_S),Darwin)\n \tCOMPAT_OBJS += compat/precompose_utf8.o\n \tBASIC_CFLAGS += -DPRECOMPOSE_UNICODE\n \tBASIC_CFLAGS += -DPROTECT_HFS_DEFAULT=1\n+\tBASIC_CFLAGS += -DFSYNC_DOESNT_FLUSH=1\n \tHAVE_BSD_SYSCTL = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n \tHAVE_NS_GET_EXECUTABLE_PATH = YesPlease\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex d6b22ede7ea..3e23eafff80 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -43,7 +43,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int core_compression_level;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b46605300ab..d14e2436276 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex 5421811273e..94a63809613 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1859,16 +1859,6 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \treturn 0;\n }\n \n-/* Finalize a file on disk, and close it. */\n-static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n-{\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n-\tif (close(fd) != 0)\n-\t\tdie_errno(_(\"error when closing loose object file\"));\n-\treturn finalize_object_file(tmpfile, filename);\n-}\n-\n /* Size of directory component, including the ending '/' */\n static inline int directory_size(const char *filename)\n {\n@@ -1982,7 +1972,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n \t}\n \n-\treturn close_loose_object(fd, tmp_file.buf, filename.buf);\n+\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf, filename.buf);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\ndiff --git a/wrapper.c b/wrapper.c\nindex 563ad590df1..37a8b61a7df 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -538,6 +538,42 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tif (action == FSYNC_WRITEOUT_ONLY) {\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on Mac OS X, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\t}\n+\n+#ifdef __APPLE__\n+\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\treturn fsync(fd);\n+#endif\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex d33e68f6abb..8f53953d4ab 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"433983","messageId":"d8843185fe46e0c4e869159f9539714c5d232810.1630108177.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"[PATCH v2 5/6] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-27T23:49:36Z","receivedAt":"2021-08-27T23:49:48Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality. This mode\nis enabled when passing paths to update-index via the --stdin flag,\nas is done by 'git stash'.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to snoop the output of --verbose to\nfind out when update-index has actually processed a given path.\nAdditionally the index is locked for the duration of the update.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex f1f16f2de52..64d025cf49e 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1152,6 +1153,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n \n \t\tsetup_work_tree();\n+\t\tplug_bulk_checkin();\n \t\twhile (getline_fn(&buf, stdin) != EOF) {\n \t\t\tchar *p;\n \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n@@ -1166,6 +1168,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n \t\t}\n+\t\tunplug_bulk_checkin(&lock_file);\n \t\tstrbuf_release(&unquoted);\n \t\tstrbuf_release(&buf);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"433984","messageId":"73b5d41be94d6e5571cbe5b8dd0d0f74edc4b474.1630108177.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"[PATCH v2 6/6] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-27T23:49:37Z","receivedAt":"2021-08-27T23:49:50Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/lib-unique-files.sh | 32 ++++++++++++++++++++++++++\n t/perf/p3700-add.sh        | 43 +++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh      | 46 ++++++++++++++++++++++++++++++++++++++\n 3 files changed, 121 insertions(+)\n create mode 100644 t/perf/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/lib-unique-files.sh b/t/perf/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..10083395ae5\n--- /dev/null\n+++ b/t/perf/lib-unique-files.sh\n@@ -0,0 +1,32 @@\n+# Helper to create files with unique contents\n+\n+test_create_unique_files_base__=$(date -u)\n+test_create_unique_files_counter__=0\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n+#\t\t\t\t    each in the current directory, all\n+#\t\t\t\t    with unique contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\" > /dev/null\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\ttest_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n+\t\t\techo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..4ca3224f364\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/perf/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..407b95c104b\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/perf/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n"},{"id":"433986","messageId":"546ad9c82e8e0c2eb4683f9f360d8f30e2136020.1630108177.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"[PATCH v2 4/6] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-08-27T23:49:35Z","receivedAt":"2021-08-27T23:49:52Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds a win32 implementation for fsync_no_flush that is\ncalled git_fsync. The 'NtFlushBuffersFileEx' function being called is\navailable since Windows 8. If the function is not available, we\nreturn -1 and Git falls back to doing a full fsync.\n\nThe operating system is told to flush data only without a hardware\nflush primitive. A later full fsync will cause the metadata log\nto be flushed and then the disk cache to be flushed on NTFS and\nReFS. Other filesystems will treat this as a full flush operation.\n\nI added a new file here for this system call so as not to conflict with\ndownstream changes in the git-for-windows repository related to fscache.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.h                      |  3 +++\n compat/win32/flush.c                | 29 +++++++++++++++++++++++++++++\n config.mak.uname                    |  2 ++\n contrib/buildsystems/CMakeLists.txt |  3 ++-\n wrapper.c                           |  4 ++++\n 5 files changed, 40 insertions(+), 1 deletion(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 87944dfec72..b5c950f1e30 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..c013920ce37\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,29 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       /* See https://docs.microsoft.com/en-us/windows-hardware/drivers/ddi/ntifs/nf-ntifs-ntflushbuffersfileex */\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 8c07f2265a8..ef1fd109b74 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -450,6 +450,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -624,6 +625,7 @@ ifneq (,$(findstring MINGW,$(uname_S)))\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 171b4124afe..b573a5ee122 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/wrapper.c b/wrapper.c\nindex 37a8b61a7df..d951306b33e 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -563,6 +563,10 @@ int git_fsync(int fd, enum fsync_action action)\n \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n #endif\n \n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n \t\terrno = ENOSYS;\n \t\treturn -1;\n \t}\n-- \ngitgitgadget\n\n"},{"id":"433989","messageId":"CANQDOdcr0gXsdXtqfN+FFRkAumNfYmr2C3qAcdzFxY26bDPWCQ@mail.gmail.com","threadId":"56371","inReplyTo":"20210826055024.GA17178@lst.de","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-08-28T00:20:44Z","receivedAt":"2021-08-28T00:20:56Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Aug 25, 2021 at 10:50 PM Christoph Hellwig <hch@lst.de> wrote:\n>\n> On Wed, Aug 25, 2021 at 05:49:45PM -0700, Neeraj Singh wrote:\n> > One conclusion from reviewing that thread is that as of then,\n> > sync_file_ranges isn't actually enough\n> > to make a hard guarantee about writeout occurring. See\n> > https://lore.kernel.org/linux-fsdevel/20190319204330.GY26298@dastard/.\n> > My hope is that the Linux FS developers have rectified that shortcoming by now.\n>\n> I'm not sure what shortcoming you mean.  sync_file_ranges is a system\n> call that only causes data writeback.  It never performs metadata write\n> back and thus is not an integrity operation at all.  That is also very\n> clearly documented in the man page.\n>\n\nYou're right. On re-read of the man page, sync_file_range is listed as\nan \"extremely dangerous\"\nsystem call.  The opportunity in the linux kernel is to offer an\nalternative set of flags or separate\nAPI that allows for an application like Git to separate a metadata\nwriteback request from the disk flush.\n\nSeparately, I'm hoping I can push from the Windows filesystem side to\nget a barrier primitive put into\nthe NVME standard so that we can offer more useful behavior to\napplications rather than these painful\nhardware flushes.\n"},{"id":"434011","messageId":"20210828065700.GA31211@lst.de","threadId":"56371","inReplyTo":"CANQDOdcr0gXsdXtqfN+FFRkAumNfYmr2C3qAcdzFxY26bDPWCQ@mail.gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-08-28T06:57:00Z","receivedAt":"2021-08-28T06:57:06Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Fri, Aug 27, 2021 at 05:20:44PM -0700, Neeraj Singh wrote:\n> You're right. On re-read of the man page, sync_file_range is listed as\n> an \"extremely dangerous\"\n> system call.  The opportunity in the linux kernel is to offer an\n> alternative set of flags or separate\n> API that allows for an application like Git to separate a metadata\n> writeback request from the disk flush.\n\nHow do you want to do that?  I metadata writeback without a cache flush\nis worse than useless, in fact it is generally actively harmful.\n\nTo take XFS as an example:  fsync and fdatasync do the following thing:\n\n 1) writeback all dirty data for file to the data device\n 2) flush the write cache of the data device to ensure they are really\n    on disk before writing back the metadata referring to them\n 3) write out the log up until the log sequence that contained the last\n    modifications to the file\n 4) flush the cache for the log device.\n    If the data device and the log device are the same (they usually are\n    for common setups) and the log device support the FUA bit that writes\n    through the cache, the log writes use that bit and this step can\n    be skipped.\n\nSo in general there are very few metadata writes, and it is absolutely\nessential to flush the cache before that, because otherwise your metadata\ncould point to data that might not actually have made it to disk.\n\nThe best way to optimize such a workload is by first batching all the\ndata writeout for multiple fils in step one, then only doing one cache\nflush and one log force (as we call it) to cover all the files.  syncfs\nwill do that, but without a good way to pick individual files.\n\n> Separately, I'm hoping I can push from the Windows filesystem side to\n> get a barrier primitive put into\n> the NVME standard so that we can offer more useful behavior to\n> applications rather than these painful\n> hardware flushes.\n\nI'm not sure what you mean with barriers, but if you mean the concept\nof implying a global ordering on I/Os as we did in Linux back in the\nbad old days the barrier bio flag, or badly reinvented by this paper:\n\n  https://www.usenix.org/conference/fast18/presentation/won\n\nthey might help a little bit with single threaded operations, but will\nheavily degrade I/O performance for multithreaded workloads.  As an\nactive member of (but not speaking for) the NVMe technical working group\nwith a bit of knowledge of SSD internals I also doubt it will be very\nwell received there.\n"},{"id":"434328","messageId":"CANQDOdfV3omEBHOAq1b2P4Wb5=FCtrtsy22VvRu+FneFx9o9Gw@mail.gmail.com","threadId":"56371","inReplyTo":"20210828065700.GA31211@lst.de","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-08-31T19:59:14Z","receivedAt":"2021-08-31T19:59:27Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Aug 27, 2021 at 11:57 PM Christoph Hellwig <hch@lst.de> wrote:\n>\n> On Fri, Aug 27, 2021 at 05:20:44PM -0700, Neeraj Singh wrote:\n> > You're right. On re-read of the man page, sync_file_range is listed as\n> > an \"extremely dangerous\"\n> > system call.  The opportunity in the linux kernel is to offer an\n> > alternative set of flags or separate\n> > API that allows for an application like Git to separate a metadata\n> > writeback request from the disk flush.\n>\n> How do you want to do that?  I metadata writeback without a cache flush\n> is worse than useless, in fact it is generally actively harmful.\n>\n> To take XFS as an example:  fsync and fdatasync do the following thing:\n>\n>  1) writeback all dirty data for file to the data device\n>  2) flush the write cache of the data device to ensure they are really\n>     on disk before writing back the metadata referring to them\n>  3) write out the log up until the log sequence that contained the last\n>     modifications to the file\n>  4) flush the cache for the log device.\n>     If the data device and the log device are the same (they usually are\n>     for common setups) and the log device support the FUA bit that writes\n>     through the cache, the log writes use that bit and this step can\n>     be skipped.\n>\n> So in general there are very few metadata writes, and it is absolutely\n> essential to flush the cache before that, because otherwise your metadata\n> could point to data that might not actually have made it to disk.\n>\n> The best way to optimize such a workload is by first batching all the\n> data writeout for multiple fils in step one, then only doing one cache\n> flush and one log force (as we call it) to cover all the files.  syncfs\n> will do that, but without a good way to pick individual files.\n\nYes, I think we want to do step (1) of your sequence for all of the files, then\nissue steps (2-4) for all files as a group.  Of course, if the log\nfills up then we\ncan flush the intermediate steps.  The unfortunate thing is that\nthere's no Linux interface\nto do step (1) and to also ensure that the relevant data is in the log\nstream or is\notherwise available to be part of the durable metadata.\n\nIt seems to me that XFS would be compatible with this sequence if the\nappropriate\nkernel API exists.\n\n>\n> > Separately, I'm hoping I can push from the Windows filesystem side to\n> > get a barrier primitive put into\n> > the NVME standard so that we can offer more useful behavior to\n> > applications rather than these painful\n> > hardware flushes.\n>\n> I'm not sure what you mean with barriers, but if you mean the concept\n> of implying a global ordering on I/Os as we did in Linux back in the\n> bad old days the barrier bio flag, or badly reinvented by this paper:\n>\n>   https://www.usenix.org/conference/fast18/presentation/won\n>\n> they might help a little bit with single threaded operations, but will\n> heavily degrade I/O performance for multithreaded workloads.  As an\n> active member of (but not speaking for) the NVMe technical working group\n> with a bit of knowledge of SSD internals I also doubt it will be very\n> well received there.\n\nI looked at that paper and definitely agree with you about the questionable\nimplementation strategy they picked. I don't (yet) have detailed knowledge of\nSSD internals, but it's surprising to me that there is little value to\nbarrier semantics\nwithin the drive as opposed to a full durability sync. At least for\nWindows, we have\na database (the Registry) for which any single-threaded latency\nimprovement would\nbe welcome.\n"},{"id":"434389","messageId":"20210901050935.GA13949@lst.de","threadId":"56371","inReplyTo":"CANQDOdfV3omEBHOAq1b2P4Wb5=FCtrtsy22VvRu+FneFx9o9Gw@mail.gmail.com","subject":"Re: [PATCH 2/2] core.fsyncobjectfiles: batch disk flushes","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-09-01T05:09:35Z","receivedAt":"2021-09-01T05:09:40Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Tue, Aug 31, 2021 at 12:59:14PM -0700, Neeraj Singh wrote:\n> > So in general there are very few metadata writes, and it is absolutely\n> > essential to flush the cache before that, because otherwise your metadata\n> > could point to data that might not actually have made it to disk.\n> >\n> > The best way to optimize such a workload is by first batching all the\n> > data writeout for multiple fils in step one, then only doing one cache\n> > flush and one log force (as we call it) to cover all the files.  syncfs\n> > will do that, but without a good way to pick individual files.\n> \n> Yes, I think we want to do step (1) of your sequence for all of the files, then\n> issue steps (2-4) for all files as a group.  Of course, if the log\n> fills up then we\n> can flush the intermediate steps.  The unfortunate thing is that\n> there's no Linux interface\n> to do step (1) and to also ensure that the relevant data is in the log\n> stream or is\n> otherwise available to be part of the durable metadata.\n\nThere is also no interface to do 2-4 separately, mostly because they\nare so hard to separate.  The only API I could envision is one that takes\nan array of file descriptors and has the semantics of doing a fsync/\nfdatasync for all of them, allowing the implementation to optimize\nthe order.  It would be implementable, but not quite as efficient\nas syncfs.  I'm also pretty sure we've seen a few attempts at it in\nthe past that ran into various issues and didn't really make it far.\n\n> I looked at that paper and definitely agree with you about the questionable\n> implementation strategy they picked. I don't (yet) have detailed knowledge of\n> SSD internals, but it's surprising to me that there is little value to\n> barrier semantics\n> within the drive as opposed to a full durability sync. At least for\n> Windows, we have\n> a database (the Registry) for which any single-threaded latency\n> improvement would\n> be welcome.\n\nThe major issue with barrier like in the paper above or as historic\nLinux 2.6 kernels had it is that it enforces a global ordering.  For\nsoftware-only implementation like the Linux one this was already bad\nenough, but a hardware/firmware implementation in nvme would mean you'd\nhave to add global serialize to a storage interface standard and its\nimplementations, while these are very much about offering parallelisms.\nIn fact we'd also have very similar issues with the modern Linux block\nlayer, which has applied some similar ideas.\n"},{"id":"434917","messageId":"CANQDOdeEic1ktyGU=dLEPi=FkU84Oqv9hDUEkfAXcS0WTwRJtQ@mail.gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-07T19:44:13Z","receivedAt":"2021-09-07T19:44:29Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Aug 27, 2021 at 4:49 PM Neeraj K. Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> Thanks to everyone for review so far! I've responded to the previous\n> feedback and changed the patch series a bit.\n>\n> Changes since v1:\n>\n>  * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n>    to dscho's suggestion, I'm still implementing the Windows version in the\n>    same patch and I'm not doing autoconf detection since this is a POSIX\n>    function.\n>\n>  * Introduce a separate preparatory patch to the bulk-checkin infrastructure\n>    to separate the 'plugged' variable and rename the 'state' variable, as\n>    suggested by dscho.\n>\n>  * Add performance numbers to the commit message of the main bulk fsync\n>    patch, as suggested by dscho.\n>\n>  * Add a comment about the non-thread-safety of the bulk-checkin\n>    infrastructure, as suggested by avarab.\n>\n>  * Rename the experimental mode to core.fsyncobjectfiles=batch, as suggested\n>    by dscho and avarab and others.\n>\n>  * Add more details to Documentation/config/core.txt about the various\n>    settings and their intended effects, as suggested by avarab.\n>\n>  * Switch to the string-list API to hold the rename state, as suggested by\n>    avarab.\n>\n>  * Create a separate update-index patch to use bulk-checkin as suggested by\n>    dscho.\n>\n>  * Add Windows support in the upstream git. This is done in a way that\n>    should not conflict with git-for-windows.\n>\n>  * Add new performance tests that shows the delta based on fsync mode.\n>\n> NOTE: Based on Christoph Hellwig's comments, the 'batch' mode is not correct\n> on Linux, since sync_file_range does not provide data integrity guarantees.\n> There is currently no kernel interface suitable to achieve disk flush\n> batching as is, but he suggested that he might implement a 'syncfs' variant\n> on top of this patchset. This code is still useful on macOS and Windows, and\n> the config documentation makes that clear.\n>\n> Neeraj Singh (6):\n>   object-file: use futimens rather than utime\n>   bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n>   core.fsyncobjectfiles: batched disk flushes\n>   core.fsyncobjectfiles: add windows support for batch mode\n>   update-index: use the bulk-checkin infrastructure\n>   core.fsyncobjectfiles: performance tests for add and stash\n>\n>  Documentation/config/core.txt       | 26 ++++++--\n>  Makefile                            |  6 ++\n>  builtin/add.c                       |  3 +-\n>  builtin/update-index.c              |  3 +\n>  bulk-checkin.c                      | 92 +++++++++++++++++++++++++----\n>  bulk-checkin.h                      |  4 +-\n>  cache.h                             |  8 ++-\n>  compat/mingw.c                      | 53 +++++++++++------\n>  compat/mingw.h                      |  5 ++\n>  compat/win32/flush.c                | 29 +++++++++\n>  config.c                            |  8 ++-\n>  config.mak.uname                    |  4 ++\n>  configure.ac                        |  8 +++\n>  contrib/buildsystems/CMakeLists.txt |  3 +-\n>  environment.c                       |  2 +-\n>  git-compat-util.h                   |  7 +++\n>  object-file.c                       | 23 ++------\n>  t/perf/lib-unique-files.sh          | 32 ++++++++++\n>  t/perf/p3700-add.sh                 | 43 ++++++++++++++\n>  t/perf/p3900-stash.sh               | 46 +++++++++++++++\n>  wrapper.c                           | 40 +++++++++++++\n>  write-or-die.c                      |  2 +-\n>  22 files changed, 389 insertions(+), 58 deletions(-)\n>  create mode 100644 compat/win32/flush.c\n>  create mode 100644 t/perf/lib-unique-files.sh\n>  create mode 100755 t/perf/p3700-add.sh\n>  create mode 100755 t/perf/p3900-stash.sh\n>\n>\n> base-commit: 225bc32a989d7a22fa6addafd4ce7dcd04675dbf\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v2\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v2\n> Pull-Request: https://github.com/git/git/pull/1076\n\nHello everyone,\nI'd like to bump this review up in people's inboxes since Patch V2\nhasn't gotten any traction in over a week.\n\nThanks in advance for taking a look,\n- Neeraj Singh\nWindows Core Filesystems Team\n"},{"id":"434921","messageId":"8735qgkvv1.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdeEic1ktyGU=dLEPi=FkU84Oqv9hDUEkfAXcS0WTwRJtQ@mail.gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-07T19:50:48Z","receivedAt":"2021-09-07T19:51:18Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Sep 07 2021, Neeraj Singh wrote:\n\n> On Fri, Aug 27, 2021 at 4:49 PM Neeraj K. Singh via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>>\n>> Thanks to everyone for review so far! I've responded to the previous\n>> feedback and changed the patch series a bit.\n>>\n>> Changes since v1:\n>>\n>>  * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n>>    to dscho's suggestion, I'm still implementing the Windows version in the\n>>    same patch and I'm not doing autoconf detection since this is a POSIX\n>>    function.\n>>\n>>  * Introduce a separate preparatory patch to the bulk-checkin infrastructure\n>>    to separate the 'plugged' variable and rename the 'state' variable, as\n>>    suggested by dscho.\n>>\n>>  * Add performance numbers to the commit message of the main bulk fsync\n>>    patch, as suggested by dscho.\n>>\n>>  * Add a comment about the non-thread-safety of the bulk-checkin\n>>    infrastructure, as suggested by avarab.\n>>\n>>  * Rename the experimental mode to core.fsyncobjectfiles=batch, as suggested\n>>    by dscho and avarab and others.\n>>\n>>  * Add more details to Documentation/config/core.txt about the various\n>>    settings and their intended effects, as suggested by avarab.\n>>\n>>  * Switch to the string-list API to hold the rename state, as suggested by\n>>    avarab.\n>>\n>>  * Create a separate update-index patch to use bulk-checkin as suggested by\n>>    dscho.\n>>\n>>  * Add Windows support in the upstream git. This is done in a way that\n>>    should not conflict with git-for-windows.\n>>\n>>  * Add new performance tests that shows the delta based on fsync mode.\n>>\n>> NOTE: Based on Christoph Hellwig's comments, the 'batch' mode is not correct\n>> on Linux, since sync_file_range does not provide data integrity guarantees.\n>> There is currently no kernel interface suitable to achieve disk flush\n>> batching as is, but he suggested that he might implement a 'syncfs' variant\n>> on top of this patchset. This code is still useful on macOS and Windows, and\n>> the config documentation makes that clear.\n>>\n>> Neeraj Singh (6):\n>>   object-file: use futimens rather than utime\n>>   bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n>>   core.fsyncobjectfiles: batched disk flushes\n>>   core.fsyncobjectfiles: add windows support for batch mode\n>>   update-index: use the bulk-checkin infrastructure\n>>   core.fsyncobjectfiles: performance tests for add and stash\n>>\n>>  Documentation/config/core.txt       | 26 ++++++--\n>>  Makefile                            |  6 ++\n>>  builtin/add.c                       |  3 +-\n>>  builtin/update-index.c              |  3 +\n>>  bulk-checkin.c                      | 92 +++++++++++++++++++++++++----\n>>  bulk-checkin.h                      |  4 +-\n>>  cache.h                             |  8 ++-\n>>  compat/mingw.c                      | 53 +++++++++++------\n>>  compat/mingw.h                      |  5 ++\n>>  compat/win32/flush.c                | 29 +++++++++\n>>  config.c                            |  8 ++-\n>>  config.mak.uname                    |  4 ++\n>>  configure.ac                        |  8 +++\n>>  contrib/buildsystems/CMakeLists.txt |  3 +-\n>>  environment.c                       |  2 +-\n>>  git-compat-util.h                   |  7 +++\n>>  object-file.c                       | 23 ++------\n>>  t/perf/lib-unique-files.sh          | 32 ++++++++++\n>>  t/perf/p3700-add.sh                 | 43 ++++++++++++++\n>>  t/perf/p3900-stash.sh               | 46 +++++++++++++++\n>>  wrapper.c                           | 40 +++++++++++++\n>>  write-or-die.c                      |  2 +-\n>>  22 files changed, 389 insertions(+), 58 deletions(-)\n>>  create mode 100644 compat/win32/flush.c\n>>  create mode 100644 t/perf/lib-unique-files.sh\n>>  create mode 100755 t/perf/p3700-add.sh\n>>  create mode 100755 t/perf/p3900-stash.sh\n>>\n>>\n>> base-commit: 225bc32a989d7a22fa6addafd4ce7dcd04675dbf\n>> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v2\n>> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v2\n>> Pull-Request: https://github.com/git/git/pull/1076\n>\n> Hello everyone,\n> I'd like to bump this review up in people's inboxes since Patch V2\n> hasn't gotten any traction in over a week.\n>\n> Thanks in advance for taking a look,\n> - Neeraj Singh\n> Windows Core Filesystems Team\n\nThanks, I've been meaning to take a look at this, and also as a\nnote-to-self: check how this interacts with the fsync()-impacted race I\nnoted in my just-sent:\nhttps://lore.kernel.org/git/cover-0.3-00000000000-20210907T193600Z-avarab@gmail.com/\n"},{"id":"434924","messageId":"003701d7a422$21c32740$654975c0$@nexbridge.com","threadId":"56371","inReplyTo":"CANQDOdeEic1ktyGU=dLEPi=FkU84Oqv9hDUEkfAXcS0WTwRJtQ@mail.gmail.com","subject":"RE: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-09-07T19:54:07Z","receivedAt":"2021-09-07T19:54:18Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On September 7, 2021 3:44 PM, Neeraj Singh wrote:\n>On Fri, Aug 27, 2021 at 4:49 PM Neeraj K. Singh via GitGitGadget <gitgitgadget@gmail.com> wrote:\n>>\n>> Thanks to everyone for review so far! I've responded to the previous\n>> feedback and changed the patch series a bit.\n>>\n>> Changes since v1:\n>>\n>>  * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n>>    to dscho's suggestion, I'm still implementing the Windows version in the\n>>    same patch and I'm not doing autoconf detection since this is a POSIX\n>>    function.\n\nWhile POSIX.1-2008, this function is not available on every single POSIX-compliant platform. Please make sure that the code will not cause a breakage on some platforms - the ones I maintain, in particular. Neither futimes nor futimens is available on either NonStop ia64 or x86. The platform only has utime, so this needs to be wrapped with an option in config.mak.uname.\n\nThanks,\nRandall\n\n\n"},{"id":"434982","messageId":"CANQDOdcKsUqrQ6K6MEBoXS1BW8_tO8mx4tcq6nvqyiuM4e2CmA@mail.gmail.com","threadId":"56371","inReplyTo":"003701d7a422$21c32740$654975c0$@nexbridge.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-08T00:54:27Z","receivedAt":"2021-09-08T00:54:39Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 7, 2021 at 12:54 PM Randall S. Becker\n<rsbecker@nexbridge.com> wrote:\n>\n> On September 7, 2021 3:44 PM, Neeraj Singh wrote:\n> >On Fri, Aug 27, 2021 at 4:49 PM Neeraj K. Singh via GitGitGadget <gitgitgadget@gmail.com> wrote:\n> >>\n> >> Thanks to everyone for review so far! I've responded to the previous\n> >> feedback and changed the patch series a bit.\n> >>\n> >> Changes since v1:\n> >>\n> >>  * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n> >>    to dscho's suggestion, I'm still implementing the Windows version in the\n> >>    same patch and I'm not doing autoconf detection since this is a POSIX\n> >>    function.\n>\n> While POSIX.1-2008, this function is not available on every single POSIX-compliant platform. Please make sure that the code will not cause a breakage on some platforms - the ones I maintain, in particular. Neither futimes nor futimens is available on either NonStop ia64 or x86. The platform only has utime, so this needs to be wrapped with an option in config.mak.uname.\n>\n> Thanks,\n> Randall\n\nUgh. Fair enough.  How do other contributors feel about me moving back\nto utime, but instead just doing the utime over in\nbuiltins/pack-objects.c?  The idea would be to eliminate the mtime\nlogic entirely from write_loose_object and just do it at the top-level\nin loosen_unused_packed_objects.\n\nThanks,\nNeeraj\n"},{"id":"434983","messageId":"CANQDOdeX-SoWnh5DJ9ZdNLfPdAW-wtp_fo99r0Rwe1DQqx4W5Q@mail.gmail.com","threadId":"56371","inReplyTo":"CANQDOdeEic1ktyGU=dLEPi=FkU84Oqv9hDUEkfAXcS0WTwRJtQ@mail.gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-08T00:55:29Z","receivedAt":"2021-09-08T00:55:41Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 7, 2021 at 12:44 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> Hello everyone,\n> I'd like to bump this review up in people's inboxes since Patch V2\n> hasn't gotten any traction in over a week.\n>\n> Thanks in advance for taking a look,\n> - Neeraj Singh\n> Windows Core Filesystems Team\n\nBTW, I updated the github PR to enable batch mode everywhere, and all\nthe tests passed, which is good news to me.\n\nThanks,\nNeeraj\n"},{"id":"434986","messageId":"87h7evkghf.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdcKsUqrQ6K6MEBoXS1BW8_tO8mx4tcq6nvqyiuM4e2CmA@mail.gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-08T01:22:38Z","receivedAt":"2021-09-08T01:23:29Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Sep 07 2021, Neeraj Singh wrote:\n\n> On Tue, Sep 7, 2021 at 12:54 PM Randall S. Becker\n> <rsbecker@nexbridge.com> wrote:\n>>\n>> On September 7, 2021 3:44 PM, Neeraj Singh wrote:\n>> >On Fri, Aug 27, 2021 at 4:49 PM Neeraj K. Singh via GitGitGadget <gitgitgadget@gmail.com> wrote:\n>> >>\n>> >> Thanks to everyone for review so far! I've responded to the previous\n>> >> feedback and changed the patch series a bit.\n>> >>\n>> >> Changes since v1:\n>> >>\n>> >>  * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n>> >>    to dscho's suggestion, I'm still implementing the Windows version in the\n>> >>    same patch and I'm not doing autoconf detection since this is a POSIX\n>> >>    function.\n>>\n>> While POSIX.1-2008, this function is not available on every single\n>> POSIX-compliant platform. Please make sure that the code will not\n>> cause a breakage on some platforms - the ones I maintain, in\n>> particular. Neither futimes nor futimens is available on either\n>> NonStop ia64 or x86. The platform only has utime, so this needs to\n>> be wrapped with an option in config.mak.uname.\n>>\n>> Thanks,\n>> Randall\n>\n> Ugh. Fair enough.  How do other contributors feel about me moving back\n> to utime, but instead just doing the utime over in\n> builtins/pack-objects.c?  The idea would be to eliminate the mtime\n> logic entirely from write_loose_object and just do it at the top-level\n> in loosen_unused_packed_objects.\n\nAside from where it lives, can't we just have a wrapper that takes both\nthe filename & fd, and then on some platforms will need to dispatch to a\nslower filename-only version, but can hopefully use the new fd-accepting\nfunction?\n"},{"id":"435029","messageId":"xmqqmton7ehn.fsf@gitster.g","threadId":"56371","inReplyTo":"CANQDOdeX-SoWnh5DJ9ZdNLfPdAW-wtp_fo99r0Rwe1DQqx4W5Q@mail.gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-08T06:44:52Z","receivedAt":"2021-09-08T06:45:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> BTW, I updated the github PR to enable batch mode everywhere, and all\n> the tests passed, which is good news to me.\n\nI doubt that fsyncObjectFiles is something we can reliably test in\nCI, either with the new batched thing or with the original \"when we\nclose one, make sure the changes hit the disk platter\" approach.  So\nI am not sure what conclusion we should draw from such an experiment,\nother than \"ok, it compiles cleanly.\"  After all, unless we cause\nsystem crashes, what we thought we have written and close(2) would\nbe seen by another process that we spawn after that, with or without\nsync, no?\n\n\n\n"},{"id":"435030","messageId":"20210908064958.GA29073@lst.de","threadId":"56371","inReplyTo":"xmqqmton7ehn.fsf@gitster.g","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-09-08T06:49:58Z","receivedAt":"2021-09-08T06:50:04Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Tue, Sep 07, 2021 at 11:44:52PM -0700, Junio C Hamano wrote:\n> I doubt that fsyncObjectFiles is something we can reliably test in\n> CI, either with the new batched thing or with the original \"when we\n> close one, make sure the changes hit the disk platter\" approach.  So\n> I am not sure what conclusion we should draw from such an experiment,\n> other than \"ok, it compiles cleanly.\"  After all, unless we cause\n> system crashes, what we thought we have written and close(2) would\n> be seen by another process that we spawn after that, with or without\n> sync, no?\n\nBasically yes.  XFS on Linux has shutdown ioctls that allow to simulate\nthat crash by shutting the file system down which really helps debugging\nthat kind of code.  A bunch of other file systems (ext4, f2fs) have\nalso picked this up now (grep for {XFS,EXT4,F2FS}_IOC_SHUTDOWN).\n"},{"id":"435091","messageId":"006d01d7a4b9$7cea64c0$76bf2e40$@nexbridge.com","threadId":"56371","inReplyTo":"20210908064958.GA29073@lst.de","subject":"RE: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-09-08T13:57:34Z","receivedAt":"2021-09-08T13:57:43Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On September 8, 2021 2:50 AM, Christoph Hellwig wrote:\n>To: Junio C Hamano <gitster@pobox.com>\n>Cc: Neeraj Singh <nksingh85@gmail.com>; Neeraj K. Singh via GitGitGadget <gitgitgadget@gmail.com>; Git List <git@vger.kernel.org>;\n>Johannes Schindelin <Johannes.Schindelin@gmx.de>; Jeff King <peff@peff.net>; Jeff Hostetler <jeffhost@microsoft.com>; Christoph\n>Hellwig <hch@lst.de>; Ævar Arnfjörð Bjarmason <avarab@gmail.com>; Neeraj K. Singh <neerajsi@microsoft.com>\n>Subject: Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles\n>\n>On Tue, Sep 07, 2021 at 11:44:52PM -0700, Junio C Hamano wrote:\n>> I doubt that fsyncObjectFiles is something we can reliably test in CI,\n>> either with the new batched thing or with the original \"when we close\n>> one, make sure the changes hit the disk platter\" approach.  So I am\n>> not sure what conclusion we should draw from such an experiment, other\n>> than \"ok, it compiles cleanly.\"  After all, unless we cause system\n>> crashes, what we thought we have written and close(2) would be seen by\n>> another process that we spawn after that, with or without sync, no?\n>\n>Basically yes.  XFS on Linux has shutdown ioctls that allow to simulate that crash by shutting the file system down which really\nhelps\n>debugging that kind of code.  A bunch of other file systems (ext4, f2fs) have also picked this up now (grep for\n>{XFS,EXT4,F2FS}_IOC_SHUTDOWN).\n\nI strongly doubt this concept will work in an MPP architecture, particularly one where \"shutting the file system down\" is not\npossible. I know of at least 3 operating systems where that is a bad plan, and if you did, you would take the test suite down while\nyou were at it.\n-Randall\n\n"},{"id":"435092","messageId":"006e01d7a4ba$72191e00$564b5a00$@nexbridge.com","threadId":"56371","inReplyTo":"87h7evkghf.fsf@evledraar.gmail.com","subject":"RE: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-09-08T14:04:26Z","receivedAt":"2021-09-08T14:04:35Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On September 7, 2021 9:23 PM, Ævar Arnfjörð Bjarmason wrote:\n>Subject: Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles\n>\n>\n>On Tue, Sep 07 2021, Neeraj Singh wrote:\n>\n>> On Tue, Sep 7, 2021 at 12:54 PM Randall S. Becker\n>> <rsbecker@nexbridge.com> wrote:\n>>>\n>>> On September 7, 2021 3:44 PM, Neeraj Singh wrote:\n>>> >On Fri, Aug 27, 2021 at 4:49 PM Neeraj K. Singh via GitGitGadget <gitgitgadget@gmail.com> wrote:\n>>> >>\n>>> >> Thanks to everyone for review so far! I've responded to the\n>>> >> previous feedback and changed the patch series a bit.\n>>> >>\n>>> >> Changes since v1:\n>>> >>\n>>> >>  * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n>>> >>    to dscho's suggestion, I'm still implementing the Windows version in the\n>>> >>    same patch and I'm not doing autoconf detection since this is a POSIX\n>>> >>    function.\n>>>\n>>> While POSIX.1-2008, this function is not available on every single\n>>> POSIX-compliant platform. Please make sure that the code will not\n>>> cause a breakage on some platforms - the ones I maintain, in\n>>> particular. Neither futimes nor futimens is available on either\n>>> NonStop ia64 or x86. The platform only has utime, so this needs to be\n>>> wrapped with an option in config.mak.uname.\n>>>\n>>> Thanks,\n>>> Randall\n>>\n>> Ugh. Fair enough.  How do other contributors feel about me moving back\n>> to utime, but instead just doing the utime over in\n>> builtins/pack-objects.c?  The idea would be to eliminate the mtime\n>> logic entirely from write_loose_object and just do it at the top-level\n>> in loosen_unused_packed_objects.\n>\n>Aside from where it lives, can't we just have a wrapper that takes both the filename & fd, and then on some platforms will need to\n>dispatch to a slower filename-only version, but can hopefully use the new fd-accepting function?\n\nI'm not really enamoured with this direction at all. It means that any platform would have to potentially skip a version of git (resulting from the broken build from a wrapper that is not compilable) after the patches were applied, unless the patches for all of those platforms are included. Even adding a Makefile option would be similar. This should be an \"Enable if supported\" feature, not a default-or-broken, feature. At best, I'd have to monitor for the time where the patch is applied and hope I can figure out the wrapper changes (around my $DAYJOB) in time to make the same release. This seems a bit counter to a \"keeping things compatible\" philosophy. Maybe there's something I'm missing here.\n-Randall\n\n"},{"id":"435093","messageId":"20210908141322.GA27146@lst.de","threadId":"56371","inReplyTo":"006d01d7a4b9$7cea64c0$76bf2e40$@nexbridge.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"'Christoph Hellwig'","fromEmail":"hch@lst.de","sentAt":"2021-09-08T14:13:22Z","receivedAt":"2021-09-08T14:13:27Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Wed, Sep 08, 2021 at 09:57:34AM -0400, Randall S. Becker wrote:\n> possible. I know of at least 3 operating systems where that is a bad plan, and if you did, you would take the test suite down while\n> you were at it.\n\nI've just mentioned a good way to write a test for this feature on a\nspecific platform.  This is absolutely no judgement if that is a good\nplan on other platforms.\n"},{"id":"435095","messageId":"007b01d7a4bd$72921730$57b64590$@nexbridge.com","threadId":"56371","inReplyTo":"20210908141322.GA27146@lst.de","subject":"RE: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-09-08T14:25:55Z","receivedAt":"2021-09-08T14:26:10Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On September 8, 2021 10:13 AM, Christoph Hellwig wrote:\n>Subject: Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles\n>\n>On Wed, Sep 08, 2021 at 09:57:34AM -0400, Randall S. Becker wrote:\n>> possible. I know of at least 3 operating systems where that is a bad\n>> plan, and if you did, you would take the test suite down while you were at it.\n>\n>I've just mentioned a good way to write a test for this feature on a specific platform.  This is absolutely no judgement if that is\na good plan\n>on other platforms.\n\nThank you for the clarification. I do appreciate it.\n-Randall\n\n"},{"id":"435119","messageId":"CANQDOddQsf4Jj+634mdnJXaPG=2idCbCHd1iXO2qm1EMGcDmXg@mail.gmail.com","threadId":"56371","inReplyTo":"xmqqmton7ehn.fsf@gitster.g","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-08T16:34:04Z","receivedAt":"2021-09-08T16:34:23Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 7, 2021 at 11:44 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > BTW, I updated the github PR to enable batch mode everywhere, and all\n> > the tests passed, which is good news to me.\n>\n> I doubt that fsyncObjectFiles is something we can reliably test in\n> CI, either with the new batched thing or with the original \"when we\n> close one, make sure the changes hit the disk platter\" approach.  So\n> I am not sure what conclusion we should draw from such an experiment,\n> other than \"ok, it compiles cleanly.\"  After all, unless we cause\n> system crashes, what we thought we have written and close(2) would\n> be seen by another process that we spawn after that, with or without\n> sync, no?\n\nThe main failure mode I was worried about is that some test or other part\nof Git is relying on a loose object being immediately available after it is\nadded to the ODB. With batch mode, the loose objects aren't actually\navailable until the bulk checkin is unplugged.\n\nI agree that it is not easy to test whether the data is actually going\nto durable\nstorage at the expected time.  FWIW, I did take a disk IO trace on Windows to\nverify that we are issuing disk writes and flushes at the right time.\nBut that's a\none-time test that would be hard to make automated.\n"},{"id":"435140","messageId":"CANQDOddSa4KguS0phCcZwK6FneSyLQQkWdrTbNM0+7dGVKVQpw@mail.gmail.com","threadId":"56371","inReplyTo":"87h7evkghf.fsf@evledraar.gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-08T19:01:10Z","receivedAt":"2021-09-08T19:01:28Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 7, 2021 at 6:23 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Tue, Sep 07 2021, Neeraj Singh wrote:\n>\n> > On Tue, Sep 7, 2021 at 12:54 PM Randall S. Becker\n> > <rsbecker@nexbridge.com> wrote:\n> >>\n> >> On September 7, 2021 3:44 PM, Neeraj Singh wrote:\n> >> >On Fri, Aug 27, 2021 at 4:49 PM Neeraj K. Singh via GitGitGadget <gitgitgadget@gmail.com> wrote:\n> >> >>\n> >> >> Thanks to everyone for review so far! I've responded to the previous\n> >> >> feedback and changed the patch series a bit.\n> >> >>\n> >> >> Changes since v1:\n> >> >>\n> >> >>  * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n> >> >>    to dscho's suggestion, I'm still implementing the Windows version in the\n> >> >>    same patch and I'm not doing autoconf detection since this is a POSIX\n> >> >>    function.\n> >>\n> >> While POSIX.1-2008, this function is not available on every single\n> >> POSIX-compliant platform. Please make sure that the code will not\n> >> cause a breakage on some platforms - the ones I maintain, in\n> >> particular. Neither futimes nor futimens is available on either\n> >> NonStop ia64 or x86. The platform only has utime, so this needs to\n> >> be wrapped with an option in config.mak.uname.\n> >>\n> >> Thanks,\n> >> Randall\n> >\n> > Ugh. Fair enough.  How do other contributors feel about me moving back\n> > to utime, but instead just doing the utime over in\n> > builtins/pack-objects.c?  The idea would be to eliminate the mtime\n> > logic entirely from write_loose_object and just do it at the top-level\n> > in loosen_unused_packed_objects.\n>\n> Aside from where it lives, can't we just have a wrapper that takes both\n> the filename & fd, and then on some platforms will need to dispatch to a\n> slower filename-only version, but can hopefully use the new fd-accepting\n> function?\n\nI had some concerns around using utime() while a file descriptor is open.\nThere's some risk of sharing violation on Windows (doesn't matter since we'd\nbe using futimens), but I was also concerned that there might be some OSes that\nupdate the mtime on close(fd), thus overwriting the effects of utime.\nMaybe that's an unwarranted concern, but it's part of why I didn't want to have\ndifferent call sequences on different OSes.\n\nI'd be happy to implement your suggestion though and see what happens. But I\nalso feel that this time update thing is pretty ancillary to the real\ngoal of my change.\nI'm only doing it because it's in the same area. The effects of\ngetting mtime wrong\nwould be pretty subtle -- I think we'd just not be deleting some\nunpacked unreachable\nobjects as soon as expected.  Do you have a strong objection to\nlifting the time update\nlogic out?\n"},{"id":"435144","messageId":"xmqqr1dy3mq5.fsf@gitster.g","threadId":"56371","inReplyTo":"CANQDOddQsf4Jj+634mdnJXaPG=2idCbCHd1iXO2qm1EMGcDmXg@mail.gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-08T19:12:50Z","receivedAt":"2021-09-08T19:12:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> On Tue, Sep 7, 2021 at 11:44 PM Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Neeraj Singh <nksingh85@gmail.com> writes:\n>>\n>> > BTW, I updated the github PR to enable batch mode everywhere, and all\n>> > the tests passed, which is good news to me.\n>>\n>> I doubt that fsyncObjectFiles is something we can reliably test in\n>> CI, either with the new batched thing or with the original \"when we\n>> close one, make sure the changes hit the disk platter\" approach.  So\n>> I am not sure what conclusion we should draw from such an experiment,\n>> other than \"ok, it compiles cleanly.\"  After all, unless we cause\n>> system crashes, what we thought we have written and close(2) would\n>> be seen by another process that we spawn after that, with or without\n>> sync, no?\n>\n> The main failure mode I was worried about is that some test or other part\n> of Git is relying on a loose object being immediately available after it is\n> added to the ODB. With batch mode, the loose objects aren't actually\n> available until the bulk checkin is unplugged.\n\nAh, I see.  If there are two processes that communicate over pipes\nto decide whose turn it is (perhaps a producer of data that feeds\nfast-import may wait for fast-import to say \"I gave this label to\nthe object you requested\" and goes ahead to use that object), and at\nthe point that the \"other\" process takes its turn, if the objects\nare not \"flushed\" yet, things can break.  That's a valid concern.\n"},{"id":"435147","messageId":"CANQDOddtfO20dysG=p2g2CVHZCAVWAr7=T-srSristJ=VGirGw@mail.gmail.com","threadId":"56371","inReplyTo":"xmqqr1dy3mq5.fsf@gitster.g","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-08T19:20:14Z","receivedAt":"2021-09-08T19:20:30Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Sep 8, 2021 at 12:12 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > On Tue, Sep 7, 2021 at 11:44 PM Junio C Hamano <gitster@pobox.com> wrote:\n> >>\n> >> Neeraj Singh <nksingh85@gmail.com> writes:\n> >>\n> >> > BTW, I updated the github PR to enable batch mode everywhere, and all\n> >> > the tests passed, which is good news to me.\n> >>\n> >> I doubt that fsyncObjectFiles is something we can reliably test in\n> >> CI, either with the new batched thing or with the original \"when we\n> >> close one, make sure the changes hit the disk platter\" approach.  So\n> >> I am not sure what conclusion we should draw from such an experiment,\n> >> other than \"ok, it compiles cleanly.\"  After all, unless we cause\n> >> system crashes, what we thought we have written and close(2) would\n> >> be seen by another process that we spawn after that, with or without\n> >> sync, no?\n> >\n> > The main failure mode I was worried about is that some test or other part\n> > of Git is relying on a loose object being immediately available after it is\n> > added to the ODB. With batch mode, the loose objects aren't actually\n> > available until the bulk checkin is unplugged.\n>\n> Ah, I see.  If there are two processes that communicate over pipes\n> to decide whose turn it is (perhaps a producer of data that feeds\n> fast-import may wait for fast-import to say \"I gave this label to\n> the object you requested\" and goes ahead to use that object), and at\n> the point that the \"other\" process takes its turn, if the objects\n> are not \"flushed\" yet, things can break.  That's a valid concern.\n\nThat's right. This appears to be a possibility in the existing bulk\ncheckin code that produces packfiles for large objects as well, but\nmy change makes the situation much more common.\n"},{"id":"435169","messageId":"87k0jqhnji.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOddQsf4Jj+634mdnJXaPG=2idCbCHd1iXO2qm1EMGcDmXg@mail.gmail.com","subject":"Re: [PATCH v2 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-08T19:23:05Z","receivedAt":"2021-09-08T19:31:35Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Sep 08 2021, Neeraj Singh wrote:\n\n> On Tue, Sep 7, 2021 at 11:44 PM Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Neeraj Singh <nksingh85@gmail.com> writes:\n>>\n>> > BTW, I updated the github PR to enable batch mode everywhere, and all\n>> > the tests passed, which is good news to me.\n>>\n>> I doubt that fsyncObjectFiles is something we can reliably test in\n>> CI, either with the new batched thing or with the original \"when we\n>> close one, make sure the changes hit the disk platter\" approach.  So\n>> I am not sure what conclusion we should draw from such an experiment,\n>> other than \"ok, it compiles cleanly.\"  After all, unless we cause\n>> system crashes, what we thought we have written and close(2) would\n>> be seen by another process that we spawn after that, with or without\n>> sync, no?\n>\n> The main failure mode I was worried about is that some test or other part\n> of Git is relying on a loose object being immediately available after it is\n> added to the ODB. With batch mode, the loose objects aren't actually\n> available until the bulk checkin is unplugged.\n>\n> I agree that it is not easy to test whether the data is actually going\n> to durable\n> storage at the expected time.  FWIW, I did take a disk IO trace on Windows to\n> verify that we are issuing disk writes and flushes at the right time.\n> But that's a\n> one-time test that would be hard to make automated.\n\nI have some semi-related patches I need to dig up and finish sometime\nwhich add a \"git gc\" test mode to the test suite, i.e. any time we call\n\"git gc --auto\" it will go ahead and actually run, and some adversarial\noptions to run always, right away, prune with --expire=now. It found\nsome false positives, but also some genuine races and bugs at the time.\n\nSimilarly, I think a good longer term goal for better fsync() and data\nintegrity in git is to refactor the various codepaths where we write to\ndisk (grepping for fsync_or_die() is a good start to find those) to all\nlive in one place, we could then easily instrument that code to run in a\nhostile test mode.\n\nE.g. make anything that expects to write out a \"foo\" file actually write\nout \"foo.not-synced-yet\" as long as fsync() etc. hasn't been called, or\nwith signals/timers/atexit() handlers fake up known FS edge cases such\nas a write of \"foo\" only renaming \"foo.not-synced-yet\" to \"foo\" 1s after\nthe last close() call not followed by an fsync, etc.\n\nAnyway, I expect given your occupation that you may have better ideas in\nthat area, presumably needing to instrument and test behavior under I/O\npressure, deferred syncs etc. is something mature FS's need to deal with\nas part of their own regression tests...\n\n1. https://lore.kernel.org/git/cover-v2-0.4-0000000000-20210908T003631Z-avarab@gmail.com/\n"},{"id":"435808","messageId":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v2.git.git.1630108177.gitgitgadget@gmail.com","subject":"[PATCH v3 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-14T03:38:39Z","receivedAt":"2021-09-14T03:38:49Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far!\n\nChanges since v2:\n\n * Removed an unused Makefile define (FSYNC_DOESNT_FLUSH) that slipped in\n   from an intermediate change.\n\n * Drop the futimens part of the patch and return to just calling utime, now\n   within the new bulk_checkin code. The utime to futimens change seemed to\n   be problematic for some platforms (thanks Randall Becker), and is really\n   orthogonal to the rest of the patch series.\n\n * (Optional commit) Enable batch mode by default so that we can shake loose\n   any issues relating to deferring the renames until the\n   unplug_bulk_checkin.\n\nChanges since v1:\n\n * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n   to dscho's suggestion, I'm still implementing the Windows version in the\n   same patch and I'm not doing autoconf detection since this is a POSIX\n   function.\n\n * Introduce a separate preparatory patch to the bulk-checkin infrastructure\n   to separate the 'plugged' variable and rename the 'state' variable, as\n   suggested by dscho.\n\n * Add performance numbers to the commit message of the main bulk fsync\n   patch, as suggested by dscho.\n\n * Add a comment about the non-thread-safety of the bulk-checkin\n   infrastructure, as suggested by avarab.\n\n * Rename the experimental mode to core.fsyncobjectfiles=batch, as suggested\n   by dscho and avarab and others.\n\n * Add more details to Documentation/config/core.txt about the various\n   settings and their intended effects, as suggested by avarab.\n\n * Switch to the string-list API to hold the rename state, as suggested by\n   avarab.\n\n * Create a separate update-index patch to use bulk-checkin as suggested by\n   dscho.\n\n * Add Windows support in the upstream git. This is done in a way that\n   should not conflict with git-for-windows.\n\n * Add new performance tests that shows the delta based on fsync mode.\n\nNOTE: Based on Christoph Hellwig's comments, the 'batch' mode is not correct\non Linux, since sync_file_range does not provide data integrity guarantees.\nThere is currently no kernel interface suitable to achieve disk flush\nbatching as is, but he suggested that he might implement a 'syncfs' variant\non top of this patchset. This code is still useful on macOS and Windows, and\nthe config documentation makes that clear.\n\nNeeraj Singh (6):\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  core.fsyncobjectfiles: add windows support for batch mode\n  update-index: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: performance tests for add and stash\n  core.fsyncobjectfiles: enable batch mode for testing\n\n Documentation/config/core.txt       |  26 +++++--\n Makefile                            |   6 ++\n builtin/add.c                       |   3 +-\n builtin/update-index.c              |   3 +\n bulk-checkin.c                      | 103 +++++++++++++++++++++++++---\n bulk-checkin.h                      |   5 +-\n cache.h                             |   8 ++-\n compat/mingw.h                      |   3 +\n compat/win32/flush.c                |  29 ++++++++\n config.c                            |   8 ++-\n config.mak.uname                    |   3 +\n configure.ac                        |   8 +++\n contrib/buildsystems/CMakeLists.txt |   3 +-\n environment.c                       |   2 +-\n git-compat-util.h                   |   7 ++\n object-file.c                       |  22 +-----\n t/perf/lib-unique-files.sh          |  32 +++++++++\n t/perf/p3700-add.sh                 |  43 ++++++++++++\n t/perf/p3900-stash.sh               |  46 +++++++++++++\n wrapper.c                           |  40 +++++++++++\n write-or-die.c                      |   2 +-\n 21 files changed, 358 insertions(+), 44 deletions(-)\n create mode 100644 compat/win32/flush.c\n create mode 100644 t/perf/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: 8b7c11b8668b4e774f81a9f0b4c30144b818f1d1\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v3\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v2:\n\n 1:  fc3d5a7b635 < -:  ----------- object-file: use futimens rather than utime\n 2:  49f72800bfb = 1:  d5893e28df1 bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n 3:  2c1c907b12a ! 2:  f8b5b709e9e core.fsyncobjectfiles: batched disk flushes\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      +}\n      +\n      +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n     -+\t\t\t\t\t      const char *filename)\n     ++\t\t\t\t\t      const char *filename, time_t mtime)\n      +{\n     ++\tint do_finalize = 1;\n     ++\tint ret = 0;\n     ++\n      +\tif (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n      +\t\t/*\n      +\t\t * If we have a plugged bulk checkin, we issue a call that\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      +\t\t    fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n      +\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n      +\t\t\tadd_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n     -+\t\t\tif (close(fd))\n     -+\t\t\t\tdie_errno(_(\"error when closing loose object file\"));\n     -+\n     -+\t\t\treturn 0;\n     ++\t\t\tdo_finalize = 0;\n      +\n      +\t\t} else {\n      +\t\t\tfsync_or_die(fd, \"loose object file\");\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      +\tif (close(fd))\n      +\t\tdie_errno(_(\"error when closing loose object file\"));\n      +\n     -+\treturn finalize_object_file(tmpfile, filename);\n     ++\tif (mtime) {\n     ++\t\tstruct utimbuf utb;\n     ++\t\tutb.actime = mtime;\n     ++\t\tutb.modtime = mtime;\n     ++\t\tif (utime(tmpfile, &utb) < 0)\n     ++\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n     ++\t}\n     ++\n     ++\tif (do_finalize)\n     ++\t\tret = finalize_object_file(tmpfile, filename);\n     ++\n     ++\treturn ret;\n      +}\n      +\n       int index_bulk_checkin(struct object_id *oid,\n     @@ bulk-checkin.h\n       \n       #include \"cache.h\"\n       \n     -+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile, const char *filename);\n     ++int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n     ++\t\t\t\t\t      const char *filename, time_t mtime);\n      +\n       int index_bulk_checkin(struct object_id *oid,\n       \t\t       int fd, size_t size, enum object_type type,\n     @@ config.mak.uname: ifeq ($(uname_S),Linux)\n       \tHAVE_GETDELIM = YesPlease\n       \tSANE_TEXT_GREP=-a\n       \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n     -@@ config.mak.uname: ifeq ($(uname_S),Darwin)\n     - \tCOMPAT_OBJS += compat/precompose_utf8.o\n     - \tBASIC_CFLAGS += -DPRECOMPOSE_UNICODE\n     - \tBASIC_CFLAGS += -DPROTECT_HFS_DEFAULT=1\n     -+\tBASIC_CFLAGS += -DFSYNC_DOESNT_FLUSH=1\n     - \tHAVE_BSD_SYSCTL = YesPlease\n     - \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n     - \tHAVE_NS_GET_EXECUTABLE_PATH = YesPlease\n      \n       ## configure.ac ##\n      @@ configure.ac: AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n     @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void\n       }\n       \n      -/* Finalize a file on disk, and close it. */\n     --static int close_loose_object(int fd, const char *tmpfile, const char *filename)\n     +-static void close_loose_object(int fd)\n      -{\n      -\tif (fsync_object_files)\n      -\t\tfsync_or_die(fd, \"loose object file\");\n      -\tif (close(fd) != 0)\n      -\t\tdie_errno(_(\"error when closing loose object file\"));\n     --\treturn finalize_object_file(tmpfile, filename);\n      -}\n      -\n       /* Size of directory component, including the ending '/' */\n       static inline int directory_size(const char *filename)\n       {\n      @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n     - \t\t\twarning_errno(_(\"failed futimes() on %s\"), tmp_file.buf);\n     - \t}\n     + \t\tdie(_(\"confused by unstable object source data for %s\"),\n     + \t\t    oid_to_hex(oid));\n       \n     --\treturn close_loose_object(fd, tmp_file.buf, filename.buf);\n     -+\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf, filename.buf);\n     +-\tclose_loose_object(fd);\n     +-\n     +-\tif (mtime) {\n     +-\t\tstruct utimbuf utb;\n     +-\t\tutb.actime = mtime;\n     +-\t\tutb.modtime = mtime;\n     +-\t\tif (utime(tmp_file.buf, &utb) < 0)\n     +-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n     +-\t}\n     +-\n     +-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n     ++\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf,\n     ++\t\t\t\t\t\t\t filename.buf, mtime);\n       }\n       \n       static int freshen_loose_object(const struct object_id *oid)\n 4:  546ad9c82e8 = 3:  815a862e229 core.fsyncobjectfiles: add windows support for batch mode\n 5:  d8843185fe4 = 4:  6b576038986 update-index: use the bulk-checkin infrastructure\n 6:  73b5d41be94 = 5:  b7ca3ba9302 core.fsyncobjectfiles: performance tests for add and stash\n -:  ----------- > 6:  55a40fc8fd5 core.fsyncobjectfiles: enable batch mode for testing\n\n-- \ngitgitgadget\n"},{"id":"435809","messageId":"d5893e28df152ce9b843d4709ed9349156633b3d.1631590725.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"[PATCH v3 1/6] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-14T03:38:40Z","receivedAt":"2021-09-14T03:39:03Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b023d9959aa..f117d62c908 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_bulk_checkin(struct bulk_checkin_state *state)\n {\n@@ -260,21 +260,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"435811","messageId":"f8b5b709e9edc363b2de7d4afa443deec0120ca0.1631590725.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"[PATCH v3 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-14T03:38:41Z","receivedAt":"2021-09-14T03:39:03Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nmacOS, and Linux each offer mechanisms to write data from the filesystem\npage cache without initiating a hardware flush.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\ntakes advantage of the bulk-checkin infrastructure to batch up hardware\nflushes.\n\nWhen the new mode is enabled we do the following for new objects:\n\n1. Create a tmp_obj_XXXX file and write the object data to it.\n2. Issue a pagecache writeback request and wait for it to complete.\n3. Record the tmp name and the final name in the bulk-checkin state for\n   later rename.\n\nAt the end of the entire transaction we:\n1. Issue a fsync against the lock file to flush the hardware writeback\n   cache, which should by now have processed the tmp file writes.\n2. Rename all of the temp files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\nwould expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 26 ++++++++---\n Makefile                      |  6 +++\n builtin/add.c                 |  3 +-\n bulk-checkin.c                | 81 ++++++++++++++++++++++++++++++++++-\n bulk-checkin.h                |  5 ++-\n cache.h                       |  8 +++-\n config.c                      |  8 +++-\n config.mak.uname              |  1 +\n configure.ac                  |  8 ++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 +++\n object-file.c                 | 22 +---------\n wrapper.c                     | 36 ++++++++++++++++\n write-or-die.c                |  2 +-\n 14 files changed, 182 insertions(+), 33 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..0006d90980d 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,26 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols the object store, so updates to any refs or the\n+\tindex may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write object data with a minimal set of FLUSH CACHE\n+  (or equivalent) commands sent to the storage controller. If the operating\n+  system interfaces are not available, this mode behaves the same as `true`.\n+  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n+  filesystems and on Windows for repos stored on NTFS or ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex 429c276058d..326c7607e0f 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1896,6 +1898,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 2244311d485..dda4bf093a0 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -678,7 +678,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\n+\tunplug_bulk_checkin(&lock_file);\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f117d62c908..ddbab5e5c8c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,15 +3,19 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n \n+static struct string_list bulk_fsync_state = STRING_LIST_INIT_DUP;\n+\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -62,6 +66,32 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n+{\n+\tif (fsync_state->nr) {\n+\t\tstruct string_list_item *rename;\n+\n+\t\t/*\n+\t\t * Issue a full hardware flush against the lock file to ensure\n+\t\t * that all objects are durable before any renames occur.\n+\t\t * The code in fsync_and_close_loose_object_bulk_checkin has\n+\t\t * already ensured that writeout has occurred, but it has not\n+\t\t * flushed any writeback cache in the storage hardware.\n+\t\t */\n+\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n+\n+\t\tfor_each_string_list_item(rename, fsync_state) {\n+\t\t\tconst char *src = rename->string;\n+\t\t\tconst char *dst = rename->util;\n+\n+\t\t\tif (finalize_object_file(src, dst))\n+\t\t\t\tdie_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n+\t\t}\n+\n+\t\tstring_list_clear(fsync_state, 1);\n+\t}\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -256,6 +286,53 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+static void add_rename_bulk_checkin(struct string_list *fsync_state,\n+\t\t\t\t    const char *src, const char *dst)\n+{\n+\tstring_list_insert(fsync_state, src)->util = xstrdup(dst);\n+}\n+\n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n+\t\t\t\t\t      const char *filename, time_t mtime)\n+{\n+\tint do_finalize = 1;\n+\tint ret = 0;\n+\n+\tif (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n+\t\t/*\n+\t\t * If we have a plugged bulk checkin, we issue a call that\n+\t\t * cleans the filesystem page cache but avoids a hardware flush\n+\t\t * command. Later on we will issue a single hardware flush\n+\t\t * before renaming files as part of do_sync_and_rename.\n+\t\t */\n+\t\tif (bulk_checkin_plugged &&\n+\t\t    fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n+\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\t\tadd_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n+\t\t\tdo_finalize = 0;\n+\n+\t\t} else {\n+\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\t}\n+\t}\n+\n+\tif (close(fd))\n+\t\tdie_errno(_(\"error when closing loose object file\"));\n+\n+\tif (mtime) {\n+\t\tstruct utimbuf utb;\n+\t\tutb.actime = mtime;\n+\t\tutb.modtime = mtime;\n+\t\tif (utime(tmpfile, &utb) < 0)\n+\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n+\t}\n+\n+\tif (do_finalize)\n+\t\tret = finalize_object_file(tmpfile, filename);\n+\n+\treturn ret;\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -273,10 +350,12 @@ void plug_bulk_checkin(void)\n \tbulk_checkin_plugged = 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void unplug_bulk_checkin(struct lock_file *lock_file)\n {\n \tassert(bulk_checkin_plugged);\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_sync_and_rename(&bulk_fsync_state, lock_file);\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..4a3309c1531 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,11 +6,14 @@\n \n #include \"cache.h\"\n \n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n+\t\t\t\t\t      const char *filename, time_t mtime);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+void unplug_bulk_checkin(struct lock_file *);\n \n #endif\ndiff --git a/cache.h b/cache.h\nindex d23de693680..39b3a88181a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum FSYNC_OBJECT_FILES_MODE {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex cb4a8058bff..9fe3602e1c4 100644\n--- a/config.c\n+++ b/config.c\n@@ -1509,7 +1509,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tif (!strcasecmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse\n+\t\t\tfsync_object_files = git_config_bool(var, value)\n+\t\t\t\t? FSYNC_OBJECT_FILES_ON : FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 76516aaa9a5..e6d482fbcc6 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex d6b22ede7ea..3e23eafff80 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -43,7 +43,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int core_compression_level;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b46605300ab..d14e2436276 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex a8be8994814..ea14c3a3483 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1859,15 +1859,6 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \treturn 0;\n }\n \n-/* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd)\n-{\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n-\tif (close(fd) != 0)\n-\t\tdie_errno(_(\"error when closing loose object file\"));\n-}\n-\n /* Size of directory component, including the ending '/' */\n static inline int directory_size(const char *filename)\n {\n@@ -1973,17 +1964,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n-\n-\tif (mtime) {\n-\t\tstruct utimbuf utb;\n-\t\tutb.actime = mtime;\n-\t\tutb.modtime = mtime;\n-\t\tif (utime(tmp_file.buf, &utb) < 0)\n-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n-\t}\n-\n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf,\n+\t\t\t\t\t\t\t filename.buf, mtime);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\ndiff --git a/wrapper.c b/wrapper.c\nindex 7c6586af321..cffe24d307a 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -540,6 +540,42 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tif (action == FSYNC_WRITEOUT_ONLY) {\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on Mac OS X, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\t}\n+\n+#ifdef __APPLE__\n+\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\treturn fsync(fd);\n+#endif\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex d33e68f6abb..8f53953d4ab 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"435810","messageId":"815a862e22940690b3db9a6fbbbda35029c88f66.1631590725.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"[PATCH v3 3/6] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-14T03:38:42Z","receivedAt":"2021-09-14T03:39:06Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds a win32 implementation for fsync_no_flush that is\ncalled git_fsync. The 'NtFlushBuffersFileEx' function being called is\navailable since Windows 8. If the function is not available, we\nreturn -1 and Git falls back to doing a full fsync.\n\nThe operating system is told to flush data only without a hardware\nflush primitive. A later full fsync will cause the metadata log\nto be flushed and then the disk cache to be flushed on NTFS and\nReFS. Other filesystems will treat this as a full flush operation.\n\nI added a new file here for this system call so as not to conflict with\ndownstream changes in the git-for-windows repository related to fscache.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.h                      |  3 +++\n compat/win32/flush.c                | 29 +++++++++++++++++++++++++++++\n config.mak.uname                    |  2 ++\n contrib/buildsystems/CMakeLists.txt |  3 ++-\n wrapper.c                           |  4 ++++\n 5 files changed, 40 insertions(+), 1 deletion(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..c013920ce37\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,29 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       /* See https://docs.microsoft.com/en-us/windows-hardware/drivers/ddi/ntifs/nf-ntifs-ntflushbuffersfileex */\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex e6d482fbcc6..34c93314a50 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -451,6 +451,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -626,6 +627,7 @@ ifneq (,$(findstring MINGW,$(uname_S)))\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 171b4124afe..b573a5ee122 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/wrapper.c b/wrapper.c\nindex cffe24d307a..a9647018b68 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -565,6 +565,10 @@ int git_fsync(int fd, enum fsync_action action)\n \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n #endif\n \n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n \t\terrno = ENOSYS;\n \t\treturn -1;\n \t}\n-- \ngitgitgadget\n\n"},{"id":"435812","messageId":"6b5760389863d86fc15c69cfb31bafce5ad636e1.1631590725.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"[PATCH v3 4/6] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-14T03:38:43Z","receivedAt":"2021-09-14T03:39:10Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality. This mode\nis enabled when passing paths to update-index via the --stdin flag,\nas is done by 'git stash'.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to snoop the output of --verbose to\nfind out when update-index has actually processed a given path.\nAdditionally the index is locked for the duration of the update.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 187203e8bb5..b0689f2cdf6 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1150,6 +1151,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n \n \t\tsetup_work_tree();\n+\t\tplug_bulk_checkin();\n \t\twhile (getline_fn(&buf, stdin) != EOF) {\n \t\t\tchar *p;\n \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n@@ -1164,6 +1166,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n \t\t}\n+\t\tunplug_bulk_checkin(&lock_file);\n \t\tstrbuf_release(&unquoted);\n \t\tstrbuf_release(&buf);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"435813","messageId":"b7ca3ba9302dbf4b2bde87fdd5f663f7790720f0.1631590725.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"[PATCH v3 5/6] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-14T03:38:44Z","receivedAt":"2021-09-14T03:39:11Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/lib-unique-files.sh | 32 ++++++++++++++++++++++++++\n t/perf/p3700-add.sh        | 43 +++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh      | 46 ++++++++++++++++++++++++++++++++++++++\n 3 files changed, 121 insertions(+)\n create mode 100644 t/perf/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/lib-unique-files.sh b/t/perf/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..10083395ae5\n--- /dev/null\n+++ b/t/perf/lib-unique-files.sh\n@@ -0,0 +1,32 @@\n+# Helper to create files with unique contents\n+\n+test_create_unique_files_base__=$(date -u)\n+test_create_unique_files_counter__=0\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n+#\t\t\t\t    each in the current directory, all\n+#\t\t\t\t    with unique contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\" > /dev/null\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\ttest_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n+\t\t\techo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..4ca3224f364\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/perf/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..407b95c104b\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/perf/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"435814","messageId":"55a40fc8fd59df6180c8a87d93fcc9a232ff8d0a.1631590725.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"[PATCH v3 6/6] core.fsyncobjectfiles: enable batch mode for testing","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-14T03:38:45Z","receivedAt":"2021-09-14T03:39:18Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n environment.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/environment.c b/environment.c\nindex 3e23eafff80..27d5e11267e 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -43,7 +43,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int core_compression_level;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n+enum FSYNC_OBJECT_FILES_MODE fsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\n-- \ngitgitgadget\n"},{"id":"435828","messageId":"20210914054926.GA26190@lst.de","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Christoph Hellwig","fromEmail":"hch@lst.de","sentAt":"2021-09-14T05:49:26Z","receivedAt":"2021-09-14T05:49:37Z","isPatch":true,"sender":{"key":"hch@lst.de","avatar":null},"body":"On Tue, Sep 14, 2021 at 03:38:39AM +0000, Neeraj K. Singh via GitGitGadget wrote:\n> NOTE: Based on Christoph Hellwig's comments, the 'batch' mode is not correct\n> on Linux, since sync_file_range does not provide data integrity guarantees.\n> There is currently no kernel interface suitable to achieve disk flush\n> batching as is, but he suggested that he might implement a 'syncfs' variant\n> on top of this patchset. This code is still useful on macOS and Windows, and\n> the config documentation makes that clear.\n\nIf this series lands I can give the syncfs variant a spin.  It might not\nbe the best option for gt hosting services, but I think it will be very\nhelpful for typical developer workstations.\n"},{"id":"435840","messageId":"b992c83d-747f-55b2-cf58-f39f4ef734aa@gmail.com","threadId":"56371","inReplyTo":"f8b5b709e9edc363b2de7d4afa443deec0120ca0.1631590725.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-09-14T10:39:45Z","receivedAt":"2021-09-14T10:40:07Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 14/09/21 10.38, Neeraj Singh via GitGitGadget wrote:\n> _Performance numbers_:\n> \n> Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\n> Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\n> Windows - Same host as Linux, a preview version of Windows 11.\n> \t  This number is from a patch later in the series.\n> \n> Adding 500 files to the repo with 'git add' Times reported in seconds.\n> \n> core.fsyncObjectFiles | Linux | Mac   | Windows\n> ----------------------|-------|-------|--------\n>                  false | 0.06  |  0.35 | 0.61\n>                  true  | 1.88  | 11.18 | 2.47\n>                  batch | 0.15  |  0.41 | 1.53\n\nInteresting here the performance.\n\nYou said that core.fsyncObjectFiles=batch performed 2.5x slower than \ncore.fsyncObjectFile=false on Linux and Windows, why?\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"435907","messageId":"CANQDOddbWCkRroeQF62GeXTE-fsSWCiXPXVfrwOFcC3zNYX7iA@mail.gmail.com","threadId":"56371","inReplyTo":"b992c83d-747f-55b2-cf58-f39f4ef734aa@gmail.com","subject":"Re: [PATCH v3 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-14T19:05:21Z","receivedAt":"2021-09-14T19:05:37Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 14, 2021 at 3:39 AM Bagas Sanjaya <bagasdotme@gmail.com> wrote:\n>\n> On 14/09/21 10.38, Neeraj Singh via GitGitGadget wrote:\n> > _Performance numbers_:\n> >\n> > Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\n> > Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\n> > Windows - Same host as Linux, a preview version of Windows 11.\n> >         This number is from a patch later in the series.\n> >\n> > Adding 500 files to the repo with 'git add' Times reported in seconds.\n> >\n> > core.fsyncObjectFiles | Linux | Mac   | Windows\n> > ----------------------|-------|-------|--------\n> >                  false | 0.06  |  0.35 | 0.61\n> >                  true  | 1.88  | 11.18 | 2.47\n> >                  batch | 0.15  |  0.41 | 1.53\n>\n> Interesting here the performance.\n>\n> You said that core.fsyncObjectFiles=batch performed 2.5x slower than\n> core.fsyncObjectFile=false on Linux and Windows, why?\n>\n\nThe goal of batch mode is to minimize the number of disk cache flush operations.\nWe still have to issue writes to the disk (and wait for them to\ncomplete) in batch mode,\nand on my test system those writes have to cross the VM boundary.  The\nMac is running\nmacOS natively, so performance of the writes is probably a little better.\n"},{"id":"435914","messageId":"xmqqfsu70x58.fsf@gitster.g","threadId":"56371","inReplyTo":"f8b5b709e9edc363b2de7d4afa443deec0120ca0.1631590725.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-14T19:34:11Z","receivedAt":"2021-09-14T19:34:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/config.c b/config.c\n> index cb4a8058bff..9fe3602e1c4 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1509,7 +1509,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t}\n>  \n>  \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> -\t\tfsync_object_files = git_config_bool(var, value);\n> +\t\tif (!value)\n> +\t\t\treturn config_error_nonbool(var);\n> +\t\tif (!strcasecmp(value, \"batch\"))\n> +\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n> +\t\telse\n> +\t\t\tfsync_object_files = git_config_bool(var, value)\n> +\t\t\t\t? FSYNC_OBJECT_FILES_ON : FSYNC_OBJECT_FILES_OFF;\n>  \t\treturn 0;\n\nThe original code used to allow the short-and-sweet valueless true\n\n\t[core]\n\t\tfsyncobjectfiles\n\nbut it no longer does by calling it a nonbool error.  This breaks\nexisting users' repositories that have been happily working, doesn't\nit?\n\nPerhaps\n\n\tif (value && !strcmp(value, \"batch\"))\n\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n\telse if (git_config_bool(var, value))\n\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n\telse\n\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n\n> -/* Finalize a file on disk, and close it. */\n> -static void close_loose_object(int fd)\n> -{\n> -\tif (fsync_object_files)\n> -\t\tfsync_or_die(fd, \"loose object file\");\n> -\tif (close(fd) != 0)\n> -\t\tdie_errno(_(\"error when closing loose object file\"));\n> -}\n> -\n>  /* Size of directory component, including the ending '/' */\n>  static inline int directory_size(const char *filename)\n>  {\n> @@ -1973,17 +1964,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>  \n> -\tclose_loose_object(fd);\n> -\n> -\tif (mtime) {\n> -\t\tstruct utimbuf utb;\n> -\t\tutb.actime = mtime;\n> -\t\tutb.modtime = mtime;\n> -\t\tif (utime(tmp_file.buf, &utb) < 0)\n> -\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n> -\t}\n> -\n> -\treturn finalize_object_file(tmp_file.buf, filename.buf);\n> +\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf,\n> +\t\t\t\t\t\t\t filename.buf, mtime);\n>  }\n\nThis block of code looked familiar and I was about to complain \"why\nadd it in one step and remove it in another?\"\n\nBut it is a different instance from the one that was added in one of\nthe previous patches ;-).  \n\n> +int git_fsync(int fd, enum fsync_action action)\n> +{\n> +\tif (action == FSYNC_WRITEOUT_ONLY) {\n> +#ifdef __APPLE__\n> +\t\t/*\n> +\t\t * on Mac OS X, fsync just causes filesystem cache writeback but does not\n> +\t\t * flush hardware caches.\n> +\t\t */\n> +\t\treturn fsync(fd);\n> +#endif\n> +\n> +#ifdef HAVE_SYNC_FILE_RANGE\n> +\t\t/*\n> +\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n> +\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n> +\t\t * indicates writeout of the entire file and the wait flags ensure that all\n> +\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n> +\t\t * before we continue.\n> +\t\t */\n> +\n> +\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n> +#endif\n> +\n> +\t\terrno = ENOSYS;\n> +\t\treturn -1;\n> +\t}\n\nThis allows the caller that can take advantage of writeout-only mode\nto naturally fall back on the full sync per each file if we cannot do\na writeout-only sync.  OK.\n\n> +#ifdef __APPLE__\n> +\treturn fcntl(fd, F_FULLFSYNC);\n> +#else\n> +\treturn fsync(fd);\n> +#endif\n> +}\n\nIf we are introducing \"enum fsync_action\", we should have some way\nto make it clear that we are covering all the possible values of\n\"action\".\n\nSwitching on action, i.e.\n\n\tswitch (action) {\n\tcase FSYNC_WRITEOUT_ONLY:\n\t\t...\n\t\tbreak;\n\tcase FSYNC_HARDWARE_FLUSH:\n\t\t...\n\t\tbreak;\n\tdefault:\n\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n\t}\n\nwould be one way to do so.\n\nThanks.\n"},{"id":"435915","messageId":"xmqqbl4v0x3e.fsf@gitster.g","threadId":"56371","inReplyTo":"6b5760389863d86fc15c69cfb31bafce5ad636e1.1631590725.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 4/6] update-index: use the bulk-checkin infrastructure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-14T19:35:17Z","receivedAt":"2021-09-14T19:35:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> The update-index functionality is used internally by 'git stash push' to\n> setup the internal stashed commit.\n\nNice.\n\n> This change enables bulk-checkin for update-index infrastructure to\n> speed up adding new objects to the object database by leveraging the\n> pack functionality and the new bulk-fsync functionality. This mode\n> is enabled when passing paths to update-index via the --stdin flag,\n> as is done by 'git stash'.\n>\n> There is some risk with this change, since under batch fsync, the object\n> files will not be available until the update-index is entirely complete.\n> This usage is unlikely, since any tool invoking update-index and\n> expecting to see objects would have to snoop the output of --verbose to\n> find out when update-index has actually processed a given path.\n> Additionally the index is locked for the duration of the update.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  builtin/update-index.c | 3 +++\n>  1 file changed, 3 insertions(+)\n>\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index 187203e8bb5..b0689f2cdf6 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -5,6 +5,7 @@\n>   */\n>  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n>  #include \"cache.h\"\n> +#include \"bulk-checkin.h\"\n>  #include \"config.h\"\n>  #include \"lockfile.h\"\n>  #include \"quote.h\"\n> @@ -1150,6 +1151,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t\tstruct strbuf unquoted = STRBUF_INIT;\n>  \n>  \t\tsetup_work_tree();\n> +\t\tplug_bulk_checkin();\n>  \t\twhile (getline_fn(&buf, stdin) != EOF) {\n>  \t\t\tchar *p;\n>  \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n> @@ -1164,6 +1166,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t\t\t\tchmod_path(set_executable_bit, p);\n>  \t\t\tfree(p);\n>  \t\t}\n> +\t\tunplug_bulk_checkin(&lock_file);\n>  \t\tstrbuf_release(&unquoted);\n>  \t\tstrbuf_release(&buf);\n>  \t}\n"},{"id":"435919","messageId":"xmqqr1dqzyl7.fsf@gitster.g","threadId":"56371","inReplyTo":"xmqqfsu70x58.fsf@gitster.g","subject":"Re: [PATCH v3 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-14T20:33:40Z","receivedAt":"2021-09-14T20:33:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Perhaps\n>\n> \tif (value && !strcmp(value, \"batch\"))\n> \t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n> \telse if (git_config_bool(var, value))\n> \t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n> \telse\n> \t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n\nBy the way, in case it wasn't clear, I do mean strcmp and not\nstrcasecmp.  Making these things that are meant to be machine\nreadable tokens to be spelled in different ways in the name of\n\"friendliness\" is a disease.\n\nThanks.\n"},{"id":"435961","messageId":"CANQDOdfnWV3EMhumjtin9z2RB36Ew4vrLPZjzMNmvrJ3=RN20Q@mail.gmail.com","threadId":"56371","inReplyTo":"xmqqfsu70x58.fsf@gitster.g","subject":"Re: [PATCH v3 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-15T04:55:30Z","receivedAt":"2021-09-15T04:55:43Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 14, 2021 at 12:34 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > diff --git a/config.c b/config.c\n> > index cb4a8058bff..9fe3602e1c4 100644\n> > --- a/config.c\n> > +++ b/config.c\n> > @@ -1509,7 +1509,13 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >       }\n> >\n> >       if (!strcmp(var, \"core.fsyncobjectfiles\")) {\n> > -             fsync_object_files = git_config_bool(var, value);\n> > +             if (!value)\n> > +                     return config_error_nonbool(var);\n> > +             if (!strcasecmp(value, \"batch\"))\n> > +                     fsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n> > +             else\n> > +                     fsync_object_files = git_config_bool(var, value)\n> > +                             ? FSYNC_OBJECT_FILES_ON : FSYNC_OBJECT_FILES_OFF;\n> >               return 0;\n>\n> The original code used to allow the short-and-sweet valueless true\n>\n>         [core]\n>                 fsyncobjectfiles\n>\n> but it no longer does by calling it a nonbool error.  This breaks\n> existing users' repositories that have been happily working, doesn't\n> it?\n>\n> Perhaps\n>\n>         if (value && !strcmp(value, \"batch\"))\n>                 fsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n>         else if (git_config_bool(var, value))\n>                 fsync_object_files = FSYNC_OBJECT_FILES_ON;\n>         else\n>                 fsync_object_files = FSYNC_OBJECT_FILES_OFF;\n\nI'll take your suggestion, including the change to case-sensitive.\n\n> > +#ifdef __APPLE__\n> > +     return fcntl(fd, F_FULLFSYNC);\n> > +#else\n> > +     return fsync(fd);\n> > +#endif\n> > +}\n>\n> If we are introducing \"enum fsync_action\", we should have some way\n> to make it clear that we are covering all the possible values of\n> \"action\".\n>\n> Switching on action, i.e.\n>\n>         switch (action) {\n>         case FSYNC_WRITEOUT_ONLY:\n>                 ...\n>                 break;\n>         case FSYNC_HARDWARE_FLUSH:\n>                 ...\n>                 break;\n>         default:\n>                 BUG(\"unexpected git_fsync(%d) call\", action);\n>         }\n>\n> would be one way to do so.\n>\n\nWill do.\n\nThanks for reviewing my changes. I've updated the github PR.\nI'll wait for a few more days to see if anyone has more feedback\nbefore sending out another round of patches.\n"},{"id":"435985","messageId":"xmqqtuilyfls.fsf@gitster.g","threadId":"56371","inReplyTo":"55a40fc8fd59df6180c8a87d93fcc9a232ff8d0a.1631590725.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 6/6] core.fsyncobjectfiles: enable batch mode for testing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-15T16:21:19Z","receivedAt":"2021-09-15T16:21:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  environment.c | 2 +-\n>  1 file changed, 1 insertion(+), 1 deletion(-)\n>\n> diff --git a/environment.c b/environment.c\n> index 3e23eafff80..27d5e11267e 100644\n> --- a/environment.c\n> +++ b/environment.c\n> @@ -43,7 +43,7 @@ const char *git_hooks_path;\n>  int zlib_compression_level = Z_BEST_SPEED;\n>  int core_compression_level;\n>  int pack_compression_level = Z_DEFAULT_COMPRESSION;\n> -enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n> +enum FSYNC_OBJECT_FILES_MODE fsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n>  size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n>  size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n>  size_t delta_base_cache_limit = 96 * 1024 * 1024;\n\nDespite what the title of the change claims, this is not \"enable for\ntesting\", but \"enable for everybody even in production\", isn't it?\n\nI'd prefer we do not do this, certainly not for \"testing\".\n\nIf setting the variable to \"batch\" were meant to eventually improve\nperformance for all different flavours of workload, I do not think\nwe would mind if we set it to \"batch\" for those who opt into the\n\"experimental\" set of features by setting the feature.experimental\nconfiguration variable to true.  And after a few development cycles\nwhen the feature proves to be useful for everybody, we may want to\napply this patch under a justification that is different from \"for\ntesting\".\n\nOn the other hand, if this is meant to help 85% of people while\ndegrading the remainder of workflow, I do not think we would want to\nsee this change without a warning that says something along the\nlines of \"under rare circumstances (e.g. if you employ such and such\nworkflow), the new default value used for the core.fsyncObjectFiles\nconfiguration variable will hurt performance.\"\n\nSince this is about answering the question \"between performance and\ncrash resilience, where do you as an end user strike the balance for\nyour needs?\", I do not think it falls into either of the above two\ncategories.  \n\nThe only plausible justification I can think of to apply a \"we\ndefault to 'batch' for everybody\" patch with is something like:\n\n    Now with the 'batch' setting for core.fsyncObjectFiles, unlike\n    'true' that paid very high overhead, the overhead to ensure our\n    writes hit the disk platters has so greatly been reduced that it\n    hurts the performance only negligibly.  Let's switch the default\n    from the unsafe value of 'false' to safer and performant value\n    of 'batch'.\n\nI however doubt with the current round of patches, we are there yet.\n"},{"id":"436057","messageId":"CANQDOdc8F7a3ZeTDpUWrt8uUntnX4jHYxyj96SPwH-P=kMrneg@mail.gmail.com","threadId":"56371","inReplyTo":"xmqqtuilyfls.fsf@gitster.g","subject":"Re: [PATCH v3 6/6] core.fsyncobjectfiles: enable batch mode for testing","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-15T22:43:38Z","receivedAt":"2021-09-15T22:43:53Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Sep 15, 2021 at 9:21 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  environment.c | 2 +-\n> >  1 file changed, 1 insertion(+), 1 deletion(-)\n> >\n> > diff --git a/environment.c b/environment.c\n> > index 3e23eafff80..27d5e11267e 100644\n> > --- a/environment.c\n> > +++ b/environment.c\n> > @@ -43,7 +43,7 @@ const char *git_hooks_path;\n> >  int zlib_compression_level = Z_BEST_SPEED;\n> >  int core_compression_level;\n> >  int pack_compression_level = Z_DEFAULT_COMPRESSION;\n> > -enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n> > +enum FSYNC_OBJECT_FILES_MODE fsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n> >  size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n> >  size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n> >  size_t delta_base_cache_limit = 96 * 1024 * 1024;\n>\n> Despite what the title of the change claims, this is not \"enable for\n> testing\", but \"enable for everybody even in production\", isn't it?\n>\n> I'd prefer we do not do this, certainly not for \"testing\".\n>\n> If setting the variable to \"batch\" were meant to eventually improve\n> performance for all different flavours of workload, I do not think\n> we would mind if we set it to \"batch\" for those who opt into the\n> \"experimental\" set of features by setting the feature.experimental\n> configuration variable to true.  And after a few development cycles\n> when the feature proves to be useful for everybody, we may want to\n> apply this patch under a justification that is different from \"for\n> testing\".\n>\n> On the other hand, if this is meant to help 85% of people while\n> degrading the remainder of workflow, I do not think we would want to\n> see this change without a warning that says something along the\n> lines of \"under rare circumstances (e.g. if you employ such and such\n> workflow), the new default value used for the core.fsyncObjectFiles\n> configuration variable will hurt performance.\"\n>\n> Since this is about answering the question \"between performance and\n> crash resilience, where do you as an end user strike the balance for\n> your needs?\", I do not think it falls into either of the above two\n> categories.\n>\n> The only plausible justification I can think of to apply a \"we\n> default to 'batch' for everybody\" patch with is something like:\n>\n>     Now with the 'batch' setting for core.fsyncObjectFiles, unlike\n>     'true' that paid very high overhead, the overhead to ensure our\n>     writes hit the disk platters has so greatly been reduced that it\n>     hurts the performance only negligibly.  Let's switch the default\n>     from the unsafe value of 'false' to safer and performant value\n>     of 'batch'.\n>\n> I however doubt with the current round of patches, we are there yet.\n\nSorry for being unclear here (and perhaps including an improper patch).\nThis commit is mainly to ensure that we get coverage of batch mode on all\nplatforms in the CI infrastructure.  I don't believe it should be included in\nmainline git without significantly more discussion and experimentation.\n\nHowever, I'd hope that Git for Windows would be able to adopt batch mode\nby default when they pull this series in. They are currently enabling fsync\nby default.\n\nBatch mode does have more cost, particularly on rotational media.\nI think git should eventually enable batch mode by default with the proviso that\nmaintainers and people running ephemeral CI infrastructure should turn\nfsync off\nif they care more about speed than durability.\n\nDo you think that feature.experimental is a good place to put this right away,\nor should we just leave this as an option that Git for Windows can pick up and\nleave the other platforms alone?\n\nThanks,\nNeeraj\n"},{"id":"436058","messageId":"xmqqy27xv3g7.fsf@gitster.g","threadId":"56371","inReplyTo":"CANQDOdc8F7a3ZeTDpUWrt8uUntnX4jHYxyj96SPwH-P=kMrneg@mail.gmail.com","subject":"Re: [PATCH v3 6/6] core.fsyncobjectfiles: enable batch mode for testing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-15T23:12:08Z","receivedAt":"2021-09-15T23:12:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> This commit is mainly to ensure that we get coverage of batch mode on all\n> platforms in the CI infrastructure.  I don't believe it should be included in\n> mainline git without significantly more discussion and experimentation.\n\nAm I incorrect to say that only just a handful of code paths can\ntake advantage of the bulk checkin \"plugging-unplugging\" feature to\nbegin with, so running _all_ the existing tests that cover\neverything with this core.fsyncobjectfiles=batch setting is rather\npointless?\n\nIf so, perhaps instead of 6/6, you should identify key code paths\nthat would be affected by this feature (perhaps \"git add\" is one of\nthem), and either write a new test script dedicated for this feature\nor piggy-back on existing test scripts that already tests the code\npaths and adding new test pieces there that exercise this new feature.\n\nIf it is a good idea to run all the tests with core.fsyncobjectfiles\nset to batch, however, it probalby is easiest to invent a new\nenvironment variable GIT_TEST_FORCE_CORE_FSYNCOBJECTFILES and have\nit honored as the default when it is set, and add a NEW CI job that\nexports the environment with the value \"batch\".  Other people\n(including the ones from Microsoft, I think) are much more familiar\nthan I am on how to make this kind of thing work in GitHub Actions.\n\n> Do you think that feature.experimental is a good place to put this right away,\n\nI think feature.experimental should be used for something that we\nhope would benefit \"everybody\", not \"most of the users\".  This is a\npromise to our testers, who opt into \"early preview\" of upcoming\nfeatures should not be subjected to \"this may or may not give better\nexperiences depending on your workflow\".  They may already be\nenjoying and even relying on other experimental features by opting\nin, and we should strive not to add a reason for them to turn the\nfeature.experimental bit off by saying \"this new experimental feature\nthat recently joined does not work for my use case.\"\n\n\n"},{"id":"436085","messageId":"xmqqczp9ujo3.fsf@gitster.g","threadId":"56371","inReplyTo":"xmqqy27xv3g7.fsf@gitster.g","subject":"Re: [PATCH v3 6/6] core.fsyncobjectfiles: enable batch mode for testing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-16T06:19:24Z","receivedAt":"2021-09-16T06:19:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n>> This commit is mainly to ensure that we get coverage of batch mode on all\n>> platforms in the CI infrastructure.  I don't believe it should be included in\n>> mainline git without significantly more discussion and experimentation.\n>\n> Am I incorrect to say that only just a handful of code paths can\n> take advantage of the bulk checkin \"plugging-unplugging\" feature to\n> begin with, so running _all_ the existing tests that cover\n> everything with this core.fsyncobjectfiles=batch setting is rather\n> pointless?\n>\n> If so, perhaps instead of 6/6, you should identify key code paths\n> that would be affected by this feature (perhaps \"git add\" is one of\n> them), and either write a new test script dedicated for this feature\n> or piggy-back on existing test scripts that already tests the code\n> paths and adding new test pieces there that exercise this new feature.\n>\n> If it is a good idea to run all the tests with core.fsyncobjectfiles\n> set to batch, however, it probalby is easiest to invent a new\n> environment variable GIT_TEST_FORCE_CORE_FSYNCOBJECTFILES and have\n> it honored as the default when it is set, and add a NEW CI job that\n> exports the environment with the value \"batch\".  \n\nI have to take a part of this back.  A new environment variable that\nis honored in the absense of core.fsyncobjectfiles would be needed\nif you need to run all tests, but you do not necessarily have to add\na new CI job---instead you should be able to piggyback on an\nexisting job, by mimicking the way how ci/run-build-and-tests.sh\nenables various test options on one of the jobs.\n\n> Other people\n> (including the ones from Microsoft, I think) are much more familiar\n> than I am on how to make this kind of thing work in GitHub Actions.\n\nThis part still stands ;-)  There might be a better way than adding\nyet another environment variable.\n\n"},{"id":"436501","messageId":"afb0028e79648c1f7be8d77df5c6d675bd27d983.1632176111.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","subject":"[PATCH v4 5/6] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-20T22:15:10Z","receivedAt":"2021-09-21T02:20:17Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for 'git add'\nand 'git stash'. These tests ensure that the added\ndata winds up in the object database.\n\nI verified the tests by introducing an incorrect rename\nin do_sync_and_rename.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh | 34 ++++++++++++++++++++++++++++++++++\n t/t3700-add.sh        | 11 +++++++++++\n t/t3903-stash.sh      | 14 ++++++++++++++\n 3 files changed, 59 insertions(+)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a8a25eba61d\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,34 @@\n+# Helper to create files with unique contents\n+\n+test_create_unique_files_base__=$(date -u)\n+test_create_unique_files_counter__=0\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n+#\t\t\t\t    each in the specified directory, all\n+#\t\t\t\t    with unique contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\n+\trm -rf $basedir >/dev/null\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\" > /dev/null\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\ttest_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n+\t\t\techo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex 4086e1ebbc9..2122acc3e9e 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -7,6 +7,8 @@ test_description='Test of git add, including the -- option.'\n \n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -33,6 +35,15 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+test_expect_success 'git add: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tcat fsynced_files | awk '{print \\$2}' | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 873aa56e359..0b4e8bb55b8 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n diff_cmp () {\n \tfor i in \"$1\" \"$2\"\n@@ -1293,6 +1294,19 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tcat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n \texpected=\"stash.useBuiltin support has been removed\" &&\n \n-- \ngitgitgadget\n\n"},{"id":"436503","messageId":"12cad737635663ed596e52f89f0f4f22f58bfe38.1632176111.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","subject":"[PATCH v4 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-20T22:15:07Z","receivedAt":"2021-09-21T02:20:18Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nmacOS, and Linux each offer mechanisms to write data from the filesystem\npage cache without initiating a hardware flush.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\ntakes advantage of the bulk-checkin infrastructure to batch up hardware\nflushes.\n\nWhen the new mode is enabled we do the following for new objects:\n\n1. Create a tmp_obj_XXXX file and write the object data to it.\n2. Issue a pagecache writeback request and wait for it to complete.\n3. Record the tmp name and the final name in the bulk-checkin state for\n   later rename.\n\nAt the end of the entire transaction we:\n1. Issue a fsync against the lock file to flush the hardware writeback\n   cache, which should by now have processed the tmp file writes.\n2. Rename all of the temp files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\nwould expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 26 ++++++++---\n Makefile                      |  6 +++\n builtin/add.c                 |  3 +-\n bulk-checkin.c                | 81 ++++++++++++++++++++++++++++++++++-\n bulk-checkin.h                |  5 ++-\n cache.h                       |  8 +++-\n config.c                      |  7 ++-\n config.mak.uname              |  1 +\n configure.ac                  |  8 ++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 +++\n object-file.c                 | 22 +---------\n wrapper.c                     | 44 +++++++++++++++++++\n write-or-die.c                |  2 +-\n 14 files changed, 189 insertions(+), 33 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..0006d90980d 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,26 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols the object store, so updates to any refs or the\n+\tindex may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write object data with a minimal set of FLUSH CACHE\n+  (or equivalent) commands sent to the storage controller. If the operating\n+  system interfaces are not available, this mode behaves the same as `true`.\n+  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n+  filesystems and on Windows for repos stored on NTFS or ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex 429c276058d..326c7607e0f 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1896,6 +1898,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 2244311d485..dda4bf093a0 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -678,7 +678,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\n+\tunplug_bulk_checkin(&lock_file);\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f117d62c908..ddbab5e5c8c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,15 +3,19 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n \n+static struct string_list bulk_fsync_state = STRING_LIST_INIT_DUP;\n+\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -62,6 +66,32 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n+{\n+\tif (fsync_state->nr) {\n+\t\tstruct string_list_item *rename;\n+\n+\t\t/*\n+\t\t * Issue a full hardware flush against the lock file to ensure\n+\t\t * that all objects are durable before any renames occur.\n+\t\t * The code in fsync_and_close_loose_object_bulk_checkin has\n+\t\t * already ensured that writeout has occurred, but it has not\n+\t\t * flushed any writeback cache in the storage hardware.\n+\t\t */\n+\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n+\n+\t\tfor_each_string_list_item(rename, fsync_state) {\n+\t\t\tconst char *src = rename->string;\n+\t\t\tconst char *dst = rename->util;\n+\n+\t\t\tif (finalize_object_file(src, dst))\n+\t\t\t\tdie_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n+\t\t}\n+\n+\t\tstring_list_clear(fsync_state, 1);\n+\t}\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -256,6 +286,53 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+static void add_rename_bulk_checkin(struct string_list *fsync_state,\n+\t\t\t\t    const char *src, const char *dst)\n+{\n+\tstring_list_insert(fsync_state, src)->util = xstrdup(dst);\n+}\n+\n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n+\t\t\t\t\t      const char *filename, time_t mtime)\n+{\n+\tint do_finalize = 1;\n+\tint ret = 0;\n+\n+\tif (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n+\t\t/*\n+\t\t * If we have a plugged bulk checkin, we issue a call that\n+\t\t * cleans the filesystem page cache but avoids a hardware flush\n+\t\t * command. Later on we will issue a single hardware flush\n+\t\t * before renaming files as part of do_sync_and_rename.\n+\t\t */\n+\t\tif (bulk_checkin_plugged &&\n+\t\t    fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n+\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\t\tadd_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n+\t\t\tdo_finalize = 0;\n+\n+\t\t} else {\n+\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\t}\n+\t}\n+\n+\tif (close(fd))\n+\t\tdie_errno(_(\"error when closing loose object file\"));\n+\n+\tif (mtime) {\n+\t\tstruct utimbuf utb;\n+\t\tutb.actime = mtime;\n+\t\tutb.modtime = mtime;\n+\t\tif (utime(tmpfile, &utb) < 0)\n+\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n+\t}\n+\n+\tif (do_finalize)\n+\t\tret = finalize_object_file(tmpfile, filename);\n+\n+\treturn ret;\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -273,10 +350,12 @@ void plug_bulk_checkin(void)\n \tbulk_checkin_plugged = 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void unplug_bulk_checkin(struct lock_file *lock_file)\n {\n \tassert(bulk_checkin_plugged);\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_sync_and_rename(&bulk_fsync_state, lock_file);\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..4a3309c1531 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,11 +6,14 @@\n \n #include \"cache.h\"\n \n+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n+\t\t\t\t\t      const char *filename, time_t mtime);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+void unplug_bulk_checkin(struct lock_file *);\n \n #endif\ndiff --git a/cache.h b/cache.h\nindex d23de693680..39b3a88181a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum FSYNC_OBJECT_FILES_MODE {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex cb4a8058bff..1b403e00241 100644\n--- a/config.c\n+++ b/config.c\n@@ -1509,7 +1509,12 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (value && !strcmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse if (git_config_bool(var, value))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n+\t\telse\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 76516aaa9a5..e6d482fbcc6 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex d6b22ede7ea..3e23eafff80 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -43,7 +43,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int core_compression_level;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b46605300ab..d14e2436276 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex a8be8994814..ea14c3a3483 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1859,15 +1859,6 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \treturn 0;\n }\n \n-/* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd)\n-{\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n-\tif (close(fd) != 0)\n-\t\tdie_errno(_(\"error when closing loose object file\"));\n-}\n-\n /* Size of directory component, including the ending '/' */\n static inline int directory_size(const char *filename)\n {\n@@ -1973,17 +1964,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n-\n-\tif (mtime) {\n-\t\tstruct utimbuf utb;\n-\t\tutb.actime = mtime;\n-\t\tutb.modtime = mtime;\n-\t\tif (utime(tmp_file.buf, &utb) < 0)\n-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n-\t}\n-\n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf,\n+\t\t\t\t\t\t\t filename.buf, mtime);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\ndiff --git a/wrapper.c b/wrapper.c\nindex 7c6586af321..bb4f9f043ce 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -540,6 +540,50 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync(fd);\n+#endif\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex d33e68f6abb..8f53953d4ab 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"436504","messageId":"d5893e28df152ce9b843d4709ed9349156633b3d.1632176111.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","subject":"[PATCH v4 1/6] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-20T22:15:06Z","receivedAt":"2021-09-21T02:20:23Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b023d9959aa..f117d62c908 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_bulk_checkin(struct bulk_checkin_state *state)\n {\n@@ -260,21 +260,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"436505","messageId":"f7f756f3932cdbca587de397598758c685bac29a.1632176111.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","subject":"[PATCH v4 4/6] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-20T22:15:09Z","receivedAt":"2021-09-21T02:20:26Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality. This mode\nis enabled when passing paths to update-index via the --stdin flag,\nas is done by 'git stash'.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to snoop the output of --verbose to\nfind out when update-index has actually processed a given path.\nAdditionally the index is locked for the duration of the update.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 187203e8bb5..b0689f2cdf6 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1150,6 +1151,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n \n \t\tsetup_work_tree();\n+\t\tplug_bulk_checkin();\n \t\twhile (getline_fn(&buf, stdin) != EOF) {\n \t\t\tchar *p;\n \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n@@ -1164,6 +1166,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n \t\t}\n+\t\tunplug_bulk_checkin(&lock_file);\n \t\tstrbuf_release(&unquoted);\n \t\tstrbuf_release(&buf);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"436506","messageId":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v3.git.git.1631590725.gitgitgadget@gmail.com","subject":"[PATCH v4 0/6] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-20T22:15:05Z","receivedAt":"2021-09-21T02:21:58Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far! Changes since v3:\n\n * Fix core.fsyncobjectfiles option parsing as suggested by Junio: We now\n   accept no value to mean \"true\" and we require 'batch' to be lowercase.\n\n * Leave the default fsync mode as 'false'. Git for windows can change its\n   default when this series makes it over to that fork.\n\n * Use a switch statement in git_fsync, as suggested by Junio.\n\n * Add regression test cases for core.fsyncobjectfiles=batch. This should\n   keep the batch functionality basically working in upstream git even if\n   few users adopt batch mode initially. I expect git-for-windows will\n   provide a good baking area for the new mode.\n\nChanges since v2:\n\n * Removed an unused Makefile define (FSYNC_DOESNT_FLUSH) that slipped in\n   from an intermediate change.\n\n * Drop the futimens part of the patch and return to just calling utime, now\n   within the new bulk_checkin code. The utime to futimens change seemed to\n   be problematic for some platforms (thanks Randall Becker), and is really\n   orthogonal to the rest of the patch series.\n\n * (Optional commit) Enable batch mode by default so that we can shake loose\n   any issues relating to deferring the renames until the\n   unplug_bulk_checkin.\n\nChanges since v1:\n\n * Switch from futimes(2) to futimens(2), which is in POSIX.1-2008. Contrary\n   to dscho's suggestion, I'm still implementing the Windows version in the\n   same patch and I'm not doing autoconf detection since this is a POSIX\n   function.\n\n * Introduce a separate preparatory patch to the bulk-checkin infrastructure\n   to separate the 'plugged' variable and rename the 'state' variable, as\n   suggested by dscho.\n\n * Add performance numbers to the commit message of the main bulk fsync\n   patch, as suggested by dscho.\n\n * Add a comment about the non-thread-safety of the bulk-checkin\n   infrastructure, as suggested by avarab.\n\n * Rename the experimental mode to core.fsyncobjectfiles=batch, as suggested\n   by dscho and avarab and others.\n\n * Add more details to Documentation/config/core.txt about the various\n   settings and their intended effects, as suggested by avarab.\n\n * Switch to the string-list API to hold the rename state, as suggested by\n   avarab.\n\n * Create a separate update-index patch to use bulk-checkin as suggested by\n   dscho.\n\n * Add Windows support in the upstream git. This is done in a way that\n   should not conflict with git-for-windows.\n\n * Add new performance tests that shows the delta based on fsync mode.\n\nNOTE: Based on Christoph Hellwig's comments, the 'batch' mode is not correct\non Linux, since sync_file_range does not provide data integrity guarantees.\nThere is currently no kernel interface suitable to achieve disk flush\nbatching as is, but he suggested that he might implement a 'syncfs' variant\non top of this patchset. This code is still useful on macOS and Windows, and\nthe config documentation makes that clear.\n\nNeeraj Singh (6):\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  core.fsyncobjectfiles: add windows support for batch mode\n  update-index: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: tests for batch mode\n  core.fsyncobjectfiles: performance tests for add and stash\n\n Documentation/config/core.txt       |  26 +++++--\n Makefile                            |   6 ++\n builtin/add.c                       |   3 +-\n builtin/update-index.c              |   3 +\n bulk-checkin.c                      | 103 +++++++++++++++++++++++++---\n bulk-checkin.h                      |   5 +-\n cache.h                             |   8 ++-\n compat/mingw.h                      |   3 +\n compat/win32/flush.c                |  29 ++++++++\n config.c                            |   7 +-\n config.mak.uname                    |   3 +\n configure.ac                        |   8 +++\n contrib/buildsystems/CMakeLists.txt |   3 +-\n environment.c                       |   2 +-\n git-compat-util.h                   |   7 ++\n object-file.c                       |  22 +-----\n t/lib-unique-files.sh               |  34 +++++++++\n t/perf/p3700-add.sh                 |  43 ++++++++++++\n t/perf/p3900-stash.sh               |  46 +++++++++++++\n t/t3700-add.sh                      |  11 +++\n t/t3903-stash.sh                    |  14 ++++\n wrapper.c                           |  48 +++++++++++++\n write-or-die.c                      |   2 +-\n 23 files changed, 392 insertions(+), 44 deletions(-)\n create mode 100644 compat/win32/flush.c\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: 8b7c11b8668b4e774f81a9f0b4c30144b818f1d1\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v4\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v3:\n\n 1:  d5893e28df1 = 1:  d5893e28df1 bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n 2:  f8b5b709e9e ! 2:  12cad737635 core.fsyncobjectfiles: batched disk flushes\n     @@ config.c: static int git_default_core_config(const char *var, const char *value,\n       \n       \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n      -\t\tfsync_object_files = git_config_bool(var, value);\n     -+\t\tif (!value)\n     -+\t\t\treturn config_error_nonbool(var);\n     -+\t\tif (!strcasecmp(value, \"batch\"))\n     ++\t\tif (value && !strcmp(value, \"batch\"))\n      +\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n     ++\t\telse if (git_config_bool(var, value))\n     ++\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n      +\t\telse\n     -+\t\t\tfsync_object_files = git_config_bool(var, value)\n     -+\t\t\t\t? FSYNC_OBJECT_FILES_ON : FSYNC_OBJECT_FILES_OFF;\n     ++\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n       \t\treturn 0;\n       \t}\n       \n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n       \n      +int git_fsync(int fd, enum fsync_action action)\n      +{\n     -+\tif (action == FSYNC_WRITEOUT_ONLY) {\n     ++\tswitch (action) {\n     ++\tcase FSYNC_WRITEOUT_ONLY:\n     ++\n      +#ifdef __APPLE__\n      +\t\t/*\n     -+\t\t * on Mac OS X, fsync just causes filesystem cache writeback but does not\n     ++\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n      +\t\t * flush hardware caches.\n      +\t\t */\n      +\t\treturn fsync(fd);\n     @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n      +\n      +\t\terrno = ENOSYS;\n      +\t\treturn -1;\n     -+\t}\n     ++\n     ++\tcase FSYNC_HARDWARE_FLUSH:\n      +\n      +#ifdef __APPLE__\n     -+\treturn fcntl(fd, F_FULLFSYNC);\n     ++\t\treturn fcntl(fd, F_FULLFSYNC);\n      +#else\n     -+\treturn fsync(fd);\n     ++\t\treturn fsync(fd);\n      +#endif\n     ++\n     ++\tdefault:\n     ++\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n     ++\t}\n     ++\n      +}\n      +\n       static int warn_if_unremovable(const char *op, const char *file, int rc)\n 3:  815a862e229 ! 3:  a5b3e21b762 core.fsyncobjectfiles: add windows support for batch mode\n     @@ wrapper.c: int git_fsync(int fd, enum fsync_action action)\n      +\n       \t\terrno = ENOSYS;\n       \t\treturn -1;\n     - \t}\n     + \n 4:  6b576038986 = 4:  f7f756f3932 update-index: use the bulk-checkin infrastructure\n -:  ----------- > 5:  afb0028e796 core.fsyncobjectfiles: tests for batch mode\n 5:  b7ca3ba9302 ! 6:  3e6b80b5fa2 core.fsyncobjectfiles: performance tests for add and stash\n     @@ Commit message\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     - ## t/perf/lib-unique-files.sh (new) ##\n     -@@\n     -+# Helper to create files with unique contents\n     -+\n     -+test_create_unique_files_base__=$(date -u)\n     -+test_create_unique_files_counter__=0\n     -+\n     -+# Create multiple files with unique contents. Takes the number of\n     -+# directories, the number of files in each directory, and the base\n     -+# directory.\n     -+#\n     -+# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n     -+#\t\t\t\t    each in the current directory, all\n     -+#\t\t\t\t    with unique contents.\n     -+\n     -+test_create_unique_files() {\n     -+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n     -+\n     -+\tlocal dirs=$1\n     -+\tlocal files=$2\n     -+\tlocal basedir=$3\n     -+\n     -+\tfor i in $(test_seq $dirs)\n     -+\tdo\n     -+\t\tlocal dir=$basedir/dir$i\n     -+\n     -+\t\tmkdir -p \"$dir\" > /dev/null\n     -+\t\tfor j in $(test_seq $files)\n     -+\t\tdo\n     -+\t\t\ttest_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n     -+\t\t\techo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n     -+\t\tdone\n     -+\tdone\n     -+}\n     -\n       ## t/perf/p3700-add.sh (new) ##\n      @@\n      +#!/bin/sh\n     @@ t/perf/p3700-add.sh (new)\n      +\n      +. ./perf-lib.sh\n      +\n     -+. $TEST_DIRECTORY/perf/lib-unique-files.sh\n     ++. $TEST_DIRECTORY/lib-unique-files.sh\n      +\n      +test_perf_default_repo\n      +test_checkout_worktree\n     @@ t/perf/p3700-add.sh (new)\n      +# We need to create the files each time we run the perf test, but\n      +# we do not want to measure the cost of creating the files, so run\n      +# the tet once.\n     -+if test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n     ++if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n      +then\n      +\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n      +\tGIT_PERF_REPEAT_COUNT=1\n     @@ t/perf/p3900-stash.sh (new)\n      +\n      +. ./perf-lib.sh\n      +\n     -+. $TEST_DIRECTORY/perf/lib-unique-files.sh\n     ++. $TEST_DIRECTORY/lib-unique-files.sh\n      +\n      +test_perf_default_repo\n      +test_checkout_worktree\n     @@ t/perf/p3900-stash.sh (new)\n      +# We need to create the files each time we run the perf test, but\n      +# we do not want to measure the cost of creating the files, so run\n      +# the tet once.\n     -+if test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n     ++if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n      +then\n      +\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n      +\tGIT_PERF_REPEAT_COUNT=1\n 6:  55a40fc8fd5 < -:  ----------- core.fsyncobjectfiles: enable batch mode for testing\n\n-- \ngitgitgadget\n"},{"id":"436507","messageId":"a5b3e21b76208fba130e3313a3e70df45ab392af.1632176111.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","subject":"[PATCH v4 3/6] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-20T22:15:08Z","receivedAt":"2021-09-21T02:21:59Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds a win32 implementation for fsync_no_flush that is\ncalled git_fsync. The 'NtFlushBuffersFileEx' function being called is\navailable since Windows 8. If the function is not available, we\nreturn -1 and Git falls back to doing a full fsync.\n\nThe operating system is told to flush data only without a hardware\nflush primitive. A later full fsync will cause the metadata log\nto be flushed and then the disk cache to be flushed on NTFS and\nReFS. Other filesystems will treat this as a full flush operation.\n\nI added a new file here for this system call so as not to conflict with\ndownstream changes in the git-for-windows repository related to fscache.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.h                      |  3 +++\n compat/win32/flush.c                | 29 +++++++++++++++++++++++++++++\n config.mak.uname                    |  2 ++\n contrib/buildsystems/CMakeLists.txt |  3 ++-\n wrapper.c                           |  4 ++++\n 5 files changed, 40 insertions(+), 1 deletion(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..c013920ce37\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,29 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       /* See https://docs.microsoft.com/en-us/windows-hardware/drivers/ddi/ntifs/nf-ntifs-ntflushbuffersfileex */\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex e6d482fbcc6..34c93314a50 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -451,6 +451,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -626,6 +627,7 @@ ifneq (,$(findstring MINGW,$(uname_S)))\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 171b4124afe..b573a5ee122 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/wrapper.c b/wrapper.c\nindex bb4f9f043ce..1a1e2fba9c9 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -567,6 +567,10 @@ int git_fsync(int fd, enum fsync_action action)\n \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n #endif\n \n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n \t\terrno = ENOSYS;\n \t\treturn -1;\n \n-- \ngitgitgadget\n\n"},{"id":"436510","messageId":"3e6b80b5fa25c5f1dfdbe299e088323c86dc8587.1632176111.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","subject":"[PATCH v4 6/6] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-20T22:15:11Z","receivedAt":"2021-09-21T02:21:59Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 43 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 46 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 89 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..e93c08a2e70\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..c9fcd0c03eb\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n"},{"id":"436691","messageId":"87mto58pkc.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"12cad737635663ed596e52f89f0f4f22f58bfe38.1632176111.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-21T23:16:24Z","receivedAt":"2021-09-21T23:41:13Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n\n> When the new mode is enabled we do the following for new objects:\n>\n> 1. Create a tmp_obj_XXXX file and write the object data to it.\n> 2. Issue a pagecache writeback request and wait for it to complete.\n> 3. Record the tmp name and the final name in the bulk-checkin state for\n>    later rename.\n>\n> At the end of the entire transaction we:\n> 1. Issue a fsync against the lock file to flush the hardware writeback\n>    cache, which should by now have processed the tmp file writes.\n> 2. Rename all of the temp files to their final names.\n> 3. When updating the index and/or refs, we assume that Git will issue\n>    another fsync internal to that operation.\n\nPerhaps note too that:\n\n4. For loose objects, refs etc. we may or may not create directories,\n   and most certainly will be updating metadata on the immediate\n   directory containing the file, but none of that's fsync()'d.\n\n> On a filesystem with a singular journal that is updated during name\n> operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n> would expect the fsync to trigger a journal writeout so that this\n> sequence is enough to ensure that the user's data is durable by the time\n> the git command returns.\n>\n> This change also updates the macOS code to trigger a real hardware flush\n> via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n> macOS there was no guarantee of durability since a simple fsync(2) call\n> does not flush any hardware caches.\n\nThere's no discussion of whether this is or isn't known to also work\nsome Linux FS's, and for these OS's where this does work is this only\nfor the object files themselves, or does metadata also \"ride along\"?\n\n> _Performance numbers_:\n>\n> Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\n> Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\n> Windows - Same host as Linux, a preview version of Windows 11.\n> \t  This number is from a patch later in the series.\n>\n> Adding 500 files to the repo with 'git add' Times reported in seconds.\n>\n> core.fsyncObjectFiles | Linux | Mac   | Windows\n> ----------------------|-------|-------|--------\n>                 false | 0.06  |  0.35 | 0.61\n>                 true  | 1.88  | 11.18 | 2.47\n>                 batch | 0.15  |  0.41 | 1.53\n\nPer my https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com\nand 6/6 in this series we've got perf tests for add/stash, but it would\nbe really interesting to see how this is impacted by\ntransfer.unpackLimit in cases where we may be writing packs or loose\nobjects.\n\n> [...]\n>  core.fsyncObjectFiles::\n> -\tThis boolean will enable 'fsync()' when writing object files.\n> -+\n> -This is a total waste of time and effort on a filesystem that orders\n> -data writes properly, but can be useful for filesystems that do not use\n> -journalling (traditional UNIX filesystems) or that only journal metadata\n> -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n> +\tA value indicating the level of effort Git will expend in\n> +\ttrying to make objects added to the repo durable in the event\n> +\tof an unclean system shutdown. This setting currently only\n> +\tcontrols the object store, so updates to any refs or the\n> +\tindex may not be equally durable.\n\nAll these mentions of \"object\" should really clarify that it's \"loose\nobjects\", i.e. we always fsync pack files. \n\n> +* `false` allows data to remain in file system caches according to\n> +  operating system policy, whence it may be lost if the system loses power\n> +  or crashes.\n\nAs noted in point #4 of\nhttps://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com/ while\nthis direction is overall an improvement over the previously flippant\ndocs, they at least alluded to the context that the assumption behind\n\"false\" is that you don't really care about loose objects, you care\nabout loose objects *and* the ref update or whatever.\n\nAs I think (this is from memory) we've covered already this may have\nbeen all based on some old ext3 assumption, but it's probably worth\nsummarizing that here, i.e. if you've got an FS with global ordered\noperations you can probably skip this, but probably not etc.\n\n> +* `true` triggers a data integrity flush for each object added to the\n> +  object store. This is the safest setting that is likely to ensure durability\n> +  across all operating systems and file systems that honor the 'fsync' system\n> +  call. However, this setting comes with a significant performance cost on\n> +  common hardware.\n\nThis is really overpromising things by omitting the fact that eve if\nwe're getting this feature you've hacked up right, we're still not\nfsyncing dir entries etc (also noted above).\n\nSo something that describes the narrow scope here, along with \"loose\nobjects\" etc....\n\n> +* `batch` enables an experimental mode that uses interfaces available in some\n> +  operating systems to write object data with a minimal set of FLUSH CACHE\n> +  (or equivalent) commands sent to the storage controller. If the operating\n> +  system interfaces are not available, this mode behaves the same as `true`.\n> +  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n> +  filesystems and on Windows for repos stored on NTFS or ReFS.\n\nAgain, even if it's called \"core.fsyncObjectFiles\" if we're going to say\n\"safe\" we really need to say safe in what sense. Having written and\nfsync()'d the file is helping nobody if the metadata never arrives....\n\n> +static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n> +{\n> +\tif (fsync_state->nr) {\n\nI think less indentation here would be nice:\n\n    if (!fsync_state->nr)\n        return;\n    /* rest of unindented body */\n\nOr better yet do this check in unplug_bulk_checkin(), then here:\n\n    fsync_or_die();\n    for_each_string_list_item() { ...}\n    string_list_clear(....);\n\n\n> +\t\tstruct string_list_item *rename;\n> +\n> +\t\t/*\n> +\t\t * Issue a full hardware flush against the lock file to ensure\n> +\t\t * that all objects are durable before any renames occur.\n> +\t\t * The code in fsync_and_close_loose_object_bulk_checkin has\n> +\t\t * already ensured that writeout has occurred, but it has not\n> +\t\t * flushed any writeback cache in the storage hardware.\n> +\t\t */\n> +\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n> +\n> +\t\tfor_each_string_list_item(rename, fsync_state) {\n> +\t\t\tconst char *src = rename->string;\n> +\t\t\tconst char *dst = rename->util;\n> +\n> +\t\t\tif (finalize_object_file(src, dst))\n> +\t\t\t\tdie_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n> +\t\t}\n> +\n> +\t\tstring_list_clear(fsync_state, 1);\n> +\t}\n> +}\n> +\n>  static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n>  {\n>  \tint i;\n> @@ -256,6 +286,53 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>  \treturn 0;\n>  }\n>  \n> +static void add_rename_bulk_checkin(struct string_list *fsync_state,\n> +\t\t\t\t    const char *src, const char *dst)\n> +{\n> +\tstring_list_insert(fsync_state, src)->util = xstrdup(dst);\n> +}\n\nJust has one caller, why not just inline the string_list_insert()\ncall...\n\n> +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n> +\t\t\t\t\t      const char *filename, time_t mtime)\n> +{\n> +\tint do_finalize = 1;\n> +\tint ret = 0;\n> +\n> +\tif (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n\nLet's do postive enum comparisons, and with switch() statements, so the\ncompiler helps us to see if we've covered them all.\n\n> +\t\t/*\n> +\t\t * If we have a plugged bulk checkin, we issue a call that\n> +\t\t * cleans the filesystem page cache but avoids a hardware flush\n> +\t\t * command. Later on we will issue a single hardware flush\n> +\t\t * before renaming files as part of do_sync_and_rename.\n> +\t\t */\n> +\t\tif (bulk_checkin_plugged &&\n> +\t\t    fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n> +\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n> +\t\t\tadd_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n> +\t\t\tdo_finalize = 0;\n> +\n> +\t\t} else {\n> +\t\t\tfsync_or_die(fd, \"loose object file\");\n> +\t\t}\n> +\t}\n\nSo nothing ever explicitly checks FSYNC_OBJECT_FILES_ON...?\n\n> -extern int fsync_object_files;\n> +enum FSYNC_OBJECT_FILES_MODE {\n> +    FSYNC_OBJECT_FILES_OFF,\n> +    FSYNC_OBJECT_FILES_ON,\n> +    FSYNC_OBJECT_FILES_BATCH\n> +};\n\nStyle: We don't use ALL_CAPS for type names in this codebase, just the\nenum labels themselves....\n\n> +extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n\n...to the point where I had to rub my eyes to see what was going on here\n... :)\n\n\n> -\t\tfsync_object_files = git_config_bool(var, value);\n> +\t\tif (value && !strcmp(value, \"batch\"))\n> +\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n> +\t\telse if (git_config_bool(var, value))\n> +\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n> +\t\telse\n> +\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n\nSince the point of this setting is safety, let's explicitly check\ntrue/false here, use git_config_maybe_bool(), and perhaps issue a\nwarning on unknown values, but maybe that would get too verbose...\n\nIf we have a future \"supersafe\" mode, it'll get mapped to \"false\" on\nolder versions of git, probably not a good idea...\n\n>  \t\treturn 0;\n>  \t}\n>  \n> diff --git a/config.mak.uname b/config.mak.uname\n> index 76516aaa9a5..e6d482fbcc6 100644\n> --- a/config.mak.uname\n> +++ b/config.mak.uname\n> @@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n>  \tHAVE_CLOCK_MONOTONIC = YesPlease\n>  \t# -lrt is needed for clock_gettime on glibc <= 2.16\n>  \tNEEDS_LIBRT = YesPlease\n> +\tHAVE_SYNC_FILE_RANGE = YesPlease\n>  \tHAVE_GETDELIM = YesPlease\n>  \tSANE_TEXT_GREP=-a\n>  \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\n> diff --git a/configure.ac b/configure.ac\n> index 031e8d3fee8..c711037d625 100644\n> --- a/configure.ac\n> +++ b/configure.ac\n> @@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n>  \t[AC_MSG_RESULT([no])\n>  \tHAVE_CLOCK_MONOTONIC=])\n>  GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n> +\n> +#\n> +# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n> +GIT_CHECK_FUNC(sync_file_range,\n> +\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n> +\t[HAVE_SYNC_FILE_RANGE])\n> +GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n> +\n>  #\n>  # Define NO_SETITIMER if you don't have setitimer.\n>  GIT_CHECK_FUNC(setitimer,\n> diff --git a/environment.c b/environment.c\n> index d6b22ede7ea..3e23eafff80 100644\n> --- a/environment.c\n> +++ b/environment.c\n> @@ -43,7 +43,7 @@ const char *git_hooks_path;\n>  int zlib_compression_level = Z_BEST_SPEED;\n>  int core_compression_level;\n>  int pack_compression_level = Z_DEFAULT_COMPRESSION;\n> -int fsync_object_files;\n> +enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n>  size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n>  size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n>  size_t delta_base_cache_limit = 96 * 1024 * 1024;\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index b46605300ab..d14e2436276 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n> @@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n>  void BUG(const char *fmt, ...);\n>  #endif\n>  \n> +enum fsync_action {\n> +    FSYNC_WRITEOUT_ONLY,\n> +    FSYNC_HARDWARE_FLUSH\n> +};\n> +\n> +int git_fsync(int fd, enum fsync_action action);\n> +\n>  /*\n>   * Preserves errno, prints a message, but gives no warning for ENOENT.\n>   * Returns 0 on success, which includes trying to unlink an object that does\n> diff --git a/object-file.c b/object-file.c\n> index a8be8994814..ea14c3a3483 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1859,15 +1859,6 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  \treturn 0;\n>  }\n>  \n> -/* Finalize a file on disk, and close it. */\n> -static void close_loose_object(int fd)\n> -{\n> -\tif (fsync_object_files)\n> -\t\tfsync_or_die(fd, \"loose object file\");\n> -\tif (close(fd) != 0)\n> -\t\tdie_errno(_(\"error when closing loose object file\"));\n> -}\n> -\n>  /* Size of directory component, including the ending '/' */\n>  static inline int directory_size(const char *filename)\n>  {\n> @@ -1973,17 +1964,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>  \n> -\tclose_loose_object(fd);\n> -\n> -\tif (mtime) {\n> -\t\tstruct utimbuf utb;\n> -\t\tutb.actime = mtime;\n> -\t\tutb.modtime = mtime;\n> -\t\tif (utime(tmp_file.buf, &utb) < 0)\n> -\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n> -\t}\n> -\n> -\treturn finalize_object_file(tmp_file.buf, filename.buf);\n> +\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf,\n> +\t\t\t\t\t\t\t filename.buf, mtime);\n>  }\n>  \n>  static int freshen_loose_object(const struct object_id *oid)\n> diff --git a/wrapper.c b/wrapper.c\n> index 7c6586af321..bb4f9f043ce 100644\n> --- a/wrapper.c\n> +++ b/wrapper.c\n> @@ -540,6 +540,50 @@ int xmkstemp_mode(char *filename_template, int mode)\n>  \treturn fd;\n>  }\n>  \n> +int git_fsync(int fd, enum fsync_action action)\n> +{\n> +\tswitch (action) {\n> +\tcase FSYNC_WRITEOUT_ONLY:\n> +\n> +#ifdef __APPLE__\n> +\t\t/*\n> +\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n> +\t\t * flush hardware caches.\n> +\t\t */\n> +\t\treturn fsync(fd);\n> +#endif\n> +\n> +#ifdef HAVE_SYNC_FILE_RANGE\n> +\t\t/*\n> +\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n> +\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n> +\t\t * indicates writeout of the entire file and the wait flags ensure that all\n> +\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n> +\t\t * before we continue.\n> +\t\t */\n> +\n> +\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n> +\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n> +#endif\n> +\n> +\t\terrno = ENOSYS;\n> +\t\treturn -1;\n> +\n> +\tcase FSYNC_HARDWARE_FLUSH:\n> +\n> +#ifdef __APPLE__\n> +\t\treturn fcntl(fd, F_FULLFSYNC);\n> +#else\n> +\t\treturn fsync(fd);\n> +#endif\n> +\n> +\tdefault:\n> +\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n> +\t}\n> +\n> +}\n> +\n>  static int warn_if_unremovable(const char *op, const char *file, int rc)\n>  {\n>  \tint err;\n> diff --git a/write-or-die.c b/write-or-die.c\n> index d33e68f6abb..8f53953d4ab 100644\n> --- a/write-or-die.c\n> +++ b/write-or-die.c\n> @@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n>  \n>  void fsync_or_die(int fd, const char *msg)\n>  {\n> -\twhile (fsync(fd) < 0) {\n> +\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n>  \t\tif (errno != EINTR)\n>  \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n>  \t}\n\n"},{"id":"436692","messageId":"87ilyt8per.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"a5b3e21b76208fba130e3313a3e70df45ab392af.1632176111.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 3/6] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-21T23:42:52Z","receivedAt":"2021-09-21T23:44:33Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n\n> +int win32_fsync_no_flush(int fd)\n> +{\n> +       IO_STATUS_BLOCK io_status;\n> +\n> +#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n> +\n> +       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n> +\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n> +\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n> +\n> +       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n> +\t\terrno = ENOSYS;\n> +\t\treturn -1;\n> +       }\n> +\n> +       /* See https://docs.microsoft.com/en-us/windows-hardware/drivers/ddi/ntifs/nf-ntifs-ntflushbuffersfileex */\n> +       memset(&io_status, 0, sizeof(io_status));\n\nSee just an informative link to the API docs, or is the comemnt on the\nmemset() in particular. This comment seems like it's just doing a\nGoogle/Bing search for you, so maybe better without it?\n"},{"id":"436693","messageId":"87ee9h8p0a.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"f7f756f3932cdbca587de397598758c685bac29a.1632176111.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 4/6] update-index: use the bulk-checkin infrastructure","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-21T23:46:57Z","receivedAt":"2021-09-21T23:53:14Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> The update-index functionality is used internally by 'git stash push' to\n> setup the internal stashed commit.\n>\n> This change enables bulk-checkin for update-index infrastructure to\n> speed up adding new objects to the object database by leveraging the\n> pack functionality and the new bulk-fsync functionality. This mode\n> is enabled when passing paths to update-index via the --stdin flag,\n> as is done by 'git stash'.\n>\n> There is some risk with this change, since under batch fsync, the object\n> files will not be available until the update-index is entirely complete.\n> This usage is unlikely, since any tool invoking update-index and\n> expecting to see objects would have to snoop the output of --verbose to\n> find out when update-index has actually processed a given path.\n> Additionally the index is locked for the duration of the update.\n\nWould you really need to sniff the verbose output? If I'm streaming data\nto update-index now it looks like I could assume before that\nupdate-index would have done the work if I managed to fflush() to it,\nsince it's processing a line at a time and doing the work in that\nline-at-a-time loop.\n\nI.e. you could print lines to it, and then do concurrent object lookups\nknowing the data was written already...\n\nI think this is probably fine, but that case seems way likelier than\nsomeone sniffing back the verbose output, presumably for the \"add\" in\nupdate_one(), but that's called in the getline_fn() loop...\n\nAll of this makes me wonder why this isn't using tmp-objdir.c, i.e. we\ncould have our cake and eat it too by writing the \"real\" objects, and\nthen just renaming them between directories instead. But perhaps the\nanswer has something to do with the metadata issues I raised.\n\nAnd well, tmp-objdir.c isn't going to help someone in practice that's\nrelying on this \"update-index --stdin\" behavior, as they won't know\nwhere we staged the temporary files...\n\n"},{"id":"436694","messageId":"87a6k58or9.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"afb0028e79648c1f7be8d77df5c6d675bd27d983.1632176111.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 5/6] core.fsyncobjectfiles: tests for batch mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-21T23:54:16Z","receivedAt":"2021-09-21T23:58:39Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Add test cases to exercise batch mode for 'git add'\n> and 'git stash'. These tests ensure that the added\n> data winds up in the object database.\n>\n> I verified the tests by introducing an incorrect rename\n> in do_sync_and_rename.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  t/lib-unique-files.sh | 34 ++++++++++++++++++++++++++++++++++\n>  t/t3700-add.sh        | 11 +++++++++++\n>  t/t3903-stash.sh      | 14 ++++++++++++++\n>  3 files changed, 59 insertions(+)\n>  create mode 100644 t/lib-unique-files.sh\n>\n> diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n> new file mode 100644\n> index 00000000000..a8a25eba61d\n> --- /dev/null\n> +++ b/t/lib-unique-files.sh\n> @@ -0,0 +1,34 @@\n> +# Helper to create files with unique contents\n> +\n> +test_create_unique_files_base__=$(date -u)\n> +test_create_unique_files_counter__=0\n> +\n> +# Create multiple files with unique contents. Takes the number of\n> +# directories, the number of files in each directory, and the base\n> +# directory.\n> +#\n> +# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n> +#\t\t\t\t    each in the specified directory, all\n> +#\t\t\t\t    with unique contents.\n> +\n> +test_create_unique_files() {\n> +\ttest \"$#\" -ne 3 && BUG \"3 param\"\n> +\n> +\tlocal dirs=$1\n> +\tlocal files=$2\n> +\tlocal basedir=$3\n> +\n> +\trm -rf $basedir >/dev/null\n\nWhy the >/dev/null? It's not a \"-rfv\", and any errors would go to\nstderr.\n\n> +\t\tmkdir -p \"$dir\" > /dev/null\n\nDitto.\n\n> +\t\tfor j in $(test_seq $files)\n> +\t\tdo\n> +\t\t\ttest_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n> +\t\t\techo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n\nWould be much more readable if we these variables were shorter.\n\nBut actually, why are we trying to create files as a function of \"date\n-u\" at all? This is all in the trash directory, which is rm -rf'd beween\nruns, why aren't names created with test_seq or whatever OK? I.e. just\n1.txt, 2.txt....\n\n> +test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n> +\ttest_create_unique_files 2 4 fsync-files &&\n> +\tgit -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n> +\trm -f fsynced_files &&\n> +\n> +\t# The files were untracked, so use the third parent,\n> +\t# which contains the untracked files\n> +\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n> +\ttest_line_count = 8 fsynced_files &&\n> +\tcat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n> +\"\n> +\n> +\n>  test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n>  \texpected=\"stash.useBuiltin support has been removed\" &&\n\nWe really prefer our tests to create the same data each time if\npossible, but as noted with the \"date -u\" comment above you're\nexplicitly bypassing that, but I still can't see why...\n"},{"id":"436697","messageId":"CANQDOdc1bNwDYhJ8ck2cwUfKmr3064uBHFDACphW+cGZRd-6EQ@mail.gmail.com","threadId":"56371","inReplyTo":"87mto58pkc.fsf@evledraar.gmail.com","subject":"Re: [PATCH v4 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-22T01:23:12Z","receivedAt":"2021-09-22T01:23:28Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 21, 2021 at 4:41 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > When the new mode is enabled we do the following for new objects:\n> >\n> > 1. Create a tmp_obj_XXXX file and write the object data to it.\n> > 2. Issue a pagecache writeback request and wait for it to complete.\n> > 3. Record the tmp name and the final name in the bulk-checkin state for\n> >    later rename.\n> >\n> > At the end of the entire transaction we:\n> > 1. Issue a fsync against the lock file to flush the hardware writeback\n> >    cache, which should by now have processed the tmp file writes.\n> > 2. Rename all of the temp files to their final names.\n> > 3. When updating the index and/or refs, we assume that Git will issue\n> >    another fsync internal to that operation.\n>\n> Perhaps note too that:\n>\n> 4. For loose objects, refs etc. we may or may not create directories,\n>    and most certainly will be updating metadata on the immediate\n>    directory containing the file, but none of that's fsync()'d.\n>\n> > On a filesystem with a singular journal that is updated during name\n> > operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n> > would expect the fsync to trigger a journal writeout so that this\n> > sequence is enough to ensure that the user's data is durable by the time\n> > the git command returns.\n> >\n> > This change also updates the macOS code to trigger a real hardware flush\n> > via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n> > macOS there was no guarantee of durability since a simple fsync(2) call\n> > does not flush any hardware caches.\n>\n> There's no discussion of whether this is or isn't known to also work\n> some Linux FS's, and for these OS's where this does work is this only\n> for the object files themselves, or does metadata also \"ride along\"?\n>\n\nI unfortunately can't examine Linux kernel source code and the details\nof metadata\nconsistency behavior across files is not something that anyone in that\ngroup wants\nto pin down. As far as I can tell, the only thing that's really\nguaranteed is fsyncing\nevery single file you write down and its parent directory if you're\ncreating a new file\n(which we always are).  As came up in conversation with Christoph\nHellwig elsewhere\non thread, Linux doesn't have any set of syscalls to make batch mode\nsafe.  It does look\nlike XFS would be safe if sync_file_ranges actually promised to wait\nfor all pagecache\nwriteback definitively, since it would do a \"log force\" to push all\nthe dirty metadata to\ndisk when we do our final fsync.\n\nI really didn't want to say something definitive about what Linux can\nor will do, since I'm\nnot in a position to really know or influence them.  Christoph did say\nthat he would be\ninterested in contributing a variant to this patch that would be\ndefinitively safe on filesystems\nthat honor syncfs.\n\n> > _Performance numbers_:\n> >\n> > Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\n> > Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\n> > Windows - Same host as Linux, a preview version of Windows 11.\n> >         This number is from a patch later in the series.\n> >\n> > Adding 500 files to the repo with 'git add' Times reported in seconds.\n> >\n> > core.fsyncObjectFiles | Linux | Mac   | Windows\n> > ----------------------|-------|-------|--------\n> >                 false | 0.06  |  0.35 | 0.61\n> >                 true  | 1.88  | 11.18 | 2.47\n> >                 batch | 0.15  |  0.41 | 1.53\n>\n> Per my https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com\n> and 6/6 in this series we've got perf tests for add/stash, but it would\n> be really interesting to see how this is impacted by\n> transfer.unpackLimit in cases where we may be writing packs or loose\n> objects.\n\nI'm having trouble understanding how unpackLimit is related to 'git stash'\nor 'git add'. From code inspection, it doesn't look like we're using\nthose settings\nfor adding objects except from across a transport.\n\nAre you proposing that we have a similar setting for adding objects\nvia 'add' using\na packfile?  I think that would be a good goal, but it might be a bit\ntricky since we've\nlikely done a lot of the work to buffer the input objects in order to\ncompute their OIDs,\nbefore we know how many objects there are to add. If the policy were\nto \"always add to\na packfile\", it would be easier.\n\n>\n> > [...]\n> >  core.fsyncObjectFiles::\n> > -     This boolean will enable 'fsync()' when writing object files.\n> > -+\n> > -This is a total waste of time and effort on a filesystem that orders\n> > -data writes properly, but can be useful for filesystems that do not use\n> > -journalling (traditional UNIX filesystems) or that only journal metadata\n> > -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n> > +     A value indicating the level of effort Git will expend in\n> > +     trying to make objects added to the repo durable in the event\n> > +     of an unclean system shutdown. This setting currently only\n> > +     controls the object store, so updates to any refs or the\n> > +     index may not be equally durable.\n>\n> All these mentions of \"object\" should really clarify that it's \"loose\n> objects\", i.e. we always fsync pack files.\n>\n> > +* `false` allows data to remain in file system caches according to\n> > +  operating system policy, whence it may be lost if the system loses power\n> > +  or crashes.\n>\n> As noted in point #4 of\n> https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com/ while\n> this direction is overall an improvement over the previously flippant\n> docs, they at least alluded to the context that the assumption behind\n> \"false\" is that you don't really care about loose objects, you care\n> about loose objects *and* the ref update or whatever.\n>\n> As I think (this is from memory) we've covered already this may have\n> been all based on some old ext3 assumption, but it's probably worth\n> summarizing that here, i.e. if you've got an FS with global ordered\n> operations you can probably skip this, but probably not etc.\n>\n> > +* `true` triggers a data integrity flush for each object added to the\n> > +  object store. This is the safest setting that is likely to ensure durability\n> > +  across all operating systems and file systems that honor the 'fsync' system\n> > +  call. However, this setting comes with a significant performance cost on\n> > +  common hardware.\n>\n> This is really overpromising things by omitting the fact that eve if\n> we're getting this feature you've hacked up right, we're still not\n> fsyncing dir entries etc (also noted above).\n>\n> So something that describes the narrow scope here, along with \"loose\n> objects\" etc....\n>\n> > +* `batch` enables an experimental mode that uses interfaces available in some\n> > +  operating systems to write object data with a minimal set of FLUSH CACHE\n> > +  (or equivalent) commands sent to the storage controller. If the operating\n> > +  system interfaces are not available, this mode behaves the same as `true`.\n> > +  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n> > +  filesystems and on Windows for repos stored on NTFS or ReFS.\n>\n> Again, even if it's called \"core.fsyncObjectFiles\" if we're going to say\n> \"safe\" we really need to say safe in what sense. Having written and\n> fsync()'d the file is helping nobody if the metadata never arrives....\n>\n\nMy concern with your feedback here is that this is user-facing documentation.\nI'd assume that people who are not intimately familiar with both their\nfilesystem\nand Git's internals would just be completely mystified by a long commentary on\nthe specifics in the Config documentation. I think over time Git should focus on\nmaking this setting really guarantee durability in a meaningful way\nacross the entire\nrepository.\n\n> > +static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n> > +{\n> > +     if (fsync_state->nr) {\n>\n> I think less indentation here would be nice:\n>\n>     if (!fsync_state->nr)\n>         return;\n>     /* rest of unindented body */\n>\n\nWill fix.\n\n> Or better yet do this check in unplug_bulk_checkin(), then here:\n>\n>     fsync_or_die();\n>     for_each_string_list_item() { ...}\n>     string_list_clear(....);\n>\n>\n\nI'd prefer to put it in the callee for reasons of\nseparation-of-concerns.  I don't want\nto have the caller and callee partially implement the contract. The\ncompiler should\ndo a good enough job, since it's only one caller and will probably get\ntotally inilined.\n\n> > +             struct string_list_item *rename;\n> > +\n> > +             /*\n> > +              * Issue a full hardware flush against the lock file to ensure\n> > +              * that all objects are durable before any renames occur.\n> > +              * The code in fsync_and_close_loose_object_bulk_checkin has\n> > +              * already ensured that writeout has occurred, but it has not\n> > +              * flushed any writeback cache in the storage hardware.\n> > +              */\n> > +             fsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n> > +\n> > +             for_each_string_list_item(rename, fsync_state) {\n> > +                     const char *src = rename->string;\n> > +                     const char *dst = rename->util;\n> > +\n> > +                     if (finalize_object_file(src, dst))\n> > +                             die_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n> > +             }\n> > +\n> > +             string_list_clear(fsync_state, 1);\n> > +     }\n> > +}\n> > +\n> >  static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n> >  {\n> >       int i;\n> > @@ -256,6 +286,53 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n> >       return 0;\n> >  }\n> >\n> > +static void add_rename_bulk_checkin(struct string_list *fsync_state,\n> > +                                 const char *src, const char *dst)\n> > +{\n> > +     string_list_insert(fsync_state, src)->util = xstrdup(dst);\n> > +}\n>\n> Just has one caller, why not just inline the string_list_insert()\n> call...\n>\n\nI thought about doing that before.  I'll do it.\n\n> > +int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n> > +                                           const char *filename, time_t mtime)\n> > +{\n> > +     int do_finalize = 1;\n> > +     int ret = 0;\n> > +\n> > +     if (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n>\n> Let's do postive enum comparisons, and with switch() statements, so the\n> compiler helps us to see if we've covered them all.\n>\n\nOk, will switch to switch.\n\n> > +             /*\n> > +              * If we have a plugged bulk checkin, we issue a call that\n> > +              * cleans the filesystem page cache but avoids a hardware flush\n> > +              * command. Later on we will issue a single hardware flush\n> > +              * before renaming files as part of do_sync_and_rename.\n> > +              */\n> > +             if (bulk_checkin_plugged &&\n> > +                 fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n> > +                 git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n> > +                     add_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n> > +                     do_finalize = 0;\n> > +\n> > +             } else {\n> > +                     fsync_or_die(fd, \"loose object file\");\n> > +             }\n> > +     }\n>\n> So nothing ever explicitly checks FSYNC_OBJECT_FILES_ON...?\n>\n\nYeah, I did it this way to avoid any code duplication, but I can change to\na switch if it doesn't require too much repetition.\n\n> > -extern int fsync_object_files;\n> > +enum FSYNC_OBJECT_FILES_MODE {\n> > +    FSYNC_OBJECT_FILES_OFF,\n> > +    FSYNC_OBJECT_FILES_ON,\n> > +    FSYNC_OBJECT_FILES_BATCH\n> > +};\n>\n> Style: We don't use ALL_CAPS for type names in this codebase, just the\n> enum labels themselves....\n>\n> > +extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n>\n> ...to the point where I had to rub my eyes to see what was going on here\n> ... :)\n>\n\nSorry, Windows Developer :). Will fix.\n\n\n> > -             fsync_object_files = git_config_bool(var, value);\n> > +             if (value && !strcmp(value, \"batch\"))\n> > +                     fsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n> > +             else if (git_config_bool(var, value))\n> > +                     fsync_object_files = FSYNC_OBJECT_FILES_ON;\n> > +             else\n> > +                     fsync_object_files = FSYNC_OBJECT_FILES_OFF;\n>\n> Since the point of this setting is safety, let's explicitly check\n> true/false here, use git_config_maybe_bool(), and perhaps issue a\n> warning on unknown values, but maybe that would get too verbose...\n>\n> If we have a future \"supersafe\" mode, it'll get mapped to \"false\" on\n> older versions of git, probably not a good idea...\n>\n\n I took Junio's suggestion verbatim.  I'll try a warning if the value\nexists, and is not 'batch' or <maybe bool>.\n\n\nThanks for looking at my changes so thoroughly!\n-Neeraj\n"},{"id":"436698","messageId":"CANQDOddiQHsVOwdCgW_r13FD5zaBwo1JJ+FbEip-F6AAcCdnMQ@mail.gmail.com","threadId":"56371","inReplyTo":"87ilyt8per.fsf@evledraar.gmail.com","subject":"Re: [PATCH v4 3/6] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-22T01:23:48Z","receivedAt":"2021-09-22T01:24:02Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 21, 2021 at 4:44 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > +int win32_fsync_no_flush(int fd)\n> > +{\n> > +       IO_STATUS_BLOCK io_status;\n> > +\n> > +#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n> > +\n> > +       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n> > +                      HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n> > +                      PIO_STATUS_BLOCK IoStatusBlock);\n> > +\n> > +       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n> > +             errno = ENOSYS;\n> > +             return -1;\n> > +       }\n> > +\n> > +       /* See https://docs.microsoft.com/en-us/windows-hardware/drivers/ddi/ntifs/nf-ntifs-ntflushbuffersfileex */\n> > +       memset(&io_status, 0, sizeof(io_status));\n>\n> See just an informative link to the API docs, or is the comemnt on the\n> memset() in particular. This comment seems like it's just doing a\n> Google/Bing search for you, so maybe better without it?\n\nWill remove. Just wanted to make sure everyone knows taht this is\ndocumented somewhere :).\n"},{"id":"436699","messageId":"CANQDOdeV4JuE8jnkzLmK6VfFj1t-+EOzvn=GD-ejPdS6unc66w@mail.gmail.com","threadId":"56371","inReplyTo":"87ee9h8p0a.fsf@evledraar.gmail.com","subject":"Re: [PATCH v4 4/6] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-22T01:27:25Z","receivedAt":"2021-09-22T01:27:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 21, 2021 at 4:53 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > The update-index functionality is used internally by 'git stash push' to\n> > setup the internal stashed commit.\n> >\n> > This change enables bulk-checkin for update-index infrastructure to\n> > speed up adding new objects to the object database by leveraging the\n> > pack functionality and the new bulk-fsync functionality. This mode\n> > is enabled when passing paths to update-index via the --stdin flag,\n> > as is done by 'git stash'.\n> >\n> > There is some risk with this change, since under batch fsync, the object\n> > files will not be available until the update-index is entirely complete.\n> > This usage is unlikely, since any tool invoking update-index and\n> > expecting to see objects would have to snoop the output of --verbose to\n> > find out when update-index has actually processed a given path.\n> > Additionally the index is locked for the duration of the update.\n>\n> Would you really need to sniff the verbose output? If I'm streaming data\n> to update-index now it looks like I could assume before that\n> update-index would have done the work if I managed to fflush() to it,\n> since it's processing a line at a time and doing the work in that\n> line-at-a-time loop.\n>\n> I.e. you could print lines to it, and then do concurrent object lookups\n> knowing the data was written already...\n>\n> I think this is probably fine, but that case seems way likelier than\n> someone sniffing back the verbose output, presumably for the \"add\" in\n> update_one(), but that's called in the getline_fn() loop...\n\nDoes fflush really guarantee that the reader has picked up the input from\na pipe across all environments?  Even if a reader picks up the input, does\nthat mean that the reader is done processing it?\n\nDo you think I really need to revise this comment? Maybe leave a terser,\n'this usage is thought to be unlikely'?\n\n>\n> All of this makes me wonder why this isn't using tmp-objdir.c, i.e. we\n> could have our cake and eat it too by writing the \"real\" objects, and\n> then just renaming them between directories instead. But perhaps the\n> answer has something to do with the metadata issues I raised.\n>\n> And well, tmp-objdir.c isn't going to help someone in practice that's\n> relying on this \"update-index --stdin\" behavior, as they won't know\n> where we staged the temporary files...\n>\n\nOne motivation of the current design behind renaming the files is that\nsome networked filesystems don't seem to like cross-directory renames\nmuch.  It also so happens that ReFS on Windows also prefers renames to\nstay within the directory. Actually any filesystem would likely be\nslightly faster,\nsince fewer objects are being modified (one dir versus two).\n"},{"id":"436700","messageId":"CANQDOdc3R0H_DS9eqh9vK7JA6=2LsB6q9mANgMgrtbOwp7Fvwg@mail.gmail.com","threadId":"56371","inReplyTo":"87a6k58or9.fsf@evledraar.gmail.com","subject":"Re: [PATCH v4 5/6] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-22T01:30:55Z","receivedAt":"2021-09-22T01:31:10Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 21, 2021 at 4:58 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Add test cases to exercise batch mode for 'git add'\n> > and 'git stash'. These tests ensure that the added\n> > data winds up in the object database.\n> >\n> > I verified the tests by introducing an incorrect rename\n> > in do_sync_and_rename.\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  t/lib-unique-files.sh | 34 ++++++++++++++++++++++++++++++++++\n> >  t/t3700-add.sh        | 11 +++++++++++\n> >  t/t3903-stash.sh      | 14 ++++++++++++++\n> >  3 files changed, 59 insertions(+)\n> >  create mode 100644 t/lib-unique-files.sh\n> >\n> > diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n> > new file mode 100644\n> > index 00000000000..a8a25eba61d\n> > --- /dev/null\n> > +++ b/t/lib-unique-files.sh\n> > @@ -0,0 +1,34 @@\n> > +# Helper to create files with unique contents\n> > +\n> > +test_create_unique_files_base__=$(date -u)\n> > +test_create_unique_files_counter__=0\n> > +\n> > +# Create multiple files with unique contents. Takes the number of\n> > +# directories, the number of files in each directory, and the base\n> > +# directory.\n> > +#\n> > +# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n> > +#                                each in the specified directory, all\n> > +#                                with unique contents.\n> > +\n> > +test_create_unique_files() {\n> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n> > +\n> > +     local dirs=$1\n> > +     local files=$2\n> > +     local basedir=$3\n> > +\n> > +     rm -rf $basedir >/dev/null\n>\n> Why the >/dev/null? It's not a \"-rfv\", and any errors would go to\n> stderr.\n\nWill fix. Clearly I don't know UNIX very well.\n\n>\n> > +             mkdir -p \"$dir\" > /dev/null\n>\n> Ditto.\n\nWill fix.\n\n>\n> > +             for j in $(test_seq $files)\n> > +             do\n> > +                     test_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n> > +                     echo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n>\n> Would be much more readable if we these variables were shorter.\n>\n> But actually, why are we trying to create files as a function of \"date\n> -u\" at all? This is all in the trash directory, which is rm -rf'd beween\n> runs, why aren't names created with test_seq or whatever OK? I.e. just\n> 1.txt, 2.txt....\n>\n\nThe uniqueness is in the contents of the file.  I wanted to make sure that\nwe are really creating new objects and not reusing old ones.  Is the scope\nof the \"trash repo\" small enough that I can be guaranteed that a new one\nis created before my test since the last time I tried adding something to\nthe ODB?\n\n> > +test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n> > +     test_create_unique_files 2 4 fsync-files &&\n> > +     git -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n> > +     rm -f fsynced_files &&\n> > +\n> > +     # The files were untracked, so use the third parent,\n> > +     # which contains the untracked files\n> > +     git ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n> > +     test_line_count = 8 fsynced_files &&\n> > +     cat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n> > +\"\n> > +\n> > +\n> >  test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n> >       expected=\"stash.useBuiltin support has been removed\" &&\n>\n> We really prefer our tests to create the same data each time if\n> possible, but as noted with the \"date -u\" comment above you're\n> explicitly bypassing that, but I still can't see why...\n\nI'm trying to make sure we get new object contents. Is there a better\nway to achieve what I want without the risk of finding that the contents\nare already in the database from a previous test run?\n\nThanks again for the thorough review,\n-Neeraj\n"},{"id":"436701","messageId":"87wnn974gx.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdc3R0H_DS9eqh9vK7JA6=2LsB6q9mANgMgrtbOwp7Fvwg@mail.gmail.com","subject":"Re: [PATCH v4 5/6] core.fsyncobjectfiles: tests for batch mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-22T01:58:56Z","receivedAt":"2021-09-22T02:02:10Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Sep 21 2021, Neeraj Singh wrote:\n\n> On Tue, Sep 21, 2021 at 4:58 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n>>\n>> > From: Neeraj Singh <neerajsi@microsoft.com>\n>> >\n>> > Add test cases to exercise batch mode for 'git add'\n>> > and 'git stash'. These tests ensure that the added\n>> > data winds up in the object database.\n>> >\n>> > I verified the tests by introducing an incorrect rename\n>> > in do_sync_and_rename.\n>> >\n>> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>> > ---\n>> >  t/lib-unique-files.sh | 34 ++++++++++++++++++++++++++++++++++\n>> >  t/t3700-add.sh        | 11 +++++++++++\n>> >  t/t3903-stash.sh      | 14 ++++++++++++++\n>> >  3 files changed, 59 insertions(+)\n>> >  create mode 100644 t/lib-unique-files.sh\n>> >\n>> > diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n>> > new file mode 100644\n>> > index 00000000000..a8a25eba61d\n>> > --- /dev/null\n>> > +++ b/t/lib-unique-files.sh\n>> > @@ -0,0 +1,34 @@\n>> > +# Helper to create files with unique contents\n>> > +\n>> > +test_create_unique_files_base__=$(date -u)\n>> > +test_create_unique_files_counter__=0\n>> > +\n>> > +# Create multiple files with unique contents. Takes the number of\n>> > +# directories, the number of files in each directory, and the base\n>> > +# directory.\n>> > +#\n>> > +# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n>> > +#                                each in the specified directory, all\n>> > +#                                with unique contents.\n>> > +\n>> > +test_create_unique_files() {\n>> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n>> > +\n>> > +     local dirs=$1\n>> > +     local files=$2\n>> > +     local basedir=$3\n>> > +\n>> > +     rm -rf $basedir >/dev/null\n>>\n>> Why the >/dev/null? It's not a \"-rfv\", and any errors would go to\n>> stderr.\n>\n> Will fix. Clearly I don't know UNIX very well.\n>\n>>\n>> > +             mkdir -p \"$dir\" > /dev/null\n>>\n>> Ditto.\n>\n> Will fix.\n>\n>>\n>> > +             for j in $(test_seq $files)\n>> > +             do\n>> > +                     test_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n>> > +                     echo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n>>\n>> Would be much more readable if we these variables were shorter.\n>>\n>> But actually, why are we trying to create files as a function of \"date\n>> -u\" at all? This is all in the trash directory, which is rm -rf'd beween\n>> runs, why aren't names created with test_seq or whatever OK? I.e. just\n>> 1.txt, 2.txt....\n>>\n>\n> The uniqueness is in the contents of the file.  I wanted to make sure that\n> we are really creating new objects and not reusing old ones.  Is the scope\n> of the \"trash repo\" small enough that I can be guaranteed that a new one\n> is created before my test since the last time I tried adding something to\n> the ODB?\n>\n>> > +test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n>> > +     test_create_unique_files 2 4 fsync-files &&\n>> > +     git -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n>> > +     rm -f fsynced_files &&\n>> > +\n>> > +     # The files were untracked, so use the third parent,\n>> > +     # which contains the untracked files\n>> > +     git ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n>> > +     test_line_count = 8 fsynced_files &&\n>> > +     cat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n>> > +\"\n>> > +\n>> > +\n>> >  test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n>> >       expected=\"stash.useBuiltin support has been removed\" &&\n>>\n>> We really prefer our tests to create the same data each time if\n>> possible, but as noted with the \"date -u\" comment above you're\n>> explicitly bypassing that, but I still can't see why...\n>\n> I'm trying to make sure we get new object contents. Is there a better\n> way to achieve what I want without the risk of finding that the contents\n> are already in the database from a previous test run?\n\nYou can just do something like:\n\ntest_expect_success 'setup data' '\n\ttest_commit A &&\n\ttest_commit B\n'\n\nWhich will create files A.t, B.t etc, or create them via:\n\n    obj=$(echo foo | git hash-object -w --stdin)\n\netc.\n\nI.e. the uniqueness you're doing here seems to assume that tests are\nre-using the same object store across runs, but we create a new trash\ndirectory for each one, if you run the test with \"-d\" you can see it\nbeing left behind for inspection. This is already ensured for the test.\n\nThe only potential caveat I can imagine is that some filesystem like say\nbtrfs-like that does some COW or object de-duplication would behave\ndifferently, but other than that...\n"},{"id":"436703","messageId":"87sfxx73vm.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdc1bNwDYhJ8ck2cwUfKmr3064uBHFDACphW+cGZRd-6EQ@mail.gmail.com","subject":"Re: [PATCH v4 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-22T02:02:29Z","receivedAt":"2021-09-22T02:14:57Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Sep 21 2021, Neeraj Singh wrote:\n\n> On Tue, Sep 21, 2021 at 4:41 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n>>\n>> > When the new mode is enabled we do the following for new objects:\n>> >\n>> > 1. Create a tmp_obj_XXXX file and write the object data to it.\n>> > 2. Issue a pagecache writeback request and wait for it to complete.\n>> > 3. Record the tmp name and the final name in the bulk-checkin state for\n>> >    later rename.\n>> >\n>> > At the end of the entire transaction we:\n>> > 1. Issue a fsync against the lock file to flush the hardware writeback\n>> >    cache, which should by now have processed the tmp file writes.\n>> > 2. Rename all of the temp files to their final names.\n>> > 3. When updating the index and/or refs, we assume that Git will issue\n>> >    another fsync internal to that operation.\n>>\n>> Perhaps note too that:\n>>\n>> 4. For loose objects, refs etc. we may or may not create directories,\n>>    and most certainly will be updating metadata on the immediate\n>>    directory containing the file, but none of that's fsync()'d.\n>>\n>> > On a filesystem with a singular journal that is updated during name\n>> > operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n>> > would expect the fsync to trigger a journal writeout so that this\n>> > sequence is enough to ensure that the user's data is durable by the time\n>> > the git command returns.\n>> >\n>> > This change also updates the macOS code to trigger a real hardware flush\n>> > via fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\n>> > macOS there was no guarantee of durability since a simple fsync(2) call\n>> > does not flush any hardware caches.\n>>\n>> There's no discussion of whether this is or isn't known to also work\n>> some Linux FS's, and for these OS's where this does work is this only\n>> for the object files themselves, or does metadata also \"ride along\"?\n>>\n>\n> I unfortunately can't examine Linux kernel source code and the details\n> of metadata\n> consistency behavior across files is not something that anyone in that\n> group wants\n> to pin down. As far as I can tell, the only thing that's really\n> guaranteed is fsyncing\n> every single file you write down and its parent directory if you're\n> creating a new file\n> (which we always are).  As came up in conversation with Christoph\n> Hellwig elsewhere\n> on thread, Linux doesn't have any set of syscalls to make batch mode\n> safe.  It does look\n> like XFS would be safe if sync_file_ranges actually promised to wait\n> for all pagecache\n> writeback definitively, since it would do a \"log force\" to push all\n> the dirty metadata to\n> disk when we do our final fsync.\n>\n> I really didn't want to say something definitive about what Linux can\n> or will do, since I'm\n> not in a position to really know or influence them.  Christoph did say\n> that he would be\n> interested in contributing a variant to this patch that would be\n> definitively safe on filesystems\n> that honor syncfs.\n\n*nod*, it's fine if it's omitted. Just wondering if we knew but weren't\n saying etc.\n\n>> > _Performance numbers_:\n>> >\n>> > Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\n>> > Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\n>> > Windows - Same host as Linux, a preview version of Windows 11.\n>> >         This number is from a patch later in the series.\n>> >\n>> > Adding 500 files to the repo with 'git add' Times reported in seconds.\n>> >\n>> > core.fsyncObjectFiles | Linux | Mac   | Windows\n>> > ----------------------|-------|-------|--------\n>> >                 false | 0.06  |  0.35 | 0.61\n>> >                 true  | 1.88  | 11.18 | 2.47\n>> >                 batch | 0.15  |  0.41 | 1.53\n>>\n>> Per my https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com\n>> and 6/6 in this series we've got perf tests for add/stash, but it would\n>> be really interesting to see how this is impacted by\n>> transfer.unpackLimit in cases where we may be writing packs or loose\n>> objects.\n>\n> I'm having trouble understanding how unpackLimit is related to 'git stash'\n> or 'git add'. From code inspection, it doesn't look like we're using\n> those settings\n> for adding objects except from across a transport.\n>\n> Are you proposing that we have a similar setting for adding objects\n> via 'add' using\n> a packfile?  I think that would be a good goal, but it might be a bit\n> tricky since we've\n> likely done a lot of the work to buffer the input objects in order to\n> compute their OIDs,\n> before we know how many objects there are to add. If the policy were\n> to \"always add to\n> a packfile\", it would be easier.\n\nNo, just that in the documentation that we should be explaining to the\nreader that this mode that optimizes for loose object writing benefits\nparticular commands, but e.g. on the server-side that we'll probably\nnever write 500 objects, but stream them to one pack.\n\nWhich might also inform next steps for the commands this does help with,\ni.e. can we make more things stream to packs? I think having this mode\nis at worst a good transitory thing to have, but perhaps longer term\nwe'll want to simply write fewer individual loose objects.\n\nIn any case, pushing to a server with this configured and scaling that\nby transfer.unpackLimit should nicely demonstrate the pack v.s. loose\nobject scenario at different fsck-settings.\n\n>>\n>> > [...]\n>> >  core.fsyncObjectFiles::\n>> > -     This boolean will enable 'fsync()' when writing object files.\n>> > -+\n>> > -This is a total waste of time and effort on a filesystem that orders\n>> > -data writes properly, but can be useful for filesystems that do not use\n>> > -journalling (traditional UNIX filesystems) or that only journal metadata\n>> > -and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n>> > +     A value indicating the level of effort Git will expend in\n>> > +     trying to make objects added to the repo durable in the event\n>> > +     of an unclean system shutdown. This setting currently only\n>> > +     controls the object store, so updates to any refs or the\n>> > +     index may not be equally durable.\n>>\n>> All these mentions of \"object\" should really clarify that it's \"loose\n>> objects\", i.e. we always fsync pack files.\n>>\n>> > +* `false` allows data to remain in file system caches according to\n>> > +  operating system policy, whence it may be lost if the system loses power\n>> > +  or crashes.\n>>\n>> As noted in point #4 of\n>> https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com/ while\n>> this direction is overall an improvement over the previously flippant\n>> docs, they at least alluded to the context that the assumption behind\n>> \"false\" is that you don't really care about loose objects, you care\n>> about loose objects *and* the ref update or whatever.\n>>\n>> As I think (this is from memory) we've covered already this may have\n>> been all based on some old ext3 assumption, but it's probably worth\n>> summarizing that here, i.e. if you've got an FS with global ordered\n>> operations you can probably skip this, but probably not etc.\n>>\n>> > +* `true` triggers a data integrity flush for each object added to the\n>> > +  object store. This is the safest setting that is likely to ensure durability\n>> > +  across all operating systems and file systems that honor the 'fsync' system\n>> > +  call. However, this setting comes with a significant performance cost on\n>> > +  common hardware.\n>>\n>> This is really overpromising things by omitting the fact that eve if\n>> we're getting this feature you've hacked up right, we're still not\n>> fsyncing dir entries etc (also noted above).\n>>\n>> So something that describes the narrow scope here, along with \"loose\n>> objects\" etc....\n>>\n>> > +* `batch` enables an experimental mode that uses interfaces available in some\n>> > +  operating systems to write object data with a minimal set of FLUSH CACHE\n>> > +  (or equivalent) commands sent to the storage controller. If the operating\n>> > +  system interfaces are not available, this mode behaves the same as `true`.\n>> > +  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n>> > +  filesystems and on Windows for repos stored on NTFS or ReFS.\n>>\n>> Again, even if it's called \"core.fsyncObjectFiles\" if we're going to say\n>> \"safe\" we really need to say safe in what sense. Having written and\n>> fsync()'d the file is helping nobody if the metadata never arrives....\n>>\n>\n> My concern with your feedback here is that this is user-facing documentation.\n> I'd assume that people who are not intimately familiar with both their\n> filesystem\n> and Git's internals would just be completely mystified by a long commentary on\n> the specifics in the Config documentation. I think over time Git should focus on\n> making this setting really guarantee durability in a meaningful way\n> across the entire\n> repository.\n\nYeah, this setting though is probably going to be tweaked only by fairly\nexpert-level users of git.\n\nI think it's fine if it just explicitly punts and says something like\n'this is what it does, this may or may not work on your FS' etc., my\nmain issue with the current docs is that they give off this vibe of\nknowing a lot more than they're telling you.\n\n>> > +static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n>> > +{\n>> > +     if (fsync_state->nr) {\n>>\n>> I think less indentation here would be nice:\n>>\n>>     if (!fsync_state->nr)\n>>         return;\n>>     /* rest of unindented body */\n>>\n>\n> Will fix.\n>\n>> Or better yet do this check in unplug_bulk_checkin(), then here:\n>>\n>>     fsync_or_die();\n>>     for_each_string_list_item() { ...}\n>>     string_list_clear(....);\n>>\n>>\n>\n> I'd prefer to put it in the callee for reasons of\n> separation-of-concerns.  I don't want\n> to have the caller and callee partially implement the contract. The\n> compiler should\n> do a good enough job, since it's only one caller and will probably get\n> totally inilined.\n\n*nod*\n\nFor what it's worth I meant the \"inlined\" just in terms of avoiding the\nindirection for human readers, it won't matter to the machine,\nespecially since this is all I/O bound...\n"},{"id":"436744","messageId":"CANQDOdfQZLp--W7d4iuDO6n5YxMM6CxmbbFjfOSLN1=-Li9w-g@mail.gmail.com","threadId":"56371","inReplyTo":"87wnn974gx.fsf@evledraar.gmail.com","subject":"Re: [PATCH v4 5/6] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-22T17:55:52Z","receivedAt":"2021-09-22T17:56:07Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 21, 2021 at 7:02 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Tue, Sep 21 2021, Neeraj Singh wrote:\n>\n> > On Tue, Sep 21, 2021 at 4:58 PM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n> >>\n> >> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >> >\n> >> > Add test cases to exercise batch mode for 'git add'\n> >> > and 'git stash'. These tests ensure that the added\n> >> > data winds up in the object database.\n> >> >\n> >> > I verified the tests by introducing an incorrect rename\n> >> > in do_sync_and_rename.\n> >> >\n> >> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> >> > ---\n> >> >  t/lib-unique-files.sh | 34 ++++++++++++++++++++++++++++++++++\n> >> >  t/t3700-add.sh        | 11 +++++++++++\n> >> >  t/t3903-stash.sh      | 14 ++++++++++++++\n> >> >  3 files changed, 59 insertions(+)\n> >> >  create mode 100644 t/lib-unique-files.sh\n> >> >\n> >> > diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n> >> > new file mode 100644\n> >> > index 00000000000..a8a25eba61d\n> >> > --- /dev/null\n> >> > +++ b/t/lib-unique-files.sh\n> >> > @@ -0,0 +1,34 @@\n> >> > +# Helper to create files with unique contents\n> >> > +\n> >> > +test_create_unique_files_base__=$(date -u)\n> >> > +test_create_unique_files_counter__=0\n> >> > +\n> >> > +# Create multiple files with unique contents. Takes the number of\n> >> > +# directories, the number of files in each directory, and the base\n> >> > +# directory.\n> >> > +#\n> >> > +# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n> >> > +#                                each in the specified directory, all\n> >> > +#                                with unique contents.\n> >> > +\n> >> > +test_create_unique_files() {\n> >> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n> >> > +\n> >> > +     local dirs=$1\n> >> > +     local files=$2\n> >> > +     local basedir=$3\n> >> > +\n> >> > +     rm -rf $basedir >/dev/null\n> >>\n> >> Why the >/dev/null? It's not a \"-rfv\", and any errors would go to\n> >> stderr.\n> >\n> > Will fix. Clearly I don't know UNIX very well.\n> >\n> >>\n> >> > +             mkdir -p \"$dir\" > /dev/null\n> >>\n> >> Ditto.\n> >\n> > Will fix.\n> >\n> >>\n> >> > +             for j in $(test_seq $files)\n> >> > +             do\n> >> > +                     test_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n> >> > +                     echo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n> >>\n> >> Would be much more readable if we these variables were shorter.\n> >>\n> >> But actually, why are we trying to create files as a function of \"date\n> >> -u\" at all? This is all in the trash directory, which is rm -rf'd beween\n> >> runs, why aren't names created with test_seq or whatever OK? I.e. just\n> >> 1.txt, 2.txt....\n> >>\n> >\n> > The uniqueness is in the contents of the file.  I wanted to make sure that\n> > we are really creating new objects and not reusing old ones.  Is the scope\n> > of the \"trash repo\" small enough that I can be guaranteed that a new one\n> > is created before my test since the last time I tried adding something to\n> > the ODB?\n> >\n> >> > +test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n> >> > +     test_create_unique_files 2 4 fsync-files &&\n> >> > +     git -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n> >> > +     rm -f fsynced_files &&\n> >> > +\n> >> > +     # The files were untracked, so use the third parent,\n> >> > +     # which contains the untracked files\n> >> > +     git ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n> >> > +     test_line_count = 8 fsynced_files &&\n> >> > +     cat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n> >> > +\"\n> >> > +\n> >> > +\n> >> >  test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n> >> >       expected=\"stash.useBuiltin support has been removed\" &&\n> >>\n> >> We really prefer our tests to create the same data each time if\n> >> possible, but as noted with the \"date -u\" comment above you're\n> >> explicitly bypassing that, but I still can't see why...\n> >\n> > I'm trying to make sure we get new object contents. Is there a better\n> > way to achieve what I want without the risk of finding that the contents\n> > are already in the database from a previous test run?\n>\n> You can just do something like:\n>\n> test_expect_success 'setup data' '\n>         test_commit A &&\n>         test_commit B\n> '\n>\n> Which will create files A.t, B.t etc, or create them via:\n>\n>     obj=$(echo foo | git hash-object -w --stdin)\n>\n> etc.\n>\n> I.e. the uniqueness you're doing here seems to assume that tests are\n> re-using the same object store across runs, but we create a new trash\n> directory for each one, if you run the test with \"-d\" you can see it\n> being left behind for inspection. This is already ensured for the test.\n>\n> The only potential caveat I can imagine is that some filesystem like say\n> btrfs-like that does some COW or object de-duplication would behave\n> differently, but other than that...\n\nIt looks like the same repo is reused for each test_expect_success\nline in the top-level t*.sh script.\nSo for test_create_unique_files to be maximally useful, it should have\nsome state that is different for\neach invocation.  How about I use the test_tick mechanism to produce\nthis uniqueness?  It wouldn't\nbe globally unique like the date method, but it should be good enough\nif the repo is recycled every time\ntest-lib is reinitialized.\n\nI'm changing lib-unique-files to use test_tick and to be a little more\nreadable as you suggested. Please\nlet me know if you have any other suggestions.\n"},{"id":"436763","messageId":"CANQDOddTQZz8+LjX2wB5j+tSO9kj6S9VJGTvxKKU8--NQt7PSw@mail.gmail.com","threadId":"56371","inReplyTo":"CANQDOdc1bNwDYhJ8ck2cwUfKmr3064uBHFDACphW+cGZRd-6EQ@mail.gmail.com","subject":"Re: [PATCH v4 2/6] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-22T19:46:43Z","receivedAt":"2021-09-22T19:46:58Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 21, 2021 at 6:23 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> On Tue, Sep 21, 2021 at 4:41 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n> >\n> >\n> > On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n\n> > > -             fsync_object_files = git_config_bool(var, value);\n> > > +             if (value && !strcmp(value, \"batch\"))\n> > > +                     fsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n> > > +             else if (git_config_bool(var, value))\n> > > +                     fsync_object_files = FSYNC_OBJECT_FILES_ON;\n> > > +             else\n> > > +                     fsync_object_files = FSYNC_OBJECT_FILES_OFF;\n> >\n> > Since the point of this setting is safety, let's explicitly check\n> > true/false here, use git_config_maybe_bool(), and perhaps issue a\n> > warning on unknown values, but maybe that would get too verbose...\n> >\n> > If we have a future \"supersafe\" mode, it'll get mapped to \"false\" on\n> > older versions of git, probably not a good idea...\n> >\n>\n>  I took Junio's suggestion verbatim.  I'll try a warning if the value\n> exists, and is not 'batch' or <maybe bool>.\n\nAn update on this.  I tested out some values:\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=batch add ./\n    fsync_object_files: 2\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=0 add ./\n    fsync_object_files: 0\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=1 add ./\n    fsync_object_files: 1\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=2 add ./\n    fsync_object_files: 1\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=barf add ./\n    fatal: bad boolean config value 'barf' for 'core.fsyncobjectfiles'\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=true add ./\n    fsync_object_files: 1\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=false add ./\n    fsync_object_files: 0\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=t add ./\n    fatal: bad boolean config value 't' for 'core.fsyncobjectfiles'\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=y add ./\n    fatal: bad boolean config value 'y' for 'core.fsyncobjectfiles'\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=yes add ./\n    fsync_object_files: 1\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=no add ./\n    fsync_object_files: 0\n    nksingh@neerajsi-x1:~/src/git$ ./git -c core.fsyncobjectfiles=nope add ./\n    fatal: bad boolean config value 'nope' for 'core.fsyncobjectfiles'\n\nSo I think the code already works like you are suggesting (thanks Junio!).\n"},{"id":"436767","messageId":"877df874xn.fsf@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdfQZLp--W7d4iuDO6n5YxMM6CxmbbFjfOSLN1=-Li9w-g@mail.gmail.com","subject":"Re: [PATCH v4 5/6] core.fsyncobjectfiles: tests for batch mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-09-22T20:01:36Z","receivedAt":"2021-09-22T20:04:25Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Sep 22 2021, Neeraj Singh wrote:\n\n> On Tue, Sep 21, 2021 at 7:02 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Tue, Sep 21 2021, Neeraj Singh wrote:\n>>\n>> > On Tue, Sep 21, 2021 at 4:58 PM Ævar Arnfjörð Bjarmason\n>> > <avarab@gmail.com> wrote:\n>> >>\n>> >>\n>> >> On Mon, Sep 20 2021, Neeraj Singh via GitGitGadget wrote:\n>> >>\n>> >> > From: Neeraj Singh <neerajsi@microsoft.com>\n>> >> >\n>> >> > Add test cases to exercise batch mode for 'git add'\n>> >> > and 'git stash'. These tests ensure that the added\n>> >> > data winds up in the object database.\n>> >> >\n>> >> > I verified the tests by introducing an incorrect rename\n>> >> > in do_sync_and_rename.\n>> >> >\n>> >> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>> >> > ---\n>> >> >  t/lib-unique-files.sh | 34 ++++++++++++++++++++++++++++++++++\n>> >> >  t/t3700-add.sh        | 11 +++++++++++\n>> >> >  t/t3903-stash.sh      | 14 ++++++++++++++\n>> >> >  3 files changed, 59 insertions(+)\n>> >> >  create mode 100644 t/lib-unique-files.sh\n>> >> >\n>> >> > diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n>> >> > new file mode 100644\n>> >> > index 00000000000..a8a25eba61d\n>> >> > --- /dev/null\n>> >> > +++ b/t/lib-unique-files.sh\n>> >> > @@ -0,0 +1,34 @@\n>> >> > +# Helper to create files with unique contents\n>> >> > +\n>> >> > +test_create_unique_files_base__=$(date -u)\n>> >> > +test_create_unique_files_counter__=0\n>> >> > +\n>> >> > +# Create multiple files with unique contents. Takes the number of\n>> >> > +# directories, the number of files in each directory, and the base\n>> >> > +# directory.\n>> >> > +#\n>> >> > +# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n>> >> > +#                                each in the specified directory, all\n>> >> > +#                                with unique contents.\n>> >> > +\n>> >> > +test_create_unique_files() {\n>> >> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n>> >> > +\n>> >> > +     local dirs=$1\n>> >> > +     local files=$2\n>> >> > +     local basedir=$3\n>> >> > +\n>> >> > +     rm -rf $basedir >/dev/null\n>> >>\n>> >> Why the >/dev/null? It's not a \"-rfv\", and any errors would go to\n>> >> stderr.\n>> >\n>> > Will fix. Clearly I don't know UNIX very well.\n>> >\n>> >>\n>> >> > +             mkdir -p \"$dir\" > /dev/null\n>> >>\n>> >> Ditto.\n>> >\n>> > Will fix.\n>> >\n>> >>\n>> >> > +             for j in $(test_seq $files)\n>> >> > +             do\n>> >> > +                     test_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n>> >> > +                     echo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n>> >>\n>> >> Would be much more readable if we these variables were shorter.\n>> >>\n>> >> But actually, why are we trying to create files as a function of \"date\n>> >> -u\" at all? This is all in the trash directory, which is rm -rf'd beween\n>> >> runs, why aren't names created with test_seq or whatever OK? I.e. just\n>> >> 1.txt, 2.txt....\n>> >>\n>> >\n>> > The uniqueness is in the contents of the file.  I wanted to make sure that\n>> > we are really creating new objects and not reusing old ones.  Is the scope\n>> > of the \"trash repo\" small enough that I can be guaranteed that a new one\n>> > is created before my test since the last time I tried adding something to\n>> > the ODB?\n>> >\n>> >> > +test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n>> >> > +     test_create_unique_files 2 4 fsync-files &&\n>> >> > +     git -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n>> >> > +     rm -f fsynced_files &&\n>> >> > +\n>> >> > +     # The files were untracked, so use the third parent,\n>> >> > +     # which contains the untracked files\n>> >> > +     git ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n>> >> > +     test_line_count = 8 fsynced_files &&\n>> >> > +     cat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n>> >> > +\"\n>> >> > +\n>> >> > +\n>> >> >  test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n>> >> >       expected=\"stash.useBuiltin support has been removed\" &&\n>> >>\n>> >> We really prefer our tests to create the same data each time if\n>> >> possible, but as noted with the \"date -u\" comment above you're\n>> >> explicitly bypassing that, but I still can't see why...\n>> >\n>> > I'm trying to make sure we get new object contents. Is there a better\n>> > way to achieve what I want without the risk of finding that the contents\n>> > are already in the database from a previous test run?\n>>\n>> You can just do something like:\n>>\n>> test_expect_success 'setup data' '\n>>         test_commit A &&\n>>         test_commit B\n>> '\n>>\n>> Which will create files A.t, B.t etc, or create them via:\n>>\n>>     obj=$(echo foo | git hash-object -w --stdin)\n>>\n>> etc.\n>>\n>> I.e. the uniqueness you're doing here seems to assume that tests are\n>> re-using the same object store across runs, but we create a new trash\n>> directory for each one, if you run the test with \"-d\" you can see it\n>> being left behind for inspection. This is already ensured for the test.\n>>\n>> The only potential caveat I can imagine is that some filesystem like say\n>> btrfs-like that does some COW or object de-duplication would behave\n>> differently, but other than that...\n>\n> It looks like the same repo is reused for each test_expect_success\n> line in the top-level t*.sh script.\n> So for test_create_unique_files to be maximally useful, it should have\n> some state that is different for\n> each invocation.  How about I use the test_tick mechanism to produce\n> this uniqueness?  It wouldn't\n> be globally unique like the date method, but it should be good enough\n> if the repo is recycled every time\n> test-lib is reinitialized.\n>\n> I'm changing lib-unique-files to use test_tick and to be a little more\n> readable as you suggested. Please\n> let me know if you have any other suggestions.\n\nAh, sorry, I thought you meant you wanted uniqueness within the test\nfile, but no, by default we'll create *one* repo for you, and each\ntest_expect_success reuses that.\n\nGenerally tests that want that do one of (in each test_expect_success):\n\n# I'm making my own repo\ngit init new-repo 1 &&\n(\n\tcd new-repo-1 &&\n\t[...]\n)\n\n# Or, in the first one\n<setup the repo data>\n# Then, in a second one\ngit clone . new-repo-1\n\nI.e. just using \"git clone\" to ferry the data around, or cp -R if you'd\nlike to retain the exact file layout etc.\n\n\n\n"},{"id":"436906","messageId":"CANQDOdckziehbqqvAPhy9aJ79TKLaSifgT7ZB553FxzABn_jfA@mail.gmail.com","threadId":"56371","inReplyTo":"CANQDOdeV4JuE8jnkzLmK6VfFj1t-+EOzvn=GD-ejPdS6unc66w@mail.gmail.com","subject":"Re: [PATCH v4 4/6] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-23T22:32:11Z","receivedAt":"2021-09-23T22:32:26Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 21, 2021 at 6:27 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> On Tue, Sep 21, 2021 at 4:53 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n> >\n> > All of this makes me wonder why this isn't using tmp-objdir.c, i.e. we\n> > could have our cake and eat it too by writing the \"real\" objects, and\n> > then just renaming them between directories instead. But perhaps the\n> > answer has something to do with the metadata issues I raised.\n> >\n> > And well, tmp-objdir.c isn't going to help someone in practice that's\n> > relying on this \"update-index --stdin\" behavior, as they won't know\n> > where we staged the temporary files...\n> >\n>\n> One motivation of the current design behind renaming the files is that\n> some networked filesystems don't seem to like cross-directory renames\n> much.  It also so happens that ReFS on Windows also prefers renames to\n> stay within the directory. Actually any filesystem would likely be\n> slightly faster,\n> since fewer objects are being modified (one dir versus two).\n\nWhelp, as part of v5 I tried to make unpack-objects.c use the batch fsync\nmode and now I see a strong reason to take your tmp-objdir suggestion. As\npart of OBJ_REF_DELTA unpacking, we need access to the object while\nwe're in the plugged state. I didn't notice this at first, but got\nlucky that I tested\nthat case first and hit an error.\n\nV5 will create a tmp-objdir and add a new interface to install it as the primary\nobjdir.\n"},{"id":"437030","messageId":"95315f35a283feabe301b24d2d465a8ae141b139.1632514331.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v5 1/7] object-file.c: do not rename in a temp odb","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:05Z","receivedAt":"2021-09-24T20:12:18Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIf a temporary ODB is active, as determined by GIT_QUARANTINE_PATH\nbeing set, create object files with their final names. This avoids\nan extra rename beyond what is needed to merge the temporary ODB in\ntmp_objdir_migrate.\n\nCreating an object file with the expected final name should be okay\nsince the git process writing to the temporary object store is the\nonly writer, and it only invokes write_loose_object/create_object_file\nafter checking that the object doesn't exist.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n environment.c  |  4 ++++\n object-file.c  | 51 ++++++++++++++++++++++++++++++++++----------------\n object-store.h |  6 ++++++\n repository.c   |  2 ++\n repository.h   |  1 +\n 5 files changed, 48 insertions(+), 16 deletions(-)\n\ndiff --git a/environment.c b/environment.c\nindex d6b22ede7ea..d9ba68402e9 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -177,6 +177,10 @@ void setup_git_env(const char *git_dir)\n \targs.graft_file = getenv_safe(&to_free, GRAFT_ENVIRONMENT);\n \targs.index_file = getenv_safe(&to_free, INDEX_ENVIRONMENT);\n \targs.alternate_db = getenv_safe(&to_free, ALTERNATE_DB_ENVIRONMENT);\n+\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n+\t\targs.object_dir_is_temp = 1;\n+\t}\n+\n \trepo_set_gitdir(the_repository, git_dir, &args);\n \tstrvec_clear(&to_free);\n \ndiff --git a/object-file.c b/object-file.c\nindex a8be8994814..ab593515cec 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1800,12 +1800,17 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n }\n \n /*\n- * Move the just written object into its final resting place.\n+ * Move the just written object into its final resting place,\n+ * unless it is already there, as indicated by an empty string for\n+ * tmpfile.\n  */\n int finalize_object_file(const char *tmpfile, const char *filename)\n {\n \tint ret = 0;\n \n+\tif (!*tmpfile)\n+\t\tgoto out;\n+\n \tif (object_creation_mode == OBJECT_CREATION_USES_RENAMES)\n \t\tgoto try_rename;\n \telse if (link(tmpfile, filename))\n@@ -1878,21 +1883,37 @@ static inline int directory_size(const char *filename)\n }\n \n /*\n- * This creates a temporary file in the same directory as the final\n- * 'filename'\n+ * This creates a loose object file for the specified object id.\n+ * If we're working in a temporary object directory, the file is\n+ * created with its final filename, otherwise it is created with\n+ * a temporary name and renamed by finalize_object_file.\n+ * If no rename is required, an empty string is returned in tmp.\n  *\n  * We want to avoid cross-directory filename renames, because those\n  * can have problems on various filesystems (FAT, NFS, Coda).\n  */\n-static int create_tmpfile(struct strbuf *tmp, const char *filename)\n+static int create_objfile(const struct object_id *oid, struct strbuf *tmp,\n+\t\t\t  struct strbuf *filename)\n {\n-\tint fd, dirlen = directory_size(filename);\n+\tint fd, dirlen, is_retrying = 0;\n+\tconst char *object_name;\n+\tstatic const int object_mode = 0444;\n \n+\tloose_object_path(the_repository, filename, oid);\n+\tdirlen = directory_size(filename->buf);\n+\n+retry_create:\n \tstrbuf_reset(tmp);\n-\tstrbuf_add(tmp, filename, dirlen);\n-\tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n-\tfd = git_mkstemp_mode(tmp->buf, 0444);\n-\tif (fd < 0 && dirlen && errno == ENOENT) {\n+\tif (!the_repository->objects->odb->is_temp) {\n+\t\tstrbuf_add(tmp, filename->buf, dirlen);\n+\t\tobject_name = \"tmp_obj_XXXXXX\";\n+\t\tstrbuf_addstr(tmp, object_name);\n+\t\tfd = git_mkstemp_mode(tmp->buf, object_mode);\n+\t} else {\n+\t\tfd = open(filename->buf, O_CREAT | O_EXCL | O_RDWR, object_mode);\n+\t}\n+\n+\tif (fd < 0 && dirlen && errno == ENOENT && !is_retrying) {\n \t\t/*\n \t\t * Make sure the directory exists; note that the contents\n \t\t * of the buffer are undefined after mkstemp returns an\n@@ -1900,15 +1921,15 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \t\t * scratch.\n \t\t */\n \t\tstrbuf_reset(tmp);\n-\t\tstrbuf_add(tmp, filename, dirlen - 1);\n+\t\tstrbuf_add(tmp, filename->buf, dirlen - 1);\n \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n \t\t\treturn -1;\n \t\tif (adjust_shared_perm(tmp->buf))\n \t\t\treturn -1;\n \n \t\t/* Try again */\n-\t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n-\t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n+\t\tis_retrying = 1;\n+\t\tgoto retry_create;\n \t}\n \treturn fd;\n }\n@@ -1925,14 +1946,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n-\tloose_object_path(the_repository, &filename, oid);\n-\n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n+\tfd = create_objfile(oid, &tmp_file, &filename);\n \tif (fd < 0) {\n \t\tif (errno == EACCES)\n \t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n \t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\t\t\treturn error_errno(_(\"unable to create object file\"));\n \t}\n \n \t/* Set it up */\ndiff --git a/object-store.h b/object-store.h\nindex b4dc6668aa2..f8c883a5730 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -26,6 +26,12 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/*\n+\t * This is a temporary object store, so there is no need to\n+\t * create new objects via rename.\n+\t */\n+\tint is_temp;\n+\n \t/*\n \t * Path to the alternative object store. If this is a relative path,\n \t * it is relative to the current working directory.\ndiff --git a/repository.c b/repository.c\nindex b2bf44c6faf..a16de04dfa8 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -80,6 +80,8 @@ void repo_set_gitdir(struct repository *repo,\n \texpand_base_dir(&repo->objects->odb->path, o->object_dir,\n \t\t\trepo->commondir, \"objects\");\n \n+\trepo->objects->odb->is_temp = o->object_dir_is_temp;\n+\n \tfree(repo->objects->alternate_db);\n \trepo->objects->alternate_db = xstrdup_or_null(o->alternate_db);\n \texpand_base_dir(&repo->graft_file, o->graft_file,\ndiff --git a/repository.h b/repository.h\nindex 3740c93bc0f..d3711367a6f 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -162,6 +162,7 @@ struct set_gitdir_args {\n \tconst char *graft_file;\n \tconst char *index_file;\n \tconst char *alternate_db;\n+\tint object_dir_is_temp;\n };\n \n void repo_set_gitdir(struct repository *repo, const char *root,\n-- \ngitgitgadget\n\n"},{"id":"437032","messageId":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v4.git.git.1632176111.gitgitgadget@gmail.com","subject":"[PATCH v5 0/7] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:04Z","receivedAt":"2021-09-24T20:12:20Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far! Changes since v4, all in response to\nreview feedback from Ævar Arnfjörð Bjarmason:\n\n * Update core.fsyncobjectfiles documentation to specify 'loose' objects and\n   to add a statement about not fsyncing parent directories.\n   \n   * I still don't want to make any promises on behalf of the Linux FS developers\n     in the documentation. However, according to [v4.1] and my understanding\n     of how XFS journals are documented to work, it looks like recent versions\n     of Linux running on XFS should be as safe as Windows or macOS in 'batch'\n     mode. I don't know about ext4, since it's not clear to me when metadata\n     updates are made visible to the journal.\n   \n\n * Rewrite the core batched fsync change to use the tmp-objdir lib. As Ævar\n   pointed out, this lets us access the added loose objects immediately,\n   rather than only after unplugging the bulk checkin. This is a hard\n   requirement in unpack-objects for resolving OBJ_REF_DELTA packed objects.\n   \n   * As a preparatory patch, the object-file code now doesn't do a rename if it's in a\n     tmp objdir (as determined by the quarantine environment variable).\n   \n   * I added support to the tmp-objdir lib to replace the 'main' writable odb.\n   \n   * Instead of using a lockfile for the final full fsync, we now use a new dummy\n     temp file. Doing that makes the below unpack-objects change easier.\n   \n\n * Add bulk-checkin support to unpack-objects, which is used in fetch and\n   push. In addition to making those operations faster, it allows us to\n   directly compare performance of packfiles against loose objects. Please\n   see [v4.2] for a measurement of 'git push' to a local upstream with\n   different numbers of unique new files.\n\n * Rename FSYNC_OBJECT_FILES_MODE to fsync_object_files_mode.\n\n * Remove comment with link to NtFlushBuffersFileEx documentation.\n\n * Make t/lib-unique-files.sh a bit cleaner. We are still creating unique\n   contents, but now this uses test_tick, so it should be deterministic from\n   run to run.\n\n * Ensure there are tests for all of the modified commands. Make the\n   unpack-objects tests validate that the unpacked objects are really\n   available in the ODB.\n\nReferences for v4: [v4.1]\nhttps://lore.kernel.org/linux-fsdevel/20190419072938.31320-1-amir73il@gmail.com/#t\n\n[v4.2]\nhttps://docs.google.com/spreadsheets/d/1uxMBkEXFFnQ1Y3lXKqcKpw6Mq44BzhpCAcPex14T-QQ/edit#gid=1898936117\n\nChanges since v3:\n\n * Fix core.fsyncobjectfiles option parsing as suggested by Junio: We now\n   accept no value to mean \"true\" and we require 'batch' to be lowercase.\n\n * Leave the default fsync mode as 'false'. Git for windows can change its\n   default when this series makes it over to that fork.\n\n * Use a switch statement in git_fsync, as suggested by Junio.\n\n * Add regression test cases for core.fsyncobjectfiles=batch. This should\n   keep the batch functionality basically working in upstream git even if\n   few users adopt batch mode initially. I expect git-for-windows will\n   provide a good baking area for the new mode.\n\nNeeraj Singh (7):\n  object-file.c: do not rename in a temp odb\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: tests for batch mode\n  core.fsyncobjectfiles: performance tests for add and stash\n\n Documentation/config/core.txt |  29 +++++++--\n Makefile                      |   6 ++\n builtin/add.c                 |   1 +\n builtin/unpack-objects.c      |   3 +\n builtin/update-index.c        |   6 ++\n bulk-checkin.c                |  92 +++++++++++++++++++++++---\n bulk-checkin.h                |   2 +\n cache.h                       |   8 ++-\n config.c                      |   7 +-\n config.mak.uname              |   1 +\n configure.ac                  |   8 +++\n environment.c                 |   6 +-\n git-compat-util.h             |   7 ++\n object-file.c                 | 118 +++++++++++++++++++++++++++++-----\n object-store.h                |  22 +++++++\n object.c                      |   2 +-\n repository.c                  |   2 +\n repository.h                  |   1 +\n t/lib-unique-files.sh         |  36 +++++++++++\n t/perf/p3700-add.sh           |  43 +++++++++++++\n t/perf/p3900-stash.sh         |  46 +++++++++++++\n t/t3700-add.sh                |  20 ++++++\n t/t3903-stash.sh              |  14 ++++\n t/t5300-pack-object.sh        |  30 +++++----\n tmp-objdir.c                  |  20 +++++-\n tmp-objdir.h                  |   6 ++\n wrapper.c                     |  44 +++++++++++++\n write-or-die.c                |   2 +-\n 28 files changed, 532 insertions(+), 50 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: 8b7c11b8668b4e774f81a9f0b4c30144b818f1d1\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v5\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v4:\n\n -:  ----------- > 1:  95315f35a28 object-file.c: do not rename in a temp odb\n 1:  d5893e28df1 = 2:  df6fab94d67 bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n 2:  12cad737635 ! 3:  fe19cdfc930 core.fsyncobjectfiles: batched disk flushes\n     @@ Commit message\n      \n          One major source of the cost of fsync is the implied flush of the\n          hardware writeback cache within the disk drive. Fortunately, Windows,\n     -    macOS, and Linux each offer mechanisms to write data from the filesystem\n     -    page cache without initiating a hardware flush.\n     +    and macOS offer mechanisms to write data from the filesystem page cache\n     +    without initiating a hardware flush. Linux has the sync_file_range API,\n     +    which issues a pagecache writeback request reliably after version 5.2.\n      \n          This patch introduces a new 'core.fsyncObjectFiles = batch' option that\n     -    takes advantage of the bulk-checkin infrastructure to batch up hardware\n     -    flushes.\n     +    batches up hardware flushes. It hooks into the bulk-checkin plugging and\n     +    unplugging functionality and takes advantage of tmp-objdir.\n      \n     -    When the new mode is enabled we do the following for new objects:\n     -\n     -    1. Create a tmp_obj_XXXX file and write the object data to it.\n     +    When the new mode is enabled we do the following for each new object:\n     +    1. Create the object in a tmp-objdir.\n          2. Issue a pagecache writeback request and wait for it to complete.\n     -    3. Record the tmp name and the final name in the bulk-checkin state for\n     -       later rename.\n      \n     -    At the end of the entire transaction we:\n     -    1. Issue a fsync against the lock file to flush the hardware writeback\n     -       cache, which should by now have processed the tmp file writes.\n     -    2. Rename all of the temp files to their final names.\n     +    At the end of the entire transaction when unplugging bulk checkin we:\n     +    1. Issue an fsync against a dummy file to flush the hardware writeback\n     +       cache, which should by now have processed the tmp-objdir writes.\n     +    2. Rename all of the tmp-objdir files to their final names.\n          3. When updating the index and/or refs, we assume that Git will issue\n     -       another fsync internal to that operation.\n     +       another fsync internal to that operation. This is not the case today,\n     +       but may be a good extension to those components.\n      \n          On a filesystem with a singular journal that is updated during name\n     -    operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n     +    operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS we\n          would expect the fsync to trigger a journal writeout so that this\n          sequence is enough to ensure that the user's data is durable by the time\n          the git command returns.\n     @@ Documentation/config/core.txt: core.whitespace::\n      +\tA value indicating the level of effort Git will expend in\n      +\ttrying to make objects added to the repo durable in the event\n      +\tof an unclean system shutdown. This setting currently only\n     -+\tcontrols the object store, so updates to any refs or the\n     -+\tindex may not be equally durable.\n     ++\tcontrols loose objects in the object store, so updates to any\n     ++\trefs or the index may not be equally durable.\n      ++\n      +* `false` allows data to remain in file system caches according to\n      +  operating system policy, whence it may be lost if the system loses power\n      +  or crashes.\n     -+* `true` triggers a data integrity flush for each object added to the\n     ++* `true` triggers a data integrity flush for each loose object added to the\n      +  object store. This is the safest setting that is likely to ensure durability\n      +  across all operating systems and file systems that honor the 'fsync' system\n      +  call. However, this setting comes with a significant performance cost on\n     -+  common hardware.\n     ++  common hardware. Git does not currently fsync parent directories for\n     ++  newly-added files, so some filesystems may still allow data to be lost on\n     ++  system crash.\n      +* `batch` enables an experimental mode that uses interfaces available in some\n     -+  operating systems to write object data with a minimal set of FLUSH CACHE\n     -+  (or equivalent) commands sent to the storage controller. If the operating\n     -+  system interfaces are not available, this mode behaves the same as `true`.\n     -+  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n     -+  filesystems and on Windows for repos stored on NTFS or ReFS.\n     ++  operating systems to write loose object data with a minimal set of FLUSH\n     ++  CACHE (or equivalent) commands sent to the storage controller. If the\n     ++  operating system interfaces are not available, this mode behaves the same as\n     ++  `true`. This mode is expected to be as safe as `true` on macOS for repos\n     ++  stored on HFS+ or APFS filesystems and on Windows for repos stored on NTFS or\n     ++  ReFS.\n       \n       core.preloadIndex::\n       \tEnable parallel index preload for operations like 'git diff'\n     @@ builtin/add.c: int cmd_add(int argc, const char **argv, const char *prefix)\n       \n       \tif (chmod_arg && pathspec.nr)\n       \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n     --\tunplug_bulk_checkin();\n      +\n     -+\tunplug_bulk_checkin(&lock_file);\n     + \tunplug_bulk_checkin();\n       \n       finish:\n     - \tif (write_locked_index(&the_index, &lock_file,\n      \n       ## bulk-checkin.c ##\n      @@\n     @@ bulk-checkin.c\n       #include \"pack.h\"\n       #include \"strbuf.h\"\n      +#include \"string-list.h\"\n     ++#include \"tmp-objdir.h\"\n       #include \"packfile.h\"\n       #include \"object-store.h\"\n       \n       static int bulk_checkin_plugged;\n     - \n     -+static struct string_list bulk_fsync_state = STRING_LIST_INIT_DUP;\n     ++static int needs_batch_fsync;\n      +\n     ++static struct tmp_objdir *bulk_fsync_objdir;\n     + \n       static struct bulk_checkin_state {\n       \tchar *pack_tmp_name;\n     - \tstruct hashfile *f;\n      @@ bulk-checkin.c: clear_exit:\n       \treprepare_packed_git(the_repository);\n       }\n       \n     -+static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n     ++/*\n     ++ * Cleanup after batch-mode fsync_object_files.\n     ++ */\n     ++static void do_batch_fsync(void)\n      +{\n     -+\tif (fsync_state->nr) {\n     -+\t\tstruct string_list_item *rename;\n     -+\n     -+\t\t/*\n     -+\t\t * Issue a full hardware flush against the lock file to ensure\n     -+\t\t * that all objects are durable before any renames occur.\n     -+\t\t * The code in fsync_and_close_loose_object_bulk_checkin has\n     -+\t\t * already ensured that writeout has occurred, but it has not\n     -+\t\t * flushed any writeback cache in the storage hardware.\n     -+\t\t */\n     -+\t\tfsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n     -+\n     -+\t\tfor_each_string_list_item(rename, fsync_state) {\n     -+\t\t\tconst char *src = rename->string;\n     -+\t\t\tconst char *dst = rename->util;\n     -+\n     -+\t\t\tif (finalize_object_file(src, dst))\n     -+\t\t\t\tdie_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n     -+\t\t}\n     -+\n     -+\t\tstring_list_clear(fsync_state, 1);\n     ++\t/*\n     ++\t * Issue a full hardware flush against a temporary file to ensure\n     ++\t * that all objects are durable before any renames occur.  The code in\n     ++\t * fsync_loose_object_bulk_checkin has already issued a writeout\n     ++\t * request, but it has not flushed any writeback cache in the storage\n     ++\t * hardware.\n     ++\t */\n     ++\n     ++\tif (needs_batch_fsync) {\n     ++\t\tstruct strbuf temp_path = STRBUF_INIT;\n     ++\t\tstruct tempfile *temp;\n     ++\n     ++\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n     ++\t\ttemp = xmks_tempfile(temp_path.buf);\n     ++\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n     ++\t\tdelete_tempfile(&temp);\n     ++\t\tstrbuf_release(&temp_path);\n      +\t}\n     ++\n     ++\tif (bulk_fsync_objdir)\n     ++\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n      +}\n      +\n       static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n       \treturn 0;\n       }\n       \n     -+static void add_rename_bulk_checkin(struct string_list *fsync_state,\n     -+\t\t\t\t    const char *src, const char *dst)\n     ++void fsync_loose_object_bulk_checkin(int fd)\n      +{\n     -+\tstring_list_insert(fsync_state, src)->util = xstrdup(dst);\n     -+}\n     -+\n     -+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n     -+\t\t\t\t\t      const char *filename, time_t mtime)\n     -+{\n     -+\tint do_finalize = 1;\n     -+\tint ret = 0;\n     -+\n     -+\tif (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n     -+\t\t/*\n     -+\t\t * If we have a plugged bulk checkin, we issue a call that\n     -+\t\t * cleans the filesystem page cache but avoids a hardware flush\n     -+\t\t * command. Later on we will issue a single hardware flush\n     -+\t\t * before renaming files as part of do_sync_and_rename.\n     -+\t\t */\n     -+\t\tif (bulk_checkin_plugged &&\n     -+\t\t    fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n     -+\t\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n     -+\t\t\tadd_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n     -+\t\t\tdo_finalize = 0;\n     -+\n     -+\t\t} else {\n     -+\t\t\tfsync_or_die(fd, \"loose object file\");\n     -+\t\t}\n     -+\t}\n     -+\n     -+\tif (close(fd))\n     -+\t\tdie_errno(_(\"error when closing loose object file\"));\n     -+\n     -+\tif (mtime) {\n     -+\t\tstruct utimbuf utb;\n     -+\t\tutb.actime = mtime;\n     -+\t\tutb.modtime = mtime;\n     -+\t\tif (utime(tmpfile, &utb) < 0)\n     -+\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n     ++\tassert(fsync_object_files == FSYNC_OBJECT_FILES_BATCH);\n     ++\n     ++\t/*\n     ++\t * If we have a plugged bulk checkin, we issue a call that\n     ++\t * cleans the filesystem page cache but avoids a hardware flush\n     ++\t * command. Later on we will issue a single hardware flush\n     ++\t * before as part of do_batch_fsync.\n     ++\t */\n     ++\tif (bulk_checkin_plugged &&\n     ++\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n     ++\t\tassert(the_repository->objects->odb->is_temp);\n     ++\t\tif (!needs_batch_fsync)\n     ++\t\t\tneeds_batch_fsync = 1;\n     ++\t} else {\n     ++\t\tfsync_or_die(fd, \"loose object file\");\n      +\t}\n     -+\n     -+\tif (do_finalize)\n     -+\t\tret = finalize_object_file(tmpfile, filename);\n     -+\n     -+\treturn ret;\n      +}\n      +\n       int index_bulk_checkin(struct object_id *oid,\n       \t\t       int fd, size_t size, enum object_type type,\n       \t\t       const char *path, unsigned flags)\n     -@@ bulk-checkin.c: void plug_bulk_checkin(void)\n     +@@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n     + void plug_bulk_checkin(void)\n     + {\n     + \tassert(!bulk_checkin_plugged);\n     ++\n     ++\t/*\n     ++\t * Create a temporary object directory if the current\n     ++\t * object directory is not already temporary.\n     ++\t */\n     ++\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n     ++\t    !the_repository->objects->odb->is_temp) {\n     ++\t\tbulk_fsync_objdir = tmp_objdir_create();\n     ++\t\tif (!bulk_fsync_objdir)\n     ++\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n     ++\n     ++\t\ttmp_objdir_replace_main_odb(bulk_fsync_objdir);\n     ++\t}\n     ++\n       \tbulk_checkin_plugged = 1;\n       }\n       \n     --void unplug_bulk_checkin(void)\n     -+void unplug_bulk_checkin(struct lock_file *lock_file)\n     - {\n     - \tassert(bulk_checkin_plugged);\n     +@@ bulk-checkin.c: void unplug_bulk_checkin(void)\n       \tbulk_checkin_plugged = 0;\n       \tif (bulk_checkin_state.f)\n       \t\tfinish_bulk_checkin(&bulk_checkin_state);\n      +\n     -+\tdo_sync_and_rename(&bulk_fsync_state, lock_file);\n     ++\tdo_batch_fsync();\n       }\n      \n       ## bulk-checkin.h ##\n     @@ bulk-checkin.h\n       \n       #include \"cache.h\"\n       \n     -+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n     -+\t\t\t\t\t      const char *filename, time_t mtime);\n     ++void fsync_loose_object_bulk_checkin(int fd);\n      +\n       int index_bulk_checkin(struct object_id *oid,\n       \t\t       int fd, size_t size, enum object_type type,\n       \t\t       const char *path, unsigned flags);\n     - \n     - void plug_bulk_checkin(void);\n     --void unplug_bulk_checkin(void);\n     -+void unplug_bulk_checkin(struct lock_file *);\n     - \n     - #endif\n      \n       ## cache.h ##\n      @@ cache.h: void reset_shared_repository(void);\n     @@ cache.h: void reset_shared_repository(void);\n       extern char *git_replace_ref_base;\n       \n      -extern int fsync_object_files;\n     -+enum FSYNC_OBJECT_FILES_MODE {\n     ++enum fsync_object_files_mode {\n      +    FSYNC_OBJECT_FILES_OFF,\n      +    FSYNC_OBJECT_FILES_ON,\n      +    FSYNC_OBJECT_FILES_BATCH\n      +};\n      +\n     -+extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n     ++extern enum fsync_object_files_mode fsync_object_files;\n       extern int core_preload_index;\n       extern int precomposed_unicode;\n       extern int protect_hfs;\n     @@ environment.c: const char *git_hooks_path;\n       int core_compression_level;\n       int pack_compression_level = Z_DEFAULT_COMPRESSION;\n      -int fsync_object_files;\n     -+enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n     ++enum fsync_object_files_mode fsync_object_files;\n       size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n       size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n       size_t delta_base_cache_limit = 96 * 1024 * 1024;\n     @@ git-compat-util.h: __attribute__((format (printf, 1, 2))) NORETURN\n        * Returns 0 on success, which includes trying to unlink an object that does\n      \n       ## object-file.c ##\n     -@@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n     - \treturn 0;\n     +@@ object-file.c: void add_to_alternates_memory(const char *reference)\n     + \t\t\t     '\\n', NULL, 0);\n       }\n       \n     --/* Finalize a file on disk, and close it. */\n     --static void close_loose_object(int fd)\n     --{\n     ++struct object_directory *set_temporary_main_odb(const char *dir)\n     ++{\n     ++\tstruct object_directory *main_odb, *new_odb, *old_next;\n     ++\n     ++\t/*\n     ++\t * Make sure alternates are initialized, or else our entry may be\n     ++\t * overwritten when they are.\n     ++\t */\n     ++\tprepare_alt_odb(the_repository);\n     ++\n     ++\t/* Copy the existing object directory and make it an alternate. */\n     ++\tmain_odb = the_repository->objects->odb;\n     ++\tnew_odb = xmalloc(sizeof(*new_odb));\n     ++\t*new_odb = *main_odb;\n     ++\t*the_repository->objects->odb_tail = new_odb;\n     ++\tthe_repository->objects->odb_tail = &(new_odb->next);\n     ++\tnew_odb->next = NULL;\n     ++\n     ++\t/*\n     ++\t * Reinitialize the main odb with the specified path, being careful\n     ++\t * to keep the next pointer value.\n     ++\t */\n     ++\told_next = main_odb->next;\n     ++\tmemset(main_odb, 0, sizeof(*main_odb));\n     ++\tmain_odb->next = old_next;\n     ++\tmain_odb->is_temp = 1;\n     ++\tmain_odb->path = xstrdup(dir);\n     ++\treturn new_odb;\n     ++}\n     ++\n     ++void restore_main_odb(struct object_directory *odb)\n     ++{\n     ++\tstruct object_directory **prev, *main_odb;\n     ++\n     ++\t/* Unlink the saved previous main ODB from the list. */\n     ++\tprev = &the_repository->objects->odb->next;\n     ++\tassert(*prev);\n     ++\twhile (*prev != odb) {\n     ++\t\tprev = &(*prev)->next;\n     ++\t}\n     ++\t*prev = odb->next;\n     ++\tif (*prev == NULL)\n     ++\t\tthe_repository->objects->odb_tail = prev;\n     ++\n     ++\t/*\n     ++\t * Restore the data from the old main odb, being careful to\n     ++\t * keep the next pointer value\n     ++\t */\n     ++\tmain_odb = the_repository->objects->odb;\n     ++\tSWAP(*main_odb, *odb);\n     ++\tmain_odb->next = odb->next;\n     ++\tfree_object_directory(odb);\n     ++}\n     ++\n     + /*\n     +  * Compute the exact path an alternate is at and returns it. In case of\n     +  * error NULL is returned and the human readable error is added to `err`\n     +@@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n     + /* Finalize a file on disk, and close it. */\n     + static void close_loose_object(int fd)\n     + {\n      -\tif (fsync_object_files)\n     --\t\tfsync_or_die(fd, \"loose object file\");\n     --\tif (close(fd) != 0)\n     --\t\tdie_errno(_(\"error when closing loose object file\"));\n     --}\n     --\n     - /* Size of directory component, including the ending '/' */\n     - static inline int directory_size(const char *filename)\n     ++\tswitch (fsync_object_files) {\n     ++\tcase FSYNC_OBJECT_FILES_OFF:\n     ++\t\tbreak;\n     ++\tcase FSYNC_OBJECT_FILES_ON:\n     + \t\tfsync_or_die(fd, \"loose object file\");\n     ++\t\tbreak;\n     ++\tcase FSYNC_OBJECT_FILES_BATCH:\n     ++\t\tfsync_loose_object_bulk_checkin(fd);\n     ++\t\tbreak;\n     ++\tdefault:\n     ++\t\tBUG(\"Invalid fsync_object_files mode.\");\n     ++\t}\n     ++\n     + \tif (close(fd) != 0)\n     + \t\tdie_errno(_(\"error when closing loose object file\"));\n     + }\n     +\n     + ## object-store.h ##\n     +@@ object-store.h: void add_to_alternates_file(const char *dir);\n     +  */\n     + void add_to_alternates_memory(const char *dir);\n     + \n     ++/*\n     ++ * Replace the current main object directory with the specified temporary\n     ++ * object directory. We make a copy of the former main object directory,\n     ++ * add it as an in-memory alternate, and return the copy so that it can\n     ++ * be restored via restore_main_odb.\n     ++ */\n     ++struct object_directory *set_temporary_main_odb(const char *dir);\n     ++\n     ++/*\n     ++ * Restore a previous ODB replaced by set_temporary_main_odb.\n     ++ */\n     ++void restore_main_odb(struct object_directory *odb);\n     ++\n     + /*\n     +  * Populate and return the loose object cache array corresponding to the\n     +  * given object ID.\n     +@@ object-store.h: struct oidtree *odb_loose_cache(struct object_directory *odb,\n     + /* Empty the loose object cache for the specified object directory. */\n     + void odb_clear_loose_cache(struct object_directory *odb);\n     + \n     ++/* Clear and free the specified object directory */\n     ++void free_object_directory(struct object_directory *odb);\n     ++\n     + struct packed_git {\n     + \tstruct hashmap_entry packmap_ent;\n     + \tstruct packed_git *next;\n     +\n     + ## object.c ##\n     +@@ object.c: struct raw_object_store *raw_object_store_new(void)\n     + \treturn o;\n     + }\n     + \n     +-static void free_object_directory(struct object_directory *odb)\n     ++void free_object_directory(struct object_directory *odb)\n       {\n     -@@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n     - \t\tdie(_(\"confused by unstable object source data for %s\"),\n     - \t\t    oid_to_hex(oid));\n     + \tfree(odb->path);\n     + \todb_clear_loose_cache(odb);\n     +\n     + ## tmp-objdir.c ##\n     +@@\n     + struct tmp_objdir {\n     + \tstruct strbuf path;\n     + \tstruct strvec env;\n     ++\tstruct object_directory *prev_main_odb;\n     + };\n       \n     --\tclose_loose_object(fd);\n     --\n     --\tif (mtime) {\n     --\t\tstruct utimbuf utb;\n     --\t\tutb.actime = mtime;\n     --\t\tutb.modtime = mtime;\n     --\t\tif (utime(tmp_file.buf, &utb) < 0)\n     --\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n     --\t}\n     --\n     --\treturn finalize_object_file(tmp_file.buf, filename.buf);\n     -+\treturn fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf,\n     -+\t\t\t\t\t\t\t filename.buf, mtime);\n     + /*\n     +@@ tmp-objdir.c: static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n     + \t * freeing memory; it may cause a deadlock if the signal\n     + \t * arrived while libc's allocator lock is held.\n     + \t */\n     +-\tif (!on_signal)\n     ++\tif (!on_signal) {\n     ++\t\tif (t->prev_main_odb)\n     ++\t\t\trestore_main_odb(t->prev_main_odb);\n     + \t\ttmp_objdir_free(t);\n     ++\t}\n     ++\n     + \treturn err;\n       }\n       \n     - static int freshen_loose_object(const struct object_id *oid)\n     +@@ tmp-objdir.c: struct tmp_objdir *tmp_objdir_create(void)\n     + \tt = xmalloc(sizeof(*t));\n     + \tstrbuf_init(&t->path, 0);\n     + \tstrvec_init(&t->env);\n     ++\tt->prev_main_odb = NULL;\n     + \n     + \tstrbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n     + \n     +@@ tmp-objdir.c: int tmp_objdir_migrate(struct tmp_objdir *t)\n     + \tif (!t)\n     + \t\treturn 0;\n     + \n     ++\tif (t->prev_main_odb) {\n     ++\t\trestore_main_odb(t->prev_main_odb);\n     ++\t\tt->prev_main_odb = NULL;\n     ++\t}\n     ++\n     + \tstrbuf_addbuf(&src, &t->path);\n     + \tstrbuf_addstr(&dst, get_object_directory());\n     + \n     +@@ tmp-objdir.c: void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n     + {\n     + \tadd_to_alternates_memory(t->path.buf);\n     + }\n     ++\n     ++void tmp_objdir_replace_main_odb(struct tmp_objdir *t)\n     ++{\n     ++\tif (t->prev_main_odb)\n     ++\t\tBUG(\"the main object database is already replaced\");\n     ++\tt->prev_main_odb = set_temporary_main_odb(t->path.buf);\n     ++}\n     +\n     + ## tmp-objdir.h ##\n     +@@ tmp-objdir.h: int tmp_objdir_destroy(struct tmp_objdir *);\n     +  */\n     + void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n     + \n     ++/*\n     ++ * Replaces the main object store in the current process with the temporary\n     ++ * object directory and makes the former main object store an alternate.\n     ++ */\n     ++void tmp_objdir_replace_main_odb(struct tmp_objdir *);\n     ++\n     + #endif /* TMP_OBJDIR_H */\n      \n       ## wrapper.c ##\n      @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n 3:  a5b3e21b762 < -:  ----------- core.fsyncobjectfiles: add windows support for batch mode\n 4:  f7f756f3932 ! 4:  485b4a767df update-index: use the bulk-checkin infrastructure\n     @@ Commit message\n          There is some risk with this change, since under batch fsync, the object\n          files will not be available until the update-index is entirely complete.\n          This usage is unlikely, since any tool invoking update-index and\n     -    expecting to see objects would have to snoop the output of --verbose to\n     -    find out when update-index has actually processed a given path.\n     -    Additionally the index is locked for the duration of the update.\n     +    expecting to see objects would have to synchronize with the update-index\n     +    process after passing it a file path.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ builtin/update-index.c\n       #include \"lockfile.h\"\n       #include \"quote.h\"\n      @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n     - \t\tstruct strbuf unquoted = STRBUF_INIT;\n       \n     - \t\tsetup_work_tree();\n     -+\t\tplug_bulk_checkin();\n     - \t\twhile (getline_fn(&buf, stdin) != EOF) {\n     - \t\t\tchar *p;\n     - \t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n     + \tthe_index.updated_skipworktree = 1;\n     + \n     ++\t/* we might be adding many objects to the object database */\n     ++\tplug_bulk_checkin();\n     ++\n     + \t/*\n     + \t * Custom copy of parse_options() because we want to handle\n     + \t * filename arguments as they come.\n      @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n     - \t\t\t\tchmod_path(set_executable_bit, p);\n     - \t\t\tfree(p);\n     - \t\t}\n     -+\t\tunplug_bulk_checkin(&lock_file);\n     - \t\tstrbuf_release(&unquoted);\n       \t\tstrbuf_release(&buf);\n       \t}\n     + \n     ++\t/* by now we must have added all of the new objects */\n     ++\tunplug_bulk_checkin();\n     + \tif (split_index > 0) {\n     + \t\tif (git_config_get_split_index() == 0)\n     + \t\t\twarning(_(\"core.splitIndex is set to false; \"\n -:  ----------- > 5:  889e7668760 unpack-objects: use the bulk-checkin infrastructure\n 5:  afb0028e796 ! 6:  0f2e3b25759 core.fsyncobjectfiles: tests for batch mode\n     @@ Metadata\n       ## Commit message ##\n          core.fsyncobjectfiles: tests for batch mode\n      \n     -    Add test cases to exercise batch mode for 'git add'\n     -    and 'git stash'. These tests ensure that the added\n     -    data winds up in the object database.\n     +    Add test cases to exercise batch mode for:\n     +     * 'git add'\n     +     * 'git stash'\n     +     * 'git update-index'\n     +     * 'git unpack-objects'\n      \n     -    I verified the tests by introducing an incorrect rename\n     -    in do_sync_and_rename.\n     +    These tests ensure that the added data winds up in the object database.\n     +\n     +    In this change we introduce a new test helper lib-unique-files.sh. The\n     +    goal of this library is to create a tree of files that have different\n     +    oids from any other files that may have been created in the current test\n     +    repo. This helps us avoid missing validation of an object being added due\n     +    to it already being in the repo.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ t/lib-unique-files.sh (new)\n      @@\n      +# Helper to create files with unique contents\n      +\n     -+test_create_unique_files_base__=$(date -u)\n     -+test_create_unique_files_counter__=0\n      +\n      +# Create multiple files with unique contents. Takes the number of\n      +# directories, the number of files in each directory, and the base\n      +# directory.\n      +#\n     -+# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n     -+#\t\t\t\t    each in the specified directory, all\n     -+#\t\t\t\t    with unique contents.\n     ++# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n     ++#\t\t\t\t\t each in my_dir, all with unique\n     ++#\t\t\t\t\t contents.\n      +\n      +test_create_unique_files() {\n      +\ttest \"$#\" -ne 3 && BUG \"3 param\"\n     @@ t/lib-unique-files.sh (new)\n      +\tlocal dirs=$1\n      +\tlocal files=$2\n      +\tlocal basedir=$3\n     ++\tlocal counter=0\n     ++\ttest_tick\n     ++\tlocal basedata=$test_tick\n     ++\n      +\n     -+\trm -rf $basedir >/dev/null\n     ++\trm -rf $basedir\n      +\n      +\tfor i in $(test_seq $dirs)\n      +\tdo\n      +\t\tlocal dir=$basedir/dir$i\n      +\n     -+\t\tmkdir -p \"$dir\" > /dev/null\n     ++\t\tmkdir -p \"$dir\"\n      +\t\tfor j in $(test_seq $files)\n      +\t\tdo\n     -+\t\t\ttest_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n     -+\t\t\techo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n     ++\t\t\tcounter=$((counter + 1))\n     ++\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n      +\t\tdone\n      +\tdone\n      +}\n     @@ t/t3700-add.sh: test_expect_success \\\n      +\trm -f fsynced_files &&\n      +\tgit ls-files --stage fsync-files/ > fsynced_files &&\n      +\ttest_line_count = 8 fsynced_files &&\n     -+\tcat fsynced_files | awk '{print \\$2}' | xargs -n1 git cat-file -e\n     ++\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n     ++\"\n     ++\n     ++test_expect_success 'git update-index: core.fsyncobjectfiles=batch' \"\n     ++\ttest_create_unique_files 2 4 fsync-files2 &&\n     ++\tfind fsync-files2 ! -type d -print | xargs git -c core.fsyncobjectfiles=batch update-index --add -- &&\n     ++\trm -f fsynced_files2 &&\n     ++\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n     ++\ttest_line_count = 8 fsynced_files2 &&\n     ++\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n      +\"\n      +\n       test_expect_success \\\n     @@ t/t3903-stash.sh: test_expect_success 'stash handles skip-worktree entries nicel\n      +\t# which contains the untracked files\n      +\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n      +\ttest_line_count = 8 fsynced_files &&\n     -+\tcat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n     ++\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n      +\"\n      +\n      +\n       test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n       \texpected=\"stash.useBuiltin support has been removed\" &&\n       \n     +\n     + ## t/t5300-pack-object.sh ##\n     +@@ t/t5300-pack-object.sh: test_expect_success 'pack-objects with bogus arguments' '\n     + \n     + check_unpack () {\n     + \ttest_when_finished \"rm -rf git2\" &&\n     +-\tgit init --bare git2 &&\n     +-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n     +-\tgit -C git2 unpack-objects <\"$1\".pack &&\n     +-\t(cd .git && find objects -type f -print) |\n     +-\twhile read path\n     +-\tdo\n     +-\t\tcmp git2/$path .git/$path || {\n     +-\t\t\techo $path differs.\n     +-\t\t\treturn 1\n     +-\t\t}\n     +-\tdone\n     ++\tgit $2 init --bare git2 &&\n     ++\t(\n     ++\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n     ++\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n     ++\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n     ++\t) <obj-list >current &&\n     ++\tcmp obj-list current\n     + }\n     + \n     + test_expect_success 'unpack without delta' '\n     + \tcheck_unpack test-1-${packname_1}\n     + '\n     + \n     ++test_expect_success 'unpack without delta (core.fsyncobjectfiles=batch)' '\n     ++\tcheck_unpack test-1-${packname_1} \"-c core.fsyncobjectfiles=batch\"\n     ++'\n     ++\n     + test_expect_success 'pack with REF_DELTA' '\n     + \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n     + \tcheck_deltas stderr -gt 0\n     +@@ t/t5300-pack-object.sh: test_expect_success 'unpack with REF_DELTA' '\n     + \tcheck_unpack test-2-${packname_2}\n     + '\n     + \n     ++test_expect_success 'unpack with REF_DELTA (core.fsyncobjectfiles=batch)' '\n     ++       check_unpack test-2-${packname_2} \"-c core.fsyncobjectfiles=batch\"\n     ++'\n     ++\n     + test_expect_success 'pack with OFS_DELTA' '\n     + \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n     + \t\t\t<obj-list 2>stderr) &&\n     +@@ t/t5300-pack-object.sh: test_expect_success 'unpack with OFS_DELTA' '\n     + \tcheck_unpack test-3-${packname_3}\n     + '\n     + \n     ++test_expect_success 'unpack with OFS_DELTA (core.fsyncobjectfiles=batch)' '\n     ++       check_unpack test-3-${packname_3} \"-c core.fsyncobjectfiles=batch\"\n     ++'\n     ++\n     + test_expect_success 'compare delta flavors' '\n     + \tperl -e '\\''\n     + \t\tdefined($_ = -s $_) or die for @ARGV;\n 6:  3e6b80b5fa2 = 7:  6543564376a core.fsyncobjectfiles: performance tests for add and stash\n\n-- \ngitgitgadget\n"},{"id":"437031","messageId":"df6fab94d6727424e0e18ed1fba3c5c985032aab.1632514331.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v5 2/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:06Z","receivedAt":"2021-09-24T20:12:21Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b023d9959aa..f117d62c908 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_bulk_checkin(struct bulk_checkin_state *state)\n {\n@@ -260,21 +260,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"437033","messageId":"fe19cdfc9305a54b40b9979842d0d1d7c7dfb828.1632514331.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v5 3/7] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:07Z","receivedAt":"2021-09-24T20:12:28Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nand macOS offer mechanisms to write data from the filesystem page cache\nwithout initiating a hardware flush. Linux has the sync_file_range API,\nwhich issues a pagecache writeback request reliably after version 5.2.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\nbatches up hardware flushes. It hooks into the bulk-checkin plugging and\nunplugging functionality and takes advantage of tmp-objdir.\n\nWhen the new mode is enabled we do the following for each new object:\n1. Create the object in a tmp-objdir.\n2. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin we:\n1. Issue an fsync against a dummy file to flush the hardware writeback\n   cache, which should by now have processed the tmp-objdir writes.\n2. Rename all of the tmp-objdir files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the case today,\n   but may be a good extension to those components.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS we\nwould expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 29 ++++++++++++---\n Makefile                      |  6 +++\n builtin/add.c                 |  1 +\n bulk-checkin.c                | 70 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  2 +\n cache.h                       |  8 +++-\n config.c                      |  7 +++-\n config.mak.uname              |  1 +\n configure.ac                  |  8 ++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 ++++\n object-file.c                 | 67 ++++++++++++++++++++++++++++++++-\n object-store.h                | 16 ++++++++\n object.c                      |  2 +-\n tmp-objdir.c                  | 20 +++++++++-\n tmp-objdir.h                  |  6 +++\n wrapper.c                     | 44 ++++++++++++++++++++++\n write-or-die.c                |  2 +-\n 18 files changed, 285 insertions(+), 13 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..200b4d9f06e 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,29 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols loose objects in the object store, so updates to any\n+\trefs or the index may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each loose object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware. Git does not currently fsync parent directories for\n+  newly-added files, so some filesystems may still allow data to be lost on\n+  system crash.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write loose object data with a minimal set of FLUSH\n+  CACHE (or equivalent) commands sent to the storage controller. If the\n+  operating system interfaces are not available, this mode behaves the same as\n+  `true`. This mode is expected to be as safe as `true` on macOS for repos\n+  stored on HFS+ or APFS filesystems and on Windows for repos stored on NTFS or\n+  ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex 429c276058d..326c7607e0f 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1896,6 +1898,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 2244311d485..9d9897cf037 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -678,6 +678,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n+\n \tunplug_bulk_checkin();\n \n finish:\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f117d62c908..957a6238684 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,14 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n+static int needs_batch_fsync;\n+\n+static struct tmp_objdir *bulk_fsync_objdir;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -62,6 +68,34 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur.  The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware.\n+\t */\n+\n+\tif (needs_batch_fsync) {\n+\t\tstruct strbuf temp_path = STRBUF_INIT;\n+\t\tstruct tempfile *temp;\n+\n+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\t\ttemp = xmks_tempfile(temp_path.buf);\n+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\t\tdelete_tempfile(&temp);\n+\t\tstrbuf_release(&temp_path);\n+\t}\n+\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -256,6 +290,26 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void fsync_loose_object_bulk_checkin(int fd)\n+{\n+\tassert(fsync_object_files == FSYNC_OBJECT_FILES_BATCH);\n+\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (bulk_checkin_plugged &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\tassert(the_repository->objects->odb->is_temp);\n+\t\tif (!needs_batch_fsync)\n+\t\t\tneeds_batch_fsync = 1;\n+\t} else {\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -270,6 +324,20 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n+\n+\t/*\n+\t * Create a temporary object directory if the current\n+\t * object directory is not already temporary.\n+\t */\n+\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n+\t    !the_repository->objects->odb->is_temp) {\n+\t\tbulk_fsync_objdir = tmp_objdir_create();\n+\t\tif (!bulk_fsync_objdir)\n+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n+\n+\t\ttmp_objdir_replace_main_odb(bulk_fsync_objdir);\n+\t}\n+\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -279,4 +347,6 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..08f292379b6 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,8 @@\n \n #include \"cache.h\"\n \n+void fsync_loose_object_bulk_checkin(int fd);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex d23de693680..d1897fe9d92 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum fsync_object_files_mode {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum fsync_object_files_mode fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex cb4a8058bff..1b403e00241 100644\n--- a/config.c\n+++ b/config.c\n@@ -1509,7 +1509,12 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (value && !strcmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse if (git_config_bool(var, value))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n+\t\telse\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 76516aaa9a5..e6d482fbcc6 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex d9ba68402e9..f318d59e585 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -43,7 +43,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int core_compression_level;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_object_files_mode fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b46605300ab..d14e2436276 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex ab593515cec..ec22560dd66 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -750,6 +750,60 @@ void add_to_alternates_memory(const char *reference)\n \t\t\t     '\\n', NULL, 0);\n }\n \n+struct object_directory *set_temporary_main_odb(const char *dir)\n+{\n+\tstruct object_directory *main_odb, *new_odb, *old_next;\n+\n+\t/*\n+\t * Make sure alternates are initialized, or else our entry may be\n+\t * overwritten when they are.\n+\t */\n+\tprepare_alt_odb(the_repository);\n+\n+\t/* Copy the existing object directory and make it an alternate. */\n+\tmain_odb = the_repository->objects->odb;\n+\tnew_odb = xmalloc(sizeof(*new_odb));\n+\t*new_odb = *main_odb;\n+\t*the_repository->objects->odb_tail = new_odb;\n+\tthe_repository->objects->odb_tail = &(new_odb->next);\n+\tnew_odb->next = NULL;\n+\n+\t/*\n+\t * Reinitialize the main odb with the specified path, being careful\n+\t * to keep the next pointer value.\n+\t */\n+\told_next = main_odb->next;\n+\tmemset(main_odb, 0, sizeof(*main_odb));\n+\tmain_odb->next = old_next;\n+\tmain_odb->is_temp = 1;\n+\tmain_odb->path = xstrdup(dir);\n+\treturn new_odb;\n+}\n+\n+void restore_main_odb(struct object_directory *odb)\n+{\n+\tstruct object_directory **prev, *main_odb;\n+\n+\t/* Unlink the saved previous main ODB from the list. */\n+\tprev = &the_repository->objects->odb->next;\n+\tassert(*prev);\n+\twhile (*prev != odb) {\n+\t\tprev = &(*prev)->next;\n+\t}\n+\t*prev = odb->next;\n+\tif (*prev == NULL)\n+\t\tthe_repository->objects->odb_tail = prev;\n+\n+\t/*\n+\t * Restore the data from the old main odb, being careful to\n+\t * keep the next pointer value\n+\t */\n+\tmain_odb = the_repository->objects->odb;\n+\tSWAP(*main_odb, *odb);\n+\tmain_odb->next = odb->next;\n+\tfree_object_directory(odb);\n+}\n+\n /*\n  * Compute the exact path an alternate is at and returns it. In case of\n  * error NULL is returned and the human readable error is added to `err`\n@@ -1867,8 +1921,19 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n+\tswitch (fsync_object_files) {\n+\tcase FSYNC_OBJECT_FILES_OFF:\n+\t\tbreak;\n+\tcase FSYNC_OBJECT_FILES_ON:\n \t\tfsync_or_die(fd, \"loose object file\");\n+\t\tbreak;\n+\tcase FSYNC_OBJECT_FILES_BATCH:\n+\t\tfsync_loose_object_bulk_checkin(fd);\n+\t\tbreak;\n+\tdefault:\n+\t\tBUG(\"Invalid fsync_object_files mode.\");\n+\t}\n+\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/object-store.h b/object-store.h\nindex f8c883a5730..9bea14e7f3b 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -62,6 +62,19 @@ void add_to_alternates_file(const char *dir);\n  */\n void add_to_alternates_memory(const char *dir);\n \n+/*\n+ * Replace the current main object directory with the specified temporary\n+ * object directory. We make a copy of the former main object directory,\n+ * add it as an in-memory alternate, and return the copy so that it can\n+ * be restored via restore_main_odb.\n+ */\n+struct object_directory *set_temporary_main_odb(const char *dir);\n+\n+/*\n+ * Restore a previous ODB replaced by set_temporary_main_odb.\n+ */\n+void restore_main_odb(struct object_directory *odb);\n+\n /*\n  * Populate and return the loose object cache array corresponding to the\n  * given object ID.\n@@ -72,6 +85,9 @@ struct oidtree *odb_loose_cache(struct object_directory *odb,\n /* Empty the loose object cache for the specified object directory. */\n void odb_clear_loose_cache(struct object_directory *odb);\n \n+/* Clear and free the specified object directory */\n+void free_object_directory(struct object_directory *odb);\n+\n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n \tstruct packed_git *next;\ndiff --git a/object.c b/object.c\nindex 4e85955a941..98635bc4043 100644\n--- a/object.c\n+++ b/object.c\n@@ -513,7 +513,7 @@ struct raw_object_store *raw_object_store_new(void)\n \treturn o;\n }\n \n-static void free_object_directory(struct object_directory *odb)\n+void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex b8d880e3626..f027c49db4c 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -11,6 +11,7 @@\n struct tmp_objdir {\n \tstruct strbuf path;\n \tstruct strvec env;\n+\tstruct object_directory *prev_main_odb;\n };\n \n /*\n@@ -50,8 +51,12 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \t * freeing memory; it may cause a deadlock if the signal\n \t * arrived while libc's allocator lock is held.\n \t */\n-\tif (!on_signal)\n+\tif (!on_signal) {\n+\t\tif (t->prev_main_odb)\n+\t\t\trestore_main_odb(t->prev_main_odb);\n \t\ttmp_objdir_free(t);\n+\t}\n+\n \treturn err;\n }\n \n@@ -132,6 +137,7 @@ struct tmp_objdir *tmp_objdir_create(void)\n \tt = xmalloc(sizeof(*t));\n \tstrbuf_init(&t->path, 0);\n \tstrvec_init(&t->env);\n+\tt->prev_main_odb = NULL;\n \n \tstrbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n \n@@ -269,6 +275,11 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n \tif (!t)\n \t\treturn 0;\n \n+\tif (t->prev_main_odb) {\n+\t\trestore_main_odb(t->prev_main_odb);\n+\t\tt->prev_main_odb = NULL;\n+\t}\n+\n \tstrbuf_addbuf(&src, &t->path);\n \tstrbuf_addstr(&dst, get_object_directory());\n \n@@ -292,3 +303,10 @@ void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n {\n \tadd_to_alternates_memory(t->path.buf);\n }\n+\n+void tmp_objdir_replace_main_odb(struct tmp_objdir *t)\n+{\n+\tif (t->prev_main_odb)\n+\t\tBUG(\"the main object database is already replaced\");\n+\tt->prev_main_odb = set_temporary_main_odb(t->path.buf);\n+}\ndiff --git a/tmp-objdir.h b/tmp-objdir.h\nindex b1e45b4c75d..4b898add05b 100644\n--- a/tmp-objdir.h\n+++ b/tmp-objdir.h\n@@ -51,4 +51,10 @@ int tmp_objdir_destroy(struct tmp_objdir *);\n  */\n void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n \n+/*\n+ * Replaces the main object store in the current process with the temporary\n+ * object directory and makes the former main object store an alternate.\n+ */\n+void tmp_objdir_replace_main_odb(struct tmp_objdir *);\n+\n #endif /* TMP_OBJDIR_H */\ndiff --git a/wrapper.c b/wrapper.c\nindex 7c6586af321..bb4f9f043ce 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -540,6 +540,50 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync(fd);\n+#endif\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex d33e68f6abb..8f53953d4ab 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"437034","messageId":"485b4a767dfa54729c40b32b7fea033aedc870d1.1632514331.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v5 4/7] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:08Z","receivedAt":"2021-09-24T20:12:33Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality. This mode\nis enabled when passing paths to update-index via the --stdin flag,\nas is done by 'git stash'.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to synchronize with the update-index\nprocess after passing it a file path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 187203e8bb5..dc7368bb1ee 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1088,6 +1089,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \tthe_index.updated_skipworktree = 1;\n \n+\t/* we might be adding many objects to the object database */\n+\tplug_bulk_checkin();\n+\n \t/*\n \t * Custom copy of parse_options() because we want to handle\n \t * filename arguments as they come.\n@@ -1168,6 +1172,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/* by now we must have added all of the new objects */\n+\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"437035","messageId":"889e76687601e3a1242e57c430a1b7f64ea1d77b.1632514331.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v5 5/7] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:09Z","receivedAt":"2021-09-24T20:12:52Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling bulk-checkin when unpacking objects, we can take advantage\nof batched fsyncs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295ba..51eb4f7b531 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tplug_bulk_checkin();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"437036","messageId":"0f2e3b25759160a31c11836b72b1f3783bf1e372.1632514331.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v5 6/7] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:10Z","receivedAt":"2021-09-24T20:12:53Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added due\nto it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 36 ++++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 20 ++++++++++++++++++++\n t/t3903-stash.sh       | 14 ++++++++++++++\n t/t5300-pack-object.sh | 30 +++++++++++++++++++-----------\n 4 files changed, 89 insertions(+), 11 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a7de4ca8512\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,36 @@\n+# Helper to create files with unique contents\n+\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with unique\n+#\t\t\t\t\t contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\tlocal counter=0\n+\ttest_tick\n+\tlocal basedata=$test_tick\n+\n+\n+\trm -rf $basedir\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\"\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1))\n+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex 4086e1ebbc9..36049a53ff7 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -7,6 +7,8 @@ test_description='Test of git add, including the -- option.'\n \n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -33,6 +35,24 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+test_expect_success 'git add: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files2 &&\n+\tfind fsync-files2 ! -type d -print | xargs git -c core.fsyncobjectfiles=batch update-index --add -- &&\n+\trm -f fsynced_files2 &&\n+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n+\ttest_line_count = 8 fsynced_files2 &&\n+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 873aa56e359..2fc819e5584 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n diff_cmp () {\n \tfor i in \"$1\" \"$2\"\n@@ -1293,6 +1294,19 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n \texpected=\"stash.useBuiltin support has been removed\" &&\n \ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex e13a8842075..38663dc1393 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -162,23 +162,23 @@ test_expect_success 'pack-objects with bogus arguments' '\n \n check_unpack () {\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $2 init --bare git2 &&\n+\t(\n+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <obj-list >current &&\n+\tcmp obj-list current\n }\n \n test_expect_success 'unpack without delta' '\n \tcheck_unpack test-1-${packname_1}\n '\n \n+test_expect_success 'unpack without delta (core.fsyncobjectfiles=batch)' '\n+\tcheck_unpack test-1-${packname_1} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with REF_DELTA' '\n \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n \tcheck_deltas stderr -gt 0\n@@ -188,6 +188,10 @@ test_expect_success 'unpack with REF_DELTA' '\n \tcheck_unpack test-2-${packname_2}\n '\n \n+test_expect_success 'unpack with REF_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-2-${packname_2} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with OFS_DELTA' '\n \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n \t\t\t<obj-list 2>stderr) &&\n@@ -198,6 +202,10 @@ test_expect_success 'unpack with OFS_DELTA' '\n \tcheck_unpack test-3-${packname_3}\n '\n \n+test_expect_success 'unpack with OFS_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-3-${packname_3} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'compare delta flavors' '\n \tperl -e '\\''\n \t\tdefined($_ = -s $_) or die for @ARGV;\n-- \ngitgitgadget\n\n"},{"id":"437037","messageId":"6543564376a7b06809d51dedbbf4571c359ace3b.1632514331.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v5 7/7] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T20:12:11Z","receivedAt":"2021-09-24T20:12:54Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 43 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 46 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 89 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..e93c08a2e70\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..c9fcd0c03eb\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n"},{"id":"437050","messageId":"CANQDOdfCgdBDE3oJ385WFWQ_in7J5xe_UwkQ5BVaxgkCw15xvg@mail.gmail.com","threadId":"56371","inReplyTo":"fe19cdfc9305a54b40b9979842d0d1d7c7dfb828.1632514331.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 3/7] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-24T21:47:19Z","receivedAt":"2021-09-24T21:47:34Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Sep 24, 2021 at 1:12 PM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> diff --git a/builtin/add.c b/builtin/add.c\n> index 2244311d485..9d9897cf037 100644\n> --- a/builtin/add.c\n> +++ b/builtin/add.c\n> @@ -678,6 +678,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n>\n>         if (chmod_arg && pathspec.nr)\n>                 exit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n> +\n>         unplug_bulk_checkin();\n>\n>  finish:\n\nI'll remove this stray change on re-roll.\n"},{"id":"437051","messageId":"CANQDOdfZrn0YK0_HomzqHkqnxmjXc20aa6TPkmwpiapeGpjjyw@mail.gmail.com","threadId":"56371","inReplyTo":"485b4a767dfa54729c40b32b7fea033aedc870d1.1632514331.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 4/7] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-24T21:49:30Z","receivedAt":"2021-09-24T21:49:45Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Sep 24, 2021 at 1:12 PM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> The update-index functionality is used internally by 'git stash push' to\n> setup the internal stashed commit.\n>\n> This change enables bulk-checkin for update-index infrastructure to\n> speed up adding new objects to the object database by leveraging the\n> pack functionality and the new bulk-fsync functionality. This mode\n> is enabled when passing paths to update-index via the --stdin flag,\n> as is done by 'git stash'.\n\nThis part of the description is now inaccurate. All modes of update-index are\nnow enlightened to use bulk_checkin. I'll just remove the sentence that\nscopes the change to --stdin on reroll.\n\n>\n> There is some risk with this change, since under batch fsync, the object\n> files will not be available until the update-index is entirely complete.\n> This usage is unlikely, since any tool invoking update-index and\n> expecting to see objects would have to synchronize with the update-index\n> process after passing it a file path.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  builtin/update-index.c | 6 ++++++\n>  1 file changed, 6 insertions(+)\n"},{"id":"437053","messageId":"CANQDOddqbiiVEt7JEDmF9QsJvPv8W6nZrenpsDPeem0brd-7_A@mail.gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 0/7] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-24T23:31:35Z","receivedAt":"2021-09-24T23:31:50Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Sep 24, 2021 at 1:12 PM Neeraj K. Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> Thanks to everyone for review so far! Changes since v4, all in response to\n> review feedback from Ævar Arnfjörð Bjarmason:\n>\n>  * Update core.fsyncobjectfiles documentation to specify 'loose' objects and\n>    to add a statement about not fsyncing parent directories.\n>\n>    * I still don't want to make any promises on behalf of the Linux FS developers\n>      in the documentation. However, according to [v4.1] and my understanding\n>      of how XFS journals are documented to work, it looks like recent versions\n>      of Linux running on XFS should be as safe as Windows or macOS in 'batch'\n>      mode. I don't know about ext4, since it's not clear to me when metadata\n>      updates are made visible to the journal.\n>\n>\n>  * Rewrite the core batched fsync change to use the tmp-objdir lib. As Ævar\n>    pointed out, this lets us access the added loose objects immediately,\n>    rather than only after unplugging the bulk checkin. This is a hard\n>    requirement in unpack-objects for resolving OBJ_REF_DELTA packed objects.\n>\n>    * As a preparatory patch, the object-file code now doesn't do a rename if it's in a\n>      tmp objdir (as determined by the quarantine environment variable).\n>\n>    * I added support to the tmp-objdir lib to replace the 'main' writable odb.\n>\n>    * Instead of using a lockfile for the final full fsync, we now use a new dummy\n>      temp file. Doing that makes the below unpack-objects change easier.\n>\n>\n>  * Add bulk-checkin support to unpack-objects, which is used in fetch and\n>    push. In addition to making those operations faster, it allows us to\n>    directly compare performance of packfiles against loose objects. Please\n>    see [v4.2] for a measurement of 'git push' to a local upstream with\n>    different numbers of unique new files.\n>\n>  * Rename FSYNC_OBJECT_FILES_MODE to fsync_object_files_mode.\n>\n>  * Remove comment with link to NtFlushBuffersFileEx documentation.\n>\n>  * Make t/lib-unique-files.sh a bit cleaner. We are still creating unique\n>    contents, but now this uses test_tick, so it should be deterministic from\n>    run to run.\n>\n>  * Ensure there are tests for all of the modified commands. Make the\n>    unpack-objects tests validate that the unpacked objects are really\n>    available in the ODB.\n>\n> References for v4: [v4.1]\n> https://lore.kernel.org/linux-fsdevel/20190419072938.31320-1-amir73il@gmail.com/#t\n>\n> [v4.2]\n> https://docs.google.com/spreadsheets/d/1uxMBkEXFFnQ1Y3lXKqcKpw6Mq44BzhpCAcPex14T-QQ/edit#gid=1898936117\n>\n> Changes since v3:\n>\n>  * Fix core.fsyncobjectfiles option parsing as suggested by Junio: We now\n>    accept no value to mean \"true\" and we require 'batch' to be lowercase.\n>\n>  * Leave the default fsync mode as 'false'. Git for windows can change its\n>    default when this series makes it over to that fork.\n>\n>  * Use a switch statement in git_fsync, as suggested by Junio.\n>\n>  * Add regression test cases for core.fsyncobjectfiles=batch. This should\n>    keep the batch functionality basically working in upstream git even if\n>    few users adopt batch mode initially. I expect git-for-windows will\n>    provide a good baking area for the new mode.\n>\n> Neeraj Singh (7):\n>   object-file.c: do not rename in a temp odb\n>   bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n>   core.fsyncobjectfiles: batched disk flushes\n>   update-index: use the bulk-checkin infrastructure\n>   unpack-objects: use the bulk-checkin infrastructure\n>   core.fsyncobjectfiles: tests for batch mode\n>   core.fsyncobjectfiles: performance tests for add and stash\n>\n>  Documentation/config/core.txt |  29 +++++++--\n>  Makefile                      |   6 ++\n>  builtin/add.c                 |   1 +\n>  builtin/unpack-objects.c      |   3 +\n>  builtin/update-index.c        |   6 ++\n>  bulk-checkin.c                |  92 +++++++++++++++++++++++---\n>  bulk-checkin.h                |   2 +\n>  cache.h                       |   8 ++-\n>  config.c                      |   7 +-\n>  config.mak.uname              |   1 +\n>  configure.ac                  |   8 +++\n>  environment.c                 |   6 +-\n>  git-compat-util.h             |   7 ++\n>  object-file.c                 | 118 +++++++++++++++++++++++++++++-----\n>  object-store.h                |  22 +++++++\n>  object.c                      |   2 +-\n>  repository.c                  |   2 +\n>  repository.h                  |   1 +\n>  t/lib-unique-files.sh         |  36 +++++++++++\n>  t/perf/p3700-add.sh           |  43 +++++++++++++\n>  t/perf/p3900-stash.sh         |  46 +++++++++++++\n>  t/t3700-add.sh                |  20 ++++++\n>  t/t3903-stash.sh              |  14 ++++\n>  t/t5300-pack-object.sh        |  30 +++++----\n>  tmp-objdir.c                  |  20 +++++-\n>  tmp-objdir.h                  |   6 ++\n>  wrapper.c                     |  44 +++++++++++++\n>  write-or-die.c                |   2 +-\n>  28 files changed, 532 insertions(+), 50 deletions(-)\n>  create mode 100644 t/lib-unique-files.sh\n>  create mode 100755 t/perf/p3700-add.sh\n>  create mode 100755 t/perf/p3900-stash.sh\n>\n>\n> base-commit: 8b7c11b8668b4e774f81a9f0b4c30144b818f1d1\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v5\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v5\n> Pull-Request: https://github.com/git/git/pull/1076\n>\n> Range-diff vs v4:\n>\n>  -:  ----------- > 1:  95315f35a28 object-file.c: do not rename in a temp odb\n>  1:  d5893e28df1 = 2:  df6fab94d67 bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n>  2:  12cad737635 ! 3:  fe19cdfc930 core.fsyncobjectfiles: batched disk flushes\n>      @@ Commit message\n>\n>           One major source of the cost of fsync is the implied flush of the\n>           hardware writeback cache within the disk drive. Fortunately, Windows,\n>      -    macOS, and Linux each offer mechanisms to write data from the filesystem\n>      -    page cache without initiating a hardware flush.\n>      +    and macOS offer mechanisms to write data from the filesystem page cache\n>      +    without initiating a hardware flush. Linux has the sync_file_range API,\n>      +    which issues a pagecache writeback request reliably after version 5.2.\n>\n>           This patch introduces a new 'core.fsyncObjectFiles = batch' option that\n>      -    takes advantage of the bulk-checkin infrastructure to batch up hardware\n>      -    flushes.\n>      +    batches up hardware flushes. It hooks into the bulk-checkin plugging and\n>      +    unplugging functionality and takes advantage of tmp-objdir.\n>\n>      -    When the new mode is enabled we do the following for new objects:\n>      -\n>      -    1. Create a tmp_obj_XXXX file and write the object data to it.\n>      +    When the new mode is enabled we do the following for each new object:\n>      +    1. Create the object in a tmp-objdir.\n>           2. Issue a pagecache writeback request and wait for it to complete.\n>      -    3. Record the tmp name and the final name in the bulk-checkin state for\n>      -       later rename.\n>\n>      -    At the end of the entire transaction we:\n>      -    1. Issue a fsync against the lock file to flush the hardware writeback\n>      -       cache, which should by now have processed the tmp file writes.\n>      -    2. Rename all of the temp files to their final names.\n>      +    At the end of the entire transaction when unplugging bulk checkin we:\n>      +    1. Issue an fsync against a dummy file to flush the hardware writeback\n>      +       cache, which should by now have processed the tmp-objdir writes.\n>      +    2. Rename all of the tmp-objdir files to their final names.\n>           3. When updating the index and/or refs, we assume that Git will issue\n>      -       another fsync internal to that operation.\n>      +       another fsync internal to that operation. This is not the case today,\n>      +       but may be a good extension to those components.\n>\n>           On a filesystem with a singular journal that is updated during name\n>      -    operations (e.g. create, link, rename, etc), such as NTFS and HFS+, we\n>      +    operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS we\n>           would expect the fsync to trigger a journal writeout so that this\n>           sequence is enough to ensure that the user's data is durable by the time\n>           the git command returns.\n>      @@ Documentation/config/core.txt: core.whitespace::\n>       + A value indicating the level of effort Git will expend in\n>       + trying to make objects added to the repo durable in the event\n>       + of an unclean system shutdown. This setting currently only\n>      -+ controls the object store, so updates to any refs or the\n>      -+ index may not be equally durable.\n>      ++ controls loose objects in the object store, so updates to any\n>      ++ refs or the index may not be equally durable.\n>       ++\n>       +* `false` allows data to remain in file system caches according to\n>       +  operating system policy, whence it may be lost if the system loses power\n>       +  or crashes.\n>      -+* `true` triggers a data integrity flush for each object added to the\n>      ++* `true` triggers a data integrity flush for each loose object added to the\n>       +  object store. This is the safest setting that is likely to ensure durability\n>       +  across all operating systems and file systems that honor the 'fsync' system\n>       +  call. However, this setting comes with a significant performance cost on\n>      -+  common hardware.\n>      ++  common hardware. Git does not currently fsync parent directories for\n>      ++  newly-added files, so some filesystems may still allow data to be lost on\n>      ++  system crash.\n>       +* `batch` enables an experimental mode that uses interfaces available in some\n>      -+  operating systems to write object data with a minimal set of FLUSH CACHE\n>      -+  (or equivalent) commands sent to the storage controller. If the operating\n>      -+  system interfaces are not available, this mode behaves the same as `true`.\n>      -+  This mode is expected to be safe on macOS for repos stored on HFS+ or APFS\n>      -+  filesystems and on Windows for repos stored on NTFS or ReFS.\n>      ++  operating systems to write loose object data with a minimal set of FLUSH\n>      ++  CACHE (or equivalent) commands sent to the storage controller. If the\n>      ++  operating system interfaces are not available, this mode behaves the same as\n>      ++  `true`. This mode is expected to be as safe as `true` on macOS for repos\n>      ++  stored on HFS+ or APFS filesystems and on Windows for repos stored on NTFS or\n>      ++  ReFS.\n>\n>        core.preloadIndex::\n>         Enable parallel index preload for operations like 'git diff'\n>      @@ builtin/add.c: int cmd_add(int argc, const char **argv, const char *prefix)\n>\n>         if (chmod_arg && pathspec.nr)\n>                 exit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n>      -- unplug_bulk_checkin();\n>       +\n>      -+ unplug_bulk_checkin(&lock_file);\n>      +  unplug_bulk_checkin();\n>\n>        finish:\n>      -  if (write_locked_index(&the_index, &lock_file,\n>\n>        ## bulk-checkin.c ##\n>       @@\n>      @@ bulk-checkin.c\n>        #include \"pack.h\"\n>        #include \"strbuf.h\"\n>       +#include \"string-list.h\"\n>      ++#include \"tmp-objdir.h\"\n>        #include \"packfile.h\"\n>        #include \"object-store.h\"\n>\n>        static int bulk_checkin_plugged;\n>      -\n>      -+static struct string_list bulk_fsync_state = STRING_LIST_INIT_DUP;\n>      ++static int needs_batch_fsync;\n>       +\n>      ++static struct tmp_objdir *bulk_fsync_objdir;\n>      +\n>        static struct bulk_checkin_state {\n>         char *pack_tmp_name;\n>      -  struct hashfile *f;\n>       @@ bulk-checkin.c: clear_exit:\n>         reprepare_packed_git(the_repository);\n>        }\n>\n>      -+static void do_sync_and_rename(struct string_list *fsync_state, struct lock_file *lock_file)\n>      ++/*\n>      ++ * Cleanup after batch-mode fsync_object_files.\n>      ++ */\n>      ++static void do_batch_fsync(void)\n>       +{\n>      -+ if (fsync_state->nr) {\n>      -+         struct string_list_item *rename;\n>      -+\n>      -+         /*\n>      -+          * Issue a full hardware flush against the lock file to ensure\n>      -+          * that all objects are durable before any renames occur.\n>      -+          * The code in fsync_and_close_loose_object_bulk_checkin has\n>      -+          * already ensured that writeout has occurred, but it has not\n>      -+          * flushed any writeback cache in the storage hardware.\n>      -+          */\n>      -+         fsync_or_die(get_lock_file_fd(lock_file), get_lock_file_path(lock_file));\n>      -+\n>      -+         for_each_string_list_item(rename, fsync_state) {\n>      -+                 const char *src = rename->string;\n>      -+                 const char *dst = rename->util;\n>      -+\n>      -+                 if (finalize_object_file(src, dst))\n>      -+                         die_errno(_(\"could not rename '%s' to '%s'\"), src, dst);\n>      -+         }\n>      -+\n>      -+         string_list_clear(fsync_state, 1);\n>      ++ /*\n>      ++  * Issue a full hardware flush against a temporary file to ensure\n>      ++  * that all objects are durable before any renames occur.  The code in\n>      ++  * fsync_loose_object_bulk_checkin has already issued a writeout\n>      ++  * request, but it has not flushed any writeback cache in the storage\n>      ++  * hardware.\n>      ++  */\n>      ++\n>      ++ if (needs_batch_fsync) {\n>      ++         struct strbuf temp_path = STRBUF_INIT;\n>      ++         struct tempfile *temp;\n>      ++\n>      ++         strbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n>      ++         temp = xmks_tempfile(temp_path.buf);\n>      ++         fsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n>      ++         delete_tempfile(&temp);\n>      ++         strbuf_release(&temp_path);\n>       + }\n>      ++\n>      ++ if (bulk_fsync_objdir)\n>      ++         tmp_objdir_migrate(bulk_fsync_objdir);\n>       +}\n>       +\n>        static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n>      @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n>         return 0;\n>        }\n>\n>      -+static void add_rename_bulk_checkin(struct string_list *fsync_state,\n>      -+                             const char *src, const char *dst)\n>      ++void fsync_loose_object_bulk_checkin(int fd)\n>       +{\n>      -+ string_list_insert(fsync_state, src)->util = xstrdup(dst);\n>      -+}\n>      -+\n>      -+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n>      -+                                       const char *filename, time_t mtime)\n>      -+{\n>      -+ int do_finalize = 1;\n>      -+ int ret = 0;\n>      -+\n>      -+ if (fsync_object_files != FSYNC_OBJECT_FILES_OFF) {\n>      -+         /*\n>      -+          * If we have a plugged bulk checkin, we issue a call that\n>      -+          * cleans the filesystem page cache but avoids a hardware flush\n>      -+          * command. Later on we will issue a single hardware flush\n>      -+          * before renaming files as part of do_sync_and_rename.\n>      -+          */\n>      -+         if (bulk_checkin_plugged &&\n>      -+             fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n>      -+             git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n>      -+                 add_rename_bulk_checkin(&bulk_fsync_state, tmpfile, filename);\n>      -+                 do_finalize = 0;\n>      -+\n>      -+         } else {\n>      -+                 fsync_or_die(fd, \"loose object file\");\n>      -+         }\n>      -+ }\n>      -+\n>      -+ if (close(fd))\n>      -+         die_errno(_(\"error when closing loose object file\"));\n>      -+\n>      -+ if (mtime) {\n>      -+         struct utimbuf utb;\n>      -+         utb.actime = mtime;\n>      -+         utb.modtime = mtime;\n>      -+         if (utime(tmpfile, &utb) < 0)\n>      -+                 warning_errno(_(\"failed utime() on %s\"), tmpfile);\n>      ++ assert(fsync_object_files == FSYNC_OBJECT_FILES_BATCH);\n>      ++\n>      ++ /*\n>      ++  * If we have a plugged bulk checkin, we issue a call that\n>      ++  * cleans the filesystem page cache but avoids a hardware flush\n>      ++  * command. Later on we will issue a single hardware flush\n>      ++  * before as part of do_batch_fsync.\n>      ++  */\n>      ++ if (bulk_checkin_plugged &&\n>      ++     git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n>      ++         assert(the_repository->objects->odb->is_temp);\n>      ++         if (!needs_batch_fsync)\n>      ++                 needs_batch_fsync = 1;\n>      ++ } else {\n>      ++         fsync_or_die(fd, \"loose object file\");\n>       + }\n>      -+\n>      -+ if (do_finalize)\n>      -+         ret = finalize_object_file(tmpfile, filename);\n>      -+\n>      -+ return ret;\n>       +}\n>       +\n>        int index_bulk_checkin(struct object_id *oid,\n>                        int fd, size_t size, enum object_type type,\n>                        const char *path, unsigned flags)\n>      -@@ bulk-checkin.c: void plug_bulk_checkin(void)\n>      +@@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n>      + void plug_bulk_checkin(void)\n>      + {\n>      +  assert(!bulk_checkin_plugged);\n>      ++\n>      ++ /*\n>      ++  * Create a temporary object directory if the current\n>      ++  * object directory is not already temporary.\n>      ++  */\n>      ++ if (fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n>      ++     !the_repository->objects->odb->is_temp) {\n>      ++         bulk_fsync_objdir = tmp_objdir_create();\n>      ++         if (!bulk_fsync_objdir)\n>      ++                 die(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n>      ++\n>      ++         tmp_objdir_replace_main_odb(bulk_fsync_objdir);\n>      ++ }\n>      ++\n>         bulk_checkin_plugged = 1;\n>        }\n>\n>      --void unplug_bulk_checkin(void)\n>      -+void unplug_bulk_checkin(struct lock_file *lock_file)\n>      - {\n>      -  assert(bulk_checkin_plugged);\n>      +@@ bulk-checkin.c: void unplug_bulk_checkin(void)\n>         bulk_checkin_plugged = 0;\n>         if (bulk_checkin_state.f)\n>                 finish_bulk_checkin(&bulk_checkin_state);\n>       +\n>      -+ do_sync_and_rename(&bulk_fsync_state, lock_file);\n>      ++ do_batch_fsync();\n>        }\n>\n>        ## bulk-checkin.h ##\n>      @@ bulk-checkin.h\n>\n>        #include \"cache.h\"\n>\n>      -+int fsync_and_close_loose_object_bulk_checkin(int fd, const char *tmpfile,\n>      -+                                       const char *filename, time_t mtime);\n>      ++void fsync_loose_object_bulk_checkin(int fd);\n>       +\n>        int index_bulk_checkin(struct object_id *oid,\n>                        int fd, size_t size, enum object_type type,\n>                        const char *path, unsigned flags);\n>      -\n>      - void plug_bulk_checkin(void);\n>      --void unplug_bulk_checkin(void);\n>      -+void unplug_bulk_checkin(struct lock_file *);\n>      -\n>      - #endif\n>\n>        ## cache.h ##\n>       @@ cache.h: void reset_shared_repository(void);\n>      @@ cache.h: void reset_shared_repository(void);\n>        extern char *git_replace_ref_base;\n>\n>       -extern int fsync_object_files;\n>      -+enum FSYNC_OBJECT_FILES_MODE {\n>      ++enum fsync_object_files_mode {\n>       +    FSYNC_OBJECT_FILES_OFF,\n>       +    FSYNC_OBJECT_FILES_ON,\n>       +    FSYNC_OBJECT_FILES_BATCH\n>       +};\n>       +\n>      -+extern enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n>      ++extern enum fsync_object_files_mode fsync_object_files;\n>        extern int core_preload_index;\n>        extern int precomposed_unicode;\n>        extern int protect_hfs;\n>      @@ environment.c: const char *git_hooks_path;\n>        int core_compression_level;\n>        int pack_compression_level = Z_DEFAULT_COMPRESSION;\n>       -int fsync_object_files;\n>      -+enum FSYNC_OBJECT_FILES_MODE fsync_object_files;\n>      ++enum fsync_object_files_mode fsync_object_files;\n>        size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n>        size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n>        size_t delta_base_cache_limit = 96 * 1024 * 1024;\n>      @@ git-compat-util.h: __attribute__((format (printf, 1, 2))) NORETURN\n>         * Returns 0 on success, which includes trying to unlink an object that does\n>\n>        ## object-file.c ##\n>      -@@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>      -  return 0;\n>      +@@ object-file.c: void add_to_alternates_memory(const char *reference)\n>      +                       '\\n', NULL, 0);\n>        }\n>\n>      --/* Finalize a file on disk, and close it. */\n>      --static void close_loose_object(int fd)\n>      --{\n>      ++struct object_directory *set_temporary_main_odb(const char *dir)\n>      ++{\n>      ++ struct object_directory *main_odb, *new_odb, *old_next;\n>      ++\n>      ++ /*\n>      ++  * Make sure alternates are initialized, or else our entry may be\n>      ++  * overwritten when they are.\n>      ++  */\n>      ++ prepare_alt_odb(the_repository);\n>      ++\n>      ++ /* Copy the existing object directory and make it an alternate. */\n>      ++ main_odb = the_repository->objects->odb;\n>      ++ new_odb = xmalloc(sizeof(*new_odb));\n>      ++ *new_odb = *main_odb;\n>      ++ *the_repository->objects->odb_tail = new_odb;\n>      ++ the_repository->objects->odb_tail = &(new_odb->next);\n>      ++ new_odb->next = NULL;\n>      ++\n>      ++ /*\n>      ++  * Reinitialize the main odb with the specified path, being careful\n>      ++  * to keep the next pointer value.\n>      ++  */\n>      ++ old_next = main_odb->next;\n>      ++ memset(main_odb, 0, sizeof(*main_odb));\n>      ++ main_odb->next = old_next;\n>      ++ main_odb->is_temp = 1;\n>      ++ main_odb->path = xstrdup(dir);\n>      ++ return new_odb;\n>      ++}\n>      ++\n>      ++void restore_main_odb(struct object_directory *odb)\n>      ++{\n>      ++ struct object_directory **prev, *main_odb;\n>      ++\n>      ++ /* Unlink the saved previous main ODB from the list. */\n>      ++ prev = &the_repository->objects->odb->next;\n>      ++ assert(*prev);\n>      ++ while (*prev != odb) {\n>      ++         prev = &(*prev)->next;\n>      ++ }\n>      ++ *prev = odb->next;\n>      ++ if (*prev == NULL)\n>      ++         the_repository->objects->odb_tail = prev;\n>      ++\n>      ++ /*\n>      ++  * Restore the data from the old main odb, being careful to\n>      ++  * keep the next pointer value\n>      ++  */\n>      ++ main_odb = the_repository->objects->odb;\n>      ++ SWAP(*main_odb, *odb);\n>      ++ main_odb->next = odb->next;\n>      ++ free_object_directory(odb);\n>      ++}\n>      ++\n>      + /*\n>      +  * Compute the exact path an alternate is at and returns it. In case of\n>      +  * error NULL is returned and the human readable error is added to `err`\n>      +@@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>      + /* Finalize a file on disk, and close it. */\n>      + static void close_loose_object(int fd)\n>      + {\n>       - if (fsync_object_files)\n>      --         fsync_or_die(fd, \"loose object file\");\n>      -- if (close(fd) != 0)\n>      --         die_errno(_(\"error when closing loose object file\"));\n>      --}\n>      --\n>      - /* Size of directory component, including the ending '/' */\n>      - static inline int directory_size(const char *filename)\n>      ++ switch (fsync_object_files) {\n>      ++ case FSYNC_OBJECT_FILES_OFF:\n>      ++         break;\n>      ++ case FSYNC_OBJECT_FILES_ON:\n>      +          fsync_or_die(fd, \"loose object file\");\n>      ++         break;\n>      ++ case FSYNC_OBJECT_FILES_BATCH:\n>      ++         fsync_loose_object_bulk_checkin(fd);\n>      ++         break;\n>      ++ default:\n>      ++         BUG(\"Invalid fsync_object_files mode.\");\n>      ++ }\n>      ++\n>      +  if (close(fd) != 0)\n>      +          die_errno(_(\"error when closing loose object file\"));\n>      + }\n>      +\n>      + ## object-store.h ##\n>      +@@ object-store.h: void add_to_alternates_file(const char *dir);\n>      +  */\n>      + void add_to_alternates_memory(const char *dir);\n>      +\n>      ++/*\n>      ++ * Replace the current main object directory with the specified temporary\n>      ++ * object directory. We make a copy of the former main object directory,\n>      ++ * add it as an in-memory alternate, and return the copy so that it can\n>      ++ * be restored via restore_main_odb.\n>      ++ */\n>      ++struct object_directory *set_temporary_main_odb(const char *dir);\n>      ++\n>      ++/*\n>      ++ * Restore a previous ODB replaced by set_temporary_main_odb.\n>      ++ */\n>      ++void restore_main_odb(struct object_directory *odb);\n>      ++\n>      + /*\n>      +  * Populate and return the loose object cache array corresponding to the\n>      +  * given object ID.\n>      +@@ object-store.h: struct oidtree *odb_loose_cache(struct object_directory *odb,\n>      + /* Empty the loose object cache for the specified object directory. */\n>      + void odb_clear_loose_cache(struct object_directory *odb);\n>      +\n>      ++/* Clear and free the specified object directory */\n>      ++void free_object_directory(struct object_directory *odb);\n>      ++\n>      + struct packed_git {\n>      +  struct hashmap_entry packmap_ent;\n>      +  struct packed_git *next;\n>      +\n>      + ## object.c ##\n>      +@@ object.c: struct raw_object_store *raw_object_store_new(void)\n>      +  return o;\n>      + }\n>      +\n>      +-static void free_object_directory(struct object_directory *odb)\n>      ++void free_object_directory(struct object_directory *odb)\n>        {\n>      -@@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n>      -          die(_(\"confused by unstable object source data for %s\"),\n>      -              oid_to_hex(oid));\n>      +  free(odb->path);\n>      +  odb_clear_loose_cache(odb);\n>      +\n>      + ## tmp-objdir.c ##\n>      +@@\n>      + struct tmp_objdir {\n>      +  struct strbuf path;\n>      +  struct strvec env;\n>      ++ struct object_directory *prev_main_odb;\n>      + };\n>\n>      -- close_loose_object(fd);\n>      --\n>      -- if (mtime) {\n>      --         struct utimbuf utb;\n>      --         utb.actime = mtime;\n>      --         utb.modtime = mtime;\n>      --         if (utime(tmp_file.buf, &utb) < 0)\n>      --                 warning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n>      -- }\n>      --\n>      -- return finalize_object_file(tmp_file.buf, filename.buf);\n>      -+ return fsync_and_close_loose_object_bulk_checkin(fd, tmp_file.buf,\n>      -+                                                  filename.buf, mtime);\n>      + /*\n>      +@@ tmp-objdir.c: static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n>      +   * freeing memory; it may cause a deadlock if the signal\n>      +   * arrived while libc's allocator lock is held.\n>      +   */\n>      +- if (!on_signal)\n>      ++ if (!on_signal) {\n>      ++         if (t->prev_main_odb)\n>      ++                 restore_main_odb(t->prev_main_odb);\n>      +          tmp_objdir_free(t);\n>      ++ }\n>      ++\n>      +  return err;\n>        }\n>\n>      - static int freshen_loose_object(const struct object_id *oid)\n>      +@@ tmp-objdir.c: struct tmp_objdir *tmp_objdir_create(void)\n>      +  t = xmalloc(sizeof(*t));\n>      +  strbuf_init(&t->path, 0);\n>      +  strvec_init(&t->env);\n>      ++ t->prev_main_odb = NULL;\n>      +\n>      +  strbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n>      +\n>      +@@ tmp-objdir.c: int tmp_objdir_migrate(struct tmp_objdir *t)\n>      +  if (!t)\n>      +          return 0;\n>      +\n>      ++ if (t->prev_main_odb) {\n>      ++         restore_main_odb(t->prev_main_odb);\n>      ++         t->prev_main_odb = NULL;\n>      ++ }\n>      ++\n>      +  strbuf_addbuf(&src, &t->path);\n>      +  strbuf_addstr(&dst, get_object_directory());\n>      +\n>      +@@ tmp-objdir.c: void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n>      + {\n>      +  add_to_alternates_memory(t->path.buf);\n>      + }\n>      ++\n>      ++void tmp_objdir_replace_main_odb(struct tmp_objdir *t)\n>      ++{\n>      ++ if (t->prev_main_odb)\n>      ++         BUG(\"the main object database is already replaced\");\n>      ++ t->prev_main_odb = set_temporary_main_odb(t->path.buf);\n>      ++}\n>      +\n>      + ## tmp-objdir.h ##\n>      +@@ tmp-objdir.h: int tmp_objdir_destroy(struct tmp_objdir *);\n>      +  */\n>      + void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n>      +\n>      ++/*\n>      ++ * Replaces the main object store in the current process with the temporary\n>      ++ * object directory and makes the former main object store an alternate.\n>      ++ */\n>      ++void tmp_objdir_replace_main_odb(struct tmp_objdir *);\n>      ++\n>      + #endif /* TMP_OBJDIR_H */\n>\n>        ## wrapper.c ##\n>       @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n>  3:  a5b3e21b762 < -:  ----------- core.fsyncobjectfiles: add windows support for batch mode\n>  4:  f7f756f3932 ! 4:  485b4a767df update-index: use the bulk-checkin infrastructure\n>      @@ Commit message\n>           There is some risk with this change, since under batch fsync, the object\n>           files will not be available until the update-index is entirely complete.\n>           This usage is unlikely, since any tool invoking update-index and\n>      -    expecting to see objects would have to snoop the output of --verbose to\n>      -    find out when update-index has actually processed a given path.\n>      -    Additionally the index is locked for the duration of the update.\n>      +    expecting to see objects would have to synchronize with the update-index\n>      +    process after passing it a file path.\n>\n>           Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>\n>      @@ builtin/update-index.c\n>        #include \"lockfile.h\"\n>        #include \"quote.h\"\n>       @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n>      -          struct strbuf unquoted = STRBUF_INIT;\n>\n>      -          setup_work_tree();\n>      -+         plug_bulk_checkin();\n>      -          while (getline_fn(&buf, stdin) != EOF) {\n>      -                  char *p;\n>      -                  if (!nul_term_line && buf.buf[0] == '\"') {\n>      +  the_index.updated_skipworktree = 1;\n>      +\n>      ++ /* we might be adding many objects to the object database */\n>      ++ plug_bulk_checkin();\n>      ++\n>      +  /*\n>      +   * Custom copy of parse_options() because we want to handle\n>      +   * filename arguments as they come.\n>       @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n>      -                          chmod_path(set_executable_bit, p);\n>      -                  free(p);\n>      -          }\n>      -+         unplug_bulk_checkin(&lock_file);\n>      -          strbuf_release(&unquoted);\n>                 strbuf_release(&buf);\n>         }\n>      +\n>      ++ /* by now we must have added all of the new objects */\n>      ++ unplug_bulk_checkin();\n>      +  if (split_index > 0) {\n>      +          if (git_config_get_split_index() == 0)\n>      +                  warning(_(\"core.splitIndex is set to false; \"\n>  -:  ----------- > 5:  889e7668760 unpack-objects: use the bulk-checkin infrastructure\n>  5:  afb0028e796 ! 6:  0f2e3b25759 core.fsyncobjectfiles: tests for batch mode\n>      @@ Metadata\n>        ## Commit message ##\n>           core.fsyncobjectfiles: tests for batch mode\n>\n>      -    Add test cases to exercise batch mode for 'git add'\n>      -    and 'git stash'. These tests ensure that the added\n>      -    data winds up in the object database.\n>      +    Add test cases to exercise batch mode for:\n>      +     * 'git add'\n>      +     * 'git stash'\n>      +     * 'git update-index'\n>      +     * 'git unpack-objects'\n>\n>      -    I verified the tests by introducing an incorrect rename\n>      -    in do_sync_and_rename.\n>      +    These tests ensure that the added data winds up in the object database.\n>      +\n>      +    In this change we introduce a new test helper lib-unique-files.sh. The\n>      +    goal of this library is to create a tree of files that have different\n>      +    oids from any other files that may have been created in the current test\n>      +    repo. This helps us avoid missing validation of an object being added due\n>      +    to it already being in the repo.\n>\n>           Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>\n>      @@ t/lib-unique-files.sh (new)\n>       @@\n>       +# Helper to create files with unique contents\n>       +\n>      -+test_create_unique_files_base__=$(date -u)\n>      -+test_create_unique_files_counter__=0\n>       +\n>       +# Create multiple files with unique contents. Takes the number of\n>       +# directories, the number of files in each directory, and the base\n>       +# directory.\n>       +#\n>      -+# test_create_unique_files 2 3 . -- Creates 2 directories with 3 files\n>      -+#                                    each in the specified directory, all\n>      -+#                                    with unique contents.\n>      ++# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n>      ++#                                         each in my_dir, all with unique\n>      ++#                                         contents.\n>       +\n>       +test_create_unique_files() {\n>       + test \"$#\" -ne 3 && BUG \"3 param\"\n>      @@ t/lib-unique-files.sh (new)\n>       + local dirs=$1\n>       + local files=$2\n>       + local basedir=$3\n>      ++ local counter=0\n>      ++ test_tick\n>      ++ local basedata=$test_tick\n>      ++\n>       +\n>      -+ rm -rf $basedir >/dev/null\n>      ++ rm -rf $basedir\n>       +\n>       + for i in $(test_seq $dirs)\n>       + do\n>       +         local dir=$basedir/dir$i\n>       +\n>      -+         mkdir -p \"$dir\" > /dev/null\n>      ++         mkdir -p \"$dir\"\n>       +         for j in $(test_seq $files)\n>       +         do\n>      -+                 test_create_unique_files_counter__=$((test_create_unique_files_counter__ + 1))\n>      -+                 echo \"$test_create_unique_files_base__.$test_create_unique_files_counter__\"  >\"$dir/file$j.txt\"\n>      ++                 counter=$((counter + 1))\n>      ++                 echo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n>       +         done\n>       + done\n>       +}\n>      @@ t/t3700-add.sh: test_expect_success \\\n>       + rm -f fsynced_files &&\n>       + git ls-files --stage fsync-files/ > fsynced_files &&\n>       + test_line_count = 8 fsynced_files &&\n>      -+ cat fsynced_files | awk '{print \\$2}' | xargs -n1 git cat-file -e\n>      ++ awk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n>      ++\"\n>      ++\n>      ++test_expect_success 'git update-index: core.fsyncobjectfiles=batch' \"\n>      ++ test_create_unique_files 2 4 fsync-files2 &&\n>      ++ find fsync-files2 ! -type d -print | xargs git -c core.fsyncobjectfiles=batch update-index --add -- &&\n>      ++ rm -f fsynced_files2 &&\n>      ++ git ls-files --stage fsync-files2/ > fsynced_files2 &&\n>      ++ test_line_count = 8 fsynced_files2 &&\n>      ++ awk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n>       +\"\n>       +\n>        test_expect_success \\\n>      @@ t/t3903-stash.sh: test_expect_success 'stash handles skip-worktree entries nicel\n>       + # which contains the untracked files\n>       + git ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n>       + test_line_count = 8 fsynced_files &&\n>      -+ cat fsynced_files | awk '{print \\$3}' | xargs -n1 git cat-file -e\n>      ++ awk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n>       +\"\n>       +\n>       +\n>        test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n>         expected=\"stash.useBuiltin support has been removed\" &&\n>\n>      +\n>      + ## t/t5300-pack-object.sh ##\n>      +@@ t/t5300-pack-object.sh: test_expect_success 'pack-objects with bogus arguments' '\n>      +\n>      + check_unpack () {\n>      +  test_when_finished \"rm -rf git2\" &&\n>      +- git init --bare git2 &&\n>      +- git -C git2 unpack-objects -n <\"$1\".pack &&\n>      +- git -C git2 unpack-objects <\"$1\".pack &&\n>      +- (cd .git && find objects -type f -print) |\n>      +- while read path\n>      +- do\n>      +-         cmp git2/$path .git/$path || {\n>      +-                 echo $path differs.\n>      +-                 return 1\n>      +-         }\n>      +- done\n>      ++ git $2 init --bare git2 &&\n>      ++ (\n>      ++         git $2 -C git2 unpack-objects -n <\"$1\".pack &&\n>      ++         git $2 -C git2 unpack-objects <\"$1\".pack &&\n>      ++         git $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n>      ++ ) <obj-list >current &&\n>      ++ cmp obj-list current\n>      + }\n>      +\n>      + test_expect_success 'unpack without delta' '\n>      +  check_unpack test-1-${packname_1}\n>      + '\n>      +\n>      ++test_expect_success 'unpack without delta (core.fsyncobjectfiles=batch)' '\n>      ++ check_unpack test-1-${packname_1} \"-c core.fsyncobjectfiles=batch\"\n>      ++'\n>      ++\n>      + test_expect_success 'pack with REF_DELTA' '\n>      +  packname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n>      +  check_deltas stderr -gt 0\n>      +@@ t/t5300-pack-object.sh: test_expect_success 'unpack with REF_DELTA' '\n>      +  check_unpack test-2-${packname_2}\n>      + '\n>      +\n>      ++test_expect_success 'unpack with REF_DELTA (core.fsyncobjectfiles=batch)' '\n>      ++       check_unpack test-2-${packname_2} \"-c core.fsyncobjectfiles=batch\"\n>      ++'\n>      ++\n>      + test_expect_success 'pack with OFS_DELTA' '\n>      +  packname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n>      +                  <obj-list 2>stderr) &&\n>      +@@ t/t5300-pack-object.sh: test_expect_success 'unpack with OFS_DELTA' '\n>      +  check_unpack test-3-${packname_3}\n>      + '\n>      +\n>      ++test_expect_success 'unpack with OFS_DELTA (core.fsyncobjectfiles=batch)' '\n>      ++       check_unpack test-3-${packname_3} \"-c core.fsyncobjectfiles=batch\"\n>      ++'\n>      ++\n>      + test_expect_success 'compare delta flavors' '\n>      +  perl -e '\\''\n>      +          defined($_ = -s $_) or die for @ARGV;\n>  6:  3e6b80b5fa2 = 7:  6543564376a core.fsyncobjectfiles: performance tests for add and stash\n>\n> --\n> gitgitgadget\n\nApologies for the spam, I'll be submitting a v6 shortly since there\nwere several things wrong with\nthis version.\n"},{"id":"437054","messageId":"e4081f81f6ae4876bf3dff09b708c9ea89fffa59.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 1/8] object-file.c: do not rename in a temp odb","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:22Z","receivedAt":"2021-09-24T23:53:34Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIf a temporary ODB is active, as determined by GIT_QUARANTINE_PATH\nbeing set, create object files with their final names. This avoids\nan extra rename beyond what is needed to merge the temporary ODB in\ntmp_objdir_migrate.\n\nCreating an object file with the expected final name should be okay\nsince the git process writing to the temporary object store is the\nonly writer, and it only invokes write_loose_object/create_object_file\nafter checking that the object doesn't exist.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n environment.c  |  4 ++++\n object-file.c  | 51 ++++++++++++++++++++++++++++++++++----------------\n object-store.h |  6 ++++++\n repository.c   |  2 ++\n repository.h   |  1 +\n 5 files changed, 48 insertions(+), 16 deletions(-)\n\ndiff --git a/environment.c b/environment.c\nindex d6b22ede7ea..d9ba68402e9 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -177,6 +177,10 @@ void setup_git_env(const char *git_dir)\n \targs.graft_file = getenv_safe(&to_free, GRAFT_ENVIRONMENT);\n \targs.index_file = getenv_safe(&to_free, INDEX_ENVIRONMENT);\n \targs.alternate_db = getenv_safe(&to_free, ALTERNATE_DB_ENVIRONMENT);\n+\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n+\t\targs.object_dir_is_temp = 1;\n+\t}\n+\n \trepo_set_gitdir(the_repository, git_dir, &args);\n \tstrvec_clear(&to_free);\n \ndiff --git a/object-file.c b/object-file.c\nindex a8be8994814..ab593515cec 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1800,12 +1800,17 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n }\n \n /*\n- * Move the just written object into its final resting place.\n+ * Move the just written object into its final resting place,\n+ * unless it is already there, as indicated by an empty string for\n+ * tmpfile.\n  */\n int finalize_object_file(const char *tmpfile, const char *filename)\n {\n \tint ret = 0;\n \n+\tif (!*tmpfile)\n+\t\tgoto out;\n+\n \tif (object_creation_mode == OBJECT_CREATION_USES_RENAMES)\n \t\tgoto try_rename;\n \telse if (link(tmpfile, filename))\n@@ -1878,21 +1883,37 @@ static inline int directory_size(const char *filename)\n }\n \n /*\n- * This creates a temporary file in the same directory as the final\n- * 'filename'\n+ * This creates a loose object file for the specified object id.\n+ * If we're working in a temporary object directory, the file is\n+ * created with its final filename, otherwise it is created with\n+ * a temporary name and renamed by finalize_object_file.\n+ * If no rename is required, an empty string is returned in tmp.\n  *\n  * We want to avoid cross-directory filename renames, because those\n  * can have problems on various filesystems (FAT, NFS, Coda).\n  */\n-static int create_tmpfile(struct strbuf *tmp, const char *filename)\n+static int create_objfile(const struct object_id *oid, struct strbuf *tmp,\n+\t\t\t  struct strbuf *filename)\n {\n-\tint fd, dirlen = directory_size(filename);\n+\tint fd, dirlen, is_retrying = 0;\n+\tconst char *object_name;\n+\tstatic const int object_mode = 0444;\n \n+\tloose_object_path(the_repository, filename, oid);\n+\tdirlen = directory_size(filename->buf);\n+\n+retry_create:\n \tstrbuf_reset(tmp);\n-\tstrbuf_add(tmp, filename, dirlen);\n-\tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n-\tfd = git_mkstemp_mode(tmp->buf, 0444);\n-\tif (fd < 0 && dirlen && errno == ENOENT) {\n+\tif (!the_repository->objects->odb->is_temp) {\n+\t\tstrbuf_add(tmp, filename->buf, dirlen);\n+\t\tobject_name = \"tmp_obj_XXXXXX\";\n+\t\tstrbuf_addstr(tmp, object_name);\n+\t\tfd = git_mkstemp_mode(tmp->buf, object_mode);\n+\t} else {\n+\t\tfd = open(filename->buf, O_CREAT | O_EXCL | O_RDWR, object_mode);\n+\t}\n+\n+\tif (fd < 0 && dirlen && errno == ENOENT && !is_retrying) {\n \t\t/*\n \t\t * Make sure the directory exists; note that the contents\n \t\t * of the buffer are undefined after mkstemp returns an\n@@ -1900,15 +1921,15 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \t\t * scratch.\n \t\t */\n \t\tstrbuf_reset(tmp);\n-\t\tstrbuf_add(tmp, filename, dirlen - 1);\n+\t\tstrbuf_add(tmp, filename->buf, dirlen - 1);\n \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n \t\t\treturn -1;\n \t\tif (adjust_shared_perm(tmp->buf))\n \t\t\treturn -1;\n \n \t\t/* Try again */\n-\t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n-\t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n+\t\tis_retrying = 1;\n+\t\tgoto retry_create;\n \t}\n \treturn fd;\n }\n@@ -1925,14 +1946,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n-\tloose_object_path(the_repository, &filename, oid);\n-\n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n+\tfd = create_objfile(oid, &tmp_file, &filename);\n \tif (fd < 0) {\n \t\tif (errno == EACCES)\n \t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n \t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\t\t\treturn error_errno(_(\"unable to create object file\"));\n \t}\n \n \t/* Set it up */\ndiff --git a/object-store.h b/object-store.h\nindex b4dc6668aa2..f8c883a5730 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -26,6 +26,12 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/*\n+\t * This is a temporary object store, so there is no need to\n+\t * create new objects via rename.\n+\t */\n+\tint is_temp;\n+\n \t/*\n \t * Path to the alternative object store. If this is a relative path,\n \t * it is relative to the current working directory.\ndiff --git a/repository.c b/repository.c\nindex b2bf44c6faf..a16de04dfa8 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -80,6 +80,8 @@ void repo_set_gitdir(struct repository *repo,\n \texpand_base_dir(&repo->objects->odb->path, o->object_dir,\n \t\t\trepo->commondir, \"objects\");\n \n+\trepo->objects->odb->is_temp = o->object_dir_is_temp;\n+\n \tfree(repo->objects->alternate_db);\n \trepo->objects->alternate_db = xstrdup_or_null(o->alternate_db);\n \texpand_base_dir(&repo->graft_file, o->graft_file,\ndiff --git a/repository.h b/repository.h\nindex 3740c93bc0f..d3711367a6f 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -162,6 +162,7 @@ struct set_gitdir_args {\n \tconst char *graft_file;\n \tconst char *index_file;\n \tconst char *alternate_db;\n+\tint object_dir_is_temp;\n };\n \n void repo_set_gitdir(struct repository *repo, const char *root,\n-- \ngitgitgadget\n\n"},{"id":"437056","messageId":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com","subject":"[PATCH v6 0/8] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:21Z","receivedAt":"2021-09-24T23:53:34Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far!\n\nv5 was a bit of a dud, with some issues that I only noticed after\nsubmitting. v6 changes:\n\n * re-add Windows support\n * fix minor formatting issues\n * reset git author and commit dates which got messed up\n\nChanges since v4, all in response to review feedback from Ævar Arnfjörð\nBjarmason:\n\n * Update core.fsyncobjectfiles documentation to specify 'loose' objects and\n   to add a statement about not fsyncing parent directories.\n   \n   * I still don't want to make any promises on behalf of the Linux FS developers\n     in the documentation. However, according to [v4.1] and my understanding\n     of how XFS journals are documented to work, it looks like recent versions\n     of Linux running on XFS should be as safe as Windows or macOS in 'batch'\n     mode. I don't know about ext4, since it's not clear to me when metadata\n     updates are made visible to the journal.\n   \n\n * Rewrite the core batched fsync change to use the tmp-objdir lib. As Ævar\n   pointed out, this lets us access the added loose objects immediately,\n   rather than only after unplugging the bulk checkin. This is a hard\n   requirement in unpack-objects for resolving OBJ_REF_DELTA packed objects.\n   \n   * As a preparatory patch, the object-file code now doesn't do a rename if it's in a\n     tmp objdir (as determined by the quarantine environment variable).\n   \n   * I added support to the tmp-objdir lib to replace the 'main' writable odb.\n   \n   * Instead of using a lockfile for the final full fsync, we now use a new dummy\n     temp file. Doing that makes the below unpack-objects change easier.\n   \n\n * Add bulk-checkin support to unpack-objects, which is used in fetch and\n   push. In addition to making those operations faster, it allows us to\n   directly compare performance of packfiles against loose objects. Please\n   see [v4.2] for a measurement of 'git push' to a local upstream with\n   different numbers of unique new files.\n\n * Rename FSYNC_OBJECT_FILES_MODE to fsync_object_files_mode.\n\n * Remove comment with link to NtFlushBuffersFileEx documentation.\n\n * Make t/lib-unique-files.sh a bit cleaner. We are still creating unique\n   contents, but now this uses test_tick, so it should be deterministic from\n   run to run.\n\n * Ensure there are tests for all of the modified commands. Make the\n   unpack-objects tests validate that the unpacked objects are really\n   available in the ODB.\n\nReferences for v4: [v4.1]\nhttps://lore.kernel.org/linux-fsdevel/20190419072938.31320-1-amir73il@gmail.com/#t\n\n[v4.2]\nhttps://docs.google.com/spreadsheets/d/1uxMBkEXFFnQ1Y3lXKqcKpw6Mq44BzhpCAcPex14T-QQ/edit#gid=1898936117\n\nChanges since v3:\n\n * Fix core.fsyncobjectfiles option parsing as suggested by Junio: We now\n   accept no value to mean \"true\" and we require 'batch' to be lowercase.\n\n * Leave the default fsync mode as 'false'. Git for windows can change its\n   default when this series makes it over to that fork.\n\n * Use a switch statement in git_fsync, as suggested by Junio.\n\n * Add regression test cases for core.fsyncobjectfiles=batch. This should\n   keep the batch functionality basically working in upstream git even if\n   few users adopt batch mode initially. I expect git-for-windows will\n   provide a good baking area for the new mode.\n\nNeeraj Singh (8):\n  object-file.c: do not rename in a temp odb\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  core.fsyncobjectfiles: add windows support for batch mode\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: tests for batch mode\n  core.fsyncobjectfiles: performance tests for add and stash\n\n Documentation/config/core.txt       |  29 +++++--\n Makefile                            |   6 ++\n builtin/unpack-objects.c            |   3 +\n builtin/update-index.c              |   6 ++\n bulk-checkin.c                      |  92 +++++++++++++++++++---\n bulk-checkin.h                      |   2 +\n cache.h                             |   8 +-\n compat/mingw.h                      |   3 +\n compat/win32/flush.c                |  28 +++++++\n config.c                            |   7 +-\n config.mak.uname                    |   3 +\n configure.ac                        |   8 ++\n contrib/buildsystems/CMakeLists.txt |   3 +-\n environment.c                       |   6 +-\n git-compat-util.h                   |   7 ++\n object-file.c                       | 118 ++++++++++++++++++++++++----\n object-store.h                      |  22 ++++++\n object.c                            |   2 +-\n repository.c                        |   2 +\n repository.h                        |   1 +\n t/lib-unique-files.sh               |  36 +++++++++\n t/perf/p3700-add.sh                 |  43 ++++++++++\n t/perf/p3900-stash.sh               |  46 +++++++++++\n t/t3700-add.sh                      |  20 +++++\n t/t3903-stash.sh                    |  14 ++++\n t/t5300-pack-object.sh              |  30 ++++---\n tmp-objdir.c                        |  20 ++++-\n tmp-objdir.h                        |   6 ++\n wrapper.c                           |  48 +++++++++++\n write-or-die.c                      |   2 +-\n 30 files changed, 570 insertions(+), 51 deletions(-)\n create mode 100644 compat/win32/flush.c\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: 8b7c11b8668b4e774f81a9f0b4c30144b818f1d1\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v6\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v6\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v5:\n\n 1:  95315f35a28 = 1:  e4081f81f6a object-file.c: do not rename in a temp odb\n 2:  df6fab94d67 = 2:  ebba65e040c bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n 3:  fe19cdfc930 ! 3:  543ea356934 core.fsyncobjectfiles: batched disk flushes\n     @@ Makefile: ifdef HAVE_CLOCK_MONOTONIC\n       \tEXTLIBS += -lrt\n       endif\n      \n     - ## builtin/add.c ##\n     -@@ builtin/add.c: int cmd_add(int argc, const char **argv, const char *prefix)\n     - \n     - \tif (chmod_arg && pathspec.nr)\n     - \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n     -+\n     - \tunplug_bulk_checkin();\n     - \n     - finish:\n     -\n       ## bulk-checkin.c ##\n      @@\n        */\n -:  ----------- > 4:  bdb99822f8c core.fsyncobjectfiles: add windows support for batch mode\n 4:  485b4a767df ! 5:  92e18cedab0 update-index: use the bulk-checkin infrastructure\n     @@ Commit message\n      \n          This change enables bulk-checkin for update-index infrastructure to\n          speed up adding new objects to the object database by leveraging the\n     -    pack functionality and the new bulk-fsync functionality. This mode\n     -    is enabled when passing paths to update-index via the --stdin flag,\n     -    as is done by 'git stash'.\n     +    pack functionality and the new bulk-fsync functionality.\n      \n          There is some risk with this change, since under batch fsync, the object\n          files will not be available until the update-index is entirely complete.\n 5:  889e7668760 = 6:  e3c5a11f225 unpack-objects: use the bulk-checkin infrastructure\n 6:  0f2e3b25759 = 7:  385199354fa core.fsyncobjectfiles: tests for batch mode\n 7:  6543564376a = 8:  504bcc95c56 core.fsyncobjectfiles: performance tests for add and stash\n\n-- \ngitgitgadget\n"},{"id":"437055","messageId":"ebba65e040cfa0b9d157e0dc383a6213c81f1943.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 2/8] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:23Z","receivedAt":"2021-09-24T23:53:36Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex b023d9959aa..f117d62c908 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_bulk_checkin(struct bulk_checkin_state *state)\n {\n@@ -260,21 +260,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"437057","messageId":"543ea3569342165363c1602ce36683a54dce7a0b.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 3/8] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:24Z","receivedAt":"2021-09-24T23:53:39Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nand macOS offer mechanisms to write data from the filesystem page cache\nwithout initiating a hardware flush. Linux has the sync_file_range API,\nwhich issues a pagecache writeback request reliably after version 5.2.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\nbatches up hardware flushes. It hooks into the bulk-checkin plugging and\nunplugging functionality and takes advantage of tmp-objdir.\n\nWhen the new mode is enabled we do the following for each new object:\n1. Create the object in a tmp-objdir.\n2. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin we:\n1. Issue an fsync against a dummy file to flush the hardware writeback\n   cache, which should by now have processed the tmp-objdir writes.\n2. Rename all of the tmp-objdir files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the case today,\n   but may be a good extension to those components.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS we\nwould expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 29 ++++++++++++---\n Makefile                      |  6 +++\n bulk-checkin.c                | 70 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  2 +\n cache.h                       |  8 +++-\n config.c                      |  7 +++-\n config.mak.uname              |  1 +\n configure.ac                  |  8 ++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 ++++\n object-file.c                 | 67 ++++++++++++++++++++++++++++++++-\n object-store.h                | 16 ++++++++\n object.c                      |  2 +-\n tmp-objdir.c                  | 20 +++++++++-\n tmp-objdir.h                  |  6 +++\n wrapper.c                     | 44 ++++++++++++++++++++++\n write-or-die.c                |  2 +-\n 17 files changed, 284 insertions(+), 13 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..200b4d9f06e 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,29 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols loose objects in the object store, so updates to any\n+\trefs or the index may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each loose object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware. Git does not currently fsync parent directories for\n+  newly-added files, so some filesystems may still allow data to be lost on\n+  system crash.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write loose object data with a minimal set of FLUSH\n+  CACHE (or equivalent) commands sent to the storage controller. If the\n+  operating system interfaces are not available, this mode behaves the same as\n+  `true`. This mode is expected to be as safe as `true` on macOS for repos\n+  stored on HFS+ or APFS filesystems and on Windows for repos stored on NTFS or\n+  ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex 429c276058d..326c7607e0f 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1896,6 +1898,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex f117d62c908..957a6238684 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,14 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n+static int needs_batch_fsync;\n+\n+static struct tmp_objdir *bulk_fsync_objdir;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -62,6 +68,34 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur.  The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware.\n+\t */\n+\n+\tif (needs_batch_fsync) {\n+\t\tstruct strbuf temp_path = STRBUF_INIT;\n+\t\tstruct tempfile *temp;\n+\n+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\t\ttemp = xmks_tempfile(temp_path.buf);\n+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\t\tdelete_tempfile(&temp);\n+\t\tstrbuf_release(&temp_path);\n+\t}\n+\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -256,6 +290,26 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void fsync_loose_object_bulk_checkin(int fd)\n+{\n+\tassert(fsync_object_files == FSYNC_OBJECT_FILES_BATCH);\n+\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (bulk_checkin_plugged &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\tassert(the_repository->objects->odb->is_temp);\n+\t\tif (!needs_batch_fsync)\n+\t\t\tneeds_batch_fsync = 1;\n+\t} else {\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -270,6 +324,20 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n+\n+\t/*\n+\t * Create a temporary object directory if the current\n+\t * object directory is not already temporary.\n+\t */\n+\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n+\t    !the_repository->objects->odb->is_temp) {\n+\t\tbulk_fsync_objdir = tmp_objdir_create();\n+\t\tif (!bulk_fsync_objdir)\n+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n+\n+\t\ttmp_objdir_replace_main_odb(bulk_fsync_objdir);\n+\t}\n+\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -279,4 +347,6 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..08f292379b6 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,8 @@\n \n #include \"cache.h\"\n \n+void fsync_loose_object_bulk_checkin(int fd);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex d23de693680..d1897fe9d92 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum fsync_object_files_mode {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum fsync_object_files_mode fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex cb4a8058bff..1b403e00241 100644\n--- a/config.c\n+++ b/config.c\n@@ -1509,7 +1509,12 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (value && !strcmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse if (git_config_bool(var, value))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n+\t\telse\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 76516aaa9a5..e6d482fbcc6 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex d9ba68402e9..f318d59e585 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -43,7 +43,7 @@ const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int core_compression_level;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_object_files_mode fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b46605300ab..d14e2436276 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1210,6 +1210,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex ab593515cec..ec22560dd66 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -750,6 +750,60 @@ void add_to_alternates_memory(const char *reference)\n \t\t\t     '\\n', NULL, 0);\n }\n \n+struct object_directory *set_temporary_main_odb(const char *dir)\n+{\n+\tstruct object_directory *main_odb, *new_odb, *old_next;\n+\n+\t/*\n+\t * Make sure alternates are initialized, or else our entry may be\n+\t * overwritten when they are.\n+\t */\n+\tprepare_alt_odb(the_repository);\n+\n+\t/* Copy the existing object directory and make it an alternate. */\n+\tmain_odb = the_repository->objects->odb;\n+\tnew_odb = xmalloc(sizeof(*new_odb));\n+\t*new_odb = *main_odb;\n+\t*the_repository->objects->odb_tail = new_odb;\n+\tthe_repository->objects->odb_tail = &(new_odb->next);\n+\tnew_odb->next = NULL;\n+\n+\t/*\n+\t * Reinitialize the main odb with the specified path, being careful\n+\t * to keep the next pointer value.\n+\t */\n+\told_next = main_odb->next;\n+\tmemset(main_odb, 0, sizeof(*main_odb));\n+\tmain_odb->next = old_next;\n+\tmain_odb->is_temp = 1;\n+\tmain_odb->path = xstrdup(dir);\n+\treturn new_odb;\n+}\n+\n+void restore_main_odb(struct object_directory *odb)\n+{\n+\tstruct object_directory **prev, *main_odb;\n+\n+\t/* Unlink the saved previous main ODB from the list. */\n+\tprev = &the_repository->objects->odb->next;\n+\tassert(*prev);\n+\twhile (*prev != odb) {\n+\t\tprev = &(*prev)->next;\n+\t}\n+\t*prev = odb->next;\n+\tif (*prev == NULL)\n+\t\tthe_repository->objects->odb_tail = prev;\n+\n+\t/*\n+\t * Restore the data from the old main odb, being careful to\n+\t * keep the next pointer value\n+\t */\n+\tmain_odb = the_repository->objects->odb;\n+\tSWAP(*main_odb, *odb);\n+\tmain_odb->next = odb->next;\n+\tfree_object_directory(odb);\n+}\n+\n /*\n  * Compute the exact path an alternate is at and returns it. In case of\n  * error NULL is returned and the human readable error is added to `err`\n@@ -1867,8 +1921,19 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n+\tswitch (fsync_object_files) {\n+\tcase FSYNC_OBJECT_FILES_OFF:\n+\t\tbreak;\n+\tcase FSYNC_OBJECT_FILES_ON:\n \t\tfsync_or_die(fd, \"loose object file\");\n+\t\tbreak;\n+\tcase FSYNC_OBJECT_FILES_BATCH:\n+\t\tfsync_loose_object_bulk_checkin(fd);\n+\t\tbreak;\n+\tdefault:\n+\t\tBUG(\"Invalid fsync_object_files mode.\");\n+\t}\n+\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/object-store.h b/object-store.h\nindex f8c883a5730..9bea14e7f3b 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -62,6 +62,19 @@ void add_to_alternates_file(const char *dir);\n  */\n void add_to_alternates_memory(const char *dir);\n \n+/*\n+ * Replace the current main object directory with the specified temporary\n+ * object directory. We make a copy of the former main object directory,\n+ * add it as an in-memory alternate, and return the copy so that it can\n+ * be restored via restore_main_odb.\n+ */\n+struct object_directory *set_temporary_main_odb(const char *dir);\n+\n+/*\n+ * Restore a previous ODB replaced by set_temporary_main_odb.\n+ */\n+void restore_main_odb(struct object_directory *odb);\n+\n /*\n  * Populate and return the loose object cache array corresponding to the\n  * given object ID.\n@@ -72,6 +85,9 @@ struct oidtree *odb_loose_cache(struct object_directory *odb,\n /* Empty the loose object cache for the specified object directory. */\n void odb_clear_loose_cache(struct object_directory *odb);\n \n+/* Clear and free the specified object directory */\n+void free_object_directory(struct object_directory *odb);\n+\n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n \tstruct packed_git *next;\ndiff --git a/object.c b/object.c\nindex 4e85955a941..98635bc4043 100644\n--- a/object.c\n+++ b/object.c\n@@ -513,7 +513,7 @@ struct raw_object_store *raw_object_store_new(void)\n \treturn o;\n }\n \n-static void free_object_directory(struct object_directory *odb)\n+void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex b8d880e3626..f027c49db4c 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -11,6 +11,7 @@\n struct tmp_objdir {\n \tstruct strbuf path;\n \tstruct strvec env;\n+\tstruct object_directory *prev_main_odb;\n };\n \n /*\n@@ -50,8 +51,12 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \t * freeing memory; it may cause a deadlock if the signal\n \t * arrived while libc's allocator lock is held.\n \t */\n-\tif (!on_signal)\n+\tif (!on_signal) {\n+\t\tif (t->prev_main_odb)\n+\t\t\trestore_main_odb(t->prev_main_odb);\n \t\ttmp_objdir_free(t);\n+\t}\n+\n \treturn err;\n }\n \n@@ -132,6 +137,7 @@ struct tmp_objdir *tmp_objdir_create(void)\n \tt = xmalloc(sizeof(*t));\n \tstrbuf_init(&t->path, 0);\n \tstrvec_init(&t->env);\n+\tt->prev_main_odb = NULL;\n \n \tstrbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n \n@@ -269,6 +275,11 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n \tif (!t)\n \t\treturn 0;\n \n+\tif (t->prev_main_odb) {\n+\t\trestore_main_odb(t->prev_main_odb);\n+\t\tt->prev_main_odb = NULL;\n+\t}\n+\n \tstrbuf_addbuf(&src, &t->path);\n \tstrbuf_addstr(&dst, get_object_directory());\n \n@@ -292,3 +303,10 @@ void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n {\n \tadd_to_alternates_memory(t->path.buf);\n }\n+\n+void tmp_objdir_replace_main_odb(struct tmp_objdir *t)\n+{\n+\tif (t->prev_main_odb)\n+\t\tBUG(\"the main object database is already replaced\");\n+\tt->prev_main_odb = set_temporary_main_odb(t->path.buf);\n+}\ndiff --git a/tmp-objdir.h b/tmp-objdir.h\nindex b1e45b4c75d..4b898add05b 100644\n--- a/tmp-objdir.h\n+++ b/tmp-objdir.h\n@@ -51,4 +51,10 @@ int tmp_objdir_destroy(struct tmp_objdir *);\n  */\n void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n \n+/*\n+ * Replaces the main object store in the current process with the temporary\n+ * object directory and makes the former main object store an alternate.\n+ */\n+void tmp_objdir_replace_main_odb(struct tmp_objdir *);\n+\n #endif /* TMP_OBJDIR_H */\ndiff --git a/wrapper.c b/wrapper.c\nindex 7c6586af321..bb4f9f043ce 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -540,6 +540,50 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync(fd);\n+#endif\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex d33e68f6abb..8f53953d4ab 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"437058","messageId":"92e18cedab0f9dccb6d64ab93c779d72be4d4cf1.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 5/8] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:26Z","receivedAt":"2021-09-24T23:53:40Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to synchronize with the update-index\nprocess after passing it a file path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 187203e8bb5..dc7368bb1ee 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1088,6 +1089,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \tthe_index.updated_skipworktree = 1;\n \n+\t/* we might be adding many objects to the object database */\n+\tplug_bulk_checkin();\n+\n \t/*\n \t * Custom copy of parse_options() because we want to handle\n \t * filename arguments as they come.\n@@ -1168,6 +1172,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/* by now we must have added all of the new objects */\n+\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"437060","messageId":"bdb99822f8c45a8b2855ee2ab38c4460e4b5e22e.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 4/8] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:25Z","receivedAt":"2021-09-24T23:53:41Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds a win32 implementation for fsync_no_flush that is\ncalled git_fsync. The 'NtFlushBuffersFileEx' function being called is\navailable since Windows 8. If the function is not available, we\nreturn -1 and Git falls back to doing a full fsync.\n\nThe operating system is told to flush data only without a hardware\nflush primitive. A later full fsync will cause the metadata log\nto be flushed and then the disk cache to be flushed on NTFS and\nReFS. Other filesystems will treat this as a full flush operation.\n\nI added a new file here for this system call so as not to conflict with\ndownstream changes in the git-for-windows repository related to fscache.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.h                      |  3 +++\n compat/win32/flush.c                | 28 ++++++++++++++++++++++++++++\n config.mak.uname                    |  2 ++\n contrib/buildsystems/CMakeLists.txt |  3 ++-\n wrapper.c                           |  4 ++++\n 5 files changed, 39 insertions(+), 1 deletion(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..75324c24ee7\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex e6d482fbcc6..34c93314a50 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -451,6 +451,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -626,6 +627,7 @@ ifneq (,$(findstring MINGW,$(uname_S)))\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 171b4124afe..b573a5ee122 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/wrapper.c b/wrapper.c\nindex bb4f9f043ce..1a1e2fba9c9 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -567,6 +567,10 @@ int git_fsync(int fd, enum fsync_action action)\n \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n #endif\n \n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n \t\terrno = ENOSYS;\n \t\treturn -1;\n \n-- \ngitgitgadget\n\n"},{"id":"437059","messageId":"e3c5a11f2252e0045801b3b5cd4e78e1b279e7c7.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 6/8] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:27Z","receivedAt":"2021-09-24T23:53:42Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling bulk-checkin when unpacking objects, we can take advantage\nof batched fsyncs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295ba..51eb4f7b531 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tplug_bulk_checkin();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"437061","messageId":"504bcc95c562c43b149d3dccf95f6a04761a9321.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 8/8] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:29Z","receivedAt":"2021-09-24T23:53:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 43 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 46 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 89 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..e93c08a2e70\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..c9fcd0c03eb\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n"},{"id":"437062","messageId":"385199354fafd65f7822fa13445a94afbe2a4dce.1632527609.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v6 7/8] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-24T23:53:28Z","receivedAt":"2021-09-24T23:53:46Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added due\nto it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 36 ++++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 20 ++++++++++++++++++++\n t/t3903-stash.sh       | 14 ++++++++++++++\n t/t5300-pack-object.sh | 30 +++++++++++++++++++-----------\n 4 files changed, 89 insertions(+), 11 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a7de4ca8512\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,36 @@\n+# Helper to create files with unique contents\n+\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with unique\n+#\t\t\t\t\t contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\tlocal counter=0\n+\ttest_tick\n+\tlocal basedata=$test_tick\n+\n+\n+\trm -rf $basedir\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\"\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1))\n+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex 4086e1ebbc9..36049a53ff7 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -7,6 +7,8 @@ test_description='Test of git add, including the -- option.'\n \n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -33,6 +35,24 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+test_expect_success 'git add: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files2 &&\n+\tfind fsync-files2 ! -type d -print | xargs git -c core.fsyncobjectfiles=batch update-index --add -- &&\n+\trm -f fsynced_files2 &&\n+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n+\ttest_line_count = 8 fsynced_files2 &&\n+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 873aa56e359..2fc819e5584 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n diff_cmp () {\n \tfor i in \"$1\" \"$2\"\n@@ -1293,6 +1294,19 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n \texpected=\"stash.useBuiltin support has been removed\" &&\n \ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex e13a8842075..38663dc1393 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -162,23 +162,23 @@ test_expect_success 'pack-objects with bogus arguments' '\n \n check_unpack () {\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $2 init --bare git2 &&\n+\t(\n+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <obj-list >current &&\n+\tcmp obj-list current\n }\n \n test_expect_success 'unpack without delta' '\n \tcheck_unpack test-1-${packname_1}\n '\n \n+test_expect_success 'unpack without delta (core.fsyncobjectfiles=batch)' '\n+\tcheck_unpack test-1-${packname_1} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with REF_DELTA' '\n \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n \tcheck_deltas stderr -gt 0\n@@ -188,6 +188,10 @@ test_expect_success 'unpack with REF_DELTA' '\n \tcheck_unpack test-2-${packname_2}\n '\n \n+test_expect_success 'unpack with REF_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-2-${packname_2} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with OFS_DELTA' '\n \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n \t\t\t<obj-list 2>stderr) &&\n@@ -198,6 +202,10 @@ test_expect_success 'unpack with OFS_DELTA' '\n \tcheck_unpack test-3-${packname_3}\n '\n \n+test_expect_success 'unpack with OFS_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-3-${packname_3} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'compare delta flavors' '\n \tperl -e '\\''\n \t\tdefined($_ = -s $_) or die for @ARGV;\n-- \ngitgitgadget\n\n"},{"id":"437065","messageId":"e8244ef1-dbd3-d56d-b9db-1e67114538fa@gmail.com","threadId":"56371","inReplyTo":"543ea3569342165363c1602ce36683a54dce7a0b.1632527609.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 3/8] core.fsyncobjectfiles: batched disk flushes","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-09-25T03:15:27Z","receivedAt":"2021-09-25T03:15:34Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 25/09/21 06.53, Neeraj Singh via GitGitGadget wrote:\n> At the end of the entire transaction when unplugging bulk checkin we:\n> 1. Issue an fsync against a dummy file to flush the hardware writeback\n>     cache, which should by now have processed the tmp-objdir writes.\n> 2. Rename all of the tmp-objdir files to their final names.\n> 3. When updating the index and/or refs, we assume that Git will issue\n>     another fsync internal to that operation. This is not the case today,\n>     but may be a good extension to those components.\n\nThe 'we' can be stripped because only point 1 and 2 that are \nsubject-inferred, so that subject needs to be explicitly mentioned, like:\n\n```\nAt the end of ... <snip>.:\n1. We issue an fsync ... <snip>.\n2. We rename ... <snip>.\n3. When ... <snip>, we assume <snip>. (stays same)\n```\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"437098","messageId":"CANQDOdfN9y+iVS=KuWkcVaFkOoeWPT7ACQWmmgxe=uxftduq+Q@mail.gmail.com","threadId":"56371","inReplyTo":"e8244ef1-dbd3-d56d-b9db-1e67114538fa@gmail.com","subject":"Re: [PATCH v6 3/8] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-27T00:27:27Z","receivedAt":"2021-09-27T00:27:41Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Fri, Sep 24, 2021 at 8:15 PM Bagas Sanjaya <bagasdotme@gmail.com> wrote:\n>\n> On 25/09/21 06.53, Neeraj Singh via GitGitGadget wrote:\n> > At the end of the entire transaction when unplugging bulk checkin we:\n> > 1. Issue an fsync against a dummy file to flush the hardware writeback\n> >     cache, which should by now have processed the tmp-objdir writes.\n> > 2. Rename all of the tmp-objdir files to their final names.\n> > 3. When updating the index and/or refs, we assume that Git will issue\n> >     another fsync internal to that operation. This is not the case today,\n> >     but may be a good extension to those components.\n>\n> The 'we' can be stripped because only point 1 and 2 that are\n> subject-inferred, so that subject needs to be explicitly mentioned, like:\n>\n> ```\n> At the end of ... <snip>.:\n> 1. We issue an fsync ... <snip>.\n> 2. We rename ... <snip>.\n> 3. When ... <snip>, we assume <snip>. (stays same)\n> ```\n\nI'll fix this in the github PR so that it will ride along with any\nother re-roll.\n\nThanks,\nNeeraj\n"},{"id":"437179","messageId":"xmqq1r59rde8.fsf@gitster.g","threadId":"56371","inReplyTo":"bdb99822f8c45a8b2855ee2ab38c4460e4b5e22e.1632527609.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 4/8] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-27T20:07:11Z","receivedAt":"2021-09-27T20:07:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/compat/mingw.h b/compat/mingw.h\n> index c9a52ad64a6..6074a3d3ced 100644\n> --- a/compat/mingw.h\n> +++ b/compat/mingw.h\n> @@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n>  #define getpagesize mingw_getpagesize\n>  #endif\n>  \n> +int win32_fsync_no_flush(int fd);\n> +#define fsync_no_flush win32_fsync_no_flush\n\n...\n\n> diff --git a/wrapper.c b/wrapper.c\n> index bb4f9f043ce..1a1e2fba9c9 100644\n> --- a/wrapper.c\n> +++ b/wrapper.c\n> @@ -567,6 +567,10 @@ int git_fsync(int fd, enum fsync_action action)\n>  \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n>  #endif\n>  \n> +#ifdef fsync_no_flush\n> +\t\treturn fsync_no_flush(fd);\n> +#endif\n> +\n>  \t\terrno = ENOSYS;\n>  \t\treturn -1;\n\nThis almost makes me wonder if we want to have a fallback\nimplementation of fsync_no_flush() that does\n\n   int fsync_no_flush(int unused)\n   {\n\terrno = ENOSYS;\n\treturn -1;\n   }\n\nwhen nobody (like Windows) define their own fsync_no_flush().  That\nway, this codepath does not have to have #ifdef/#endif here.\n\nThis function is already #ifdef ridden anyway, so reducing just one\ninstance may not make much difference, but since I noticed it ...\n\nThanks.\n"},{"id":"437191","messageId":"CANQDOdfxZP1GSb29LfcLQ2U84hRsgSq4kfzshPHFU8=o9+BSjg@mail.gmail.com","threadId":"56371","inReplyTo":"xmqq1r59rde8.fsf@gitster.g","subject":"Re: [PATCH v6 4/8] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-27T20:55:10Z","receivedAt":"2021-09-27T20:55:29Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Sep 27, 2021 at 1:07 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > diff --git a/compat/mingw.h b/compat/mingw.h\n> > index c9a52ad64a6..6074a3d3ced 100644\n> > --- a/compat/mingw.h\n> > +++ b/compat/mingw.h\n> > @@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n> >  #define getpagesize mingw_getpagesize\n> >  #endif\n> >\n> > +int win32_fsync_no_flush(int fd);\n> > +#define fsync_no_flush win32_fsync_no_flush\n>\n> ...\n>\n> > diff --git a/wrapper.c b/wrapper.c\n> > index bb4f9f043ce..1a1e2fba9c9 100644\n> > --- a/wrapper.c\n> > +++ b/wrapper.c\n> > @@ -567,6 +567,10 @@ int git_fsync(int fd, enum fsync_action action)\n> >                                                SYNC_FILE_RANGE_WAIT_AFTER);\n> >  #endif\n> >\n> > +#ifdef fsync_no_flush\n> > +             return fsync_no_flush(fd);\n> > +#endif\n> > +\n> >               errno = ENOSYS;\n> >               return -1;\n>\n> This almost makes me wonder if we want to have a fallback\n> implementation of fsync_no_flush() that does\n>\n>    int fsync_no_flush(int unused)\n>    {\n>         errno = ENOSYS;\n>         return -1;\n>    }\n>\n> when nobody (like Windows) define their own fsync_no_flush().  That\n> way, this codepath does not have to have #ifdef/#endif here.\n>\n> This function is already #ifdef ridden anyway, so reducing just one\n> instance may not make much difference, but since I noticed it ...\n>\n> Thanks.\n\nI'll make your suggested change on Github so that it will be available if\nwe do another re-roll.\n\nThanks,\nNereaj\n"},{"id":"437192","messageId":"CANQDOdd5T=6HR5hr1VOom0b+x0ps0yrQOeo8XD5QkSv2BuUNGA@mail.gmail.com","threadId":"56371","inReplyTo":"CANQDOdfxZP1GSb29LfcLQ2U84hRsgSq4kfzshPHFU8=o9+BSjg@mail.gmail.com","subject":"Re: [PATCH v6 4/8] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-27T21:03:09Z","receivedAt":"2021-09-27T21:03:38Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Sep 27, 2021 at 1:55 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> On Mon, Sep 27, 2021 at 1:07 PM Junio C Hamano <gitster@pobox.com> wrote:\n> >\n> > \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> >\n> > > diff --git a/compat/mingw.h b/compat/mingw.h\n> > > index c9a52ad64a6..6074a3d3ced 100644\n> > > --- a/compat/mingw.h\n> > > +++ b/compat/mingw.h\n> > > @@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n> > >  #define getpagesize mingw_getpagesize\n> > >  #endif\n> > >\n> > > +int win32_fsync_no_flush(int fd);\n> > > +#define fsync_no_flush win32_fsync_no_flush\n> >\n> > ...\n> >\n> > > diff --git a/wrapper.c b/wrapper.c\n> > > index bb4f9f043ce..1a1e2fba9c9 100644\n> > > --- a/wrapper.c\n> > > +++ b/wrapper.c\n> > > @@ -567,6 +567,10 @@ int git_fsync(int fd, enum fsync_action action)\n> > >                                                SYNC_FILE_RANGE_WAIT_AFTER);\n> > >  #endif\n> > >\n> > > +#ifdef fsync_no_flush\n> > > +             return fsync_no_flush(fd);\n> > > +#endif\n> > > +\n> > >               errno = ENOSYS;\n> > >               return -1;\n> >\n> > This almost makes me wonder if we want to have a fallback\n> > implementation of fsync_no_flush() that does\n> >\n> >    int fsync_no_flush(int unused)\n> >    {\n> >         errno = ENOSYS;\n> >         return -1;\n> >    }\n> >\n> > when nobody (like Windows) define their own fsync_no_flush().  That\n> > way, this codepath does not have to have #ifdef/#endif here.\n> >\n> > This function is already #ifdef ridden anyway, so reducing just one\n> > instance may not make much difference, but since I noticed it ...\n> >\n> > Thanks.\n>\n> I'll make your suggested change on Github so that it will be available if\n> we do another re-roll.\n>\n> Thanks,\n> Nereaj\n\nActually, while trying your suggestion, my conclusion is that we'd\neither have the\ninverse ifdef around the fsync_no_flush fallback or an #undef, or some\nother confusing\nstate.  The current ifdeffery is unpleasant to read but not too long\nand also pretty direct.\nWin32 has an extra level of indirection, but the unix platforms\nsyscalls are directly written\nin one place.\n\nThanks,\nNeeraj\n"},{"id":"437207","messageId":"xmqqzgrxmv8i.fsf@gitster.g","threadId":"56371","inReplyTo":"CANQDOdd5T=6HR5hr1VOom0b+x0ps0yrQOeo8XD5QkSv2BuUNGA@mail.gmail.com","subject":"Re: [PATCH v6 4/8] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-09-27T23:53:01Z","receivedAt":"2021-09-27T23:53:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> ....  The current ifdeffery is unpleasant to read but not too long\n> and also pretty direct.\n> Win32 has an extra level of indirection, but the unix platforms\n> syscalls are directly written\n> in one place.\n\nYes, that is exactly why I concluded that reducing just one instance\nwould not make that much difference ;-)\n\nThanks.\n"},{"id":"437381","messageId":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v6.git.git.1632527609.gitgitgadget@gmail.com","subject":"[PATCH v7 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:42Z","receivedAt":"2021-09-28T23:32:56Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far!\n\nThe patch is now at version 7: changes since v6:\n\n * Rebased onto current upstream master\n\n * Separate the tmp-objdir changes and move to the beginning of the series\n   so that Elijah Newren's similar changes can be merged.\n\n * Use some of Elijah's implementation for replacing the primary ODB. I was\n   doing some unnecessarily complex copying for no good reason.\n\n * Make the tmp objdir code use a name beginning with tmp_ and having a\n   operation-specific prefix.\n\n * Add git-prune support for removing a stale object directory.\n\nv5 was a bit of a dud, with some issues that I only noticed after\nsubmitting. v6 changes:\n\n * re-add Windows support\n * fix minor formatting issues\n * reset git author and commit dates which got messed up\n\nChanges since v4, all in response to review feedback from Ævar Arnfjörð\nBjarmason:\n\n * Update core.fsyncobjectfiles documentation to specify 'loose' objects and\n   to add a statement about not fsyncing parent directories.\n   \n   * I still don't want to make any promises on behalf of the Linux FS developers\n     in the documentation. However, according to [v4.1] and my understanding\n     of how XFS journals are documented to work, it looks like recent versions\n     of Linux running on XFS should be as safe as Windows or macOS in 'batch'\n     mode. I don't know about ext4, since it's not clear to me when metadata\n     updates are made visible to the journal.\n   \n\n * Rewrite the core batched fsync change to use the tmp-objdir lib. As Ævar\n   pointed out, this lets us access the added loose objects immediately,\n   rather than only after unplugging the bulk checkin. This is a hard\n   requirement in unpack-objects for resolving OBJ_REF_DELTA packed objects.\n   \n   * As a preparatory patch, the object-file code now doesn't do a rename if it's in a\n     tmp objdir (as determined by the quarantine environment variable).\n   \n   * I added support to the tmp-objdir lib to replace the 'main' writable odb.\n   \n   * Instead of using a lockfile for the final full fsync, we now use a new dummy\n     temp file. Doing that makes the below unpack-objects change easier.\n   \n\n * Add bulk-checkin support to unpack-objects, which is used in fetch and\n   push. In addition to making those operations faster, it allows us to\n   directly compare performance of packfiles against loose objects. Please\n   see [v4.2] for a measurement of 'git push' to a local upstream with\n   different numbers of unique new files.\n\n * Rename FSYNC_OBJECT_FILES_MODE to fsync_object_files_mode.\n\n * Remove comment with link to NtFlushBuffersFileEx documentation.\n\n * Make t/lib-unique-files.sh a bit cleaner. We are still creating unique\n   contents, but now this uses test_tick, so it should be deterministic from\n   run to run.\n\n * Ensure there are tests for all of the modified commands. Make the\n   unpack-objects tests validate that the unpacked objects are really\n   available in the ODB.\n\nReferences for v4: [v4.1]\nhttps://lore.kernel.org/linux-fsdevel/20190419072938.31320-1-amir73il@gmail.com/#t\n\n[v4.2]\nhttps://docs.google.com/spreadsheets/d/1uxMBkEXFFnQ1Y3lXKqcKpw6Mq44BzhpCAcPex14T-QQ/edit#gid=1898936117\n\nChanges since v3:\n\n * Fix core.fsyncobjectfiles option parsing as suggested by Junio: We now\n   accept no value to mean \"true\" and we require 'batch' to be lowercase.\n\n * Leave the default fsync mode as 'false'. Git for windows can change its\n   default when this series makes it over to that fork.\n\n * Use a switch statement in git_fsync, as suggested by Junio.\n\n * Add regression test cases for core.fsyncobjectfiles=batch. This should\n   keep the batch functionality basically working in upstream git even if\n   few users adopt batch mode initially. I expect git-for-windows will\n   provide a good baking area for the new mode.\n\nNeeraj Singh (9):\n  object-file.c: do not rename in a temp odb\n  tmp-objdir: new API for creating temporary writable databases\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  core.fsyncobjectfiles: add windows support for batch mode\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: tests for batch mode\n  core.fsyncobjectfiles: performance tests for add and stash\n\n Documentation/config/core.txt       |  29 ++++++--\n Makefile                            |   6 ++\n builtin/prune.c                     |  22 ++++--\n builtin/receive-pack.c              |   2 +-\n builtin/unpack-objects.c            |   3 +\n builtin/update-index.c              |   6 ++\n bulk-checkin.c                      |  92 +++++++++++++++++++++---\n bulk-checkin.h                      |   2 +\n cache.h                             |   8 ++-\n compat/mingw.h                      |   3 +\n compat/win32/flush.c                |  28 ++++++++\n config.c                            |   7 +-\n config.mak.uname                    |   3 +\n configure.ac                        |   8 +++\n contrib/buildsystems/CMakeLists.txt |   3 +-\n environment.c                       |   6 +-\n git-compat-util.h                   |   7 ++\n object-file.c                       | 106 +++++++++++++++++++++++-----\n object-store.h                      |  25 +++++++\n object.c                            |   2 +-\n repository.c                        |   2 +\n repository.h                        |   1 +\n t/lib-unique-files.sh               |  36 ++++++++++\n t/perf/p3700-add.sh                 |  43 +++++++++++\n t/perf/p3900-stash.sh               |  46 ++++++++++++\n t/t3700-add.sh                      |  20 ++++++\n t/t3903-stash.sh                    |  14 ++++\n t/t5300-pack-object.sh              |  30 +++++---\n tmp-objdir.c                        |  30 +++++++-\n tmp-objdir.h                        |  14 +++-\n wrapper.c                           |  48 +++++++++++++\n write-or-die.c                      |   2 +-\n 32 files changed, 592 insertions(+), 62 deletions(-)\n create mode 100644 compat/win32/flush.c\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: cefe983a320c03d7843ac78e73bd513a27806845\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v7\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v7\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v6:\n\n  1:  e4081f81f6a =  1:  6e65f68fd6d object-file.c: do not rename in a temp odb\n  -:  ----------- >  2:  6ce72a709a1 tmp-objdir: new API for creating temporary writable databases\n  2:  ebba65e040c !  3:  c272f8776fa bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n     @@ bulk-checkin.c: static struct bulk_checkin_state {\n      -} state;\n      +} bulk_checkin_state;\n       \n     - static void finish_bulk_checkin(struct bulk_checkin_state *state)\n     - {\n     + static void finish_tmp_packfile(struct strbuf *basename,\n     + \t\t\t\tconst char *pack_tmp_name,\n      @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n       \t\t       int fd, size_t size, enum object_type type,\n       \t\t       const char *path, unsigned flags)\n  3:  543ea356934 !  4:  55556bb3e90 core.fsyncobjectfiles: batched disk flushes\n     @@ Commit message\n          batches up hardware flushes. It hooks into the bulk-checkin plugging and\n          unplugging functionality and takes advantage of tmp-objdir.\n      \n     -    When the new mode is enabled we do the following for each new object:\n     +    When the new mode is enabled do the following for each new object:\n          1. Create the object in a tmp-objdir.\n          2. Issue a pagecache writeback request and wait for it to complete.\n      \n     -    At the end of the entire transaction when unplugging bulk checkin we:\n     +    At the end of the entire transaction when unplugging bulk checkin:\n          1. Issue an fsync against a dummy file to flush the hardware writeback\n             cache, which should by now have processed the tmp-objdir writes.\n          2. Rename all of the tmp-objdir files to their final names.\n     @@ Commit message\n             but may be a good extension to those components.\n      \n          On a filesystem with a singular journal that is updated during name\n     -    operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS we\n     -    would expect the fsync to trigger a journal writeout so that this\n     +    operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\n     +    we would expect the fsync to trigger a journal writeout so that this\n          sequence is enough to ensure that the user's data is durable by the time\n          the git command returns.\n      \n     @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n      +\t */\n      +\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n      +\t    !the_repository->objects->odb->is_temp) {\n     -+\t\tbulk_fsync_objdir = tmp_objdir_create();\n     ++\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n      +\t\tif (!bulk_fsync_objdir)\n      +\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n      +\n     -+\t\ttmp_objdir_replace_main_odb(bulk_fsync_objdir);\n     ++\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n      +\t}\n      +\n       \tbulk_checkin_plugged = 1;\n     @@ configure.ac: AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n       GIT_CHECK_FUNC(setitimer,\n      \n       ## environment.c ##\n     -@@ environment.c: const char *git_hooks_path;\n     +@@ environment.c: const char *git_attributes_file;\n     + const char *git_hooks_path;\n       int zlib_compression_level = Z_BEST_SPEED;\n     - int core_compression_level;\n       int pack_compression_level = Z_DEFAULT_COMPRESSION;\n      -int fsync_object_files;\n      +enum fsync_object_files_mode fsync_object_files;\n     @@ git-compat-util.h: __attribute__((format (printf, 1, 2))) NORETURN\n        * Returns 0 on success, which includes trying to unlink an object that does\n      \n       ## object-file.c ##\n     -@@ object-file.c: void add_to_alternates_memory(const char *reference)\n     - \t\t\t     '\\n', NULL, 0);\n     - }\n     - \n     -+struct object_directory *set_temporary_main_odb(const char *dir)\n     -+{\n     -+\tstruct object_directory *main_odb, *new_odb, *old_next;\n     -+\n     -+\t/*\n     -+\t * Make sure alternates are initialized, or else our entry may be\n     -+\t * overwritten when they are.\n     -+\t */\n     -+\tprepare_alt_odb(the_repository);\n     -+\n     -+\t/* Copy the existing object directory and make it an alternate. */\n     -+\tmain_odb = the_repository->objects->odb;\n     -+\tnew_odb = xmalloc(sizeof(*new_odb));\n     -+\t*new_odb = *main_odb;\n     -+\t*the_repository->objects->odb_tail = new_odb;\n     -+\tthe_repository->objects->odb_tail = &(new_odb->next);\n     -+\tnew_odb->next = NULL;\n     -+\n     -+\t/*\n     -+\t * Reinitialize the main odb with the specified path, being careful\n     -+\t * to keep the next pointer value.\n     -+\t */\n     -+\told_next = main_odb->next;\n     -+\tmemset(main_odb, 0, sizeof(*main_odb));\n     -+\tmain_odb->next = old_next;\n     -+\tmain_odb->is_temp = 1;\n     -+\tmain_odb->path = xstrdup(dir);\n     -+\treturn new_odb;\n     -+}\n     -+\n     -+void restore_main_odb(struct object_directory *odb)\n     -+{\n     -+\tstruct object_directory **prev, *main_odb;\n     -+\n     -+\t/* Unlink the saved previous main ODB from the list. */\n     -+\tprev = &the_repository->objects->odb->next;\n     -+\tassert(*prev);\n     -+\twhile (*prev != odb) {\n     -+\t\tprev = &(*prev)->next;\n     -+\t}\n     -+\t*prev = odb->next;\n     -+\tif (*prev == NULL)\n     -+\t\tthe_repository->objects->odb_tail = prev;\n     -+\n     -+\t/*\n     -+\t * Restore the data from the old main odb, being careful to\n     -+\t * keep the next pointer value\n     -+\t */\n     -+\tmain_odb = the_repository->objects->odb;\n     -+\tSWAP(*main_odb, *odb);\n     -+\tmain_odb->next = odb->next;\n     -+\tfree_object_directory(odb);\n     -+}\n     -+\n     - /*\n     -  * Compute the exact path an alternate is at and returns it. In case of\n     -  * error NULL is returned and the human readable error is added to `err`\n      @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n     - /* Finalize a file on disk, and close it. */\n       static void close_loose_object(int fd)\n       {\n     --\tif (fsync_object_files)\n     -+\tswitch (fsync_object_files) {\n     -+\tcase FSYNC_OBJECT_FILES_OFF:\n     -+\t\tbreak;\n     -+\tcase FSYNC_OBJECT_FILES_ON:\n     - \t\tfsync_or_die(fd, \"loose object file\");\n     -+\t\tbreak;\n     -+\tcase FSYNC_OBJECT_FILES_BATCH:\n     -+\t\tfsync_loose_object_bulk_checkin(fd);\n     -+\t\tbreak;\n     -+\tdefault:\n     -+\t\tBUG(\"Invalid fsync_object_files mode.\");\n     -+\t}\n     -+\n     - \tif (close(fd) != 0)\n     - \t\tdie_errno(_(\"error when closing loose object file\"));\n     - }\n     -\n     - ## object-store.h ##\n     -@@ object-store.h: void add_to_alternates_file(const char *dir);\n     -  */\n     - void add_to_alternates_memory(const char *dir);\n     - \n     -+/*\n     -+ * Replace the current main object directory with the specified temporary\n     -+ * object directory. We make a copy of the former main object directory,\n     -+ * add it as an in-memory alternate, and return the copy so that it can\n     -+ * be restored via restore_main_odb.\n     -+ */\n     -+struct object_directory *set_temporary_main_odb(const char *dir);\n     -+\n     -+/*\n     -+ * Restore a previous ODB replaced by set_temporary_main_odb.\n     -+ */\n     -+void restore_main_odb(struct object_directory *odb);\n     -+\n     - /*\n     -  * Populate and return the loose object cache array corresponding to the\n     -  * given object ID.\n     -@@ object-store.h: struct oidtree *odb_loose_cache(struct object_directory *odb,\n     - /* Empty the loose object cache for the specified object directory. */\n     - void odb_clear_loose_cache(struct object_directory *odb);\n     - \n     -+/* Clear and free the specified object directory */\n     -+void free_object_directory(struct object_directory *odb);\n     -+\n     - struct packed_git {\n     - \tstruct hashmap_entry packmap_ent;\n     - \tstruct packed_git *next;\n     -\n     - ## object.c ##\n     -@@ object.c: struct raw_object_store *raw_object_store_new(void)\n     - \treturn o;\n     - }\n     + \tif (!the_repository->objects->odb->will_destroy) {\n     +-\t\tif (fsync_object_files)\n     ++\t\tswitch (fsync_object_files) {\n     ++\t\tcase FSYNC_OBJECT_FILES_OFF:\n     ++\t\t\tbreak;\n     ++\t\tcase FSYNC_OBJECT_FILES_ON:\n     + \t\t\tfsync_or_die(fd, \"loose object file\");\n     ++\t\t\tbreak;\n     ++\t\tcase FSYNC_OBJECT_FILES_BATCH:\n     ++\t\t\tfsync_loose_object_bulk_checkin(fd);\n     ++\t\t\tbreak;\n     ++\t\tdefault:\n     ++\t\t\tBUG(\"Invalid fsync_object_files mode.\");\n     ++\t\t}\n     + \t}\n       \n     --static void free_object_directory(struct object_directory *odb)\n     -+void free_object_directory(struct object_directory *odb)\n     - {\n     - \tfree(odb->path);\n     - \todb_clear_loose_cache(odb);\n     + \tif (close(fd) != 0)\n      \n       ## tmp-objdir.c ##\n     -@@\n     - struct tmp_objdir {\n     - \tstruct strbuf path;\n     - \tstruct strvec env;\n     -+\tstruct object_directory *prev_main_odb;\n     - };\n     - \n     - /*\n     -@@ tmp-objdir.c: static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n     - \t * freeing memory; it may cause a deadlock if the signal\n     - \t * arrived while libc's allocator lock is held.\n     - \t */\n     --\tif (!on_signal)\n     -+\tif (!on_signal) {\n     -+\t\tif (t->prev_main_odb)\n     -+\t\t\trestore_main_odb(t->prev_main_odb);\n     - \t\ttmp_objdir_free(t);\n     -+\t}\n     -+\n     - \treturn err;\n     - }\n     - \n     -@@ tmp-objdir.c: struct tmp_objdir *tmp_objdir_create(void)\n     - \tt = xmalloc(sizeof(*t));\n     - \tstrbuf_init(&t->path, 0);\n     - \tstrvec_init(&t->env);\n     -+\tt->prev_main_odb = NULL;\n     - \n     - \tstrbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n     - \n      @@ tmp-objdir.c: int tmp_objdir_migrate(struct tmp_objdir *t)\n       \tif (!t)\n       \t\treturn 0;\n       \n     -+\tif (t->prev_main_odb) {\n     -+\t\trestore_main_odb(t->prev_main_odb);\n     -+\t\tt->prev_main_odb = NULL;\n     -+\t}\n     -+\n     - \tstrbuf_addbuf(&src, &t->path);\n     - \tstrbuf_addstr(&dst, get_object_directory());\n     - \n     -@@ tmp-objdir.c: void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n     - {\n     - \tadd_to_alternates_memory(t->path.buf);\n     - }\n     -+\n     -+void tmp_objdir_replace_main_odb(struct tmp_objdir *t)\n     -+{\n     -+\tif (t->prev_main_odb)\n     -+\t\tBUG(\"the main object database is already replaced\");\n     -+\tt->prev_main_odb = set_temporary_main_odb(t->path.buf);\n     -+}\n     -\n     - ## tmp-objdir.h ##\n     -@@ tmp-objdir.h: int tmp_objdir_destroy(struct tmp_objdir *);\n     -  */\n     - void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n     - \n     -+/*\n     -+ * Replaces the main object store in the current process with the temporary\n     -+ * object directory and makes the former main object store an alternate.\n     -+ */\n     -+void tmp_objdir_replace_main_odb(struct tmp_objdir *);\n     -+\n     - #endif /* TMP_OBJDIR_H */\n     +-\n     +-\n     + \tif (t->prev_odb) {\n     + \t\tif (the_repository->objects->odb->will_destroy)\n     + \t\t\tBUG(\"migrating and ODB that was marked for destruction\");\n      \n       ## wrapper.c ##\n      @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n  4:  bdb99822f8c =  5:  6c33e79d6f0 core.fsyncobjectfiles: add windows support for batch mode\n  5:  92e18cedab0 =  6:  09dbff1004e update-index: use the bulk-checkin infrastructure\n  6:  e3c5a11f225 =  7:  1eced9f9f9a unpack-objects: use the bulk-checkin infrastructure\n  7:  385199354fa =  8:  7aaa08d5f5f core.fsyncobjectfiles: tests for batch mode\n  8:  504bcc95c56 =  9:  ff286fb461a core.fsyncobjectfiles: performance tests for add and stash\n\n-- \ngitgitgadget\n"},{"id":"437379","messageId":"6e65f68fd6d4d90b0a7bca2e2e57ace9ad749266.1632871971.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 1/9] object-file.c: do not rename in a temp odb","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:43Z","receivedAt":"2021-09-28T23:32:57Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIf a temporary ODB is active, as determined by GIT_QUARANTINE_PATH\nbeing set, create object files with their final names. This avoids\nan extra rename beyond what is needed to merge the temporary ODB in\ntmp_objdir_migrate.\n\nCreating an object file with the expected final name should be okay\nsince the git process writing to the temporary object store is the\nonly writer, and it only invokes write_loose_object/create_object_file\nafter checking that the object doesn't exist.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n environment.c  |  4 ++++\n object-file.c  | 51 ++++++++++++++++++++++++++++++++++----------------\n object-store.h |  6 ++++++\n repository.c   |  2 ++\n repository.h   |  1 +\n 5 files changed, 48 insertions(+), 16 deletions(-)\n\ndiff --git a/environment.c b/environment.c\nindex b4ba4fa22db..30fca67e6d6 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -176,6 +176,10 @@ void setup_git_env(const char *git_dir)\n \targs.graft_file = getenv_safe(&to_free, GRAFT_ENVIRONMENT);\n \targs.index_file = getenv_safe(&to_free, INDEX_ENVIRONMENT);\n \targs.alternate_db = getenv_safe(&to_free, ALTERNATE_DB_ENVIRONMENT);\n+\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n+\t\targs.object_dir_is_temp = 1;\n+\t}\n+\n \trepo_set_gitdir(the_repository, git_dir, &args);\n \tstrvec_clear(&to_free);\n \ndiff --git a/object-file.c b/object-file.c\nindex be4f94ecf3b..49c53f801f7 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1826,12 +1826,17 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n }\n \n /*\n- * Move the just written object into its final resting place.\n+ * Move the just written object into its final resting place,\n+ * unless it is already there, as indicated by an empty string for\n+ * tmpfile.\n  */\n int finalize_object_file(const char *tmpfile, const char *filename)\n {\n \tint ret = 0;\n \n+\tif (!*tmpfile)\n+\t\tgoto out;\n+\n \tif (object_creation_mode == OBJECT_CREATION_USES_RENAMES)\n \t\tgoto try_rename;\n \telse if (link(tmpfile, filename))\n@@ -1904,21 +1909,37 @@ static inline int directory_size(const char *filename)\n }\n \n /*\n- * This creates a temporary file in the same directory as the final\n- * 'filename'\n+ * This creates a loose object file for the specified object id.\n+ * If we're working in a temporary object directory, the file is\n+ * created with its final filename, otherwise it is created with\n+ * a temporary name and renamed by finalize_object_file.\n+ * If no rename is required, an empty string is returned in tmp.\n  *\n  * We want to avoid cross-directory filename renames, because those\n  * can have problems on various filesystems (FAT, NFS, Coda).\n  */\n-static int create_tmpfile(struct strbuf *tmp, const char *filename)\n+static int create_objfile(const struct object_id *oid, struct strbuf *tmp,\n+\t\t\t  struct strbuf *filename)\n {\n-\tint fd, dirlen = directory_size(filename);\n+\tint fd, dirlen, is_retrying = 0;\n+\tconst char *object_name;\n+\tstatic const int object_mode = 0444;\n \n+\tloose_object_path(the_repository, filename, oid);\n+\tdirlen = directory_size(filename->buf);\n+\n+retry_create:\n \tstrbuf_reset(tmp);\n-\tstrbuf_add(tmp, filename, dirlen);\n-\tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n-\tfd = git_mkstemp_mode(tmp->buf, 0444);\n-\tif (fd < 0 && dirlen && errno == ENOENT) {\n+\tif (!the_repository->objects->odb->is_temp) {\n+\t\tstrbuf_add(tmp, filename->buf, dirlen);\n+\t\tobject_name = \"tmp_obj_XXXXXX\";\n+\t\tstrbuf_addstr(tmp, object_name);\n+\t\tfd = git_mkstemp_mode(tmp->buf, object_mode);\n+\t} else {\n+\t\tfd = open(filename->buf, O_CREAT | O_EXCL | O_RDWR, object_mode);\n+\t}\n+\n+\tif (fd < 0 && dirlen && errno == ENOENT && !is_retrying) {\n \t\t/*\n \t\t * Make sure the directory exists; note that the contents\n \t\t * of the buffer are undefined after mkstemp returns an\n@@ -1926,15 +1947,15 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \t\t * scratch.\n \t\t */\n \t\tstrbuf_reset(tmp);\n-\t\tstrbuf_add(tmp, filename, dirlen - 1);\n+\t\tstrbuf_add(tmp, filename->buf, dirlen - 1);\n \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n \t\t\treturn -1;\n \t\tif (adjust_shared_perm(tmp->buf))\n \t\t\treturn -1;\n \n \t\t/* Try again */\n-\t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n-\t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n+\t\tis_retrying = 1;\n+\t\tgoto retry_create;\n \t}\n \treturn fd;\n }\n@@ -1951,14 +1972,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n-\tloose_object_path(the_repository, &filename, oid);\n-\n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n+\tfd = create_objfile(oid, &tmp_file, &filename);\n \tif (fd < 0) {\n \t\tif (errno == EACCES)\n \t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n \t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\t\t\treturn error_errno(_(\"unable to create object file\"));\n \t}\n \n \t/* Set it up */\ndiff --git a/object-store.h b/object-store.h\nindex c5130d8baea..551639f173d 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -27,6 +27,12 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/*\n+\t * This is a temporary object store, so there is no need to\n+\t * create new objects via rename.\n+\t */\n+\tint is_temp;\n+\n \t/*\n \t * Path to the alternative object store. If this is a relative path,\n \t * it is relative to the current working directory.\ndiff --git a/repository.c b/repository.c\nindex 710a3b4bf87..75966153b75 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -80,6 +80,8 @@ void repo_set_gitdir(struct repository *repo,\n \texpand_base_dir(&repo->objects->odb->path, o->object_dir,\n \t\t\trepo->commondir, \"objects\");\n \n+\trepo->objects->odb->is_temp = o->object_dir_is_temp;\n+\n \tfree(repo->objects->alternate_db);\n \trepo->objects->alternate_db = xstrdup_or_null(o->alternate_db);\n \texpand_base_dir(&repo->graft_file, o->graft_file,\ndiff --git a/repository.h b/repository.h\nindex 3740c93bc0f..d3711367a6f 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -162,6 +162,7 @@ struct set_gitdir_args {\n \tconst char *graft_file;\n \tconst char *index_file;\n \tconst char *alternate_db;\n+\tint object_dir_is_temp;\n };\n \n void repo_set_gitdir(struct repository *repo, const char *root,\n-- \ngitgitgadget\n\n"},{"id":"437380","messageId":"c272f8776fa26a1f52e04e45886643bbce3246bc.1632871972.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 3/9] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:45Z","receivedAt":"2021-09-28T23:32:58Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..6ae18401e04 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -277,21 +277,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"437382","messageId":"6ce72a709a11686b9082439a257fd5f58e5eb0f7.1632871971.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 2/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:44Z","receivedAt":"2021-09-28T23:33:01Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis patch is based on work by Elijah Newren. Any bugs however are my\nown.\n\nThe tmp_objdir API provides the ability to create temporary object\ndirectories, but was designed with the goal of having subprocesses\naccess these object stores, followed by the main process migrating\nobjects from it to the main object store or just deleting it.  The\nsubprocesses would view it as their primary datastore and write to it.\n\nHere we add the tmp_objdir_replace_primary_odb function that replaces\nthe current process's writable \"main\" object directory with the\nspecified one. The previous main object directory is restored in either\ntmp_objdir_migrate or tmp_objdir_destroy.\n\nFor the --remerge-diff usecase, add a new `will_destroy` flag in `struct\nobject_database` to mark ephemeral object databases that do not require\nfsync durability.\n\nAdd 'git prune' support for removing temporary object databases, and\nmake sure that they have a name starting with tmp_ and containing an\noperation-specific name.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/prune.c        | 22 +++++++++++++++++----\n builtin/receive-pack.c |  2 +-\n object-file.c          | 45 ++++++++++++++++++++++++++++++++++++++++--\n object-store.h         | 21 +++++++++++++++++++-\n object.c               |  2 +-\n tmp-objdir.c           | 32 +++++++++++++++++++++++++++---\n tmp-objdir.h           | 14 ++++++++++---\n 7 files changed, 123 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/prune.c b/builtin/prune.c\nindex 02c6ab7cbaa..9c72ecf5a58 100644\n--- a/builtin/prune.c\n+++ b/builtin/prune.c\n@@ -18,6 +18,7 @@ static int show_only;\n static int verbose;\n static timestamp_t expire;\n static int show_progress = -1;\n+static struct strbuf remove_dir_buf = STRBUF_INIT;\n \n static int prune_tmp_file(const char *fullpath)\n {\n@@ -26,10 +27,19 @@ static int prune_tmp_file(const char *fullpath)\n \t\treturn error(\"Could not stat '%s'\", fullpath);\n \tif (st.st_mtime > expire)\n \t\treturn 0;\n-\tif (show_only || verbose)\n-\t\tprintf(\"Removing stale temporary file %s\\n\", fullpath);\n-\tif (!show_only)\n-\t\tunlink_or_warn(fullpath);\n+\tif (S_ISDIR(st.st_mode)) {\n+\t\tif (show_only || verbose)\n+\t\t\tprintf(\"Removing stale temporary directory %s\\n\", fullpath);\n+\t\tif (!show_only) {\n+\t\t\tstrbuf_addstr(&remove_dir_buf, fullpath);\n+\t\t\tremove_dir_recursively(&remove_dir_buf, 0);\n+\t\t}\n+\t} else {\n+\t\tif (show_only || verbose)\n+\t\t\tprintf(\"Removing stale temporary file %s\\n\", fullpath);\n+\t\tif (!show_only)\n+\t\t\tunlink_or_warn(fullpath);\n+\t}\n \treturn 0;\n }\n \n@@ -97,6 +107,9 @@ static int prune_cruft(const char *basename, const char *path, void *data)\n \n static int prune_subdir(unsigned int nr, const char *path, void *data)\n {\n+\tif (verbose)\n+\t\tprintf(\"Removing directory %s\\n\", path);\n+\n \tif (!show_only)\n \t\trmdir(path);\n \treturn 0;\n@@ -185,5 +198,6 @@ int cmd_prune(int argc, const char **argv, const char *prefix)\n \t\tprune_shallow(show_only ? PRUNE_SHOW_ONLY : 0);\n \t}\n \n+\tstrbuf_release(&remove_dir_buf);\n \treturn 0;\n }\ndiff --git a/builtin/receive-pack.c b/builtin/receive-pack.c\nindex 48960a9575b..418a42ca069 100644\n--- a/builtin/receive-pack.c\n+++ b/builtin/receive-pack.c\n@@ -2208,7 +2208,7 @@ static const char *unpack(int err_fd, struct shallow_info *si)\n \t\tstrvec_push(&child.args, alt_shallow_file);\n \t}\n \n-\ttmp_objdir = tmp_objdir_create();\n+\ttmp_objdir = tmp_objdir_create(\"incoming\");\n \tif (!tmp_objdir) {\n \t\tif (err_fd > 0)\n \t\t\tclose(err_fd);\ndiff --git a/object-file.c b/object-file.c\nindex 49c53f801f7..1a3ad558c45 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -751,6 +751,44 @@ void add_to_alternates_memory(const char *reference)\n \t\t\t     '\\n', NULL, 0);\n }\n \n+struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy)\n+{\n+\tstruct object_directory *new_odb;\n+\n+\t/*\n+\t * Make sure alternates are initialized, or else our entry may be\n+\t * overwritten when they are.\n+\t */\n+\tprepare_alt_odb(the_repository);\n+\n+\t/*\n+\t * Make a new primary odb and link the old primary ODB in as an\n+\t * alternate\n+\t */\n+\tnew_odb = xcalloc(1, sizeof(*new_odb));\n+\tnew_odb->path = xstrdup(dir);\n+\tnew_odb->is_temp = 1;\n+\tnew_odb->will_destroy = will_destroy;\n+\tnew_odb->next = the_repository->objects->odb;\n+\tthe_repository->objects->odb = new_odb;\n+\treturn new_odb->next;\n+}\n+\n+void restore_primary_odb(struct object_directory *restore_odb, const char *old_path)\n+{\n+\tstruct object_directory *cur_odb = the_repository->objects->odb;\n+\n+\tif (strcmp(old_path, cur_odb->path))\n+\t\tBUG(\"expected %s as primary object store; found %s\",\n+\t\t    old_path, cur_odb->path);\n+\n+\tif (cur_odb->next != restore_odb)\n+\t\tBUG(\"we expect the old primary object store to be the first alternate\");\n+\n+\tthe_repository->objects->odb = restore_odb;\n+\tfree_object_directory(cur_odb);\n+}\n+\n /*\n  * Compute the exact path an alternate is at and returns it. In case of\n  * error NULL is returned and the human readable error is added to `err`\n@@ -1893,8 +1931,11 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\tif (!the_repository->objects->odb->will_destroy) {\n+\t\tif (fsync_object_files)\n+\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/object-store.h b/object-store.h\nindex 551639f173d..5bc9da6634e 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -31,7 +31,12 @@ struct object_directory {\n \t * This is a temporary object store, so there is no need to\n \t * create new objects via rename.\n \t */\n-\tint is_temp;\n+\tint is_temp : 8;\n+\n+\t/*\n+\t * This object store is ephemeral, so there is no need to fsync.\n+\t */\n+\tint will_destroy : 8;\n \n \t/*\n \t * Path to the alternative object store. If this is a relative path,\n@@ -64,6 +69,17 @@ void add_to_alternates_file(const char *dir);\n  */\n void add_to_alternates_memory(const char *dir);\n \n+/*\n+ * Replace the current writable object directory with the specified temporary\n+ * object directory; returns the former primary object directory.\n+ */\n+struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy);\n+\n+/*\n+ * Restore a previous ODB replaced by set_temporary_main_odb.\n+ */\n+void restore_primary_odb(struct object_directory *restore_odb, const char *old_path);\n+\n /*\n  * Populate and return the loose object cache array corresponding to the\n  * given object ID.\n@@ -74,6 +90,9 @@ struct oidtree *odb_loose_cache(struct object_directory *odb,\n /* Empty the loose object cache for the specified object directory. */\n void odb_clear_loose_cache(struct object_directory *odb);\n \n+/* Clear and free the specified object directory */\n+void free_object_directory(struct object_directory *odb);\n+\n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n \tstruct packed_git *next;\ndiff --git a/object.c b/object.c\nindex 4e85955a941..98635bc4043 100644\n--- a/object.c\n+++ b/object.c\n@@ -513,7 +513,7 @@ struct raw_object_store *raw_object_store_new(void)\n \treturn o;\n }\n \n-static void free_object_directory(struct object_directory *odb)\n+void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex b8d880e3626..366ffe28511 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -11,6 +11,7 @@\n struct tmp_objdir {\n \tstruct strbuf path;\n \tstruct strvec env;\n+\tstruct object_directory *prev_odb;\n };\n \n /*\n@@ -38,6 +39,9 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \tif (t == the_tmp_objdir)\n \t\tthe_tmp_objdir = NULL;\n \n+\tif (!on_signal && t->prev_odb)\n+\t\trestore_primary_odb(t->prev_odb, t->path.buf);\n+\n \t/*\n \t * This may use malloc via strbuf_grow(), but we should\n \t * have pre-grown t->path sufficiently so that this\n@@ -52,6 +56,7 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \t */\n \tif (!on_signal)\n \t\ttmp_objdir_free(t);\n+\n \treturn err;\n }\n \n@@ -121,7 +126,7 @@ static int setup_tmp_objdir(const char *root)\n \treturn ret;\n }\n \n-struct tmp_objdir *tmp_objdir_create(void)\n+struct tmp_objdir *tmp_objdir_create(const char *prefix)\n {\n \tstatic int installed_handlers;\n \tstruct tmp_objdir *t;\n@@ -129,11 +134,16 @@ struct tmp_objdir *tmp_objdir_create(void)\n \tif (the_tmp_objdir)\n \t\tBUG(\"only one tmp_objdir can be used at a time\");\n \n-\tt = xmalloc(sizeof(*t));\n+\tt = xcalloc(1, sizeof(*t));\n \tstrbuf_init(&t->path, 0);\n \tstrvec_init(&t->env);\n \n-\tstrbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n+\t/*\n+\t * Use a string starting with tmp_ so that the builtin/prune.c code\n+\t * can recognize any stale objdirs left behind by a crash and delete\n+\t * them.\n+\t */\n+\tstrbuf_addf(&t->path, \"%s/tmp_objdir-%s-XXXXXX\", get_object_directory(), prefix);\n \n \t/*\n \t * Grow the strbuf beyond any filename we expect to be placed in it.\n@@ -269,6 +279,15 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n \tif (!t)\n \t\treturn 0;\n \n+\n+\n+\tif (t->prev_odb) {\n+\t\tif (the_repository->objects->odb->will_destroy)\n+\t\t\tBUG(\"migrating and ODB that was marked for destruction\");\n+\t\trestore_primary_odb(t->prev_odb, t->path.buf);\n+\t\tt->prev_odb = NULL;\n+\t}\n+\n \tstrbuf_addbuf(&src, &t->path);\n \tstrbuf_addstr(&dst, get_object_directory());\n \n@@ -292,3 +311,10 @@ void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n {\n \tadd_to_alternates_memory(t->path.buf);\n }\n+\n+void tmp_objdir_replace_primary_odb(struct tmp_objdir *t, int will_destroy)\n+{\n+\tif (t->prev_odb)\n+\t\tBUG(\"the primary object database is already replaced\");\n+\tt->prev_odb = set_temporary_primary_odb(t->path.buf, will_destroy);\n+}\ndiff --git a/tmp-objdir.h b/tmp-objdir.h\nindex b1e45b4c75d..75754cbfba6 100644\n--- a/tmp-objdir.h\n+++ b/tmp-objdir.h\n@@ -10,7 +10,7 @@\n  *\n  * Example:\n  *\n- *\tstruct tmp_objdir *t = tmp_objdir_create();\n+ *\tstruct tmp_objdir *t = tmp_objdir_create(\"incoming\");\n  *\tif (!run_command_v_opt_cd_env(cmd, 0, NULL, tmp_objdir_env(t)) &&\n  *\t    !tmp_objdir_migrate(t))\n  *\t\tprintf(\"success!\\n\");\n@@ -22,9 +22,10 @@\n struct tmp_objdir;\n \n /*\n- * Create a new temporary object directory; returns NULL on failure.\n+ * Create a new temporary object directory with the specified prefix;\n+ * returns NULL on failure.\n  */\n-struct tmp_objdir *tmp_objdir_create(void);\n+struct tmp_objdir *tmp_objdir_create(const char *prefix);\n \n /*\n  * Return a list of environment strings, suitable for use with\n@@ -51,4 +52,11 @@ int tmp_objdir_destroy(struct tmp_objdir *);\n  */\n void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n \n+/*\n+ * Replaces the main object store in the current process with the temporary\n+ * object directory and makes the former main object store an alternate.\n+ * If will_destroy is nonzero, the object directory may not be migrated.\n+ */\n+void tmp_objdir_replace_primary_odb(struct tmp_objdir *, int will_destroy);\n+\n #endif /* TMP_OBJDIR_H */\n-- \ngitgitgadget\n\n"},{"id":"437386","messageId":"55556bb3e90225263fa39d8813d1831eda135eb5.1632871972.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 4/9] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:46Z","receivedAt":"2021-09-28T23:33:02Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nand macOS offer mechanisms to write data from the filesystem page cache\nwithout initiating a hardware flush. Linux has the sync_file_range API,\nwhich issues a pagecache writeback request reliably after version 5.2.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\nbatches up hardware flushes. It hooks into the bulk-checkin plugging and\nunplugging functionality and takes advantage of tmp-objdir.\n\nWhen the new mode is enabled do the following for each new object:\n1. Create the object in a tmp-objdir.\n2. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1. Issue an fsync against a dummy file to flush the hardware writeback\n   cache, which should by now have processed the tmp-objdir writes.\n2. Rename all of the tmp-objdir files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the case today,\n   but may be a good extension to those components.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 29 ++++++++++++---\n Makefile                      |  6 +++\n bulk-checkin.c                | 70 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  2 +\n cache.h                       |  8 +++-\n config.c                      |  7 +++-\n config.mak.uname              |  1 +\n configure.ac                  |  8 ++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 ++++\n object-file.c                 | 12 +++++-\n tmp-objdir.c                  |  2 -\n wrapper.c                     | 44 ++++++++++++++++++++++\n write-or-die.c                |  2 +-\n 14 files changed, 187 insertions(+), 13 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..200b4d9f06e 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,29 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols loose objects in the object store, so updates to any\n+\trefs or the index may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each loose object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware. Git does not currently fsync parent directories for\n+  newly-added files, so some filesystems may still allow data to be lost on\n+  system crash.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write loose object data with a minimal set of FLUSH\n+  CACHE (or equivalent) commands sent to the storage controller. If the\n+  operating system interfaces are not available, this mode behaves the same as\n+  `true`. This mode is expected to be as safe as `true` on macOS for repos\n+  stored on HFS+ or APFS filesystems and on Windows for repos stored on NTFS or\n+  ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex a9f9b689f0c..313b3dc7cd6 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1874,6 +1876,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6ae18401e04..e6c830f9c0f 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,14 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n+static int needs_batch_fsync;\n+\n+static struct tmp_objdir *bulk_fsync_objdir;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -79,6 +85,34 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur.  The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware.\n+\t */\n+\n+\tif (needs_batch_fsync) {\n+\t\tstruct strbuf temp_path = STRBUF_INIT;\n+\t\tstruct tempfile *temp;\n+\n+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\t\ttemp = xmks_tempfile(temp_path.buf);\n+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\t\tdelete_tempfile(&temp);\n+\t\tstrbuf_release(&temp_path);\n+\t}\n+\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -273,6 +307,26 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void fsync_loose_object_bulk_checkin(int fd)\n+{\n+\tassert(fsync_object_files == FSYNC_OBJECT_FILES_BATCH);\n+\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (bulk_checkin_plugged &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\tassert(the_repository->objects->odb->is_temp);\n+\t\tif (!needs_batch_fsync)\n+\t\t\tneeds_batch_fsync = 1;\n+\t} else {\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -287,6 +341,20 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n+\n+\t/*\n+\t * Create a temporary object directory if the current\n+\t * object directory is not already temporary.\n+\t */\n+\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n+\t    !the_repository->objects->odb->is_temp) {\n+\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\t\tif (!bulk_fsync_objdir)\n+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n+\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+\t}\n+\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -296,4 +364,6 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..08f292379b6 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,8 @@\n \n #include \"cache.h\"\n \n+void fsync_loose_object_bulk_checkin(int fd);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex f6295f3b048..1ed8137b5e6 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -984,7 +984,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum fsync_object_files_mode {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum fsync_object_files_mode fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex 2edf835262f..8315d020eeb 100644\n--- a/config.c\n+++ b/config.c\n@@ -1506,7 +1506,12 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (value && !strcmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse if (git_config_bool(var, value))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n+\t\telse\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 76516aaa9a5..e6d482fbcc6 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex 30fca67e6d6..371a73c1e30 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,7 +42,7 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_object_files_mode fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 7c99eef6612..9daee873782 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1213,6 +1213,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex 1a3ad558c45..8ea1348f0db 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1932,8 +1932,18 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n static void close_loose_object(int fd)\n {\n \tif (!the_repository->objects->odb->will_destroy) {\n-\t\tif (fsync_object_files)\n+\t\tswitch (fsync_object_files) {\n+\t\tcase FSYNC_OBJECT_FILES_OFF:\n+\t\t\tbreak;\n+\t\tcase FSYNC_OBJECT_FILES_ON:\n \t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\t\tbreak;\n+\t\tcase FSYNC_OBJECT_FILES_BATCH:\n+\t\t\tfsync_loose_object_bulk_checkin(fd);\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tBUG(\"Invalid fsync_object_files mode.\");\n+\t\t}\n \t}\n \n \tif (close(fd) != 0)\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex 366ffe28511..c26cb5eafee 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -279,8 +279,6 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n \tif (!t)\n \t\treturn 0;\n \n-\n-\n \tif (t->prev_odb) {\n \t\tif (the_repository->objects->odb->will_destroy)\n \t\t\tBUG(\"migrating and ODB that was marked for destruction\");\ndiff --git a/wrapper.c b/wrapper.c\nindex 7c6586af321..bb4f9f043ce 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -540,6 +540,50 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync(fd);\n+#endif\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 0b1ec8190b6..cc8291d9794 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"437383","messageId":"6c33e79d6f0499bc24e4b01c44bb290556f61eef.1632871972.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 5/9] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:47Z","receivedAt":"2021-09-28T23:33:04Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds a win32 implementation for fsync_no_flush that is\ncalled git_fsync. The 'NtFlushBuffersFileEx' function being called is\navailable since Windows 8. If the function is not available, we\nreturn -1 and Git falls back to doing a full fsync.\n\nThe operating system is told to flush data only without a hardware\nflush primitive. A later full fsync will cause the metadata log\nto be flushed and then the disk cache to be flushed on NTFS and\nReFS. Other filesystems will treat this as a full flush operation.\n\nI added a new file here for this system call so as not to conflict with\ndownstream changes in the git-for-windows repository related to fscache.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.h                      |  3 +++\n compat/win32/flush.c                | 28 ++++++++++++++++++++++++++++\n config.mak.uname                    |  2 ++\n contrib/buildsystems/CMakeLists.txt |  3 ++-\n wrapper.c                           |  4 ++++\n 5 files changed, 39 insertions(+), 1 deletion(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..75324c24ee7\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex e6d482fbcc6..34c93314a50 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -451,6 +451,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -626,6 +627,7 @@ ifneq (,$(findstring MINGW,$(uname_S)))\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 171b4124afe..b573a5ee122 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/wrapper.c b/wrapper.c\nindex bb4f9f043ce..1a1e2fba9c9 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -567,6 +567,10 @@ int git_fsync(int fd, enum fsync_action action)\n \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n #endif\n \n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n \t\terrno = ENOSYS;\n \t\treturn -1;\n \n-- \ngitgitgadget\n\n"},{"id":"437384","messageId":"09dbff1004ed9b7ae501f7b1cfc91cd743195fd6.1632871972.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 6/9] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:48Z","receivedAt":"2021-09-28T23:33:05Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to synchronize with the update-index\nprocess after passing it a file path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 187203e8bb5..dc7368bb1ee 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1088,6 +1089,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \tthe_index.updated_skipworktree = 1;\n \n+\t/* we might be adding many objects to the object database */\n+\tplug_bulk_checkin();\n+\n \t/*\n \t * Custom copy of parse_options() because we want to handle\n \t * filename arguments as they come.\n@@ -1168,6 +1172,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/* by now we must have added all of the new objects */\n+\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"437385","messageId":"7aaa08d5f5fc2b3c83bc48de8dac25dea90d956b.1632871972.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 8/9] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:50Z","receivedAt":"2021-09-28T23:33:06Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added due\nto it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 36 ++++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 20 ++++++++++++++++++++\n t/t3903-stash.sh       | 14 ++++++++++++++\n t/t5300-pack-object.sh | 30 +++++++++++++++++++-----------\n 4 files changed, 89 insertions(+), 11 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a7de4ca8512\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,36 @@\n+# Helper to create files with unique contents\n+\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with unique\n+#\t\t\t\t\t contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\tlocal counter=0\n+\ttest_tick\n+\tlocal basedata=$test_tick\n+\n+\n+\trm -rf $basedir\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\"\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1))\n+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex 4086e1ebbc9..36049a53ff7 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -7,6 +7,8 @@ test_description='Test of git add, including the -- option.'\n \n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -33,6 +35,24 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+test_expect_success 'git add: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files2 &&\n+\tfind fsync-files2 ! -type d -print | xargs git -c core.fsyncobjectfiles=batch update-index --add -- &&\n+\trm -f fsynced_files2 &&\n+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n+\ttest_line_count = 8 fsynced_files2 &&\n+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 873aa56e359..2fc819e5584 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n diff_cmp () {\n \tfor i in \"$1\" \"$2\"\n@@ -1293,6 +1294,19 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n \texpected=\"stash.useBuiltin support has been removed\" &&\n \ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex e13a8842075..38663dc1393 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -162,23 +162,23 @@ test_expect_success 'pack-objects with bogus arguments' '\n \n check_unpack () {\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $2 init --bare git2 &&\n+\t(\n+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <obj-list >current &&\n+\tcmp obj-list current\n }\n \n test_expect_success 'unpack without delta' '\n \tcheck_unpack test-1-${packname_1}\n '\n \n+test_expect_success 'unpack without delta (core.fsyncobjectfiles=batch)' '\n+\tcheck_unpack test-1-${packname_1} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with REF_DELTA' '\n \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n \tcheck_deltas stderr -gt 0\n@@ -188,6 +188,10 @@ test_expect_success 'unpack with REF_DELTA' '\n \tcheck_unpack test-2-${packname_2}\n '\n \n+test_expect_success 'unpack with REF_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-2-${packname_2} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with OFS_DELTA' '\n \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n \t\t\t<obj-list 2>stderr) &&\n@@ -198,6 +202,10 @@ test_expect_success 'unpack with OFS_DELTA' '\n \tcheck_unpack test-3-${packname_3}\n '\n \n+test_expect_success 'unpack with OFS_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-3-${packname_3} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'compare delta flavors' '\n \tperl -e '\\''\n \t\tdefined($_ = -s $_) or die for @ARGV;\n-- \ngitgitgadget\n\n"},{"id":"437387","messageId":"1eced9f9f9a882749eb4210908e6561c51c48d87.1632871972.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 7/9] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:49Z","receivedAt":"2021-09-28T23:33:09Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling bulk-checkin when unpacking objects, we can take advantage\nof batched fsyncs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295ba..51eb4f7b531 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tplug_bulk_checkin();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"437388","messageId":"ff286fb461a3cc2b798aa4e135bfdfa462235c65.1632871972.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v7 9/9] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-09-28T23:32:51Z","receivedAt":"2021-09-28T23:33:10Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 43 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 46 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 89 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..e93c08a2e70\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..c9fcd0c03eb\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n"},{"id":"437396","messageId":"YVOrikAl/u5/Vi61@coredump.intra.peff.net","threadId":"56371","inReplyTo":"6e65f68fd6d4d90b0a7bca2e2e57ace9ad749266.1632871971.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v7 1/9] object-file.c: do not rename in a temp odb","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-09-28T23:55:54Z","receivedAt":"2021-09-28T23:56:13Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 28, 2021 at 11:32:43PM +0000, Neeraj Singh via GitGitGadget wrote:\n\n> If a temporary ODB is active, as determined by GIT_QUARANTINE_PATH\n> being set, create object files with their final names. This avoids\n> an extra rename beyond what is needed to merge the temporary ODB in\n> tmp_objdir_migrate.\n\nWhat's our goal here? Is it the performance of avoiding the extra\nrename()? Or do we benefit from the simplicity of avoiding it?\n\nIf the former, do we have measurements on how much this matters?\n\nIf the latter, what does the simplicity buy us? I thought maybe it would\nmake reasoning about fsync() easier, because we don't have to worry\nabout fsyncing the rename. But we'd eventually have to rename() into the\nreal object directory anyway.\n\nThe reason I want to push back is...\n\n> Creating an object file with the expected final name should be okay\n> since the git process writing to the temporary object store is the\n> only writer, and it only invokes write_loose_object/create_object_file\n> after checking that the object doesn't exist.\n\n...this seems like a kind-of dangerous assumption. Most of the time,\nyeah, I'd expect just a single process to be writing. But one of the\nthings that happens during the receive-pack quarantine is that we run\nhooks, which can run any set of arbitrary Git commands, including\nsimultaneous readers and writers. It seems like we might be introducing\nsubtle races there.\n\n-Peff\n"},{"id":"437398","messageId":"CANQDOddgurMuxXwpjoTWHDCC0LqSiKdvCtGCikvpu0EPr0x-3g@mail.gmail.com","threadId":"56371","inReplyTo":"YVOrikAl/u5/Vi61@coredump.intra.peff.net","subject":"Re: [PATCH v7 1/9] object-file.c: do not rename in a temp odb","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-29T00:10:18Z","receivedAt":"2021-09-29T00:10:32Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Sep 28, 2021 at 4:55 PM Jeff King <peff@peff.net> wrote:\n>\n> On Tue, Sep 28, 2021 at 11:32:43PM +0000, Neeraj Singh via GitGitGadget wrote:\n>\n> > If a temporary ODB is active, as determined by GIT_QUARANTINE_PATH\n> > being set, create object files with their final names. This avoids\n> > an extra rename beyond what is needed to merge the temporary ODB in\n> > tmp_objdir_migrate.\n>\n> What's our goal here? Is it the performance of avoiding the extra\n> rename()? Or do we benefit from the simplicity of avoiding it?\n>\n> If the former, do we have measurements on how much this matters?\n>\n> If the latter, what does the simplicity buy us? I thought maybe it would\n> make reasoning about fsync() easier, because we don't have to worry\n> about fsyncing the rename. But we'd eventually have to rename() into the\n> real object directory anyway.\n>\n> The reason I want to push back is...\n>\n> > Creating an object file with the expected final name should be okay\n> > since the git process writing to the temporary object store is the\n> > only writer, and it only invokes write_loose_object/create_object_file\n> > after checking that the object doesn't exist.\n>\n> ...this seems like a kind-of dangerous assumption. Most of the time,\n> yeah, I'd expect just a single process to be writing. But one of the\n> things that happens during the receive-pack quarantine is that we run\n> hooks, which can run any set of arbitrary Git commands, including\n> simultaneous readers and writers. It seems like we might be introducing\n> subtle races there.\n>\n> -Peff\n\nYes, the main goal was to avoid an extra rename. I see your concern\nand I guess we have no way of knowing if someone is really going to\nget bitten by this or not. On the other hand, we do know in the case\nof batch_fsync that only Git is running against the objdir at that\ntime.  I'll remove this change since it's not where the real perf\nbenefit is.\n\nThanks,\nNeeraj\n"},{"id":"437443","messageId":"CABPp-BHAQU=i0K9KCtqdifECw4qQjH=6c=4-Bz45yEmbT1YABw@mail.gmail.com","threadId":"56371","inReplyTo":"6ce72a709a11686b9082439a257fd5f58e5eb0f7.1632871971.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v7 2/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2021-09-29T08:41:41Z","receivedAt":"2021-09-29T08:41:55Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Hi,\n\nThanks for working on this, and for moving this up in your series near\nthe beginning.\n\nOn Tue, Sep 28, 2021 at 4:34 PM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> This patch is based on work by Elijah Newren. Any bugs however are my\n> own.\n\nThis kind of information is often included in a commit message via a\ntrailer such as:\n    Based-on-patch-by: Elijah Newren <newren@gmail.com>\nor Helped-by: or Co-authored-by: or Contributions-by: .\n\n> The tmp_objdir API provides the ability to create temporary object\n> directories, but was designed with the goal of having subprocesses\n> access these object stores, followed by the main process migrating\n> objects from it to the main object store or just deleting it.  The\n> subprocesses would view it as their primary datastore and write to it.\n>\n> Here we add the tmp_objdir_replace_primary_odb function that replaces\n> the current process's writable \"main\" object directory with the\n> specified one. The previous main object directory is restored in either\n> tmp_objdir_migrate or tmp_objdir_destroy.\n>\n> For the --remerge-diff usecase, add a new `will_destroy` flag in `struct\n> object_database` to mark ephemeral object databases that do not require\n> fsync durability.\n>\n> Add 'git prune' support for removing temporary object databases, and\n> make sure that they have a name starting with tmp_ and containing an\n> operation-specific name.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  builtin/prune.c        | 22 +++++++++++++++++----\n>  builtin/receive-pack.c |  2 +-\n>  object-file.c          | 45 ++++++++++++++++++++++++++++++++++++++++--\n>  object-store.h         | 21 +++++++++++++++++++-\n>  object.c               |  2 +-\n>  tmp-objdir.c           | 32 +++++++++++++++++++++++++++---\n>  tmp-objdir.h           | 14 ++++++++++---\n>  7 files changed, 123 insertions(+), 15 deletions(-)\n>\n> diff --git a/builtin/prune.c b/builtin/prune.c\n> index 02c6ab7cbaa..9c72ecf5a58 100644\n> --- a/builtin/prune.c\n> +++ b/builtin/prune.c\n> @@ -18,6 +18,7 @@ static int show_only;\n>  static int verbose;\n>  static timestamp_t expire;\n>  static int show_progress = -1;\n> +static struct strbuf remove_dir_buf = STRBUF_INIT;\n>\n>  static int prune_tmp_file(const char *fullpath)\n>  {\n> @@ -26,10 +27,19 @@ static int prune_tmp_file(const char *fullpath)\n>                 return error(\"Could not stat '%s'\", fullpath);\n>         if (st.st_mtime > expire)\n>                 return 0;\n> -       if (show_only || verbose)\n> -               printf(\"Removing stale temporary file %s\\n\", fullpath);\n> -       if (!show_only)\n> -               unlink_or_warn(fullpath);\n> +       if (S_ISDIR(st.st_mode)) {\n> +               if (show_only || verbose)\n> +                       printf(\"Removing stale temporary directory %s\\n\", fullpath);\n> +               if (!show_only) {\n> +                       strbuf_addstr(&remove_dir_buf, fullpath);\n> +                       remove_dir_recursively(&remove_dir_buf, 0);\n> +               }\n> +       } else {\n> +               if (show_only || verbose)\n> +                       printf(\"Removing stale temporary file %s\\n\", fullpath);\n> +               if (!show_only)\n> +                       unlink_or_warn(fullpath);\n> +       }\n>         return 0;\n>  }\n>\n> @@ -97,6 +107,9 @@ static int prune_cruft(const char *basename, const char *path, void *data)\n>\n>  static int prune_subdir(unsigned int nr, const char *path, void *data)\n>  {\n> +       if (verbose)\n> +               printf(\"Removing directory %s\\n\", path);\n> +\n>         if (!show_only)\n>                 rmdir(path);\n>         return 0;\n> @@ -185,5 +198,6 @@ int cmd_prune(int argc, const char **argv, const char *prefix)\n>                 prune_shallow(show_only ? PRUNE_SHOW_ONLY : 0);\n>         }\n>\n> +       strbuf_release(&remove_dir_buf);\n>         return 0;\n>  }\n> diff --git a/builtin/receive-pack.c b/builtin/receive-pack.c\n> index 48960a9575b..418a42ca069 100644\n> --- a/builtin/receive-pack.c\n> +++ b/builtin/receive-pack.c\n> @@ -2208,7 +2208,7 @@ static const char *unpack(int err_fd, struct shallow_info *si)\n>                 strvec_push(&child.args, alt_shallow_file);\n>         }\n>\n> -       tmp_objdir = tmp_objdir_create();\n> +       tmp_objdir = tmp_objdir_create(\"incoming\");\n>         if (!tmp_objdir) {\n>                 if (err_fd > 0)\n>                         close(err_fd);\n> diff --git a/object-file.c b/object-file.c\n> index 49c53f801f7..1a3ad558c45 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -751,6 +751,44 @@ void add_to_alternates_memory(const char *reference)\n>                              '\\n', NULL, 0);\n>  }\n>\n> +struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy)\n> +{\n> +       struct object_directory *new_odb;\n> +\n> +       /*\n> +        * Make sure alternates are initialized, or else our entry may be\n> +        * overwritten when they are.\n> +        */\n> +       prepare_alt_odb(the_repository);\n\nThis implicit dependence on the_repository is unfortunate.  My\nversions passed the repository parameter explicitly.  While my\nremerge-diff code doesn't really make use of that currently, it could\nmake sense to have temporary object stores for a submodule and do\nremerge-diff work on them.  You've also got two more uses of\nthe_repository later in this function.\n\n> +\n> +       /*\n> +        * Make a new primary odb and link the old primary ODB in as an\n> +        * alternate\n> +        */\n> +       new_odb = xcalloc(1, sizeof(*new_odb));\n> +       new_odb->path = xstrdup(dir);\n> +       new_odb->is_temp = 1;\n> +       new_odb->will_destroy = will_destroy;\n> +       new_odb->next = the_repository->objects->odb;\n> +       the_repository->objects->odb = new_odb;\n> +       return new_odb->next;\n> +}\n> +\n> +void restore_primary_odb(struct object_directory *restore_odb, const char *old_path)\n> +{\n> +       struct object_directory *cur_odb = the_repository->objects->odb;\n\nAnother use of the_repository, and some more below.\n\n> +\n> +       if (strcmp(old_path, cur_odb->path))\n> +               BUG(\"expected %s as primary object store; found %s\",\n> +                   old_path, cur_odb->path);\n> +\n> +       if (cur_odb->next != restore_odb)\n> +               BUG(\"we expect the old primary object store to be the first alternate\");\n> +\n> +       the_repository->objects->odb = restore_odb;\n> +       free_object_directory(cur_odb);\n> +}\n> +\n>  /*\n>   * Compute the exact path an alternate is at and returns it. In case of\n>   * error NULL is returned and the human readable error is added to `err`\n> @@ -1893,8 +1931,11 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  /* Finalize a file on disk, and close it. */\n>  static void close_loose_object(int fd)\n>  {\n> -       if (fsync_object_files)\n> -               fsync_or_die(fd, \"loose object file\");\n> +       if (!the_repository->objects->odb->will_destroy) {\n> +               if (fsync_object_files)\n> +                       fsync_or_die(fd, \"loose object file\");\n> +       }\n> +\n>         if (close(fd) != 0)\n>                 die_errno(_(\"error when closing loose object file\"));\n>  }\n> diff --git a/object-store.h b/object-store.h\n> index 551639f173d..5bc9da6634e 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -31,7 +31,12 @@ struct object_directory {\n>          * This is a temporary object store, so there is no need to\n>          * create new objects via rename.\n>          */\n> -       int is_temp;\n> +       int is_temp : 8;\n> +\n> +       /*\n> +        * This object store is ephemeral, so there is no need to fsync.\n> +        */\n> +       int will_destroy : 8;\n\nWhy 8 bits wide rather than 1?  I thought these were boolean\nvalues...was I mistaken?\n\n(Also, if boolean and compressing to 1 bit, should probably be\nunsigned rather than signed.)\n\n>         /*\n>          * Path to the alternative object store. If this is a relative path,\n> @@ -64,6 +69,17 @@ void add_to_alternates_file(const char *dir);\n>   */\n>  void add_to_alternates_memory(const char *dir);\n>\n> +/*\n> + * Replace the current writable object directory with the specified temporary\n> + * object directory; returns the former primary object directory.\n> + */\n> +struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy);\n> +\n> +/*\n> + * Restore a previous ODB replaced by set_temporary_main_odb.\n> + */\n> +void restore_primary_odb(struct object_directory *restore_odb, const char *old_path);\n> +\n>  /*\n>   * Populate and return the loose object cache array corresponding to the\n>   * given object ID.\n> @@ -74,6 +90,9 @@ struct oidtree *odb_loose_cache(struct object_directory *odb,\n>  /* Empty the loose object cache for the specified object directory. */\n>  void odb_clear_loose_cache(struct object_directory *odb);\n>\n> +/* Clear and free the specified object directory */\n> +void free_object_directory(struct object_directory *odb);\n> +\n>  struct packed_git {\n>         struct hashmap_entry packmap_ent;\n>         struct packed_git *next;\n> diff --git a/object.c b/object.c\n> index 4e85955a941..98635bc4043 100644\n> --- a/object.c\n> +++ b/object.c\n> @@ -513,7 +513,7 @@ struct raw_object_store *raw_object_store_new(void)\n>         return o;\n>  }\n>\n> -static void free_object_directory(struct object_directory *odb)\n> +void free_object_directory(struct object_directory *odb)\n>  {\n>         free(odb->path);\n>         odb_clear_loose_cache(odb);\n> diff --git a/tmp-objdir.c b/tmp-objdir.c\n> index b8d880e3626..366ffe28511 100644\n> --- a/tmp-objdir.c\n> +++ b/tmp-objdir.c\n> @@ -11,6 +11,7 @@\n>  struct tmp_objdir {\n>         struct strbuf path;\n>         struct strvec env;\n> +       struct object_directory *prev_odb;\n>  };\n>\n>  /*\n> @@ -38,6 +39,9 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n>         if (t == the_tmp_objdir)\n>                 the_tmp_objdir = NULL;\n>\n> +       if (!on_signal && t->prev_odb)\n> +               restore_primary_odb(t->prev_odb, t->path.buf);\n> +\n>         /*\n>          * This may use malloc via strbuf_grow(), but we should\n>          * have pre-grown t->path sufficiently so that this\n> @@ -52,6 +56,7 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n>          */\n>         if (!on_signal)\n>                 tmp_objdir_free(t);\n> +\n>         return err;\n>  }\n>\n> @@ -121,7 +126,7 @@ static int setup_tmp_objdir(const char *root)\n>         return ret;\n>  }\n>\n> -struct tmp_objdir *tmp_objdir_create(void)\n> +struct tmp_objdir *tmp_objdir_create(const char *prefix)\n>  {\n>         static int installed_handlers;\n>         struct tmp_objdir *t;\n> @@ -129,11 +134,16 @@ struct tmp_objdir *tmp_objdir_create(void)\n>         if (the_tmp_objdir)\n>                 BUG(\"only one tmp_objdir can be used at a time\");\n>\n> -       t = xmalloc(sizeof(*t));\n> +       t = xcalloc(1, sizeof(*t));\n>         strbuf_init(&t->path, 0);\n>         strvec_init(&t->env);\n>\n> -       strbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n> +       /*\n> +        * Use a string starting with tmp_ so that the builtin/prune.c code\n> +        * can recognize any stale objdirs left behind by a crash and delete\n> +        * them.\n> +        */\n> +       strbuf_addf(&t->path, \"%s/tmp_objdir-%s-XXXXXX\", get_object_directory(), prefix);\n>\n>         /*\n>          * Grow the strbuf beyond any filename we expect to be placed in it.\n> @@ -269,6 +279,15 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n>         if (!t)\n>                 return 0;\n>\n> +\n> +\n\nWhy so many blank lines?\n\n> +       if (t->prev_odb) {\n> +               if (the_repository->objects->odb->will_destroy)\n\nAnother implicit dependence on the_repository.\n\n> +                       BUG(\"migrating and ODB that was marked for destruction\");\n> +               restore_primary_odb(t->prev_odb, t->path.buf);\n> +               t->prev_odb = NULL;\n> +       }\n> +\n>         strbuf_addbuf(&src, &t->path);\n>         strbuf_addstr(&dst, get_object_directory());\n>\n> @@ -292,3 +311,10 @@ void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n>  {\n>         add_to_alternates_memory(t->path.buf);\n>  }\n> +\n> +void tmp_objdir_replace_primary_odb(struct tmp_objdir *t, int will_destroy)\n> +{\n> +       if (t->prev_odb)\n> +               BUG(\"the primary object database is already replaced\");\n> +       t->prev_odb = set_temporary_primary_odb(t->path.buf, will_destroy);\n> +}\n> diff --git a/tmp-objdir.h b/tmp-objdir.h\n> index b1e45b4c75d..75754cbfba6 100644\n> --- a/tmp-objdir.h\n> +++ b/tmp-objdir.h\n> @@ -10,7 +10,7 @@\n>   *\n>   * Example:\n>   *\n> - *     struct tmp_objdir *t = tmp_objdir_create();\n> + *     struct tmp_objdir *t = tmp_objdir_create(\"incoming\");\n>   *     if (!run_command_v_opt_cd_env(cmd, 0, NULL, tmp_objdir_env(t)) &&\n>   *         !tmp_objdir_migrate(t))\n>   *             printf(\"success!\\n\");\n> @@ -22,9 +22,10 @@\n>  struct tmp_objdir;\n>\n>  /*\n> - * Create a new temporary object directory; returns NULL on failure.\n> + * Create a new temporary object directory with the specified prefix;\n> + * returns NULL on failure.\n>   */\n> -struct tmp_objdir *tmp_objdir_create(void);\n> +struct tmp_objdir *tmp_objdir_create(const char *prefix);\n>\n>  /*\n>   * Return a list of environment strings, suitable for use with\n> @@ -51,4 +52,11 @@ int tmp_objdir_destroy(struct tmp_objdir *);\n>   */\n>  void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n>\n> +/*\n> + * Replaces the main object store in the current process with the temporary\n> + * object directory and makes the former main object store an alternate.\n> + * If will_destroy is nonzero, the object directory may not be migrated.\n> + */\n> +void tmp_objdir_replace_primary_odb(struct tmp_objdir *, int will_destroy);\n> +\n>  #endif /* TMP_OBJDIR_H */\n> --\n> gitgitgadget\n\nOther than those minor things, I couldn't find any problems.\n"},{"id":"437467","messageId":"CANQDOddqwVtWfC4eEP3fJB4sUiszGX8bLqoEVLcMf=v+jzx19g@mail.gmail.com","threadId":"56371","inReplyTo":"CABPp-BHAQU=i0K9KCtqdifECw4qQjH=6c=4-Bz45yEmbT1YABw@mail.gmail.com","subject":"Re: [PATCH v7 2/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-09-29T16:40:58Z","receivedAt":"2021-09-29T16:41:13Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Sep 29, 2021 at 1:42 AM Elijah Newren <newren@gmail.com> wrote:\n>\n> Hi,\n>\n> Thanks for working on this, and for moving this up in your series near\n> the beginning.\n>\n> On Tue, Sep 28, 2021 at 4:34 PM Neeraj Singh via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > This patch is based on work by Elijah Newren. Any bugs however are my\n> > own.\n>\n> This kind of information is often included in a commit message via a\n> trailer such as:\n>     Based-on-patch-by: Elijah Newren <newren@gmail.com>\n> or Helped-by: or Co-authored-by: or Contributions-by: .\n\nWill fix. I didn't know what some acceptable trailers were.  I'll use:\nBased-on-patch-by: Elijah Newren <newren@gmail.com>\n\n> >\n> > +struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy)\n> > +{\n> > +       struct object_directory *new_odb;\n> > +\n> > +       /*\n> > +        * Make sure alternates are initialized, or else our entry may be\n> > +        * overwritten when they are.\n> > +        */\n> > +       prepare_alt_odb(the_repository);\n>\n> This implicit dependence on the_repository is unfortunate.  My\n> versions passed the repository parameter explicitly.  While my\n> remerge-diff code doesn't really make use of that currently, it could\n> make sense to have temporary object stores for a submodule and do\n> remerge-diff work on them.  You've also got two more uses of\n> the_repository later in this function.\n\nThe core loose object code in object-file.c is riven with\nthe_repository assumptions. I'd have to refactor that code (including\nthe alternates code) to take repository arguments.  Given the\nextensive assumptions, I'd like to push back on this suggestion and\nall of the related suggestions.\n\n> > diff --git a/object-store.h b/object-store.h\n> > index 551639f173d..5bc9da6634e 100644\n> > --- a/object-store.h\n> > +++ b/object-store.h\n> > @@ -31,7 +31,12 @@ struct object_directory {\n> >          * This is a temporary object store, so there is no need to\n> >          * create new objects via rename.\n> >          */\n> > -       int is_temp;\n> > +       int is_temp : 8;\n> > +\n> > +       /*\n> > +        * This object store is ephemeral, so there is no need to fsync.\n> > +        */\n> > +       int will_destroy : 8;\n>\n> Why 8 bits wide rather than 1?  I thought these were boolean\n> values...was I mistaken?\n>\n> (Also, if boolean and compressing to 1 bit, should probably be\n> unsigned rather than signed.)\n\nThis will go away when I drop the rename patch.  I wish we had a\nstandard bool_t type which is one char wide.  This is a\nmicrooptimization, since accessing bits usually encodes to more or\nlarger instructions than accessing bytes.\n\n> > +        */\n> > +       strbuf_addf(&t->path, \"%s/tmp_objdir-%s-XXXXXX\", get_object_directory(), prefix);\n> >\n> >         /*\n> >          * Grow the strbuf beyond any filename we expect to be placed in it.\n> > @@ -269,6 +279,15 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n> >         if (!t)\n> >                 return 0;\n> >\n> > +\n> > +\n>\n> Why so many blank lines?\n>\n\nThis was an accident, will remove.\n\n> Other than those minor things, I couldn't find any problems.\n\nThanks for the review!\n"},{"id":"437914","messageId":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v7.git.git.1632871971.gitgitgadget@gmail.com","subject":"[PATCH v8 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:38Z","receivedAt":"2021-10-04T16:57:53Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far!\n\nThis series shares the base tmp-objdir patches with my merged version of\nElijah Newren's remerge-diff series at:\nhttps://github.com/neerajsi-msft/git/tree/neerajsi/remerge-diff.\n\nThe patch is now at version 8: changes since v7:\n\n * Dropped the tmp-objdir patch to avoid renaming in a quarantine/temporary\n   objdir, as suggested by Jeff King. This wasn't a good idea because we\n   don't really know that there's only a single reader/writer. Avoiding the\n   rename was a relatively minor perf optimization so it's okay to drop.\n\n * Added disable_ref_updates logic (as a flag on the odb) which is set when\n   we're in a quarantine or when a tmp objdir is active. I believe this\n   roughly follows the strategy suggested by Jeff King.\n\nThe patch is now at version 7: changes since v6:\n\n * Rebased onto current upstream master\n\n * Separate the tmp-objdir changes and move to the beginning of the series\n   so that Elijah Newren's similar changes can be merged.\n\n * Use some of Elijah's implementation for replacing the primary ODB. I was\n   doing some unnecessarily complex copying for no good reason.\n\n * Make the tmp objdir code use a name beginning with tmp_ and having a\n   operation-specific prefix.\n\n * Add git-prune support for removing a stale object directory.\n\nv5 was a bit of a dud, with some issues that I only noticed after\nsubmitting. v6 changes:\n\n * re-add Windows support\n * fix minor formatting issues\n * reset git author and commit dates which got messed up\n\nChanges since v4, all in response to review feedback from Ævar Arnfjörð\nBjarmason:\n\n * Update core.fsyncobjectfiles documentation to specify 'loose' objects and\n   to add a statement about not fsyncing parent directories.\n   \n   * I still don't want to make any promises on behalf of the Linux FS developers\n     in the documentation. However, according to [v4.1] and my understanding\n     of how XFS journals are documented to work, it looks like recent versions\n     of Linux running on XFS should be as safe as Windows or macOS in 'batch'\n     mode. I don't know about ext4, since it's not clear to me when metadata\n     updates are made visible to the journal.\n   \n\n * Rewrite the core batched fsync change to use the tmp-objdir lib. As Ævar\n   pointed out, this lets us access the added loose objects immediately,\n   rather than only after unplugging the bulk checkin. This is a hard\n   requirement in unpack-objects for resolving OBJ_REF_DELTA packed objects.\n   \n   * As a preparatory patch, the object-file code now doesn't do a rename if it's in a\n     tmp objdir (as determined by the quarantine environment variable).\n   \n   * I added support to the tmp-objdir lib to replace the 'main' writable odb.\n   \n   * Instead of using a lockfile for the final full fsync, we now use a new dummy\n     temp file. Doing that makes the below unpack-objects change easier.\n   \n\n * Add bulk-checkin support to unpack-objects, which is used in fetch and\n   push. In addition to making those operations faster, it allows us to\n   directly compare performance of packfiles against loose objects. Please\n   see [v4.2] for a measurement of 'git push' to a local upstream with\n   different numbers of unique new files.\n\n * Rename FSYNC_OBJECT_FILES_MODE to fsync_object_files_mode.\n\n * Remove comment with link to NtFlushBuffersFileEx documentation.\n\n * Make t/lib-unique-files.sh a bit cleaner. We are still creating unique\n   contents, but now this uses test_tick, so it should be deterministic from\n   run to run.\n\n * Ensure there are tests for all of the modified commands. Make the\n   unpack-objects tests validate that the unpacked objects are really\n   available in the ODB.\n\nReferences for v4: [v4.1]\nhttps://lore.kernel.org/linux-fsdevel/20190419072938.31320-1-amir73il@gmail.com/#t\n\n[v4.2]\nhttps://docs.google.com/spreadsheets/d/1uxMBkEXFFnQ1Y3lXKqcKpw6Mq44BzhpCAcPex14T-QQ/edit#gid=1898936117\n\nChanges since v3:\n\n * Fix core.fsyncobjectfiles option parsing as suggested by Junio: We now\n   accept no value to mean \"true\" and we require 'batch' to be lowercase.\n\n * Leave the default fsync mode as 'false'. Git for windows can change its\n   default when this series makes it over to that fork.\n\n * Use a switch statement in git_fsync, as suggested by Junio.\n\n * Add regression test cases for core.fsyncobjectfiles=batch. This should\n   keep the batch functionality basically working in upstream git even if\n   few users adopt batch mode initially. I expect git-for-windows will\n   provide a good baking area for the new mode.\n\nNeeraj Singh (9):\n  tmp-objdir: new API for creating temporary writable databases\n  tmp-objdir: disable ref updates when replacing the primary odb\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  core.fsyncobjectfiles: add windows support for batch mode\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: tests for batch mode\n  core.fsyncobjectfiles: performance tests for add and stash\n\n Documentation/config/core.txt       | 29 ++++++++--\n Makefile                            |  6 ++\n builtin/prune.c                     | 22 +++++--\n builtin/receive-pack.c              |  2 +-\n builtin/unpack-objects.c            |  3 +\n builtin/update-index.c              |  6 ++\n bulk-checkin.c                      | 90 +++++++++++++++++++++++++----\n bulk-checkin.h                      |  2 +\n cache.h                             |  8 ++-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n config.c                            |  7 ++-\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n environment.c                       |  6 +-\n git-compat-util.h                   |  7 +++\n object-file.c                       | 60 ++++++++++++++++++-\n object-store.h                      | 26 +++++++++\n object.c                            |  2 +-\n refs.c                              |  2 +-\n repository.c                        |  2 +\n repository.h                        |  1 +\n t/lib-unique-files.sh               | 36 ++++++++++++\n t/perf/p3700-add.sh                 | 43 ++++++++++++++\n t/perf/p3900-stash.sh               | 46 +++++++++++++++\n t/t3700-add.sh                      | 20 +++++++\n t/t3903-stash.sh                    | 14 +++++\n t/t5300-pack-object.sh              | 30 ++++++----\n tmp-objdir.c                        | 30 +++++++++-\n tmp-objdir.h                        | 14 ++++-\n wrapper.c                           | 48 +++++++++++++++\n write-or-die.c                      |  2 +-\n 33 files changed, 562 insertions(+), 47 deletions(-)\n create mode 100644 compat/win32/flush.c\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: cefe983a320c03d7843ac78e73bd513a27806845\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v8\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v8\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v7:\n\n  2:  6ce72a709a1 !  1:  f03797fd80d tmp-objdir: new API for creating temporary writable databases\n     @@ Metadata\n       ## Commit message ##\n          tmp-objdir: new API for creating temporary writable databases\n      \n     -    This patch is based on work by Elijah Newren. Any bugs however are my\n     -    own.\n     -\n          The tmp_objdir API provides the ability to create temporary object\n          directories, but was designed with the goal of having subprocesses\n          access these object stores, followed by the main process migrating\n     @@ Commit message\n          make sure that they have a name starting with tmp_ and containing an\n          operation-specific name.\n      \n     +    Based-on-patch-by: Elijah Newren <newren@gmail.com>\n     +\n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## builtin/prune.c ##\n     @@ object-file.c: void add_to_alternates_memory(const char *reference)\n      +\t */\n      +\tnew_odb = xcalloc(1, sizeof(*new_odb));\n      +\tnew_odb->path = xstrdup(dir);\n     -+\tnew_odb->is_temp = 1;\n      +\tnew_odb->will_destroy = will_destroy;\n      +\tnew_odb->next = the_repository->objects->odb;\n      +\tthe_repository->objects->odb = new_odb;\n     @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void\n      \n       ## object-store.h ##\n      @@ object-store.h: struct object_directory {\n     - \t * This is a temporary object store, so there is no need to\n     - \t * create new objects via rename.\n     - \t */\n     --\tint is_temp;\n     -+\tint is_temp : 8;\n     -+\n     + \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n     + \tstruct oidtree *loose_objects_cache;\n     + \n      +\t/*\n      +\t * This object store is ephemeral, so there is no need to fsync.\n      +\t */\n     -+\tint will_destroy : 8;\n     - \n     ++\tint will_destroy;\n     ++\n       \t/*\n       \t * Path to the alternative object store. If this is a relative path,\n     + \t * it is relative to the current working directory.\n      @@ object-store.h: void add_to_alternates_file(const char *dir);\n        */\n       void add_to_alternates_memory(const char *dir);\n     @@ tmp-objdir.c: int tmp_objdir_migrate(struct tmp_objdir *t)\n       \tif (!t)\n       \t\treturn 0;\n       \n     -+\n     -+\n      +\tif (t->prev_odb) {\n      +\t\tif (the_repository->objects->odb->will_destroy)\n     -+\t\t\tBUG(\"migrating and ODB that was marked for destruction\");\n     ++\t\t\tBUG(\"migrating an ODB that was marked for destruction\");\n      +\t\trestore_primary_odb(t->prev_odb, t->path.buf);\n      +\t\tt->prev_odb = NULL;\n      +\t}\n  1:  6e65f68fd6d !  2:  bc085137340 object-file.c: do not rename in a temp odb\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    object-file.c: do not rename in a temp odb\n     +    tmp-objdir: disable ref updates when replacing the primary odb\n      \n     -    If a temporary ODB is active, as determined by GIT_QUARANTINE_PATH\n     -    being set, create object files with their final names. This avoids\n     -    an extra rename beyond what is needed to merge the temporary ODB in\n     -    tmp_objdir_migrate.\n     +    When creating a subprocess with a temporary ODB, we set the\n     +    GIT_QUARANTINE_ENVIRONMENT env var to tell child Git processes not\n     +    to update refs, since the tmp-objdir may go away.\n      \n     -    Creating an object file with the expected final name should be okay\n     -    since the git process writing to the temporary object store is the\n     -    only writer, and it only invokes write_loose_object/create_object_file\n     -    after checking that the object doesn't exist.\n     +    Introduce a similar mechanism for in-process temporary ODBs when\n     +    we call tmp_objdir_replace_primary_odb. Now both mechanisms set\n     +    the disable_ref_updates flag on the odb, which is queried by\n     +    the ref_transaction_prepare function.\n     +\n     +    Note: This change adds an assumption that the state of\n     +    the_repository is relevant for any ref transaction that might\n     +    be initiated. Unwinding this assumption should be straightforward\n     +    by saving the relevant repository to query in the transaction or\n     +    the ref_store.\n     +\n     +    Peff's test case was invoking ref updates via the cachetextconv\n     +    setting. That particular code silently does nothing when a ref\n     +    update is forbidden. See the call to notes_cache_put in\n     +    fill_textconv where errors are ignored.\n     +\n     +    Reported-by: Jeff King <peff@peff.net>\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ environment.c: void setup_git_env(const char *git_dir)\n       \targs.index_file = getenv_safe(&to_free, INDEX_ENVIRONMENT);\n       \targs.alternate_db = getenv_safe(&to_free, ALTERNATE_DB_ENVIRONMENT);\n      +\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n     -+\t\targs.object_dir_is_temp = 1;\n     ++\t\targs.disable_ref_updates = 1;\n      +\t}\n      +\n       \trepo_set_gitdir(the_repository, git_dir, &args);\n     @@ environment.c: void setup_git_env(const char *git_dir)\n       \n      \n       ## object-file.c ##\n     -@@ object-file.c: static void write_object_file_prepare(const struct git_hash_algo *algo,\n     - }\n     - \n     - /*\n     -- * Move the just written object into its final resting place.\n     -+ * Move the just written object into its final resting place,\n     -+ * unless it is already there, as indicated by an empty string for\n     -+ * tmpfile.\n     -  */\n     - int finalize_object_file(const char *tmpfile, const char *filename)\n     - {\n     - \tint ret = 0;\n     - \n     -+\tif (!*tmpfile)\n     -+\t\tgoto out;\n     +@@ object-file.c: struct object_directory *set_temporary_primary_odb(const char *dir, int will_des\n     + \t */\n     + \tnew_odb = xcalloc(1, sizeof(*new_odb));\n     + \tnew_odb->path = xstrdup(dir);\n      +\n     - \tif (object_creation_mode == OBJECT_CREATION_USES_RENAMES)\n     - \t\tgoto try_rename;\n     - \telse if (link(tmpfile, filename))\n     -@@ object-file.c: static inline int directory_size(const char *filename)\n     - }\n     - \n     - /*\n     -- * This creates a temporary file in the same directory as the final\n     -- * 'filename'\n     -+ * This creates a loose object file for the specified object id.\n     -+ * If we're working in a temporary object directory, the file is\n     -+ * created with its final filename, otherwise it is created with\n     -+ * a temporary name and renamed by finalize_object_file.\n     -+ * If no rename is required, an empty string is returned in tmp.\n     -  *\n     -  * We want to avoid cross-directory filename renames, because those\n     -  * can have problems on various filesystems (FAT, NFS, Coda).\n     -  */\n     --static int create_tmpfile(struct strbuf *tmp, const char *filename)\n     -+static int create_objfile(const struct object_id *oid, struct strbuf *tmp,\n     -+\t\t\t  struct strbuf *filename)\n     - {\n     --\tint fd, dirlen = directory_size(filename);\n     -+\tint fd, dirlen, is_retrying = 0;\n     -+\tconst char *object_name;\n     -+\tstatic const int object_mode = 0444;\n     - \n     -+\tloose_object_path(the_repository, filename, oid);\n     -+\tdirlen = directory_size(filename->buf);\n     -+\n     -+retry_create:\n     - \tstrbuf_reset(tmp);\n     --\tstrbuf_add(tmp, filename, dirlen);\n     --\tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n     --\tfd = git_mkstemp_mode(tmp->buf, 0444);\n     --\tif (fd < 0 && dirlen && errno == ENOENT) {\n     -+\tif (!the_repository->objects->odb->is_temp) {\n     -+\t\tstrbuf_add(tmp, filename->buf, dirlen);\n     -+\t\tobject_name = \"tmp_obj_XXXXXX\";\n     -+\t\tstrbuf_addstr(tmp, object_name);\n     -+\t\tfd = git_mkstemp_mode(tmp->buf, object_mode);\n     -+\t} else {\n     -+\t\tfd = open(filename->buf, O_CREAT | O_EXCL | O_RDWR, object_mode);\n     -+\t}\n     -+\n     -+\tif (fd < 0 && dirlen && errno == ENOENT && !is_retrying) {\n     - \t\t/*\n     - \t\t * Make sure the directory exists; note that the contents\n     - \t\t * of the buffer are undefined after mkstemp returns an\n     -@@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filename)\n     - \t\t * scratch.\n     - \t\t */\n     - \t\tstrbuf_reset(tmp);\n     --\t\tstrbuf_add(tmp, filename, dirlen - 1);\n     -+\t\tstrbuf_add(tmp, filename->buf, dirlen - 1);\n     - \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n     - \t\t\treturn -1;\n     - \t\tif (adjust_shared_perm(tmp->buf))\n     - \t\t\treturn -1;\n     - \n     - \t\t/* Try again */\n     --\t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n     --\t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n     -+\t\tis_retrying = 1;\n     -+\t\tgoto retry_create;\n     - \t}\n     - \treturn fd;\n     - }\n     -@@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n     - \tstatic struct strbuf tmp_file = STRBUF_INIT;\n     - \tstatic struct strbuf filename = STRBUF_INIT;\n     - \n     --\tloose_object_path(the_repository, &filename, oid);\n     --\n     --\tfd = create_tmpfile(&tmp_file, filename.buf);\n     -+\tfd = create_objfile(oid, &tmp_file, &filename);\n     - \tif (fd < 0) {\n     - \t\tif (errno == EACCES)\n     - \t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n     - \t\telse\n     --\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n     -+\t\t\treturn error_errno(_(\"unable to create object file\"));\n     - \t}\n     - \n     - \t/* Set it up */\n     ++\t/*\n     ++\t * Disable ref updates while a temporary odb is active, since\n     ++\t * the objects in the database may roll back.\n     ++\t */\n     ++\tnew_odb->disable_ref_updates = 1;\n     + \tnew_odb->will_destroy = will_destroy;\n     + \tnew_odb->next = the_repository->objects->odb;\n     + \tthe_repository->objects->odb = new_odb;\n      \n       ## object-store.h ##\n      @@ object-store.h: struct object_directory {\n     @@ object-store.h: struct object_directory {\n       \tstruct oidtree *loose_objects_cache;\n       \n      +\t/*\n     -+\t * This is a temporary object store, so there is no need to\n     -+\t * create new objects via rename.\n     ++\t * This is a temporary object store created by the tmp_objdir\n     ++\t * facility. Disable ref updates since the objects in the store\n     ++\t * might be discarded on rollback.\n      +\t */\n     -+\tint is_temp;\n     ++\tunsigned int disable_ref_updates : 1;\n      +\n     + \t/*\n     + \t * This object store is ephemeral, so there is no need to fsync.\n     + \t */\n     +-\tint will_destroy;\n     ++\tunsigned int will_destroy : 1;\n     + \n       \t/*\n       \t * Path to the alternative object store. If this is a relative path,\n     - \t * it is relative to the current working directory.\n     +\n     + ## refs.c ##\n     +@@ refs.c: int ref_transaction_prepare(struct ref_transaction *transaction,\n     + \t\tbreak;\n     + \t}\n     + \n     +-\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n     ++\tif (the_repository->objects->odb->disable_ref_updates) {\n     + \t\tstrbuf_addstr(err,\n     + \t\t\t      _(\"ref updates forbidden inside quarantine environment\"));\n     + \t\treturn -1;\n      \n       ## repository.c ##\n      @@ repository.c: void repo_set_gitdir(struct repository *repo,\n       \texpand_base_dir(&repo->objects->odb->path, o->object_dir,\n       \t\t\trepo->commondir, \"objects\");\n       \n     -+\trepo->objects->odb->is_temp = o->object_dir_is_temp;\n     ++\trepo->objects->odb->disable_ref_updates = o->disable_ref_updates;\n      +\n       \tfree(repo->objects->alternate_db);\n       \trepo->objects->alternate_db = xstrdup_or_null(o->alternate_db);\n     @@ repository.h: struct set_gitdir_args {\n       \tconst char *graft_file;\n       \tconst char *index_file;\n       \tconst char *alternate_db;\n     -+\tint object_dir_is_temp;\n     ++\tint disable_ref_updates;\n       };\n       \n       void repo_set_gitdir(struct repository *repo, const char *root,\n  3:  c272f8776fa =  3:  9335646ed91 bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  4:  55556bb3e90 !  4:  b9d3d874432 core.fsyncobjectfiles: batched disk flushes\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      +\t */\n      +\tif (bulk_checkin_plugged &&\n      +\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n     -+\t\tassert(the_repository->objects->odb->is_temp);\n      +\t\tif (!needs_batch_fsync)\n      +\t\t\tneeds_batch_fsync = 1;\n      +\t} else {\n     @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n       \tassert(!bulk_checkin_plugged);\n      +\n      +\t/*\n     -+\t * Create a temporary object directory if the current\n     -+\t * object directory is not already temporary.\n     ++\t * A temporary object directory is used to hold the files\n     ++\t * while they are not fsynced.\n      +\t */\n     -+\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH &&\n     -+\t    !the_repository->objects->odb->is_temp) {\n     ++\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH) {\n      +\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n      +\t\tif (!bulk_fsync_objdir)\n      +\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n     @@ object-file.c: int hash_object_file(const struct git_hash_algo *algo, const void\n       \n       \tif (close(fd) != 0)\n      \n     - ## tmp-objdir.c ##\n     -@@ tmp-objdir.c: int tmp_objdir_migrate(struct tmp_objdir *t)\n     - \tif (!t)\n     - \t\treturn 0;\n     - \n     --\n     --\n     - \tif (t->prev_odb) {\n     - \t\tif (the_repository->objects->odb->will_destroy)\n     - \t\t\tBUG(\"migrating and ODB that was marked for destruction\");\n     -\n       ## wrapper.c ##\n      @@ wrapper.c: int xmkstemp_mode(char *filename_template, int mode)\n       \treturn fd;\n  5:  6c33e79d6f0 =  5:  8df32eaaa9a core.fsyncobjectfiles: add windows support for batch mode\n  6:  09dbff1004e =  6:  15767270984 update-index: use the bulk-checkin infrastructure\n  7:  1eced9f9f9a =  7:  e88bab809a2 unpack-objects: use the bulk-checkin infrastructure\n  8:  7aaa08d5f5f =  8:  811d6d31509 core.fsyncobjectfiles: tests for batch mode\n  9:  ff286fb461a =  9:  f4fa20f591e core.fsyncobjectfiles: performance tests for add and stash\n\n-- \ngitgitgadget\n"},{"id":"437913","messageId":"f03797fd80daf519f26b93eab06142fe9a0b2b94.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 1/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:39Z","receivedAt":"2021-10-04T16:57:54Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe tmp_objdir API provides the ability to create temporary object\ndirectories, but was designed with the goal of having subprocesses\naccess these object stores, followed by the main process migrating\nobjects from it to the main object store or just deleting it.  The\nsubprocesses would view it as their primary datastore and write to it.\n\nHere we add the tmp_objdir_replace_primary_odb function that replaces\nthe current process's writable \"main\" object directory with the\nspecified one. The previous main object directory is restored in either\ntmp_objdir_migrate or tmp_objdir_destroy.\n\nFor the --remerge-diff usecase, add a new `will_destroy` flag in `struct\nobject_database` to mark ephemeral object databases that do not require\nfsync durability.\n\nAdd 'git prune' support for removing temporary object databases, and\nmake sure that they have a name starting with tmp_ and containing an\noperation-specific name.\n\nBased-on-patch-by: Elijah Newren <newren@gmail.com>\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/prune.c        | 22 +++++++++++++++++----\n builtin/receive-pack.c |  2 +-\n object-file.c          | 44 ++++++++++++++++++++++++++++++++++++++++--\n object-store.h         | 19 ++++++++++++++++++\n object.c               |  2 +-\n tmp-objdir.c           | 30 +++++++++++++++++++++++++---\n tmp-objdir.h           | 14 +++++++++++---\n 7 files changed, 119 insertions(+), 14 deletions(-)\n\ndiff --git a/builtin/prune.c b/builtin/prune.c\nindex 02c6ab7cbaa..9c72ecf5a58 100644\n--- a/builtin/prune.c\n+++ b/builtin/prune.c\n@@ -18,6 +18,7 @@ static int show_only;\n static int verbose;\n static timestamp_t expire;\n static int show_progress = -1;\n+static struct strbuf remove_dir_buf = STRBUF_INIT;\n \n static int prune_tmp_file(const char *fullpath)\n {\n@@ -26,10 +27,19 @@ static int prune_tmp_file(const char *fullpath)\n \t\treturn error(\"Could not stat '%s'\", fullpath);\n \tif (st.st_mtime > expire)\n \t\treturn 0;\n-\tif (show_only || verbose)\n-\t\tprintf(\"Removing stale temporary file %s\\n\", fullpath);\n-\tif (!show_only)\n-\t\tunlink_or_warn(fullpath);\n+\tif (S_ISDIR(st.st_mode)) {\n+\t\tif (show_only || verbose)\n+\t\t\tprintf(\"Removing stale temporary directory %s\\n\", fullpath);\n+\t\tif (!show_only) {\n+\t\t\tstrbuf_addstr(&remove_dir_buf, fullpath);\n+\t\t\tremove_dir_recursively(&remove_dir_buf, 0);\n+\t\t}\n+\t} else {\n+\t\tif (show_only || verbose)\n+\t\t\tprintf(\"Removing stale temporary file %s\\n\", fullpath);\n+\t\tif (!show_only)\n+\t\t\tunlink_or_warn(fullpath);\n+\t}\n \treturn 0;\n }\n \n@@ -97,6 +107,9 @@ static int prune_cruft(const char *basename, const char *path, void *data)\n \n static int prune_subdir(unsigned int nr, const char *path, void *data)\n {\n+\tif (verbose)\n+\t\tprintf(\"Removing directory %s\\n\", path);\n+\n \tif (!show_only)\n \t\trmdir(path);\n \treturn 0;\n@@ -185,5 +198,6 @@ int cmd_prune(int argc, const char **argv, const char *prefix)\n \t\tprune_shallow(show_only ? PRUNE_SHOW_ONLY : 0);\n \t}\n \n+\tstrbuf_release(&remove_dir_buf);\n \treturn 0;\n }\ndiff --git a/builtin/receive-pack.c b/builtin/receive-pack.c\nindex 48960a9575b..418a42ca069 100644\n--- a/builtin/receive-pack.c\n+++ b/builtin/receive-pack.c\n@@ -2208,7 +2208,7 @@ static const char *unpack(int err_fd, struct shallow_info *si)\n \t\tstrvec_push(&child.args, alt_shallow_file);\n \t}\n \n-\ttmp_objdir = tmp_objdir_create();\n+\ttmp_objdir = tmp_objdir_create(\"incoming\");\n \tif (!tmp_objdir) {\n \t\tif (err_fd > 0)\n \t\t\tclose(err_fd);\ndiff --git a/object-file.c b/object-file.c\nindex be4f94ecf3b..990381abee5 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -751,6 +751,43 @@ void add_to_alternates_memory(const char *reference)\n \t\t\t     '\\n', NULL, 0);\n }\n \n+struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy)\n+{\n+\tstruct object_directory *new_odb;\n+\n+\t/*\n+\t * Make sure alternates are initialized, or else our entry may be\n+\t * overwritten when they are.\n+\t */\n+\tprepare_alt_odb(the_repository);\n+\n+\t/*\n+\t * Make a new primary odb and link the old primary ODB in as an\n+\t * alternate\n+\t */\n+\tnew_odb = xcalloc(1, sizeof(*new_odb));\n+\tnew_odb->path = xstrdup(dir);\n+\tnew_odb->will_destroy = will_destroy;\n+\tnew_odb->next = the_repository->objects->odb;\n+\tthe_repository->objects->odb = new_odb;\n+\treturn new_odb->next;\n+}\n+\n+void restore_primary_odb(struct object_directory *restore_odb, const char *old_path)\n+{\n+\tstruct object_directory *cur_odb = the_repository->objects->odb;\n+\n+\tif (strcmp(old_path, cur_odb->path))\n+\t\tBUG(\"expected %s as primary object store; found %s\",\n+\t\t    old_path, cur_odb->path);\n+\n+\tif (cur_odb->next != restore_odb)\n+\t\tBUG(\"we expect the old primary object store to be the first alternate\");\n+\n+\tthe_repository->objects->odb = restore_odb;\n+\tfree_object_directory(cur_odb);\n+}\n+\n /*\n  * Compute the exact path an alternate is at and returns it. In case of\n  * error NULL is returned and the human readable error is added to `err`\n@@ -1888,8 +1925,11 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\tif (!the_repository->objects->odb->will_destroy) {\n+\t\tif (fsync_object_files)\n+\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/object-store.h b/object-store.h\nindex c5130d8baea..74b1b5872a6 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -27,6 +27,11 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/*\n+\t * This object store is ephemeral, so there is no need to fsync.\n+\t */\n+\tint will_destroy;\n+\n \t/*\n \t * Path to the alternative object store. If this is a relative path,\n \t * it is relative to the current working directory.\n@@ -58,6 +63,17 @@ void add_to_alternates_file(const char *dir);\n  */\n void add_to_alternates_memory(const char *dir);\n \n+/*\n+ * Replace the current writable object directory with the specified temporary\n+ * object directory; returns the former primary object directory.\n+ */\n+struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy);\n+\n+/*\n+ * Restore a previous ODB replaced by set_temporary_main_odb.\n+ */\n+void restore_primary_odb(struct object_directory *restore_odb, const char *old_path);\n+\n /*\n  * Populate and return the loose object cache array corresponding to the\n  * given object ID.\n@@ -68,6 +84,9 @@ struct oidtree *odb_loose_cache(struct object_directory *odb,\n /* Empty the loose object cache for the specified object directory. */\n void odb_clear_loose_cache(struct object_directory *odb);\n \n+/* Clear and free the specified object directory */\n+void free_object_directory(struct object_directory *odb);\n+\n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n \tstruct packed_git *next;\ndiff --git a/object.c b/object.c\nindex 4e85955a941..98635bc4043 100644\n--- a/object.c\n+++ b/object.c\n@@ -513,7 +513,7 @@ struct raw_object_store *raw_object_store_new(void)\n \treturn o;\n }\n \n-static void free_object_directory(struct object_directory *odb)\n+void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex b8d880e3626..45d42a7bcf0 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -11,6 +11,7 @@\n struct tmp_objdir {\n \tstruct strbuf path;\n \tstruct strvec env;\n+\tstruct object_directory *prev_odb;\n };\n \n /*\n@@ -38,6 +39,9 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \tif (t == the_tmp_objdir)\n \t\tthe_tmp_objdir = NULL;\n \n+\tif (!on_signal && t->prev_odb)\n+\t\trestore_primary_odb(t->prev_odb, t->path.buf);\n+\n \t/*\n \t * This may use malloc via strbuf_grow(), but we should\n \t * have pre-grown t->path sufficiently so that this\n@@ -52,6 +56,7 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \t */\n \tif (!on_signal)\n \t\ttmp_objdir_free(t);\n+\n \treturn err;\n }\n \n@@ -121,7 +126,7 @@ static int setup_tmp_objdir(const char *root)\n \treturn ret;\n }\n \n-struct tmp_objdir *tmp_objdir_create(void)\n+struct tmp_objdir *tmp_objdir_create(const char *prefix)\n {\n \tstatic int installed_handlers;\n \tstruct tmp_objdir *t;\n@@ -129,11 +134,16 @@ struct tmp_objdir *tmp_objdir_create(void)\n \tif (the_tmp_objdir)\n \t\tBUG(\"only one tmp_objdir can be used at a time\");\n \n-\tt = xmalloc(sizeof(*t));\n+\tt = xcalloc(1, sizeof(*t));\n \tstrbuf_init(&t->path, 0);\n \tstrvec_init(&t->env);\n \n-\tstrbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n+\t/*\n+\t * Use a string starting with tmp_ so that the builtin/prune.c code\n+\t * can recognize any stale objdirs left behind by a crash and delete\n+\t * them.\n+\t */\n+\tstrbuf_addf(&t->path, \"%s/tmp_objdir-%s-XXXXXX\", get_object_directory(), prefix);\n \n \t/*\n \t * Grow the strbuf beyond any filename we expect to be placed in it.\n@@ -269,6 +279,13 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n \tif (!t)\n \t\treturn 0;\n \n+\tif (t->prev_odb) {\n+\t\tif (the_repository->objects->odb->will_destroy)\n+\t\t\tBUG(\"migrating an ODB that was marked for destruction\");\n+\t\trestore_primary_odb(t->prev_odb, t->path.buf);\n+\t\tt->prev_odb = NULL;\n+\t}\n+\n \tstrbuf_addbuf(&src, &t->path);\n \tstrbuf_addstr(&dst, get_object_directory());\n \n@@ -292,3 +309,10 @@ void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n {\n \tadd_to_alternates_memory(t->path.buf);\n }\n+\n+void tmp_objdir_replace_primary_odb(struct tmp_objdir *t, int will_destroy)\n+{\n+\tif (t->prev_odb)\n+\t\tBUG(\"the primary object database is already replaced\");\n+\tt->prev_odb = set_temporary_primary_odb(t->path.buf, will_destroy);\n+}\ndiff --git a/tmp-objdir.h b/tmp-objdir.h\nindex b1e45b4c75d..75754cbfba6 100644\n--- a/tmp-objdir.h\n+++ b/tmp-objdir.h\n@@ -10,7 +10,7 @@\n  *\n  * Example:\n  *\n- *\tstruct tmp_objdir *t = tmp_objdir_create();\n+ *\tstruct tmp_objdir *t = tmp_objdir_create(\"incoming\");\n  *\tif (!run_command_v_opt_cd_env(cmd, 0, NULL, tmp_objdir_env(t)) &&\n  *\t    !tmp_objdir_migrate(t))\n  *\t\tprintf(\"success!\\n\");\n@@ -22,9 +22,10 @@\n struct tmp_objdir;\n \n /*\n- * Create a new temporary object directory; returns NULL on failure.\n+ * Create a new temporary object directory with the specified prefix;\n+ * returns NULL on failure.\n  */\n-struct tmp_objdir *tmp_objdir_create(void);\n+struct tmp_objdir *tmp_objdir_create(const char *prefix);\n \n /*\n  * Return a list of environment strings, suitable for use with\n@@ -51,4 +52,11 @@ int tmp_objdir_destroy(struct tmp_objdir *);\n  */\n void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n \n+/*\n+ * Replaces the main object store in the current process with the temporary\n+ * object directory and makes the former main object store an alternate.\n+ * If will_destroy is nonzero, the object directory may not be migrated.\n+ */\n+void tmp_objdir_replace_primary_odb(struct tmp_objdir *, int will_destroy);\n+\n #endif /* TMP_OBJDIR_H */\n-- \ngitgitgadget\n\n"},{"id":"437915","messageId":"bc08513734099decc96e7e4d8747e6e93aa435c6.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 2/9] tmp-objdir: disable ref updates when replacing the primary odb","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:40Z","receivedAt":"2021-10-04T16:57:55Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen creating a subprocess with a temporary ODB, we set the\nGIT_QUARANTINE_ENVIRONMENT env var to tell child Git processes not\nto update refs, since the tmp-objdir may go away.\n\nIntroduce a similar mechanism for in-process temporary ODBs when\nwe call tmp_objdir_replace_primary_odb. Now both mechanisms set\nthe disable_ref_updates flag on the odb, which is queried by\nthe ref_transaction_prepare function.\n\nNote: This change adds an assumption that the state of\nthe_repository is relevant for any ref transaction that might\nbe initiated. Unwinding this assumption should be straightforward\nby saving the relevant repository to query in the transaction or\nthe ref_store.\n\nPeff's test case was invoking ref updates via the cachetextconv\nsetting. That particular code silently does nothing when a ref\nupdate is forbidden. See the call to notes_cache_put in\nfill_textconv where errors are ignored.\n\nReported-by: Jeff King <peff@peff.net>\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n environment.c  | 4 ++++\n object-file.c  | 6 ++++++\n object-store.h | 9 ++++++++-\n refs.c         | 2 +-\n repository.c   | 2 ++\n repository.h   | 1 +\n 6 files changed, 22 insertions(+), 2 deletions(-)\n\ndiff --git a/environment.c b/environment.c\nindex b4ba4fa22db..46ec5072c05 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -176,6 +176,10 @@ void setup_git_env(const char *git_dir)\n \targs.graft_file = getenv_safe(&to_free, GRAFT_ENVIRONMENT);\n \targs.index_file = getenv_safe(&to_free, INDEX_ENVIRONMENT);\n \targs.alternate_db = getenv_safe(&to_free, ALTERNATE_DB_ENVIRONMENT);\n+\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n+\t\targs.disable_ref_updates = 1;\n+\t}\n+\n \trepo_set_gitdir(the_repository, git_dir, &args);\n \tstrvec_clear(&to_free);\n \ndiff --git a/object-file.c b/object-file.c\nindex 990381abee5..f16441afb93 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -767,6 +767,12 @@ struct object_directory *set_temporary_primary_odb(const char *dir, int will_des\n \t */\n \tnew_odb = xcalloc(1, sizeof(*new_odb));\n \tnew_odb->path = xstrdup(dir);\n+\n+\t/*\n+\t * Disable ref updates while a temporary odb is active, since\n+\t * the objects in the database may roll back.\n+\t */\n+\tnew_odb->disable_ref_updates = 1;\n \tnew_odb->will_destroy = will_destroy;\n \tnew_odb->next = the_repository->objects->odb;\n \tthe_repository->objects->odb = new_odb;\ndiff --git a/object-store.h b/object-store.h\nindex 74b1b5872a6..bd53bdf2f2e 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -27,10 +27,17 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/*\n+\t * This is a temporary object store created by the tmp_objdir\n+\t * facility. Disable ref updates since the objects in the store\n+\t * might be discarded on rollback.\n+\t */\n+\tunsigned int disable_ref_updates : 1;\n+\n \t/*\n \t * This object store is ephemeral, so there is no need to fsync.\n \t */\n-\tint will_destroy;\n+\tunsigned int will_destroy : 1;\n \n \t/*\n \t * Path to the alternative object store. If this is a relative path,\ndiff --git a/refs.c b/refs.c\nindex 8b9f7c3a80a..7c182607dcf 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -2126,7 +2126,7 @@ int ref_transaction_prepare(struct ref_transaction *transaction,\n \t\tbreak;\n \t}\n \n-\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n+\tif (the_repository->objects->odb->disable_ref_updates) {\n \t\tstrbuf_addstr(err,\n \t\t\t      _(\"ref updates forbidden inside quarantine environment\"));\n \t\treturn -1;\ndiff --git a/repository.c b/repository.c\nindex 710a3b4bf87..18e0526da01 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -80,6 +80,8 @@ void repo_set_gitdir(struct repository *repo,\n \texpand_base_dir(&repo->objects->odb->path, o->object_dir,\n \t\t\trepo->commondir, \"objects\");\n \n+\trepo->objects->odb->disable_ref_updates = o->disable_ref_updates;\n+\n \tfree(repo->objects->alternate_db);\n \trepo->objects->alternate_db = xstrdup_or_null(o->alternate_db);\n \texpand_base_dir(&repo->graft_file, o->graft_file,\ndiff --git a/repository.h b/repository.h\nindex 3740c93bc0f..77316367d99 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -162,6 +162,7 @@ struct set_gitdir_args {\n \tconst char *graft_file;\n \tconst char *index_file;\n \tconst char *alternate_db;\n+\tint disable_ref_updates;\n };\n \n void repo_set_gitdir(struct repository *repo, const char *root,\n-- \ngitgitgadget\n\n"},{"id":"437916","messageId":"9335646ed91d3472191b8be233839b3a583c9718.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 3/9] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:41Z","receivedAt":"2021-10-04T16:58:03Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..6ae18401e04 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -277,21 +277,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"437920","messageId":"b9d3d87443266767f00e77c967bd77357fe50484.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 4/9] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:42Z","receivedAt":"2021-10-04T16:58:05Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nand macOS offer mechanisms to write data from the filesystem page cache\nwithout initiating a hardware flush. Linux has the sync_file_range API,\nwhich issues a pagecache writeback request reliably after version 5.2.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\nbatches up hardware flushes. It hooks into the bulk-checkin plugging and\nunplugging functionality and takes advantage of tmp-objdir.\n\nWhen the new mode is enabled do the following for each new object:\n1. Create the object in a tmp-objdir.\n2. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1. Issue an fsync against a dummy file to flush the hardware writeback\n   cache, which should by now have processed the tmp-objdir writes.\n2. Rename all of the tmp-objdir files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the case today,\n   but may be a good extension to those components.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 29 +++++++++++----\n Makefile                      |  6 ++++\n bulk-checkin.c                | 68 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  2 ++\n cache.h                       |  8 ++++-\n config.c                      |  7 +++-\n config.mak.uname              |  1 +\n configure.ac                  |  8 +++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 ++++\n object-file.c                 | 12 ++++++-\n wrapper.c                     | 44 +++++++++++++++++++++++\n write-or-die.c                |  2 +-\n 13 files changed, 185 insertions(+), 11 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..200b4d9f06e 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,29 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols loose objects in the object store, so updates to any\n+\trefs or the index may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each loose object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware. Git does not currently fsync parent directories for\n+  newly-added files, so some filesystems may still allow data to be lost on\n+  system crash.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write loose object data with a minimal set of FLUSH\n+  CACHE (or equivalent) commands sent to the storage controller. If the\n+  operating system interfaces are not available, this mode behaves the same as\n+  `true`. This mode is expected to be as safe as `true` on macOS for repos\n+  stored on HFS+ or APFS filesystems and on Windows for repos stored on NTFS or\n+  ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex a9f9b689f0c..313b3dc7cd6 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1874,6 +1876,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6ae18401e04..4deee1af46e 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,14 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n+static int needs_batch_fsync;\n+\n+static struct tmp_objdir *bulk_fsync_objdir;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -79,6 +85,34 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur.  The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware.\n+\t */\n+\n+\tif (needs_batch_fsync) {\n+\t\tstruct strbuf temp_path = STRBUF_INIT;\n+\t\tstruct tempfile *temp;\n+\n+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\t\ttemp = xmks_tempfile(temp_path.buf);\n+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\t\tdelete_tempfile(&temp);\n+\t\tstrbuf_release(&temp_path);\n+\t}\n+\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -273,6 +307,25 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void fsync_loose_object_bulk_checkin(int fd)\n+{\n+\tassert(fsync_object_files == FSYNC_OBJECT_FILES_BATCH);\n+\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (bulk_checkin_plugged &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\tif (!needs_batch_fsync)\n+\t\t\tneeds_batch_fsync = 1;\n+\t} else {\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -287,6 +340,19 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n+\n+\t/*\n+\t * A temporary object directory is used to hold the files\n+\t * while they are not fsynced.\n+\t */\n+\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH) {\n+\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\t\tif (!bulk_fsync_objdir)\n+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n+\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+\t}\n+\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -296,4 +362,6 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..08f292379b6 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,8 @@\n \n #include \"cache.h\"\n \n+void fsync_loose_object_bulk_checkin(int fd);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex f6295f3b048..1ed8137b5e6 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -984,7 +984,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum fsync_object_files_mode {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum fsync_object_files_mode fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex 2edf835262f..8315d020eeb 100644\n--- a/config.c\n+++ b/config.c\n@@ -1506,7 +1506,12 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (value && !strcmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse if (git_config_bool(var, value))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n+\t\telse\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 76516aaa9a5..e6d482fbcc6 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -53,6 +53,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex 46ec5072c05..0ace9fa2167 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,7 +42,7 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_object_files_mode fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 7c99eef6612..9daee873782 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1213,6 +1213,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex f16441afb93..b035d88f309 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1932,8 +1932,18 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n static void close_loose_object(int fd)\n {\n \tif (!the_repository->objects->odb->will_destroy) {\n-\t\tif (fsync_object_files)\n+\t\tswitch (fsync_object_files) {\n+\t\tcase FSYNC_OBJECT_FILES_OFF:\n+\t\t\tbreak;\n+\t\tcase FSYNC_OBJECT_FILES_ON:\n \t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\t\tbreak;\n+\t\tcase FSYNC_OBJECT_FILES_BATCH:\n+\t\t\tfsync_loose_object_bulk_checkin(fd);\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tBUG(\"Invalid fsync_object_files mode.\");\n+\t\t}\n \t}\n \n \tif (close(fd) != 0)\ndiff --git a/wrapper.c b/wrapper.c\nindex 7c6586af321..bb4f9f043ce 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -540,6 +540,50 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync(fd);\n+#endif\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 0b1ec8190b6..cc8291d9794 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"437917","messageId":"8df32eaaa9a9c41b6d9d49a50630952f14f33495.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 5/9] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:43Z","receivedAt":"2021-10-04T16:58:06Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds a win32 implementation for fsync_no_flush that is\ncalled git_fsync. The 'NtFlushBuffersFileEx' function being called is\navailable since Windows 8. If the function is not available, we\nreturn -1 and Git falls back to doing a full fsync.\n\nThe operating system is told to flush data only without a hardware\nflush primitive. A later full fsync will cause the metadata log\nto be flushed and then the disk cache to be flushed on NTFS and\nReFS. Other filesystems will treat this as a full flush operation.\n\nI added a new file here for this system call so as not to conflict with\ndownstream changes in the git-for-windows repository related to fscache.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.h                      |  3 +++\n compat/win32/flush.c                | 28 ++++++++++++++++++++++++++++\n config.mak.uname                    |  2 ++\n contrib/buildsystems/CMakeLists.txt |  3 ++-\n wrapper.c                           |  4 ++++\n 5 files changed, 39 insertions(+), 1 deletion(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..75324c24ee7\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex e6d482fbcc6..34c93314a50 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -451,6 +451,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -626,6 +627,7 @@ ifneq (,$(findstring MINGW,$(uname_S)))\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex 171b4124afe..b573a5ee122 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/wrapper.c b/wrapper.c\nindex bb4f9f043ce..1a1e2fba9c9 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -567,6 +567,10 @@ int git_fsync(int fd, enum fsync_action action)\n \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n #endif\n \n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n \t\terrno = ENOSYS;\n \t\treturn -1;\n \n-- \ngitgitgadget\n\n"},{"id":"437918","messageId":"15767270984a87d6112ece4057dd10bb22bad1eb.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 6/9] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:44Z","receivedAt":"2021-10-04T16:58:07Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to synchronize with the update-index\nprocess after passing it a file path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 187203e8bb5..dc7368bb1ee 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1088,6 +1089,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \tthe_index.updated_skipworktree = 1;\n \n+\t/* we might be adding many objects to the object database */\n+\tplug_bulk_checkin();\n+\n \t/*\n \t * Custom copy of parse_options() because we want to handle\n \t * filename arguments as they come.\n@@ -1168,6 +1172,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/* by now we must have added all of the new objects */\n+\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"437919","messageId":"e88bab809a2cd226170f511df4e7b57a55d87ff4.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 7/9] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:45Z","receivedAt":"2021-10-04T16:58:09Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling bulk-checkin when unpacking objects, we can take advantage\nof batched fsyncs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295ba..51eb4f7b531 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tplug_bulk_checkin();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"437921","messageId":"811d6d315098924b0f588c83c74dacef829eb147.1633366667.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 8/9] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:46Z","receivedAt":"2021-10-04T16:58:10Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added due\nto it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 36 ++++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 20 ++++++++++++++++++++\n t/t3903-stash.sh       | 14 ++++++++++++++\n t/t5300-pack-object.sh | 30 +++++++++++++++++++-----------\n 4 files changed, 89 insertions(+), 11 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a7de4ca8512\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,36 @@\n+# Helper to create files with unique contents\n+\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with unique\n+#\t\t\t\t\t contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\tlocal counter=0\n+\ttest_tick\n+\tlocal basedata=$test_tick\n+\n+\n+\trm -rf $basedir\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\"\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1))\n+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex 4086e1ebbc9..36049a53ff7 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -7,6 +7,8 @@ test_description='Test of git add, including the -- option.'\n \n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -33,6 +35,24 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+test_expect_success 'git add: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files2 &&\n+\tfind fsync-files2 ! -type d -print | xargs git -c core.fsyncobjectfiles=batch update-index --add -- &&\n+\trm -f fsynced_files2 &&\n+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n+\ttest_line_count = 8 fsynced_files2 &&\n+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 873aa56e359..2fc819e5584 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n diff_cmp () {\n \tfor i in \"$1\" \"$2\"\n@@ -1293,6 +1294,19 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n \texpected=\"stash.useBuiltin support has been removed\" &&\n \ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex e13a8842075..38663dc1393 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -162,23 +162,23 @@ test_expect_success 'pack-objects with bogus arguments' '\n \n check_unpack () {\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $2 init --bare git2 &&\n+\t(\n+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <obj-list >current &&\n+\tcmp obj-list current\n }\n \n test_expect_success 'unpack without delta' '\n \tcheck_unpack test-1-${packname_1}\n '\n \n+test_expect_success 'unpack without delta (core.fsyncobjectfiles=batch)' '\n+\tcheck_unpack test-1-${packname_1} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with REF_DELTA' '\n \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n \tcheck_deltas stderr -gt 0\n@@ -188,6 +188,10 @@ test_expect_success 'unpack with REF_DELTA' '\n \tcheck_unpack test-2-${packname_2}\n '\n \n+test_expect_success 'unpack with REF_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-2-${packname_2} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with OFS_DELTA' '\n \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n \t\t\t<obj-list 2>stderr) &&\n@@ -198,6 +202,10 @@ test_expect_success 'unpack with OFS_DELTA' '\n \tcheck_unpack test-3-${packname_3}\n '\n \n+test_expect_success 'unpack with OFS_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-3-${packname_3} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'compare delta flavors' '\n \tperl -e '\\''\n \t\tdefined($_ = -s $_) or die for @ARGV;\n-- \ngitgitgadget\n\n"},{"id":"437922","messageId":"f4fa20f591e580107b961aa1ca46d844603559d6.1633366668.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v8 9/9] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-10-04T16:57:47Z","receivedAt":"2021-10-04T16:58:12Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 43 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 46 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 89 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..e93c08a2e70\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..c9fcd0c03eb\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n"},{"id":"441224","messageId":"71817cccfb9573929723a34fe9fb0db72498d4b9.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 2/9] tmp-objdir: disable ref updates when replacing the primary odb","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:50:56Z","receivedAt":"2021-11-16T03:24:21Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen creating a subprocess with a temporary ODB, we set the\nGIT_QUARANTINE_ENVIRONMENT env var to tell child Git processes not\nto update refs, since the tmp-objdir may go away.\n\nIntroduce a similar mechanism for in-process temporary ODBs when\nwe call tmp_objdir_replace_primary_odb. Now both mechanisms set\nthe disable_ref_updates flag on the odb, which is queried by\nthe ref_transaction_prepare function.\n\nNote: This change adds an assumption that the state of\nthe_repository is relevant for any ref transaction that might\nbe initiated. Unwinding this assumption should be straightforward\nby saving the relevant repository to query in the transaction or\nthe ref_store.\n\nPeff's test case was invoking ref updates via the cachetextconv\nsetting. That particular code silently does nothing when a ref\nupdate is forbidden. See the call to notes_cache_put in\nfill_textconv where errors are ignored.\n\nReported-by: Jeff King <peff@peff.net>\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n environment.c  | 4 ++++\n object-file.c  | 6 ++++++\n object-store.h | 9 ++++++++-\n refs.c         | 2 +-\n repository.c   | 2 ++\n repository.h   | 1 +\n 6 files changed, 22 insertions(+), 2 deletions(-)\n\ndiff --git a/environment.c b/environment.c\nindex 342400fcaad..2701dfeeec8 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -169,6 +169,10 @@ void setup_git_env(const char *git_dir)\n \targs.graft_file = getenv_safe(&to_free, GRAFT_ENVIRONMENT);\n \targs.index_file = getenv_safe(&to_free, INDEX_ENVIRONMENT);\n \targs.alternate_db = getenv_safe(&to_free, ALTERNATE_DB_ENVIRONMENT);\n+\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n+\t\targs.disable_ref_updates = 1;\n+\t}\n+\n \trepo_set_gitdir(the_repository, git_dir, &args);\n \tstrvec_clear(&to_free);\n \ndiff --git a/object-file.c b/object-file.c\nindex 0b6a61aeaff..659ef7623ff 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -699,6 +699,12 @@ struct object_directory *set_temporary_primary_odb(const char *dir, int will_des\n \t */\n \tnew_odb = xcalloc(1, sizeof(*new_odb));\n \tnew_odb->path = xstrdup(dir);\n+\n+\t/*\n+\t * Disable ref updates while a temporary odb is active, since\n+\t * the objects in the database may roll back.\n+\t */\n+\tnew_odb->disable_ref_updates = 1;\n \tnew_odb->will_destroy = will_destroy;\n \tnew_odb->next = the_repository->objects->odb;\n \tthe_repository->objects->odb = new_odb;\ndiff --git a/object-store.h b/object-store.h\nindex cb173e69392..9ae9262c340 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -27,10 +27,17 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/*\n+\t * This is a temporary object store created by the tmp_objdir\n+\t * facility. Disable ref updates since the objects in the store\n+\t * might be discarded on rollback.\n+\t */\n+\tunsigned int disable_ref_updates : 1;\n+\n \t/*\n \t * This object store is ephemeral, so there is no need to fsync.\n \t */\n-\tint will_destroy;\n+\tunsigned int will_destroy : 1;\n \n \t/*\n \t * Path to the alternative object store. If this is a relative path,\ndiff --git a/refs.c b/refs.c\nindex d7cc0a23a3b..27ec7d1fc64 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -2137,7 +2137,7 @@ int ref_transaction_prepare(struct ref_transaction *transaction,\n \t\tbreak;\n \t}\n \n-\tif (getenv(GIT_QUARANTINE_ENVIRONMENT)) {\n+\tif (the_repository->objects->odb->disable_ref_updates) {\n \t\tstrbuf_addstr(err,\n \t\t\t      _(\"ref updates forbidden inside quarantine environment\"));\n \t\treturn -1;\ndiff --git a/repository.c b/repository.c\nindex c5b90ba93ea..dce8e35ac20 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -80,6 +80,8 @@ void repo_set_gitdir(struct repository *repo,\n \texpand_base_dir(&repo->objects->odb->path, o->object_dir,\n \t\t\trepo->commondir, \"objects\");\n \n+\trepo->objects->odb->disable_ref_updates = o->disable_ref_updates;\n+\n \tfree(repo->objects->alternate_db);\n \trepo->objects->alternate_db = xstrdup_or_null(o->alternate_db);\n \texpand_base_dir(&repo->graft_file, o->graft_file,\ndiff --git a/repository.h b/repository.h\nindex a057653981c..7c04e99ac5c 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -158,6 +158,7 @@ struct set_gitdir_args {\n \tconst char *graft_file;\n \tconst char *index_file;\n \tconst char *alternate_db;\n+\tint disable_ref_updates;\n };\n \n void repo_set_gitdir(struct repository *repo, const char *root,\n-- \ngitgitgadget\n\n"},{"id":"441225","messageId":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v8.git.git.1633366667.gitgitgadget@gmail.com","subject":"[PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:50:54Z","receivedAt":"2021-11-16T03:24:22Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"Thanks to everyone for review so far!\n\nThis series shares the base tmp-objdir patches with my merged version of\nElijah Newren's remerge-diff series at:\nhttps://github.com/neerajsi-msft/git/tree/neerajsi/remerge-diff.\n\nChanges between v8 and v9 [1]:\n\n * Rebased onto master at tag v2.34.0\n * Fixed git-prune bug when trying to clean up multiple cruft directories.\n * Preserve the tmp-objdir around update_relative_gitdir, which is called by\n   setup_work_tree through the chdir_notify mechanism.\n * Per [2], I'm leaving the fsyncObjectFiles configuration as is with\n   'true', 'false', and 'batch'. This makes using old and new versions of\n   git with 'batch' mode a little trickier, but hopefully people will\n   generally be moving forward in versions.\n\n[1] See\nhttps://lore.kernel.org/git/pull.1067.git.1635287730.gitgitgadget@gmail.com/\n[2] https://lore.kernel.org/git/xmqqh7cimuxt.fsf@gitster.g/\n\nChanges between v7 and v8:\n\n * Dropped the tmp-objdir patch to avoid renaming in a quarantine/temporary\n   objdir, as suggested by Jeff King. This wasn't a good idea because we\n   don't really know that there's only a single reader/writer. Avoiding the\n   rename was a relatively minor perf optimization so it's okay to drop.\n\n * Added disable_ref_updates logic (as a flag on the odb) which is set when\n   we're in a quarantine or when a tmp objdir is active. I believe this\n   roughly follows the strategy suggested by Jeff King.\n\nNeeraj Singh (9):\n  tmp-objdir: new API for creating temporary writable databases\n  tmp-objdir: disable ref updates when replacing the primary odb\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncobjectfiles: batched disk flushes\n  core.fsyncobjectfiles: add windows support for batch mode\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsyncobjectfiles: tests for batch mode\n  core.fsyncobjectfiles: performance tests for add and stash\n\n Documentation/config/core.txt       | 29 ++++++++--\n Makefile                            |  6 ++\n builtin/prune.c                     | 23 ++++++--\n builtin/receive-pack.c              |  2 +-\n builtin/unpack-objects.c            |  3 +\n builtin/update-index.c              |  6 ++\n bulk-checkin.c                      | 90 +++++++++++++++++++++++++----\n bulk-checkin.h                      |  2 +\n cache.h                             |  8 ++-\n compat/mingw.h                      |  3 +\n compat/win32/flush.c                | 28 +++++++++\n config.c                            |  7 ++-\n config.mak.uname                    |  3 +\n configure.ac                        |  8 +++\n contrib/buildsystems/CMakeLists.txt |  3 +-\n environment.c                       | 11 +++-\n git-compat-util.h                   |  7 +++\n object-file.c                       | 60 ++++++++++++++++++-\n object-store.h                      | 26 +++++++++\n object.c                            |  2 +-\n refs.c                              |  2 +-\n repository.c                        |  2 +\n repository.h                        |  1 +\n t/lib-unique-files.sh               | 36 ++++++++++++\n t/perf/p3700-add.sh                 | 43 ++++++++++++++\n t/perf/p3900-stash.sh               | 46 +++++++++++++++\n t/t3700-add.sh                      | 20 +++++++\n t/t3903-stash.sh                    | 14 +++++\n t/t5300-pack-object.sh              | 30 ++++++----\n tmp-objdir.c                        | 55 +++++++++++++++++-\n tmp-objdir.h                        | 29 +++++++++-\n wrapper.c                           | 48 +++++++++++++++\n write-or-die.c                      |  2 +-\n 33 files changed, 608 insertions(+), 47 deletions(-)\n create mode 100644 compat/win32/flush.c\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: cd3e606211bb1cf8bc57f7d76bab98cc17a150bc\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1076%2Fneerajsi-msft%2Fneerajsi%2Fbulk-fsync-object-files-v9\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1076/neerajsi-msft/neerajsi/bulk-fsync-object-files-v9\nPull-Request: https://github.com/git/git/pull/1076\n\nRange-diff vs v8:\n\n  1:  f03797fd80d !  1:  6b27afa60e0 tmp-objdir: new API for creating temporary writable databases\n     @@ Commit message\n          Based-on-patch-by: Elijah Newren <newren@gmail.com>\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n     +    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n      \n       ## builtin/prune.c ##\n      @@ builtin/prune.c: static int show_only;\n     @@ builtin/prune.c: static int prune_tmp_file(const char *fullpath)\n      +\t\tif (show_only || verbose)\n      +\t\t\tprintf(\"Removing stale temporary directory %s\\n\", fullpath);\n      +\t\tif (!show_only) {\n     ++\t\t\tstrbuf_reset(&remove_dir_buf);\n      +\t\t\tstrbuf_addstr(&remove_dir_buf, fullpath);\n      +\t\t\tremove_dir_recursively(&remove_dir_buf, 0);\n      +\t\t}\n     @@ builtin/receive-pack.c: static const char *unpack(int err_fd, struct shallow_inf\n       \t\tif (err_fd > 0)\n       \t\t\tclose(err_fd);\n      \n     + ## environment.c ##\n     +@@\n     + #include \"commit.h\"\n     + #include \"strvec.h\"\n     + #include \"object-store.h\"\n     ++#include \"tmp-objdir.h\"\n     + #include \"chdir-notify.h\"\n     + #include \"shallow.h\"\n     + \n     +@@ environment.c: static void update_relative_gitdir(const char *name,\n     + \t\t\t\t   void *data)\n     + {\n     + \tchar *path = reparent_relative_path(old_cwd, new_cwd, get_git_dir());\n     ++\tstruct tmp_objdir *tmp_objdir = tmp_objdir_unapply_primary_odb();\n     + \ttrace_printf_key(&trace_setup_key,\n     + \t\t\t \"setup: move $GIT_DIR to '%s'\",\n     + \t\t\t path);\n     ++\n     + \tset_git_dir_1(path);\n     ++\tif (tmp_objdir)\n     ++\t\ttmp_objdir_reapply_primary_odb(tmp_objdir, old_cwd, new_cwd);\n     + \tfree(path);\n     + }\n     + \n     +\n       ## object-file.c ##\n      @@ object-file.c: void add_to_alternates_memory(const char *reference)\n       \t\t\t     '\\n', NULL, 0);\n     @@ object.c: struct raw_object_store *raw_object_store_new(void)\n       \todb_clear_loose_cache(odb);\n      \n       ## tmp-objdir.c ##\n     +@@\n     + #include \"cache.h\"\n     + #include \"tmp-objdir.h\"\n     ++#include \"chdir-notify.h\"\n     + #include \"dir.h\"\n     + #include \"sigchain.h\"\n     + #include \"string-list.h\"\n      @@\n       struct tmp_objdir {\n       \tstruct strbuf path;\n       \tstruct strvec env;\n      +\tstruct object_directory *prev_odb;\n     ++\tint will_destroy;\n       };\n       \n       /*\n     @@ tmp-objdir.c: void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n      +\tif (t->prev_odb)\n      +\t\tBUG(\"the primary object database is already replaced\");\n      +\tt->prev_odb = set_temporary_primary_odb(t->path.buf, will_destroy);\n     ++\tt->will_destroy = will_destroy;\n     ++}\n     ++\n     ++struct tmp_objdir *tmp_objdir_unapply_primary_odb(void)\n     ++{\n     ++\tif (!the_tmp_objdir || !the_tmp_objdir->prev_odb)\n     ++\t\treturn NULL;\n     ++\n     ++\trestore_primary_odb(the_tmp_objdir->prev_odb, the_tmp_objdir->path.buf);\n     ++\tthe_tmp_objdir->prev_odb = NULL;\n     ++\treturn the_tmp_objdir;\n     ++}\n     ++\n     ++void tmp_objdir_reapply_primary_odb(struct tmp_objdir *t, const char *old_cwd,\n     ++\t\tconst char *new_cwd)\n     ++{\n     ++\tchar *path;\n     ++\n     ++\tpath = reparent_relative_path(old_cwd, new_cwd, t->path.buf);\n     ++\tstrbuf_reset(&t->path);\n     ++\tstrbuf_addstr(&t->path, path);\n     ++\tfree(path);\n     ++\ttmp_objdir_replace_primary_odb(t, t->will_destroy);\n      +}\n      \n       ## tmp-objdir.h ##\n     @@ tmp-objdir.h: int tmp_objdir_destroy(struct tmp_objdir *);\n      + * If will_destroy is nonzero, the object directory may not be migrated.\n      + */\n      +void tmp_objdir_replace_primary_odb(struct tmp_objdir *, int will_destroy);\n     ++\n     ++/*\n     ++ * If the primary object database was replaced by a temporary object directory,\n     ++ * restore it to its original value while keeping the directory contents around.\n     ++ * Returns NULL if the primary object database was not replaced.\n     ++ */\n     ++struct tmp_objdir *tmp_objdir_unapply_primary_odb(void);\n     ++\n     ++/*\n     ++ * Reapplies the former primary temporary object database, after protentially\n     ++ * changing its relative path.\n     ++ */\n     ++void tmp_objdir_reapply_primary_odb(struct tmp_objdir *, const char *old_cwd,\n     ++\t\tconst char *new_cwd);\n     ++\n      +\n       #endif /* TMP_OBJDIR_H */\n  2:  bc085137340 !  2:  71817cccfb9 tmp-objdir: disable ref updates when replacing the primary odb\n     @@ Commit message\n          Reported-by: Jeff King <peff@peff.net>\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n     +    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n      \n       ## environment.c ##\n      @@ environment.c: void setup_git_env(const char *git_dir)\n  3:  9335646ed91 =  3:  8fd1ca4c00a bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  4:  b9d3d874432 =  4:  e1747ce00af core.fsyncobjectfiles: batched disk flushes\n  5:  8df32eaaa9a !  5:  951a559874e core.fsyncobjectfiles: add windows support for batch mode\n     @@ config.mak.uname: endif\n       \t\tcompat/win32/path-utils.o \\\n       \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n       \t\tcompat/win32/trace2_win32_process_info.o \\\n     -@@ config.mak.uname: ifneq (,$(findstring MINGW,$(uname_S)))\n     +@@ config.mak.uname: ifeq ($(uname_S),MINGW)\n       \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n       \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n       \t\tcompat/win32/trace2_win32_process_info.o \\\n  6:  15767270984 =  6:  4a40fd4a29a update-index: use the bulk-checkin infrastructure\n  7:  e88bab809a2 =  7:  cfc6a347d08 unpack-objects: use the bulk-checkin infrastructure\n  8:  811d6d31509 !  8:  270c24827d0 core.fsyncobjectfiles: tests for batch mode\n     @@ t/lib-unique-files.sh (new)\n      \n       ## t/t3700-add.sh ##\n      @@ t/t3700-add.sh: test_description='Test of git add, including the -- option.'\n     - \n     + TEST_PASSES_SANITIZE_LEAK=true\n       . ./test-lib.sh\n       \n      +. $TEST_DIRECTORY/lib-unique-files.sh\n  9:  f4fa20f591e =  9:  12d99641f4c core.fsyncobjectfiles: performance tests for add and stash\n\n-- \ngitgitgadget\n"},{"id":"441226","messageId":"6b27afa60e05c6c0b7752f1bcf6629c446ede520.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 1/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:50:55Z","receivedAt":"2021-11-16T03:24:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe tmp_objdir API provides the ability to create temporary object\ndirectories, but was designed with the goal of having subprocesses\naccess these object stores, followed by the main process migrating\nobjects from it to the main object store or just deleting it.  The\nsubprocesses would view it as their primary datastore and write to it.\n\nHere we add the tmp_objdir_replace_primary_odb function that replaces\nthe current process's writable \"main\" object directory with the\nspecified one. The previous main object directory is restored in either\ntmp_objdir_migrate or tmp_objdir_destroy.\n\nFor the --remerge-diff usecase, add a new `will_destroy` flag in `struct\nobject_database` to mark ephemeral object databases that do not require\nfsync durability.\n\nAdd 'git prune' support for removing temporary object databases, and\nmake sure that they have a name starting with tmp_ and containing an\noperation-specific name.\n\nBased-on-patch-by: Elijah Newren <newren@gmail.com>\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n builtin/prune.c        | 23 +++++++++++++++---\n builtin/receive-pack.c |  2 +-\n environment.c          |  5 ++++\n object-file.c          | 44 +++++++++++++++++++++++++++++++--\n object-store.h         | 19 +++++++++++++++\n object.c               |  2 +-\n tmp-objdir.c           | 55 +++++++++++++++++++++++++++++++++++++++---\n tmp-objdir.h           | 29 +++++++++++++++++++---\n 8 files changed, 165 insertions(+), 14 deletions(-)\n\ndiff --git a/builtin/prune.c b/builtin/prune.c\nindex 485c9a3c56f..a76e6a5f0e8 100644\n--- a/builtin/prune.c\n+++ b/builtin/prune.c\n@@ -18,6 +18,7 @@ static int show_only;\n static int verbose;\n static timestamp_t expire;\n static int show_progress = -1;\n+static struct strbuf remove_dir_buf = STRBUF_INIT;\n \n static int prune_tmp_file(const char *fullpath)\n {\n@@ -26,10 +27,20 @@ static int prune_tmp_file(const char *fullpath)\n \t\treturn error(\"Could not stat '%s'\", fullpath);\n \tif (st.st_mtime > expire)\n \t\treturn 0;\n-\tif (show_only || verbose)\n-\t\tprintf(\"Removing stale temporary file %s\\n\", fullpath);\n-\tif (!show_only)\n-\t\tunlink_or_warn(fullpath);\n+\tif (S_ISDIR(st.st_mode)) {\n+\t\tif (show_only || verbose)\n+\t\t\tprintf(\"Removing stale temporary directory %s\\n\", fullpath);\n+\t\tif (!show_only) {\n+\t\t\tstrbuf_reset(&remove_dir_buf);\n+\t\t\tstrbuf_addstr(&remove_dir_buf, fullpath);\n+\t\t\tremove_dir_recursively(&remove_dir_buf, 0);\n+\t\t}\n+\t} else {\n+\t\tif (show_only || verbose)\n+\t\t\tprintf(\"Removing stale temporary file %s\\n\", fullpath);\n+\t\tif (!show_only)\n+\t\t\tunlink_or_warn(fullpath);\n+\t}\n \treturn 0;\n }\n \n@@ -97,6 +108,9 @@ static int prune_cruft(const char *basename, const char *path, void *data)\n \n static int prune_subdir(unsigned int nr, const char *path, void *data)\n {\n+\tif (verbose)\n+\t\tprintf(\"Removing directory %s\\n\", path);\n+\n \tif (!show_only)\n \t\trmdir(path);\n \treturn 0;\n@@ -184,5 +198,6 @@ int cmd_prune(int argc, const char **argv, const char *prefix)\n \t\tprune_shallow(show_only ? PRUNE_SHOW_ONLY : 0);\n \t}\n \n+\tstrbuf_release(&remove_dir_buf);\n \treturn 0;\n }\ndiff --git a/builtin/receive-pack.c b/builtin/receive-pack.c\nindex 49b846d9605..8815e24cde5 100644\n--- a/builtin/receive-pack.c\n+++ b/builtin/receive-pack.c\n@@ -2213,7 +2213,7 @@ static const char *unpack(int err_fd, struct shallow_info *si)\n \t\tstrvec_push(&child.args, alt_shallow_file);\n \t}\n \n-\ttmp_objdir = tmp_objdir_create();\n+\ttmp_objdir = tmp_objdir_create(\"incoming\");\n \tif (!tmp_objdir) {\n \t\tif (err_fd > 0)\n \t\t\tclose(err_fd);\ndiff --git a/environment.c b/environment.c\nindex 9da7f3c1a19..342400fcaad 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -17,6 +17,7 @@\n #include \"commit.h\"\n #include \"strvec.h\"\n #include \"object-store.h\"\n+#include \"tmp-objdir.h\"\n #include \"chdir-notify.h\"\n #include \"shallow.h\"\n \n@@ -331,10 +332,14 @@ static void update_relative_gitdir(const char *name,\n \t\t\t\t   void *data)\n {\n \tchar *path = reparent_relative_path(old_cwd, new_cwd, get_git_dir());\n+\tstruct tmp_objdir *tmp_objdir = tmp_objdir_unapply_primary_odb();\n \ttrace_printf_key(&trace_setup_key,\n \t\t\t \"setup: move $GIT_DIR to '%s'\",\n \t\t\t path);\n+\n \tset_git_dir_1(path);\n+\tif (tmp_objdir)\n+\t\ttmp_objdir_reapply_primary_odb(tmp_objdir, old_cwd, new_cwd);\n \tfree(path);\n }\n \ndiff --git a/object-file.c b/object-file.c\nindex c3d866a287e..0b6a61aeaff 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -683,6 +683,43 @@ void add_to_alternates_memory(const char *reference)\n \t\t\t     '\\n', NULL, 0);\n }\n \n+struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy)\n+{\n+\tstruct object_directory *new_odb;\n+\n+\t/*\n+\t * Make sure alternates are initialized, or else our entry may be\n+\t * overwritten when they are.\n+\t */\n+\tprepare_alt_odb(the_repository);\n+\n+\t/*\n+\t * Make a new primary odb and link the old primary ODB in as an\n+\t * alternate\n+\t */\n+\tnew_odb = xcalloc(1, sizeof(*new_odb));\n+\tnew_odb->path = xstrdup(dir);\n+\tnew_odb->will_destroy = will_destroy;\n+\tnew_odb->next = the_repository->objects->odb;\n+\tthe_repository->objects->odb = new_odb;\n+\treturn new_odb->next;\n+}\n+\n+void restore_primary_odb(struct object_directory *restore_odb, const char *old_path)\n+{\n+\tstruct object_directory *cur_odb = the_repository->objects->odb;\n+\n+\tif (strcmp(old_path, cur_odb->path))\n+\t\tBUG(\"expected %s as primary object store; found %s\",\n+\t\t    old_path, cur_odb->path);\n+\n+\tif (cur_odb->next != restore_odb)\n+\t\tBUG(\"we expect the old primary object store to be the first alternate\");\n+\n+\tthe_repository->objects->odb = restore_odb;\n+\tfree_object_directory(cur_odb);\n+}\n+\n /*\n  * Compute the exact path an alternate is at and returns it. In case of\n  * error NULL is returned and the human readable error is added to `err`\n@@ -1809,8 +1846,11 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n /* Finalize a file on disk, and close it. */\n static void close_loose_object(int fd)\n {\n-\tif (fsync_object_files)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\tif (!the_repository->objects->odb->will_destroy) {\n+\t\tif (fsync_object_files)\n+\t\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+\n \tif (close(fd) != 0)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4be..cb173e69392 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -27,6 +27,11 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/*\n+\t * This object store is ephemeral, so there is no need to fsync.\n+\t */\n+\tint will_destroy;\n+\n \t/*\n \t * Path to the alternative object store. If this is a relative path,\n \t * it is relative to the current working directory.\n@@ -58,6 +63,17 @@ void add_to_alternates_file(const char *dir);\n  */\n void add_to_alternates_memory(const char *dir);\n \n+/*\n+ * Replace the current writable object directory with the specified temporary\n+ * object directory; returns the former primary object directory.\n+ */\n+struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy);\n+\n+/*\n+ * Restore a previous ODB replaced by set_temporary_main_odb.\n+ */\n+void restore_primary_odb(struct object_directory *restore_odb, const char *old_path);\n+\n /*\n  * Populate and return the loose object cache array corresponding to the\n  * given object ID.\n@@ -68,6 +84,9 @@ struct oidtree *odb_loose_cache(struct object_directory *odb,\n /* Empty the loose object cache for the specified object directory. */\n void odb_clear_loose_cache(struct object_directory *odb);\n \n+/* Clear and free the specified object directory */\n+void free_object_directory(struct object_directory *odb);\n+\n struct packed_git {\n \tstruct hashmap_entry packmap_ent;\n \tstruct packed_git *next;\ndiff --git a/object.c b/object.c\nindex 23a24e678a8..048f96a260e 100644\n--- a/object.c\n+++ b/object.c\n@@ -513,7 +513,7 @@ struct raw_object_store *raw_object_store_new(void)\n \treturn o;\n }\n \n-static void free_object_directory(struct object_directory *odb)\n+void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex b8d880e3626..3d38eeab66b 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -1,5 +1,6 @@\n #include \"cache.h\"\n #include \"tmp-objdir.h\"\n+#include \"chdir-notify.h\"\n #include \"dir.h\"\n #include \"sigchain.h\"\n #include \"string-list.h\"\n@@ -11,6 +12,8 @@\n struct tmp_objdir {\n \tstruct strbuf path;\n \tstruct strvec env;\n+\tstruct object_directory *prev_odb;\n+\tint will_destroy;\n };\n \n /*\n@@ -38,6 +41,9 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \tif (t == the_tmp_objdir)\n \t\tthe_tmp_objdir = NULL;\n \n+\tif (!on_signal && t->prev_odb)\n+\t\trestore_primary_odb(t->prev_odb, t->path.buf);\n+\n \t/*\n \t * This may use malloc via strbuf_grow(), but we should\n \t * have pre-grown t->path sufficiently so that this\n@@ -52,6 +58,7 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n \t */\n \tif (!on_signal)\n \t\ttmp_objdir_free(t);\n+\n \treturn err;\n }\n \n@@ -121,7 +128,7 @@ static int setup_tmp_objdir(const char *root)\n \treturn ret;\n }\n \n-struct tmp_objdir *tmp_objdir_create(void)\n+struct tmp_objdir *tmp_objdir_create(const char *prefix)\n {\n \tstatic int installed_handlers;\n \tstruct tmp_objdir *t;\n@@ -129,11 +136,16 @@ struct tmp_objdir *tmp_objdir_create(void)\n \tif (the_tmp_objdir)\n \t\tBUG(\"only one tmp_objdir can be used at a time\");\n \n-\tt = xmalloc(sizeof(*t));\n+\tt = xcalloc(1, sizeof(*t));\n \tstrbuf_init(&t->path, 0);\n \tstrvec_init(&t->env);\n \n-\tstrbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n+\t/*\n+\t * Use a string starting with tmp_ so that the builtin/prune.c code\n+\t * can recognize any stale objdirs left behind by a crash and delete\n+\t * them.\n+\t */\n+\tstrbuf_addf(&t->path, \"%s/tmp_objdir-%s-XXXXXX\", get_object_directory(), prefix);\n \n \t/*\n \t * Grow the strbuf beyond any filename we expect to be placed in it.\n@@ -269,6 +281,13 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n \tif (!t)\n \t\treturn 0;\n \n+\tif (t->prev_odb) {\n+\t\tif (the_repository->objects->odb->will_destroy)\n+\t\t\tBUG(\"migrating an ODB that was marked for destruction\");\n+\t\trestore_primary_odb(t->prev_odb, t->path.buf);\n+\t\tt->prev_odb = NULL;\n+\t}\n+\n \tstrbuf_addbuf(&src, &t->path);\n \tstrbuf_addstr(&dst, get_object_directory());\n \n@@ -292,3 +311,33 @@ void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n {\n \tadd_to_alternates_memory(t->path.buf);\n }\n+\n+void tmp_objdir_replace_primary_odb(struct tmp_objdir *t, int will_destroy)\n+{\n+\tif (t->prev_odb)\n+\t\tBUG(\"the primary object database is already replaced\");\n+\tt->prev_odb = set_temporary_primary_odb(t->path.buf, will_destroy);\n+\tt->will_destroy = will_destroy;\n+}\n+\n+struct tmp_objdir *tmp_objdir_unapply_primary_odb(void)\n+{\n+\tif (!the_tmp_objdir || !the_tmp_objdir->prev_odb)\n+\t\treturn NULL;\n+\n+\trestore_primary_odb(the_tmp_objdir->prev_odb, the_tmp_objdir->path.buf);\n+\tthe_tmp_objdir->prev_odb = NULL;\n+\treturn the_tmp_objdir;\n+}\n+\n+void tmp_objdir_reapply_primary_odb(struct tmp_objdir *t, const char *old_cwd,\n+\t\tconst char *new_cwd)\n+{\n+\tchar *path;\n+\n+\tpath = reparent_relative_path(old_cwd, new_cwd, t->path.buf);\n+\tstrbuf_reset(&t->path);\n+\tstrbuf_addstr(&t->path, path);\n+\tfree(path);\n+\ttmp_objdir_replace_primary_odb(t, t->will_destroy);\n+}\ndiff --git a/tmp-objdir.h b/tmp-objdir.h\nindex b1e45b4c75d..a3145051f25 100644\n--- a/tmp-objdir.h\n+++ b/tmp-objdir.h\n@@ -10,7 +10,7 @@\n  *\n  * Example:\n  *\n- *\tstruct tmp_objdir *t = tmp_objdir_create();\n+ *\tstruct tmp_objdir *t = tmp_objdir_create(\"incoming\");\n  *\tif (!run_command_v_opt_cd_env(cmd, 0, NULL, tmp_objdir_env(t)) &&\n  *\t    !tmp_objdir_migrate(t))\n  *\t\tprintf(\"success!\\n\");\n@@ -22,9 +22,10 @@\n struct tmp_objdir;\n \n /*\n- * Create a new temporary object directory; returns NULL on failure.\n+ * Create a new temporary object directory with the specified prefix;\n+ * returns NULL on failure.\n  */\n-struct tmp_objdir *tmp_objdir_create(void);\n+struct tmp_objdir *tmp_objdir_create(const char *prefix);\n \n /*\n  * Return a list of environment strings, suitable for use with\n@@ -51,4 +52,26 @@ int tmp_objdir_destroy(struct tmp_objdir *);\n  */\n void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n \n+/*\n+ * Replaces the main object store in the current process with the temporary\n+ * object directory and makes the former main object store an alternate.\n+ * If will_destroy is nonzero, the object directory may not be migrated.\n+ */\n+void tmp_objdir_replace_primary_odb(struct tmp_objdir *, int will_destroy);\n+\n+/*\n+ * If the primary object database was replaced by a temporary object directory,\n+ * restore it to its original value while keeping the directory contents around.\n+ * Returns NULL if the primary object database was not replaced.\n+ */\n+struct tmp_objdir *tmp_objdir_unapply_primary_odb(void);\n+\n+/*\n+ * Reapplies the former primary temporary object database, after protentially\n+ * changing its relative path.\n+ */\n+void tmp_objdir_reapply_primary_odb(struct tmp_objdir *, const char *old_cwd,\n+\t\tconst char *new_cwd);\n+\n+\n #endif /* TMP_OBJDIR_H */\n-- \ngitgitgadget\n\n"},{"id":"441227","messageId":"8fd1ca4c00aa94dbb75a34359727bdbd3adccf60.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 3/9] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:50:57Z","receivedAt":"2021-11-16T03:24:27Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_state'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..6ae18401e04 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -277,21 +277,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"441228","messageId":"e1747ce00af7ab3170a69955b07d995d5321d6f3.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 4/9] core.fsyncobjectfiles: batched disk flushes","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:50:58Z","receivedAt":"2021-11-16T03:24:43Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with core.fsyncObjectFiles set to\ntrue, the cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. Fortunately, Windows,\nand macOS offer mechanisms to write data from the filesystem page cache\nwithout initiating a hardware flush. Linux has the sync_file_range API,\nwhich issues a pagecache writeback request reliably after version 5.2.\n\nThis patch introduces a new 'core.fsyncObjectFiles = batch' option that\nbatches up hardware flushes. It hooks into the bulk-checkin plugging and\nunplugging functionality and takes advantage of tmp-objdir.\n\nWhen the new mode is enabled do the following for each new object:\n1. Create the object in a tmp-objdir.\n2. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1. Issue an fsync against a dummy file to flush the hardware writeback\n   cache, which should by now have processed the tmp-objdir writes.\n2. Rename all of the tmp-objdir files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the case today,\n   but may be a good extension to those components.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nThis change also updates the macOS code to trigger a real hardware flush\nvia fnctl(fd, F_FULLFSYNC) when fsync_or_die is called. Previously, on\nmacOS there was no guarantee of durability since a simple fsync(2) call\ndoes not flush any hardware caches.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\t  This number is from a patch later in the series.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\ncore.fsyncObjectFiles | Linux | Mac   | Windows\n----------------------|-------|-------|--------\n                false | 0.06  |  0.35 | 0.61\n                true  | 1.88  | 11.18 | 2.47\n                batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt | 29 +++++++++++----\n Makefile                      |  6 ++++\n bulk-checkin.c                | 68 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  2 ++\n cache.h                       |  8 ++++-\n config.c                      |  7 +++-\n config.mak.uname              |  1 +\n configure.ac                  |  8 +++++\n environment.c                 |  2 +-\n git-compat-util.h             |  7 ++++\n object-file.c                 | 12 ++++++-\n wrapper.c                     | 44 +++++++++++++++++++++++\n write-or-die.c                |  2 +-\n 13 files changed, 185 insertions(+), 11 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..200b4d9f06e 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -548,12 +548,29 @@ core.whitespace::\n   errors. The default tab width is 8. Allowed values are 1 to 63.\n \n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-+\n-This is a total waste of time and effort on a filesystem that orders\n-data writes properly, but can be useful for filesystems that do not use\n-journalling (traditional UNIX filesystems) or that only journal metadata\n-and not file contents (OS X's HFS+, or Linux ext3 with \"data=writeback\").\n+\tA value indicating the level of effort Git will expend in\n+\ttrying to make objects added to the repo durable in the event\n+\tof an unclean system shutdown. This setting currently only\n+\tcontrols loose objects in the object store, so updates to any\n+\trefs or the index may not be equally durable.\n++\n+* `false` allows data to remain in file system caches according to\n+  operating system policy, whence it may be lost if the system loses power\n+  or crashes.\n+* `true` triggers a data integrity flush for each loose object added to the\n+  object store. This is the safest setting that is likely to ensure durability\n+  across all operating systems and file systems that honor the 'fsync' system\n+  call. However, this setting comes with a significant performance cost on\n+  common hardware. Git does not currently fsync parent directories for\n+  newly-added files, so some filesystems may still allow data to be lost on\n+  system crash.\n+* `batch` enables an experimental mode that uses interfaces available in some\n+  operating systems to write loose object data with a minimal set of FLUSH\n+  CACHE (or equivalent) commands sent to the storage controller. If the\n+  operating system interfaces are not available, this mode behaves the same as\n+  `true`. This mode is expected to be as safe as `true` on macOS for repos\n+  stored on HFS+ or APFS filesystems and on Windows for repos stored on NTFS or\n+  ReFS.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\ndiff --git a/Makefile b/Makefile\nindex 12be39ac497..241dc322c09 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -406,6 +406,8 @@ all::\n #\n # Define HAVE_CLOCK_MONOTONIC if your platform has CLOCK_MONOTONIC.\n #\n+# Define HAVE_SYNC_FILE_RANGE if your platform has sync_file_range.\n+#\n # Define NEEDS_LIBRT if your platform requires linking with librt (glibc version\n # before 2.17) for clock_gettime and CLOCK_MONOTONIC.\n #\n@@ -1884,6 +1886,10 @@ ifdef HAVE_CLOCK_MONOTONIC\n \tBASIC_CFLAGS += -DHAVE_CLOCK_MONOTONIC\n endif\n \n+ifdef HAVE_SYNC_FILE_RANGE\n+\tBASIC_CFLAGS += -DHAVE_SYNC_FILE_RANGE\n+endif\n+\n ifdef NEEDS_LIBRT\n \tEXTLIBS += -lrt\n endif\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6ae18401e04..4deee1af46e 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,14 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n+static int needs_batch_fsync;\n+\n+static struct tmp_objdir *bulk_fsync_objdir;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -79,6 +85,34 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur.  The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware.\n+\t */\n+\n+\tif (needs_batch_fsync) {\n+\t\tstruct strbuf temp_path = STRBUF_INIT;\n+\t\tstruct tempfile *temp;\n+\n+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\t\ttemp = xmks_tempfile(temp_path.buf);\n+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\t\tdelete_tempfile(&temp);\n+\t\tstrbuf_release(&temp_path);\n+\t}\n+\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -273,6 +307,25 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void fsync_loose_object_bulk_checkin(int fd)\n+{\n+\tassert(fsync_object_files == FSYNC_OBJECT_FILES_BATCH);\n+\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (bulk_checkin_plugged &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\tif (!needs_batch_fsync)\n+\t\t\tneeds_batch_fsync = 1;\n+\t} else {\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -287,6 +340,19 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n+\n+\t/*\n+\t * A temporary object directory is used to hold the files\n+\t * while they are not fsynced.\n+\t */\n+\tif (fsync_object_files == FSYNC_OBJECT_FILES_BATCH) {\n+\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\t\tif (!bulk_fsync_objdir)\n+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n+\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+\t}\n+\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -296,4 +362,6 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..08f292379b6 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,8 @@\n \n #include \"cache.h\"\n \n+void fsync_loose_object_bulk_checkin(int fd);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex eba12487b99..6d6e6770ecc 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -985,7 +985,13 @@ void reset_shared_repository(void);\n extern int read_replace_refs;\n extern char *git_replace_ref_base;\n \n-extern int fsync_object_files;\n+enum fsync_object_files_mode {\n+    FSYNC_OBJECT_FILES_OFF,\n+    FSYNC_OBJECT_FILES_ON,\n+    FSYNC_OBJECT_FILES_BATCH\n+};\n+\n+extern enum fsync_object_files_mode fsync_object_files;\n extern int core_preload_index;\n extern int precomposed_unicode;\n extern int protect_hfs;\ndiff --git a/config.c b/config.c\nindex c5873f3a706..5eb36ecd77a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1491,7 +1491,12 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t}\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n-\t\tfsync_object_files = git_config_bool(var, value);\n+\t\tif (value && !strcmp(value, \"batch\"))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_BATCH;\n+\t\telse if (git_config_bool(var, value))\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_ON;\n+\t\telse\n+\t\t\tfsync_object_files = FSYNC_OBJECT_FILES_OFF;\n \t\treturn 0;\n \t}\n \ndiff --git a/config.mak.uname b/config.mak.uname\nindex 3236a4918a3..5ead1377667 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -57,6 +57,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_CLOCK_MONOTONIC = YesPlease\n \t# -lrt is needed for clock_gettime on glibc <= 2.16\n \tNEEDS_LIBRT = YesPlease\n+\tHAVE_SYNC_FILE_RANGE = YesPlease\n \tHAVE_GETDELIM = YesPlease\n \tSANE_TEXT_GREP=-a\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/configure.ac b/configure.ac\nindex 031e8d3fee8..c711037d625 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -1090,6 +1090,14 @@ AC_COMPILE_IFELSE([CLOCK_MONOTONIC_SRC],\n \t[AC_MSG_RESULT([no])\n \tHAVE_CLOCK_MONOTONIC=])\n GIT_CONF_SUBST([HAVE_CLOCK_MONOTONIC])\n+\n+#\n+# Define HAVE_SYNC_FILE_RANGE=YesPlease if sync_file_range is available.\n+GIT_CHECK_FUNC(sync_file_range,\n+\t[HAVE_SYNC_FILE_RANGE=YesPlease],\n+\t[HAVE_SYNC_FILE_RANGE])\n+GIT_CONF_SUBST([HAVE_SYNC_FILE_RANGE])\n+\n #\n # Define NO_SETITIMER if you don't have setitimer.\n GIT_CHECK_FUNC(setitimer,\ndiff --git a/environment.c b/environment.c\nindex 2701dfeeec8..aeafe80235e 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -42,7 +42,7 @@ const char *git_attributes_file;\n const char *git_hooks_path;\n int zlib_compression_level = Z_BEST_SPEED;\n int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-int fsync_object_files;\n+enum fsync_object_files_mode fsync_object_files;\n size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex d70ce142861..4defd4ab200 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1214,6 +1214,13 @@ __attribute__((format (printf, 1, 2))) NORETURN\n void BUG(const char *fmt, ...);\n #endif\n \n+enum fsync_action {\n+    FSYNC_WRITEOUT_ONLY,\n+    FSYNC_HARDWARE_FLUSH\n+};\n+\n+int git_fsync(int fd, enum fsync_action action);\n+\n /*\n  * Preserves errno, prints a message, but gives no warning for ENOENT.\n  * Returns 0 on success, which includes trying to unlink an object that does\ndiff --git a/object-file.c b/object-file.c\nindex 659ef7623ff..9d0aac792ae 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1853,8 +1853,18 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n static void close_loose_object(int fd)\n {\n \tif (!the_repository->objects->odb->will_destroy) {\n-\t\tif (fsync_object_files)\n+\t\tswitch (fsync_object_files) {\n+\t\tcase FSYNC_OBJECT_FILES_OFF:\n+\t\t\tbreak;\n+\t\tcase FSYNC_OBJECT_FILES_ON:\n \t\t\tfsync_or_die(fd, \"loose object file\");\n+\t\t\tbreak;\n+\t\tcase FSYNC_OBJECT_FILES_BATCH:\n+\t\t\tfsync_loose_object_bulk_checkin(fd);\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tBUG(\"Invalid fsync_object_files mode.\");\n+\t\t}\n \t}\n \n \tif (close(fd) != 0)\ndiff --git a/wrapper.c b/wrapper.c\nindex 36e12119d76..689288d2e31 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -546,6 +546,50 @@ int xmkstemp_mode(char *filename_template, int mode)\n \treturn fd;\n }\n \n+int git_fsync(int fd, enum fsync_action action)\n+{\n+\tswitch (action) {\n+\tcase FSYNC_WRITEOUT_ONLY:\n+\n+#ifdef __APPLE__\n+\t\t/*\n+\t\t * on macOS, fsync just causes filesystem cache writeback but does not\n+\t\t * flush hardware caches.\n+\t\t */\n+\t\treturn fsync(fd);\n+#endif\n+\n+#ifdef HAVE_SYNC_FILE_RANGE\n+\t\t/*\n+\t\t * On linux 2.6.17 and above, sync_file_range is the way to issue\n+\t\t * a writeback without a hardware flush. An offset of 0 and size of 0\n+\t\t * indicates writeout of the entire file and the wait flags ensure that all\n+\t\t * dirty data is written to the disk (potentially in a disk-side cache)\n+\t\t * before we continue.\n+\t\t */\n+\n+\t\treturn sync_file_range(fd, 0, 0, SYNC_FILE_RANGE_WAIT_BEFORE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WRITE |\n+\t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n+#endif\n+\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+\n+\tcase FSYNC_HARDWARE_FLUSH:\n+\n+#ifdef __APPLE__\n+\t\treturn fcntl(fd, F_FULLFSYNC);\n+#else\n+\t\treturn fsync(fd);\n+#endif\n+\n+\tdefault:\n+\t\tBUG(\"unexpected git_fsync(%d) call\", action);\n+\t}\n+\n+}\n+\n static int warn_if_unremovable(const char *op, const char *file, int rc)\n {\n \tint err;\ndiff --git a/write-or-die.c b/write-or-die.c\nindex 0b1ec8190b6..cc8291d9794 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -57,7 +57,7 @@ void fprintf_or_die(FILE *f, const char *fmt, ...)\n \n void fsync_or_die(int fd, const char *msg)\n {\n-\twhile (fsync(fd) < 0) {\n+\twhile (git_fsync(fd, FSYNC_HARDWARE_FLUSH) < 0) {\n \t\tif (errno != EINTR)\n \t\t\tdie_errno(\"fsync error on '%s'\", msg);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"441229","messageId":"951a559874e3e76c85e98d24c457e81a65de76e4.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 5/9] core.fsyncobjectfiles: add windows support for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:50:59Z","receivedAt":"2021-11-16T03:24:47Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit adds a win32 implementation for fsync_no_flush that is\ncalled git_fsync. The 'NtFlushBuffersFileEx' function being called is\navailable since Windows 8. If the function is not available, we\nreturn -1 and Git falls back to doing a full fsync.\n\nThe operating system is told to flush data only without a hardware\nflush primitive. A later full fsync will cause the metadata log\nto be flushed and then the disk cache to be flushed on NTFS and\nReFS. Other filesystems will treat this as a full flush operation.\n\nI added a new file here for this system call so as not to conflict with\ndownstream changes in the git-for-windows repository related to fscache.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n compat/mingw.h                      |  3 +++\n compat/win32/flush.c                | 28 ++++++++++++++++++++++++++++\n config.mak.uname                    |  2 ++\n contrib/buildsystems/CMakeLists.txt |  3 ++-\n wrapper.c                           |  4 ++++\n 5 files changed, 39 insertions(+), 1 deletion(-)\n create mode 100644 compat/win32/flush.c\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex c9a52ad64a6..6074a3d3ced 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -329,6 +329,9 @@ int mingw_getpagesize(void);\n #define getpagesize mingw_getpagesize\n #endif\n \n+int win32_fsync_no_flush(int fd);\n+#define fsync_no_flush win32_fsync_no_flush\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/compat/win32/flush.c b/compat/win32/flush.c\nnew file mode 100644\nindex 00000000000..75324c24ee7\n--- /dev/null\n+++ b/compat/win32/flush.c\n@@ -0,0 +1,28 @@\n+#include \"../../git-compat-util.h\"\n+#include <winternl.h>\n+#include \"lazyload.h\"\n+\n+int win32_fsync_no_flush(int fd)\n+{\n+       IO_STATUS_BLOCK io_status;\n+\n+#define FLUSH_FLAGS_FILE_DATA_ONLY 1\n+\n+       DECLARE_PROC_ADDR(ntdll.dll, NTSTATUS, NtFlushBuffersFileEx,\n+\t\t\t HANDLE FileHandle, ULONG Flags, PVOID Parameters, ULONG ParameterSize,\n+\t\t\t PIO_STATUS_BLOCK IoStatusBlock);\n+\n+       if (!INIT_PROC_ADDR(NtFlushBuffersFileEx)) {\n+\t\terrno = ENOSYS;\n+\t\treturn -1;\n+       }\n+\n+       memset(&io_status, 0, sizeof(io_status));\n+       if (NtFlushBuffersFileEx((HANDLE)_get_osfhandle(fd), FLUSH_FLAGS_FILE_DATA_ONLY,\n+\t\t\t\tNULL, 0, &io_status)) {\n+\t\terrno = EINVAL;\n+\t\treturn -1;\n+       }\n+\n+       return 0;\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 5ead1377667..5727fb093ca 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -455,6 +455,7 @@ endif\n \tCFLAGS =\n \tBASIC_CFLAGS = -nologo -I. -Icompat/vcbuild/include -DWIN32 -D_CONSOLE -DHAVE_STRING_H -D_CRT_SECURE_NO_WARNINGS -D_CRT_NONSTDC_NO_DEPRECATE\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n@@ -630,6 +631,7 @@ ifeq ($(uname_S),MINGW)\n \tCOMPAT_CFLAGS += -DSTRIP_EXTENSION=\\\".exe\\\"\n \tCOMPAT_OBJS += compat/mingw.o compat/winansi.o \\\n \t\tcompat/win32/trace2_win32_process_info.o \\\n+\t\tcompat/win32/flush.o \\\n \t\tcompat/win32/path-utils.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex fd1399c440f..ef0c1e4976d 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -261,7 +261,8 @@ if(CMAKE_SYSTEM_NAME STREQUAL \"Windows\")\n \t\t\t\tNOGDI OBJECT_CREATION_MODE=1 __USE_MINGW_ANSI_STDIO=0\n \t\t\t\tUSE_NED_ALLOCATOR OVERRIDE_STRDUP MMAP_PREVENTS_DELETE USE_WIN32_MMAP\n \t\t\t\tUNICODE _UNICODE HAVE_WPGMPTR ENSURE_MSYSTEM_IS_SET)\n-\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c compat/win32/path-utils.c\n+\tlist(APPEND compat_SOURCES compat/mingw.c compat/winansi.c\n+\t\tcompat/win32/flush.c compat/win32/path-utils.c\n \t\tcompat/win32/pthread.c compat/win32mmap.c compat/win32/syslog.c\n \t\tcompat/win32/trace2_win32_process_info.c compat/win32/dirent.c\n \t\tcompat/nedmalloc/nedmalloc.c compat/strdup.c)\ndiff --git a/wrapper.c b/wrapper.c\nindex 689288d2e31..ece3d2ca106 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -573,6 +573,10 @@ int git_fsync(int fd, enum fsync_action action)\n \t\t\t\t\t\t SYNC_FILE_RANGE_WAIT_AFTER);\n #endif\n \n+#ifdef fsync_no_flush\n+\t\treturn fsync_no_flush(fd);\n+#endif\n+\n \t\terrno = ENOSYS;\n \t\treturn -1;\n \n-- \ngitgitgadget\n\n"},{"id":"441230","messageId":"4a40fd4a29a468b9ce320bc7b22f19e5a526fad6.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 6/9] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:51:00Z","receivedAt":"2021-11-16T03:24:56Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\npack functionality and the new bulk-fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will not be available until the update-index is entirely complete.\nThis usage is unlikely, since any tool invoking update-index and\nexpecting to see objects would have to synchronize with the update-index\nprocess after passing it a file path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 187203e8bb5..dc7368bb1ee 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1088,6 +1089,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \tthe_index.updated_skipworktree = 1;\n \n+\t/* we might be adding many objects to the object database */\n+\tplug_bulk_checkin();\n+\n \t/*\n \t * Custom copy of parse_options() because we want to handle\n \t * filename arguments as they come.\n@@ -1168,6 +1172,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/* by now we must have added all of the new objects */\n+\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"441231","messageId":"cfc6a347d08ce465f260990010267e4e71d469eb.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 7/9] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:51:01Z","receivedAt":"2021-11-16T03:24:57Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling bulk-checkin when unpacking objects, we can take advantage\nof batched fsyncs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295ba..51eb4f7b531 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tplug_bulk_checkin();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"441232","messageId":"270c24827d0a5b246ec0db31de30c1f8c25cd0c3.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 8/9] core.fsyncobjectfiles: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:51:02Z","receivedAt":"2021-11-16T03:24:58Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added due\nto it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 36 ++++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 20 ++++++++++++++++++++\n t/t3903-stash.sh       | 14 ++++++++++++++\n t/t5300-pack-object.sh | 30 +++++++++++++++++++-----------\n 4 files changed, 89 insertions(+), 11 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a7de4ca8512\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,36 @@\n+# Helper to create files with unique contents\n+\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with unique\n+#\t\t\t\t\t contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\tlocal counter=0\n+\ttest_tick\n+\tlocal basedata=$test_tick\n+\n+\n+\trm -rf $basedir\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\"\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1))\n+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex 283a66955d6..aaecefda159 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -8,6 +8,8 @@ test_description='Test of git add, including the -- option.'\n TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -34,6 +36,24 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+test_expect_success 'git add: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files2 &&\n+\tfind fsync-files2 ! -type d -print | xargs git -c core.fsyncobjectfiles=batch update-index --add -- &&\n+\trm -f fsynced_files2 &&\n+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n+\ttest_line_count = 8 fsynced_files2 &&\n+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex f0a82be9de7..6324b52c874 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n diff_cmp () {\n \tfor i in \"$1\" \"$2\"\n@@ -1293,6 +1294,19 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+test_expect_success 'stash with core.fsyncobjectfiles=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit -c core.fsyncobjectfiles=batch stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'stash -c stash.useBuiltin=false warning ' '\n \texpected=\"stash.useBuiltin support has been removed\" &&\n \ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex e13a8842075..38663dc1393 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -162,23 +162,23 @@ test_expect_success 'pack-objects with bogus arguments' '\n \n check_unpack () {\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $2 init --bare git2 &&\n+\t(\n+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <obj-list >current &&\n+\tcmp obj-list current\n }\n \n test_expect_success 'unpack without delta' '\n \tcheck_unpack test-1-${packname_1}\n '\n \n+test_expect_success 'unpack without delta (core.fsyncobjectfiles=batch)' '\n+\tcheck_unpack test-1-${packname_1} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with REF_DELTA' '\n \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n \tcheck_deltas stderr -gt 0\n@@ -188,6 +188,10 @@ test_expect_success 'unpack with REF_DELTA' '\n \tcheck_unpack test-2-${packname_2}\n '\n \n+test_expect_success 'unpack with REF_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-2-${packname_2} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'pack with OFS_DELTA' '\n \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n \t\t\t<obj-list 2>stderr) &&\n@@ -198,6 +202,10 @@ test_expect_success 'unpack with OFS_DELTA' '\n \tcheck_unpack test-3-${packname_3}\n '\n \n+test_expect_success 'unpack with OFS_DELTA (core.fsyncobjectfiles=batch)' '\n+       check_unpack test-3-${packname_3} \"-c core.fsyncobjectfiles=batch\"\n+'\n+\n test_expect_success 'compare delta flavors' '\n \tperl -e '\\''\n \t\tdefined($_ = -s $_) or die for @ARGV;\n-- \ngitgitgadget\n\n"},{"id":"441233","messageId":"12d99641f4c3738efbc5c70ef4f8d70eccb6fc8b.1637020263.git.gitgitgadget@gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"[PATCH v9 9/9] core.fsyncobjectfiles: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-11-15T23:51:03Z","receivedAt":"2021-11-16T03:25:09Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd a basic performance test for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 43 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 46 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 89 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..e93c08a2e70\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,43 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\ttest_perf \"add $total_files files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..c9fcd0c03eb\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,46 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncObjectFiles=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the tet once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for core.fsyncObjectFiles=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash 500 files (core.fsyncObjectFiles=$m)\" \"\n+\t\tgit -c core.fsyncobjectfiles=$m stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n"},{"id":"441261","messageId":"211116.867dd8r2b1.gmgdl@evledraar.gmail.com","threadId":"56371","inReplyTo":"71817cccfb9573929723a34fe9fb0db72498d4b9.1637020263.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v9 2/9] tmp-objdir: disable ref updates when replacing the primary odb","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-11-16T07:23:21Z","receivedAt":"2021-11-16T07:24:30Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Nov 15 2021, Neeraj Singh via GitGitGadget wrote:\n\n>  \t/*\n>  \t * This object store is ephemeral, so there is no need to fsync.\n>  \t */\n> -\tint will_destroy;\n> +\tunsigned int will_destroy : 1;\n>  \n>  \t/*\n>  \t * Path to the alternative object store. If this is a relative path,\n\nWhy add this as an int in the preceding commit and turn it \"unsigned :\n1\" here?\n"},{"id":"441262","messageId":"211116.8635nwr055.gmgdl@evledraar.gmail.com","threadId":"56371","inReplyTo":"pull.1076.v9.git.git.1637020263.gitgitgadget@gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-11-16T08:02:14Z","receivedAt":"2021-11-16T08:10:55Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Nov 15 2021, Neeraj K. Singh via GitGitGadget wrote:\n\n>  * Per [2], I'm leaving the fsyncObjectFiles configuration as is with\n>    'true', 'false', and 'batch'. This makes using old and new versions of\n>    git with 'batch' mode a little trickier, but hopefully people will\n>    generally be moving forward in versions.\n>\n> [1] See\n> https://lore.kernel.org/git/pull.1067.git.1635287730.gitgitgadget@gmail.com/\n> [2] https://lore.kernel.org/git/xmqqh7cimuxt.fsf@gitster.g/\n\nI really think leaving that in-place is just being unnecessarily\ncavalier. There's a lot of mixed-version environments where git is\ndeployed in, and we almost never break the configuration in this way (I\nthink in the past always by mistake).\n\nIn this case it's easy to avoid it, and coming up with a less narrow\nconfig model[1] seems like a good idea in any case to unify the various\noutstanding work in this area.\n\nMore generally on this series, per the thread ending in [2] I really\ndon't get why we have code like this:\n\t\n\t@@ -503,10 +504,12 @@ static void unpack_all(void)\n\t \tif (!quiet)\n\t \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n\t \tCALLOC_ARRAY(obj_list, nr_objects);\n\t+\tplug_bulk_checkin();\n\t \tfor (i = 0; i < nr_objects; i++) {\n\t \t\tunpack_one(i);\n\t \t\tdisplay_progress(progress, i + 1);\n\t \t}\n\t+\tunplug_bulk_checkin();\n\t \tstop_progress(&progress);\n\t \n\t \tif (delta_list)\n\nAs opposed to doing an fsync on the last object we're\nprocessing. I.e. why do we need the step of intentionally making the\nobjects unavailable in the tmp-objdir, and creating a \"cookie\" file to\nsync at the start/end, as opposed to fsyncing on the last file (which\nwe're writing out anyway).\n\n1. https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n2. https://lore.kernel.org/git/20211111000349.GA703@neerajsi-x1.localdomain/\n"},{"id":"441360","messageId":"CANQDOddCC7+gGUy1VBxxwvN7ieP+N8mQhbxK2xx6ySqZc6U7-g@mail.gmail.com","threadId":"56371","inReplyTo":"211116.867dd8r2b1.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v9 2/9] tmp-objdir: disable ref updates when replacing the primary odb","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-16T20:38:38Z","receivedAt":"2021-11-16T20:38:52Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Nov 15, 2021 at 11:24 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Nov 15 2021, Neeraj Singh via GitGitGadget wrote:\n>\n> >       /*\n> >        * This object store is ephemeral, so there is no need to fsync.\n> >        */\n> > -     int will_destroy;\n> > +     unsigned int will_destroy : 1;\n> >\n> >       /*\n> >        * Path to the alternative object store. If this is a relative path,\n>\n> Why add this as an int in the preceding commit and turn it \"unsigned :\n> 1\" here?\n\nGood catch.  I'll fix this. This was an artifact of a previous version\nwhere there was another variable here as well.\n\nThanks,\nNeeraj\n"},{"id":"441406","messageId":"CANQDOdcEtOMMOLcHrnTKReRS23PvjOGp58VdpEkV_6iZuSPXaw@mail.gmail.com","threadId":"56371","inReplyTo":"211116.8635nwr055.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-17T07:06:25Z","receivedAt":"2021-11-17T07:06:41Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Nov 16, 2021 at 12:10 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Nov 15 2021, Neeraj K. Singh via GitGitGadget wrote:\n>\n> >  * Per [2], I'm leaving the fsyncObjectFiles configuration as is with\n> >    'true', 'false', and 'batch'. This makes using old and new versions of\n> >    git with 'batch' mode a little trickier, but hopefully people will\n> >    generally be moving forward in versions.\n> >\n> > [1] See\n> > https://lore.kernel.org/git/pull.1067.git.1635287730.gitgitgadget@gmail.com/\n> > [2] https://lore.kernel.org/git/xmqqh7cimuxt.fsf@gitster.g/\n>\n> I really think leaving that in-place is just being unnecessarily\n> cavalier. There's a lot of mixed-version environments where git is\n> deployed in, and we almost never break the configuration in this way (I\n> think in the past always by mistake).\n\n> In this case it's easy to avoid it, and coming up with a less narrow\n> config model[1] seems like a good idea in any case to unify the various\n> outstanding work in this area.\n>\n> More generally on this series, per the thread ending in [2] I really\n\nMy primary goal in all of these changes is to move git-for-windows over to\na default of batch fsync so that it can get closer to other platforms\nin performance\nof 'git add' while still retaining the same level of data integrity.\nI'm hoping that\nmost end-users are just sticking to defaults here.\n\nI'm happy to change the configuration schema again if there's a\nconsensus from the Git\ncommunity that backwards-compatibility of the configuration is\nactually important to someone.\n\nAlso, if we're doing a deeper rethink of the fsync configuration (as\nprompted by this work and\nEric Wong's and Patrick Steinhardts work), do we want to retain a mode\nwhere we fsync some\nparts of the persistent repo data but not others?  If we add fsyncing\nof the index in addition to the refs,\nI believe we would have covered all of the critical data structures\nthat would be needed to find the\ndata that a user has added to the repo if they complete a series of\ngit commands and then experience\na system crash.\n\n> don't get why we have code like this:\n>\n>         @@ -503,10 +504,12 @@ static void unpack_all(void)\n>                 if (!quiet)\n>                         progress = start_progress(_(\"Unpacking objects\"), nr_objects);\n>                 CALLOC_ARRAY(obj_list, nr_objects);\n>         +       plug_bulk_checkin();\n>                 for (i = 0; i < nr_objects; i++) {\n>                         unpack_one(i);\n>                         display_progress(progress, i + 1);\n>                 }\n>         +       unplug_bulk_checkin();\n>                 stop_progress(&progress);\n>\n>                 if (delta_list)\n>\n> As opposed to doing an fsync on the last object we're\n> processing. I.e. why do we need the step of intentionally making the\n> objects unavailable in the tmp-objdir, and creating a \"cookie\" file to\n> sync at the start/end, as opposed to fsyncing on the last file (which\n> we're writing out anyway).\n>\n> 1. https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n> 2. https://lore.kernel.org/git/20211111000349.GA703@neerajsi-x1.localdomain/\n\nIt's important to not expose an object's final name until its contents\nhave been fsynced\nto disk. We want to ensure that wherever we crash, we won't have a\nloose object that\nGit may later try to open where the filename doesn't match the content\nhash. I believe it's\nokay for a given OID to be missing, since a later command could\nrecreate it, but an object\nwith a wrong hash looks like it would persist until we do a git-fsck.\n\nI thought about figuring out how to sync the last object rather than some random\n\"cookie\" file, but it wasn't clear to me how I'd figure out which\nobject is actually last\nfrom library code in a way that doesn't burden each command with\nsomehow figuring\nout its last object and communicating that. The 'cookie' approach\nseems to lead to a cleaner\ninterface for callers.\n\nThanks,\nNeeraj\n"},{"id":"441407","messageId":"211117.86ee7f8cm4.gmgdl@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdcEtOMMOLcHrnTKReRS23PvjOGp58VdpEkV_6iZuSPXaw@mail.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-11-17T07:24:49Z","receivedAt":"2021-11-17T07:28:54Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Nov 16 2021, Neeraj Singh wrote:\n\n> On Tue, Nov 16, 2021 at 12:10 AM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Mon, Nov 15 2021, Neeraj K. Singh via GitGitGadget wrote:\n>>\n>> >  * Per [2], I'm leaving the fsyncObjectFiles configuration as is with\n>> >    'true', 'false', and 'batch'. This makes using old and new versions of\n>> >    git with 'batch' mode a little trickier, but hopefully people will\n>> >    generally be moving forward in versions.\n>> >\n>> > [1] See\n>> > https://lore.kernel.org/git/pull.1067.git.1635287730.gitgitgadget@gmail.com/\n>> > [2] https://lore.kernel.org/git/xmqqh7cimuxt.fsf@gitster.g/\n>>\n>> I really think leaving that in-place is just being unnecessarily\n>> cavalier. There's a lot of mixed-version environments where git is\n>> deployed in, and we almost never break the configuration in this way (I\n>> think in the past always by mistake).\n>\n>> In this case it's easy to avoid it, and coming up with a less narrow\n>> config model[1] seems like a good idea in any case to unify the various\n>> outstanding work in this area.\n>>\n>> More generally on this series, per the thread ending in [2] I really\n>\n> My primary goal in all of these changes is to move git-for-windows over to\n> a default of batch fsync so that it can get closer to other platforms\n> in performance\n> of 'git add' while still retaining the same level of data integrity.\n> I'm hoping that\n> most end-users are just sticking to defaults here.\n>\n> I'm happy to change the configuration schema again if there's a\n> consensus from the Git\n> community that backwards-compatibility of the configuration is\n> actually important to someone.\n>\n> Also, if we're doing a deeper rethink of the fsync configuration (as\n> prompted by this work and\n> Eric Wong's and Patrick Steinhardts work), do we want to retain a mode\n> where we fsync some\n> parts of the persistent repo data but not others?  If we add fsyncing\n> of the index in addition to the refs,\n> I believe we would have covered all of the critical data structures\n> that would be needed to find the\n> data that a user has added to the repo if they complete a series of\n> git commands and then experience\n> a system crash.\n\nJust talking about it is how we'll find consensus, maybe you & Junio\nwould like to keep it as-is. I don't see why we'd expose this bad edge\ncase in configuration handling to users when it's entirely avoidable,\nand we're still in the design phase.\n\n>> don't get why we have code like this:\n>>\n>>         @@ -503,10 +504,12 @@ static void unpack_all(void)\n>>                 if (!quiet)\n>>                         progress = start_progress(_(\"Unpacking objects\"), nr_objects);\n>>                 CALLOC_ARRAY(obj_list, nr_objects);\n>>         +       plug_bulk_checkin();\n>>                 for (i = 0; i < nr_objects; i++) {\n>>                         unpack_one(i);\n>>                         display_progress(progress, i + 1);\n>>                 }\n>>         +       unplug_bulk_checkin();\n>>                 stop_progress(&progress);\n>>\n>>                 if (delta_list)\n>>\n>> As opposed to doing an fsync on the last object we're\n>> processing. I.e. why do we need the step of intentionally making the\n>> objects unavailable in the tmp-objdir, and creating a \"cookie\" file to\n>> sync at the start/end, as opposed to fsyncing on the last file (which\n>> we're writing out anyway).\n>>\n>> 1. https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n>> 2. https://lore.kernel.org/git/20211111000349.GA703@neerajsi-x1.localdomain/\n>\n> It's important to not expose an object's final name until its contents\n> have been fsynced\n> to disk. We want to ensure that wherever we crash, we won't have a\n> loose object that\n> Git may later try to open where the filename doesn't match the content\n> hash. I believe it's\n> okay for a given OID to be missing, since a later command could\n> recreate it, but an object\n> with a wrong hash looks like it would persist until we do a git-fsck.\n\nYes, we handle that rather badly, as I mentioned in some other threads,\nbut not doing the fsync on the last object v.s. a \"cookie\" file right\nafterwards seems like a hail-mary at best, no?\n\n> I thought about figuring out how to sync the last object rather than some random\n> \"cookie\" file, but it wasn't clear to me how I'd figure out which\n> object is actually last\n> from library code in a way that doesn't burden each command with\n> somehow figuring\n> out its last object and communicating that. The 'cookie' approach\n> seems to lead to a cleaner\n> interface for callers.\n\nThe above quoted code is looping through nr_objects isn't it? Can't a\n\"do fsync\" be passed down to unpack_one() when we process the last loose\nobject?\n"},{"id":"441566","messageId":"CANQDOdcKzxM+M7wgxUz831SbpwGWR7gcUC8xLFM14BcCJ+60sA@mail.gmail.com","threadId":"56371","inReplyTo":"211117.86ee7f8cm4.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-18T05:03:16Z","receivedAt":"2021-11-18T05:03:34Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Nov 16, 2021 at 11:28 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Tue, Nov 16 2021, Neeraj Singh wrote:\n>\n> > On Tue, Nov 16, 2021 at 12:10 AM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Mon, Nov 15 2021, Neeraj K. Singh via GitGitGadget wrote:\n> >>\n> >> >  * Per [2], I'm leaving the fsyncObjectFiles configuration as is with\n> >> >    'true', 'false', and 'batch'. This makes using old and new versions of\n> >> >    git with 'batch' mode a little trickier, but hopefully people will\n> >> >    generally be moving forward in versions.\n> >> >\n> >> > [1] See\n> >> > https://lore.kernel.org/git/pull.1067.git.1635287730.gitgitgadget@gmail.com/\n> >> > [2] https://lore.kernel.org/git/xmqqh7cimuxt.fsf@gitster.g/\n> >>\n> >> I really think leaving that in-place is just being unnecessarily\n> >> cavalier. There's a lot of mixed-version environments where git is\n> >> deployed in, and we almost never break the configuration in this way (I\n> >> think in the past always by mistake).\n> >\n> >> In this case it's easy to avoid it, and coming up with a less narrow\n> >> config model[1] seems like a good idea in any case to unify the various\n> >> outstanding work in this area.\n> >>\n> >> More generally on this series, per the thread ending in [2] I really\n> >\n> > My primary goal in all of these changes is to move git-for-windows over to\n> > a default of batch fsync so that it can get closer to other platforms\n> > in performance\n> > of 'git add' while still retaining the same level of data integrity.\n> > I'm hoping that\n> > most end-users are just sticking to defaults here.\n> >\n> > I'm happy to change the configuration schema again if there's a\n> > consensus from the Git\n> > community that backwards-compatibility of the configuration is\n> > actually important to someone.\n> >\n> > Also, if we're doing a deeper rethink of the fsync configuration (as\n> > prompted by this work and\n> > Eric Wong's and Patrick Steinhardts work), do we want to retain a mode\n> > where we fsync some\n> > parts of the persistent repo data but not others?  If we add fsyncing\n> > of the index in addition to the refs,\n> > I believe we would have covered all of the critical data structures\n> > that would be needed to find the\n> > data that a user has added to the repo if they complete a series of\n> > git commands and then experience\n> > a system crash.\n>\n> Just talking about it is how we'll find consensus, maybe you & Junio\n> would like to keep it as-is. I don't see why we'd expose this bad edge\n> case in configuration handling to users when it's entirely avoidable,\n> and we're still in the design phase.\n\nAfter trying to figure out an implementation, I have a new proposal,\nwhich I've shared on the other thread [1].\n\n[1] https://lore.kernel.org/git/CANQDOdcdhfGtPg0PxpXQA5gQ4x9VknKDKCCi4HEB0Z1xgnjKzg@mail.gmail.com/\n\n>\n> >> don't get why we have code like this:\n> >>\n> >>         @@ -503,10 +504,12 @@ static void unpack_all(void)\n> >>                 if (!quiet)\n> >>                         progress = start_progress(_(\"Unpacking objects\"), nr_objects);\n> >>                 CALLOC_ARRAY(obj_list, nr_objects);\n> >>         +       plug_bulk_checkin();\n> >>                 for (i = 0; i < nr_objects; i++) {\n> >>                         unpack_one(i);\n> >>                         display_progress(progress, i + 1);\n> >>                 }\n> >>         +       unplug_bulk_checkin();\n> >>                 stop_progress(&progress);\n> >>\n> >>                 if (delta_list)\n> >>\n> >> As opposed to doing an fsync on the last object we're\n> >> processing. I.e. why do we need the step of intentionally making the\n> >> objects unavailable in the tmp-objdir, and creating a \"cookie\" file to\n> >> sync at the start/end, as opposed to fsyncing on the last file (which\n> >> we're writing out anyway).\n> >>\n> >> 1. https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n> >> 2. https://lore.kernel.org/git/20211111000349.GA703@neerajsi-x1.localdomain/\n> >\n> > It's important to not expose an object's final name until its contents\n> > have been fsynced\n> > to disk. We want to ensure that wherever we crash, we won't have a\n> > loose object that\n> > Git may later try to open where the filename doesn't match the content\n> > hash. I believe it's\n> > okay for a given OID to be missing, since a later command could\n> > recreate it, but an object\n> > with a wrong hash looks like it would persist until we do a git-fsck.\n>\n> Yes, we handle that rather badly, as I mentioned in some other threads,\n> but not doing the fsync on the last object v.s. a \"cookie\" file right\n> afterwards seems like a hail-mary at best, no?\n>\n\nI'm not quite grasping what you're saying here. Are you saying that\nusing a dummy\nfile instead of one of the actual objects is less likely to produce\nthe desired outcome\non actual filesystem implementations?\n\n> > I thought about figuring out how to sync the last object rather than some random\n> > \"cookie\" file, but it wasn't clear to me how I'd figure out which\n> > object is actually last\n> > from library code in a way that doesn't burden each command with\n> > somehow figuring\n> > out its last object and communicating that. The 'cookie' approach\n> > seems to lead to a cleaner\n> > interface for callers.\n>\n> The above quoted code is looping through nr_objects isn't it? Can't a\n> \"do fsync\" be passed down to unpack_one() when we process the last loose\n> object?\n\nAre you proposing that we do something different for unpack_objects\nversus update_index\nand git-add?  I was hoping to keep all of the users of the batch fsync\nfunctionality equivalent.\nFor the git-add workflow and update-index, we'd need to track the most\nrecent file so that we\ncan go back and fsync it.  I don't believe that syncing the last\nobject composes well with the existing\nimplementation of those commands.\n\nThanks,\nNeeraj\n"},{"id":"442692","messageId":"CABPp-BH6m4q_EoX77bqLcpCN1HRfJ_XayeCV2O0sRybX53rPrw@mail.gmail.com","threadId":"56371","inReplyTo":"6b27afa60e05c6c0b7752f1bcf6629c446ede520.1637020263.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v9 1/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2021-11-30T21:27:18Z","receivedAt":"2021-11-30T21:27:45Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 15, 2021 at 3:51 PM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> The tmp_objdir API provides the ability to create temporary object\n> directories, but was designed with the goal of having subprocesses\n> access these object stores, followed by the main process migrating\n> objects from it to the main object store or just deleting it.  The\n> subprocesses would view it as their primary datastore and write to it.\n>\n> Here we add the tmp_objdir_replace_primary_odb function that replaces\n> the current process's writable \"main\" object directory with the\n> specified one. The previous main object directory is restored in either\n> tmp_objdir_migrate or tmp_objdir_destroy.\n>\n> For the --remerge-diff usecase, add a new `will_destroy` flag in `struct\n> object_database` to mark ephemeral object databases that do not require\n> fsync durability.\n>\n> Add 'git prune' support for removing temporary object databases, and\n> make sure that they have a name starting with tmp_ and containing an\n> operation-specific name.\n>\n> Based-on-patch-by: Elijah Newren <newren@gmail.com>\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  builtin/prune.c        | 23 +++++++++++++++---\n>  builtin/receive-pack.c |  2 +-\n>  environment.c          |  5 ++++\n>  object-file.c          | 44 +++++++++++++++++++++++++++++++--\n>  object-store.h         | 19 +++++++++++++++\n>  object.c               |  2 +-\n>  tmp-objdir.c           | 55 +++++++++++++++++++++++++++++++++++++++---\n>  tmp-objdir.h           | 29 +++++++++++++++++++---\n>  8 files changed, 165 insertions(+), 14 deletions(-)\n>\n> diff --git a/builtin/prune.c b/builtin/prune.c\n> index 485c9a3c56f..a76e6a5f0e8 100644\n> --- a/builtin/prune.c\n> +++ b/builtin/prune.c\n> @@ -18,6 +18,7 @@ static int show_only;\n>  static int verbose;\n>  static timestamp_t expire;\n>  static int show_progress = -1;\n> +static struct strbuf remove_dir_buf = STRBUF_INIT;\n>\n>  static int prune_tmp_file(const char *fullpath)\n>  {\n> @@ -26,10 +27,20 @@ static int prune_tmp_file(const char *fullpath)\n>                 return error(\"Could not stat '%s'\", fullpath);\n>         if (st.st_mtime > expire)\n>                 return 0;\n> -       if (show_only || verbose)\n> -               printf(\"Removing stale temporary file %s\\n\", fullpath);\n> -       if (!show_only)\n> -               unlink_or_warn(fullpath);\n> +       if (S_ISDIR(st.st_mode)) {\n> +               if (show_only || verbose)\n> +                       printf(\"Removing stale temporary directory %s\\n\", fullpath);\n> +               if (!show_only) {\n> +                       strbuf_reset(&remove_dir_buf);\n> +                       strbuf_addstr(&remove_dir_buf, fullpath);\n> +                       remove_dir_recursively(&remove_dir_buf, 0);\n\nWhy not just define remove_dir_buf here rather than as a global?  It'd\nnot only make the code more readable by keeping everything localized,\nit would have prevented the forgotten strbuf_reset() bug from the\nearlier round of this patch.\n\nSure, that'd be an extra memory allocation/free for each directory you\nhit, which should be negligible compared to the cost of\nremove_dir_recursively()...and I'm not sure this is performance\ncritical anyway (I don't see why we'd expect more than O(1) cruft\ntemporary directories).\n\n> +               }\n> +       } else {\n> +               if (show_only || verbose)\n> +                       printf(\"Removing stale temporary file %s\\n\", fullpath);\n> +               if (!show_only)\n> +                       unlink_or_warn(fullpath);\n> +       }\n>         return 0;\n>  }\n>\n> @@ -97,6 +108,9 @@ static int prune_cruft(const char *basename, const char *path, void *data)\n>\n>  static int prune_subdir(unsigned int nr, const char *path, void *data)\n>  {\n> +       if (verbose)\n\nShouldn't this be\n    if (show_only || verbose)\n?\n\n> +               printf(\"Removing directory %s\\n\", path);\n> +\n>         if (!show_only)\n>                 rmdir(path);\n>         return 0;\n> @@ -184,5 +198,6 @@ int cmd_prune(int argc, const char **argv, const char *prefix)\n>                 prune_shallow(show_only ? PRUNE_SHOW_ONLY : 0);\n>         }\n>\n> +       strbuf_release(&remove_dir_buf);\n>         return 0;\n>  }\n> diff --git a/builtin/receive-pack.c b/builtin/receive-pack.c\n> index 49b846d9605..8815e24cde5 100644\n> --- a/builtin/receive-pack.c\n> +++ b/builtin/receive-pack.c\n> @@ -2213,7 +2213,7 @@ static const char *unpack(int err_fd, struct shallow_info *si)\n>                 strvec_push(&child.args, alt_shallow_file);\n>         }\n>\n> -       tmp_objdir = tmp_objdir_create();\n> +       tmp_objdir = tmp_objdir_create(\"incoming\");\n>         if (!tmp_objdir) {\n>                 if (err_fd > 0)\n>                         close(err_fd);\n> diff --git a/environment.c b/environment.c\n> index 9da7f3c1a19..342400fcaad 100644\n> --- a/environment.c\n> +++ b/environment.c\n> @@ -17,6 +17,7 @@\n>  #include \"commit.h\"\n>  #include \"strvec.h\"\n>  #include \"object-store.h\"\n> +#include \"tmp-objdir.h\"\n>  #include \"chdir-notify.h\"\n>  #include \"shallow.h\"\n>\n> @@ -331,10 +332,14 @@ static void update_relative_gitdir(const char *name,\n>                                    void *data)\n>  {\n>         char *path = reparent_relative_path(old_cwd, new_cwd, get_git_dir());\n> +       struct tmp_objdir *tmp_objdir = tmp_objdir_unapply_primary_odb();\n>         trace_printf_key(&trace_setup_key,\n>                          \"setup: move $GIT_DIR to '%s'\",\n>                          path);\n> +\n>         set_git_dir_1(path);\n> +       if (tmp_objdir)\n> +               tmp_objdir_reapply_primary_odb(tmp_objdir, old_cwd, new_cwd);\n>         free(path);\n>  }\n>\n> diff --git a/object-file.c b/object-file.c\n> index c3d866a287e..0b6a61aeaff 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -683,6 +683,43 @@ void add_to_alternates_memory(const char *reference)\n>                              '\\n', NULL, 0);\n>  }\n>\n> +struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy)\n> +{\n> +       struct object_directory *new_odb;\n> +\n> +       /*\n> +        * Make sure alternates are initialized, or else our entry may be\n> +        * overwritten when they are.\n> +        */\n> +       prepare_alt_odb(the_repository);\n> +\n> +       /*\n> +        * Make a new primary odb and link the old primary ODB in as an\n> +        * alternate\n> +        */\n> +       new_odb = xcalloc(1, sizeof(*new_odb));\n> +       new_odb->path = xstrdup(dir);\n> +       new_odb->will_destroy = will_destroy;\n> +       new_odb->next = the_repository->objects->odb;\n> +       the_repository->objects->odb = new_odb;\n> +       return new_odb->next;\n> +}\n> +\n> +void restore_primary_odb(struct object_directory *restore_odb, const char *old_path)\n> +{\n> +       struct object_directory *cur_odb = the_repository->objects->odb;\n> +\n> +       if (strcmp(old_path, cur_odb->path))\n> +               BUG(\"expected %s as primary object store; found %s\",\n> +                   old_path, cur_odb->path);\n> +\n> +       if (cur_odb->next != restore_odb)\n> +               BUG(\"we expect the old primary object store to be the first alternate\");\n> +\n> +       the_repository->objects->odb = restore_odb;\n> +       free_object_directory(cur_odb);\n> +}\n> +\n>  /*\n>   * Compute the exact path an alternate is at and returns it. In case of\n>   * error NULL is returned and the human readable error is added to `err`\n> @@ -1809,8 +1846,11 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  /* Finalize a file on disk, and close it. */\n>  static void close_loose_object(int fd)\n>  {\n> -       if (fsync_object_files)\n> -               fsync_or_die(fd, \"loose object file\");\n> +       if (!the_repository->objects->odb->will_destroy) {\n> +               if (fsync_object_files)\n> +                       fsync_or_die(fd, \"loose object file\");\n> +       }\n> +\n>         if (close(fd) != 0)\n>                 die_errno(_(\"error when closing loose object file\"));\n>  }\n> diff --git a/object-store.h b/object-store.h\n> index 952efb6a4be..cb173e69392 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -27,6 +27,11 @@ struct object_directory {\n>         uint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n>         struct oidtree *loose_objects_cache;\n>\n> +       /*\n> +        * This object store is ephemeral, so there is no need to fsync.\n> +        */\n> +       int will_destroy;\n> +\n>         /*\n>          * Path to the alternative object store. If this is a relative path,\n>          * it is relative to the current working directory.\n> @@ -58,6 +63,17 @@ void add_to_alternates_file(const char *dir);\n>   */\n>  void add_to_alternates_memory(const char *dir);\n>\n> +/*\n> + * Replace the current writable object directory with the specified temporary\n> + * object directory; returns the former primary object directory.\n> + */\n> +struct object_directory *set_temporary_primary_odb(const char *dir, int will_destroy);\n> +\n> +/*\n> + * Restore a previous ODB replaced by set_temporary_main_odb.\n> + */\n> +void restore_primary_odb(struct object_directory *restore_odb, const char *old_path);\n> +\n>  /*\n>   * Populate and return the loose object cache array corresponding to the\n>   * given object ID.\n> @@ -68,6 +84,9 @@ struct oidtree *odb_loose_cache(struct object_directory *odb,\n>  /* Empty the loose object cache for the specified object directory. */\n>  void odb_clear_loose_cache(struct object_directory *odb);\n>\n> +/* Clear and free the specified object directory */\n> +void free_object_directory(struct object_directory *odb);\n> +\n>  struct packed_git {\n>         struct hashmap_entry packmap_ent;\n>         struct packed_git *next;\n> diff --git a/object.c b/object.c\n> index 23a24e678a8..048f96a260e 100644\n> --- a/object.c\n> +++ b/object.c\n> @@ -513,7 +513,7 @@ struct raw_object_store *raw_object_store_new(void)\n>         return o;\n>  }\n>\n> -static void free_object_directory(struct object_directory *odb)\n> +void free_object_directory(struct object_directory *odb)\n>  {\n>         free(odb->path);\n>         odb_clear_loose_cache(odb);\n> diff --git a/tmp-objdir.c b/tmp-objdir.c\n> index b8d880e3626..3d38eeab66b 100644\n> --- a/tmp-objdir.c\n> +++ b/tmp-objdir.c\n> @@ -1,5 +1,6 @@\n>  #include \"cache.h\"\n>  #include \"tmp-objdir.h\"\n> +#include \"chdir-notify.h\"\n>  #include \"dir.h\"\n>  #include \"sigchain.h\"\n>  #include \"string-list.h\"\n> @@ -11,6 +12,8 @@\n>  struct tmp_objdir {\n>         struct strbuf path;\n>         struct strvec env;\n> +       struct object_directory *prev_odb;\n> +       int will_destroy;\n>  };\n>\n>  /*\n> @@ -38,6 +41,9 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n>         if (t == the_tmp_objdir)\n>                 the_tmp_objdir = NULL;\n>\n> +       if (!on_signal && t->prev_odb)\n> +               restore_primary_odb(t->prev_odb, t->path.buf);\n> +\n>         /*\n>          * This may use malloc via strbuf_grow(), but we should\n>          * have pre-grown t->path sufficiently so that this\n> @@ -52,6 +58,7 @@ static int tmp_objdir_destroy_1(struct tmp_objdir *t, int on_signal)\n>          */\n>         if (!on_signal)\n>                 tmp_objdir_free(t);\n> +\n>         return err;\n>  }\n>\n> @@ -121,7 +128,7 @@ static int setup_tmp_objdir(const char *root)\n>         return ret;\n>  }\n>\n> -struct tmp_objdir *tmp_objdir_create(void)\n> +struct tmp_objdir *tmp_objdir_create(const char *prefix)\n>  {\n>         static int installed_handlers;\n>         struct tmp_objdir *t;\n> @@ -129,11 +136,16 @@ struct tmp_objdir *tmp_objdir_create(void)\n>         if (the_tmp_objdir)\n>                 BUG(\"only one tmp_objdir can be used at a time\");\n>\n> -       t = xmalloc(sizeof(*t));\n> +       t = xcalloc(1, sizeof(*t));\n>         strbuf_init(&t->path, 0);\n>         strvec_init(&t->env);\n>\n> -       strbuf_addf(&t->path, \"%s/incoming-XXXXXX\", get_object_directory());\n> +       /*\n> +        * Use a string starting with tmp_ so that the builtin/prune.c code\n> +        * can recognize any stale objdirs left behind by a crash and delete\n> +        * them.\n> +        */\n> +       strbuf_addf(&t->path, \"%s/tmp_objdir-%s-XXXXXX\", get_object_directory(), prefix);\n>\n>         /*\n>          * Grow the strbuf beyond any filename we expect to be placed in it.\n> @@ -269,6 +281,13 @@ int tmp_objdir_migrate(struct tmp_objdir *t)\n>         if (!t)\n>                 return 0;\n>\n> +       if (t->prev_odb) {\n> +               if (the_repository->objects->odb->will_destroy)\n> +                       BUG(\"migrating an ODB that was marked for destruction\");\n> +               restore_primary_odb(t->prev_odb, t->path.buf);\n> +               t->prev_odb = NULL;\n> +       }\n> +\n>         strbuf_addbuf(&src, &t->path);\n>         strbuf_addstr(&dst, get_object_directory());\n>\n> @@ -292,3 +311,33 @@ void tmp_objdir_add_as_alternate(const struct tmp_objdir *t)\n>  {\n>         add_to_alternates_memory(t->path.buf);\n>  }\n> +\n> +void tmp_objdir_replace_primary_odb(struct tmp_objdir *t, int will_destroy)\n> +{\n> +       if (t->prev_odb)\n> +               BUG(\"the primary object database is already replaced\");\n> +       t->prev_odb = set_temporary_primary_odb(t->path.buf, will_destroy);\n> +       t->will_destroy = will_destroy;\n> +}\n> +\n> +struct tmp_objdir *tmp_objdir_unapply_primary_odb(void)\n> +{\n> +       if (!the_tmp_objdir || !the_tmp_objdir->prev_odb)\n> +               return NULL;\n> +\n> +       restore_primary_odb(the_tmp_objdir->prev_odb, the_tmp_objdir->path.buf);\n> +       the_tmp_objdir->prev_odb = NULL;\n> +       return the_tmp_objdir;\n> +}\n> +\n> +void tmp_objdir_reapply_primary_odb(struct tmp_objdir *t, const char *old_cwd,\n> +               const char *new_cwd)\n> +{\n> +       char *path;\n> +\n> +       path = reparent_relative_path(old_cwd, new_cwd, t->path.buf);\n> +       strbuf_reset(&t->path);\n> +       strbuf_addstr(&t->path, path);\n> +       free(path);\n> +       tmp_objdir_replace_primary_odb(t, t->will_destroy);\n> +}\n> diff --git a/tmp-objdir.h b/tmp-objdir.h\n> index b1e45b4c75d..a3145051f25 100644\n> --- a/tmp-objdir.h\n> +++ b/tmp-objdir.h\n> @@ -10,7 +10,7 @@\n>   *\n>   * Example:\n>   *\n> - *     struct tmp_objdir *t = tmp_objdir_create();\n> + *     struct tmp_objdir *t = tmp_objdir_create(\"incoming\");\n>   *     if (!run_command_v_opt_cd_env(cmd, 0, NULL, tmp_objdir_env(t)) &&\n>   *         !tmp_objdir_migrate(t))\n>   *             printf(\"success!\\n\");\n> @@ -22,9 +22,10 @@\n>  struct tmp_objdir;\n>\n>  /*\n> - * Create a new temporary object directory; returns NULL on failure.\n> + * Create a new temporary object directory with the specified prefix;\n> + * returns NULL on failure.\n>   */\n> -struct tmp_objdir *tmp_objdir_create(void);\n> +struct tmp_objdir *tmp_objdir_create(const char *prefix);\n>\n>  /*\n>   * Return a list of environment strings, suitable for use with\n> @@ -51,4 +52,26 @@ int tmp_objdir_destroy(struct tmp_objdir *);\n>   */\n>  void tmp_objdir_add_as_alternate(const struct tmp_objdir *);\n>\n> +/*\n> + * Replaces the main object store in the current process with the temporary\n> + * object directory and makes the former main object store an alternate.\n> + * If will_destroy is nonzero, the object directory may not be migrated.\n> + */\n> +void tmp_objdir_replace_primary_odb(struct tmp_objdir *, int will_destroy);\n> +\n> +/*\n> + * If the primary object database was replaced by a temporary object directory,\n> + * restore it to its original value while keeping the directory contents around.\n> + * Returns NULL if the primary object database was not replaced.\n> + */\n> +struct tmp_objdir *tmp_objdir_unapply_primary_odb(void);\n> +\n> +/*\n> + * Reapplies the former primary temporary object database, after protentially\n> + * changing its relative path.\n> + */\n> +void tmp_objdir_reapply_primary_odb(struct tmp_objdir *, const char *old_cwd,\n> +               const char *new_cwd);\n> +\n> +\n>  #endif /* TMP_OBJDIR_H */\n> --\n> gitgitgadget\n"},{"id":"442708","messageId":"CANQDOdd7EHUqD_JBdO9ArpvOQYUnU9GSL6EVR7W7XXgNASZyhQ@mail.gmail.com","threadId":"56371","inReplyTo":"CABPp-BH6m4q_EoX77bqLcpCN1HRfJ_XayeCV2O0sRybX53rPrw@mail.gmail.com","subject":"Re: [PATCH v9 1/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2021-11-30T21:52:12Z","receivedAt":"2021-11-30T21:52:28Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Nov 30, 2021 at 1:27 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Mon, Nov 15, 2021 at 3:51 PM Neeraj Singh via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > The tmp_objdir API provides the ability to create temporary object\n> > directories, but was designed with the goal of having subprocesses\n> > access these object stores, followed by the main process migrating\n> > objects from it to the main object store or just deleting it.  The\n> > subprocesses would view it as their primary datastore and write to it.\n> >\n> > Here we add the tmp_objdir_replace_primary_odb function that replaces\n> > the current process's writable \"main\" object directory with the\n> > specified one. The previous main object directory is restored in either\n> > tmp_objdir_migrate or tmp_objdir_destroy.\n> >\n> > For the --remerge-diff usecase, add a new `will_destroy` flag in `struct\n> > object_database` to mark ephemeral object databases that do not require\n> > fsync durability.\n> >\n> > Add 'git prune' support for removing temporary object databases, and\n> > make sure that they have a name starting with tmp_ and containing an\n> > operation-specific name.\n> >\n> > Based-on-patch-by: Elijah Newren <newren@gmail.com>\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> > ---\n> >  builtin/prune.c        | 23 +++++++++++++++---\n> >  builtin/receive-pack.c |  2 +-\n> >  environment.c          |  5 ++++\n> >  object-file.c          | 44 +++++++++++++++++++++++++++++++--\n> >  object-store.h         | 19 +++++++++++++++\n> >  object.c               |  2 +-\n> >  tmp-objdir.c           | 55 +++++++++++++++++++++++++++++++++++++++---\n> >  tmp-objdir.h           | 29 +++++++++++++++++++---\n> >  8 files changed, 165 insertions(+), 14 deletions(-)\n> >\n> > diff --git a/builtin/prune.c b/builtin/prune.c\n> > index 485c9a3c56f..a76e6a5f0e8 100644\n> > --- a/builtin/prune.c\n> > +++ b/builtin/prune.c\n> > @@ -18,6 +18,7 @@ static int show_only;\n> >  static int verbose;\n> >  static timestamp_t expire;\n> >  static int show_progress = -1;\n> > +static struct strbuf remove_dir_buf = STRBUF_INIT;\n> >\n> >  static int prune_tmp_file(const char *fullpath)\n> >  {\n> > @@ -26,10 +27,20 @@ static int prune_tmp_file(const char *fullpath)\n> >                 return error(\"Could not stat '%s'\", fullpath);\n> >         if (st.st_mtime > expire)\n> >                 return 0;\n> > -       if (show_only || verbose)\n> > -               printf(\"Removing stale temporary file %s\\n\", fullpath);\n> > -       if (!show_only)\n> > -               unlink_or_warn(fullpath);\n> > +       if (S_ISDIR(st.st_mode)) {\n> > +               if (show_only || verbose)\n> > +                       printf(\"Removing stale temporary directory %s\\n\", fullpath);\n> > +               if (!show_only) {\n> > +                       strbuf_reset(&remove_dir_buf);\n> > +                       strbuf_addstr(&remove_dir_buf, fullpath);\n> > +                       remove_dir_recursively(&remove_dir_buf, 0);\n>\n> Why not just define remove_dir_buf here rather than as a global?  It'd\n> not only make the code more readable by keeping everything localized,\n> it would have prevented the forgotten strbuf_reset() bug from the\n> earlier round of this patch.\n>\n> Sure, that'd be an extra memory allocation/free for each directory you\n> hit, which should be negligible compared to the cost of\n> remove_dir_recursively()...and I'm not sure this is performance\n> critical anyway (I don't see why we'd expect more than O(1) cruft\n> temporary directories).\n\nI'll take this suggestion.\n\n> > +               }\n> > +       } else {\n> > +               if (show_only || verbose)\n> > +                       printf(\"Removing stale temporary file %s\\n\", fullpath);\n> > +               if (!show_only)\n> > +                       unlink_or_warn(fullpath);\n> > +       }\n> >         return 0;\n> >  }\n> >\n> > @@ -97,6 +108,9 @@ static int prune_cruft(const char *basename, const char *path, void *data)\n> >\n> >  static int prune_subdir(unsigned int nr, const char *path, void *data)\n> >  {\n> > +       if (verbose)\n>\n> Shouldn't this be\n>     if (show_only || verbose)\n> ?\n\nDoing that breaks one of the tests, since we print extra stuff that's\nunexpected. I think I'm going to just revert this change, since it\nappears that we call this callback and try to remove the directory\neven if it's non-empty.\n\nDo you have any comments or thoughts on how we want to allow the user\nto configure fsync settings?\n\nThanks,\nNeeraj\n"},{"id":"442713","messageId":"CABPp-BELu8xzJeSoDtMYFXj6Zw6JSm4qgBSXnSFmQbGzwjiD_Q@mail.gmail.com","threadId":"56371","inReplyTo":"CANQDOdd7EHUqD_JBdO9ArpvOQYUnU9GSL6EVR7W7XXgNASZyhQ@mail.gmail.com","subject":"Re: [PATCH v9 1/9] tmp-objdir: new API for creating temporary writable databases","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2021-11-30T22:36:10Z","receivedAt":"2021-11-30T22:36:25Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Tue, Nov 30, 2021 at 1:52 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> On Tue, Nov 30, 2021 at 1:27 PM Elijah Newren <newren@gmail.com> wrote:\n> >\n> > On Mon, Nov 15, 2021 at 3:51 PM Neeraj Singh via GitGitGadget\n> > <gitgitgadget@gmail.com> wrote:\n> > >\n> > > From: Neeraj Singh <neerajsi@microsoft.com>\n> > >\n> > > The tmp_objdir API provides the ability to create temporary object\n> > > directories, but was designed with the goal of having subprocesses\n> > > access these object stores, followed by the main process migrating\n> > > objects from it to the main object store or just deleting it.  The\n> > > subprocesses would view it as their primary datastore and write to it.\n> > >\n> > > Here we add the tmp_objdir_replace_primary_odb function that replaces\n> > > the current process's writable \"main\" object directory with the\n> > > specified one. The previous main object directory is restored in either\n> > > tmp_objdir_migrate or tmp_objdir_destroy.\n> > >\n> > > For the --remerge-diff usecase, add a new `will_destroy` flag in `struct\n> > > object_database` to mark ephemeral object databases that do not require\n> > > fsync durability.\n> > >\n> > > Add 'git prune' support for removing temporary object databases, and\n> > > make sure that they have a name starting with tmp_ and containing an\n> > > operation-specific name.\n> > >\n> > > Based-on-patch-by: Elijah Newren <newren@gmail.com>\n> > >\n> > > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > > Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> > > ---\n> > >  builtin/prune.c        | 23 +++++++++++++++---\n> > >  builtin/receive-pack.c |  2 +-\n> > >  environment.c          |  5 ++++\n> > >  object-file.c          | 44 +++++++++++++++++++++++++++++++--\n> > >  object-store.h         | 19 +++++++++++++++\n> > >  object.c               |  2 +-\n> > >  tmp-objdir.c           | 55 +++++++++++++++++++++++++++++++++++++++---\n> > >  tmp-objdir.h           | 29 +++++++++++++++++++---\n> > >  8 files changed, 165 insertions(+), 14 deletions(-)\n> > >\n> > > diff --git a/builtin/prune.c b/builtin/prune.c\n> > > index 485c9a3c56f..a76e6a5f0e8 100644\n> > > --- a/builtin/prune.c\n> > > +++ b/builtin/prune.c\n> > > @@ -18,6 +18,7 @@ static int show_only;\n> > >  static int verbose;\n> > >  static timestamp_t expire;\n> > >  static int show_progress = -1;\n> > > +static struct strbuf remove_dir_buf = STRBUF_INIT;\n> > >\n> > >  static int prune_tmp_file(const char *fullpath)\n> > >  {\n> > > @@ -26,10 +27,20 @@ static int prune_tmp_file(const char *fullpath)\n> > >                 return error(\"Could not stat '%s'\", fullpath);\n> > >         if (st.st_mtime > expire)\n> > >                 return 0;\n> > > -       if (show_only || verbose)\n> > > -               printf(\"Removing stale temporary file %s\\n\", fullpath);\n> > > -       if (!show_only)\n> > > -               unlink_or_warn(fullpath);\n> > > +       if (S_ISDIR(st.st_mode)) {\n> > > +               if (show_only || verbose)\n> > > +                       printf(\"Removing stale temporary directory %s\\n\", fullpath);\n> > > +               if (!show_only) {\n> > > +                       strbuf_reset(&remove_dir_buf);\n> > > +                       strbuf_addstr(&remove_dir_buf, fullpath);\n> > > +                       remove_dir_recursively(&remove_dir_buf, 0);\n> >\n> > Why not just define remove_dir_buf here rather than as a global?  It'd\n> > not only make the code more readable by keeping everything localized,\n> > it would have prevented the forgotten strbuf_reset() bug from the\n> > earlier round of this patch.\n> >\n> > Sure, that'd be an extra memory allocation/free for each directory you\n> > hit, which should be negligible compared to the cost of\n> > remove_dir_recursively()...and I'm not sure this is performance\n> > critical anyway (I don't see why we'd expect more than O(1) cruft\n> > temporary directories).\n>\n> I'll take this suggestion.\n>\n> > > +               }\n> > > +       } else {\n> > > +               if (show_only || verbose)\n> > > +                       printf(\"Removing stale temporary file %s\\n\", fullpath);\n> > > +               if (!show_only)\n> > > +                       unlink_or_warn(fullpath);\n> > > +       }\n> > >         return 0;\n> > >  }\n> > >\n> > > @@ -97,6 +108,9 @@ static int prune_cruft(const char *basename, const char *path, void *data)\n> > >\n> > >  static int prune_subdir(unsigned int nr, const char *path, void *data)\n> > >  {\n> > > +       if (verbose)\n> >\n> > Shouldn't this be\n> >     if (show_only || verbose)\n> > ?\n>\n> Doing that breaks one of the tests, since we print extra stuff that's\n> unexpected. I think I'm going to just revert this change, since it\n> appears that we call this callback and try to remove the directory\n> even if it's non-empty.\n\nMakes sense.\n\n> Do you have any comments or thoughts on how we want to allow the user\n> to configure fsync settings?\n\nI don't; sorry.\n"},{"id":"442777","messageId":"211201.864k7sbdjt.gmgdl@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdcKzxM+M7wgxUz831SbpwGWR7gcUC8xLFM14BcCJ+60sA@mail.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-01T14:15:47Z","receivedAt":"2021-12-01T14:35:52Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Nov 17 2021, Neeraj Singh wrote:\n\n[Very late reply, sorry]\n\n> On Tue, Nov 16, 2021 at 11:28 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Tue, Nov 16 2021, Neeraj Singh wrote:\n>>\n>> > On Tue, Nov 16, 2021 at 12:10 AM Ævar Arnfjörð Bjarmason\n>> > <avarab@gmail.com> wrote:\n>> >>\n>> >>\n>> >> On Mon, Nov 15 2021, Neeraj K. Singh via GitGitGadget wrote:\n>> >>\n>> >> >  * Per [2], I'm leaving the fsyncObjectFiles configuration as is with\n>> >> >    'true', 'false', and 'batch'. This makes using old and new versions of\n>> >> >    git with 'batch' mode a little trickier, but hopefully people will\n>> >> >    generally be moving forward in versions.\n>> >> >\n>> >> > [1] See\n>> >> > https://lore.kernel.org/git/pull.1067.git.1635287730.gitgitgadget@gmail.com/\n>> >> > [2] https://lore.kernel.org/git/xmqqh7cimuxt.fsf@gitster.g/\n>> >>\n>> >> I really think leaving that in-place is just being unnecessarily\n>> >> cavalier. There's a lot of mixed-version environments where git is\n>> >> deployed in, and we almost never break the configuration in this way (I\n>> >> think in the past always by mistake).\n>> >\n>> >> In this case it's easy to avoid it, and coming up with a less narrow\n>> >> config model[1] seems like a good idea in any case to unify the various\n>> >> outstanding work in this area.\n>> >>\n>> >> More generally on this series, per the thread ending in [2] I really\n>> >\n>> > My primary goal in all of these changes is to move git-for-windows over to\n>> > a default of batch fsync so that it can get closer to other platforms\n>> > in performance\n>> > of 'git add' while still retaining the same level of data integrity.\n>> > I'm hoping that\n>> > most end-users are just sticking to defaults here.\n>> >\n>> > I'm happy to change the configuration schema again if there's a\n>> > consensus from the Git\n>> > community that backwards-compatibility of the configuration is\n>> > actually important to someone.\n>> >\n>> > Also, if we're doing a deeper rethink of the fsync configuration (as\n>> > prompted by this work and\n>> > Eric Wong's and Patrick Steinhardts work), do we want to retain a mode\n>> > where we fsync some\n>> > parts of the persistent repo data but not others?  If we add fsyncing\n>> > of the index in addition to the refs,\n>> > I believe we would have covered all of the critical data structures\n>> > that would be needed to find the\n>> > data that a user has added to the repo if they complete a series of\n>> > git commands and then experience\n>> > a system crash.\n>>\n>> Just talking about it is how we'll find consensus, maybe you & Junio\n>> would like to keep it as-is. I don't see why we'd expose this bad edge\n>> case in configuration handling to users when it's entirely avoidable,\n>> and we're still in the design phase.\n>\n> After trying to figure out an implementation, I have a new proposal,\n> which I've shared on the other thread [1].\n>\n> [1] https://lore.kernel.org/git/CANQDOdcdhfGtPg0PxpXQA5gQ4x9VknKDKCCi4HEB0Z1xgnjKzg@mail.gmail.com/\n\nThis LGTM, or something simpler as Junio points out with his \"too\nfine-grained?\" comment as a follow-up. I'm honestly quite apathetic\nabout what we end up with exactly as long as:\n\n 1. We get the people who are adding these config settings to talk & see if they make\n    sense in combination.\n\n 2. We avoid the trap of hard dying on older versions.\n\n>>\n>> >> don't get why we have code like this:\n>> >>\n>> >>         @@ -503,10 +504,12 @@ static void unpack_all(void)\n>> >>                 if (!quiet)\n>> >>                         progress = start_progress(_(\"Unpacking objects\"), nr_objects);\n>> >>                 CALLOC_ARRAY(obj_list, nr_objects);\n>> >>         +       plug_bulk_checkin();\n>> >>                 for (i = 0; i < nr_objects; i++) {\n>> >>                         unpack_one(i);\n>> >>                         display_progress(progress, i + 1);\n>> >>                 }\n>> >>         +       unplug_bulk_checkin();\n>> >>                 stop_progress(&progress);\n>> >>\n>> >>                 if (delta_list)\n>> >>\n>> >> As opposed to doing an fsync on the last object we're\n>> >> processing. I.e. why do we need the step of intentionally making the\n>> >> objects unavailable in the tmp-objdir, and creating a \"cookie\" file to\n>> >> sync at the start/end, as opposed to fsyncing on the last file (which\n>> >> we're writing out anyway).\n>> >>\n>> >> 1. https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n>> >> 2. https://lore.kernel.org/git/20211111000349.GA703@neerajsi-x1.localdomain/\n>> >\n>> > It's important to not expose an object's final name until its contents\n>> > have been fsynced\n>> > to disk. We want to ensure that wherever we crash, we won't have a\n>> > loose object that\n>> > Git may later try to open where the filename doesn't match the content\n>> > hash. I believe it's\n>> > okay for a given OID to be missing, since a later command could\n>> > recreate it, but an object\n>> > with a wrong hash looks like it would persist until we do a git-fsck.\n>>\n>> Yes, we handle that rather badly, as I mentioned in some other threads,\n>> but not doing the fsync on the last object v.s. a \"cookie\" file right\n>> afterwards seems like a hail-mary at best, no?\n>>\n>\n> I'm not quite grasping what you're saying here. Are you saying that\n> using a dummy\n> file instead of one of the actual objects is less likely to produce\n> the desired outcome\n> on actual filesystem implementations?\n\n[...covered below...]\n\n>> > I thought about figuring out how to sync the last object rather than some random\n>> > \"cookie\" file, but it wasn't clear to me how I'd figure out which\n>> > object is actually last\n>> > from library code in a way that doesn't burden each command with\n>> > somehow figuring\n>> > out its last object and communicating that. The 'cookie' approach\n>> > seems to lead to a cleaner\n>> > interface for callers.\n>>\n>> The above quoted code is looping through nr_objects isn't it? Can't a\n>> \"do fsync\" be passed down to unpack_one() when we process the last loose\n>> object?\n>\n> Are you proposing that we do something different for unpack_objects\n> versus update_index\n> and git-add?  I was hoping to keep all of the users of the batch fsync\n> functionality equivalent.\n> For the git-add workflow and update-index, we'd need to track the most\n> recent file so that we\n> can go back and fsync it.  I don't believe that syncing the last\n> object composes well with the existing\n> implementation of those commands.\n\nThere's probably cases where we need the cookie. I just mean instead of\nthe API being (as seen above in the quoted part), pseudocode:\n\n    # A\n    bulk_checkin_start_make_cookie():\n    n = 10\n    for i in 1..n:\n        write_nth(i, fsync: 0);\n    bulk_checkin_end_commit_cookie();\n\nTo have it be:\n\n    # B\n    bulk_checkin_start(do_cookie: 0);\n    n = 10\n    for i in 1..n:\n        write_nth(i, fsync: (i == n));\n    bulk_checkin_end();\n\nOr actually, presumably simpler as:\n\n    # C\n    all_fsync = bulk_checkin_mode() ? 0 : fsync_turned_on_in_general();\n    end_fsync = bulk_checkin_mode() ? 1 : all_fsync;\n    n = 10;\n    for i in 1..n:\n        write_nth(i, fsync: (i == n) ? end_fsync : all_fsync);\n\nI.e. maybe there are cases where you really do need \"A\", but we're\nusually (always?) writing out N objects, and we usually know it at the\nsame point where you'd want the plug_bulk_checkin/unplug_bulk_checkin,\nso just fsyncing the last object/file/ref/whatever means we don't need\nthe whole ceremony of the cookie file.\n\nI don't mind it per-se, but \"B\" and \"C\" just seem a lot simpler,\nparticulary since as those examples show we'll presumably want to pass\ndown a \"do fsync?\" to these in general, and we even usually have a\ndisplay_progress() in there.\n\nSo doesn't just doing \"B\" or \"C\" eliminate the need for a cookie\nentirely?\n\nAnother advantage of that is that you'll presumably want such tracking\nanyway even for the case of \"A\".\n\nBecause as soon as you have say a batch operation of writing X objects\nand Y refs you'd want to track this anyway. I.e. either only fsync() on\nthe ref write (particularly if there's just the one ref), or on the last\nref, or for each ref and no object syncs. So this (like \"C\", except for\nthe \"do_batch\" in the \"end_fsync\" case):\n\n    # D\n    do_batch = in_existing_bulk_checkin() ? 1 : 0;\n    all_fsync = bulk_checkin_mode() ? 0 : fsync_turned_on_in_general();\n    end_fsync = bulk_checkin_mode() ? do_batch : all_fsync;\n    n = 10;\n    for i in 1..n:\n        write_nth(i, fsync: (i == n) ? end_fsync : all_fsync);\n\nI mean, usually we'd want the \"all refs\", I'm just thinking of a case\nlike \"git fast-import\" or other known-to-the-user batch operation.\n\nOr, as in the case of my 4bc1fd6e394 (pack-objects: rename .idx files\ninto place after .bitmap files, 2021-09-09) we'd want to know that we're\nwriting all of say *.bitmap, *.rev where we currently fsync() all of\nthem, write *.bitmap, *.rev and *.pack (not sure that one is safe)\nwithout fsync(), and then only fsync (or that and in-place move) the\n*.idx.\n"},{"id":"450954","messageId":"220310.86lexilo3d.gmgdl@evledraar.gmail.com","threadId":"56371","inReplyTo":"211201.864k7sbdjt.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-09T23:02:49Z","receivedAt":"2022-03-09T23:10:06Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Dec 01 2021, Ævar Arnfjörð Bjarmason wrote:\n\n> On Wed, Nov 17 2021, Neeraj Singh wrote:\n>\n> [Very late reply, sorry]\n>\n>> On Tue, Nov 16, 2021 at 11:28 PM Ævar Arnfjörð Bjarmason\n>> <avarab@gmail.com> wrote:\n>>>\n>>>\n>>> On Tue, Nov 16 2021, Neeraj Singh wrote:\n>>>\n>>> > On Tue, Nov 16, 2021 at 12:10 AM Ævar Arnfjörð Bjarmason\n>>> > <avarab@gmail.com> wrote:\n>>> >>\n>>> >>\n>>> >> On Mon, Nov 15 2021, Neeraj K. Singh via GitGitGadget wrote:\n>>> >>\n>>> >> >  * Per [2], I'm leaving the fsyncObjectFiles configuration as is with\n>>> >> >    'true', 'false', and 'batch'. This makes using old and new versions of\n>>> >> >    git with 'batch' mode a little trickier, but hopefully people will\n>>> >> >    generally be moving forward in versions.\n>>> >> >\n>>> >> > [1] See\n>>> >> > https://lore.kernel.org/git/pull.1067.git.1635287730.gitgitgadget@gmail.com/\n>>> >> > [2] https://lore.kernel.org/git/xmqqh7cimuxt.fsf@gitster.g/\n>>> >>\n>>> >> I really think leaving that in-place is just being unnecessarily\n>>> >> cavalier. There's a lot of mixed-version environments where git is\n>>> >> deployed in, and we almost never break the configuration in this way (I\n>>> >> think in the past always by mistake).\n>>> >\n>>> >> In this case it's easy to avoid it, and coming up with a less narrow\n>>> >> config model[1] seems like a good idea in any case to unify the various\n>>> >> outstanding work in this area.\n>>> >>\n>>> >> More generally on this series, per the thread ending in [2] I really\n>>> >\n>>> > My primary goal in all of these changes is to move git-for-windows over to\n>>> > a default of batch fsync so that it can get closer to other platforms\n>>> > in performance\n>>> > of 'git add' while still retaining the same level of data integrity.\n>>> > I'm hoping that\n>>> > most end-users are just sticking to defaults here.\n>>> >\n>>> > I'm happy to change the configuration schema again if there's a\n>>> > consensus from the Git\n>>> > community that backwards-compatibility of the configuration is\n>>> > actually important to someone.\n>>> >\n>>> > Also, if we're doing a deeper rethink of the fsync configuration (as\n>>> > prompted by this work and\n>>> > Eric Wong's and Patrick Steinhardts work), do we want to retain a mode\n>>> > where we fsync some\n>>> > parts of the persistent repo data but not others?  If we add fsyncing\n>>> > of the index in addition to the refs,\n>>> > I believe we would have covered all of the critical data structures\n>>> > that would be needed to find the\n>>> > data that a user has added to the repo if they complete a series of\n>>> > git commands and then experience\n>>> > a system crash.\n>>>\n>>> Just talking about it is how we'll find consensus, maybe you & Junio\n>>> would like to keep it as-is. I don't see why we'd expose this bad edge\n>>> case in configuration handling to users when it's entirely avoidable,\n>>> and we're still in the design phase.\n>>\n>> After trying to figure out an implementation, I have a new proposal,\n>> which I've shared on the other thread [1].\n>>\n>> [1] https://lore.kernel.org/git/CANQDOdcdhfGtPg0PxpXQA5gQ4x9VknKDKCCi4HEB0Z1xgnjKzg@mail.gmail.com/\n>\n> This LGTM, or something simpler as Junio points out with his \"too\n> fine-grained?\" comment as a follow-up. I'm honestly quite apathetic\n> about what we end up with exactly as long as:\n>\n>  1. We get the people who are adding these config settings to talk & see if they make\n>     sense in combination.\n>\n>  2. We avoid the trap of hard dying on older versions.\n>\n>>>\n>>> >> don't get why we have code like this:\n>>> >>\n>>> >>         @@ -503,10 +504,12 @@ static void unpack_all(void)\n>>> >>                 if (!quiet)\n>>> >>                         progress = start_progress(_(\"Unpacking objects\"), nr_objects);\n>>> >>                 CALLOC_ARRAY(obj_list, nr_objects);\n>>> >>         +       plug_bulk_checkin();\n>>> >>                 for (i = 0; i < nr_objects; i++) {\n>>> >>                         unpack_one(i);\n>>> >>                         display_progress(progress, i + 1);\n>>> >>                 }\n>>> >>         +       unplug_bulk_checkin();\n>>> >>                 stop_progress(&progress);\n>>> >>\n>>> >>                 if (delta_list)\n>>> >>\n>>> >> As opposed to doing an fsync on the last object we're\n>>> >> processing. I.e. why do we need the step of intentionally making the\n>>> >> objects unavailable in the tmp-objdir, and creating a \"cookie\" file to\n>>> >> sync at the start/end, as opposed to fsyncing on the last file (which\n>>> >> we're writing out anyway).\n>>> >>\n>>> >> 1. https://lore.kernel.org/git/211110.86r1bogg27.gmgdl@evledraar.gmail.com/\n>>> >> 2. https://lore.kernel.org/git/20211111000349.GA703@neerajsi-x1.localdomain/\n>>> >\n>>> > It's important to not expose an object's final name until its contents\n>>> > have been fsynced\n>>> > to disk. We want to ensure that wherever we crash, we won't have a\n>>> > loose object that\n>>> > Git may later try to open where the filename doesn't match the content\n>>> > hash. I believe it's\n>>> > okay for a given OID to be missing, since a later command could\n>>> > recreate it, but an object\n>>> > with a wrong hash looks like it would persist until we do a git-fsck.\n>>>\n>>> Yes, we handle that rather badly, as I mentioned in some other threads,\n>>> but not doing the fsync on the last object v.s. a \"cookie\" file right\n>>> afterwards seems like a hail-mary at best, no?\n>>>\n>>\n>> I'm not quite grasping what you're saying here. Are you saying that\n>> using a dummy\n>> file instead of one of the actual objects is less likely to produce\n>> the desired outcome\n>> on actual filesystem implementations?\n>\n> [...covered below...]\n>\n>>> > I thought about figuring out how to sync the last object rather than some random\n>>> > \"cookie\" file, but it wasn't clear to me how I'd figure out which\n>>> > object is actually last\n>>> > from library code in a way that doesn't burden each command with\n>>> > somehow figuring\n>>> > out its last object and communicating that. The 'cookie' approach\n>>> > seems to lead to a cleaner\n>>> > interface for callers.\n>>>\n>>> The above quoted code is looping through nr_objects isn't it? Can't a\n>>> \"do fsync\" be passed down to unpack_one() when we process the last loose\n>>> object?\n>>\n>> Are you proposing that we do something different for unpack_objects\n>> versus update_index\n>> and git-add?  I was hoping to keep all of the users of the batch fsync\n>> functionality equivalent.\n>> For the git-add workflow and update-index, we'd need to track the most\n>> recent file so that we\n>> can go back and fsync it.  I don't believe that syncing the last\n>> object composes well with the existing\n>> implementation of those commands.\n>\n> There's probably cases where we need the cookie. I just mean instead of\n> the API being (as seen above in the quoted part), pseudocode:\n>\n>     # A\n>     bulk_checkin_start_make_cookie():\n>     n = 10\n>     for i in 1..n:\n>         write_nth(i, fsync: 0);\n>     bulk_checkin_end_commit_cookie();\n>\n> To have it be:\n>\n>     # B\n>     bulk_checkin_start(do_cookie: 0);\n>     n = 10\n>     for i in 1..n:\n>         write_nth(i, fsync: (i == n));\n>     bulk_checkin_end();\n>\n> Or actually, presumably simpler as:\n>\n>     # C\n>     all_fsync = bulk_checkin_mode() ? 0 : fsync_turned_on_in_general();\n>     end_fsync = bulk_checkin_mode() ? 1 : all_fsync;\n>     n = 10;\n>     for i in 1..n:\n>         write_nth(i, fsync: (i == n) ? end_fsync : all_fsync);\n>\n> I.e. maybe there are cases where you really do need \"A\", but we're\n> usually (always?) writing out N objects, and we usually know it at the\n> same point where you'd want the plug_bulk_checkin/unplug_bulk_checkin,\n> so just fsyncing the last object/file/ref/whatever means we don't need\n> the whole ceremony of the cookie file.\n>\n> I don't mind it per-se, but \"B\" and \"C\" just seem a lot simpler,\n> particulary since as those examples show we'll presumably want to pass\n> down a \"do fsync?\" to these in general, and we even usually have a\n> display_progress() in there.\n>\n> So doesn't just doing \"B\" or \"C\" eliminate the need for a cookie\n> entirely?\n>\n> Another advantage of that is that you'll presumably want such tracking\n> anyway even for the case of \"A\".\n>\n> Because as soon as you have say a batch operation of writing X objects\n> and Y refs you'd want to track this anyway. I.e. either only fsync() on\n> the ref write (particularly if there's just the one ref), or on the last\n> ref, or for each ref and no object syncs. So this (like \"C\", except for\n> the \"do_batch\" in the \"end_fsync\" case):\n>\n>     # D\n>     do_batch = in_existing_bulk_checkin() ? 1 : 0;\n>     all_fsync = bulk_checkin_mode() ? 0 : fsync_turned_on_in_general();\n>     end_fsync = bulk_checkin_mode() ? do_batch : all_fsync;\n>     n = 10;\n>     for i in 1..n:\n>         write_nth(i, fsync: (i == n) ? end_fsync : all_fsync);\n>\n> I mean, usually we'd want the \"all refs\", I'm just thinking of a case\n> like \"git fast-import\" or other known-to-the-user batch operation.\n>\n> Or, as in the case of my 4bc1fd6e394 (pack-objects: rename .idx files\n> into place after .bitmap files, 2021-09-09) we'd want to know that we're\n> writing all of say *.bitmap, *.rev where we currently fsync() all of\n> them, write *.bitmap, *.rev and *.pack (not sure that one is safe)\n> without fsync(), and then only fsync (or that and in-place move) the\n> *.idx.\n\nReplying to an old-ish E-Mail of mine with some more thought that came\nto mind after[1] (another recently resurrected fsync() thread).\n\nI wonder if there's another twist on the plan outlined in [2] that would\nbe both portable & efficient, i.e. the \"slow\" POSIX way to write files\nA..Z is to open/write/close/fsync each one, so we'll trigger a HW flush\nN times.\n\nAnd as we've discussed, doing it just on Z will implicitly flush A..Y on\ncommon OS's in the wild, which we're taking advantage of here.\n\nBut aside from the rename() dance in[2], what do those OS's do if you\nwrite A..Z, fsync() the \"fd\" for Z, and then fsync A..Y (or, presumably\nequivalently, in reverse order: Y..A).\n\nI'd think they'd be smart enough to know that they already implicitly\nflushed that data since Z was flushend, and make those fsync()'s a\nrather cheap noop.\n\nBut I don't know, hence the question.\n\nIf that's true then perhaps it's a path towards having our cake and\neating it too in some cases?\n\nI.e. an FS that would flush A..Y if we flush Z would do so quickly and\nreliably, whereas a FS that doesn't have such an optimization might be\njust as slow for all of A..Y, but at least it'll be safe.\n\n1. https://lore.kernel.org/git/220309.867d93lztw.gmgdl@evledraar.gmail.com/\n2. https://lore.kernel.org/git/e1747ce00af7ab3170a69955b07d995d5321d6f3.1637020263.git.gitgitgadget@gmail.com/\n"},{"id":"450971","messageId":"CANQDOdcJX9bYAJN4_M5-k_Ssg+kK+CVOsanXr+Xnu7B+nzfqSw@mail.gmail.com","threadId":"56371","inReplyTo":"220310.86lexilo3d.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T01:16:07Z","receivedAt":"2022-03-10T01:16:26Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 9, 2022 at 3:10 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n> Replying to an old-ish E-Mail of mine with some more thought that came\n> to mind after[1] (another recently resurrected fsync() thread).\n>\n> I wonder if there's another twist on the plan outlined in [2] that would\n> be both portable & efficient, i.e. the \"slow\" POSIX way to write files\n> A..Z is to open/write/close/fsync each one, so we'll trigger a HW flush\n> N times.\n>\n> And as we've discussed, doing it just on Z will implicitly flush A..Y on\n> common OS's in the wild, which we're taking advantage of here.\n>\n> But aside from the rename() dance in[2], what do those OS's do if you\n> write A..Z, fsync() the \"fd\" for Z, and then fsync A..Y (or, presumably\n> equivalently, in reverse order: Y..A).\n>\n> I'd think they'd be smart enough to know that they already implicitly\n> flushed that data since Z was flushend, and make those fsync()'s a\n> rather cheap noop.\n>\n> But I don't know, hence the question.\n>\n> If that's true then perhaps it's a path towards having our cake and\n> eating it too in some cases?\n>\n> I.e. an FS that would flush A..Y if we flush Z would do so quickly and\n> reliably, whereas a FS that doesn't have such an optimization might be\n> just as slow for all of A..Y, but at least it'll be safe.\n>\n> 1. https://lore.kernel.org/git/220309.867d93lztw.gmgdl@evledraar.gmail.com/\n> 2. https://lore.kernel.org/git/e1747ce00af7ab3170a69955b07d995d5321d6f3.1637020263.git.gitgitgadget@gmail.com/\n\nThe important angle here is that we need some way to indicate to the\nOS what A..Y is before we fsync on Z.  I.e. the OS will cache any\nwrites in memory until some sync-ish operation is done on *that\nspecific file*.  Syncing just 'Z' with no sync operations on A..Y\ndoesn't indicate that A..Y would get written out.  Apparently the bad\nold ext3 behavior was similar to what you're proposing where a sync on\n'Z' would imply something about independent files.\n\nHere's an interesting paper I recently came across that proposes the\ninterface we'd really want, 'syncv':\nhttps://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.924.1168&rep=rep1&type=pdf.\n\nThanks,\nNeeraj\n"},{"id":"451016","messageId":"220310.86r179ki38.gmgdl@evledraar.gmail.com","threadId":"56371","inReplyTo":"CANQDOdcJX9bYAJN4_M5-k_Ssg+kK+CVOsanXr+Xnu7B+nzfqSw@mail.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-10T14:01:34Z","receivedAt":"2022-03-10T14:20:45Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Mar 09 2022, Neeraj Singh wrote:\n\n> On Wed, Mar 9, 2022 at 3:10 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>\n>> Replying to an old-ish E-Mail of mine with some more thought that came\n>> to mind after[1] (another recently resurrected fsync() thread).\n>>\n>> I wonder if there's another twist on the plan outlined in [2] that would\n>> be both portable & efficient, i.e. the \"slow\" POSIX way to write files\n>> A..Z is to open/write/close/fsync each one, so we'll trigger a HW flush\n>> N times.\n>>\n>> And as we've discussed, doing it just on Z will implicitly flush A..Y on\n>> common OS's in the wild, which we're taking advantage of here.\n>>\n>> But aside from the rename() dance in[2], what do those OS's do if you\n>> write A..Z, fsync() the \"fd\" for Z, and then fsync A..Y (or, presumably\n>> equivalently, in reverse order: Y..A).\n>>\n>> I'd think they'd be smart enough to know that they already implicitly\n>> flushed that data since Z was flushend, and make those fsync()'s a\n>> rather cheap noop.\n>>\n>> But I don't know, hence the question.\n>>\n>> If that's true then perhaps it's a path towards having our cake and\n>> eating it too in some cases?\n>>\n>> I.e. an FS that would flush A..Y if we flush Z would do so quickly and\n>> reliably, whereas a FS that doesn't have such an optimization might be\n>> just as slow for all of A..Y, but at least it'll be safe.\n>>\n>> 1. https://lore.kernel.org/git/220309.867d93lztw.gmgdl@evledraar.gmail.com/\n>> 2. https://lore.kernel.org/git/e1747ce00af7ab3170a69955b07d995d5321d6f3.1637020263.git.gitgitgadget@gmail.com/\n>\n> The important angle here is that we need some way to indicate to the\n> OS what A..Y is before we fsync on Z.  I.e. the OS will cache any\n> writes in memory until some sync-ish operation is done on *that\n> specific file*.  Syncing just 'Z' with no sync operations on A..Y\n> doesn't indicate that A..Y would get written out.  Apparently the bad\n> old ext3 behavior was similar to what you're proposing where a sync on\n> 'Z' would imply something about independent files.\n\nIt's certainly starting to sound like I'm misunderstanding this whole\nthing, but just to clarify again I'm talking about the sort of loops\nmentioned upthread in my [1]. I.e. you have (to copy from that E-Mail):\n\n    bulk_checkin_start_make_cookie():\n    n = 10\n    for i in 1..n:\n        write_nth(i, fsync: 0);\n    bulk_checkin_end_commit_cookie();\n\nI.e. we have a \"cookie\" file in a given dir (where, in this example,\nwe'd also write files A..Z). I.e. we write:\n\n    cookie\n    {A..Z}\n    cookie\n\nAnd then only fsync() on the \"cookie\" at the end, which \"flushes\" the\nA..Z updates on some FS's (again, all per my possibly-incorrect\nunderstanding).\n\nWhich is why I proposed that in many/all cases we could do this,\ni.e. just the same without the \"cookie\" file (which AFAICT isn't needed\nper-se, but was just added to make the API a bit simpler in not needing\nto modify the relevant loops):\n\n    all_fsync = bulk_checkin_mode() ? 0 : fsync_turned_on_in_general();\n    end_fsync = bulk_checkin_mode() ? 1 : all_fsync;\n    n = 10;\n    for i in 1..n:\n        write_nth(i, fsync: (i == n) ? end_fsync : all_fsync);\n\nI.e. we don't pay the cost of the fsync() as we're in the loop, but just\nfor the last file, which \"flushes\" the rest.\n\nSo far all of that's a paraphrasing of existing exchanges, but what I\nwas wondering now in[2] is if we add this to this last example above:\n\n    for i in 1..n-1:\n        fsync_nth(i)\n\nWouldn't those same OS's that are being clever about deferring the\nsyncing of A..Z as a \"batch\" be clever enough to turn that (re-)syncing\ninto a NOOP?\n\nOf course in this case we'd need to keep the fd's open and be clever\nabout E[MN]FILE (i.e. \"Too many open...\"), or do an fsync() every Nth\nfor some reasonable Nth, e.g. somewhere in the 2^10..2^12 range.\n\nBut *if* this works it seems to me to be something we might be able to\nenable when \"core.fsyncObjectFiles\" is configured on those systems.\n\nI.e. the implicit assumption with that configuration was that if we sync\nN loose objects and then update and fsync the ref that the FS would\nqueue up the ref update after the syncing of the loose objects.\n\nThis new \"cookie\" (or my suggested \"fsync last of N\") is basically\nmaking the same assumption, just with the slight twist that some OSs/FSs\nare known to behave like that on a per-subdir basis, no?\n\n> Here's an interesting paper I recently came across that proposes the\n> interface we'd really want, 'syncv':\n> https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.924.1168&rep=rep1&type=pdf.\n\n1. https://lore.kernel.org/git/211201.864k7sbdjt.gmgdl@evledraar.gmail.com/\n2. https://lore.kernel.org/git/220310.86lexilo3d.gmgdl@evledraar.gmail.com/\n"},{"id":"451051","messageId":"CANQDOdf1pE+PUv_XqLobGq8Wvan-iH28RhBJFYM-NfxHKBjU+Q@mail.gmail.com","threadId":"56371","inReplyTo":"220310.86r179ki38.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T17:52:55Z","receivedAt":"2022-03-10T17:53:13Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 10, 2022 at 6:17 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Wed, Mar 09 2022, Neeraj Singh wrote:\n>\n> > On Wed, Mar 9, 2022 at 3:10 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> >>\n> >> Replying to an old-ish E-Mail of mine with some more thought that came\n> >> to mind after[1] (another recently resurrected fsync() thread).\n> >>\n> >> I wonder if there's another twist on the plan outlined in [2] that would\n> >> be both portable & efficient, i.e. the \"slow\" POSIX way to write files\n> >> A..Z is to open/write/close/fsync each one, so we'll trigger a HW flush\n> >> N times.\n> >>\n> >> And as we've discussed, doing it just on Z will implicitly flush A..Y on\n> >> common OS's in the wild, which we're taking advantage of here.\n> >>\n> >> But aside from the rename() dance in[2], what do those OS's do if you\n> >> write A..Z, fsync() the \"fd\" for Z, and then fsync A..Y (or, presumably\n> >> equivalently, in reverse order: Y..A).\n> >>\n> >> I'd think they'd be smart enough to know that they already implicitly\n> >> flushed that data since Z was flushend, and make those fsync()'s a\n> >> rather cheap noop.\n> >>\n> >> But I don't know, hence the question.\n> >>\n> >> If that's true then perhaps it's a path towards having our cake and\n> >> eating it too in some cases?\n> >>\n> >> I.e. an FS that would flush A..Y if we flush Z would do so quickly and\n> >> reliably, whereas a FS that doesn't have such an optimization might be\n> >> just as slow for all of A..Y, but at least it'll be safe.\n> >>\n> >> 1. https://lore.kernel.org/git/220309.867d93lztw.gmgdl@evledraar.gmail.com/\n> >> 2. https://lore.kernel.org/git/e1747ce00af7ab3170a69955b07d995d5321d6f3.1637020263.git.gitgitgadget@gmail.com/\n> >\n> > The important angle here is that we need some way to indicate to the\n> > OS what A..Y is before we fsync on Z.  I.e. the OS will cache any\n> > writes in memory until some sync-ish operation is done on *that\n> > specific file*.  Syncing just 'Z' with no sync operations on A..Y\n> > doesn't indicate that A..Y would get written out.  Apparently the bad\n> > old ext3 behavior was similar to what you're proposing where a sync on\n> > 'Z' would imply something about independent files.\n>\n> It's certainly starting to sound like I'm misunderstanding this whole\n> thing, but just to clarify again I'm talking about the sort of loops\n> mentioned upthread in my [1]. I.e. you have (to copy from that E-Mail):\n>\n>     bulk_checkin_start_make_cookie():\n>     n = 10\n>     for i in 1..n:\n>         write_nth(i, fsync: 0);\n>     bulk_checkin_end_commit_cookie();\n>\n> I.e. we have a \"cookie\" file in a given dir (where, in this example,\n> we'd also write files A..Z). I.e. we write:\n>\n>     cookie\n>     {A..Z}\n>     cookie\n>\n> And then only fsync() on the \"cookie\" at the end, which \"flushes\" the\n> A..Z updates on some FS's (again, all per my possibly-incorrect\n> understanding).\n>\n> Which is why I proposed that in many/all cases we could do this,\n> i.e. just the same without the \"cookie\" file (which AFAICT isn't needed\n> per-se, but was just added to make the API a bit simpler in not needing\n> to modify the relevant loops):\n>\n>     all_fsync = bulk_checkin_mode() ? 0 : fsync_turned_on_in_general();\n>     end_fsync = bulk_checkin_mode() ? 1 : all_fsync;\n>     n = 10;\n>     for i in 1..n:\n>         write_nth(i, fsync: (i == n) ? end_fsync : all_fsync);\n>\n> I.e. we don't pay the cost of the fsync() as we're in the loop, but just\n> for the last file, which \"flushes\" the rest.\n>\n> So far all of that's a paraphrasing of existing exchanges, but what I\n> was wondering now in[2] is if we add this to this last example above:\n>\n>     for i in 1..n-1:\n>         fsync_nth(i)\n>\n> Wouldn't those same OS's that are being clever about deferring the\n> syncing of A..Z as a \"batch\" be clever enough to turn that (re-)syncing\n> into a NOOP?\n>\n> Of course in this case we'd need to keep the fd's open and be clever\n> about E[MN]FILE (i.e. \"Too many open...\"), or do an fsync() every Nth\n> for some reasonable Nth, e.g. somewhere in the 2^10..2^12 range.\n>\n> But *if* this works it seems to me to be something we might be able to\n> enable when \"core.fsyncObjectFiles\" is configured on those systems.\n>\n> I.e. the implicit assumption with that configuration was that if we sync\n> N loose objects and then update and fsync the ref that the FS would\n> queue up the ref update after the syncing of the loose objects.\n>\n> This new \"cookie\" (or my suggested \"fsync last of N\") is basically\n> making the same assumption, just with the slight twist that some OSs/FSs\n> are known to behave like that on a per-subdir basis, no?\n>\n> > Here's an interesting paper I recently came across that proposes the\n> > interface we'd really want, 'syncv':\n> > https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.924.1168&rep=rep1&type=pdf.\n>\n> 1. https://lore.kernel.org/git/211201.864k7sbdjt.gmgdl@evledraar.gmail.com/\n> 2. https://lore.kernel.org/git/220310.86lexilo3d.gmgdl@evledraar.gmail.com/\n\nOn the actual FS implementations in the three common OSes I'm familiar with\n(macOS, Windows, Linux), each file has its own independent data caching in OS\nmemory.  Fsyncing one of them doesn't necessarily imply writing out\nthe OS cache for\nany other file.  Except, apparently, on ext3 in data=ordered mode, but\nthat FS is no\nlonger common.  On Linux, we use sync_file_range to get the OS to\nwrite the in-memory\ncache to the storage hardware, which is what makes the data\n'available' to fsync.\n\nNow, we could consider an implementation where we call sync_file_range\nwithout the\nwait flags (i.e. without SYNC_FILE_RANGE_WAIT_BEFORE and\nSYNC_FILE_RANGE_WAIT_AFTER). Then we could later fsync every file (or batch of\nfiles), which might be more efficient if the OS coalesces the disk\ncache flushes.  I expect\nthat this method is less likely to give us the desired performance on\ncommon linux FSes,\nhowever.\n\nThe macOS and Windows APIs are defined a bit differently from Linux.\nIn both those OSes,\nwe're actually calling fsync-equivalent APIs that are defined to write\nback all the relevant data and\nmetadata, just without the storage cache flush.\n\nSo to summarize:\n1. We need to do write(2) to get the data out of Git and into the OS\nfilesystem cache.\n2. We need some API (macOS-fsync, Windows-NtFlushBuffersFileEx,\nLinux-sync_file_range)\n   to transfer the data per-file to the storage controller, but\nwithout flushing the storage controller.\n3. We need some api (macOS-F_FULLFSYNC, Windows-NtFlushBuffersFile, Linux-fsync)\n   to push the storage controller cache to durable media. This only\nneeds to be done once\n   at the end to push out the data made available in step (2).\n\n\nThanks,\nNeeraj\n"},{"id":"451053","messageId":"00ae01d834a9$d443a530$7ccaef90$@nexbridge.com","threadId":"56371","inReplyTo":"CANQDOdf1pE+PUv_XqLobGq8Wvan-iH28RhBJFYM-NfxHKBjU+Q@mail.gmail.com","subject":"RE: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-10T18:08:16Z","receivedAt":"2022-03-10T18:08:29Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 10, 2022 12:53 PM, Neeraj Singh wrote:\n>On Thu, Mar 10, 2022 at 6:17 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>wrote:\n>>\n>>\n>> On Wed, Mar 09 2022, Neeraj Singh wrote:\n>>\n>> > On Wed, Mar 9, 2022 at 3:10 PM Ævar Arnfjörð Bjarmason\n><avarab@gmail.com> wrote:\n>> >>\n>> >> Replying to an old-ish E-Mail of mine with some more thought that\n>> >> came to mind after[1] (another recently resurrected fsync() thread).\n>> >>\n>> >> I wonder if there's another twist on the plan outlined in [2] that\n>> >> would be both portable & efficient, i.e. the \"slow\" POSIX way to\n>> >> write files A..Z is to open/write/close/fsync each one, so we'll\n>> >> trigger a HW flush N times.\n>> >>\n>> >> And as we've discussed, doing it just on Z will implicitly flush\n>> >> A..Y on common OS's in the wild, which we're taking advantage of here.\n>> >>\n>> >> But aside from the rename() dance in[2], what do those OS's do if\n>> >> you write A..Z, fsync() the \"fd\" for Z, and then fsync A..Y (or,\n>> >> presumably equivalently, in reverse order: Y..A).\n>> >>\n>> >> I'd think they'd be smart enough to know that they already\n>> >> implicitly flushed that data since Z was flushend, and make those\n>> >> fsync()'s a rather cheap noop.\n>> >>\n>> >> But I don't know, hence the question.\n>> >>\n>> >> If that's true then perhaps it's a path towards having our cake and\n>> >> eating it too in some cases?\n>> >>\n>> >> I.e. an FS that would flush A..Y if we flush Z would do so quickly\n>> >> and reliably, whereas a FS that doesn't have such an optimization\n>> >> might be just as slow for all of A..Y, but at least it'll be safe.\n>> >>\n>> >> 1.\n>> >> https://lore.kernel.org/git/220309.867d93lztw.gmgdl@evledraar.gmail\n>> >> .com/ 2.\n>> >> https://lore.kernel.org/git/e1747ce00af7ab3170a69955b07d995d5321d6f\n>> >> 3.1637020263.git.gitgitgadget@gmail.com/\n>> >\n>> > The important angle here is that we need some way to indicate to the\n>> > OS what A..Y is before we fsync on Z.  I.e. the OS will cache any\n>> > writes in memory until some sync-ish operation is done on *that\n>> > specific file*.  Syncing just 'Z' with no sync operations on A..Y\n>> > doesn't indicate that A..Y would get written out.  Apparently the\n>> > bad old ext3 behavior was similar to what you're proposing where a\n>> > sync on 'Z' would imply something about independent files.\n>>\n>> It's certainly starting to sound like I'm misunderstanding this whole\n>> thing, but just to clarify again I'm talking about the sort of loops\n>> mentioned upthread in my [1]. I.e. you have (to copy from that E-Mail):\n>>\n>>     bulk_checkin_start_make_cookie():\n>>     n = 10\n>>     for i in 1..n:\n>>         write_nth(i, fsync: 0);\n>>     bulk_checkin_end_commit_cookie();\n>>\n>> I.e. we have a \"cookie\" file in a given dir (where, in this example,\n>> we'd also write files A..Z). I.e. we write:\n>>\n>>     cookie\n>>     {A..Z}\n>>     cookie\n>>\n>> And then only fsync() on the \"cookie\" at the end, which \"flushes\" the\n>> A..Z updates on some FS's (again, all per my possibly-incorrect\n>> understanding).\n>>\n>> Which is why I proposed that in many/all cases we could do this, i.e.\n>> just the same without the \"cookie\" file (which AFAICT isn't needed\n>> per-se, but was just added to make the API a bit simpler in not\n>> needing to modify the relevant loops):\n>>\n>>     all_fsync = bulk_checkin_mode() ? 0 : fsync_turned_on_in_general();\n>>     end_fsync = bulk_checkin_mode() ? 1 : all_fsync;\n>>     n = 10;\n>>     for i in 1..n:\n>>         write_nth(i, fsync: (i == n) ? end_fsync : all_fsync);\n>>\n>> I.e. we don't pay the cost of the fsync() as we're in the loop, but\n>> just for the last file, which \"flushes\" the rest.\n>>\n>> So far all of that's a paraphrasing of existing exchanges, but what I\n>> was wondering now in[2] is if we add this to this last example above:\n>>\n>>     for i in 1..n-1:\n>>         fsync_nth(i)\n>>\n>> Wouldn't those same OS's that are being clever about deferring the\n>> syncing of A..Z as a \"batch\" be clever enough to turn that\n>> (re-)syncing into a NOOP?\n>>\n>> Of course in this case we'd need to keep the fd's open and be clever\n>> about E[MN]FILE (i.e. \"Too many open...\"), or do an fsync() every Nth\n>> for some reasonable Nth, e.g. somewhere in the 2^10..2^12 range.\n>>\n>> But *if* this works it seems to me to be something we might be able to\n>> enable when \"core.fsyncObjectFiles\" is configured on those systems.\n>>\n>> I.e. the implicit assumption with that configuration was that if we\n>> sync N loose objects and then update and fsync the ref that the FS\n>> would queue up the ref update after the syncing of the loose objects.\n>>\n>> This new \"cookie\" (or my suggested \"fsync last of N\") is basically\n>> making the same assumption, just with the slight twist that some\n>> OSs/FSs are known to behave like that on a per-subdir basis, no?\n>>\n>> > Here's an interesting paper I recently came across that proposes the\n>> > interface we'd really want, 'syncv':\n>> >\n>https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.924.1168&rep=rep1\n>&type=pdf.\n>>\n>> 1.\n>> https://lore.kernel.org/git/211201.864k7sbdjt.gmgdl@evledraar.gmail.co\n>> m/ 2.\n>> https://lore.kernel.org/git/220310.86lexilo3d.gmgdl@evledraar.gmail.co\n>> m/\n>\n>On the actual FS implementations in the three common OSes I'm familiar with\n>(macOS, Windows, Linux), each file has its own independent data caching in OS\n>memory.  Fsyncing one of them doesn't necessarily imply writing out the OS cache\n>for any other file.  Except, apparently, on ext3 in data=ordered mode, but that FS\n>is no longer common.  On Linux, we use sync_file_range to get the OS to write the\n>in-memory cache to the storage hardware, which is what makes the data\n>'available' to fsync.\n>\n>Now, we could consider an implementation where we call sync_file_range\n>without the wait flags (i.e. without SYNC_FILE_RANGE_WAIT_BEFORE and\n>SYNC_FILE_RANGE_WAIT_AFTER). Then we could later fsync every file (or batch\n>of files), which might be more efficient if the OS coalesces the disk cache flushes.  I\n>expect that this method is less likely to give us the desired performance on\n>common linux FSes, however.\n>\n>The macOS and Windows APIs are defined a bit differently from Linux.\n>In both those OSes,\n>we're actually calling fsync-equivalent APIs that are defined to write back all the\n>relevant data and metadata, just without the storage cache flush.\n>\n>So to summarize:\n>1. We need to do write(2) to get the data out of Git and into the OS filesystem\n>cache.\n>2. We need some API (macOS-fsync, Windows-NtFlushBuffersFileEx,\n>Linux-sync_file_range)\n>   to transfer the data per-file to the storage controller, but without flushing the\n>storage controller.\n>3. We need some api (macOS-F_FULLFSYNC, Windows-NtFlushBuffersFile, Linux-\n>fsync)\n>   to push the storage controller cache to durable media. This only needs to be\n>done once\n>   at the end to push out the data made available in step (2).\n\nWhile this might not be a surprise, on some platforms fsync is a thread-blocking operation. When the OS has kernel threads, fsync can potentially cause multiple processes (if implemented that way) to block, particularly where an fd is shared across threads (and thus processes), which may end up causing a deadlock. We might need to keep an eye out for this type of situation in the future and at least try to test for it. I cannot actually see a situation where this would occur in git, but that does not mean it is impossible. Food for thought.\n--Randall\n\n"},{"id":"451059","messageId":"CANQDOdeFdTzB6GKEeNPJm0-j0qyD+n8+e=+Qn98PvRg8N5wdEQ@mail.gmail.com","threadId":"56371","inReplyTo":"00ae01d834a9$d443a530$7ccaef90$@nexbridge.com","subject":"Re: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-10T18:43:08Z","receivedAt":"2022-03-10T18:43:24Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 10, 2022 at 10:08 AM <rsbecker@nexbridge.com> wrote:\n> While this might not be a surprise, on some platforms fsync is a thread-blocking operation. When the OS has kernel threads, fsync can potentially cause multiple processes (if implemented that way) to block, particularly where an fd is shared across threads (and thus processes), which may end up causing a deadlock. We might need to keep an eye out for this type of situation in the future and at least try to test for it. I cannot actually see a situation where this would occur in git, but that does not mean it is impossible. Food for thought.\n> --Randall\n\nfsync is expected to block the calling thread until the underlying\ndata is durable.  Unless the OS somehow depends on the git process to\nmake progress before fsync can complete, there should be no deadlock,\nsince there would be no cycle in the waiting graph.  This could be a\nproblem for FUSE implementations that are backed by Git, but they\nalready have to deal with that possiblity today and this patch series\ndoesn't change anything.\n"},{"id":"451061","messageId":"00b501d834af$794cd830$6be68890$@nexbridge.com","threadId":"56371","inReplyTo":"CANQDOdeFdTzB6GKEeNPJm0-j0qyD+n8+e=+Qn98PvRg8N5wdEQ@mail.gmail.com","subject":"RE: [PATCH v9 0/9] Implement a batched fsync option for core.fsyncObjectFiles","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-10T18:48:41Z","receivedAt":"2022-03-10T18:48:53Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 10, 2022 1:43 PM, Neeraj Singh wrote:\n>On Thu, Mar 10, 2022 at 10:08 AM <rsbecker@nexbridge.com> wrote:\n>> While this might not be a surprise, on some platforms fsync is a thread-blocking\n>operation. When the OS has kernel threads, fsync can potentially cause multiple\n>processes (if implemented that way) to block, particularly where an fd is shared\n>across threads (and thus processes), which may end up causing a deadlock. We\n>might need to keep an eye out for this type of situation in the future and at least\n>try to test for it. I cannot actually see a situation where this would occur in git, but\n>that does not mean it is impossible. Food for thought.\n>> --Randall\n>\n>fsync is expected to block the calling thread until the underlying data is durable.\n>Unless the OS somehow depends on the git process to make progress before\n>fsync can complete, there should be no deadlock, since there would be no cycle in\n>the waiting graph.  This could be a problem for FUSE implementations that are\n>backed by Git, but they already have to deal with that possiblity today and this\n>patch series doesn't change anything.\n\nThat assumption is based on a specific threading model. In cooperative user-thread models, fsync is process-blocking. While fsync, by spec is required to block the thread, there are no limitations on blocking everything else. In some systems, an fsync can block the entire file system. Just pointing that out. \n\n"}]}